Here are
2 public repositories
matching this topic...
Open recipes, engine patches, and benchmark harnesses for LLM inference on Intel Arc Pro B60/B70 (Battlemage, Xe2). MoE 35B at 160 t/s decode / 7.5K t/s prefill single-stream, 27B at 50~ t/s decode / 1.7K t/s prefill single stream. vLLM XPU MTP unlocked. Muse Glimmer recipe added!!
Updated
Aug 15, 2026
Python
Field-tested guide: multi-GPU vLLM tensor-parallel (TP=2/TP=4) on Intel Arc Pro B70 (Battlemage BMG-G31, Xe2) on Linux. Driver setup (xe force_probe=e223), bare-metal vLLM + oneAPI 2025.3, the compute-runtime multi-root USM + triton-xpu init_devices fixes, FP8/int4-AutoRound quant, root-cause error reports. AI-agent readable (AGENTS.md).
Updated
Jun 13, 2026
Shell
Improve this page
Add a description, image, and links to the
intel-arc-b70
topic page so that developers can more easily learn about it.
Curate this topic
Add this topic to your repo
To associate your repository with the
intel-arc-b70
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.