GPU for local LLM · budget
Intel Arc B580
12GB for around the price of an 8GB card, if you'll ride Intel's improving-but-young AI stack.
- VRAM
- 12 GB
- TDP
- 190 W
- Bandwidth
- 456 GB/s
- £ / GB
- £21
What it runs
With 12 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly
~20B params
Verdict
The cheapest sensible way onto 12GB, if you treat the software as a project rather than a product.
BEST FOR
- + first cheap AI card
- + 7B to 8B models
- + low-power tinker boxes
NOT FOR
- - fast token speed
- - plug-and-play setups
- - 32B ambitions
The Intel Arc B580 is the budget card with a twist worth noticing. Twelve gigabytes of memory for roughly the price of an 8GB competitor, which for local AI is the number that matters. It won’t be quick and the software is still growing up, but as a first cheap way onto 12GB it’s a tempting little card. Treat the setup as a weekend project and it rewards you.
The take
The appeal is VRAM per pound at the bottom of the market. Most cards this cheap ship with 8GB, which local models outgrow fast, so the B580’s 12GB lets you run a 7B or 8B properly rather than squeezed. The trade-off is Intel’s AI software, which is younger than CUDA and moving quickly but not there yet. You’ll use IPEX-LLM, OpenVINO or a recent Ollama build, and you’ll do a bit more reading than an NVIDIA owner. Go in expecting that and it’s a fair deal.
What it’ll actually run
Twelve gigabytes is an entry ticket, and a useful one. A 7B or 8B at Q4 sits comfortably with a decent context, which covers assistants, chat and light coding. A 13B to 14B fits at a tighter quant. Beyond that you’re offloading to system RAM and watching speed fall away, so a 32B isn’t the job for this card. Match your ambitions to 12GB and it does what you ask.
Speed is modest. At 456 GB/s the bandwidth is well below the pricier cards, and inference is bandwidth bound, so tokens come out at a steady pace rather than a rush. For a personal assistant that’s fine. For heavy batch work or long generations, you’ll feel the wait. This is a card you buy for capacity and value, not throughput.
The hardware is easy to live with. One hundred and ninety watts, a single 8-pin, a compact dual-slot card that runs cool and quiet. It slots into a small, cheap, low-power build without fuss, which suits an always-on box tucked away somewhere.
Who should buy it, and who shouldn’t
Buy it if you want the cheapest sensible route onto 12GB and you see the software as part of the hobby. It suits first-time local-AI builders on a tight budget, people who mainly want a 7B or 8B assistant, and anyone building a quiet low-power machine. The improving Intel stack is a bet on the platform getting better, and it has been.
Skip it if you want speed, a big-model machine, or software that just works out of the box. A used 3060 12GB gives you the same VRAM on mature CUDA if hassle-free matters more than watts and warranty. But for the money, the B580 gets you onto 12GB cheaper than almost anything, and that counts for a lot at the entry level.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Everyday small model
Llama 3.1 8B or Qwen2.5 7B at Q4 with IPEX-LLM or a recent Ollama
Fits inside 12GB with a decent context; the win over an 8GB card is running these without a tight squeeze.
Intel-native path
IPEX-LLM or the OpenVINO runtime rather than CUDA-based tooling
Intel's own stack is the smoother route on Arc; generic CUDA-first tools will not use the card.
Low-power box
Stock 190W, single 8-pin
Modest draw and a compact dual-slot card make it easy to build a quiet, cheap always-on assistant.
What owners report
Real first-hand experience gathered from owners and the community.
- “
Builders have run DeepSeek and other models on the B580 through Intel's IPEX-LLM and OpenVINO paths, and even paired two cards for 24GB, though the setup takes more effort than a CUDA card.
1 claim(s) we couldn't fully verify
- · indicative price around GBP 250 / USD 260 (mid-2026) - No primary source (Intel or TechPowerUp) tracks street pricing; figure is indicative only.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
▶Xiao Yang
Running DeepSeek R1 on Intel Arc B580: Setup and Performance
▶YourAvgDev
Gen AI on Intel Arc GPUs - Building a Dual Arc B580 LLM Inference Server
- Intel Arc B580 SpecsTechPowerUp · primary source
Common questions
Is the Arc B580 any good for local LLMs?+
For the money, it's interesting. Twelve gigabytes at around the price of an 8GB card means you can run 7B to 8B models with a decent context, which a cheaper card can't. The software is the catch. You'll go through IPEX-LLM, OpenVINO or a recent Ollama build rather than the CUDA path everything else assumes.
How does it compare to a used 3060 12GB?+
Same 12GB, different trade-offs. The 3060 rides mature CUDA drivers, so it just works and has a huge software base. The B580 is newer, more efficient, and often cheaper new with a warranty, but its AI stack is younger. For zero hassle, the 3060. For a new card and lower power, the B580, if you'll tolerate the setup.
What models fit in 12GB?+
A 7B or 8B at Q4 sits comfortably with a good context, and that covers most assistant and light coding work. A 13B to 14B fits at a tighter quant. A 32B is out without heavy offload. It's an entry card, and 12GB is the point of it.