whichaipc

GPU for local LLM · upper-mid

NVIDIA GeForce RTX 5080

Blistering speed on models that fit, boxed in by 16GB when they don't.

VRAM
16 GB
TDP
360 W
Bandwidth
960 GB/s
£ / GB
£63

What it runs

With 16 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~28B params

Verdict

One of the fastest new cards for inference, on the same 16GB as cards costing a third as much.

BEST FOR

  • + fast 7B to 14B work
  • + new-with-warranty buyers
  • + people whose models fit 16GB

NOT FOR

  • - big-model headroom
  • - VRAM-per-pound value
  • - anyone who'll want 24GB

The RTX 5080 is one of the fastest cards you can buy new, and for local LLMs it arrives with a frustrating asterisk: 16GB. The GDDR7 memory and Blackwell silicon are seriously quick, but the capacity is the same as cards costing a third as much. Great for models that fit, limiting for the ones that don’t.

The take

Buy it for blistering speed on models up to about 14B, on a warranty, if the 16GB ceiling suits your work. The 960 GB/s of GDDR7 is roughly the fastest inference bandwidth at this price, and it shows in every session. But at around £1000, paying flagship-adjacent money for 16GB stings when a used 3090 gives you 24GB for a third of that. This is a card you buy because it’s fast and new, not because it’s sensible VRAM value. If your models fit in 16GB, though, very few new cards feel quicker.

What it’ll actually run

Sixteen gigabytes comfortably runs 8B to 14B models at Q4 with a generous context, and it runs them very fast. A 32B at a tight quant will fit if you keep the context modest, but you’re managing memory rather than relaxing into it. The moment you want a big model with room to breathe, 16GB is the wall you meet.

Where models fit, the speed is the point. Expect something like 70 to 100 tokens/sec on a 13B at Q4, among the quickest you’ll see on a single consumer card, and it holds that pace through long context and heavier prompts where slower cards start to sag. On value, roughly £1000 for 16GB is about £63 per gigabyte, the steepest £/GB-VRAM figure here by a distance; you’re paying for pace and a warranty, full stop. Software is current Blackwell CUDA, well supported across the major inference stacks, though a brand-new architecture occasionally wants an up-to-date build of your tools before everything clicks into place.

Who should buy it, and who shouldn’t

Buy it if you want the fastest new 16GB card, you value a warranty and current drivers, and your models really do live under 16GB. For quick 7B to 14B work it’s superb, and it doubles happily as a top-tier gaming card when you’re not running models. If new and fast is the brief, this delivers.

Skip it if VRAM is your priority; a used 3090 or 4090 gives you 24GB and runs models the 5080 simply can’t hold, and for big-model work capacity beats speed every time. If you want new and you can stretch further, the 5090’s 32GB is the far more future-proof buy. But for pure single-card speed on models that fit, the 5080 is hard to beat.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

Fast small-model inference (fits 16GB)

Llama 3.1 8B in LM Studio

One owner measured around 132 tokens/sec on the desktop 5080, averaging only about 141W of the 360W budget.

13B to 14B at Q4

Qwen2.5 14B Q4 with generous context

Community benchmarks land near 45 tokens/sec and sit comfortably inside 16GB; this is the sensible fast ceiling.

32B model (the 16GB wall)

Qwen2.5-Coder 32B Q4 (~20GB) with partial CPU offload

Will not fit 16GB; one tester saw it spill to the CPU and collapse to around 5.6 tokens/sec. Keep to 14B or smaller for full-speed work.

What owners report

Real first-hand experience gathered from owners and the community.

  • On the desktop 5080 an owner measured about 132 tokens/sec on Llama 3.1 8B, faster than the same test on a mobile 5090, while drawing only around 141W average of the 360W limit.

    Alex Ziskind

  • The same tester loaded a roughly 20GB Qwen 32B Q4 that would not fit the 16GB card; part of it ran on the CPU and throughput collapsed to about 5.6 tokens/sec, a direct demonstration of the 16GB ceiling.

    Alex Ziskind

  • Reviewers repeatedly note that a 24GB card such as a used 3090 holds larger models the 5080 cannot, so for big-model work the cheaper card can be the more capable one even if it runs slower.

    Alex Ziskind

Fact-checked 18 Jul 20268 claims verified against primary sources.
3 claim(s) we couldn't fully verify
  • · slots: 2 (Founders Edition reference) - Set to the FE dual-slot spec (TechPowerUp c4217), consistent with the 5090 FE; many AIB triple-fan 5080s are 2.5-3 slots, so check the height of the specific card you buy.
  • · indicative price around GBP 1000 / USD 1100 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks street pricing; figure is indicative only.
  • · single-card tokens/sec figures (132 t/s on 8B, ~45 t/s on 14B Q4, ~5.6 t/s on a 32B that overflows VRAM) - Community and owner benchmarks; vary with quant, context, driver and inference stack.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Is 16GB enough on a card this expensive?+

That's the frustrating question. For 8B to 14B models with generous context, 16GB is plenty and the speed is superb. But paying flagship-adjacent money for the same capacity as a much cheaper card stings the moment you want a bigger model. If your work fits in 16GB, it's brilliant; if it doesn't, the price hurts.

RTX 5080 or a used RTX 4090 for local LLMs?+

It depends on what you value. The 5080 is new, fast, and warrantied, but only 16GB. The used 4090 has 24GB and runs models the 5080 can't hold, at a similar or lower price. For pure inference headroom the 4090 wins; for a new card with support, the 5080 does.

Will the 5080 run a 32B model?+

At a tight quant with a modest context, yes, but you're managing memory to do it. 14B at Q4 is the comfortable, fast ceiling. Treat 32B as possible rather than roomy on 16GB.