AI · GPU comparison
The best GPUs for local AI, ranked
For running models at home, VRAM decides what fits and bandwidth decides how fast it answers - gaming benchmarks barely come into it. Here's the whole field, ranked by our own weighted score, with price per gigabyte alongside as the value check.
How we rank. Top down: VRAM first, because it decides what you can load at all, then memory bandwidth, because that sets how fast the answers come. Our own weighted score - VRAM value, tokens per second, efficiency, drivers, thermals, longevity - sits in its own column, so sort by that, or by price per gig, whenever you'd rather.
Score = our weighted rubric out of 5. ~Q4 fit = largest dense model, roughly, at 4-bit. Prices indicative; cards link through to the full review.
The datacentre ceiling
The workstation cards - the A6000 and the 96GB PRO 6000 - now sit in the table above, ranked with the rest. Above even those are the datacentre parts. Nobody's buying an H200 for the study, but they're the reference ceiling: this is what the frontier runs on.
| Card | VRAM | Memory | Bandwidth | Notes |
|---|---|---|---|---|
| NVIDIA H100 reference | 80GB | HBM3 | 3350 GB/s | Datacentre part. Enormous bandwidth, rarely a sane home buy. |
| NVIDIA H200 reference | 141GB | HBM3e | 4800 GB/s | Frontier inference. Reference only - the top of the mountain. |
What will it actually run?
Try the VRAM calculator →| Model size | VRAM at Q4 | Cheapest card that clears it |
|---|---|---|
| 3-4B | ~3-4 GB | Any 8GB+ card |
| 7-8B | ~5-6 GB | RTX 3060 12GB |
| 13-14B | ~9-10 GB | RTX 3060 12GB / 4060 Ti 16GB |
| 27-32B | ~18-20 GB | RTX 3090 / 4090 / 5090 (24GB+) |
| 70B | ~40 GB | Dual 24GB, RTX PRO 6000, or a 128GB AI PC |
| 120B+ MoE | ~60-70 GB | RTX PRO 6000 96GB / multi-GPU / big unified memory |
Rule of thumb: a model needs roughly half a gigabyte of VRAM per billion parameters at 4-bit, plus a little headroom for context. Go over your VRAM and the overflow spills into system RAM, which works but slows right down. If a used RTX 3090 fits your models, it's still the value pick of the whole table.