Benchmarks
How fast each machine answers
Generating a token means reading the whole model out of memory, so for a dense model the bandwidth-to-size ratio sets the ceiling. Nothing beats it, and real output lands somewhere below it. The table below works that ceiling out for an 8B model at 4-bit, which is the yardstick most people compare on.
Read this first
These are calculated ceilings, not stopwatch numbers. Your serving stack, batch size, context length and quantisation all pull the real figure down, usually to somewhere between half and three-quarters of the ceiling. Measured runs from my own bench are being added card by card, and each product page states which of its numbers are mine and which are sourced.
Measured, not calculated
36 stopwatch runs on dual RTX 4090s
Twelve model and quantisation configurations timed on the same harness: decode speed, four-way concurrency, prefill rate and time to first token. An 80B coder at 118 tokens per second across two cards, 4-bit beating FP8 four times out of four, and the one comparison we will not publish yet because the run was contaminated.
See the measured numbers →
| Machine | Type | Memory | Bandwidth | Fits @Q4 | 8B ceiling | Price |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | GPU | 32 GB | 1792 GB/s | ~60B | ~448 tok/s | £1,900 |
| NVIDIA RTX PRO 6000 Blackwell | GPU | 96 GB | 1792 GB/s | ~188B | ~448 tok/s | £7,500 |
| NVIDIA GeForce RTX 3090 Ti | GPU | 24 GB | 1008 GB/s | ~44B | ~252 tok/s | £700 |
| NVIDIA GeForce RTX 4090 | GPU | 24 GB | 1008 GB/s | ~44B | ~252 tok/s | £1,400 |
| AMD Radeon RX 7900 XTX | GPU | 24 GB | 960 GB/s | ~44B | ~240 tok/s | £850 |
| NVIDIA GeForce RTX 5080 | GPU | 16 GB | 960 GB/s | ~28B | ~240 tok/s | £1,000 |
| NVIDIA GeForce RTX 3090 | GPU | 24 GB | 936 GB/s | ~44B | ~234 tok/s | £650 |
| NVIDIA GeForce RTX 5070 Ti | GPU | 16 GB | 896 GB/s | ~28B | ~224 tok/s | £700 |
| Apple Mac Studio (M3 Ultra) | AI PC | 512 GB | 819 GB/s | ~1020B | ~205 tok/s | £9,499 |
| NVIDIA RTX A6000 | GPU | 48 GB | 768 GB/s | ~92B | ~192 tok/s | £3,200 |
| NVIDIA GeForce RTX 3080 | GPU | 10 GB | 760 GB/s | ~16B | ~190 tok/s | £350 |
| NVIDIA GeForce RTX 4070 Ti SUPER | GPU | 16 GB | 672 GB/s | ~28B | ~168 tok/s | £600 |
| Apple Mac Studio (M4 Max) | AI PC | 128 GB | 546 GB/s | ~252B | ~137 tok/s | £3,399 |
| Intel Arc B580 | GPU | 12 GB | 456 GB/s | ~20B | ~114 tok/s | £250 |
| NVIDIA GeForce RTX 3060 12GB | GPU | 12 GB | 360 GB/s | ~20B | ~90 tok/s | £230 |
| NVIDIA Tesla P40 | GPU | 24 GB | 347 GB/s | ~44B | ~87 tok/s | £250 |
| NVIDIA GeForce RTX 4060 Ti 16GB | GPU | 16 GB | 288 GB/s | ~28B | ~72 tok/s | £400 |
| ASUS Ascent GX10 (NVIDIA GB10) | AI PC | 128 GB | 273 GB/s | ~252B | ~68 tok/s | £2,999 |
| Apple Mac Mini (M4 Pro) | AI PC | 64 GB | 273 GB/s | ~124B | ~68 tok/s | £1,999 |
| NVIDIA DGX Spark | AI PC | 128 GB | 273 GB/s | ~252B | ~68 tok/s | £3,399 |
| Framework Desktop (Ryzen AI Max+ 395) | AI PC | 128 GB | 256 GB/s | ~252B | ~64 tok/s | £1,699 |
| Beelink GTR9 Pro (Ryzen AI Max+ 395) | AI PC | 128 GB | 256 GB/s | ~252B | ~64 tok/s | £3,999 |
| GMKtec EVO-X2 (Ryzen AI Max+ 395) | AI PC | 128 GB | 256 GB/s | ~252B | ~64 tok/s | £1,699 |