whichaipc

GPU for local LLM · value

NVIDIA GeForce RTX 3090 Ti

The fastest 24GB Ampere card, and the hungriest, for anyone chasing VRAM per pound on the used market.

VRAM
24 GB
TDP
450 W
Bandwidth
1008 GB/s
£ / GB
£29

What it runs

With 24 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~44B params

Verdict

The quickest way onto 24GB of Ampere, at the cost of 450W and used-market risk.

BEST FOR

  • + single-card 32B models
  • + fastest Ampere memory
  • + VRAM-per-pound hunters

NOT FOR

  • - low-power always-on boxes
  • - quiet small builds
  • - warranty buyers

The 3090 Ti is the plain 3090’s louder, thirstier sibling. Same 24GB, the fastest memory Ampere ever shipped, and a 450W appetite to match. On the used market it’s a niche pick. You buy it when you want the quickest 24GB Ampere card and you’re willing to feed it. For most people the ordinary 3090 is the wiser call, but the Ti has its buyers.

The take

The extra over a 3090 is small and comes at a cost. Memory bandwidth rises to 1008 GB/s, which nudges token speed up a touch, but board power jumps to 450W and used prices usually sit above a plain 3090. So the value maths is worse even though the card is faster. If you power limit it, which you should, the gap to a 3090 narrows further. Buy it for the speed, with eyes open about the power bill.

What it’ll actually run

Twenty-four gigabytes is the reason to be here. A 32B at Q4 sits in VRAM with a workable context, and that’s the class of model most home users actually want. Drop to a 13B or an 8B and it flies, fully resident with room for a long context. The 3090 Ti runs these a shade quicker than a 3090 thanks to the extra bandwidth, though you’d struggle to feel it in normal use.

Speed sits comfortably ahead of the 16GB cards on anything memory bound, and the 24GB lets you run models they can’t hold at all. A 70B is the usual story. One card can’t hold it without offload, so pair two for 48GB and split the layers. That works well, but two 450W triple-slot cards is a serious build.

The catch is everywhere you look at power and heat. Four hundred and fifty watts is a lot for the performance on offer, and the GDDR6X runs hot. Cap it at 300W and you lose next to no token speed while saving real watts. I would never run one of these stock.

Who should buy it, and who shouldn’t

Buy it if you specifically want the fastest 24GB Ampere card and you find one priced sensibly, and you’re happy to power limit and provide cooling. Used-market rules apply. Check seller history, ask for proof it works, test the memory on arrival, and expect to re-pad a card that lived in a mining rig.

Skip it if you want value, in which case a plain 3090 gives you the same 24GB for less money and less heat. Skip it too if you want low idle power, a quiet build, or a warranty, because 450W and triple-slot coolers are none of those things. But for the person chasing the quickest Ampere 24GB, this is it.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

Single-card 32B

Qwen2.5 32B at Q4_K_M, full 24GB

Holds in VRAM with usable context; expect roughly 20-30 tokens/sec depending on quant and context.

Efficient inference

Power limit 300W (from 450W default)

Generation is bandwidth bound, so capping core power saves around 150W with negligible token-speed loss.

70B across two cards

Two 3090 Tis, 48GB pooled, Q4_K_M, layers split over PCIe

48GB holds a 70B at Q4 with context; plan cooling and PSU headroom for two 450W cards.

What owners report

Real first-hand experience gathered from owners and the community.

  • Dual 24GB Ampere setups are widely described as the value route to 70B models, pooling 48GB across two cards split over PCIe, with NVLink optional.

    r/LocalLLaMA

Fact-checked 19 Jul 20269 claims verified against primary sources.
2 claim(s) we couldn't fully verify
  • · typical used price around GBP 700 / USD 850 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks used-market pricing; figure is indicative only.
  • · single-card 32B tokens/sec figures (roughly 20-30 t/s) - Community benchmarks; vary with quant, context, driver and inference stack.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Is the 3090 Ti worth more than a plain 3090 for AI?+

Only a little. Both carry 24GB and run the same models. The 3090 Ti has higher memory bandwidth, 1008 GB/s against 936, so it's marginally faster, but it also pulls 450W against 350 and usually costs more used. For pure value the plain 3090 is the smarter buy; the Ti is for people who want the fastest Ampere going.

How much power does it really draw?+

A lot. Board power is 450W stock, and it'll take it. The good news is inference is memory-bandwidth bound, so a power limit around 300W costs almost no token speed while saving real heat and watts. I'd run one capped, not stock.

Can two 3090 Tis run a 70B model?+

Yes. Two cards pool 48GB, which holds a 70B at Q4 with context. Just plan the build carefully, because two 450W triple-slot cards need airflow, PCIe spacing and a large PSU. This is not a tidy setup.