whichaipc

AI PC · desktop

Framework Desktop (Ryzen AI Max+ 395)

128GB of unified memory on plain x86, for roughly half what the rivals cost.

Memory
128 GB
Bandwidth
256 GB/s
AI compute
50 TOPS
£ / GB
£13

What it runs

With 128 GB of memory you can load, at 4-bit, a model up to roughly

~252B params

Verdict

The value pick of the unified-memory machines: same 128GB, plain x86, half the price.

BEST FOR

  • + most model per pound
  • + people who want Windows or Linux

NOT FOR

  • - raw speed
  • - anyone who needs CUDA

The Framework Desktop takes the interesting bit of the Strix Halo chip - up to 128GB of memory the graphics side can treat as its own - and drops it into a proper little x86 desktop you can actually tinker with. For anyone who wants big-model headroom without leaving Windows or Linux behind, this is the friendly option.

The take

This is the value pick of the unified-memory crowd, and it’s not close. For roughly half the price of a DGX Spark or a 128GB Mac Studio you get the same headline 128GB to load models into, on hardware that runs everything x86 does. It isn’t the fastest of the bunch - the memory tops out around 256 GB/s, so it shares the same bandwidth ceiling as the Spark - but for the money, the amount of model you can load is frankly brilliant. If you want the most capacity per pound and you don’t fancy learning a new operating system to get it, start here.

What it’ll actually run

The Radeon 8060S iGPU borrows from that shared pool, so with the memory split set generously in the BIOS you can hand it most of the 128GB. That means a 70B model at Q4 fits with room to spare, and you can sit comfortably in the 30B to 70B range that most people top out wanting at home. Smaller models in the 8B to 14B bracket run nicely and leave you plenty of memory for context.

Speed is the caveat, same as its rivals. At around 256 GB/s the bandwidth is the bottleneck, so a large model answers at a measured pace rather than instantly - expect single figures to low double figures of tokens per second on a 70B, quicker as you drop to smaller models. There’s a 50 TOPS NPU on board too, though for local LLMs it’s not much use yet - no inference tooling had made proper use of it at the time of writing, so the graphics side does the actual work. The software picture is the reassuring part: it’s plain x86, so Windows and Linux both run, and the ROCm stack keeps improving for the AMD graphics side. You will need to nudge the memory allocation in the BIOS to get the most out of it, but that’s a one-time bit of fiddling.

Who should buy it

Buy it if value and flexibility are what you’re after. It’s the machine for the tinkerer who wants a real desktop that runs anything, loads big models, and doesn’t cost flagship money to do it. The Framework side of things means it’s more serviceable than most mini machines, too, which suits this site’s crowd nicely.

Look elsewhere if you need raw speed or CUDA. If your models fit in 24GB, a used 3090 will be faster for less, and if you’re wedded to Nvidia’s tooling this isn’t your box. But for the most model you can load without spending silly money, on an OS you already know, the Framework Desktop is hard to beat.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

Biggest model that fits

gpt-oss-120b (MoE, ~60GB); BIOS memory split set high

Around 33 tokens/s and comfortably usable; up to 112GB of the 128GB can be handed to the GPU.

Dense large model

Llama 70B Q4 (~40GB)

About 5 tokens/s; it fits, but a dense 70B is slow here.

Everyday use

8B-32B Q4 (Qwen2.5 14B, Llama 8B)

Qwen2.5 14B around 23 tokens/s, Llama 8B around 21; comfortable for daily work.

Backend tip

llama.cpp with the Vulkan backend

Vulkan beat ROCm on Strix Halo in testing; Ollama defaults can run CPU-only, so force the iGPU.

What owners report

Real first-hand experience gathered from owners and the community.

  • Single Framework Desktop 128GB tests: Llama 3.1 70B Q4 dense about 5 tokens/s, gpt-oss-120b (MoE) about 33, gpt-oss-20b about 45, Qwen2.5 14B about 23. Power draw sat around 97-140W through these runs.

    Jeff Geerling / ai-benchmarks

  • As of testing, the 50 TOPS Strix Halo NPU could not be used for local LLM inference (no tooling had cracked it), and the Vulkan backend outran ROCm.

    Jeff Geerling / ai-benchmarks

  • The 256 GB/s is the theoretical figure; real-world measured bandwidth on Strix Halo comes in lower, around 210-215 GB/s, and it is what caps token speed on large models.

    Framework community forum

Fact-checked 18 Jul 20264 claims verified against primary sources.
2 claim(s) we couldn't fully verify
  • · Memory bandwidth of 256 GB/s in practice - 256 GB/s is the theoretical maximum; community measurements land nearer 210-215 GB/s, so real-world bandwidth is lower than the quoted figure.
  • · Power draw of 140W (power_w) - Reflects measured LLM-load draw (Geerling saw 97-153W); the unit ships with a larger PSU, so 140W is a load figure rather than the supply rating.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

How much memory can the Framework Desktop give to models?+

Up to 128GB is shared between the CPU and the Radeon 8060S graphics. You set the split in the BIOS, and with it set generously the graphics side can take the lion's share, which is what lets a 70B model at Q4 fit with room for context.

Is the Framework Desktop fast for local LLMs?+

It loads big models well but answers at a measured pace. Memory bandwidth tops out around 256 GB/s, the same ceiling as the DGX Spark, so a 70B lands single figures to low double figures of tokens per second. Smaller models feel quicker.

Does it need CUDA to run AI models?+

No. It's an AMD machine, so it runs on ROCm and standard tools like llama.cpp and Ollama on Windows or Linux. If your workflow depends specifically on Nvidia CUDA, this isn't the box for you.