AI PC · workstation
Apple Mac Studio (M3 Ultra)
The only desk-side box that loads 400B-class models: 512GB of unified memory at 819 GB/s, on macOS.
- Memory
- 512 GB
- Bandwidth
- 819 GB/s
- AI compute
- -
- £ / GB
- £19
What it runs
With 512 GB of memory you can load, at 4-bit, a model up to roughly
~1020B params
Verdict
The big-model gold standard: 512GB of fast unified memory in a silent desktop, if you can stomach the price.
BEST FOR
- + running very large models locally
- + people already on macOS
NOT FOR
- - value per pound
- - anyone who needs CUDA
The Mac Studio with the M3 Ultra is the machine you buy when nothing else will hold your model. 512GB of unified memory, fed at 819 GB/s - roughly triple the bandwidth of the Strix Halo and DGX boxes - in a silent desktop that draws less power than a games console. It’s the closest thing to a personal AI workstation that just works, and it’s priced to match.
The take
This is the gold standard for loading big, and the numbers back it up. 512GB means you can run 400B-class models locally that no other desk-side box comes near, and the 819 GB/s of bandwidth is the fastest of any machine here by a distance, so once a model fits, it answers quicker than the PC-based rivals could dream of. The trade is money and prompt processing. It’s around ten grand kitted out like this, and while token generation is quick, feeding it a very long prompt takes its time - owners running huge models notice the wait before the first token. Weigh that carefully before the money leaves your account.
What it’ll actually run
This is where the 512GB does something no rival can. A 70B at Q4 is trivial, it barely notices, and you can climb all the way to 400B-plus mixture-of-experts models like the full DeepSeek R1, which people really do run on this machine at home. Load a dense 70B at high precision, keep several big models resident at once, or hand a model a huge context window: the capacity is simply in a different league.
Speed splits two ways. Token generation rides that 819 GB/s and is quick even on large models. Prompt processing - the bit where it reads your input before replying - is the M3 Ultra’s soft spot, and on a very long prompt or a massive model you’ll wait for the first token. MLX is the stack to use, it’s matured nicely on Apple silicon, and the machine stays cool and near silent through all of it. The usual Apple caveat applies: no CUDA, macOS only, so Nvidia-specific tooling is out.
Who should buy it
Buy it if you need to run models that won’t fit anywhere else, and you’re happy on a Mac. For a researcher or developer prototyping frontier-size models at home, it’s the one box that does the job in silence, on your desk, without a cloud meter running.
Give it a miss if value is the deciding factor, or if your models fit in 128GB. A Strix Halo box loads a 70B for a fifth of the price, and if you don’t need more than 128GB you’re paying thousands for capacity you’ll never touch. But for the very biggest models, run locally and quietly, this is the machine. There isn’t really a second place.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Biggest model that fits
400B-class MoE (full DeepSeek R1, ~404B) at Q4
Runs locally where no rival can; token generation is quick, prompt processing is the wait.
Comfortable large model
70B Q4/Q8 dense, or a 120B MoE
Barely troubles the memory; leaves masses of headroom for context.
Fast inference
MLX quants via LM Studio
MLX beats GGUF/llama.cpp on Apple silicon; use it.
Order tip
The 32-core M3 Ultra is required for the 512GB option
The base 28-core M3 Ultra caps at 256GB; you need the top chip for 512GB.
What owners report
Real first-hand experience gathered from owners and the community.
- “
Owners run the full DeepSeek R1 (a ~404B mixture-of-experts model) locally on the 512GB M3 Ultra, which no other desk-side machine can hold. Token generation is usable, but prompt processing is slow on very long inputs.
- “
Several testers conclude the 512GB M3 Ultra is superb for holding huge models but not the value pick for inference, with prompt-processing speed and the roughly ten-grand price the main gripes.
- “
819 GB/s is the M3 Ultra figure; don't confuse it with the M4 Max's 410/546 GB/s, which is a different, lower-memory chip in the same chassis.
3 claim(s) we couldn't fully verify
- · power_w 240 - Apple only publishes a 480W maximum continuous whole-unit rating; 240W is an approximate GPU-inference-load figure, not a vendor number, and real draw varies with the model.
- · Price around GBP 9,499 / USD 9,499 - Indicative for a 512GB configuration; a fully maxed 512GB + 16TB unit runs about USD 14,099. Apple pricing varies by store and storage.
- · compute_tops null - Apple does not headline an AI TOPS figure for the M3 Ultra, and the Neural Engine isn't used for LLM inference here, so it's left null.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
- Mac Studio (2025) - Tech Specs (vendor spec)Apple · primary source
- Mac Studio with M3 Ultra runs massive DeepSeek R1 locallyMacRumors · forum
- Owner thread: 512GB M3 Ultra for local LLMsr/LocalLLaMA · thread
Common questions
What size model can the Mac Studio M3 Ultra run?+
This is the machine's whole reason to exist. 512GB of unified memory holds 400B-class mixture-of-experts models like the full DeepSeek R1 locally, which no other desk-side box comes near. A 70B at Q4 is trivial, and you can keep several large models resident at once.
Why is prompt processing slow on the M3 Ultra?+
Token generation rides the 819 GB/s bandwidth and is quick even on big models. Prompt processing - the bit where it reads your input before replying - is the M3 Ultra's soft spot, so on a very long prompt or a massive model you'll wait for the first token.
Is 512GB worth it over a 128GB machine?+
Only if your models exceed 128GB. If they fit in 128GB, a Strix Halo box loads the same models for a fraction of the price. The 512GB M3 Ultra earns its money purely when nothing smaller can hold what you're running.
Can it run CUDA models?+
No. There's no Nvidia CUDA on Apple silicon. You run models through MLX or llama.cpp, and anything that specifically needs CUDA won't run here.
