M5 Ultra Mac Studio: Apple Claimed 4.3x, Reviewers Measured 2.5x
M5 Ultra Mac Studio: Apple Claimed 4.3x, Reviewers Measured 2.5x
The review embargo on the M5 Ultra Mac Studio lifted yesterday, the machine starts shipping today, and I've spent the last several hours reading every hands-on I could find — mostly to answer one question: does the 1.2 TB/s of memory bandwidth actually change what you can run on your desk?
My August post about this machine ended with a promise to do a real hands-on the moment I could get a 512 GB unit. I still don't have one. What I have instead is better than a spec sheet: four reviewers who measured the thing properly, on the exact local-model workloads I care about, with numbers I can cross-check against each other.
So let me put the answer up front. Apple's headline number — up to 4.3x faster AI — is a peak-compute figure that no local model will ever actually hit. The number that does the real work is memory bandwidth, and that part is true: reviewers measured prompt processing about 2.5x faster than the M3 Ultra and token generation around 70% faster. The CPU lands almost exactly on Apple's claim. The weak spot isn't the silicon at all — it's the receipt.
What reviewers actually got their hands on
The units that went out are not the base model, and that matters when you read any of these numbers. The Verge tested a 36-core CPU / 80-core GPU configuration with 256 GB of memory and a 4 TB SSD at $12,299. Engadget had the same chip with a 2 TB drive for $11,299. Tom's Hardware ran local models on the 256 GB / 4 TB machine, and MacStories' Federico Viticci did the deepest local-AI work of the group on the 256 GB model, in MLX, against both his M3 Ultra and his RTX 5090 PC.
The Verge's review unit — identically sized to every Mac Studio since 2022, with the M5 Ultra's front Thunderbolt 5 ports and an SDXC slot (image: The Verge / Antonio G. Di Benedetto).
| Config | What you get | US price |
|---|---|---|
| M5 Ultra, base | 30-core CPU / 64-core GPU, 96GB, 1TB SSD | $5,499 |
| M5 Ultra, 36-core | 36-core CPU / 80-core GPU, 96GB, 1TB | from $6,799 |
| M5 Ultra, 256GB | same chip, 96GB → 256GB memory | +$4,000 |
| Reviewed units | 36-core CPU / 80-core GPU, 256GB, 2–4TB | $11,299–$12,299 |
| M5 Ultra, 512GB | the configuration that makes this machine famous | late October, no price yet |
Apple's US pricing for the M5 Ultra Mac Studio, as published on August 25 and itemized by reviewers on September 21.
The 512 GB option — the one everybody actually wants to talk about, and the one that can hold a frontier-class open model plus its context — still isn't orderable. It arrives in late October with no announced price, and given that 96 GB to 256 GB costs $4,000, I would not plan a budget around a pleasant surprise.
The 4.3x is not a tokens-per-second number
Here's the claim-versus-reality chart I wish Apple had published alongside its own. Apple told us the M5 Ultra is up to 1.25x faster single-core, 1.3x faster multi-core, 1.8x faster at graphics, twice as fast at storage, and up to 4.3x at AI compute versus the M3 Ultra it replaces.
Reviewers put numbers on every one of those, and the boring claims held up almost exactly. Single-core: 1.26x in The Verge's Geekbench 7 run. Multi-core: 1.30x, essentially a bullseye. Storage read: 2.08x. Even the graphics claim is close enough at 1.63x on Cinebench 2026's GPU test.
The AI number is where the marketing splits from the measurement. "Up to 4.3x" is peak matrix-multiply throughput with every GPU core busy and the data already where it needs to be. A running language model never looks like that. It re-reads its weights from memory for every single token, which is why MacStories measured the real-world gain at 2.5x on prompt processing and 1.7x on generation — a big win, and nowhere near 4.3x.
Apple's claims versus what reviewers measured, both against the M3 Ultra. The CPU and storage claims are honest; the AI claim is peak compute, not tokens per second (chart: Open Source Factory; data: Apple, The Verge, MacStories).
| Benchmark | M5 Ultra (36C/80C) | M3 Ultra (32C/80C) | Gain |
|---|---|---|---|
| Geekbench 7 CPU single | 3,763 | 2,975 | 1.26x |
| Geekbench 7 CPU multi | 52,441 | 40,375 | 1.30x |
| Cinebench 2026 GPU | 133,119 | 81,731 | 1.63x |
| Blender classroom scene | 11 sec | 15 sec | 1.36x |
| Sustained SSD read | 14,868 MB/s | 7,160 MB/s | 2.08x |
| Prompt processing (MLX, avg) | 2.5x | 1.0x | +150% |
| Token generation (MLX, avg) | 1.7x | 1.0x | +70% |
Both machines at 256 GB of memory, both tested by The Verge except the last two rows, which are MacStories' averaged MLX results indexed to the M3 Ultra at 1.0x — averages, not single runs.
One more honest data point the marketing doesn't include: PCMag's HandBrake conversion finished in 87 seconds, three seconds slower than the M3 Ultra comparison unit. It's one test out of dozens, and everything else on that machine was faster, but it's a useful reminder that "newer and more expensive" isn't a benchmark.
What 1.2 TB/s actually buys you
If you only remember one thing about local AI hardware, make it this: capacity decides which models fit, bandwidth decides how fast they answer. A 4-bit 120B model is roughly 60 GB of weights, and the machine has to pull those weights through the memory bus for every token it writes. That's why the M5 Ultra's headline spec isn't its TOPS or its core count — it's 1.2 TB/s, 50% above the M3 Ultra and roughly 4.4x the DGX Spark.
Memory bandwidth across the boxes people actually compare in September 2026. The M5 Ultra isn't the fastest chip here — it's the fastest thing with 256 GB to 512 GB behind it (chart: Open Source Factory; data: vendor specs).
The measured results line up with the theory. MacStories clocked Qwen3.8-Flash-Next above 100 tokens per second on short prompts on the 256 GB M5 Ultra, and still writing at 60 to 85 tokens per second with 64K to 256K of context loaded behind it. Their prompt-processing figure for a 6,000-token prompt was around 1,700 tokens per second, versus about 3,000 for the RTX 5090.
That last comparison is the one to sit with, because it also defines the limit. The 5090 has more bandwidth (1.79 TB/s) and stays about 25% ahead on generation at every prompt size — until the model doesn't fit in its 32 GB of VRAM. Then it's offloading layers over PCIe to system memory, and the Mac keeps going. Tom's Hardware's review makes the same point from the other direction: on Qwen 3.8-27B at Q4, the M5 Ultra generated tokens nearly four times as fast as Nvidia's DGX Spark and about twice as fast as the outgoing M4 Max, with prompt processing also ahead of the Spark.
"If you've been skeptical of testing OpenClaw or Hermes Agent with local models because they'd never be even remotely near the intelligence and speed of cloud ones, this Mac will change your mind about that." — Federico Viticci, MacStories
The practical version, from the same review: a 99-day background agent run over 310 documents and PDFs, all of it local, for zero token cost. Engadget added a transcription number that made me laugh — a 70-minute podcast transcribed in 18 seconds, under a quarter of the time the M4 Max took. Tom's Hardware's reviewer went further and ran models much larger than their benchmark set through Hermes and LM Studio on the 256 GB machine, reporting strong performance, and noted the whole thing is untested in an MLX cluster so far. Quietly, that's the actual pitch for this machine: an always-on local agent box that doesn't heat your room or bill you per token.
The back: four Thunderbolt 5 ports, 10Gb Ethernet, two USB-A, HDMI 2.1, and the power inlet — plus two more Thunderbolt 5 ports behind the front panel on the Ultra (image: The Verge / Antonio G. Di Benedetto).
The price is the honest part of this review
Now the uncomfortable half. The M3 Ultra Mac Studio launched at $3,999 in 2025. Apple's June 25 price increase — the one driven by memory and storage costs, not by features — took that same machine to $5,299, the largest single jump in the Mac lineup. The M5 Ultra inherited that ladder and added $200 on top for a base price of $5,499, with the 96 GB to 256 GB step alone costing $4,000.
The configuration ladder, drawn from Apple's published prices and the reviewers' itemized upgrades. The 512 GB rung — the one this whole product exists for — has no price yet (chart: Open Source Factory).
A few harder facts to go with it. Nothing inside is upgradeable — memory is soldered, and the SSD is paired to the machine. There's no memory tier between 96 GB and 256 GB, so the jump from "comfortable" to "runs the big MoE" is a $4,000 cliff. A maxed 256 GB / 16 TB configuration runs $18,299, and the editors at PCMag expect the 512 GB option to land well over $20,000. Tom's Hardware also noted that the configuration they reviewed was backordered 16 to 18 weeks at review time — so "available September 22" is doing some heavy lifting.
Two more things I'd want to see before I call the cluster story real. Apple claims four Mac Studios linked over Thunderbolt 5 with RDMA deliver up to 3x the AI inference of a single machine; Engadget asked for four review units, was declined, and tested none of it. And PCMag's complaint is worth quoting in spirit: Apple ships a machine pitched at local AI and still publishes no playbooks or guides for it, so the entire software story — MLX builds, quantizations, offload tricks — comes from the community.
One more reality check, because a $12,299 desktop invites the wrong fantasy. In Tom's Hardware's CPU rendering tests the M5 Ultra lost to 64-core Threadripper workstations, and in Cyberpunk 2077 at 1080p with ray tracing it ran at 66 fps — roughly where an RTX 5070 gaming PC sits, and not playable at 4K. Buy it for the memory and the quiet, not for frames.
So should you buy one?
If your work is 8K timelines, high-end 3D rendering, or training on datasets that don't fit in 128 GB, yes — this is the only Mac that does it, and reviewers agree it does it fast and quietly. If what you actually want is a local model answering your agents without a cloud bill, buy the 36-core/256 GB configuration or don't buy one: the 96 GB base with 64 GPU cores is a different, weaker machine than the one in every benchmark you've read this week.
And if you're like me — running a 27B on a 1080 Ti and doing RAM math before dinner — the honest answer is that this machine is priced for people who bill for it. The interesting question isn't whether the M5 Ultra is good. It's whether 256 GB of unified memory at $12,299 pays for itself against your token bill, and that's an arithmetic question about your own usage, not a review score. I did that arithmetic for the cheaper boxes in my local AI value ranking, and the M5 Ultra sits in a different league because nothing else in the list can hold the model at all.
A review setup that's identical to the one from 2022 — which is the point. Apple changed the silicon, not the desk (image: CNET / Josh Goldman).
My take
I came into this week expecting the 4.3x to be the story and the price to be the caveat. It's the reverse. The bandwidth jump is real and it shows up exactly where local AI users feel pain — prompt processing, long contexts, and running an agent all day without listening to a fan. That's a meaningful upgrade for a machine class that has been memory-starved for two years, and it's why I wrote in August that this was the first Mac Studio that actually competes for local model work.
What I can't verify yet is the part I care about most. Every number in this post comes from machines with 256 GB of memory; the 512 GB configuration that makes "frontier-class open model on your desk" literally true has not shipped to a single reviewer. When it does — and when someone independent publishes cluster numbers instead of repeating Apple's 3x — I'll write that post. This one already has a sequel, and I'd rather promise that than pretend a $12,299 review unit arrived at my desk.
Sources: Apple Newsroom (Aug 25, 2026) announcements for Mac Studio and for M5 Ultra/M6; The Verge's M5 Ultra benchmark testing (Sep 21); Engadget's Mac Studio review; Tom's Hardware's M5 Ultra review; MacStories' M5 Ultra review; PCMag's Mac Studio (M5 Ultra) review; CNET's Mac Studio review; MacRumors and AppleInsider roundups; EveryMac and coverage of Apple's June 25, 2026 price increase; community discussion in r/LocalLLaMA. Every performance multiplier Apple published is vendor-reported; every measured figure is attributed to the outlet that ran it. I have not tested this machine — no review unit was provided to this blog, and this post contains no first-party benchmarks. The two charts and the config table are my own arithmetic from the published prices. Related reading on this blog: the August announcement analysis, the local AI value ranking, and how much RAM an MoE actually needs.
Comments
Post a Comment