
M5 Ultra Mac Studio Review: A Leap Forward for Local AI Agents
MacStories spent four days testing the 256 GB M5 Ultra Mac Studio against the M3 Ultra and an RTX 5090 PC. The review reports prompt processing up 150% on average, which makes locally hosted agents practical.
MacStories has published a review of the M5 Ultra Mac Studio, the machine that currently tops Apple's desktop line. Author Federico Viticci tested the 256 GB configuration for four days against its predecessor, an M3 Ultra with 512 GB of memory, and his own gaming PC with an RTX 5090 inside. His conclusion: this is the first Mac that makes locally hosted AI assistants feel practical rather than experimental.
A new architecture in an unchanged shell
The M5 Ultra looks identical to the M3 Ultra model it replaces, but its silicon is new: Apple uses UltraFusion to join two dual-die M5 Max chips into a quad-die design, a first for the company. The GPU has 80 cores, each paired with a Neural Accelerator, giving up to 4.5 times the peak AI compute of the M3 Ultra. Unified memory, which local MLX models rely on, is still offered at 256 GB and 512 GB, and bandwidth has grown from 819 GB/s to 1.2 TB/s, or 50% more. The 512 GB model is expected in late October.
The measured gains
Across four days of tests on macOS 27 with oMLX 0.7.0.dev2 and Qwen3.8-Flash-Next, prompt processing was up 150% on average, roughly 2.5 times what the reviewer's previous setup could do: a 16,000-token prompt was read at 2,887 tokens per second against 1,143 on the M3 Ultra. Generation speed in the same comparison was about 70% higher, and the charts show 91 to 75 tokens per second as the context grows from 4K to 256K. On a cold-cache 256K prompt, the M3 Ultra needed about 245 seconds to reach its first visible token; the M5 Ultra needed 102. With three requests at once (~6.5K tokens each), the newer machine produced 81.5 tokens per second of combined output, 23% more than a single request, while the M3 Ultra gained 4%.
The RTX 5090 still leads, with limits
NVIDIA's card remains faster as long as the model fits in its memory: on a 6,000-token prompt it read at about 3,000 tokens per second against roughly 1,700 for the M5 Ultra, and held a 25% lead in generation thanks to 1.79 TB/s versus 1.2 TB/s of bandwidth. Its ceiling is 32 GB of VRAM: with an 8-bit attention cache it finished long contexts at 49.6, 40.4 and 30 tokens per second at 64K, 128K and 256K, but once it had to borrow system memory over PCIe those rates fell to 4.6, 2.9 and 1.5. The reviewer also notes that the PC is far larger, louder and hotter, while the Mac Studio's fan stays inaudible in ordinary local-model use.
Why local agents matter to the author
Viticci's interest is practical. His research setup for this summer's iOS and iPadOS 27 review relied on agents that ran around the clock for 99 days over a project of 310 documents, at a total cost of $0 — a workload he says would have been prohibitive through cloud APIs. He has made the local model his default in the Open Minis app for iOS. He calls 5-bit quantization the sweet spot on the 256 GB model.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.