Skip to content mac maxi Usable local AI is a lot of fun, but how much would you pay for it? Apple's M5 Ultra Mac Studio. Credit: Andrew Cunningham Apple's M5 Ultra Mac Studio.
Credit: Andrew Cunningham This spring, Apple made it official: The Mac Pro is dead . When Apple was using Intel processors and AMD GPUs, the idea of an expandable and upgradeable tower made sense. But Apple Silicon chips bring the CPU, GPU, and memory interface all under one roof in a way that makes that sort of upgradeability not just superfluous but impossible.
At least not without breaking some of the chips’ key benefits. And as it happens, Apple Silicon’s combination of a fast CPU, solid GPU, and unified memory architecture has made it a good fit for a modern “pro” workload that really pushes high-end hardware: running language models and agents locally. That’s really the core audience for the M5 Ultra Mac Studio, Apple’s tippity-top config that is jumping two processor generations in a single refresh (there was no M4 Ultra, if you recall).
The highest-end Studios have always felt like a bit of overkill for most people, but now there’s a particular kind of AI-pilled coder or designer who can legitimately benefit from having this kind of system on their desk. Apple was caught off guard by the demand for these systems from people running software like OpenClaw and running open-weight models locally, which, along with the AI-fueled memory shortage, is why it has been nearly impossible to buy most of these systems for months . That’s the context that the new Mac Studio is launching into.
Same design, new hardware Same design language as the Mac mini, but less mini. Andrew Cunningham The new Mac Studio looks the same as it did when it launched with the M1 Max and M1 Ultra back in 2022. It’s a small, squat box, with the same 7.7-by-7.7-inch footprint as the old Mac mini, but it’s taller.
Both the Max and Ultra variants of the Studio look the same, but the M5 Ultra version weighs two pounds more because it has heavier but more-conductive copper in its heatsink to help cool the more powerful chip. Both Studio models also come with nearly all the same ports. On the back, you’ll find four 120Gbps Thunderbolt 5 ports, a 10Gbps Ethernet port, an HDMI 2.1 port, and a pair of 5Gbps USB-A ports.
On the front, both Studios have a UHS-II SD card reader, but the Max comes with two 10Gbps USB-C ports, and the Ultra comes with two more 120Gbps Thunderbolt 5 ports. The Ultra’s additional oomph extends to external display support. The M5 Max supports up to five external displays, and the M5 Ultra supports up to eight , though both come with caveats depending on the resolution and refresh rates you’re using.
Display support details for the M5 Max and M5 Ultra Mac Studios. Credit: Apple Display support details for the M5 Max and M5 Ultra Mac Studios. Credit: Apple The M5 Max is a known quantity at this point, and Apple has been shipping it since March in the MacBook Pro .
It improved over the last-generation M4 Max by introducing new “super” cores for the CPU and nudging the maximum memory bandwidth up from 546GB/s to 614GB/s. The number of GPU cores stayed the same (40), but new neural accelerators built into each GPU core helped improve its performance for ML/AI-assisted technologies like MetalFX upscaling and frame generation. The M5 Ultra is a bigger departure for a few reasons.
The first is that there was no M4 Ultra, so we’re effectively hopping two processor generations despite only jumping a single product generation. CPU S/P/E-cores GPU cores RAM options Memory bandwidth Apple M5 Ultra (low) 10S/20P 64 96GB/256GB 1,228.8GB/s Apple M5 Ultra (high) 12S/24P 80 96/256GB/512GB (planned) 1,228.8GB/s Apple M3 Ultra (high) 24P/8E 80 128GB/256GB/512GB 819.2GB/s Apple M2 Ultra (high) 16P/8E 76 Up to 192GB 819.2GB/s Apple M1 Ultra (high) 16P/4E 32 Up to 128GB 819.2GB/s The second is that the way the chip is built has actually changed quite a bit. The M1 Ultra, M2 Ultra, and M3 Ultra were all essentially a pair of Max chips strapped together using a high-speed interconnect to get most of the benefits of making one huge chip without all of the manufacturing difficulties of making one huge chip.
The M5 Ultra is still basically a pair of M5 Max chips stitched together; that much is the same. The chip includes up to 12 super CPU cores, up to 24 P-cores, up to 80 GPU cores, 32 Neural Engine cores, two video decoding engines and four encoding engines, and a little over 1.2TB/s (yes, that’s TB with a T) of memory bandwidth. That’s a neat doubling of all the M5 Max’s major specs.
What’s different is that this generation’s M5 Pro and M5 Max are both already a pair of chiplets attached together via silicon interconnect, which means the M5 Ultra is actually four distinct bits of silicon packaged together. The potential downside for this arrangement is that die-to-die interconnects often can’t communicate quite as fast as components all housed on the same silicon. Overall, the Fusion Architecture hasn’t kept previous Ultra chips from being fast, and it doesn’t keep the M5 Ultra from being fast, either.
But performance on the Ultra chips has never scaled perfectly linearly with the number of cores, and the M5 Ultra is the same way. Also: New prices :( If you forget what the computer is called, look at the bottom. Credit: Andrew Cunningham If you forget what the computer is called, look at the bottom.
Credit: Andrew Cunningham Normally we could just cover the hardware improvements in a new Mac while assuming that all else was equal on the pricing front—that the new hardware would automatically be a better deal than the old hardware because it was being sold for the same price. But Apple has raised prices across the board this year because of the ongoing AI-fueled memory crunch. As its highest-end computer with its highest-end memory and storage configurations, the Mac Studio is getting hit even harder than other Macs.
At its introduction , the basic M4 Max Mac Studio ran $1,999, which got you a slightly cut-down version of the chip, 36GB of RAM, and 512GB of storage. The M5 Max Studio starts at $2,499, a $500 increase. An M3 Ultra Mac Studio with a cut-down chip, 96GB of RAM, and 1TB of storage started at $3,999; the same M5 Ultra config will run you a whopping $5,499, a $1,500 increase.
And that’s for the base models. Going for a fully enabled M5 Max, 64GB of RAM, and 1TB of storage—what I’d call the price/performance sweet spot for dabbling with local AI—will run you $3,799, $900 more than the M4 Max Studio would have cost when it launched. As for the Ultra?
A 256GB RAM/1TB storage model that would have cost $5,599 18 months ago now costs $9,499—not quite twice the price but close enough to feel like it. There will be a 512GB version of the M5 Ultra Studio, but Apple hasn’t said what it will cost. The M3 Ultra version started at $9,499, so I feel pretty confident in saying that the M5 Ultra version will be a $20,000-and-up computer.
To be interested in high-end local AI, you already have to be willing to pay more up front for hardware than you would to just use some company’s Nvidia-powered data center. At these prices, buying a cluster of Mac Studios feels like paying to build a data center out of your own pocket. Bear that in mind as we talk about performance.
General computing performance The M5 Ultra is, by almost every measure, the fastest Apple Silicon processor built so far. The only place it loses is in single-core CPU performance, where the Mac mini’s humble M6 very narrowly beats it. But the gap is small—smaller than the gap between the M3 and M4 generations, for example—which helps keep it from feeling like as big a deal as it did when the last-generation Studio launched with the M4 Max and M3 Ultra.
In many of our general-purpose CPU and GPU tests, the Ultra’s CPU outruns the M5 Max by 80 or 90 percent, and the GPU is between 50 and 80 percent faster. That’s essentially in keeping with what we’ve observed in past Ultra chips—the CPU comes closer to 2x scaling than the GPU does. Compared to the outgoing M3 Ultra, M5 Ultra usually posts around 30 percent faster single-core CPU speeds, 50 percent faster multi-core CPU speeds, and GPU performance that’s anywhere from 33 to 66 percent faster, depending on the test.
As you’d expect for a two-generation upgrade, it’s a big one. But we also observed some less-expected behavior. The Ultra’s Geekbench multicore performance is only around 26 percent faster than the Max, and in our CPU-based Handbrake video encoding test, the Max is actually faster to complete the H.264 encode (and barely slower at H.265).
Looking at the power consumption numbers offers a possible explanation. According to the powermetrics tool, both the M5 Max and M5 Ultra consume about 75 W of power on average during the video transcoding test. That’s also about the amount of power that the M3 Ultra Mac Studio used.
But the Max Mac Studios have historically used less power than the Ultra chips. To me, this suggests that the M5 Ultra is being power-limited to keep it within the power/cooling envelope of the existing Studio design. But it could just as easily be a bug.
When checking Activity Monitor during the video encoding job, you don’t see every single CPU core being 100 percent utilized, which is the typical behavior for this test. And powermetrics reports that all of the Ultra’s CPU cores are running at the same sustained clock speeds as the M5 Max’s for the duration of the encoding test. If there were power- or temperature-related throttling going on, you’d expect those clock speeds to be lower.
I’ve reported my findings to Apple and will report back if I get a response. Why AI people like Macs I have learned virtually every firsthand thing I know about vibe coding in the last week, so bear with me if I get any of this wrong. Many AI model providers, including Apple, have also released models that can work their way around some of these Hard Facts.
But broadly, this should be an accurate local-AI crash course. When you’re talking about running AI models locally, there are two hardware numbers that loom the largest. The first is the amount of GPU memory you have.
The second is the amount of memory bandwidth you have, or how quickly that GPU can communicate with the rest of the system. Your graphics RAM decides both the size of the models you can run—these need to be loaded pretty much entirely into GPU memory because anything after this will either spill over into your main system memory (slow) or, in a worst-case scenario, to your disk (even slower). On top of this, particularly for coding, you need additional memory for “context,” or the amount of information an individual agent can recall before running out of memory.
For a basic question-and-answer chatbot interaction, you don’t need a ton of context. For coding a full app, you’ll want to have a bunch of RAM available for context so the agent doesn’t get halfway through a complex task and “forget” what it was doing. There are handoff mechanisms—one agent can condense a session to a single file that contains most of the relevant information from a session and pass it to another, effectively resetting the context.
But things inevitably fall through the cracks with this mechanism, especially if there’s not all that much context to condense in the first place. The memory bandwidth isn’t the only thing that determines how quickly the model will be able to spit out new words (“tokens”), but it’s probably the single most important thing. Even the speed and capabilities of the GPU that the RAM is attached to don’t matter that much, compared to the amount of memory you have and how quickly the computer can move data into and out of it.
Apple Silicon Macs have become popular for running these models because they’ve hit a sweet spot. They don’t provide as much memory bandwidth as a dedicated desktop GPU (a 5-year-old Nvidia GeForce RTX 3060 12GB offers 360GB/s of bandwidth, almost 20 percent more than the M5 Pro; an RTX 5070 12GB offers 672GB/s, more than M5 Max). But a 64GB Apple Silicon Mac offers more GPU memory than any single consumer GPU you can buy, its memory bandwidth is fast enough , and it’s dramatically more power-efficient than an RTX 5090 or a pair of RTX 3090s.
And 96GB, 128GB, 256GB, and 512GB Apple Silicon Macs open the door to even larger, frontier-class models. Obviously, those top-tier systems aren’t cheap right now, but this hopefully explains part of the reason it’s suddenly difficult to buy a Mac mini, of all things. And partly because the hardware is so well-suited to these workloads, Apple’s MLX framework has become reasonably well-optimized and supported by different language models and local AI-focused apps like LM Studio Bionic .
On top of these stats, you also have something called “time to first token” (TTFT)—the amount of time that passes between you handing a prompt to a model and the model processing it and responding. This is one place where the M5 generation improves on M4 and older: The neural accelerators Apple added to the GPU dramatically speed up the TTFT, making any locally run model feel more responsive. Add this to the fact that the M5 Ultra dramatically boosts memory bandwidth compared to the M3 Ultra, and you get an idea of why this particular Mac Studio is well-suited to this kind of work.
The fact that it’s a two-generation jump over last year’s M3 Ultra makes it look even better than the M5 Max Studio or the M5 Pro Mac mini. I am become vibe coder This is Bankee Doodle, the impeccably designed front-end for my vibe-coded small-business bank-statement-PDF-to-spreadsheet converter. Credit: Andrew Cunningham This is Bankee Doodle, the impeccably designed front-end for my vibe-coded small-business bank-statement-PDF-to-spreadsheet converter.
Credit: Andrew Cunningham Looking beyond the standard suite of benchmarks, I have been trying some local vibe coding on the Studio, the first time I’ve tried it. On the recommendation of a few people who are a bit further down this road, I downloaded LM Studio Bionic and the Qwen 3.8 27B model. When it comes to coding and development, I barely know enough basic HTML and CSS to tweak an existing template, but my desire to learn more has always exceeded my ability.
Blame the ADHD, blame the lack of free time, blame whatever it is in my brain that started bouncing off the one coding class I tried to take just a couple of lessons after “Hello world!” I could just never make it stick. Using this M5 Ultra Mac Studio and LM Studio Bionic to run an all-local coding-capable language model is the most fun I’ve had with any piece of new technology in a long time. And I say this as someone who loathes generative AI in a creative context—in general, I think generative AI’s capabilities have been massively oversold and that it was built in an extractive, exploitative way, cribbing the work of millions of artists, journalists, researchers, and regular old Internet posters.
But there is a reason why vibe coding has taken off in the last year or two. If you know enough to tell the system what you want, you can make some genuinely useful things that actually function . I’ve been part of a book podcast for almost a decade and a half, and I handle our accounting—a side gig to my side gig.
For years, this has involved the hand-entry of figures from bank statements into a spreadsheet I built to track our income, expenses, and estimated tax burden. It was too light a job to be worth paying for QuickBooks for the rest of eternity—the software does too much we don’t need. It’s one of those probably solvable problems that’s consistently irritating when you’re thinking about it, but you run into it just infrequently enough to never get around to fixing it.
And now it’s fixed! First, I had the Qwen model write a Python script to convert bank statement PDFs into spreadsheets and format and sort them the way I wanted. I then built a web app around it so I could easily upload future sheets instead of messing with the LLM or the Python script.
I’m pretty conflicted about this, but there is an undeniable appeal to building a little piece of software to solve some intractable, specific-to-you problem in the space of a couple of afternoons. And in this case, you can do it on hardware in your own home, which your data never leaves, and it consumes less power than a single PC graphics card. I probably could have solved this problem in an afternoon with a few dollars’ worth of tokens, but sending years of financial data to Anthropic or whoever is something I just couldn’t countenance.
This way, I didn’t have to. The Qwen 3.8 27B model is “open-weight” and has been released under an Apache 2.0 license, so there aren’t limits on its use. And it offers different “quantizations” (basically, trading a little precision for a smaller size) that help it span Macs with different amounts of RAM.
I’d say 32GB is the smallest amount of RAM you could use to run the reasonably competent 4K quantization with enough context for very small projects or narrow fixes for existing projects, but 64GB is a better target for a system with a large window for context and enough memory to actually use the Mac as a regular computer at the same time. Using the 4K quantization’s default settings in LM Studio Bionic running on macOS 27.0, my daily-driver M2 Mac Studio managed a token rate of roughly 18 to 20 per second on a sample prompt about spinning up a new website project. A Framework Desktop (which is based on the Strix Point Ryzen AI Max+ platform and is fairly popular among local AI enthusiasts because of its large pool of reasonably fast unified memory) managed roughly the same 18-to-20 tokens-per-second rate on the same prompt, using Ubuntu 26.06 and AMD’s ROCm backend.
This token-per-second rate is by no means terrible, and it means the system can generate text at just about the same rate that I can read it. But for a longer coding project (especially if you leave “reasoning” on, letting the bot spit out a bunch of sequences of words before it settles on a course of action) it means a whole lot of waiting, and that tokens-per-second rate does begin to slow even more as you get deeper into a long context window. The M5 Max Mac Studio had a tokens-per-second rate of around 31 on the same prompt.
The M5 Ultra manages just over 50. (This is, incidentally, why a Mac mini with an M5 Pro is less usable for this kind of work, even with 64GB of RAM—it has half the memory bandwidth of the M5 Max, and memory bandwidth is generally the spec that these workloads are the most sensitive to.) Both of the Mac Studio chips are genuinely usable for this kind of work, as long as your ambitions are modest and you have a mind for methodical troubleshooting and debugging (the agent will virtually always get some thing wrong the first time, and a lot of the time and effort I’ve expended while vibe coding has been in service of knocking features into shape, one fix at a time.) The workloads, they are a-changing The M5 Ultra Mac Studio. Credit: Andrew Cunningham The M5 Ultra Mac Studio. Credit: Andrew Cunningham The Ultra version of the Mac Studio remains a niche product for a small audience: People with thousands of dollars to spend and who want to work with AI models, who are OK with good-but-not-cutting-edge speeds, and who don’t want to pay OpenAI or Anthropic or whoever for all the tokens they might burn with playful experiments or by absorbing and editing a large codebase.
But that audience does exist. Almost against my will, I find myself in it. It’s genuinely fun and freeing to be able to write hobby-project code at usable speeds without my data ever leaving my control.
It’s just too bad I’ve made this discovery as Apple has instituted 25-percent-and-up price increases across the entire Mac Studio line. Considered against the wider Mac and PC market, where prices are awful everywhere you look, the Mac Studio can still be a reasonably good deal. The M5 Max version is a great machine for the Studio’s core audience of photographers, video editors, streamers, and developers, and if you’re primarily using cloud AI models, sticking to 36GB or 48GB of RAM won’t feel like a hardship (you could and possibly should just go with an M5 Pro Mac mini instead, though).
Even entry-level local coding agents don’t require a top-end config; the 64GB, 96GB, and 128GB configurations each cost between $3,500 and $5,500, which is still a lot , but nowhere close to five figures—and still conceivably justifiable if you’re buying a machine you intend to use as your primary workstation for a few years. But the high-end versions of these machines are priced well outside the range of what most people want to spend on a desktop computer. As reviewed, this M5 Ultra Mac Studio costs $12,299.
I have an M2 Max Mac Studio, an M3 MacBook Air, a Windows gaming tower, a PC under my TV, and a MacBook Neo. Granted, I bought these all in the Before Times, when memory prices were sane. But I don’t need to do the math to know that this computer costs more than all of those computers combined.
The good The Mac Studio retains its best selling points: It’s small, fast, quiet, and power efficient M5 generation is a solid upgrade over M4 Max and M3 Ultra Nearly ideal machines for usably fast, power-efficient local ML and AI workloads, including competent coding models The bad Price hikes start at 25 percent generation-over-generation and go up from there M5 Ultra has 2x the computing resources of M5 Max but usually can’t get you 2x the speeds The ugly The prospect of a $20,000 consumer desktop computer Andrew is a Senior Technology Reporter at Ars Technica, with a focus on consumer tech including computer hardware and in-depth reviews of operating systems like Windows and macOS. Andrew lives in Philadelphia and co-hosts a weekly book podcast called Overdue . 6 Comments
Source: Ars Technica
Insider · Berlins Today



