August 28, 2026
Apple caught up to the language. The 512GB Studio is still a waitlist.
Apple now sells Mac Studio on token bills and frontier-class inference. I ordered the 256GB M5 Ultra and retired the laptop. The configuration that holds GLM-5.2 properly is the one they will not take an order for.

Two months ago I opened the M3 Ultra in the other room and loaded GLM-5.2. It sat in unified memory with no meter running. The honest part of that post was the wait: Apple gives you a huge pool of memory and then makes you stand in line for bandwidth.
On Tuesday Apple posted three press releases, no keynote, and called Mac Studio “the ultimate desktop for on-device AI.” The newsroom says “run enormous LLMs entirely on device.” The product page names Ollama and LM Studio and prints “no cloud tokens needed” under a chip.
I ordered one. Then I read the fine print on the configuration I wanted.
The copy caught up
This is not the first time Apple typed LLM. They did it for the M3 Ultra in March 2025, with a 600-billion-parameter claim nobody outside the local-model crowd read.
What changed is the register. Johny Srouji sells inference now. The Mac mini page calls the small box an “always-on agentic device.” The Studio page counts tokens you will not pay for. At WWDC I wrote that Apple’s bet was architectural: let the computer you bought do the thinking first. This week they said it like a workstation vendor, with a price list.
Mini for the agents. Studio for the weights.
The new Mac mini is the first Mac with M6, Apple’s first 2-nanometer chip. It starts at $899 and still tops out at 32GB. The M5 Pro configuration goes to 64GB for $1,699. That is the box for an agent, and I already run one: the crew behind my studio lives on a mini today. A 64GB mini runs a good small model and a lot of tool calls. It does not run GLM-5.2.
For almost everyone else the mini is the answer. Buy the monster when you can name the model that does not fit.
Mac Studio is the other product. M5 Max starts at $2,499 with up to 128GB. M5 Ultra starts at $5,499 with up to 512GB and 1.2TB/s. Most configurations ship 22 September. 256GB you can configure today. 512GB is “coming in late October,” and Apple will not take the order yet.
Nvidia priced this machine
Theo’s video this week walks through why the Studio lands where it does, and the reason is Nvidia’s own price list. An RTX 5090 has fast silicon and 32GB. The RTX Pro 6000 has roughly the same silicon, 96GB, and costs several times more. The DGX Spark gives you 128GB for $4,000, but it is slow LPDDR5 at around 270GB/s on an ARM box nobody uses as a computer. Fast chip, or enough memory, or both at a price that only makes sense on someone else’s budget.
The M5 Ultra breaks that split. 256GB of unified memory at 1.2TB/s is more than four times the Spark’s bandwidth and about a third slower than a 5090, in a box that holds a model the 5090 never will. The Spark category is dead the day this ships.
The other half of Theo’s argument is the model itself. The anonymous “Ox Alpha” that flooded OpenRouter with free tokens turned out to be GLM-5.3 Flash, served entirely on Huawei silicon. The model I am about to run on Apple silicon was served at scale without a single Nvidia chip. Nvidia’s moat is a language, and my stack no longer speaks it.
Bandwidth was the complaint
In June the problem was movement. A big model on an M3 Ultra is smart and slow because the pool is large and the pipe is not. People who have not run this think RAM is the whole story. People who have, wait on prefill.
1.2TB/s is about 50 percent more than the 819GB/s I have now. Neural Accelerators land in every GPU core of an Ultra chip for the first time, and Apple’s July tests put prompt processing in LM Studio at four times M3 Ultra. Apple’s own M5 versus M4 research showed the same shape: time to first token jumps, generation inches. Agentic coding is prefill. Time to first token is the job.
Apple also sells a cluster: four Studios over Thunderbolt 5 and RDMA for 3x the inference. Four machines, three times the speed, so the interconnect is not NVLink. I argued against buying a rack on a fear bet in June. This is a rack with better furniture.

What I ordered
The 256GB M5 Ultra. Desktop only. The MacBook Pro M4 Max retires to the shelf: 128GB was its ceiling, the fans knew when a model was loaded, and I never worked on a train anyway.
256GB is the right size because the model I run changed. GLM-5.3 Flash is 320 billion parameters with 18 billion active. A good 4-bit lands around 200GB. That fits with room for cache, and 18B active means it moves like a much smaller model. It is what the studio crew runs today, through Ollama’s cloud. A 256GB Studio brings the team home.
The correction I owe from June: the GLM-5.2 I called “roughly 250GB” was a 2-bit quant. A proper 4-bit is closer to 380GB. 256GB never held that model well. 512GB is the first Apple configuration sized for it. It is late, it has no price, and MacRumors expects it well above $10,000. I will buy it the day I can name a model that does not fit in 256GB and that I want at home instead of rented. Today Opus and Fable cover that work, and that is not enough to pay the price of a small car for a tier without a price.
My read
The interesting part is not that Apple made a faster Mac. It is that Apple, Huawei and OpenAI’s own silicon all showed up in the same month, and Nvidia’s answer was a Jim Cramer interview.
They already had the letters LLM. They already had 512GB, then took it away. What they added is bandwidth, Neural Accelerators on Ultra, and copy that talks like the people who run the models.
The most useful model is still the one you can actually use. Apple just started selling that sentence. The machine that makes it fully true is the one on backorder.