2026-10-11 16:38 UTC

Independent benchmarks will determine whether Apple’s M5 Ultra Mac Studio, with up to 512GB unified memory and roughly 1.2TB/s memory bandwidth, provides materially better capacity and economics for local LLM inference.

state: corroboratedheat: lowuncertainty: lowconvergesscott: mediumapple-silicon local-inference unified-memory inference-economicsApple
Surfaced 2026-09-21T15:42:50Z β€” Apple announced the M6 and M5 Ultra as new chips offering a substantial increase in performance and AI compute. β€” A hands-on launch-window review adds the first consequential comparative local-inference results, shifting the case from speculative specifications toward a testable hardware-buying proposition. Its figures appear to reuse the earlier oMLX dataset, however, so broad Reddit/HN circulation raises immediate attention without yet supplying independent corroboration of economics.

What is this?

Apple's 2026 Mac Studio refresh introduces the M5 Ultra, a quad-die chip offering up to 512GB of unified memory at a claimed 1.2TB/s bandwidth and aimed squarely at local AI inference; configurations run from $5,499 (96GB) to a maxed $18,299, with base models shipping September 22 and the 512GB tier still promised for late October. The supplied web coverage is mostly spec- and estimate-level β€” ModelFit figures that explicitly disclaim themselves as non-measured, Geekbench runs, one YouTube Llama test, and a '90% faster' headline the snippet never substantiates β€” so independent measurement evidence remains thin in the snippets, with price details also varying across sources. The case's own accumulated record goes further: hands-on artifacts (a 112M-token owner run, a GLM-5.3-Flash agentic report, a matched same-harness 262K-context 256GB shootout, config-tuning PSAs) have settled the verdict into a corroborated tradeoff β€” best-in-class capacity-per-dollar and single-user fit, but slower per-stream decode at double the power and poor multi-user concurrency versus CUDA-class GPU rigs β€” with 256GB hardened as the community sweet spot and the 512GB tier reframed as likely compute-bottlenecked.

Why it matters to Scott

Independent testing has settled exactly the question Scott's canon frames β€” Apple Silicon vs CUDA rig β€” landing on the verdict his own build already implied: M5 Ultra buys capacity-per-dollar and single-user fit (256GB sweet spot, 512GB likely compute-bottlenecked), not per-stream speed, power efficiency, or multi-user concurrency, so gamepc/CUDA stays his substrate and this is a dated-receipts convergence with usable-mass-over-unusable-power rather than a reason to buy. The tuning/matched-optimization stream (prefill 8192, MTP/dflash, oMLX 4.7bpw quants) extends his hardware-aware-local-inference and MLX practice, with the October 512GB delivery and a matched GPU-rig decode comparison the remaining re-wake catalysts.
dev:project.gamepcdev:concept.hardware-aware-local-inferencedev:technology.mlxdev:technology.cudaip:concept.usable-mass-over-unusable-powerradar:aa-agentperf-local-benchmarkradar:adaptive-kv-cache-streamingradar:2027-memory-capacity-selloutradar:ante-offline-coding-agent
queries asked of Scott's wikis
  • unified memory vs multi-GPU sharding tradeoff for large local models
  • capacity-per-dollar versus throughput in local inference hardware choices
  • MLX vs GGUF quantization and prefill/decode config-tuning gains
  • local coding-agent serving concurrency tokens-per-second requirements
  • local vs cloud API inference cost-per-token economics thresholds
  • Apple Silicon viability as AI workstation substrate versus CUDA rig

Measured heat

now 0 pts/hpeak 56 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1131h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

08-25 13:26 (minted)⭐ origin echo-reconstructedApple announced the M6 and M5 Ultra as new chips offering a substantial increase in performance and AI compute.
Apple on blog (echo) Β· attributed from reddit.post.1vxzgyt, hn.story.49433292, reddit.post.1vxzg6v, hn.story.49433316 Β· published time unknown
β€”
08-25 13:01first on hacker news Β· published Β· lag ?Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute
interpol_p
β€”
08-25 13:11first on r/LocalLLaMA Β· published Β· lag ?Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
themixtergames
β€”
08-25 13:01amplified on hacker news πŸ‘‘hn.story.49433292
interpol_p
peak 1312 Β· 1295 comments Β· 23% of case engagement
08-25 13:03amplified on hacker newshn.story.49433316
interpol_p
peak 823 Β· 554 comments Β· 12% of case engagement
08-25 13:11amplified on r/LocalLLaMAreddit.post.1vxzg6v
themixtergames
peak 1661 Β· 778 comments Β· 12% of case engagement
08-25 13:12amplified on r/LocalLLaMAreddit.post.1vxzgyt
Last-Owl-8342
peak 882 Β· 207 comments Β· 5% of case engagement
08-25 23:16amplified on r/LocalLLaMAreddit.post.1vyfved
Mxmtm
peak 76 Β· 105 comments Β· 1% of case engagement
08-26 03:21amplified on r/LocalLLaMAreddit.post.1vylfc0
dreamingwell
peak 0 Β· 15 comments Β· 0% of case engagement
28 more amplifiers in ainews.case_chain
08-25 13:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1vxzgytβ€”
09-21 15:42reached heat=high Β· lag ? Β· via ledgerβ€”β€”

Evidence (35) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditApple releases M5 ultra at 1.2TB/s bandwith
LocalLLaMA
Last-Owl-8342882207
🟧 hnApple introduces M6 and M5 Ultra for a big leap in performance and AI computeinterpol_p13121295
🟠 redditApple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
LocalLLaMA
themixtergames1661774
🟧 hnApple Introduces New Mac Studio with M5 Max and M5 Ultrainterpol_p823554
🟧 echo.blog ⭐Apple announced the M6 and M5 Ultra as new chips offering a substantial increase in performance and AI compute.Appleβ€”β€”
🟠 redditM5 Ultra 96GB vs M5 Max 128GB β€” is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow?
LocalLLaMA
Mxmtm76105
🟠 redditI built a tracker for which open-weight models actually run on Apple Silicon β€” most of the good ones don't (yet)
LocalLLaMA
dreamingwell015
🟠 redditM7 Ultra might be GLM 5.3-flash monster (native FP8, Apple Silicon M6)
LocalLLaMA
Brilliant-Hall1387020
🟠 reddit5090 now officially cost 5090
LocalLLaMA
Sadge4041642399
🟠 redditMac Studio m5 ultra - 2x96gb or 1x256gb?
LocalLLaMA
anonmt57435
🟠 redditExo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering
LocalLLaMA
anonmt57169138
🟠 redditModels to download for M5 Ultra 512GB
LocalLLaMA
Ok_Warning21462056
🟧 hnApple Caught Off Guard by AI Demand for Mac Mini and Mac Studiothm501564
🟧 hnApple Is Suddenly an AI Infra Stock as OpenAI Buys 10k+ Macsprabal973924
🟠 redditThe DGX Spark joins the 5090 in its price increase.
LocalLLaMA
Sadge4049697
🟠 redditPlanning to serve multiple user with mac studio
LocalLLaMA
Interesting-Print366028
🟠 redditMac Heads: Is there any point to MLX in September 2026?
LocalLLaMA
MrPecunius1932
🟧 hnGetting 50 GB/S Back Out of the Neural Engineeiln11
🟧 hnGetting 50 GB/S Back from the Apple Neural Engineeiln21736
🟠 redditMac Studio M5 Ultra
LocalLLaMA
Blues5201432
🟠 redditShould I sell my RTX 5090 for a Mac Studio M5 Ultra 96GB?
LocalLLaMA
unchikuso242288
🟠 redditFirst M5 Ultra benchmarks
LocalLLaMA
Ashefromapex122149
🟠 redditM5 Ultra and M6 Chip Benchmark Results Reveal Graphics Performance
LocalLLaMA
DustNearby2848269153
🟠 redditM5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories
LocalLLaMA
themixtergames351173
🟧 hnM5 Ultra Mac Studio Review: The Dream Mac for Local AI Agentspiotrgrabowski269253
🟠 redditM5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents
LocalLLaMA
pscoutou1219
🟠 redditM5 ultra AI test results
LocalLLaMA
DigitalguyCH4771
🟠 redditM5U base 96GB inference numbers for Q3.8FN after 112M tokens
LocalLLaMA
Every-Fortune-31511725
🟠 redditFolks, have you purchased the Mac M5 Ultra with 256GB yet? We need serious benchmarks, because we only get YouTube clowns influencers results
LocalLLaMA
More-Curious816290221
🟠 redditM5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds
LocalLLaMA
dreamingwell194111
🟠 redditPSA for M5Ultra owners running LLMs: set your prefill step to 8192
LocalLLaMA
bakawolf1232117
🟠 redditMinisforum MS-S1 MAX-P495 @ €7.799,00
LocalLLaMA
streppelchen5859
🟠 redditM5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1
LocalLLaMA
nonlinearsystems923
🟠 redditM5 Ultra vs 2 DGX Sparks
LocalLLaMA
kaggie021
🟠 redditSwift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra
LocalLLaMA
DankpawsDev127

Interpretation history

Decision trace