2026-10-11 17:11 UTC

Independent benchmarks will determine whether Slipstream's SSD expert streaming enables 35B–480B MoE coding models to run usefully on memory-constrained Macs without prohibitive latency, swapping, or SSD wear.

state: expiredheat: lowuncertainty: mediumnovelscott: lowmoe-inference weight-streaming local-inferenceSlipstream

What is this?

Slipstream appears to be a self-contained implementation for running very large mixture-of-experts coding models—claimed at 35B–480B parameters—on a 36 GB MacBook by streaming experts from SSD rather than keeping all weights in memory. The evidence titles say its repository predates the associated Reddit post and includes the implementation, app, benchmarks, and failed experiments. However, the supplied web results are unrelated and provide no independent confirmation of performance, latency, swapping behavior, SSD wear, or even the repository’s details, so the core claims remain unverified here.

Why it matters to Scott

No intersection found: there are no Scott wiki hits connecting this claim to his positions or projects, and no radar hits showing that Slipstream or this development is already tracked.
queries asked of Scott's wikis
  • SSD weight streaming for local MoE inference
  • local inference under unified-memory constraints
  • MoE expert paging and inference latency
  • storage wear tradeoffs in model offloading
  • independent benchmarking of local AI runtimes
  • large coding models on Apple Silicon

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI run 35B–480B coding models on my 36 GB MacBook by streaming MoE experts from SSD — self-contained app, and I publish the benchmarks that *failed* too
LocalLLaMA
Illustrious-Cup-5895221
🟧 echo.github ⭐The repository predates the Reddit post and contains the primary implementation, app, benchmarks, and failed experiments. Its README says SlSchero D. (GitHub: Schero94)——
🟧 hnShow HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Macgitpusher42896330
🟧 hnShow HN: A new engine to run Kimi K3 on a laptopmarcobambini73
🟠 redditSsd stream models on strix halo
LocalLLaMA
lawanda12307
🟧 hnShow HN: RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktopcoderredlab70
🟠 redditTurbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon
LocalLLaMA
minefew9922
🟠 redditI ported TurboFieldfare to Qwen 3.6 35B and it runs in 1.4 GB of RAM
LocalLLaMA
Blahblahblakha8825

Interpretation history

Decision trace