2026-10-11 18:04 UTC

Independent testing will determine whether ExpertCache can run the full 63GB GPT-OSS 120B model on a 16GB M1 Pro at practically useful speed and output quality through expert caching.

state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference open-models expert-caching apple-siliconAMOS Labs

What is this?

ExpertCache is presented by AMOS Labs as a system for running the full 63GB GPT-OSS 120B model on a 16GB M1 Pro by caching model experts rather than keeping the entire model resident in memory. The cited primary artifact is an ExpertCache commit documenting a claimed physical 16GB M1 Pro result, but the supplied search snippets provide no independent benchmark of its speed, output quality, or full-model fidelity. The surrounding evidence instead shows GPT-OSS 120B conventionally associated with substantially larger memory configurations, so the practical value of the claimed result remains unverified.

Why it matters to Scott

The radar already tracks essentially the same unresolved claim in `radar:hotpin-lossless-moe-streaming` and `radar:slipstream-ssd-moe-streaming`: expert paging or streaming may enable oversized MoE models on memory-constrained consumer hardware, but practical speed and fidelity require independent validation. This directly touches Scott’s hardware-aware local-inference and capability-audit work, yet ExpertCache currently adds only another unverified implementation rather than evidence that would change what he builds or argues.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditradar:hotpin-lossless-moe-streamingradar:slipstream-ssd-moe-streamingradar:concept.expert-streamingradar:concept.local-inference
queries asked of Scott's wikis
  • expert caching and sparse MoE inference
  • local inference under severe memory constraints
  • Apple Silicon out-of-core model serving
  • full-model fidelity versus expert offloading
  • local open-model hardware economics
  • independent benchmarking of inference optimizations

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: ExpertCache – Run the Full GPT-OSS 120B (63GB) on a 16GB M1 Prorick_barkley10
🟧 echo.github ⭐The earliest public primary artifact I found for the claimed 16GB result is the ExpertCache commit documenting β€œPhysical 16 GB M1 Pro resultRick Barkleyβ€”β€”

Interpretation history

Decision trace