Independent testing will determine whether ExpertCache can run the full 63GB GPT-OSS 120B model on a 16GB M1 Pro at practically useful speed and output quality through expert caching.
state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference open-models expert-caching apple-siliconAMOS Labs
What is this?
ExpertCache is presented by AMOS Labs as a system for running the full 63GB GPT-OSS 120B model on a 16GB M1 Pro by caching model experts rather than keeping the entire model resident in memory. The cited primary artifact is an ExpertCache commit documenting a claimed physical 16GB M1 Pro result, but the supplied search snippets provide no independent benchmark of its speed, output quality, or full-model fidelity. The surrounding evidence instead shows GPT-OSS 120B conventionally associated with substantially larger memory configurations, so the practical value of the claimed result remains unverified.
Why it matters to Scott
The radar already tracks essentially the same unresolved claim in `radar:hotpin-lossless-moe-streaming` and `radar:slipstream-ssd-moe-streaming`: expert paging or streaming may enable oversized MoE models on memory-constrained consumer hardware, but practical speed and fidelity require independent validation. This directly touches Scottβs hardware-aware local-inference and capability-audit work, yet ExpertCache currently adds only another unverified implementation rather than evidence that would change what he builds or argues.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditradar:hotpin-lossless-moe-streamingradar:slipstream-ssd-moe-streamingradar:concept.expert-streamingradar:concept.local-inference
queries asked of Scott's wikis
- expert caching and sparse MoE inference
- local inference under severe memory constraints
- Apple Silicon out-of-core model serving
- full-model fidelity versus expert offloading
- local open-model hardware economics
- independent benchmarking of inference optimizations
Measured heat
no measured readings yet β the hourly heat pass fills this in
How the heat travelled
no chain yet β the hourly chain pass fills this in
Evidence (2) β β canonical anchor
Interpretation history
2026-08-08T23:34:18Z
No independent benchmark, implementation report, or practical speed-and-quality evidence has appeared; the original claim remains unvalidated and this low-attention episode has faded without changing the broader expert-streaming thesis.
2026-08-08T23:31:46Z
grounded: known/low β The radar already tracks essentially the same unresolved claim in `radar:hotpin-lossless-moe-streaming` and `radar:slipstream-ssd-moe-streaming`: expert paging
2026-08-08T23:29:23Z
origin walked (codex/luna, conf 0.96): anchor hn.story.49226502 -> echo.github.46dd132d56 by Rick Barkley
2026-08-08T23:27:22Z
case created β The first-party implementation makes an unusually concrete and reproducible constrained-memory inference claim for a large open model.
Decision trace
- 08-09 09:34expireNo independent benchmark, implementation report, or practical speed-and-quality evidence has appeared; the original claim remains unvalidated and this low-attention episode has faded without changing
- 08-09 09:34alert_silentThe only delta is a routine re-evaluation with unchanged evidence and engagement, so there is nothing consequential to surface before a future independent test appears.
- 08-09 09:34alert_routeThe only delta is a routine re-evaluation with unchanged evidence and engagement, so there is nothing consequential to surface before a future independent test appears.
- 08-09 09:32alert_silentA primary project artifact now reports that the complete 63.4GB checkpoint executed on a physical 16GB M1 Pro, but it provides no independent benchmark or practical speed and output-quality evidence.
- 08-09 09:32surface_candidateA primary project artifact now reports that the complete 63.4GB checkpoint executed on a physical 16GB M1 Pro, but it provides no independent benchmark or practical speed and output-quality evidence.
- 08-09 09:32alert_routeA primary project artifact now reports that the complete 63.4GB checkpoint executed on a physical 16GB M1 Pro, but it provides no independent benchmark or practical speed and output-quality evidence.
- 08-09 09:31groundThe radar already tracks essentially the same unresolved claim in `radar:hotpin-lossless-moe-streaming` and `radar:slipstream-ssd-moe-streaming`: expert paging or streaming may enable oversized MoE mo
- 08-09 09:29promote_anchororigin walk conf 0.96
- 08-09 09:27createThe first-party implementation makes an unusually concrete and reproducible constrained-memory inference claim for a large open model.