2026-10-11 17:09 UTC

expert-caching

band: warmmomentum: stable score: 0.546
temperature history

Episodes (2)

Independent testing will determine whether ExpertCache can run the full 63GB GPT-OSS 120B model on a 16GB M1 Pro at practically useful speed and output quality through expert caching.
expiredknownscott: low
Fork author neuralll claims his released llama.cpp fork's VRAM-filling hot-expert cache (built on csantiago78's PR #27861) roughly doubles decode throughput for GLM-5.3-Flash and MiMo MoE models far larger than total VRAM on two RTX 3090s with unchanged perplexity, and projects further gains per added GPU โ€” if independent multi-GPU users reproduce it, hot-expert caching becomes a practical standard path for memory-constrained local MoE inference.
acceleratingconvergesscott: high