2026-10-11 18:04 UTC

Independent benchmarks will determine whether FreeToken can run 290B-plus sparse-MoE models on gaming PCs with practically useful correctness, throughput, and memory efficiency.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference sparse-moe inference-economicsFlashML

What is this?

FreeToken is a research system for serving sparse mixture-of-experts models larger than available GPU memory on consumer hardware, with claims spanning roughly 35B models on laptops to 284B models on gaming desktops. The reported evaluation compares it with llama.cpp, Ollama, KTransformers, and MoE-Infinity across six GPUs and four agentic workloads, claiming more stable decode throughput and lower tail time-to-first-token through bandwidth-adaptive execution and pipelined prefill overlap. The case attributes the work to FlashML, but the supplied snippets do not independently establish the authorship, and they provide no clearly independent FreeToken benchmark confirming correctness or the headline performance claims.

Why it matters to Scott

The validation-first position is already explicit in Scott’s Capability Audit and Evaluation-Driven Development pages, while the claimed consumer-hardware deployment directly bears on his gamepc model-serving substrate and hardware-aware local-inference work. If independently verified, FreeToken could materially expand what he can run locally; for now it remains another unconfirmed implementation in a territory the radar already tracks through Slipstream, AirLLM, HotPin, and DeepSeek expert-streaming cases.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.moe-inferenceradar:concept.expert-streamingradar:slipstream-ssd-moe-streamingradar:airllm-low-vram-model-streamingradar:hotpin-lossless-moe-streamingradar:deepseek-v4-flash-expert-streaming
queries asked of Scott's wikis
  • local inference economics on consumer GPUs
  • sparse MoE offloading and memory bandwidth
  • gaming-PC inference for coding agents
  • benchmarking practical correctness of local models
  • local model sovereignty versus hosted inference
  • agent workload latency and time-to-first-token

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (9) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnRun 290B+ frontier MoE models locally on your gaming PCshenli3514403
🟧 echo.paper ⭐The original primary artifact is the authors’ arXiv paper, submitted 17 Aug 2026. Its abstract says FreeToken supports models “from a 35B moShuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, and Ion Stoica——
🟠 redditFreeToken: Frontier models at interactive speed using official checkpoints without extreme quantizations.
LocalLLaMA
SysPsych010
🟠 redditFreetokens project is impressive
LocalLLaMA
ViRROOO5889
🟠 redditAnyone trying out Freetoken and stuck on weight conversion?
LocalLLaMA
T_rex270028
🟠 reddit[2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
LocalLLaMA
SteppenAxolotl5614
🟠 redditIs this a fluke?
LocalLLaMA
Gianniarrenzetti124
🟧 hnFreeToken to Speed Up Moe on low VRAMgitowiec10
🟧 hnShow HN: PulsarForge – run a 744B MoE model on 32GB RAM with zero GPU (pure C)siris947610

Interpretation history

Decision trace