NakliTechie claims his released MIT LocalMind โ a static, no-install WebGPU page whose engines stream MoE expert weights from disk while generating โ runs Gemma 4 26B-A4B and a 37GB Qwen3.6 MoE larger than a 24GB Mac's RAM entirely in a browser tab at llama.cpp-comparable output, and independent replication and adoption would establish the zero-install browser tab as a practical tier for over-RAM local inference.
state: seedheat: lowuncertainty: mediumconvergesscott: highlocal-inference webgpu moe-streamingNakliTechie
What is this?
The models at the center of this claim are real and current per the supplied snippets: Google's Gemma 4 26B-A4B is a recently released Apache-2.0 mixture-of-experts model (25.2B total / ~3.8B active parameters, 128 experts with top-8 routing) with day-0 llama.cpp, LM Studio, Ollama and vLLM support, and browser-local WebGPU/transformers.js demos of it already circulate (noted by xenovacom, amplified by ClementDelangue); a Qwen3.6 MoE (35B-A3B) also appears in current local-inference and MoE-offloading discussion. The claimed mechanism โ streaming expert weights from disk while generating, so a 37GB MoE can run on a 24GB-RAM Mac โ is coherent with active MoE-offload threads on Hacker News, but none of the supplied snippets mention NakliTechie, LocalMind, or the specific static-browser-page disk-streaming release. The project's distinguishing claims (MIT code, zero-install page, over-RAM operation in a tab, llama.cpp-comparable output) therefore rest on the case's own assertion and are uncorroborated by this search; note also that Gemma 4 26B-A4B at Q4 is ~14.4GB per Google's docs, so the over-RAM claim primarily hinges on the 37GB Qwen3.6 size, which the snippets do not confirm.
Why it matters to Scott
Converges with Scott's hardware-aware-local-inference position โ streaming experts from disk while generating is precisely memory pressure and placement treated as explicit runtime policy โ and with his zero-install static-page, no-backend distribution instinct carried in the CMS-unbundling pages, while breaking the small-model premise his in-browser-student-model concept rests on: if replication holds, the browser tab becomes a practical over-RAM tier that competes with the Ollama/MLX/llama.cpp endpoints he actually operates. This is also the first radar case fusing the two tracked lineages โ expert-streaming (Edge0, Slipstream, SSD-LLaMA) and browser-WebGPU (Mentria, Fleet) โ into one artifact, which is what makes it consequential rather than another streaming variant.
dev:concept.hardware-aware-local-inferencedev:concept.in-browser-student-modeldev:technology.ollamaip:framework.cms-unbundlingradar:concept.browser-inferenceradar:concept.webgpuradar:concept.expert-streamingradar:concept.weight-streamingradar:edge0-streaming-moe-releaseradar:mentria-bonsai27b-webgpu-inferenceradar:fleet-webgpu-kernel-benchmark
queries asked of Scott's wikis
- browser inference WebGPU transformers.js WebLLM local model in-browser
- MoE expert weights streaming offload disk mmap over-RAM inference
- zero-install static page app distribution no backend no server
- local inference Mac Apple Silicon RAM ceiling quantization consumer hardware
- llama.cpp baseline performance comparison local inference engines
- MoE active parameter efficiency sparse activation local model economics
Measured heat
now 0 pts/hpeak 6 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 155h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p53 vs 1247 stories at the 96h mark (now 155h old) โ ahead of ai-vuln-reports-oss-disclosure (1.1x), behind agent-iap-credential-brokering (0.9x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-06T08:05:33Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wyvo42 -> echo.x.35f11508be by Chirag (Chirag Patnaik, @chirag on X; same person as Reddit u/naklitechie and GitHub NakliTechie)
2026-10-06T07:34:57Z
grounded: converges/high โ Converges with Scott's hardware-aware-local-inference position โ streaming experts from disk while generating is precisely memory pressure and placement treated
2026-10-06T07:26:13Z
case created โ A maintainer-released, inspectable artifact (live demo plus MIT code) adding disk-streamed over-RAM MoE to the browser-inference tier โ distinct from the native-streaming (Edge0) and small-model-WebGPU (Mentria, Fleet) cases already open.
Decision trace
- 10-11 12:07review_screenjev screen: no material development (noul=0.19)
- 10-06 19:20sensor_dirtycomment_update
- 10-06 19:05promote_anchororigin walk conf 0.85
- 10-06 18:34groundConverges with Scott's hardware-aware-local-inference position โ streaming experts from disk while generating is precisely memory pressure and placement treated as explicit runtime policy โ and w
- 10-06 18:26createA maintainer-released, inspectable artifact (live demo plus MIT code) adding disk-streamed over-RAM MoE to the browser-inference tier โ distinct from the native-streaming (Edge0) and small-model-WebGP