2026-10-11 16:37 UTC

NakliTechie claims his released MIT LocalMind โ€” a static, no-install WebGPU page whose engines stream MoE expert weights from disk while generating โ€” runs Gemma 4 26B-A4B and a 37GB Qwen3.6 MoE larger than a 24GB Mac's RAM entirely in a browser tab at llama.cpp-comparable output, and independent replication and adoption would establish the zero-install browser tab as a practical tier for over-RAM local inference.

state: seedheat: lowuncertainty: mediumconvergesscott: highlocal-inference webgpu moe-streamingNakliTechie

What is this?

The models at the center of this claim are real and current per the supplied snippets: Google's Gemma 4 26B-A4B is a recently released Apache-2.0 mixture-of-experts model (25.2B total / ~3.8B active parameters, 128 experts with top-8 routing) with day-0 llama.cpp, LM Studio, Ollama and vLLM support, and browser-local WebGPU/transformers.js demos of it already circulate (noted by xenovacom, amplified by ClementDelangue); a Qwen3.6 MoE (35B-A3B) also appears in current local-inference and MoE-offloading discussion. The claimed mechanism โ€” streaming expert weights from disk while generating, so a 37GB MoE can run on a 24GB-RAM Mac โ€” is coherent with active MoE-offload threads on Hacker News, but none of the supplied snippets mention NakliTechie, LocalMind, or the specific static-browser-page disk-streaming release. The project's distinguishing claims (MIT code, zero-install page, over-RAM operation in a tab, llama.cpp-comparable output) therefore rest on the case's own assertion and are uncorroborated by this search; note also that Gemma 4 26B-A4B at Q4 is ~14.4GB per Google's docs, so the over-RAM claim primarily hinges on the 37GB Qwen3.6 size, which the snippets do not confirm.

Why it matters to Scott

Converges with Scott's hardware-aware-local-inference position โ€” streaming experts from disk while generating is precisely memory pressure and placement treated as explicit runtime policy โ€” and with his zero-install static-page, no-backend distribution instinct carried in the CMS-unbundling pages, while breaking the small-model premise his in-browser-student-model concept rests on: if replication holds, the browser tab becomes a practical over-RAM tier that competes with the Ollama/MLX/llama.cpp endpoints he actually operates. This is also the first radar case fusing the two tracked lineages โ€” expert-streaming (Edge0, Slipstream, SSD-LLaMA) and browser-WebGPU (Mentria, Fleet) โ€” into one artifact, which is what makes it consequential rather than another streaming variant.
dev:concept.hardware-aware-local-inferencedev:concept.in-browser-student-modeldev:technology.ollamaip:framework.cms-unbundlingradar:concept.browser-inferenceradar:concept.webgpuradar:concept.expert-streamingradar:concept.weight-streamingradar:edge0-streaming-moe-releaseradar:mentria-bonsai27b-webgpu-inferenceradar:fleet-webgpu-kernel-benchmark
queries asked of Scott's wikis
  • browser inference WebGPU transformers.js WebLLM local model in-browser
  • MoE expert weights streaming offload disk mmap over-RAM inference
  • zero-install static page app distribution no backend no server
  • local inference Mac Apple Silicon RAM ceiling quantization consumer hardware
  • llama.cpp baseline performance comparison local inference engines
  • MoE active parameter efficiency sparse activation local model economics

Measured heat

now 0 pts/hpeak 6 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 155h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-05 04:41โญ origin echo-reconstructedOrigin tweet thread root (1/): "Run a model larger than your GPU memory in a browser tab. diskformer.js 0.2 keeps the weights on disk and gi
Chirag (Chirag Patnaik, @chirag on X; same person as Reddit u/naklitechie and GitHub NakliTechie) on x (echo) ยท attributed from reddit.post.1wyvo42
โ€”
10-06 06:41first on r/LocalLLaMA ยท published ยท +26.0hGemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac โ€” experts streamed from disk, output matches llama.cpp
naklitechie
โ€”
10-06 06:41amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyvo42
naklitechie
peak 7 ยท 12 comments ยท 101% of case engagement
10-06 07:20our radar first saw it ยท +26.6hdiscovery anchor: reddit.post.1wyvo42โ€”
pace: p53 vs 1247 stories at the 96h mark (now 155h old) โ€” ahead of ai-vuln-reports-oss-disclosure (1.1x), behind agent-iap-credential-brokering (0.9x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditGemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac โ€” experts streamed from disk, output matches llama.cpp
LocalLLaMA
naklitechie712
๐ŸŸง echo.x โญOrigin tweet thread root (1/): "Run a model larger than your GPU memory in a browser tab. diskformer.js 0.2 keeps the weights on disk and giChirag (Chirag Patnaik, @chirag on X; same person as Reddit u/naklitechie and GitHub NakliTechie)โ€”โ€”

Interpretation history

Decision trace