2026-10-11 16:38 UTC

Mentria.ai’s creator claims its WebGPU engine runs Prism ML’s one-bit Bonsai-27B at 25–30 tokens per second on a 6GB RTX 3060 Laptop GPU entirely in Chrome, potentially enabling responsive 27B inference without installation or hosted processing.

state: watchingheat: highuncertainty: highconvergesscott: mediumlocal-inference browser-inference webgpu quantizationmentria-aiMentria.aiPrism ML
Surfaced 2026-09-19T09:23:41Z — 1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) — Bonsai 2’s top-decile spread across Reddit and Hacker News warrants attention to the surrounding low-memory inference episode, despite adding no new implementation result. This raises heat, not confidence: successor publicity still does not validate Mentria’s original throughput claim or establish useful coding reliability.

What is this?

Mentria.ai’s creator claims a community WebGPU engine runs PrismML’s one-bit Bonsai-27B entirely in Chrome at 25–30 tokens per second on a 6GB RTX 3060 Laptop GPU. PrismML’s announcement identifies Bonsai 27B as based on Qwen3.6-27B, and its Hugging Face page reports a 3.9GB footprint for the binary model. The supplied snippets support browser-local execution in general, but do not independently establish Mentria.ai’s specific hardware-throughput claim; the model-card benchmarks use native Metal/CUDA or MLX runtimes, not WebGPU. No installation does not mean no download: the application and model assets must first be loaded onto the device.

Why it matters to Scott

Mentria.ai’s claimed low-memory browser runtime converges with Scott’s Hardware-aware local inference approach and offers a concrete alternative to test alongside his gamepc/Ollama serving setup, rather than merely illustrating cheaper cognition. The specific development is not present in the supplied radar hits, but its value remains conditional: the creator’s throughput claim lacks independent WebGPU verification, and the supplied evidence does not establish useful task quality or end-to-end responsiveness.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.browser-inferenceradar:concept.webgpuradar:concept.quantizationradar:browser-llm-fit-hardware-selectionradar:hashagent-browser-local-agents
queries asked of Scott's wikis
  • browser-local inference WebGPU zero-install AI products
  • low-bit quantization capability retention agent reliability
  • local inference economics consumer GPU memory budgets
  • private on-device RAG knowledge systems
  • inference benchmarks decode throughput prefill context limits

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 770h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-09 13:49⭐ origin directly observed1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)
mentria-ai on r/LocalLLaMA
—
09-17 21:05first on r/LocalLLaMA · published · +199.3hTernary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.
xenovatech
—
09-17 21:13first on hacker news · published · +199.4hBonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
JonSchneider
—
09-09 13:49amplified on r/LocalLLaMAreddit.post.1wbm50k
mentria-ai
peak 73 · 28 comments · 3% of case engagement
09-17 21:05amplified on r/LocalLLaMA 👑reddit.post.1wj6c4l
xenovatech
peak 1677 · 344 comments · 57% of case engagement
09-17 21:13amplified on hacker newshn.story.49746618
JonSchneider
peak 589 · 200 comments · 40% of case engagement
09-18 14:08amplified on r/LocalLLaMAreddit.post.1wjr8oo
BeyondRedline
peak 3 · 3 comments · 0% of case engagement
09-09 14:20our radar first saw it · +0.5hdiscovery anchor: reddit.post.1wbm50k—
09-19 09:23reached heat=high · +235.6h · via ledger——
pace: p96 vs 519 stories at the 720h mark (now 770h old) — ahead of openai-chatgpt-weekly-prompt-caps (1.1x), behind k2-horizon-open-models (0.9x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)
LocalLLaMA
mentria-ai7328
🟠 redditTernary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.
LocalLLaMA
xenovatech1677344
🟧 hnBonsai 2 27B: Near-Lossless Compression in a 9x Smaller FootprintJonSchneider589200
🟠 redditPrismML Bonsai 2 released
LocalLLaMA
BeyondRedline03

Interpretation history

Decision trace