2026-10-11 17:10 UTC

PrismML's Bonsai binary and ternary models will retain usable quality and fine-tunability on consumer Apple hardware at roughly 1.1–1.7 bits per weight.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumextreme-quantization ternary-weights local-inference local-finetuningPrismML

What is this?

PrismML has released Bonsai, a family of ultra-low-bit language models intended for local deployment, including 1-bit and ternary variants; its demo supports Mac/Metal and other backends, while the company says the 3.9GB 1-bit 27B model can run on an iPhone. PrismML claims its 1.58-bit ternary models preserve comparatively strong benchmark accuracy at roughly one-ninth the memory footprint of 16-bit models, with a 27B version described as about 1.7 bits per weight. The supplied snippets establish local inference and company-reported quality claims, but they do not independently establish usable fine-tuning on Apple hardware or the full claimed 1.1–1.7-bit range.

Why it matters to Scott

Bonsai potentially extends Scott’s hardware-aware local-inference work and “usable mass” thesis by making much larger models deployable within consumer-device memory constraints. It could change local runtime and model-routing choices if validated, but the supplied evidence does not independently establish usable quality or Metal fine-tuning at the claimed bit rates.
dev:concept.hardware-aware-local-inferenceip:concept.usable-mass-over-unusable-power
queries asked of Scott's wikis
  • sub-2-bit quantization quality tradeoffs
  • ternary weights versus post-training quantization
  • Apple Silicon local inference economics
  • Metal-based local model fine-tuning
  • on-device models and model sovereignty
  • memory bandwidth limits for local LLMs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (15) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
LocalLLaMA
ElmBark781118
🟠 redditI tried fine-tuning a ternary model, Bonsai 8b, on metal
LocalLLaMA
professormunchies154
🟠 redditDon't want to be that guy, but... Bonsai (1-bit and Ternary) vs ThinkingCap (@2.5-bit) - Pareto of generated tokens vs accuracy, model size vs accuracy.
LocalLLaMA
JLeonsarmiento152
🟧 hn1-Bit LLM in the Browsersimonebrunozzi11947
🟠 redditI ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
LocalLLaMA
Creative-Regular679926675
🟠 redditHas anyone tried running PrismML Bonsai 27B yet?
LocalLLaMA
tony10000010
🟠 reddityou can now fine tune Prism-ML's ternary Bonsai models
LocalLLaMA
terminoid_304
🟠 redditBuilt a from-scratch BitNet inference engine in pure C — 1.8× faster than bitnet.cpp on Xeon (36 tok/s), zero dependencies [BitNet & Bonsai CPU testers wanted]
LocalLLaMA
shifu_legend3831
🟠 redditTernary Bonsai 27B?
LocalLLaMA
leo-k7v117
🟧 hnShow HN: Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rulespcdeni164
🟧 hnShow HN: Avoiding the Memory Wall by computing LLM inference directly inside RAMpcdeni10
🟠 redditI hope ternary will eventually work but ... sigh
LocalLLaMA
quadra-lab2039
🟠 redditUsing the Bonsai 27b 1b quant locally - regularly.
LocalLLaMA
fuckAIbruhIhateCorps4829
🟠 redditGot a 27B model running locally on a Jetson Orin NX 16GB (1-bit). still kind of amazed it works
LocalLLaMA
Clean-Mention65431731
🟠 redditTried PrismML’s Bonsai 27B (ternary) on an RX 9070 XT — impressions on a real AMD setup
LocalLLaMA
blakok14623

Interpretation history

Decision trace