2026-10-11 18:02 UTC

Independent testing will determine whether llama.cpp’s merged Bonsai and ternary-model support enables correct, performant local inference for 1-bit and 1.58-bit Bonsai checkpoints across common backends.

state: expiredheat: lowuncertainty: highknownscott: mediumopen-models local-inference llama-cppPrism AIllama.cpp

What is this?

PrismML’s Bonsai family uses extremely low-bit weights—1-bit and ternary/approximately 1.58-bit—to reduce model size for local inference, with PrismML reporting competitive benchmark retention. Community snippets show promising performance, including roughly 106 tokens/second in one configuration, but also weak results on some reasoning tests. Backend compatibility remains unclear and format-specific: one llama.cpp discussion says a Bonsai Q2_0 checkpoint works only with PrismML’s fork while a g64 variant is needed elsewhere, and the supplied snippets do not firmly establish that complete support has merged into stock llama.cpp or works correctly across common backends.

Why it matters to Scott

The radar already tracks the underlying Bonsai extreme-quantization claim in `radar:bonsai-extreme-quantization`; this case extends that open story into stock llama.cpp compatibility and backend validation rather than introducing a new position. It matters to Scott because verified support could make sub-2-bit checkpoints deployable on his hardware-aware local inference substrate, while format-specific failures would reinforce the need for repeatable capability audits before adoption.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditip:concept.evaluation-driven-developmentdev:project.gamepcradar:bonsai-extreme-quantizationradar:concept.llama-cppradar:concept.quantization
queries asked of Scott's wikis
  • sub-2-bit local inference economics
  • llama.cpp backend compatibility and quantization
  • independent evaluation of low-bit model quality
  • local model format fragmentation
  • on-device inference performance tradeoffs
  • grammar-constrained decoding for small local models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking
LocalLLaMA
pmttyji6313
🟠 redditBeating the vendor's official runtime on free ARM cores: a from-scratch engine for ternary 8B models (decode +14%, prefill +55%). Live demo included.
LocalLLaMA
Annual_Manner_590143

Interpretation history

Decision trace