PrismML’s Bonsai family uses extremely low-bit weights—1-bit and ternary/approximately 1.58-bit—to reduce model size for local inference, with PrismML reporting competitive benchmark retention. Community snippets show promising performance, including roughly 106 tokens/second in one configuration, but also weak results on some reasoning tests. Backend compatibility remains unclear and format-specific: one llama.cpp discussion says a Bonsai Q2_0 checkpoint works only with PrismML’s fork while a g64 variant is needed elsewhere, and the supplied snippets do not firmly establish that complete support has merged into stock llama.cpp or works correctly across common backends.
The radar already tracks the underlying Bonsai extreme-quantization claim in `radar:bonsai-extreme-quantization`; this case extends that open story into stock llama.cpp compatibility and backend validation rather than introducing a new position. It matters to Scott because verified support could make sub-2-bit checkpoints deployable on his hardware-aware local inference substrate, while format-specific failures would reinforce the need for repeatable capability audits before adoption.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditip:concept.evaluation-driven-developmentdev:project.gamepcradar:bonsai-extreme-quantizationradar:concept.llama-cppradar:concept.quantization
queries asked of Scott's wikis
- sub-2-bit local inference economics
- llama.cpp backend compatibility and quantization
- independent evaluation of low-bit model quality
- local model format fragmentation
- on-device inference performance tradeoffs
- grammar-constrained decoding for small local models
2026-08-24T06:22:26Z
The validation window has faded without a reproducible stock llama.cpp run, backend matrix, or controlled benchmark. Preserve the merge as historical implementation context, but reopen only if concrete independent testing appears.
2026-08-22T05:30:51Z
The refreshed comment adds methodological scrutiny rather than validation: key benchmark controls such as thread count, batch size, and prompt length remain unspecified. It reinforces that neither the ARM performance claim nor stock llama.cpp support is independently reproducible yet.
2026-08-21T21:29:47Z
The ARM benchmark shows that Bonsai ternary inference can be optimized beyond PrismML’s fork, but it neither validates stock llama.cpp’s merged support nor provides an independently reproducible artifact. The case remains a backend-compatibility and correctness question awaiting a documented stock llama.cpp run.
2026-08-21T20:23:15Z
evidence attached: reddit.post.1vuq0oa — A from-scratch runtime benchmark supplies useful independent context on the practical performance of Bonsai ternary models on low-cost ARM hardware.
2026-08-21T08:32:20Z
No new evidence arrived within the staleness window, so the case remains an unvalidated implementation claim rather than a developing adoption signal. Keep it open but cold until a reproducible stock llama.cpp run reports checkpoint format, backend, correctness, quality, and speed.
2026-08-19T08:25:31Z
Anecdotal use of Bonsai 27B suggests some practical utility, but it does not identify stock llama.cpp, a backend, checkpoint format, correctness, or measured performance. The case still lacks independent validation of the merged support and its meaning is unchanged.
2026-08-19T03:31:28Z
The refreshed discussion adds skepticism and deployment questions but no independent run confirming correctness, backend compatibility, quality, or speed. The merge remains a concrete test opportunity, while the validation hypothesis is unchanged and now cooler pending actual results.
2026-08-18T19:39:35Z
No independent compatibility, correctness, quality, or performance testing has arrived; the score increase is only amplification of the already-known merge claim. The case remains a fresh implementation event awaiting substantive validation across common backends.
2026-08-18T19:30:34Z
grounded: known/medium — The radar already tracks the underlying Bonsai extreme-quantization claim in `radar:bonsai-extreme-quantization`; this case extends that open story into stock l
2026-08-18T19:26:00Z
case created — Merged mainline support is a concrete local-inference event whose compatibility, speed, and model quality remain to be validated.