On September 17, 2026, PrismML released Ternary Bonsai 2 27B, an Apache-2.0 open-weight, ternary-compressed version of Qwen3.8-27B, with GGUF artifacts and local deployment positioned as its main use case. PrismML reports a roughly 5.8–5.9 GB model and an aggregate benchmark score of 84.78 versus 86.32 for the 54 GB full-precision baseline, but those are publisher-derived figures and the supplied community evidence reports mixed quality, looping, overthinking, and long completion times. Runtime support is also qualified: an Ollama mirror says the packs require PrismML’s llama.cpp fork and do not run in stock Ollama, while listed sizes vary by artifact and format.
2026-10-11T14:24:26Z
llama.cpp PR #29600 merging the ternary Bonsai 2 27B release resolves the stock runtime support question — the fork requirement is lifting. Model-quality verdict remains settled (low agentic utility, real low-memory footprint, unreconciled headline claims). The DFlash2 drafter and now upstream llama.cpp support are concrete ecosystem mitigations; the proposed UkisAI Swift derivative has not materialized. Attention remains collapsed (0.33 pts/h vs 27 peak, peer percentile 33).
2026-10-11T13:39:44Z
evidence attached: reddit.post.1x37kxn — Links the llama.cpp PR #29600 that merges the ternary Bonsai 2 27B release tracked in the corroborated case.
2026-10-11T04:42:04Z
The jev screen flagged 'material development (noul=0.78)' but the only changes are negligible engagement bumps (+2 score, +1 comment) on the existing DFlash2 drafter post — no new evidence, implementations, or ecosystem developments. The case remains settled: model quality verdict is established (low agentic utility, real low-memory footprint, unreconciled headline claims), ecosystem questions (stock llama.cpp/Ollama/LM Studio support, MLX, UkisAI Swift derivative) are unchanged, and attention has collapsed to near-zero (0.5 pts/h vs 29 peak, peer percentile ~0). The noul signal appears to be noise on a cooling case.
2026-10-07T08:30:50Z
The retrained DFlash2 drafter is the first concrete ecosystem mitigation — third-party decode speedups (2.2x L4, 1.5x Mac, 1.2x Chrome, published weights) partially answer the slow-completion complaint without touching the settled low-agentic-quality verdict — moving the case from 'evaluation settled, ecosystem questions open' to 'evaluation settled, ecosystem self-help underway.' Heat stays low despite magnitude-valve eligibility: that spread reading reflects the launch window, while current velocity is ~2% of peak (0.5 pts/h vs 28.8), comments/h is ~0, peer percentile is median, and the periphery is contracting (recent additions are isolated single-digit posts), not expanding.
2026-10-07T08:25:40Z
evidence attached: reddit.post.1wzqyfl — A purpose-trained DFlash2 drafter for Bonsai 2 27B with cross-runtime speedups (2.2x L4, 1.5x Mac, 1.2x Chrome) is concrete ecosystem/practicality evidence for the ternary release.
2026-10-04T21:27:28Z
The 429-trial Toolery benchmark converts weeks of anecdote into a structured (single-run, contested) verdict — last of 15 for agent/tool use — consolidating the case's meaning from contested evaluation question to established limited-value release: real low-memory footprint, unreconciled headline claims, well-supported low agentic prior. The launch-time magnitude-valve spread no longer reflects reality — velocity is ~3% of peak, recent additions are isolated low-traction posts, and the periphery is contracting rather than expanding — so heat drops to low; the live remaining questions are ecosystem ones (stock llama.cpp/Ollama/LM Studio support, MLX, the proposed UkisAI Swift derivative), not model quality.
2026-10-04T20:26:26Z
evidence attached: reddit.post.1wxoiug — A 429-trial Toolery agent/tool-use benchmark ranks Bonsai 27B last of 15 local models — material usability context for the release, though the single-run methodology is contested in comments.
2026-09-24T01:12:11Z
grounded: converges/medium — Community testing converges with Scott’s hardware-aware inference and evaluation-driven development positions: model residency and headline benchmark retention
2026-09-24T01:08:46Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-23T13:25:49Z
evidence attached: reddit.post.1wo4ydh — Community follow-up on the same ternary Bonsai 2 release: tool-calling weakness on 8GB hardware reportedly fixed by a Qwen chat template, material context for judging its practical local-inference value.
2026-09-21T14:20:39Z
The latest verdict-seeking and LM Studio threads show continued demand among memory-constrained users but add no measured result or confirmed access change. They reinforce that fork/runtime support and inconsistent reliability remain adoption barriers while the still-expanding, unusually loud evaluation periphery sustains high attention.
2026-09-21T10:22:29Z
evidence attached: reddit.post.1wm84og — It provides user-side evidence about availability and hardware requirements for the open Bonsai 27B release.
2026-09-21T06:21:24Z
evidence attached: reddit.post.1wm3yvy — User reports provide early adoption evidence that may contradict the case's promise of a useful low-memory 27B option.
2026-09-20T05:23:34Z
A user's report of immediate looping in the official WebGPU demo extends the reliability concern to a previously untested deployment route, although missing prompts and runtime details prevent attribution to the model rather than its configuration. Positive reactions to the file-size-matched comparison do not supply its missing results; broad spread and an expanding deployment-test periphery sustain high attention without establishing practical utility.
2026-09-19T21:39:49Z
The newest independent test introduces a more relevant file-size-matched Qwen IQ2_XXS comparison, but its supplied excerpt contains no results, so the reassuring title does not establish competitive quality. The evaluation periphery continues expanding and the multi-platform spread signal remains loud, warranting high attention without upgrading confidence in useful inference economics.
2026-09-19T21:22:33Z
evidence attached: reddit.post.1wkwz69 — Independent testing materially contextualizes the released ternary Bonsai 2 27B model’s practical quality against similarly sized local alternatives.
2026-09-19T15:26:01Z
A new independent comparison against Qwen3.8 IQ3_XXS reports that Bonsai's smaller artifact delivers somewhat worse results while consuming more tokens and taking substantially longer, strengthening the case that weight savings do not imply cheaper useful inference. The expanding evaluation periphery and multi-platform spread justify retaining high attention, without establishing dependable task quality or promoting maturity.
2026-09-19T15:21:43Z
evidence attached: reddit.post.1wkorli — Independent local testing adds useful evidence on the released Bonsai model’s quality, memory, and speed tradeoffs versus Qwen quantization.
2026-09-19T10:23:40Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-19T02:27:11Z
The audiobook-pipeline comparison introduces a relevant structured-attribution workload, but the supplied excerpt stops before results and does not establish an 8 GB deployment success or a quality advantage. It sharpens the evaluation question without changing the assessment: low weight memory must translate into reliable task completion, not merely model residency.
2026-09-19T02:21:36Z
evidence attached: reddit.post.1wk94rq — Independent task-level testing of the released ternary Bonsai 2 27B provides useful corroboration about its practical quality and memory tradeoffs.
2026-09-18T19:57:54Z
The GTX 1080 Ti post adds a claimed Bonsai-family coding experiment, but its model name does not establish Bonsai 2 identity and the excerpt supplies no performance or task outcome. It does not yet extend the established deployment result to aging GPUs; the practical assessment remains unchanged and does not warrant urgent attention.
2026-09-18T19:22:26Z
evidence attached: reddit.post.1wjy4k4 — This independent hardware test supports the case that ternary Bonsai-27B makes 27B-class local inference feasible on aging GPUs.
2026-09-18T18:42:02Z
An RTX 5090 comparison extends deployment evidence beyond Macs and suggests that attractive task output can coexist with a substantial completion-time penalty; the claimed 15× slowdown is commenter testimony, not a verified measurement here. This strengthens the case for evaluating time-to-acceptable-result against smaller models, rather than treating low weight memory or decoding throughput as sufficient evidence of practical value.
2026-09-18T18:22:40Z
evidence attached: reddit.post.1wjwups — Independent local testing corroborates the Bonsai 2 27B release and adds practical quality and reasoning-token observations.
2026-09-18T16:42:37Z
Separate evaluator and deployment reports now corroborate Bonsai 2 as a runnable low-memory candidate rather than just a publisher announcement. They soften the blanket failure narrative but expose a practical distinction: fitting on a 16 GB Mac does not establish useful coding quality or acceptable completion latency.
2026-09-18T16:22:51Z
evidence attached: reddit.post.1wjt5f3 — A real local-use example supplies limited deployment evidence for the ternary Bonsai 2 release, though not independent quality validation.
2026-09-18T16:22:51Z
evidence attached: reddit.post.1wju8ky — Independent benchmarking provides useful evidence on whether Bonsai 2 delivers a practical quality-throughput tradeoff.
2026-09-18T12:36:32Z
UkisAI's reported testing adds a second firsthand account of overthinking problems, making the earlier looping failure harder to dismiss as an isolated configuration issue, though no reproducible comparison is supplied. A document-based critique also challenges the breadth of the 98.2% retention headline; this is now a low-memory evaluation candidate with substantive quality warnings, not an established practical upgrade.
2026-09-18T12:22:38Z
evidence attached: reddit.post.1wjnklv — The document-based critique materially challenges Bonsai's headline retention claims and should affect assessment of the existing release case.
2026-09-18T12:22:38Z
evidence attached: reddit.post.1wjocnh — The proposed Swifted Bonsai 2 release adds practitioner evidence that overthinking and token usage may be important weaknesses in the current ternary model.
2026-09-18T05:28:16Z
A firsthand report of repetitive thinking on a coding task adds a concrete usability warning, shifting this from an unvalidated low-memory release lead to a configuration-sensitive evaluation candidate. Suggestions to use Prism's llama.cpp fork and adjust thinking parameters are unverified troubleshooting, not evidence that the failure is fixed or intrinsic to the model.
2026-09-18T05:21:47Z
evidence attached: reddit.post.1wjgnok — This independent user report contradicts the practical usability implied by the ternary Bonsai 2 27B release.
2026-09-18T03:29:55Z
The discussion adds a concrete WebGPU demo link and claimed llama.cpp support, making this a more actionable evaluation lead rather than a validated local-inference upgrade. The performance figures remain publisher-derived testimony, with no independent run establishing compatibility or useful quality.
2026-09-17T22:30:32Z
grounded: known/medium — The grounded portion falls within the Bonsai low-bit deployment story already tracked in radar:bonsai-extreme-quantization and radar:llama-cpp-bonsai-ternary-su
2026-09-17T22:25:03Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1wj84nz -> echo.blog.4a65016fd0 by PrismML
2026-09-17T22:23:27Z
case created — The linked owner-hosted model collection makes this a bounded release lead, but the observation establishes neither runtime support nor quality.