2026-10-11 16:35 UTC

Redditor No-Name-Person111 reports a ternary Bonsai 2 27B release in Prism ML's model collection, potentially expanding lower-memory options for local 27B-class inference.

state: corroboratedheat: lowuncertainty: lowconvergesscott: mediumternary-models local-inferencePrism ML
Surfaced 2026-09-19T10:23:40Z — PrismML’s original announcement says: “Today, we’re releasing Ternary Bonsai 2 27B, our most capable model yet.” It describes a Qwen3.8-27B- — The spread signal and sustained discussion around evaluation and a proposed Swift derivative justify moderate attention, but the supplied changes contain no released fix or new measured result. Bonsai remains a credible low-memory evaluation candidate whose practical value is constrained by unresolved reliability and completion-time tradeoffs, not an established substitute for larger-memory deployments.

What is this?

On September 17, 2026, PrismML released Ternary Bonsai 2 27B, an Apache-2.0 open-weight, ternary-compressed version of Qwen3.8-27B, with GGUF artifacts and local deployment positioned as its main use case. PrismML reports a roughly 5.8–5.9 GB model and an aggregate benchmark score of 84.78 versus 86.32 for the 54 GB full-precision baseline, but those are publisher-derived figures and the supplied community evidence reports mixed quality, looping, overthinking, and long completion times. Runtime support is also qualified: an Ollama mirror says the packs require PrismML’s llama.cpp fork and do not run in stock Ollama, while listed sizes vary by artifact and format.

Why it matters to Scott

Community testing converges with Scott’s hardware-aware inference and evaluation-driven development positions: model residency and headline benchmark retention are insufficient without testing runtime compatibility, task quality, tool calling, latency, and total memory on the actual stack. It is an actionable candidate for his gamepc/Ollama/Ask evaluation path, but fork requirements and unreliable completion behavior make it a test target rather than a deployment upgrade.
dev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisondev:project.gamepcdev:technology.ollamadev:project.askradar:bonsai-extreme-quantizationradar:llama-cpp-bonsai-ternary-supportradar:qwen38-27b-default-reasoning-costradar:qwen38-27b-16gb-quant-benchmarkradar:concept.ternary-modelsradar:concept.inference-economics
queries asked of Scott's wikis
  • ternary and sub-2-bit model deployment strategy
  • local inference memory versus time-to-useful-result
  • quantization quality evaluation for coding and agents
  • llama.cpp forks and unsupported model kernels
  • chat templates as agent reliability configuration
  • low-memory local models for tool calling

Measured heat

now 2 pts/hpeak 27 pts/hcomments 1/hpeers p46momentum: steady2 platformsage 602h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 14:00⭐ origin echo-reconstructedPrismML’s original announcement says: “Today, we’re releasing Ternary Bonsai 2 27B, our most capable model yet.” It describes a Qwen3.8-27B-
PrismML on blog (echo) · attributed from reddit.post.1wj84nz
—
09-17 22:16first on r/LocalLLaMA · published · +32.3hTernary Bonsai 2 27B
No-Name-Person111
—
09-17 22:16amplified on r/LocalLLaMAreddit.post.1wj84nz
No-Name-Person111
peak 96 · 38 comments · 7% of case engagement
09-18 04:50amplified on r/LocalLLaMAreddit.post.1wjgnok
notadithyabhat
peak 122 · 67 comments · 9% of case engagement
09-18 11:26amplified on r/LocalLLaMAreddit.post.1wjnklv
KURD_1_STAN
peak 149 · 64 comments · 10% of case engagement
09-18 12:04amplified on r/LocalLLaMA 👑reddit.post.1wjocnh
Secure_Recording_472
peak 148 · 182 comments · 16% of case engagement
09-18 15:21amplified on r/LocalLLaMAreddit.post.1wjt5f3
Thatisverytrue54321
peak 52 · 24 comments · 4% of case engagement
09-18 16:01amplified on r/LocalLLaMAreddit.post.1wju8ky
ali_byteshape
peak 165 · 65 comments · 11% of case engagement
11 more amplifiers in ainews.case_chain
09-17 22:20our radar first saw it · +32.3hdiscovery anchor: reddit.post.1wj84nz—
09-19 10:23reached heat=high · +68.4h · via ledger——
pace: p94 vs 1032 stories at the 336h mark (now 602h old) — ahead of ukisai-swift-family-release (1.0x), behind minetrials-hour-long-agent-evaluation (1.0x)

Evidence (18) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditTernary Bonsai 2 27B
LocalLLaMA
No-Name-Person1119638
🟧 echo.blog ⭐PrismML’s original announcement says: “Today, we’re releasing Ternary Bonsai 2 27B, our most capable model yet.” It describes a Qwen3.8-27B-PrismML——
🟠 redditTernary Bonsai is a headless chicken
LocalLLaMA
notadithyabhat11867
🟠 redditQuestion: UkisAI Swift Ternary Bonsai 2 27B?
LocalLLaMA
Secure_Recording_472148182
🟠 redditbonsai's document reveal how much cherry picked their headlines are
LocalLLaMA
KURD_1_STAN14964
🟠 redditPrism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison
LocalLLaMA
ali_byteshape16565
🟠 redditBonsai 2 27b Q2 - Donkey making coffee svg and a mushroom riding a donkey
LocalLLaMA
Thatisverytrue543215224
🟠 redditRTX 5090 Bonsai 2 27B vs Gemma 4 12B vs Qwen 3.5 9B Japanese voxel pagoda
LocalLLaMA
Fun-Meaning-647420185
🟠 redditTesting Ternary-Bonsai 27B Q2 on a GTX 1080 Ti
LocalLLaMA
chinto80015
🟠 redditTernary Bonsai 2 27B (1.75bpw) vs. Gemma 26B-A4B MoE
LocalLLaMA
autonoma_20421510
🟠 redditI tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)
LocalLLaMA
Danmoreng7995
🟠 redditTernary-Bonsai-2-27B-PQ2_0 is not completely lobotomized
LocalLLaMA
Fancy-Snow77240
🟠 redditWhat's the verdict on Ternary Bonsai 2 27B?
LocalLLaMA
PotterSkxawng3691
🟠 redditwhen will the new bonsai 27b be available on LM Studio?
LocalLLaMA
_maverick9806
🟠 redditTernary Bonsai 2 with Qwen Sharp Chat Template ?
LocalLLaMA
needthosepylons25
🟠 redditTested 15 local models for agent/tool use.. Bonsai 27B was last!
LocalLLaMA
Reno0vacio027
🟠 redditI re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome
LocalLLaMA
naklitechie3212
🟠 redditNow you can grow Bonsai on your potato
LocalLLaMA
jacek202398

Interpretation history

Decision trace