PrismML has released Bonsai, a family of ultra-low-bit language models intended for local deployment, including 1-bit and ternary variants; its demo supports Mac/Metal and other backends, while the company says the 3.9GB 1-bit 27B model can run on an iPhone. PrismML claims its 1.58-bit ternary models preserve comparatively strong benchmark accuracy at roughly one-ninth the memory footprint of 16-bit models, with a 27B version described as about 1.7 bits per weight. The supplied snippets establish local inference and company-reported quality claims, but they do not independently establish usable fine-tuning on Apple hardware or the full claimed 1.1–1.7-bit range.
2026-07-29T07:25:27Z
The case resolves as absorbed: Bonsai's extreme quantization is validated for deployment and fine-tuning on consumer hardware, but usable quality is sharply workload-specific — adequate for casual chat and tutoring, poor for reasoning, coding, and agentic tool use. Apple-specific fine-tuning remains unvalidated but the broader pattern is clear.
2026-07-29T07:21:02Z
evidence attached: reddit.post.1v9nwzx — Real AMD tool-use testing supports Bonsai’s usable quality at extreme quantization while also exposing current agentic-reliability limits.
2026-07-27T19:25:10Z
The nominal new-evidence trigger exposes no substantive result and continues the long pattern of repetitive amplification. Compact sub-2-bit deployment and ternary trainability are corroborated, but quality remains workload-specific and reproducible Apple-side fine-tuning unresolved.
2026-07-27T18:26:06Z
The only change is modest engagement growth on the already-priced Jetson deployment, not new technical evidence. Compact sub-2-bit inference and ternary trainability remain corroborated, while quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved.
2026-07-26T17:26:32Z
The trigger contains no identifiable new result and continues the engagement-only amplification pattern. Compact sub-2-bit inference and ternary trainability are corroborated, but useful quality remains workload-specific and reproducible Apple-side fine-tuning unresolved.
2026-07-26T11:22:19Z
The apparent new-evidence trigger is another engagement-only refresh and does not change the case. Compact sub-2-bit inference and ternary trainability are corroborated, but quality remains workload-specific and reproducible Apple-side fine-tuning unresolved; pause frequent checks pending substantive benchmarks or training results.
2026-07-26T05:21:24Z
The nominal new-evidence trigger contains no identifiable substantive result, continuing a long run of repetitive amplification. Compact sub-2-bit deployment and ternary trainability are corroborated, but useful quality remains workload-specific and reproducible Apple-side fine-tuning is unresolved.
2026-07-26T03:21:36Z
The nominal attachment exposes no substantive result beyond the already-priced implementations and mixed quality reports. Compact sub-2-bit inference and ternary trainability remain corroborated, but quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved; repeated refreshes are noise.
2026-07-26T02:21:39Z
The refreshed comments add no substantive evidence beyond already-priced implementations and mixed quality reports. Compact sub-2-bit deployment and ternary trainability remain corroborated, while useful quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved.
2026-07-25T23:23:37Z
The nominal new-evidence trigger contains no identifiable substantive result beyond the already-priced implementations and mixed quality reports. Compact deployment and ternary trainability remain corroborated, but quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved; further engagement-only refreshes are noise.
2026-07-25T22:28:17Z
No substantive evidence has arrived beyond the already-priced Jetson deployment; the minor engagement change is repetitive amplification. Extreme memory savings and ternary trainability remain corroborated, while useful quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved.
2026-07-25T18:28:23Z
The Jetson implementation further establishes that roughly 1-bit Bonsai can fit and run on constrained consumer-class hardware, extending deployment validation beyond Apple devices. It does not change the core uncertainty: quality remains workload-specific and reproducible Apple-side fine-tuning is still unproven.
2026-07-25T18:21:46Z
evidence attached: reddit.post.1v6evbe — Independent real-hardware deployment reports Bonsai 27B running at roughly 1-bit on a 16GB Jetson, materially supporting practical extreme-quantization claims.
2026-07-25T17:22:57Z
The only change is modest engagement growth on an already-priced negative usage report, not new technical evidence. Compact inference and ternary trainability remain corroborated, while quality is confined to lighter workloads and reproducible Apple-specific fine-tuning remains unresolved.
2026-07-25T16:22:30Z
The nominal new-evidence trigger contains no substantive result, extending repetitive amplification rather than changing the case. Compact deployment and ternary trainability remain corroborated, while quality is limited to lighter workloads and reproducible Apple-specific fine-tuning remains unresolved.
2026-07-25T09:28:04Z
No substantive new evidence beyond the already-priced deployment, benchmark, and fine-tuning reports; continued engagement-only refreshes. Sub-2-bit deployment and ternary trainability remain corroborated, but broadly usable quality stays workload-specific (weak on reasoning/coding, adequate for casual/tutoring use) and reproducible Apple-specific fine-tuning remains unresolved.
2026-07-25T00:22:00Z
The nominal evidence trigger contains no substantive new result and continues repetitive amplification. Compact inference and ternary trainability are corroborated, but quality remains limited to lighter workloads and reproducible Apple-specific fine-tuning remains unresolved.
2026-07-24T21:23:42Z
The nominal new-evidence trigger exposes no substantive result beyond the already-priced deployment, training, and mixed quality reports. Bonsai remains validated for compact inference and ternary trainability, but useful quality is limited to lighter workloads and reproducible Apple-specific fine-tuning remains unresolved.
2026-07-24T20:25:58Z
The trigger contains no identifiable new substantive evidence, only a negligible discussion refresh. Compact deployment and ternary trainability remain corroborated, but useful quality is confined to lighter workloads and reproducible Apple-specific fine-tuning remains unsettled.
2026-07-24T19:26:14Z
The latest trigger is only a negligible engagement refresh and adds no substantive evidence. Compact cross-platform deployment and ternary trainability remain corroborated, while quality is useful only for lighter workloads and reproducible Apple-specific fine-tuning remains unsettled.
2026-07-24T17:28:09Z
No substantive evidence has arrived beyond the already-priced MacBook usage report; this is another engagement-only refresh. Bonsai is established as compact and trainable, but useful quality remains limited to lighter workloads while broad capability retention and reproducible Apple-side fine-tuning remain unsettled.
2026-07-24T16:27:48Z
Firsthand MacBook use supports a narrow form of usable quality for casual conversation, tutoring, and lightweight literature review, partially balancing failures on coding and reasoning. It still does not validate broad capability retention or reproducible Apple-side fine-tuning, so the hypothesis remains workload-specific and unsettled.
2026-07-24T16:21:50Z
evidence attached: reddit.post.1v5etch — Independent firsthand use on a 16GB MacBook provides practical, though anecdotal, evidence that Bonsai's extreme quantization retains useful quality.
2026-07-24T15:23:23Z
The trigger exposes no substantive evidence beyond the already-priced implementations, benchmarks, and hands-on failures. Bonsai’s compact deployment and ternary trainability remain corroborated, but broadly usable quality is increasingly contradicted for reasoning and coding, while reproducible Apple-specific fine-tuning remains unresolved.
2026-07-24T14:30:12Z
No substantive new evidence beyond minor engagement refreshes on already-priced items; the picture remains stable: sub-2-bit deployment and ternary trainability are established, while broad usable quality and reproducible Apple-side fine-tuning stay unresolved and workload-specific. Continued repetitive amplification without new benchmarks.
2026-07-24T13:24:01Z
The new hands-on failure adds to independent benchmark and real-use evidence that Bonsai’s sub-2-bit memory advantage does not preserve broadly usable quality, especially for reasoning and coding. Ternary trainability remains corroborated, but reproducible Apple-specific fine-tuning is still unresolved, so the hypothesis is increasingly narrowed rather than disproved outright.
2026-07-24T13:21:16Z
evidence attached: reddit.post.1v5a1br — This hands-on report questions whether the 27B ternary Bonsai model retains usable quality, directly bearing on the extreme-quantization hypothesis.
2026-07-23T15:25:22Z
The latest discussion explains why memory-side execution could amplify Bonsai’s deployment advantage, but it does not validate retained quality or reproducible Apple-side fine-tuning. The case remains established for compact cross-platform inference and ternary trainability, not for broadly usable capability at extreme bit rates.
2026-07-23T15:21:40Z
evidence attached: hn.story.49022097 — The discussion adds contextual support for why extreme quantization and memory-bandwidth changes matter for practical on-device deployment, but offers no new validation.
2026-07-23T10:31:32Z
The new DRAM demonstration adds another deployment path but no meaningful evidence on retained quality or reproducible Apple-side fine-tuning. Cross-platform sub-2-bit operation and ternary trainability remain corroborated, while practical quality is sharply workload-dependent and the broader hypothesis stays unsettled.
2026-07-23T10:21:45Z
evidence attached: hn.story.49019271 — Describes a method to run Bonsai extreme-quantized models on edge devices, corroborating the feasibility of Bonsai deployment.
2026-07-23T06:27:55Z
The latest activity is only a minor engagement refresh on already-priced evidence, not new validation. Cross-platform sub-2-bit deployment and ternary trainability remain corroborated, while broad quality retention and reproducible Apple-side fine-tuning remain unresolved.
2026-07-23T03:25:28Z
The trigger adds no identifiable substantive evidence beyond the already-priced deployment, training, and quality reports. Sub-2-bit operation and ternary trainability remain corroborated, but severe workload-dependent capability loss and unresolved Apple-specific fine-tuning keep the broader hypothesis unsettled.
2026-07-23T02:25:07Z
Another real-use report aligns with existing benchmarks that Bonsai’s extreme compression imposes severe coding and agentic capability loss, further confining “usable quality” to basic or selected workloads. The report is too thin to resolve the hypothesis, and reproducible Apple-specific fine-tuning remains unsettled.
2026-07-23T02:21:00Z
evidence attached: reddit.post.1v3zech — A real-use coding report weakly contradicts Bonsai's claimed practical quality, despite limited detail and engagement.
2026-07-22T22:25:25Z
The nominal new-evidence trigger contains no usable result beyond the already-priced implementations and benchmarks, so it is repetitive amplification. Cross-platform sub-2-bit deployment and ternary trainability are corroborated, but broad quality retention and reproducible Apple-side fine-tuning remain unresolved.
2026-07-22T21:25:59Z
The trigger exposes no identifiable evidence beyond the already-priced deployment, benchmark, runtime, and ternary-training reports, so it is repetitive amplification. Cross-platform sub-2-bit operation and trainability are corroborated, but broad quality retention and reproducible Apple-side fine-tuning remain unresolved.
2026-07-22T20:33:18Z
No substantive evidence arrived beyond the already-priced CPU implementation; this is another engagement-only refresh. Cross-platform deployment and ternary trainability are corroborated, while broad quality retention and reproducible Apple-side fine-tuning remain unresolved.
2026-07-22T19:27:57Z
Project Zero adds an independent CPU runtime and performance evidence, making Bonsai’s deployment advantage more robust across consumer hardware. It does not address the remaining load-bearing questions: workload-level quality retention and reproducible Apple-side fine-tuning.
2026-07-22T19:21:17Z
evidence attached: reddit.post.1v3pn1w — The independent Project Zero implementation provides practical CPU inference evidence and benchmarks for Bonsai's extreme-quantization claims.
2026-07-22T09:26:28Z
No substantive new evidence beyond the already-priced deployment, benchmark, and ternary fine-tuning reports; this is another engagement-only refresh. Extreme memory savings and ternary trainability remain corroborated; broad quality retention and reproducible Apple-specific fine-tuning remain unsettled and workload-specific.
2026-07-22T08:27:53Z
The latest activity is only a minor comment-count refresh on already-priced evidence, with no new substantive result. Sub-2-bit deployment and ternary trainability remain corroborated; broad quality retention and reproducible Apple-specific fine-tuning remain unsettled and workload-specific.
2026-07-22T03:22:03Z
The trigger adds no substantive evidence beyond existing deployment, benchmark, and ternary-training reports. Extreme memory savings and trainability remain corroborated, but broad quality retention and reproducible Apple-specific fine-tuning remain unsettled; repeated engagement is noise.
2026-07-21T23:31:43Z
The trigger contains no identifiable new result beyond existing deployment, benchmark, and ternary-training evidence, so it is repetitive amplification. Extreme memory savings and trainability remain corroborated, while broad quality retention and reproducible Apple-specific fine-tuning remain unsettled.
2026-07-21T21:28:48Z
The attachment adds no substantive result beyond existing deployment, benchmark, and ternary-training evidence, so this remains repetitive amplification. Extreme memory savings and trainability are corroborated, but broad quality retention and reproducible Apple-specific fine-tuning remain unsettled; revisit only when independent benchmarks or training results arrive.
2026-07-21T20:30:11Z
No substantive new evidence accompanies the trigger; it is further repetitive amplification of already-established deployment and ternary trainability. Broad quality retention and reproducible Apple-specific fine-tuning remain unsettled, so wait for independent benchmarks or training results.
2026-07-21T19:26:14Z
The supposed new evidence is only a negligible engagement change and adds no substantive validation. Extreme memory savings and ternary trainability remain corroborated, while broad quality retention and reproducible Apple-specific fine-tuning remain unsettled; stop frequent checks until reproducible results appear.
2026-07-21T18:27:27Z
The apparent new-evidence trigger yields only a minor engagement increase, not fresh validation. Extreme memory savings and ternary trainability remain corroborated, while broad quality retention and reproducible Apple-side fine-tuning remain unsettled and workload-specific.
2026-07-21T16:33:07Z
The nominal new-evidence trigger contains no identifiable substantive result, continuing repetitive amplification rather than advancing the core quality or Apple-specific fine-tuning claims. Extreme memory savings and ternary trainability remain corroborated, but usefulness is workload-specific and dependable Apple-side tuning remains unsettled.
2026-07-21T15:32:19Z
The trigger adds no substantive evidence beyond the existing deployment, benchmark, and ternary fine-tuning reports. Extreme memory savings and trainability are corroborated, but broadly usable quality and reproducible Apple-specific fine-tuning remain unsettled; continued engagement is repetitive amplification.
2026-07-21T14:28:47Z
The trigger exposes no substantive new result beyond the already-priced deployment, benchmark, and fine-tuning reports. Extreme memory savings and ternary trainability remain corroborated, but broadly usable quality and reproducible Apple-specific fine-tuning are still unsettled; repeated engagement is noise.
2026-07-21T13:23:04Z
The trigger adds no identifiable substantive evidence beyond the already-priced deployment, benchmark, and ternary fine-tuning reports. Extreme memory savings and trainability are corroborated, but broadly usable quality and reproducible Apple-specific fine-tuning remain unsettled; further engagement alone is repetitive amplification.
2026-07-21T12:23:04Z
No material evidence has arrived beyond the already-priced ternary fine-tuning implementation; the latest activity is repetitive amplification. Trainability and extreme memory savings are corroborated, but Apple-specific reproducibility and broadly usable quality remain unsettled.
2026-07-21T11:27:23Z
A second hands-on fine-tuning implementation materially strengthens the claim that Bonsai’s ternary weights remain trainable, rather than merely deployable. It still does not establish reproducible Apple-specific tuning or overcome benchmark evidence that useful quality is sharply workload-dependent.
2026-07-21T11:20:59Z
evidence attached: reddit.post.1v2egi3 — Hands-on evidence that PrismML's ternary Bonsai models are fine-tunable materially strengthens the case's consumer-hardware usability claim.
2026-07-21T10:23:37Z
The nominal new-evidence trigger contains no identifiable substantive result, continuing repetitive amplification rather than validating retained quality or reproducible Apple-side fine-tuning. Deployment and extreme memory savings are established, but usefulness remains workload-specific and the fine-tunability claim unresolved.
2026-07-21T09:25:12Z
The latest trigger exposes no substantive evidence beyond prior deployment demonstrations and benchmarks, continuing repetitive amplification. Extreme memory savings are established, but quality remains workload-specific and reproducible Apple-side fine-tuning is still unvalidated.
2026-07-21T08:23:40Z
The nominal new-evidence trigger adds no substantive result, extending the repetitive amplification around already-established memory and deployment gains. Usable quality remains sharply workload-dependent, while reproducible Apple-side fine-tuning is still unvalidated.
2026-07-21T07:25:13Z
The attachment exposes no substantive new result beyond prior deployment reports and benchmarks, so it is further repetitive amplification. Extreme memory savings are established, but usable quality remains workload-specific and reproducible Apple-side fine-tuning is still unvalidated.
2026-07-21T06:28:14Z
The refreshed discussion adds only repetitive attention, not new validation of retained quality or dependable Apple-side fine-tuning. Keep the case cold: deployment and memory savings are established, while practical capability remains narrow and the fine-tunability claim unresolved.
2026-07-21T05:30:48Z
The nominal new-evidence trigger contains no identifiable substantive result, continuing repetitive amplification of already-corroborated sub-2-bit deployment. Usable quality remains workload-specific, and reproducible Apple-side fine-tuning is still unvalidated.
2026-07-21T04:24:48Z
The trigger adds no identifiable evidence beyond the already-priced deployment reports and benchmarks. Extreme memory savings remain corroborated, but usable quality is workload-specific and reproducible Apple-side fine-tuning remains unresolved; further engagement is repetitive amplification.
2026-07-21T03:26:59Z
The low-end-hardware report modestly broadens confirmation that Bonsai’s memory savings translate into accessible deployment, but it adds nothing on retained quality or reproducible Apple-side fine-tuning. The case remains a validated deployment result with workload-specific capability loss and an unresolved fine-tunability claim.
2026-07-21T03:21:07Z
evidence attached: reddit.post.1v25zkd — Provides an independent usage report that Bonsai's roughly one-bit quantization is runnable on very low-end hardware, supporting its practical consumer-device claims.
2026-07-21T02:23:48Z
The trigger supplies no identifiable new evidence beyond the already-priced deployment and benchmark reports. Extreme memory savings are corroborated, but useful quality remains workload-specific and reproducible Apple-side fine-tuning is still unvalidated; repeated engagement is noise until substantive results appear.
2026-07-21T01:26:06Z
The attachment yields no identifiable new evidence on retained quality or reproducible Apple-side fine-tuning, continuing the pattern of repetitive amplification. Sub-2-bit deployment is corroborated, but practical quality remains workload-specific and the fine-tunability claim unresolved.
2026-07-21T00:22:51Z
No identifiable new evidence advances retained quality or reproducible Apple-side fine-tuning; this is further repetitive amplification of already-corroborated sub-2-bit deployment. Keep the case cold pending substantive benchmarks or training results.
2026-07-20T23:24:09Z
The trigger adds no identifiable evidence beyond the already-priced deployment and Terminal-Bench reports. Extreme memory savings remain corroborated, but retained quality is narrow and dependable Apple-side fine-tuning remains unvalidated; further engagement without reproducible results is noise.
2026-07-20T22:25:04Z
The latest trigger adds no usable evidence beyond the already-priced Terminal-Bench result, so it does not change the narrowed interpretation: extreme memory savings are real, but usable quality remains workload-specific and Apple-side fine-tunability unvalidated. Further engagement is repetitive amplification until reproducible quality or fine-tuning results appear.
2026-07-20T21:22:09Z
Independent Terminal-Bench testing adds practical evidence that Bonsai’s memory advantage survives deployment but carries substantial capability loss on agentic coding tasks. This further narrows “usable quality” to selected workloads while leaving reliable Apple-side fine-tuning unresolved.
2026-07-20T21:21:05Z
evidence attached: reddit.post.1v1ya97 — Independent Terminal-Bench testing gives useful but weak evidence that 2-bit Bonsai runs on 8GB VRAM while trailing larger models substantially.
2026-07-20T15:37:39Z
The new attachment contains no usable evidence and continues the repetitive amplification pattern. Sub-2-bit deployment is established, but retained quality and dependable Apple-side fine-tuning remain workload-specific and unvalidated.
2026-07-20T14:27:31Z
The latest trigger contains no usable new evidence, extending the pattern of repetitive amplification rather than advancing quality or Apple fine-tuning validation. Keep the case cold until reproducible benchmarks or training results address its workload-specific weaknesses.
2026-07-20T13:25:14Z
The new attachment contains no usable independent evidence and does not advance the unresolved claims about retained quality or dependable Apple-side fine-tuning. Repeated engagement updates are now noise around already-corroborated deployment, so the case should remain cold until substantive benchmarks or reproducible training results appear.
2026-07-20T12:26:42Z
The nominally new attachment adds no discernible evidence beyond the already-priced deployment demonstrations and limited Apple fine-tuning report. Sub-2-bit operation remains corroborated, but retained quality and reliable fine-tunability are still workload-specific and unsettled.
2026-07-20T11:24:40Z
The latest attachment provides no discernible new evidence on retained quality or reliable Apple-side fine-tuning, so it is repetitive amplification rather than a substantive advance. Cross-platform sub-2-bit deployment remains corroborated, while practical usefulness is still narrow and workload-dependent.
2026-07-20T10:22:22Z
The nominally new attachment adds no discernible independent evidence beyond the already-priced deployment demonstrations and limited fine-tuning report. Sub-2-bit operation remains corroborated, while broadly usable quality and reliable Apple-side fine-tunability remain unsettled and workload-specific.
2026-07-20T09:24:14Z
No material new evidence strengthens the unresolved quality or Apple fine-tuning claims; the latest activity only repeats that sub-2-bit deployment works. Bonsai remains promising for narrow instruct workloads, with reasoning, multilingual performance, and stability still limiting broader usefulness.
2026-07-20T08:22:23Z
The browser WebGPU demo broadens confirmation that Bonsai’s extreme quantization is deployable across consumer hardware, but it does not strengthen the core claims about retained quality or Apple-side fine-tunability. The case remains workload-specific and technically promising rather than generally validated.
2026-07-20T08:20:35Z
evidence attached: hn.story.48936994 — This browser WebGPU demo is independent corroboration that Bonsai's highly quantized models can run on consumer hardware.
2026-07-20T04:23:03Z
grounded: converges/medium — Bonsai potentially extends Scott’s hardware-aware local-inference work and “usable mass” thesis by making much larger models deployable within consumer-device m
2026-07-20T02:20:59Z
The attached evidence adds no material independent confirmation beyond the already-priced fine-tuning and benchmark reports. Bonsai remains operational on Apple hardware, but usable quality is workload-specific and broad retention claims remain unsettled.
2026-07-20T01:36:24Z
Independent fine-tuning and benchmark reports now substantiate that Bonsai is operational at extreme bit rates, but narrow the quality claim: instruct use may remain viable while reasoning, multilingual behavior, and stability degrade sharply. The case is now about workload-specific usefulness rather than broadly retaining near-baseline quality.
2026-07-20T01:35:44Z
evidence attached: reddit.post.1v164ee — Direct benchmark comparison supplies evidence about Bonsai's quality loss under reasoning and extreme quantization.
2026-07-19T11:24:18Z
case created — A mobile demonstration and a separate Metal fine-tuning report make Bonsai's practical extreme-quantization claims worth verifying.