2026-10-11 17:09 UTC

ternary-models

band: warmmomentum: stable score: 0.394
temperature history

Episodes (4)

Evangelos Georganas and coauthors claim BITCOS losslessly exploits ternary-weight zero density to reach 1.485 bits per weight and improve decode throughput by up to 1.18Γ— on CPUs and 1.27Γ— on tested GPUs, potentially reducing local-inference memory and bandwidth costs.
seedconvergesscott: low
Independent testing will determine whether Maple-Preview’s ternary 20B MoE sustains roughly 120 tokens per second on an iPhone while retaining practically useful model quality.
expiredknownscott: low
Redditor No-Name-Person111 reports a ternary Bonsai 2 27B release in Prism ML's model collection, potentially expanding lower-memory options for local 27B-class inference.
corroboratedconvergesscott: medium
Prism ML (SkyIsNotGreen) claims its released Scion-35B-A3B β€” a 35B-A3B MoE shipped as one 11.3GB GGUF with ternary PQ2_0 expert banks plus embedded trained corrections at 2.61 bpw, a bundled MTP drafter, and a required llama.cpp fork β€” achieves Q4-class task retention at roughly half Q4_K_M's size, making ternary MoE a practically servable local tier; independent adoption and reproduced benchmarks confirm it, quiet fade closes it.
seednovelscott: medium