pmttyji claims to have introduced B3S, a base-3 packing format for ternary GGUF model weights that stores them losslessly at 1.75 bits per weight, reportedly reducing weight memory by about 22%. The supplied GGUF documentation confirms that GGUF supports ternary quantization types, but the snippets do not independently establish B3S’s implementation, benchmarks, adoption, or runtime support; its practical value therefore remains contingent on compatible inference kernels and tooling.
B3S extends Scott’s hardware-aware local-inference practice with a specific lossless ternary representation that could increase model density on his self-hosted GPU/Ollama stack. The connection is actionable but contingent: supplied evidence does not establish compatible kernels, benchmarks, or adoption, so it would matter only if GGUF runtimes implement it efficiently.
dev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcradar:concept.quantizationradar:concept.ggufradar:concept.extreme-quantizationradar:llama-cpp-bonsai-ternary-supportradar:glm-lossless-weight-compression
queries asked of Scott's wikis
- ternary model packing and sub-2-bit inference
- local inference memory-bandwidth economics
- GGUF runtime and custom quantization support
- lossless weight encoding versus quantization
- local model density and hardware constraints
- quantization format adoption and kernel support
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 885h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-09-30T02:30:50Z
The Sherry 3:4 WebGPU demo is the first runnable sub-2bpw ternary result attached to this case, but it implements a different format (forced-sparsity, 1.375 bpw) on a tiny non-LLM scorer outside GGUF — it neither validates B3S nor corroborates its lossless claim, and it shows the ternary density frontier already sits below B3S's 1.75 bpw. B3S itself is unchanged: still an author-reported proposal with no code, no scale-preserving round-trip proof, and no runtime measurements.
2026-09-30T00:38:31Z
evidence attached: reddit.post.1wtp8h5 — A denser ternary format (Sherry 3:4 at 1.375 bpw) running on WebGPU contextualizes the ternary-packing density episode, though only demonstrated on a tiny non-LLM scorer.
2026-09-11T08:33:24Z
The refreshed discussion adds a plausible unpacking-cost explanation for why denser storage might not improve inference speed, but supplies no measurements or independent validation. B3S remains an author-reported representation improvement, not yet a demonstrated runtime-memory or throughput benefit for Scott’s local stack.
2026-09-09T08:28:54Z
No new evidence changes B3S from an author-reported packing proposal into a usable local-inference improvement. Its value remains contingent on inspectable code, scale-preserving round-trip tests, and runtime measurements; the unanswered scale objection neither validates nor disproves the claim.
2026-09-07T07:32:07Z
The stale review adds no technical evidence: B3S remains an author-reported packing proposal, not a demonstrated improvement to Scott’s local-inference stack. Keep it on a slower cadence pending inspectable code, scale-preserving round-trip tests, or runtime memory and throughput measurements.
2026-09-05T07:25:50Z
The refreshed comments add no substantive evidence beyond the already-known upstream-support dependency and an unresolved objection about scale preservation. B3S remains an author-reported packing format, with neither lossless equivalence nor practical inference-memory savings independently demonstrated.
2026-09-04T23:31:38Z
The refreshed discussion reinforces that upstream llama.cpp support is the decisive adoption path while adding skepticism about the claimed representation details; it provides no implementation, benchmark, or independent validation.
2026-09-04T19:39:18Z
The modest engagement increase adds no independent technical evidence; B3S remains an unverified author proposal awaiting code, runtime integration, or reproducible memory and performance results.
2026-09-04T19:36:15Z
grounded: converges/medium — B3S extends Scott’s hardware-aware local-inference practice with a specific lossless ternary representation that could increase model density on his self-hosted
2026-09-04T19:32:50Z
case created — The packing proposal is technically specific and economically relevant to local inference, but currently has only one lightly observed source and no linked implementation.