2026-10-11 17:15 UTC

pmttyji claims the B3S base-3 GGUF format losslessly packs ternary model weights at 1.75 bits per weight, cutting weight memory by about 22% and potentially making ternary local models denser if runtime support follows.

state: seedheat: lowuncertainty: highconvergesscott: mediumlocal-inference model-quantization ggufpmttyji

What is this?

pmttyji claims to have introduced B3S, a base-3 packing format for ternary GGUF model weights that stores them losslessly at 1.75 bits per weight, reportedly reducing weight memory by about 22%. The supplied GGUF documentation confirms that GGUF supports ternary quantization types, but the snippets do not independently establish B3S’s implementation, benchmarks, adoption, or runtime support; its practical value therefore remains contingent on compatible inference kernels and tooling.

Why it matters to Scott

B3S extends Scott’s hardware-aware local-inference practice with a specific lossless ternary representation that could increase model density on his self-hosted GPU/Ollama stack. The connection is actionable but contingent: supplied evidence does not establish compatible kernels, benchmarks, or adoption, so it would matter only if GGUF runtimes implement it efficiently.
dev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcradar:concept.quantizationradar:concept.ggufradar:concept.extreme-quantizationradar:llama-cpp-bonsai-ternary-supportradar:glm-lossless-weight-compression
queries asked of Scott's wikis
  • ternary model packing and sub-2-bit inference
  • local inference memory-bandwidth economics
  • GGUF runtime and custom quantization support
  • lossless weight encoding versus quantization
  • local model density and hardware constraints
  • quantization format adoption and kernel support

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 885h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-04 19:04⭐ origin directly observed~22% less weight VRAM, lossless: base-3 packing for ternary GGUFs
pmttyji on r/LocalLLaMA
—
09-29 23:16first on r/LocalLLaMA · published · +604.2hSherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version
Brilliant-Hall1387
—
09-04 19:04amplified on r/LocalLLaMA 👑reddit.post.1w7dlo5
pmttyji
peak 35 · 11 comments · 76% of case engagement
09-29 23:16amplified on r/LocalLLaMAreddit.post.1wtp8h5
Brilliant-Hall1387
peak 12 · 2 comments · 23% of case engagement
09-04 19:20our radar first saw it · +0.2hdiscovery anchor: reddit.post.1w7dlo5—
pace: p63 vs 519 stories at the 720h mark (now 885h old) — ahead of ctx-agent-code-provenance (1.1x), behind otodock-self-hosted-agent-workspace (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐~22% less weight VRAM, lossless: base-3 packing for ternary GGUFs
LocalLLaMA
pmttyji3511
🟠 redditSherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version
LocalLLaMA
Brilliant-Hall1387122

Interpretation history

Decision trace