2026-10-11 18:04 UTC

Daxfortuna reports that llama.cpp quantization fallbacks leave some GGUF files labeled as lower-bit formats than their tensors actually use, affecting 64 of 443 audited files and undermining reproducible local-model packaging.

state: expiredheat: lowuncertainty: highconvergesscott: mediumquantization gguf local-inference llama-cppDaxfortunallama.cpp

What is this?

Daxfortuna reports an audit of 443 GGUF files across 25 repositories, finding 64 whose filenames claim a lower-bit quantization than some tensors actually use. The supplied llama.cpp discussion confirms that mixed tensor types can be intentional—for example, 1D tensors may remain at 16/32-bit and K-quants may assign higher precision to selected tensors—so the evidence establishes a mismatch between shorthand labels and tensor contents, but not necessarily malformed files. The snippets are too thin to verify the audit methodology, identify affected repositories, or determine whether llama.cpp fallback behavior rather than accepted quantization conventions caused every mismatch.

Why it matters to Scott

The audit independently supports Scott’s provenance-coupled-work position and could justify tensor-level validation in his active local-model stack, since misleading quantization labels may affect memory planning and reproducible packaging. The supplied evidence does not establish malformed GGUFs or prove that llama.cpp fallback behavior caused all mismatches, so this extends his verification practice rather than overturning an existing claim.
ip:framework.provenance-coupled-workdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.ggufradar:concept.quantizationradar:concept.model-provenanceradar:concept.llama-cppradar:llama-cpp-gguf-loader-hardening
queries asked of Scott's wikis
  • local-model artifact provenance and reproducible packaging
  • quantization labels versus effective tensor precision
  • GGUF validation and model supply-chain auditing
  • llama.cpp conversion and quantization workflows
  • reproducible builds for local inference artifacts
  • machine-readable model metadata and integrity checks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI audited 443 GGUF quants across 25 repos. 64 of them can't be the quant their filename claims.
LocalLLaMA
Daxfortuna7830
🟧 echo.github ⭐The first complete primary artifact for the Reddit post is JoshBolding's initial ggufaudit commit. It explicitly says it “includes a live ceJoshBolding——
🟠 redditNemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
LocalLLaMA
Daxfortuna216
🟠 redditDeceptive model quantization from AtomicChat?
LocalLLaMA
po_stulate5140

Interpretation history

Decision trace