2026-10-11 17:10 UTC

Independent benchmarks will determine whether tensor-level bit allocation materially improves reasoning quality in ultra-low-bit Qwen3.5-4B quantizations at effectively unchanged model size.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference quantization qwenByteOtter

What is this?

ByteOtter created a Hugging Face quantization of Qwen3.5-4B whose model card reportedly claims a 16.67% reasoning gain from allocating bits at tensor level while keeping model size effectively unchanged. The supplied web results support the broader premise that tensor-sensitive allocation and importance matrices can materially affect very-low-bit quantization quality, while 2–3-bit performance depends heavily on implementation details and may incur inference-speed tradeoffs. However, none of the snippets independently benchmarks ByteOtter’s exact artifact or verifies its claimed reasoning improvement, so the result remains provisional pending third-party evaluation.

Why it matters to Scott

The radar already tracks essentially the same fixed-size, tensor-aware extreme-quantization validation question in `radar:gemma-tensor-level-iq2-quantization` and the broader dynamic-allocation claim in `radar:unsloth-dynamic-3-gguf-validation`; this is a Qwen3.5-4B instance rather than a new thesis. It still bears on Scott’s hardware-aware local inference and gamepc model-selection policy because credible independent benchmarks could establish a better quality-per-memory deployment option, with his Evaluation-Driven Development doctrine requiring repeatable tests rather than the model card’s claim.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentradar:gemma-tensor-level-iq2-quantizationradar:unsloth-dynamic-3-gguf-validationradar:concept.quantizationradar:concept.extreme-quantizationradar:concept.local-inference
queries asked of Scott's wikis
  • tensor-aware mixed-precision quantization strategy
  • local inference quality versus memory economics
  • independent benchmark design for reasoning models
  • ultra-low-bit model quality thresholds
  • importance matrices and selective precision
  • quantization effects on reasoning token usage

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation
LocalLLaMA
devildip911
🟧 echo.other ⭐The earliest primary artifact is the Hugging Face model repository, created by ByteOtter. Its model card says QLAB produced a reasoning-direByteOtter / QLAB——

Interpretation history

Decision trace