2026-10-11 17:11 UTC

Independent evaluations will determine whether preserving internal representation geometry during quantization-aware distillation improves NVFP4 model quality over KL-only distillation without reducing low-precision efficiency.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference nvfp4 quantization distillation

What is this?

A new arXiv paper proposes CKA-QAD, adding a lightweight layerwise alignment loss to NVFP4 quantization-aware distillation so that a student model preserves the BF16 teacher’s internal representation geometry rather than matching only its output distribution via KL divergence. The authors report that, on Qwen3-4B-Thinking-2507, the method recovers more benchmark accuracy than KL-only QAD—for example, 72.3% versus 68.5% on AIME25—after NVFP4 quantization. The supplied results do not identify the authors or provide independent replication, and they do not directly establish that the added training objective preserves the same deployed inference efficiency, beyond both approaches producing NVFP4 models.

Why it matters to Scott

CKA-QAD gives a concrete model-compression analogue of Scott’s Derivational Provenance position: matching the final answer does not establish that the underlying route or representation remains intact. It also bears directly on his hardware-aware local-inference work by proposing a potentially better quality-preservation objective for NVFP4 deployment, but independent evaluation and unchanged runtime efficiency remain unestablished.
ip:framework.derivational-provenanceip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferenceradar:compressed-llm-fidelity-safety-gapradar:concept.quantizationradar:concept.model-compressionradar:concept.model-distillationradar:concept.model-evaluationradar:concept.inference-efficiency
queries asked of Scott's wikis
  • quantization quality beyond output accuracy
  • internal representations as model evaluation signals
  • local inference precision-quality tradeoffs
  • distillation objectives for compressed models
  • independent evaluation of quantized models
  • NVFP4 hardware and inference strategy

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation
LocalLLaMA
Aaaaaaaaaeeeee245
🟧 echo.paper ⭐The authors’ original arXiv paper introduces CKA-QAD, arguing that “output matching alone can mask internal degradation” and proposing CKA rFangbo Tu, Junhua Zhao, Chi Liu, Xin Chen, Haifeng Wu, Jian Wan, Srinivasan Manoharan——

Interpretation history

Decision trace