2026-10-11 17:10 UTC

Independent benchmarks will determine whether tensor-level precision allocation materially improves Gemma reasoning quality over conventional IQ2_XXS quantization at the same 3.3 GB memory budget.

state: expiredheat: lowuncertainty: highknownscott: mediumquantization local-inference inference-economicsByteOtterGoogle

What is this?

Google’s Gemma 4 is a model family designed with reasoning and inference efficiency in mind, while its quantization-aware variants target reduced memory use on consumer devices. The cited ByteOtter evidence claims tensor-specific precision allocation raises reasoning performance by 140.54% over conventional IQ2_XXS quantization without exceeding a 3.3 GB budget; a separate search result reports a smaller 8.55% coding gain from tensor-level allocation at Q3. However, the supplied snippets do not independently verify the 140.54% result, its methodology, or the same-budget comparison, so broader benchmarking remains necessary.

Why it matters to Scott

Scott’s Evaluation-Driven Development and Hardware-aware Local Inference pages already establish both the need for repeatable independent evaluation and precision/memory policy as part of local deployment. The specific tensor-level allocation claim is unverified but could affect quantization choices on his gamepc/Ollama stack if same-budget tests show materially better reasoning quality.
ip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.quantizationradar:concept.extreme-quantizationradar:concept.model-evaluationradar:concept.local-inferenceradar:concept.inference-economics
queries asked of Scott's wikis
  • tensor-level mixed-precision quantization
  • quality per VRAM for local models
  • benchmarking quantized reasoning models
  • local inference memory economics
  • task-aware precision allocation
  • quantization benchmark validity

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Gemma 4 E4B IQ2_XXS: + 140.54% Reasoning Performance From Tensor Level Quantization Allocation
LocalLLaMA
devildip7625

Interpretation history

Decision trace