2026-10-11 17:13 UTC

Independent testing will determine whether the English-focused Kimi K3 IQ2-XXS GGUF reduces storage from roughly 711GB to 478GB while preserving useful English-language capability.

state: expiredheat: lowuncertainty: highknownscott: mediumkimi-k3 local-inference model-compressionUnsloth

What is this?

A model upload by hellohazime, described as a Kimi K3 IQ2-XXS GGUF, claims to reduce storage from roughly 711 GB to 478.5 GB by retaining 576 of 896 experts and removing multilingual capability while targeting English use. One evidence title associates the release with Unsloth, but the supplied material does not establish Unsloth’s exact role. The search results provide no relevant corroboration or benchmarks, so preservation of useful English-language capability remains an unverified claim pending independent testing.

Why it matters to Scott

The radar already tracks this development under “Independent testing will determine whether Unsloth’s Kimi K3 GGUF quantizations enable stable local inference,” while this artifact adds a specific 478.5 GB expert-pruning and English-only capability claim. It bears on Scott’s hardware-aware local-inference work and evaluation doctrine because the storage gain is meaningful only if representative benchmarks verify retained capability.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferenceradar:unsloth-kimi-k3-gguf-local-validationradar:concept.model-compressionradar:compressed-llm-fidelity-safety-gap
queries asked of Scott's wikis
  • expert pruning versus quantization tradeoffs
  • local inference storage and hardware economics
  • language-specific model compression
  • capability preservation benchmarks for compressed models
  • sparse mixture-of-experts deployment
  • GGUF large-model inference tooling

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditKimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size
LocalLLaMA
Hannibalj2ca226100
🟧 echo.other ⭐The primary artifact is hellohazime’s REAP576-IQ2_XXS model upload. Its model card says it retains 576/896 experts, is 478.5 GB, and targetshellohazime——

Interpretation history

Decision trace