2026-10-11 16:38 UTC

Xyntetik (Reddit's ZenZombie117) claims his released Xyntetik-Kvist-14B โ€” Muse-Glimmer 30B halved by width and distilled back with Ornith-1.0-9B as policy teacher, no RL โ€” retains near-full tool-task competence (57/60 held-out tasks vs the parent's 60) on a 24GB card; independent replication or adoption would establish distill-and-narrow as a practical route to small agent-capable local models.

state: watchingheat: mediumuncertainty: highconvergesscott: highlocal-models model-distillation agent-tool-useZenZombie117Xyntetik

What is this?

Meta shipped Muse Glimmer 30B on 10 August 2026: a 29.6B dense multimodal model under Apache 2.0, distilled from its larger Muse Spark family and architected (16:1 GQA, sliding-window attention, a DFlash speculative-decoding drafter) specifically so a 128K-token agent loop fits a 24GB consumer GPU. The case's claimant, Reddit user ZenZombie117 posting as Xyntetik, reports releasing Xyntetik-Kvist-14B โ€” that 30B halved by width and distilled back to competence using Ornith-1.0-9B as policy teacher with no RL โ€” retaining 57/60 of held-out tool tasks the parent scores 60/60 on. The supplied snippets corroborate the parent model and the claimant's ecosystem position (Ornith-1.0-9B appears on third-party benchmark-comparison pages as a real, benchmarked 9B model, maker not named in these snippets; a comment in Glimmer's official Hugging Face GGUF discussion promotes 'Xyntetik Runner' as a llama.cpp/ollama alternative with built-in schema support, consistent with an active local-inference tooling builder), but none of the snippets mention Kvist or test the 57/60 number. Replication is genuinely unresolved in this whole corner โ€” AI Weekly notes every published Glimmer benchmark is self-reported and OpenClaw tells readers to wait for independent numbers โ€” so the case's stated resolution path (independent replication or adoption) matches the actual evidentiary state.

Why it matters to Scott

A hobbyist builder independently optimized for exactly what usable-mass-over-unusable-power and the model barbell already argue โ€” trading 3 of 60 tool tasks for half the memory and a model that fits the 24GB card gamepc actually runs โ€” and the claim is replication-resolvable with assets Scott already owns: the trace-backed agent comparison fixture could test the 57/60 number directly, ask's local path (which currently suppresses native tools) is the deployment that a competent 14B would reopen, and his synthetic-finetuning/data-factory projects could rerun the distill-and-narrow recipe. Until independently replicated it stays seller-class evidence on his own evidence-class ladder, and the actor already carries a radar edge via the Xyntetik Runner truncation-safe tool-call case.
ip:concept.usable-mass-over-unusable-powerip:concept.model-barbellip:concept.model-dividendip:concept.evidence-class-ladderdev:project.gamepcdev:project.askdev:concept.trace-backed-agent-comparisondev:concept.synthetic-finetuning-datasetradar:xyntetik-truncation-safe-tool-callsradar:concept.model-distillationradar:concept.model-compressionradar:concept.tool-callingradar:concept.benchmark-integrityradar:haar-wavelet-llm-pruningradar:compressed-llm-fidelity-safety-gapradar:distillation-censorship-transfer
queries asked of Scott's wikis
  • width pruning then distillation โ€” halving a model and recovering capability
  • distillation without RL โ€” can a small policy teacher transfer tool-use competence
  • compression-to-capability tradeoff: quantization vs pruning vs distillation
  • local agent-capable models on a 24GB single consumer GPU
  • schema-constrained tool calling support in local model runners and harnesses
  • self-reported model evals vs independent replication practice

Measured heat

now 0 pts/hpeak 19 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 362h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-26 14:00โญ origin echo-reconstructedThe HF model card for Xyntetik-Kvist-14B is the primary artifact the Reddit post summarizes ("Numbers, from the card"). It states: "Muse-Gli
Joakimpalm-Zen (HuggingFace user; almost certainly the same person as Reddit poster ZenZombie117) on github (echo) ยท attributed from reddit.post.1ws6bmh
โ€”
09-28 05:45first on r/LocalLLaMA ยท published ยท +39.8hLiked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of
ZenZombie117
โ€”
09-28 05:45amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1ws6bmh
ZenZombie117
peak 60 ยท 62 comments ยท 100% of case engagement
09-28 06:20our radar first saw it ยท +40.3hdiscovery anchor: reddit.post.1ws6bmhโ€”
pace: p70 vs 1032 stories at the 336h mark (now 362h old) โ€” ahead of doltlite-2000-agent-pr-build (1.0x), behind never-give-up-adaptive-rl (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditLiked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of
LocalLLaMA
ZenZombie1175962
๐ŸŸง echo.github โญThe HF model card for Xyntetik-Kvist-14B is the primary artifact the Reddit post summarizes ("Numbers, from the card"). It states: "Muse-GliJoakimpalm-Zen (HuggingFace user; almost certainly the same person as Reddit poster ZenZombie117)โ€”โ€”

Interpretation history

Decision trace