Xyntetik (Reddit's ZenZombie117) claims his released Xyntetik-Kvist-14B โ Muse-Glimmer 30B halved by width and distilled back with Ornith-1.0-9B as policy teacher, no RL โ retains near-full tool-task competence (57/60 held-out tasks vs the parent's 60) on a 24GB card; independent replication or adoption would establish distill-and-narrow as a practical route to small agent-capable local models.
state: watchingheat: mediumuncertainty: highconvergesscott: highlocal-models model-distillation agent-tool-useZenZombie117Xyntetik
What is this?
Meta shipped Muse Glimmer 30B on 10 August 2026: a 29.6B dense multimodal model under Apache 2.0, distilled from its larger Muse Spark family and architected (16:1 GQA, sliding-window attention, a DFlash speculative-decoding drafter) specifically so a 128K-token agent loop fits a 24GB consumer GPU. The case's claimant, Reddit user ZenZombie117 posting as Xyntetik, reports releasing Xyntetik-Kvist-14B โ that 30B halved by width and distilled back to competence using Ornith-1.0-9B as policy teacher with no RL โ retaining 57/60 of held-out tool tasks the parent scores 60/60 on. The supplied snippets corroborate the parent model and the claimant's ecosystem position (Ornith-1.0-9B appears on third-party benchmark-comparison pages as a real, benchmarked 9B model, maker not named in these snippets; a comment in Glimmer's official Hugging Face GGUF discussion promotes 'Xyntetik Runner' as a llama.cpp/ollama alternative with built-in schema support, consistent with an active local-inference tooling builder), but none of the snippets mention Kvist or test the 57/60 number. Replication is genuinely unresolved in this whole corner โ AI Weekly notes every published Glimmer benchmark is self-reported and OpenClaw tells readers to wait for independent numbers โ so the case's stated resolution path (independent replication or adoption) matches the actual evidentiary state.
Why it matters to Scott
A hobbyist builder independently optimized for exactly what usable-mass-over-unusable-power and the model barbell already argue โ trading 3 of 60 tool tasks for half the memory and a model that fits the 24GB card gamepc actually runs โ and the claim is replication-resolvable with assets Scott already owns: the trace-backed agent comparison fixture could test the 57/60 number directly, ask's local path (which currently suppresses native tools) is the deployment that a competent 14B would reopen, and his synthetic-finetuning/data-factory projects could rerun the distill-and-narrow recipe. Until independently replicated it stays seller-class evidence on his own evidence-class ladder, and the actor already carries a radar edge via the Xyntetik Runner truncation-safe tool-call case.
ip:concept.usable-mass-over-unusable-powerip:concept.model-barbellip:concept.model-dividendip:concept.evidence-class-ladderdev:project.gamepcdev:project.askdev:concept.trace-backed-agent-comparisondev:concept.synthetic-finetuning-datasetradar:xyntetik-truncation-safe-tool-callsradar:concept.model-distillationradar:concept.model-compressionradar:concept.tool-callingradar:concept.benchmark-integrityradar:haar-wavelet-llm-pruningradar:compressed-llm-fidelity-safety-gapradar:distillation-censorship-transfer
queries asked of Scott's wikis
- width pruning then distillation โ halving a model and recovering capability
- distillation without RL โ can a small policy teacher transfer tool-use competence
- compression-to-capability tradeoff: quantization vs pruning vs distillation
- local agent-capable models on a 24GB single consumer GPU
- schema-constrained tool calling support in local model runners and harnesses
- self-reported model evals vs independent replication practice
Measured heat
now 0 pts/hpeak 19 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 362h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p70 vs 1032 stories at the 336h mark (now 362h old) โ ahead of doltlite-2000-agent-pr-build (1.0x), behind never-give-up-adaptive-rl (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-09-28T10:53:02Z
The release moved from a quiet first-party claim to an actively discussed artifact: commenters are requesting 4-bit/8GB quants and anchoring it against gemma4 12B and qwen3.5 9B, while others challenge the eval's sensitivity โ so the case's live question shifts from 'is this release real' (it is) to 'will community quants and third-party runs produce the independent numbers the claim still lacks.' Attention now precedes verification; the resolution path is live but unexercised.
2026-09-28T06:38:26Z
origin walked (opencode/cheap-glm, conf 0.9): anchor reddit.post.1ws6bmh -> echo.github.e622760d07 by Joakimpalm-Zen (HuggingFace user; almost certainly the same person as Reddit poster ZenZombie117)
2026-09-28T06:34:05Z
grounded: converges/high โ A hobbyist builder independently optimized for exactly what usable-mass-over-unusable-power and the model barbell already argue โ trading 3 of 60 tool tasks for
2026-09-28T06:23:57Z
case created โ Concrete first-party model release with a bounded, replication-resolvable capability claim at the intersection of Scott's local-inference and agent-tool-use interests, and no existing open case covers distill-and-narrow agent capability.
Decision trace
- 10-07 20:26review_dormantscheduled targets exhausted or 28 quiet days
- 10-07 20:26drop_targetsquiet through full ladder or over cap 8
- 09-29 19:54review_screenOnly comment churn: a speculative suggestion about depth pruning, a style complaint about LLM-written prose, and removal of the REAP question already documented in the assessment. No new facts, quants
- 09-29 19:54review_screenjev screen borderline (noul=0.34) โ luna review
- 09-29 01:21sensor_dirtycomment_update
- 09-28 21:21sensor_dirtycomment_update
- 09-28 20:53repriceThe release moved from a quiet first-party claim to an actively discussed artifact: commenters are requesting 4-bit/8GB quants and anchoring it against gemma4 12B and qwen3.5 9B, while others challeng
- 09-28 17:21sensor_dirtyvelocity_spike
- 09-28 16:38promote_anchororigin walk conf 0.9
- 09-28 16:34groundA hobbyist builder independently optimized for exactly what usable-mass-over-unusable-power and the model barbell already argue โ trading 3 of 60 tool tasks for half the memory and a model that fits t
- 09-28 16:23createConcrete first-party model release with a bounded, replication-resolvable capability claim at the intersection of Scott's local-inference and agent-tool-use interests, and no existing open case c