LocalLLaMA builder AdventurousTwo6445 claims closed-form trajectory weight surgery โ solving SwiGLU MLP weight updates directly from layer-to-layer hidden-state trajectories on a handful of calibration prompts โ transfers a 4B teacher's capabilities into 0.8B students and survives cross-architecture edits on fragile GPT-2 small; replication on other model pairs would establish editing-based capability transfer as a practical alternative to billion-token distillation.
state: watchingheat: lowuncertainty: mediumnovelscott: highweight-editing model-distillation small-local-modelsAdventurousTwo6445
What is this?
A hobbyist builder (Reddit: AdventurousTwo6445, GitHub repo: dsadawq3/DynamicTune โ the supplied material doesn't explicitly tie the two handles, but the repo's technique details match the post) claims a training-free capability-transfer method: treat the transformer as a residual dynamical system h_{l+1} = h_l + f_l(h_l), align teacher and student layer-to-layer hidden-state trajectories through a local orthogonal Procrustes atlas, and solve closed-form SwiGLU MLP weight updates from ~8 calibration prompts โ no backprop, no billion-token training run. Reported results: Qwen3.5 4B teacher into 0.8B student on an 8GB RX 580, with new capabilities (e.g. writing non-degenerate code) appearing on held-out prompts, plus cross-architecture edits surviving on GPT-2 small where small weight perturbations usually collapse the model. This sits between two established lineages visible in the search results: conventional knowledge distillation (teacher-student training on large datasets, per the distillation overviews) and MEMIT-style closed-form MLP weight editing (which targets factual associations, not trajectory/capability transfer). The evidence is self-reported, first-party, and unreplicated; nothing in the supplied material establishes independent verification.
Why it matters to Scott
Not a convergence with or challenge to any position Scott's canon holds โ it's a new third-party technique claim โ but if replication holds it bears directly on him: training-free 4Bโ0.8B capability transfer would upgrade the cheap rung of his model barbell and offer a closed-form alternative to the corpus-plus-fine-tune upgrade path his synthetic fine-tuning data-factory work embodies, and since the claimant ran it on an 8GB RX 580 he could test it on his own gamepc/Ollama stack rather than just read about it. As an unreplicated first-party claim it remains a watch-and-try item, distinct from the radar's Hirundo and grafting weight-editing episodes in targeting capability transfer rather than censorship removal or architecture conversion.
ip:concept.model-barbelldev:concept.synthetic-finetuning-datasetdev:project.gamepcdev:technology.ollamaradar:concept.model-distillationradar:concept.model-adaptationradar:concept.small-modelsradar:concept.activation-steeringradar:hirundo-westernized-qwenradar:echoleak-immunized-llama-weightsradar:qwen-model-grafting-encoder-decoder
queries asked of Scott's wikis
- small local models for coding agents and agent harnesses โ minimum viable model size
- weight editing / model grafting / model surgery prior claims and positions
- distillation vs fine-tuning vs editing โ cost of upgrading small open-weight models
- consumer GPU local inference economics โ 8GB VRAM, AMD ROCm constraints
- open-weights capability gap โ techniques for closing small-model deficits cheaply
- transformer as dynamical system / residual stream trajectories โ interpretability views
Measured heat
now 0 pts/hpeak 11 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 410h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p52 vs 1032 stories at the 336h mark (now 410h old) โ ahead of crowdstrike-safemind-security-agents (1.1x), behind agentdrive-persistent-shared-storage (0.9x)
Evidence (3) โ โญ canonical anchor
Interpretation history
2026-10-06T12:26:23Z
TPN Bench's third-party L4 run of the released FP16 checkpoint moves the case past pure self-report โ the headline numbers now rest on a public artifact someone else evaluated โ but it verifies one benchmark on the author's own 4Bโ0.8B pair and reaches us only through the author's post, so it is verification of results, not replication of the method. The hypothesis stays open pending someone re-running the surgery on a different teacher-student pair, which is its own stated resolution condition.
2026-10-06T11:33:25Z
evidence attached: reddit.post.1wyycx6 โ Independent corroboration: TPN Bench's third-party L4 verification of DynamicTune's 0.8B-beats-2B numbers is the external replication evidence the seed case awaits.
2026-10-06T10:11:21Z
origin walked (opencode/cheap-glm, conf 0.95): anchor reddit.post.1wyx2w9 -> echo.github.051a80c31f by dsadawq3 (GitHub; same project published on HuggingFace as org "F-Labs" and on Hacker News as "fermlon30000")
2026-10-06T09:38:33Z
grounded: novel/high โ Not a convergence with or challenge to any position Scott's canon holds โ it's a new third-party technique claim โ but if replication holds it bears directly on
2026-10-06T09:29:30Z
case created โ Concrete first-party technique claim for cheap small-model capability transfer that resolves on replication and is distinct from the existing grafting and Westernized-Qwen weight-editing cases.
Decision trace
- 10-11 12:07review_screenjev screen: no material development (noul=0.22)
- 10-08 00:59attention_routeThe editor compared this story and chose to keep watching.
- 10-06 23:26repriceTPN Bench's third-party L4 run of the released FP16 checkpoint moves the case past pure self-report โ the headline numbers now rest on a public artifact someone else evaluated โ but it verifies o
- 10-06 22:33attachIndependent corroboration: TPN Bench's third-party L4 verification of DynamicTune's 0.8B-beats-2B numbers is the external replication evidence the seed case awaits.
- 10-06 22:26propose_attachIndependent corroboration: TPN Bench's third-party L4 verification of DynamicTune's 0.8B-beats-2B numbers is the external replication evidence the seed case awaits.
- 10-06 21:11promote_anchororigin walk conf 0.95
- 10-06 20:38groundNot a convergence with or challenge to any position Scott's canon holds โ it's a new third-party technique claim โ but if replication holds it bears directly on him: training-free 4Bโ0.8B ca
- 10-06 20:29createConcrete first-party technique claim for cheap small-model capability transfer that resolves on replication and is distinct from the existing grafting and Westernized-Qwen weight-editing cases.