2026-10-11 17:13 UTC

LocalLLaMA builder AdventurousTwo6445 claims closed-form trajectory weight surgery โ€” solving SwiGLU MLP weight updates directly from layer-to-layer hidden-state trajectories on a handful of calibration prompts โ€” transfers a 4B teacher's capabilities into 0.8B students and survives cross-architecture edits on fragile GPT-2 small; replication on other model pairs would establish editing-based capability transfer as a practical alternative to billion-token distillation.

state: watchingheat: lowuncertainty: mediumnovelscott: highweight-editing model-distillation small-local-modelsAdventurousTwo6445

What is this?

A hobbyist builder (Reddit: AdventurousTwo6445, GitHub repo: dsadawq3/DynamicTune โ€” the supplied material doesn't explicitly tie the two handles, but the repo's technique details match the post) claims a training-free capability-transfer method: treat the transformer as a residual dynamical system h_{l+1} = h_l + f_l(h_l), align teacher and student layer-to-layer hidden-state trajectories through a local orthogonal Procrustes atlas, and solve closed-form SwiGLU MLP weight updates from ~8 calibration prompts โ€” no backprop, no billion-token training run. Reported results: Qwen3.5 4B teacher into 0.8B student on an 8GB RX 580, with new capabilities (e.g. writing non-degenerate code) appearing on held-out prompts, plus cross-architecture edits surviving on GPT-2 small where small weight perturbations usually collapse the model. This sits between two established lineages visible in the search results: conventional knowledge distillation (teacher-student training on large datasets, per the distillation overviews) and MEMIT-style closed-form MLP weight editing (which targets factual associations, not trajectory/capability transfer). The evidence is self-reported, first-party, and unreplicated; nothing in the supplied material establishes independent verification.

Why it matters to Scott

Not a convergence with or challenge to any position Scott's canon holds โ€” it's a new third-party technique claim โ€” but if replication holds it bears directly on him: training-free 4Bโ†’0.8B capability transfer would upgrade the cheap rung of his model barbell and offer a closed-form alternative to the corpus-plus-fine-tune upgrade path his synthetic fine-tuning data-factory work embodies, and since the claimant ran it on an 8GB RX 580 he could test it on his own gamepc/Ollama stack rather than just read about it. As an unreplicated first-party claim it remains a watch-and-try item, distinct from the radar's Hirundo and grafting weight-editing episodes in targeting capability transfer rather than censorship removal or architecture conversion.
ip:concept.model-barbelldev:concept.synthetic-finetuning-datasetdev:project.gamepcdev:technology.ollamaradar:concept.model-distillationradar:concept.model-adaptationradar:concept.small-modelsradar:concept.activation-steeringradar:hirundo-westernized-qwenradar:echoleak-immunized-llama-weightsradar:qwen-model-grafting-encoder-decoder
queries asked of Scott's wikis
  • small local models for coding agents and agent harnesses โ€” minimum viable model size
  • weight editing / model grafting / model surgery prior claims and positions
  • distillation vs fine-tuning vs editing โ€” cost of upgrading small open-weight models
  • consumer GPU local inference economics โ€” 8GB VRAM, AMD ROCm constraints
  • open-weights capability gap โ€” techniques for closing small-model deficits cheaply
  • transformer as dynamical system / residual stream trajectories โ€” interpretability views

Measured heat

now 0 pts/hpeak 11 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 410h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-24 14:00โญ origin echo-reconstructedREADME: "Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectur
dsadawq3 (GitHub; same project published on HuggingFace as org "F-Labs" and on Hacker News as "fermlon30000") on github (echo) ยท attributed from reddit.post.1wyx2w9
โ€”
10-06 08:15first on r/LocalLLaMA ยท published ยท +282.2hClosed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)
AdventurousTwo6445
โ€”
10-06 08:15amplified on r/LocalLLaMAreddit.post.1wyx2w9
AdventurousTwo6445
peak 3 ยท 10 comments ยท 44% of case engagement
10-06 09:41amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyycx6
AdventurousTwo6445
peak 14 ยท 3 comments ยท 56% of case engagement
10-06 08:20our radar first saw it ยท +282.3hdiscovery anchor: reddit.post.1wyx2w9โ€”
pace: p52 vs 1032 stories at the 336h mark (now 410h old) โ€” ahead of crowdstrike-safemind-security-agents (1.1x), behind agentdrive-persistent-shared-storage (0.9x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditClosed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)
LocalLLaMA
AdventurousTwo6445310
๐ŸŸง echo.github โญREADME: "Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architecturdsadawq3 (GitHub; same project published on HuggingFace as org "F-Labs" and on Hacker News as "fermlon30000")โ€”โ€”
๐ŸŸ  redditA 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)
LocalLLaMA
AdventurousTwo6445143

Interpretation history

Decision trace