Hugging Face's post-training team (Lewis/lewtun) claims its published TRL + Harbor recipe makes multi-harness RL β training open models inside arbitrary coding-agent harnesses like Pi and its extensions β a standard, repeatable method; third-party teams adopting the recipe to train on their own harnesses confirm it, while the guide going uncited and unreplicated refutes it.
state: seedheat: lowuncertainty: mediumconvergesscott: highrl-training agent-harnesses open-model-traininglewtun (Lewis, Hugging Face post-training)
What is this?
Hugging Face's post-training team released 'The ultimate guide to multi-harness RL,' an open recipe pairing its TRL post-training library with a Harbor integration that makes a matrix of coding-agent harnesses (Claude Code, Pi, Codex, OpenCode, Cursor) trainable through one path β HF's own blog names Harbor as the restructuring that replaces per-harness hand-wiring. The premise, echoed in a Thomas Wolf repost, is that identical model weights score differently in every agent harness, so the harness is part of the training environment; the recipe uses a proxy that captures token IDs and logprobs between the model and OpenAI/Anthropic/Gemini APIs without modifying the harnesses. Reported results include OpenCode-only training lifting that harness from 34% to 58% while multi-harness training improved everywhere. The snippets confirm the release and its headline claims (and Pi itself as a real, extension-driven harness ecosystem), but show no third-party team yet training on their own harnesses β the adoption test is unresolved β and lewtun's specific authorship is asserted by the case rather than confirmed by the supplied results.
Why it matters to Scott
HF's post-training team independently arrives at β and operationalizes β the training-side corollary of Scott's model-plus-harness benchmark-unit position: if identical weights score differently in every harness, the harness is part of the environment, so train inside it, with the Harbor token proxy making the exact harnesses he A/B-benchmarks (OpenCode, Codex, Pi) into trainable targets. That is a dated-receipts moment for the Workshop ebook's multiplicative modelΓharness product, and the still-unresolved third-party adoption test gives him a concrete claim to track rather than just another harness-variance data point.
ip:concept.model-plus-harness-benchmark-unitip:source.give-the-agent-a-workshop-ebookdev:concept.trace-backed-agent-comparisonradar:harbor-token-proxy-agentic-rlradar:concept.agent-harnessesradar:concept.agentic-rlradar:frontierharness-17x-cost-variationradar:trained-harness-cross-model-transferradar:earendil-pi-1-0-release
queries asked of Scott's wikis
- harness sensitivity same weights different scores eval variance
- agent harness architecture agent loop tool schemas
- Pi harness minimal design extensions
- open-weight model RL post-training GRPO local training economics
- training coding agents on real task feedback traces
- reproducible open training recipe third-party adoption replication
Measured heat
now 0 pts/hpeak 7 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 266h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p65 vs 1188 stories at the 168h mark (now 266h old) β ahead of paddock-native-llm-runtime (1.0x), behind maccconc-kernel-race-testing (1.0x)
Evidence (2) β β canonical anchor
| source | object | author | score | comments |
| π reddit | The ultimate guide to multi-harness RL LocalLLaMA Retrieved article excerptOpen article Β· Retrieved 2026-10-03T12:24:58.800582+00:00 # Prove your humanity
Weβre committed to safety and security. But not for bots. Complete the challenge below and let us know youβre
a real person.
[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)
[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us) | lewtun | 82 | 12 |
| π§ echo.other β | The article itself (frontmatter): title "The ultimate guide to multi-harness RL", description: "The same model performs differently in every | Hugging Face post-training team (authors: Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall, Leandro von Werra); hosted on the FineEnvs HF Space org, source in github.com/adithya-s-k/FineEnvs | β | β |
Interpretation history
2026-10-04T15:32:39Z
Attention arc peaked and flattened inside four days: the velocity blip decayed to ~0 pts/h (peer percentile 17.5), leaving a well-received but non-spreading first-party writeup. Comments corroborate the harness-variance premise as practitioner folklore (one reader expressing intent to try in months) but the hypothesis's real test β third-party teams training on their own harnesses β has zero takers yet; this is now a cooled claim-owner watch, not a moving story.
2026-10-03T12:44:46Z
origin walked (opencode/cheap-glm, conf 0.95): anchor reddit.post.1wwk49n -> echo.other.5b84123ccc by Hugging Face post-training team (authors: Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall, Leandro von Werra); hosted on the FineEnvs HF Space org, source in github.com/adithya-s-k/FineEnvs
2026-10-03T12:42:24Z
grounded: converges/high β HF's post-training team independently arrives at β and operationalizes β the training-side corollary of Scott's model-plus-harness benchmark-unit position: if i
2026-10-03T12:35:51Z
case created β A first-party engineering writeup from HF's post-training team naming its own artifact (TRL + Harbor recipe) is a concrete claim-owner episode under the hot agent-harnesses topic, with adoption as the resolvable test.
Decision trace
- 10-05 02:32repriceAttention arc peaked and flattened inside four days: the velocity blip decayed to ~0 pts/h (peer percentile 17.5), leaving a well-received but non-spreading first-party writeup. Comments corroborate t
- 10-04 08:21sensor_dirtycomment_update
- 10-04 03:20sensor_dirtyvelocity_spike
- 10-04 01:21sensor_dirtycomment_update
- 10-03 22:44promote_anchororigin walk conf 0.95
- 10-03 22:42groundHF's post-training team independently arrives at β and operationalizes β the training-side corollary of Scott's model-plus-harness benchmark-unit position: if identical weights score differe
- 10-03 22:35createA first-party engineering writeup from HF's post-training team naming its own artifact (TRL + Harbor recipe) is a concrete claim-owner episode under the hot agent-harnesses topic, with adoption a