2026-10-11 17:09 UTC

Microsoft's SkillOpt claims a training loop for agent skills — running a frozen agent on scored batches, having an optimizer model propose structured add/delete/replace edits, and accepting candidates only when held-out validation improves — establishing automatically optimized skill libraries as a method beyond hand-maintained prompts.

state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-skills agent-harnesses agent-memoryMicrosoft

What is this?

SkillOpt is a Microsoft Research method and open-source release (MIT-licensed repo plus arXiv paper, May 2026) that treats a natural-language skill document — a plain markdown file — as the 'trainable parameter' of a frozen LLM agent. A separate optimizer model runs a deep-learning-style loop in text space: the frozen agent executes scored task batches (rollout), the optimizer reflects over success and failure trajectories and proposes bounded add/delete/replace edits to the skill, candidate edits are clipped by a per-step 'textual learning rate' budget, and a candidate skill is accepted only if it strictly improves a held-out validation score, with rejected edits buffered as negative feedback. Microsoft reports best or tied-best results across 6 benchmarks × 7 models (52 evaluation cells), often 20–40 points over no-skill baselines, and frames it as the first systematic controllable text-space optimizer for agent skills. The mechanism is consistently described across the MSR blog, repo, paper, and third-party writeups; the performance numbers come from Microsoft's own evaluation and have not been independently verified in the supplied material.

Why it matters to Scott

Microsoft Research has independently formalized the position Scott's canon argues: a markdown skill treated as 'the trainable parameter' is soft weights / the third substrate turned into an official training method, the frozen-agent-plus-optimized-scaffolding design is the scaffolding hypothesis in the wild, and strict held-out acceptance mechanizes evaluation-driven development and the Generate→Evaluate→Extract→Encode loop — a first-party dated receipt for the Third Substrate and Self-Equipping Agent theses (numbers still Microsoft's own, unverified). Its bounded add/delete/replace edits with a textual learning rate also speak directly to the skill-bloat episode and offer a validation-gated alternative to Backpass's human-gated memory edits — a mechanism Scott could adopt or critique in his own skill and harness maintenance.
ip:concept.three-substratesip:source.the-third-substrate-ebookip:concept.soft-weightsip:concept.scaffolding-hypothesisip:concept.self-improving-loopsip:concept.evaluation-driven-developmentip:framework.worldview-recursive-compressionradar:concept.agent-skillsradar:backpass-evidence-gated-memory-editsradar:agent-skill-bloat-gradingradar:dspy-factorio-rlm-gepa-agentsradar:harnessopt-agent-harness-optimization-benchmarkradar:trained-harness-cross-model-transfer
queries asked of Scott's wikis
  • agent skill file SKILL.md skill library harness
  • agent-maintained wiki memory consolidation
  • validation-gated evals for agent prompt changes
  • automatic prompt optimization DSPy position
  • skills as external state vs fine-tuning local models
  • self-improving agent experience distillation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p37momentum: steady2 platformsage 403h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-24 23:43 (minted)⭐ origin echo-reconstructed"A skill is external state for an agent... SkillOpt runs the frozen agent on scored batches, asks an optimizer model to propose structured e
Microsoft (SkillOpt) on github (echo) · attributed from hn.story.49836602 · published time unknown
—
09-24 20:55first on hacker news · published · lag ?SkillOpt: Training Loop for Agent Skills
DaveFr
—
09-24 20:55amplified on hacker news 👑hn.story.49836602
DaveFr
peak 11 · 3 comments · 94% of case engagement
10-10 18:46amplified on hacker newshn.story.50035864
jeffreyip
peak 1 · 0 comments · 7% of case engagement
09-24 23:21our radar first saw it · lag ?discovery anchor: hn.story.49836602—
pace: p50 vs 1032 stories at the 336h mark (now 403h old) — ahead of aipass-false-success-fixes (1.1x), behind android-editable-graph-agent (0.9x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSkillOpt: Training Loop for Agent Skills
Retrieved article excerpt

Open article · Retrieved 2026-09-24T23:34:00.438878+00:00

### A skill is external state for an agent.

Instead of fine-tuning a model or hand-maintaining prompts, SkillOpt runs
the frozen agent on scored batches, asks an optimizer model to
propose structured edits, and accepts a candidate only when validation
performance improves.

Frozen target model
Optimizer model
Add / delete / replace edits
Held-out gate
DaveFr113
🟧 echo.github ⭐"A skill is external state for an agent... SkillOpt runs the frozen agent on scored batches, asks an optimizer model to propose structured eMicrosoft (SkillOpt)——
🟧 hnShow HN: Eval-skills for automatic agent improvementjeffreyip10

Interpretation history

Decision trace