Microsoft's SkillOpt claims a training loop for agent skills — running a frozen agent on scored batches, having an optimizer model propose structured add/delete/replace edits, and accepting candidates only when held-out validation improves — establishing automatically optimized skill libraries as a method beyond hand-maintained prompts.
state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-skills agent-harnesses agent-memoryMicrosoft
What is this?
SkillOpt is a Microsoft Research method and open-source release (MIT-licensed repo plus arXiv paper, May 2026) that treats a natural-language skill document — a plain markdown file — as the 'trainable parameter' of a frozen LLM agent. A separate optimizer model runs a deep-learning-style loop in text space: the frozen agent executes scored task batches (rollout), the optimizer reflects over success and failure trajectories and proposes bounded add/delete/replace edits to the skill, candidate edits are clipped by a per-step 'textual learning rate' budget, and a candidate skill is accepted only if it strictly improves a held-out validation score, with rejected edits buffered as negative feedback. Microsoft reports best or tied-best results across 6 benchmarks × 7 models (52 evaluation cells), often 20–40 points over no-skill baselines, and frames it as the first systematic controllable text-space optimizer for agent skills. The mechanism is consistently described across the MSR blog, repo, paper, and third-party writeups; the performance numbers come from Microsoft's own evaluation and have not been independently verified in the supplied material.
Why it matters to Scott
Microsoft Research has independently formalized the position Scott's canon argues: a markdown skill treated as 'the trainable parameter' is soft weights / the third substrate turned into an official training method, the frozen-agent-plus-optimized-scaffolding design is the scaffolding hypothesis in the wild, and strict held-out acceptance mechanizes evaluation-driven development and the Generate→Evaluate→Extract→Encode loop — a first-party dated receipt for the Third Substrate and Self-Equipping Agent theses (numbers still Microsoft's own, unverified). Its bounded add/delete/replace edits with a textual learning rate also speak directly to the skill-bloat episode and offer a validation-gated alternative to Backpass's human-gated memory edits — a mechanism Scott could adopt or critique in his own skill and harness maintenance.
ip:concept.three-substratesip:source.the-third-substrate-ebookip:concept.soft-weightsip:concept.scaffolding-hypothesisip:concept.self-improving-loopsip:concept.evaluation-driven-developmentip:framework.worldview-recursive-compressionradar:concept.agent-skillsradar:backpass-evidence-gated-memory-editsradar:agent-skill-bloat-gradingradar:dspy-factorio-rlm-gepa-agentsradar:harnessopt-agent-harness-optimization-benchmarkradar:trained-harness-cross-model-transfer
queries asked of Scott's wikis
- agent skill file SKILL.md skill library harness
- agent-maintained wiki memory consolidation
- validation-gated evals for agent prompt changes
- automatic prompt optimization DSPy position
- skills as external state vs fine-tuning local models
- self-improving agent experience distillation
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p37momentum: steady2 platformsage 403h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p50 vs 1032 stories at the 336h mark (now 403h old) — ahead of aipass-false-success-fixes (1.1x), behind android-editable-graph-agent (0.9x)
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-10-10T22:25:41Z
New Show HN post (eval-skills) adds another project in the automated skill-optimization space but with negligible engagement (1 point, 0 comments) — no independent verification of SkillOpt's claims, no replications, no framework adoption signals. The case remains a first-party Microsoft artifact worth watching for independent evaluation, not for chatter.
2026-10-10T21:38:14Z
evidence attached: hn.story.50035864 — Show HN release of eval-skills for automatic agent improvement independently corroborates the automated skill-optimization trend.
2026-09-26T01:41:15Z
The debut thread died quietly (velocity ~0, 20th percentile) and its only substantive addition is a critique that Microsoft's headline charts lack units/baselines, sharpening the single-source caveat on the claimed gains; the artifact itself is established and first-party, so the case moves from discovery to watching for independent verification and adoption.
2026-09-24T23:50:55Z
grounded: converges/high — Microsoft Research has independently formalized the position Scott's canon argues: a markdown skill treated as 'the trainable parameter' is soft weights / the t
2026-09-24T23:43:27Z
case created — A first-party Microsoft artifact treating skills as trainable, optimizable objects is a concrete new method in a hot harness/memory area, distinct from the open skill-bloat and backpass claims.
Decision trace
- 10-11 10:03attention_routeThe editor compared this story and chose to keep watching.
- 10-11 09:25repriceNew Show HN post (eval-skills) adds another project in the automated skill-optimization space but with negligible engagement (1 point, 0 comments) — no independent verification of SkillOpt's clai
- 10-11 08:42attention_routeFirst-party dated receipt for Scott's third-substrate and self-equipping-agent theses; validation-gated skill editing directly relevant to his harness maintenance. New case, high relevance, brief
- 10-11 08:38attention_candidateattach
- 10-11 08:38attachShow HN release of eval-skills for automatic agent improvement independently corroborates the automated skill-optimization trend.
- 10-11 08:38propose_attachShow HN release of eval-skills for automatic agent improvement independently corroborates the automated skill-optimization trend.
- 10-06 21:42review_dormantscheduled targets exhausted or 28 quiet days
- 10-06 21:42drop_targetsquiet through full ladder or over cap 8
- 09-26 11:41repriceThe debut thread died quietly (velocity ~0, 20th percentile) and its only substantive addition is a critique that Microsoft's headline charts lack units/baselines, sharpening the single-source ca
- 09-26 11:40review_screenNew HN thread includes a first-party GitHub release of Microsoft's SkillOpt (public code availability), plus a substantive critique that the headline results lack units/baselines, which bears on
- 09-26 11:40review_screenjev screen borderline (noul=0.39) — luna review
- 09-25 09:50groundMicrosoft Research has independently formalized the position Scott's canon argues: a markdown skill treated as 'the trainable parameter' is soft weights / the third substrate turned int
- 09-25 09:43createA first-party Microsoft artifact treating skills as trainable, optimizable objects is a concrete new method in a hot harness/memory area, distinct from the open skill-bloat and backpass claims.