2026-10-11 16:38 UTC

Google Research claims regularized search over an open harness edit space โ€” annealed edit budgets, history-conditioned proposing, leakage screening, noise floors, and token-cost rules โ€” makes agent harnesses improve themselves with out-of-distribution gains (+6.0 Terminal-Bench 2.1, +1.8 SWE-bench Verified OOD, across three domains and two policy families), and independent replication or adoption would establish controlled recursive harness self-improvement as a working method.

state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-harnesses recursive-self-improvement agent-evaluationGoogle ResearchPeng Xia

What is this?

RRSI (Regularized Recursive Self-Improvement) is a 21 September 2026 paper from Google Cloud AI Research with UNC and Stanford โ€” first author Peng Xia, with Rujun Han, Zifeng Wang and others, code at github.com/google-research/rrsi โ€” claiming that an agent harness (prompts, tools, control flow, memory around a frozen LLM) can be improved by a search loop over its own source. Instead of constraining what the harness may contain, it regularizes the search itself โ€” annealed edit budgets, history-conditioned proposing, pre-evaluation leakage screening, noise-calibrated acceptance floors, cost-tied-to-gain acceptance โ€” which the authors credit for out-of-distribution gains (+6.0 Terminal-Bench 2.1 and +1.8 SWE-bench Verified with Claude Opus 4.8; +14.1 Terminal-Bench with Gemini 3.5 Flash; no held-out split regresses across three domains, and an excluded Gemini 3.1 Flash Lite also improves). Coverage remains paper-adjacent and partly critical: a skeptical StackSweep read notes the OOD gains are modest, 'OOD' here means a different benchmark in the same domain rather than a different kind of work, and absolute token cost exceeds a static harness (2.42M vs 1.56M per trial); secondary sources corroborate '30% fewer policy tokens than unregularized evolution' and none in the supplied coverage mentions a '58% less compute' figure, so that discrepancy stands uncorroborated. No independent replication or third-party adoption of the recipe appears anywhere in the supplied material; HN attention has decayed to near zero with zero comments, and the only new method-class activity is a one-author 'Proteus' Show HN pair with no content, results, or stated RRSI linkage.

Why it matters to Scott

Google Research's RRSI is a first-party arrival โ€” with paper, Apache-2.0 code, and measured OOD gains โ€” at where Scott's canon already sits: compounding capability lives in the owned harness around a frozen model (Scaffolding Hypothesis / Frozen Model Paradox), and its specific mechanisms (history-conditioned proposing, noise-calibrated acceptance floors, cost-tied-to-gain acceptance, pre-evaluation leakage screening) are directly adoptable in Synthetic Futures promotion gates. Nothing since the last diff changes the meaning โ€” Proteus is faint periphery that doesn't touch the replication condition, and the new skeptical coverage (modest same-domain 'OOD', higher absolute token cost than a static harness, still zero independent replication) actually sharpens the dated-receipts reading: Exemplar Ratchet and Future-Leakage Rule predict exactly the failure mode being flagged (transfer must be tested, the loop cannot certify itself), and the publishing window stays open until replication or refutation appears.
ip:concept.scaffolding-hypothesisip:concept.frozen-model-paradoxip:concept.self-improving-loopsip:framework.exemplar-ratchetip:concept.future-leakage-ruleip:framework.replay-driven-design-evolutiondev:project.synthetic-futuresradar:concept.agent-harnessesradar:concept.recursive-self-improvementradar:harnessopt-agent-harness-optimization-benchmarkradar:skillopt-agent-skill-training-loopradar:evoharnessrl-self-evolving-agent-harnessradar:autodesign-meta-harness-optimizationradar:trained-harness-cross-model-transferradar:dream-rsi-replay-exploration
queries asked of Scott's wikis
  • scaffolding hypothesis frozen model paradox โ€” where compounding learning lives in agent systems
  • exemplar ratchet โ€” reusing accumulated successful edits / edit-history conditioning in harnesses
  • future leakage rule โ€” benchmark contamination and pre-evaluation leakage screening
  • promotion gates โ€” noise floors, acceptance thresholds, evaluation variance in Synthetic Futures
  • recursive self-improvement / self-modifying agents โ€” prior essays or ebook territory
  • automated prompt or harness search โ€” dev projects doing search-over-scaffolding

Measured heat

now 0 pts/hpeak 14 pts/hcomments 0/hpeers p14momentum: steady1 platformsage 310h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-28 17:54โญ origin directly observedRRSI: Regularized Recursive Self-Improvement of Agent Harnesses
simonpure on hacker news
โ€”
10-01 23:03first on hacker news ยท published ยท +77.2hShow HN: Proteus โ€“ A framework for harnesses that evolve their own source
easonlu1017
โ€”
09-28 17:54amplified on hacker news ๐Ÿ‘‘hn.story.49881797
simonpure
peak 10 ยท 0 comments ยท 77% of case engagement
10-01 23:03amplified on hacker newshn.story.49928073
easonlu1017
peak 1 ยท 0 comments ยท 8% of case engagement
10-01 23:05amplified on hacker newshn.story.49928092
easonlu1017
peak 2 ยท 0 comments ยท 15% of case engagement
09-28 18:21our radar first saw it ยท +0.5hdiscovery anchor: hn.story.49881797โ€”
pace: p46 vs 1188 stories at the 168h mark (now 310h old) โ€” ahead of aipass-false-success-fixes (1.1x), behind agentgit-accountless-agent-handoffs (0.9x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hn โญRRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Retrieved article excerpt

Open article ยท Retrieved 2026-09-28T18:37:06.086716+00:00

# RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Check out our [paper](https://arxiv.org/abs/2609.24972) and [project page](https://regularized-rsi.com/) for more details.

## ๐Ÿ”ฅ Updates

- [09/21/2026] Our [paper](https://arxiv.org/abs/2609.24972) is out! Check it out! [[project page]](https://regularized-rsi.com/)

## ๐Ÿงฌ Overview

[RRSI overview: proposal-side and selection-side regularization of the harness search](https://github.com/google-research/rrsi/blob/main/assets/rrsi_overview.png)

An LLM agent's capability is largely set by its harness: the prompts, control
flow, tools, memory and context management around a frozen model. Evolving the
harness against a fixed evolve set is effective but overfits: the harness
memorizes the training tasks, and large in-distribution gains shrink or vanish
out of distribution. RRSI keeps the harness edit space open and regularizes the
search trajectory through it instead.

On the proposal side, an annealed budget caps how many independent edits one
candidate may bundle, the proposer is conditioned on the full edit history so a
falsified hypothesis is not redrawn, and a stalled run is redirected toward
components it has never exercised. On the selection side, a critic screens
every candidate for suite-specific logic before it is evaluated, a
noise-adjusted floor blocks gains within evaluation variance, a cost rule
requires added inference tokens to be paid for by measured gain, and
components that stop helping are pruned.

### โœจ Key features

- **Open edit space, regularized search.** Prompts, control flow, configuration, context management, tools, skills, memory and sub-agents may all be modified; the constraints act on how the search moves, not on what the harness may contain.
- **One method, three instances.** The same loop drives a terminal agent (Terminal-Bench 2.1), a document-work agent (Harvey LAB) and an engineering-design agent (EngDesign); each instance is a `Domain` adapter plus its starting harness.
- **Candidates in git worktrees.** Every candidate harness is drafted, screened and evaluated in its own worktree on a branch off `evolve/<domain>`; accepting one fast-forwards the branch, so the incumbent is always a commit.
- **Evidence you can audit.** The edit history records, per edit, the component, the hypothesis, the measured score and cost change and the verdict; the prompts the proposer, analyst and critic receive are plain files in `domains/<name>/`.

## ๐Ÿงฉ Method to code

| Paper | Code |
| --- | --- |
| Empirical score and cost estimate | `rrsi/evaluate.py: aggregate` (weighted per-trial rewards; a missing trial counts 0 with the full denominator) |
| Annealed edit budget b\_t | `rrsi/schedule.py: edit_budget`, enforced in the proposer's done() |
| Edit history L\_t, tried set T\_t, recent yield g\_t | `rrsi/history.py: History` (one JSONL record per edit; a = 1 only for the edits of the candidate that became H\_{t+1}) |
| Stall flag, untried components, exploration directives | `rrsi/history.py: stall_flag, exploration`; reserved slots enforced in `rrsi/propose.py` |
| Analyze(H\_t, D) | `rrsi/analyst.py` dispatching `rrsi/digester.py` |
| Proposer with (component, hypothesis, diff) tags | `rrsi/propose.py`; tags validated against the diff by `rrsi/components.py` |
| Critic (leakage screen before evaluation) | `rrsi/critic.py` (domain regex denylist plus LLM review, bounded repair) |
| Evaluate in parallel | `rrsi/evaluate.py`, `Run.round` thread pool |
| Noise-adjusted floor, cost rule, within-band rule, argmax | `rrsi/selection.py` |
| Novelty nu\_t (structural component types never in a winning edit) | `rrsi/components.py: novelty` over K\_str = client\_tool, skill, memory, subagent |
| Prune set B\_t | `History.prune_set`, handed to the proposer with the accepted machinery to remove |
| Noise band delta | fixed per instance in `rrsi.json` (0.017 / 0.004 / 0.020); `rrsi/calibrate.py` re-estimates it when `delta` is `null` (bootstrap over trials of the base evaluation, or repeated base evaluations) |
| Non-compensatory domain criteria | `Domain.guards` (engineering: valid-rate drop, no-submission rise) |



---

## โšก๏ธ Quickstart

### 0. Install

```
git clone https://github.com/google-research/rrsi.git && cd rrsi
pip install -e ".[dev]"            # the search core (Python 3.10 or newer)
python3 -m pytest tests
```

The benchmark runners live in their own environments: harbor for the coding instance (`domains/coding/.venv`), and a Python 3.11 environment with `pip install -e ".[agentic]"` for the workspace and engineering instances (`RRSI_AGENT_PYTHON`).

### 1. LLM configuration

The proposer, the analyst, the critic and the frozen policy are Claude Opus 4.8 on Vertex AI (`policy_model` in `domains/coding/rrsi.json`, `ORCHESTRATOR_MODEL` for the other two instances; any LiteLLM model string works). The Harvey LAB judge is Gemini 3.5 Flash.

```
gcloud auth application-default login
export VERTEX_PROJECT="your-project-id" VERTEXAI_PROJECT="your-project-id"
export VERTEX_LOCATION=global VERTEXAI_LOCATION=global
export RRSI_VERTEX_PROJECTS="your-project-id"
```

### 2. Run an instance

Every instance follows the same shape:

```
python3 rrsi.py --domain <coding|workspace|eng> smoke     # liveness: compile, construct, a couple of tasks
python3 rrsi.py --domain <name> baseline                   # Evaluate(H_0), seed runs/<name>/frontier.json
python3 rrsi.py --domain <name> run                        # rounds 0..T-1, resumable; touch runs/<name>/STOP to stop
python3 rrsi.py --domain <name> status
```

Each round drafts two candidates in their own git worktrees, screens them, evaluates both on the full evolve set and fast-forwards `evolve/<name>` to the winner. `runs/<name>/` holds the frontier, the edit history and the raw trials. Hyperparameters live in `domains/<name>/rrsi.json` and can be overridden on the command line (`--T`, `--k`, `--delta`, `--beta1`, ...); `readjudicate --t <t>` re-applies Algorithm 2 to a stored round and `reevaluate --t <t>` re-measures one after an infrastructure failure.

Please refer to the specific document for the instance you want to run for its environment, its evaluation protocol and the out-of-distribution runs:

- [`domains/coding`](https://github.com/google-research/rrsi/blob/main/domains/coding/README.md): Terminal-Bench 2.1, then SWE-bench Verified
- [`domains/workspace`](https://github.com/google-research/rrsi/blob/main/domains/workspace/README.md): Harvey LAB, then JobBench, GDPval and APEX-Agents
- [`domains/eng`](https://github.com/google-research/rrsi/blob/main/domains/eng/README.md): EngDesign, then EngDesign v1 and Frontier-Eng

The short version of each:

```
# coding: Docker + harbor
python3 -m venv domains/coding/.venv && domains/coding/.venv/bin/pip install "harbor>=0.18"
python3 rrsi.py --domain coding baseline && python3 rrsi.py --domain coding run
bash domains/coding/scripts/swe_eval.sh                       # H_0 and the incumbent on SWE-bench Verified

# workspace: a Harvey LAB checkout at the pinned commit; the split is generated from it on first use
git clone https://github.com/harveyai/harvey-labs.git && (cd harvey-labs && git checkout 1da4750 && uv sync)
export HARVEY_LAB_ROOT=$PWD/harvey-labs RRSI_AGENT_PYTHON=~/venvs/rrsi-agentic/bin/python
python3 rrsi.py --domain workspace baseline && python3 rrsi.py --domain workspace run
python3 rrsi.py --domain workspace heldout --label champ      # the 40 held-out tasks; ood/run_{jobbench,gdpval,apex}.sh for the rest

# eng: the official EngDesign tasks in the verifier layout, a grading venv, a jailed tool gateway
git clone https://github.com/AGI4Engineering/EngDesign.git
python3 domains/eng/scripts/engdesign/build_engdesign_bench.py --engdesign-open EngDesign/EngDesign-Open --out domains/eng/engdesign_bench
python3 -m venv domains/eng/.venvs/engdesign && domains/eng/.venvs/engdesign/bin/pip install -r domains/eng/scripts/engdesign/requirements.txt
bash domains/eng/scripts/preflight.sh
python3 rrsi.py --domain eng baseline && python3 rrsi.py --domain eng run
bash domains/eng/scripts/final_eval.sh frontier               # Frontier-Eng, from a Frontier-Engineering checkout
```

## ๐Ÿ“Š Results

Numbers from the paper, with Claude Opus 4.8 as the frozen policy in every instance and every number measured against the unevolved harness H\_0 in the same window. "Evolve" is the split the harness was searched on; the other rows never entered selection. Terminal-Bench, SWE-bench, JobBench, GDPval, APEX-Agents and EngDesign report pass rate, Harvey LAB the fraction of rubric criteria passed and Frontier-Eng Medal points.

| Domain | Benchmark | Role | H\_0 | RRSI | ฮ” |
| --- | --- | --- | --- | --- | --- |
| Coding | Terminal-Bench 2.1 | evolve | 74.2 | **80.2** | +6.0 |
| Coding | SWE-bench Verified | OOD | 82.0 | **83.8** | +1.8 |
| Agentic workspace | Harvey LAB | evolve | 89.4 | **90.5** | +1.1 |
| Agentic workspace | Harvey LAB | ID held-out | 86.9 | **89.2** | +2.3 |
| Agentic workspace | JobBench | OOD | 36.0 | **40.7** | +4.7 |
| Agentic workspace | GDPval | OOD | 48.8 | **52.3** | +3.5 |
| Agentic workspace | APEX-Agents | OOD | 34.2 | **37.9** | +3.7 |
| Engineering design | EngDesign | evolve | 50.0 | **54.9** | +4.9 |
| Engineering design | Frontier-Eng | OOD | 17.7 | **22.0** | +4.3 |

The search is not tied to one policy family: with Gemini 3.5 Flash as the frozen policy, the same coding instance goes from 64.6 to 78.7 on Terminal-Bench 2.1 and from 76.8 to 79.0 on SWE-bench Verified.

## ๐Ÿงฑ Adding a domain

A domain is one module, `domains/<name>/adapter.py`, exporting `DOMAIN`, an
instance of `rrsi.domain.Domain` that implements:

- `evolve_ids`, `heldout_ids`, `smoke_ids`: the task splits;
- `run(root, runs_dir, job, ids, k)` and `score(runs_dir, job, ids, k)`: run the harness checked out under `root` and return per-task trial rewards (Evaluate);
- `load_trial`, `render_trace`, `task_row`: the evidence the analyst, digester and proposer read;
- `smoke`: a liveness check of a candidate before it is evaluated;
- `critic_patterns`, `component_signals`, `briefs`, `guards`: the domain's leakage denylist, diff-to-component signals, role prompts and non-compensatory acceptance criteria;

plus `harness_path` (the evolvable directory), `SKILL.md` and `PATTERNS.md`
(the proposer's constitution) and `rrsi.json` (hyperparameters). The core
never reads a trajectory format or a benchmark directory itself.

## ๐Ÿงช Tests

```
python3 -m pytest tests            # or: python3 tests/test_core.py
```

## ๐Ÿ™ Acknowledgements

The starting harnesses are the Terminus-2 agent from [harbor](https://github.com/laude-institute/harbor) and the react\_toolbelt agent and runner from [archipelago](https://github.com/Mercor-Intelligence/archipelago). The instances evaluate on [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), [SWE-bench Verified](https://github.com/SWE-bench/SWE-bench), [Harvey LAB](https://github.com/harveyai/harvey-labs), [JobBench](https://github.com/Job-Bench/job-bench-eval), [GDPval](https://openai.com/index/gdpval/), [APEX-Agents](https://www.mercor.com/apex/apex-agents-leaderboard/), [EngDesign](https://github.com/AGI4Engineering/EngDesign) and [Frontier-Eng](https://github.com/Einsia/Frontier-Engineering).

## ๐Ÿ’ฌ Citation

```
@article{xia2026rrsi,
  title={RRSI: Regularized Recursive Self-Improvement of Agent Harnesses},
  author={Xia, Peng and Han, Rujun and Wang, Zifeng and Chen, Yanfei and Zhuang, Yufan and Lee, Yoonho and Huang, Chengsong and Yu, Han and CuiZhu, Zhongying and Ming, Yifei and Yao, Huaxiu and Gokturk, Burak and Pfister, Tomas and Lee, Chen-Yu},
  journal={arXiv preprint arXiv:2609.24972},
  year={2026}
}
```

## Contributing

See [CONTRIBUTING.md](https://github.com/google-research/rrsi/blob/main/CONTRIBUTING.md).

## License

Apache 2.0; see [LICENSE](https://github.com/google-research/rrsi/blob/main/LICENSE). Third-party code under `third_party/` carries its own li
simonpure100
๐ŸŸง hnShow HN: Proteus โ€“ Can DeepSeek Harness evolve itself to handle audio?easonlu101720
๐ŸŸง hnShow HN: Proteus โ€“ A framework for harnesses that evolve their own sourceeasonlu101710

Interpretation history

Decision trace