2026-10-11 16:38 UTC

Tong Zheng and coauthors claim Dream-RSI uses historical discovery trees to cheaply refine exploration policies around an unchanged coding agent, reducing discovery costs while maintaining or improving results in algorithm, mathematical-optimization, and GPU-kernel tasks.

state: watchingheat: mediumuncertainty: highconvergesscott: highagent-harnesses recursive-self-improvement research-agents inference-economicsTong Zheng

What is this?

Dream-RSI is a research framework presented by Tong Zheng and coauthors affiliated with Google, Google DeepMind, the University of Maryland, and the University of Virginia; the supplied alphaXiv listing dates its submission to 14 September 2026. It leaves the underlying coding agent unchanged and instead refines exploration policies in a lightweight orchestration layer, using historical discovery trees as replay simulators before redeploying improved policies online. The project site and abstract claim competitive or better discovery quality at substantially lower cost across algorithm engineering, mathematical optimization, and GPU kernel engineering. The snippets provide neither quantitative results nor independent validation, so the cost and quality improvements remain author-reported claims.

Why it matters to Scott

Dream-RSI’s Google/DeepMind-affiliated authors independently converge with Scott’s Replay-Driven Design Evolution and Scaffolding Hypothesis: retained discovery history becomes a simulator for improving orchestration around an unchanged coding agent, offering a concrete mechanism to investigate for his replay-tuned dev-wiki navigation and a publishing comparison opportunity. The supplied radar hits track related harness improvement, not Dream-RSI itself; its cost and quality gains remain author-reported, and the hits do not establish dated priority for Scott’s position.
ip:framework.replay-driven-design-evolutionip:concept.scaffolding-hypothesisip:concept.model-plus-harness-benchmark-unitdev:concept.replay-tuned-graph-halodev:project.dev-wikiradar:evoharnessrl-self-evolving-agent-harnessradar:penguin-harness-recursive-improvementradar:harnessopt-agent-harness-optimization-benchmarkradar:concept.agent-harnessesradar:concept.recursive-self-improvement
queries asked of Scott's wikis
  • coding agent harness optimization unchanged base model
  • agent execution history replay offline policy evaluation
  • persistent agent memory accumulated experience improving exploration
  • recursive self-improvement orchestration versus model training
  • research agent evaluation costs search budgets discovery quality

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 674h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-13 14:00⭐ origin echo-reconstructedDream-RSI constructs replay simulators from historical discovery trees to evaluate and refine exploration policies with low-cost off-policy
Tong Zheng and coauthors on paper (echo) · attributed from hn.story.49726955
—
09-16 13:44first on hacker news · published · +71.7hDeepMind Paper: Dream-RSI: Recursive Self-Improvement Through Evolving Worlds
bananaflag
—
09-16 15:24first on r/LocalLLaMA · published · +73.4hDream-RSI: Recursive Self-Improvement through Evolving Worlds
No-Name-Person111
—
09-16 13:44amplified on hacker news 👑hn.story.49726955
bananaflag
peak 213 · 52 comments · 98% of case engagement
09-16 15:24amplified on r/LocalLLaMAreddit.post.1wi0c78
No-Name-Person111
peak 8 · 1 comments · 2% of case engagement
09-16 14:21our radar first saw it · +72.3hdiscovery anchor: hn.story.49726955—
pace: p78 vs 1032 stories at the 336h mark (now 674h old) — ahead of nvidia-cmp-vram-unlock (1.0x), behind experiential-open-model-gateway (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDeepMind Paper: Dream-RSI: Recursive Self-Improvement Through Evolving Worlds
Retrieved article excerpt

Open article · Retrieved 2026-09-16T14:23:03.360590+00:00

# Computer Science > Computation and Language

**arXiv:2609.14858** (cs)

[Submitted on 14 Sep 2026]

# Title:Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Authors:[Tong Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+T), [Xidong Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+X), [Zheng Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z), [Zhankui He](https://arxiv.org/search/cs?searchtype=author&query=He,+Z), [Chaoyi Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+C), [Benjamin Coleman](https://arxiv.org/search/cs?searchtype=author&query=Coleman,+B), [Ruoqiao Wei](https://arxiv.org/search/cs?searchtype=author&query=Wei,+R), [Di Bai](https://arxiv.org/search/cs?searchtype=author&query=Bai,+D), [Haolin Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+H), [Rui Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+R), [Xue Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+X), [Yue Zhuan](https://arxiv.org/search/cs?searchtype=author&query=Zhuan,+Y), [Wang-Cheng Kang](https://arxiv.org/search/cs?searchtype=author&query=Kang,+W), [Renkai Xiang](https://arxiv.org/search/cs?searchtype=author&query=Xiang,+R), [Heng Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+H), [Xinwu Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+X), [Yunsong Guo](https://arxiv.org/search/cs?searchtype=author&query=Guo,+Y)

View a PDF of the paper titled Dream-RSI: Recursive Self-Improvement through Evolving Worlds, by Tong Zheng and 16 other authors

[View PDF](https://arxiv.org/pdf/2609.14858)
[HTML (experimental)](https://arxiv.org/html/2609.14858v1)
> Abstract:Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, \textsc{Dream-RSI} secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, \textsc{Dream-RSI} achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.

|  |
| --- |
| Comments: |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2609.14858](https://arxiv.org/abs/2609.14858) [cs.CL] |
|  | (or  [arXiv:2609.14858v1](https://arxiv.org/abs/2609.14858v1) [cs.CL] for this version) |
|  | <https://doi.org/10.48550/arXiv.2609.14858> Focus to learn more  arXiv-issued DOI via DataCite (pending registration) |

## Submission history

From: Tong Zheng [[view email](https://arxiv.org/show-email/bdccc64e/2609.14858)]   
 **[v1]**
Mon, 14 Sep 2026 00:10:47 UTC (777 KB)

Full-text links:

## Access Paper:

View a PDF of the paper titled Dream-RSI: Recursive Self-Improvement through Evolving Worlds, by Tong Zheng and 16 other authors

- [View PDF](https://arxiv.org/pdf/2609.14858)
- [HTML (experimental)](https://arxiv.org/html/2609.14858v1)
- [TeX Source](https://arxiv.org/src/2609.14858)

[view license](http://arxiv.org/licenses/nonexclusive-distrib/1.0/ "Rights to this article")

### Current browse context:

cs.CL

[< prev](https://arxiv.org/prevnext?id=2609.14858&function=prev&context=cs.CL "previous in cs.CL (accesskey p)")
  |   
[next >](https://arxiv.org/prevnext?id=2609.14858&function=next&context=cs.CL "next in cs.CL (accesskey n)")

[new](https://arxiv.org/list/cs.CL/new)
 | 
[recent](https://arxiv.org/list/cs.CL/recent)
 | [2026-09](https://arxiv.org/list/cs.CL/2026-09)

Change to browse by:

[cs](https://arxiv.org/abs/2609.14858?context=cs)

### References & Citations

- [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2609.14858)
- [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2609.14858)
- [Semantic Scholar](https://api.semanticscholar.org/arXiv:2609.14858)

export BibTeX citation
Loading...

## BibTeX formatted citation

×

loading...

Data provided by:

### Bookmark

[BibSonomy](http://www.bibsonomy.org/BibtexHandler?requTask=upload&url=https://arxiv.org/abs/2609.14858&description=Dream-RSI: Recursive Self-Improvement through Evolving Worlds "Bookmark on BibSonomy")
[Reddit](https://reddit.com/submit?url=https://arxiv.org/abs/2609.14858&title=Dream-RSI: Recursive Self-Improvement through Evolving Worlds "Bookmark on Reddit")



Bibliographic Tools

# Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers Toggle

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps Toggle

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite.ai Toggle

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

Code, Data, Media

# Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

Links to Code Toggle

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com))*

DagsHub Toggle

DagsHub *([What is DagsHub?](https://dagshub.com/))*

GotitPub Toggle

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Huggingface Toggle

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast Toggle

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

Demos

# Demos

Replicate Toggle

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Spaces Toggle

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

Spaces Toggle

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai))*

Related Papers

# Recommenders and Search Tools

Link to Influence Flower

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

Core recommender toggle

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

- Author
- Venue
- Institution
- Topic


About arXivLabs

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.14858) |
Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
bananaflag21352
🟧 echo.paper ⭐Dream-RSI constructs replay simulators from historical discovery trees to evaluate and refine exploration policies with low-cost off-policy Tong Zheng and coauthors——
🟠 redditDream-RSI: Recursive Self-Improvement through Evolving Worlds
LocalLLaMA
No-Name-Person11181

Interpretation history

Decision trace