2026-10-11 16:37 UTC

Martin Bertran Lopez and Aaron Roth claim successful ML research-agent strategies retain performance when compressed to as few as 16 tokens while overfit gains disappear, making compression a practical diagnostic for benchmark generalization.

state: seedheat: lowuncertainty: highconvergesscott: mediumresearch-agents evaluation overfittingMartin Bertran LopezAaron RothAmazon Science

What is this?

An Amazon Science article describes experiments in which successful ML research-agent strategies were compressed into short prompts and handed to a fresh reproducer agent. Across eight datasets covering five task areas, it reports that 32-token prompts matched the explorer's adaptively optimized models on the large majority of problems; one language-modeling strategy preserved held-out performance with just 16 tokens. The case attributes the work to Martin Bertran Lopez and Aaron Roth, but the supplied article snippet does not confirm authorship. The snippet supports strategy compressibility as a proposed explanation for generalization, but does not establish that overfit gains disappear under compression or that compression is a validated practical diagnostic.

Why it matters to Scott

The Amazon Science account supplies an empirical convergence with Scott’s Kernel Doctrine and long-running-agents architecture: reusable learning compressed upstream can regenerate successful work in a fresh agent, suggesting a concrete compression-and-reproduction test for his agent-authored checkpoints and scraper playbooks. The reported results extend that position into ML research strategies, but do not establish compression as an overfitting diagnostic; related radar pages track compression fidelity and benchmark integrity, not this development.
ip:concept.kernel-doctrineip:framework.long-running-agentsdev:concept.agent-authored-context-compactiondev:concept.scraper-playbook-memoryradar:distil-decision-equivalent-context-compressionradar:harnessopt-agent-harness-optimization-benchmarkradar:concept.research-agents
queries asked of Scott's wikis
  • agent evaluation benchmark overfitting held-out validation
  • agent memory compression strategy transfer fresh context
  • research agent harness reproducibility experiment design
  • compressed knowledge generalization reusable recipes
  • adaptive benchmark optimization evaluation leakage

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 647h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 17:31 (minted)⭐ origin echo-reconstructedThe authors' account of 'What fits (into few tokens) doesn't overfit' reports that a fresh agent can recover successful strategies through a
Martin Bertran Lopez, Aaron Roth, and coauthors on paper (echo) · attributed from hn.story.49699648 · published time unknown
—
09-14 16:32first on hacker news · published · lag ?Why don't machine learning research agents overfit?
Betelbuddy
—
09-14 16:32amplified on hacker news 👑hn.story.49699648
Betelbuddy
peak 136 · 94 comments · 100% of case engagement
09-14 17:22our radar first saw it · lag ?discovery anchor: hn.story.49699648—
pace: p76 vs 1032 stories at the 336h mark (now 647h old) — ahead of minimax-code-open-source (1.0x), behind gguf-quant-filename-mismatch (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWhy don't machine learning research agents overfit?
Retrieved article excerpt

Open article · Retrieved 2026-09-14T17:24:43.869404+00:00

CompressionModels-03-16x9.png

The more your listener already knows, the shorter the message you need to send. An expert ML engineer needs only a few sentences; a newcomer needs the whole manual.

[Machine learning](https://www.amazon.science/research-areas/machine-learning)

# Why don’t machine learning research agents overfit?

## New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.

By [Martin Bertran Lopez](https://www.amazon.science/author/martin-bertran-lopez), [Aaron Roth](https://www.amazon.science/author/aaron-roth)

September 10, 2026

11 min read

Share

Share

- Copy link
- Email
- [X](https://twitter.com/intent/tweet?url=https%3A%2F%2Fwww.amazon.science%2Fblog%2Fwhy-dont-machine-learning-research-agents-overfit&text=Why%20don%E2%80%99t%20machine%20learning%20research%20agents%20overfit%3F)
- [LinkedIn](https://www.linkedin.com/shareArticle?url=https%3A%2F%2Fwww.amazon.science%2Fblog%2Fwhy-dont-machine-learning-research-agents-overfit&mini=true&title=Why%20don%E2%80%99t%20machine%20learning%20research%20agents%20overfit%3F&summary=New%20research%20indicates%20that%20AI%20agents%20learn%20compressible%20models%20of%20data%2C%20which%20don%E2%80%99t%20have%20enough%20space%20to%20enable%20memorization.&source=Amazon%20Science)
- [Facebook](https://www.facebook.com/dialog/share?app_id=1024652704536162&display=popup&href=https%3A%2F%2Fwww.amazon.science%2Fblog%2Fwhy-dont-machine-learning-research-agents-overfit)
- [Line](https://social-plugins.line.me/lineit/share?url=https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit)
- [Reddit](https://www.reddit.com/submit?url=https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit&title=Why don’t machine learning research agents overfit?)
- [QZone](http://sns.qzone.qq.com/cgi-bin/qzshare/cgi_qzshare_onekey?url=https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit&title=Why don’t machine learning research agents overfit?&summary=New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.)
- [Sina Weibo](https://service.weibo.com/share/share.php?url=https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit)
- [WeChat](https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit "Share on wechat")
- [WhatsApp](https://api.whatsapp.com/send?text=Why don’t machine learning research agents overfit?%20https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit)

分享到微信

x

Key takeaways

- ML models don't overfit benchmarks, even after many rounds of iterative improvement. This contradicts textbook predictions that repeatedly evaluating against the same held-out data should lead to overfitting.
- Experiments with ML research agents indicate that successful strategies are highly compressible. When a successful agent's strategy is squeezed through an information bottleneck (as few as 16 tokens), a fresh agent with no memory can reproduce the original agent's performance, meaning the strategy captured real structure, not memorized data.
- Compression provides both an explanation and a diagnostic tool. Strategies that genuinely overfit fail the compression test: their validation-specific gains vanish when passed through the bottleneck.
- LLMs are powerful compression decoders. Because they carry vast world knowledge, they can reconstruct full ML pipelines from terse, expert-shorthand prompts, which is a concrete way of understanding why they're so capable.

Was this answer helpful?

Machine learning, at its core, is about generalization, not memorization. You hand your learning algorithm a pile of training examples and use them to fit a model. But the goal is not to perform well on the training examples — that's easy, you could just memorize the answers. The goal is to perform well on *new* examples that you have never before seen. If a model does well on the data it was trained on but poorly on fresh data, it hasn’t actually learned anything; you have only fooled yourself into thinking it has. This failure mode has a name: overfitting.

Anyone who has taken an introductory statistics or machine learning class knows the standard defense. You hold out some of your data and refuse to train on it. In practice, this held-out data plays two roles. A *validation set* is one you consult repeatedly while building the model — to compare candidates, tune hyperparameters, and decide what to try next. A final *test set* (or *holdout*) is meant to be touched only once, at the very end: because the training procedure never saw it, strong performance there is a correct proxy for the new examples you will encounter in the wild.

> Machine learning, at its core, is about generalization, not memorization

The “holdout” condition is crucial, though. The correct-proxy guarantee holds if the held-out set stays genuinely unseen. If you check your performance on it, tweak your training procedure in response, recheck, and iterate, chasing better and better numbers, that set is no longer unseen; it has become part of your training procedure. Do this enough times, and you can overfit it just as you might have overfit the training set, and you have lost your proxy for unseen data. This is true of any held-out set you reuse this way, including a validation set, which is reused by design.

## A puzzle at the heart of machine learning

Real machine learning research looks *exactly* like the iterative improvement loop we just described. Everyone gauges performance using a handful of benchmark datasets that go unrevised for years. The research community repeats an enormous, distributed loop: evaluate a model on the benchmark, revise the training procedure, re-evaluate, publish, and let the next group eke out a little more improvement.

This is precisely the kind of hill-climbing against a held-out set that, by the textbook account, ought to produce rampant overfitting. By now, the leaderboards should be saturated with models that look great on the benchmark and mediocre everywhere else.

And yet that is not what happens. Studies that build entirely fresh test sets for old, heavily reused benchmarks have found that improvements largely *transfer*: on the new data, models demonstrate the same gains they did on the old benchmark. Benchmark-driven machine learning, against the textbook's prediction, has produced rapid and largely *real* progress. Why?

There is no shortage of hypotheses, but they have been hard to test empirically, because the "subject" of the experiment is the entire human research community. You cannot reset a field, wipe its memory, and rerun the last decade under controlled conditions.

But we can do something similar. We now have capable, LLM-based research agents that can autonomously run the same machine-learning optimization loops that human communities run. They engage in the same benchmark hill-climbing — and, intriguingly, they too seem not to overfit. The difference is that an agent, unlike a research community, is something you *can* reset. You can clear its memory, control exactly what information it sees, and run the experiment again. In a recent paper, "[What fits (into few tokens) doesn't overfit: Compression and generalization in ML research agents](https://arxiv.org/abs/2606.11045)", we do exactly that — and in the process offer a concrete explanation for the long-standing mystery.

## Occam's razor, made precise

The explanation begins with a very old idea. Occam's razor says that among hypotheses that explain the data equally well, the simpler one is more likely to be correct. It turns out this intuition has a precise mathematical form, and it is what underlies the whole story.

Suppose you can describe your hypothesis — your model, your strategy — in a small number of bits, far fewer than it would take to memorize the training data. If that compact hypothesis performs very well on the training data, it must also perform well on new data.

[

](https://cdn.amazon.science/72/eb/d021ebfc445d8f4acd2fe2ac211c/compressionmodels-01-1x1.mp4)

Occam's razor, formalized: among hypotheses that explain the data equally well, the simpler one — describable in fewer bits — is more likely to generalize to new examples.

The reasoning runs through a counting argument. There simply are not very many *short* descriptions, because there are not very many short strings. The fewer candidate hypotheses there are, the less likely it is that any one of them fooled you on the training set by luck — even though you used the training set to guide your search.

Another way to get the intuition: if your compressed description is too small to secretly record the training data, then when it performs well on the training data, it cannot be because it memorized the answers — it didn't have space to do that. It must be because it captured something true about the data's structure. Short descriptions cannot cheat because there isn't room.

Here is an attractive hypothesis: *successful machine learning strategies are highly compressible.* A researcher might stare at thousands of benchmark scores over the course of a project, but the strategy that ultimately survives is usually a short list of familiar choices — an architecture family, an optimizer, a learning-rate schedule, a data-handling recipe, a regularization scheme. If that final recipe can be communicated in just a few bits, then the model's true dependence on the benchmark is far smaller than the long, winding transcript of experiments would suggest. The hill-climbing was extensive, but the thing that came out the other end was — or could have been — tiny.

## Compression, intelligence, and the power of a knowledgeable listener

Imagine trying to explain a specific machine learning pipeline to a bright high-school student, in enough detail that they could actually reproduce it. It would be a long, laborious conversation. You would have to explain what gradient descent is, what a neural network is, what PyTorch or JAX or TensorFlow does, what a learning rate is, and on and on. Almost none of that is specific to *your* problem; it is general background about how machine learning works.

Now imagine explaining the same pipeline to an expert ML engineer. The conversation now collapses to a few sentences. You skip everything that counts as common knowledge and communicate only what is genuinely specific to *this* problem: the architecture choice, the batch size, the optimizer, a couple of hyperparameters. The more your listener already knows about the world, the shorter the message you need to send — and the more aggressively you can compress. None of this "world knowledge" counts against you in the Occam's-razor argument, because you could have written all of that down without having looked at the training set.

This is where large language models enter the picture. Modern LLMs carry an enormous amount of world knowledge. They know how ML tooling works; they know the standard optimization algorithms; they know the conventional hyperparameter choices and the common defaults. If a detail is left unspecified, they can fill in a plausible value. That makes them extraordinarily good *compression decoders*: hand an LLM a terse, expert-to-expert message, and it can unpack it into a full, working procedure. If you think about it, this is exactly why they are so powerful.

## The experiment: Squeezing a strategy through a bottleneck

This suggests a clean experiment. Have an ML research agent — *the explorer* — try to solve a new machine learning problem. Give it full access to a validation set and let it experiment and iterate freely, chasing better validation performance over hundreds of rounds. Here the validation set plays the role of the benchmark: a *reusable holdout* the agent queries again and again. This is the hill-climbing loo
Betelbuddy13694
🟧 echo.paper ⭐The authors' account of 'What fits (into few tokens) doesn't overfit' reports that a fresh agent can recover successful strategies through aMartin Bertran Lopez, Aaron Roth, and coauthors——

Interpretation history

Decision trace