2026-10-11 16:38 UTC

Evolutionary LLM program search improves upper bounds for square packing, breaking 25 records including a 47-year-old record, demonstrating LLMs as optimization tools for combinatorial problems.

state: watchingheat: mediumuncertainty: mediumconvergesscott: highai-assisted-mathematics llm-optimization combinatorial-searchryan (slightlysupervised)

What is this?

Ryan (slightlysupervised) published a Substack post demonstrating evolutionary LLM program search for square packing, breaking 25 upper-bound records including Walter Trump's 47-year-old 1979 result for 11 unit squares. The method uses LLM-driven evolutionary program search similar to Google DeepMind's AlphaEvolve, seeded with hand-coded primitives and lightly guided over 2 days, 7.5k CPU-hours, and ~$125 in token costs. A GitHub repository with reproducible results is referenced. Separately, OpenAI's Astra and Anthropic's Claude were used to formalize Trump's 1979 packing as optimal in Lean.

Why it matters to Scott

Ryan's evolutionary LLM program search for square packing is a concrete, reproducible demonstration of the AlphaEvolve-style search pattern Scott has built frameworks around (Agent Loop, Agentic Exploration, Iterative Deepening, Replay-Driven Design Evolution, Search-Space Collapse, Bounded Cognitive-Worker Ladder). The $125/7.5k CPU-hour token economics directly instantiate Scott's Token Economics, Cost of Cognition, Model Dividend, and Compound Returns positions. The 2-day long-running search with GitHub reproducibility and Lean formalization mirrors Scott's Long-Running Agents, Durable External State, Checkpoint Discipline, and Verification Loops/Proof-Carrying Transformation requirements. Scott's own AMA project (Dialectical Tree Search) applies chess-engine search to LLM reasoning — the same architectural pattern now validated on a 47-year-old combinatorial record.
ip:framework.agent-loopip:concept.agentic-explorationip:concept.iterative-deepeningip:framework.replay-driven-design-evolutionip:concept.search-space-collapseip:concept.token-economicsip:concept.cost-of-cognitionip:concept.model-dividendip:concept.compound-returnsip:framework.long-running-agentsip:concept.durable-external-stateip:concept.checkpoint-disciplineip:concept.agent-mortalityip:concept.verification-loopsip:framework.proof-carrying-transformationip:concept.spec-driven-developmentip:concept.spec-as-assetip:concept.delete-testip:concept.deterministic-verification-before-assertionip:dev:project.amaip:dev:concept.dialectical-tree-searchip:dev:concept.bounded-cognitive-worker-ladderip:dev:concept.progressive-resolution-orchestrationradar:concept.ai-assisted-mathematicsradar:alphaevolve-matrix-exponent-improvementradar:artificium-covering-design-searchradar:dots-swarm-covering-recordradar:concept.automated-theorem-provingradar:concept.formal-verificationradar:concept.lean-formalizationradar:concept.agentic-researchradar:concept.agent-memoryradar:concept.long-running-agentsradar:concept.persistent-agentsradar:concept.inference-economicsradar:concept.token-economicsradar:concept.open-modelsradar:concept.program-synthesis
queries asked of Scott's wikis
  • evolutionary program search LLM alpha-evolve pattern
  • llm-driven combinatorial optimization agent loop
  • reproducible ai-assisted mathematics workflow
  • token compute economics evolutionary search
  • open-weight models program synthesis search
  • agent memory evolutionary search continuation

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p16momentum: steady1 platformsage 46h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-09 18:08⭐ origin directly observedImproving 25 square packing upper bounds through evolutionary LLM program search
ryanxu on hacker news
—
10-09 18:08amplified on hacker news 👑hn.story.50024473
ryanxu
peak 2 · 0 comments · 98% of case engagement
10-09 19:34our radar first saw it · +1.4hdiscovery anchor: hn.story.50024473—
pace: p33 vs 968 stories at the 24h mark (now 46h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentlane-git-native-coordination (0.7x)

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Improving 25 square packing upper bounds through evolutionary LLM program search
Retrieved article excerpt

Open article · Retrieved 2026-10-09T19:59:01.348083+00:00

# Improving 25 square packing upper bounds through evolutionary LLM program search

[ryan's avatar](https://substack.com/@slightlysupervised)

[ryan](https://substack.com/@slightlysupervised)

Oct 09, 2026

1

Share

#### What is the smallest square that encloses n unit squares?

For some values of n, the answer is clear. 16 squares pack into a 4x4.

In contrast, here is the beautiful, [recently proven to be optimal](https://github.com/Queuingtheorydotcom/11SquaresFormalized), answer for n=11.

s = 3.87708

Records for different values of n were historically [maintained](https://web.archive.org/web/20230530194618/https://erich-friedman.github.io/packing/squinsqu/) by Erich Friedman, and [more recently](https://kingbird.myphotos.cc/packing/squares_in_squares.html) by David Ellsworth. Driven by a flurry of recent excitement around square packing, records are now also nicely [maintained](https://jlevy.github.io/squares/) by Joshua Levy.

#### How did we break these records?

TL;DR: We achieve [these results](https://github.com/ry-xu/square_packing) through LLM-driven evolutionary program search, similar to [AlphaEvolve](https://arxiv.org/abs/2506.13131). The programs were seeded with primitives vibe-coded by yours truly, and the search was lightly guided between meetings and insufficient sleep.

---

### Motivating program search

Following recent state-of-the-art approaches to solving open problems, I began by asking claude to break a record and then let it crunch away for an hour to no success.

Seared into my brain from years of ML is to **always look at the data**. Looking at some of the resulting packings, it became apparent that the style of search was inefficient, often landing in uninteresting minima. To build some intuition, I vibe-coded a web app and decided to try for the records myself.

Over the course of many hours and many packed squares, I requested various features that I thought may help the algorithm to break out of degenerate solutions, such as:

- being able to round the corners of the squares
- a control for shaking intensity
- “wind” - restricting the shaking direction
- “fade” - shaking less as the box shrinks

This packing lab is memorialized [here](https://ryanxu.net/packing_lab/).

While packing, I often found myself repeating certain actions that consistently led to interesting packings (e.g. rounding the squares, then letting them sharpen over and over again). Eventually it occurred to me that I was simply executing a program on top of these vibe-coded toggles, and that an LLM could search through the space of programs to rediscover mine, and hopefully many better.

The initial evolutionary search was implemented and within an hour, the record for n=51 had been broken on my little MacBook Air without a fan. Soon after, n=103 and n=105 had also fallen.

The next morning, as I ran around sharing my packings with the office, I was kindly gifted a *bit* more compute by [the boss](https://x.com/Lifrordi) and encouraged to take the day to scale up the approach.

---

### The run

The run began with a seeded baseline program. This program places n squares at random in a box, then slowly shrinks the box. Whenever squares overlap, an optimizer (L-BFGS) finds a local arrangement where they no longer overlap. When this converges, we have a packing! Alongside this code was a set of unused functions developed for the packing lab for future programs to use.

Every generation, we tasked a fleet of claude haiku agents with generating mutations of the previous generation’s best programs. Best is decided by a fitness score—the distance from the best known packing for a fixed sample of n values.

We ran this evolutionary loop for 128 generations.

In the end, a total of 25 records were broken over the course of 2 days, using 7.5k CPU-hours and $125 of tokens. The oldest record broken had stood for 47 years.

The run

Although the records were broken in one continuous evolutionary search, a few modifications were made throughout:

1. Scaling up from 16 candidates per generation to 64
2. Sampling instructions from a set of roles for candidate generation, since the diversity of evolved programs appeared low

   1. default: try to improve the program
   2. crossover: merge two candidates
   3. invent: create and use a new primitive
   4. moonshot: make a high risk high reward change
3. Starting to include n values > 200
4. Increasing max runtime 10s → 60s, as I noticed that large n were timing out
5. Upweighting unbeaten n in the fitness score

In many cases, we broke our own records multiple times!

### The evolved programs

In the first half of the search, evolved programs kept the initial random search and squeeze engine, opting to add increasingly strong and diverse packing-refining algorithms on top of standard basin hopping.

Once the fitness score was modified to include larger values of n, programs began to initialize squares in more structured patterns, nicely matching what real world approaches might look like.

high level program descriptions for select record breakers

Much of the fun in working with RL or evolutionary algorithms is seeing what novel behaviors are discovered. Here we list a few:

- Initially rounding the squares, then sharpening them as the box shrinks
- Writing the trivial grid immediately, so that an incomplete packing is not penalized—some light reward hacking
- Tricks during polishing

  - Moving a square into the most empty spot—some programs choose random squares, others choose edge squares
  - Swapping the angles or positions of two squares.
- Better initialization

  - Initializing squares at nice angles such as atan(1/2), atan(1/3), atan(2/3)
  - Initializing with structures such as staircases or diagonal columns
- Fixing squares during squeeze

  - For large n, fixing staircases and optimizing only the rest
  - Squares that cause overlaps during perturbations are jittered less, leading to edges stabilizing fast

In general, we see that the evolution agents had a tendency to add code and complexity. We see an explosion of program diversity around generation 20 after expanding the set of evolution prompts.

---

Around ten years ago I spent an embarrassing number of hours trying to pack n=17 on paper. While I’m unsure what the future of human involvement in math looks like, I’m content knowing that this push began with an insight from a human in a coffee shop while packing squares in squares.

I’d like to thank Mamacoffee Vodičkova, all the square packers out there, and [Equilibre Technologies](https://equilibre.ai/) for the time and a slice of the company cluster.

more to come :)

n = 295, the oldest record broken, set in 1979 by Frits Göbel

n = 86, set in 1997 by the shape packer himself, Erich Friedman

n=126 - just funny

1

Share
ryanxu20

Interpretation history

Decision trace