2026-10-11 17:13 UTC

Volotat claims mini-AGI's disk-paged, dynamically growing model trains from a batch-1 stream on an 8GB GPU while a reduced trunk learning rate sharply limits measured forgetting, potentially enabling consumer-hardware continual-learning experiments without a frozen pretrained base.

state: watchingheat: highuncertainty: highnovelscott: lowcontinual-learning small-model-training local-inferencevolotat
Surfaced 2026-09-22T04:30:30Z โ€” Mini-AGI โ€“ dynamic continual learning model trained from scratch on 8GB VRAM โ€” The project now has loud, fast-moving attention across Hacker News and Reddit, warranting high heat, but the discussion adds questions rather than independent validation. The case remains a single-author toy experiment awaiting published weights, reproduction, and meaningful generalization tests.

What is this?

Volotat's mini-AGI is a GitHub-hosted experiment (volotat/mini-AGI) that trains a byte-level, dynamically growing MoE language model from a single batch-1 data stream on an 8 GB consumer GPU, paging weights and optimizer state to disk so parameter count is limited by storage not VRAM. The author reports that lowering the shared trunk learning rate to 0.1ร— reduced measured forgetting on seven held-out subjects from +2.23 to +0.0067 nats, but the checkpoint weights remain unpublished (promised after the first corpus pass), no independent reproduction exists, and community discussion (HN, Reddit) notes the training transcripts show no coherent output and no generalization benchmarks have been run. The project explicitly labels itself a 'toy-level' demonstration, not a frontier capability.

Why it matters to Scott

The case is a single-author toy experiment (Volotat/mini-AGI) demonstrating disk-paged, dynamically growing MoE training on an 8GB GPU with a reported forgetting reduction via trunk LR scaling. It illustrates patterns Scott tracks โ€” local training, continual learning, memory-efficient MoE โ€” but adds no independent validation, no coherent output, no benchmarks, and no published weights. It does not challenge, extend, or converge on any specific claim in Scott's canon (scaffolding hypothesis, wiki-kernel, sovereign software assurance, generative pendulum). It is another data point in the space, not a bearing event.
radar:concept.continual-learningradar:concept.local-inferenceradar:concept.memory-efficient-trainingradar:concept.moe-streamingradar:thomson-1-continual-learningradar:rwkv8-m1-low-memory-trainingradar:sub2bit-disk-context-local-modelradar:deepseek-v4-flash-expert-streamingradar:llama-cpp-lazy-tensor-loadingradar:llama-cpp-hot-expert-offload
queries asked of Scott's wikis
  • continual learning catastrophic forgetting agent memory
  • local training consumer GPU 8GB VRAM disk offloading
  • dynamic model growth MoE expert paging from scratch
  • single stream batch-1 training online learning
  • open weight model sovereignty local inference economics
  • agent maintained wiki memory systems continual adaptation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 493h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-21 04:42โญ origin directly observedMini-AGI โ€“ dynamic continual learning model trained from scratch on 8GB VRAM
volotat on hacker news
โ€”
09-21 03:26first on r/LocalLLaMA ยท published ยท +-1.3hmini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.
Another__one
โ€”
09-21 05:23first on github (echo) ยท first seen by us ยท +0.7hPublishes a toy-level continual-learning implementation and training samples, reporting forgetting of +0.0067 rather than +2.2300 nats with
volotat
โ€”
09-21 03:26amplified on r/LocalLLaMAreddit.post.1wm1gab
Another__one
peak 235 ยท 48 comments ยท 28% of case engagement
09-21 04:42amplified on hacker news ๐Ÿ‘‘hn.story.49783133
volotat
peak 277 ยท 80 comments ยท 63% of case engagement
09-22 17:46amplified on r/LocalLLaMAreddit.post.1wnggz5
returnity
peak 76 ยท 21 comments ยท 9% of case engagement
09-21 04:20our radar first saw it ยท +-0.4hdiscovery anchor: reddit.post.1wm1gabโ€”
09-22 04:30reached heat=high ยท +23.8h ยท via ledgerโ€”โ€”
pace: p87 vs 1032 stories at the 336h mark (now 493h old) โ€” ahead of linear-ci-throughput-redesign (1.0x), behind world-labs-atlas-spatial-model (1.0x)

Evidence (4) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditmini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-09-21T04:21:48.980815+00:00

# mini-AGI

mini-AGI - is a **continual learning** byte-level language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads.
It stores its weights as ordinary files on disk and pages them onto the card as it needs them, so the parameter count is bounded by free disk space rather than by VRAM. It grows new capacity while training when it runs short, prunes what nothing asks for, and reads through exactly the same code path it serves on. Targeted at a PC or laptop with at least an 8 GB VRAM GPU on the board.

**NOTE: as of now this is a small toy-level model.** Do not expect a frontier level capabilities. This is rather a small experiment to show, that continual learning from the single stream of data without catastrophic forgetting is possible. Furthermore it is possible on a modest hardware. Which means that almost everyone could train their own version of the model (or simply continue training this one) exactly as they see it fit. And the capabilities would be bounded by the actual hardware, scale and quality of the data available and the amount of time one willing to spend on training the model.

[dashboard](https://github.com/volotat/mini-AGI/blob/main/assets/dashboard.png)
*Here is how min-run dashboard looks like. The model is pointed to the corpus to constantly read and learn from.*

[History](https://github.com/volotat/mini-AGI/blob/main/runs/samples.txt) - here is the samples from the whole training run history so far. You can inspect them yourself to see how the model improved over the course of training/reading the corpus.

The weights are **not published yet**. The run is still reading its first pass over the corpus, the weights go up once it has been through all of it, which is a couple of weeks away at the current rate.

## Motivation

Every language model you can actually own today is a model somebody else trained and then froze. You can fine-tune around the edges of it, but you cannot train one from scratch on your own hardware, and you cannot keep training it on what you do day to day - the moment you try, it forgets what it knew before. The result is that a personal model is always somebody else's model with a thin layer of you on top, and it stops learning the day it ships.

**mini-AGI** model has small enough GPU footprint that it is possible to train end-to-end on one consumer card, and it is built so that training never has to stop. It reads a stream of characters one chunk at a time, takes a gradient step on each, and the same path serves generation. There is no separate fine-tuning regime and no frozen base: reading and being trained are the same event.

Three constraints shape everything else in the design:

- **It has to fit on 8 GB.** Not with quantisation - training needs gradients and optimiser state, which is roughly three times the weights again. So the weights live on disk and only the working set is resident.
- **It has to not forget.** A model that learns continually and overwrites itself is worse than one that does not learn at all.
- **It has to be able to read anything.** The alphabet is the 256 byte values, so there is no tokenizer to fit and no data type that needs a new vocabulary.

The model is genuinely yours: trained on your hardware, on your data, that keeps learning from every conversation you have with it, and that nobody else can take it away or switch it off.

## How the architecture works

Characters (bytes) does not pass through a fixed stack of layers as it would be in a traditional LLM. Instead, it passes through **two dense prelude blocks** and then through **one recurrent block applied up to 24 times**, each application choosing its own experts from a shared pool. The latent state between applications is never decoded - it is merged with the embedded input by an adapter each time round, so the loop cannot drift away from the text it is reading.

Three distinct blocks, up to 26 block-applications per character.

- **Adaptive depth.** A halting head scores every character at every row, and the character stops as soon as another row would not change the answer. Easy characters take one row, hard ones take many. This is the PonderNet recipe: while training, every depth is computed and weighted by its halting probability, so the halting head learns through those weights.
- **Routing per block-application, not per character.** Each of the 26 applications picks its own top-8 experts, so one character touches far more of the pool than "top-8" suggests, and the same expert can be selected several times at different depths. What varies is *which* eight at each point.
- **No expert is assigned a subject.** There are no labels anywhere. Soft top-k routing distributes capability across the pool by itself, and a character can combine fragments from several experts. The cost is that capabilities share parameters and so *can* interfere.

[how the model processes one character](https://github.com/volotat/mini-AGI/blob/main/assets/shape.gif)

**This is the architecture assembling itself, one character at a time**, captured from the live model - nothing here is drawn by hand.

Each tile on the left is one expert; colour is expert identity and stays the same for the whole clip. A **row** is one application of the recurrent block, and the eight tiles in it are the eight experts that row actually ran. The stack grows downward as the model keeps going, and the amber line is where halting stopped it - **the grey rows below are computation the model declined to spend.**

The trace on the right is how many rows each character took. It moves constantly between 4 and 14 against a ceiling of 24, and the caret under the text shows which character is being read.

Positions are rotary and carry no learned parameters, which is why the context window can be extended by continued training rather than by re-initialising anything.

### ...and the same thing while it writes

[how the model generates text](https://github.com/volotat/mini-AGI/blob/main/assets/generate.gif)

The clip above is the model **reading** - every character is held-out text it is being shown. This one is the model **writing**: it was primed with 2,500 characters of a held-out story and then continued on its own, so the grey text is what it was given and **the green text is entirely its own**. Greedy decoding, no sampling anywhere - run it twice and you get the same sentence.

Two things are worth watching. The stack behaves the same way, because generating and reading are the same forward pass in this model - the only difference is whether the next character comes from a file or from the model's own argmax. And **writing costs more depth than reading**: about 9.9 rows a character against 8.0 on the same subject. The dotted lines mark where the working set was re-chosen, which happens every 64 characters; in this clip nothing swapped, because the prompt had already pulled the right experts onto the card.

What it produced, continuing a story about a cherry tree:

> They worked together and saw their favorite shore. One day, they wanted to play with their favorite shore. They wanted to play with it, but

Grammatically correct and on-topic. It does repeats itself for now - which is a fair picture of where the model is at 243M characters.

## How paging works

Every expert is a file on disk holding its weights and its Adam moments. Above disk sit two caches and a working set:

|  | key | what it is |
| --- | --- | --- |
| disk | โ€” | every expert the model has; bounded by free space |
| RAM | `ram_cache` | recently wanted experts, least-recently-used evicted |
| VRAM | `resident` | the working set - what a character may route through |

Before every chunk the model is asked what the text about to be read wants, and the answer becomes the working set. Demand is scored on the hidden states the call sites actually routed on while reading the previous chunk - an embedding carries no context, so scoring on raw embeddings would have every subject asking for the same experts.

Two rules the project holds to:

- **Adam's moments travel with the expert.** They belong to the expert, not to the slot of VRAM it happened to occupy. Leaving them behind would hand one expert's momentum to whatever took its place, and training would carry on looking healthy while every swapped expert inherited a stranger's history.
- **An expert already on the card stays in the slot it is in.** Demand comes back sorted, so the order churns while the set itself barely moves. Matching by identity rather than by position is what keeps the number of loads equal to how much the set really changed.

Because the choice is made from the previous chunk, it cannot see the text it is about to predict. What keeps the working set from churning on noise is hysteresis - a candidate has to beat a resident by `margin` to displace it, and a newcomer is safe for `dwell_chars` of reading.

## How growth and pruning work

The pool grows when it is short of capacity and shrinks when parts of it stop being asked for.

New experts are added on speculation, at a small gate so they change almost nothing, and kept only if something goes on asking for them. A new expert is built by **recombination** - whole hidden units taken from several existing experts - because a clone of one parent is not novel enough to be worth routing to, and a random expert computes nothing worth routing to. What works is novelty assembled from trained parts.

Growth is refused unless every brake agrees:

- **room** - disk and VRAM can take it
- **used** - the capacity already added is being asked for
- **earning** - the previous cohort survived its trial
- **fits** - not too many experts are already inside their trial
- **honest** - train and held-out have not separated

**Dead means unaddressed.** Both the growth brake and the pruner read how long it has been since anything asked for an expert, and never its gate. This is the single most useful finding in the repository: the gate is not merely uninformative here, it is anti-predictive. The smallest gates belong to the *busiest* experts - one that behaves as a sink, chosen constantly and contributing little per character, reads as dead on a gate test, while a high-gate expert nothing has wanted in hundreds of thousands of segments reads as alive.

A new expert is safe for a full survival window no matter what, so it cannot be judged before it has had a chance to be chosen. When the model grows an expert a new file appears; when it prunes one, that file is deleted.

## How continual learning works

Training on a single stream, one subject at a time, is the classic recipe for catastrophic forgetting. Reading half a million characters of chess at the experts' own learning rate takes the other seven subjects from 1.12 to 3.73 nats.

**The trunk learning rate is the mechanism.** The trunk - embeddings, attention, routers, the halting head - is the part every character passes through, and it carries 97.6% of the squared gradient norm. Running it at 0.1x the experts' rate takes forgetting from +2.2300 to +0.0067 nats, which is 99.84% of progress retained against chance.

| configuration | unread subjects | retained vs chance |
| --- | --- | --- |
| working set frozen, trunk LR = expert LR | +2.5871 | 42.88% |
| swapping, trunk LR = expert LR | +2.2300 | 50.68% |
| **swapping, trunk at 0.1x - what the run uses** | **+0.0067** | **99.84%** |
| *control: all seven subjects read* | *-0.0077* | *-* |

[Forgetting under three configurations](https://github.com/volotat/mini-AGI/blob/main/assets/mitigations.png)

**This is the measurement the whole design rests on.** The model reads 524,000 characters of chess and nothing else, at batch 1, and the y-axis on the left is what happened to the **seven subjects it did not read** - zero means nothing was forgotten, up means worse. Three lines, one variable each. Two of them climb to +2.2 and +2.6 nats, which is t
Another__one23548
๐ŸŸง hn โญMini-AGI โ€“ dynamic continual learning model trained from scratch on 8GB VRAM
Retrieved article excerpt

Open article ยท Retrieved 2026-09-21T05:21:54.199542+00:00

# mini-AGI

mini-AGI - is a **continual learning** byte-level language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads.
It stores its weights as ordinary files on disk and pages them onto the card as it needs them, so the parameter count is bounded by free disk space rather than by VRAM. It grows new capacity while training when it runs short, prunes what nothing asks for, and reads through exactly the same code path it serves on. Targeted at a PC or laptop with at least an 8 GB VRAM GPU on the board.

**NOTE: as of now this is a small toy-level model.** Do not expect a frontier level capabilities. This is rather a small experiment to show, that continual learning from the single stream of data without catastrophic forgetting is possible. Furthermore it is possible on a modest hardware. Which means that almost everyone could train their own version of the model (or simply continue training this one) exactly as they see it fit. And the capabilities would be bounded by the actual hardware, scale and quality of the data available and the amount of time one willing to spend on training the model.

[dashboard](https://github.com/volotat/mini-AGI/blob/main/assets/dashboard.png)
*Here is how min-run dashboard looks like. The model is pointed to the corpus to constantly read and learn from.*

[History](https://github.com/volotat/mini-AGI/blob/main/runs/samples.txt) - here is the samples from the whole training run history so far. You can inspect them yourself to see how the model improved over the course of training/reading the corpus.

The weights are **not published yet**. The run is still reading its first pass over the corpus, the weights go up once it has been through all of it, which is a couple of weeks away at the current rate.

## Motivation

Every language model you can actually own today is a model somebody else trained and then froze. You can fine-tune around the edges of it, but you cannot train one from scratch on your own hardware, and you cannot keep training it on what you do day to day - the moment you try, it forgets what it knew before. The result is that a personal model is always somebody else's model with a thin layer of you on top, and it stops learning the day it ships.

**mini-AGI** model has small enough GPU footprint that it is possible to train end-to-end on one consumer card, and it is built so that training never has to stop. It reads a stream of characters one chunk at a time, takes a gradient step on each, and the same path serves generation. There is no separate fine-tuning regime and no frozen base: reading and being trained are the same event.

Three constraints shape everything else in the design:

- **It has to fit on 8 GB.** Not with quantisation - training needs gradients and optimiser state, which is roughly three times the weights again. So the weights live on disk and only the working set is resident.
- **It has to not forget.** A model that learns continually and overwrites itself is worse than one that does not learn at all.
- **It has to be able to read anything.** The alphabet is the 256 byte values, so there is no tokenizer to fit and no data type that needs a new vocabulary.

The model is genuinely yours: trained on your hardware, on your data, that keeps learning from every conversation you have with it, and that nobody else can take it away or switch it off.

## How the architecture works

Characters (bytes) does not pass through a fixed stack of layers as it would be in a traditional LLM. Instead, it passes through **two dense prelude blocks** and then through **one recurrent block applied up to 24 times**, each application choosing its own experts from a shared pool. The latent state between applications is never decoded - it is merged with the embedded input by an adapter each time round, so the loop cannot drift away from the text it is reading.

Three distinct blocks, up to 26 block-applications per character.

- **Adaptive depth.** A halting head scores every character at every row, and the character stops as soon as another row would not change the answer. Easy characters take one row, hard ones take many. This is the PonderNet recipe: while training, every depth is computed and weighted by its halting probability, so the halting head learns through those weights.
- **Routing per block-application, not per character.** Each of the 26 applications picks its own top-8 experts, so one character touches far more of the pool than "top-8" suggests, and the same expert can be selected several times at different depths. What varies is *which* eight at each point.
- **No expert is assigned a subject.** There are no labels anywhere. Soft top-k routing distributes capability across the pool by itself, and a character can combine fragments from several experts. The cost is that capabilities share parameters and so *can* interfere.

[how the model processes one character](https://github.com/volotat/mini-AGI/blob/main/assets/shape.gif)

**This is the architecture assembling itself, one character at a time**, captured from the live model - nothing here is drawn by hand.

Each tile on the left is one expert; colour is expert identity and stays the same for the whole clip. A **row** is one application of the recurrent block, and the eight tiles in it are the eight experts that row actually ran. The stack grows downward as the model keeps going, and the amber line is where halting stopped it - **the grey rows below are computation the model declined to spend.**

The trace on the right is how many rows each character took. It moves constantly between 4 and 14 against a ceiling of 24, and the caret under the text shows which character is being read.

Positions are rotary and carry no learned parameters, which is why the context window can be extended by continued training rather than by re-initialising anything.

### ...and the same thing while it writes

[how the model generates text](https://github.com/volotat/mini-AGI/blob/main/assets/generate.gif)

The clip above is the model **reading** - every character is held-out text it is being shown. This one is the model **writing**: it was primed with 2,500 characters of a held-out story and then continued on its own, so the grey text is what it was given and **the green text is entirely its own**. Greedy decoding, no sampling anywhere - run it twice and you get the same sentence.

Two things are worth watching. The stack behaves the same way, because generating and reading are the same forward pass in this model - the only difference is whether the next character comes from a file or from the model's own argmax. And **writing costs more depth than reading**: about 9.9 rows a character against 8.0 on the same subject. The dotted lines mark where the working set was re-chosen, which happens every 64 characters; in this clip nothing swapped, because the prompt had already pulled the right experts onto the card.

What it produced, continuing a story about a cherry tree:

> They worked together and saw their favorite shore. One day, they wanted to play with their favorite shore. They wanted to play with it, but

Grammatically correct and on-topic. It does repeats itself for now - which is a fair picture of where the model is at 243M characters.

## How paging works

Every expert is a file on disk holding its weights and its Adam moments. Above disk sit two caches and a working set:

|  | key | what it is |
| --- | --- | --- |
| disk | โ€” | every expert the model has; bounded by free space |
| RAM | `ram_cache` | recently wanted experts, least-recently-used evicted |
| VRAM | `resident` | the working set - what a character may route through |

Before every chunk the model is asked what the text about to be read wants, and the answer becomes the working set. Demand is scored on the hidden states the call sites actually routed on while reading the previous chunk - an embedding carries no context, so scoring on raw embeddings would have every subject asking for the same experts.

Two rules the project holds to:

- **Adam's moments travel with the expert.** They belong to the expert, not to the slot of VRAM it happened to occupy. Leaving them behind would hand one expert's momentum to whatever took its place, and training would carry on looking healthy while every swapped expert inherited a stranger's history.
- **An expert already on the card stays in the slot it is in.** Demand comes back sorted, so the order churns while the set itself barely moves. Matching by identity rather than by position is what keeps the number of loads equal to how much the set really changed.

Because the choice is made from the previous chunk, it cannot see the text it is about to predict. What keeps the working set from churning on noise is hysteresis - a candidate has to beat a resident by `margin` to displace it, and a newcomer is safe for `dwell_chars` of reading.

## How growth and pruning work

The pool grows when it is short of capacity and shrinks when parts of it stop being asked for.

New experts are added on speculation, at a small gate so they change almost nothing, and kept only if something goes on asking for them. A new expert is built by **recombination** - whole hidden units taken from several existing experts - because a clone of one parent is not novel enough to be worth routing to, and a random expert computes nothing worth routing to. What works is novelty assembled from trained parts.

Growth is refused unless every brake agrees:

- **room** - disk and VRAM can take it
- **used** - the capacity already added is being asked for
- **earning** - the previous cohort survived its trial
- **fits** - not too many experts are already inside their trial
- **honest** - train and held-out have not separated

**Dead means unaddressed.** Both the growth brake and the pruner read how long it has been since anything asked for an expert, and never its gate. This is the single most useful finding in the repository: the gate is not merely uninformative here, it is anti-predictive. The smallest gates belong to the *busiest* experts - one that behaves as a sink, chosen constantly and contributing little per character, reads as dead on a gate test, while a high-gate expert nothing has wanted in hundreds of thousands of segments reads as alive.

A new expert is safe for a full survival window no matter what, so it cannot be judged before it has had a chance to be chosen. When the model grows an expert a new file appears; when it prunes one, that file is deleted.

## How continual learning works

Training on a single stream, one subject at a time, is the classic recipe for catastrophic forgetting. Reading half a million characters of chess at the experts' own learning rate takes the other seven subjects from 1.12 to 3.73 nats.

**The trunk learning rate is the mechanism.** The trunk - embeddings, attention, routers, the halting head - is the part every character passes through, and it carries 97.6% of the squared gradient norm. Running it at 0.1x the experts' rate takes forgetting from +2.2300 to +0.0067 nats, which is 99.84% of progress retained against chance.

| configuration | unread subjects | retained vs chance |
| --- | --- | --- |
| working set frozen, trunk LR = expert LR | +2.5871 | 42.88% |
| swapping, trunk LR = expert LR | +2.2300 | 50.68% |
| **swapping, trunk at 0.1x - what the run uses** | **+0.0067** | **99.84%** |
| *control: all seven subjects read* | *-0.0077* | *-* |

[Forgetting under three configurations](https://github.com/volotat/mini-AGI/blob/main/assets/mitigations.png)

**This is the measurement the whole design rests on.** The model reads 524,000 characters of chess and nothing else, at batch 1, and the y-axis on the left is what happened to the **seven subjects it did not read** - zero means nothing was forgotten, up means worse. Three lines, one variable each. Two of them climb to +2.2 and +2.6 nats, which is t
volotat27780
๐ŸŸง echo.githubPublishes a toy-level continual-learning implementation and training samples, reporting forgetting of +0.0067 rather than +2.2300 nats with volotatโ€”โ€”
๐ŸŸ  redditmini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)
LocalLLaMA
returnity7621

Interpretation history

Decision trace