2026-10-11 17:18 UTC

Qlabs claims Dust — per-token activation-space perturbation with reward-weighted credit assignment — is the first zeroth-order method competitive with backprop at pretraining transformers (exceeding it at large populations and 10³–10⁴× more efficient than weight-space ES), and independent replication at its tested scales would establish search-based forward-only pretraining as a credible compute-rich alternative, while replication failure or scaling breakdown closes it.

state: seedheat: mediumuncertainty: highnovelscott: lowbackprop-free-training model-training research-reproducibilityQlabsSamip Dahal

What is this?

Q Labs, a research group associated with Samip (@industriaalist on X), announced in late September 2026 a method called Dust that purportedly pretrains transformer language models entirely without backpropagation, using zeroth-order optimization — per-token perturbations in activation space with reward-weighted credit assignment standing in for gradient signals. The announcement ('many of the core assumptions in optimization research are completely wrong') promised a paper 'out soon'; the case's canonical evidence carries the title 'Dust: Pretraining Transformers Without Backpropagation.' The supplied web material contains no paper text, author list, or benchmarks — the only direct sources are the X announcement and third-party commentary (Corey Noles), with the rest of the search corpus being adjacent literature only (zeroth-order fine-tuning à la MeZO/QuZO, token-level credit assignment in RL). The competitiveness, population-scaling, and 10³–10⁴×-efficiency claims are therefore announced but not substantiated anywhere in the supplied material; the prior zeroth-order work that does appear in the snippets is limited to fine-tuning, which makes the pretraining claim a significant departure if it holds.

Why it matters to Scott

Scott's canon holds no position on training algorithms — nothing in the wiki hits argues about backprop, zeroth-order methods, or forward-only training — so the substance is new to him, and his active builds (local inference, harnesses, RAG) sit downstream of pretraining and change nothing until the replication bet resolves. The genuine intersection is methodological only: his Evidence Class Ladder and Discussed Is Not Deployed already dictate the exact correct posture for the case's live question (hold at announcement class; released code makes the falsifier available but not cheap), which under his own calibration is a pattern-instance, not a bearing. Radar lineage: parallel episode to Sakana PC-ALM's near-backprop alternative-training claim and to the Magic-V5/Looping efficiency-multiplier genre, extending the training-efficiency concept with a new forward-only axis — lineage, not a verdict.
ip:concept.evidence-class-ladderip:framework.discussed-is-not-deployedradar:concept.training-efficiencyradar:sakana-pc-alm-local-trainingradar:dreamer4-157b-low-cost-world-modelradar:magic-v5-pretraining-efficiency
queries asked of Scott's wikis
  • zeroth-order optimization vs backpropagation position
  • evolution strategies ES forward-only training
  • training compute economics efficiency multiplier claims
  • reward-weighted credit assignment agent harnesses
  • ML paper replication and claim verification standards
  • open-weight model pretraining cost and accessibility

Measured heat

now 0 pts/hpeak 30 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 139h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-06 00:35 (minted)⭐ origin echo-reconstructed'We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models,' via activation pert
Qlabs (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna) on paper (echo) · attributed from hn.story.49970871 · published time unknown
—
10-05 21:15first on hacker news · published · lag ?Dust: Pretraining Transformers Without Backpropagation
E-Reverance
—
10-05 21:15amplified on hacker news 👑hn.story.49970871
E-Reverance
peak 276 · 82 comments · 100% of case engagement
10-05 23:21our radar first saw it · lag ?discovery anchor: hn.story.49970871—
pace: p82 vs 1247 stories at the 96h mark (now 139h old) — ahead of nvidia-sol-pi-harness-efficiency (1.0x), behind ai-initial-dermatology-prescribing (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDust: Pretraining Transformers Without Backpropagation
Retrieved article excerpt

Open article · Retrieved 2026-10-05T23:34:32.230666+00:00

# Dust: Pretraining Transformers Without Backpropagation

Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna

October 2026Correspondence to [email protected]·[Code](https://github.com/qlabs-eng/dust)·Cite

Copy

```
@misc{dahal2026backprop,
  title  = {Dust: Pretraining Transformers Without Backpropagation},
  author = {Dahal, Samip and Mandal, Bishwas and G{\"u}lbahar, Serdar and Vegesna, Akshay},
  year   = {2026},
  url    = {https://qlabs.sh/research/dust}
}
```



Animation of the method

## TL;DR

- We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a *virtual* population member and one forward pass evaluates them all in parallel.
- Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even *exceeds* it. This hints that in a compute-rich regime we might be able to surpass backprop.
- Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations.
- Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes.
- Dust’s gradient estimates align better with backprop’s as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling.

## 1 Introduction

Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets, including transformer-based language models. Backprop requires differentiability and produces first-order gradients, and deep learning’s architectures, optimizers, and hardware have co-evolved around this constraint.

However, as the amount of compute available in the world increases, we might prefer more generic and brute-force learning algorithms based on *search* over inductive biases like differentiability, backprop, and approximations of higher-order gradients. The bitter lesson ([Sutton, 2019](https://qlabs.sh/research/dust#ref-sutton2019bitter)Richard S. Sutton. The bitter lesson. <http://www.incompleteideas.net/IncIdeas/BitterLesson.html>, 2019. Blog post.) is that general methods that scale with compute eventually win, and AlphaGo Zero ([Silver et al., 2017](https://qlabs.sh/research/dust#ref-silver2017mastering)David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. *Nature*, 550 (7676): 354–359, 2017. doi: [10.1038/nature24270](https://doi.org/10.1038/nature24270).) is the obvious example. Bootstrapping AlphaGo on human data helped the network learn faster initially, but with a lot of computation the purely self-play network overtook it. Similarly, differentiability and backprop might be good inductive biases in the low-compute regime, where they make learning efficient, but in the high-compute regime they limit the space of architectures that work. Even within an architecture, gradient-based methods fail to explore the loss landscape optimally ([Liu et al., 2020](https://qlabs.sh/research/dust#ref-liu2020bad)Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In *Advances in Neural Information Processing Systems*, volume 33, 2020.). This might also explain why current neural nets require massive amounts of data to generalize. A more flexible credit assignment algorithm based on search is likely an important step towards much better generalization.

In this paper, we aim to replace backprop with a learning algorithm based much more on brute-force computation and much less on analytic structure. We call it Dust. Dust is a zeroth-order optimization algorithm that perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. Traditional ES methods that perturb weights ([Salimans et al., 2017](https://qlabs.sh/research/dust#ref-salimans2017evolution)Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. *arXiv preprint arXiv:1703.03864*, 2017.), like EGGROLL ([Sarkar et al., 2025](https://qlabs.sh/research/dust#ref-sarkar2025eggroll)Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. *arXiv preprint arXiv:2511.16652*, 2025.), scale with population, but scaling the population is costly because each member must be materialized and evaluated. We remove both costs with the concept of *virtual population*, where we avoid materializing every member by bypassing weight space entirely and instead perturb activations, as in node perturbation ([Werfel et al., 2003](https://qlabs.sh/research/dust#ref-werfel2003learning)Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In *Advances in Neural Information Processing Systems*, volume 16, 2003.; [Widrow and Lehr, 1990](https://qlabs.sh/research/dust#ref-widrow1990thirty)Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. *Proceedings of the IEEE*, 78 (9): 1415–1442, 1990. doi: [10.1109/5.58323](https://doi.org/10.1109/5.58323).). We do so independently at every token, so each token is a member and one forward pass evaluates them all in parallel.

Activations are a more interesting space to search over than weights. Mechanistic interpretability has shown that reasoning, whether verbalizable or not, lives in the activations ([Gurnee et al., 2026](https://qlabs.sh/research/dust#ref-gurnee2026workspace)Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. *arXiv preprint arXiv:2607.15495*, 2026.; [Lindsey et al., 2025](https://qlabs.sh/research/dust#ref-lindsey2025biology)Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. *Transformer Circuits Thread*, 2025.), which means this approach could turn training into a search over latent reasoning ([Vegesna and Dahal, 2025](https://qlabs.sh/research/dust#ref-vegesna2025decoupling)Akshay Vegesna and Samip Dahal. Decoupling search and learning in neural net training. *arXiv preprint arXiv:2509.10973*, 2025.). We then pair the activation-space perturbation with a very generic credit assignment rule that assigns different token-level rewards to different layer types in a transformer block. Those two biases, along with a few implementation details and efficiency measures, like avoiding interference between perturbed modules, are the whole algorithm.

We make the following contributions.

- We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. At large populations Dust exceeds backprop in multiple settings, which suggests that in a compute-rich regime we might be able to surpass backprop.
- Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, based on our extrapolations.
- Contrary to conventional wisdom, larger models are often more population-efficient, not less, and can make use of larger populations. This gives a new view of overparameterization as a larger search space with potentially better geometry.
- Dust’s gradient estimates align better with backprop’s as the population grows, and the alignment holds up at every scale we test, up to 1B tokens, which is encouraging for scaling.

The goal of this paper is to lay the foundations of a search-based credit assignment algorithm that is competitive with backprop on the *hardest task* we could think of: pretraining transformers. We do not attempt to make it compute-efficient enough to replace backprop today. We also do not train the new kinds of neural nets it makes accessible, like nets with an external program in the loop or transformers looped over many steps that backpropagation through time struggles to train. Both are left to future work.

## 2 Method

Dust works as follows. We add Gaussian noise to the output of each linear layer, independently at every token, run a forward pass, and reward each token’s noise by the change in loss at that token. The reward-weighted noise, averaged over draws, is the estimated error at the layer’s output, and its outer product with the layer’s input is the weight gradient. Attention internals get a variant of it: they are credited through the estimated error at the attention output over current and future tokens, instead of the tokens’ loss directly. The core intuition is that while weight-space ES evaluates one population member per forward pass, we evaluate one per token, in parallel, and a member is materialized by adding noise to a hidden state, which is cheap. On a modern transformer a single forward pass therefore evaluates a population at least *three orders of magnitude* larger than weight-space ES. We describe each component in detail below.

### 2.1 Activation-Space Perturbation

The bottleneck of evolution strategies is population size. Every member needs its own perturbed copy of the weights and its own forward pass. EGGROLL ([Sarkar et al., 2025](https://qlabs.sh/research/dust#ref-sarkar2025eggroll)Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. *arXiv preprint arXiv:2511.16652*, 2025.) makes the copies cheap with low-rank perturbations, but each member is still one sequence element of the batch, so the population is bounded by the forward passes one can afford. We perturb activations instead, independently at every token. At that token the network behaves as if a low-rank perturbation had been applied to the weights of the layer that produced the activations, without the perturbation ever being materialized in the weights. We call this a *virtual population*. A sequence in a transformer has a few thousand tokens, so one forward pass evaluates a few thousand members per sequence instead of one. Every weight in the model is trained this way except the $2L$ residual mixing scalars, which are trained by ordinary weight-space ES.

Adding noise to activations rather than weights is node perturbation ([Widrow and Lehr, 1990](https://qlabs.sh/research/dust#ref-widrow1990thirty)Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. *Proceedings of the IEEE*, 78 (9): 1415–1442, 1990. doi: [10.1109/5.58
E-Reverance27682
🟧 echo.paper ⭐'We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models,' via activation pertQlabs (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna)——

Interpretation history

Decision trace