Retrieved article excerpt
Open article · Retrieved 2026-10-05T23:34:32.230666+00:00
# Dust: Pretraining Transformers Without Backpropagation
Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna
October 2026Correspondence to [email protected]·[Code](https://github.com/qlabs-eng/dust)·Cite
Copy
```
@misc{dahal2026backprop,
title = {Dust: Pretraining Transformers Without Backpropagation},
author = {Dahal, Samip and Mandal, Bishwas and G{\"u}lbahar, Serdar and Vegesna, Akshay},
year = {2026},
url = {https://qlabs.sh/research/dust}
}
```
Animation of the method
## TL;DR
- We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a *virtual* population member and one forward pass evaluates them all in parallel.
- Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even *exceeds* it. This hints that in a compute-rich regime we might be able to surpass backprop.
- Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations.
- Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes.
- Dust’s gradient estimates align better with backprop’s as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling.
## 1 Introduction
Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets, including transformer-based language models. Backprop requires differentiability and produces first-order gradients, and deep learning’s architectures, optimizers, and hardware have co-evolved around this constraint.
However, as the amount of compute available in the world increases, we might prefer more generic and brute-force learning algorithms based on *search* over inductive biases like differentiability, backprop, and approximations of higher-order gradients. The bitter lesson ([Sutton, 2019](https://qlabs.sh/research/dust#ref-sutton2019bitter)Richard S. Sutton. The bitter lesson. <http://www.incompleteideas.net/IncIdeas/BitterLesson.html>, 2019. Blog post.) is that general methods that scale with compute eventually win, and AlphaGo Zero ([Silver et al., 2017](https://qlabs.sh/research/dust#ref-silver2017mastering)David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. *Nature*, 550 (7676): 354–359, 2017. doi: [10.1038/nature24270](https://doi.org/10.1038/nature24270).) is the obvious example. Bootstrapping AlphaGo on human data helped the network learn faster initially, but with a lot of computation the purely self-play network overtook it. Similarly, differentiability and backprop might be good inductive biases in the low-compute regime, where they make learning efficient, but in the high-compute regime they limit the space of architectures that work. Even within an architecture, gradient-based methods fail to explore the loss landscape optimally ([Liu et al., 2020](https://qlabs.sh/research/dust#ref-liu2020bad)Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In *Advances in Neural Information Processing Systems*, volume 33, 2020.). This might also explain why current neural nets require massive amounts of data to generalize. A more flexible credit assignment algorithm based on search is likely an important step towards much better generalization.
In this paper, we aim to replace backprop with a learning algorithm based much more on brute-force computation and much less on analytic structure. We call it Dust. Dust is a zeroth-order optimization algorithm that perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. Traditional ES methods that perturb weights ([Salimans et al., 2017](https://qlabs.sh/research/dust#ref-salimans2017evolution)Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. *arXiv preprint arXiv:1703.03864*, 2017.), like EGGROLL ([Sarkar et al., 2025](https://qlabs.sh/research/dust#ref-sarkar2025eggroll)Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. *arXiv preprint arXiv:2511.16652*, 2025.), scale with population, but scaling the population is costly because each member must be materialized and evaluated. We remove both costs with the concept of *virtual population*, where we avoid materializing every member by bypassing weight space entirely and instead perturb activations, as in node perturbation ([Werfel et al., 2003](https://qlabs.sh/research/dust#ref-werfel2003learning)Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In *Advances in Neural Information Processing Systems*, volume 16, 2003.; [Widrow and Lehr, 1990](https://qlabs.sh/research/dust#ref-widrow1990thirty)Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. *Proceedings of the IEEE*, 78 (9): 1415–1442, 1990. doi: [10.1109/5.58323](https://doi.org/10.1109/5.58323).). We do so independently at every token, so each token is a member and one forward pass evaluates them all in parallel.
Activations are a more interesting space to search over than weights. Mechanistic interpretability has shown that reasoning, whether verbalizable or not, lives in the activations ([Gurnee et al., 2026](https://qlabs.sh/research/dust#ref-gurnee2026workspace)Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. *arXiv preprint arXiv:2607.15495*, 2026.; [Lindsey et al., 2025](https://qlabs.sh/research/dust#ref-lindsey2025biology)Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. *Transformer Circuits Thread*, 2025.), which means this approach could turn training into a search over latent reasoning ([Vegesna and Dahal, 2025](https://qlabs.sh/research/dust#ref-vegesna2025decoupling)Akshay Vegesna and Samip Dahal. Decoupling search and learning in neural net training. *arXiv preprint arXiv:2509.10973*, 2025.). We then pair the activation-space perturbation with a very generic credit assignment rule that assigns different token-level rewards to different layer types in a transformer block. Those two biases, along with a few implementation details and efficiency measures, like avoiding interference between perturbed modules, are the whole algorithm.
We make the following contributions.
- We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. At large populations Dust exceeds backprop in multiple settings, which suggests that in a compute-rich regime we might be able to surpass backprop.
- Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, based on our extrapolations.
- Contrary to conventional wisdom, larger models are often more population-efficient, not less, and can make use of larger populations. This gives a new view of overparameterization as a larger search space with potentially better geometry.
- Dust’s gradient estimates align better with backprop’s as the population grows, and the alignment holds up at every scale we test, up to 1B tokens, which is encouraging for scaling.
The goal of this paper is to lay the foundations of a search-based credit assignment algorithm that is competitive with backprop on the *hardest task* we could think of: pretraining transformers. We do not attempt to make it compute-efficient enough to replace backprop today. We also do not train the new kinds of neural nets it makes accessible, like nets with an external program in the loop or transformers looped over many steps that backpropagation through time struggles to train. Both are left to future work.
## 2 Method
Dust works as follows. We add Gaussian noise to the output of each linear layer, independently at every token, run a forward pass, and reward each token’s noise by the change in loss at that token. The reward-weighted noise, averaged over draws, is the estimated error at the layer’s output, and its outer product with the layer’s input is the weight gradient. Attention internals get a variant of it: they are credited through the estimated error at the attention output over current and future tokens, instead of the tokens’ loss directly. The core intuition is that while weight-space ES evaluates one population member per forward pass, we evaluate one per token, in parallel, and a member is materialized by adding noise to a hidden state, which is cheap. On a modern transformer a single forward pass therefore evaluates a population at least *three orders of magnitude* larger than weight-space ES. We describe each component in detail below.
### 2.1 Activation-Space Perturbation
The bottleneck of evolution strategies is population size. Every member needs its own perturbed copy of the weights and its own forward pass. EGGROLL ([Sarkar et al., 2025](https://qlabs.sh/research/dust#ref-sarkar2025eggroll)Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. *arXiv preprint arXiv:2511.16652*, 2025.) makes the copies cheap with low-rank perturbations, but each member is still one sequence element of the batch, so the population is bounded by the forward passes one can afford. We perturb activations instead, independently at every token. At that token the network behaves as if a low-rank perturbation had been applied to the weights of the layer that produced the activations, without the perturbation ever being materialized in the weights. We call this a *virtual population*. A sequence in a transformer has a few thousand tokens, so one forward pass evaluates a few thousand members per sequence instead of one. Every weight in the model is trained this way except the $2L$ residual mixing scalars, which are trained by ordinary weight-space ES.
Adding noise to activations rather than weights is node perturbation ([Widrow and Lehr, 1990](https://qlabs.sh/research/dust#ref-widrow1990thirty)Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. *Proceedings of the IEEE*, 78 (9): 1415–1442, 1990. doi: [10.1109/5.58