Dust: Pretraining Transformers Without Backpropagation
Summary
Dust introduces a zeroth-order method that pretrains transformers by perturbing activations instead of weights. It achieves competitive, and sometimes superior, performance to backprop on large compute budgets, and offers orders-of-magnitude efficiency gains over weight-space ES. The work also analyzes how activation perturbations can scale with model size and compute, and discusses implications for future architecture search and training dynamics.