DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Dust: Pretraining Transformers Without Backpropagation

Quality: 8/10 Relevance: 9/10

Summary

Dust introduces a zeroth-order method that pretrains transformers by perturbing activations instead of weights. It achieves competitive, and sometimes superior, performance to backprop on large compute budgets, and offers orders-of-magnitude efficiency gains over weight-space ES. The work also analyzes how activation perturbations can scale with model size and compute, and discusses implications for future architecture search and training dynamics.

🚀 Service construit par Johan Denoyer