>10x More Efficient Pretraining
Summary
Magic AI Labs reports a >10x improvement in compute efficiency for pretraining, achieving substantially lower FLOPs than leading open-weight base models while maintaining strong perplexity and generalization. The post presents scaling laws, bits-per-byte metrics, and heldout evaluations across domains, and outlines plans to scale toward trillion-parameter models, long-context RL, and alignment research.