Compute-Optimal Is Not Cluster-Optimal
Summary
The article introduces MOSAIC, a framework that folds the systems optimization into scaling laws for large-scale model pretraining, showing that compute-optimal configurations are not always cluster-optimal. It emphasizes hardware-aware pricing (MFU, goodput) alongside model-FLOPs, and demonstrates that sparsity in mixture-of-experts (MoE) can yield different optimal choices when deliverable FLOPs are considered. The result challenges conventional two-stage workflows and argues for a hybrid ML-and-systems approach to architecture selection.