How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Summary
The arXiv paper reevaluates frontier models on physics benchmarks with expert review and finds many initial failures were due to benchmarking issues rather than model limitations. Expert corrections to reference solutions and flawed questions substantially raise measured performance, suggesting current benchmarks understate frontier models' physics problem-solving abilities. The work argues for more demanding, expert-validated evaluations.