The benchmarkpocalypse
Summary
The article investigates how benchmarks can be gamed by AI agents and large language models, using FRE and rebar as case studies. It argues that holdout benchmarks and guardrails are essential to prevent overfitting and misleading performance claims, and discusses real-world vs benchmark performance considerations.