The Two MMLU Scores: What a Benchmark Name Does Not Fix
Summary
Dmitrii Zatona analyzes two MMLU accuracy numbers reported for the same model family but different builds, arguing that the benchmark name alone does not ensure comparability. He details how splits, implementations, prompts, graders, and runner environments affect results, and notes that comparability is a property of the reference rather than the number. The article also discusses how the AI-Eval verifier can report incomparability for a score-delta when no bridge is supplied.