DigiNews

Tech Watch by Johan Denoyer

← Back to articles

The Two MMLU Scores: What a Benchmark Name Does Not Fix

Quality: 8/10 Relevance: 9/10

Summary

Dmitrii Zatona analyzes two MMLU accuracy numbers reported for the same model family but different builds, arguing that the benchmark name alone does not ensure comparability. He details how splits, implementations, prompts, graders, and runner environments affect results, and notes that comparability is a property of the reference rather than the number. The article also discusses how the AI-Eval verifier can report incomparability for a score-delta when no bridge is supplied.

🚀 Service construit par Johan Denoyer