DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Exercises in benchmarking and evals, part 7: DeepSWE, Senior SWE-Bench, napkin math, and winter tires

Quality: 8/10 Relevance: 9/10

Summary

The article critiques popular AI benchmarking suites DeepSWE and Senior SWE-Bench, discusses napkin-math style benchmarks, and questions representativeness and methodology. It also covers how such benchmarks impact perception and decisions, with caveats about value for interview prep and for developers. It includes references and a call-out to a tire-benchmark discussion and disk performance notes.

🚀 Service construit par Johan Denoyer