DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Quality: 8/10 Relevance: 9/10

Summary

Terminal-Bench-Science 0.1 introduces a continuous benchmark for evaluating AI agents on real scientific workflows, spanning 70 tasks across life, physical, Earth, mathematical, and engineering sciences. The project emphasizes verifiable, reproducible results and invites the research community to contribute tasks and iterate as frontier AI evolves. Claude Opus 5 achieves the highest resolution at 30% on the initial release.

🚀 Service construit par Johan Denoyer