DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Big Pickle on SWE Atlas – Codebase QnA

Quality: 9/10 Relevance: 9/10

Summary

The page documents a SWE Atlas Codebase QnA benchmark run using the mini-SWE-Agent scaffold with the big-pickle model. It reports 50.81% task resolution, placing it above many leaderboard entries within the Mini-SWE-Agent class, and provides language- and category-level breakdowns, along with methodology, caveats, and reproducibility notes.

🚀 Service construit par Johan Denoyer