Big Pickle on SWE Atlas – Codebase QnA
Summary
The page documents a SWE Atlas Codebase QnA benchmark run using the mini-SWE-Agent scaffold with the big-pickle model. It reports 50.81% task resolution, placing it above many leaderboard entries within the Mini-SWE-Agent class, and provides language- and category-level breakdowns, along with methodology, caveats, and reproducibility notes.