DoGBench: The first user-facing docs generation benchmark. No model scores >50%
Summary
DoGBench evaluates AI agents' ability to write and update user-facing documentation. It uses 292 tasks drawn from real open-source projects; the held-out 117-item set shows the best scores around the mid-40s, with cloud agents peaking at 54.8, highlighting that current AI docs often miss prerequisites and crucial guidance, necessitating expert review.