DigiNews

Tech Watch by Johan Denoyer

← Back to articles

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Quality: 8/10 Relevance: 9/10

Summary

The paper HANDBOOK.md introduces a long-context agentic instruction-following benchmark. It tests whether enterprise-style handbooks effectively constrain AI agents, revealing that policy documents often fail to constrain actions over long horizons. It provides 65 tasks across domains, with a fully deterministic rubric; results show only a minority pass; aims to release the tasks and evaluation harness on GitHub.

🚀 Service construit par Johan Denoyer