HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Summary
The paper HANDBOOK.md introduces a long-context agentic instruction-following benchmark. It tests whether enterprise-style handbooks effectively constrain AI agents, revealing that policy documents often fail to constrain actions over long horizons. It provides 65 tasks across domains, with a fully deterministic rubric; results show only a minority pass; aims to release the tasks and evaluation harness on GitHub.