Harvey LAB: The Legal Agent Benchmark
Summary
Harvey LAB is an open-source benchmark for evaluating AI agents on legal work in realistic settings. It comprises a dataset of tasks, an execution harness for running and scoring agents, and documentation guiding evaluation methodologies and participation. The project aims to standardize benchmarking in legal AI agent use, with an emphasis on reproducibility and ongoing improvement.