Benchmarking retrieval for agents on messy real-world company knowledge
Summary
The article introduces Company Knowledge Bench, a benchmark for evaluating how agents retrieve information from messy real-world company data. It compares seven retrievers across 1,000 eval cases, spanning fixed pipelines and agentic retrievers, and reports metrics like time per query, precision, and tokens returned. The findings suggest optimized agentic retrievers offer strong performance with manageable latency and cost, and the piece details how the benchmark was constructed from production data and human-labelled guidance.