ORCA-bench: How Ready Are Language Model Agents for Oncall?
Summary
ORCA-bench introduces a production-fidelity benchmark for language model agents performing on-call root-cause analysis, using a live microservice system with six days of telemetry exposed via Prometheus, Jaeger, and OpenSearch. The study uses 1,079 RCA tasks and human evaluation to show current frontier models struggle with accurate RCA, highlighting a substantial gap before deploying LLM-driven agents in production reliability scenarios.