Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Summary
Real-SWE introduces a benchmark to evaluate frontier AI models on private, real-world enterprise codebases. It showcases model comparisons, task types from production code, and cost-per-rollout analyses, highlighting the importance of company-specific context and cross-functional tooling in enterprise AI workflows.