Show HN: JevBench, a reproducible benchmark for typed decision models
Summary
This article introduces JevBench 1.3.0, Benchmark Heaven's reproducible benchmark for Jev-class typed decision models. It evaluates 52 systems across 534 decisions, detailing the JevBench Score (a geometric mean of Intelligence, Calibration, Speed, and Cost) and how rankings vary with different weightings. The piece also covers public vs held-out tasks, extensive methodology, cost modeling, and open-source resources, including a MIT-licensed harness and GitHub repo.