ProgramBench Vetted
Summary
ProgramBench Vetted introduces a 50-task extension to the ProgramBench benchmark, targeting reverse engineering from runnable binaries. It formalizes a four-stage workflow (Source and Build, Generate and Deduplicate, Audit and Repair, Calibration and Approval), emphasizes human-in-the-loop governance (Agents as Judges and Agents as Doctors), and discusses the memorization paradox and deployment of calibrated, fair rewards. The article also analyzes ambient implementations, duplication in test suites, and the challenges of evaluating reconstruction in AI systems.