Jev can't be calibrated
Summary
The post analyzes Jev, TypeSafe’s System One Model, which outputs typed decisions with calibrated-like scores. The author argues that, despite claims of calibrated probabilities, Jev’s outputs may not be truly calibrated across different data distributions and that practitioners should treat its outputs as scores rather than reliable probabilities. Recalibration on real data is suggested if precise probabilistic interpretation is required.