How accurately calibrated is Jev?
Summary
The post argues that Jev, a System One-style classifier built on top of transformer models, provides probability distributions for multiple-choice prompts rather than direct answers. It reports experiments calibrating Jev against known distributions (Gaussian, Maxwell-Boltzmann, etc.), finding that Jev tends to produce overly peaky distributions and struggles with multi-step arithmetic, though some distributions (like certain Lorentzian cases) fare better. The piece also surveys related literature and discusses the limitations of using automated judges for calibration, offering practical insights and caveats for experimental design.