How to Solve Hallucination
Summary
The article explains that a model's stated confidence is not a true measurement but a byproduct of next-token predictions. It introduces RLCD (Reinforcement Learning with Calibrated Distributions) to calibrate forecasts against actual outcomes, distinguishing epistemic and aleatoric uncertainty, and demonstrates how calibrated distributions can inform decision-making (e.g., weather forecasts and stock betting using the Kelly criterion). It also discusses applying RLCD to open models to tailor calibration to specific domains and evaluating models via calibration loops rather than bare confidence.