What Is RLCD? The Secret Behind Jev
Summary
Di Zhang’s blog post explains RLCD as a calibrated, schema-conditioned extension of reward modeling, culminating in Jev, a productized evaluator that returns typed, calibrated decisions. It outlines the progression from scalar rewards to pairwise preferences, then to multiway Plackett–Luce distributions, and finally to calibrated decision outputs (Noul, Choice, Score). The article also covers the decision head, parallel inference via packing and masking, and testable predictions for RLCD behavior.