When LLM judges agree, should we believe them?
Summary
A Amazon Science blog post exploring how to aggregate judgments from multiple LLMs when they evaluate the same content. It argues that simple vote counts can be misleading if judge outputs are correlated, and proposes dependence-aware label aggregation using Ising models to account for inter-judge dependencies in unsupervised settings.