DigiNews

Tech Watch by Johan Denoyer

← Back to articles

When LLM judges agree, should we believe them?

Quality: 9/10 Relevance: 9/10

Summary

A Amazon Science blog post exploring how to aggregate judgments from multiple LLMs when they evaluate the same content. It argues that simple vote counts can be misleading if judge outputs are correlated, and proposes dependence-aware label aggregation using Ising models to account for inter-judge dependencies in unsupervised settings.

🚀 Service construit par Johan Denoyer