LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It
Summary
An arXiv paper investigates whether LLM judges can reliably detect omissions in AI-generated clinical notes. It compares a per-fact pipeline with a single-call GEPA-prompt approach across eight judge designs, showing that structured, fact-by-fact checks reduce omissions and false alarms, while one-shot prompts offer different trade-offs. The study provides benchmarks, prompts, and judgments to enable reproducibility.