2026-10-11 17:12 UTC

The paper’s authors claim LLM judges detect facts that are present but systematically miss clinically important omissions, making them unreliable as sole evaluators for completeness-sensitive agent outputs.

state: expiredheat: lowuncertainty: highconvergesscott: highllm-evaluation ai-judges reliability

What is this?

The case concerns a paper claiming that LLM-based judges evaluating clinical notes are substantially better at confirming information that is present than detecting clinically important information that has been omitted—a failure mode termed “omission blindness.” If supported, this means such judges should not serve as the sole evaluators of completeness-sensitive agent outputs, particularly in medicine. The supplied snippets broadly support concerns about bias and poor agreement with clinicians in autonomous medical evaluation, but they do not identify this paper’s authors or directly establish its specific presence-versus-absence findings.

Why it matters to Scott

The claimed omission-blindness result independently supports Scott’s load-bearing position that LLM judges cannot alone certify completeness and must be paired with mechanically different, deterministic coverage gates—especially for consequential outputs. It creates a strong dated-receipts and evaluation-harness design opportunity, although the supplied grounding does not directly verify the paper’s specific findings or authorship.
ip:concept.mechanically-different-verifiersip:concept.sufficient-coverageip:concept.correlated-checkers-pitfallip:concept.deterministic-coredev:concept.validation-gated-llm-extractionradar:concept.llm-evaluationradar:concept.agent-verificationradar:tracelint-deterministic-agent-trace-checksradar:microsoft-agent-validation-framework
queries asked of Scott's wikis
  • completeness testing for agent outputs
  • LLM judges versus rubric and invariant checks
  • negative evidence and omission detection
  • evaluation harnesses for high-stakes agents
  • multi-layer verification of generated outputs
  • recall-oriented evaluation for RAG and agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnLLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notessbulaev2010
🟧 echo.paper ⭐The paper reports omission blindness in LLM evaluation of clinical notes, with judges verifying present information more reliably than missipaper authors——
🟧 hnWhat Is Your LLM Judge Measuring?splxai20

Interpretation history

Decision trace