2026-10-11 17:12 UTC

The paper’s authors claim that exposing LLM judges to prior scores materially anchors their ratings and can distort model comparisons based on those evaluations.

state: expiredheat: lowuncertainty: mediumconvergesscott: mediumllm-evaluation benchmark-reliability llm-judges

What is this?

The case concerns a paper claiming that showing an LLM judge prior scores anchors its subsequent ratings, systematically biasing score-based evaluations and potentially altering model comparisons. The supplied results support the broader concern that LLM-as-a-Judge scoring is vulnerable to systematic bias and requires calibration or debiasing, but they do not clearly identify this paper’s authors or substantiate the cited 192,000-evaluation experiment. One result discusses bias from judges’ prior beliefs, which is related but not clearly the same prior-score anchoring effect.

Why it matters to Scott

The claimed prior-score anchoring effect directly supports Scott’s implemented “Verdict-free evidence reuse” pattern, which withholds old scores and conclusions from subsequent reviewers, and reinforces his insistence on independently failing verification mechanisms. It could inform evaluation-harness design and provide an empirical receipt for that position, although the supplied grounding does not substantiate the paper’s authorship or reported experimental scale.
dev:concept.verdict-free-evidence-reuseip:concept.mechanically-different-verifiersip:concept.correlated-failureradar:concept.llm-evaluationradar:concept.benchmark-integrityradar:concept.model-evaluation
queries asked of Scott's wikis
  • LLM judge anchoring and evaluation contamination
  • benchmark reliability and biased model rankings
  • calibration protocols for automated evaluators
  • independent scoring versus shared evaluation context
  • evaluation harnesses for detecting judge bias
  • human review thresholds for LLM evaluations

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAnchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluationsbulaev10
🟧 echo.paper ⭐The authors’ arXiv paper reports that prior scores systematically anchor LLM-as-a-Judge ratings. Across 192,000 attempted evaluations, sevenAnte Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, and Emanuel Lacic——

Interpretation history

Decision trace