The paper’s authors claim that exposing LLM judges to prior scores materially anchors their ratings and can distort model comparisons based on those evaluations.
state: expiredheat: lowuncertainty: mediumconvergesscott: mediumllm-evaluation benchmark-reliability llm-judges
What is this?
The case concerns a paper claiming that showing an LLM judge prior scores anchors its subsequent ratings, systematically biasing score-based evaluations and potentially altering model comparisons. The supplied results support the broader concern that LLM-as-a-Judge scoring is vulnerable to systematic bias and requires calibration or debiasing, but they do not clearly identify this paper’s authors or substantiate the cited 192,000-evaluation experiment. One result discusses bias from judges’ prior beliefs, which is related but not clearly the same prior-score anchoring effect.
Why it matters to Scott
The claimed prior-score anchoring effect directly supports Scott’s implemented “Verdict-free evidence reuse” pattern, which withholds old scores and conclusions from subsequent reviewers, and reinforces his insistence on independently failing verification mechanisms. It could inform evaluation-harness design and provide an empirical receipt for that position, although the supplied grounding does not substantiate the paper’s authorship or reported experimental scale.
dev:concept.verdict-free-evidence-reuseip:concept.mechanically-different-verifiersip:concept.correlated-failureradar:concept.llm-evaluationradar:concept.benchmark-integrityradar:concept.model-evaluation
queries asked of Scott's wikis
- LLM judge anchoring and evaluation contamination
- benchmark reliability and biased model rankings
- calibration protocols for automated evaluators
- independent scoring versus shared evaluation context
- evaluation harnesses for detecting judge bias
- human review thresholds for LLM evaluations
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-30T07:27:11Z
The paper remains a useful empirical lead for verdict-free evaluation design, but no independent validation, implementation uptake, or discussion emerged within the case horizon. The episode has faded without advancing beyond its original uncorroborated claim.
2026-08-28T06:31:46Z
No new evidence or uptake changes the interpretation: this remains a directly relevant but uncorroborated paper claim about prior-score contamination in LLM judging. It supports verdict-free evaluation design, but independent replication or implementation evidence is still absent.
2026-08-28T06:27:22Z
grounded: converges/medium — The claimed prior-score anchoring effect directly supports Scott’s implemented “Verdict-free evidence reuse” pattern, which withholds old scores and conclusions
2026-08-28T06:25:22Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49474905 -> echo.paper.34402451c0 by Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, and Emanuel Lacic
2026-08-28T06:24:38Z
case created — The linked research paper presents a specific, testable evaluation failure distinct from general benchmark saturation.
Decision trace
- 08-30 17:27expireThe paper remains a useful empirical lead for verdict-free evaluation design, but no independent validation, implementation uptake, or discussion emerged within the case horizon. The episode has faded
- 08-30 17:27alert_silentThis look contains only unchanged engagement and elapsed staleness, with no consequential new evidence or uptake to interrupt Scott about.
- 08-30 17:27alert_routeThis look contains only unchanged engagement and elapsed staleness, with no consequential new evidence or uptake to interrupt Scott about.
- 08-28 16:31repriceNo new evidence or uptake changes the interpretation: this remains a directly relevant but uncorroborated paper claim about prior-score contamination in LLM judging. It supports verdict-free evaluatio
- 08-28 16:31alert_silentThe paper’s findings were already routed, and this look adds no material evidence, confirmation, or consequential uptake that warrants another interruption.
- 08-28 16:31alert_routeThe paper’s findings were already routed, and this look adds no material evidence, confirmation, or consequential uptake that warrants another interruption.
- 08-28 16:30alert_shadowThe authors report a large, directly implementation-relevant study in which prior-score metadata shifted seven of eight tested judges, blocked 48% of error corrections, and flipped 10.18% of correct c
- 08-28 16:30alert_routeThe authors report a large, directly implementation-relevant study in which prior-score metadata shifted seven of eight tested judges, blocked 48% of error corrections, and flipped 10.18% of correct c
- 08-28 16:27groundThe claimed prior-score anchoring effect directly supports Scott’s implemented “Verdict-free evidence reuse” pattern, which withholds old scores and conclusions from subsequent reviewers, and reinforc
- 08-28 16:25promote_anchororigin walk conf 0.99
- 08-28 16:24createThe linked research paper presents a specific, testable evaluation failure distinct from general benchmark saturation.