2026-10-11 16:38 UTC

Viridian researcher Derry Mitchell claims a frozen 20-pair replication found GPT-5.6 Sol pairwise judging preserves the underlying winner through exact order reversal in 19/20 comparisons โ€” 'not sufficient evidence for a systematic or directional position bias' โ€” against the widely assumed LLM-judge position-bias failure mode, and larger-scale replication or methodological rebuttal settles whether position bias remains a live eval-design concern under current judge configurations.

state: seedheat: lowuncertainty: mediumnovelscott: mediumllm-judge-reliability agent-evaluationDerry MitchellViridian

What is this?

The case centers on a claim by Viridian researcher Derry Mitchell (Technical Note 001) that a frozen 20-pair replication using GPT-5.6 Sol as a pairwise judge found the underlying winner preserved in 19/20 comparisons under exact order reversal โ€” interpreted as 'not sufficient evidence for a systematic or directional position bias.' This challenges the widely assumed LLM-judge position-bias failure mode. The web search returned no direct coverage of this Technical Note, Viridian, or Derry Mitchell; results were dominated by unrelated statistical-methods content and a likely-hallucinated 'GPT-5.6 Sol' statistics claim. The primary evidence remains the case's own cited artifact (Technical Note 001), which has not been independently verified in the supplied snippets.

Why it matters to Scott

The case is a first-party, artifact-pinned replication (20 frozen pairs, GPT-5.6 Sol) that fails to detect the widely assumed position-bias failure mode in pairwise LLM judging. Scott's canon does not hold a position on position bias specifically โ€” his load-bearing claims concern correlated-checker pitfalls and the need for mechanically different verifiers, not order effects. However, the finding directly bears on evaluation harnesses he actively builds (Synthetic Futures judge/repair gates, LLM rubric grading, brief A/B testing): if position bias is negligible for current frontier judges, order-counterbalancing and position-aware aggregation become lower priority, simplifying the verification loops he advocates. The actor (Derry Mitchell) and the concept (llm-judges) are already tracked in the radar.
ip:concept.correlated-checkers-pitfallip:concept.mechanically-different-verifiersdev:concept.llm-rubric-gradingdev:project.synthetic-futuresradar:person.derry-mitchell-viridian-ai-quality-assuranceradar:concept.llm-judgesradar:llm-judge-prior-score-anchoringradar:llm-judge-omission-blindnessradar:concept.reproducibility
queries asked of Scott's wikis
  • llm-judge position bias evaluation methodology
  • agent evaluation frameworks pairwise judging reliability
  • eval design patterns position bias mitigation
  • llm-as-judge systematic error modes
  • replication crisis llm evaluation benchmarks
  • Viridian research org technical notes

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 171h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-04 13:00โญ origin echo-reconstructedTechnical Note 001, 'We Tried to Reproduce LLM Judge Position Bias. The Evidence Wasn't Strong Enough': same-order control agreed 20/20, the
Derry Mitchell (Viridian AI Quality Assurance) on github (echo) ยท attributed from hn.story.49990944
โ€”
10-07 10:51first on hacker news ยท published ยท +69.9hWe tried to reproduce LLM judge position bias
viridianassuran
โ€”
10-07 10:51amplified on hacker news ๐Ÿ‘‘hn.story.49990944
viridianassuran
peak 1 ยท 0 comments ยท 106% of case engagement
10-07 11:20our radar first saw it ยท +70.3hdiscovery anchor: hn.story.49990944โ€”
pace: p8 vs 1188 stories at the 168h mark (now 171h old) โ€” behind addom-local-coding-harness (0.5x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnWe tried to reproduce LLM judge position bias
Retrieved article excerpt

Open article ยท Retrieved 2026-10-07T11:24:20.143215+00:00

# Technical Note 001 โ€” Position Sensitivity in Pairwise LLM Judging

**We Tried to Reproduce LLM Judge Position Bias. The Evidence Wasn't Strong Enough.**

A controlled replication and experimental-assurance study of response-order sensitivity in pairwise LLM judging.

Author: Derry Mitchell  
Viridian  
5 October 2026  
Status: revised technical manuscript for public release; not peer reviewed.

[View the v1 GitHub release](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/releases/tag/study-001-v1)

## Research question

Does reversing the presentation order of two frozen candidate responses change the preferred underlying response under the tested GPT-5.6 Sol judge configuration?

The study uses a frozen 20-prompt response set. Two local candidate models generated one response per prompt, all forty response strings were frozen and hashed, and the same response pairs were then judged in an original orientation and an exact A/B reversal. A same-order repeat control was added before any judge outcome was observed.

## Result

- Same-order control agreement with Pass 1: **20/20 (100%)**
- Same underlying winner after exact reversal: **19/20 (95%)**
- Underlying winner changed after reversal: **1/20 (5%)**
- First-position patterns: **0/20**
- Second-position patterns: **1/20**
- Across the 40 formal orientation judgments, first position won **19** times and second position won **21** times.

The single changed pair was **P007**. The judge selected surface position B in both orientations; because the candidate strings were swapped, the preferred underlying response changed. This is a genuine second-position pattern in the paired record, but it is not sufficient evidence for a systematic or directional position bias.

## Maximum defensible claim

> Under the tested GPT-5.6 Sol judging configuration and frozen 20-prompt response set, reversing candidate presentation order preserved the underlying preferred response in 19 of 20 comparisons. One comparison exhibited a second-position preference pattern. This small study does not establish a systematic or directional position bias.

The evidence also does not establish the opposite claim that GPT-5.6 Sol is generally position-robust.

## Statistical boundary

The observed reversal switch rate was **1/20 = 5%**. The exact 95% Clopper-Pearson interval was approximately **0.1%โ€“24.9%**, which is too wide to support a precise population-level estimate. The same-order control reproduced all 20 Pass 1 winner labels, but one repeat per prompt is not enough to estimate the full within-condition variability of the judge.

## Experimental identity

Candidate models:

- Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B โ€” architecture Qwen3\_5ForConditionalGeneration
- K2 Horizon 3.7B โ€” architecture K2HorizonForCausalLM

Candidate generation used local model-native chat templates, greedy decoding, max\_new\_tokens=256, BF16 with bitsandbytes 8-bit quantisation, SDPA attention, and device\_map=auto.

The evaluator was GPT-5.6 Sol through ChatGPT using the frozen pairwise rubric. The judge-facing packet exposed only the evaluation identifier, user prompt, Response A, and Response B.

## Important limitations

- Only 20 prompt pairs were tested.
- Only one same-order repeat and one reversed judgment were obtained per prompt.
- The judge was accessed through ChatGPT rather than a pinned API snapshot with an exposed deterministic serving seed.
- Several candidate outputs were truncated by the 256-token generation ceiling.
- Some candidate outputs exposed meta-reasoning text.
- Candidate generation was not designed to support a general model-capability comparison.

These limitations constrain the claims they affect; they do not invalidate the exact response-order contrast because the same frozen response strings were reused across orientations.

## Osmium assurance execution

After the evidence package was frozen, Viridian Osmium evaluated the study using its qualified position-sensitivity procedure and returned **PARTIALLY\_SUPPORTED**. That execution preserved the same maximum defensible claim.

This is a reproducibility/product-behaviour check on Viridian's assurance implementation, **not** an independent scientific replication and not additional evidence about GPT-5.6 Sol.

## Study materials

- [Frozen judge rubric](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/blob/main/studies/001-position-sensitivity/judge_rubric.md)
- [Selected artifact identities](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/blob/main/studies/001-position-sensitivity/integrity.md)
- [Release notes](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/blob/main/studies/001-position-sensitivity/RELEASE_NOTES.md)
- [Technical Note 002 โ€” Evaluator Contract Stability](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/blob/main/studies/002-evaluator-contract-stability)
- [Technical Note 003 โ€” Presentation / Serialization Sensitivity](https://github.com/ViridianAIQualityAssurance/viridian-assurance-research/blob/main/studies/003-presentation-serialization-sensitivity)

The source manuscript is Viridian Technical Note 001 dated 5 October 2026. The original DOCX is intentionally not included in this repository.
viridianassuran10
๐ŸŸง echo.github โญTechnical Note 001, 'We Tried to Reproduce LLM Judge Position Bias. The Evidence Wasn't Strong Enough': same-order control agreed 20/20, theDerry Mitchell (Viridian AI Quality Assurance)โ€”โ€”

Interpretation history

Decision trace