2026-10-11 16:37 UTC

llm-judges

band: coolmomentum: stable score: 0.006
temperature history

Episodes (2)

Krishna Balasubramanian, Sasha Podkopaev, and Shiva Kasiviswanathan claim their unsupervised Ising-model aggregator improves ten-judge panel accuracy by 9–14% over weighted voting on three tasks, enabling better automated evaluation without human reference labels for training.
seedconvergesscott: medium
The paper’s authors claim that exposing LLM judges to prior scores materially anchors their ratings and can distort model comparisons based on those evaluations.
expiredconvergesscott: medium