2026-10-11 16:38 UTC

A first-person AAAI 2027 reviewer reports Phase 1 advanced an unblinded off-template paper and a visibly incomplete paper while the other 'human' reviews mirrored the AI-generated ones, marking a concrete peer-review quality regression at a major ML venue if corroborated or acknowledged by AAAI.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumpeer-review-integrity llm-generated-reviews ml-conferencesAAAI

What is this?

AAAI has been moving LLM reviews into its pipeline: the AAAI-26 pilot gave every main-track paper one clearly labeled, non-decisional AI review alongside human reviews (~23k papers processed in under a day at under $1/paper, per the pilot paper's self-report, with survey respondents preferring AI reviews on technical accuracy), and the AAAI-27 CFP continues the same Phase 1 structure of two human reviews plus one identified AI review β€” with Phase 1 acting as a rejection gate where rejected authors get no response opportunity. The specific claim here β€” a first-person reviewer reporting that Phase 1 advanced an unblinded, off-template submission and a visibly incomplete one while the other 'human' reviews read as echoes of the AI review β€” appears nowhere in the supplied search material: the snippets cover only the CFPs, the pilot's self-evaluation, and reviewer-assignment mechanics, and include neither the reviewer's post nor any AAAI acknowledgment. So the incident rests entirely on a single community discussion thread until further reviewer reports or an AAAI response surface. The broader snippets do show venues treating reviewer-integrity as live (IJCAI's CFP explicitly bars reviewers from LLM assessment; AAAI-26 assignment work targeted collusion risk), but that is policy context, not evidence about this report.

Why it matters to Scott

If corroborated, this is the world independently producing exactly the failure pattern Scott's canon argues: human reviews echoing the AI review would make the 'non-decisional' AI review decisional through anchoring β€” correlated checkers and a violation of mechanically-different-verifiers at a rejection gate β€” while visibly incomplete and unblinded papers advancing shows gate criteria without authority (no deterministic completeness/blinding check before judgment), i.e. a dated-receipts opportunity for his gate-criteria and LLM-judge pages. But the incident rests on a single uncorroborated community thread, so it's a watch-case: relevance upgrades only if further reviewer reports or an AAAI response land.
ip:concept.correlated-checkers-pitfallip:concept.mechanically-different-verifiersip:framework.gate-criteria-frameworkip:concept.sufficient-coverageradar:llm-judge-prior-score-anchoringradar:concept.llm-judgesradar:concept.verification
queries asked of Scott's wikis
  • LLM-as-judge reliability and failure modes
  • detecting AI-generated text, provenance and stylometric signals
  • pipeline quality gates, validation before downstream stages
  • automation offloading, human-in-the-loop degradation
  • batch LLM inference economics, cost per document
  • peer review and epistemic infrastructure, provenance in knowledge systems

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 400h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-25 00:09⭐ origin directly observedWhat's up with AAAI reviewers and organizers? [D]
OutsideSimple4854 on r/MachineLearning
β€”
09-25 11:26first on r/MachineLearning Β· published Β· +11.3hiclr 2027 de anonymization [D]
Striking-Warning9533
β€”
09-25 00:09amplified on r/MachineLearning πŸ‘‘reddit.post.1wphteu
OutsideSimple4854
peak 20 Β· 24 comments Β· 52% of case engagement
09-25 11:26amplified on r/MachineLearningreddit.post.1wptsvx
Striking-Warning9533
peak 29 Β· 6 comments Β· 42% of case engagement
09-30 08:38amplified on r/MachineLearningreddit.post.1wtznk0
Foreign_Tonight_7584
peak 0 Β· 5 comments Β· 6% of case engagement
09-25 01:20our radar first saw it Β· +1.2hdiscovery anchor: reddit.post.1wphteuβ€”
pace: p65 vs 1032 stories at the 336h mark (now 400h old) β€” ahead of gfx906-expert-pool-admission-fix (1.0x), behind ordewell-editable-multi-runner-plans (1.0x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐What's up with AAAI reviewers and organizers? [D]
MachineLearning
OutsideSimple48542024
🟠 redditiclr 2027 de anonymization [D]
MachineLearning
Striking-Warning9533296
🟠 redditPapers Are Now Written for Machines, Not Humans [D]
MachineLearning
Foreign_Tonight_758405

Interpretation history

Decision trace