A first-person AAAI 2027 reviewer reports Phase 1 advanced an unblinded off-template paper and a visibly incomplete paper while the other 'human' reviews mirrored the AI-generated ones, marking a concrete peer-review quality regression at a major ML venue if corroborated or acknowledged by AAAI.
state: watchingheat: lowuncertainty: mediumconvergesscott: mediumpeer-review-integrity llm-generated-reviews ml-conferencesAAAI
What is this?
AAAI has been moving LLM reviews into its pipeline: the AAAI-26 pilot gave every main-track paper one clearly labeled, non-decisional AI review alongside human reviews (~23k papers processed in under a day at under $1/paper, per the pilot paper's self-report, with survey respondents preferring AI reviews on technical accuracy), and the AAAI-27 CFP continues the same Phase 1 structure of two human reviews plus one identified AI review β with Phase 1 acting as a rejection gate where rejected authors get no response opportunity. The specific claim here β a first-person reviewer reporting that Phase 1 advanced an unblinded, off-template submission and a visibly incomplete one while the other 'human' reviews read as echoes of the AI review β appears nowhere in the supplied search material: the snippets cover only the CFPs, the pilot's self-evaluation, and reviewer-assignment mechanics, and include neither the reviewer's post nor any AAAI acknowledgment. So the incident rests entirely on a single community discussion thread until further reviewer reports or an AAAI response surface. The broader snippets do show venues treating reviewer-integrity as live (IJCAI's CFP explicitly bars reviewers from LLM assessment; AAAI-26 assignment work targeted collusion risk), but that is policy context, not evidence about this report.
Why it matters to Scott
If corroborated, this is the world independently producing exactly the failure pattern Scott's canon argues: human reviews echoing the AI review would make the 'non-decisional' AI review decisional through anchoring β correlated checkers and a violation of mechanically-different-verifiers at a rejection gate β while visibly incomplete and unblinded papers advancing shows gate criteria without authority (no deterministic completeness/blinding check before judgment), i.e. a dated-receipts opportunity for his gate-criteria and LLM-judge pages. But the incident rests on a single uncorroborated community thread, so it's a watch-case: relevance upgrades only if further reviewer reports or an AAAI response land.
ip:concept.correlated-checkers-pitfallip:concept.mechanically-different-verifiersip:framework.gate-criteria-frameworkip:concept.sufficient-coverageradar:llm-judge-prior-score-anchoringradar:concept.llm-judgesradar:concept.verification
queries asked of Scott's wikis
- LLM-as-judge reliability and failure modes
- detecting AI-generated text, provenance and stylometric signals
- pipeline quality gates, validation before downstream stages
- automation offloading, human-in-the-loop degradation
- batch LLM inference economics, cost per document
- peer review and epistemic infrastructure, provenance in knowledge systems
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 400h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
| 09-25 00:09 | β origin directly observed | What's up with AAAI reviewers and organizers? [D] OutsideSimple4854 on r/MachineLearning | β |
| 09-25 11:26 | first on r/MachineLearning Β· published Β· +11.3h | iclr 2027 de anonymization [D] Striking-Warning9533 | β |
| 09-25 00:09 | amplified on r/MachineLearning π | reddit.post.1wphteu OutsideSimple4854 | peak 20 Β· 24 comments Β· 52% of case engagement |
| 09-25 11:26 | amplified on r/MachineLearning | reddit.post.1wptsvx Striking-Warning9533 | peak 29 Β· 6 comments Β· 42% of case engagement |
| 09-30 08:38 | amplified on r/MachineLearning | reddit.post.1wtznk0 Foreign_Tonight_7584 | peak 0 Β· 5 comments Β· 6% of case engagement |
| 09-25 01:20 | our radar first saw it Β· +1.2h | discovery anchor: reddit.post.1wphteu | β |
pace: p65 vs 1032 stories at the 336h mark (now 400h old) β ahead of gfx906-expert-pool-admission-fix (1.0x), behind ordewell-editable-multi-runner-plans (1.0x)
Evidence (3) β β canonical anchor
Interpretation history
2026-09-30T11:37:45Z
The new 'papers written for machines' post adds a third failure mode to the ambient 2027 review-degradation context (style-adaptation feedback loop β conceptually close to the correlated-checkers canon) but is itself weakly received (0 pts, 0.19 ratio) and does not corroborate the AAAI Phase 1 claim, which remains single-source with no AAAI response; the case is drifting toward a pattern-context watch while its specific hypothesis sits still, and the numbers line agrees (~0 pts/h, 25th percentile, 131h old) β low heat, no material reset.
2026-09-30T11:24:57Z
evidence attached: reddit.post.1wtznk0 β Independent firsthand account of AI-parseable paper style and an AI-assisted-reviewing feedback loop, materially contextualising the peer-review degradation episode.
2026-09-25T12:31:39Z
The first-party ICLR 2027 PC-exposure statement makes review-integrity failures a documented multi-venue pattern this cycle, but it is a different failure mode (blinding, not gate quality or anchoring) and does not corroborate the AAAI Phase 1 claim, which still rests on one cooling thread; same-thread comments add anecdotes (hallucinated references, prompt-flavored paired reviews), not independent lines. Graduating seedβwatching reflects a checkable first-person claim tracked against a live sibling-venue context, not corroboration.
2026-09-25T12:24:45Z
evidence attached: reddit.post.1wptsvx β First-party OpenReview statement that ICLR 2027 submissions were exposed to PC members is a second venue's review-integrity failure in the same cycle, indicating peer-review regression spans venues beyond AAAI.
2026-09-25T01:34:53Z
grounded: converges/medium β If corroborated, this is the world independently producing exactly the failure pattern Scott's canon argues: human reviews echoing the AI review would make the
2026-09-25T01:27:14Z
case created β Concrete first-person reviewer evidence of a checkable review-gate failure at a major venue, resolvable by further reviewer reports or an AAAI response.
Decision trace
- 10-05 06:52review_screenOnly delta is a new generic grievance comment (mrpkeya, no new facts β opinion/popularity) and the deletion of fmeneguzzi's truncated process-context comment, which neither corroborated nor contr
- 10-05 06:51review_screenjev screen borderline (noul=0.37) β luna review
- 09-30 21:37repriceThe new 'papers written for machines' post adds a third failure mode to the ambient 2027 review-degradation context (style-adaptation feedback loop β conceptually close to the correlated-che
- 09-30 21:24attachIndependent firsthand account of AI-parseable paper style and an AI-assisted-reviewing feedback loop, materially contextualising the peer-review degradation episode.
- 09-30 21:23propose_attachIndependent firsthand account of AI-parseable paper style and an AI-assisted-reviewing feedback loop, materially contextualising the peer-review degradation episode.
- 09-28 02:38review_screenjev screen: no material development (noul=0.08)
- 09-28 02:21sensor_dirtycomment_update
- 09-26 13:27review_screenjev screen: no material development (noul=0.12)
- 09-26 11:21sensor_dirtycomment_update
- 09-25 22:31repriceThe first-party ICLR 2027 PC-exposure statement makes review-integrity failures a documented multi-venue pattern this cycle, but it is a different failure mode (blinding, not gate quality or anchoring
- 09-25 22:24attachFirst-party OpenReview statement that ICLR 2027 submissions were exposed to PC members is a second venue's review-integrity failure in the same cycle, indicating peer-review regression spans venu
- 09-25 22:23propose_attachFirst-party OpenReview statement that ICLR 2027 submissions were exposed to PC members is a second venue's review-integrity failure in the same cycle, indicating peer-review regression spans venu
- 09-25 13:21sensor_dirtycomment_update
- 09-25 11:34groundIf corroborated, this is the world independently producing exactly the failure pattern Scott's canon argues: human reviews echoing the AI review would make the 'non-decisional' AI revie
- 09-25 11:27createConcrete first-person reviewer evidence of a checkable review-gate failure at a major venue, resolvable by further reviewer reports or an AAAI response.