2026-10-11 16:38 UTC

Australian GP trainee radeon2000 claims his process-scored, tool-constrained medical consultation game put 13 AI models through 195 consults and every one reached the correct diagnosis β€” with safety behavior, not diagnostic accuracy, the only separator β€” a finding that, if replicated or adopted by evaluators, would shift medical-agent differentiation from outcome benchmarks to process/safety scoring.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation medical-ai-safetyradeon2000

What is this?

An Australian GP trainee posting under the handle 'radeon2000' reports building a process-scored, tool-constrained medical consultation game and running 13 AI models through 195 simulated consults; per the post's own title, every consult reached the correct diagnosis, leaving safety behavior as the only differentiator between models. The supplied web results do not surface the original post or any independent coverage, so the claim rests entirely on that first-party artifact. The surrounding coverage establishes context rather than corroboration: Google's AMIE program already evaluates diagnostic dialogue on multi-axis consultation rubrics (history-taking, clinical management, communication, empathy) that go well beyond accuracy, Harvard's CaBot is explicitly built to expose reasoning process rather than just answers, and the study covered by Topol reports frontier diagnostic accuracy already running ahead of physicians-with-AI (92% vs 76%). A direct complication to 'safety as the score' comes from Stanford HAI (Jul 2026): in mental-health chatbot evaluation, expert human raters rarely agree on what counts as safe β€” which makes process/safety scoring both the plausible next frontier and a methodologically contested one.

Why it matters to Scott

Converges on the Flatline Tell: a 195/195 diagnostic ceiling is outcome metrics stopping differentiation, pushing evaluation into the process/safety dimension Scott's Cognition Dimension Ladder already predicts β€” an independent domain-expert dated receipt for his saturation argument, landed in the medical vertical adjacent to his dental escalation/duty-of-care canon and built in the same shape as his own trace-backed, fixed-fixture model comparisons. It stays medium because it is one unreplicated, low-traction artifact, but the Stanford rater-disagreement complication makes it feed rather than merely confirm his judge-reliability position β€” the radar's own omission-blindness episode shows LLM judges miss clinically important omissions, which bears directly on whether safety scoring can serve as the new arbiter.
ip:framework.cognition-dimension-ladderdev:concept.trace-backed-agent-comparisonip:framework.voice-ai-readiness-the-13-pillars-framework-ebookradar:concept.agent-evaluationradar:concept.benchmark-saturationradar:concept.agent-safetyradar:concept.healthcare-airadar:gpt6-luna-puppy-kill-benchradar:yhahn-agent-escalation-ergonomicsradar:ai-benchmark-saturation-distortionradar:llm-judge-omission-blindness
queries asked of Scott's wikis
  • process-based evaluation vs outcome benchmarks for agents
  • tool-constrained agent harness evals and safety rails
  • benchmark saturation β€” what happens when outcome metrics stop differentiating
  • rubric / LLM-as-judge scoring of agent behavior
  • domain-expert built DIY evaluation artifacts and eval games
  • medical vertical agent product patterns and escalation behavior

Measured heat

now 0 pts/hpeak 20 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 177h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-04 06:44⭐ origin directly observedI made 13 AI models play the doctor in my medical consultation game. All 195 consults got the diagnosis right; what separated them was safety.
radeon2000 on r/artificial
β€”
10-07 12:10first on r/ClaudeAI Β· published Β· +77.4hAsking Sonnet 5 to β€œbe brief” reduced both necessary and unnecessary medical follow-up questions
sadra_blog
β€”
10-04 06:44amplified on r/artificial πŸ‘‘reddit.post.1wx8vyi
radeon2000
peak 34 Β· 12 comments Β· 87% of case engagement
10-07 12:10amplified on r/ClaudeAIreddit.post.1wzuvaz
sadra_blog
peak 3 Β· 4 comments Β· 13% of case engagement
10-04 07:20our radar first saw it Β· +0.6hdiscovery anchor: reddit.post.1wx8vyiβ€”
pace: p60 vs 1188 stories at the 168h mark (now 177h old) β€” ahead of local-kv-cache-pressure-probe (1.0x), behind artificium-covering-design-search (1.0x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I made 13 AI models play the doctor in my medical consultation game. All 195 consults got the diagnosis right; what separated them was safety.
artificial
Retrieved article excerpt

Open article Β· Retrieved 2026-10-04T07:23:46.980012+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
radeon20003312
🟠 redditAsking Sonnet 5 to β€œbe brief” reduced both necessary and unnecessary medical follow-up questions
ClaudeAI
sadra_blog34

Interpretation history

Decision trace