Australian GP trainee radeon2000 claims his process-scored, tool-constrained medical consultation game put 13 AI models through 195 consults and every one reached the correct diagnosis β with safety behavior, not diagnostic accuracy, the only separator β a finding that, if replicated or adopted by evaluators, would shift medical-agent differentiation from outcome benchmarks to process/safety scoring.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation medical-ai-safetyradeon2000
What is this?
An Australian GP trainee posting under the handle 'radeon2000' reports building a process-scored, tool-constrained medical consultation game and running 13 AI models through 195 simulated consults; per the post's own title, every consult reached the correct diagnosis, leaving safety behavior as the only differentiator between models. The supplied web results do not surface the original post or any independent coverage, so the claim rests entirely on that first-party artifact. The surrounding coverage establishes context rather than corroboration: Google's AMIE program already evaluates diagnostic dialogue on multi-axis consultation rubrics (history-taking, clinical management, communication, empathy) that go well beyond accuracy, Harvard's CaBot is explicitly built to expose reasoning process rather than just answers, and the study covered by Topol reports frontier diagnostic accuracy already running ahead of physicians-with-AI (92% vs 76%). A direct complication to 'safety as the score' comes from Stanford HAI (Jul 2026): in mental-health chatbot evaluation, expert human raters rarely agree on what counts as safe β which makes process/safety scoring both the plausible next frontier and a methodologically contested one.
Why it matters to Scott
Converges on the Flatline Tell: a 195/195 diagnostic ceiling is outcome metrics stopping differentiation, pushing evaluation into the process/safety dimension Scott's Cognition Dimension Ladder already predicts β an independent domain-expert dated receipt for his saturation argument, landed in the medical vertical adjacent to his dental escalation/duty-of-care canon and built in the same shape as his own trace-backed, fixed-fixture model comparisons. It stays medium because it is one unreplicated, low-traction artifact, but the Stanford rater-disagreement complication makes it feed rather than merely confirm his judge-reliability position β the radar's own omission-blindness episode shows LLM judges miss clinically important omissions, which bears directly on whether safety scoring can serve as the new arbiter.
ip:framework.cognition-dimension-ladderdev:concept.trace-backed-agent-comparisonip:framework.voice-ai-readiness-the-13-pillars-framework-ebookradar:concept.agent-evaluationradar:concept.benchmark-saturationradar:concept.agent-safetyradar:concept.healthcare-airadar:gpt6-luna-puppy-kill-benchradar:yhahn-agent-escalation-ergonomicsradar:ai-benchmark-saturation-distortionradar:llm-judge-omission-blindness
queries asked of Scott's wikis
- process-based evaluation vs outcome benchmarks for agents
- tool-constrained agent harness evals and safety rails
- benchmark saturation β what happens when outcome metrics stop differentiating
- rubric / LLM-as-judge scoring of agent behavior
- domain-expert built DIY evaluation artifacts and eval games
- medical vertical agent product patterns and escalation behavior
Measured heat
now 0 pts/hpeak 20 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 177h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p60 vs 1188 stories at the 168h mark (now 177h old) β ahead of local-kv-cache-pressure-probe (1.0x), behind artificium-covering-design-search (1.0x)
Evidence (2) β β canonical anchor
Interpretation history
2026-10-07T12:28:33Z
Independent corroboration landed: sadra_blog's 900-conversation HealthBench experiment converges on the same separator β escalation/safety behavior, not accuracy, differentiates medical agents, and judge choice flips results β taking the case from an unreplicated seed to a two-source corroborated finding. Traction is near-zero and single-platform (0.33 pts/h vs a 20-pt peak), so it prices as a quiet corroborated case whose next signal is evaluator adoption, not social spread.
2026-10-07T12:26:29Z
evidence attached: reddit.post.1wzuvaz β Independent 900-conversation HealthBench experiment showing escalation/safety behavior, not accuracy, differentiates medical agents and that judge choice flips results β direct supporting evidence for the process/safety-scoring thesis.
2026-10-04T07:34:06Z
grounded: converges/medium β Converges on the Flatline Tell: a 195/195 diagnostic ceiling is outcome metrics stopping differentiation, pushing evaluation into the process/safety dimension S
2026-10-04T07:24:52Z
case created β First-party domain-expert evaluation artifact carrying a crisp, unclaimed finding (diagnosis parity, safety-behavior divergence) in a hot topic, but near-zero traction so it stays a low-heat seed rather than watching.
Decision trace
- 10-08 12:22attention_routeThe editor compared this story and chose to keep watching.
- 10-08 00:17attention_routeFirst delivery of this development, and nothing lands within six hours that would improve it β no replication or evaluator adoption is imminent, so a hold buys nothing β and no action is needed before
- 10-07 23:33attention_routeFirst notification on this development, and what makes tonight the moment is the arrival of a second, independent line (the HealthBench 'be brief' result) turning one hobbyist artifact into
- 10-07 23:28attention_candidatematerial_reprice
- 10-07 23:28repriceIndependent corroboration landed: sadra_blog's 900-conversation HealthBench experiment converges on the same separator β escalation/safety behavior, not accuracy, differentiates medical agents, a
- 10-07 23:26attention_candidateattach
- 10-07 23:26attachIndependent 900-conversation HealthBench experiment showing escalation/safety behavior, not accuracy, differentiates medical agents and that judge choice flips results β direct supporting evidence for
- 10-07 23:24propose_attachIndependent 900-conversation HealthBench experiment showing escalation/safety behavior, not accuracy, differentiates medical agents and that judge choice flips results β direct supporting evidence for
- 10-05 18:25review_screenjev screen: no material development (noul=0.07)
- 10-05 05:21sensor_dirtycomment_update
- 10-04 22:22sensor_dirtycomment_update
- 10-04 18:34groundConverges on the Flatline Tell: a 195/195 diagnostic ceiling is outcome metrics stopping differentiation, pushing evaluation into the process/safety dimension Scott's Cognition Dimension Ladder a
- 10-04 18:24createFirst-party domain-expert evaluation artifact carrying a crisp, unclaimed finding (diagnosis parity, safety-behavior divergence) in a hot topic, but near-zero traction so it stays a low-heat seed rath