2026-10-11 17:11 UTC

AIStupidLevel’s developer claims production LLM benchmark performance varies materially across hours and days, making single-point evaluations unreliable for comparing models and APIs.

state: expiredheat: lowuncertainty: highknownscott: mediummodel-evaluation llm-reliability benchmark-stabilityAIStupidLevelionutvi

What is this?

AI Stupid Level is an independent benchmarking platform operated by Studio Platforms in Romania, running real-time, hourly, and daily evaluations of LLM coding, reasoning, tool use, speed, and performance drift. Its developer reports that an analysis of 31,352 hourly scores found 2.80-point variation within a day versus 8.43 points between days, arguing that one-off comparisons of production models or APIs can be misleading. The supplied snippets support the platform’s continuous-evaluation methodology and the broader concern about benchmark variance, but they do not expose the underlying analysis well enough to verify those figures or establish ionutvi’s identity and role.

Why it matters to Scott

Scott already holds the core position in “Nightly AI Decision Builds,” “Drift Monitoring,” and “Trace-backed agent comparison”: production model behavior requires repeated, time-series evaluation rather than one-shot validation. The reported within-day and between-day variance could materially improve his provider benchmarks and routing tests by requiring multi-day sampling, but the supplied evidence does not independently verify the figures or introduce a well-established external actor.
ip:framework.nightly-ai-decision-buildsip:concept.drift-monitoringdev:concept.trace-backed-agent-comparisondev:project.remote-execradar:concept.model-evaluationradar:concept.llm-reliabilityradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • continuous evaluation versus one-shot LLM benchmarks
  • production model drift and API reliability
  • time-series evaluation of hosted model APIs
  • statistical confidence in LLM model comparisons
  • model routing under changing benchmark performance
  • evaluation harnesses for detecting capability regressions

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P]
MachineLearning
ionutvi08
🟧 echo.github ⭐The commit’s calibration note reports: “within-day noise is 2.80pts vs 8.43pts between-day,” measured on 31,352 real hourly scores across 49StudioPlatforms——
🟠 redditI built a free continuous Claude benchmark - Opus 5 currently ranks #1 across 22 active models
ClaudeAI
ionutvi08
🟠 redditSame prompt, same model, ten runs: scores from 0.30 to 0.81. This week my skill-testing tool refused to publish its own results, and I shipped the refusal as the report.
ClaudeAI
maverick_man111108

Interpretation history

Decision trace