AIStupidLevel’s developer claims production LLM benchmark performance varies materially across hours and days, making single-point evaluations unreliable for comparing models and APIs.
state: expiredheat: lowuncertainty: highknownscott: mediummodel-evaluation llm-reliability benchmark-stabilityAIStupidLevelionutvi
What is this?
AI Stupid Level is an independent benchmarking platform operated by Studio Platforms in Romania, running real-time, hourly, and daily evaluations of LLM coding, reasoning, tool use, speed, and performance drift. Its developer reports that an analysis of 31,352 hourly scores found 2.80-point variation within a day versus 8.43 points between days, arguing that one-off comparisons of production models or APIs can be misleading. The supplied snippets support the platform’s continuous-evaluation methodology and the broader concern about benchmark variance, but they do not expose the underlying analysis well enough to verify those figures or establish ionutvi’s identity and role.
Why it matters to Scott
Scott already holds the core position in “Nightly AI Decision Builds,” “Drift Monitoring,” and “Trace-backed agent comparison”: production model behavior requires repeated, time-series evaluation rather than one-shot validation. The reported within-day and between-day variance could materially improve his provider benchmarks and routing tests by requiring multi-day sampling, but the supplied evidence does not independently verify the figures or introduce a well-established external actor.
ip:framework.nightly-ai-decision-buildsip:concept.drift-monitoringdev:concept.trace-backed-agent-comparisondev:project.remote-execradar:concept.model-evaluationradar:concept.llm-reliabilityradar:concept.benchmark-integrity
queries asked of Scott's wikis
- continuous evaluation versus one-shot LLM benchmarks
- production model drift and API reliability
- time-series evaluation of hosted model APIs
- statistical confidence in LLM model comparisons
- model routing under changing benchmark performance
- evaluation harnesses for detecting capability regressions
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (4) — ⭐ canonical anchor
Interpretation history
2026-08-31T20:40:18Z
No new evidence arrived within the case’s horizon, so the specific hour-to-day production drift claim remains unverified and has faded as an active episode. The broader warning against one-shot evaluation remains useful but is already absorbed in Scott’s practice.
2026-08-29T20:27:12Z
Driftproof adds an independent implementation showing that single-draw variance can invalidate an evaluation and justify abstention, strengthening the practical case against one-shot measurements. It does not corroborate AIStupidLevel’s specific hour-to-day production drift claim or resolve whether that pattern reflects provider changes, batching, or sampling noise.
2026-08-29T19:25:08Z
evidence attached: reddit.post.1w1ugn5 — Receipt-backed repeated testing finds scores from 0.30 to 0.81 for the same prompt and model, independently reinforcing concerns about unstable LLM evaluation.
2026-08-29T14:26:52Z
Refreshed discussion adds no independent evidence and chiefly reiterates unresolved batching and stochastic-sampling confounds. The case remains a first-party longitudinal claim whose methodology is not inspectable enough to attribute variance to production API drift.
2026-08-29T13:26:44Z
The attached benchmark post confirms that the developer operates a continuous-evaluation artifact, but it remains the same first-party evidence line rather than independent corroboration. Discussion also sharpens unresolved stochastic-sampling and batching confounds, so the reported variance still cannot be attributed to production-model drift.
2026-08-29T12:23:31Z
evidence attached: reddit.post.1w1k4wq — The continuous benchmark explicitly measures model stability over time, providing a relevant first-party artifact for the temporal-variance hypothesis.
2026-08-29T11:33:23Z
No substantive evidence has arrived beyond negligible discussion activity. The claim remains a potentially useful but self-reported illustration of temporal variance, without inspectable methodology or independent corroboration separating model drift from benchmark and sampling effects.
2026-08-29T11:29:23Z
grounded: known/medium — Scott already holds the core position in “Nightly AI Decision Builds,” “Drift Monitoring,” and “Trace-backed agent comparison”: production model behavior requir
2026-08-29T11:27:28Z
origin walked (codex/luna, conf 0.96): anchor reddit.post.1w1jp1j -> echo.github.da6bb41bda by StudioPlatforms
2026-08-29T11:25:19Z
case created — The reported 31,352-score longitudinal dataset presents a specific, testable claim about temporal instability in production model APIs.
Decision trace
- 09-01 06:40expireNo new evidence arrived within the case’s horizon, so the specific hour-to-day production drift claim remains unverified and has faded as an active episode. The broader warning against one-shot evalua
- 09-01 06:40alert_silentThis look contains only a staleness trigger and no consequential new evidence; there is nothing for Scott to act on or learn before the next briefing.
- 09-01 06:40alert_routeThis look contains only a staleness trigger and no consequential new evidence; there is nothing for Scott to act on or learn before the next briefing.
- 08-30 06:27repriceDriftproof adds an independent implementation showing that single-draw variance can invalidate an evaluation and justify abstention, strengthening the practical case against one-shot measurements. It
- 08-30 06:27alert_silentThe new evidence is useful corroboration of single-run evaluation risk, but it does not establish temporal API drift and does not change Scott’s existing repeated-sampling practice enough to warrant i
- 08-30 06:27alert_routeThe new evidence is useful corroboration of single-run evaluation risk, but it does not establish temporal API drift and does not change Scott’s existing repeated-sampling practice enough to warrant i
- 08-30 06:21sensor_dirtycomment_update
- 08-30 05:25alert_silentA second small open-source evaluator reports that generation-draw variance overwhelmed judge noise and caused its drift report to refuse publication, offering a useful transferable lesson about repeat
- 08-30 05:25surface_candidateA second small open-source evaluator reports that generation-draw variance overwhelmed judge noise and caused its drift report to refuse publication, offering a useful transferable lesson about repeat
- 08-30 05:25alert_routeA second small open-source evaluator reports that generation-draw variance overwhelmed judge noise and caused its drift report to refuse publication, offering a useful transferable lesson about repeat
- 08-30 05:25attachReceipt-backed repeated testing finds scores from 0.30 to 0.81 for the same prompt and model, independently reinforcing concerns about unstable LLM evaluation.
- 08-30 05:23propose_attachReceipt-backed repeated testing finds scores from 0.30 to 0.81 for the same prompt and model, independently reinforcing concerns about unstable LLM evaluation.
- 08-30 00:26repriceRefreshed discussion adds no independent evidence and chiefly reiterates unresolved batching and stochastic-sampling confounds. The case remains a first-party longitudinal claim whose methodology is n
- 08-30 00:26alert_silentThe new delta is only repetitive commentary, with no dataset, controls, replication, or consequential participant; it does not change Scott's existing practice or merit attention before the next
- 08-30 00:26alert_routeThe new delta is only repetitive commentary, with no dataset, controls, replication, or consequential participant; it does not change Scott's existing practice or merit attention before the next
- 08-30 00:21sensor_dirtycomment_update
- 08-30 00:21sensor_dirtycomment_update
- 08-29 23:26repriceThe attached benchmark post confirms that the developer operates a continuous-evaluation artifact, but it remains the same first-party evidence line rather than independent corroboration. Discussion a
- 08-29 23:26alert_silentThe delta adds an implementation artifact but no inspectable dataset, statistical controls, or independent replication; it reinforces Scott's existing time-series evaluation practice without crea
- 08-29 23:26alert_routeThe delta adds an implementation artifact but no inspectable dataset, statistical controls, or independent replication; it reinforces Scott's existing time-series evaluation practice without crea
- 08-29 22:24alert_silentThe new post adds self-reported platform scale and point-in-time model rankings, but no reproducible dataset, methodology artifact, or independent evidence that materially strengthens the temporal-var
- 08-29 22:24alert_routeThe new post adds self-reported platform scale and point-in-time model rankings, but no reproducible dataset, methodology artifact, or independent evidence that materially strengthens the temporal-var
- 08-29 22:23attachThe continuous benchmark explicitly measures model stability over time, providing a relevant first-party artifact for the temporal-variance hypothesis.
- 08-29 22:22propose_attachThe continuous benchmark explicitly measures model stability over time, providing a relevant first-party artifact for the temporal-variance hypothesis.
- 08-29 22:21sensor_dirtycomment_update
- 08-29 21:33repriceNo substantive evidence has arrived beyond negligible discussion activity. The claim remains a potentially useful but self-reported illustration of temporal variance, without inspectable methodology o
- 08-29 21:33alert_silentThe new delta is only one additional comment with no supplied substantive content; it neither validates the analysis nor changes Scott's existing time-series evaluation practice, so normal monito
- 08-29 21:33alert_routeThe new delta is only one additional comment with no supplied substantive content; it neither validates the analysis nor changes Scott's existing time-series evaluation practice, so normal monito
- 08-29 21:32alert_silentA self-reported analysis claims materially greater between-day than within-day variance across 31,352 hourly scores, but the supplied evidence does not expose enough methodology, raw data, task stabil
- 08-29 21:32surface_candidateA self-reported analysis claims materially greater between-day than within-day variance across 31,352 hourly scores, but the supplied evidence does not expose enough methodology, raw data, task stabil
- 08-29 21:32alert_routeA self-reported analysis claims materially greater between-day than within-day variance across 31,352 hourly scores, but the supplied evidence does not expose enough methodology, raw data, task stabil
- 08-29 21:29groundScott already holds the core position in “Nightly AI Decision Builds,” “Drift Monitoring,” and “Trace-backed agent comparison”: production model behavior requires repeated, time-series evaluation rath
- 08-29 21:27promote_anchororigin walk conf 0.96
- 08-29 21:25createThe reported 31,352-score longitudinal dataset presents a specific, testable claim about temporal instability in production model APIs.