AI Stupid Level describes itself on Hugging Face as an independent benchmarking project built by Ionut Visan and the Studio Platforms ecosystem in Romania, focused on LLM evaluation, model drift, and real-time benchmarks. The case attributes to founder ionutvi a claim that 31,352 repeated benchmark measurements reveal performance changes that static scores can miss. The supplied web snippets establish the project's identity and focus, but do not verify that measurement count, its methodology, the reported findings, or the link between the handle and Visan; general evaluation articles discuss ongoing reliability checks but do not corroborate this particular study.
2026-09-24T02:28:50Z
The case's premise — that API model names are not stable reliability guarantees and can drift behind identical labels — has moved from an unverified founder claim to broadly echoed community consensus: the Astra stealth-nerf wave (828-point thread plus X spread), multiple independent same-name drift anecdotes, and dedicated monitoring tooling (TrackLLM) all treat silent behind-the-name change as real. But that premise is already established practice in Scott's drift-monitoring and nightly-build pages and is carried by the sibling radar case on temporal variance, so this episode is absorbed with no distinct actionable finding left here; ionutvi's specific 31,352-measurement methodology remains unverified and quixotic.
2026-09-23T21:46:36Z
evidence attached: reddit.post.1wof7oy — Independent report of a same-named model silently changing behaviour within hours, directly supporting the case's premise that model names are not stable reliability guarantees.
2026-09-20T15:33:17Z
The new contamination critique concerns benchmark validity, not longitudinal changes behind stable API endpoints, and does not corroborate AI Stupid Level’s measurements. The spread signal reflects accumulated discussion across several distinct evaluation issues rather than an expanding temporal-drift episode, so attention remains low.
2026-09-20T15:22:21Z
evidence attached: reddit.post.1wlimaj — The post adds substantive evaluation criticism by arguing that self-reported decontamination cannot independently establish benchmark validity.
2026-09-17T21:40:49Z
The new classification-task anecdote is more directly relevant to temporal degradation, but the supplied excerpt establishes only reported quality decline with an unchanged prompt—not a stable endpoint or controlled evidence of provider-side drift. It does not independently validate the founder’s measurements or change Scott’s evaluation decisions.
2026-09-17T21:22:00Z
evidence attached: reddit.post.1wj5hp5 — The report gives additional user evidence that unchanged model endpoints can drift on edge-case behavior, directly supporting the open case about ongoing reliability measurement.
2026-09-16T17:56:30Z
The new Qwen-versus-Muse report concerns workload-specific differences between models, not changes in the same model over time; it supplies neither reproducible tests nor controlled configurations. It does not corroborate the longitudinal-drift claim or change the operational implications for Scott.
2026-09-16T17:24:40Z
evidence attached: reddit.post.1wi1zjk — This user report supplies workload-level anecdotal evidence that leaderboard scores can diverge from long-context and multi-step behavior, materially contextualising the benchmark-drift case.
2026-09-16T14:38:05Z
The newly attached release-age and training-cutoff tracker concerns knowledge freshness, not measured behavioral drift behind a stable API model name. Its headline and discussion neither establish a validated implementation nor corroborate AI Stupid Level’s longitudinal findings; the attachment rationale overstates its support.
2026-09-16T14:22:54Z
evidence attached: hn.story.49726343 — A usable tool tracking model release age and training cutoffs materially supports the case that model identity and static benchmark scores are insufficient freshness signals.
2026-09-15T07:32:45Z
The expanded benchmark discussion adds practitioner concerns about scoring and task representativeness, not controlled evidence of temporal API drift. It neither corroborates AI Stupid Level’s measurements nor changes the operational implications for Scott.
2026-09-14T07:21:54Z
TrackLLM adds a separate author’s announcement about API-stability monitoring, but the supplied headline and authorship comment establish neither a usable implementation nor measured drift. It is a potentially relevant evidence lead, not independent validation of AI Stupid Level’s findings or a reason to change Scott’s evaluation practice.
2026-09-14T07:21:33Z
evidence attached: hn.story.49692798 — A released tool for monitoring LLM API stability directly supports the case's hypothesis that API model behavior drifts and requires ongoing evaluation rather than treating model names as stable.
2026-09-14T06:24:44Z
The latest attachment supplies only a benchmark-criticism headline and an incidental comment, not the essay’s methodology or any longitudinal measurements. It does not independently corroborate the founder’s drift claim or change Scott’s evaluation decisions.
2026-09-14T06:21:34Z
evidence attached: hn.story.49655621 — The essay provides independent methodological criticism supporting the case that benchmark scores can be misleading and require stronger real-task evaluation.
2026-09-13T13:26:52Z
The new discussion concerns workload-sensitive throughput metrics, not measured performance drift under a stable API model name. It adds no validation of the founder’s study or operational finding that changes Scott’s existing evaluation practice.
2026-09-13T13:21:36Z
evidence attached: reddit.post.1wf6w41 — Suggests task-complexity-per-time is a more meaningful and workload-sensitive metric than raw token speed, materially informing the evaluation hypothesis.
2026-09-11T01:30:07Z
The Astra complaint adds a specific allegation of post-release degradation, but provides no matched before-and-after tests or verified serving change; claimed widespread agreement does not validate the founder’s study. The Artificial Analysis discussion still concerns benchmark construction rather than fixed-model drift, leaving Scott without a new operational finding.
2026-09-10T23:22:51Z
evidence attached: reddit.post.1wcxxm8 — The discussion and comments provide relevant context on benchmark reweighting and whether model rankings remain stable over time.
2026-09-10T20:23:43Z
evidence attached: reddit.post.1wcu0fj — Independent user reports of Astra capability degradation materially support the open case that static model benchmarks can miss post-release behavior changes.
2026-09-10T18:24:04Z
evidence attached: reddit.post.1wcqpsk — Reports repeated changes to Artificial Analysis scoring and growing private-test weighting, materially contextualizing concerns about benchmark stability and comparability.
2026-09-10T00:25:40Z
The two new HN titles concern evaluation false alarms and LLM-judge economics, but supply no findings or methods; they do not independently validate the founder’s temporal-drift study. This broadens the surrounding evaluation discussion without establishing a fixed-model regression or a new operational decision for Scott.
2026-09-09T14:24:00Z
evidence attached: hn.story.49626461 — The account of cost and methodology errors in an LLM judge materially contextualizes the open case on evaluation instability.
2026-09-09T14:24:00Z
evidence attached: hn.story.49626606 — Reported measurements of an unreliable LLM evaluation reinforce the open case that static or flawed evals can misrepresent model performance.
2026-09-09T09:28:14Z
Refreshed discussion adds no controlled longitudinal results or independent validation of the founder’s study. Local-model complaints remain confounded by runtime configuration, while benchmark revisions and harness-dependent scores address evaluation setup rather than drift behind an unchanged API model name.
2026-09-09T03:25:51Z
The new HN title concerns harness-dependent scores, not longitudinal changes to an unchanged API model, and supplies no underlying findings that validate the founder’s study. It appears to repeat an already-alerted claim; neither it nor the refreshed Claude anecdotes establishes a new regression or changes Scott’s evaluation decisions.
2026-09-09T03:22:19Z
evidence attached: hn.story.49620445 — The report that OpenAI's headline score depended heavily on a harness materially reinforces the case that benchmark numbers reflect evaluation setups as well as model capability.
2026-09-08T23:26:04Z
The Claude discussion adds anecdotal reports of contradictory explanations, including one where the requested action was performed correctly, but no controlled before-and-after evidence of model drift. Further benchmark-ranking commentary concerns measurement changes rather than fixed-model deterioration; neither validates the founder’s study or changes Scott’s evaluation decisions.
2026-09-08T21:22:50Z
evidence attached: reddit.post.1wb17wn — The anecdote is weak evidence of apparent model-behavior drift, though it lacks controlled measurements.
2026-09-08T17:23:23Z
evidence attached: reddit.post.1wathnm — The reported benchmark revisions provide independent contextual evidence that static evaluations can lag meaningful model capability changes.
2026-09-08T16:23:24Z
evidence attached: reddit.post.1warjmn — Another frontier ranking snapshot illustrates why static benchmark tables may not remain reliable, though it offers no independent measurement.
2026-09-08T15:38:21Z
The new local-model complaints concern benchmark-to-practice mismatch, with runtime configuration unresolved, rather than deterioration of a fixed API model over time. They do not validate the founder’s longitudinal study or establish a reproducible regression that would change Scott’s evaluation practice.
2026-09-08T15:23:05Z
evidence attached: reddit.post.1war50q — The report and comment provide weak anecdotal evidence that frequently updated rankings can diverge from observed model behavior.
2026-09-08T07:34:52Z
The refreshed discussion adds a ranking observation under the revised index, not evidence of performance drift in an unchanged API model. The founder’s longitudinal study remains unvalidated; further ranking commentary does not change Scott’s evaluation decisions.
2026-09-08T05:25:31Z
Refreshed comments remain speculation about benchmark rankings and incentives, not validation of the temporal-drift study. Changes to an evaluation suite still do not establish changes to a fixed API model; the case needs controlled longitudinal results or independent replication to gain operational meaning for Scott.
2026-09-08T01:25:47Z
Refreshed discussion adds speculation about benchmark rankings and incentives, not reproducible findings or validation of the founder’s temporal-drift study. Benchmark-suite changes remain distinct from fixed-model performance drift, leaving this case without a new operational implication for Scott.
2026-09-07T21:31:37Z
The second report of Artificial Analysis’s index revision repeats a benchmark-methodology change; it does not supply independent evidence that a fixed API model deteriorated over time. The original measurement claim remains unvalidated, and the added discussion establishes neither a reproducible regression nor a new operational decision for Scott.
2026-09-07T21:22:56Z
evidence attached: reddit.post.1wa3o5j — The benchmark-suite revision is independent evidence that model evaluations and rankings are changing enough to require scrutiny of measurement stability.
2026-09-07T20:40:05Z
The attached Artificial Analysis update concerns changes to the benchmark itself, not changes to a fixed API model, so it does not independently corroborate AI Stupid Level’s temporal-drift claim. Refreshed comments add suspicions about rankings rather than verifiable findings; the original study remains unvalidated and offers no new operational decision for Scott.
2026-09-07T19:22:50Z
evidence attached: reddit.post.1wa0nrk — The revised index adds newer agentic benchmarks and private tests, materially contextualizing the case for evolving evaluations over static scores.
2026-09-07T08:26:31Z
This recheck adds no substantive evidence: the three same-author posts remain one unvalidated study claim, not independent corroboration of operationally meaningful drift. The case still reinforces Scott’s existing monitoring practice without establishing a new regression or actionable evaluation method.
2026-09-07T08:25:46Z
grounded: known/low — The radar already tracks AIStupidLevel’s same temporal-variance claim in radar:production-llm-temporal-variance, and ongoing baseline-relative evaluation is alr
2026-09-07T08:23:01Z
case created — The reported measurement study is a bounded evaluation episode distinct from endpoint outages, but these same-author posts provide neither independent corroboration nor visible results establishing material drift or tool-use degradation.