2026-10-11 17:11 UTC

AI Stupid Level founder ionutvi claims observations from 31,352 repeated benchmark measurements show that static LLM scores can miss performance changes over time, potentially requiring ongoing evaluation rather than treating API model names as stable reliability guarantees.

state: resolvedheat: lowuncertainty: lowknownscott: lowllm-evaluation model-drift api-reliabilityionutviAI Stupid Level

What is this?

AI Stupid Level describes itself on Hugging Face as an independent benchmarking project built by Ionut Visan and the Studio Platforms ecosystem in Romania, focused on LLM evaluation, model drift, and real-time benchmarks. The case attributes to founder ionutvi a claim that 31,352 repeated benchmark measurements reveal performance changes that static scores can miss. The supplied web snippets establish the project's identity and focus, but do not verify that measurement count, its methodology, the reported findings, or the link between the handle and Visan; general evaluation articles discuss ongoing reliability checks but do not corroborate this particular study.

Why it matters to Scott

The radar already tracks AIStupidLevel’s same temporal-variance claim in radar:production-llm-temporal-variance, and ongoing baseline-relative evaluation is already explicit in Scott’s Drift Monitoring and Nightly AI Decision Builds pages. The claimed 31,352 measurements remain unverified in the supplied material, so this adds no established finding or actionable method that would change his evaluation practice.
ip:concept.drift-monitoringip:framework.nightly-ai-decision-buildsradar:production-llm-temporal-variance
queries asked of Scott's wikis
  • continuous LLM evaluation regression monitoring
  • API model version stability provider drift
  • coding agent harness repeatability reliability tests
  • static benchmark scores versus production reliability
  • repeated evaluations statistical variance drift detection

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-07 07:44⭐ origin directly observedMeasuring LLM performance drift: observations and methodology from 31,352 repeated benchmark measurements [D]
ionutvi on r/MachineLearning
—
09-07 07:49first on r/artificial · published · +0.1hStatic LLM benchmarks can miss performance changes over time - observations from 31,352 repeated measurements
ionutvi
—
09-07 07:57first on r/OpenAI · published · +0.2hHow should developers detect LLM performance drift when the API model name stays the same?
ionutvi
—
09-07 18:44first on r/LocalLLaMA · published · +11.0hArtificial Analysis Intelligence Index v4.3
SteppenAxolotl
—
09-07 20:35first on r/singularity · published · +12.8hArtificial Analysis updates its Intelligence Index to version 4.3
Profanion
—
09-08 21:08first on r/ClaudeAI · published · +37.4hHas Claude suddenly started contradicting itself, saying the opposite of what it means and just generally being awful for you too?
nerbertsalad
—
09-09 03:17first on hacker news · published · +43.5hOpenAI's AGI number came from a harness, not the model
mjshashank
—
09-20 14:31first on r/MachineLearning · published · +318.8hWhy decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]
NoahPersaud
—
09-07 07:44amplified on r/MachineLearningreddit.post.1w9llr4
ionutvi
peak 2 · 1 comments · 0% of case engagement
09-07 07:49amplified on r/artificialreddit.post.1w9lon7
ionutvi
peak 1 · 1 comments · 0% of case engagement
09-07 18:44amplified on r/LocalLLaMAreddit.post.1wa0nrk
SteppenAxolotl
peak 0 · 43 comments · 2% of case engagement
09-07 20:35amplified on r/singularityreddit.post.1wa3o5j
Profanion
peak 170 · 42 comments · 9% of case engagement
09-08 15:11amplified on r/LocalLLaMAreddit.post.1war50q
Tall_Abrocoma_3533
peak 7 · 18 comments · 1% of case engagement
09-08 15:26amplified on r/LocalLLaMAreddit.post.1warjmn
Tall_Abrocoma_3533
peak 34 · 45 comments · 3% of case engagement
15 more amplifiers in ainews.case_chain
09-07 08:20our radar first saw it · +0.6hdiscovery anchor: reddit.post.1w9llr4—

Evidence (23) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Measuring LLM performance drift: observations and methodology from 31,352 repeated benchmark measurements [D]
MachineLearning
ionutvi21
🟠 redditHow should developers detect LLM performance drift when the API model name stays the same?
OpenAI
ionutvi00
🟠 redditStatic LLM benchmarks can miss performance changes over time - observations from 31,352 repeated measurements
artificial
ionutvi11
🟠 redditArtificial Analysis Intelligence Index v4.3
LocalLLaMA
SteppenAxolotl043
🟠 redditArtificial Analysis updates its Intelligence Index to version 4.3
singularity
Profanion17042
🟠 redditAA updated yet again, shaking up the rankings.
LocalLLaMA
Tall_Abrocoma_3533617
🟠 redditAA updated yet again, here's how the frontier ranks.
LocalLLaMA
Tall_Abrocoma_35333445
🟠 redditAstra is a quiet force. Artificial Analysis has updated their benchmark twice in 4 days to reflect its real strength
OpenAI
py-net1211
🟠 redditHas Claude suddenly started contradicting itself, saying the opposite of what it means and just generally being awful for you too?
ClaudeAI
nerbertsalad1318
🟧 hnOpenAI's AGI number came from a harness, not the modelmjshashank10
🟧 hnMy LLM eval cried wolf. Here's what I measuredalexpran60
🟧 hnThe expensive model was cheaper: six other things I got wrong on an LLM judgeAdamWendrich20
🟠 redditDoesn't it seem strange to you that AA changes twice in one week?
LocalLLaMA
Altruistic_Plate10905624
🟠 redditAfter Astra's stealth nerf last night, we really need benchmarks to do a re-bench 1 week after any model release. This is ridiculous.
singularity
Flope828178
🟠 redditArtificial Analysis is not "broken", and they prove it.
LocalLLaMA
Antblue234167
🟠 redditllm performance community metric
LocalLLaMA
AleksandrNikitin57
🟧 hnBad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tiresluu5646
🟧 hnShow HN: TrackLLM: Are the LLM APIs you rely on stable?timotheechauvin11
🟧 hnShow HN: How Stale Is Your AI? Release age and training cutoff for 20 modelsjoozio8450
🟠 redditBenchmarks are misleading?
LocalLLaMA
octagoncat23020
🟠 redditSpent two days convinced our prompt had a bug. Prompt hadn't changed in months.
OpenAI
ClickOk581174
🟠 redditWhy decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]
MachineLearning
NoahPersaud03
🟠 redditThe AI you test in the afternoon may not be the AI you test at night, even with the same name
artificial
FishingCharming5604835

Interpretation history

Decision trace