2026-10-11 17:12 UTC

Basalt Labs' claimed 99.44% HLE result will be shown to rely on a misleading model identity, evaluation setup, or served-model substitution.

state: expiredheat: lowuncertainty: highconvergesscott: mediumbenchmark-integrity model-identity ai-scams hleBasalt LabsDeepSeekQwen

What is this?

Basalt Labs is alleged to have claimed a 99.44% score on Humanity’s Last Exam (HLE) using tools; HLE is a 3,000-question, multimodal academic benchmark whose reported frontier scores remain below 60% in the supplied results. The evidence title further alleges that Basalt released a Qwen2.5-7B-Instruct-based model while serving DeepSeek on its website, raising possible model-substitution and evaluation-integrity concerns. However, the supplied web snippets neither independently document Basalt Labs, its claimed result, nor the alleged identity mismatch, so the scam hypothesis remains unverified here.

Why it matters to Scott

If verified, the alleged served-model substitution would reinforce Scott’s requirement to declare model versions, preserve evaluation/production parity, and audit real execution rather than trust vendor demos. It also bears directly on his provider-side execution benchmark and multi-provider gateways by suggesting that served-model identity should become an explicit verification control, but the allegation is currently uncorroborated.
ip:framework.12-factor-agents-frameworkip:concept.capability-auditdev:project.remote-execdev:technology.litellmdev:technology.openrouter
queries asked of Scott's wikis
  • benchmark integrity and reproducible model evaluations
  • served-model identity and provenance verification
  • benchmark contamination tool use and evaluation leakage
  • AI product claims audit and scam detection
  • model substitution in hosted inference APIs
  • HLE benchmark limitations and evaluation methodology

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐"Basalt Labs" pulling a generationally dumb scam. Incredibly stupid lmao. Claiming 99.44% on HLE with tools. Model they released is based on Qwen2.5-7B-Instruct and the model they're serving on their website is DeepSeek.
LocalLLaMA
WithoutReason172928587

Interpretation history

Decision trace