Paradigma's linked Limite-1B-Violetto announcement is reported to claim 94% AIME accuracy with one billion parameters, potentially raising the mathematical-reasoning capability available from compact models.
state: expiredheat: lowuncertainty: highnovelscott: lowsmall-model-reasoning model-evaluationParadigma
What is this?
The supplied case describes a Hacker News submission linking an announcement attributed to Paradigma for Limite-1B-Violetto, with the headline claim of 94% AIME accuracy using one billion parameters. None of the supplied web snippets directly covers this model or announcement, so they do not independently establish Paradigma's identity or substantiate the claim. The AIME edition, evaluation protocol, inference budget, parameter accounting, and model availability remain unestablished; the material therefore does not yet demonstrate a compact-model reasoning advance.
Why it matters to Scott
The claim has a potential connection to Scott’s Model Barbell and task-aware model routing, but an unsubstantiated AIME headline with unknown inference budget and availability does not establish a cheaper viable worker or justify changing his model choices. No supplied radar page tracks this specific announcement, and the evidence establishes neither a credible challenge to Scott’s positions nor independent convergence on them.
ip:concept.model-barbelldev:concept.task-aware-model-routingradar:concept.small-modelsradar:concept.model-evaluation
queries asked of Scott's wikis
- compact reasoning models local inference economics
- benchmark validity contamination reproducible evaluation
- test-time compute reasoning accuracy cost tradeoffs
- specialist small models versus general-purpose agents
- model selection evaluation harnesses deployment thresholds
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-09-23T23:22:58Z
The second submission (Limite 1BVioletto) landed with zero comments and the first stalled at 2 points plus a title-editorialization complaint; after ~51h there is no independent coverage, no evaluation protocol, no weights, and no sign the claim will receive scrutiny. The episode faded without ever substantiating the 94% AIME headline.
2026-09-22T17:28:44Z
evidence attached: hn.story.49803871 — shared external link with case evidence
2026-09-21T20:57:49Z
grounded: novel/low — The claim has a potential connection to Scott’s Model Barbell and task-aware model routing, but an unsubstantiated AIME headline with unknown inference budget a
2026-09-21T20:53:01Z
case created — The named announcement and unusually strong numerical claim warrant a narrow seed without assuming benchmark comparability or weight availability.
Decision trace
- 09-24 09:22expireThe second submission (Limite 1BVioletto) landed with zero comments and the first stalled at 2 points plus a title-editorialization complaint; after ~51h there is no independent coverage, no evaluatio
- 09-23 03:28attachshared external link with case evidence
- 09-23 03:22propose_attachshared external link with case evidence
- 09-22 06:57groundThe claim has a potential connection to Scott’s Model Barbell and task-aware model routing, but an unsubstantiated AIME headline with unknown inference budget and availability does not establish a che
- 09-22 06:53createThe named announcement and unusually strong numerical claim warrant a narrow seed without assuming benchmark comparability or weight availability.