Independent evaluations will determine whether Mistral's Leanstral materially improves Lean proof development and formal-verification workflows over general-purpose reasoning models.
state: expiredheat: lowuncertainty: highconvergesscott: mediumleanstral formal-verification lean-provingMistral AI
What is this?
Leanstral is Mistral AI’s open-source, Lean 4–focused model/code agent for constructing mathematical proofs and verifying software specifications. Mistral claims Leanstral 1.5 is Apache-2.0 licensed with 6B active parameters and reports state-of-the-art results on several formal-verification benchmarks, including FATE-H and FATE-X. However, the supplied independent evaluation names Gemini 3.1 Pro and Claude Opus 4.7 as leaders and does not establish that Leanstral was included, so the claimed advantage over general-purpose reasoning models remains unverified here.
Why it matters to Scott
Mistral’s Lean-specialized open model converges with Scott’s use of deterministic verification and task-aware model routing, while its unverified benchmark claims make his vendor-neutral Capability Audit directly applicable. Independent comparison with frontier generalists could materially test whether specialization lowers the formalisation bottleneck; until then, this is a relevant evaluation target rather than validated progress.
ip:concept.capability-auditip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersip:concept.formalisation-bottleneckdev:concept.task-aware-model-routingradar:concept.open-modelsradar:concept.ai-benchmarksradar:concept.benchmark-integrityradar:concept.coding-modelsradar:proofatlas-collatz-formalizationradar:llm-verified-linux-nftables
queries asked of Scott's wikis
- formal verification in coding-agent harnesses
- verifier-guided agent loops and proof-carrying code
- Lean 4 theorem proving workflows
- specialized open models versus general-purpose reasoning models
- benchmark saturation and independent model evaluation
- AI-generated code trust beyond human review
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-08-03T14:25:23Z
No independent evaluation or implementation evidence emerged within the case’s horizon, and repeated checks found only the original vendor claim plus unrelated kernel context. The signal has faded and can be reopened if a substantive comparison appears.
2026-07-31T13:21:53Z
The added attention remains negligible and no independent evaluation has tested Leanstral’s claimed advantage. The kernel issue is contextual rather than corroborating, so this remains a cold, open evaluation target.
2026-07-28T12:22:43Z
The Lean kernel soundness issue raises the bar for evaluating Leanstral: practical assessments must verify the trusted kernel and environment, not merely proof-generation rates. It does not independently test Leanstral or corroborate its claimed advantage over general-purpose models.
2026-07-28T12:21:42Z
evidence attached: hn.story.49082495 — A kernel soundness issue materially contextualizes the reliability of Lean-based AI verification workflows.
2026-07-26T21:22:08Z
No new evidence since grounding; HN discussion never materialized (0 comments) and no independent evaluation has yet tested Leanstral against frontier generalists. Remains an open evaluation target, not yet corroborated.
2026-07-23T02:23:44Z
grounded: converges/medium — Mistral’s Lean-specialized open model converges with Scott’s use of deterministic verification and task-aware model routing, while its unverified benchmark clai
2026-07-23T02:21:12Z
case created — A first-party technical report introduces a specialized model whose practical advantage remains independently testable.
Decision trace
- 08-04 00:25expireNo independent evaluation or implementation evidence emerged within the case’s horizon, and repeated checks found only the original vendor claim plus unrelated kernel context. The signal has faded and
- 07-31 23:21repriceThe added attention remains negligible and no independent evaluation has tested Leanstral’s claimed advantage. The kernel issue is contextual rather than corroborating, so this remains a cold, open ev
- 07-28 22:22repriceThe Lean kernel soundness issue raises the bar for evaluating Leanstral: practical assessments must verify the trusted kernel and environment, not merely proof-generation rates. It does not independen
- 07-28 22:21attachA kernel soundness issue materially contextualizes the reliability of Lean-based AI verification workflows.
- 07-28 22:21propose_attachA kernel soundness issue materially contextualizes the reliability of Lean-based AI verification workflows.
- 07-27 07:22repriceNo new evidence since grounding; HN discussion never materialized (0 comments) and no independent evaluation has yet tested Leanstral against frontier generalists. Remains an open evaluation target, n
- 07-23 12:23groundMistral’s Lean-specialized open model converges with Scott’s use of deterministic verification and task-aware model routing, while its unverified benchmark claims make his vendor-neutral Capability Au
- 07-23 12:21createA first-party technical report introduces a specialized model whose practical advantage remains independently testable.