LinearSolveBench's maintainer claims the released benchmark measures whether coding agents can produce fast, accurate, and general C solvers for large sparse linear systems, extending agent evaluation beyond conventional repository tasks.
state: seedheat: lowuncertainty: mediumknownscott: lowagent-evaluation ai-assisted-mathematics coding-agentsLinearSolveBenchhgarud
What is this?
LinearSolveBench is a newly released benchmark (surfaced via a Show HN post by user hgarud and a Reddit r/ScientificComputing thread, hosted at autodidakt.ai) that measures a model or agent harness's ability to write fast, accurate, and general numerical solvers in C for large sparse linear systems. The stated goal is to spur algorithmic advances in numerical linear algebra by using AI systems as the search/optimization engine, and it positions itself as an agent evaluation beyond the usual repository-task benchmarks (SWE-bench-style patch-and-test suites). The supplied snippets are thin: nothing establishes the maintainer's identity or affiliation, the benchmark's size/methodology, any leaderboard results, or whether anyone has actually run coding agents on it โ the HN post had 1 point and no discussion at scrape time.
Why it matters to Scott
This is another entry in the beyond-SWE-bench agent-benchmark wave the radar already tracks (radar:concept.coding-agent-benchmarks, with AWS-bench and Terminal-Bench-Science as the same story-shape in open cases), and it exercises positions Scott's wiki already holds via ip:concept.model-plus-harness-benchmark-unit โ capability measured as model-plus-harness against an external oracle, not SWE-bench-style patches. With no leaderboard results, methodology detail, adoption, or even HN traction in the supplied evidence, it's the world providing another example of his evaluation pattern rather than a challenge, extension, or publishing opportunity; if the benchmark later shows real adoption it could become a fresh fixture for his trace-backed harness comparisons, but on current evidence there is nothing to act on.
ip:concept.model-plus-harness-benchmark-unitdev:project.remote-execradar:concept.coding-agent-benchmarksradar:aws-bench-cloud-agent-evaluationradar:terminal-bench-science-workflowsradar:concept.ai-assisted-mathematics
queries asked of Scott's wikis
- coding agent evaluation beyond SWE-bench โ his position on what agent benchmarks actually measure
- benchmark design for agents: fast/accurate code, performance-based scoring, verification harnesses
- numerical/scientific code generation โ sparse linear algebra, C kernels, performance-critical code
- agent harness comparison methodology โ how he evaluates his own coding agents and tools
- AI-driven algorithm discovery โ models beating or finding known numerical algorithms
- autodidakt.ai or hgarud โ any prior radar tracking of this actor or project
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 456h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p42 vs 1032 stories at the 336h mark (now 456h old) โ ahead of agent-memory-add-search-evaluation (1.2x), behind agenticos-self-hosted-governance (0.9x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-23T19:57:26Z
grounded: known/low โ This is another entry in the beyond-SWE-bench agent-benchmark wave the radar already tracks (radar:concept.coding-agent-benchmarks, with AWS-bench and Terminal-
2026-09-23T02:22:37Z
case created โ The linked benchmark is a concrete, distinct evaluation artifact for numerical programming.
Decision trace
- 09-24 05:57groundThis is another entry in the beyond-SWE-bench agent-benchmark wave the radar already tracks (radar:concept.coding-agent-benchmarks, with AWS-bench and Terminal-Bench-Science as the same story-shape in
- 09-24 04:15createThe linked benchmark is a concrete, distinct evaluation artifact for numerical programming.