Neocatalyst Labs' SemLayer evaluation finds AI agents produce syntactically valid SQL but return wrong answers 58% of the time on a real data warehouse, exposing a semantic correctness gap beyond syntax validity.
state: seedheat: mediumuncertainty: mediumconvergesscott: mediumagent-evaluation text-to-sql semantic-correctnessrajdevnathNeocatalyst Labs
What is this?
Neocatalyst Labs (associated with rajdevnath) published an evaluation called SemLayer testing AI agents on a real data warehouse, finding that while agents produce syntactically valid SQL, 58% of answers are semantically wrong. The web snippets confirm this pattern is widely recognized โ Colrows documents an 'accuracy cliff' where models score 86-91% on academic benchmarks but 10-21% on real enterprise schemas; dbt's 2026 benchmark shows semantic layers achieve 100% correctness within scope vs. text-to-SQL's fragility; multiple practitioners report silent semantic errors (wrong joins, wrong metric definitions) that pass syntax checks. The specific Neocatalyst Labs study and the 58% figure appear only in the case's evidence title; the snippets corroborate the phenomenon but not this particular evaluation.
Why it matters to Scott
Neocatalyst's production-warehouse result independently arrives where Scott's canon already argues: evaluation-driven development, capability-audit's 'representative production data, not demo conditions', and benchmarking-the-wrong-unit all predict the academic-benchmark-to-production collapse, and the 58% silent-semantic-failure figure is a dated receipt for validation-gated extraction. It touches his own Manager project directly โ a SOQL-generating Salesforce agent whose failure mode is exactly wrong-answers-behind-valid-syntax โ and arms his LeverageAI consulting line on agent evaluation with a citable number; the value is the receipt, not a surprise.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.benchmarking-the-wrong-unitdev:concept.validation-gated-llm-extractiondev:project.managerip:concept.answer-failure-classesradar:schemagate-denied-versus-emptyradar:anchor-reviewable-semantic-modelsradar:qorl-small-model-postgres-planningradar:eval-tool-scoring-audit
queries asked of Scott's wikis
- agent-evaluation frameworks semantic-correctness vs syntactic-validity
- text-to-sql agent reliability production benchmarks
- semantic-layer infrastructure agent-governance determinism
- evaluation-harnesses agent-memory wiki-maintained
- local-model text-to-sql fine-tuning semantic-accuracy
- agent-output-validation silent-failure-detection
Measured heat
now 0 pts/hpeak 15 pts/hcomments 0/hpeers p14momentum: steady1 platformsage 73h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p34 vs 1243 stories at the 72h mark (now 73h old) โ ahead of 3jsbench-llm-3d-generation-benchmark (1.5x), behind acs-local-skill-risk-catalog (0.8x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-08T19:42:20Z
grounded: converges/medium โ Neocatalyst's production-warehouse result independently arrives where Scott's canon already argues: evaluation-driven development, capability-audit's 'represent
2026-10-08T19:29:59Z
case created โ Concrete evaluation result on a real warehouse showing high semantic error rate despite valid SQL; directly relevant to agent-evaluation hot topic.
Decision trace
- 10-09 07:59attention_routeThe editor compared this story and chose to keep watching.
- 10-09 07:54attention_candidatecreate
- 10-09 06:42groundNeocatalyst's production-warehouse result independently arrives where Scott's canon already argues: evaluation-driven development, capability-audit's 'representative production dat
- 10-09 06:30createConcrete evaluation result on a real warehouse showing high semantic error rate despite valid SQL; directly relevant to agent-evaluation hot topic.