2026-10-11 17:10 UTC

Independent evaluations will determine whether Snowflake’s Data-eng-bench provides reproducible, realistic, and decision-useful measurements of AI agents performing data-engineering tasks.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-evaluation data-engineering coding-agentsSnowflake

What is this?

Snowflake AI Research released Data-eng-bench, an open-source benchmark for testing coding agents on data-engineering work in a large retail-warehouse dbt project. Agents receive ticket-style tasks, edit or create dbt models in a containerized environment, and are scored by hidden pytest verifiers that compare materialized tables row by row with reference solutions; it runs on Harbor to support multiple agents. Snowflake describes it as realistic and hard to saturate, but the supplied results do not establish broad adoption or independent validation of its reproducibility, realism, or decision usefulness.

Why it matters to Scott

Snowflake’s benchmark independently adopts several positions Scott already operationalizes: evaluate the model-plus-harness unit in a live task environment, use objective hidden acceptance checks, and preserve comparable execution traces across agents. This creates a dated-receipts and hands-on comparison opportunity for his trace-backed evaluation work, although the supplied material does not yet show that Data-eng-bench itself is reproducible, realistic, or decision-useful.
ip:concept.model-plus-harness-benchmark-unitip:framework.hidden-gates-frameworkdev:concept.trace-backed-agent-comparisondev:concept.rubric-blind-agent-reviewradar:concept.coding-agent-benchmarksradar:concept.benchmark-integrityradar:concept.agent-evaluation
queries asked of Scott's wikis
  • realistic benchmark design for coding agents
  • hidden verifiers and benchmark reproducibility
  • agent harness versus model performance
  • evaluation contamination and benchmark saturation
  • task-level evaluation for autonomous software agents
  • data-engineering agents and dbt workflows

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnData-eng-bench, Snowflake's data engineering benchmark for agentsmarcociavarella10
🟧 echo.github ⭐Snowflake’s released benchmark for evaluating AI agents on data-engineering tasks.Snowflake-Labs——

Interpretation history

Decision trace