Independent evaluations will determine whether Snowflake’s Data-eng-bench provides reproducible, realistic, and decision-useful measurements of AI agents performing data-engineering tasks.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-evaluation data-engineering coding-agentsSnowflake
What is this?
Snowflake AI Research released Data-eng-bench, an open-source benchmark for testing coding agents on data-engineering work in a large retail-warehouse dbt project. Agents receive ticket-style tasks, edit or create dbt models in a containerized environment, and are scored by hidden pytest verifiers that compare materialized tables row by row with reference solutions; it runs on Harbor to support multiple agents. Snowflake describes it as realistic and hard to saturate, but the supplied results do not establish broad adoption or independent validation of its reproducibility, realism, or decision usefulness.
Why it matters to Scott
Snowflake’s benchmark independently adopts several positions Scott already operationalizes: evaluate the model-plus-harness unit in a live task environment, use objective hidden acceptance checks, and preserve comparable execution traces across agents. This creates a dated-receipts and hands-on comparison opportunity for his trace-backed evaluation work, although the supplied material does not yet show that Data-eng-bench itself is reproducible, realistic, or decision-useful.
ip:concept.model-plus-harness-benchmark-unitip:framework.hidden-gates-frameworkdev:concept.trace-backed-agent-comparisondev:concept.rubric-blind-agent-reviewradar:concept.coding-agent-benchmarksradar:concept.benchmark-integrityradar:concept.agent-evaluation
queries asked of Scott's wikis
- realistic benchmark design for coding agents
- hidden verifiers and benchmark reproducibility
- agent harness versus model performance
- evaluation contamination and benchmark saturation
- task-level evaluation for autonomous software agents
- data-engineering agents and dbt workflows
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-18T20:36:12Z
Repeated checks have produced no independent reproduction, implementation, or comparative evaluation, so this release no longer merits active episode tracking. The benchmark remains available, but its realism and decision usefulness are unresolved rather than disproved.
2026-08-16T20:28:57Z
No independent evaluation, implementation, or substantive discussion has emerged; the minor engagement increase is repetition rather than validation. The benchmark remains an open artifact whose realism and reproducibility may need a longer evaluation horizon.
2026-08-14T19:46:08Z
The benchmark remains a credible first-party artifact, but this reobservation adds no independent evaluation, implementation, or discussion bearing on realism or reproducibility. With the release already surfaced and no substantive follow-through, the case cools while its core hypothesis remains open.
2026-08-14T19:35:19Z
grounded: converges/medium — Snowflake’s benchmark independently adopts several positions Scott already operationalizes: evaluate the model-plus-harness unit in a live task environment, use
2026-08-14T19:33:24Z
case created — This is a first-party benchmark artifact for an important production workload, with usefulness and durability still awaiting external validation.
Decision trace
- 08-19 06:36expireRepeated checks have produced no independent reproduction, implementation, or comparative evaluation, so this release no longer merits active episode tracking. The benchmark remains available, but its
- 08-19 06:36alert_silentThe only trigger is staleness, with no consequential new evidence; attention should wait for an independent reproduction, comparative result, or material benchmark revision.
- 08-19 06:36alert_routeThe only trigger is staleness, with no consequential new evidence; attention should wait for an independent reproduction, comparative result, or material benchmark revision.
- 08-17 06:28repriceNo independent evaluation, implementation, or substantive discussion has emerged; the minor engagement increase is repetition rather than validation. The benchmark remains an open artifact whose reali
- 08-17 06:28alert_silentThere is no consequential new delta to route; revisit only if an independent reproduction, comparative result, or material benchmark revision appears.
- 08-17 06:28alert_routeThere is no consequential new delta to route; revisit only if an independent reproduction, comparative result, or material benchmark revision appears.
- 08-15 05:46repriceThe benchmark remains a credible first-party artifact, but this reobservation adds no independent evaluation, implementation, or discussion bearing on realism or reproducibility. With the release alre
- 08-15 05:46alert_silentThere is no new consequential delta beyond an unchanged reobservation; the established release was already routed, and no external validation now warrants renewed attention.
- 08-15 05:46alert_routeThere is no new consequential delta beyond an unchanged reobservation; the established release was already routed, and no external validation now warrants renewed attention.
- 08-15 05:40alert_shadowThe first-party release is established and directly overlaps Scott’s trace-backed evaluation approach: model-plus-harness testing, executable task environments, hidden acceptance checks, and comparabl
- 08-15 05:40alert_routeThe first-party release is established and directly overlaps Scott’s trace-backed evaluation approach: model-plus-harness testing, executable task environments, hidden acceptance checks, and comparabl
- 08-15 05:35groundSnowflake’s benchmark independently adopts several positions Scott already operationalizes: evaluate the model-plus-harness unit in a live task environment, use objective hidden acceptance checks, and
- 08-15 05:33createThis is a first-party benchmark artifact for an important production workload, with usefulness and durability still awaiting external validation.