2026-10-11 17:11 UTC

Independent evaluations will determine whether DFAH-Bench’s trajectory-agreement metrics reveal consequential agent nondeterminism that outcome-only evaluations miss.

state: expiredheat: lowuncertainty: highconvergesscott: highagent-benchmarks tool-use agent-reliabilityIBM Client Engineering

What is this?

DFAH-Bench is a benchmark and assurance harness for financial AI agents that compares repeated runs reaching the same decision while measuring changes in tool choice, call order, arguments, and evidence. The supplied arXiv and IBM Client Engineering GitHub snippets frame this as observable execution fidelity under replay, arguing that outcome-only pass/fail evaluation misses process variation relevant to replayability and change control. The material identifies IBM Client Engineering as the associated organization, but it does not establish that independent evaluations have yet validated the benchmark’s metrics or shown the observed nondeterminism to be consequential in production.

Why it matters to Scott

IBM Client Engineering’s DFAH-Bench independently operationalizes Scott’s existing Path Testing and Cognitive Provenance position: correct outcomes are insufficient when the agent’s actual tool and evidence path varies or cannot be replayed. This creates a strong dated-receipts and potential benchmark-comparison opportunity, although the supplied evidence does not yet show that its trajectory-agreement metrics predict consequential production failures.
ip:concept.path-testingip:concept.cognitive-provenanceip:concept.non-determinismip:framework.reflexive-agent-designdev:concept.deterministic-agent-control-planeradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:coding-agent-self-report-failure-blindness
queries asked of Scott's wikis
  • trajectory-level evaluation versus outcome-only metrics
  • agent replayability and deterministic tool execution
  • tool-call trace observability and change control
  • agent nondeterminism in production harnesses
  • execution faithfulness versus answer correctness
  • reproducible agent runs and audit trails

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: DFAH-Bench – same agent decision, different tool pathsraffisk10
🟧 echo.paper ⭐The earliest primary artifact for the specific same-decision/different-tool-path finding is “Replayable Financial Agents,” submitted 17 JanuRaffi Khatchadourian——

Interpretation history

Decision trace