2026-10-11 17:11 UTC

Independent use will determine whether Replaybook can reproducibly evaluate infrastructure agents through realistic incident replays and produce useful results beyond its author’s own training workflow.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-benchmarks infrastructure-agents agent-evaluation

What is this?

Replaybook is presented in the supplied evidence titles as an infrastructure-agent evaluation framework, originating from a repository whose initial concept was a terminal game in which an agent is paged to repair realistic failures. The broader snippets support replayable, controlled simulations as a method for reproducible multi-step agent evaluation, but they do not establish Replaybook’s specific implementation, authorship, deterministic-branching capabilities, or results. No independent use or validation is shown, so the hypothesis remains an open test rather than an established outcome.

Why it matters to Scott

Replaybook independently applies Scott’s replay-driven evaluation position to infrastructure-agent incident response, creating a potential dated-receipts and practical harness-comparison opportunity. However, the supplied evidence establishes neither implementation quality nor independent results, and the radar tracks only adjacent benchmarks such as Orca-Bench—not this development itself.
ip:framework.replay-driven-design-evolutionip:framework.reflexive-agent-designip:concept.counterfactual-design-replaydev:project.routerradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.agent-harnessesradar:orca-bench-oncall-agent-readiness
queries asked of Scott's wikis
  • realistic incident replays for coding-agent evaluation
  • reproducible harnesses for stateful multi-step agents
  • infrastructure agents and autonomous incident response
  • trajectory-level grading versus outcome-based evaluation
  • benchmark contamination and author-built agent workflows
  • deterministic replay and diff-based agent experiments

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Replaybook, an Infrastructure Agent Evaluation Frameworkducks_10
🟧 echo.github ⭐The repository's first commit, “Initial spec,” contains the project concept: “A terminal game where you get paged and have to fix real brokeJake Goldsborough——

Interpretation history

Decision trace