2026-10-11 16:37 UTC

Firedrill's maintainers claim their released framework combines stateful synthetic tools, fault injection, virtual time, and state assertions across existing agent interfaces, enabling reproducible workflow regression tests without changing production agent logic.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumagent-evaluation agent-harnesses tool-simulationFiredrill

What is this?

Firedrill is an open-source framework, released by an independent maintainer (a Reddit post in r/learnAIAgents appears to be the primary announcement), for testing AI agents against simulated tools instead of production accounts or real APIs. Its claimed differentiators are stateful synthetic tools — mocks that hold and mutate state across calls rather than returning canned responses — plus fault injection, virtual time, snapshots, and state assertions, exposed through Python and TypeScript interfaces and CI reporting, with the pitch that it wraps existing agent interfaces rather than requiring changes to production agent logic. Independent corroboration in the supplied snippets is thin: the surrounding web results are mostly about the adjacent durable-execution/fault-tolerance space (Temporal, LangGraph, Restate), not about Firedrill itself, so the specific feature claims rest on the maintainers' own descriptions.

Why it matters to Scott

Firedrill is essentially Scott's deterministic-test-seams doctrine productized for agents by another party: stateful fake tool implementations, virtual time, fault injection and snapshot/reset in CI — the same 'clock, payment, LLM behind a Protocol with a real and a fake impl' pattern he already codified, extended with agent-specific state assertions that complement his trace-backed agent comparison practice. It's a dated receipt from an independent builder, not a mere illustration, and it's a tool he could evaluate directly in his agent-harness projects — though its thin provenance (single Reddit announcement, no independent corroboration in the snippets) caps it below high.
dev:concept.deterministic-test-seamsdev:concept.trace-backed-agent-comparisonip:framework.12-factor-agents-frameworkdev:concept.transparent-protocol-capture-for-replacementradar:514-coding-agent-simulation-infra
queries asked of Scott's wikis
  • agent testing and evaluation harness patterns — how does Scott currently test agent tool calls, and does he distinguish schema-correct mocks from stateful simulation?
  • fault injection and deterministic replay in his agent harness or coding-agent projects — has he built or wanted corrupted-fixture / adversarial tool-response testing?
  • virtual time and snapshot/state reset in CI for agent workflows — what does his test isolation story look like for multi-step agent runs?
  • non-invasive instrumentation: does he have a position on test tooling that wraps existing agent interfaces vs. requiring production code changes?
  • open-source agent infrastructure tools he tracks or contributes to — is there an existing entry where a simulation/testing framework would slot into his stack?
  • regression testing for LLM workflows — has he written about reproducibility or flakiness of agent runs as a blocker to shipping?

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 458h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-22 18:26 (minted)⭐ origin echo-reconstructedReleased a stateful agent simulation and testing framework with Python and TypeScript interfaces, scenario faults, snapshots, and CI reports
Firedrill on github (echo) · attributed from hn.story.49801073 · published time unknown
—
09-22 13:35first on hacker news · published · lag ?Firedrill: Stateful tool simulation for AI agents
newton_reload
—
09-22 13:35amplified on hacker news 👑hn.story.49801073
newton_reload
peak 2 · 0 comments · 98% of case engagement
09-22 14:21our radar first saw it · lag ?discovery anchor: hn.story.49801073—
pace: p23 vs 1032 stories at the 336h mark (now 458h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFiredrill: Stateful tool simulation for AI agents
Retrieved article excerpt

Open article · Retrieved 2026-09-23T14:28:02.152285+00:00

# Firedrill

Firedrill is a simulation and testing framework for AI agents. Define synthetic
tools and data, run your agent against them, and assert on tool calls, state
changes, and events.

- Stateful tools with HTTP, MCP, CLI, and function bindings.
- Scenario-based tests with faults, response overrides, and virtual time.
- Isolated world state, seeded data, snapshots, and resets.
- HTML, JSON, and JUnit reports with timelines and optional browser captures.
- Repository-defined tools, including independently distributed packages.

[Quickstart](https://github.com/firedrill-tools/firedrill#quickstart) · [Python](https://github.com/firedrill-tools/firedrill/blob/main/python/README.md) · [TypeScript](https://github.com/firedrill-tools/firedrill#using-the-sdk) ·
[Documentation](https://docs.firedrill.run) · [Neutral example](https://github.com/firedrill-tools/firedrill/blob/main/examples/quickstart/README.md) ·
[Gmail Agent example](https://github.com/firedrill-tools/firedrill-example-gmail-agent)

## Installation

### Python

Install the current release from [PyPI](https://pypi.org/project/firedrill-run/)
into your virtual environment with Python 3.10 or later:

```
python -m pip install "firedrill-run[pytest]"
firedrill --help
```

Import the local SDK with `from firedrill import World, run_drills`. The wheel
includes the runtime, CLI, inspector, and report engine. No separate Node.js or
npm installation is required. Choose Tools with `firedrill init`; they install
on demand and use the same packages as TypeScript. See the
[Python guide](https://github.com/firedrill-tools/firedrill/blob/main/python/README.md) for pytest, async agents, mocks, and browser tests.

### TypeScript and JavaScript

Requires Node.js 20.19 or later.

Install the current release in your project:

```
npm install --save-dev @firedrill-run/cli @firedrill-run/sdk
npx firedrill --help
```

Or run the CLI without adding a project dependency:

```
npx @firedrill-run/cli init
```

For a source checkout, use pnpm 9.15–10:

```
pnpm install --frozen-lockfile
pnpm build

# Use the built CLI in this terminal.
export FIREDRILL_CLI="$PWD/packages/cli/dist/bin.js"
firedrill() { node "$FIREDRILL_CLI" "$@"; }
```

The examples below use `npx firedrill` inside a project that has
`@firedrill-run/cli` installed. Never run bare `npx firedrill` elsewhere: outside
such a project npm resolves an unrelated package with that name. Python
installations expose `firedrill` directly, so drop the `npx` prefix. From source,
use the shell function above or invoke the CLI directly with
`node /path/to/firedrill/packages/cli/dist/bin.js`.

The programmatic API is `@firedrill-run/sdk`. To prepare installable archives of the
CLI, SDK, and other packages from this checkout, run
`pnpm pack:artifacts -- --output /absolute/path/to/an/empty/directory`.
The output includes a package manifest.

## Quickstart

Create a project with a synthetic record store:

```
mkdir firedrill-example
cd firedrill-example
npm init -y
npm install --save-dev @firedrill-run/cli
npx firedrill init --custom records
npx firedrill serve
```

`init` creates a Tool declaration, a behavior module, and starting data.
`serve` starts the backend and opens the inspector.

Open **Tools** to inspect the implementation or call an operation.
**State & activity** shows records and calls; **Connect agent** provides the
connection settings. Tools with a bundled UI also have an **Open app** action.
Browser actions and API calls use the same state.

The server listens on loopback using available ports. Keep the terminal open;
Ctrl+C stops it. Use `--no-open` to skip opening the inspector automatically.

Run `firedrill init` in an existing project for guided setup, or select one of
the published packages in the [Tool catalog](https://github.com/firedrill-tools/firedrill/blob/main/registry/README.md), for example
`npx firedrill init --tool gmail --install` or
`npx firedrill tool add gmail --install`. Both add the catalog's exact
`@firedrill-tools/<id>` npm release as an ordinary dev dependency and refuse an
archive whose integrity differs from the catalog. The CLI shows the exact
package and version before it asks for installation consent.

Tools can run independently of tests. To check an agent's behavior, add a drill.

For a model-backed project, see the
[Gmail Agent example](https://github.com/firedrill-tools/firedrill-example-gmail-agent):
an existing Claude Agent SDK assistant runs three drills against a pinned
stateful Gmail Tool, with its synthetic mailbox and reports kept in the project.

## Writing drills

A **drill** defines an agent task, starting conditions, and assertions about the
result. To try one, stop the server, return to the parent directory (`cd ..`),
and create the example test project:

```
mkdir firedrill-first-drill
cd firedrill-first-drill
npm init -y
npm install --save-dev @firedrill-run/cli
npx firedrill init --path template
npx firedrill validate
npx firedrill plan
npx firedrill run changes-resource
npx firedrill inspect
```

The template contains a deterministic example agent that writes `7` to a record.
Its drill checks that the write succeeded once and that the final value is `7`.
Replace the example target with your agent when adding your own tests.

To inspect a failing result, edit
`firedrill/drills/changes-resource.drill.yaml`: change the `value-changed`
assertion's expected value to `8`, keeping `task.input.value` at `7`.
Rerun the drill. The report shows expected `8` and actual `7`; the process exits
with code `1`. Restore the expectation afterwards.

| Command | Purpose |
| --- | --- |
| `firedrill validate` | Check source definitions |
| `firedrill plan` | List tools, scenarios, targets, and drills |
| `firedrill run <id>` | Run one drill |
| `firedrill` | Run all drills |
| `firedrill serve` | Start a standalone synthetic backend |
| `firedrill inspect` | Browse definitions and saved results |
| `firedrill ci init github` | Add drills to pull requests, pushes, schedules, or manual CI jobs |

## Continuous integration

Run the same repository-owned drill command locally and in CI. A pull request is
one useful trigger, not a requirement: pushes, scheduled suites, manual jobs, and
other CI providers use the same command and reports.

```
firedrill ci describe
firedrill ci init github
```

The initializer detects the project's package manager or Python setup, preserves
an existing `test:firedrill`/`drills` script when present, and writes
`.github/workflows/firedrill.yml`. It uploads self-contained HTML, JSON, and JUnit
reports even when a drill fails. Use `--run` for a caller-owned SDK harness and
repeat `--trigger` to select exact events. See the
[continuous-integration guide](https://github.com/firedrill-tools/firedrill/blob/main/docs/continuous-integration.md).

### Put drills on a pull request

Add a focused safety suite to pull requests without moving the agent into
Firedrill or changing production agent code:

```
firedrill ci init github \
  --run "npm run test:firedrill" \
  --trigger pull-request \
  --trigger manual
```

Commit the generated workflow. GitHub runs the repository's existing command at
the exact pull-request revision and shows its exit status as a normal check.
Firedrill retains `.firedrill/reports/` as an Actions artifact even when a drill
fails, so a reviewer can open the HTML result, JUnit output, changed state, and
ordered evidence behind the verdict. Requiring that check in branch protection
is optional.

The hosted workflow adds managed worlds, a base-versus-head behavior comparison,
a durable evidence link, and a Firedrill check and pull-request summary. Local CI
remains complete and account-free.

### Definitions

| Term | Meaning |
| --- | --- |
| Tool | A synthetic dependency with callable operations, input/output schemas, and an implementation |
| World | Tools, starting data, identities, permissions, and a clock |
| Scenario | A variation of the starting data, permissions, faults, or scheduled events |
| Target | Configuration for invoking the agent under test |
| Drill | A task and its assertions |
| Run | A recorded execution result, including checks, calls, and state changes |

Actors identify who is calling a tool and which operations they may use.
Personas provide descriptions of those identities. See
[people and permissions](https://github.com/firedrill-tools/firedrill/blob/main/docs/world-authoring.md#people-and-permissions).

## Connecting an agent

Configure the agent's dependencies in test setup:

| Dependency | Integration |
| --- | --- |
| HTTP client | Point its base URL and authentication at the Tool's declared HTTP routes |
| MCP server | Use the world's MCP endpoint and actor token |
| CLI tool | Use a test-side command adapter or the Firedrill world CLI |
| Function or SDK method | Use `mock_tool` with Python's `unittest.mock` / pytest, or `mockTool` with JavaScript runner mocks |
| Web interface | Use a Playwright harness or the optional browser-test package |

Bindings use existing configuration or test-side adapters, leaving production
agent logic unchanged. Hardcoded dependencies need an interceptable boundary or
an explicit adapter. See [binding recipes](https://github.com/firedrill-tools/firedrill/blob/main/docs/quickstart.md#3-keep-the-agent-integration-at-one-seam)
and [test-side mocking](https://github.com/firedrill-tools/firedrill/blob/main/docs/test-mocking.md).

Targets can invoke a module, start a command, call an HTTP endpoint, or use an
`external` callback supplied by a test harness. External targets run through the
SDK; module, command, and HTTP targets can also run through the CLI.

Model credentials belong to the agent process. For command targets, pass model
credentials and other required host variables through `environmentFromHost`.
Set target timeouts for the complete model/tool loop.

## Using the SDK

Use `runDrills` from an existing test runner:

```
import { runDrills } from "@firedrill-run/sdk";

const result = await runDrills({
  root: process.cwd(),
  drill: "my-drill",
  agent: ({ task, binding, signal }) =>
    runMyAgent({ task, environment: binding.environment, signal }),
});

expect(result.verdict).toBe("passed");
```

This example assumes a declared `my-drill` with an `external` target.
`runMyAgent` is your test adapter; it applies the supplied connection values to
your agent. `expect` comes from your test runner.

`runDrills({ setup })` supports per-test data, fault, and Tool overrides.
`createLocalWorld()` provides direct control over calls, state, time, and resets.
See the [SDK reference](https://github.com/firedrill-tools/firedrill/blob/main/packages/sdk/README.md) for lifecycle hooks, concurrency,
capture, and report APIs.

## Project structure

The example template uses the following layout:

```
your-project/
  firedrill.json                         # source location and selected packages
  firedrill/
    world.yaml                          # starting data, identities, access, time
    tools/resource-store/
      resource-store.tool.yaml          # operations and state schemas
      behavior.mjs                      # operation implementations
    scenarios/baseline.scenario.yaml    # starting conditions
    targets/starter-agent.target.yaml  # agent invocation
    drills/changes-resource.drill.yaml  # task and assertions
    suites/resource-store-conformance.suite.yaml
  firedrill-example/agent.mjs           # example agent
  .firedrill-tools/                     # vendored Tool dependencies
  .firedrill/                           # generated state, builds, and reports
```

Definitions support JSON or YAML; use either consistently or mix them.
Resource suffixes identify file types, such as `.tool.json` and `.drill.yaml`.
References use IDs inside the files, so you can organize folders as needed.
Tool implementations are JavaScript or TypeScript.

Commit definitions, behavior modules, test code, package manifests, lockfiles,
and referenced `.firedrill-
newton_reload20
🟧 echo.github ⭐Released a stateful agent simulation and testing framework with Python and TypeScript interfaces, scenario faults, snapshots, and CI reportsFiredrill——

Interpretation history

Decision trace