2026-10-11 16:38 UTC

Replay maintainer Daniel Saito claims its released local CLI identifies the turns, causes, and token costs of prompt-cache breaks in agent transcripts, exposing 42.9 million re-billed tokens in his 119-session corpus and enabling turn-level inference-cost diagnosis.

state: seedheat: lowuncertainty: highconvergesscott: mediumprompt-caching inference-economics agent-observabilityDaniel SaitoReplay Doctor

What is this?

The case describes Replay (also called Replay Doctor) as a released local CLI for auditing prompt-cache misses in AI agent transcripts, attributed to maintainer Daniel Saito. Saito reportedly claims it identifies the turns, causes, and token costs of cache breaks, finding 42.9 million re-billed tokens across his 119-session corpus. The supplied web search returned no results, so the tool’s release, capabilities, and measurements remain uncorroborated case claims; the evidence title’s dollar-cost framing is also unclear.

Why it matters to Scott

Replay’s claimed turn-level cache-break diagnosis converges with Scott’s Agent Observability position and offers a concrete way to test the stable-prefix assumption behind Prefix-Caching Economics against his coding-session archive. The radar already tracks related work in cache-hunter-prompt-cache-debugging and tool-schema-prompt-cache-invalidation, but not this Replay development; its capabilities, 42.9-million-token finding, billing interpretation, and compatibility with Scott’s archive remain unverified.
ip:concept.agent-observabilityip:concept.prefix-caching-economicsdev:project.search-conversationsradar:cache-hunter-prompt-cache-debuggingradar:tool-schema-prompt-cache-invalidationradar:agentmeasure-token-accounting-auditradar:concept.prompt-caching
queries asked of Scott's wikis
  • agent harness prompt-cache stability and invalidation
  • turn-level token accounting and inference-cost observability
  • local agent transcript auditing and debugging
  • agent memory and tool-schema changes affecting cached prefixes
  • measured versus estimated API billing costs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 661h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 03:28 (minted)⭐ origin echo-reconstructedReplay Doctor offers a local transcript-auditing CLI and reports 42.9 million re-billed tokens, valued at $199.77 in list prices or 5% of it
Daniel Saito on blog (echo) · attributed from hn.story.49691257 · published time unknown
—
09-14 02:38first on hacker news · published · lag ?Show HN: Replay – Audit silent prompt cache misses in AI agent transcripts
danielsaito
—
09-14 05:11first on r/ClaudeAI · published · lag ?I built a tool to search my agent logs. Across 3000+ sessions and 30 billion tokens, only 0.3% of that was the model actually writing anything
Rare_Guide_9830
—
09-16 15:56first on r/LocalLLaMA · published · lag ?Is there a coding harness that will not invalidate the entire context cache at each prompt?
Academic-Tea6729
—
09-14 02:38amplified on hacker newshn.story.49691257
danielsaito
peak 1 · 0 comments · 1% of case engagement
09-14 05:11amplified on r/ClaudeAIreddit.post.1wfu8eh
Rare_Guide_9830
peak 1 · 5 comments · 2% of case engagement
09-16 09:27amplified on r/ClaudeAIreddit.post.1whsbju
Marmelab
peak 70 · 32 comments · 28% of case engagement
09-16 12:23amplified on r/ClaudeAIreddit.post.1whvqcl
Schadz
peak 11 · 13 comments · 7% of case engagement
09-16 15:56amplified on r/LocalLLaMAreddit.post.1wi17z6
Academic-Tea6729
peak 8 · 35 comments · 12% of case engagement
09-17 19:55amplified on r/ClaudeAI 👑reddit.post.1wj4gs0
Schadz
peak 134 · 58 comments · 52% of case engagement
09-14 03:20our radar first saw it · lag ?discovery anchor: hn.story.49691257—
pace: p81 vs 1032 stories at the 336h mark (now 661h old) — ahead of google-weathernext-3-release (1.0x), behind fan-coding-harness-component-study (1.0x)

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Replay – Audit silent prompt cache misses in AI agent transcripts
Retrieved article excerpt

Open article · Retrieved 2026-09-14T03:21:29.767669+00:00

# If the agent bill went up and nothing errored, a prompt cache broke.

Replay Doctor reads the transcripts already on your disk and names the turn it broke on, the cause, and the tokens re-billed at write prices. Expiry is only one of the causes it can name.

$ curl -fsSL https://replay.doctor/replay.sh | lessCopy

Read it, then | sh. Or go install github.com/RedRobotKK/Replay/cmd/replay@latest, or a [signed release](https://github.com/RedRobotKK/Replay/releases). macOS and Linux. Nothing leaves your machine.

42.9Mtokens re-read at write prices

$199.77 at list, which is the price for someone billed per token. 119 sessions · 1,881 agent lanes · one machine, one operator · read 2026-09-12 by replay cost, build 0.5.4. A method demonstration, not a market.

pool: 1 machine, 121 sessions, pooled 2026-09-12. The binary cannot send; you post one file yourself. [Add yours](https://replay.doctor/contribute/) · [roster](https://replay.doctor/pool/)

$ replay cost$ replay diff$ replay context

Cost per task, across 119 sessions (1,881 agent lanes)

at list prices, read 2026-09-12 by replay cost, build 0.5.4.

total

$4039.23

median task

$0.79

p90 task

$7.79

avoidable

$199.77  (5% of the total)

42.9M tokens re-billed

Avoidable is the part nobody chose: tokens re-billed because a prompt cache broke. It is not a forecast, it is what was already spent twice.

one machine, one account, one operator

$ replay diff <session>

added 3 tool(s): mcp\_\_claude\_ai\_Otter\_ai\_\_otter\_fetch, otter\_get\_user\_info, otter\_search; removed 1 tool(s): WaitForMcpServers. That break cost 157,080 tokens.

One line per break, with its cause. A re-render is frequent and small; a TTL expiry is rare and enormous. Shares are not printed: the ranking flipped between two readings on 2026-09-11.

$ replay context <session>

system prompt: 8.0% here, 14.0% across 13.5M sessions, 760.5M LLM calls (arXiv:2608.00101)

The population travels with the figure on the same line. There is no "high", no "average" and no advice anywhere in the vocabulary, and a test asserts there never will be.

119

sessions read

1,881 lanes · read by replay cost 0.5.4 on 2026-09-12

$4,039.23

at list prices

one machine, one operator

$199.77

already spent twice

5% of the total · 42.9M tokens · replay cost 0.5.4, read 2026-09-12

3

retractions, all still visible

[98.8% became 4.2%, files became sessions, fetches were not installs](https://replay.doctor/docs/trust/corrections/)

On a subscription seat the dollars are list price for someone else. The tokens are yours: a re-billed token is context the work did not get. One machine proves the method and nothing about the market; the pool below is how that changes.

Do you need this?

On a Pro seat with short sessions

Set promptCacheTtl to 1h in Claude Code settings, keep the tool list stable mid-session, and you are done. This page is not for you, and that is fine.

On Max, hitting the usage window

Run it once. The diff tells you whether the wall is output, compaction, or silent re-reads.

Metered API, unattended runs, sub-agents, or you ship a harness

This is for you. Every re-bill is an invoice line, and the tool names the turn.

The population is one machine. Change that.

## Add your machine to the pool. Under 600 bytes, nothing about you.

Replay writes one file: sixteen counts and ratios, three strings naming the build that priced them, no paths and no account. You read it, then you post it yourself; the binary cannot. The pool is summed by replay pool in public, and every row on the [roster](https://replay.doctor/pool/) is a file you can download and re-hash.

1 machine · 121 sessions · pooled 2026-09-12

[How it works, and what it discloses](https://replay.doctor/contribute/)[See the roster](https://replay.doctor/pool/)

three commands, one file, one postCopy

```
mkdir -p ~/.config/replay && printf 'corpus_opt_in = true\n' > ~/.config/replay/corpus-consent.toml
replay cost --contribute launch-2026-09 --contribute-dir . ~/.claude/projects
curl -sS -X POST --data-binary @replay-corpus-*.json "https://replay.doctor/api/contribute?campaign=launch-2026-09"
```

--contribute needs 0.6.0 or later: go install github.com/RedRobotKK/Replay/cmd/replay@latest, or a signed release. Read the file before the third line; the reply names your digest.

## Four commands. Each one prints what it measured, and refuses what it did not.

[Every command, generated from the binary](https://replay.doctor/docs/reference/)

replay diff

Names the turn the cache broke on, what changed, and the tokens it cost.

added 3 tool(s): mcp\_\_claude\_ai\_Otter\_ai\_\_otter\_fetch, otter\_get\_user\_info, otter\_search; removed 1 tool(s): WaitForMcpServers. That break cost 157,080 tokens.

[Run it](https://replay.doctor/docs/reference/diff/)

replay context

What is filling your context, ranked, and by how much the ranking overstates after a compaction.

39 compactions · median retention 2.55%

[Run it](https://replay.doctor/docs/reference/context/)

replay route --to

Prices the switch itself and names the turn where the cheaper model becomes cheaper.

refused: pair not measured on the wire

[See the refusal](https://replay.doctor/docs/reference/route/)

replay cost --contribute

Writes under 600 bytes: 16 numbers and 3 strings naming the build, no paths. This is how one machine stops being the sample.

replay.corpus.v1 · 19 fields · under 600 bytes

[Post yours](https://replay.doctor/contribute/)

The number

This much was paid twice, because a prompt cache broke and nothing said so. The meter beside it is the same reading, taken once, on one machine.

read by replay cost 0.5.4 on 2026-09-12 · 119 sessions · one machine, one operator

[Contribute under 600 bytes](https://replay.doctor/contribute/)

RE‑BILLEDUSD

0019977

TOKENS RE‑BILLED42.9M

OF SPEND5%

SESSIONS119

REPLAY COST 0.5.4READ 2026-09-12

Billed per token, running agents unattended?

I will read your corpus with you, on your machine. Nothing leaves it. One week, one operator.

Diagnosis is free and stays free. The week is the only thing for sale: you get the invoice annotated turn by turn, every break named with its cause, the layout changes worth making and the ones that are not, and the write-up, which is yours. As of 2026-09-13 the week is $22,000, quoted at $18,000 to $25,000 by scope, and none has been sold yet; that sentence changes the day one is.

[Talk to Daniel](https://replay.doctor/cdn-cgi/l/email-protection#2440454a4d414864564140564b464b500a4e541b5751464e4147501956415448455d0a404b47504b56)

What it got wrong, in public

~~98.8% of re-billed tokens from one cause~~ 2026-09-06 → 4.2% same day, per lane

~~1,363 sessions~~ files counted as sessions → 78 2026-09-06

~~26 installs~~ fetches, not installs → 53 fetches, 2 IPs 2026-09-07

[Every correction, dated, with the wrong number still on the page](https://replay.doctor/docs/trust/corrections/)

**A note from Daniel, who maintains Replay Doctor.** Replay is free and stays free: no account, no key, no telemetry.

The measurement behind it is not free. The figures are calibrated against **32,188 real requests** from 115 of my own sessions, read 2026-09-07, and each new provider costs the same again. I was on a subscription, so at list prices that corpus is a four-figure sum the funding page states with its date. If Replay found something on your machine you had already paid for once, a share of that back is what keeps it maintained.

[Support the work](https://buymeacoffee.com/saitodaniel)[What it pays for](https://github.com/RedRobotKK/Replay/blob/main/FUNDING.md)Maybe later

Read the script. Then run it.

No account, no telemetry, no first-run prompt. Checksum verified, refuses to build from source when it cannot.

$ curl -fsSL https://replay.doctor/replay.sh | shCopy

[Star on GitHub](https://github.com/RedRobotKK/Replay)[go install](https://replay.doctor/docs/getting-started/)[Releases](https://github.com/RedRobotKK/Replay/releases)

## Hear what happens next.

Replay is free and installs in one line, so there is nothing here to wait for. This list is for the two things that are not available yet: word when the tool gains something worth knowing about, and the forensic engagement, which is one person's calendar and is therefore genuinely finite.

No schedule, because there has never been one. One confirmation email, then nothing until there is something. Leaving is one click and needs no reply.

Or sign in, which skips the confirmation email because your provider has already verified the address.

Only your verified email address is read. No profile, no name, no repositories, and the token is used once and never stored.

### Who has recommended it
danielsaito10
🟧 echo.blog ⭐Replay Doctor offers a local transcript-auditing CLI and reports 42.9 million re-billed tokens, valued at $199.77 in list prices or 5% of itDaniel Saito——
🟠 redditI built a tool to search my agent logs. Across 3000+ sessions and 30 billion tokens, only 0.3% of that was the model actually writing anything
ClaudeAI
Rare_Guide_983005
🟠 redditI tested 3 more Claude Code plugins to cut costs. Here’s my verdict
ClaudeAI
Marmelab7032
🟠 redditQuota draining? Claude Code re-bills a finished sub-agent's whole context on a follow-up
ClaudeAI
Schadz913
🟠 redditIs there a coding harness that will not invalidate the entire context cache at each prompt?
LocalLLaMA
Academic-Tea6729735
🟠 redditSub-agents burning your Claude Code 5-hour window? Check their 5-minute prompt cache
ClaudeAI
Schadz13158

Interpretation history

Decision trace