2026-10-11 16:38 UTC

Callwitness's maintainer claims the released transparent MCP proxy records byte-exact tool traffic in hash-chained logs without blocking or delaying calls, potentially giving operators auditable evidence of what agent tools returned and where data went.

state: seedheat: lowuncertainty: mediumknownscott: mediumagent-observability tool-call-auditing mcpCallwitness

What is this?

Callwitness is presented in the case as a released transparent MCP proxy that records byte-exact tool traffic in hash-chained logs while forwarding calls without blocking or delay. The supplied search results do not directly identify Callwitness or its maintainer, so those implementation and performance claims remain unverified here. The broader architecture is established by related projects: MCP proxies can capture requests and responses into tamper-evident audit trails, although hash chains alone may not detect wholesale rewriting or truncation without an external anchor.

Why it matters to Scott

Scott already holds the core pattern in “Transparent protocol capture for local replacement,” Agent Observability, and Agent Receipts, and has built transparent MCP proxies. Callwitness could be a directly testable implementation for that active work, but its byte-exact capture, tamper evidence, and no-delay claims remain unverified and the radar already tracks closely analogous signed-log and MCP-observability tools.
dev:concept.transparent-protocol-capture-for-replacementip:concept.agent-observabilityip:concept.agent-receiptsdev:project.aws-bedrockdev:technology.mcpradar:agentgate-signed-agent-receiptsradar:traceseal-signed-agent-receiptsradar:agent-lens-v030-tracingradar:trustnotch-verifiable-agent-logsradar:fentaris-governed-mcp-proxy
queries asked of Scott's wikis
  • MCP proxy observability and tool-call tracing
  • byte-exact agent execution provenance
  • hash-chained logs external anchoring and truncation
  • non-blocking interception latency for agent tools
  • agent data exfiltration and tool-response lineage
  • coding-agent harness audit trails

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady1 platformsage 483h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-21 12:37⭐ origin directly observedShow HN: Callwitness – record what your AI agent's tools return
aditichaudharyy on hacker news
—
09-21 12:37amplified on hacker news 👑hn.story.49786506
aditichaudharyy
peak 1 · 1 comments · 98% of case engagement
09-21 13:20our radar first saw it · +0.7hdiscovery anchor: hn.story.49786506—
pace: p23 vs 1032 stories at the 336h mark (now 483h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Show HN: Callwitness – record what your AI agent's tools return
Retrieved article excerpt

Open article · Retrieved 2026-09-21T13:22:14.470329+00:00

callwitness

**Record what your AI agent's tools actually return.**  
Forwards every byte. Blocks nothing. Hash-chains the log.

[tests](https://github.com/AditiChaudharyy14/callwitness/actions/workflows/tests.yml)
[PyPI](https://pypi.org/project/callwitness/)
[Python 3.8+](https://pypi.org/project/callwitness/)
[Dependencies: zero](https://camo.githubusercontent.com/c60c09614c8f97cb5b07e69a494f65db8f8b0d1a1a0664c59413db5fa7bc5112/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570656e64656e636965732d7a65726f2d313832363436)
[License: MIT](https://github.com/AditiChaudharyy14/callwitness/blob/main/LICENSE)

[Website](https://callwitness.tech) ·
[PyPI](https://pypi.org/project/callwitness/) ·
[Security](https://github.com/AditiChaudharyy14/callwitness/blob/main/SECURITY.md) ·
[Issues](https://github.com/AditiChaudharyy14/callwitness/issues)

A recorder for AI agents. It sits between an agent and its tools, forwards every
byte unchanged, blocks nothing, and hash-chains every record it writes -- so a log
its operator could have edited still proves what happened. Works with any MCP
server today.

No dependencies. Python 3.8+. MIT.

Your agent talks to MCP servers through Callwitness, which forwards every byte unchanged and writes a copy to a hash-chained log on your machine

---

## Thirty seconds

```
pip install callwitness
callwitness demo
```

If `pip` is not a command on your machine, `python -m pip install callwitness`
does the same thing.

If the install finishes with a warning that the script went somewhere **not on
PATH**, the `callwitness` command will not exist. Run it as a module instead --
same tool, same output:

```
python -m callwitness demo
python -m callwitness last
```

No agent, no API key, nothing to configure. It starts a real MCP server through
the recorder, calls the tools that read like reads, refuses the ones that don't,
shows what came back, and tells you where your responses sit against 140 calls
measured across 65 public servers.

```
  npx -y @modelcontextprotocol/server-everything declared 8 tools.
  6 read like reads; 2 refused by the safety rule and never called.

    echo                              41 B
    get-resource-reference           369 B
    get-structured-content           187 B

  6 calls recorded. Nothing was blocked, nothing was altered.

  Your calls against the public baseline (65 servers, 140 calls)

    server-everything/echo            41 B   p12  vs same tool    3x smaller
    server-everything/get-resource   369 B   p61  vs same tool    about typical
```

Point it at your own server instead:

```
callwitness demo -- npx -y @your/server
```

---

## The thing it shows you

Same tool. Same permission. Two very different actions:

```
      51B  send_email   [email protected]
    20085B send_email   exfil.example.net, [email protected]
```

An allowlist cannot tell those apart — the agent is permitted to send email in
both cases. The difference is *how much* is leaving and *where it is going*, and
those are the two signals Callwitness records on every call.

## Why it blocks nothing

Because it should be installable in production on a Tuesday afternoon.

Callwitness cannot corrupt what an agent sends or receives: it relays every message
whether or not it can parse it, and every write to storage is wrapped so a
recorder bug can't reach the stream. That property is tested, not asserted —
see `tests/test_passthrough.py`, which asserts the proxied output is
byte-identical to running the server directly.

**It cannot delay, either.** Observation runs on its own thread behind a bounded
queue, so the relay only ever does a non-blocking hand-off. A deliberately
half-second-slow observer moves the gap between two forwarded messages by 0.05ms
— it used to move it by 4.2 seconds. The queue drops rather than growing without
limit under load, and counts what it dropped: unrecorded data nobody can see is
worse than data that was never collected.

It matters because the security industry is currently writing rules against
agent failures nobody has measured. Enforcement without data is guessing with
extra steps. Collect first.

## Use

Wrap any stdio MCP server:

```
callwitness run --echo -- npx -y @modelcontextprotocol/server-filesystem /data
```

Or let it wrap the servers you already have. It finds your client's config,
shows you exactly what would change, and writes nothing until you say so:

```
$ callwitness install

/Users/you/Library/Application Support/Claude/claude_desktop_config.json
  filesystem
    - npx -y @modelcontextprotocol/server-filesystem /data
    + callwitness run --label filesystem -- npx -y @modelcontextprotocol/server-filesystem /data
  git
    - uvx mcp-server-git --repository /repo
    + callwitness run --label git -- uvx mcp-server-git --repository /repo
  remote-api  SKIPPED: remote server -- needs `callwitness proxy --upstream
              https://mcp.acme.com/mcp --port <port>` and a port you choose

2 servers would be wrapped. Nothing has been changed.
```

`--apply` writes it, after a timestamped backup. `callwitness uninstall --apply`
puts everything back. Running install twice does nothing the second time.

Dry-run is the default because this edits a file you did not write and a broken
MCP config means a broken agent — the one outcome this whole tool promises not
to cause. Anything it does not recognise is skipped and named rather than
guessed at.

Knows about Claude Desktop, Cursor, Windsurf, Claude Code, and project-local
`.mcp.json` / `.vscode/mcp.json`. A single config can hold more than one server
map — Claude Code keeps a global one and a separate one per project under
`projects.<path>.mcpServers` — and every one of them is walked. When the answer
is "nothing to wrap", install prints every location it checked, because *no
config exists*, *no servers in it* and *already wrapped* are three different
situations and only one of them means you are finished.

If yours lives somewhere else:
`callwitness install --config /path/to/mcp.json`.

### Remote servers

Production agents mostly talk to remote MCP servers over Streamable HTTP. Put
Callwitness in front of one and point the client at the local address instead:

```
callwitness proxy --upstream https://mcp.example.com/mcp --port 8100 --echo
```

```
{
  "mcpServers": {
    "example": { "url": "http://127.0.0.1:8100/mcp" }
  }
}
```

POST, the SSE response stream, the server-initiated `GET` stream and session
teardown are all relayed verbatim, headers included, so the `Mcp-Session-Id`
handshake works without Callwitness understanding it. Both transports share one
recorder (`CallTracker`), so a row looks the same whichever produced it.

The agent behaves exactly as before. Then look at what it did:

```
callwitness last             # what happened in the last run: failures, the biggest responses
callwitness cost --since 7d  # what your tools returned, estimated in tokens
callwitness tail --errors    # only the calls that failed
callwitness stats            # per-tool volume, errors, latency, destinations
callwitness verify           # check nothing has been altered since it was written
callwitness export out.jsonl # everything, for analysis
```

### What happened in that run

`stats` answers which tools exist, which is nobody's question. The two people
actually have are *what broke* and *what was enormous*, and both are about one
run rather than every run ever recorded.

```
$ callwitness last

  fetch   17 Sep 05:56 -> 06:02   open, idle 1h
  2 calls, 0 failed, 71.0 KB returned, 3.4 s in tools

  nothing failed

  biggest responses
    fetch                          68.6 KB      2.1 s  17 Sep 06:01
    fetch                           2.4 KB      1.2 s  17 Sep 06:02

  where it went
    callwitness.tech                 2

  Every call:  callwitness tail --session 6d4e47f7
```

Failures come first, then the biggest responses, then anything slow that was
not already listed. If nothing failed it says so in one line and moves on.

`callwitness last --runs` lists recent runs, and is careful about what it
claims. A session with no recorded end reads *still running* only while calls
are still arriving; after an hour of silence it reads *open, idle 3d*. That is
the honest form — the process may have been killed, or it may be sitting there
alive and unused, and the record cannot tell you which. Saying *no end
recorded* invited the first conclusion; saying *still running* asserted the
second. Knowing for certain would mean storing a pid and testing liveness, and
on Windows the obvious test terminates the process rather than checking it.

The arrow points at the last call rather than at now, for the same reason:
nothing is known to have happened after it.

Then drill in. `--session` takes the id `last` prints, and matches on a prefix:

```
callwitness tail --session 6d4e47f7
callwitness tail --errors
```

`--errors` is the one worth running on a bad day:

```
  13 Sep 08:19  ERR  tool_get_definition     in=23   out=0 B    16.2 s
  13 Sep 08:21  ERR  feed                    in=38   out=0 B    37.3 s
  13 Sep 08:25  ERR  get_gcores_new          in=2    out=0 B    21.2 s
```

Three calls that took between sixteen and thirty-seven seconds to return
nothing. An agent waits that out and moves on without saying anything, and the
same rows sit invisibly in the middle of a 180-row `stats` table. Those three
are from the census sweep of 86 public servers rather than from my own agent --
`--errors` spans the whole database, which is the point of it.

### What it is costing you

The recorder measures bytes because bytes are a fact. Nobody budgets in bytes.

```
$ callwitness cost --since 7d

  tool                          calls   returned      ~tokens   worst call
  deepwiki_fetch                    3     1.4 MB     ~366,246     684.9 KB
  list_directory                    4     1.0 MB     ~251,216     981.3 KB
  fetch                             5    71.8 KB      ~18,374      68.6 KB
  ... and 130 more tools

  total                           222     3.8 MB   ~1,001,351
```

The ratio is an estimate, it is printed on every run, and it is a flag
(`--bytes-per-token`), because a constant nobody can see is a constant nobody
can correct. Four bytes per token is a middle figure for JSON — worse with
dense punctuation or non-ASCII, better for prose.

Failed calls are excluded: returning nothing costs no context, however long it
took. Those belong in `tail --errors`, where they are.

And it counts what tools returned and nothing else — not the schemas your
client sends at startup, not your prompts, not the model's replies. Someone
will hold this next to their bill, so it says so itself.

### What it cannot see

Callwitness is a proxy for MCP traffic. That is the whole of what it records.

An agent doing work through its own built-in tools is invisible here. Claude
Code, for instance, reads files and runs commands natively and only reaches for
an MCP server when it has no built-in equivalent — so wrapping its servers
records the edges of what it does, not the middle. Agents that are MCP-driven
by construction, which is most custom agents and most Cursor or Windsurf setups
with real servers wired up, are recorded completely.

Direct API calls made inside agent code are the same blind spot and would need
an SDK wrapper, which is deliberately not in v1.

### Where your numbers sit

A recording tells you what happened. It cannot tell you whether what happened is
normal — a number is only unusual relative to something, and on day one your own
history is empty, which is exactly when you are deciding whether this tool is
worth keeping.

```
callwitness baseline --compare
```

```
  npm:mcp-deepwiki/deepwiki_fetch      684.9 KB  p100  vs same tool  11x the median for this tool
  npm:mcp-trends-hub/get_bbc_news       13.5 KB  p100  vs same tool  high end
  npm:@primeng/mcp/version               1.0 KB  p100  vs same tool  high end
```

The reference is the published census — the
aditichaudharyy11

Interpretation history

Decision trace