Retrieved article excerpt
Open article ยท Retrieved 2026-09-30T18:42:50.623003+00:00
# RuleReceipt
[CI](https://github.com/rulereceipt/rulereceipt/actions/workflows/ci.yml)
[CodeQL](https://github.com/rulereceipt/rulereceipt/actions/workflows/codeql.yml)
[OpenSSF Scorecard](https://scorecard.dev/viewer/?uri=github.com/rulereceipt/rulereceipt)
[npm](https://www.npmjs.com/package/rulereceipt)
[provenance](https://www.npmjs.com/package/rulereceipt#provenance)
**Check if your AI coding agent followed the rules in your CLAUDE.md, with the exact line as proof.**
[RuleReceipt checking an agent session against your rules โ three rules broken, each with the quoted line](https://rulereceipt.dev)
```
npx rulereceipt
```
Runs entirely on your machine. Plain `rulereceipt check` makes zero network
calls โ [Trust, privacy and licensing](https://github.com/rulereceipt/rulereceipt#trust-privacy-and-licensing) has the full
detail, including the three off-by-default opt-ins. Works with Claude Code today
(OpenAI Codex CLI in testing); reads rules from CLAUDE.md, AGENTS.md, Cursor
(`.cursor/rules`), GitHub Copilot, Windsurf, Gemini (`GEMINI.md`), Google's
`.agents/rules`, and Claude Code memory. [Accuracy](https://rulereceipt.dev/accuracy)
ยท [Known gaps](https://github.com/rulereceipt/rulereceipt/blob/main/KNOWN-GAPS.md) ยท Source-available, not OSI โ see [LICENSE](https://github.com/rulereceipt/rulereceipt/blob/main/LICENSE).
## See it in 10 seconds
```
npx rulereceipt demo
```
No install, no config, no API key, no real session needed โ prints a sample
report so you can see the output shape immediately.
Then, in a project you actually use an agent in:
```
npx rulereceipt
```
With no arguments it runs **history mode**: it checks *every* session for this
project in the last 30 days and leads with the rules broken most โ each with a
count, the last date, and one quoted line from the session. The headline counts
only proven breaks (a structured check with evidence); judgment rules stay on
their own line, so the number never overstates. `rulereceipt check` still checks
one session in full.
To stop it happening again:
```
npx rulereceipt protect
```
Adds a PreToolUse guard and a Stop hook to `.claude/settings.json` โ after
showing you exactly what it will add and asking. `protect --undo` restores the
file byte-for-byte. It's the only place RuleReceipt writes settings, and only
with your yes.
To share the result:
```
npx rulereceipt card
```
Writes a small SVG summary and prints ready-to-post links for X, LinkedIn,
Bluesky and Reddit, plus copy-text for Slack or a PR. It carries **counts
only** โ no code, paths, or rule text (add rule names to the copy-text with
`--show-rules`). Nothing is posted for you and nothing is uploaded.
## Status
Published and live on npm, actively developed.
## How it works
1. Reads your rules and extracts individual ones โ from CLAUDE.md / AGENTS.md,
Cursor / Copilot / Windsurf rule files, and Claude Code memory, across the
current project directory and your global rules file.
2. Reads your most recent agent session transcript โ Claude Code today
(including hosted/enterprise variants under a different directory), and
OpenAI Codex CLI (in testing); newest session across tools wins.
3. Routes each rule to the narrowest check that can actually answer it:
- **Structured checks** read what the session really did โ an actual
git command's branch argument, actual file edits, actual file
operations. These are the only checks that report a confident FAIL,
because they can tell an action from a mention.
- **Literal checks** look for a specific string named in the rule.
Absence is real evidence, so a clean session PASSes. A match reports
UNCLEAR with the text quoted, because a text match alone cannot
distinguish doing the forbidden thing from grepping for it, quoting
it, or naming it in a commit message.
- **Judgment** rules need real understanding (e.g. "surface bad news
first"). With `--llm` each is graded individually using *your own*
Claude key; without it they report UNCLEAR rather than guessing.
- Lines containing no instruction at all โ directory listings,
reference tables, examples โ aren't rules, and are reported as such
instead of being checked. This step is a heuristic over English
instruction words, so it can be wrong in both directions: run
`rulereceipt check --show-skipped` once on your rules file to see
exactly what it excluded. A rule phrased unusually, or written in
another language, can land there โ and a rule dropped silently is
worse than one reported wrongly.
When it gets one wrong, `rulereceipt rules --include <handle>` fixes it
permanently. The handle is a hash of the rule's own text, not its
position, so the correction survives edits elsewhere in the file. That
matters more than making the classifier smarter: imperative verbs are
not a closed class and the word list is English-only, so it will keep
being wrong โ it just needs to be correctable.
4. Prints a report โ terminal table by default, `--markdown` for pasting
into a PR or Slack message, or `--html` for a shareable single file โ
showing what passed, what failed, and a quoted line of evidence for
each. Every report includes a SHA-256 hash of the session file it
checked, so anyone with that file can confirm the report describes
that exact file. (It proves the report matches the file, not that the
file is an unmodified record โ see SECURITY.md.)
## Usage
```
rulereceipt audit # score your rules file for checkability โ NO session needed
rulereceipt check # check the latest session in this project
rulereceipt check --markdown # same, formatted for pasting into a PR/Slack
rulereceipt check --html # write a shareable single-file HTML report you can send
rulereceipt check --html report.html # ...to a specific path
rulereceipt check --show-skipped # list what was treated as documentation and not checked
rulereceipt check --require-session # fail if there's no session, instead of passing silently
rulereceipt check --exit-zero # report failures without failing the build
rulereceipt check --llm # opt-in: grade judgment rules with your own Claude key
rulereceipt check --share # opt-in: send anonymous pass/fail/unclear counts
rulereceipt check --telemetry # opt-in: send one random per-machine ID
rulereceipt check --transcript <path> # check a specific session file
rulereceipt rules # show corrections you've made to what counts as a rule
rulereceipt rules --include <handle> # "this IS a rule" โ check it from now on
rulereceipt rules --exclude <handle> # "this isn't" โ stop reporting it
rulereceipt rules --coverage # which rules a configured hook might actually enforce
rulereceipt doctor # list hooks/auto-run tasks configured on this machine
rulereceipt hook # run AS a Claude Code Stop hook โ block Claude finishing on a broken rule
rulereceipt guard # run AS a Claude Code PreToolUse hook โ refuse a call before it runs
rulereceipt lint # find contradictions between CLAUDE.md and AGENTS.md
rulereceipt digest # summarise recent checks; --email to send it
rulereceipt config # set up email sending (stays on your machine)
rulereceipt demo # sample output, no setup needed
rulereceipt demo --markdown
rulereceipt --version # print the installed version
rulereceipt verify <session-file> <hash> # spot-check a report you received against the real session file
```
`verify` isn't a routine check โ trust your team day to day, same as any status update. It's there for the rare case it actually matters (a dispute, an incident review): give it the session file and the hash printed in the report, and it confirms whether they really match.
### Claims of having read something
A session that writes `PAGES READ: 1-20`, `STATUS: READ IN FULL` or "confirmed
at source" while never opening a file is asserting provenance it does not
have. Reported by a user in [anthropics/claude-code#92505](https://github.com/anthropics/claude-code/issues/92505), where those headers
went into tracked files and commit messages for material the model had never
read.
The check is narrow on purpose. It fires only when **nothing at all** was read
in the session โ no `Read`, no `Grep`, no `cat`. That much a transcript can
prove, and it contradicts any claim of reading. It cannot tell you *which*
document was read when reads did happen, so a session that read the wrong
thing is still beyond it, and the report says so rather than guessing.
"I will read the filing next" is a plan, not a claim, and does not fire.
## Blocking, not just reporting
`rulereceipt check` tells you afterwards. `rulereceipt hook` refuses to let the
session end.
Add this to `.claude/settings.json` โ you add it, we never do:
```
{
"hooks": {
"Stop": [
{ "hooks": [ { "type": "command", "command": "npx rulereceipt hook" } ] }
]
}
}
```
When Claude tries to finish, it reads the session that just happened. If a rule
was broken it hands Claude the rule, the evidence, and instructions to keep
working, so the session cannot end on a claim that isn't backed.
The one it is actually for: *"Done โ all tests pass"* when the last run of
`npm test` returned two failures. It checks the claim against what ran, which
is the part a model cannot talk its way around.
Three properties worth knowing before you wire it in:
- **It blocks two things, both narrow.** A claim a recorded run contradicts,
and a claim of done that nothing in the session verified. Never a judgment
rule, never an LLM opinion. Run against thirteen real sessions it stopped
two, and both were read by hand.
- **The report and the gate disagree in exactly one place.** When a session
claims work is done and nothing recorded verifies it, the report says
"couldn't tell" โ the tests may have run in another terminal, and a
transcript cannot see that. The gate refuses the exit anyway, because it is
not saying the claim is false. It is declining to let "done" end a session
with nothing behind it.
- **It cannot loop.** Claude Code sets `stop_hook_active` when a session is
already continuing because of a block; the hook returns immediately in that
case. One interruption per stop.
- **It fails open.** Unreadable transcript, missing rules file, a bug in us โ
it allows the stop and writes a line to stderr. Failing closed would mean our
bug locks you out of finishing your own session. That is a deliberate
weakening, and it is why `check` in CI stays the backstop.
It runs when Claude stops, so it catches a finished session, not a command
mid-flight. For that, use a `PreToolUse` hook of your own โ `rulereceipt doctor` will show you what you already have.
### Refusing a command before it runs
`rulereceipt guard` runs as a `PreToolUse` hook and refuses a call outright:
```
{
"hooks": {
"PreToolUse": [
{ "hooks": [ { "type": "command", "command": "npx rulereceipt guard" } ] }
]
}
}
```
Read the limit before wiring it in, because it is most of the story. It
enforces rules naming a **file** or a **branch** โ "never modify `.env`",
"never commit to `main`" โ and nothing else.
It does **not** block a banned command unless you have said which command is
banned. That was the point of building it, and the automatic version did not
survive measurement: replaying 16,336 real tool calls against every forbidding
rule in a 559-file corpus, blocking on command literals refused 62.8% of them.
Narrowing twice reached 2.5%, and the residue was still wrong in a way no
matcher fixes โ one rule refused `npm run build` 112 times, because it forbids
running Playwright unprompted and *recommends* `npm run build`, which is its
only command-shaped literal.
Nothing in a rules file marks which backtick is the prohibition