2026-10-11 17:14 UTC

LinusInnovator claims the released WTF CLI separates mechanical diff churn from review-worthy changes, flags risky patterns, and records locally executed verification results, reducing human review effort for coding-agent changes without uploading code or requiring an AI model.

state: seedheat: lowuncertainty: highconvergesscott: mediumcoding-agent-observability change-validation local-toolsLinusInnovator

What is this?

The supplied case describes WTF as a CLI released by LinusInnovator to inspect Git changes left by coding agents, separate mechanical churn from changes needing review, flag risky patterns, and record local verification results. Its evidence titles describe a Show HN announcement and a distinction between reported, observed, verified, and unknown evidence; the case claims it needs neither code uploads nor an AI model. None of the supplied web results directly covers WTF or LinusInnovator, so the release, capabilities, and claimed reduction in review effort remain uncorroborated here; the snippets instead describe other tools and approaches to agent-change review and verification.

Why it matters to Scott

WTF’s claimed evidence-state separation converges with Scott’s Discussed Is Not Deployed framework, while its Git-change filtering and local verification offer a concrete candidate to evaluate alongside Superlever’s deterministic validation of Codex worktrees—not evidence that review effort has actually fallen. The release and capabilities remain uncorroborated in the supplied material; the radar already tracks related approaches in ProofRun and CodeEraser, but no supplied page tracks WTF itself.
ip:framework.discussed-is-not-deployedip:concept.deterministic-ai-pendulumdev:project.superleverradar:proofrun-local-agent-verification-receiptsradar:codeeraser-deterministic-code-judgeradar:concept.agent-verificationradar:concept.code-review
queries asked of Scott's wikis
  • coding-agent harness independent change verification
  • reported observed verified evidence provenance
  • human review bottleneck mechanical diff filtering
  • deterministic local tools versus model-based review
  • agent change risk routing verification artifacts

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 510h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-20 10:22 (minted)⭐ origin echo-reconstructedWTF inspects resulting Git changes rather than agents, distinguishes reported, observed, verified, and unknown evidence, and offers local te
LinusInnovator on github (echo) · attributed from hn.story.49774320 · published time unknown
—
09-20 10:05first on hacker news · published · lag ?Show HN: WTF > Auto-check what your coding agent changed
Linusinnovator
—
09-20 10:37first on r/LocalLLaMA · published · lag ?WTF is for me, to know what the agent just did.
Disrupt-Linus
—
09-20 10:05amplified on hacker newshn.story.49774320
Linusinnovator
peak 3 · 1 comments · 21% of case engagement
09-20 10:37amplified on r/LocalLLaMA 👑reddit.post.1wldnu2
Disrupt-Linus
peak 0 · 15 comments · 45% of case engagement
09-21 09:48amplified on r/LocalLLaMAreddit.post.1wm892x
Disrupt-Linus
peak 3 · 8 comments · 33% of case engagement
09-20 10:21our radar first saw it · lag ?discovery anchor: hn.story.49774320—
pace: p58 vs 1032 stories at the 336h mark (now 510h old) — ahead of agent-iap-credential-brokering (1.0x), behind anthropic-fourth-cyber-incident-review-miss (0.9x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: WTF > Auto-check what your coding agent changed
Retrieved article excerpt

Open article · Retrieved 2026-09-20T10:22:20.811587+00:00

# WTF

[License: MIT](https://github.com/LinusInnovator/wtf/blob/main/LICENSE)
[Security Policy](https://github.com/LinusInnovator/wtf/blob/main/SECURITY.md)
[Zero Dependencies](https://github.com/LinusInnovator/wtf#local-by-default-paranoid-by-design)
[Gauntlet: 100/100](https://github.com/LinusInnovator/wtf#contributing)

### Your coding agent says it’s done. WTF checks.

Any agent. Any Git repo. Local. No account. No AI required.

WTF works with any coding agent (Claude Code, Cursor, Copilot, Codex, Aider) because it inspects the resulting software change, not the agent.

```
npx agent-wtf
```

---

```
$ npx agent-wtf
WTF — what just happened?
2 files changed · +7 / -5

VERIFIED
  ○ Tests not yet run · Run wtf verify to validate
    tests (npm test)

PAY ATTENTION
  1. TESTS
     Test skipped or disabled
     test/charge.test.js:9
     > test.skip('handles VIP coupon cap calculation', () => {

ALSO
  ⚠ 1 skipped test
  ⚠ 1 debug statement (console.log)

Review surface:
  12 lines to review across 2 files
```

Then:

```
$ npx agent-wtf verify
WTF — what just happened?
2 files changed · +6 / -5

VERIFIED
  ✓ tests       (254ms)

Review surface:
  11 lines to review across 2 files
```

```
Agent finishes
     ↓
    WTF
     ↓
  Evidence
     ↓
Agent fixes
     ↓
WTF verify
     ↓
Human gets receipt
```

---

## Why WTF?

Coding agents can generate more code in two minutes than you can review in an afternoon.

When an agent claims: *"Done! Refactored the billing module and all tests pass."* — humans are left with three questions:

1. **What actually happened?**
2. **Did it actually work, or did the agent just say it did?**
3. **Where do I actually need to look?**

WTF gives you the answers in under a second.

---

## Compression is the Product

Most agent diffs are dominated by lockfiles, minified bundles, snapshots, and generated boilerplate.

WTF separates mechanical churn from code that actually deserves human attention:

```
  3,812 changed lines
          ↓
         WTF
          ↓
  94 meaningful lines to review (97.5% compressed)
```

You review what matters. WTF accounts for the rest.

---

## Why Not Just `git diff`?

| Standard Tooling | What Happens with Coding Agents | How WTF Solves It |
| --- | --- | --- |
| **`git status`** | Lists modified files, but treats a 2,000-line lockfile the same as an auth timeout modification. | **Separates signal from churn**: Classifies mechanical lines vs. meaningful lines deserving human review. |
| **`git diff`** | Floods your terminal with generated boilerplate, snapshots, and minified bundles. | **Focuses human attention**: Automatically highlights high-risk patterns (auth, DB migrations, env vars, debug leftovers). |
| **Agent Claims** | Believes the agent when it claims *"refactored billing module and all tests pass"*. | **Verifies independently**: Flags skipped or disabled tests (`test.skip`) and generates an unforgeable local execution receipt (`wtf verify`). |



---

## Try It in 5 Seconds

No installation required:

```
npx agent-wtf
```

Or install globally:

```
npm install -g agent-wtf
```

### The 5 Essential Commands

| Command | What it does |
| --- | --- |
| `wtf` | See what changed and what needs attention in your working tree (~50ms). |
| `wtf init-agent` | Automatically configure your repository for autonomous agent self-auditing. |
| `wtf verify` | Discover and run your tests/builds to produce an independently verified receipt. |
| `wtf show` | View exact diff snippets and line evidence for every finding. |
| `wtf --json` | Machine-readable evidence schema (`wtf/0.1`) for coding agents. |



---

## Make Your Agent Check Itself Autonomously

Run this once in any repository:

```
npx agent-wtf init-agent
```

This automatically configures your repository's agent rules (`AGENTS.md`, `CLAUDE.md`, `.cursorrules`, and `.github/copilot-instructions.md`).

From that moment on, whenever Claude Code, Cursor, Copilot, or Cline works in your repo, the agent autonomously:

1. **Runs WTF** before declaring completion.
2. **Catches shortcuts**: Detects its own skipped tests (`test.skip`), debug leftovers (`console.log`), and schema risks.
3. **Executes tests**: Runs `wtf verify` to independently validate your test suite.
4. **Hands you proof**: Attaches the unforgeable verification receipt directly to its final reply before you review.

*(To view the markdown template without modifying files, pass `npx agent-wtf init-agent --print`).*

---

## Local by Default. Paranoid by Design.

WTF is designed to inspect machine-generated changes, so it treats repository content as untrusted input.

- **No code uploads**: Zero code or diffs ever leave your machine.
- **No telemetry**: Works completely offline with zero tracking or background pings.
- **No account or API key**: No signup, no LLM tokens, no monthly bill.
- **No required AI model**: Fast, local deterministic analysis.
- **No shell-based Git commands**: Direct binary spawning (`shell: false`) with baseline Git configuration overrides.
- **Repository filesystem containment**: Enforces realpath containment to prevent symlinks from escaping the repository.
- **Terminal control-sequence sanitization**: Strips ANSI cursor escapes, OSC sequences, and Unicode Bidi controls.
- **Zero runtime npm dependencies**: Pure ESM package with 0 runtime dependencies, reducing third-party supply-chain exposure.

Normal `wtf` analysis does not intentionally execute project code.

`wtf verify` is different: it runs your project’s verification commands locally with your user permissions and is not sandboxed. Only use it on code you trust to execute.

See [SECURITY.md](https://github.com/LinusInnovator/wtf/blob/main/SECURITY.md) for details.

---

## Epistemic Integrity

WTF is an evidence ledger, not an oracle.

We strictly avoid fabricated confidence scores (e.g. "87% safe" or "clean code guarantee"). Instead, WTF categorizes facts into four strict evidence tiers:

- **`REPORTED`**: What something claims happened (e.g., an agent summary).
- **`OBSERVED`**: What WTF directly confirmed in the Git diff (e.g., session timeout altered, `.env` introduced).
- **`VERIFIED`**: What WTF independently executed and validated (e.g., test runner exited code 0).
- **`UNKNOWN`**: What available evidence cannot prove (e.g., tests exist but have not been run).

WTF does not claim to catch every bug or replace human judgment. It eliminates the blind spots between what the machine claimed and what the machine actually did.

---

## Contributing

```
git clone https://github.com/LinusInnovator/wtf.git
cd wtf
npm install
npm run build
npm test
npm run gauntlet
```

---

Built by [@LinusInnovator](https://github.com/LinusInnovator). Explored in depth at [Great Delights](https://great.delights.pro/ai-patterns).

## License

[MIT](https://github.com/LinusInnovator/wtf/blob/main/LICENSE)
Linusinnovator31
🟧 echo.github ⭐WTF inspects resulting Git changes rather than agents, distinguishes reported, observed, verified, and unknown evidence, and offers local teLinusInnovator——
🟠 redditWTF is for me, to know what the agent just did.
LocalLLaMA
Disrupt-Linus015
🟠 redditWTF is a small overview check-in remedy (attempt) for lying agents. And more
LocalLLaMA
Disrupt-Linus28

Interpretation history

Decision trace