Retrieved article excerpt
Open article · Retrieved 2026-09-16T06:22:05.837708+00:00
# PatchWing
**Verified bug fixing — not vulnerability discovery.**
PatchWing takes a *known* bug — an advisory, a Semgrep hit, a fuzzer crash on code you
already own — and produces a fix a maintainer can merge in a few minutes: a reproducer that
**fails before the fix and passes after**, the patch itself, the project's own test suite
still green, a **byte-exact rollback proof**, and a container **anyone can re-run** to check
the whole thing by hand.
The patch is the cheap part — a model will write a plausible one on demand. **The product is
the evidence** that the patch closes the specific bug and that removing it brings the bug
back. Nobody merges an AI patch on faith; PatchWing is built to make the review cheaper than
ignoring the finding.
> Finding bugs at scale is a solved-enough problem. Fixing them at the same scale, with proof,
> is not. PatchWing is built for the fixing half.
**What it costs:** the `decompress` CVE below was reproduced, patched, and verified for **≈ $1.28**
on a hosted endpoint (~4¢ on a cheap model tier, ~$0 self-hosted) — under a $25 ceiling it never
approached. The average data breach costs **$4.99M** ([IBM 2026](https://www.ibm.com/reports/data-breach)).
Full breakdown: [Why PatchWing → The economics are lopsided](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md#the-economics-are-lopsided).
[The PatchWing pipeline console — real findings with per-stage progress bars; two closed CVEs fully green.](https://github.com/jaymunshi/patchwing/blob/main/docs/img/ui-console.jpg)
## Documentation
New here? Start with **[Why PatchWing](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md)** for the case, then the
**[User Manual](https://github.com/jaymunshi/patchwing/blob/main/docs/PatchWing-User-Manual.docx)** to install and run.
| Document | What it covers |
| --- | --- |
| **[Why PatchWing](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md)** | The case: finding isn't the bottleneck, verified fixing is — and why nothing else covers it. Read this first. |
| **[The economics](https://github.com/jaymunshi/patchwing/blob/main/ECONOMICS.md)** | The business case on one page: what a fix really costs (~$1.28) against what a breach costs ($4.99M), with an illustrative calculator. |
| **[User Manual (Word)](https://github.com/jaymunshi/patchwing/blob/main/docs/PatchWing-User-Manual.docx)** | The full manual with screenshots — install, configure, run, read the output. |
| [Tutorial](https://github.com/jaymunshi/patchwing/blob/main/docs/TUTORIAL.md) | Install → point it at your own endpoint → add & run a finding → read the bundle. |
| [Worked example](https://github.com/jaymunshi/patchwing/blob/main/docs/WORKED-EXAMPLE.md) | A real CVE (`decompress` Zip-Slip) closed end-to-end, with the actual diff and red→green→revert→red chain. |
| [Evidence package](https://github.com/jaymunshi/patchwing/blob/main/docs/EVIDENCE-PACKAGE.md) | Anatomy of the output bundle — the provenance and the `apply`/`verify`/`rollback` scripts a maintainer runs. |
| [Release notes & field report](https://github.com/jaymunshi/patchwing/blob/main/RELEASE-NOTES.md) | What works today, honest model-performance notes, a failed run written up in full, and what to expect operating PatchWing through an AI agent. |
| [Example bundle](https://github.com/jaymunshi/patchwing/blob/main/examples/decompress-CVE-2026-10732-evidence-bundle) | A complete, real evidence bundle (126 files) you can inspect and re-run offline. |
---
## Status — September 2026
Read this before anything else. It is deliberately un-varnished.
- **The pipeline runs end-to-end.** Three real CVEs have been carried
`ingest → localize → reproduce → patch → verify → package`, each producing a full evidence
bundle (see [Worked examples](https://github.com/jaymunshi/patchwing#worked-examples)).
- **All three completed the full chain** — `red → patch → green → rollback → red again`.
Vite got there last: the verifier first **refused to certify a rollback it could not
actually observe** (`harness_fault`, not a fake red), which surfaced a genuine harness bug
— a long-lived dev-server reproducer reused a stale server across the rollback leg. Once
that was fixed (a pre-rollback daemon reset), Vite's chain closed too. The refusal doing its
job — exposing a real defect instead of papering over it — is the system working as designed.
- **It is still early.** Single-operator development deployment. All three model seats
currently run the *same* model, so the "independent second opinion" is advisory only — and
the tool says so, in the evidence itself. **No upstream pull request has been submitted
anywhere yet**, and the disclosure policy is not finalised. (The licence is: released under
[Apache 2.0](https://github.com/jaymunshi/patchwing/blob/main/LICENSE).)
- **What it does not do yet:** prove a patch is *globally* correct (only that the frozen
reproducer flips and the suite stays green), run a genuinely different-family verifier, or
operate unattended against repositories you do not control.
If a claim here is not backed by an artifact you can recompute, it is a bug in this document.
---
## What it is *not*
- **Not exploit discovery.** Reproducers stay at **Definition A**: a trigger input plus an
*observable boundary violation*. It stops at "the bad outcome happened" and never proceeds
to "and here is what an attacker does next."
- **Not a scanner.** Discovery is an *input*. You run your own fuzzers (AFL++) and sanitiser
builds (ASan/MSan/UBSan) on code you own; PatchWing consumes the crash or the advisory and
fixes it. It never goes hunting on systems the operator does not control.
- **Not autonomous.** A machine never signs off its own patch. The terminal stage is a human
reading a prepared package.
- **Not "auto-remediation."** That term now means dependency version-bumping (GitLab, Snyk,
Dependabot) — swapping to a fixed release *when one already exists*. PatchWing's whole
reason to exist is the bugs where **there is no version to bump to** because nobody has
written the fix yet.
---
## The pipeline
```
ingest validate the finding has enough to work with
localize narrow to the files holding the flaw
reproduce build a reproducer that FAILS on the unpatched code, then FREEZE it
patch generate the fix (a minimal diff) plus, where possible, a regression test
verify in a sandbox: apply, reproducer passes, project suite stays green,
then roll back and prove the bug returns
package assemble the evidence bundle a human will read
review human sign-off — a machine never signs its own work off
```
Each finding is a row in a SQLite database; each stage is a function; progress is a loop.
There is no workflow engine. The deployment target is a customer's VPC with no outbound
network, and shipping a server, a database and a UI into a security review is a liability.
The bar for the output is a single sentence: **a PatchWing PR must be cheaper to review than
to ignore.** If review takes more than a few minutes, we have added load to the maintainers
who are the real bottleneck, not removed it.
---
## How a "green" is actually decided
This is the load-bearing part of PatchWing and the reason its "pass" is worth anything. The
intuitive version — "run the reproducer, check the exit code" — is wrong in two ways that
matter.
### 1. The verdict is read from the reproducer's **output**, never its exit code
`verify` runs the reproducer inside the pod, captures stdout+stderr, and classifies *that
text*. Exit code is secondary, consulted only when there is no marker — because real targets
lie about exit codes: an ASan build under one harness exits `1` on everything, MemorySanitizer
with `abort_on_error=0` exits `0` on a genuine failure, and a missing mount also exits `1`.
Under a naive "non-zero = bug reproduced" rule, *every malfunction reads as a successful
reproduction* — and at rollback the assertion literally *is* "red again," so any malfunction
would satisfy it. A check that cannot fail.
So a red is: **a sanitizer marker plus a matching crash identity** (`DEDUP_TOKEN`), or for an
HTTP-shaped finding, the operator's declared evidence rule firing against the parsed wire
output.
### 2. There are **four** states, never two
| state | meaning |
| --- | --- |
| `confirmed_red` | marker/rule present **and** the crash identity matches the pristine reproducer |
| `confirmed_green` | clean, no marker, and every declared rule was affirmatively reached with a non-vulnerable reading |
| `different_bug` | a marker fired, but the crash identity does **not** match — this is not our bug |
| `harness_fault` | no marker + non-zero exit, or a timeout, or a reading too fast to be real |
`different_bug` and `harness_fault` are **neither red nor green and must never fold into
either.** The whole module exists to prevent one failure mode, stated four different ways in
four different files:
> **Absence must never wear failure's face.** A timeout is not a rejection. Exit-1-with-no-marker
> is not a reproduction. A broken build is not a verdict on the patch. An infrastructure error
> is not a judgement.
Every terminal path that is *not* a real verdict carries `patch_evaluated: false`.
### 3. The chain: red → apply → green → revert → **red again**
All inside one persistent container (a deliberate, narrow exception to
one-container-per-attempt — a fresh pod would reintroduce every variable the rollback exists
to hold still):
1. hash the target file (**before\_patch**)
2. apply the fix, hash again (**after\_patch**)
3. re-hash the *frozen* reproducer files — any change **voids the run** (the "wall")
4. build, run the reproducer, classify → must be `confirmed_green`
5. **only then**, the rollback leg: revert to the exact pre-patch bytes, hash (**after\_revert**),
and assert `after_revert == before_patch` **byte for byte** (the *§6a hash triple*);
rebuild, re-run → must read `confirmed_red` **again**.
There is deliberately **no "re-apply and leave it green" step.** The proof is the chain
itself: it shows both that the fix closes the bug *and* that the fix is *what* closed it
(removing it brings the bug back). The green fix lives in the patch artifact and the evidence
bundle, never in the container's final state, which is torn down.
**Rollback is evidence, not a gate.** If the revert does *not* come back red, that does **not**
reject the patch — the container already ruled green. A green-then-not-red result indicts the
*undo mechanism or the oracle*, not the fix, and is recorded honestly as
`verify_green_rollback_not_red`. The one outcome the whole design exists to produce is
`verify_green_rollback_red` — the full chain.
### 4. The wall
The reproducer's files are frozen under sha256 *before any patch exists* and re-checked after
the patch is applied and again after the suite runs. If any byte changed, the run is void. So
a "pass" cannot have been obtained by weakening the test. Every evidence bundle prints the
frozen hashes so a reviewer can recompute them.
---
## Worked examples
These are real findings closed by the running pipeline. Each is reproduced in `examples/`
with the full evidence bundle; the diffs and readings below are copied verbatim from the
database.
### 1. `decompress` — CVE-2026-10732, Zip Slip via symlink race (full chain) ✅
**CWE-22 · npm · [`kevva/decompress`](https://github.com/kevva/decompress)** — an archive
with two entries at the same path (a symlink to an arbitrary target, then a regular file)
writes the file's content *through* the symlink, outside the output directory. Root cause: the
symlink guard runs concurrently via `Promise.all(files.map(...))`, so it checks for the
symlink *before* the earlier entry has finished creating it.
The fix serialises extraction so each entry is fully written before the next entry's safety
checks run:
```
--- a/node_modules/decompress/index.js
+++ b/node_modules/de