2026-10-11 16:37 UTC

PatchWing maintainer jaymunshi claims the released pipeline packages AI-generated fixes for known bugs with frozen reproducers, passing project tests, and byte-exact rollback checks, potentially reducing maintainer review effort without claiming global patch correctness.

state: seedheat: mediumuncertainty: mediumknownscott: lowagentic-security vulnerability-remediation patch-verificationjaymunshiPatchWing

What is this?

The supplied case describes PatchWing as a bring-your-own-model pipeline whose maintainer, jaymunshi, claims it packages AI-generated fixes for known vulnerabilities with frozen reproducers, project-test results, and rerunnable evidence bundles. It reports three known CVEs completing a fail–patch–pass–rollback–fail chain, including byte-exact rollback checks, without claiming overall patch correctness. None of the supplied web results directly identifies PatchWing or corroborates its release or results; related snippets establish concern about AI submissions consuming maintainer review time, not that PatchWing reduces it.

Why it matters to Scott

PatchWing’s claimed workflow repeats positions Scott already holds in Evaluation-Driven Development and Provenance-Coupled Work, with validation and exact-release restoration already implemented in Superlever. The rollback-to-failing-reproducer check is a concrete comparison point, but the supplied evidence establishes neither reduced review effort nor a reason to change his implementation; related radar episodes track verification receipts and patch-review burden, not PatchWing itself.
ip:concept.evaluation-driven-developmentip:framework.provenance-coupled-workdev:project.superleverradar:proofrun-local-agent-verification-receiptsradar:five-bugs-test-spec-blindspotradar:linux-ai-patch-review-overload
queries asked of Scott's wikis
  • coding agent harness deterministic verification gates
  • frozen tests independent validation agent generated patches
  • reproducible evidence bundles audit trails
  • rollback integrity red green regression testing
  • AI code review burden passing tests versus correctness

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 611h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 06:22 (minted)⭐ origin echo-reconstructedPatchWing reports three known CVEs completing a red–patch–green–rollback–red verification chain with rerunnable evidence bundles; it explici
jaymunshi on github (echo) · attributed from hn.story.49722321 · published time unknown
—
09-16 05:07first on hacker news · published · lag ?PatchWing – verified, reproducible fixes for known CVEs (bring your own model)
jaymunshi
—
09-16 05:07amplified on hacker news 👑hn.story.49722321
jaymunshi
peak 3 · 2 comments · 100% of case engagement
09-16 05:20our radar first saw it · lag ?discovery anchor: hn.story.49722321—
pace: p39 vs 1032 stories at the 336h mark (now 611h old) — ahead of agentsec-static-config-auditing (1.2x), behind anthropic-meta-lawsuit (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnPatchWing – verified, reproducible fixes for known CVEs (bring your own model)
Retrieved article excerpt

Open article · Retrieved 2026-09-16T06:22:05.837708+00:00

# PatchWing

**Verified bug fixing — not vulnerability discovery.**

PatchWing takes a *known* bug — an advisory, a Semgrep hit, a fuzzer crash on code you
already own — and produces a fix a maintainer can merge in a few minutes: a reproducer that
**fails before the fix and passes after**, the patch itself, the project's own test suite
still green, a **byte-exact rollback proof**, and a container **anyone can re-run** to check
the whole thing by hand.

The patch is the cheap part — a model will write a plausible one on demand. **The product is
the evidence** that the patch closes the specific bug and that removing it brings the bug
back. Nobody merges an AI patch on faith; PatchWing is built to make the review cheaper than
ignoring the finding.

> Finding bugs at scale is a solved-enough problem. Fixing them at the same scale, with proof,
> is not. PatchWing is built for the fixing half.

**What it costs:** the `decompress` CVE below was reproduced, patched, and verified for **≈ $1.28**
on a hosted endpoint (~4¢ on a cheap model tier, ~$0 self-hosted) — under a $25 ceiling it never
approached. The average data breach costs **$4.99M** ([IBM 2026](https://www.ibm.com/reports/data-breach)).
Full breakdown: [Why PatchWing → The economics are lopsided](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md#the-economics-are-lopsided).

[The PatchWing pipeline console — real findings with per-stage progress bars; two closed CVEs fully green.](https://github.com/jaymunshi/patchwing/blob/main/docs/img/ui-console.jpg)

## Documentation

New here? Start with **[Why PatchWing](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md)** for the case, then the
**[User Manual](https://github.com/jaymunshi/patchwing/blob/main/docs/PatchWing-User-Manual.docx)** to install and run.

| Document | What it covers |
| --- | --- |
| **[Why PatchWing](https://github.com/jaymunshi/patchwing/blob/main/WHY-PATCHWING.md)** | The case: finding isn't the bottleneck, verified fixing is — and why nothing else covers it. Read this first. |
| **[The economics](https://github.com/jaymunshi/patchwing/blob/main/ECONOMICS.md)** | The business case on one page: what a fix really costs (~$1.28) against what a breach costs ($4.99M), with an illustrative calculator. |
| **[User Manual (Word)](https://github.com/jaymunshi/patchwing/blob/main/docs/PatchWing-User-Manual.docx)** | The full manual with screenshots — install, configure, run, read the output. |
| [Tutorial](https://github.com/jaymunshi/patchwing/blob/main/docs/TUTORIAL.md) | Install → point it at your own endpoint → add & run a finding → read the bundle. |
| [Worked example](https://github.com/jaymunshi/patchwing/blob/main/docs/WORKED-EXAMPLE.md) | A real CVE (`decompress` Zip-Slip) closed end-to-end, with the actual diff and red→green→revert→red chain. |
| [Evidence package](https://github.com/jaymunshi/patchwing/blob/main/docs/EVIDENCE-PACKAGE.md) | Anatomy of the output bundle — the provenance and the `apply`/`verify`/`rollback` scripts a maintainer runs. |
| [Release notes & field report](https://github.com/jaymunshi/patchwing/blob/main/RELEASE-NOTES.md) | What works today, honest model-performance notes, a failed run written up in full, and what to expect operating PatchWing through an AI agent. |
| [Example bundle](https://github.com/jaymunshi/patchwing/blob/main/examples/decompress-CVE-2026-10732-evidence-bundle) | A complete, real evidence bundle (126 files) you can inspect and re-run offline. |



---

## Status — September 2026

Read this before anything else. It is deliberately un-varnished.

- **The pipeline runs end-to-end.** Three real CVEs have been carried
  `ingest → localize → reproduce → patch → verify → package`, each producing a full evidence
  bundle (see [Worked examples](https://github.com/jaymunshi/patchwing#worked-examples)).
- **All three completed the full chain** — `red → patch → green → rollback → red again`.
  Vite got there last: the verifier first **refused to certify a rollback it could not
  actually observe** (`harness_fault`, not a fake red), which surfaced a genuine harness bug
  — a long-lived dev-server reproducer reused a stale server across the rollback leg. Once
  that was fixed (a pre-rollback daemon reset), Vite's chain closed too. The refusal doing its
  job — exposing a real defect instead of papering over it — is the system working as designed.
- **It is still early.** Single-operator development deployment. All three model seats
  currently run the *same* model, so the "independent second opinion" is advisory only — and
  the tool says so, in the evidence itself. **No upstream pull request has been submitted
  anywhere yet**, and the disclosure policy is not finalised. (The licence is: released under
  [Apache 2.0](https://github.com/jaymunshi/patchwing/blob/main/LICENSE).)
- **What it does not do yet:** prove a patch is *globally* correct (only that the frozen
  reproducer flips and the suite stays green), run a genuinely different-family verifier, or
  operate unattended against repositories you do not control.

If a claim here is not backed by an artifact you can recompute, it is a bug in this document.

---

## What it is *not*

- **Not exploit discovery.** Reproducers stay at **Definition A**: a trigger input plus an
  *observable boundary violation*. It stops at "the bad outcome happened" and never proceeds
  to "and here is what an attacker does next."
- **Not a scanner.** Discovery is an *input*. You run your own fuzzers (AFL++) and sanitiser
  builds (ASan/MSan/UBSan) on code you own; PatchWing consumes the crash or the advisory and
  fixes it. It never goes hunting on systems the operator does not control.
- **Not autonomous.** A machine never signs off its own patch. The terminal stage is a human
  reading a prepared package.
- **Not "auto-remediation."** That term now means dependency version-bumping (GitLab, Snyk,
  Dependabot) — swapping to a fixed release *when one already exists*. PatchWing's whole
  reason to exist is the bugs where **there is no version to bump to** because nobody has
  written the fix yet.

---

## The pipeline

```
ingest     validate the finding has enough to work with
localize   narrow to the files holding the flaw
reproduce  build a reproducer that FAILS on the unpatched code, then FREEZE it
patch      generate the fix (a minimal diff) plus, where possible, a regression test
verify     in a sandbox: apply, reproducer passes, project suite stays green,
           then roll back and prove the bug returns
package    assemble the evidence bundle a human will read
review     human sign-off — a machine never signs its own work off
```

Each finding is a row in a SQLite database; each stage is a function; progress is a loop.
There is no workflow engine. The deployment target is a customer's VPC with no outbound
network, and shipping a server, a database and a UI into a security review is a liability.

The bar for the output is a single sentence: **a PatchWing PR must be cheaper to review than
to ignore.** If review takes more than a few minutes, we have added load to the maintainers
who are the real bottleneck, not removed it.

---

## How a "green" is actually decided

This is the load-bearing part of PatchWing and the reason its "pass" is worth anything. The
intuitive version — "run the reproducer, check the exit code" — is wrong in two ways that
matter.

### 1. The verdict is read from the reproducer's **output**, never its exit code

`verify` runs the reproducer inside the pod, captures stdout+stderr, and classifies *that
text*. Exit code is secondary, consulted only when there is no marker — because real targets
lie about exit codes: an ASan build under one harness exits `1` on everything, MemorySanitizer
with `abort_on_error=0` exits `0` on a genuine failure, and a missing mount also exits `1`.
Under a naive "non-zero = bug reproduced" rule, *every malfunction reads as a successful
reproduction* — and at rollback the assertion literally *is* "red again," so any malfunction
would satisfy it. A check that cannot fail.

So a red is: **a sanitizer marker plus a matching crash identity** (`DEDUP_TOKEN`), or for an
HTTP-shaped finding, the operator's declared evidence rule firing against the parsed wire
output.

### 2. There are **four** states, never two

| state | meaning |
| --- | --- |
| `confirmed_red` | marker/rule present **and** the crash identity matches the pristine reproducer |
| `confirmed_green` | clean, no marker, and every declared rule was affirmatively reached with a non-vulnerable reading |
| `different_bug` | a marker fired, but the crash identity does **not** match — this is not our bug |
| `harness_fault` | no marker + non-zero exit, or a timeout, or a reading too fast to be real |

`different_bug` and `harness_fault` are **neither red nor green and must never fold into
either.** The whole module exists to prevent one failure mode, stated four different ways in
four different files:

> **Absence must never wear failure's face.** A timeout is not a rejection. Exit-1-with-no-marker
> is not a reproduction. A broken build is not a verdict on the patch. An infrastructure error
> is not a judgement.

Every terminal path that is *not* a real verdict carries `patch_evaluated: false`.

### 3. The chain: red → apply → green → revert → **red again**

All inside one persistent container (a deliberate, narrow exception to
one-container-per-attempt — a fresh pod would reintroduce every variable the rollback exists
to hold still):

1. hash the target file (**before\_patch**)
2. apply the fix, hash again (**after\_patch**)
3. re-hash the *frozen* reproducer files — any change **voids the run** (the "wall")
4. build, run the reproducer, classify → must be `confirmed_green`
5. **only then**, the rollback leg: revert to the exact pre-patch bytes, hash (**after\_revert**),
   and assert `after_revert == before_patch` **byte for byte** (the *§6a hash triple*);
   rebuild, re-run → must read `confirmed_red` **again**.

There is deliberately **no "re-apply and leave it green" step.** The proof is the chain
itself: it shows both that the fix closes the bug *and* that the fix is *what* closed it
(removing it brings the bug back). The green fix lives in the patch artifact and the evidence
bundle, never in the container's final state, which is torn down.

**Rollback is evidence, not a gate.** If the revert does *not* come back red, that does **not**
reject the patch — the container already ruled green. A green-then-not-red result indicts the
*undo mechanism or the oracle*, not the fix, and is recorded honestly as
`verify_green_rollback_not_red`. The one outcome the whole design exists to produce is
`verify_green_rollback_red` — the full chain.

### 4. The wall

The reproducer's files are frozen under sha256 *before any patch exists* and re-checked after
the patch is applied and again after the suite runs. If any byte changed, the run is void. So
a "pass" cannot have been obtained by weakening the test. Every evidence bundle prints the
frozen hashes so a reviewer can recompute them.

---

## Worked examples

These are real findings closed by the running pipeline. Each is reproduced in `examples/`
with the full evidence bundle; the diffs and readings below are copied verbatim from the
database.

### 1. `decompress` — CVE-2026-10732, Zip Slip via symlink race (full chain) ✅

**CWE-22 · npm · [`kevva/decompress`](https://github.com/kevva/decompress)** — an archive
with two entries at the same path (a symlink to an arbitrary target, then a regular file)
writes the file's content *through* the symlink, outside the output directory. Root cause: the
symlink guard runs concurrently via `Promise.all(files.map(...))`, so it checks for the
symlink *before* the earlier entry has finished creating it.

The fix serialises extraction so each entry is fully written before the next entry's safety
checks run:

```
--- a/node_modules/decompress/index.js
+++ b/node_modules/de
jaymunshi32
🟧 echo.github ⭐PatchWing reports three known CVEs completing a red–patch–green–rollback–red verification chain with rerunnable evidence bundles; it explicijaymunshi——

Interpretation history

Decision trace