2026-10-11 16:36 UTC

Kyle Clouthier claims RunBoth’s released Python behavior-diff tool detects reproducible changes across seven observation channels without a test suite or AI model, potentially adding a practical regression gate for AI-generated edits while explicitly abstaining on uncheckable functions.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumcoding-agents software-testing regression-testingKyle ClouthierClouthier Simulation Labs

What is this?

The supplied case describes RunBoth as a released Python behavior-diff tool attributed to Kyle Clouthier: it reportedly runs old and new functions on generated inputs, reports behavioral differences or budget-limited no-change results, and abstains on functions it cannot check. The sole web snippet is a selected-work page for Clouthier that broadly emphasizes audited tools and documented test methods; it does not mention RunBoth. Consequently, the release, seven observation channels, no-test-suite/no-model claims, and Clouthier Simulation Labs’ involvement remain unverified by the supplied web results.

Why it matters to Scott

RunBoth’s claimed old/new execution comparison converges with Scott’s Characterisation Testing and AI Legacy Takeover position: if validated, generating behavioral witnesses without an existing suite could lower the cost of establishing regression checks for AI-driven replacements, while abstention preserves explicit coverage limits. This warrants a bounded tool trial rather than adoption or a correctness claim: the supplied web results do not verify its capabilities, and the related radar pages track other verification tools, not this development.
ip:concept.characterisation-testingip:source.ai-legacy-takeoverip:concept.mechanically-different-verifiersip:concept.typed-uncertaintyradar:vise-deterministic-refactor-gatesradar:codeeraser-deterministic-code-judge
queries asked of Scott's wikis
  • coding agent regression gates behavioral verification
  • generated inputs differential testing without test suites
  • AI code edits reproducible failure witnesses
  • verification budgets abstention versus correctness guarantees
  • agent harness deterministic validation without LLM judges

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 670h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-13 18:29 (minted)⭐ origin echo-reconstructedRunBoth executes old and new Python functions on generated inputs and reports behavioral witnesses, budget-limited no-change results, or abs
Kyle Clouthier / Clouthier Simulation Labs on github (echo) · attributed from hn.story.49686484 · published time unknown
—
09-13 17:40first on hacker news · published · lag ?Show HN: RunBoth, a behaviour diff for code an AI changed
kyleclouthier
—
09-13 17:40amplified on hacker news 👑hn.story.49686484
kyleclouthier
peak 2 · 0 comments · 98% of case engagement
09-13 18:21our radar first saw it · lag ?discovery anchor: hn.story.49686484—
pace: p9 vs 1032 stories at the 336h mark (now 670h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: RunBoth, a behaviour diff for code an AI changed
Retrieved article excerpt

Open article · Retrieved 2026-09-13T18:22:26.392664+00:00

[RunBoth](https://raw.githubusercontent.com/runboth/runboth/main/brand/mark-128.png)

# RunBoth

**An AI changed your code. RunBoth runs both versions and tells you what actually behaves
differently, including the functions nobody touched.**

[runboth.dev](https://runboth.dev) ·
[How it works](https://github.com/runboth/runboth/blob/main/HOW_IT_WORKS.md) ·
[Red team results](https://github.com/runboth/runboth/blob/main/RED_TEAM_2026-09-12.md)

---

```
$ git commit -m "refactor: tidy up the rates module"

  BLOCKED: the behaviour changed and your message does not say so.

    rate(100)
      used to:  return 0.1
      now:      return 0.0

  AND 1 function you did NOT touch now behaves differently,
  because it calls what you changed:

    total(2.5, 100)          pkg/invoice.py
      used to:  return 225.0
      now:      return 250.0
```

## What it does

It checks out both versions of your code, generates inputs for every changed function from its
signature and from the constants mined out of its own bytecode, runs both versions in separate
sandboxed subprocesses, and compares seven observation channels. When they disagree it hands you
the exact input that separates them.

No test suite required. No network calls. No AI model. No dependencies.

## Install

```
pip install runboth                                  # once the first release is on PyPI
pip install git+https://github.com/runboth/runboth   # works today
runboth install-hook          # a commit-msg gate, silent unless behaviour moved
```

As a GitHub Action, running on your own runners:

```
- uses: runboth/[email protected]
  with:
    budget: 60
```

## Three verdicts, never two

| verdict | meaning |
| --- | --- |
| `changed` | with a witness: the arguments, the old result, the new result |
| `no_change at budget N` | N generated inputs found no difference across seven channels |
| `abstained` | it could not be checked, and here is the reason |

"Cannot tell" and "no difference" are different claims, and collapsing them into a green check is
how tools end up lying. **RunBoth never says safe.**

## The seven channels

Return value · exception raised · warnings · stdout · stderr · argument mutation · object state.

A narrow definition of behaviour does not under-report, it lies, because whatever sits outside the
definition comes back as `no_change`.

## What it is not

**Not a model checker.** Kani and CBMC translate code into logic, let inputs be unconstrained
symbols, and ask a solver whether a bad state is reachable within a bound. They return a proof.
RunBoth executes real code on concrete values. It finds differences and reproduces them; it
cannot prove absence, and never claims to.

**Not mutation testing.** Mutation testing damages your code to score your test suite. RunBoth
damages nothing; both versions come from your git history, and no test suite is needed.

## Measured

Red-teamed against eight public repositories it had never been tuned on, with an automated oracle
built to catch the tool lying. **2,548 functions, zero false positives.** Full method and numbers
in [RED\_TEAM\_2026-09-12.md](https://github.com/runboth/runboth/blob/main/RED_TEAM_2026-09-12.md).

| repo | layout | functions | abstained | median/commit |
| --- | --- | --- | --- | --- |
| boltons | flat | 268 | 0.0% | 7.9s |
| sqlparse | flat | 230 | 0.9% | 14.3s |
| arrow | flat | 283 | 1.4% | 118s |
| cachetools | src/ | 325 | 1.8% | 60.9s |
| more-itertools | flat | 843 | 4.4% | 102.5s |
| packaging | src/ | 78 | 5.1% | 0.2s |
| pluggy | src/ | 157 | 6.4% | 33.5s |
| tenacity | async | 364 | 14.3% | 192.8s |

An adversarial corpus of 22 functions written specifically to induce false positives (object
addresses in default `repr`, `datetime.now`, unseeded `random`, `uuid4`, `os.getpid`, set
iteration order, mutable defaults, generators, `__file__` paths) produced none.

## Environment variables

| variable | effect |
| --- | --- |
| `RUNBOTH_SKIP=1` | let a commit through without checking it |
| `RUNBOTH_BUDGET` | generated inputs per function (gate default 80) |
| `RUNBOTH_WORKERS` | parallel adjudications, default 4 |
| `RUNBOTH_ALL_PATHS=1` | also check tests, benchmarks, docs and task runners |
| `RUNBOTH_ENGINE` | engine directory, if the hook cannot resolve it |

`git commit --no-verify` also bypasses the gate, and the gate says so itself when it blocks.

## Honest limits

- Function-level checking is **Python only**. Changed files in other languages are named
  explicitly rather than passed over quietly.
- Sampling finds differences; it cannot prove their absence.
- Nondeterministic, too-slow, or unconstructible functions abstain **with a reason**, and are
  never counted as passing.
- The sandbox contains accidents: resource limits, network blocked, filesystem writes blocked. It
  is **not** a security boundary against hostile code, and no pure-Python sandbox is.

## Development

```
pip install -e .
pytest tests/ -q
runboth selftest          # the control suites, half of which must fail
```

## Licence

[FSL-1.1-Apache-2.0](https://github.com/runboth/runboth/blob/main/LICENSE.md). Free for every use except building a competing product, and it
converts to plain Apache 2.0 two years after each release.

Built by Kyle Clouthier at Clouthier Simulation Labs.
kyleclouthier20
🟧 echo.github ⭐RunBoth executes old and new Python functions on generated inputs and reports behavioral witnesses, budget-limited no-change results, or absKyle Clouthier / Clouthier Simulation Labs——

Interpretation history

Decision trace