2026-10-11 16:37 UTC

HarnessRouter claims System One Harness 0.3.1 turns Jev into a traced agent loop using finite typed actions and confidence gates, enabling low-latency automation without generated control text.

state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-harnesses typed-decisions agent-runtimeHarnessRouterTypeSafe

What is this?

System One Harness 0.3.1 is an open-source agent runtime from HarnessRouter that wraps Jev β€” a non-generative 'System One' model from TypeSafe AI (released Sept 2026) which returns calibrated probabilities over a finite, environment-declared action space instead of generating text. The harness executes a traced loop: observe state, compile typed questions, gate each decision by configurable confidence thresholds, execute the chosen action, and record a complete structured trace across Python, MCP, browser, and real-time environments. The publisher reports ~200 ms per model step on small deterministic order-fulfilment tasks at very low cost, but acknowledges narrow benchmarking; no independent reproduction, broader adoption, or evidence on ambiguous workloads exists yet.

Why it matters to Scott

HarnessRouter's System One Harness 0.3.1 is a third-party open-source controller implementing the exact typed-decision harness pattern Scott has been advocating (cheap-model-front-door) and running in production across multiple projects (all-in-one-software's guarded mail front door, Venture World's stage manager, dev-wiki's review screener). It independently arrives at the architecture Scott's frameworks prescribe: finite typed actions, confidence gates per irreversibility gradient, complete traces as decision attestation packages, and code-first control flow over generated text. This gives Scott a dated-receipts opportunity and a concrete artifact to evaluate against his trace-backed agent comparison framework.
dev:project.jevdev:technology.typesafe-jevdev:concept.cheap-model-front-doordev:concept.trace-backed-agent-comparisondev:concept.deterministic-agent-control-planeip:framework.decision-authority-infrastructureip:framework.code-first-architectureip:framework.agent-native-computingdev:project.all-in-one-softwaredev:project.venture-worldradar:typesafe-jev-structured-decisionsradar:system-one-lite-typed-decisionsradar:verdict-local-jev-compatible-decisionsradar:typesafe-macos-computer-useradar:jevman-pacman-decision-model-benchmarkradar:proofrun-local-agent-verification-receiptsradar:traceseal-signed-agent-receiptsradar:conduct-tool-call-guardrails
queries asked of Scott's wikis
  • agent-harness runtime patterns: typed-action loops vs generated-control-text
  • system-one vs system-two model routing: classifier-as-reflex-layer under LLM planner
  • confidence-gate design: per-action-type thresholds (read/write/destructive/finish)
  • local-inference economics: sub-cent per decision, ~200ms latency on small models
  • verification workflows: complete traces as reproducible policy decisions
  • open-weight/non-generative model integration: Jev/TypeSafe in agent stack

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 502h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-20 18:10⭐ origin directly observedShow HN: System One Harness (SOH), the harness for System One models
kuanzema on hacker news
β€”
09-20 21:23first on github (echo) Β· first seen by us Β· +3.2hThe released harness uses one model call per step, enumerable actions, confidence gates, and complete traces; its published live measurement
HarnessRouter
β€”
09-20 18:10amplified on hacker news πŸ‘‘hn.story.49778358
kuanzema
peak 1 Β· 1 comments Β· 98% of case engagement
09-20 18:20our radar first saw it Β· +0.2hdiscovery anchor: hn.story.49778358β€”
pace: p23 vs 1032 stories at the 336h mark (now 502h old) β€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Show HN: System One Harness (SOH), the harness for System One models
Retrieved article excerpt

Open article Β· Retrieved 2026-09-20T18:22:55.734858+00:00

# [System One Harness](https://github.com/HarnessRouter/SystemOneHarness/blob/main/.github/images/systemone-harness-logo.png) The harness for System One models.

[GitHub stars](https://github.com/HarnessRouter/SystemOneHarness "Star System One Harness on GitHub")
[License: Apache 2.0](https://github.com/HarnessRouter/SystemOneHarness/blob/main/LICENSE)
[Version 0.3.1](https://github.com/HarnessRouter/SystemOneHarness/blob/main/pyproject.toml)
[UHP conformance: Core](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/uhp-conformance-core-2026-09-19.json)
[Python 3.10 or newer](https://github.com/HarnessRouter/SystemOneHarness/blob/main/pyproject.toml)

**Turn a System One decision model into an agent loop.** System One Harness observes an environment, compiles its finite action space into typed questions, gates each decision by confidence, executes the chosen action, and records the complete trace.

One model call per step. No generated actions. A probability on every transition.

[Jev uses System One Harness to play a live browser game by choosing one typed action per step.](https://github.com/HarnessRouter/SystemOneHarness/blob/main/.github/images/system-one-browser-gameplay-demo.gif)
  
**Jev playing a live browser game through System One Harness** β€” one typed decision per step, with no generated control text.

The first supported model is [Jev](https://typesafe.ai) by TypeSafe, available through OpenRouter or TypeSafe directly.

[Help build the System One ecosystem. Star this repo.](https://github.com/HarnessRouter/SystemOneHarness "Star System One Harness on GitHub")

Tip

**Start here:** [Run the example](https://github.com/HarnessRouter/SystemOneHarness#quickstart) Β· [Understand the loop](https://github.com/HarnessRouter/SystemOneHarness#how-it-works) Β· [Connect an environment](https://github.com/HarnessRouter/SystemOneHarness#connect-an-environment) Β· [Read the design](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/design.md)

## Quickstart

```
git clone https://github.com/HarnessRouter/SystemOneHarness.git
cd SystemOneHarness
pip install -e .

export OPENROUTER_API_KEY=sk-or-...   # or TYPESAFE_API_KEY=...
s1 run --env order:ship_fastest_gift
```

This runs the built-in order fulfilment environment against the live model:

```
goal: Order B-220 is a gift: note it, then ship it by the fastest carrier.
model: ~typesafe/jev-latest
    0  add_note(note='gift')              p=0.96  241 ms
    1  pick_item(item='scarf')            p=1.00  269 ms
    2  pack()                             p=0.99  166 ms
    3  choose_carrier(carrier='express')  p=0.98  151 ms
    4  ship()                             p=0.93  152 ms

status=completed reason=environment_terminal steps=5 wall=0.98s cost=$0.000208
```

Each step shows the selected action, its weakest required probability, and the model round trip. The final line records how the run ended, how long it took, and what it cost.

## What it provides

|  |  |
| --- | --- |
| **Finite actions** | The model chooses only from actions and parameter values declared by the environment. |
| **Confidence gates** | Read, write, and destructive actions can require different probability thresholds. |
| **Explicit outcomes** | Every run ends as completed, incomplete, failed, or cancelled with a structured reason. |
| **Complete traces** | State, questions, distributions, verdicts, results, latency, and usage are recorded step by step. |
| **Pluggable environments** | Drive Python processes, MCP servers, or web pages with the same controller. |
| **UHP compatibility** | Serve the loop through the Unified Harness Protocol for streaming, continuation, cancellation, and discovery. |

### Real-time environments

Games, live feeds, and other moving environments can return `"realtime": true`. The controller then treats a refused or repeated decision as a clock tick, keeps the model's history short, and lets the last action remain active until it changes.

## How it works

```
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚                        controller                        β”‚
 goal ────► β”‚ observe ─► compile ─► encode ─► decide ─► gate ─► execute β”‚ ────► trace
            β”‚    β–²                                           β”‚         β”‚
            β”‚    └────────────── environment β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

1. **Observe.** The environment reports text, structured fields, candidates, and terminal state.
2. **Compile.** The available actions become typed `choice`, `noul`, and `score` questions.
3. **Encode.** Goal, observation, bounded history, and memory become a state within the model budget.
4. **Decide.** The model answers the action, its parameters, and the goal check in one request.
5. **Gate.** The weakest required probability must clear the selected action's risk threshold.
6. **Execute.** The environment applies the action and returns the next state.

`finish` and `escalate` are actions, not generated prose. The controller always knows why it stopped.

[Read the measured design and architecture β†’](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/design.md)

## Connect an environment

Choose the smallest boundary that fits your system.

| Environment | Use it when | Start with |
| --- | --- | --- |
| **Action space + process** | You own a local program or service loop. | `s1 run --actions actions.yaml --env-cmd "python3 env.py" --goal "..."` |
| **MCP server** | Your tools already expose enumerable inputs over MCP. | `s1 run --mcp "python -m your_server" --goal "..."` |
| **Browser** | The task is expressed through DOM controls in Chrome. | `s1 run --browser --headless --start-url https://example.com --goal "..."` |
| **Python** | You want an in-process integration. | Subclass `Environment` and implement `observe()` and `execute()`. |

### Declare an action space

An action space is YAML or the same structure in Python. Every parameter must be enumerable.

```
instructions: >-
  Move the order to shipped, or cancel it when the goal says so.

actions:
  choose_carrier:
    description: Select a carrier for the packed order.
    risk: write
    params:
      carrier:
        from: available_carriers
  ship:
    description: Hand the packed order to the selected carrier.
    risk: destructive

gate:
  read: 0.5
  write: 0.6
  destructive: 0.8
  finish: 0.5
```

Parameters can use fixed `choices`, observation `candidates`, a boolean `flag`, or ordered `levels`. Free text is rejected because a System One model does not generate text.

[Open the complete example β†’](https://github.com/HarnessRouter/SystemOneHarness/blob/main/systemone_harness/examples/order_fulfilment.yaml)

### Use an MCP server

The harness lists an MCP server's tools, compiles supported schemas into actions, and explains every unsupported tool instead of silently dropping it.

```
pip install -e ".[mcp]"
s1 tools --mcp "python -m systemone_harness.envs.order_mcp"
s1 run --mcp "python -m systemone_harness.envs.order_mcp --scenario ship_fastest_gift" \
  --goal "Order B-220 is a gift. Ship it by the fastest carrier."
```

An `observe` tool provides state. An optional `reset` tool starts a run. Every other compatible tool becomes an action. MCP annotations determine whether the action is read, write, or destructive.

### Drive a browser

The browser environment uses [Browser Use](https://github.com/browser-use/browser-use) to turn visible DOM controls into a finite action space. Text comes from named values supplied by the caller. The model chooses values by name and never writes them.

```
pip install -e ".[browser]"
s1 run --browser --headless --start-url https://example.com/book \
  --text name=Customer --text [email protected] \
  --goal "Book a table at 19:30 with a window seat."
```

[Browser setup, measurements, and limits β†’](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/browser-use.md)

## Serve over UHP

Expose any configured loop as a [Unified Harness Protocol](https://unifiedharnessprotocol.org) server:

```
export OPENROUTER_API_KEY=sk-or-...
s1 serve --api-key choose-a-secret --port 8710
```

```
input                  β†’ goal
function_call          β†’ selected action
function_call_output   β†’ environment result
reasoning              β†’ distribution and gate verdict
previous_response_id   β†’ continued environment and history
```

Streaming emits each item as it happens. Cancellation lets the current model step finish and records it. The included report passes all 40 checks in the UHP `core` conformance class.

[View the conformance report β†’](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/uhp-conformance-core-2026-09-19.json)

## Measured, not implied

Five live runs per scenario on 2026-09-19 with `typesafe/jev-1.13-20260917` through OpenRouter:

| Scenario | Goal met | Mean steps | Mean model latency | Mean wall time | Cost per run |
| --- | --- | --- | --- | --- | --- |
| Ship by cheapest carrier | 5/5 | 6.0 | 241 ms | 1.45 s | $0.000265 |
| Ship fastest and add gift note | 5/5 | 5.0 | 199 ms | 0.99 s | $0.000214 |
| Cancel a fraudulent order | 5/5 | 1.0 | 197 ms | 0.20 s | $0.000044 |

The benchmark proves the controller, compiler, gate, and model can complete these small deterministic tasks. It does not claim the same result for ambiguous state, arithmetic, dates, or long irrelevant context.

[Inspect the raw benchmark rows β†’](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/bench-2026-09-19.json)

## Command line

```
s1 run    Run one goal against a built-in, process, MCP, or browser environment
s1 tools  Inspect how an MCP server compiles into supported actions
s1 serve  Expose a configured loop as a UHP server
s1 bench  Run the built-in live benchmark
```

Use `s1 <command> --help` for every option. Add `--json trace.json` to `run`, `tools`, or `bench` when you need machine-readable output.

## Test

```
pip install -e . pytest
pytest -q tests
```

The suite covers action compilation, unsupported inputs, state truncation, confidence gates, every terminal reason, cancellation, continuation, MCP, browser actions, and the UHP server. Recorded model answers keep the default suite deterministic and keyless.

## Project status

|  |  |
| --- | --- |
| Version | `0.3.1` |
| Models | Jev through OpenRouter or TypeSafe directly |
| UHP | `core`, 40 of 40 checks |
| Environments | Python, stdio, MCP, browser, and real-time loops |
| Python | 3.10 or newer |

Next milestones are a `systemone` base in [HarnessRouter](https://github.com/HarnessRouter/harnessrouter), skills as loadable actions, the Chrome side panel as a plain UHP client, and the UHP `extended` conformance class.

## Resources

| Goal | Resource |
| --- | --- |
| Understand the model and architecture | [Design of record](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/design.md) |
| Configure the browser environment | [Browser guide](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/browser-use.md) |
| Try the protocol client | [Chrome extension](https://github.com/HarnessRouter/SystemOneHarness/blob/main/extension/README.md) |
| Review benchmark evidence | [Benchmark report](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/bench-2026-09-19.json) |
| Review protocol evidence | [UHP conformance report](https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/uhp-conformance-core-2026-09-19.json) |

## License

System One Harness is licensed under [Apache 2.0](https://github.com/HarnessRouter/SystemOneHarness/blob/main/LICENSE).
kuanzema11
🟧 echo.githubThe released harness uses one model call per step, enumerable actions, confidence gates, and complete traces; its published live measurementHarnessRouterβ€”β€”

Interpretation history

Decision trace