2026-10-11 17:14 UTC

Orcrist maintainer simone20a claims its released desktop coding agent lets a stronger model author a validated per-task finite-state-machine harness for a smaller or local executor, making workflow checks, retry budgets, and failure paths explicit rather than relying on the executor's reasoning.

state: seedheat: mediumuncertainty: mediumconvergesscott: highagent-harnesses coding-agents local-inferencesimone20a

What is this?

Orcrist is an open-source project by maintainer simone20a (GitHub), consisting of a DSL for state machines whose states are executed by an LLM, plus a desktop coding agent that runs on it. Per the README, given a task, the agent first writes an Orcrist machine for that specific task โ€” grounded in the repo's grammar and authoring guide โ€” then executes it state by state, with the machine's guards and declared reports determining transitions until a final state. The stated motivation is that explicit machines force failure paths, retry budgets, and escalation states to be declared up front, unlike a to-do list. Note: the supplied snippets directly confirm the per-task FSM-authored-harness design, but the specific hypothesis claim that a *stronger authoring model* writes the harness for a *smaller or local executor* is not established by the visible snippet text โ€” that split is part of the case's framing and would need the repo docs or the evidence objects themselves to verify.

Why it matters to Scott

Orcrist is an independent builder arriving almost exactly where Scott's Generative Pendulum and Model Barbell arguments land: a stronger model pays the judgment cost once by authoring a per-task deterministic FSM, then a cheaper executor runs inside its declared guards, retry budgets, and failure paths โ€” the scout-senior split realized as a control-flow artifact rather than a transcript handoff. That makes it a dated receipt for positions he has published, and a useful live contrast for his own ask terminal agent, whose complex-work approval is behavioural/prompted rather than mechanically gated the way an FSM harness would be; the one open question is that the strong-author/local-executor split is the case's framing, not snippet-verified, so the local-inference economics angle needs the repo docs before he cites it.
ip:source.the-generative-pendulum-ebookip:concept.model-barbellip:framework.scout-senior-splitip:concept.deterministic-ai-pendulumdev:project.askradar:concept.agent-harnessesradar:concept.local-inferenceradar:agent6-jailed-state-machine-harnessradar:grapharc-runtime-agent-graph-gates
queries asked of Scott's wikis
  • agent harness as deterministic scaffold vs model reasoning โ€” my position on explicit control flow for coding agents
  • strong-model-author / weak-or-local-model-execute split โ€” architecture notes, cost/latency arguments, local inference projects
  • finite state machine DSL or runtime experiments โ€” anything I've built or argued about declarative agent workflows, guards, retry budgets
  • coding agent projects and harness code โ€” my own agent runtime, tool-loop design, where I put failure handling
  • local inference economics โ€” when a small local model is viable as an executor given the right scaffolding
  • agent memory / agent-maintained wikis โ€” compiling learned structure into reusable, validated artifacts vs re-deriving plans per task

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 463h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-22 10:22 (minted)โญ origin echo-reconstructedOrcrist is a desktop coding agent whose harness is written fresh for every task as a finite state machine by a stronger authoring model, the
simone20a on github (echo) ยท attributed from hn.story.49798500 ยท published time unknown
โ€”
09-22 09:24first on hacker news ยท published ยท lag ?Orcrist: A desktop coding agent whose harness is written fresh for every task
simone20a
โ€”
09-22 09:24amplified on hacker news ๐Ÿ‘‘hn.story.49798500
simone20a
peak 2 ยท 0 comments ยท 98% of case engagement
09-22 10:20our radar first saw it ยท lag ?discovery anchor: hn.story.49798500โ€”
pace: p23 vs 1032 stories at the 336h mark (now 463h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnOrcrist: A desktop coding agent whose harness is written fresh for every task
Retrieved article excerpt

Open article ยท Retrieved 2026-09-23T10:22:06.283186+00:00

[Orcrist](https://github.com/simone20a/Orcrist/blob/main/brand/orcrist-logo.png)

**Orcrist** is a desktop coding agent whose harness is written fresh for every task, as a finite state machine, by a stronger model than the one that carries the work out.

Every coding agent has a harness: the loop around the model that hands it tools, decides when a step is done, and says what happens when one fails. Usually that loop lives in the app's source code and is the same for every task, so the only thing that varies with the job is the prompt. Here it is the other way round. The loop itself is the per-task artefact, and it is an object you can read, check and correct before any code is touched.

You give the agent a task. Before touching anything, the **authoring model** writes a machine for it, grounded in the grammar and the authoring guide in this repo. The **execution model** then runs that machine one state at a time: each state's instruction goes to it on its own, it works with real tools, the runtime measures what it can and records what the state is declared to report, and the machine's guards decide where the run goes next. The run ends in a `final` state, success or otherwise.

Splitting the two is the point. Deciding what the phases are, what would prove each one works, and what each is allowed to touch is judgment, and it is exactly what a smaller or local model lacks. Writing that down once, as a machine, is what lets the smaller model be the one that does the work.

A machine also has to say what happens when things go wrong, which a prompt does not. Failure paths, retry budgets and escalation states are all declared before the run starts, and the validator rejects a machine that cannot terminate, that leaves a state unreachable, or that lets the model grade its own work where a command could have settled it. That is what the DSL underneath is for: it is the language the harness is specified in, not the product.

[A run in progress: the transcript on the left, the machine and its store on the right](https://github.com/simone20a/Orcrist/blob/main/screens/session.png)

*A run in progress. The transcript carries the states as they happen; the panel on the right holds the harness itself: the machine, the state now executing with the instruction it was given, and the store, each location marked by who writes it, the agent, a measurement, or a `set`.*

---

## What's in here

```
orcrist/
โ”œโ”€โ”€ src/core/                  the agent: the authoring loop, the executor, the tools, the providers
โ”œโ”€โ”€ src/orcrist/               the language the harness is written in: parser, validator, evaluator
โ”œโ”€โ”€ src/renderer/              the desktop UI
โ”œโ”€โ”€ electron/                  the main process and the preload bridge
โ”œโ”€โ”€ metamodel/
โ”‚   โ”œโ”€โ”€ orcrist.langium        the grammar, ground truth for the authoring step
โ”‚   โ””โ”€โ”€ authoring-guide.md     the criteria the authoring model is told to follow
โ”œโ”€โ”€ examples/*.orc             canonical syntax, shown to the authoring model
โ”œโ”€โ”€ scripts/                   the self-test
โ”œโ”€โ”€ prompts/                   project prompts to run the agent against
โ”œโ”€โ”€ brand/                     the mark, the lockup, the icons
โ”œโ”€โ”€ screens/                   the screenshots in this README
โ””โ”€โ”€ docs/                      how the app works, in depth
```

The grammar, the guide and the examples are read **at runtime**, not compiled in: the app walks up from its own folder until it finds `metamodel/orcrist.langium`. Editing the language therefore changes what the authoring step is grounded in without rebuilding anything, and a checkout with `metamodel/` missing has nothing to author machines against.

## Requirements

- **Node 20 or newer**, with npm
- macOS, Linux or Windows
- An API key for Anthropic or OpenAI, **or** a local [Ollama](https://ollama.com) with a model that supports tool calling

## Install

```
git clone <this repo>
cd orcrist
npm install
```

`npm install` downloads the Electron binary in a postinstall step. If that step fails (see [Troubleshooting](https://github.com/simone20a/Orcrist#troubleshooting)), the rest still installs, and `npm test` and `npm run build` work without it.

## Run

```
npm start
```

That compiles the main process, bundles the renderer and launches the app. For iterative work:

```
npm run dev
```

which runs the Vite dev server with hot reload and points Electron at it.

## Build

```
npm run build          # both halves
npm run build:main     # main process + core, via tsc
npm run build:renderer # renderer, via Vite
```

Output goes to `dist/`: `dist/main/` for the Electron side, `dist/renderer/` for the UI. There is no packaging step yet; `npm start` runs the built app in place.

## Test

```
npm test
```

This parses and validates every `.orc` in `examples/`, checks that a set of deliberately broken machines is rejected for the right reasons, runs a machine end to end against a scripted mock provider, and checks the things a real run depends on: that cancelling a run actually stops the request, that a measured value beats a model's claim, that a state restricted to no tools is handed none, and that the store travels with every state instruction.

It needs no API key and makes no network calls.

## First run

1. Open **Settings โ†’ Providers** and put in an API key. For Ollama there is no key: the models installed on the machine are listed for you under whichever role you point at it.
2. In **Settings โ†’ Models**, choose the two models. They are the two halves described above:
   - **Authoring** writes the harness. Give it the strongest model you have, because it decides the shape of the whole run, what counts as proof that each part works, and what each state is allowed to touch.
   - **Execution** runs each state. This is where a smaller or local model is affordable, because the machine supplies the structure it would otherwise have to hold in its head.
3. Create a project. A project is a name plus a workspace folder; every shell command and file operation is sandboxed to that folder, and the session history lives inside it under `.orcrist-agent/`.
4. Send a task.

[Settings: the authoring and execution model roles, with the locally installed Ollama models listed under the execution role](https://github.com/simone20a/Orcrist/blob/main/screens/settings.png)

*The two roles, set separately. Point one at Ollama and the models installed on the machine are listed underneath it, and clicking one uses it.*

**Settings โ†’ Palette** changes the whole app's colours. Six palettes ship; each is six seed colours and everything else is derived from them.

## Troubleshooting

**"Electron failed to install correctly."** The postinstall download failed on its own (a network hiccup, a proxy, a GitHub rate limit) while the rest of the install reported success. The symptom is `node_modules/electron/` with no `dist/` inside. Re-run just that step:

```
npm run fix:electron
```

If it fails again the error says why. Behind a proxy, set `HTTPS_PROXY` and retry. If GitHub is unreachable or rate-limiting, use a mirror:

```
ELECTRON_MIRROR="https://npmmirror.com/mirrors/electron/" npm run fix:electron
```

**"Could not find metamodel/orcrist.langium."** `metamodel/` is missing from the checkout, or the app was copied somewhere on its own. It looks up the folder chain from its own location, so `metamodel/` and `examples/` belong at the root of the repo, beside `package.json`.

**A run stops at "no machine for this message".** The authoring model decided the task is a single question with no process in it, and said so rather than wrapping one step in ceremony. The reason is printed in the transcript.

## Going deeper

| document | what it covers |
| --- | --- |
| [`docs/how-it-works.md`](https://github.com/simone20a/Orcrist/blob/main/docs/how-it-works.md) | how the harness works: the authoring loop, the state boundary, claims versus measurements, the tools, the type scale and the palette system |
| [`metamodel/authoring-guide.md`](https://github.com/simone20a/Orcrist/blob/main/metamodel/authoring-guide.md) | how to turn a prompt into a machine: the document the authoring model is given verbatim |
| [`metamodel/orcrist.langium`](https://github.com/simone20a/Orcrist/blob/main/metamodel/orcrist.langium) | the language a harness is written in: the grammar, and the constraints the validator enforces on top of it |
| [`brand/README.md`](https://github.com/simone20a/Orcrist/blob/main/brand/README.md) | the mark: why those six colours, and which file to use where |
simone20a20
๐ŸŸง echo.github โญOrcrist is a desktop coding agent whose harness is written fresh for every task as a finite state machine by a stronger authoring model, thesimone20aโ€”โ€”

Interpretation history

Decision trace