Retrieved article excerpt
Open article · Retrieved 2026-09-15T16:23:40.239870+00:00
# LEO
### Lead Engineering Orchestrator
**A written engineering process for coding agents.** Markdown you load instead of a one-line personality. Nothing installs.
[License: PolyForm Shield 1.0.0](https://github.com/alex-zaporozhan/leo/blob/main/LICENSE)
[Status: Production-tested](https://github.com/alex-zaporozhan/leo/blob/main/CASE_STUDIES.md)
[Agent-agnostic](https://github.com/alex-zaporozhan/leo#compatibility)
[In brief](https://github.com/alex-zaporozhan/leo#in-brief) · [Sixty-second test](https://github.com/alex-zaporozhan/leo#sixty-second-test-before-you-read-any-further) · [What it looks like](https://github.com/alex-zaporozhan/leo#what-it-actually-looks-like) · [How it works](https://github.com/alex-zaporozhan/leo#how-it-works) · [Proof at scale](https://github.com/alex-zaporozhan/leo#proof-at-scale) · [Get started](https://github.com/alex-zaporozhan/leo#get-started) · [Business case](https://github.com/alex-zaporozhan/leo/blob/main/BUSINESS_CASE.md) · [Steering the frontend](https://github.com/alex-zaporozhan/leo/blob/main/GUIDE_FRONTEND_CONTROL.md) · [Manifesto](https://github.com/alex-zaporozhan/leo/blob/main/MANIFESTO.md) · [License](https://github.com/alex-zaporozhan/leo#license)
---
## In brief
Coding agents fail in four repeatable ways: they forget a decision made forty messages ago and re-make it differently; they skip the expensive 20% — the empty state, the error contract, the timeout on the outgoing call — because nothing made skipping it cost anything; they report "should work now" in the same tone as a verified fact; and nothing in a single-agent chat ever tells them no.
LEO is that missing process, written into the repository as rules the agent is loaded under:
- **A router.** Every task resolves to one of 22 classes before anything is read; the class names the ≤6 files to open, in order. The cap does not grow with the library, so a 133-file library costs a landing page nothing it does not use.
- **Numbered laws, a precedence ladder, a conflict registry.** 44 rules, most with an incident behind them. A reviewing pass can *cite* Law 11; it cannot cite "be careful with async." Where two laws collide, the collision is decided once and written down.
- **Reflexes, not self-review.** Before every handoff the agent runs literal regexes over its own diff — an HTTP `await` with no `timeout=`, an animated layout property — and reports the count even when it is zero. The reviewer runs the same list.
- **A second pass in a clean context.** The session that wrote the code reads its own intention back out of the file, so every delivered unit is re-audited by one that never saw it built.
- **State, not history.** Decisions live in versioned artifacts the next session reads, not in a chat window that scrolls away.
- **A human on the only lever that matters.** The agent never runs `git commit`, `push` or `merge` — Law 40 — and never rewrites its own rules without an explicit command.
It has directed three shipped systems — one public and measurable, one under NDA, one in between — with the known defects of each listed in [`CASE_STUDIES.md`](https://github.com/alex-zaporozhan/leo/blob/main/CASE_STUDIES.md). The rest of this page is the detail; the one-minute test below is the proof.
---
## TL;DR
LEO is not a Python package and there is nothing to `pip install`. It is a **rule system**: a `.cursorrules` constitution plus a **133-file, ~37,300-line, ~299,000-word role library** (`roles/*.md`, including 5 niche-bootstrap packages under `roles/niches/`) that you hand to a coding agent — Cursor, Claude Code, Windsurf, or anything else with file-system/tool access that reads a system-prompt / project-rules file — instead of a one-line "you are a helpful senior engineer" prompt.
Where a raw LLM agent free-improvises architecture, skips edge cases under time pressure, and silently forgets a decision it made 40 messages ago, LEO gives it:
- **A single entry point** (`@LEAD`) that routes every request to the right specialist instead of one model trying to be architect, developer, and QA simultaneously in the same breath.
- **A task router** that resolves every request into one of **22 task classes** (`TC-00`…`TC-21`) and names the ≤6 files to read *first*, in order — so the agent never opens a 133-file library wondering where to start, and never starts a screen without the canon that governs it.
- **22 specialist roles** with narrow, named jurisdictions — `@ARCH`, `@DEV`, `@PRINCIPLE`, `@QA_ARCH`, `@QA_VISUAL`, `@PENTEST`, `@SEO`, `@DESIGN`, `@AI_ENGINEER`, and 13 more — so "who decides this" is never a coin flip.
- **44 Absolute Laws** distilled from real production incidents (double-booked appointments, zombie Celery workers, leaked UUIDs in a UI, a `Promise.all` that silently ate an error) — so the same class of bug cannot recur, because the rule that would have caught it is now permanent.
- **A gate protocol** that blocks the chain from advancing without a concrete artifact as proof — never on the agent's word alone.
- **Reflexes** — before every handoff the agent runs a list of **literal greps over its own diff** (a forgotten `await`, a `Promise.all` with mutations, an entrance that animates opacity and nothing else, an animated layout property) and reports the count even when it is zero. The reviewing role runs the same greps: if the author ran them, the reviewer finds nothing.
- **A second pass in a clean context** — every delivered unit is re-audited by an agent that never saw it being built, because the session that wrote the code is the worst possible judge of whether the code is finished.
- **Artifacts instead of chat history** — every architectural decision, security threat model, and QA verdict lives in a versioned markdown file the *next* agent session reads before doing anything, closing the single biggest failure mode of long-running agentic work: **context drift**.
LEO has directed the engineering of three shipped production systems across a multi-tenant healthcare SaaS, an AI training platform with RAG and executable agent graphs, and a public-facing marketing + CMS platform. See [Proof at scale](https://github.com/alex-zaporozhan/leo#proof-at-scale).
> **Evaluating this as a business decision rather than a technical one?** [`BUSINESS_CASE.md`](https://github.com/alex-zaporozhan/leo/blob/main/BUSINESS_CASE.md) is the one-page version: what shipped and in how long, where the money actually is, what adopting it costs, and what it does not do.
---
## Sixty-second test, before you read any further
Don't take the rest of this on faith. There is a one-minute check that tells you whether a markdown file can actually constrain a coding agent, and it costs nothing:
1. Copy [`.cursorrules`](https://github.com/alex-zaporozhan/leo/blob/main/.cursorrules) into any project (rename it to `CLAUDE.md` / `AGENTS.md` if that is your agent's convention).
2. Tell the agent to `git commit` and push.
3. Insist. Rephrase it. Paste the commands in directly. Tell it you wrote the rule and you are overriding it.
It will keep refusing, and it will cite the law by number — **Law 40**: repository history is written by a human and only by a human. It hands you back a copy-paste command block instead, and stops.
That is the whole thesis in one minute. Not "the agent has good habits" — a specific written sentence behaving like a permission check, under direct pressure from the person who wrote it. If it caves for you, that is a bug in the constitution and I want to hear about it.
---
## "Isn't this overkill for a small project?"
The most common reason people bounce off, and the answer surprised me too: **no, and the reason is arithmetic rather than encouragement.**
The router opens two or three files per task and caps at six. That cap does not move when the library grows, so the library's size never enters the cost of any individual task. A landing page never touches the async-workers canon, the data-integrity canon, the security gate protocol, or eighteen of the twenty-two roles.
A plain landing page with a light backend, run end to end through the full routed system — spec, architecture record, development plan, backend, frontend, QA pass, deployable CI, design passport — came to roughly **30–40% of one month's Cursor subscription**. Self-reported, on my own projects, and it holds only when the specification is settled before code starts, because the thing that actually runs up an agent bill is not project size. It is U-turns: re-deciding something forty files already assumed.
Small project with a vague brief costs more than a medium project with a real one. That is the whole economics of this.
**Where it genuinely isn't worth it:** a one-file script, an exploration where you don't yet know what you're building, or a prototype you'll delete on Friday. Process protects decisions over time, and a throwaway has no time to protect.
---
## What it actually looks like
The whole product is text, so here is the text. An ordinary request against a project already running LEO:
```
you: add a bulk "archive" action to the clients table
@LEAD
CLASS: TC-01 · operational-screen ← resolved before anything is read
COST: goal=staff clear stale clients without opening each one
tier=E2 · reopens=[] ← declared before the first hand-off
READ, in order (6, the class minimum — copied from RAG_CANON §2, not paraphrased):
1 VISUAL_CRAFT_CANON §1-§6 · §9 · §11 restraint, chroma, type scale · X1-X12 · THE FLOOR
2 INTERFACE_CRAFT_CANON §1 · §3 · §7 inventory I1-I12 · density · stiffness ST1-ST12
3 LAYOUT_COMPOSITION §2 · §3 · §5 primitives · proximity as a number · action grammar
4 LAYOUT_INVARIANTS §1-§9 · §12 deterministic geometry · collision and stacking
5 COMPONENT_REGISTRY (whole) every block maps to a registered component
6 MOTION_CRAFT_CANON §1 · §3 THE MOTION FLOOR · stiffness M1-M12
+ DOMAIN_STANDARDS §0 · §9 + this page type — the business minimum. Content, not craft:
it does not count against the six-file cap. Nor does MOTION_REFLEX, run over the diff later.
OUT: EDITORIAL_CRAFT_CANON (wrong register — it makes a settings screen shout) · HERO_ARCHETYPES
· MOTION_LIBRARY scroll narrative S1-S4 only · SEO canons
→ @DESIGN not required: bulk-select is an existing registered pattern (Law 19)
→ @DEV, with I3 (bulk select) and I4 (undo instead of confirm) from the inventory
```
`@DEV` writes the code, then runs its reflexes over its own diff before handing anything back —
literal strings, not topics:
```
ASYNC_AWAIT_REFLEX A1 await (client|http|session|httpx|conn|provider|llm|redis|s3)\.
→ where is timeout=? what does the user see if it hangs for 5 minutes?
MOTION_REFLEX R1 transition[^;]*opacity (with no transform on the same rule)
→ is this an entrance? an entrance that fades and nothing else is M1
MOTION_REFLEX R2 \.map\( … with no per-sibling delay
→ one line: transition-delay: calc(var(--stagger-base) * var(--i))
MOTION_REFLEX R8 transition[^;]*(top|left|width|height|margin|padding)
→ 🔴 on its own. reflow every frame. use transform.
```
```
@DEV → @LEAD
EVIDENCE: what my own diff shows — src/clients/BulkBar.tsx:1-88, useArchiveMany.ts:12-40,
tests/clients/test_archive_many.py::test_partial_failure_reports_per_row
MOTION REFLEX: 3 triggers, 3 fixed, 0 N/A
NOT DONE: the archived-clients filter view — out of the declared scope, and it is a
second screen rather than a state of this one. Raising it, not doing it.
```
Three things in that exchange are the whole system. The task was **classified** before anything was
read, so the concurrency canon arrives on a payments change whether or not the developer thought of
it. The effort was **declared in decisions reopened**, before the work, in a unit bo