2026-10-11 17:09 UTC

VirtusLab claims its released Orca orchestrator enforces coding and review stages through Scala workflows and commits progress alongside code, enabling resumable multi-agent development without relying on prompts to enforce workflow order.

state: resolvedheat: lowuncertainty: mediumconvergesscott: mediumcoding-agents agent-harnesses durable-orchestration software-workflowsVirtusLab

What is this?

VirtusLab's Orca is an open-source CLI for 'deterministic, AI-driven development flows': the developer defines the workflow (plan → implement → review) as a Scala 3 script run via scala-cli, while coding agents (Claude Code, Codex, OpenCode, Pi, Gemini) do the actual coding. Workflow requirements are expressed in code rather than agent prompts — the docs' 'Review is code' and 'stages commit' mechanics make each stage's code changes plus a progress-log entry a single atomic git commit, so interrupted runs resume from the last commit with completed stages skipped. It is developed and released by VirtusLab (a Scala-ecosystem consultancy; Adam Warski has publicized it on LinkedIn), and its mechanics are corroborated by an independent HN writeup running Codex and Pi through it side by side. Notably, search for it is heavily polluted by Stably AI's unrelated 'Orca' agent development environment, a worktree-parallel multi-agent IDE.

Why it matters to Scott

Orca independently ships what Scott's deterministic-agent-control-plane and 'can't beats shouldn't' doctrine prescribe — workflow order and review enforced in Scala code rather than prompts, with each stage committing code plus a progress-log entry as one atomic checkpoint — making it a released reference implementation, not merely an example, of positions his canon already holds. Its commit-as-unit-of-atomicity recovery is a direct comparison point for the Proposal Compiler's externalised job-resumption plane (git history vs persisted PID/session state), and the third-party Codex-vs-Pi run corroborates the harness-agnostic backend swap he cares about. Negligible attention keeps this a cold artifact worth hands-on comparison rather than a story to track.
dev:concept.deterministic-agent-control-planedev:concept.resumable-agent-job-control-planedev:project.proposalip:framework.architecture-not-vibesip:concept.checkpoint-disciplineip:framework.long-running-agentsradar:agentlane-git-native-coordinationradar:aws-aidlc-multi-harness-workflowsradar:claramap-cross-harness-orchestrationradar:trigora-continuation-recoveryradar:gantree-chat-independent-long-horizon-harness
queries asked of Scott's wikis
  • deterministic control plane for coding agents
  • Proposal Compiler job recovery checkpoint
  • commit-as-checkpoint resume interrupted agent run
  • agent-agnostic harness interchangeable backends claude codex
  • code-enforced review gate vs prompt-enforced workflow
  • workflow-as-code orchestration framework comparison

Measured heat

now 1 pts/hpeak 3 pts/hcomments 0/hpeers p71momentum: steady2 platformsage 439h
points/hour across evidence · reading as of 2026-10-07 09:04:59.495596+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 16:34 (minted)⭐ origin echo-reconstructedOrca defines development flows programmatically, delegates coding to existing agent backends, and commits stage results with repository chan
VirtusLab on github (echo) · attributed from hn.story.49755856 · published time unknown
—
09-18 15:32first on hacker news · published · lag ?Orca – deterministic, AI-driven development flows
mihau
—
09-18 15:32amplified on hacker news 👑hn.story.49755856
mihau
peak 4 · 0 comments · 45% of case engagement
10-06 15:07amplified on hacker newshn.story.49979636
piscaries
peak 2 · 1 comments · 33% of case engagement
10-06 17:23amplified on hacker newshn.story.49981502
sijieg
peak 2 · 0 comments · 22% of case engagement
09-18 16:20our radar first saw it · lag ?discovery anchor: hn.story.49755856—

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOrca – deterministic, AI-driven development flows
Retrieved article excerpt

Open article · Retrieved 2026-09-18T16:23:00.332109+00:00

# Orca

Deterministic, AI-driven development flows.

Orca allows you to programmatically define software development workflows where
AI agents perform the coding. If you want AI-generated code to always be
reviewed by another agent, don't try to coerce the agents; just express that
requirement in code. Don't waste tokens on formatting, committing, or creating
PRs - all of this can be handled by an ordinary script.

Orca comes with an `orca` cli, which can be used interactively by humans, or
headlessly by humans and agents alike. A number of built-in flows, implementing
e.g. a plan-implement-review loop, allow you to start using Orca right away.

Orca flow scripts are written in Scala, and can be run with a single command
through [scala-cli](https://scala-cli.virtuslab.org), which is installed by the
`orca` installer. No other dependencies are needed - everything is automatically
bootstrapped. Scala 3 looks like Python, but with types - so you get quick
feedback if your flow script has any problems.

Orca's development flows are resumable, so that if work is interrupted mid-flow
for any reason, it can be continued from the last commit.

You can use Orca to orchestrate development in any language and ecosystem.

Orca assumes that it has configured, logged-in access to Claude, Codex,
OpenCode, or Pi (depending which backend you use), as well as `gh` and `git`.

Install with one command, which installs `scala-cli` (via its official
installer) if you don't have it already, and writes the `orca` executable to
`~/.local/bin/orca`:

```
curl -fsSL https://raw.githubusercontent.com/VirtusLab/orca/master/install.sh | bash
```

See [Orca Shell](https://github.com/VirtusLab/orca#orca-shell) for the details and the full command-line
reference, or just run `orca` / `orca help`.

## Three ways to work with Orca

**Interactively**: install the CLI, run `orca`, pick a flow (`implement.sc`
comes first in the list) and enter your task. Non-interactively, use `orca run <flow> "<task>"`. See [Orca Shell](https://github.com/VirtusLab/orca#orca-shell) for installation and the full
command-line reference.

> [!WARNING] **Orca is designed to work in a sandboxed environment!** Coding
> agent tool usage is auto-approved by default (`tools = ToolSet.Full`,
> `autoApprove = AutoApprove.All`): write-capable turns let the agent edit files
> and run shell commands without prompting. This can be changed by changing the
> flow's options in code. Alternatively, use a VPS or local sandbox such as
> [Sandcat](https://github.com/VirtusLab/sandcat), [Docker
> Sandboxes](https://docs.docker.com/ai/sandboxes/), or any other.

**Driven by an agent (headless)**: a coding agent or harness invokes the CLI
non-interactively to implement a task, e.g. from CI or as a sub-task of another
agent:

```
orca run implement.sc "add a rate limiter to /login"
```

Useful flags: `--skip-branch` (continue on the current branch instead of
creating one), `--keep-changes` (leave uncommitted files in place instead of
stashing them) and `--worktree` (run in a git worktree of this repository
instead of the current checkout).

In every mode, which agent (and model) handles the planning, coding, and review
roles comes from `settings.properties` — written for you by the shell's
first-run wizard or `orca config`, hand-editable too; see [Settings](https://github.com/VirtusLab/orca#settings).

Agents can load [`skills/using-orca`](https://github.com/VirtusLab/orca/blob/master/skills/using-orca/SKILL.md) to know when
and how to delegate here — installable as a Claude Code plugin, a Pi package, or
by symlinking into any harness's skills directory; see [its
README](https://github.com/VirtusLab/orca/blob/master/skills/using-orca/README.md) for specifics.

**As a script**: run a flow directly with `scala-cli`, no install required — see
[An example flow](https://github.com/VirtusLab/orca#an-example-flow).

```
scala-cli run implement.sc -- "add a rate limiter to /login"
```

## An example flow

Save this as `implement.sc` and run it with your task:

```
//> using scala 3.9.0
//> using dep "org.virtuslab::orca:0.1.7"
//> using jvm 21

import orca.{*, given}

// Roles (planning / coding / review) come from settings.properties —
// per-project `.orca/settings.properties`, else ~/.config/orca/settings.properties,
// else claude for everything. Bodies can still name a concrete harness
// (`claude`, `codex.mini`, …) where a flow wants one — details under "Coding
// agent tools".
flow(OrcaArgs(args)):
  // `stage` is the committing, resumable unit of work. The plan is produced in
  // one agentic turn and recorded in the stage log; a re-run with the same
  // prompt skips this stage and reads the stored Plan back.
  val plan = stage("Plan"):
    Plan.autonomous.from(userPrompt, planningAgent).value  

  // One stage per task: each stage commits its work + a progress-log entry as
  // one commit. Completed stages are skipped on resume — re-running the same
  // prompt picks up from the first incomplete task. Each task gets its own
  // session, keyed by that task and seeded with the plan's brief (which primes
  // it on first use, and is replayed if the backend session is lost on resume).
  for (task, n) <- plan.tasks.zipWithIndex do
    stage(s"Task: ${task.title}"):
      val session = codingAgent.session(
        "implementer",
        detail = s"task ${n + 1}: ${task.title}",
        seed = plan.brief
      )
      session.run(task.description)
      reviewThenFix(
        coderSession = session,
        reviewers = allReviewers(reviewAgent),
        // One review round, one fix turn. Reviewers are picked by a picker LLM
        // on reviewAgent.cheap (see "Review utilities"); format and lint
        // default to the project's stack settings
        // (`.orca/settings.properties`, auto-discovered on first run) — see
        // "Settings" below. The whole task goes in: reviewers are shown its
        // title and description, plus the run's prompt, each labelled.
        task = task
      )

  // Each task's single pass took the fixer's word for its own fixes; this loop
  // over everything the run changed is what checks them.
  val openFindings = stage("Final review"):
    reviewAndFixLoop(
      coderSession = session,
      reviewers = allReviewers(reviewAgent),
      task = Task(Title("The whole planned change"), plan.brief),
      diff = ReviewDiff.WholeRun,
      maxIterations = 5
    )

  // Best effort: opens a PR when the checkout is on a GitHub `gh` can reach,
  // and says why in one line when it isn't. What the loop left open is listed
  // in the PR body.
  openPrIfGitHub(
    summarisingAgent = codingAgent.cheap,
    openFindings = openFindings
  )
```

```
scala-cli run implement.sc -- "Add a rate-limiter to the /login endpoint"
```

Each flow starts by creating a feature branch, named by a short
cheap-model-generated label derived from the prompt (slugged; pass `branchNaming = ...` to override). On success the flow opens a PR when the repository is on a
GitHub `gh` can reach, and hands you back the branch you started on — the work
is on the PR. Otherwise it says so in one line and leaves you on the feature
branch, ready to test or open a PR by hand — see [The flow
lifecycle](https://github.com/VirtusLab/orca#the-flow-lifecycle) for the full success/failure/resume behavior.

If the flow is interrupted — user intervention or an intermittent error — just
run the same command again: it resumes from the last committed set of changes,
so only a small amount of work is repeated. Orca borrows ideas from durable
computing: which stages have completed, and with what results, is tracked in a
progress file committed alongside the modified code, making commits the unit of
atomicity — the progress log can't drift from the changes in the repository.
When the flow is done, the progress log is removed from the branch in one last
commit, which is pushed too if the flow had already pushed the branch.

There are two runnable examples under
[`examples/runnable/`](https://github.com/VirtusLab/orca/blob/master/examples/runnable):

- [01-simple](https://github.com/VirtusLab/orca/blob/master/examples/runnable/01-simple) (in-memory plan + review, autonomous
  planner),
- [02-interactive](https://github.com/VirtusLab/orca/blob/master/examples/runnable/02-interactive) (same shape as 01, but the
  planner can ask clarifying questions via `ask_user`).

More flow scripts — `issue-pr.sc`, `issue-pr-bugfix.sc`,
`implement-enhanced.sc`, `review.sc` — live in [`flows/`](https://github.com/VirtusLab/orca/blob/master/flows); run them
against your own git repo.

For convenient editing of Orca flow scripts, with code-completion, you can try
the [Metals](https://scalameta.org/metals/) VSCode extension.

## Built-in tools

The following are available inside a `flow(...) { ... }`.

The five coding agents — `claude`, `codex`, `opencode`, `pi`, `gemini` — share
one call surface. Durable: `session(name, detail, seed): FlowSession` →
`.run(prompt)` / `.resultAs[O].run(input)`. One-shot: `run(prompt)`,
`resultAs[O].{autonomous,interactive}.run(input)`. Ephemeral multi-turn:
`chat(): Chat` → `.run(prompt)` / `.resultAs[O]...run(input)`. Common tuning:
`withModel`, `withCheapModel`, `withConfig`, `withSystemPrompt`, `withName`,
`withReadOnly`, `withNetworkOnly`, `withSelfManagedGit`. The table lists each
backend's model accessors and backend-specific extras:

| Tool | Backend-specific methods | Purpose |
| --- | --- | --- |
| `claude` | `haiku`/`sonnet`/`opus`/`fable`, `cheap` (→ haiku), `withModel(Model)`, `withNetworkTools` | Claude Code coding/reviewing agent. Bare `claude` is **Opus with the 1M-token context window** (the coder; reviewers share it); use `claude.sonnet`/`claude.haiku` for cheap one-shot calls, or `claude.fable` for the hardest ones. `interactive` mode lives only on `resultAs[O]`. See [Sessions](https://github.com/VirtusLab/orca#sessions) for durable (`session`) vs ephemeral (`run`/`chat`). |
| `codex` | `mini`, `cheap` (→ mini), `withModel(Model)` | OpenAI Codex coding/reviewing agent. Bare `codex` pins **GPT-5.6 Sol**; use `codex.mini` for cheap one-shot calls. |
| `opencode` | `anthropicOpus`/`anthropicSonnet`/`anthropicHaiku`, `openaiSol`/`openaiTerra`/`openaiLuna`, `cheap` (provider-matched: openai→luna, else anthropicHaiku), `withModel(providerModel)` / `withModel(provider, modelId)` | [OpenCode](https://opencode.ai) coding/reviewing agent, driven over HTTP+SSE against a headless `opencode serve` (started lazily, shared for the run; sessions survive it — see [Sessions](https://github.com/VirtusLab/orca#sessions)). Spans providers, so models are provider-qualified: use an accessor (`opencode.openaiLuna`) or `opencode.withModel("openai/gpt-5-mini")` / `opencode.withModel("ollama", "llama3.1")`. Inherits the user's configured `opencode` providers/auth. |
| `pi` | `withModel(Model)` | [Pi](https://pi.dev/) coding agent backend, driven through `pi --mode rpc`. Pi handles provider/model selection through its own CLI configuration; pin a model with `pi.withModel(Model("provider/model"))`. Interactive calls can ask clarifying questions via Orca's `ask_user` bridge. |
| `gemini` | `flash`, `cheap` (→ flash), `withModel(Model)` | Google Gemini CLI coding/reviewing agent, driven via `gemini --output-format stream-json`. Bare `gemini` pins **Gemini 2.5 Pro**; use `gemini.flash` for cheaper one-shot calls. Structured output is prompt-enforced (Gemini has no schema flag); `withReadOnly` maps to `--approval-mode plan`. See [ADR 0015](https://github.com/VirtusLab/orca/blob/master/adr/0015-gemini-stream-json-driver.md). |
| `git` | `createBranch`, `checkout`, `ensureClean`, `commit`, `forceAdd`, `push`, `currentBranch`, `headCommit`, `uncommittedDiff`, `changedFiles`, `reviewChanges`, `pendingChanges`, `diffVsBase`, `defaultBase`, `discardUncommitted`, `deleteBranch`, `branchHasChangesExcludingOrca` | Git operations against the working tree. Recoverable failures (`BranchA
mihau40
🟧 echo.github ⭐Orca defines development flows programmatically, delegates coding to existing agent backends, and commits stage results with repository chanVirtusLab——
🟧 hnI had Codex and Pi build the same app in Orca, then compared them side by sidepiscaries21
🟧 hnOrca – Open-source runtime that turns any agent harness into a managed agentsijieg20

Interpretation history

Decision trace