Independent evaluations will determine whether Bough’s program-per-turn architecture improves coding-agent tool efficiency or reliability over conventional iterative tool-call loops.
state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agents agent-harnesses program-synthesisBough
What is this?
Bough is presented in a Show HN launch and an initial Gleam-monorepo commit as a sandboxed coding agent that writes a program on each turn rather than issuing conventional iterative tool calls. The supplied material does not identify its creator clearly or provide independent evaluations of Bough itself. The search snippets discuss general agent-evaluation metrics and unrelated reductions in tool calls, so they do not establish that Bough improves efficiency, reliability, speed, or plan adherence.
Why it matters to Scott
Bough independently instantiates Scott’s Code-First Architecture claim that agents can replace repeated model-facing tool calls with code executed inside a sandbox, creating a concrete dated-receipts and benchmarking opportunity. The relevance is capped at medium because the supplied evidence provides no independent results showing that Bough actually improves efficiency or reliability.
ip:framework.code-first-architectureip:concept.code-as-step-between-model-runsip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:concept.agent-harnessesradar:concept.coding-agentsradar:concept.coding-agent-benchmarks
queries asked of Scott's wikis
- program-per-turn coding agent architecture
- program synthesis versus iterative tool calling
- coding-agent trajectory efficiency benchmarks
- sandboxed agent harness design
- agent plan execution and reliability
- tool-call overhead and parallelism
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-16T16:29:56Z
Repeated rechecks produced no independent benchmark, replication, adoption, or comparative result, and there is no near-term confirming event to watch; the launch remains a useful architectural example but no longer warrants an active episode.
2026-08-14T15:49:06Z
The small engagement increase is repetitive amplification, not evidence of comparative performance, adoption, or independent replication. Bough remains a concrete benchmarking target whose architectural claims are unvalidated.
2026-08-12T15:41:32Z
Reobservation adds no independent evaluation, implementation uptake, or comparative results; Bough remains a testable architecture example rather than evidence that program-per-turn execution improves agent performance.
2026-08-12T15:37:27Z
grounded: converges/medium — Bough independently instantiates Scott’s Code-First Architecture claim that agents can replace repeated model-facing tool calls with code executed inside a sand
2026-08-12T15:34:58Z
origin walked (codex/luna, conf 0.96): anchor hn.story.49273308 -> echo.github.15f7165ed5 by Andrey Lukin
2026-08-12T15:33:45Z
case created — Bough is a usable implementation of a distinct agent-harness architecture with a directly testable systems claim.
Decision trace
- 08-17 02:29expireRepeated rechecks produced no independent benchmark, replication, adoption, or comparative result, and there is no near-term confirming event to watch; the launch remains a useful architectural exampl
- 08-17 02:29alert_silentThe only delta is scheduled staleness with no new evidence, so there is nothing consequential to surface or hold for.
- 08-17 02:29alert_routeThe only delta is scheduled staleness with no new evidence, so there is nothing consequential to surface or hold for.
- 08-15 01:49repriceThe small engagement increase is repetitive amplification, not evidence of comparative performance, adoption, or independent replication. Bough remains a concrete benchmarking target whose architectur
- 08-15 01:49alert_silentNo consequential new fact has emerged; wait for an independent benchmark, replication, or meaningful implementation uptake.
- 08-15 01:49alert_routeNo consequential new fact has emerged; wait for an independent benchmark, replication, or meaningful implementation uptake.
- 08-13 01:41repriceReobservation adds no independent evaluation, implementation uptake, or comparative results; Bough remains a testable architecture example rather than evidence that program-per-turn execution improves
- 08-13 01:41alert_silentNothing consequential changed beyond an unchanged launch observation, so the case can wait for independent benchmarks, replication, or meaningful adoption.
- 08-13 01:41alert_routeNothing consequential changed beyond an unchanged launch observation, so the case can wait for independent benchmarks, replication, or meaningful adoption.
- 08-13 01:38alert_silentBough is a concrete open-source implementation of program-per-turn agent execution and is relevant as a benchmarking target, but the only new evidence is its low-visibility project launch; no evaluati
- 08-13 01:38surface_candidateBough is a concrete open-source implementation of program-per-turn agent execution and is relevant as a benchmarking target, but the only new evidence is its low-visibility project launch; no evaluati
- 08-13 01:38alert_routeBough is a concrete open-source implementation of program-per-turn agent execution and is relevant as a benchmarking target, but the only new evidence is its low-visibility project launch; no evaluati
- 08-13 01:37groundBough independently instantiates Scott’s Code-First Architecture claim that agents can replace repeated model-facing tool calls with code executed inside a sandbox, creating a concrete dated-receipts
- 08-13 01:34promote_anchororigin walk conf 0.96
- 08-13 01:33createBough is a usable implementation of a distinct agent-harness architecture with a directly testable systems claim.