2026-10-11 17:10 UTC

Canary (YC) claims its released service independently verifies AI-generated code by booting the app in remote sandboxes and break-testing each change with agent swarms β€” adoption by coding-agent workflows would establish independent verification as a standard post-generation gate.

state: corroboratedheat: highuncertainty: mediumconvergesscott: highagent-code-verification agent-evaluation coding-agent-workflowsCanaryY Combinator
Surfaced 2026-09-27T03:54:40Z β€” The linked site is the primary artifact β€” a product landing page whose own words: "Canary is the independent tester for code your agents wri β€” Nom Army's release ('don't trust coding agents when they say they're done') supplies a third independent implementation of the post-generation verification gate within a single week β€” the periphery Keel opened is now a pattern, and implementation diversity (not engagement) tips the case to corroborated. Demand and independent efficacy remain silent across all three tools, so this is a quietly forming category, not a moving story; heat stays low.

What is this?

Canary is a YC W26 startup founded by ex-Windsurf and Google engineers that bills itself as 'the validation layer for AI-generated code' β€” an 'AI QA engineer' that connects to your codebase, reads each diff, maps the blast radius across routes and API schemas, then runs the app and tries to break the change before human review or merge, reporting breakage rated P0/P1. It ships as a CLI (@runcanary/cli) with integrations for Claude Code, Cursor, and Codex, so the verification loop runs from inside existing coding agents and on every PR. Its core argument is that agents write both the code and the tests, so both inherit the agent's blind spots, and independent adversarial checking must sit outside the generating agent. The supplied snippets confirm the product, team, and positioning but do not establish the specific remote-sandbox/agent-swarm mechanics; the AI-QA space is visibly crowding (fellow YC company TesterArmy raised €1M for black-box user-journey test agents), and the name collides with at least one unrelated agent-evaluation project ('Code Canary' by Fred Benenson).

Why it matters to Scott

Converges on the core of mechanically-different-verifiers and correlated-checkers-pitfall β€” Canary's pitch that agent-written code and agent-written tests inherit the same blind spots, so the gate must be an external mechanism that can fail differently, is Scott's own argument now operationalized as a YC-backed product integrating with the Claude Code/Cursor/Codex loops he actively drives (cf. his own claim-bounded adversarial verification gates). It is also a dated market receipt for the verification-cost thesis β€” verification itself becoming the purchased product, with the space visibly crowding (Argus, Kery, TesterArmy) β€” while raising a live tension for Don't Buy Software, Build AI Instead: does hosted third-party verification strengthen the trust-technology pattern the ebook describes, or displace the buyer-owned harness it prescribes?
ip:concept.mechanically-different-verifiersip:concept.correlated-checkers-pitfallip:concept.verification-costip:source.dont-buy-software-build-ai-instead-ebookip:concept.adversarial-closerdev:concept.claim-bounded-adversarial-verificationradar:argus-agentic-qa-validationradar:argus-agent-web-app-testingradar:kery-browser-pr-validationradar:modelpeer-cross-model-code-reviewradar:concept.agent-verificationradar:concept.agent-evaluation
queries asked of Scott's wikis
  • independent verification gate for agent-written code
  • agent-authored tests inherit generator blind spots
  • coding agent harness verification loop design
  • evals for coding agents post-generation checks
  • verification-as-a-service AI QA product pattern
  • human review bottleneck in agent coding workflows

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 434h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-23 14:00⭐ origin echo-reconstructedThe linked site is the primary artifact β€” a product landing page whose own words: "Canary is the independent tester for code your agents wri
Canaries, Inc. (founded 2026 by Aakash Mahalingam and Viswesh N G, San Francisco; YC-backed) on other (echo) Β· attributed from hn.story.49833435
β€”
09-24 16:51first on hacker news Β· published Β· +26.9hShow HN: Canary (YC) – Independent verification for AI code
Visweshyc
β€”
09-24 16:51amplified on hacker newshn.story.49833435
Visweshyc
peak 1 Β· 1 comments Β· 11% of case engagement
09-24 20:57amplified on hacker news πŸ‘‘hn.story.49836632
Visweshyc
peak 8 Β· 0 comments Β· 44% of case engagement
09-26 13:31amplified on hacker newshn.story.49856381
danebalia
peak 2 Β· 2 comments Β· 22% of case engagement
09-27 03:17amplified on hacker newshn.story.49863042
RaysonTech
peak 3 Β· 1 comments Β· 22% of case engagement
09-24 20:21our radar first saw it Β· +30.4hdiscovery anchor: hn.story.49833435β€”
09-27 03:42reached heat=high Β· +85.7h Β· via queue+ledgerβ€”β€”
pace: p50 vs 1032 stories at the 336h mark (now 434h old) β€” ahead of agentgit-accountless-agent-handoffs (1.1x), behind anthropic-pentagon-blacklist-ruling (0.9x)

Evidence (5) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Canary (YC) – Independent verification for AI code
Retrieved article excerpt

Open article Β· Retrieved 2026-09-24T20:37:54.616789+00:00

Backed by Y Combinator

# ClaudeCodexCursorCopilotClaude writes your code.code.code.code.code.Canary tests it.

Canary is the independent tester for code your agents write. It runs your app, tries to break every change, and reports what would have reached production.

Copy the prompt One line for your agent.

Claude Codeβ€” pocketbaseClaude CodeCodex

Ready for review 1

Database-aware health endpoint

+46 -1 Β· apis/health.go

Database-aware health endpoint

Add a database-aware health endpoint: GET /api/health should ping the data and auxiliary databases and answer 503 when either cannot run SELECT 1.

apis/health.goUpdated

Done. GET /api/health runs SELECT 1 against the data and the auxiliary database before answering. If either probe fails, the endpoint logs the error and returns HTTP 503 with code 503 and message "API is unhealthy." Each probe is bounded by a 3 second timeout derived from the request context. The successful response, including the superuser-only fields, is unchanged.

∞ Agent Claude Code

apis/health.go

@@ -1,9 +1,12 @@

package apis

import (

+ "context"

"net/http"

"slices"

+ "time"

+ "github.com/pocketbase/dbx"

"github.com/pocketbase/pocketbase/core"

"github.com/pocketbase/pocketbase/tools/router"

)

@@ -14,7 +17,12 @@ func bindHealthApi(app core.App, rg \*router.RouterGroup[\*core.RequestEvent]) {

subGroup.GET("", healthCheck)

}

-// healthCheck returns a 200 OK response if the server is healthy.

+// healthCheckDBTimeout is the max time the health check waits for each database ping.

+const healthCheckDBTimeout = 3 \* time.Second

+

+// healthCheck returns a 200 OK response if the server is healthy

+// (aka. the HTTP server is up and the data and aux databases answer a trivial query),

+// otherwise a 503 Service Unavailable response.

func healthCheck(e \*core.RequestEvent) error {

resp := struct {

Message string `json:"message"`

@@ -25,6 +33,16 @@ func healthCheck(e \*core.RequestEvent) error {

Message: "API is healthy.",

}

+ if err := pingHealthDBs(e); err != nil {

+ e.App.Logger().Error("Health check database ping failed", "error", err)

+

+ resp.Code = http.StatusServiceUnavailable

+ resp.Message = "API is unhealthy."

+ resp.Data = map[string]any{}

+

+ return e.JSON(http.StatusServiceUnavailable, resp)

+ }

+

// @todo evaluate whether it is worth removing the extra info from the health endpoint

if e.HasSuperuserAuth() {

resp.Data = make(map[string]any, 3)

@@ -52,3 +70,30 @@ func healthCheck(e \*core.RequestEvent) error {

return e.JSON(http.StatusOK, resp)

}

+

+// pingHealthDBs runs a trivial query against the data and aux databases

+// and returns the first error it encounters.

+//

+// Each database gets its own healthCheckDBTimeout budget (derived from

+// the request context so that a client cancellation still stops the probe).

+func pingHealthDBs(e \*core.RequestEvent) error {

+ if err := pingHealthDB(e.Request.Context(), e.App.DB()); err != nil {

+ return err

+ }

+

+ if err := pingHealthDB(e.Request.Context(), e.App.AuxDB()); err != nil {

+ return err

+ }

+

+ return nil

+}

+

+// pingHealthDB runs SELECT 1 against the provided db within healthCheckDBTimeout.

+func pingHealthDB(parent context.Context, db dbx.Builder) error {

+ ctx, cancel := context.WithTimeout(parent, healthCheckDBTimeout)

+ defer cancel()

+

+ var result int

+

+ return db.NewQuery("SELECT 1").WithContext(ctx).Row(&result)

+}

0K+

bugs found before merge

0K+

engineer hours saved

0%+

fix rate

0%

rated P0 / P1

01 / Install

Get started in under 2 minutes.

Works with Claude Code, Cursor, Codex and any of your agents

01

Paste the prompt

02

Connect your stack

03

Catch failures

Paste this into your coding agent

`` Install the Canary CLI with `npm i -g @runcanary/cli`, then run `canary skills` and follow its instructions to onboard this repository. ``

Copy

02 / The problem

## The blast radius is the unknown unknown.The blast radius is the unknown unknown.

1. 01

   Your agents raise more PRs than your team can review. Review turned into a skim.
2. 02

   The tests pass. The same agent wrote them, and nobody on your team reads them. Nor should they.
3. 03

   The bug shows up as a Slack message from a customer, or a page that wakes your on-call at 2am.
4. Teams on Canary have seen bug reports on new code go down, and on-call pages with them.

03 / How it works

## Release the flock.Release the flock.

One loop. From your coding agent, and on every pull request.

Reviewers read the diff. Canary runs the app, tries to break the diff, and reports the blind spots you missed.

01Read

02Run

03Break

04Verify

01

Canary reads the diff and your codebase and maps every flow the change can reach.

β–Έ diff: apis/health.go Β· +46 βˆ’1

β–Έ reach: GET /api/health Β· backups page Β· superuser page Β· proxy header

β–Έ 5 tests queued

$

04 / Where it runs

## Runs where you write code.Runs where you write code.

Two triggers. One loop. Every run ends in a report.

01In your agent

02On every PR

03Fix loop

04Report

01

Before your agent calls a task done. Canary snapshots the uncommitted tree, boots it in a clean sandbox and tests it. No commit, no PR.

$ canary verify

β–Έ snapshot: working tree Β· 4 files Β· nothing committed

β–Έ sandbox: clean boot Β· pocketbase

β–Έ 5 tests queued … running

$

05 / Integrations

## Works where you work.Works where you work.

Canary reads the context you already have, and files every break back to your coding agent.

GitHub

Linear

Sentry

Datadog

Notion

Slack

CANARYtests

Claude Code

PR Β· CLI Β· MCP

06 / Examples

## Every failure arrives with proof.Every failure arrives with proof.

Four failures Canary caught on real customer pull requests.

01Wrong org's file02One click, two records03Fetches any URL04Delete wipes all

test replay Β· api/documentsREPLAY

test Β· api/documents

β–Έ auth as Org A Β· request a review image by storageId

β–Έ swap storageId β†’ a document owned by Org B

βœ— server resolves it Β· Org B's private file returned to Org A

Β· one customer can open another customer's file

$

Case file Β· api/documentsP0

staging Β· found autonomously in 6m 09s

β–Έ

Test

swap storageId to another org's document

βœ—

Break

another org's private file opens for the caller

βŒ–

Root cause

image-URL resolver looks up by storageId with no org-ownership check

+

Suggested fix

βˆ’ resolveImage(storageId)

+ resolveImage(storageId, { orgId: caller.orgId })

βœ“

Regression armed

cross-org.storageId reruns on every PR

07 / Research

## Field notes from the frontier.Field notes from the frontier.

We publish what we learn measuring AI against real, messy codebases.

[Benchmark15 min read Β· March 2026

### QA-Bench v0: measuring how AI models handle code verification

Given a real pull request on a production codebase, can a model find every affected user flow and catch what breaks? We put a purpose-built agent against the frontier models across 35 PRs on four production-scale repos.

Read the benchmark β†’

Overall accuracy Β· QA-Bench v0

Canary83.1

GPT-5.480.2

Claude Code78

Sonnet 4.673.2](https://www.runcanary.ai/blog/qa-bench-v0)

08 / Your move

## Find it before your users do.Find it before your users do.

Put the flock on your next pull request. It runs on every one after.

Copy the prompt One line for your agent.
Visweshyc11
🟧 echo.other ⭐The linked site is the primary artifact β€” a product landing page whose own words: "Canary is the independent tester for code your agents wriCanaries, Inc. (founded 2026 by Aakash Mahalingam and Viswesh N G, San Francisco; YC-backed)β€”β€”
🟧 hnShow HN: Canary (YC) – Independent verification for AI codeVisweshyc80
🟧 hnShow HN: AI coding agents that prove their workdanebalia22
🟧 hnNom Army: Don't trust coding agents when they say they're doneRaysonTech31

Interpretation history

Decision trace