2026-10-11 16:37 UTC

Mupt AI claims SelfBench — which converts a repository's merged PRs into Harbor-gated tasks with hidden tests and publishes accuracy-vs-cost leaderboards — becomes a standard private gate teams use to benchmark coding agents on their own codebases; external teams running it and releasing results confirm it, quiet fade closes it.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation agentic-coding coding-agent-harnessesMupt AIavyayv

What is this?

SelfBench is an open-source tool from the GitHub org mupt-ai that builds private coding-agent benchmarks out of a repository's own merged pull requests — each PR is converted into a task with hidden tests so teams can compare agents and models on their own codebase rather than on public suites. The supplied snippets confirm only the repo's existence and this high-level mechanism; they do not identify who or what Mupt AI is beyond the org, and they say nothing about the 'Harbor-gated' pipeline detail, leaderboard publishing, or any external team adoption. The surrounding climate is real, though: Sigmabench argues agent performance varies too much per codebase for public leaderboards to transfer, and a March 2026 METR note documents SWE-bench-passing patches that would not actually be merged — both independently motivating private, repo-native evals, which is also adjacent to Turing's repo-derived Code Review Bench.

Why it matters to Scott

SelfBench is a shipped productization of positions Scott's own wikis already argue: eval tasks mined from a repo's frozen git history as the behavioural oracle (counterfactual design replay / replay-driven design evolution, with future-leakage-rule as the validity caveat for any public-repo use), hidden tests as held-out acceptance checks (Hidden Gates), CI-bound gating (evaluation-driven development), and accuracy-vs-cost leaderboards as exactly the per-codebase measurement his model-barbell selection doctrine assumes. Mupt AI is an unidentified small org with an untested adoption claim, so this is wave-confirmation of the private repo-native evals cluster rather than a load-bearing challenge — but it's a dated-receipts publishing beat and a tool he could run on his own repos to ground agent-selection decisions.
ip:concept.counterfactual-design-replayip:concept.replay-driven-design-evolutionip:concept.future-leakage-ruleip:concept.characterisation-testingip:framework.hidden-gates-frameworkip:concept.evaluation-driven-developmentip:concept.model-barbelldev:concept.trace-backed-agent-comparisonradar:specific-real-swe-releaseradar:artificial-analysis-optimaradar:pairmark-blind-coding-agent-racesradar:agent-review-studio-local-evaluationradar:harbor-token-proxy-agentic-rlradar:frontierharness-17x-cost-variation
queries asked of Scott's wikis
  • private eval harness for coding agents on own codebase
  • SWE-bench validity critique benchmark contamination
  • evals as CI gate merged PR hidden tests
  • Harbor eval harness task format
  • agent selection accuracy vs cost per codebase
  • repo-derived eval task generation from git history

Measured heat

now 0 pts/hpeak 5 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 143h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-05 18:01 (minted)⭐ origin echo-reconstructedSelfBench: private coding-agent benchmarks built from a repository's own merged PRs — each PR rebuilt into an instruction, hidden tests, and
mupt-ai (avyayv) on github (echo) · attributed from hn.story.49967408 · published time unknown
—
10-05 16:58first on hacker news · published · lag ?Show HN: Self-bench – benchmark coding agents on real-world software
avyayv
—
10-05 16:58amplified on hacker news 👑hn.story.49967408
avyayv
peak 4 · 0 comments · 98% of case engagement
10-05 17:21our radar first saw it · lag ?discovery anchor: hn.story.49967408—
pace: p24 vs 1247 stories at the 96h mark (now 143h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Self-bench – benchmark coding agents on real-world software
Retrieved article excerpt

Open article · Retrieved 2026-10-05T17:32:37.152661+00:00

[SelfBench](https://selfbench.dev)

# SelfBench

**Find the best models for your repo.**  
Private coding-agent benchmarks built from your repository's own merged pull requests.

[Leaderboards at selfbench.dev](https://selfbench.dev)
[Benchmark your repo at app.selfbench.dev](https://app.selfbench.dev)

[CI](https://github.com/mupt-ai/self-bench/actions/workflows/ci.yml)
[License](https://github.com/mupt-ai/self-bench/blob/main/LICENSE)

[Browsing selfbench.dev: searching for a repository, opening its accuracy vs cost chart, and reading every model setting's score](https://selfbench.dev)

Public benchmarks tell you how a model does on someone else's code. SelfBench tells you how it does on yours: it turns your merged PRs into tasks with hidden tests, runs agents and models on them, and plots accuracy against cost.

## Results

Every open-source repository released on **[selfbench.dev](https://selfbench.dev)** gets a live leaderboard: each model, harness, and reasoning setting placed by accuracy and cost per task, with the Pareto frontier drawn through the settings nothing else beats on both.

**Browse the leaderboards:** [vercel/next.js](https://selfbench.dev/vercel/next.js) · [supabase/supabase](https://selfbench.dev/supabase/supabase) · [earendil-works/pi](https://selfbench.dev/earendil-works/pi) · [getsentry/sentry](https://selfbench.dev/getsentry/sentry) · [PostHog/posthog](https://selfbench.dev/PostHog/posthog) · [pingdotgg/t3code](https://selfbench.dev/pingdotgg/t3code) · [vercel/vercel](https://selfbench.dev/vercel/vercel) · **[all repositories →](https://selfbench.dev)**

## How it works

How SelfBench works: a merged PR is rebuilt from the commit before the change into an instruction, hidden tests, and a reference solution; Harbor's smoke, nop, oracle, and determinism gates and an independent review accept it; agents and models attempt every task; and accuracy is plotted against cost with the Pareto frontier

For each merged PR, SelfBench rebuilds the task from the commit before the change: the PR's own request becomes the instruction, and an authoring agent writes hidden tests and a reference solution. A task is accepted only if the tests fail without a solution, pass with the original implementation, pass again on a rerun, and survive an independent review. Every accepted task is a native [Harbor](https://harborframework.com/) task.

## Using SelfBench

Everything happens in the web app at [app.selfbench.dev](https://app.selfbench.dev):

1. **Sign in** with GitHub and **connect a repository**.
2. **Batch Generation**: choose how many easy, medium, and hard tasks to build, or use **Add PRs** on the Dataset page to build one task from each pull request you pick. A batch takes hours; it keeps running after you close the page.
3. **Dataset**: inspect each task (instruction, environment, hidden tests, reference patch, pipeline artifacts) and approve or reject it.
4. **Run**: pick models, harnesses, and a sandbox, and run them on the approved tasks.
5. **Results**: compare accuracy against cost, and open any trial's transcript and scores.
6. **Releases**: publish a public repository's results to [selfbench.dev](https://selfbench.dev).

Models and sandboxes run on your organization's own keys under **Credentials**. Read released results through the [Public Results API](https://github.com/mupt-ai/self-bench/blob/main/docs/api-reference/overview.mdx), or automate your workspace with an [API key](https://github.com/mupt-ai/self-bench/blob/main/docs/api-reference/workspace.mdx).

## Self-hosting

SelfBench's reference deployment runs on GCP: Cloud Run serves the API, while GKE Autopilot runs the Temporal workflow worker and KEDA-scaled Harbor jobs, with Cloud SQL and GCS. A Cloud Run worker pool remains available as an alternative.

1. Provision a GCP project, billing, Terraform state bucket, and GitHub Actions Workload Identity Federation.
2. Configure Terraform inputs and store each runtime secret value in its own Secret Manager secret.
3. Apply the environment with Terraform, or configure the protected GitHub `dev` and `prod` environments to deploy through Actions.
4. Point your domain at the provisioned load balancer and configure GitHub OAuth for the app URL.

For prerequisites, exact Terraform commands, runtime configuration, GitHub Actions setup, and GKE worker setup, see the [self-hosting and infrastructure guide](https://github.com/mupt-ai/self-bench/blob/main/infra/README.md).

## Development

Requires [Bun](https://bun.sh/) 1.3.14+ and Docker with Compose.

```
bun install --frozen-lockfile
bun run validate
```

Run the whole stack (API serving the app, worker, Temporal, Postgres, local Docker sandboxes) from a checkout:

```
cp .env.example .env   # GitHub OAuth app, session secret, credential key, managed keys
SELFBENCH_PUBLIC_URL=https://your-tunnel.example docker compose --profile sandbox up -d --build
docker compose port api 8080
```

Compose names the project after the checkout directory and publishes ephemeral host ports, so worktrees run side by side. Set `SELFBENCH_PUBLIC_URL` to the origin the browser opens (a tunnel or reverse proxy) before `up`, and register `<origin>/auth/github/callback` on the OAuth app.

For frontend work, `bun run dev:site` runs the API and Vite with hot reload (secrets in `.env.site`; see `scripts/dev-site.ts`).

## Documentation

- [Mintlify docs](https://github.com/mupt-ai/self-bench/blob/main/docs): guides, public results API, and workspace API (`cd docs && npm ci && npm run dev`)
- [How it works](https://github.com/mupt-ai/self-bench/blob/main/docs/concepts/task-generation.mdx): the pipeline and what makes a task valid
- [Public Results API](https://github.com/mupt-ai/self-bench/blob/main/docs/api-reference/overview.mdx) · [Workspace API](https://github.com/mupt-ai/self-bench/blob/main/docs/api-reference/workspace.mdx)
- [Infrastructure](https://github.com/mupt-ai/self-bench/blob/main/infra/README.md)

## License

[MIT](https://github.com/mupt-ai/self-bench/blob/main/LICENSE) © 2026 Mupt AI.
avyayv40
🟧 echo.github ⭐SelfBench: private coding-agent benchmarks built from a repository's own merged PRs — each PR rebuilt into an instruction, hidden tests, andmupt-ai (avyayv)——

Interpretation history

Decision trace