2026-10-11 16:37 UTC

Specific Labs claims its newly released Real-SWE benchmark finds tested model-and-harness combinations resolve at most 38.8% of private enterprise tasks, exposing a company-context and cross-service reliability gap relevant to production coding-agent deployment.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumcoding-agents agent-harnesses enterprise-software benchmarkingSpecific Labs

What is this?

Specific Labs announces Real-SWE as a benchmark evaluating frontier AI models on private production codebases licensed from real companies, using tasks drawn from those companies’ engineering work. Its release snippet emphasizes navigating proprietary systems and existing-product context rather than public repositories. The supplied snippets do not establish the claimed 38.8% maximum resolution rate, the tested model-and-harness combinations, or cross-service failures; those details and the proposed explanation for the performance gap remain unverified here.

Why it matters to Scott

Specific Labs’ move to tasks from private production codebases converges with Scott’s argument in Think in Whole Stories that coding agents need existing-product context, offering a concrete evaluation setting for that claim rather than just another public-repository score. No supplied radar hit tracks Real-SWE itself; however, the 38.8% ceiling, harness comparisons and cross-service failure explanation remain unverified, so this is a benchmark-methodology follow-up opportunity, not yet evidence against his production-coding playbook.
ip:source.think-in-whole-stories-why-ai-coding-agents-write-better-code-when-they-see-the-complete-picture-ebookip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:concept.coding-agent-benchmarksradar:github-copilot-production-trace-findingsradar:ship-harness-bench
queries asked of Scott's wikis
  • coding-agent harness evaluation model versus system performance
  • private repository benchmarks production readiness task resolution
  • company context agent memory codebase knowledge retrieval
  • cross-service coding tasks integration testing reliability
  • benchmark validity contamination representative workloads

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 692h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-12 21:28 (minted)⭐ origin echo-reconstructedSpecific Labs introduces Real-SWE using licensed private production codebases and native agent harnesses, reporting Fable 5.1 with Claude Co
Specific Labs on blog (echo) · attributed from hn.story.49676820 · published time unknown
—
09-12 20:25first on hacker news · published · lag ?Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
theanonymousone
—
09-12 20:25amplified on hacker news 👑hn.story.49676820
theanonymousone
peak 275 · 158 comments · 100% of case engagement
09-12 21:21our radar first saw it · lag ?discovery anchor: hn.story.49676820—
pace: p83 vs 1032 stories at the 336h mark (now 692h old) — ahead of claude-code-cache-ttl-analyzer (1.0x), behind bartowski-gguf-tensor-layouts (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnReal-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Retrieved article excerpt

Open article · Retrieved 2026-09-12T21:22:01.454733+00:00

September 2026

# Introducing Real-SWE

Benchmarking frontier AI models on private, real-world, enterprise codebases.

```
                                                                              .                         
;+;+;:;+;+;.#                                                                                           
;           ;                                                                                           
; .   . . . ;                                    .                                                      
+           ;                                                                                           
; .     .   : @;;+;+#                            +;+;+;+;:;+:#                                          
:           : ;     .       ;                    ;           ;              @;+;+.+;:;+;+*              
; . .       : ; . . :       ;       @++:;:;+.;;@ ;. . .   . .:     .        ; .     . . .;  @++;;;;+;:;@
;           : ;     ; @;+++:@;++;:+ ;          : :           . :+.+;;;+:;@  ;            :  +          .
; . .   .   ; ;   . ; : . . . .     ; . . . . .; ;  . . . . .+ +  . . . .;  ; . .   . . .;  ;     . .  :
;           : ;     ; ;           ; :          ; .           ; .         ;  ;            ;  ;          :
;       . . %   . . ; : . .   . . ; ; ... . . .; +. .   . . .: ;. . :.. .:  : . . ::. .  .  ; .   . . .;
:        *. @ ;     ; +.@     .   ; : +:    *+ : ;    :.     %+;    +.   ;  *.    *#     @: :.*.       ;
; .   . .;::@;; . . ; :#@;. ..@:. . ;.@%: .:@%:: :. . #+.   :@%;. .:@*. .:  @+  . #*. . ;@#.;.*   . . .:
.::: ..:%@@+@+#... :+:%@@+..:@@@::*.:;@@:..+@@:; *.::#@@#:..:@@#...+@#: :*:%@@+ .:.+::..*@@:@@@#;.. ::.#
        +@#.@         .*@:  .*@%:     *:   :@%.      ;@@:   .@%.    @     .;@@:   .:    :@#..#@*.       
        +@#.@          .@   .*@%:.    *:    %+       ;@@:    #;     +#    .;@@:   ..     @: .#@*.       
         @. ..         ;@   .*@%:     *:    @@       ;@@:    %#    * #    .;@@:          @+  .@         
        :*#..          @::    @.       .   ::;.       @@     ..     .       @%          .*@  ;@.        
        %.@            ..    .@;            .         @@     ..            .@@.          ..  @:#        
        ...             .    ;%#             .        @@                   +;#:          .. .@ @        
         .                   ...                      ...                   ..               ..         
                              ...                     ..                   ...                .
```

## 01Introduction

Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models on private, real-world, enterprise codebases. Each task comes from a private production codebase that we licensed from a real-world company. These are problems their engineers work on, with all the context and complexity that comes with an existing product.

- **Private codebases.** Agents must navigate proprietary systems whose code and solutions aren’t available on the public internet.
- **Work with business consequences.** Getting billing right, calculating taxes, migrating customers. Changes that affect how a business runs, often across multiple services.
- **Company-specific complexity.** Every company has its own rules and ways of writing code. Agents have to understand those conventions and make changes that work with what’s already there.

**Can a coding agent actually do the work of a software engineer in the real world?**

1. 1

   Fable 5.1

   Claude Code

   Resolution rate: 38.8%
2. 2

   GPT-6 Astra

   Codex CLI

   Resolution rate: 33.8%
3. 3

   Gemini 3.8 Flash

   Gemini CLI

   Resolution rate: 31.2%
4. 4

   GLM 5.3

   Claude Code

   Resolution rate: 28.8%
5. =5

   Grok 4.6

   Grok Build

   Resolution rate: 23.8%
6. =5

   Muse Spark 1.3

   Muse Code

   Resolution rate: 23.8%
7. 7

   Kimi K3

   Kimi Code

   Resolution rate: 18.8%
8. 8

   GPT-5.6 Sol

   Codex CLI

   Resolution rate: 16.2%

| # | Model | Harness | Resolution rate |
| --- | --- | --- | --- |
| 1 | Fable 5.1 | Claude Code | 38.8% |
| 2 | GPT-6 Astra | Codex CLI | 33.8% |
| 3 | Gemini 3.8 Flash | Gemini CLI | 31.2% |
| 4 | GLM 5.3 | Claude Code | 28.8% |
| =5 | Grok 4.6 | Grok Build | 23.8% |
| =5 | Muse Spark 1.3 | Muse Code | 23.8% |
| 7 | Kimi K3 | Kimi Code | 18.8% |
| 8 | GPT-5.6 Sol | Codex CLI | 16.2% |

Resolution rate is equivalent to pass@1, averaged over eight independent runs per task. 95% confidence intervals are shown.

Expert-generated or synthetic tasks can be well designed, but they aren’t the verbatim, actual tasks that engineers in real companies need to do. Our tasks differ on two axes: the underlying coding artifact and specificity of the instruction. Both add complexities that challenge today’s frontier models.

We use native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation.

### Real company tasks require company-specific context

### Correct billing depends on business rules and external services

Fix invoice billing so each business charges the right tax and exempt customers aren't taxed.

View full instructionHide full instruction▾

Billing reopens on Monday and every invoice this service issues is coming out untaxed. Each business on the platform settles its tax a different way: some maintain a rate themselves, some want each invoice priced against the buyer's destination by our tax authority provider, and some collect nothing at all, while a customer we hold an exemption for is charged nothing whichever way its business is configured. Pricing a destination means going to the authority with both addresses, the priced lines and the product category that business sells under, on the sandbox or the production authority according to the account the business is on; an address the authority refuses must be reported without stopping the invoice. The rate, the tax and the gross belong on the issued invoice, and once an invoice is settled the sale is filed back to the authority under that invoice's number so the returns reconcile. Invoices between European parties show both sides' VAT registrations. The authority and ledger are available at TAX\_JAR\_URL, PROD\_TAX\_JAR\_URL and INFLUX\_URL.

Services in the sandbox

- TJTaxJar sandbox
- TJTaxJar production
- InfluxDB ledger
- NestJS service
- TypeScript

### Agents work across code, infrastructure, and business tools

Tools and services across Real-SWE task environments. Each task exposes only the services its workflow needs.

- AWS emulator
- Docker
- Kubernetes
- GitHub
- Linear MCP
- PostgreSQL
- MySQL
- MongoDB
- GeGel
- Redis
- Go
- Python
- Node.js
- Vitest
- Slack
- Intercom
- Google Drive
- Email
- ClickUp

### Codebase Selection

We selected codebases through a rigorous screening process, focusing on real companies with substantial usage, strong engineering teams, and demanding production workloads. The sample tasks analyzed below come from these codebases, including:

- A Luma/Partiful competitor with **200K+ users** and a **top 100 App Store ranking**
- A consumer fintech platform processing **100K+ bank statements**
- Enterprise AI sales platforms supporting complex business workflows

We prioritize code written to meet an actual user or business need over code written solely to create a benchmark task. Production engineering requires understanding existing architecture, preserving behavior that users rely on, and making changes within real operational constraints.

### Brief instructions can require changes across many files

Our tasks describe the change needed, leaving agents to discover implementation details in the codebase and surrounding tools. Any behavior required by the verifier must be stated or reasonably discoverable. This leads to our prompts being slightly underspecified, about par with DeepSWE and Terminal Bench, but specific enough to not omit instructions.

The work is cross-functional and complex: a single change can span multiple parts of the application. Agents must understand existing business logic and company coding patterns while keeping the surrounding system working.

Prompt length · median

A typical Real-SWE instruction is 1,742 characters.

- FrontierCode2,056 chars
- DeepSWE1,975 chars
- Terminal-Bench 31,584 chars
- FrontierSWE v2992 chars
- Real-SWE1,742 chars

Files edited by the reference solution · median

11 files in Real-SWE, compared with 6 in FrontierCode and DeepSWE.

- FrontierCode6
- DeepSWE6
- Real-SWE11

All figures are medians. FrontierCode and DeepSWE use [Cognition's published comparison](https://cognition.com/blog/frontier-code); FrontierCode includes task descriptions and codebase guidelines. We measured instruction files from [Terminal-Bench 3's 74 tasks](https://github.com/harbor-framework/terminal-bench/tree/2b0442c3c583b710ca8da14c8e601b99f2f1f244/tasks), [FrontierSWE v2's 34 tasks](https://github.com/Proximal-Labs/frontier-swe-v2/tree/9e3f71cac38ef3d7e14a41b361c7b2b54c59899b/tasks), and Real-SWE's eight repository-backed sample tasks. Character counts are rounded to the nearest whole character. No comparable files-edited figure is included for Terminal-Bench 3 or FrontierSWE v2.

### Models fail even in short rollouts.

71.4% of rollouts under 10 minutes failed, compared with 73.4% of longer rollouts.

Triaging multiple systems and understanding requirements in codebases riddled with existing business logic and coding patterns is difficult.

Under 10 min

Under 10 min: 70 failed (71.4%) and 28 passed (28.6%), out of 98 rollouts.71.4%28.6%

70/98 failed

10 min or longer

10 min or longer: 398 failed (73.4%) and 144 passed (26.6%), out of 542 rollouts.73.4%26.6%

398/542 failed

- Failed
- Passed

Every task is inspired or lifted verbatim from a private, real-world codebase. We find these types of tasks super interesting for three reasons:

1. **Tasks on private codebases are natively out of distribution.** These types of coding tasks are not available anywhere on the internet and are unlikely to have ever been trained on by any other ai model. 99% of tokens in real-world enterprises are hidden away from the frontier models.
2. **These tasks are economically viable work.** Each task here has a direct relationship to spend and was assigned to an engineer earning a salary. Most benchmarks test interesting, experimental capabilities that are often unlikely to be widespread in the real-world.
3. **Company-specific engineering patterns matter.** Does AI code match the bar of a real-world enterprise? Our results show us that we're far from that reality. Many enterprises care about code standards and patterns. We've found that today's models are weaker at understanding company coding patterns and frequently miss requirements or don't verify their assumptions.

## 02Analysis

Here's an analysis of a small sample of tasks from our benchmark. If you're interested in the sample, [request access here](https://withspecific.com/benchmarks/real-swe/request-access).

### 6 of 10 tasks have resolution rates below 15%

Select a task to view model results. Percentages show the overall resolution rate.

Multi-region sweep67.2%⌄

- Fable 5.17/8 passed
- GPT-6 Astra8/8 passed
- Gemini 3.8 Flash8/8 passed
- GLM 5.32/8 passed
- Grok 4.63/8 passed
- Muse Spark 1.38/8 passed
- Kimi K32/8 passed
- GPT-5.6 Sol5/8 passed
API keys & environments65.6%⌄

- Fable 5.18/8 passed
- GPT-6 Astra5/8 passed
- Gemini 3.8 Flash7/8 passed
- GLM 5.35/8 passed
- Grok 4.64/8 passed
- Muse Spark 1.36/8 passed
- Kimi K30/8 passed
- GPT-5.6 Sol7/8 passed
Entitlement overage lines50.0%⌄

- Fable 5.18/8 passed
- GPT-6 Astra7/8 passed
- Gemini 3.8 Flash5/8 passed
- GLM 5.33/8 passed
- Grok 4.61/8 passed
- Muse Spark 1.31/8 passed
- Kimi K36/8 passed
- GPT-5.6 Sol1/8 passed
Customer identity migration40.6%⌄

- Fable 5.13/8 passed
- GPT-6 Astra1/8 passed
- Gemini 3.8 Flash3/8 passed
- GLM 5.34/8 passed
- Grok 4.68/8 passed
- 
theanonymousone275158
🟧 echo.blog ⭐Specific Labs introduces Real-SWE using licensed private production codebases and native agent harnesses, reporting Fable 5.1 with Claude CoSpecific Labs——

Interpretation history

Decision trace