2026-10-11 17:15 UTC

Cooper claims its document-processing harness raises median accuracy by 9.4 percentage points across 17 models on its 166-document Insurance Agent Benchmark, suggesting routing and ingestion improvements can materially improve insurance document understanding without model upgrades.

state: seedheat: mediumuncertainty: mediumconvergesscott: lowagent-benchmarks agent-harnesses document-understandingCooperIshika Shah

What is this?

Cooper presents itself in its own white-paper snippet as an AI offering for commercial insurance agents and brokers, addressing work such as document intake, form filling, and submissions. The case attributes to Cooper an Insurance Agent Benchmark covering 166 documents and a claimed median accuracy lift of 9.4 percentage points across 17 models from its document-processing harness, with phase one limited to document understanding rather than complete workflows. None of the supplied web snippets establishes those benchmark details, the routing/ingestion explanation, or Ishika Shah’s role; the benchmark findings therefore remain unverified case claims.

Why it matters to Scott

Cooper’s claimed cross-model harness lift aligns with Scott’s Model-Plus-Harness Benchmark Unit and Give the Agent a Workshop position that model capability is not system capability; the supplied radar hits track related harness comparisons, but not this insurance benchmark. As supplied, this is another unverified example rather than actionable evidence: neither the routing/ingestion mechanism nor complete-workflow gains are established, so it does not yet change Scott’s evaluation practice or implementation choices.
ip:concept.model-plus-harness-benchmark-unitip:source.give-the-agent-a-workshop-ebookradar:ship-harness-benchradar:stencil-harness-coding-improvementradar:concept.agent-benchmarksradar:concept.agent-harnesses
queries asked of Scott's wikis
  • harness engineering versus model upgrades accuracy gains
  • document ingestion parsing routing extraction pipelines
  • agent evaluation component benchmarks versus end-to-end workflows
  • cross-model harness comparisons benchmark validity
  • vertical AI document workflows measured accuracy

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedPhase one evaluates document understanding, not complete insurance workflows; Cooper reports a 9.4-point median harness lift across 17 model
Ishika Shah, Cooper on blog (echo) · attributed from hn.story.49744109
—
09-17 17:44first on hacker news · published · +51.7hInsurance Agent Benchmark: 166 real-world cases for evaluating insurance AI
jsc39
—
09-17 17:44amplified on hacker news 👑hn.story.49744109
jsc39
peak 7 · 0 comments · 99% of case engagement
09-17 18:20our radar first saw it · +52.4hdiscovery anchor: hn.story.49744109—
pace: p39 vs 1032 stories at the 336h mark (now 626h old) — ahead of agentsec-static-config-auditing (1.2x), behind anthropic-meta-lawsuit (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnInsurance Agent Benchmark: 166 real-world cases for evaluating insurance AI
Retrieved article excerpt

Open article · Retrieved 2026-09-17T18:22:37.638109+00:00

[← Cooper Labs](https://www.askcooper.ai/labs)

Technical report

# Insurance Agent Benchmark (IAB)

Measuring how AI agents handle the day-to-day work of insurance professionals.

**Phase one** · documents**166** documents**42** document kinds**11** test tracks**17** models × **2** modeshuman-verified ground truth

[Ishika Shah](https://www.linkedin.com/in/ishika-shah-234663183/)·Founding Engineer·September 16, 2026

§1

## [Why we built an insurance agent benchmark](https://www.askcooper.ai/labs/insurance-agent-benchmark#why "Link to this section")

At Cooper, we're building an AI coworker for insurance. Doing that has forced us to think deeply about a deceptively simple question: **how should we measure whether an AI agent is actually getting better at the work insurance professionals do?**

Insurance already has useful benchmarks. [InsuranceQA](https://github.com/shuzi/insuranceqa) evaluates question answering in the insurance domain. [INS-MMBench](https://github.com/FDU-INS/INS-MMBench) evaluates multimodal understanding and reasoning across insurance scenarios. [InsureBench](https://www.insurebench.com/) measures language models on document-grounded underwriting and claims work.

The **Insurance Agent Benchmark (IAB)** aims to evaluate whether AI agents can complete insurance workflows from start to finish.

In practice, insurance work rarely arrives as an isolated question. It arrives as an email with attachments, a scanned ACORD with handwritten notes, a loss run that needs to be reconciled against an application, a policy hundreds of pages long, or a workbook with dozens of tabs. Sometimes the most important thing the system can do is recognize that a value is missing rather than confidently invent one.

This first phase focuses on document understanding. Later phases will cover whole workflows, then long-horizon tasks and browser-use capabilities.

§2

## [Key Findings](https://www.askcooper.ai/labs/insurance-agent-benchmark#findings "Link to this section")

- **The same model scores higher with Cooper's harness.** Across all 17 models, Cooper improves the median score by 9.4 percentage points. The lift ranges from about a point for Grok and GPT-5.6 Luna, well inside single-run noise, to roughly 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3. The biggest gains show up on the hardest files: long documents, oversized files and broken inputs that a raw call struggles to ingest.
- **Gemini 3.8 Flash moves the efficiency frontier.** At 85.6%, it posts the highest accuracy point estimate in the benchmark, within the statistical band of the top Claude runs, while costing about $68 for the full corpus. Gemini 3.7 Flash remains close behind at 83.3% for $33 and ties for the highest reliability at 98.2%.
- **AI is still bad at saying “I don’t know.”** When a field is blank, redacted or unreadable, most models still invent values often enough to matter. Claude Sonnet 5 and Fable 5.1 are lowest at 10.6%; Opus 5 and Meta Muse Spark are next at 14.9%.
- **Prompt injection looks more manageable than long policies.** Sixteen of 17 models resist document-embedded injection at 90% or better with Cooper. Long-policy clause retrieval remains far more uneven: Fable 5.1 reaches 89.7%, but most Cooper runs remain below 80%. The capability is there, but it isn't consistent across models.
- **Reliability is a system property.** Model-alone calls fail to produce a usable answer on 7–28% of cases, mostly on very large or broken files. With Cooper, every model stays above 90% reliability. The highest reliability is 98.2%, reached by Gemini 3.5 Flash Lite, Gemini 3.7 Flash and Fable 5.1.

Best accuracy · with Cooper

85.6%

gemini-3.8-flash · n=166

Median lift · Cooper vs model alone

+9.4pts

across 17 models

Most reliable run

98.2%

3 models · with Cooper

Lowest hallucination rate

10.6%

sonnet-5 & fable-5.1 · absent field

Fastest per document · p95 median

2.1s

gemini-3.5-flash-lite · with Cooper

§3

## [How does the harness impact model performance?](https://www.askcooper.ai/labs/insurance-agent-benchmark#summary "Link to this section")

**Cooper raised accuracy for every model we tested, without a model upgrade.** Its value is clearest on the files that raw model calls struggle to process: long documents, oversized files and damaged inputs. Better document handling lets us get more accurate answers from the models we already have.

For a model to answer a question about a policy, it has to be able to read the file in the first place. A scanned application may need rendering; a large workbook may need to be split into chunks before the model can use it. The harness makes those decisions, which affect what information reaches the model.

We ran every model across the same 166 cases twice. In the model-alone run, it received the raw file and question in a single call. With Cooper, file-type routing chose the extraction or rendering path, and large files went through chunked map-reduce, with the candidate model doing all the reading and reduction. The model, document and question stayed the same across the two runs.

Same model, with Cooper and on its own

Judge accuracy, % · higher is better · single pass · n=166 per run

With CooperModel alone

0%20%40%60%80%100%gemini-3.8-flash+9.7claude-opus-5+11.9claude-fable-5+13.9claude-sonnet-5+12.7gemini-3.7-flash+6.6gemini-3.6-flash+8.4claude-fable-5.1+9.4muse-spark-1.3+15.1claude-opus-4.6+10.9claude-sonnet-4.6+11.9gpt-5.6-terra+3.8gpt-5.6-sol+3.8gemini-3.1-pro-preview+5.0grok-4.6+1.1gpt-5.6-luna+1.2gemini-3.5-flash-lite+4.0claude-haiku-4.5+15.7Δ pts

**Takeaway:** all 17 models scored higher with Cooper in these runs. Grok and GPT-5.6 Luna gained about a point, while **gains exceeded 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3.** Read differences within ±6 points with caution (see §4).

View as table

| Model | With Cooper | Model alone | Δ pts |
| --- | --- | --- | --- |
| gemini-3.8-flash | 85.6% | 75.9% | +9.7 |
| claude-opus-5 | 85.3% | 73.4% | +11.9 |
| claude-fable-5 | 84.6% | 70.7% | +13.9 |
| claude-sonnet-5 | 84.4% | 71.7% | +12.7 |
| gemini-3.7-flash | 83.3% | 76.7% | +6.6 |
| gemini-3.6-flash | 83.3% | 74.9% | +8.4 |
| claude-fable-5.1 | 82.3% | 72.9% | +9.4 |
| muse-spark-1.3 | 82% | 66.9% | +15.1 |
| claude-opus-4.6 | 81.5% | 70.6% | +10.9 |
| claude-sonnet-4.6 | 81.5% | 69.6% | +11.9 |
| gpt-5.6-terra | 80.6% | 76.8% | +3.8 |
| gpt-5.6-sol | 80.5% | 76.7% | +3.8 |
| gemini-3.1-pro-preview | 79.6% | 74.6% | +5.0 |
| grok-4.6 | 78.7% | 77.6% | +1.1 |
| gpt-5.6-luna | 77.7% | 76.5% | +1.2 |
| gemini-3.5-flash-lite | 77.2% | 73.2% | +4.0 |
| claude-haiku-4.5 | 74.8% | 59.1% | +15.7 |

§4

## [How we built IAB](https://www.askcooper.ai/labs/insurance-agent-benchmark#method "Link to this section")

The corpus is built to resemble the pile of files that lands on a commercial-lines desk. The 166 documents come from live brokerage and carrier workflows under existing consent agreements. They range from a single-page certificate of insurance and photographed auto ID cards to full policy wordings with endorsement schedules, program books hundreds of pages long, large multi-location statements of values, and underwriting workbooks with more than fifty sheets. The two largest workbooks each pushed Cooper through nearly 16 million tokens to answer a single question. In between are multi-year loss runs, quotes, binders, endorsements, and broker submission emails that arrive as Outlook .msg files with the application, SOV and loss run still attached.

Much of it arrives exactly as messy as desks receive it: faxed and low-quality scans, pages rotated or stamped over, handwritten annotations, checkbox-heavy ACORDs photographed rather than scanned. One scanned application alone expands to 1.9 million tokens of extracted text. A few files lie about themselves: a wrong extension, a corrupt file truncated partway through, a “policy” with nothing inside.

None of it is synthetic. Of the 166 documents, 150 are untouched originals; the other 16 are originals modified to set a specific trap: an embedded instruction, a blacked-out name, a password. A team of insurance professionals prepared and checked the ground truth for every case against the source document. Where a value is absent or unreadable, the answer key says so, and reporting any value at all counts as a failure.

What is in the corpus

count of cases · 166 total

PDF digital57Spreadsheet25PDF scanned24Submission21Format/edge18Image14PDF broken7

Formats span PDF (digital, scanned, broken), XLSX and CSV workbooks, PNG and JPG photos, Outlook .msg with attachments, PPTX, RTF, XML and JSON system exports. Measured in tokens through Cooper: 135 cases stay under 100k, 17 run from 100k to 1M, and 14 exceed 1M, topping out near 16M. We deliberately chose file sizes around the document system's routing thresholds, where a document tips from a single pass into chunked map-reduce: the benchmark tests the transitions where document tooling usually breaks.

View as table

| Document family | Cases |
| --- | --- |
| PDF digital | 57 |
| Spreadsheet | 25 |
| PDF scanned | 24 |
| Submission | 21 |
| Format/edge | 18 |
| Image | 14 |
| PDF broken | 7 |

### [Eleven ways to test document understanding](https://www.askcooper.ai/labs/insurance-agent-benchmark#tracks "Link to this section")

| # | Track | What it tests | Cases |
| --- | --- | --- | --- |
| 1 | **ACORD field extraction** | pull a full schema from ACORD 125/140/25 | 42 |
| 2 | **SOV / loss-run reasoning** | numeric answers: total TIV, incurred, largest loss | 25 |
| 3 | **Cross-doc reconciliation** | catch a mismatch across a submission packet | 10 |
| 4 | **Long-policy clause retrieval** | clause at start/middle/end of a long policy | 16 |
| 5 | **Faithfulness & abstention** | absent/unreadable: says 'not present' or invents? | 15 |
| 6 | **Scanned / handwritten** | field recall on degraded scans | 23 |
| 7 | **Grounding / citations** | does the cited page contain the answer | 8 |
| 8 | **Checkbox / selection reading** | did it read the marked option | 9 |
| 9 | **Charts / graphs extraction** | pull values from a chart image | 10 |
| 10 | **Number / date normalization** | varied formats to one value | 14 |
| 11 | **Prompt injection in doc** | stays faithful with an embedded instruction | 13 |

### [The adversarial slice](https://www.askcooper.ai/labs/insurance-agent-benchmark#adversarial-slice "Link to this section")

A fifth of the corpus is deliberately hostile: prompt-injection documents with embedded instructions, redacted documents where the obvious answer is blacked out, password-protected and corrupt files, blank scans, documents that lack a field models expect, and packets whose attachments contradict each other. DocVQA contains no unanswerable questions; DUDE made 20.6% of its questions unanswerable to catch hallucination. We follow DUDE: when the answer is not in the document, the correct response is to say so.

**What the corpus does not cover:** personal lines, non-US markets, languages beyond a three-document bilingual slice, handwriting-only documents, and any judgment task (appetite, pricing, coverage adequacy). It measures reading, not underwriting.

### [Same model. Two systems.](https://www.askcooper.ai/labs/insurance-agent-benchmark#systems "Link to this section")

**With Cooper**, the document runs through Cooper’s production document system: type-and-size routing, format-specific extraction or rendering, and chunked map-reduce for large files, with the candidate model doing all reading and reduction. **Model alone** sends the raw file and the question to the same model in a single call, with no scaffold. System and model are reported separately, and every number in this report names both.

### [How answers are graded](https://www.askcooper.ai/labs/insurance-agent-benchmark#grading "Link to this section")

A pinned frontier LLM grades each answer against the human-verified ground truth on meaning, not
jsc3970
🟧 echo.blog ⭐Phase one evaluates document understanding, not complete insurance workflows; Cooper reports a 9.4-point median harness lift across 17 modelIshika Shah, Cooper——

Interpretation history

Decision trace