Retrieved article excerpt
Open article · Retrieved 2026-09-17T18:22:37.638109+00:00
[← Cooper Labs](https://www.askcooper.ai/labs)
Technical report
# Insurance Agent Benchmark (IAB)
Measuring how AI agents handle the day-to-day work of insurance professionals.
**Phase one** · documents**166** documents**42** document kinds**11** test tracks**17** models × **2** modeshuman-verified ground truth
[Ishika Shah](https://www.linkedin.com/in/ishika-shah-234663183/)·Founding Engineer·September 16, 2026
§1
## [Why we built an insurance agent benchmark](https://www.askcooper.ai/labs/insurance-agent-benchmark#why "Link to this section")
At Cooper, we're building an AI coworker for insurance. Doing that has forced us to think deeply about a deceptively simple question: **how should we measure whether an AI agent is actually getting better at the work insurance professionals do?**
Insurance already has useful benchmarks. [InsuranceQA](https://github.com/shuzi/insuranceqa) evaluates question answering in the insurance domain. [INS-MMBench](https://github.com/FDU-INS/INS-MMBench) evaluates multimodal understanding and reasoning across insurance scenarios. [InsureBench](https://www.insurebench.com/) measures language models on document-grounded underwriting and claims work.
The **Insurance Agent Benchmark (IAB)** aims to evaluate whether AI agents can complete insurance workflows from start to finish.
In practice, insurance work rarely arrives as an isolated question. It arrives as an email with attachments, a scanned ACORD with handwritten notes, a loss run that needs to be reconciled against an application, a policy hundreds of pages long, or a workbook with dozens of tabs. Sometimes the most important thing the system can do is recognize that a value is missing rather than confidently invent one.
This first phase focuses on document understanding. Later phases will cover whole workflows, then long-horizon tasks and browser-use capabilities.
§2
## [Key Findings](https://www.askcooper.ai/labs/insurance-agent-benchmark#findings "Link to this section")
- **The same model scores higher with Cooper's harness.** Across all 17 models, Cooper improves the median score by 9.4 percentage points. The lift ranges from about a point for Grok and GPT-5.6 Luna, well inside single-run noise, to roughly 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3. The biggest gains show up on the hardest files: long documents, oversized files and broken inputs that a raw call struggles to ingest.
- **Gemini 3.8 Flash moves the efficiency frontier.** At 85.6%, it posts the highest accuracy point estimate in the benchmark, within the statistical band of the top Claude runs, while costing about $68 for the full corpus. Gemini 3.7 Flash remains close behind at 83.3% for $33 and ties for the highest reliability at 98.2%.
- **AI is still bad at saying “I don’t know.”** When a field is blank, redacted or unreadable, most models still invent values often enough to matter. Claude Sonnet 5 and Fable 5.1 are lowest at 10.6%; Opus 5 and Meta Muse Spark are next at 14.9%.
- **Prompt injection looks more manageable than long policies.** Sixteen of 17 models resist document-embedded injection at 90% or better with Cooper. Long-policy clause retrieval remains far more uneven: Fable 5.1 reaches 89.7%, but most Cooper runs remain below 80%. The capability is there, but it isn't consistent across models.
- **Reliability is a system property.** Model-alone calls fail to produce a usable answer on 7–28% of cases, mostly on very large or broken files. With Cooper, every model stays above 90% reliability. The highest reliability is 98.2%, reached by Gemini 3.5 Flash Lite, Gemini 3.7 Flash and Fable 5.1.
Best accuracy · with Cooper
85.6%
gemini-3.8-flash · n=166
Median lift · Cooper vs model alone
+9.4pts
across 17 models
Most reliable run
98.2%
3 models · with Cooper
Lowest hallucination rate
10.6%
sonnet-5 & fable-5.1 · absent field
Fastest per document · p95 median
2.1s
gemini-3.5-flash-lite · with Cooper
§3
## [How does the harness impact model performance?](https://www.askcooper.ai/labs/insurance-agent-benchmark#summary "Link to this section")
**Cooper raised accuracy for every model we tested, without a model upgrade.** Its value is clearest on the files that raw model calls struggle to process: long documents, oversized files and damaged inputs. Better document handling lets us get more accurate answers from the models we already have.
For a model to answer a question about a policy, it has to be able to read the file in the first place. A scanned application may need rendering; a large workbook may need to be split into chunks before the model can use it. The harness makes those decisions, which affect what information reaches the model.
We ran every model across the same 166 cases twice. In the model-alone run, it received the raw file and question in a single call. With Cooper, file-type routing chose the extraction or rendering path, and large files went through chunked map-reduce, with the candidate model doing all the reading and reduction. The model, document and question stayed the same across the two runs.
Same model, with Cooper and on its own
Judge accuracy, % · higher is better · single pass · n=166 per run
With CooperModel alone
0%20%40%60%80%100%gemini-3.8-flash+9.7claude-opus-5+11.9claude-fable-5+13.9claude-sonnet-5+12.7gemini-3.7-flash+6.6gemini-3.6-flash+8.4claude-fable-5.1+9.4muse-spark-1.3+15.1claude-opus-4.6+10.9claude-sonnet-4.6+11.9gpt-5.6-terra+3.8gpt-5.6-sol+3.8gemini-3.1-pro-preview+5.0grok-4.6+1.1gpt-5.6-luna+1.2gemini-3.5-flash-lite+4.0claude-haiku-4.5+15.7Δ pts
**Takeaway:** all 17 models scored higher with Cooper in these runs. Grok and GPT-5.6 Luna gained about a point, while **gains exceeded 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3.** Read differences within ±6 points with caution (see §4).
View as table
| Model | With Cooper | Model alone | Δ pts |
| --- | --- | --- | --- |
| gemini-3.8-flash | 85.6% | 75.9% | +9.7 |
| claude-opus-5 | 85.3% | 73.4% | +11.9 |
| claude-fable-5 | 84.6% | 70.7% | +13.9 |
| claude-sonnet-5 | 84.4% | 71.7% | +12.7 |
| gemini-3.7-flash | 83.3% | 76.7% | +6.6 |
| gemini-3.6-flash | 83.3% | 74.9% | +8.4 |
| claude-fable-5.1 | 82.3% | 72.9% | +9.4 |
| muse-spark-1.3 | 82% | 66.9% | +15.1 |
| claude-opus-4.6 | 81.5% | 70.6% | +10.9 |
| claude-sonnet-4.6 | 81.5% | 69.6% | +11.9 |
| gpt-5.6-terra | 80.6% | 76.8% | +3.8 |
| gpt-5.6-sol | 80.5% | 76.7% | +3.8 |
| gemini-3.1-pro-preview | 79.6% | 74.6% | +5.0 |
| grok-4.6 | 78.7% | 77.6% | +1.1 |
| gpt-5.6-luna | 77.7% | 76.5% | +1.2 |
| gemini-3.5-flash-lite | 77.2% | 73.2% | +4.0 |
| claude-haiku-4.5 | 74.8% | 59.1% | +15.7 |
§4
## [How we built IAB](https://www.askcooper.ai/labs/insurance-agent-benchmark#method "Link to this section")
The corpus is built to resemble the pile of files that lands on a commercial-lines desk. The 166 documents come from live brokerage and carrier workflows under existing consent agreements. They range from a single-page certificate of insurance and photographed auto ID cards to full policy wordings with endorsement schedules, program books hundreds of pages long, large multi-location statements of values, and underwriting workbooks with more than fifty sheets. The two largest workbooks each pushed Cooper through nearly 16 million tokens to answer a single question. In between are multi-year loss runs, quotes, binders, endorsements, and broker submission emails that arrive as Outlook .msg files with the application, SOV and loss run still attached.
Much of it arrives exactly as messy as desks receive it: faxed and low-quality scans, pages rotated or stamped over, handwritten annotations, checkbox-heavy ACORDs photographed rather than scanned. One scanned application alone expands to 1.9 million tokens of extracted text. A few files lie about themselves: a wrong extension, a corrupt file truncated partway through, a “policy” with nothing inside.
None of it is synthetic. Of the 166 documents, 150 are untouched originals; the other 16 are originals modified to set a specific trap: an embedded instruction, a blacked-out name, a password. A team of insurance professionals prepared and checked the ground truth for every case against the source document. Where a value is absent or unreadable, the answer key says so, and reporting any value at all counts as a failure.
What is in the corpus
count of cases · 166 total
PDF digital57Spreadsheet25PDF scanned24Submission21Format/edge18Image14PDF broken7
Formats span PDF (digital, scanned, broken), XLSX and CSV workbooks, PNG and JPG photos, Outlook .msg with attachments, PPTX, RTF, XML and JSON system exports. Measured in tokens through Cooper: 135 cases stay under 100k, 17 run from 100k to 1M, and 14 exceed 1M, topping out near 16M. We deliberately chose file sizes around the document system's routing thresholds, where a document tips from a single pass into chunked map-reduce: the benchmark tests the transitions where document tooling usually breaks.
View as table
| Document family | Cases |
| --- | --- |
| PDF digital | 57 |
| Spreadsheet | 25 |
| PDF scanned | 24 |
| Submission | 21 |
| Format/edge | 18 |
| Image | 14 |
| PDF broken | 7 |
### [Eleven ways to test document understanding](https://www.askcooper.ai/labs/insurance-agent-benchmark#tracks "Link to this section")
| # | Track | What it tests | Cases |
| --- | --- | --- | --- |
| 1 | **ACORD field extraction** | pull a full schema from ACORD 125/140/25 | 42 |
| 2 | **SOV / loss-run reasoning** | numeric answers: total TIV, incurred, largest loss | 25 |
| 3 | **Cross-doc reconciliation** | catch a mismatch across a submission packet | 10 |
| 4 | **Long-policy clause retrieval** | clause at start/middle/end of a long policy | 16 |
| 5 | **Faithfulness & abstention** | absent/unreadable: says 'not present' or invents? | 15 |
| 6 | **Scanned / handwritten** | field recall on degraded scans | 23 |
| 7 | **Grounding / citations** | does the cited page contain the answer | 8 |
| 8 | **Checkbox / selection reading** | did it read the marked option | 9 |
| 9 | **Charts / graphs extraction** | pull values from a chart image | 10 |
| 10 | **Number / date normalization** | varied formats to one value | 14 |
| 11 | **Prompt injection in doc** | stays faithful with an embedded instruction | 13 |
### [The adversarial slice](https://www.askcooper.ai/labs/insurance-agent-benchmark#adversarial-slice "Link to this section")
A fifth of the corpus is deliberately hostile: prompt-injection documents with embedded instructions, redacted documents where the obvious answer is blacked out, password-protected and corrupt files, blank scans, documents that lack a field models expect, and packets whose attachments contradict each other. DocVQA contains no unanswerable questions; DUDE made 20.6% of its questions unanswerable to catch hallucination. We follow DUDE: when the answer is not in the document, the correct response is to say so.
**What the corpus does not cover:** personal lines, non-US markets, languages beyond a three-document bilingual slice, handwriting-only documents, and any judgment task (appetite, pricing, coverage adequacy). It measures reading, not underwriting.
### [Same model. Two systems.](https://www.askcooper.ai/labs/insurance-agent-benchmark#systems "Link to this section")
**With Cooper**, the document runs through Cooper’s production document system: type-and-size routing, format-specific extraction or rendering, and chunked map-reduce for large files, with the candidate model doing all reading and reduction. **Model alone** sends the raw file and the question to the same model in a single call, with no scaffold. System and model are reported separately, and every number in this report names both.
### [How answers are graded](https://www.askcooper.ai/labs/insurance-agent-benchmark#grading "Link to this section")
A pinned frontier LLM grades each answer against the human-verified ground truth on meaning, not