2026-10-11 17:14 UTC

SecondState's ex-EY team claims its released FAB benchmark โ€” 50 tasks, 160 documents and 231 grading criteria in a synthetic data room, with published traces โ€” shows frontier agents finding relevant financial facts but failing to carry them through to complete, reliable due-diligence analysis; adoption by evaluators and labs, or expansion to more companies and models, would make FAB the reference benchmark for long-horizon financial agent work.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation financial-diligence-agents benchmarksSecondStateELAmrani

What is this?

FAB is an open-source benchmark launched via Show HN by SecondState, a company fielding an ex-EY team, for evaluating LLM agents on financial due diligence: agents work through 50 tasks inside a synthetic data room of 160 documents, graded against 231 criteria, with execution traces published for inspection. The team's headline finding is that frontier agents can locate the relevant financial facts but fail to carry them through to complete, reliable diligence analyses. Notably, the supplied web results never mention FAB, SecondState, or ELAmrani directly โ€” the 'FAB' that does surface is Vals AI's separate proprietary 'Finance Agent Benchmark' (best model ~60% partial credit), a distinct artifact which, alongside UC Berkeley's Agents' Last Exam (0% on the hardest professional tier), maps the crowded long-horizon agent-eval landscape FAB is entering. So the general 'agents find facts but fail at sustained professional work' pattern is independently echoed by adjacent benchmarks, but FAB's own existence beyond the Show HN launch and any traction toward becoming the reference benchmark rest on the claim owner's positioning, not corroborated coverage.

Why it matters to Scott

An ex-EY team independently lands where Scott's canon already sits โ€” FAB's graded traces show navigation and fact-finding succeeding while carried-through synthesis fails, the navigation/synthesis split his answer-failure-classes taxonomy and long-running-agents framework predict โ€” and its 231-criterion rubric grading plus published traces mirror his llm-rubric-grading and agent-receipts practices, in the exact professional-services diligence territory his FDE practice factory targets. Medium rather than high because it is the second finance-agent failure benchmark after the radar's ATLAS-Finance episode and its claim to become the reference eval rests on adoption that hasn't happened; the dated-receipts play is checking whether FAB's grading separates failure classes the way his taxonomy does.
ip:concept.answer-failure-classesip:framework.long-running-agentsdev:concept.fde-practice-factorydev:concept.llm-rubric-gradingip:concept.agent-receiptsradar:atlas-finance-workplace-benchmarkradar:concept.agent-benchmarksradar:concept.long-horizon-agentsradar:concept.agent-evaluationradar:concept.agent-observability
queries asked of Scott's wikis
  • agent evaluation harness rubric-based grading
  • long-horizon agent failures dropped intermediate findings
  • traces as first-class artifacts agent debugging
  • RAG evaluation retrieval quality vs final synthesis
  • vertical domain evals moat product pattern
  • synthetic document corpus for testing agents

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 362h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-26 14:00โญ origin echo-reconstructedFAB is an open-source benchmark for LLM agents doing financial due diligence with 50 tasks, 160 documents and 231 grading criteria; its publ
SecondState (ELAmrani, ex-EY team) on github (echo) ยท attributed from hn.story.49874360
โ€”
09-28 06:34first on hacker news ยท published ยท +40.6hShow HN: FAB โ€“ A benchmark for AI agents doing financial due diligence
ELAmrani
โ€”
09-28 06:34amplified on hacker news ๐Ÿ‘‘hn.story.49874360
ELAmrani
peak 2 ยท 0 comments ยท 98% of case engagement
09-28 07:20our radar first saw it ยท +41.4hdiscovery anchor: hn.story.49874360โ€”
pace: p9 vs 1032 stories at the 336h mark (now 362h old) โ€” behind addom-local-coding-harness (0.5x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: FAB โ€“ A benchmark for AI agents doing financial due diligence
Retrieved article excerpt

Open article ยท Retrieved 2026-09-28T07:25:14.832128+00:00

# FAB โ€” Finance Agents Benchmark

FAB is an open-source project for benchmarking LLM agents' ability to perform
financial due diligence in a synthetic company data room.

FAB consists of two parts: a dataset of *tasks* containing agent instructions,
documents and rubrics, and an *execution harness* for running and evaluating agents
against those tasks. The current release contains **50 tasks, 160 documents and
231 grading criteria** for one company, Meridian Industrial Supply LLC.

## Getting Started

Start with [the walkthrough](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/docs/tutorial.md) for setup, task inspection, running
an agent and reviewing its scores. Requires Python 3.11+, [uv](https://docs.astral.sh/uv/),
Docker or Podman, and model API credentials.

```
git clone https://github.com/SecondState-ai/finance-agents-benchmark.git
cd finance-agents-benchmark
uv sync --locked
cp -n .env.example .env.local
# Fill in your API keys in .env.local before running.
./scripts/run-task --task 001 --model gpt-6-luna
./scripts/grade --run results/001/gpt-6-luna/<timestamp>
```

`OPENAI_API_KEY` is required for the judge and OpenAI agents; `FW_API_KEY` is only
needed for Fireworks agents. Replace `<timestamp>` with the run directory printed
by the runner.

## Results

Four models, three trials on **all 50 tasks (600 answers)**. The judge is
`gpt-6-luna` at maximum reasoning. A task passes only when every criterion passes.

| Model | Task pass rate | Criterion pass rate | Passed at least once | Passed all three |
| --- | --- | --- | --- | --- |
| DeepSeek V4.1 Flash | 60.0% | 81.0% | 38/50 | 23/50 |
| GPT-6 Sol | 58.7% | 83.4% | 35/50 | 24/50 |
| GPT-6 Luna | 50.7% | 79.5% | 33/50 | 19/50 |
| GLM 5.3 Flash | 47.3% | 76.2% | 30/50 | 18/50 |

Results describe one synthetic company. See the report for
per-trial scores, usage, recovery, task assumptions and evaluation limitations.

## Documentation and Data

| Resource | Contents |
| --- | --- |
| [Walkthrough](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/docs/tutorial.md) | Setup, dataset, task format, sandbox, running and grading |
| [Results report](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/reports/meridian/results-2026-09-27.json) | Criterion verdicts, judge reasoning, usage and recovery |
| [Trial CSV](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/reports/meridian/results-2026-09-27-trials.csv) | One row per answer |
| [Model answers and traces](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/results/published) | All 600 completed runs, final grades, metadata and tool traces |
| [Hugging Face](https://huggingface.co/datasets/secondstate/finance-agents-benchmark) | Versioned data room, tasks and question index |
| [Agent traces dataset](https://huggingface.co/datasets/secondstate/finance-agents-benchmark-traces) | Loadable traces, answers, rubrics and final grades for all 600 runs |

## License and Citation

Code: [MIT](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/LICENSE). Data, tasks, rubrics and results: [CC BY 4.0](https://github.com/SecondState-ai/finance-agents-benchmark/blob/main/LICENSE-DATA).

```
@misc{secondstatefab2026,
  title = {FAB: Finance Agents Benchmark},
  author = {{SecondState}},
  year = {2026},
  url = {https://github.com/SecondState-ai/finance-agents-benchmark}
}
```

Include the code and dataset revisions when reporting results.
ELAmrani20
๐ŸŸง echo.github โญFAB is an open-source benchmark for LLM agents doing financial due diligence with 50 tasks, 160 documents and 231 grading criteria; its publSecondState (ELAmrani, ex-EY team)โ€”โ€”

Interpretation history

Decision trace