2026-10-11 18:02 UTC

Independent use will determine whether ExtractBench provides reproducible schema-extraction evaluations that reveal meaningful reliability differences among models and extraction systems.

state: expiredheat: lowuncertainty: highknownscott: mediumllm-evaluation structured-extraction benchmarksLlamaIndex

What is this?

ExtractBench is presented as an open-source benchmark for evaluating schema extraction and comparing reliability across models or extraction systems. The case associates it with LlamaIndex, but the supplied evidence does not establish LlamaIndex’s exact role, benchmark methodology, datasets, metrics, or any demonstrated model differences. The web results are unrelated, so reproducibility and practical value remain ungrounded pending independent use and fuller documentation.

Why it matters to Scott

The core position is already held in “Evaluation-Driven Development” and “Model-Plus-Harness Benchmark Unit”: schema extraction reliability should be measured with repeatable evaluations of the complete extraction system, not model claims alone. ExtractBench could provide reusable fixtures for Scott’s validation-gated extraction and Scrape work, but absent methodology, datasets, or independent results, it does not yet extend or challenge those positions.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitdev:concept.validation-gated-llm-extractiondev:project.scraperadar:concept.model-evaluationradar:concept.ai-benchmarks
queries asked of Scott's wikis
  • structured-output and schema-extraction reliability
  • LLM evaluation harnesses and reproducible benchmarks
  • benchmark contamination and evaluation validity
  • JSON schema validation and extraction failure modes
  • model comparison for structured data pipelines
  • production evals for RAG and knowledge ingestion

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: ExtractBench, an open-source schema extraction benchmarkcheesyFish50
🟧 echo.github ⭐Released an open-source benchmark for evaluating schema extraction.run-llama——

Interpretation history

Decision trace