2026-10-11 17:09 UTC

Independent comparisons will determine whether Chandra is a leading practical local PDF parser across tables, mathematics, handwriting, typography, and complex layouts.

state: expiredheat: lowuncertainty: highknownscott: lowdocument-parsing ocr rag local-inference

What is this?

Chandra is a Datalab-developed OCR/document-parsing model for converting complex documents into Markdown, HTML, or JSON while preserving layout information. Datalab says Chandra 2 is a 4B-parameter model supporting 90+ languages, tables, forms, mathematics, handwriting, and complex layouts, with an 85.9% score on the external olmOCR benchmark; its GitHub page notes that commercial self-hosting requires a license. The supplied snippets contain strong vendor claims and one favorable third-party comparison with Tesseract, but they do not include enough detail from the cited 14-capability benchmark to establish that Chandra is the leading practical local parser overall.

Why it matters to Scott

Scott already holds the relevant position in Capability Audit and Evaluation-Driven Development: practical parser claims should be settled through vendor-neutral tests on representative production documents. Chandra could affect his local OCR/RAG stack, but the supplied evidence does not establish the 14-capability results or a practical advantage, so this currently adds no actionable finding beyond the validation pattern already tracked for Nemotron Parse 2.0.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferencedev:concept.source-native-semantic-chunkingradar:nemotron-parse-2-validationradar:concept.document-parsingradar:concept.model-evaluationradar:concept.ai-benchmarks
queries asked of Scott's wikis
  • PDF parsing quality for RAG ingestion
  • local document intelligence and OCR stack
  • layout-aware chunking for tables and mathematics
  • benchmark methodology for document parsers
  • self-hosted model licensing and local inference
  • structured document extraction into Markdown or JSON

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI compared even more parsers on 14 PDF-parsing capabilities using different types
LocalLLaMA
LowerGears17922
🟧 echo.github ⭐The primary artifact is the author's benchmark repository and commit, not an upstream article. Its README says: “Eight open-source PDF parseAlaa Mroue (alaamroue)——

Interpretation history

Decision trace