Skip to content
CoreNovusPublic

About

Deterministic evaluation for document parsing and semantic extraction: markdown fidelity (NED, TEDS, heading-tree F1, Spearman reading order) and extraction accuracy with explicit hallucination / over-abstention rates. Stdlib core, pytest plugin, three corpus tiers.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

DocEval

Check how well a document was converted — and read the score without fooling yourself.

Licence: Apache-2.0 CI Python 3.10+ Core dependencies: none


You built something that turns PDFs, Word files or HTML into Markdown — or that pulls fields out of invoices and reports. How good is it? "It looks fine" does not survive a refactor, and a made-up percentage is worse than no number.

DocEval gives you real numbers and tells you what they do and do not mean — including which of them come from the research benchmarks and which are ours. Four are published metrics computed the published way; the assertion axis follows a published methodology; the rest are DocEval's own, and are marked as such. A house metric is not a worse metric, but it is not comparable with a paper's figures, and you should know which you are quoting.

  • No model, no network, no API key. Every score is a plain calculation over your output and your reference. Same input, same score, on any machine.
  • The core has no dependencies at all. Standard library only.
  • It refuses to flatter you. A score that cannot be measured is reported as n/a, never as zero. Averages print how many documents they cover. A blank output scores nothing instead of passing by default.

Two questions it answers

Did the conversion keep the document? Headings, tables, text, reading order, bold/italic/code.

Are the extracted fields right — and did the model invent anything? Plus the question most tools skip: did it refuse to answer something the document does answer?

Both work on your own documents. You do not need a benchmark corpus to start, and you do not need a reference document either — see below.

Quickest start: inside Claude Code

/plugin marketplace add CoreNovus/doc-eval
/plugin install doc-eval@doc-eval

That is the whole setup. Then just ask:

Is my converter losing the tables in these PDFs?

No pip install, no configuration. The plugin carries its own scorer, because the core needs nothing but Python.

Or as a Python package

pip install "doc-eval[cli,corpus-gen] @ git+https://github.com/CoreNovus/doc-eval"
from doc_eval import assert_a1_fidelity, assert_a2_accuracy

result = assert_a1_fidelity(my_markdown, reference_markdown)
result.require(min_heading_tree_f1=0.90, min_teds_struct=0.95)

risk = assert_a2_accuracy(extracted_fields, expected_fields)
risk.require(max_inventing_on_absent=0.0)  # zero tolerance for invented values

Use require() rather than assert result.teds_struct >= 0.95. A document with no table has no table score, and comparing that to a number raises TypeError — technically true, and no help at all. require() skips what it cannot measure and says which metric it skipped.

What gets measured

Conversion

Metric In plain words Source
Syntax rules Is the Markdown even well-formed — table rows aligned, code fences closed, links complete ours
NED How much of the text survived, character by character OmniDocBench
TEDS How close the tables are, contents included PubTabNet
TEDS-Struct How close the tables are in shape, ignoring the cell text PubTabNet
Heading tree F1 Did the section hierarchy survive, levels included ours
Reading order Did the sections come out in the right order OmniDocBench
Inline F1 Did bold, italic, code and inline maths survive ours
Content checks Did the specific things you declared survive olmOCR-Bench methodology

Extraction

Metric In plain words Source
ANLS Text match that forgives small OCR slips ST-VQA / DocVQA
ANLS* List and object match that credits near-misses inside an item ANLS* paper
Normalized-exact Strict match after tidying case, spacing and date format ours
Made things up Answered a question the document does not answer ours
Refused to answer Left blank a question the document does answer ours

The last two are reported separately on purpose: their fixes pull in opposite directions, so one combined "risk score" would let one get worse while the other improved and show no change at all.

You do not need a reference document

Writing reference Markdown for a real PDF takes hours, and two careful people produce different answers. So DocEval also scores against checks — small, unarguable statements about the document:

[
  {"kind": "must_contain",     "value": "383,285"},
  {"kind": "must_not_contain", "value": "Page 3 of 12"},
  {"kind": "heading_level",    "level": 2, "value": "Revenue"},
  {"kind": "order",            "first": "Introduction", "then": "Conclusion"}
]

"The total appears" and "the running footer does not" are cheap to write and impossible to argue with. This is how the published OCR benchmarks do it.

Test corpora, if you want a benchmark

doc-eval corpora lists three kinds, and for each one says where its reference came from — the property that decides how far to trust it:

Size Reference comes from Good for
synth-v1 25 files: 6 source documents, spread across up to 8 formats and 3 languages each (not every source document ships every format) The same declaration that writes the file, so it cannot drift Regression: did my change alter anything?
wild-v1 3 real documents Hand-written, checked against the real bytes Smoke-testing on genuinely messy input
olmocr-textual-v1 657 documents, 2,605 checks Ships with the benchmark (ODC-BY-1.0) Comparing two converters

Every downloaded file is pinned by SHA-256, so a run today and a run next year measure the same bytes — or fail loudly.

Getting one

No test document ships with this project. Not in the repository, not in the package. What ships is a manifest per corpus — a URL, a SHA-256 and a licence per document — and the CLI fetches or generates the bytes on demand into a cache outside your project. Recording a document's licence is a much smaller claim than redistributing it, which is why the corpora can name real published documents at all.

doc-eval corpora                          # what exists, and where each reference came from
doc-eval build-corpus synth-v1            # generates locally; needs the corpus-gen extra
doc-eval fetch-corpus wild-v1             # downloads, and verifies every SHA-256
doc-eval fetch-corpus olmocr-textual-v1   # the olmOCR-Bench importer
doc-eval verify-corpus synth-v1           # re-reads each generated file with another reader

synth-v1 is the one that needs no network: it is written from declarations inside the package, so the [cli,corpus-gen] install above is all it takes. The other two are real published documents and have to be downloaded. verify-corpus applies to the generated corpus only — it exists to catch a file that is on disk and unreadable, which is a failure mode a corpus you did not write cannot have.

A digest that does not match is a hard failure that leaves nothing behind, never a warning: the bytes land in a temporary file and are moved into place only once they hash correctly. Reporting a mismatch and keeping the file would be the worst of both, since every later run would score against unverified input while believing it pinned.

How much can each one actually prove?

Worth stating plainly, because "we scored higher" is a claim about statistics and these are the numbers behind it. Smallest difference each corpus can distinguish from noise, comparing two converters on the same documents (95% confidence, 80% power):

Corpus Independent documents Smallest real difference
synth-v1 6 ~63 percentage points
wild-v1 3 ~89 pp
olmocr-textual-v1 657 ~6 pp

synth-v1 counts 6, not 25: the same six source documents rendered into eight formats are not 25 independent observations, and treating them as such would roughly halve the number above and be wrong.

So: use synth-v1 to catch regressions — it is deterministic, so any change at all shows up, and no statistics are needed for that. Use olmocr-textual-v1 to compare two converters. Do not claim a 3-point win on six documents.

Detecting a 1-point difference would need roughly 23,000 independent documents. That is what the published benchmarks are for, and why this ships an importer for one rather than pretending 25 files can do the job.

Your output stays out of your project

Results go to ~/.local/state/doc-eval/runs/. DocEval refuses to write anywhere inside a git repository and tells you which one it found. Your working tree stays clean and nothing lands in a commit by accident.

Reading the score honestly

The numbers are the easy part. Four rules, each learned the hard way while building this:

  1. A bad first score is usually a broken measurement. Open one output file before blaming the converter. One slice here scored 33% while every output file was 0 bytes — no OCR, scanned pages. A missing capability, not a bad algorithm.
  2. n/a is not zero. Averaging "no table in this document" in as 0 punishes a fault that cannot exist.
  3. Quote the n. 0.99 over 2 of 25 documents is not 0.99.
  4. Compare like with like. A partial import of a public benchmark is not comparable to that benchmark's published figures. Say what you ran.

Documentation

Licence

Apache-2.0. See LICENSE and NOTICE.

No third-party test document is stored in this repository, or in the package — only source URLs, checksums and licences, with each third-party document's licence recorded individually in corpus_manifest/. synth-v1 has no such record and needs none: those 25 documents are written from declarations in this repository and are covered by the licence above, like the rest of it.

About

Deterministic evaluation for document parsing and semantic extraction: markdown fidelity (NED, TEDS, heading-tree F1, Spearman reading order) and extraction accuracy with explicit hallucination / over-abstention rates. Stdlib core, pytest plugin, three corpus tiers.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages