Check how well a document was converted — and read the score without fooling yourself.
You built something that turns PDFs, Word files or HTML into Markdown — or that pulls fields out of invoices and reports. How good is it? "It looks fine" does not survive a refactor, and a made-up percentage is worse than no number.
DocEval gives you real numbers and tells you what they do and do not mean — including which of them come from the research benchmarks and which are ours. Four are published metrics computed the published way; the assertion axis follows a published methodology; the rest are DocEval's own, and are marked as such. A house metric is not a worse metric, but it is not comparable with a paper's figures, and you should know which you are quoting.
- No model, no network, no API key. Every score is a plain calculation over your output and your reference. Same input, same score, on any machine.
- The core has no dependencies at all. Standard library only.
- It refuses to flatter you. A score that cannot be measured is reported as
n/a, never as zero. Averages print how many documents they cover. A blank output scores nothing instead of passing by default.
Did the conversion keep the document? Headings, tables, text, reading order, bold/italic/code.
Are the extracted fields right — and did the model invent anything? Plus the question most tools skip: did it refuse to answer something the document does answer?
Both work on your own documents. You do not need a benchmark corpus to start, and you do not need a reference document either — see below.
/plugin marketplace add CoreNovus/doc-eval
/plugin install doc-eval@doc-eval
That is the whole setup. Then just ask:
Is my converter losing the tables in these PDFs?
No pip install, no configuration. The plugin carries its own scorer, because the
core needs nothing but Python.
pip install "doc-eval[cli,corpus-gen] @ git+https://github.com/CoreNovus/doc-eval"from doc_eval import assert_a1_fidelity, assert_a2_accuracy
result = assert_a1_fidelity(my_markdown, reference_markdown)
result.require(min_heading_tree_f1=0.90, min_teds_struct=0.95)
risk = assert_a2_accuracy(extracted_fields, expected_fields)
risk.require(max_inventing_on_absent=0.0) # zero tolerance for invented valuesUse require() rather than assert result.teds_struct >= 0.95. A document with no
table has no table score, and comparing that to a number raises TypeError —
technically true, and no help at all. require() skips what it cannot measure and
says which metric it skipped.
Conversion
| Metric | In plain words | Source |
|---|---|---|
| Syntax rules | Is the Markdown even well-formed — table rows aligned, code fences closed, links complete | ours |
| NED | How much of the text survived, character by character | OmniDocBench |
| TEDS | How close the tables are, contents included | PubTabNet |
| TEDS-Struct | How close the tables are in shape, ignoring the cell text | PubTabNet |
| Heading tree F1 | Did the section hierarchy survive, levels included | ours |
| Reading order | Did the sections come out in the right order | OmniDocBench |
| Inline F1 | Did bold, italic, code and inline maths survive | ours |
| Content checks | Did the specific things you declared survive | olmOCR-Bench methodology |
Extraction
| Metric | In plain words | Source |
|---|---|---|
| ANLS | Text match that forgives small OCR slips | ST-VQA / DocVQA |
| ANLS* | List and object match that credits near-misses inside an item | ANLS* paper |
| Normalized-exact | Strict match after tidying case, spacing and date format | ours |
| Made things up | Answered a question the document does not answer | ours |
| Refused to answer | Left blank a question the document does answer | ours |
The last two are reported separately on purpose: their fixes pull in opposite directions, so one combined "risk score" would let one get worse while the other improved and show no change at all.
Writing reference Markdown for a real PDF takes hours, and two careful people produce different answers. So DocEval also scores against checks — small, unarguable statements about the document:
[
{"kind": "must_contain", "value": "383,285"},
{"kind": "must_not_contain", "value": "Page 3 of 12"},
{"kind": "heading_level", "level": 2, "value": "Revenue"},
{"kind": "order", "first": "Introduction", "then": "Conclusion"}
]"The total appears" and "the running footer does not" are cheap to write and impossible to argue with. This is how the published OCR benchmarks do it.
doc-eval corpora lists three kinds, and for each one says where its reference came
from — the property that decides how far to trust it:
| Size | Reference comes from | Good for | |
|---|---|---|---|
| synth-v1 | 25 files: 6 source documents, spread across up to 8 formats and 3 languages each (not every source document ships every format) | The same declaration that writes the file, so it cannot drift | Regression: did my change alter anything? |
| wild-v1 | 3 real documents | Hand-written, checked against the real bytes | Smoke-testing on genuinely messy input |
| olmocr-textual-v1 | 657 documents, 2,605 checks | Ships with the benchmark (ODC-BY-1.0) | Comparing two converters |
Every downloaded file is pinned by SHA-256, so a run today and a run next year measure the same bytes — or fail loudly.
No test document ships with this project. Not in the repository, not in the package. What ships is a manifest per corpus — a URL, a SHA-256 and a licence per document — and the CLI fetches or generates the bytes on demand into a cache outside your project. Recording a document's licence is a much smaller claim than redistributing it, which is why the corpora can name real published documents at all.
doc-eval corpora # what exists, and where each reference came from
doc-eval build-corpus synth-v1 # generates locally; needs the corpus-gen extra
doc-eval fetch-corpus wild-v1 # downloads, and verifies every SHA-256
doc-eval fetch-corpus olmocr-textual-v1 # the olmOCR-Bench importer
doc-eval verify-corpus synth-v1 # re-reads each generated file with another readersynth-v1 is the one that needs no network: it is written from declarations inside the
package, so the [cli,corpus-gen] install above is all it takes. The other two are
real published documents and have to be downloaded. verify-corpus applies to the
generated corpus only — it exists to catch a file that is on disk and unreadable, which
is a failure mode a corpus you did not write cannot have.
A digest that does not match is a hard failure that leaves nothing behind, never a warning: the bytes land in a temporary file and are moved into place only once they hash correctly. Reporting a mismatch and keeping the file would be the worst of both, since every later run would score against unverified input while believing it pinned.
Worth stating plainly, because "we scored higher" is a claim about statistics and these are the numbers behind it. Smallest difference each corpus can distinguish from noise, comparing two converters on the same documents (95% confidence, 80% power):
| Corpus | Independent documents | Smallest real difference |
|---|---|---|
| synth-v1 | 6 | ~63 percentage points |
| wild-v1 | 3 | ~89 pp |
| olmocr-textual-v1 | 657 | ~6 pp |
synth-v1 counts 6, not 25: the same six source documents rendered into eight
formats are not 25 independent observations, and treating them as such would roughly
halve the number above and be wrong.
So: use synth-v1 to catch regressions — it is deterministic, so any change at all
shows up, and no statistics are needed for that. Use olmocr-textual-v1 to compare
two converters. Do not claim a 3-point win on six documents.
Detecting a 1-point difference would need roughly 23,000 independent documents. That is what the published benchmarks are for, and why this ships an importer for one rather than pretending 25 files can do the job.
Results go to ~/.local/state/doc-eval/runs/. DocEval refuses to write anywhere
inside a git repository and tells you which one it found. Your working tree stays
clean and nothing lands in a commit by accident.
The numbers are the easy part. Four rules, each learned the hard way while building this:
- A bad first score is usually a broken measurement. Open one output file before blaming the converter. One slice here scored 33% while every output file was 0 bytes — no OCR, scanned pages. A missing capability, not a bad algorithm.
n/ais not zero. Averaging "no table in this document" in as 0 punishes a fault that cannot exist.- Quote the
n.0.99over 2 of 25 documents is not0.99. - Compare like with like. A partial import of a public benchmark is not comparable to that benchmark's published figures. Say what you ran.
- Writing checks for your own documents
- Every metric: exact definition and where it comes from
- Design notes and the defects that shaped them
- Contributing
Apache-2.0. See LICENSE and NOTICE.
No third-party test document is stored in this repository, or in the package — only
source URLs, checksums and licences, with each third-party document's licence recorded
individually in corpus_manifest/. synth-v1 has no such record and needs none: those
25 documents are written from declarations in this repository and are covered by the
licence above, like the rest of it.