Ground Truth Dataset Construction for OCR Evaluation
How to build reliable OCR benchmarks without breaking the budget.

Every OCR benchmark rests on a single fragile assumption: that the reference transcription paired with each scanned image is correct enough to tell a good output from a bad one. When that assumption fails, the benchmark doesn't measure the OCR system at all. It measures the artifacts of whoever built the reference set. Most claims about model accuracy in this field rest on the assumption that the reference set itself is sound, and that assumption rarely gets the scrutiny given to the models themselves.
The tension driving everything downstream is cost against coverage. Manual transcription is reliable, but it's slow and expensive at any real scale, and that pressure pushes practitioners toward shortcuts, some of which quietly invalidate the results they produce. This is a structural cost that resists optimization. It is the reason the field has split into three structurally different approaches to building ground truth, each one resting on a different idea of what "authoritative" even means. Manual transcription treats human reading as the final word. Semi-automatic alignment treats a prior digitization as authoritative, provided it lines up. Synthetic generation sidesteps the question of authority entirely by constructing documents where the answer is known in advance.
None of these pathways is a neutral technical choice. Each one shapes which document types, which error patterns, and which languages end up represented in a benchmark, and therefore shapes what a model trained or scored against that benchmark actually learns to handle well. The rest of this piece works through each pathway in turn, the metrics used to score against them, and where each one breaks.
How CER and WER translate transcription into a measurable signal
Before asking how ground truth gets built, ask what it gets measured against, because the metric shapes which construction method even makes sense. Character Error Rate and Word Error Rate are the dominant tools, and both turn transcription fidelity into a form of edit distance. That's a useful simplification, but it carries assumptions that quietly warp the evaluation whenever the reference itself is imperfect or, worse, treated as the only possible correct answer.
CER counts the substitutions, deletions, and insertions needed to turn an OCR output into the ground truth, divided by the number of characters in that reference, using Levenshtein distance as the underlying calculation. WER applies the identical logic at the word level. A single wrong character inside a word is enough to fail the entire word. On the same document, WER nearly always comes out harsher than CER.
The deeper issue is what happens when there's only one canonical reference to score against. Benchmarks such as OmniDocBench use continuous metrics, including BLEU and edit distance, and these metrics punish small, harmless differences in punctuation, spacing, and line breaks just as readily as they punish genuine misreadings.
A study of manual scene-text annotation against the Total-Text dataset put a number on how much of this matters: removing case normalization alone dropped CER from 0.223 to 0.013. By that study's accounting, 94% of what looked like reading error was actually a disagreement over annotation convention rather than a genuine transcription failure. That's a substantial gap. It means the overwhelming majority of the apparent signal in that dataset was noise generated by inconsistent labeling rules, not by the OCR system under test.
The implication carries forward regardless of which pathway produces the ground truth: shared transcription guidelines are a precondition built into the process from the start. They're the precondition for the metric to mean anything at all. And the critique goes further than case conventions. CER and WER reduce an entire document to a string-edit calculation. A single misread character and a collapsed newspaper column register as the same category of error. That equivalence becomes acute the moment you're evaluating historical newspapers, dense tables, or multi-column layouts, where the failure that matters most isn't a wrong letter at all.
Manual transcription: the reference standard that becomes the bottleneck
Manual transcription earns its status as the reference standard for one simple reason: its authority comes from a human reading the page, not from any assumption that a matching digital text happens to exist somewhere already. Standard practice uses stratified sampling to make sure the sample covers the range of page layouts in a collection, and domain experts then transcribe those pages against a shared set of guidelines to produce the reference.
Two recent examples show what this looks like when the material resists shortcuts. For Patrologia Graeca, a 30-page test set was randomly sampled and manually transcribed for its Greek text regions in 2025, and the result was blunt: existing public models for Greek couldn't get below a 5% average CER on it. The corpus combines complex typography, heterogeneous layouts, and degraded scan quality, and no automated pipeline could substitute for a human transcriber establishing what the baseline even looked like. In the ICDAR 2026 HIPE-OCRepair challenge, organizers introduced controlled image degradation at three separate noise levels and then had the development and test set ground truth manually re-corrected. Even in a pipeline that was otherwise automated end to end, manual correction remained the last word.
The bottleneck appears once the collection scales past a few hundred pages. For historical newspaper archives, legal records, or community-produced document collections running into the millions of pages, the per-page cost of expert transcription makes full manual ground truth construction a non-starter economically.
A second, quieter failure sits alongside the cost of manual transcription: saturation without coverage. OmniDocBench built its reputation on manual screening, layered annotation review, and expert plus large-model quality checks, and its top models are now saturating the leaderboard. A survey by Beyene and Dancy found that real-world document parsing, especially on historical and community archives, is nowhere near solved. The gap isn't random noise in the sampling. Leading benchmarks, OmniDocBench included, lean heavily on scientific papers, corporate forms, and current-generation digital PDFs, which leaves historical and community archives systematically underrepresented, including Black historical newspapers such as Freedom's Journal from 1827, The North Star from 1847, and the Chicago Defender from 1905.
The practical takeaway is that manual transcription remains essential in the right circumstances, even as it resists scaling. It's that manual transcription is non-negotiable for novel document types, degraded or marginalized archives, and any collection with no existing digital text to lean on, but it needs to be scoped through deliberate stratified sampling rather than applied wall to wall across a collection.
Semi-automatic alignment: when existing digital text can anchor the ground truth
Semi-automatic alignment takes a different bet: if a digital transcription of the same text already exists somewhere, why re-transcribe it by hand? The workflow pairs a scanned image with that pre-existing digital text, uses automatic layout detection to locate and match text regions, and brings in manual correction only where the alignment breaks down. Platforms such as Calfa Vision run exactly this combination of automatic layout detection and targeted manual correction. But the entire approach is conditional. It only works when a digital version of the same text is sitting in a library catalog or came out of an earlier digitization project.
The digital text has to be a genuine transcription of the same source, not a related edition, a translation, or a parallel text that merely resembles it. Get that pairing wrong and the resulting ground truth is quietly incorrect in a way that neither an automated check nor a casual human glance is likely to catch. Even when the pairing is correct, layout detection still makes mistakes: missed text blocks, merged columns, reading order scrambled by an unusual page design. Those errors propagate straight into the aligned ground truth unless the manual correction step is thorough enough to catch them, which brings the labor cost right back in through the side door.
The VERITAS pipeline, built by Bassanini and collaborators at the Università degli Studi di Milano, shows a more disciplined version of this approach. It treats archival document analysis as a sequence of independently checkable stages: transcription, layout analysis, and semantic enrichment. According to the paper Quid est VERITAS?, that structure delivered a 67.6% relative reduction in word error rate against a commercial OCR baseline. The design supports semi-automatic alignment well because it isolates exactly where a human needs to step in, rather than forcing a full re-transcription every time something looks off.
Semi-automatic alignment is strongest for well-cataloged collections with a clean prior digitization sitting behind them. It's weakest for degraded, multilingual, or heterogeneous historical corpora, which happen to be exactly the collections where trustworthy ground truth is needed most urgently.
Synthetic ground truth generation: exact labels without transcription, but on a constrained document population
Synthetic generation removes transcription from the process entirely. There's no ambiguity to resolve after the fact, because the ground truth was fixed at construction time.
Horn and Keuper, working at Offenburg University, applied this to table extraction: they pulled real tables out of arXiv papers to keep the content realistically complex and varied, then embedded those tables into synthetic PDFs to get exact LaTeX ground truth, combining realistic table content with zero transcription labor. Scaling produced the payoff. That framework made it possible to evaluate a large number of contemporary PDF parsers across documents containing hundreds of tables, at a level of reproducibility that manual annotation simply can't reach.
The boundary on this approach is hard. Synthetic generation only works for digitally native content. It has nothing to offer for handwritten documents, degraded scans, historical typography, or any material whose visual difficulty comes from physical wear and age rather than digital layout rules. SynthTabNet and similar large synthetic table datasets have pushed table structure recognition forward, but they work on cropped table images, not full pages, which makes them a poor fit for benchmarking end-to-end PDF parsing pipelines that have to find the table inside the page before they can even parse it.
A second limitation runs alongside the domain boundary: distribution. Synthetic PDFs don't reliably reproduce the font rendering quirks, spacing artifacts, scan noise, and layout irregularities that actually cause OCR systems to fail on real documents. A parser that scores well against synthetic ground truth can still underperform badly once it meets the document population it was actually built to process. None of this makes synthetic generation a lesser method. It's the right tool for its domain, provided the boundary is respected rather than assumed away.
How annotation formats structure the ground truth regardless of which pathway produces it
Whichever pathway produces the transcription, the ground truth still has to encode more than the text itself. Layout regions, reading order, and element types all need to be captured, because without them, evaluation collapses back down to character strings and loses the structural information that makes a document a document. The annotation format is what makes that richer encoding possible.
PAGE XML has become a widely used answer to this problem. OmniDocBench's schema shows how far this expectation has moved. It localizes block-level elements like paragraphs, headings, and tables, alongside span-level elements like text lines, inline formulas, and subscripts, adds reading order annotations, and tags attributes at both the page and block level. Recognition output gets expressed differently depending on what it is: plain text for prose, LaTeX for formulas, and both LaTeX and HTML for tables.
That last detail isn't cosmetic. It determines which metric even applies. Table extraction needs TEDS or a grid-based comparison. Formula extraction needs CDM. Plain text blocks are better served by normalized edit distance, BLEU, or METEOR. A ground truth format that flattens everything into one uniform string can't support any of that element-specific evaluation, no matter how carefully the transcription itself was done.
Some mismatches, though, live in evaluation logic rather than in the annotation format itself. OmniDocBench v1.6 introduced Multi-Granularity Adaptive Matching to fix a bias that occurred whenever a parser's output segmentation didn't match the reference's segmentation. Rather than rewriting the ground truth, MGAM leaves it untouched and searches the prediction side for the best matching granularity, which removes a distortion the annotation format alone had no way to solve. For archival material specifically, the VERITAS pipeline's separation of layout analysis, transcription, and structural annotation into independently checkable stages shows how a format can be built to hold the genuine heterogeneity of historical sources instead of flattening it into something more convenient but less true.
Where all three pathways break down: contamination, saturation, and the disagreement-as-signal problem
Contamination erodes the independence of the reference. Saturation decouples leaderboard scores from what a system can actually do in the field. And collapsing annotator disagreement into a single gold label throws away information the ground truth ought to be preserving.
Standard benchmarks are frequently present in LLM pretraining corpora, and prior detection methods using n-gram overlap, membership inference attacks, and surprisal-based probes are limited, especially for historical material. LLM-as-judge evaluation adds a second vector for the same problem: models trained on synthetic data built from architecturally similar foundations end up favored unfairly by judges built the same way. The response taking shape is to build entirely new datasets from sources that were never digitized before, and to keep them offline through every known pretraining window, a direct consequence of just how much web text modern models ingest.
Saturation appears differently but points at the same underlying weakness. OmniDocBench's top models now clear 94% accuracy on its leaderboard, and yet the Beyene and Dancy survey from 2026 documents that Black historical newspapers and other community-produced historical documents are structurally missing from both OCR training data and evaluation benchmarks. That gap between leaderboard performance and what a system can actually do on real archival material has become a central concern. It's a structural critique of what static ground truth datasets can tell anyone. ICDAR 2026 HIPE-OCRepair ran into the same wall from the opposite direction: OCR on the original images turned out too accurate to make a meaningful post-correction challenge, so organizers had to introduce controlled degradation at three noise levels, using both a matched strategy (identical pages across levels) and an unmatched one (different pages at each level, to keep information from leaking across the split). Difficulty had to be engineered on purpose, because the dataset wasn't hard enough on its own.
The third failure is quieter and easier to overlook. Standard practice collapses several annotators' transcriptions into one gold label, and in ambiguous or high-variability domains, that disagreement carries information the averaging discards. It's genuine uncertainty, task subjectivity, or a sign the guidelines weren't specific enough. Treating it as noise instead of signal hands the field a reference that's more confident than it has any right to be.
Annotation-free and unit-test evaluation as responses to static ground truth's limits
Two genuinely different alternatives have started to emerge in response, and both matter precisely because they don't try to build a better static reference. They change what evaluation is grounded in altogether. One is annotation-free evaluation, which uses consensus across multiple large multimodal models to approximate a ranking without any reference labels at all. The other is unit-test-driven evaluation, which swaps reference strings for programmatic pass/fail checks. Both are a signal that manual ground truth is no longer treated as the only defensible foundation for evaluation.
DocOCR-Eval, published in 2026 by Xu and colleagues, was built directly in response to the cost of manual annotation. Manual transcription, semi-automatic alignment, and synthetic generation still matter enormously for training data, for fine-grained diagnostic work, and for any collection where a system's failure mode needs to be pinned down precisely rather than approximated. What these newer methods offer instead is a way to keep evaluating at scale and cadence once a fixed reference set has been mined for everything it can teach, without pretending that a saturated leaderboard is still telling anyone something new.
Sources
- Benchmarking PDF Parsers on Table Extraction with LLM-based Semantic Evaluation
- ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents
- GitHub - opendatalab/OmniDocBench: [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation · GitHub
- DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
- A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
- Quid est VERITAS? A Modular Framework for Archival Document Analysis


