Layout Parsing Benchmark Datasets and Their Coverage Gaps
Academic layout datasets sacrifice breadth for depth or scale.

Layout parsing, which divides a document page into labeled zones like paragraphs, charts, images, and titles, has shifted from rule-driven methods to supervised learning. That change gives the data underlying the models a dual role: they teach the system how documents appear, while also setting the standard for success after training ends. Training benchmarks and evaluation benchmarks are seldom designed with matching priorities, so when a single dataset fills both roles, each annotation decision, page count, category count, document variety, reflects what the field considers important. The forces that push a dataset toward training-ready scale usually shrink its scope and simplify its categories, whereas the forces favoring detailed, elaborate annotation usually yield collections lacking the scale needed for broad applicability. A benchmark's missing coverage is less an oversight to fix than the price of the design decision that brought the dataset into being, and it should be judged by its intended purpose before receiving credit or blame for its omissions.
How PubLayNet and DocBank achieved scale by narrowing scope
Both collections that enabled layout supervision at scale achieved their size by narrowing the domain until annotation could be automated. PubLayNet offers over 360,000 document images to detect layouts, a volume that let teams build models far beyond what manual labeling could afford on a similar budget. Since PubLayNet draws exclusively from scholarly papers, the resulting detectors often struggle in real-world scenarios involving invoices, forms, or government filings that look nothing like academic layouts. Li and colleagues from Beihang University alongside Microsoft unveiled DocBank at COLING 2020, matching PubLayNet's volume by relying on the same machine-generated labels. Yet because DocBank's automated pipeline lacks character-level annotations for evaluating retrieval accuracy, it can only teach region placement rather than verify whether the extracted text is correct. Weak supervision let PubLayNet grow alongside DocBank by turning LaTeX sources and PDF metadata into annotations automatically. This approach brought six-figure corpora within budget, while quietly defining a “document” as a polished, single-language research article from one domain, the only setting in which LaTeX source and tidy metadata can usually be harvested. Both corpora remain valuable resources. At early-2020s labeling prices, constraining the domain and generating labels automatically was about the only way to build a corpus large enough for training. Problems emerged only after researchers let academic-paper layouts represent document layouts overall.
How DocLayNet balanced label richness against manual annotation economics
DocLayNet, published by Pfitzmann and colleagues in 2022, went back to manual annotation while still reaching for scale, so it could hold both ends of the tradeoff together. It remains the largest benchmark for PDF layout segmentation built on human-labeled pages, covering 11 layout element types across 80,863 manually annotated pages, a page count that no prior hand-labeled dataset had matched and a label set meaningfully richer than PubLayNet's. That combination is why you still see DocLayNet used widely as a baseline years after release. But the dataset predates the arrival of multimodal, vision-language models as the dominant architecture for document understanding, and it was built before those models exposed how much a spatial-overlap score can miss about logical structure. Its eleven element types may seem extensive beside earlier efforts, but they remain far short of later benchmark baselines, which expected 28 categories for blocks and 4 for spans, while DocLayNet omits reading order, leaves page-element logic unmodeled, and has no account of cross-page structure. DocLayNet’s label set stops before the capabilities extraction systems require, such as semantic figure and table types, formula recognition, and column-aware reading order; that limitation is better understood as a boundary than a flaw. It defines how far manual labeling could realistically go at that scale, and later dataset efforts made that shortfall their target, trading scale for richer annotation.
The depth-for-size tradeoff in M⁶Doc, D⁴LA, and OmniDocBench
Collections assembled to address that shortfall deliberately prioritized depth over scale. Although M⁶Doc alongside D⁴LA surpass PubLayNet and DocBank in annotation richness and domain breadth, neither supports training a robust model from scratch or resolves the multilingual gap pervading the field. Introduced at CVPR 2025 and brought to v1.7 by April 2026, OmniDocBench stands as the boldest effort so far to expand both annotation categories and document coverage simultaneously. Spanning 1,651 PDF pages over 10 document categories such as academic papers, financial reports, newspapers, textbooks, plus handwritten notes, the corpus covers 5 layout and 5 language varieties, offering 28 block-level annotation classes alongside 4 span-level ones, with sequence and attribute labels recorded per page and per block. Such fine-grained annotation surpasses what PubLayNet, DocBank, and DocLayNet ever provided. That still leaves the problem unsolved. Although it includes 5 language types, OmniDocBench still offers limited coverage, drawing most pages from relatively tidy scans, a constraint that matters more than it may seem once deployment conditions are involved. The dilemma is now out in the open: systems built with DocBank alongside PubLayNet use incompatible labels for a fair OmniDocBench test, while training on a corpus of that size cannot match the generalization made possible by PubLayNet's 360,000 images. OmniDocBench makes that compromise sharper. It leaves the tradeoff in place.
The scale-depth tradeoff compounded by multilingual coverage and physical-world conditions
Even the richest benchmarks quietly assume that documents are written mainly in one script and originate as clean digital files. As a result, many real-life document types are left out. The multilingual side of the problem is clear and measurable. Today’s datasets are still concentrated around English plus a few well-supported languages, leaving low-resource languages behind, especially the many-script document traditions of the Indian subcontinent. Still, there is real progress: after tuning English-trained systems with IndicDLP, accuracy rises sharply for Indic documents, while systems built with IndicDLP also transfer effectively to non-Indic layouts. The multilingual gap is therefore solvable if the data is chosen well, but getting there calls for dedicated datasets larger than any multilingual benchmark now available. A second gap falls along a different axis: the condition of physical capture. Benchmarks built only from pristine PDFs fail to capture how documents appear after leaving digital workflows, and models excelling on standard academic tests frequently falter when processing real-world institutional files. Real5-OmniDocBench, unveiled during the ECCV 2026 conference by Baidu Inc. is the first to physically reconstruct the complete OmniDocBench v1.5 test set at full scale and one-to-one correspondence, rebuilding it across five real-world scenarios: scanning, warping, screen-capture, lighting, and skew, which allows researchers to attribute performance loss to specific physical factors. Even under these distortions, small distortion-aware models kept their accuracy above a high level, but general-purpose pipelines dropped sharply. This confirms that the divide between test scores and real-world results stems from how datasets get built for training and evaluation, not from an accidental byproduct. Handling multiple languages and surviving physical capture are distinct challenges, and conflating them obscures the specific solutions each demands. What links them is a root issue: each demands labeling investment the field has deprioritized, since such investment vies directly with identical resources already strained by competing demands of scale and depth.
How conventional metrics hide logical-structure failures
You can have a dataset with plenty of scale, rich labels, and broad domain coverage, and it will still steer you wrong when the evaluation method merely verifies box overlap. Both IoU and mAP only capture spatial alignment, checking if a detected box occupies the same area as its reference. Such metrics ignore logical coherence, letting a model achieve top marks while producing results that downstream extraction cannot actually use. Modern models constantly produce structural mistakes like incorrect merges, improper splits, or total omissions, yet overlap-based metrics never catch them. Dividing one region into a pair of distinct boxes may yield strong IoU results, yet such fragmentation disrupts reading order and destroys the extracted text. Published during 2025 and 2026, the LED benchmark evaluated structural reasoning within layout predictions firsthand rather than depending on surface accuracy, revealing a stark divide in model performance on this task. While the GPT-4o family handled basic two-class fault spotting capably, it struggled badly with detailed layout logic, and both DeepSeek alongside LLaMA 4 lagged at every size, suggesting that design choices matter more than parameter count. Reading order exposes how superficial those measurements can be. Most models do best on single-column pages, while multi-column formats sharply increase mistakes and text turned sideways, especially at 90 degrees, reveals orientation blind spots that mAP does not measure. In Parser-Oriented Structural Refinement, Liu's Unisound and Chinese Academy of Sciences team showed that largely accurate text recognition can still yield broken Markdown or JSON when layout retention or sequencing goes wrong, a problem outside page-level mAP scoring. In 2026, Dr.DocBench measured the same kind of breakdown in multi-page documents: as pages pile up, nearly every model loses reading-order accuracy, since it must arrange many more blocks and distinguish within-page material from text carried over the page break, which single-page tests miss. Taken together, these findings show that a model may score well on OmniDocBench for normalized edit distance as well as mAP, yet still break down severely on logical structure across multiple pages, so topping out on the usual metrics is not proof that the core problem is solved.
Restructuring evaluation around error taxonomy and end-to-end tasks
Faced with both kinds of gap, the field has stopped treating a benchmark as a single-score exercise and started using it as a way to identify which failure occurred and at what point. Released in May 2026, PureDocBench stretches across 4,425 samples and 66 document types that range from clean to degraded to real-world settings, and it exceeds all previous benchmarks in sample total and variety of types while anchoring itself in a structured error taxonomy. A quiet 2026 consensus now tests OmniDocBench v1.6 and Wild-OmniDocBench alongside PureDocBench, a tacit recognition that any single benchmark falls short of full coverage. READoc, an ACL Findings 2025 paper, charts a fresh course toward a closely related aim, recasting layout-sensitive work as whole-pipeline transformation of lengthy PDFs into organized Markdown, then auto-generating numerous paired documents so that any layout mistake is punished through its consequences when the extracted result is checked against the bounding box. PureDocBench and READoc converge on the same underlying difficulty through two separate lines of attack. PureDocBench identifies the specific category of mistake a system made. READoc determines if the mistake had any real impact on your desired result. By earning a spot at ECCV 2026, Real5-OmniDocBench signals that purely digital assessment falls short and establishes attributing degradation by factor as the bar future evaluations must clear. Yet the earlier tension between scale and depth remains unresolved. Such instruments merely expose the shortfall in training data more clearly, intensifying demands for a remedy.
The scale-depth tradeoff in practice for practitioners choosing or interpreting a benchmark
Because no current benchmark combines very large scale with deep labeling, teams should choose one for the deployment their model must handle, rather than just choosing whichever dataset is cited most. For scientific-publishing layout work, a team can often draw on PubLayNet’s or DocBank’s breadth and get solid performance, since its target documents closely resemble the training data. For mixed-document deployments covering items like forms, financial statements, and handwriting, OmniDocBench-style label richness is the better fit, but with only 1,651 pages, the team should budget for extra training material to shoulder a production model’s full training load. For teams handling Indic scripts or similarly underrepresented languages, purpose-built resources such as IndicDLP are the right starting point, since studies indicate that focused fine-tuning bridges the transfer shortfall where broad training corpora fail. Teams working with degraded document images should benchmark specifically on a dataset like Real5-OmniDocBench, since high scores on pristine PDFs reveal little about robustness to artifacts like scan noise, geometric distortion, camera capture, lighting variation, or tilt. Anyone looking at a reported score should ask what failure mode it was designed to detect, since strong mAP or near-perfect edit-distance results can sit alongside broken reading order across pages, structural mistakes that spatial-overlap measures miss, or a label scheme too coarse for the downstream extraction task. Treating these scores as solutions to a broader problem than they were meant to address is the source of most confusion in the field over what counts as "state of the art" in document parsing.


