Character Error Rate vs Word Error Rate for Document Benchmarking
Choose the right metric—CER penalizes character mistakes while WER penalizes whole words.

Word Error Rate and Character Error Rate share a single piece of arithmetic: count the swaps, drops, and extra symbols needed to make a system's output match the reference, then divide by how long that reference is. What separates them is the unit of counting. CER scales that count against the reference's character total, working beneath the word level entirely, with no sense of where one word stops and another starts. WER runs the identical formula over whole words, and by that rule one mistyped character is enough to mark the whole word wrong, however close the system got.
Using distinct units yields two interpretations of one document that may differ greatly. Because of this asymmetry, WER usually matches or exceeds CER when evaluating identical content, even if the underlying formulas do not ensure that outcome in every instance. As Docsumo's OCR accuracy guide likewise notes, WER typically penalizes the same passage more severely since it cannot award partial credit.
This is not a small footnote. The same page may show a low CER while its WER is several times larger, without implying a faulty metric or messy data. It follows from using two counting bases. When a benchmark publishes just CER or WER, or blurs the distinction between them, it quietly decides which errors count most. The remaining sections spell out that tradeoff.
How the CER-to-WER gap scales with document characteristics
CER and WER do not maintain a constant distance from each other. This gap expands or contracts depending on particular, recognizable qualities of the evaluated text. Because of that, comparing these two values yields a helpful diagnostic clue instead of mere noise to be smoothed out.
The length of words drives the largest effect. On a page with typical word lengths, even a modest CER inflates into a much higher WER, driven purely by the text rather than any shift in how well the system actually performs. Studies of languages with rich morphology reveal the same skew: brief tokens suffer harsher WER penalties than extended ones, while results vary unpredictably by language and genre, since segmentation rules alter how accurate a system seems without touching its actual output.
Error distribution creates another separate effect. Studies of historical OCR demonstrate the pattern: errors are nearly uniform throughout some volumes but clustered around particular words in others, a skew that raises WER far more sharply than CER.
Under severe degradation, the scaling can fail completely. When the recognizer hallucinates content absent from the original, these spurious additions and omissions compound until WER surpasses what should be its upper bound. Once a measure breaks past its own ceiling in certain situations, it stops functioning as a limited ratio and cannot reliably contrast texts of varying quality.
What a metric that ignores structure misses
CER and WER evaluate only the sequential stream of extracted characters or words. Neither metric recognizes tabular structure, columns, or fields. A document can therefore score well even when its extraction has the wrong structure.
A model may score extremely well by CER on the test data yet still misunderstand the page structure, for example by splitting a spanned cell into separate rows or rendering the columns of an invoice as one plain text sequence. Those mistakes are invisible to CER and WER, which score only the output text sequence rather than its source location or page arrangement. Newer OCR benchmarks have surfaced errors of the same sort, including omitted strikethrough passages, broken handling of two-page tables, and scrambled multi-column reading order, with serious consequences for retrieval despite almost flawless character recognition.
In business documents, structure often matters above all else. An invoice, bank statement, or structured intake form is only as useful as the accuracy of its individual entries, and a single mistake in that sequence undermines the entire file rather than balancing out across the text. Docsumo advises tracking per-field precision alongside per-file correctness metrics, which respectively measure how many extracted entries match ground truth exactly and what proportion of files have every entry right, to supplement CER plus WER rather than replace them. An acceptable CER in isolation becomes unacceptable once it affects a policy number, dollar figure, or medication dosage, since this aggregate metric reveals nothing about where errors occur within the document.
Because CER ignores how language is built and whether meaning holds, text can show a low CER and still read like nonsense. WER improves on that point by checking words as units, so a meaning-changing substitution is flagged once its spelling no longer matches the reference. Still, WER’s fix for this problem is fairly crude. It gives identical weight to "the" and to a decisive account number.
When CER is the more honest metric
CER is the more trustworthy of the two numbers for specific document types and languages, and those conditions are easy to list.
Logographic scripts and tongues with complex word-formation systems make this clearest. WER presupposes a firm, shared boundary between words, but that foundation crumbles when symbols represent entire lexical units or meaning-bearing chunks rather than individual characters, or when dividing text into words becomes inherently uncertain. A recent NAACL paper from 2025 extends this argument, contending that CER sidesteps numerous structural flaws inherent to WER, remains steadier across diverse scripts, and aligns more tightly with how people judge transcription quality, even for English.
Old or damaged records illustrate the same point differently. When pages are torn, scanning jumbles line order, or sentence marks and letter case are unclear, WER can swing wildly, while CER still measures the share of source content that came through. In tests on Leopardi historical Italian handwriting, Claude Sonnet 3.5 led the evaluated large language models, yet its CER remained above one-fifth of all characters. By itself, the figure fixed the practical lower bound across the whole benchmark, while WER added little because the handwriting was so irregular.
CER is especially useful for making close distinctions among rival systems tested on comparable data, because the gaps between models can be too subtle for word-level scores to catch. Using the benchmark IAM for offline handwriting, Crosilla and co-authors use CER to pinpoint the residual mistakes, because tiny character-by-character gaps between rival systems can disappear in word-based evaluation.
When WER is the more useful signal
CER's advantages don't mean it should be the automatic choice in every setting. WER has a domain where it belongs, and researchers' near-universal preference for it when working with alphabetic languages shows a genuine match to real-world usage rather than mere inertia.
Where word boundaries are stable and clear, WER directly reflects what human editors or downstream parsers end up doing. They examine each word, fix any mistakes by typing anew, making that token the basic measure of editing work. Such practical alignment is why WER stays the standard for evaluating OCR or spoken-language systems in alphabetic scripts: reviewing Interspeech publications from 2023 to 2025 showed 86.6% of eligible studies used it for scoring. The convention tracks something real. When the goal is identifying correct tokens, such as spotting named entities, searching for keywords, or extracting data from English form fields, WER tracks real-world outcomes better than CER since those later stages process entire words instead of individual characters.
By demanding that the entire word be correct, WER offers a coarse yet real guard against meaning drift beyond CER’s reach. A word can match at the character level yet carry the wrong meaning, an error WER flags while CER lets slip.
Benchmark results and the gap in practice
Results in many document types back the ideas described above with real consistency, and, as document complexity rises, the pattern sharpens.
When evaluating contemporary script, IAM and RIMES reveal the smallest performance difference. According to Crosilla et al. distinct networks tested on the IAM and RIMES datasets each produced CER-to-WER proportions near 1:2, aligning with what average character counts per word would suggest for legible text where mistakes are scarce. In such scenarios, large language models surpass the established handwriting recognition model from Transkribus, yet this advantage falters with historical documents, where CER best exposes the gap by revealing each system's lower limit as quality declines.
Slightly more challenging are the typewritten Leiden University archives, produced from 1983 through 1985 to catalogue professors' and curators' biographical details reaching as far back as the year 1575. Despite such precise character recognition, converting the OCR results into JSON yielded only moderate success. Such results highlight the previously noted discrepancy between accurate letter identification and dependable structured data.
Hard-to-read historical scripts stretch the divide further still. On 18th-century Russian Civil documents, TrOCR after fine-tuning kept CER very low while WER ran several times higher, and Surya erred much more on both measures. PyLaia from Transkribus, however, kept CER and WER much nearer to each other on this corpus, pointing to mistakes clustered in a few long or structurally odd words instead of spread uniformly over the page, the distribution effect described earlier.
At the hardest end of the range are old scripts with few resources, whether Arabic, Cyrillic, or Latin. Recent work on historical documents in many languages shows that large multimodal language models reach CERs that differ a great deal depending on which script they are reading. WER would largely hide those differences, because these three writing systems each define the word in their own way, and that is exactly the situation where CER was called the more honest metric earlier in this piece.
Why current benchmarks are under pressure from LLM-era OCR
Document-recognition LLMs are now exposing assumptions in CER and WER that those measures were never designed to cover. These metrics assume a system will copy the page’s text in sequence, following its reading order and reflecting where that text sits on the original page. But multimodal LLMs often take a different path. They may move material around, condense it instead of copying it verbatim, or present a table differently without losing the information inside it. Those choices can be penalized heavily by scores that demand exact character and word alignment, although the output may serve a downstream system better than verbatim transcription.
The benchmark results already reveal this pressure. Systems excelling with present-day English handwriting, like some tested against IAM along with RIMES, tend to falter unpredictably when facing historical texts or scripts outside English, which suggests their training background explains this more than any flaw in CER and WER metrics. When a system rearranges structure freely rather than maintaining it exactly, edit distance evaluation can unfairly punish the adaptable comprehension that draws people to LLM-based OCR for challenging practical texts.
Purpose-built extraction platforms sidestep this problem entirely, since their architecture natively supports evaluating both individual fields and whole documents rather than shoehorning outputs into line-by-line transcription scores. Phantom Farm's work on a document follows that same logic, focusing on field-level verification instead of raw transcription fidelity. Therefore, anyone comparing an insurer's OCR tool alongside a generic LLM API or specialized extraction service should ask which option places the correct data into its proper location on the forms that truly count, rather than which achieves the lowest standalone CER and WER scores. As a starting point for raw recognition performance, CER and WER still serve their purpose. Such measures fall short once the objective shifts to pulling a particular value out of a designated field on a given document, meaning evaluation frameworks must develop structural criteria that match current model capabilities.
Sources
- Benchmarking Large Language Models for Handwritten Text Recognition
- Enriching Historical Records: An OCR and AI-Driven Approach for Database Integration
- Deciphering the Underserved: Benchmarking LLM OCR for Low-Resource Scripts
- Advocating Character Error Rate for Multilingual ASR ...
- Advocating Character Error Rate for Multilingual ASR Evaluation
- When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation


