LLM-as-Judge Reliability for Document Extraction Evaluation
Consistency in LLM judges doesn't guarantee they measure what actually matters.

Document extraction teams have widely embraced LLM-as-Judge practices, yet a straightforward finding is forcing them to reconsider that approach: perfect consistency in scoring does not guarantee a judge is actually measuring what matters. A judge's consistency and its accuracy are distinct traits, mathematically independent of each other, so improving one never ensures the other. The stakes are highest in the exact settings where judges are used most heavily: lengthy, structured outputs with many fields, little ground truth, and errors that can be indistinguishable from correct values.
What document extraction teams are buying when they reach for LLM judges
Using LLM-as-Judge for document extraction solves a concrete engineering problem. Such pipelines fail at one of two stages: fetching irrelevant material, or the generator overlooking what was correctly retrieved. Teams usually choose from three evaluation approaches based on their available resources. The options are absolute rating for a standalone metric, head-to-head ranking to compare outputs, and binary verification when labeled answers exist. Selecting the appropriate method depends on having ground truth and needing standalone metrics versus comparative ones.
What made this shift from handy option to standard practice was the money. A sector handbook from 2026 reports that judges built on frontier models cost only a tiny sliver of what human reviewers charge, a margin big enough for judge-driven evaluation to scale to the throughput real production extraction pipelines demand. What comes back is more than a number: most judges pair their verdict with an explanation, so extraction teams can see at a glance which fields break, which documents misbehave, and where along a retrieval chain things went wrong. Teams value this approach largely because its readable output turns a score into something diagnostic. That point is not oversold. This piece asks what a steady judge score really proves for teams, and whether steadiness was the right target all along.
The metric teams use to validate judges (and why it misleads them)
When teams audit a judge, they almost always default to raw agreement, the fraction of cases in which the judge's call lines up with a reference call or agrees across runs. That number carries a familiar weakness: it ignores the agreement chance alone would produce, and that weakness is not minor.
The most thorough systematic probe of this issue to date comes from a UC Berkeley trio made up of Hughes, Norman, and Rivera, who ran judges from nine providers through MT-Bench, RewardBench, and JudgeBench, collecting about 541,000 separate judgments. Raw agreement ran tens of percentage points above kappa, Cohen's chance-corrected measure of discrimination, and the pattern was consistent for all models, whatever the provider, the size, or the generation. The gap separating raw agreement from kappa is big enough that one batch of judgments yields wildly divergent accuracy numbers, depending purely on which metric a research group chooses to report.
The problem gets harder when the first yardsticks practitioners reach for are MT-Bench alongside RewardBench and JudgeBench, benchmarks that put raw agreement in the spotlight. Teams that vet these benchmarks during diligence carry forward the inflation already built into them. In document extraction work, the issue is not just a theoretical measurement dispute. Across legal, finance, healthcare, and online retail, teams that put a judge in the gatekeeper role, sorting successful extractions from those sent to review, often rely on a score that looks much more discriminating than it is.
A consistent judge can still be wrong about what it measures
The deeper problem appears once you separate two questions that sound similar but aren't: does the judge give the same answer every time, and does the judge's answer actually track what it's supposed to be measuring. Reliability is whether the judge gives the same answer every time, and validity is whether the judge's answer tracks what it's supposed to be measuring. They can diverge sharply, and the Norman, Rivera, and Hughes data shows how.
In two live judges, the researchers measured reliability over 0.95 across reruns, meaning their scores barely changed, plus position bias over 0.10, meaning placement in a pairwise comparison could strongly change the outcome. They describe the zone as "consistent but biased"; its clearest example was a compact Qwen 3 model, scoring 0.992 for test-retest reliability, 0.192 position bias, and only 0.289 on JudgeBench kappa. Put simply, the judge repeats itself with near-perfect steadiness, often switches its pick when the two answers trade places, and only weakly distinguishes truly strong outputs from weak ones under chance-adjusted scoring. Gemini 2.5 Flash follows that profile less sharply, with 0.988 test-retest reliability and 0.125 position bias.
Position bias and within-judge agreement are orthogonal properties. Neither constrains the other, so a judge can be maximally consistent and maximally biased at the same time, and consistency gives no information at all about bias. For document extraction work, where a judge is frequently asked to compare a reference extraction against a candidate extraction field by field, this has a direct consequence: the order in which fields happen to appear in the prompt, not the accuracy of the fields themselves, can decide the verdict.
The same dataset also contains one real bit of good news. The tendency to give extra credit for length, regardless of accuracy, proved more modest than researchers had expected. For the 21-model pool scored under the same head-to-head rubric, the study put verbosity bias below 0.011, far lower than earlier reports, which points to real progress on that failure mode even while position bias stays stubborn. The results for Qwen and Gemini do not mark them as broadly poor judges; instead, they reveal reproducibility as something most validation checklists measure, yet unable to tell whether a judge’s decision rests on the right cause.
How judge rankings fall apart when the benchmark changes
There's a second reliability problem that compounds the original issue, and it isn't about any individual judge but about how they're ranked in comparison. A judge who appears near the top of the pack on one benchmark can come across as middling on a different one, and the swing between them doesn't come down to fine-tuning at the edges.
This happens partly because each benchmark makes judge-to-judge gaps visible at a different level of detail. JudgeBench draws much finer distinctions among judges than MT-Bench: switching benchmarks not only reshuffles the leaderboard, but also alters how much genuine distance among judges is visible at all.
For teams building extractors, picking a judge for its strong general-benchmark score still does not ensure comparable behavior on the targeted extraction job; calibration is not automatically portable between domains. In addition, when pairwise scoring pits a candidate result against one fixed anchor, the anchor’s idiosyncrasies become an unseen evaluation variable, steering which answers appear stronger or weaker.
Reliability when the output being judged is long and structured
Document extraction output, with its extended length, multiple fields, and structural complexity, is precisely the case that puts the heaviest strain on LLM judges, standing in stark relief to the brief open-ended answers such benchmarks typically feature. Across 32 tested combinations of models and settings, a minority alone achieved correlations exceeding 0.60 against ground truth, while the overall mean landed just above the random pairwise guessing baseline. Reliable judgment of long-form output remains an open issue: the structured outputs that define document extraction pose the toughest challenge in LLM evaluation, with accuracy barely exceeding chance, so the gaps recorded in briefer evaluations widen rather than shrink precisely where extraction teams rely on these tools.
The structural problem is clear. When an extraction spans several fields, the evaluator must verify numerous distinct facts simultaneously, allowing a single error to slip past amid adjacent correct entries. An evaluator focused on general readability instead of verifying every item separately might overlook a fabricated detail while still issuing a confident high score. This problem is compounded by the observation of (2026)](https://arxiv.org/pdf/2606.19544) that assessors scoring extracted fields have no baseline facts for comparison and depend entirely on prompt instructions. Consequently, there is no way to flag a convincing yet incorrect figure.
A live deployment shows what the repair changes in practice, not just in theory. A Frontiers study of sahibinden query parsing showed that using a category-specific prompt for each extraction type instead of a single general judge for all outputs brought results closer to human assessments while avoiding any fine-tuning expense. The lesson extends beyond this deployment: extraction tasks that cover multiple schemas need more tailored evaluation than one generic judge prompt can provide.
Reliability and validity are independent problems
The empirical failures documented, kappa deflation, quadrant that's consistent but biased, benchmark fragility, and long-form degradation all show one underlying structural truth, proven formally rather than just seen.
In a 2026 preprint, Chen and colleagues formalize what makes a judge valid, and they do it with two separate coordinates. One of these, invariance, they label S: the chance that a judge's ruling holds steady when an output gets an edit that leaves its underlying quality alone. The other, construct sensitivity, gets the label R: the chance that a verdict moves when an edit genuinely alters the output's underlying quality. To be good, a judge has to score high on both: it must shrug off edits that don't matter and react to those that do. Chen's team also proves, rather than just noticing, the mathematical independence of S and R: scoring well on one in no way promises scoring well on the other.
A thermometer makes this easy to picture. If you lock the gauge at one number, it ignores every shift in real heat. Such a device clears all reliability checks by staying absolutely fixed, yet reveals zero about actual heat. So a judge can keep S close to its maximum even as its power to sense what it claims to track sits near zero.
The empirical numbers show this directly. With invariance held at 0.90 or above, the study's judges reached just R=0.319 in construct sensitivity. In the document-extraction setting, a judge may label each extraction it views as "good," every time, regardless of where it falls in a comparison and regardless of response length, yet remain incapable at a structural level of distinguishing an extraction whose fields are all traceable to the source text from one containing fabricated field values in plain view.
This result overturns the idea behind most mitigation work so far, which assumes that eliminating biases tied to verbosity and position solves the problem. Removing such biases satisfies a necessary requirement for sound evaluation, yet that alone cannot guarantee it. Instead of assuming it emerges automatically after eliminating familiar distortions, construct sensitivity demands direct measurement and deliberate design as a distinct challenge.
When structured, schema-constrained evaluation reduces the validity gap
One real, well-documented case breaks this rule, and its value is telling teams where their judges can be trusted more, not less. When a rigid yes/no schema is the yardstick for pulling fields out of documents, the task parts ways with open-ended grading: each field only has to be checked for presence and for falling within an approved range of values.
Research published in 2026 on assessing regulatory compliance revealed no favorable bias toward a model's own outputs when judges used a structured checklist, overturning the field's standard assumption that systems tend to prefer responses resembling their typical style. While open-ended assessments consistently reveal this tendency, switching to binary scoring might eliminate it entirely. Checklists probably work by reframing evaluation from a subjective judgment of quality into verifying whether each mandatory field exists and its value meets the specified constraints. Such a reframing eliminates the very fluency-driven mechanism responsible for both verbosity bias and self-preference.
A pipeline for extracting legal judgments, set out in the 2026 paper arXiv 2607.03325, ran DeepSeek V3-0324, producing XML output; citations were verified against the source judgment's text, and any reference present in the result yet absent from the source was flagged as a potential hallucination to be eliminated. The schema defined the required fields and their allowed values, so what the judge aimed to gauge was baked directly into the system's working rules instead of being described in flowery prompt language. In separate work on 5W1H news extraction, a dual structured prompting design drew on Function Calling plus Pydantic schema validation, forcing deterministic, JSON-formatted output and converting open-ended interpretive calls into quantitative data that varies far less in scoring.
Applying a schema does not by itself resolve questions of validity. It shuts down the particular route by which smooth phrasing biases evaluations toward acceptance. Chen's team defined construct sensitivity as the R coordinate, and it still requires separate validation. While structural rules prevent evaluators from mistaking polish for accuracy, they offer no assurance that all incorrect values within those boundaries get detected. Embedding the construct directly into evaluative parameters shrinks this shortfall noticeably compared to relying on prompt wording, yet a residual gap remains.
Requirements for a Minimum Viable Validation Protocol for extraction judges
Practically speaking, teams ought to continue relying on LLM judges. Rather than accepting a lone raw-agreement figure as sufficient validation, groups should embrace the handful of design and reporting habits they typically ignore.
The first habit is to start by reporting the right metric. In their 2026)](https://arxiv.org/pdf/2606.19544) recommend that production judge deployments include Cohen's κ adjusted for chance, plus clear numbers for position bias as well as test-retest, under the Minimum Viable Validation Protocol. On its own, raw agreement misleads teams in the deployment decisions with the highest stakes.
The second habit concerns architecture, not statistics: assemble a judging panel from diverse model providers instead of depending on a single evaluator or multiple models sharing one provider family. Including models from two or more separate provider lineages neutralizes idiosyncratic biases, turning any resulting split among evaluators into a meaningful indicator rather than random noise. When evaluators diverge, that split highlights an unclear extraction needing direct human inspection. Panel-based evaluation systems, such as what Openlayer operates, embed this principle intentionally: divergent verdicts across distinct lineages prompt manual review instead of yielding a figure meant for averaging.
Combining chance-adjusted measures, transparent bias and consistency disclosures, schema-limited grading when feasible, plus cross-provider panel reviews yields a framework treating validity alongside reliability as the distinct engineering challenges that research confirms they are. Extraction teams adopting this approach maintain proper control over their LLM evaluators. At last, they capture the actual work those evaluators have been performing.
Sources
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- LLM-as-judge: A complete guide to evaluation best practices in March 2026
- Frontiers
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
- Published as a conference paper at ICLR 2025 JUDGEBENCH: A BENCHMARK
- LLM as a Judge: A 2026 Guide to Automated Model Assessment


