Est.

RAG Evaluation Frameworks Compared for Document QA

Faithfulness scores miss half the problem—here's which metrics actually catch retrieval failures.

Contributing Editor · · 10 min read
Cover illustration for “RAG Evaluation Frameworks Compared for Document QA”
Evaluation, benchmarks and ground truth · October 1, 2026 · 10 min read · 2,314 words

RAG systems break down at two distinct stages, yet the measurements teams typically rely on first obscure both issues. Take one legal-research tool earning a 0.91 faithfulness rating in pre-launch testing, which appeared ready for release. Customers still said about one in six responses left out a key statute. That rating held up since the model assembled logically sound replies grounded in its supplied material, regardless of gaps. Pre-launch metrics never caught the retrieval gap, because that rating merely confirms alignment with fetched passages rather than their sufficiency.

Evaluating RAG is tougher than standard NLP evaluation, since the pipeline has two separate places where things can go wrong and overlap-based scores spot neither one dependably. ROUGE and BLEU check how many words match the gold answer, a technique meant to grade translations, not to verify that an LLM kept to the facts the retriever supplied or that the retriever pulled the correct documents at all. A RAG pipeline runs in two distinct phases, first retrieval and then generation, and either can break while the other keeps going. The retrieval step may overlook the crucial document, yet the generator keeps producing polished, assured answers built on whatever scraps it received. The generator, meanwhile, may invent content far beyond the documents the retriever delivered, even when retrieval itself performed flawlessly. Neither breakdown is caught dependably by overlap metrics, which were simply never designed to tell the two phases apart.

Conventional benchmarks treat the index as reliable and up-to-date, leaving issues such as stale indices, damaged data lineage, or obsolete document versions beyond their intended scope. Comparing frameworks must begin with that premise if individual tool scores are to carry any weight.

The four metrics that form the evaluation vocabulary every framework shares

A practical RAG evaluation baseline now centers on four measures spanning generation and retrieval alike, with checks for grounded answers, useful responses, precise retrieved context, and complete context coverage as the core set. Because those four metrics span both pipeline stages, no framework comparison means much without a clear view of the errors each metric flags and the ones it overlooks. TruLens sets itself apart from RAGAS and DeepEval through always-on production oversight, pairing feedback functions with OpenTelemetry traces, while AIMultiple ranks it highest for separating context relevance, even though TruLens still cannot tell factually wrong contexts from factually correct ones.

Faithfulness verifies that a response stays grounded in the sourced material rather than introducing unsupported assertions, thereby preventing fabrication during output creation. Answer relevance ensures the reply tackles the user's prompt directly, flagging instances where generation wanders despite relying on valid source material. Context precision evaluates how effectively sourced segments are ordered, ensuring the most pertinent details appear upfront instead of being obscured by irrelevant data. Context recall examines a separate concern: whether the gathered sources hold all necessary information to address the query, independent of their ordering.

Faithfulness hit 0.91, yet context recall reached just 0.62. On multi-hop queries requiring facts drawn from several documents, the retriever consistently overlooked a second statute. Such a shortfall escapes faithfulness, since that score merely confirms the response aligns with what was fetched rather than verifying nothing was left out. Only context recall exposes this specific blind spot, so evaluating frameworks demands clarity on which among those four measurements they truly publish.

What RAGAS does well for document QA

RAGAS was the first to let RAG pipelines be judged without reference answers, and document QA teams working on LangChain still do best to begin there, though when a system enters production its tight customization limits become genuine blind spots. It started out as a research write-up in early 2023 that scored RAG systems without hand-written answers, and after OpenAI name-checked it during the November 2023 event for developers, it drew wide attention.

RAGAS stands out for document QA because it covers all four canonical metrics with no ground-truth labels required, offering LangChain-stack teams a setup that feels lightweight in Python and smooth to script. That is what makes it the go-to choice when teams need a quick first evaluation.

The critique rests on concrete failure modes, not broad misgivings. EncouRAGe’s authors say current evaluation toolkits leave LLM-as-a-judge metrics with too little tailoring space and no prompt-level control for case-specific evaluations, creating a serious hole once domain calibration is needed. In specialized document QA, teams usually need domain-specific calibration, so that gap becomes hard to ignore. RAGAS is thin on interpretability, produces unwieldy logs, and offers narrow error handling, leaving teams with a score but little guidance on what to repair. The strongest critique is that, although RAGAS is framed as reference-free and broadly applicable, enterprise document QA often violates the premises behind broad benchmarks: standalone questions and retrieval that is both exhaustive and accurate. In corpora for law, healthcare, or company policy, that tidy fit is uncommon.

How DeepEval extends the metric surface

DeepEval covers more evaluation ground than any other option, spanning agentic processes, conversational turns, safety verifications, plus multimodal capabilities, and it hooks straight into CI/CD via Pytest, but a medical AI study documented a rank inversion proving that metric breadth alone cannot guard against systematic judgment errors. Unlike RAGAS and its notebook-centered, dataset-driven workflow, DeepEval embeds directly into deployment pipelines where evaluation acts as a gate rather than an isolated analytical phase. Of the three frameworks, its metric catalog is the most extensive, covering agentic processes, conversational turns, MCP checks, safety verifications, plus multimodal capabilities.

Breadth of coverage does not guarantee correctness of judgment, though, and a medical AI study makes that point concretely. The medical AI study is the clearest evidence of the judgment inversion risk: DeepEval ranked Arm B above Arm A (0.818 vs. 0.743) while an outcome-quality metric ranked Arm A above Arm B, a direct rank inversion among RAG baselines. That's a direct rank inversion between two systems doing the same job, produced by a framework with far more metric dimensions than RAGAS. Metric count did not prevent it.

For QA over documents, the practical risk is clear: when groups rely on metric outputs to select retrieval methods or weigh pipeline variants, inverted rankings steer them toward inferior setups no matter the breadth of criteria purportedly assessed. Covering more failure modes throughout a pipeline helps, yet no individual assessment, be it for relevance, faithfulness, or another criterion, ensures two rival systems are ranked correctly.

TruLens as a production monitor

TruLens solves a different problem than RAGAS or DeepEval. It's built for continuous monitoring in production, scoring live traffic through feedback functions and tracing behavior with OpenTelemetry, rather than running as a one-time evaluation pass over a fixed dataset. Its Triad, context relevance, groundedness, and answer relevance, is evaluated continuously via feedback functions rather than as a batch eval run. The OpenTelemetry tracing is what sets it apart operationally: it ties evaluation scores directly to specific production traces, so a low score can be traced back to the exact query and retrieval call that produced it.

In a March 2026 AIMultiple evaluation comparing how well leading platforms distinguish relevant context, TruLens achieved the top ratio among those tested. It correctly identified contextual alignment with greater consistency than its peers while producing the fewest backwards assessments. This genuine advantage most clearly shows that TruLens handles contextual relevance in ways rival frameworks cannot match.

The same benchmark revealed that this weakness ran through the full test set, with TruLens among them: all five systems failed to judge whether a context’s facts were wrong or correct. On that AIMultiple test, TruLens and the other systems likewise failed to tell false-evidence contexts from accurate ones, and each ranked hard negatives above partial contexts, reversing the relevance order that should have held. For these tools, a context-relevance score reflects topic match, not whether the content is true. Systems can therefore earn strong context-relevance marks across every major framework while retrieving on-topic documents whose facts are false, leaving document QA with a gap.

The shared reliability problem in LLM-as-a-judge evaluation

RAGAS, alongside DeepEval and TruLens, derives its scoring from one basic pattern: one LLM evaluates the response produced by a second LLM. Sandler and colleagues’ work with the RAND Corporation’s Judge Reliability Harness showed that every judge model had reliability gaps, with agreement breaking even after routine edits such as reformatting, rewording, or altering answer length.

LLM-as-a-judge won out as the go-to method since it scales beyond the reach of human reviewers and picks up on semantic subtleties that metrics like ROUGE and BLEU simply were not designed to detect. The validation behind these judges, though, usually hinges on whether their verdicts match a human rater's exactly, a figure that ignores agreement by luck and therefore makes the judge seem sharper than it truly is. Familiar biases make things worse: judges favor the answer listed first, a position bias; they award points for length even when quality suffers, a verbosity bias; and they score replies higher the more those replies echo their own style, a self-enhancement bias.

RAND's benchmark suite put four judge models through safety, influence, abuse-risk, and autonomous-task evaluations, yet none proved reliable on consistency and discrimination alike. Practically, a judge tried only on chat should not be relied on to assess RAG systems, review code, or grade agent tasks unless it is tuned for that domain, which is precisely the work most teams leave out. LLM judging is most defensible as a rough signal, and in many document QA settings, movement over time can still help even if the exact scores cannot be trusted. That argument works for monitoring direction over time, not for release gates or pipeline head-to-heads, since reversals such as the medical AI study's rank flip can carry real consequences.

What ARES and RAGChecker add

Two specialized tools address shortcomings that leading frameworks overlook, with each tackling a distinct challenge. ARES enables teams to evaluate retrieval pipelines under simulated load prior to live deployment. Presented during NeurIPS 2024, RAGChecker supplies the granular analysis required to determine if an error originated from the retrieval or generation stage.

ARES works by fine-tuning compact language-model judges to score context relevance, answer faithfulness, and answer relevance, and it can generate synthetic queries and answers straight from a team's own documents. That synthetic-generation capability lets a team build an evaluation set and get real signal before a single user has touched the system. The tradeoff is calibration cost. ARES needs a meaningful amount of human labeling to fine-tune its judges properly, which is a real expense for any team that picked RAGAS in the first place specifically to avoid ground-truth annotation.

RAGChecker takes another route, benchmarking the retrieval-generation handoff so individual failure points can be pinpointed. Such a check might have surfaced the context recall failures in legal-research workflows long before customers realized statutes were missing.

EncouRAGe, a newer framework whose preprint dropped in October 2025, tackles the customization gap head-on. Built as an open-source library in Python, it bundles ten RAG approaches, a manifest of object-oriented types, and a wide set of metrics, checked on four datasets spanning 25,000 question-answer pairs plus a large collection of documents. By prioritizing running locally, reproducible results, and user-defined LLM-as-a-judge evaluation prompts, it answers the very customization shortcoming RAGAS shows. One concrete result, showing Hybrid BM25 turned in top retrieval performance on every dataset of the four it evaluated, hands document QA teams something real to weigh as they choose a retrieval method, not just generic advice.

At EvalLLM 2026, Orange Research presented a rare direct comparison using real enterprise data rather than academic benchmarks. The team evaluated both the model’s answers and the passages it pulled back, using Ragas, DeepEval, Opik, and RAGChecker to measure a QA dataset that human annotators assembled from actual business data. Its value lies in the rarity of such enterprise-data comparisons against fabricated academic material, a point to note without giving one study more weight than it can bear.

The evaluation gaps that no current framework covers out of the box

Index trustworthiness leads that list. Even when a system rates 0.95 for faithfulness by merely comparing generated text to fetched passages, outdated indexes can produce incorrect results. All the measurements covered here, including faithfulness alongside context recall and context precision, take for granted that the queried index remains up to date with accurate origins. None of them examine data ownership, recency, or lineage integrity, leaving a fifth evaluation dimension entirely unaddressed by the frameworks discussed.

Questions lacking any valid response reveal a similar shortcoming. Because standard approaches presume every question can be resolved from the provided texts, they struggle when no valid response is present. Dedicated tools such as UAEval4RAG tackle this exact blind spot, handling cases where standard methods falter by presuming the corpus holds the solution.

Generic RAG evaluation presupposes queries that need no context, retrieval that is both full and accurate, and ground truth checkable from outside. A search may return mostly correct content yet omit one essential procedural action, sources may blend superseded policy phrasing with its up-to-date form, and many questions hingeon internal processes beyond any public benchmark's scope.

The test questions themselves introduce a separate coverage gap. Boston Consulting Group designed its semantic coverage metric to address this precise blind spot: while each framework discussed here evaluates a model's accuracy on provided prompts, not one verifies that those prompts collectively represent the underlying knowledge domain. Models may excel against a narrow, skewed evaluation suite yet collapse in real-world deployment without any conventional benchmark flagging the risk ahead of time.

GroUSE also flags faithfulness, a top concern for document QA teams: even leading judges can miss serious grounding errors in QA scores. And RAG-QA Arena uses LFRQA’s seven-domain query set to show that results tied to a single domain, language, and document type do not reliably carry over to another. Ultimately, teams are not selecting a flawless framework; they are deciding which gaps they can accept.

Sources

  1. Best Open Source RAG Frameworks in 2026: Comparison and Guide
  2. EncouRAGe: Evaluating RAG Local, Fast, and Reliable
  3. RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
  4. Methodological Framework for Quantifying Semantic Test Coverage in RAG Systems

More in Evaluation, benchmarks and ground truth