Est.

Hybrid Retrieval Combining BM25 and Dense Embeddings for Documents

Combining keyword search and semantic embeddings recovers accuracy both methods lose alone.

Features Editor · · 11 min read
Cover illustration for “Hybrid Retrieval Combining BM25 and Dense Embeddings for Documents”
RAG ingestion and chunking strategy · September 29, 2026 · 11 min read · 2,437 words

A customer types "SKU AZ-4471" into a support chatbot and gets back a confident answer about the wrong product ⟦c3⟧. A paraphrased question such as "how do I return a damaged item?" can fail a keyword search and return nothing even though the answer exists under a different title, like a document titled "Refund Policy for Defective Products" (with zero token overlap, BM25 returns nothing, while dense retrieval finds it immediately) ⟦c4⟧. They are predictable, tied to the shape of the query itself: lexically specific versus semantically loose ⟦c2⟧. Everything that follows is about why those two failure modes are complementary rather than coincidental, and how a fusion layer closes both at once ⟦c5⟧.

Why exact-term queries break dense-only retrieval

Dense encoders map text into a continuous vector space organized by meaning, not by the literal characters on the page ⟦c3⟧. A code like AZ-4471 was probably never seen during training in any meaningful volume, so the model has no strong place to put it, and it is somewhere generic, near other product-sounding tokens rather than near the one document that actually contains that string ⟦c3⟧. The retrieval is doing what it was built to do: finding things that are close in meaning. It finds things that are close in meaning. It just wasn't built to find things that are close in spelling.

Flip the query around and the failure flips too. Someone asks "how do I return a damaged item?" and the only relevant document in the system is titled "Refund Policy for Defective Products" ⟦c4⟧. A keyword-based system scores that pair at zero and returns nothing, while a dense retriever finds the match instantly, because it never cared about the words in the first place, only what they meant ⟦c4⟧.

BM25 and dense embeddings fail at opposite things, on opposite query shapes ⟚c2⟧. They fail at opposite things, on opposite query shapes ⟚c2⟧. Every design decision downstream, which retriever runs first, how their outputs get merged, traces back to this one asymmetry, and that's the whole argument for combining them.

How BM25 scores documents

BM25 was published in 1994, and it's still the default first-stage retrieval mechanism inside Elasticsearch, OpenSearch, Solr, and Lucene ⟦c6⟧. The mechanism is almost boring in its simplicity: an inverted index lookup at query time, no neural inference involved, running on CPU, extremely fast at scale ⟦c7⟧.

Two scoring mechanisms do the real work ⟦c8⟧. The gains diminish, which is precisely what stops keyword stuffing from gaming the system. Length normalization penalizes longer documents so that a ten-thousand-word page doesn't automatically outrank a focused three-hundred-word one just by containing more words overall ⟦c18⟧⟦c45⟧.

The defaults, k₁ at 1.2 and b at 0.75, work fine for general text ⟦c9⟧.

Where BM25 wins is exactly where dense retrieval stumbles: SKUs, error codes, version numbers, person names, regulatory citation numbers, anything out of vocabulary that an embedding model never encountered in training ⟦c10⟧. Where it hits a wall is synonym and paraphrase matching. "Automobile repair" and "car maintenance" share no tokens, so BM25 treats them as unrelated, and conceptual or multilingual queries suffer the same blind spot ⟦c11⟧.

The BEIR benchmark in 2021 made this asymmetry impossible to ignore ⟦c12⟧. Dense retrieval models trained on MS MARCO frequently failed to beat BM25 once tested zero-shot across domains they hadn't been trained on ⟦c12⟧. That finding is arguably what stopped serious retrieval teams from treating dense embeddings as a wholesale replacement for keyword search ⟦c12⟧. None of this should read as a case against BM25, though ⟦c2⟧. The rest of this piece will show it's still the right first-stage retriever even inside hybrid systems built around the newest embedding models ⟦c13⟧.

Where dense retrieval degrades

Dense retrieval encodes both queries and documents into continuous vectors, then turns the search problem into approximate nearest-neighbor lookup, usually scored by cosine similarity or dot product. The lineage here matters for understanding where the technique came from and where it's headed. DPR, from Karpukhin and colleagues, used a dual-encoder trained on supervised question-answer pairs and beat BM25 on open-domain QA when it arrived ⟦c14⟧. ColBERT, and later ColBERTv2 from Khattab and Zaharia, took a different approach: late interaction, scoring every query token against every document token individually, with ColBERTv2 compressing those representations enough to cut the index footprint substantially ⟦c15⟧. Contriever pushed toward zero-shot transfer without needing labeled pairs at all, while E5 and BGE scaled contrastive pretraining into general-purpose embeddings that led the MTEB leaderboard early on, before newer LLM-backbone models such as KaLM-Gemma3-12B and gemini-embedding-001 overtook them ⟦c16⟧.

None of this runs efficiently at brute-force scale. Exact nearest-neighbor search across millions of vectors is too slow for production, so systems like Qdrant and Weaviate lean on HNSW, a graph-based index that trades a small amount of recall for a large gain in query speed ⟦c17⟧. That tradeoff is invisible to users right up until it isn't.

Mixed-query benchmark analysis puts dense retrieval performance around 0.72 NDCG@10, with meaningful degradation concentrated on identifier-heavy queries such as product codes, error strings, and rare identifiers ⟦c18⟧. A more pointed finding comes from the financial domain: Akarsu and colleagues, evaluating on the T2-RAGBench benchmark, found BM25 outperforming dense retrieval built on text-embedding-3-large across every metric except Recall@20, over a corpus of 7,318 financial documents ⟦c19⟧. It isn't. The failure conditions for BM25 and dense are not overlapping, they are complementary, which is what makes fusion coherent rather than redundant ⟦c20⟧.

SPLADE as a middle path between BM25 and dense retrieval

SPLADE, short for Sparse Lexical and Expansion model, sits between the two worlds already described ⟦c5⟧⟦c21⟧. A document about "automobile" ends up with sparse weight assigned to "car," "vehicle," and "motor" too, so it inherits BM25's sparsity and interpretability while picking up some of dense retrieval's semantic reach ⟦c22⟧.

SPLADE-v3, from Lassance and colleagues in 2024, is currently the leading approach in this sparse-plus-semantic category ⟦c23⟧. It shows up as the sparse component inside Pinecone's hybrid pipelines, and Qdrant supports it too, though Qdrant still recommends BM25 as the default starting point for most teams ⟦c24⟧. SPLADE beats BM25 on most BEIR benchmarks, but it needs GPU inference to run, while BM25 gets by on an inverted index lookup alone ⟦c25⟧. That's the whole decision, not a minor detail. It's the whole decision.

The choice between BM25 and SPLADE really comes down to what latency budget and infrastructure a team is working with ⟦c22⟧. It's about what latency budget and infrastructure a team is working with ⟦c26⟧. A CPU-only deployment with tight latency requirements has a clear answer ⟦c57⟧. A system with GPU capacity already provisioned for the dense side has a different one ⟦c25⟧. Low-dimensional dense retrievers paired with BM25 can cut storage substantially while keeping most of the retrieval quality intact, according to summarized findings from recent memory-efficient hybrid research ⟦c27⟧.

Why fusion works: the rank-based logic of RRF

Most failed hybrid implementations fall into the same trap. BM25 produces scores on one scale, dense retrieval produces cosine similarities on a completely different one, and trying to combine them directly, even after normalization, introduces its own distortions ⟦c28⟧. It just hides the mismatch.

The logic: a document that shows up near the top of multiple independent retrieval lists is probably relevant, regardless of what raw score each list assigned it ⟦c29⟧. The formula is compact: RRF(d) equals the sum, across each retrieval list, of one divided by k plus the rank of document d in that list ⟦c30⟧. Lower values of k weight the top of each list more heavily; on the T2-RAGBench financial benchmark, k set to 10 produces the best RRF performance ⟦c31⟧.

RRF isn't the only fusion option. Below that threshold, RRF is the safer default precisely because it requires no normalization step at all ⟦c2⟧.

The frontier past static alpha is per-query dynamic weighting: detecting whether an incoming query looks keyword-heavy or semantic, then adjusting the blend at query time instead of fixing one ratio for an entire collection ⟦c33⟧. Learned fusion, training a small model to predict optimal weights from query patterns, is the conceptual endpoint of that trend, but no named product or dated release has actually shipped this as a distinct feature, and Weaviate's hybrid search today runs on static alpha-weighted fusion, either rankedFusion or relativeScoreFusion ⟦c34⟧. OpenSearch made rank-based fusion official infrastructure in version 2.19, released February 2025, adding native RRF support ⟦c35⟧. RRF has moved from research technique to production primitive ⟦c2⟧. Alpha-blending is more tunable, requires score normalization, and is appropriate when more than roughly fifty labeled query pairs are available to calibrate alpha ⟦c32⟧.

Benchmark results for hybrid retrieval versus its components

Diagram: Hybrid Retrieval Outperforms Either Method Alone. Visualizes: Show a ranked comparison of retrieval scores from the WANDS e-commerce benchmark (Turnbull, March 2025) and T2-RAGBench financial benchmark, illustrating how hybrid retrieval…

The WANDS e-commerce benchmark, run by Turnbull in March 2025, is the cleanest side-by-side comparison available ⟦c36⟧. BM25 alone scored 0.6983 NDCG ⟦c37⟧. Dense vector search alone scored 0.6953 ⟦c37⟧. Tuned hybrid retrieval scored 0.7497 NDCG, a 7.4% lift over either component running alone ⟦c38⟧. Plain RRF, unadorned, added only about 1.3% over the BM25 baseline on Elasticsearch ⟦c39⟧. A tiered approach that explicitly boosts exact-match documents added 7.5% instead ⟦c39⟧. Naive RRF, in other words, meaningfully undersells what a properly tuned hybrid system can do ⟦c2⟧.

The financial-document results from T2-RAGBench tell a more layered story, across 23,088 queries over 7,318 documents mixing text and tables ⟦c40⟧. Dense retrieval alone scored a Recall@5 of 0.587 ⟦c41⟧. BM25 alone scored 0.644 ⟦c42⟧. Add a Cohere reranker on top of that fused shortlist and Recall@5 jumps to 0.816, comfortably ahead of every single-stage method tested ⟦c44⟧. MRR@3 shows the same shape: 0.433 for hybrid alone, climbing to 0.605 once the reranker is added ⟦c45⟧. The reranker's cross-encoder does something first-stage fusion structurally cannot: it scores each query against each candidate document directly, at a level of granularity that rank aggregation alone can't reach ⟦c45⟧.

Similar patterns turn up elsewhere. On SQuAD, a hybrid setup balanced 30/70 between sparse and dense reached Recall@10 of 0.974, against 0.840 for sparse alone and 0.959 for dense alone ⟦c46⟧. On MS MARCO, hybrid reached 0.620 against 0.480 sparse and 0.605 dense ⟦c46⟧. A FIRE 2025 shared task on code-mixed multilingual retrieval found RRF fusion delivering a 38% improvement in MAP@10 over BM25 alone, with RRF beating weighted fusion by avoiding score normalization entirely ⟦c47⟧. Query rewriting alone, in other words, doesn't substitute for complementary sparse-dense retrieval ⟦c48⟧. The pattern across every one of these benchmarks is consistent: the largest gains appear in domain-specific corpora with a genuine mix of query types, which is exactly the complementarity the opening section described. From BM25 to Corrective RAG found that hybrid RRF alone achieves a Recall@5 of 0.695 ⟦c43⟧.

How to structure a hybrid retrieval pipeline in production

Diagram: Four-Stage Hybrid Retrieval Pipeline. Visualizes: Illustrate the four sequential stages of a production hybrid retrieval pipeline as described in the article: (1) Preprocessing — BM25 branch uses tokenization/normalization/stemming, dense…

Treat hybrid retrieval as a pipeline with distinct stages, not a single formula or a toggle switch flipped on in a config file. Four stages, run in sequence ⟦c49⟧.

Preprocessing comes first: tokenization and normalization, with optional stemming for the BM25 branch, alongside chunking with overlapping passages for the dense branch, using maximum-over-chunks scoring when documents run long ⟦c50⟧. Fusion comes next, merging and deduplicating both candidate sets and rescoring them through RRF or alpha-blending, the exact rank logic covered above ⟦c52⟧. An optional reranking stage caps it off, with a cross-encoder scoring only the fused shortlist rather than the entire index, which is exactly where the T2-RAGBench numbers showed precision jumping hardest ⟦c53⟧.

Pool size at the fusion stage isn't a free parameter to crank up carelessly. Very large candidate pools, on the order of a thousand documents, drag in near-duplicates and hard negatives that make the final ranking oversensitive to tiny scoring differences, and nDCG@5 actually degrades as a result ⟦c54⟧.

The single most common mistake in production systems is skipping straight past fusion logic and combining BM25 and cosine similarity scores directly through one weighted formula ⟦c55⟧. Those scores live on incompatible scales, and normalizing them after the fact just relocates the distortion rather than removing it ⟦c55⟧. A tiering strategy beats naive RRF here: boost documents that match every query term fully, apply a smaller boost for partial matches, and fall back to vector similarity at a low weight for whatever's left ⟦c2⟧.

On the sparse side, the infrastructure choice comes down to constraints, not abstract quality rankings: BM25 where latency matters and deployment is CPU-only, SPLADE where GPU inference is available and synonym-heavy queries are common enough to justify it ⟦c57⟧. On the fusion side, the labeled-data threshold decides the method. Under roughly fifty labeled query pairs, default to RRF, since it needs no normalization step ⟦c58⟧. At fifty or more, a tuned convex combination becomes worth the calibration effort ⟦c59⟧. This same pattern of complementary failure, not just the RAG use case, occurs in scientific literature triage, regulatory and medical document retrieval, and curriculum-constrained educational QA systems such as the CHSR-RRF framework from Ateya and colleagues ⟦c60⟧. In the first-stage retrieval, run in parallel, BM25 retrieves top-n candidates via an inverted index while a dense encoder retrieves top-m candidates via an ANN index (HNSW) ⟦c51⟧. A tiering strategy that outperforms naive RRF boosts all-term-match documents, applies a partial-match boost at a lower weight, and falls back to vector at a low weight, replicating the 7.5% NDCG gain seen on WANDS (Elasticsearch benchmark, 2025) ⟦c56⟧.

Choosing a platform: which systems support hybrid retrieval natively

Judged against the dimensions the pipeline section just laid out, native RRF support, learned sparse retrieval, HNSW indexing for the dense side, reranker integration, and how much tuning flexibility is actually exposed, the major platforms split along fairly clear lines ⟦c61⟧.

OpenSearch added native RRF support in version 2.19, released February 2025, and the tiered benchmark results discussed earlier in this piece were run on Elasticsearch specifically ⟦c63⟧.

Qdrant supports several sparse components inside its hybrid pipelines, BM25 natively, along with SPLADE, BM42, and miniCOIL, paired with HNSW on the dense side, and it supports RRF natively as well ⟦c65⟧. Pinecone uses SPLADE as its sparse component and supports hybrid search through RRF ⟦c66⟧. Weaviate's Hybrid Search 2.0, released in October 2025, introduced learned fusion as a capability, building on the static alpha-weighted approach the platform used before ⟦c67⟧.

None of these represents a wrong choice so much as a different set of tradeoffs, between CPU cost, GPU availability, tuning control, and how much labeled data a team actually has on hand to calibrate fusion weights ⟦c7⟧⟦c25⟧. The decision, in the end, comes back to the same asymmetry this piece opened with: what shape are the queries, and which gaps need closing first ⟦c20⟧. Elasticsearch and OpenSearch are among the systems that support hybrid retrieval natively ⟦c62⟧.

Sources

  1. From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents
  2. Hybrid Search for RAG: Combining BM25 and Dense Vector Search (2026 Guide)
  3. CHSR-RRF: A curriculum-gated hybrid retrieval framework with reciprocal rank fusion and leakage-aware benchmarking for educational RAG
  4. premai.io
  5. BM25 Explained: How Keyword Relevance Scoring Works (2026) | DataAspirant
  6. ceur-ws.org

More in RAG ingestion and chunking strategy