Est.

Hybrid Retrieval Combining Vector Search and Live Web Fetch

Combining static indexes with live web search routes time-sensitive queries to current data.

Contributing Editor · · 11 min read
Cover illustration for “Hybrid Retrieval Combining Vector Search and Live Web Fetch”
RAG & Data Freshness · September 29, 2026 · 11 min read · 2,583 words

Hybrid retrieval reaches its full potential when it extends beyond combining BM25 and dense vectors to also routing time-sensitive queries to a live web fetch, giving AI pipelines both the precision of indexed knowledge and the freshness of the current web. This piece walks through the architecture for that three-mode setup and gives the routing logic for deciding which mode answers which query.

Why single-mode retrieval breaks before it gets to production

Vector embeddings are approximation machines. They're excellent at catching semantic similarity: they know that "car maintenance" and "automobile repair" are talking about the same thing, but they're systematically bad at holding onto exact strings, like a version number, an error code, or a feature flag name. BM25 is the mirror image of that weakness: it nails exact-match terms because it's built on inverted indexes and term frequency, but it has zero semantic representation, so a query for "automobile repair" against a document that only says "car maintenance" comes back empty.

Most real queries don't sit cleanly in one camp or the other. A production system fields questions that need both meaning and exact match in the same breath, which is exactly the condition under which a single retrieval method loses.

That's the argument for combining sparse and dense search. An index reflects the world exactly as it looked the moment it was crawled, and nothing in a hybrid retrieval score tells the end user that the confident-sounding answer they just got is built on six-month-old data. This opens the three-mode framing: the article is about extending hybrid retrieval to include a live-fetch path, not just tuning the sparse/dense blend.

BM25 and dense vector search combined into a working retrieval layer

Dense retrieval works completely differently. Every chunk of text gets pre-encoded into a fixed-size embedding vector ahead of time, the incoming query gets embedded into that same vector space, and the top candidates come back through approximate nearest-neighbor search, using something like FAISS.

BM25 produces unbounded positive real numbers, while cosine similarity is bounded in [-1, 1].

Reciprocal Rank Fusion solves the mismatch by ignoring the raw scores entirely and working off rank position instead Parallel Elasticsearch. Since it never touches the actual score values, there's nothing to normalize Parallel Elasticsearch. Elasticsearch ships k=60 as its production default, a number that traces back to the 2009 Cormack, Clarke, and Büttcher paper at SIGIR, where the authors picked 60 empirically and noted, almost in passing, that the exact value didn't matter much Parallel.

Vendors don't agree on how to implement this, and that disagreement is worth knowing before picking a stack. Elasticsearch keeps native RRF behind an Enterprise license, which pushes teams on the free tier toward a client-side library called ranx as the workaround Lucene.

Fusion gets you a decent candidate list, but reranking is what sharpens it. Cross-encoders score a query and a document together, jointly, so they can't be run across millions of documents at once. The right shape for this is a first-stage ANN retrieval that narrows candidates down to somewhere between top-100 and top-1000, and only then does the cross-encoder rerank that shortlist Voyage. Voyage's benchmarking of its rerank-2.5 model, released in August 2025, reported that adding instruction-following to this step improved accuracy by 7.94% over Cohere's Rerank v3.5 across a 93-dataset test suite.

None of this is theoretical. BM25 mechanics rely on an inverted index, inverse document frequency (IDF) weights that emphasize rare distinguishing tokens, term-frequency saturation controlled by k1 (Lucene default k1 = 1.2), and document-length normalization controlled by b (default b = 0.75) (Lucene). A concrete example from the MARK system (arxiv 2506.23026, published June 2025) shows that hybrid BM25 + dense vector retrieval improved robustness across both general and domain-specific student queries in a deployed classroom QA system.

Static index failure points despite perfect hybrid tuning

Assume the hybrid retrieval layer is tuned as well as it can be. RRF is dialed in, reranking is in place, NDCG numbers look good on every benchmark available WANDS e-commerce dataset. None of that touches the fact that the index only knows what existed at the moment it was built.

Nightly batch re-indexing, still common in a lot of production stacks, is a leftover habit from before streaming pipelines were practical, and by 2026 standards it counts as a design flaw rather than a reasonable tradeoff. Streaming RAG keeps embeddings updated within seconds of a document changing, and the cost of that scales with how often things change, not with how big the corpus is.

Coverage gaps make staleness worse, not just slower to fix. A vector store can only ever answer questions about what's already inside it. A breaking change published an hour ago, or a competitor's pricing update posted this morning, returns nothing at all, and the LLM sitting on top of that empty retrieval either hallucinates an answer or admits it doesn't know. Neither outcome is acceptable in a product a user actually depends on.

The cost side is easy to undercount, too. Managed vector database instances start around $70 a month and climb into the thousands as usage grows, and that number doesn't include the engineering hours that go into ingestion pipelines, chunking strategy, and keeping tabs on whether the index is actually fresh parallel.ai. That's a real, ongoing expense, layered on top of infrastructure that still can't answer questions about anything published after the last crawl.

The scale of this problem is about to get bigger, not smaller. Gartner projected that by the end of 2026, 40% of enterprise applications would include task-specific AI agents, up from under 5% not long before National Assessment of Educational Progress (NAEP). Agents, by their nature, operate against current conditions: current inventory, current pricing, current regulation. A static index structurally cannot meet that requirement, no matter how well it's tuned National Assessment of Educational Progress (NAEP). Finance, retail, travel, and healthcare all generate new, retrieval-relevant content constantly, and grounding an answer in a weeks-old crawl in any of those domains isn't a minor inaccuracy, it's the wrong architecture for the job. For time-sensitive queries, live fetch is the only mode that can actually be correct.

What a live-web fetch does inside a retrieval pipeline

A live-web RAG call runs the whole retrieval process at query time instead of ahead of it. A search API finds candidate sources, those pages get fetched and cleaned, and the content gets chunked and embedded on the spot, against a tiny, disposable, per-query corpus rather than anything resembling a crawl of the open web.

Markdown headings give a chunker natural boundaries that line up with how the document is actually organized, something raw HTML simply doesn't offer, and that keeps each chunk self-contained enough to mean something on its own. Crawl4AI's design draws a useful distinction here between fit_markdown and raw_markdown: fit_markdown strips out navigation bars, footers, sidebars, and other boilerplate, and for feeding an LLM, it's almost always the better choice on token economics alone. The tradeoff is trusting the boilerplate detector doing that stripping, which works well on standard article layouts and less reliably on pages that don't follow that structure.

Getting the raw page in the first place is its own problem. JavaScript-heavy frontends, bot detection systems like Cloudflare or DataDome, infinite scroll, and CAPTCHAs are the default obstacle course.

The output format chosen at this stage shapes everything downstream. Raw HTML loaded with boilerplate creates noisy, badly-bounded chunks; clean markdown or structured JSON is far easier to chunk, store, retrieve, and eventually hand to the LLM.

Routing logic that decides which retrieval mode fires for a given query

Mode two is the live web fetch, for anything time-sensitive or sitting outside the corpus entirely. Mode three sits in between: a vector cache running under a freshness TTL, with live fetch kicking in as the fallback whenever the cache misses or has gone stale.

A handful of signals do most of the routing work in practice. Freshness language in the query itself, words like "latest," "current," "today," "just released," or "as of," is a strong trigger toward live fetch. A weak or scattered confidence score coming back from the vector retriever, where the top result doesn't stand out clearly from the rest, is a sign the indexed corpus doesn't have a good answer and the system should reach outside it. Certain topic categories, like pricing, regulatory filings, competitive intelligence, or release notes, should default to live fetch regardless of what keywords the user typed, simply because those domains move too fast for any index to keep up. And queries carrying specific entities, version numbers, product SKUs, or error codes the embedding model likely never saw during training, call for a hybrid approach with BM25 weighted more heavily, since exact string matching is what actually catches those terms.

None of this requires a heavyweight classifier bolted onto the front of the pipeline. A handful of explicit signal checks handles routing correctly for most production systems, and that simplicity is a feature, not a shortcut.

The TTL governing the cache leg deserves its own attention, because the hybrid middle-ground mode uses a vector cache under a freshness TTL with live fetch as the fallback when the cache misses or is stale. Finance and news content need short TTLs, sometimes measured in minutes. Product documentation can tolerate something much longer. The routing layer should treat this as a configuration exposed per domain. A latency budget causes this: live fetch adds real time and real per-query cost, while cached vector retrieval stays sub-millisecond at any reasonable corpus size. A user-facing chat interface and a background research agent have very different tolerances here, and that difference belongs in the routing decision, not as an afterthought bolted on later.

Structured data extraction and its impact on the live-fetch path

Once a live fetch pulls content back, the next question is what shape that content needs to take before it's useful. Modern LLM APIs support structured output natively now: the model can be constrained to return valid JSON matching a schema a developer defines ahead of time, and Pydantic in Python or Zod in JavaScript have become the standard tools for writing those schemas. The objective, what to look for, and the schema, how to format what's found, are separate concerns. The same extraction logic can serve more than one downstream consumer without being rebuilt each time.

Cost is a real constraint here. A February 2026 paper out of Cairo University, on a system called AXE, showed that pruning the DOM intelligently before it ever reaches the LLM cut input tokens by 97.9% while holding an F1 score of 88.1%. That's a meaningful result: the bottleneck for good extraction isn't always a bigger model, sometimes it's a smarter preprocessing step.

On the SWDE benchmark, Amazon's PARSE framework achieved up to a 64.7% improvement in extraction accuracy and reduced extraction errors by 92% within the first retry. For any team whose schemas need to serve both a human engineer reading the code and an LLM agent calling the API, that's a meaningful design choice.

Schema validation earns its keep in a second way too: it catches hallucination. When a model invents a plausible-sounding value that isn't actually present anywhere in the source document, enforcing the schema exposes that mismatch in validation instead of quietly passing it downstream. That matters most in agent pipelines, where a tool call further down the chain is going to act on whatever value the extraction step handed it.

The WebLists benchmark found that state-of-the-art web agents managed only 31% recall on structured extraction tasks, while a much simpler record-and-replay system built on CSS selectors hit 66%. Pure LLM agents cap out at a few thousand tokens of output, which makes large-scale extraction impractical without wrapping the LLM call in a programmatic loop. Structured extraction pipelines need a looping or pagination layer around the model call, not a single one-shot prompt and a hope.

Infrastructure choices that determine whether the three-mode system holds up at scale

A live-fetch pipeline is only as reliable as its ability to actually reach the page. A large share of the content worth retrieving sits behind geo-gating, rate limits, or anti-bot systems, and getting to it reliably at scale requires proxy infrastructure. In RAG, a missed retrieval means one specific user question comes back stale or empty, and that failure is visible. High success rates and consistency aren't nice extras on top of a fetch layer, they're the baseline requirement.

Redis has become a common answer to the caching leg of this architecture, offering semantic caching (through Redis LangCache for a managed option, or RedisVL for teams running it themselves), alongside vectors, metadata, and general operational data, all inside one system with sub-millisecond latency. That's relevant twice over: once for the cached-vector leg of the hybrid retrieval layer, and again for agent memory in pipelines that run over long stretches of time.

The broader market is voting with its investment on how central this infrastructure is becoming. The AI-based web scraping market is projected to hit $3.16 billion by 2029, growing at a 39.4% compound annual rate, which says something about how much of the industry now treats programmatic web access as durable infrastructure rather than a one-off utility bytetunnels.com.

When an agent needs to walk across a site rather than fetch one known URL, a crawl API earns its cost by handling queues, retries, throttling, and result pipelines, and teams should not build this logic themselves for production workloads. For coding assistants and research agents built on an LLM co-pilot pattern, wrapping a fetch tool as a CLI skill the agent can invoke directly keeps the tool-use loop short and avoids a bespoke integration layer. Taken together, these point toward a build-versus-buy tradeoff that isn't close: maintaining custom scrapers, crawlers, and proxy pools is undifferentiated work sitting off the critical path of the actual product, and a single well-designed API removes that maintenance burden entirely.

Wiring live fetch, hybrid search, and caching together as one retrieval layer

Reach for live web fetch when the query carries freshness signals, when the topic moves faster than any re-indexing schedule could keep pace with, or when the answer requires a source the index has simply never encountered. Reach for the hybrid middle, a cached layer under a TTL with live fetch as fallback, when the corpus is semi-stable: not frozen, but not changing every minute either, where most queries can ride the speed of the cache and only the edge cases need to pay the cost of a live fetch.

The live-fetch path's results get chunked, embedded, and handed to the LLM using the same downstream conventions as the indexed path, so the model doesn't need to treat the two differently once the context reaches it https://github.com/sakshirautela/skin-disease-docai/issues/2.

Three modes, one routing layer, and a clear sense of which failure each mode is built to prevent, exact-match misses, stale answers, or an index that never saw the source at all, is the architecture that combines the exact-match, stale-answer, and unindexed-source safeguards described above. Getting any one piece wrong, tuning the hybrid brilliantly while ignoring staleness, or building a live fetch without proxy reliability behind it, reintroduces the exact failure the other two pieces were built to close. Use BM25+vector hybrid when the retrieval target is a stable internal corpus (documentation, compliance archives, private knowledge bases), latency must be sub-100ms, and content is not public-web sourced. BM25 and vector search run in parallel, with RRF merging the results.

Sources

  1. Machine Assistant with Reliable Knowledge: Enhancing Student Learning via RAG-based Retrieval
  2. How to Build a Web RAG Pipeline Without Vector Databases
  3. Hybrid Search: BM25, Vector & Reranking Reference 2026
  4. [Core RAG] Vector Embeddings & Hybrid Retrieval Engine (Dense + BM25) · Issue #2 · sakshirautela/skin-disease-docai

More in RAG & Data Freshness