Back
Think you need a better reranker? Check your retrieval first
SiTech AI Team2 წთ. საკითხავი

Think you need a better reranker? Check your retrieval first

The New Stack published an analysis arguing that teams unhappy with poor search results often reach for a bigger reranker before checking whether the right documents ever reach the candidate pool.

The New Stack published an analysis by Kimberly Fessel on September 28, 2026, arguing that teams unhappy with their search results usually invest in a bigger reranking model before checking whether the right documents reach the candidate pool at all. The piece, published with the support of Vespa.ai, proposes a multi-stage retrieval funnel instead of a single expensive reranking step.

Two problems with scaling the reranker

Fessel identifies two costs that come with pouring resources into a larger reranker. First, large-scale reranking gets expensive quickly and raises latency, because the system has to comb through a mountain of candidates before it produces a ranking. Second, a reranker cannot surface better results it never received: if the retrieval stage never returned the right document, no amount of reordering will bring it back.

Her recommendation is to check two things first: whether the correct results are making it into the candidate pool at all, and what it costs to rerank everything in that pool.

The retrieval funnel

Instead of one expensive step, the article suggests treating retrieval as a funnel. It starts with a large corpus and a broad candidate pool built with relatively inexpensive lexical, vector or hybrid search. The pool is then narrowed before a select group of likely answers is passed to reranking based on machine learning inference. Broad coverage happens early, and costly operations are limited to the end of the process.

Designing that funnel means balancing recall, relevance, latency and cost. Retrieving more candidates can improve recall, but it also creates more results to evaluate and potentially rerank, which pushes both cost and latency up.

Where RAG fits

The same logic applies to retrieval-augmented generation. As RAG systems scale, retrieval carries more weight in getting the right information to the model, and even after the right documents arrive, someone has to decide which passages earn space in the context window. Progressive narrowing helps at that stage too.

Vespa.ai’s Bonnie Chase, director of product marketing, and Jenny Morris, senior principal solutions architect, will cover the approach in a live webinar, “Why the Best Reranker Can’t Fix Bad Retrieval,” on October 13. The session is set to address how many candidates to carry between stages, when lexical retrieval is enough, when vectors or hybrid search add value, how to tell whether reranking is actually improving results, and how the funnel extends to RAG context selection.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.