ArticleRetrieval & Evaluation

RAG Reranking: When Better Ordering Improves Answers

The right chunk can be retrieved and still miss the final context. Compare candidate recall, ordering and answer support before paying for another stage.

Last reviewed: 2026-10-04

Rows of translucent rectangular tiles on a cyan-lit platform beneath dark slabs, with orange light shining through a narrow opening.
Reranking changes which retrieved passages reach the answer context; it cannot recover evidence missing from the candidate pool.
TL;DR

Reranking earns its place when useful evidence is already in the candidate pool but gets cut from the final context. Test that diagnosis before adding another model call. Compare ordering at the same context budget, then check whether the answers actually use the evidence. Reject the extra stage when supported answers do not improve enough to justify its measured latency and cost.

Candidate Recall Sets the Ceiling

Your support assistant finds a document about a deployment failure, yet answers from a generic troubleshooting page. Inspect the retrieved list before changing the generator. Three different failures can look identical in the final response: the specific document was never retrieved; it was retrieved but excluded from context; or it was included and the answer ignored it.

Here is a hypothetical example, not a benchmark. Retrieve 30 candidates and admit five equally sized chunks. The paragraph explaining the deployment failure is at position 14. Reordering could move it to position three and make it available within the existing budget. If the paragraph is absent from all 30 candidates, reordering those candidates cannot recover it. That ceiling follows from the fixed candidate set, rather than from a model performance claim.

Label the evidence needed for each question before inspecting the proposed ranking. A question about a rollback may need both the rollback procedure and a version-specific exception. Mark both as required; one generic document containing the word “rollback” should not count as complete evidence. Include questions whose answer is absent from the corpus, so the evaluation also checks whether the system admits missing support. The Retrieval and RAG guide covers the broader retrieval pipeline; this article focuses on deciding what to change after retrieval.

What Reranking Changes

Cohere's Rerank overview describes sorting a supplied document list by relevance to a query. Its model documentation describes considering query and document tokens together. In this experiment, that becomes a second ordering stage between candidate retrieval and context assembly. Keep document IDs intact so you can map the new ordering back to the original chunks.

A high relevance score does not establish that an answer is supported. Cohere's best practices explicitly describe scores as query dependent and caution against reading too much into their magnitudes. Treat them as ranking signals, not universal probabilities of correctness. If you want a filtering threshold, tune and validate it on your domain's labelled questions rather than copying a score from a demonstration.

Preserve the text actually seen by the scorer. The API reference says that top_n limits returned results and that long document text can be truncated by max_tokens_per_doc. A returned-result count is not a context-token budget. Verify the chosen model, endpoint and document-length handling when moving to production; do not assume the exception near the end of a long page was scored just because the page ID appears in the response.

Compare Equal Context Budgets

Separate the ordering experiment from the retrieval experiment. First freeze a corpus snapshot, labelled question set and candidate list per question. Compare the original ranking with the reranked version of that exact list. Hold the answer model, prompt, context assembly, output-token limit and sampling settings constant. Only the candidate order changes.

Use a token budget as well as a chunk-count cutoff. Five long chunks are a different workload from five short ones. Apply the same deduplication and packing rule in both arms, logging the included IDs and actual input tokens. A valid illustrative rule is “take candidates in rank order, skip any that do not fit, and stop after five admitted chunks or the evidence-token limit.” Record both limits and preserve the rule across arms.

Then run a separate broader-retrieval arm with the existing context budget. Its candidate pool is deliberately different, so it measures a retrieval change rather than isolating ordering. Finally test a larger-context arm using the frozen candidate pool and original ordering. That arm deliberately changes the evidence-token limit. You now have separate comparisons for finding evidence, selecting it and making room for it.

Observed failureNext comparisonWhat would justify the change?
Required evidence absent from candidatesBroader or different retrieval; same final context budgetMore required evidence enters candidates and survives packing
Evidence present but outside admitted contextRerank the frozen pool; same token and chunk limitsMore required evidence enters context and supports answers
Several required sources cannot fit togetherLarger context; same pool and original rankingComplete evidence and supported answers improve within cost limits
Required evidence already included; answer failsInspect grounding instructions and generation separatelyThe answer correctly uses supplied evidence without unsupported additions

Read Retrieval and Answer Metrics Separately

For this worksheet, candidate recall is the fraction of labelled required chunks found in the candidate pool. Context recall is the fraction admitted after packing. Precision at the chosen chunk cutoff is the fraction of admitted chunks labelled relevant. State how you handle questions with no required chunks; exclude them from recall denominators and report their abstention results separately. When chunks duplicate the same evidence, also track required-fact coverage so duplication does not masquerade as completeness.

For generation, use a fixed request: answer the question using the supplied excerpts, cite the excerpt ID for each factual claim, and say when the excerpts lack the answer. Score the same contract: required facts answered correctly, factual claims supported by the cited excerpt, and appropriate admission of missing evidence. A citation attached to the wrong paragraph fails support. Fluent wording and a higher rerank score cannot override that judgment. The AI Evals guide explains how to keep evaluation cases representative.

The following worksheet is deliberately unfilled: no retrieval or generation experiments were executed for this article. Fill it with paired results for the same questions, recording corpus version, question-set version, candidate IDs, scorer version, packing rule, evidence-token budget and evaluation rubric alongside it.

ArmCandidate / context recall; precision at cutoffSupported answers / evaluated questionsEnd-to-end p95 (ms)Cost per request / supported answer
Original order, fixed pool and budgetNot measuredNot measuredNot measuredNot measured
Rerank, identical pool and budgetNot measuredNot measuredNot measuredNot measured
Broader retrieval, original budgetNot measuredNot measuredNot measuredNot measured
Original pool and order, larger budgetNot measuredNot measuredNot measuredNot measured

Report numerators and denominators, not just percentage changes. Look at per-question wins and regressions, including version-sensitive questions and those needing multiple excerpts. Reserve held-out questions for the final decision after tuning. If the observed difference rests on only a few disputed labels, resolve those labels and expand the evaluation before presenting the difference as reliable.

When Latency Outweighs the Gain

Set acceptance limits before looking at results: the minimum useful increase in supported answers, maximum end-to-end p95 latency and maximum cost per supported answer. These are product choices, not vendor guarantees. Calculate cost per supported answer as total evaluation-request cost divided by supported answers under the fixed rubric; if none qualify, report it as undefined. Include retrieval, scoring, generation and retries, state the currency and measurement period, and separately record the reranking stage's latency.

Exercise the failure path too. When reranking times out, does the system use the original order, retry within the request budget, or return an explicit failure? Evaluate the chosen fallback as its own outcome and include it in end-to-end measurements under realistic concurrency. Otherwise the worksheet describes only the successful scoring calls.

Reject reranking when candidate recall is the main bottleneck, when better context ordering leaves answer support unchanged, when gains disappear on held-out questions, or when latency and cost exceed the agreed limits. If larger context gives comparable supported answers at acceptable cost with fewer operational dependencies, that is a valid choice. Keep reranking when the paired evaluation shows a meaningful support gain within your limits, and re-run the comparison when documents, questions or models change.

Primary sources checked 2026-10-04

The diagnosis table, worksheet, formulas and acceptance process are this article's proposed method. The ranking example is hypothetical. No benchmark results, tested improvements or current price claims are presented.