
Reranking earns its place when useful evidence is already in the candidate pool but gets cut from the final context. Test that diagnosis before adding another model call. Compare ordering at the same context budget, then check whether the answers actually use the evidence. Reject the extra stage when supported answers do not improve enough to justify its measured latency and cost.
Candidate Recall Sets the Ceiling
Your support assistant finds a document about a deployment failure, yet answers from a generic troubleshooting page. Inspect the retrieved list before changing the generator. Three different failures can look identical in the final response: the specific document was never retrieved; it was retrieved but excluded from context; or it was included and the answer ignored it.
Here is a hypothetical example, not a benchmark. Retrieve 30 candidates and admit five equally sized chunks. The paragraph explaining the deployment failure is at position 14. Reordering could move it to position three and make it available within the existing budget. If the paragraph is absent from all 30 candidates, reordering those candidates cannot recover it. That ceiling follows from the fixed candidate set, rather than from a model performance claim.
Label the evidence needed for each question before inspecting the proposed ranking. A question about a rollback may need both the rollback procedure and a version-specific exception. Mark both as required; one generic document containing the word “rollback” should not count as complete evidence. Include questions whose answer is absent from the corpus, so the evaluation also checks whether the system admits missing support. The Retrieval and RAG guide covers the broader retrieval pipeline; this article focuses on deciding what to change after retrieval.
What Reranking Changes
Cohere's Rerank overview describes sorting a supplied document list by relevance to a query. Its model documentation describes considering query and document tokens together. In this experiment, that becomes a second ordering stage between candidate retrieval and context assembly. Keep document IDs intact so you can map the new ordering back to the original chunks.
A high relevance score does not establish that an answer is supported. Cohere's best practices explicitly describe scores as query dependent and caution against reading too much into their magnitudes. Treat them as ranking signals, not universal probabilities of correctness. If you want a filtering threshold, tune and validate it on your domain's labelled questions rather than copying a score from a demonstration.
Preserve the text actually seen by the scorer. The API reference says that top_n limits returned results and that long document text can be truncated by max_tokens_per_doc. A returned-result count is not a context-token budget. Verify the chosen model, endpoint and document-length handling when moving to production; do not assume the exception near the end of a long page was scored just because the page ID appears in the response.
Compare Equal Context Budgets
Separate the ordering experiment from the retrieval experiment. First freeze a corpus snapshot, labelled question set and candidate list per question. Compare the original ranking with the reranked version of that exact list. Hold the answer model, prompt, context assembly, output-token limit and sampling settings constant. Only the candidate order changes.
Use a token budget as well as a chunk-count cutoff. Five long chunks are a different workload from five short ones. Apply the same deduplication and packing rule in both arms, logging the included IDs and actual input tokens. A valid illustrative rule is “take candidates in rank order, skip any that do not fit, and stop after five admitted chunks or the evidence-token limit.” Record both limits and preserve the rule across arms.
Then run a separate broader-retrieval arm with the existing context budget. Its candidate pool is deliberately different, so it measures a retrieval change rather than isolating ordering. Finally test a larger-context arm using the frozen candidate pool and original ordering. That arm deliberately changes the evidence-token limit. You now have separate comparisons for finding evidence, selecting it and making room for it.
| Observed failure | Next comparison | What would justify the change? |
|---|---|---|
| Required evidence absent from candidates | Broader or different retrieval; same final context budget | More required evidence enters candidates and survives packing |
| Evidence present but outside admitted context | Rerank the frozen pool; same token and chunk limits | More required evidence enters context and supports answers |
| Several required sources cannot fit together | Larger context; same pool and original ranking | Complete evidence and supported answers improve within cost limits |
| Required evidence already included; answer fails | Inspect grounding instructions and generation separately | The answer correctly uses supplied evidence without unsupported additions |
Read Retrieval and Answer Metrics Separately
For this worksheet, candidate recall is the fraction of labelled required chunks found in the candidate pool. Context recall is the fraction admitted after packing. Precision at the chosen chunk cutoff is the fraction of admitted chunks labelled relevant. State how you handle questions with no required chunks; exclude them from recall denominators and report their abstention results separately. When chunks duplicate the same evidence, also track required-fact coverage so duplication does not masquerade as completeness.
For generation, use a fixed request: answer the question using the supplied excerpts, cite the excerpt ID for each factual claim, and say when the excerpts lack the answer. Score the same contract: required facts answered correctly, factual claims supported by the cited excerpt, and appropriate admission of missing evidence. A citation attached to the wrong paragraph fails support. Fluent wording and a higher rerank score cannot override that judgment. The AI Evals guide explains how to keep evaluation cases representative.
The following worksheet is deliberately unfilled: no retrieval or generation experiments were executed for this article. Fill it with paired results for the same questions, recording corpus version, question-set version, candidate IDs, scorer version, packing rule, evidence-token budget and evaluation rubric alongside it.
| Arm | Candidate / context recall; precision at cutoff | Supported answers / evaluated questions | End-to-end p95 (ms) | Cost per request / supported answer |
|---|---|---|---|---|
| Original order, fixed pool and budget | Not measured | Not measured | Not measured | Not measured |
| Rerank, identical pool and budget | Not measured | Not measured | Not measured | Not measured |
| Broader retrieval, original budget | Not measured | Not measured | Not measured | Not measured |
| Original pool and order, larger budget | Not measured | Not measured | Not measured | Not measured |
Report numerators and denominators, not just percentage changes. Look at per-question wins and regressions, including version-sensitive questions and those needing multiple excerpts. Reserve held-out questions for the final decision after tuning. If the observed difference rests on only a few disputed labels, resolve those labels and expand the evaluation before presenting the difference as reliable.
When Latency Outweighs the Gain
Set acceptance limits before looking at results: the minimum useful increase in supported answers, maximum end-to-end p95 latency and maximum cost per supported answer. These are product choices, not vendor guarantees. Calculate cost per supported answer as total evaluation-request cost divided by supported answers under the fixed rubric; if none qualify, report it as undefined. Include retrieval, scoring, generation and retries, state the currency and measurement period, and separately record the reranking stage's latency.
Exercise the failure path too. When reranking times out, does the system use the original order, retry within the request budget, or return an explicit failure? Evaluate the chosen fallback as its own outcome and include it in end-to-end measurements under realistic concurrency. Otherwise the worksheet describes only the successful scoring calls.
Reject reranking when candidate recall is the main bottleneck, when better context ordering leaves answer support unchanged, when gains disappear on held-out questions, or when latency and cost exceed the agreed limits. If larger context gives comparable supported answers at acceptable cost with fewer operational dependencies, that is a valid choice. Keep reranking when the paired evaluation shows a meaningful support gain within your limits, and re-run the comparison when documents, questions or models change.
- Cohere Rerank overview: ordering supplied documents against a query.
- Cohere Rerank model details: query/document processing and model-dependent context handling.
- Cohere reranking best practices: interpreting query-dependent relevance scores.
- Cohere v2 Rerank API: returned-result limits and document truncation.
The diagnosis table, worksheet, formulas and acceptance process are this article's proposed method. The ranking example is hypothetical. No benchmark results, tested improvements or current price claims are presented.