Guide

Embedding Model Migration: Reindex Without Losing Recall

Build a parallel vector index, compare recall on labelled queries, switch a small traffic slice and rehearse rollback.

2026-09-29

Two dark server blocks inside a transparent enclosure, connected by an orange light above cyan circuit traces.
A parallel index keeps the original retrieval path available while the replacement is measured.

1. Record the current index and a question set

Changing an embedding model changes the representation used for search. Treat the migration as a change to both indexing and querying: build a new index while the current one continues serving requests, measure both paths against the same labelled questions, then decide whether to switch. This guide assumes you control the index build and a routing setting in your application. Adapt the operations to your search provider before touching production.

Freeze a snapshot of the existing configuration: corpus revision and document count, chunking rules, filters, index schema, embedding model and version, output dimension, distance metric, query preprocessing, top-k value, reranker settings, and application revision. Keep the old index and its query model deployable. Save the result IDs for a small set of real questions from several categories, including filters, exact terms, short and long queries, and queries expected to have no answer. Remove sensitive content from the evaluation export.

Label relevant document or chunk IDs for each question from the source material, independently of either search result. A useful starter set might contain 30 questions, but 30 is a proposed exercise size, not a statistical guarantee. Mark ambiguous cases for review. Keep a stable ID mapping when chunk boundaries change, or compare at the document level; otherwise a renamed chunk can look like a retrieval loss even when its content is present.

2. Build a parallel index from the same corpus snapshot

In Azure AI Search's index guide, a vector field declares its dimensions and a vector search profile. The dimensions must match the embedding model output. Azure also requires a unique document key and supports human-readable fields alongside vectors. These are concrete checks to translate into your provider's schema. Do not reuse the old vectors with the new model, even if both models happen to output the same number of values: equal length does not establish a shared embedding space.

Create a new index name, such as docs-embed-v2, and retain docs-embed-v1. Generate embeddings for every chunk in the frozen corpus with the new model. Assert vector length before upload; log and stop on any mismatch. Preserve stable IDs, text, metadata, access filters and the same chunking policy for the first comparison. If chunking or filtering also changes, record that as a second experiment because it prevents attribution of a recall change to the embedding model alone. Azure notes that changes to existing vector fields can require a rebuild; the separate index also gives this procedure a simple rollback target.

Check coverage before running relevance tests. Count source chunks, successfully embedded chunks, rejected chunks and searchable chunks; investigate missing or duplicate IDs. Query a few known IDs and confirm that the returned text and permissions match the source. Use the new model to embed new-index queries and the old model for old-index queries. Azure's documentation describes the index as an embedding space populated by one model and pairs indexing skills with equivalent query vectorizers. Never send an old-model query vector to the new index or compare raw similarity scores across the two paths.

3. Compare labelled results using recall@k

Run each labelled question through both paths with the same k, filters, corpus snapshot and downstream retrieval settings. Save the ordered IDs, empty results, errors and latency. For every question that has at least one labelled relevant ID, calculate recall@k = relevant labelled IDs in the top k / all labelled relevant IDs. Average the per-question fractions so a question with many labels does not silently dominate. Report how many questions were eligible and keep unanswered questions in a separate no-answer check. This recall formula is the guide's proposed measure, not a claim about a provider's built-in metric.

For example, if the labels for one question are A and B, and the first five new-index results contain A but not B, its recall@5 is 1/2. If the old path finds both, this case regressed. Inspect the actual text before deciding whether the label or retrieval failed. Microsoft's RAG evaluator guide treats labelled document retrieval as a process evaluation and generated-response quality as a separate system evaluation. A recall gain alone therefore does not establish a better answer. Run an answer-level sample with the same prompts and grading rule if the index feeds a RAG application.

Use this worksheet as a decision record. The thresholds are examples for a local acceptance exercise; set production limits from your own baseline and service objectives before rollout.

CheckProposed pass ruleRecord
Corpus coverageAll expected IDs searchable; zero unexplained rejectsOld / new counts and reject log
Recall@5New mean at least old mean; no critical question loses its required documentMeans, eligible count, regressions
No-answer casesNo new unsafe or irrelevant citations in reviewed casesCase IDs and reviewer notes
Answer sampleNo unacceptable groundedness or completeness regressionPrompt, answer, evidence, decision
Latency and failuresWithin your existing service objectivesp95 by path and error rate

4. Switch a bounded share of traffic

Only after the offline checks pass, route a small, fixed share of eligible requests to the new index. For an exercise, start with 1% for a defined observation window; this is a suggested rollout setting, not a vendor requirement. Assign traffic by a stable user or session key so repeat requests do not bounce between paths. Tag logs with index version, query-model version, corpus revision and application revision. Compare errors, latency, empty-result rate and reviewed answer quality against the old path during the same window. Do not treat different user mixes as a controlled recall experiment.

Hold at each exposure step until the prewritten checks have enough observations for your service. If a required document disappears, permissions change, or the error budget is breached, stop the rollout and route the affected traffic back. This guide's RAG overview explains how retrieval feeds generation; the reliability guide provides a wider incident and fallback context. Keep the old index receiving necessary source updates during the window, or document any data lag that would make rollback return stale content.

5. Rehearse rollback and write the decision

Before full cutover, trigger a rollback in staging using the same routing switch you plan to use in production. Confirm three things with test questions: the route again uses the old index and old query model, required documents still appear, and no cache or queued request continues to return new-path results after the agreed drain period. Time the reversal and name the person who can execute it. A rollback that restores only the old index while leaving the new query encoder active is incomplete.

Write a one-page decision: corpus and code revisions; model versions; dimension and coverage checks; labelled question count; old and new recall@5 with regression IDs; answer sample outcome; latency and error evidence; traffic share and observation window; and the rollback result. The correct answer to the exercise is hold if any required check fails or rollback cannot be demonstrated. It is proceed to the next bounded step only when every stated rule passes. A final production switch remains an operational decision using your own limits and live evidence. No index build, evaluation, traffic switch or rollback has been run for this article.

Sources checked 2026-09-29