Guide

RAG Document Deletion: Test Every Retrieval Path, Including Late Ingestion

Run a reproducible RAG deletion drill: inventory versioned chunks, clear cached context, contain a late ingestion write and keep a control document live.

2026-10-08

A dark cutaway model of a building on a reflective floor, with glowing cyan blocks in a glass shaft, orange light on the left and an orange-lit doorway on the right.
A deletion drill follows a document into every stored copy, not only the first index it was written to.

1. Decide what “deleted” has to mean

A document is removed from your source repository. Your assistant still quotes it. You delete its vectors, check once and close the ticket. The next morning an ingestion worker finishes a job it picked up before the deletion and writes the document back. A useful deletion drill has to cover that last event, not only the first successful delete request.

This guide builds a small, executable model of that failure, then shows how to carry its assertions over to your own pipeline. You need Node.js and a non-production ingestion environment. The fixture uses in-memory maps, not embeddings or a hosted index: it tests your application’s deletion contract, not search quality or a provider’s consistency behaviour. For the wider pipeline, start with Retrieval and RAG for Developers.

Write down two separate completion states. Serving blocked means no newly assembled model context can contain the document. Derived copies removed means every inventoried chunk and cache entry is gone. The first can happen before the second. A delete acknowledgement is evidence for neither: Pinecone’s delete documentation states that the service is eventually consistent, so there can be a slight delay before changed records are visible to queries. Do not describe either state as erasing what a model has learned, removing backups or meeting a legal deletion obligation; this drill tests none of those.

2. Build two synthetic documents and enumerate chunks

The fixture uses two invented documents. removed is the deletion target; control must keep working. The target is edited once, so it has a version 1 and a version 2, each split into two chunks. Chunk IDs carry the source version (removed@v1#0), and the update path deliberately leaves the version 1 chunks in storage, as a careless re-ingestion would. After setup the target has four stored chunks but only two served ones.

That gap is the reason to inventory by document, not by what retrieval currently returns. A search for the document title, or a list filtered to the current version, misses the leftovers. In the fixture, changing inventory to count only servable rows makes the setup assertion fail at once. In your own system, include the tenant or namespace in every identity; the fixture has a single synthetic scope and is not an authorization example.

Before deleting anything, make the target retrievable and prime the cached context. Otherwise an empty result afterwards could simply mean ingestion never worked. Then check the deletion unit your provider offers. Pinecone documents deletion by ID, by metadata filter, or of everything in a namespace, and uses different APIs for indexes with a document schema and for vector indexes. Its vector API accepts at most 1,000 IDs per delete request. The provider deletes the records you name; mapping a source document to all of its chunk IDs and to your own caches is your application’s job.

3. Record the deletion boundary and the source version

Deletion in the fixture first turns the source record into an inactive version 3, then deletes every inventoried chunk and the cached context. The inactive record is a deletion marker holding a version number, not the document text. Every serving read checks rows against it, and every ingestion job compares its own version with it before writing.

Save the code as deletion-drill.mjs and run node deletion-drill.mjs. It needs no packages, credentials or network access. The expected output is PASS: inventory, cache, late write, stale jobs, cleanup, control; any failed assertion exits with an error.

import assert from 'node:assert/strict';

const source = new Map(); // docId -> { version, active }
const chunks = new Map(); // chunkId -> { doc, version, text }
const cache = new Map();  // docId -> rows from an earlier retrieval

const live = row => {
  const s = source.get(row.doc);
  return Boolean(s?.active) && s.version === row.version;
};
const inventory = doc => [...chunks.keys()].filter(id => chunks.get(id).doc === doc);
const retrieve = doc => [...chunks.values()].filter(r => r.doc === doc && live(r));
function context(doc) {
  if (!cache.has(doc)) cache.set(doc, retrieve(doc));
  return cache.get(doc).filter(live);
}
function publish(doc, version) {
  source.set(doc, { version, active: true });
  return { doc, version, parts: ['intro', 'detail'] };
}
async function ingest(job, beforeWrite = async () => {}) {
  const s = source.get(job.doc);
  if (!s?.active || s.version !== job.version) return 'rejected';
  await beforeWrite(); // where a real worker awaits its index client
  job.parts.forEach((text, n) => chunks.set(`${job.doc}@v${job.version}#${n}`,
    { doc: job.doc, version: job.version, text }));
  return 'written';
}
function remove(doc) {
  const s = source.get(doc);
  if (s?.active) source.set(doc, { version: s.version + 1, active: false });
  for (const id of inventory(doc)) chunks.delete(id);
  cache.delete(doc);
}
function reconcile() {
  let removed = 0;
  for (const [id, row] of chunks) if (!live(row)) { chunks.delete(id); removed++; }
  return removed;
}

// 1. Two synthetic documents. The target is edited once, so v1 and v2 chunks coexist.
const v1 = publish('removed', 1);
await ingest(v1);
const v2 = publish('removed', 2);
await ingest(v2);
await ingest(publish('control', 1));
assert.equal(inventory('removed').length, 4); // stale v1 chunks are still stored
assert.equal(context('removed').length, 2);   // only v2 is served; cache is primed
assert.equal(context('control').length, 2);

// 2. A worker holding the v2 job passes its check, then pauses before writing.
let release, reachedPause;
const gate = new Promise(r => { release = r; });
const paused = new Promise(r => { reachedPause = r; });
const late = ingest(v2, () => { reachedPause(); return gate; });
await paused;

// 3. Delete while that worker is paused, and retry the deletion once.
remove('removed');
remove('removed');
assert.deepEqual(source.get('removed'), { version: 3, active: false });
assert.deepEqual(inventory('removed'), []);
assert.equal(cache.has('removed'), false);

// 4. Release the late write, then replay both old jobs.
release();
assert.equal(await late, 'written');          // the race really recreates chunks
assert.equal(inventory('removed').length, 2);
assert.deepEqual(retrieve('removed'), []);     // the serving gate hides them
assert.deepEqual(context('removed'), []);
assert.equal(await ingest(v1), 'rejected');
assert.equal(await ingest(v2), 'rejected');
assert.equal(reconcile(), 2);                  // cleanup removes the late copies
assert.deepEqual(inventory('removed'), []);

// 5. Positive control: the pipeline still serves the retained document.
assert.equal(inventory('control').length, 2);
assert.equal(context('control').length, 2);
console.log('PASS: inventory, cache, late write, stale jobs, cleanup, control');

4. Replay a delayed ingestion job

The interesting part is step 2 of the script. A worker holding the version 2 job passes its version check and then waits at a named barrier, the point where a real worker would await its index client. The deletion runs while the worker is paused, and is retried once to show that a second delete keeps the marker inactive. When the barrier opens, the late write succeeds and recreates two chunks under the deleted document. The script asserts that this happens instead of pretending the race is impossible.

Three mechanisms then contain it. The serving check hides the late rows from retrieval and from context. Replaying the old version 1 and version 2 jobs is rejected because neither matches the current source version. A reconciliation sweep deletes every stored chunk that fails the serving check, and the inventory returns to empty. The control document still returns its two chunks.

Use barriers, not sleeps, when you move this to real workers. A delay may or may not produce the interleaving; a barrier produces it on every run. Log the document ID, source version, job ID and affected chunk IDs so you can reconstruct the order without logging document content.

ScheduleRequired observationIn this fixture
Prime retrieval and cache, then deleteInventoried chunk IDs and cache entry are goneExecuted
Pause a worker after its check, delete, releaseLate rows are not served; reconciliation removes themExecuted
Replay old queued jobs after deletionEach job is rejected and writes nothingExecuted for v1 and v2
Retry the deletionThe marker stays inactive at version 3Executed; partial failures are not modelled
Pause between context assembly and the model callDocumented decision on whether dispatch rechecksNot modelled
Read the control documentTwo chunks and its context remainExecuted

The fixture runs in one process, so nothing else can change the source between the check and the write except the deletion you schedule. A distributed worker has no such protection. Without a transaction spanning source state and index writes, keep an authoritative serving check and reconcile after in-flight writers drain. A lock helps only if every writer and every deletion path honours it. Keep the deletion marker for as long as old jobs can still be replayed; purging it early lets a stale job look like a new document.

5. Invalidate derived serving paths and check negative results

Probe each path separately: direct lookup by the inventoried IDs, semantic retrieval with several document-specific questions, keyword search if you have it, cached assembled context, and the final context handed to the model. Check IDs and content markers at those boundaries. Asking the assistant whether it remembers the document is a weak test, because a generated answer can omit retrieved context or resemble a deleted fact without retrieving it.

For asynchronous cleanup, poll until a written deadline and keep the observations. On Pinecone serverless indexes, the data freshness guide describes a better signal than a fixed sleep: a write response, including a delete, carries an x-pinecone-request-lsn header, and a query response carries x-pinecone-max-indexed-lsn. When the query value is greater than or equal to the delete’s value, that query reflects the delete. If the deadline passes, keep serving blocked and leave the cleanup job open.

List the paths you do not test, such as exports, support tools and backup restores. The conversation-memory guide covers stored conversation records; this drill covers source-document ingestion. During an embedding-model migration, run the probes against both the active and the rollback index.

6. Prove the test can catch a regression

A deletion test that never fails proves little. Each row below was run against a disposable copy of the script with one change; every mutant exited with status 1 at the assertion named.

Change to the fixtureAssertion that failed
Remove the version check in ingestReplay of the version 1 job is not rejected
Remove cache.delete(doc)The cache entry still exists after deletion
Drop the serving check in retrieveRetrieval returns the late-written chunks
Disable the reconciliation sweepThe sweep reports 0 instead of 2 removed chunks
Inventory only servable chunksSetup finds 2 stored chunks instead of 4

What was actually run: the script above and those five mutants, with Node.js 24.13.0 on 8 October 2026, entirely in memory. No Pinecone index, embedding model or LLM was called, and the provider behaviour cited in sections 2 and 5 comes from documentation, not from a test. Your own release evidence should name the tested index, caches, application revision, worker schedules, completion deadline and the control result, and your status display should not report a deletion as complete until those checks pass for every path you operate.

Sources checked 2026-10-08