ArticleAI Architecture

RAG vs Fine-Tuning: Which One Your AI Feature Needs

Retrieval changes what the model can see on each request. Fine-tuning changes how a particular model behaves. They fix different problems, carry different costs, and age differently when the base model is retired.

Last reviewed: 2026-10-03

Two dark workbenches on a reflective floor: a small machine and a stack of files inside a glass case on the left, an open lathe with a small calendar on the right.
One setup keeps its reference material close at hand; the other reshapes the part itself. Retrieval and fine-tuning differ in the same way.
TL;DR

Start with a better prompt and an eval set. If the model lacks facts, add retrieval. Consider fine-tuning only when the remaining failures are about format or behavior and you have a vendor that still trains models for you. In 2026 that last condition is the hard one: OpenAI is winding down its fine-tuning platform, and the Claude API does not currently offer fine-tuning.

What Each Approach Actually Changes

Suppose you are building a support assistant that routes tickets and answers questions about your product. Two kinds of failure show up. The assistant does not know that plan limits changed last week, and it writes free-form answers when your ticket system expects a fixed label plus a one-line reason. These are different problems.

The first is a knowledge problem. Retrieval-augmented generation fetches relevant documents at request time and places them in the context window, as Anthropic's glossary describes it. The weights stay the same; what changes is the evidence the model reads. Update the document store and the next request sees the new plan limits.

The second is a behavior problem. Fine-tuning further trains a model on example inputs and outputs. OpenAI's model optimization guide lists classification, nuanced translation, generating content in a specific format and correcting instruction-following failures among the uses of supervised fine-tuning. Those are all about shape and consistency, not about facts that change every week. A fine-tuned model that memorized last month's prices still needs retraining when the prices move.

Prompt First: the Cheapest Baseline

Both vendors put measurement before technique. Anthropic's prompt engineering overview assumes you already have success criteria, a way to test against them empirically, and a first draft prompt. It also notes that not every failing eval is best fixed with prompting; sometimes a different model is the easier route to lower latency and cost. OpenAI's guide describes the same loop: write evals, prompt with relevant context, fine-tune only "for some use cases", then measure again.

For the ticket router, a reasonable baseline is a prompt with the label list, three or four worked examples and a strict output format, scored against a few dozen real tickets you have labeled by hand. That set size is a starting suggestion, not a statistical threshold. If the format failures disappear here, you have solved the behavior problem without training anything. The AI Evals guide covers how to build that set so it reflects production inputs rather than easy cases.

When Retrieval Is the Answer

Retrieval is the right tool when the failures are missing, private or fast-changing facts. OpenAI's guide makes the same point for prompting: include content the model needs from outside its training data, such as data from private databases or current information. Retrieval is how you do that at scale.

Check whether you need a retrieval pipeline at all. In its 2024 Contextual Retrieval post, Anthropic suggested that a knowledge base under 200,000 tokens, about 500 pages, could go into the prompt directly, with prompt caching to reduce the repeated cost. Treat that as a threshold to test against your model's current context window and pricing, not as a fixed rule.

Above that size, retrieval quality is the thing to measure. The same post reports that in Anthropic's experiments, averaged across codebases, fiction, arXiv papers and science papers, adding chunk-level context and BM25 cut the top-20 retrieval failure rate from 5.7% to 2.9%, and to 1.9% with reranking (metric: 1 minus recall@20). Those are their corpora and settings, not a forecast for yours. Two lessons transfer: lexical search such as BM25 helps when users type exact identifiers like error codes, and reranking adds latency and cost per request. The Retrieval and RAG guide walks through chunking, hybrid search and evaluation in detail.

The maintenance cost of retrieval sits in the data: re-indexing when documents change, deciding who may see which chunks, and re-running retrieval evals when you change the embedding model. None of it depends on a particular generation model, which matters in the next two sections.

When Fine-Tuning Still Earns Its Cost

OpenAI lists the benefits of fine-tuning over prompting alone: you can show more examples than fit in one request, use shorter prompts that save tokens and latency at scale, train on proprietary data without sending it with every request, and train a smaller model to handle a task where a larger one is not cost-effective. Those are real advantages for high-volume, narrow tasks such as the ticket router, where the label set is stable and the prompt would otherwise carry many examples.

The costs are different in kind. You need a curated dataset of correct outputs, a training run per iteration, and an eval that compares the tuned model with the prompted baseline on the same tickets. If the tuned model does not beat a well-prompted base model on that set, the shorter prompt alone does not justify the training work. And the result is tied to one base model, as How AI Models Are Trained explains in more depth: fine-tuning adjusts an existing model, so it cannot outlive it.

The 2026 Vendor Reality

The base-model dependency is no longer theoretical. OpenAI's deprecations page lists these steps for its self-serve fine-tuning platform:

DateOpenAI's listed change
May 7, 2026Creating fine-tuning jobs or training is not available to organizations that have not previously run fine-tuning.
July 2, 2026Creating fine-tuning jobs is no longer available to organizations that have not run inference on a fine-tuned model in the past 60 days.
January 6, 2027Active existing customers can no longer create new fine-tuning jobs.

Inference on existing fine-tuned models continues until the underlying base model is deprecated. The same page shows what that means in practice: in its April 22, 2026 deprecation list, the fine-tuned snapshot ft-gpt-4.1-nano-2025-04-14 is scheduled to shut down on October 23, 2026, the same date as its base model, with gpt-5.6-luna named as the recommended replacement base model. Your training data survives; the tuned model does not.

On the Anthropic side, the Claude API glossary says the API does not currently offer fine-tuning and suggests asking your Anthropic contact if you want to explore it. This article does not cover fine-tuning on other platforms or vendors.

For planning, that makes retrieval and prompts the portable layer: they move to a new model with re-testing, while a tuned model has to be rebuilt, if a vendor still lets you. When Your AI Model Gets Pulled covers the migration side of that risk.

Decision Table and Three Questions

This table is the order this article proposes, not a vendor recommendation:

StepUse it when the failure isWhat you maintainExposure when the base model is retired
1. Prompt and examplesUnclear instructions, wrong format, missing examplesPrompt versions and an eval setRe-test the prompt on the replacement model
2. RetrievalMissing, private or frequently changing factsDocument store, index, access rules, retrieval evalsRe-test generation; the index can stay unless you change embedding model
3. Fine-tuningBehavior or format that prompting cannot hold at your volumeTraining dataset, training runs, tuned-model evalsTuned model stops with its base; retrain only if a vendor still offers it

Before you choose, answer three questions for your own feature:

  1. Is the failure about facts or about behavior? Read twenty failing outputs and tag each one. If most are missing or stale facts, fine-tuning is the wrong tool.
  2. Did a measured prompt baseline fail? Without an eval score for the prompt-only version, you cannot show that retrieval or training improved anything.
  3. Can you rebuild it when the model goes away? Check whether your vendor still accepts new training jobs for your organization, and whether the base model has a shutdown date. If the answer is no, put that effort into prompts and retrieval.
Primary sources checked 2026-10-03

The decision table, the three questions, the ticket-router example and the suggested eval set sizes are this article's proposals. No training runs, retrieval benchmarks or cost measurements were performed for it.