RAG vs Fine-Tuning: Which One Should You Use?

A plain-English way to choose between RAG and fine-tuning. What each one actually changes, a three-question test, a worked example, and the mistake most teams make first.

People ask "should we use RAG or fine-tune?" as if one of them wins. They change different things, so the useful question is which problem you have.

RAG changes what the model sees. Fine-tuning changes the model itself. Once that is clear, most of the decision follows.

What each one actually does

RAG (retrieval augmented generation) looks things up before the model answers. When a question arrives, the system finds the passages in your documents that matter and puts them in the prompt. The model is not changed. Edit a document, and the next answer changes.

Fine-tuning trains the model further on examples of the inputs and outputs you want. This changes the model's weights, and you end up with a new version of the model. Anything you want it to learn has to be in the training data, and learning something new means training again.

The short answer

  • Facts go in RAG. Your refund policy, product docs, prices, last week's meeting notes, anything that changes or that only you have.
  • Behaviour goes in fine-tuning. A tone of voice, an output format you need every time, a narrow task you want a smaller, cheaper model to do consistently.
  • If you need both, use both. They do not compete.

Why fine-tuning is a weak way to add facts

It looks attractive, because "train it on our data" sounds like the obvious route. In practice it has four problems.

  1. Updates are slow. A new fact means new training data and a new training run. With RAG you add or edit a document and re-index it.
  2. No source. A fine-tuned model cannot show where it learned something. A RAG system can show the passages it used, so a person can check the answer.
  3. No guarantee of recall. Training on a fact does not guarantee the model will state it correctly when asked a different way. It can still answer with confidence and be wrong.
  4. No per-user permissions. With RAG you control which documents a given user can retrieve. Knowledge baked into weights is available to everyone who can use the model.

There is a fair case on the other side. Some companies fine-tune a small model on private data, mostly so the data never leaves their own infrastructure. If your knowledge is static and private, that is reasonable. If it changes often, as support tickets, news, docs and prices do, RAG is the better default.

Why RAG is not a cure-all

RAG has its own ways to fail:

  • It only works if the right passage is found. Chunking, embeddings and the wording of the question all affect that, so many RAG bugs start in retrieval.
  • The model can be handed the right passage and ignore it.
  • It adds work to every question: a search, and a longer prompt.
  • It is a poor tool for changing how the model writes. Tone and format are a job for the prompt or for fine-tuning.

A worked example: a support assistant

Say you are building an assistant for a software product.

  • The facts are pricing, plan limits and setup steps. They change, and customers will want to know where an answer came from. That is RAG over your help articles.
  • The behaviour is "answer in three short steps, stay friendly, never promise a refund, and return JSON that the ticket system can read". Try a clear prompt with a few examples first. If the format is still inconsistent across thousands of real questions, a fine-tuned small model is a good fit for that part.

Notice the order. The facts and the behaviour are separate problems and get separate tools.

Three questions to decide

  1. Does the answer depend on information that changes, or that only you have? Use RAG.
  2. Does the model answer in the wrong style or format even when it has the facts? Improve the prompt and add examples. If that still fails at scale, fine-tune.
  3. Do you need to show where an answer came from? Use RAG, because a fine-tuned model has no source to show.

Start with the cheapest thing

Each option costs more to build and to keep running than the one before it:

  1. A better prompt, with a few examples.
  2. RAG.
  3. Fine-tuning.

It is tempting to jump to fine-tuning first because it sounds more serious. Often a clear prompt and a working retrieval step are enough, and you keep the ability to change your answers by editing a document.

Where to go from here

The LogicWiz GenAI course covers this in the lesson Why Nova needs RAG, which includes a section on why not just retrain a model on your own data, and then goes through embeddings, chunking and retrieval. If you want to see the retrieval side in code, the RAG tutorial for beginners builds it from scratch in plain Python. The whole course is completely free, with no card. Chapters one to three open without an account, and from chapter four a free account keeps you going.