RAG Tutorial for Beginners: Build a Retrieval Pipeline in Python

What RAG actually is, explained in plain English, then built in about 40 lines of Python with no framework. Chunk, embed, retrieve, augment, generate, and the three places it quietly goes wrong.

RAG stands for retrieval augmented generation. Strip the jargon and it is one idea: the model does not know your documents, so before you ask it a question, you look up the relevant passage yourself and paste it into the prompt. Retrieve, augment, generate.

That is the whole trick. Everything else in a RAG pipeline (embeddings, chunking, vector databases, rerankers) exists to make the "look up the relevant passage" step work on thousands of pages instead of one.

This tutorial builds the pipeline from scratch in Python. No LangChain, no vector database, one file. By the end you will know what those tools are doing when you do reach for them.

Why the model needs help in the first place

A language model answers from what it saw during training. Ask it about your company's refund policy, last week's meeting notes, or a PDF you were sent this morning, and it has two options: admit it does not know, or produce something confident and wrong. Most models pick the second.

You could paste the whole document into every prompt. That works until the document is longer than the context window, or you have five hundred documents, or you are paying per token and the document is 40 pages. RAG is the fix: send only the few paragraphs that matter for this question.

The pipeline in one picture

your documents
   │
   ▼
1. chunk      split each document into passages of a few hundred words
   │
   ▼
2. embed      turn each passage into a list of numbers that captures its meaning
   │
   ▼
3. index      store the passages and their numbers so you can search them
   │
   ▼            user question
4. retrieve   embed the question, find the passages whose numbers are closest
   │
   ▼
5. augment    put those passages into the prompt, above the question
   │
   ▼
6. generate   the model answers using the passages it was just shown

Steps 1 to 3 happen once, when you load the documents. Steps 4 to 6 happen on every question.

Step 1: chunk the documents

A passage has to be small enough that a handful of them fit in the prompt, and large enough to still make sense on its own. A few hundred words is the usual compromise.

def chunk(text: str, size: int = 300, overlap: int = 50) -> list[str]:
    words = text.split()
    chunks = []
    start = 0
    while start < len(words):
        chunks.append(" ".join(words[start:start + size]))
        start += size - overlap
    return chunks

The overlap matters. Without it, a sentence that straddles a boundary is cut in half and neither chunk contains the full thought.

Step 2: embed each chunk

An embedding is a list of a few hundred or a few thousand numbers. Passages that mean similar things get lists that point in similar directions. That is what lets you search by meaning rather than by exact words.

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from the environment

def embed(texts: list[str]) -> list[list[float]]:
    response = client.embeddings.create(
        model="text-embedding-3-small",
        input=texts,
    )
    return [item.embedding for item in response.data]

Step 3: index

For a tutorial, the index is a Python list. A vector database is this list plus fast search over millions of rows. You do not need one until the list gets slow.

documents = {
    "refunds.md": open("refunds.md").read(),
    "shipping.md": open("shipping.md").read(),
}

index = []  # each entry: (source, chunk_text, embedding)
for source, text in documents.items():
    chunks = chunk(text)
    for chunk_text, vector in zip(chunks, embed(chunks)):
        index.append((source, chunk_text, vector))

Step 4: retrieve

Embed the question with the same model, then find the chunks whose vectors point the same way. Cosine similarity is the standard measure.

import math

def cosine(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x * x for x in a))
    norm_b = math.sqrt(sum(y * y for y in b))
    return dot / (norm_a * norm_b)

def retrieve(question: str, k: int = 3) -> list[tuple[str, str]]:
    q_vector = embed([question])[0]
    scored = [
        (cosine(q_vector, vector), source, chunk_text)
        for source, chunk_text, vector in index
    ]
    scored.sort(reverse=True)
    return [(source, chunk_text) for _, source, chunk_text in scored[:k]]

Steps 5 and 6: augment and generate

Put the retrieved passages in the prompt, tell the model to answer from them, and ask.

def answer(question: str) -> str:
    passages = retrieve(question)
    context = "\n\n".join(f"[{source}]\n{text}" for source, text in passages)
    prompt = (
        "Answer the question using only the passages below. "
        "If the passages do not contain the answer, say so.\n\n"
        f"{context}\n\nQuestion: {question}"
    )
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    )
    return response.choices[0].message.content

print(answer("How many days do customers have to request a refund?"))

That is a working RAG pipeline. Run it against two small markdown files and you will see the model quote your refund policy instead of inventing one.

The three places it goes wrong

Every RAG bug in production is one of these.

The right passage was never retrieved. The question used different words from the document ("money back" versus "refund"), or the chunk boundary split the answer, or the answer needed two passages and you only fetched one. Print the retrieved passages before you blame the model. Nine times out of ten the model never saw the answer.

The right passage was retrieved and the model ignored it. Usually the prompt is too soft. "Use the passages below" is a suggestion; "answer only from the passages, and say 'not in the documents' otherwise" is an instruction. Test with a question the documents cannot answer and check that the model says so.

The passages are stale. Documents changed and the index did not. The index is a cache, and it needs the same care as any cache: rebuild it when the source changes.

Where to go from here

Once the basic loop works, the upgrades are all in the retrieve step: hybrid search (keyword match plus embeddings), reranking the top 20 down to the best 3 with a second model, and metadata filters so a question about shipping never retrieves a refund passage. Each one fixes a specific failure you will hit, so add them when you hit it, not before.

The LogicWiz GenAI course covers this chapter with an animation you can step through and a lab that runs in the browser: Why Nova needs RAG, embeddings and semantic search, chunking, and indexing and retrieval. The whole course is completely free, with no card. Chapters one to three open without an account, and from chapter four a free account keeps you going.