RAG Tutorial for Beginners: Build a Retrieval Pipeline in Python
What RAG actually is, explained in plain English, then built in about 40 lines of Python with no framework. Chunk, embed, retrieve, augment, generate, and the three places it quietly goes wrong.
RAG stands for retrieval augmented generation. Strip the jargon and it is one idea: the model does not know your documents, so before you ask it a question, you look up the relevant passage yourself and paste it into the prompt. Retrieve, augment, generate.
That is the whole trick. Everything else in a RAG pipeline (embeddings, chunking, vector databases, rerankers) exists to make the "look up the relevant passage" step work on thousands of pages instead of one.
This tutorial builds the pipeline from scratch in Python. No LangChain, no vector database, one file. By the end you will know what those tools are doing when you do reach for them.
Why the model needs help in the first place
A language model answers from what it saw during training. Ask it about your company's refund policy, last week's meeting notes, or a PDF you were sent this morning, and it has two options: admit it does not know, or produce something confident and wrong. Most models pick the second.
You could paste the whole document into every prompt. That works until the document is longer than the context window, or you have five hundred documents, or you are paying per token and the document is 40 pages. RAG is the fix: send only the few paragraphs that matter for this question.
The pipeline in one picture
your documents
│
▼
1. chunk split each document into passages of a few hundred words
│
▼
2. embed turn each passage into a list of numbers that captures its meaning
│
▼
3. index store the passages and their numbers so you can search them
│
▼ user question
4. retrieve embed the question, find the passages whose numbers are closest
│
▼
5. augment put those passages into the prompt, above the question
│
▼
6. generate the model answers using the passages it was just shown
Steps 1 to 3 happen once, when you load the documents. Steps 4 to 6 happen on every question.
Step 1: chunk the documents
A passage has to be small enough that a handful of them fit in the prompt, and large enough to still make sense on its own. A few hundred words is the usual compromise.
def chunk(text: str, size: int = 300, overlap: int = 50) -> list[str]:
words = text.split()
chunks = []
start = 0
while start < len(words):
chunks.append(" ".join(words[start:start + size]))
start += size - overlap
return chunks
The overlap matters. Without it, a sentence that straddles a boundary is cut in half and neither chunk contains the full thought.
Step 2: embed each chunk
An embedding is a list of a few hundred or a few thousand numbers. Passages that mean similar things get lists that point in similar directions. That is what lets you search by meaning rather than by exact words.
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
def embed(texts: list[str]) -> list[list[float]]:
response = client.embeddings.create(
model="text-embedding-3-small",
input=texts,
)
return [item.embedding for item in response.data]
Step 3: index
For a tutorial, the index is a Python list. A vector database is this list plus fast search over millions of rows. You do not need one until the list gets slow.
documents = {
"refunds.md": open("refunds.md").read(),
"shipping.md": open("shipping.md").read(),
}
index = [] # each entry: (source, chunk_text, embedding)
for source, text in documents.items():
chunks = chunk(text)
for chunk_text, vector in zip(chunks, embed(chunks)):
index.append((source, chunk_text, vector))
Step 4: retrieve
Embed the question with the same model, then find the chunks whose vectors point the same way. Cosine similarity is the standard measure.
import math
def cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
norm_a = math.sqrt(sum(x * x for x in a))
norm_b = math.sqrt(sum(y * y for y in b))
return dot / (norm_a * norm_b)
def retrieve(question: str, k: int = 3) -> list[tuple[str, str]]:
q_vector = embed([question])[0]
scored = [
(cosine(q_vector, vector), source, chunk_text)
for source, chunk_text, vector in index
]
scored.sort(reverse=True)
return [(source, chunk_text) for _, source, chunk_text in scored[:k]]
Steps 5 and 6: augment and generate
Put the retrieved passages in the prompt, tell the model to answer from them, and ask.
def answer(question: str) -> str:
passages = retrieve(question)
context = "\n\n".join(f"[{source}]\n{text}" for source, text in passages)
prompt = (
"Answer the question using only the passages below. "
"If the passages do not contain the answer, say so.\n\n"
f"{context}\n\nQuestion: {question}"
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
return response.choices[0].message.content
print(answer("How many days do customers have to request a refund?"))
That is a working RAG pipeline. Run it against two small markdown files and you will see the model quote your refund policy instead of inventing one.
The three places it goes wrong
Every RAG bug in production is one of these.
The right passage was never retrieved. The question used different words from the document ("money back" versus "refund"), or the chunk boundary split the answer, or the answer needed two passages and you only fetched one. Print the retrieved passages before you blame the model. Nine times out of ten the model never saw the answer.
The right passage was retrieved and the model ignored it. Usually the prompt is too soft. "Use the passages below" is a suggestion; "answer only from the passages, and say 'not in the documents' otherwise" is an instruction. Test with a question the documents cannot answer and check that the model says so.
The passages are stale. Documents changed and the index did not. The index is a cache, and it needs the same care as any cache: rebuild it when the source changes.
Where to go from here
Once the basic loop works, the upgrades are all in the retrieve step: hybrid search (keyword match plus embeddings), reranking the top 20 down to the best 3 with a second model, and metadata filters so a question about shipping never retrieves a refund passage. Each one fixes a specific failure you will hit, so add them when you hit it, not before.
The LogicWiz GenAI course covers this chapter with an animation you can step through and a lab that runs in the browser: Why Nova needs RAG, embeddings and semantic search, chunking, and indexing and retrieval. The whole course is completely free, with no card. Chapters one to three open without an account, and from chapter four a free account keeps you going.