LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1527 min read

Building RAG (Retrieval-Augmented Generation) Pipelines

Combine embeddings, a vector database, and an LLM provider SDK into a complete, working RAG pipeline from scratch.

Introduction

RAG has come up implicitly in nearly every lesson since lesson 10 — LlamaIndex's query engine, a vector database's .query() call, and embeddings all exist to support it. This lesson assembles all three pieces into one complete, explicit pipeline built from raw calls, so the pattern is fully visible end to end.

What You Will Learn
  • The four steps every RAG pipeline follows, regardless of which tools implement them.
  • How to build a complete RAG pipeline from embeddings, Chroma, and an LLM call.
  • How to think about chunking, and when RAG is the right choice over fine-tuning.

A Real-Life Analogy First

Think about the difference between a closed-book exam and an open-book exam. In a closed-book exam, a student answers purely from what they memorized beforehand — if they never studied a fact, they cannot know it, no matter how smart they are. In an open-book exam, the same student can flip to the exact right page of a textbook before answering, so even a fact they never memorized can be answered correctly and precisely, as long as it is somewhere in the book they are allowed to consult.

That Is Exactly What RAG Does

A plain LLM call is a closed-book exam — the model can only answer from what it learned during training, which does not include your company's internal documents or last week's data. RAG turns it into an open-book exam: before answering, the system looks up the most relevant "pages" (retrieved chunks) from your own documents and hands them to the model alongside the question, so it can answer accurately about things it was never trained on.

The Four Steps of RAG

StepWhat HappensTool From Earlier LessonsOpen-Book Exam Analogy
1. IndexSplit source documents into chunks and embed each oneEmbeddings (lesson 14)Organizing the textbook into labeled sections
2. StoreSave the chunks and their embeddings for fast lookupA vector database (lessons 12–13)Putting the textbook on the desk, ready to consult
3. RetrieveEmbed the incoming question and find the most similar chunksSame vector databaseFlipping to the exact right page for this question
4. GeneratePass the question plus retrieved chunks to an LLM to produce a grounded answerA provider SDK (lesson 5)Writing the answer, using only what that page says

A Complete RAG Pipeline in Action

Use case: this example builds all four steps explicitly with Chroma and the OpenAI SDK — the same pattern LlamaIndex's query_engine.query() automated behind the scenes in lesson 10.

import chromadb
from openai import OpenAI
client_openai = OpenAI()
client_chroma = chromadb.PersistentClient(path="./chroma-db")
collection = client_chroma.get_or_create_collection("support-docs")
# Step 1 & 2: Index and store
collection.add(
ids=["doc-1", "doc-2"],
documents=[
"Refunds are processed within 5-7 business days of approval.",
"Shipping takes 3-5 business days for domestic orders.",
],
)
def answer_question(question: str) -> str:
# Step 3: Retrieve the most relevant chunk(s)
results = collection.query(query_texts=[question], n_results=2)
context = "\n".join(results["documents"][0])
# Step 4: Generate a grounded answer
response = client_openai.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Answer only using the provided context. If the answer isn't in the context, say you don't know."},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
],
)
return response.choices[0].message.content
print(answer_question("How long until I get my refund?"))
Terminal Output

Click Run to see what this code prints.

The System Prompt Is the Guardrail

"Answer only using the provided context" is what keeps the model from making up an answer when the retrieved chunks do not actually contain one — the equivalent of instructing a student "if it isn't in the book, write 'not covered' instead of guessing" — a critical guardrail against hallucination in any RAG system.

Chunking Strategy

The two example documents above happened to be short enough to embed whole. Real documents need to be split into smaller chunks first (a few hundred words each, often with some overlap between consecutive chunks) — too large a chunk dilutes its embedding's specificity, while too small a chunk loses surrounding context the model would need to answer well. This is similar to how a well-organized textbook is split into digestible sections rather than one giant unbroken wall of text — a section that is too broad forces you to read too much to find the answer, while a section that is too narrow might cut off half of the explanation you actually needed.

RAG vs Fine-Tuning

A common question is whether to use RAG or fine-tune a model (lesson 17) to "teach" it your data. RAG is almost always the right starting point: it is faster to set up, keeps data current without retraining, and lets you cite exactly which source grounded an answer — fine-tuning is better reserved for teaching a model a specific style, format, or narrow skill rather than injecting knowledge. Continuing the exam analogy: RAG is giving a student an open book; fine-tuning is more like giving a student months of extra practice tests on one narrow topic so their instincts improve — the two solve genuinely different problems.

Common Mistakes

Avoid These Mistakes
  • Skipping a "answer only from context" instruction, which leaves the model free to hallucinate when retrieval comes back empty or irrelevant.
  • Retrieving too few chunks (missing relevant context) or too many (drowning the model in irrelevant text and wasting tokens).
  • Reaching for fine-tuning to "add knowledge" when RAG solves that problem more directly and stays current as source documents change.

Best Practices

  • Always instruct the model explicitly to answer only from the provided context, and to say when it does not know.
  • Tune chunk size and the number of retrieved chunks (top_k / n_results) empirically against your actual documents, not by guessing.
  • Show retrieved sources alongside the answer in your UI so users can verify where an answer came from.

Frequently Asked Questions

No — this lesson shows RAG built from raw calls, which is worth understanding even if you use a framework's higher-level abstraction in practice.

Re-run the indexing step whenever your source documents change, or set up a scheduled job to re-sync the vector database periodically.

No — it substantially reduces it by grounding answers in real text, but a model can still misread or misstate what the retrieved context says, so evaluation (lesson 16) matters even with RAG in place.

Yes. RAG is the single most common real-world pattern in generative AI applications, and it directly reuses every tool from lessons 5, 12/13, and 14 — if any earlier lesson still felt abstract, seeing all of them work together here, solving one clear "open book exam" problem, is often the point where it fully clicks.

Key Takeaways

  • RAG has four steps: index, store, retrieve, and generate — like organizing a textbook, then flipping to the right page for an open-book exam.
  • A complete RAG pipeline can be built from just an embedding call, a vector database, and a provider SDK call.
  • Chunk size and count directly affect answer quality — tune both against real documents.
  • RAG is the default choice for grounding a model in your own data; fine-tuning solves a different problem.

Summary

RAG is the single most common pattern in real generative AI applications, and now that you have built one from scratch, every framework's higher-level RAG abstraction (LlamaIndex's query engine, LangChain's retrievers) will make a lot more sense.

Lesson 15 Completed
  • You know the four steps every RAG pipeline follows.
  • You built a complete RAG pipeline from raw embeddings, Chroma, and an LLM call.
  • You understand when to choose RAG over fine-tuning.
Next Lesson →

Prompt Engineering & Evaluation Tools