Prompt Orchestration Frameworks: LlamaIndex
Learn LlamaIndex — purpose-built for connecting LLMs to your own data — with a working document-indexing example.
Introduction
While LangChain is a general-purpose orchestration toolkit, LlamaIndex was built specifically around one problem: connecting an LLM to your own data. Loading documents, splitting them, indexing them, and querying that index is LlamaIndex's core loop, and it does that loop with less setup code than a general-purpose framework.
- What problem LlamaIndex specifically optimizes for.
- How to index a small set of documents and query them.
- When to reach for LlamaIndex instead of LangChain.
A Real-Life Analogy First
Imagine a brand-new librarian who, on their very first day, is handed every single book in a library and asked a question. They'd have to search shelf by shelf. Now imagine a librarian who has already read every book cover to cover, has a mental index of which page covers which topic, and can walk straight to the right paragraph the moment you ask a question. LlamaIndex is the tooling that turns "a pile of documents" into that second, already-prepared librarian — automatically, without you personally reading and memorizing anything.
What Problem LlamaIndex Solves
Use case: LlamaIndex's VectorStoreIndex handles loading raw documents, splitting them into chunks, generating embeddings (lesson 14), storing them, and answering natural-language questions against them — the entire "chat with your documents" workflow — in a handful of lines rather than assembling each piece by hand.
pip install llama-indexLlamaIndex in Action
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Loads every file in the ./docs folder (txt, pdf, md, and more)documents = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(documents)query_engine = index.as_query_engine()
response = query_engine.query("What is our refund policy?")print(response)Click Run to see what this code prints.
from_documents() split each file into chunks, embedded each chunk (lesson 14), and stored the result in an in-memory vector index — like the librarian speed-reading every book once and jotting a note about what each page covers. as_query_engine().query() then embedded the question, found the most relevant chunks, and passed them to an LLM alongside the question — the full RAG pattern covered in depth in lesson 15.
LangChain vs LlamaIndex
| Aspect | LangChain | LlamaIndex |
|---|---|---|
| Core focus | General-purpose orchestration — chains, agents, tools | Data indexing and retrieval — "chat with your data" |
| Setup for a basic RAG query | More explicit steps, more control | Fewer lines, more built-in defaults |
| Best for | Complex, custom multi-step pipelines | Document Q&A and retrieval-heavy applications |
| Can they be combined? | Yes — many projects use LlamaIndex for retrieval inside a LangChain chain | Yes — same as left |
Common Mistakes
- Rebuilding the vector index from scratch on every query instead of persisting it to disk with index.storage_context.persist().
- Loading an entire large document collection into memory during development instead of testing against a small sample first.
- Treating LangChain and LlamaIndex as mutually exclusive when many real projects use both together.
Best Practices
- Reach for LlamaIndex first when the core problem is "answer questions from documents," not a complex multi-tool workflow.
- Persist an index to disk (or a real vector database, lessons 12–13) once you move past local experimentation.
- Inspect response.source_nodes on a query result to see exactly which chunks were used to ground an answer.
Frequently Asked Questions
No — many production RAG systems use LlamaIndex specifically for the retrieval layer and LangChain (or raw SDK calls) for the surrounding application logic.
By default the example above is in-memory only. For anything beyond a demo, persist the index to disk or connect it to a dedicated vector database (lessons 12–13).
Plain text, Markdown, PDF, Word documents, and several other common formats out of the box, with additional loaders available for more specialized sources.
Yes, essentially — many of those tools are built on exactly this pattern: load a document, index it, and query it with an LLM grounding its answer in the retrieved text. You have now seen the core mechanism behind a whole category of products you may have already used.
Key Takeaways
- LlamaIndex specializes in connecting an LLM to your own documents with minimal setup code.
- VectorStoreIndex.from_documents() and as_query_engine() cover the full load-embed-retrieve-answer loop.
- LlamaIndex and LangChain solve overlapping but distinct problems, and are often combined.
- A query response's source_nodes show exactly which chunks grounded the answer.
Summary
LlamaIndex is the fastest path from "a folder of documents" to "a working question-answering system," and pairs naturally with the vector databases and RAG concepts covered in the next few lessons.
- You understand what problem LlamaIndex optimizes for.
- You can index a set of documents and query them.
- You know when to reach for LlamaIndex versus LangChain.