Local Inference & Open-Weight Models
Run open-weight models entirely on your own hardware with Ollama and llama.cpp, with no API key or hosted service required.
Introduction
Every provider SDK so far has sent your data to someone else's server. Local inference tools remove that dependency entirely — an open-weight model, downloaded once, runs fully on your own machine, with no API key, no per-token cost, and no data leaving your device.
- Why a team might choose local inference over a hosted API.
- How to run a model locally with Ollama in just a few commands.
- What llama.cpp is, and how it relates to Ollama under the hood.
A Real-Life Analogy First
This lesson is the direct extension of lesson 1's restaurant analogy: every provider SDK so far has been ordering delivery. Local inference is cooking the meal yourself, at home. You need your own kitchen and ingredients (hardware and a downloaded model), the meal might not taste quite as refined as a top restaurant's (a smaller local model versus a huge hosted one), but nothing leaves your house, there's no delivery fee per meal, and you can cook at 2am with no restaurant open.
Why Run a Model Locally?
Data Privacy
Sensitive data never leaves your machine — important for regulated industries or confidential documents.
No Per-Token Cost
Once downloaded, a local model can be called unlimited times with no ongoing API bill.
Offline Availability
Local models keep working with no internet connection, useful for edge devices or unreliable networks.
Full Control
You control exactly which model version runs, with no risk of a provider silently updating or deprecating it.
Ollama: The Easiest On-Ramp
Use case: Ollama packages downloading, quantizing, and running an open-weight model behind a single command and a local HTTP API that mirrors the OpenAI SDK's shape — the simplest way to go from zero to a running local model.
# after installing Ollama from ollama.comollama pull llama3.2ollama run llama3.2 "Explain what an API is in one sentence."Click Run to see what this code prints.
Ollama in Action (from Python)
pip install ollamaimport ollama
response = ollama.chat( model="llama3.2", messages=[{"role": "user", "content": "Explain what an API is in one sentence."}],)
print(response["message"]["content"])Click Run to see what this code prints.
Notice the messages list is the same role-based shape from lesson 5 — most orchestration frameworks (LangChain, LlamaIndex) also have an Ollama integration, so you can swap between a hosted and a local model with minimal code changes.
llama.cpp: The Engine Underneath
Use case: llama.cpp is a highly optimized C/C++ library for running LLM inference on CPUs and consumer GPUs with minimal dependencies — Ollama actually uses llama.cpp internally for much of its heavy lifting. Reaching for llama.cpp directly makes sense when you need lower-level control (custom quantization settings, embedding it in another application) than Ollama's simpler interface exposes. If Ollama is a ready-to-use kitchen appliance, llama.cpp is the raw stovetop and gas line it is built on top of — more work to use directly, but more control if you need it.
pip install llama-cpp-pythonfrom llama_cpp import Llama
llm = Llama(model_path="./models/llama-3.2-3b.Q4_K_M.gguf")
output = llm("Explain what an API is in one sentence.", max_tokens=50)print(output["choices"][0]["text"])Click Run to see what this code prints.
Choosing Between Them
| Tool | Ease of Use | Best For |
|---|---|---|
| Ollama | Very high — one command to pull and run a model | Getting started quickly, local development, most projects |
| llama.cpp | Lower-level, more manual setup | Custom quantization, embedding inference directly into another app |
Common Mistakes
- Expecting a local model on modest hardware to match a large hosted model's quality — local models are typically smaller and less capable, a real tradeoff for their privacy and cost benefits.
- Downloading a model larger than your available RAM/VRAM comfortably supports, leading to extremely slow or failed inference.
- Assuming Ollama and llama.cpp are competing alternatives rather than layered — Ollama is often the friendlier interface built on top of llama.cpp.
Best Practices
- Start with Ollama for local development and prototyping — reach for llama.cpp directly only once you need control it doesn't expose.
- Check a model's parameter count and quantization level against your available RAM/VRAM before pulling it.
- Use local models for development and privacy-sensitive tasks, and hosted APIs when you need maximum capability — many teams use both.
Frequently Asked Questions
No — both Ollama and llama.cpp run on CPU-only machines, though a GPU significantly speeds up inference, especially for larger models.
Yes — both frameworks have dedicated Ollama integrations, so a local model can slot into any chain or query engine built in earlier lessons.
It refers to storing a model's weights at lower numeric precision (like 4-bit instead of 16-bit) to reduce memory and disk usage, at a small, usually acceptable cost to output quality.
The model itself and the software are free, but it is not truly "free" in a deeper sense — your own electricity and hardware do the work a provider's server would otherwise do, so there is still a real cost, just paid differently (upfront hardware and ongoing electricity) instead of per API call.
Key Takeaways
- Local inference trades some model capability for privacy, zero per-token cost, and offline availability.
- Ollama is the easiest way to download and run an open-weight model locally.
- llama.cpp is the optimized engine Ollama builds on, useful directly when you need lower-level control.
- Local and hosted models can be mixed within the same project depending on the task.
Summary
Local inference tools close the loop on the "own your own data" side of this course's toolkit — a real, practical alternative to hosted APIs whenever privacy, cost, or offline access matter more than squeezing out the absolute best model quality.
- You understand why teams choose local inference over hosted APIs.
- You can pull and run a model locally with Ollama.
- You know what llama.cpp is and how it relates to Ollama.