Embeddings & Semantic Search
Understand what embeddings actually are, generate them with the OpenAI embeddings API, and compare vectors with cosine similarity.
Introduction
Every vector database lesson so far has referenced "embeddings" without fully unpacking them. This lesson closes that gap: what an embedding actually is, how to generate one, and the math behind comparing two of them — the foundation semantic search and RAG are both built on.
- What an embedding is and why it captures meaning, not just keywords.
- How to generate embeddings with the OpenAI embeddings API.
- How cosine similarity measures how "close" two pieces of text are in meaning.
A Real-Life Analogy First
Think of GPS coordinates. "Eiffel Tower, Paris" and "Louvre Museum, Paris" are two completely different strings of text, sharing no words in common with a search for "famous Paris landmark" — but their GPS coordinates are numerically close to each other, because they are physically close in the real world. An embedding does the same trick for meaning instead of geography: it converts a piece of text into a set of "coordinates" in an imaginary space where texts with similar meaning end up numerically close together, regardless of which exact words they use.
What an Embedding Actually Is
An embedding model converts a piece of text into a fixed-length list of numbers (a vector) such that texts with similar meaning end up as nearby points in that numeric space — regardless of whether they share any of the same words. This is why a search for "how do I get my money back" can match a document about "refund policy" even though the two share no words in common, the same way "Eiffel Tower" and "Louvre" share no words but are close on a map.
Generating Embeddings in Action
Use case: the embeddings endpoint on the openai SDK (from lesson 5) converts text into a vector — the same SDK you already installed for chat completions.
from openai import OpenAI
client = OpenAI()
response = client.embeddings.create( model="text-embedding-3-small", input="How do I get my money back?",)
vector = response.data[0].embeddingprint("Dimensions:", len(vector))print("First 5 values:", vector[:5])Click Run to see what this code prints.
Cosine Similarity: How "Closeness" Is Measured
Use case: cosine similarity measures the angle between two vectors rather than their raw distance, returning a score from -1 (opposite meaning) to 1 (identical meaning) — the standard way vector databases rank how relevant a stored chunk is to a query. If GPS coordinates measure "how many miles apart," cosine similarity measures "how aligned in direction" — for embeddings, direction turns out to capture meaning better than raw distance does.
import numpy as np
def cosine_similarity(a, b): a, b = np.array(a), np.array(b) return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
refund_vector = client.embeddings.create(model="text-embedding-3-small", input="Refund policy").data[0].embeddingquestion_vector = client.embeddings.create(model="text-embedding-3-small", input="How do I get my money back?").data[0].embeddingunrelated_vector = client.embeddings.create(model="text-embedding-3-small", input="Weather forecast for tomorrow").data[0].embedding
print("Refund vs money-back question:", cosine_similarity(refund_vector, question_vector))print("Refund vs unrelated text:", cosine_similarity(refund_vector, unrelated_vector))Click Run to see what this code prints.
Every .query() call in lessons 12 and 13 is doing exactly this cosine similarity comparison internally, just optimized to run across millions of stored vectors instead of a handful compared by hand.
Choosing an Embedding Model
| Model | Provider | Notes |
|---|---|---|
| text-embedding-3-small | OpenAI | 1536 dimensions, low cost, strong general-purpose default |
| text-embedding-3-large | OpenAI | Higher dimensionality, better accuracy, higher cost |
| voyage-3 | Voyage AI | Popular alternative, competitive on retrieval benchmarks |
| all-MiniLM-L6-v2 | Open-source (via Hugging Face, lesson 8) | Runs locally, no API cost, smaller and less accurate |
Common Mistakes
- Comparing embeddings generated by two different models — similarity scores are only meaningful within the same model's vector space.
- Re-embedding the same unchanged text repeatedly instead of caching the result, wasting API calls.
- Assuming a high similarity score guarantees factual relevance — embeddings capture semantic closeness, not truth.
Best Practices
- Pick one embedding model per project and use it consistently for everything you store and query.
- Cache embeddings for content that does not change, rather than regenerating them on every request.
- Chunk long documents into smaller pieces before embedding — a single embedding for an entire book loses too much specific detail to be useful for search.
Frequently Asked Questions
No — most modern embedding models, including OpenAI's, are multilingual and can compare meaning across different languages to some degree.
Yes, with a multimodal embedding model — the concept (map content to a vector space by meaning) extends beyond text, covered further in lesson 22.
Not necessarily — it usually improves accuracy at the cost of more storage and slower search, so the right tradeoff depends on your project's scale and accuracy needs.
No — you never need to write that formula by hand in practice, since every vector database and framework in this course computes it internally for you. It is shown here once so "similarity search" stops being a mysterious black box and becomes a specific, understandable calculation.
Key Takeaways
- An embedding converts text into a vector where semantic similarity corresponds to numeric closeness — like GPS coordinates, but for meaning.
- The OpenAI embeddings API generates a vector for any text with a single call.
- Cosine similarity measures how close two vectors are, from -1 to 1.
- Every vector database's "similarity search" is this same comparison, automated at scale.
Summary
Embeddings and cosine similarity are the mathematical foundation underneath every vector database and RAG pipeline in this course — understanding them makes every tool built on top far less like a black box.
- You understand what an embedding captures and why it enables semantic search.
- You can generate embeddings with the OpenAI API.
- You understand how cosine similarity ranks relevance.