LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1424 min read

Embeddings & Semantic Search

Understand what embeddings actually are, generate them with the OpenAI embeddings API, and compare vectors with cosine similarity.

Introduction

Every vector database lesson so far has referenced "embeddings" without fully unpacking them. This lesson closes that gap: what an embedding actually is, how to generate one, and the math behind comparing two of them — the foundation semantic search and RAG are both built on.

What You Will Learn
  • What an embedding is and why it captures meaning, not just keywords.
  • How to generate embeddings with the OpenAI embeddings API.
  • How cosine similarity measures how "close" two pieces of text are in meaning.

A Real-Life Analogy First

Think of GPS coordinates. "Eiffel Tower, Paris" and "Louvre Museum, Paris" are two completely different strings of text, sharing no words in common with a search for "famous Paris landmark" — but their GPS coordinates are numerically close to each other, because they are physically close in the real world. An embedding does the same trick for meaning instead of geography: it converts a piece of text into a set of "coordinates" in an imaginary space where texts with similar meaning end up numerically close together, regardless of which exact words they use.

What an Embedding Actually Is

An embedding model converts a piece of text into a fixed-length list of numbers (a vector) such that texts with similar meaning end up as nearby points in that numeric space — regardless of whether they share any of the same words. This is why a search for "how do I get my money back" can match a document about "refund policy" even though the two share no words in common, the same way "Eiffel Tower" and "Louvre" share no words but are close on a map.

Generating Embeddings in Action

Use case: the embeddings endpoint on the openai SDK (from lesson 5) converts text into a vector — the same SDK you already installed for chat completions.

from openai import OpenAI
client = OpenAI()
response = client.embeddings.create(
model="text-embedding-3-small",
input="How do I get my money back?",
)
vector = response.data[0].embedding
print("Dimensions:", len(vector))
print("First 5 values:", vector[:5])
Terminal Output

Click Run to see what this code prints.

Cosine Similarity: How "Closeness" Is Measured

Use case: cosine similarity measures the angle between two vectors rather than their raw distance, returning a score from -1 (opposite meaning) to 1 (identical meaning) — the standard way vector databases rank how relevant a stored chunk is to a query. If GPS coordinates measure "how many miles apart," cosine similarity measures "how aligned in direction" — for embeddings, direction turns out to capture meaning better than raw distance does.

import numpy as np
def cosine_similarity(a, b):
a, b = np.array(a), np.array(b)
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
refund_vector = client.embeddings.create(model="text-embedding-3-small", input="Refund policy").data[0].embedding
question_vector = client.embeddings.create(model="text-embedding-3-small", input="How do I get my money back?").data[0].embedding
unrelated_vector = client.embeddings.create(model="text-embedding-3-small", input="Weather forecast for tomorrow").data[0].embedding
print("Refund vs money-back question:", cosine_similarity(refund_vector, question_vector))
print("Refund vs unrelated text:", cosine_similarity(refund_vector, unrelated_vector))
Terminal Output

Click Run to see what this code prints.

This Is What a Vector Database Automates

Every .query() call in lessons 12 and 13 is doing exactly this cosine similarity comparison internally, just optimized to run across millions of stored vectors instead of a handful compared by hand.

Choosing an Embedding Model

ModelProviderNotes
text-embedding-3-smallOpenAI1536 dimensions, low cost, strong general-purpose default
text-embedding-3-largeOpenAIHigher dimensionality, better accuracy, higher cost
voyage-3Voyage AIPopular alternative, competitive on retrieval benchmarks
all-MiniLM-L6-v2Open-source (via Hugging Face, lesson 8)Runs locally, no API cost, smaller and less accurate

Common Mistakes

Avoid These Mistakes
  • Comparing embeddings generated by two different models — similarity scores are only meaningful within the same model's vector space.
  • Re-embedding the same unchanged text repeatedly instead of caching the result, wasting API calls.
  • Assuming a high similarity score guarantees factual relevance — embeddings capture semantic closeness, not truth.

Best Practices

  • Pick one embedding model per project and use it consistently for everything you store and query.
  • Cache embeddings for content that does not change, rather than regenerating them on every request.
  • Chunk long documents into smaller pieces before embedding — a single embedding for an entire book loses too much specific detail to be useful for search.

Frequently Asked Questions

No — most modern embedding models, including OpenAI's, are multilingual and can compare meaning across different languages to some degree.

Yes, with a multimodal embedding model — the concept (map content to a vector space by meaning) extends beyond text, covered further in lesson 22.

Not necessarily — it usually improves accuracy at the cost of more storage and slower search, so the right tradeoff depends on your project's scale and accuracy needs.

No — you never need to write that formula by hand in practice, since every vector database and framework in this course computes it internally for you. It is shown here once so "similarity search" stops being a mysterious black box and becomes a specific, understandable calculation.

Key Takeaways

  • An embedding converts text into a vector where semantic similarity corresponds to numeric closeness — like GPS coordinates, but for meaning.
  • The OpenAI embeddings API generates a vector for any text with a single call.
  • Cosine similarity measures how close two vectors are, from -1 to 1.
  • Every vector database's "similarity search" is this same comparison, automated at scale.

Summary

Embeddings and cosine similarity are the mathematical foundation underneath every vector database and RAG pipeline in this course — understanding them makes every tool built on top far less like a black box.

Lesson 14 Completed
  • You understand what an embedding captures and why it enables semantic search.
  • You can generate embeddings with the OpenAI API.
  • You understand how cosine similarity ranks relevance.
Next Lesson →

Building RAG (Retrieval-Augmented Generation) Pipelines