LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 2019 min read

Embedding vs Referencing

The single most important MongoDB modeling decision: when to nest related data inside a document, and when to reference it separately.

Two Ways to Model a Relationship

MongoDB gives you two fundamental tools for representing that one piece of data relates to another: embedding it directly inside the parent document, or referencing it by storing the related document's _id and looking it up separately. Nearly every schema design decision comes down to choosing between these two.

Embedding

Embedding nests the related data directly as a sub-document or array within the parent.

{
"_id": "order1",
"customerName": "Ada Lovelace",
"items": [
{ "product": "Keyboard", "price": 45, "quantity": 1 },
{ "product": "Mouse", "price": 20, "quantity": 2 }
]
}

Embedding's Strengths

  • One document read retrieves everything
  • Naturally atomic — one document, one write
  • No joins needed at query time

Embedding's Weaknesses

  • Can lead to duplicated data if the same info appears elsewhere
  • Documents can grow toward the 16MB limit for unbounded arrays
  • Harder to query the embedded data independently of its parent

Referencing

Referencing stores just the related document's _id, keeping the two pieces of data in separate collections — closer to how a foreign key works in a relational database.

// users collection
{ "_id": "u1", "name": "Ada Lovelace" }
// orders collection
{ "_id": "order1", "userId": "u1", "items": [ ... ] }

Referencing's Strengths

  • No duplication — one source of truth per entity
  • Related data can grow unboundedly without bloating the parent
  • Related data is independently queryable

Referencing's Weaknesses

  • Requires a second query (or a $lookup) to fetch related data
  • No automatic referential integrity — a deleted user leaves orphaned orderIds
  • Slightly more application logic to keep things consistent

Decision Factors

QuestionFavors EmbeddingFavors Referencing
Is the related data always read together?YesNo
Does the related data grow unboundedly?No (bounded, small)Yes (could be thousands of items)
Is the related data queried independently?RarelyOften
Is the related data shared across many parents?NoYes (e.g. a shared "category" document)

A Hybrid Approach

A common, pragmatic middle ground: reference the full related document, but also embed a small, denormalized summary of it for fast, join-free display in common cases.

{
"_id": "order1",
"customer": { "id": "u1", "name": "Ada Lovelace" }, // enough to display without a lookup
"items": [ ... ]
}
This Trades Consistency for Speed

If the customer's name changes, this embedded copy becomes stale until explicitly updated — a deliberate, common tradeoff for read-heavy, display-focused fields that rarely change.

Common Beginner Mistakes

Embedding an unboundedly growing array

A "comments" array on a viral post, or "orders" embedded in a customer, can grow toward the 16MB document limit — reference collections that can grow without a natural bound.

Referencing everything out of pure SQL habit

This throws away one of MongoDB's biggest advantages — retrieving related, always-read-together data in a single document lookup.

FAQs

No single rule covers every case — the decision factors table above is a strong starting heuristic, but real-world modeling often involves genuine tradeoffs judged case by case.

Yes, but it requires a migration — writing a script to restructure existing documents and updating application code — so it's worth thinking carefully upfront, even though it's not literally impossible to change.

Summary

Embedding favors fast, atomic, single-document reads for tightly coupled, bounded data; referencing favors independence and unbounded growth. Next, you'll apply this directly to modeling one-to-many relationships.

Next Lesson →

Modeling One-to-Many Relationships