Embedding vs Referencing
The single most important MongoDB modeling decision: when to nest related data inside a document, and when to reference it separately.
Two Ways to Model a Relationship
MongoDB gives you two fundamental tools for representing that one piece of data relates to another: embedding it directly inside the parent document, or referencing it by storing the related document's _id and looking it up separately. Nearly every schema design decision comes down to choosing between these two.
Embedding
Embedding nests the related data directly as a sub-document or array within the parent.
{ "_id": "order1", "customerName": "Ada Lovelace", "items": [ { "product": "Keyboard", "price": 45, "quantity": 1 }, { "product": "Mouse", "price": 20, "quantity": 2 } ]}Embedding's Strengths
- One document read retrieves everything
- Naturally atomic — one document, one write
- No joins needed at query time
Embedding's Weaknesses
- Can lead to duplicated data if the same info appears elsewhere
- Documents can grow toward the 16MB limit for unbounded arrays
- Harder to query the embedded data independently of its parent
Referencing
Referencing stores just the related document's _id, keeping the two pieces of data in separate collections — closer to how a foreign key works in a relational database.
// users collection{ "_id": "u1", "name": "Ada Lovelace" }
// orders collection{ "_id": "order1", "userId": "u1", "items": [ ... ] }Referencing's Strengths
- No duplication — one source of truth per entity
- Related data can grow unboundedly without bloating the parent
- Related data is independently queryable
Referencing's Weaknesses
- Requires a second query (or a $lookup) to fetch related data
- No automatic referential integrity — a deleted user leaves orphaned orderIds
- Slightly more application logic to keep things consistent
Decision Factors
| Question | Favors Embedding | Favors Referencing |
|---|---|---|
| Is the related data always read together? | Yes | No |
| Does the related data grow unboundedly? | No (bounded, small) | Yes (could be thousands of items) |
| Is the related data queried independently? | Rarely | Often |
| Is the related data shared across many parents? | No | Yes (e.g. a shared "category" document) |
A Hybrid Approach
A common, pragmatic middle ground: reference the full related document, but also embed a small, denormalized summary of it for fast, join-free display in common cases.
{ "_id": "order1", "customer": { "id": "u1", "name": "Ada Lovelace" }, // enough to display without a lookup "items": [ ... ]}If the customer's name changes, this embedded copy becomes stale until explicitly updated — a deliberate, common tradeoff for read-heavy, display-focused fields that rarely change.
Common Beginner Mistakes
A "comments" array on a viral post, or "orders" embedded in a customer, can grow toward the 16MB document limit — reference collections that can grow without a natural bound.
This throws away one of MongoDB's biggest advantages — retrieving related, always-read-together data in a single document lookup.
FAQs
No single rule covers every case — the decision factors table above is a strong starting heuristic, but real-world modeling often involves genuine tradeoffs judged case by case.
Yes, but it requires a migration — writing a script to restructure existing documents and updating application code — so it's worth thinking carefully upfront, even though it's not literally impossible to change.
Summary
Embedding favors fast, atomic, single-document reads for tightly coupled, bounded data; referencing favors independence and unbounded growth. Next, you'll apply this directly to modeling one-to-many relationships.