Introduction to the Aggregation Framework
Understand the aggregation pipeline model — a sequence of stages that progressively transform documents.
What Is the Aggregation Framework?
The aggregation framework is MongoDB's tool for complex data transformation, analysis, and reporting — think "group sales by month," "compute an average rating per product," or "join and reshape data from two collections." find() handles filtering and simple queries; aggregation handles genuinely computed, multi-step results.
The Pipeline Model
An aggregation is expressed as a pipeline — an array of stages, each one transforming the documents flowing through it and passing the result to the next stage, much like a Unix pipeline of chained commands.
Your First Aggregation
db.orders.aggregate([ { $match: { status: "completed" } }, { $group: { _id: "$customerId", totalSpent: { $sum: "$total" } } }, { $sort: { totalSpent: -1 } },]);This pipeline: filters to completed orders, groups them by customer while summing their totals, then sorts customers by how much they spent, highest first.
aggregate() vs find()
| find() | aggregate() | |
|---|---|---|
| Purpose | Filter and retrieve documents as-is | Transform, group, and compute across documents |
| Reshaping data | Limited to projection | Can compute entirely new fields and structures |
| Combining collections | Not supported directly | Supported via $lookup |
| Complexity | Simple, fast to write | More powerful, more to learn |
If you find yourself wanting to compute a sum, average, or count grouped by some field — or combine data from two collections — that's exactly the signal to reach for the aggregation framework.
FAQs
Very much so conceptually — $group plays a similar role to SQL's GROUP BY, and the overall pipeline model covers what SQL expresses through a combination of WHERE, GROUP BY, HAVING, and JOIN.
Entirely on the server — this is a major performance advantage, since only the final, transformed result is sent back, not the raw underlying documents.
Summary
The aggregation framework processes documents through a pipeline of stages, each transforming the data before passing it to the next. Next, you'll go deep on the two most fundamental stages: $match and $group.