LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 2917 min read

Introduction to the Aggregation Framework

Understand the aggregation pipeline model — a sequence of stages that progressively transform documents.

What Is the Aggregation Framework?

The aggregation framework is MongoDB's tool for complex data transformation, analysis, and reporting — think "group sales by month," "compute an average rating per product," or "join and reshape data from two collections." find() handles filtering and simple queries; aggregation handles genuinely computed, multi-step results.

The Pipeline Model

An aggregation is expressed as a pipeline — an array of stages, each one transforming the documents flowing through it and passing the result to the next stage, much like a Unix pipeline of chained commands.

Documents enter the pipeline
↓
Stage 1: $match (filter)
↓
Stage 2: $group (summarize)
↓
Stage 3: $sort (order)
↓
Final transformed result

Your First Aggregation

db.orders.aggregate([
{ $match: { status: "completed" } },
{ $group: { _id: "$customerId", totalSpent: { $sum: "$total" } } },
{ $sort: { totalSpent: -1 } },
]);

This pipeline: filters to completed orders, groups them by customer while summing their totals, then sorts customers by how much they spent, highest first.

aggregate() vs find()

find()aggregate()
PurposeFilter and retrieve documents as-isTransform, group, and compute across documents
Reshaping dataLimited to projectionCan compute entirely new fields and structures
Combining collectionsNot supported directlySupported via $lookup
ComplexitySimple, fast to writeMore powerful, more to learn
Reach for aggregate() When find() Isn't Enough

If you find yourself wanting to compute a sum, average, or count grouped by some field — or combine data from two collections — that's exactly the signal to reach for the aggregation framework.

FAQs

Very much so conceptually — $group plays a similar role to SQL's GROUP BY, and the overall pipeline model covers what SQL expresses through a combination of WHERE, GROUP BY, HAVING, and JOIN.

Entirely on the server — this is a major performance advantage, since only the final, transformed result is sent back, not the raw underlying documents.

Summary

The aggregation framework processes documents through a pipeline of stages, each transforming the data before passing it to the next. Next, you'll go deep on the two most fundamental stages: $match and $group.

Next Lesson →

$match and $group