LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 2624 min read

Real-World Project: Building a Data Science Dependency Stack

Bring the whole course together by assembling a realistic dependency stack for a churn-prediction pipeline, plus a decision framework for choosing libraries on any new project.

Introduction

Over the last 25 lessons, you have gone through dozens of Python packages one at a time — a library for loading data, a library for plotting it, a library for modeling it, a library for serving it. In a real project, none of those decisions happen in isolation. You have to look at what you are actually building and pick a coherent stack, not just a shopping list of every interesting tool this course covered.

This final lesson works through exactly that, using one realistic project as the example: a churn-prediction pipeline for a subscription business. By the end, you will have a complete requirements.txt for the project, and — more importantly — a repeatable way of deciding what belongs in requirements.txt on any future project.

What You Will Learn
  • How to walk through a realistic project and choose a dependency for each stage, using libraries from across this course.
  • What a complete, production-ready requirements.txt looks like for that project.
  • A decision-framework checklist for picking libraries on any new project.
  • How to recognize and avoid dependency bloat.

The Project: Churn Prediction

The scenario: a subscription business wants to predict which customers are likely to cancel next month, so the retention team can reach out first. The data lives in a database. Nobody on the team wants to run a Spark cluster for this — the dataset is a few hundred thousand rows, well within reach of a single machine.

Walking through the pipeline stage by stage, and pulling the right tool for each stage from earlier in this course:

StageLibraryWhy This One
Pull data from the databasesqlalchemy + pandasThe dataset fits on one machine — no need for PySpark.
Clean and reshape the datapandasThe standard, well-understood tool for tabular cleaning.
Explore and visualize patternsmatplotlib + seabornQuick statistical plots to understand churn drivers before modeling.
Train the modelscikit-learnA classic tabular classification problem — no need for a deep learning framework.
Track experimentsmlflowSelf-hosted tracking is enough; no need to pay for a hosted dashboard yet.
Serve the modelfastapiThe retention team's internal tools need a real API to call.
Keep secrets out of the codepython-dotenvThe database URL and any API keys should never be hardcoded.
Keep the codebase cleanblack + ruffTwo engineers are collaborating on this repo.

Notice what is missing: no TensorFlow or PyTorch, no Scrapy, no Airflow, no PySpark. Every one of those is a legitimate tool covered in this course — but none of them is needed for this specific project. That absence is the whole point of this lesson.

pipeline_sketch.py
import pandas as pd
from sqlalchemy import create_engine
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
import mlflow
import mlflow.sklearn
engine = create_engine("postgresql://user:pass@localhost/churn_db")
df = pd.read_sql("SELECT * FROM customers", engine)
df = df.dropna(subset=["tenure_months", "monthly_charges"])
X = df[["tenure_months", "monthly_charges", "support_tickets"]]
y = df["churned"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
with mlflow.start_run():
model = RandomForestClassifier(n_estimators=150, max_depth=8)
model.fit(X_train, y_train)
accuracy = accuracy_score(y_test, model.predict(X_test))
mlflow.log_metric("accuracy", accuracy)
mlflow.sklearn.log_model(model, "churn_model")
print(f"Model trained with accuracy: {accuracy:.3f}")
Output

Click Run to see what this code prints.

The Final requirements.txt

Pulling it all together, this is the complete, production-ready dependency list for the project — every line earns its place for a specific stage of the pipeline above.

requirements.txt
# Data access
sqlalchemy==2.0.30
psycopg2-binary==2.9.9
# Data manipulation
pandas==2.2.2
# Visualization
matplotlib==3.8.4
seaborn==0.13.2
# Modeling
scikit-learn==1.4.2
# Experiment tracking
mlflow==2.13.0
# Serving
fastapi==0.111.0
uvicorn==0.29.0
# Config & secrets
python-dotenv==1.0.1
# Dev tooling
black==24.4.2
ruff==0.4.4
Pin Your Versions

Notice every dependency is pinned to an exact version rather than left open-ended. This is what makes the pipeline reproducible — the same requirements.txt installs the same behavior on any machine, today or a year from now.

A Decision Framework for Any Project

The table above was not arbitrary — it came from asking the same small set of questions about every candidate library. Use this checklist the next time you are tempted to pip install something new.

  • Do I need this right now, or am I speculatively adding it for a feature that does not exist yet?
  • Is there a lighter-weight tool already in my stack that solves the same problem well enough?
  • Does this library solve a problem I actually have (data too large for memory, a real API to serve, an actual multi-page crawl) — or am I reaching for it because it is the "advanced" or exciting option?
  • What is the cost of being wrong — how hard is it to swap this library out later if the project outgrows it, or if it turns out to be overkill?
  • Does adding this pull in a chain of heavy transitive dependencies, slower installs, and a larger attack surface?
Applying the Framework

For this project: PySpark failed question 1 (the data fits on one machine right now) and question 3 (no real cluster-scale problem exists). Airflow failed question 1 for the same reason — a single daily retraining script does not yet need a scheduler with a UI. If the business scales to millions of customers and the pipeline grows to ten dependent steps, that answer would change, and the framework would say yes instead.

Avoiding Dependency Bloat

Dependency bloat is what happens when a project accumulates libraries that solve problems it does not actually have — usually because someone reached for the well-known, heavyweight tool instead of asking whether a lighter one would do. It slows down installs, increases the surface area for security vulnerabilities, and makes the codebase harder for new contributors to understand, since every dependency is a concept someone has to learn.

Heavyweight DefaultLighter AlternativeWhen the Lighter Tool Is Enough
TensorFlow or PyTorchscikit-learnThe problem is tabular classification/regression, not deep learning.
PySparkpandas or PolarsThe dataset fits comfortably on one machine.
apache-airflowprefect, or a simple cron jobThe pipeline is small, or the team wants a gentler learning curve.
Scrapyrequests + beautifulsoup4The task is a single page or a handful of pages, not a large crawl.
Don't Reach for TensorFlow Out of Habit

It is tempting to install TensorFlow or PyTorch on every ML project because they are the most famous names in the space. For a tabular problem like churn prediction, scikit-learn is not a lesser choice — it is simply the correct one, and it avoids pulling in a GPU-oriented deep learning framework the project will never actually use.

Frequently Asked Questions

For production projects, yes — exact pins make behavior reproducible. For quick personal experiments, minimum-version constraints are sometimes fine, but pinning is the safer default once other people or production systems depend on the project.

A practical signal is memory: if pandas operations start crashing with out-of-memory errors, or a single machine cannot hold the dataset even with Polars/Dask, that is when distributed processing starts to earn its complexity.

Generally yes — it is easier to add a dependency when you actually hit the problem it solves than to carry the install size, security surface, and cognitive overhead of an unused library indefinitely.

Key Takeaways From the Course

  • Data manipulation: pandas, Polars, and NumPy form the foundation of nearly every data science project.
  • Visualization: Matplotlib and Seaborn cover statistical plotting; Plotly adds interactivity when that is worth the tradeoff.
  • Classical ML: scikit-learn is the default for tabular classification, regression, and clustering problems.
  • Deep learning: TensorFlow/Keras and PyTorch are reserved for problems that genuinely need neural networks — not a default first choice.
  • NLP and computer vision: specialized libraries exist for text and image problems, and are worth their complexity only when the problem actually calls for them.
  • Data access, scraping, and big data: SQLAlchemy, requests, BeautifulSoup, Scrapy, and PySpark each match a different scale and source of data — choosing the smallest tool that fits the job matters more than reaching for the most powerful one.
  • Serving and orchestration: FastAPI, Gradio, Airflow, and Prefect turn a trained model or pipeline into something other people and systems can actually use and rely on.
  • MLOps: MLflow and Weights & Biases make model development reproducible and comparable across runs.
  • Across all of it, the same question keeps coming back: what does this project actually need right now, and what is the lightest tool that gets it done well.

Summary

The Data Science Dependencies course was never really about memorizing 26 lessons of individual pip packages. It was about building the judgment to look at a real project and know which of those packages actually belong in it — and just as importantly, which do not. That judgment is what separates a clean, maintainable dependency stack from a bloated one, and it is the skill this final lesson was built to practice.

Course Complete!

You've completed the Data Science Dependencies course — you now know what to reach for, why, and how to wire it in, across data manipulation, visualization, classical ML, deep learning, NLP, computer vision, and MLOps.

Explore More Courses