Real-World Project: Building a Data Science Dependency Stack
Bring the whole course together by assembling a realistic dependency stack for a churn-prediction pipeline, plus a decision framework for choosing libraries on any new project.
Introduction
Over the last 25 lessons, you have gone through dozens of Python packages one at a time — a library for loading data, a library for plotting it, a library for modeling it, a library for serving it. In a real project, none of those decisions happen in isolation. You have to look at what you are actually building and pick a coherent stack, not just a shopping list of every interesting tool this course covered.
This final lesson works through exactly that, using one realistic project as the example: a churn-prediction pipeline for a subscription business. By the end, you will have a complete requirements.txt for the project, and — more importantly — a repeatable way of deciding what belongs in requirements.txt on any future project.
- How to walk through a realistic project and choose a dependency for each stage, using libraries from across this course.
- What a complete, production-ready requirements.txt looks like for that project.
- A decision-framework checklist for picking libraries on any new project.
- How to recognize and avoid dependency bloat.
The Project: Churn Prediction
The scenario: a subscription business wants to predict which customers are likely to cancel next month, so the retention team can reach out first. The data lives in a database. Nobody on the team wants to run a Spark cluster for this — the dataset is a few hundred thousand rows, well within reach of a single machine.
Walking through the pipeline stage by stage, and pulling the right tool for each stage from earlier in this course:
| Stage | Library | Why This One |
|---|---|---|
| Pull data from the database | sqlalchemy + pandas | The dataset fits on one machine — no need for PySpark. |
| Clean and reshape the data | pandas | The standard, well-understood tool for tabular cleaning. |
| Explore and visualize patterns | matplotlib + seaborn | Quick statistical plots to understand churn drivers before modeling. |
| Train the model | scikit-learn | A classic tabular classification problem — no need for a deep learning framework. |
| Track experiments | mlflow | Self-hosted tracking is enough; no need to pay for a hosted dashboard yet. |
| Serve the model | fastapi | The retention team's internal tools need a real API to call. |
| Keep secrets out of the code | python-dotenv | The database URL and any API keys should never be hardcoded. |
| Keep the codebase clean | black + ruff | Two engineers are collaborating on this repo. |
Notice what is missing: no TensorFlow or PyTorch, no Scrapy, no Airflow, no PySpark. Every one of those is a legitimate tool covered in this course — but none of them is needed for this specific project. That absence is the whole point of this lesson.
import pandas as pdfrom sqlalchemy import create_enginefrom sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import accuracy_scoreimport mlflowimport mlflow.sklearn
engine = create_engine("postgresql://user:pass@localhost/churn_db")df = pd.read_sql("SELECT * FROM customers", engine)
df = df.dropna(subset=["tenure_months", "monthly_charges"])X = df[["tenure_months", "monthly_charges", "support_tickets"]]y = df["churned"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
with mlflow.start_run(): model = RandomForestClassifier(n_estimators=150, max_depth=8) model.fit(X_train, y_train)
accuracy = accuracy_score(y_test, model.predict(X_test)) mlflow.log_metric("accuracy", accuracy) mlflow.sklearn.log_model(model, "churn_model")
print(f"Model trained with accuracy: {accuracy:.3f}")Click Run to see what this code prints.
The Final requirements.txt
Pulling it all together, this is the complete, production-ready dependency list for the project — every line earns its place for a specific stage of the pipeline above.
# Data accesssqlalchemy==2.0.30psycopg2-binary==2.9.9
# Data manipulationpandas==2.2.2
# Visualizationmatplotlib==3.8.4seaborn==0.13.2
# Modelingscikit-learn==1.4.2
# Experiment trackingmlflow==2.13.0
# Servingfastapi==0.111.0uvicorn==0.29.0
# Config & secretspython-dotenv==1.0.1
# Dev toolingblack==24.4.2ruff==0.4.4Notice every dependency is pinned to an exact version rather than left open-ended. This is what makes the pipeline reproducible — the same requirements.txt installs the same behavior on any machine, today or a year from now.
A Decision Framework for Any Project
The table above was not arbitrary — it came from asking the same small set of questions about every candidate library. Use this checklist the next time you are tempted to pip install something new.
- Do I need this right now, or am I speculatively adding it for a feature that does not exist yet?
- Is there a lighter-weight tool already in my stack that solves the same problem well enough?
- Does this library solve a problem I actually have (data too large for memory, a real API to serve, an actual multi-page crawl) — or am I reaching for it because it is the "advanced" or exciting option?
- What is the cost of being wrong — how hard is it to swap this library out later if the project outgrows it, or if it turns out to be overkill?
- Does adding this pull in a chain of heavy transitive dependencies, slower installs, and a larger attack surface?
For this project: PySpark failed question 1 (the data fits on one machine right now) and question 3 (no real cluster-scale problem exists). Airflow failed question 1 for the same reason — a single daily retraining script does not yet need a scheduler with a UI. If the business scales to millions of customers and the pipeline grows to ten dependent steps, that answer would change, and the framework would say yes instead.
Avoiding Dependency Bloat
Dependency bloat is what happens when a project accumulates libraries that solve problems it does not actually have — usually because someone reached for the well-known, heavyweight tool instead of asking whether a lighter one would do. It slows down installs, increases the surface area for security vulnerabilities, and makes the codebase harder for new contributors to understand, since every dependency is a concept someone has to learn.
| Heavyweight Default | Lighter Alternative | When the Lighter Tool Is Enough |
|---|---|---|
| TensorFlow or PyTorch | scikit-learn | The problem is tabular classification/regression, not deep learning. |
| PySpark | pandas or Polars | The dataset fits comfortably on one machine. |
| apache-airflow | prefect, or a simple cron job | The pipeline is small, or the team wants a gentler learning curve. |
| Scrapy | requests + beautifulsoup4 | The task is a single page or a handful of pages, not a large crawl. |
It is tempting to install TensorFlow or PyTorch on every ML project because they are the most famous names in the space. For a tabular problem like churn prediction, scikit-learn is not a lesser choice — it is simply the correct one, and it avoids pulling in a GPU-oriented deep learning framework the project will never actually use.
Frequently Asked Questions
For production projects, yes — exact pins make behavior reproducible. For quick personal experiments, minimum-version constraints are sometimes fine, but pinning is the safer default once other people or production systems depend on the project.
A practical signal is memory: if pandas operations start crashing with out-of-memory errors, or a single machine cannot hold the dataset even with Polars/Dask, that is when distributed processing starts to earn its complexity.
Generally yes — it is easier to add a dependency when you actually hit the problem it solves than to carry the install size, security surface, and cognitive overhead of an unused library indefinitely.
Key Takeaways From the Course
- Data manipulation: pandas, Polars, and NumPy form the foundation of nearly every data science project.
- Visualization: Matplotlib and Seaborn cover statistical plotting; Plotly adds interactivity when that is worth the tradeoff.
- Classical ML: scikit-learn is the default for tabular classification, regression, and clustering problems.
- Deep learning: TensorFlow/Keras and PyTorch are reserved for problems that genuinely need neural networks — not a default first choice.
- NLP and computer vision: specialized libraries exist for text and image problems, and are worth their complexity only when the problem actually calls for them.
- Data access, scraping, and big data: SQLAlchemy, requests, BeautifulSoup, Scrapy, and PySpark each match a different scale and source of data — choosing the smallest tool that fits the job matters more than reaching for the most powerful one.
- Serving and orchestration: FastAPI, Gradio, Airflow, and Prefect turn a trained model or pipeline into something other people and systems can actually use and rely on.
- MLOps: MLflow and Weights & Biases make model development reproducible and comparable across runs.
- Across all of it, the same question keeps coming back: what does this project actually need right now, and what is the lightest tool that gets it done well.
Summary
The Data Science Dependencies course was never really about memorizing 26 lessons of individual pip packages. It was about building the judgment to look at a real project and know which of those packages actually belong in it — and just as importantly, which do not. That judgment is what separates a clean, maintainable dependency stack from a bloated one, and it is the skill this final lesson was built to practice.