Overview
The Data Science Dependencies course is deliberately organized as a categorized reference — one lesson per library, each with its own use case and example — because that is how the ecosystem is actually learned: one tool at a time. But no real analysis ever uses just one library. This project is where that changes: you will take three libraries this course covers in separate lessons — pandas for wrangling, matplotlib/seaborn for visualization, and scikit-learn for modeling — and chain them into a single pipeline where each library's output is the next library's input.
The goal is not to build the most sophisticated model possible. It is to practice the workflow every real analysis follows: load data, understand its shape and its problems, look at it visually before trusting any number a model produces, then model it and honestly evaluate how good that model actually is. Each step below explains not just what each library call does, but why that library was the right tool for that specific step — the same "why this, not that" framing this course uses lesson by lesson, now applied inside one connected project.
- A small, self-contained housing-price dataset generated inline and loaded with `pandas.read_csv()`.
- A cleaning step that finds and fills missing values using `.info()`, `.isna()`, and `.fillna()`.
- Two exploratory charts — a scatter plot and a correlation heatmap — built with `matplotlib` and `seaborn`.
- A `scikit-learn` `LinearRegression` model trained with a proper `train_test_split`.
- An evaluation step reporting Mean Absolute Error, R², and the model's learned coefficients.
Prerequisites
- Core Python — functions, f-strings, and working with lists and dicts.
- pandas basics — what a DataFrame is and how `df["column"]` selects a column (this course's Data Manipulation Libraries lesson).
- Familiarity with installing packages via `pip install` (covered in the pip vs conda and Virtual Environments lessons).
- No prior visualization or machine learning experience required — every call to `matplotlib`, `seaborn`, and `scikit-learn` below is explained as it appears.
Project Structure
Everything lives in one file, `analysis.py`, plus the files it produces when run: `housing.csv` (the sample dataset, generated by the script itself so this project needs no external download), and two chart images, `price_vs_sqft.png` and `correlation_heatmap.png`. Before running it, install the four libraries this pipeline depends on:
pip install pandas matplotlib seaborn scikit-learnThe dataset itself is a small, made-up table of ten houses with four numeric features — `square_feet`, `bedrooms`, `age_years`, and `neighborhood_score` — and a `price` column that is the value you will try to predict. Two of the ten rows are missing a `neighborhood_score`, which is intentional: Step 2 exists specifically to handle that gap before it can quietly break Step 3's charts or Step 4's model.
| square_feet | bedrooms | age_years | neighborhood_score | price |
|---|---|---|---|---|
| 1200 | 2 | 15 | 7.2 | 245000 |
| 1550 | 3 | 8 | 8.1 | 312000 |
| 980 | 2 | 32 | 5.4 | 178000 |
| 2100 | 4 | 5 | 9.0 | 415000 |
| 1400 | 3 | 20 | (missing) | 268000 |
| 1750 | 3 | 12 | 7.8 | 335000 |
| 2300 | 4 | 3 | 9.4 | 455000 |
| 1100 | 2 | 25 | (missing) | 205000 |
| 1600 | 3 | 18 | 6.9 | 298000 |
| 1900 | 4 | 10 | 8.5 | 372000 |
Step 1: Set Up and Load the Sample Dataset with pandas
The script writes its own sample CSV to disk before reading it back, so this whole project is reproducible without asking you to download anything first — in a real job, this same `pandas.read_csv()` call would instead point at a file exported from a database or spreadsheet. `.info()` and `.describe()` are run immediately after loading, before any other library gets involved, because pandas is the right tool for a first-pass structural check: it can report dtypes, non-null counts, and summary statistics in three method calls, work `matplotlib` or `scikit-learn` are not designed to do at all.
import pandas as pd # pandas is the library every step below builds on: it turns raw rows into a DataFrame you can filter, group, and clean
# In a real project this data would already exist as a CSV exported from a database or spreadsheet.# It's generated inline here so this tutorial is fully self-contained and reproducible with no external download.SAMPLE_CSV = """square_feet,bedrooms,age_years,neighborhood_score,price1200,2,15,7.2,2450001550,3,8,8.1,312000980,2,32,5.4,1780002100,4,5,9.0,4150001400,3,20,,2680001750,3,12,7.8,3350002300,4,3,9.4,4550001100,2,25,,2050001600,3,18,6.9,2980001900,4,10,8.5,372000""" # note the two blank entries in the neighborhood_score column — those become NaN values pandas has to handle in Step 2
with open("housing.csv", "w") as f: # write the sample data to disk once, so read_csv() below behaves exactly like it would on a real file f.write(SAMPLE_CSV)
df = pd.read_csv("housing.csv") # read_csv infers each column's dtype automatically (int64, float64, etc.) from the text it parses
print(df.head()) # first 5 rows — a quick sanity check that every column parsed into the right placeprint(df.info()) # column dtypes plus non-null counts per column — this is where the 2 missing neighborhood_score values surfaceprint(df.describe()) # count/mean/std/min/max per numeric column — a fast way to spot outliers or unexpected scale before modelingStep 2: Clean the Data
This is the step a rushed analysis skips, and it is the step that makes the rest of the pipeline trustworthy. `df.isna().sum()` confirms precisely which column has missing values and how many, rather than guessing from `.info()`'s output. The fill value is deliberately the median, not the mean: the mean is pulled toward outliers, so a single unusually high or low `neighborhood_score` would distort every row that gets filled with it, while the median is robust to exactly that.
print(df.isna().sum()) # confirms exactly which column(s) are missing values and how many — here, neighborhood_score: 2
# Why median, not mean: the mean is pulled toward outliers, so one unusually high or low# neighborhood_score would skew every row that gets filled with it. The median doesn't have that problem.median_score = df["neighborhood_score"].median()df["neighborhood_score"] = df["neighborhood_score"].fillna(median_score) # fill, don't drop — dropping wastes otherwise-good rows
# Why fillna() over dropna(): with only 10 rows, dropping the 2 rows missing neighborhood_score# would throw away 20% of the dataset over a single blemished column, discarding perfectly good# square_feet/bedrooms/age_years/price data along with it. Filling keeps every row usable.print(df.isna().sum()) # every column should now report 0 missing valuesStep 3: Visualize Relationships with Matplotlib and Seaborn
`matplotlib` is the underlying plotting engine nearly every other Python visualization library, including `seaborn`, is built on top of. `seaborn` is used here specifically because it wraps `matplotlib` with statistical-plot defaults — better default styling, and a one-line `heatmap()` that would otherwise take several lines of raw `matplotlib` calls to reproduce. The two charts below exist to answer a real question each, not just to decorate the analysis: does square footage actually relate to price, and which of the four features correlates with price the most strongly? Both questions are answered before Step 4 picks which features to feed into a model, not after.
import matplotlib.pyplot as plt # the plotting engine; pyplot gives a simple, stateful interface for building one figure at a timeimport seaborn as sns # wraps matplotlib with statistical-plot defaults — cleaner styling and a built-in heatmap() call
# Scatter plot: does square footage actually predict price? This is the question Step 4 needs# answered before it's worth trusting square_feet as a model feature at all.sns.scatterplot(data=df, x="square_feet", y="price")plt.title("Price vs. Square Footage")plt.xlabel("Square Feet")plt.ylabel("Price ($)")plt.savefig("price_vs_sqft.png") # saved to a file instead of plt.show() so this script also runs unattended in a terminal or CI jobplt.close() # closes the current figure so the next plt call starts a fresh, empty figure
# Heatmap: of the four numeric features, which correlates most strongly with price? This# guides feature selection in Step 4 instead of guessing which columns actually matter.plt.figure(figsize=(6, 5))sns.heatmap(df.corr(numeric_only=True), annot=True, cmap="coolwarm") # annot=True prints each correlation value directly on the cellplt.title("Feature Correlation Heatmap")plt.savefig("correlation_heatmap.png")plt.close()Running this step produces `price_vs_sqft.png` and `correlation_heatmap.png` on disk. The scatter plot shows the ten points trending clearly upward and to the right — bigger houses cost more, with no obvious outliers breaking that pattern in this small dataset. The heatmap shows `square_feet` as the feature with the strongest correlation to `price`, `bedrooms` and `neighborhood_score` as moderate positive correlations, and `age_years` as a mild negative correlation — older houses tend to be worth slightly less, holding size roughly constant. That ranking is exactly what Step 4 uses `square_feet` as the model's single strongest predictor for.
Step 4: Train a Model with Scikit-learn
`scikit-learn` handles both the train/test split and the model itself through one consistent API, which is the main reason it is the default choice for a first model rather than a deep learning framework: `train_test_split` holds back 20% of the rows purely for evaluation, so Step 5 can honestly report how the model performs on data it never saw during training, not just how well it memorized the data it trained on.
from sklearn.model_selection import train_test_split # splits data so evaluation happens on rows the model never trained onfrom sklearn.linear_model import LinearRegression # a simple, interpretable baseline — see the note below for why it fits here
FEATURES = ["square_feet", "bedrooms", "age_years", "neighborhood_score"] # every numeric column except the target itselfX = df[FEATURES]y = df["price"]
# test_size=0.2 holds back 20% of rows purely for evaluation in Step 5; random_state=42 makes the# split reproducible, so re-running this script doesn't silently change which rows land in the test set.X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Why LinearRegression and not a more powerful model like RandomForestRegressor: Step 3's scatter# plot already showed price trending roughly linearly with square footage, and with only 10 rows# total, a linear model is both sufficient and far more interpretable — its coefficients say# directly "how many dollars per extra square foot," which a tree ensemble's output can't.model = LinearRegression()model.fit(X_train, y_train) # learns one coefficient per feature plus an intercept from X_train/y_trainStep 5: Evaluate the Model and Interpret Results
A model is only as useful as an honest measurement of how wrong it is. Mean Absolute Error is reported first because it is in the same units as the target — dollars — which makes it easy to explain to someone who has never heard of R²: "the model is off by about this many dollars on average." R² is reported second as the more standard machine-learning metric, representing the fraction of price variance the model's features explain.
from sklearn.metrics import mean_absolute_error, r2_score # two complementary ways to judge fit: error in dollars, and variance explained
predictions = model.predict(X_test) # predict on the held-out rows the model never saw during fit()
mae = mean_absolute_error(y_test, predictions) # average absolute dollar error — easy to explain to a non-technical stakeholderr2 = r2_score(y_test, predictions) # fraction of price variance the model explains; 1.0 would be a perfect fit
print(f"Mean Absolute Error: ${mae:,.2f}")print(f"R^2 Score: {r2:.3f}")
# Printing each coefficient alongside its feature name turns the model's internals into a plain-# English sentence: "every extra square foot is worth about $X in this data," which is the kind# of interpretability that was the whole justification for choosing LinearRegression in Step 4.for feature, coef in zip(FEATURES, model.coef_): print(f" {feature}: ${coef:,.2f}")print(f"Intercept: ${model.intercept_:,.2f}")pandas owned data loading, inspection, and cleaning — the structural work no other library here does. matplotlib and seaborn owned exploration, turning columns of numbers into a shape a human can judge before trusting a model built on them. scikit-learn owned splitting, fitting, and evaluating — and only had a job to do because the first two libraries had already made sure the data going into it was clean and worth modeling. That handoff, not any single library, is what "end-to-end" means in this project's title.
Complete Code
Here is the full pipeline assembled in order, ready to run with `python analysis.py` after installing the four required libraries.
import pandas as pdimport matplotlib.pyplot as pltimport seaborn as snsfrom sklearn.model_selection import train_test_splitfrom sklearn.linear_model import LinearRegressionfrom sklearn.metrics import mean_absolute_error, r2_score
# --- Step 1: Set up and load the sample dataset ---SAMPLE_CSV = """square_feet,bedrooms,age_years,neighborhood_score,price1200,2,15,7.2,2450001550,3,8,8.1,312000980,2,32,5.4,1780002100,4,5,9.0,4150001400,3,20,,2680001750,3,12,7.8,3350002300,4,3,9.4,4550001100,2,25,,2050001600,3,18,6.9,2980001900,4,10,8.5,372000"""
with open("housing.csv", "w") as f: f.write(SAMPLE_CSV)
df = pd.read_csv("housing.csv")print(df.head())print(df.info())print(df.describe())
# --- Step 2: Clean the data ---print(df.isna().sum())median_score = df["neighborhood_score"].median()df["neighborhood_score"] = df["neighborhood_score"].fillna(median_score)print(df.isna().sum())
# --- Step 3: Visualize relationships ---sns.scatterplot(data=df, x="square_feet", y="price")plt.title("Price vs. Square Footage")plt.xlabel("Square Feet")plt.ylabel("Price ($)")plt.savefig("price_vs_sqft.png")plt.close()
plt.figure(figsize=(6, 5))sns.heatmap(df.corr(numeric_only=True), annot=True, cmap="coolwarm")plt.title("Feature Correlation Heatmap")plt.savefig("correlation_heatmap.png")plt.close()
# --- Step 4: Train a model ---FEATURES = ["square_feet", "bedrooms", "age_years", "neighborhood_score"]X = df[FEATURES]y = df["price"]X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LinearRegression()model.fit(X_train, y_train)
# --- Step 5: Evaluate the model ---predictions = model.predict(X_test)mae = mean_absolute_error(y_test, predictions)r2 = r2_score(y_test, predictions)
print(f"Mean Absolute Error: ${mae:,.2f}")print(f"R^2 Score: {r2:.3f}")for feature, coef in zip(FEATURES, model.coef_): print(f" {feature}: ${coef:,.2f}")print(f"Intercept: ${model.intercept_:,.2f}")print("Saved price_vs_sqft.png and correlation_heatmap.png")Sample Run
Click Run to see what this code prints.
Extend This Project
- Add a `distance_to_city_center` feature to the dataset and see whether R² improves or the correlation heatmap ranks it above `neighborhood_score`.
- Swap `LinearRegression` for `RandomForestRegressor` and compare its MAE/R² against the linear baseline — is the extra complexity actually worth it on data this small?
- Replace the manual `fillna()` step with a `scikit-learn` `Pipeline` and `ColumnTransformer` so imputation and scaling happen automatically before every fit.
- Use `cross_val_score` with 5-fold cross-validation instead of a single `train_test_split`, which is more reliable on a dataset this size.
- Load a larger, real housing dataset (e.g. from Kaggle) through the same pipeline unchanged, and see which assumptions from this small sample no longer hold.
Summary
You built one connected pipeline out of three libraries this course otherwise teaches separately: pandas loaded and cleaned the data, matplotlib and seaborn made its relationships visible before any model touched it, and scikit-learn split, trained, and honestly evaluated a model on top of that groundwork. The specific dataset here is small and made up, but the sequence — load, inspect, clean, visualize, model, evaluate — is the same one every real analysis follows, no matter which specific libraries fill in each step.