LearnAI ToolsCareerPractice BuildsPlayContact
Data Science DependenciesBeginner~1.5 hours

Dependency Audit

Review a sample requirements.txt, identify redundant or missing libraries, and justify each change.

pipVirtual EnvironmentsLibrary Selection

Overview

This course catalogs 30+ real data science libraries, and the most common way that catalog goes wrong in a real project is not "using the wrong library" — it's a `requirements.txt` file that has quietly accumulated dead weight: a plotting library nobody removed after switching to a different one, a version pin nobody remembers the reason for, an import that works locally because it happens to already be installed globally, but is not actually declared anywhere. This project is a deliberate exercise in reading a dependency list the way an experienced data scientist does — skeptically, one line at a time.

Rather than writing a data science pipeline, you will audit one: a small, intentionally messy `requirements.txt` sits alongside two short Python source files. Your job is to work out, with evidence rather than guesswork, which listed packages the code actually uses, which used package is missing from the list entirely, and which listed packages can be safely removed — then produce a cleaned, version-pinned file with a comment justifying every single line.

What You'll Build
  • An isolated virtual environment created with `python -m venv`, so the audit reflects only what the project actually declares.
  • A messy sample `requirements.txt` reviewed against two real source files it is supposed to support.
  • A small `audit_imports.py` script that scans `.py` files for `import` statements and cross-checks them against `requirements.txt`.
  • A final, cleaned `requirements.txt` with every remaining entry pinned to a version and commented with why it's there.

Prerequisites

  • pip basics — `pip install`, `pip install -r requirements.txt`, and `pip freeze`.
  • Basic command-line comfort — running a script and activating a virtual environment from a terminal.
  • Core Python — reading files, splitting strings, and using a `set` to compare two collections.
  • Understanding what an `import` statement actually does — it is what makes a package "used," not just installed.
  • No prior data science library knowledge required — this project is about dependency hygiene, not modeling or visualization.

Project Structure

The audited project is small on purpose: two source files, `analyze.py` and `report.py`, and the `requirements.txt` that is supposed to describe everything they need. Your own tooling — the virtual environment and the `audit_imports.py` script you'll write in Step 3 — sits alongside them but is not part of what is being audited.

project/
requirements.txt // the messy file being audited — shown in full in Step 2
analyze.py // source file 1 — its imports are the ground truth for what's actually used
report.py // source file 2 — same ground truth, a second file
audit_imports.py // your tool: scans analyze.py + report.py and cross-checks against requirements.txt
venv/ // created in Step 1, isolated from your system Python — never committed to version control

Here are the two source files exactly as they exist before the audit starts. Nothing about them needs to change — they are the evidence, not the thing being edited.

analyze.py
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns # used here, but — check requirements.txt in Step 2 — it is never listed
from sklearn.linear_model import LinearRegression
import requests
def load_and_model(path):
df = pd.read_csv(path)
X = df[["square_feet"]].values
y = df["price"].values
model = LinearRegression().fit(X, y)
return model
report.py
import pandas as pd
import matplotlib.pyplot as plt
def summarize(df):
print(df.describe())
df.plot(kind="bar")
plt.savefig("summary.png")

Step 1: Create and Activate a Virtual Environment

Auditing against your system Python would give a false picture: if `seaborn` already happens to be installed globally from some other project, `import seaborn` would work locally even though `requirements.txt` never declares it — and the moment someone else clones the project into a clean environment, it breaks. A virtual environment removes that blind spot entirely by starting from nothing.

python -m venv venv # creates an isolated Python + pip inside ./venv, completely separate from your system Python
source venv/bin/activate # macOS/Linux — on Windows, run venv\Scripts\activate instead
pip install -r requirements.txt # install exactly what requirements.txt currently declares, messy version and all
pip freeze > installed.txt # snapshot of everything pip actually resolved, including sub-dependencies pulled in for you

That last command matters more than it looks: `pip freeze` shows every package actually installed, including ones nobody listed directly because another package depends on them. Cross-referencing that against `requirements.txt` is a fast way to notice, for example, that `scikit-learn` silently pulled in `scipy` and `joblib` — packages the audited project relies on indirectly without ever needing to list them itself.

Step 2: Read the Messy requirements.txt

Here is the file exactly as it exists before the audit. Read it the way you would review a diff: for every line, ask "is this actually used, and if so, is the version pin sensible?"

requirements.txt (before)
pandas
numpy==1.21.0
matplotlib
plotly
scikit-learn>=0.24
requests
beautifulsoup4
flask
pytest

Even before checking a single import, a few lines are worth being suspicious of on sight: `plotly` is a second, entirely separate charting library sitting next to `matplotlib` — is the project actually using both, or did `plotly` get added once and never removed? `numpy` is pinned to an exact version (`==1.21.0`) while `scikit-learn` uses a loose floor (`>=0.24`) and everything else has no pin at all — that inconsistency itself is a signal nobody has reviewed this file as a whole. `flask` is a web framework; nothing about `analyze.py` or `report.py` looks like it serves HTTP requests. Suspicion is a starting point, not a conclusion — Step 3 replaces it with evidence.

Step 3: Scan the Source Code for What's Actually Imported

Rather than manually re-reading both files every time the audit changes, `audit_imports.py` automates the cross-check: it scans every `.py` file in the project for top-level `import x` and `from x import y` lines, extracts just the package name being imported, and compares that set against what `requirements.txt` declares. The result is three lists — used-and-listed, used-but-missing, and listed-but-unused — which turns "I think this is unused" into "here is the evidence this is unused."

import re # regular expressions are enough for this — a full Python parser (the ast module) would be overkill for a two-file audit
import glob # finds every .py file in the current directory without hardcoding filenames
# Some packages are imported under a different name than the one pip installs them under —
# scikit-learn installs as "scikit-learn" but is imported as "sklearn". Without this map, the
# audit would wrongly flag scikit-learn as unused just because "sklearn" never appears verbatim.
IMPORT_TO_PACKAGE = {
"sklearn": "scikit-learn",
"bs4": "beautifulsoup4",
"cv2": "opencv-python",
}
# Matches "import pandas", "import pandas as pd", and "from pandas import DataFrame" —
# group(1) or group(2) captures just the top-level package name in either form.
IMPORT_PATTERN = re.compile(r"^(?:import|from)\s+([a-zA-Z0-9_]+)")
def find_imports(filenames):
"""Scan the given .py files and return the set of top-level packages they import."""
imports = set()
for filename in filenames:
with open(filename, "r") as f:
for line in f:
match = IMPORT_PATTERN.match(line.strip()) # only matches lines starting with import/from, ignoring indented ones
if match:
module = match.group(1)
package = IMPORT_TO_PACKAGE.get(module, module) # translate to the pip package name if it differs from the import name
imports.add(package)
return imports
def parse_requirements(path):
"""Read requirements.txt and return the set of package names it declares, ignoring version pins."""
declared = set()
with open(path, "r") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"): # skip blank lines and comment lines
continue
# strips off ==, >=, or any other version specifier so "numpy==1.21.0" becomes just "numpy"
name = re.split(r"[=<>]", line)[0].strip()
declared.add(name)
return declared
if __name__ == "__main__":
source_files = glob.glob("*.py")
source_files = [f for f in source_files if f != "audit_imports.py"] # don't scan the audit tool's own imports
used = find_imports(source_files)
declared = parse_requirements("requirements.txt")
print(f"Scanning {len(source_files)} source files for imports...\n")
print("Used in code AND declared in requirements.txt (keep):")
for name in sorted(used & declared): # set intersection: present in both
print(f" {name}")
print("\nUsed in code but MISSING from requirements.txt (add these):")
for name in sorted(used - declared): # set difference: used, but not declared
print(f" {name}")
print("\nListed in requirements.txt but NOT imported anywhere (candidates to remove):")
for name in sorted(declared - used): # set difference: declared, but never used
print(f" {name}")

Step 4: Pin Versions Responsibly

The audit script tells you *which* packages belong in the file; it deliberately does not decide version numbers for you, because that judgment call depends on context the script can't see. An exact pin (`pandas==2.1.4`) guarantees every install is byte-for-byte identical, which is the right default for a project other people will clone and run — nobody wants "works on my machine" bugs from a silent minor-version upgrade. A loose floor (`requests>=2.31`) is reasonable for a package with a stable, rarely-breaking API where you actively want small security and bugfix updates to flow in automatically. No pin at all is rarely the right choice outside of a throwaway experiment — it means the exact same `pip install -r requirements.txt` command can install different code on different days.

Use `pip show <package>` inside the activated virtual environment from Step 1 to find the exact version currently installed and resolved as compatible with everything else, and use that as your pin — it is the version you have actually tested against, not a guess.

pip show pandas # prints the exact installed version, e.g. Version: 2.1.4 — use this as the pin, not a guess
pip show scikit-learn
pip show numpy

Step 5: Write the Cleaned requirements.txt

Combining what `audit_imports.py` reported with the version-pinning judgment from Step 4 produces the final file below. Every line carries a comment explaining why it survived the audit — including `pytest`, which the import scanner correctly reports as "unused" in `analyze.py`/`report.py` because tests are not application code, but which still belongs in the file for a different, equally valid reason.

requirements.txt (after)
pandas==2.1.4 # used in analyze.py and report.py for loading and describing data
numpy==1.26.2 # used indirectly via pandas/scikit-learn, and directly for X/y array shaping in analyze.py — bumped from the stale 1.21.0 pin
matplotlib==3.8.2 # used in report.py to render and save the summary chart
seaborn==0.13.0 # used in analyze.py for the scatter plot — was imported in code but MISSING from the original file entirely
scikit-learn==1.3.2 # used in analyze.py for LinearRegression — pinned exactly instead of the original loose >=0.24 floor
requests==2.31.0 # used in analyze.py to fetch remote data before loading it into pandas
pytest==7.4.3 # not imported by application code, but required to run this project's test suite — kept, with this comment explaining why
# Removed, with reasons:
# plotly -> never imported anywhere; matplotlib already covers this project's one chart
# beautifulsoup4 -> never imported anywhere; no HTML scraping happens in this project
# flask -> never imported anywhere; this project has no HTTP server or API

Complete Code

Here is the complete `audit_imports.py` tool from Step 3, ready to drop into any Python project directory and run with `python audit_imports.py`.

import re
import glob
IMPORT_TO_PACKAGE = {
"sklearn": "scikit-learn",
"bs4": "beautifulsoup4",
"cv2": "opencv-python",
}
IMPORT_PATTERN = re.compile(r"^(?:import|from)\s+([a-zA-Z0-9_]+)")
def find_imports(filenames):
"""Scan the given .py files and return the set of top-level packages they import."""
imports = set()
for filename in filenames:
with open(filename, "r") as f:
for line in f:
match = IMPORT_PATTERN.match(line.strip())
if match:
module = match.group(1)
package = IMPORT_TO_PACKAGE.get(module, module)
imports.add(package)
return imports
def parse_requirements(path):
"""Read requirements.txt and return the set of package names it declares, ignoring version pins."""
declared = set()
with open(path, "r") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"):
continue
name = re.split(r"[=<>]", line)[0].strip()
declared.add(name)
return declared
if __name__ == "__main__":
source_files = glob.glob("*.py")
source_files = [f for f in source_files if f != "audit_imports.py"]
used = find_imports(source_files)
declared = parse_requirements("requirements.txt")
print(f"Scanning {len(source_files)} source files for imports...\n")
print("Used in code AND declared in requirements.txt (keep):")
for name in sorted(used & declared):
print(f" {name}")
print("\nUsed in code but MISSING from requirements.txt (add these):")
for name in sorted(used - declared):
print(f" {name}")
print("\nListed in requirements.txt but NOT imported anywhere (candidates to remove):")
for name in sorted(declared - used):
print(f" {name}")

Sample Run

Sample Run

Click Run to see what this code prints.

Extend This Project

  • Split the cleaned file into `requirements.txt` (runtime) and `requirements-dev.txt` (`pytest`, linters), matching how `pytest` was flagged as a special case above.
  • Extend `audit_imports.py` to recurse into subdirectories with `glob.glob("**/*.py", recursive=True)` instead of only scanning the top-level folder.
  • Add a check that flags any `requirements.txt` line with no version pin at all, not just genuinely unused packages.
  • Run `pip list --outdated` inside the virtual environment and decide, package by package, which updates are safe to pull in.
  • Wire `audit_imports.py` into a pre-commit hook or CI step so a genuinely unused or undeclared import gets caught automatically before it ships.

Summary

You audited a dependency list the way experienced data scientists actually do it: isolate the environment first so nothing outside the project can mask a problem, gather evidence about what is genuinely imported instead of guessing, and only then decide what to keep, pin, add, or remove — with a written reason for every line. That process matters more than any specific library in this file: the same virtual-environment-plus-import-scan approach applies to any Python project's `requirements.txt`, data science or otherwise.