Overview
This course catalogs 30+ real data science libraries, and the most common way that catalog goes wrong in a real project is not "using the wrong library" — it's a `requirements.txt` file that has quietly accumulated dead weight: a plotting library nobody removed after switching to a different one, a version pin nobody remembers the reason for, an import that works locally because it happens to already be installed globally, but is not actually declared anywhere. This project is a deliberate exercise in reading a dependency list the way an experienced data scientist does — skeptically, one line at a time.
Rather than writing a data science pipeline, you will audit one: a small, intentionally messy `requirements.txt` sits alongside two short Python source files. Your job is to work out, with evidence rather than guesswork, which listed packages the code actually uses, which used package is missing from the list entirely, and which listed packages can be safely removed — then produce a cleaned, version-pinned file with a comment justifying every single line.
- An isolated virtual environment created with `python -m venv`, so the audit reflects only what the project actually declares.
- A messy sample `requirements.txt` reviewed against two real source files it is supposed to support.
- A small `audit_imports.py` script that scans `.py` files for `import` statements and cross-checks them against `requirements.txt`.
- A final, cleaned `requirements.txt` with every remaining entry pinned to a version and commented with why it's there.
Prerequisites
- pip basics — `pip install`, `pip install -r requirements.txt`, and `pip freeze`.
- Basic command-line comfort — running a script and activating a virtual environment from a terminal.
- Core Python — reading files, splitting strings, and using a `set` to compare two collections.
- Understanding what an `import` statement actually does — it is what makes a package "used," not just installed.
- No prior data science library knowledge required — this project is about dependency hygiene, not modeling or visualization.
Project Structure
The audited project is small on purpose: two source files, `analyze.py` and `report.py`, and the `requirements.txt` that is supposed to describe everything they need. Your own tooling — the virtual environment and the `audit_imports.py` script you'll write in Step 3 — sits alongside them but is not part of what is being audited.
project/ requirements.txt // the messy file being audited — shown in full in Step 2 analyze.py // source file 1 — its imports are the ground truth for what's actually used report.py // source file 2 — same ground truth, a second file audit_imports.py // your tool: scans analyze.py + report.py and cross-checks against requirements.txt venv/ // created in Step 1, isolated from your system Python — never committed to version controlHere are the two source files exactly as they exist before the audit starts. Nothing about them needs to change — they are the evidence, not the thing being edited.
import pandas as pdimport numpy as npimport matplotlib.pyplot as pltimport seaborn as sns # used here, but — check requirements.txt in Step 2 — it is never listedfrom sklearn.linear_model import LinearRegressionimport requests
def load_and_model(path): df = pd.read_csv(path) X = df[["square_feet"]].values y = df["price"].values model = LinearRegression().fit(X, y) return modelimport pandas as pdimport matplotlib.pyplot as plt
def summarize(df): print(df.describe()) df.plot(kind="bar") plt.savefig("summary.png")Step 1: Create and Activate a Virtual Environment
Auditing against your system Python would give a false picture: if `seaborn` already happens to be installed globally from some other project, `import seaborn` would work locally even though `requirements.txt` never declares it — and the moment someone else clones the project into a clean environment, it breaks. A virtual environment removes that blind spot entirely by starting from nothing.
python -m venv venv # creates an isolated Python + pip inside ./venv, completely separate from your system Pythonsource venv/bin/activate # macOS/Linux — on Windows, run venv\Scripts\activate insteadpip install -r requirements.txt # install exactly what requirements.txt currently declares, messy version and allpip freeze > installed.txt # snapshot of everything pip actually resolved, including sub-dependencies pulled in for youThat last command matters more than it looks: `pip freeze` shows every package actually installed, including ones nobody listed directly because another package depends on them. Cross-referencing that against `requirements.txt` is a fast way to notice, for example, that `scikit-learn` silently pulled in `scipy` and `joblib` — packages the audited project relies on indirectly without ever needing to list them itself.
Step 2: Read the Messy requirements.txt
Here is the file exactly as it exists before the audit. Read it the way you would review a diff: for every line, ask "is this actually used, and if so, is the version pin sensible?"
pandasnumpy==1.21.0matplotlibplotlyscikit-learn>=0.24requestsbeautifulsoup4flaskpytestEven before checking a single import, a few lines are worth being suspicious of on sight: `plotly` is a second, entirely separate charting library sitting next to `matplotlib` — is the project actually using both, or did `plotly` get added once and never removed? `numpy` is pinned to an exact version (`==1.21.0`) while `scikit-learn` uses a loose floor (`>=0.24`) and everything else has no pin at all — that inconsistency itself is a signal nobody has reviewed this file as a whole. `flask` is a web framework; nothing about `analyze.py` or `report.py` looks like it serves HTTP requests. Suspicion is a starting point, not a conclusion — Step 3 replaces it with evidence.
Step 3: Scan the Source Code for What's Actually Imported
Rather than manually re-reading both files every time the audit changes, `audit_imports.py` automates the cross-check: it scans every `.py` file in the project for top-level `import x` and `from x import y` lines, extracts just the package name being imported, and compares that set against what `requirements.txt` declares. The result is three lists — used-and-listed, used-but-missing, and listed-but-unused — which turns "I think this is unused" into "here is the evidence this is unused."
import re # regular expressions are enough for this — a full Python parser (the ast module) would be overkill for a two-file auditimport glob # finds every .py file in the current directory without hardcoding filenames
# Some packages are imported under a different name than the one pip installs them under —# scikit-learn installs as "scikit-learn" but is imported as "sklearn". Without this map, the# audit would wrongly flag scikit-learn as unused just because "sklearn" never appears verbatim.IMPORT_TO_PACKAGE = { "sklearn": "scikit-learn", "bs4": "beautifulsoup4", "cv2": "opencv-python",}
# Matches "import pandas", "import pandas as pd", and "from pandas import DataFrame" —# group(1) or group(2) captures just the top-level package name in either form.IMPORT_PATTERN = re.compile(r"^(?:import|from)\s+([a-zA-Z0-9_]+)")
def find_imports(filenames): """Scan the given .py files and return the set of top-level packages they import.""" imports = set() for filename in filenames: with open(filename, "r") as f: for line in f: match = IMPORT_PATTERN.match(line.strip()) # only matches lines starting with import/from, ignoring indented ones if match: module = match.group(1) package = IMPORT_TO_PACKAGE.get(module, module) # translate to the pip package name if it differs from the import name imports.add(package) return imports
def parse_requirements(path): """Read requirements.txt and return the set of package names it declares, ignoring version pins.""" declared = set() with open(path, "r") as f: for line in f: line = line.strip() if not line or line.startswith("#"): # skip blank lines and comment lines continue # strips off ==, >=, or any other version specifier so "numpy==1.21.0" becomes just "numpy" name = re.split(r"[=<>]", line)[0].strip() declared.add(name) return declared
if __name__ == "__main__": source_files = glob.glob("*.py") source_files = [f for f in source_files if f != "audit_imports.py"] # don't scan the audit tool's own imports
used = find_imports(source_files) declared = parse_requirements("requirements.txt")
print(f"Scanning {len(source_files)} source files for imports...\n")
print("Used in code AND declared in requirements.txt (keep):") for name in sorted(used & declared): # set intersection: present in both print(f" {name}")
print("\nUsed in code but MISSING from requirements.txt (add these):") for name in sorted(used - declared): # set difference: used, but not declared print(f" {name}")
print("\nListed in requirements.txt but NOT imported anywhere (candidates to remove):") for name in sorted(declared - used): # set difference: declared, but never used print(f" {name}")Step 4: Pin Versions Responsibly
The audit script tells you *which* packages belong in the file; it deliberately does not decide version numbers for you, because that judgment call depends on context the script can't see. An exact pin (`pandas==2.1.4`) guarantees every install is byte-for-byte identical, which is the right default for a project other people will clone and run — nobody wants "works on my machine" bugs from a silent minor-version upgrade. A loose floor (`requests>=2.31`) is reasonable for a package with a stable, rarely-breaking API where you actively want small security and bugfix updates to flow in automatically. No pin at all is rarely the right choice outside of a throwaway experiment — it means the exact same `pip install -r requirements.txt` command can install different code on different days.
Use `pip show <package>` inside the activated virtual environment from Step 1 to find the exact version currently installed and resolved as compatible with everything else, and use that as your pin — it is the version you have actually tested against, not a guess.
pip show pandas # prints the exact installed version, e.g. Version: 2.1.4 — use this as the pin, not a guesspip show scikit-learnpip show numpyStep 5: Write the Cleaned requirements.txt
Combining what `audit_imports.py` reported with the version-pinning judgment from Step 4 produces the final file below. Every line carries a comment explaining why it survived the audit — including `pytest`, which the import scanner correctly reports as "unused" in `analyze.py`/`report.py` because tests are not application code, but which still belongs in the file for a different, equally valid reason.
pandas==2.1.4 # used in analyze.py and report.py for loading and describing datanumpy==1.26.2 # used indirectly via pandas/scikit-learn, and directly for X/y array shaping in analyze.py — bumped from the stale 1.21.0 pinmatplotlib==3.8.2 # used in report.py to render and save the summary chartseaborn==0.13.0 # used in analyze.py for the scatter plot — was imported in code but MISSING from the original file entirelyscikit-learn==1.3.2 # used in analyze.py for LinearRegression — pinned exactly instead of the original loose >=0.24 floorrequests==2.31.0 # used in analyze.py to fetch remote data before loading it into pandaspytest==7.4.3 # not imported by application code, but required to run this project's test suite — kept, with this comment explaining why
# Removed, with reasons:# plotly -> never imported anywhere; matplotlib already covers this project's one chart# beautifulsoup4 -> never imported anywhere; no HTML scraping happens in this project# flask -> never imported anywhere; this project has no HTTP server or APIComplete Code
Here is the complete `audit_imports.py` tool from Step 3, ready to drop into any Python project directory and run with `python audit_imports.py`.
import reimport glob
IMPORT_TO_PACKAGE = { "sklearn": "scikit-learn", "bs4": "beautifulsoup4", "cv2": "opencv-python",}
IMPORT_PATTERN = re.compile(r"^(?:import|from)\s+([a-zA-Z0-9_]+)")
def find_imports(filenames): """Scan the given .py files and return the set of top-level packages they import.""" imports = set() for filename in filenames: with open(filename, "r") as f: for line in f: match = IMPORT_PATTERN.match(line.strip()) if match: module = match.group(1) package = IMPORT_TO_PACKAGE.get(module, module) imports.add(package) return imports
def parse_requirements(path): """Read requirements.txt and return the set of package names it declares, ignoring version pins.""" declared = set() with open(path, "r") as f: for line in f: line = line.strip() if not line or line.startswith("#"): continue name = re.split(r"[=<>]", line)[0].strip() declared.add(name) return declared
if __name__ == "__main__": source_files = glob.glob("*.py") source_files = [f for f in source_files if f != "audit_imports.py"]
used = find_imports(source_files) declared = parse_requirements("requirements.txt")
print(f"Scanning {len(source_files)} source files for imports...\n")
print("Used in code AND declared in requirements.txt (keep):") for name in sorted(used & declared): print(f" {name}")
print("\nUsed in code but MISSING from requirements.txt (add these):") for name in sorted(used - declared): print(f" {name}")
print("\nListed in requirements.txt but NOT imported anywhere (candidates to remove):") for name in sorted(declared - used): print(f" {name}")Sample Run
Click Run to see what this code prints.
Extend This Project
- Split the cleaned file into `requirements.txt` (runtime) and `requirements-dev.txt` (`pytest`, linters), matching how `pytest` was flagged as a special case above.
- Extend `audit_imports.py` to recurse into subdirectories with `glob.glob("**/*.py", recursive=True)` instead of only scanning the top-level folder.
- Add a check that flags any `requirements.txt` line with no version pin at all, not just genuinely unused packages.
- Run `pip list --outdated` inside the virtual environment and decide, package by package, which updates are safe to pull in.
- Wire `audit_imports.py` into a pre-commit hook or CI step so a genuinely unused or undeclared import gets caught automatically before it ships.
Summary
You audited a dependency list the way experienced data scientists actually do it: isolate the environment first so nothing outside the project can mask a problem, gather evidence about what is genuinely imported instead of guessing, and only then decide what to keep, pin, add, or remove — with a written reason for every line. That process matters more than any specific library in this file: the same virtual-environment-plus-import-scan approach applies to any Python project's `requirements.txt`, data science or otherwise.