Core Data Manipulation Libraries (NumPy & pandas)
Learn NumPy for fast numerical arrays and pandas for labeled tabular data, with working examples of each.
Introduction
NumPy and pandas are the two libraries almost every other data science tool in Python is built on top of. NumPy provides fast, memory-efficient numerical arrays; pandas builds labeled, spreadsheet-like tables on top of those arrays. Nearly every catalog entry later in this course either depends on one of them directly or shares their conventions.
- What NumPy solves and how to install and use its ndarray.
- What pandas solves and how to filter and group data with a DataFrame.
- A direct comparison of when to reach for a raw NumPy array versus a pandas DataFrame.
NumPy: Fast Numerical Arrays
Use case: NumPy provides the ndarray, a fixed-type, contiguous-memory array that supports fast vectorized math — element-wise operations, linear algebra, and broadcasting — without writing explicit Python loops.
pip install numpyNumPy in Action
import numpy as np
prices = np.array([19.99, 45.50, 12.00, 89.99, 5.25])quantities = np.array([3, 1, 10, 2, 20])
# Vectorized element-wise multiplication -- no loop neededtotals = prices * quantitiesrevenue = totals.sum()
print("Line totals:", totals)print("Total revenue:", round(revenue, 2))print("Average price:", round(prices.mean(), 2))Click Run to see what this code prints.
The multiplication of prices and quantities happens as a single vectorized operation across the whole array in compiled C code, instead of Python looping element by element. This is the same principle behind the 150x speedup shown in lesson 2.
pandas: Labeled Tabular Data
Use case: pandas provides the DataFrame, a table of rows and columns with labels, mixed data types per column, and built-in methods for filtering, grouping, merging, and reshaping — the standard tool for working with structured, spreadsheet-like data in Python.
pip install pandaspandas in Action
import pandas as pd
df = pd.DataFrame({ 'region': ['North', 'South', 'North', 'East', 'South', 'East'], 'product': ['Widget', 'Widget', 'Gadget', 'Widget', 'Gadget', 'Gadget'], 'sales': [1200, 950, 1800, 600, 1100, 750],})
# Filter: only Widget saleswidgets = df[df['product'] == 'Widget']
# Group + aggregate: total sales per regionby_region = df.groupby('region')['sales'].sum().sort_values(ascending=False)
print(widgets)print()print(by_region)Click Run to see what this code prints.
NumPy Arrays vs pandas DataFrames: When to Use Which
| Aspect | NumPy ndarray | pandas DataFrame |
|---|---|---|
| Best for | Pure numerical data, matrices, math-heavy code | Labeled, mixed-type, tabular data |
| Column labels | No | Yes |
| Mixed data types per column | No — single dtype for the whole array | Yes |
| Missing value handling | Manual (NaN as float) | Built in (isna, fillna, dropna) |
| Performance for raw math | Slightly faster, less overhead | Slightly more overhead due to indexing/labels |
| Typical use | Under the hood of ML models, image data, matrices | Loading, cleaning, and exploring real-world datasets |
In practice, pandas DataFrames are built on top of NumPy arrays internally — when you call .values or .to_numpy() on a DataFrame, you get the underlying NumPy array back. Most projects use both together: pandas for loading and cleaning, NumPy for the numerical heavy lifting underneath.
Common Mistakes
- Using a pandas DataFrame for pure numerical computation where a NumPy array would be simpler and faster.
- Looping over DataFrame rows with .iterrows() instead of using vectorized column operations.
- Forgetting that NumPy arrays require a single consistent data type, which can silently upcast integers to floats.
Best Practices
- Use pandas for anything that arrives as a CSV, spreadsheet, or database table with named columns.
- Drop down to NumPy arrays for tight numerical loops or when feeding data into libraries that expect raw arrays.
- Prefer df.groupby(), df.merge(), and vectorized boolean filtering over manual iteration in pandas.
Frequently Asked Questions
Not strictly, but it helps — pandas is built on NumPy, and understanding vectorization in NumPy makes pandas' behavior easier to reason about.
Yes — DataFrame.to_numpy() converts to a NumPy array, and pd.DataFrame(array) converts a NumPy array back into a DataFrame.
It can be, since it is single-threaded and memory-bound. Lesson 7 covers Polars and Dask, which address exactly this limitation.
Key Takeaways
- NumPy provides fast, vectorized numerical arrays via the ndarray.
- pandas provides labeled, tabular data structures via the DataFrame, built on top of NumPy.
- Use NumPy for pure numerical/matrix work; use pandas for labeled, mixed-type, real-world tabular data.
- Most real projects use both together, with pandas handling the outer layer and NumPy underneath.
Summary
NumPy and pandas together form the foundation of the entire Python data science ecosystem. Nearly everything else in this course, from visualization to machine learning, either consumes their data structures directly or borrows their conventions.
- You can install and use NumPy's ndarray for vectorized math.
- You can install and use pandas' DataFrame to filter and group data.
- You know when to reach for each one.