LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 621 min read

Core Data Manipulation Libraries (NumPy & pandas)

Learn NumPy for fast numerical arrays and pandas for labeled tabular data, with working examples of each.

Introduction

NumPy and pandas are the two libraries almost every other data science tool in Python is built on top of. NumPy provides fast, memory-efficient numerical arrays; pandas builds labeled, spreadsheet-like tables on top of those arrays. Nearly every catalog entry later in this course either depends on one of them directly or shares their conventions.

What You Will Learn
  • What NumPy solves and how to install and use its ndarray.
  • What pandas solves and how to filter and group data with a DataFrame.
  • A direct comparison of when to reach for a raw NumPy array versus a pandas DataFrame.

NumPy: Fast Numerical Arrays

Use case: NumPy provides the ndarray, a fixed-type, contiguous-memory array that supports fast vectorized math — element-wise operations, linear algebra, and broadcasting — without writing explicit Python loops.

pip install numpy

NumPy in Action

import numpy as np
prices = np.array([19.99, 45.50, 12.00, 89.99, 5.25])
quantities = np.array([3, 1, 10, 2, 20])
# Vectorized element-wise multiplication -- no loop needed
totals = prices * quantities
revenue = totals.sum()
print("Line totals:", totals)
print("Total revenue:", round(revenue, 2))
print("Average price:", round(prices.mean(), 2))
Terminal Output

Click Run to see what this code prints.

Why This Is Fast

The multiplication of prices and quantities happens as a single vectorized operation across the whole array in compiled C code, instead of Python looping element by element. This is the same principle behind the 150x speedup shown in lesson 2.

pandas: Labeled Tabular Data

Use case: pandas provides the DataFrame, a table of rows and columns with labels, mixed data types per column, and built-in methods for filtering, grouping, merging, and reshaping — the standard tool for working with structured, spreadsheet-like data in Python.

pip install pandas

pandas in Action

import pandas as pd
df = pd.DataFrame({
'region': ['North', 'South', 'North', 'East', 'South', 'East'],
'product': ['Widget', 'Widget', 'Gadget', 'Widget', 'Gadget', 'Gadget'],
'sales': [1200, 950, 1800, 600, 1100, 750],
})
# Filter: only Widget sales
widgets = df[df['product'] == 'Widget']
# Group + aggregate: total sales per region
by_region = df.groupby('region')['sales'].sum().sort_values(ascending=False)
print(widgets)
print()
print(by_region)
Terminal Output

Click Run to see what this code prints.

NumPy Arrays vs pandas DataFrames: When to Use Which

AspectNumPy ndarraypandas DataFrame
Best forPure numerical data, matrices, math-heavy codeLabeled, mixed-type, tabular data
Column labelsNoYes
Mixed data types per columnNo — single dtype for the whole arrayYes
Missing value handlingManual (NaN as float)Built in (isna, fillna, dropna)
Performance for raw mathSlightly faster, less overheadSlightly more overhead due to indexing/labels
Typical useUnder the hood of ML models, image data, matricesLoading, cleaning, and exploring real-world datasets

In practice, pandas DataFrames are built on top of NumPy arrays internally — when you call .values or .to_numpy() on a DataFrame, you get the underlying NumPy array back. Most projects use both together: pandas for loading and cleaning, NumPy for the numerical heavy lifting underneath.

Common Mistakes

Avoid These Mistakes
  • Using a pandas DataFrame for pure numerical computation where a NumPy array would be simpler and faster.
  • Looping over DataFrame rows with .iterrows() instead of using vectorized column operations.
  • Forgetting that NumPy arrays require a single consistent data type, which can silently upcast integers to floats.

Best Practices

  • Use pandas for anything that arrives as a CSV, spreadsheet, or database table with named columns.
  • Drop down to NumPy arrays for tight numerical loops or when feeding data into libraries that expect raw arrays.
  • Prefer df.groupby(), df.merge(), and vectorized boolean filtering over manual iteration in pandas.

Frequently Asked Questions

Not strictly, but it helps — pandas is built on NumPy, and understanding vectorization in NumPy makes pandas' behavior easier to reason about.

Yes — DataFrame.to_numpy() converts to a NumPy array, and pd.DataFrame(array) converts a NumPy array back into a DataFrame.

It can be, since it is single-threaded and memory-bound. Lesson 7 covers Polars and Dask, which address exactly this limitation.

Key Takeaways

  • NumPy provides fast, vectorized numerical arrays via the ndarray.
  • pandas provides labeled, tabular data structures via the DataFrame, built on top of NumPy.
  • Use NumPy for pure numerical/matrix work; use pandas for labeled, mixed-type, real-world tabular data.
  • Most real projects use both together, with pandas handling the outer layer and NumPy underneath.

Summary

NumPy and pandas together form the foundation of the entire Python data science ecosystem. Nearly everything else in this course, from visualization to machine learning, either consumes their data structures directly or borrows their conventions.

Lesson 6 Completed
  • You can install and use NumPy's ndarray for vectorized math.
  • You can install and use pandas' DataFrame to filter and group data.
  • You know when to reach for each one.
Next Lesson →

Modern & High-Performance Data Libraries (Polars & Dask)