LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1020 min read

Scientific Computing & Statistics (SciPy & Statsmodels)

Learn how SciPy extends NumPy with optimization, linear algebra, and statistical functions, and how Statsmodels provides R-style statistical modeling with detailed summary tables.

Introduction

NumPy gives you fast arrays and basic linear algebra, but most real data science work needs more: optimization routines, statistical tests, signal processing, and interpretable statistical models with p-values and confidence intervals. That is where SciPy and Statsmodels come in.

In this lesson you will install both libraries, run a real hypothesis test with scipy.stats, and fit a linear regression with statsmodels that prints a full R-style summary table.

What You Will Learn
  • What SciPy adds on top of NumPy, and when you reach for it.
  • How to run a two-sample t-test with scipy.stats.
  • What Statsmodels is used for and how it differs from scikit-learn.
  • How to fit and interpret an OLS linear regression summary.
  • When to pick SciPy, Statsmodels, or scikit-learn for a given task.

What is SciPy?

SciPy (Scientific Python) is a library built directly on top of NumPy arrays. Where NumPy gives you the array data structure and basic operations, SciPy adds entire submodules of specialized algorithms: scipy.optimize for optimization and curve fitting, scipy.linalg for advanced linear algebra, scipy.signal for signal processing, scipy.integrate for numerical integration, and scipy.stats for probability distributions and statistical tests.

In practice, if you find yourself needing a statistical test, an optimizer, or an algorithm from a numerical methods textbook, SciPy almost certainly already has it implemented and tested.

Installing SciPy

SciPy is installed with pip like any other package. It depends on NumPy, which pip will install automatically if it is not already present.

pip install scipy

Example: A t-test with scipy.stats

A very common real-world task is comparing two groups of numbers to see if their averages are meaningfully different, or if the difference could just be random noise. An independent two-sample t-test answers exactly this question.

from scipy import stats
# Test scores for two different teaching methods
method_a = [78, 82, 88, 91, 74, 86, 79, 84]
method_b = [85, 90, 95, 92, 88, 91, 89, 94]
t_statistic, p_value = stats.ttest_ind(method_a, method_b)
print(f"t-statistic: {t_statistic:.3f}")
print(f"p-value: {p_value:.4f}")
if p_value < 0.05:
print("The difference between the groups is statistically significant.")
else:
print("No statistically significant difference was found.")
Output

Click Run to see what this code prints.

A p-value below 0.05 conventionally means the observed difference is unlikely to be pure chance, so here method B's scores are significantly higher than method A's. scipy.stats has dozens of similar tests: chi-square, ANOVA, correlation tests, and more, all following the same pattern of returning a statistic and a p-value.

What is Statsmodels?

Statsmodels focuses on statistical modeling rather than prediction accuracy alone. Where scikit-learn is built for training a model and generating predictions efficiently, statsmodels is built for understanding a model: it produces detailed output tables with coefficients, standard errors, t-statistics, p-values, and confidence intervals, similar to what you would get from R.

This makes statsmodels the go-to library when the goal is explaining a relationship in data (for example, in economics, social science, or A/B test analysis) rather than only maximizing predictive accuracy.

Installing Statsmodels

pip install statsmodels

Example: OLS Linear Regression

Ordinary Least Squares (OLS) regression is the classic starting point for statistical modeling. The example below fits hours studied against exam scores and prints the full summary table.

import statsmodels.api as sm
import numpy as np
hours_studied = np.array([1, 2, 3, 4, 5, 6, 7, 8])
exam_score = np.array([52, 58, 63, 68, 74, 78, 85, 90])
# statsmodels requires an explicit constant (intercept) term
X = sm.add_constant(hours_studied)
model = sm.OLS(exam_score, X)
results = model.fit()
print(results.summary())
Output (abridged)

Click Run to see what this code prints.

The coefficient on x1 (5.4762) means each additional hour studied is associated with roughly 5.48 more exam points, and the near-zero P>|t| value tells you this relationship is highly statistically significant.

SciPy vs Statsmodels vs Scikit-learn

LibraryPrimary FocusTypical Output
SciPyGeneral scientific computing: optimization, integration, individual stat testsA single statistic and p-value, or an optimized array
StatsmodelsStatistical modeling and inferenceDetailed summary tables with coefficients, p-values, confidence intervals
Scikit-learnPredictive machine learningA fitted model object used mainly to call .predict()

Common Mistakes

Avoid These Mistakes
  • Forgetting sm.add_constant() before fitting an OLS model in statsmodels — without it, the regression is forced through the origin.
  • Treating a low p-value as proof of a large or important effect, rather than just statistical significance.
  • Reaching for statsmodels when the actual goal is prediction accuracy on new data — scikit-learn is usually the better fit there.

Best Practices

  • Use scipy.stats for quick, single-purpose statistical tests.
  • Use statsmodels when you need to explain a relationship, not just predict an outcome.
  • Always read the confidence interval alongside the p-value, not just the p-value alone.
  • Pin exact versions of scipy and statsmodels in requirements.txt since statistical APIs occasionally change between major versions.

Frequently Asked Questions

Yes, if you need anything beyond basic array math. NumPy handles arrays and simple linear algebra; SciPy adds optimization, advanced statistics, signal processing, and more.

No. They solve different problems. Statsmodels is for statistical inference and interpretation; scikit-learn is for building predictive models efficiently. Many real projects use both.

Yes. Statsmodels accepts pandas DataFrames and Series directly, and its formula API (statsmodels.formula.api) even lets you write regressions using R-style formula strings like 'score ~ hours_studied'.

Key Takeaways

  • SciPy extends NumPy with optimization, linear algebra, and statistical test functions.
  • scipy.stats.ttest_ind() runs a two-sample t-test and returns a statistic and a p-value.
  • Statsmodels specializes in statistical modeling with detailed, interpretable summary output.
  • sm.OLS() fits linear regression and results.summary() prints coefficients, p-values, and confidence intervals.
  • Choose SciPy for individual tests, statsmodels for interpretable models, and scikit-learn for predictive machine learning.

Summary

SciPy and statsmodels round out the classic Python statistics toolkit: SciPy for general scientific computing and quick statistical tests, and statsmodels for detailed, explainable statistical models. Together with NumPy and pandas, they form the foundation most other data science libraries are built on.

Lesson 10 Completed
  • You installed and used scipy.stats to run a hypothesis test.
  • You fit an OLS regression with statsmodels and read its summary table.
  • You are ready to move into machine learning with scikit-learn.
Next Lesson →

Classical Machine Learning (Scikit-learn)