LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1122 min read

Classical Machine Learning (Scikit-learn)

Learn scikit-learn, the standard Python toolkit for classical machine learning covering classification, regression, clustering, and preprocessing, with a full train/predict/evaluate example.

Introduction

If pandas is how data scientists load and clean data, scikit-learn is how they turn it into predictions. It is the single most widely used machine learning library in Python for classical (non-deep-learning) algorithms, and its API design has been copied by nearly every library that came after it.

This lesson covers what scikit-learn provides, how it is organized into modules, and walks through a complete, realistic example: splitting data, training a model, generating predictions, and measuring accuracy.

What You Will Learn
  • What scikit-learn is and the range of problems it solves.
  • How its module system is organized (model_selection, preprocessing, metrics, ensemble, and more).
  • How to split data into training and test sets.
  • How to train a RandomForestClassifier and generate predictions.
  • How to measure model accuracy with accuracy_score.

What is Scikit-learn?

Scikit-learn (imported as sklearn) is a general-purpose machine learning library covering the three classic problem types: classification (predicting a category), regression (predicting a number), and clustering (finding groups in unlabeled data). It also provides the surrounding tooling every ML project needs: splitting data, scaling features, encoding categories, and scoring model performance.

It is built on top of NumPy and SciPy, and it deliberately does not handle deep learning — for neural networks you would reach for TensorFlow or PyTorch instead, covered later in this course. Scikit-learn's strength is the huge range of well-tested, well-documented classical algorithms it gives you through one consistent interface.

Installing Scikit-learn

The pip package name is scikit-learn (with hyphens), but you import it in Python as sklearn.

pip install scikit-learn

Key Modules

Scikit-learn is organized into focused submodules rather than one flat namespace. Knowing the main ones makes the documentation much easier to navigate.

ModulePurpose
sklearn.model_selectionSplitting data (train_test_split), cross-validation, and hyperparameter search (GridSearchCV)
sklearn.preprocessingScaling and encoding features: StandardScaler, MinMaxScaler, OneHotEncoder, LabelEncoder
sklearn.metricsEvaluating models: accuracy_score, precision_score, confusion_matrix, mean_squared_error
sklearn.ensembleEnsemble algorithms that combine many models: RandomForestClassifier, GradientBoostingClassifier

Example: Training a Classifier

The example below uses the classic Iris dataset (bundled with scikit-learn) to train a random forest that classifies flowers into species based on petal and sepal measurements.

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# Load a built-in dataset: features (X) and labels (y)
iris = load_iris()
X, y = iris.data, iris.target
# Split into 80% training data, 20% test data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# Create and train the model
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)
# Generate predictions on data the model has never seen
predictions = model.predict(X_test)
# Compare predictions to the true labels
accuracy = accuracy_score(y_test, predictions)
print(f"Model accuracy: {accuracy:.2%}")
Output

Click Run to see what this code prints.

The Iris dataset is small and clean, so near-perfect accuracy here is expected — real-world datasets are almost always messier. What matters is the pattern: load data, split it, fit a model, predict, and score.

The Scikit-learn API Pattern

Nearly every scikit-learn model, whether it is a RandomForestClassifier, a LinearRegression, or a KMeans clusterer, follows the exact same three-method pattern, which is a big part of why the library is so easy to learn once.

  • .fit(X, y) — train the model on labeled data (for clustering, just .fit(X)).
  • .predict(X) — generate predictions for new, unseen data.
  • .score(X, y) — a quick built-in accuracy or R² measurement, separate from the metrics module.

Common Mistakes

Avoid These Mistakes
  • Evaluating a model on the same data it was trained on — always hold out a test set with train_test_split.
  • Forgetting to scale features (with StandardScaler) for algorithms that are sensitive to feature scale, like SVMs or k-NN.
  • Not setting random_state, which makes results impossible to reproduce between runs.

Best Practices

  • Always split data into train and test sets before fitting a model.
  • Use cross-validation (sklearn.model_selection.cross_val_score) for a more reliable accuracy estimate than a single train/test split.
  • Wrap preprocessing and modeling steps together with sklearn.pipeline.Pipeline to avoid data leakage.
  • Check sklearn.metrics for the right evaluation metric for your problem — accuracy alone can be misleading on imbalanced datasets.

Frequently Asked Questions

It is a historical naming quirk. The project started as 'scikits.learn', and while the pip package name kept the scikit-learn form, the import name settled on the shorter sklearn.

No, scikit-learn runs entirely on CPU. For GPU-accelerated training on very large datasets you would look at libraries like XGBoost or LightGBM (covered in the next lesson) or a deep learning framework.

No. It is intentionally focused on classical machine learning. For neural networks, use TensorFlow/Keras or PyTorch, both covered later in this course.

Key Takeaways

  • Scikit-learn is the standard library for classical machine learning: classification, regression, clustering, and preprocessing.
  • It is organized into modules like model_selection, preprocessing, metrics, and ensemble.
  • Every model follows the same .fit() / .predict() / .score() pattern.
  • train_test_split() and accuracy_score() are used in almost every classification workflow.
  • Scikit-learn does not handle deep learning — that is the job of TensorFlow or PyTorch.

Summary

Scikit-learn is the workhorse of classical machine learning in Python: a huge, well-tested collection of algorithms behind one consistent, predictable API. Once you understand fit/predict/score, you can apply almost any algorithm in the library with minimal code changes.

Lesson 11 Completed
  • You understand what scikit-learn is used for and how it is organized.
  • You trained a RandomForestClassifier and evaluated it with accuracy_score.
  • You are ready to explore gradient boosting libraries that often outperform scikit-learn's built-in models.
Next Lesson →

Gradient Boosting Libraries (XGBoost, LightGBM & CatBoost)