Classical Machine Learning (Scikit-learn)
Learn scikit-learn, the standard Python toolkit for classical machine learning covering classification, regression, clustering, and preprocessing, with a full train/predict/evaluate example.
Introduction
If pandas is how data scientists load and clean data, scikit-learn is how they turn it into predictions. It is the single most widely used machine learning library in Python for classical (non-deep-learning) algorithms, and its API design has been copied by nearly every library that came after it.
This lesson covers what scikit-learn provides, how it is organized into modules, and walks through a complete, realistic example: splitting data, training a model, generating predictions, and measuring accuracy.
- What scikit-learn is and the range of problems it solves.
- How its module system is organized (model_selection, preprocessing, metrics, ensemble, and more).
- How to split data into training and test sets.
- How to train a RandomForestClassifier and generate predictions.
- How to measure model accuracy with accuracy_score.
What is Scikit-learn?
Scikit-learn (imported as sklearn) is a general-purpose machine learning library covering the three classic problem types: classification (predicting a category), regression (predicting a number), and clustering (finding groups in unlabeled data). It also provides the surrounding tooling every ML project needs: splitting data, scaling features, encoding categories, and scoring model performance.
It is built on top of NumPy and SciPy, and it deliberately does not handle deep learning — for neural networks you would reach for TensorFlow or PyTorch instead, covered later in this course. Scikit-learn's strength is the huge range of well-tested, well-documented classical algorithms it gives you through one consistent interface.
Installing Scikit-learn
The pip package name is scikit-learn (with hyphens), but you import it in Python as sklearn.
pip install scikit-learnKey Modules
Scikit-learn is organized into focused submodules rather than one flat namespace. Knowing the main ones makes the documentation much easier to navigate.
| Module | Purpose |
|---|---|
| sklearn.model_selection | Splitting data (train_test_split), cross-validation, and hyperparameter search (GridSearchCV) |
| sklearn.preprocessing | Scaling and encoding features: StandardScaler, MinMaxScaler, OneHotEncoder, LabelEncoder |
| sklearn.metrics | Evaluating models: accuracy_score, precision_score, confusion_matrix, mean_squared_error |
| sklearn.ensemble | Ensemble algorithms that combine many models: RandomForestClassifier, GradientBoostingClassifier |
Example: Training a Classifier
The example below uses the classic Iris dataset (bundled with scikit-learn) to train a random forest that classifies flowers into species based on petal and sepal measurements.
from sklearn.datasets import load_irisfrom sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import accuracy_score
# Load a built-in dataset: features (X) and labels (y)iris = load_iris()X, y = iris.data, iris.target
# Split into 80% training data, 20% test dataX_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42)
# Create and train the modelmodel = RandomForestClassifier(n_estimators=100, random_state=42)model.fit(X_train, y_train)
# Generate predictions on data the model has never seenpredictions = model.predict(X_test)
# Compare predictions to the true labelsaccuracy = accuracy_score(y_test, predictions)print(f"Model accuracy: {accuracy:.2%}")Click Run to see what this code prints.
The Iris dataset is small and clean, so near-perfect accuracy here is expected — real-world datasets are almost always messier. What matters is the pattern: load data, split it, fit a model, predict, and score.
The Scikit-learn API Pattern
Nearly every scikit-learn model, whether it is a RandomForestClassifier, a LinearRegression, or a KMeans clusterer, follows the exact same three-method pattern, which is a big part of why the library is so easy to learn once.
- .fit(X, y) — train the model on labeled data (for clustering, just .fit(X)).
- .predict(X) — generate predictions for new, unseen data.
- .score(X, y) — a quick built-in accuracy or R² measurement, separate from the metrics module.
Common Mistakes
- Evaluating a model on the same data it was trained on — always hold out a test set with train_test_split.
- Forgetting to scale features (with StandardScaler) for algorithms that are sensitive to feature scale, like SVMs or k-NN.
- Not setting random_state, which makes results impossible to reproduce between runs.
Best Practices
- Always split data into train and test sets before fitting a model.
- Use cross-validation (sklearn.model_selection.cross_val_score) for a more reliable accuracy estimate than a single train/test split.
- Wrap preprocessing and modeling steps together with sklearn.pipeline.Pipeline to avoid data leakage.
- Check sklearn.metrics for the right evaluation metric for your problem — accuracy alone can be misleading on imbalanced datasets.
Frequently Asked Questions
It is a historical naming quirk. The project started as 'scikits.learn', and while the pip package name kept the scikit-learn form, the import name settled on the shorter sklearn.
No, scikit-learn runs entirely on CPU. For GPU-accelerated training on very large datasets you would look at libraries like XGBoost or LightGBM (covered in the next lesson) or a deep learning framework.
No. It is intentionally focused on classical machine learning. For neural networks, use TensorFlow/Keras or PyTorch, both covered later in this course.
Key Takeaways
- Scikit-learn is the standard library for classical machine learning: classification, regression, clustering, and preprocessing.
- It is organized into modules like model_selection, preprocessing, metrics, and ensemble.
- Every model follows the same .fit() / .predict() / .score() pattern.
- train_test_split() and accuracy_score() are used in almost every classification workflow.
- Scikit-learn does not handle deep learning — that is the job of TensorFlow or PyTorch.
Summary
Scikit-learn is the workhorse of classical machine learning in Python: a huge, well-tested collection of algorithms behind one consistent, predictable API. Once you understand fit/predict/score, you can apply almost any algorithm in the library with minimal code changes.
- You understand what scikit-learn is used for and how it is organized.
- You trained a RandomForestClassifier and evaluated it with accuracy_score.
- You are ready to explore gradient boosting libraries that often outperform scikit-learn's built-in models.