LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1423 min read

Deep Learning Frameworks (PyTorch)

Learn PyTorch, the research-friendly, dynamic-graph deep learning framework, and build a minimal neural network with a manual training loop.

Introduction

PyTorch, developed by Meta AI, is the other dominant deep learning framework alongside TensorFlow, and today it is the default choice in most academic research and a large share of production systems, including much of the recent wave of large language models.

Where Keras hides most of the training loop behind .fit(), PyTorch expects you to write the loop yourself, giving you full visibility and control over every training step. This lesson builds a small network and trains it step by step so you can see exactly what is happening under the hood.

What You Will Learn
  • What PyTorch is and why it is popular in research.
  • How to install torch with pip.
  • How to define a model as a class using nn.Module.
  • How a manual training loop works: forward pass, loss, backward pass, optimizer step.
  • How PyTorch compares to TensorFlow/Keras.

What is PyTorch?

PyTorch is an open-source deep learning framework built around dynamic computation graphs — the network's structure is defined by the actual Python code that runs, rather than a graph you build separately and then execute. This makes debugging far more natural: you can set breakpoints, print tensor values mid-computation, and use ordinary Python control flow (if statements, loops) inside a model's forward pass.

This flexibility is a big reason PyTorch became the standard in academic research and is now heavily used in production too, particularly for natural language processing and generative AI, where model architectures change frequently.

Installing PyTorch

The pip package name is torch. For most learning and CPU-based work, the standard pip install is enough; GPU-accelerated builds are also available and are recommended by PyTorch's own install selector for your specific CUDA version.

pip install torch

Defining a Model with nn.Module

In PyTorch, a neural network is defined as a Python class that inherits from torch.nn.Module. The __init__ method declares the layers, and the forward method defines how data flows through them.

import torch
import torch.nn as nn
class SimpleNet(nn.Module):
def __init__(self):
super().__init__()
self.layer1 = nn.Linear(2, 16)
self.layer2 = nn.Linear(16, 8)
self.layer3 = nn.Linear(8, 1)
self.relu = nn.ReLU()
self.sigmoid = nn.Sigmoid()
def forward(self, x):
x = self.relu(self.layer1(x))
x = self.relu(self.layer2(x))
x = self.sigmoid(self.layer3(x))
return x
model = SimpleNet()
print(model)
Output

Click Run to see what this code prints.

Example: A Manual Training Loop

Unlike Keras, PyTorch does not have a built-in .fit(). You write the loop explicitly: run data through the model (forward pass), measure how wrong it was (loss), compute gradients (backward pass), and nudge the weights (optimizer step).

from sklearn.datasets import make_moons
import torch.optim as optim
# Toy dataset, converted to PyTorch tensors
X, y = make_moons(n_samples=500, noise=0.2, random_state=42)
X_tensor = torch.tensor(X, dtype=torch.float32)
y_tensor = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
loss_fn = nn.BCELoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
for epoch in range(20):
optimizer.zero_grad() # clear gradients from the last step
predictions = model(X_tensor) # forward pass
loss = loss_fn(predictions, y_tensor) # compute how wrong the predictions are
loss.backward() # backward pass: compute gradients
optimizer.step() # update weights using those gradients
if (epoch + 1) % 5 == 0:
print(f"Epoch {epoch + 1}, Loss: {loss.item():.4f}")
Output

Click Run to see what this code prints.

The steadily decreasing loss shows the model is learning: each epoch, the optimizer nudges the weights slightly in the direction that reduces prediction error, based on the gradients computed by loss.backward().

TensorFlow/Keras vs PyTorch

Choosing Between Them
  • Keras: training loop is hidden behind .fit() — faster to write, less to debug manually, great for standard architectures.
  • PyTorch: training loop is explicit — more code, but full control and easier debugging for custom or research architectures.
  • Both support GPU acceleration and can deploy to production; the choice often comes down to team convention and the surrounding ecosystem (many NLP/research libraries default to PyTorch).
  • It is common for data scientists to be comfortable in both, since a model trained in a paper's code is often written in one or the other.

Common Mistakes

Avoid These Mistakes
  • Forgetting optimizer.zero_grad() at the start of each loop iteration — gradients accumulate by default instead of resetting.
  • Mismatched tensor shapes, such as forgetting .unsqueeze(1) to turn a 1D label tensor into the 2D shape a loss function expects.
  • Leaving the model in training mode during evaluation, which affects layers like dropout — call model.eval() before running inference.

Best Practices

  • Wrap the four training-loop steps (zero_grad, forward, backward, step) into a small reusable function once you have several models to train.
  • Use torch.no_grad() when running inference, so PyTorch does not waste memory tracking gradients you will not use.
  • Move both the model and the data to the same device (model.to('cuda') and tensor.to('cuda')) when training on a GPU.
  • Save and load models with torch.save(model.state_dict(), ...) rather than pickling the entire model object.

Frequently Asked Questions

It has a slightly steeper learning curve initially because you write the training loop yourself, but many developers find it more intuitive long-term since every step is visible ordinary Python code rather than hidden inside a framework method.

Yes. Moving a model and its tensors to a GPU is done explicitly with .to('cuda'), and PyTorch's dynamic graph approach works the same way on GPU as on CPU.

Either is a reasonable start. Keras is often recommended first because it gets you a working model in fewer lines of code; PyTorch is worth learning next since it dominates current research and much of the NLP and generative AI ecosystem.

Key Takeaways

  • PyTorch uses dynamic computation graphs, making models easier to debug with ordinary Python tools.
  • Models are defined as classes inheriting from nn.Module, with layers in __init__ and data flow in forward().
  • Training is an explicit loop: zero_grad(), forward pass, loss.backward(), optimizer.step().
  • PyTorch dominates in research and much of NLP/generative AI; Keras favors speed of development for standard architectures.
  • Both frameworks support GPU acceleration and production deployment.

Summary

PyTorch trades some of Keras's convenience for full visibility into the training process, which is exactly why it has become the standard in research and a large share of production deep learning, especially for NLP and generative models. Next, you will see how Hugging Face builds on top of both frameworks to make state-of-the-art pretrained models usable in just a few lines of code.

Lesson 14 Completed
  • You understand PyTorch's dynamic-graph approach and how it differs from Keras.
  • You defined a model with nn.Module and wrote a manual training loop.
  • You are ready to use pretrained models with Hugging Face instead of training from scratch.
Next Lesson →

Pretrained Models & Transfer Learning (Hugging Face)