The ML Engines Under Generative AI: PyTorch & TensorFlow
Understand, at an applied level, how PyTorch and TensorFlow power the models behind every provider SDK in this course.
Introduction
Every model behind every provider SDK in this course — GPT, Claude, Gemini, Llama — was trained using a deep learning framework, almost always PyTorch or TensorFlow. You will not train a foundation model yourself, but recognizing these two names and what they represent explains a lot of vocabulary you'll see across the rest of this ecosystem: "tensors," "checkpoints," ".safetensors" files, and more.
- What PyTorch and TensorFlow are, at the level needed to understand the rest of this course.
- A minimal tensor example in each, just enough to recognize the vocabulary elsewhere.
- Where to go for a full, dedicated deep dive into either framework.
A Real-Life Analogy First
Think of a car manufacturer's factory versus the car you actually drive. PyTorch and TensorFlow are the factory floor — the heavy machinery, assembly lines, and engineering processes used to build an engine (the model). You, as a driver (a developer using a provider SDK), never see the factory floor at all. You just get in the finished car and turn the key. This lesson is a short tour of the factory, purely so the parts list in the finished car's manual — "tensor," "gradient," "checkpoint" — stops sounding like a foreign language.
PyTorch: The Dominant Research Framework
Use case: PyTorch is the framework almost every recent foundation model (including the open-weight models you'll run locally in lesson 18) was built and trained with. It represents data as tensors — multi-dimensional arrays similar to NumPy's, but with automatic differentiation and GPU acceleration built in.
pip install torchPyTorch in Action
import torch
# A tensor is PyTorch's core data structure -- like a NumPy array,# but it can track gradients and run on a GPU.weights = torch.tensor([0.2, 0.5, 0.3], requires_grad=True)inputs = torch.tensor([10.0, 20.0, 5.0])
output = (weights * inputs).sum()output.backward() # computes how much each weight contributed
print("Output:", output.item())print("Gradients:", weights.grad)Click Run to see what this code prints.
.backward() is the mechanism behind every model's training loop: it automatically computes how much each weight contributed to the output, which is exactly what gets used to adjust the weights during training. This single mechanism, run billions of times, is how a model like GPT learns — similar to a student redoing a practice problem, checking exactly which step they got wrong, and adjusting only that step next time, repeated billions of times across a huge stack of practice problems.
TensorFlow: The Production-Focused Framework
Use case: TensorFlow (often used through its high-level Keras API) covers the same core idea as PyTorch — tensors, automatic differentiation, GPU acceleration — but has historically leaned more toward production deployment tooling (TensorFlow Serving, TensorFlow Lite for mobile/edge).
pip install tensorflowTensorFlow in Action
import tensorflow as tf
weights = tf.Variable([0.2, 0.5, 0.3])inputs = tf.constant([10.0, 20.0, 5.0])
with tf.GradientTape() as tape: output = tf.reduce_sum(weights * inputs)
grads = tape.gradient(output, weights)
print("Output:", output.numpy())print("Gradients:", grads.numpy())Click Run to see what this code prints.
Why This Lesson Stays Shallow on Purpose
This course is about the generative AI application layer, not model training — so this lesson intentionally stops at "recognize the vocabulary and see the shape of a tensor operation." If you want the full, dedicated treatment of PyTorch and TensorFlow, including real training loops, this platform's Data Science Dependencies course covers both in depth across two full lessons.
Common Mistakes
- Assuming you need to learn full model training to build with generative AI — the vast majority of application work never touches a training loop.
- Installing the GPU-enabled build of either framework without a compatible GPU driver set up, which leads to confusing installation errors.
- Confusing "tensor" (the data structure) with "Transformer" (the model architecture, covered in the next lesson) — the names are unrelated in origin.
Best Practices
- Learn just enough of one framework's vocabulary (tensors, gradients, checkpoints) to read a model's documentation, unless training is actually your goal.
- Install the CPU-only build of either framework unless you specifically know you have a compatible GPU set up.
- Treat this lesson as a bridge, not a destination — the next lesson picks the application-layer thread back up.
Frequently Asked Questions
No — those SDKs only make HTTP calls. You would only touch PyTorch or TensorFlow directly if you were fine-tuning or running an open-weight model yourself (lessons 17–18).
PyTorch is the more common choice for generative AI specifically, since most open-weight models (Llama, Mistral, and others) ship PyTorch weights first.
Keras is TensorFlow's official high-level API — you'll usually see them used together, with tf.keras providing the friendlier interface shown conceptually above.
It is a measurement of "if I nudge this one number slightly, how much does the final answer change, and in which direction." Training a model is just repeatedly measuring that for every weight and nudging each one slightly in the direction that improves the result — the .backward() call above computes exactly this.
Key Takeaways
- PyTorch and TensorFlow are the deep learning frameworks nearly every model behind this course's provider SDKs was built with.
- Both represent data as tensors and support automatic differentiation for computing gradients.
- PyTorch is the more common choice specifically within the open-weight generative AI ecosystem.
- This course intentionally stays applied — see Data Science Dependencies for a full deep dive into either framework.
Summary
You do not need to train models to build with generative AI, but recognizing PyTorch and TensorFlow — and the vocabulary they introduced — makes the rest of this ecosystem's documentation far less intimidating.
- You understand what PyTorch and TensorFlow are used for.
- You have seen a minimal tensor and gradient example in each.
- You know where to go for a full deep dive if you want one.