LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 2320 min read

Model Serving & APIs (FastAPI & Gradio)

Learn how to wrap a trained model in a production REST API with FastAPI, and how to build a shareable demo UI in minutes with Gradio.

Introduction

A trained model that only exists inside a Jupyter notebook has not delivered any value yet — someone still has to be able to use it. This lesson covers the two most common ways Python data scientists turn a trained model into something other people or systems can actually call: a production REST API with FastAPI, or a quick shareable demo with Gradio.

What You Will Learn
  • How to wrap a trained scikit-learn model in a FastAPI prediction endpoint.
  • How to build a shareable web demo for a model in a few lines with Gradio.
  • When to reach for a production API versus a quick demo UI.

fastapi: Production Model APIs

FastAPI is a modern web framework for building REST APIs in Python. It is the standard choice for serving a trained model in production because it is fast, generates interactive API documentation automatically, and uses type hints to validate incoming request data before your code ever sees it — catching malformed requests before they reach your model.

pip install fastapi uvicorn scikit-learn joblib

Assume a churn-prediction model has already been trained and saved to disk with joblib. The API below loads it once at startup and exposes a /predict endpoint.

main.py
from fastapi import FastAPI
from pydantic import BaseModel
import joblib
import numpy as np
app = FastAPI()
model = joblib.load("churn_model.pkl")
class CustomerFeatures(BaseModel):
tenure_months: float
monthly_charges: float
support_tickets: int
@app.post("/predict")
def predict(features: CustomerFeatures):
X = np.array([[
features.tenure_months,
features.monthly_charges,
features.support_tickets,
]])
prediction = model.predict(X)[0]
probability = model.predict_proba(X)[0][1]
return {
"will_churn": bool(prediction),
"churn_probability": round(float(probability), 3),
}
uvicorn main:app --reload
Example request
curl -X POST "http://127.0.0.1:8000/predict" \
-H "Content-Type: application/json" \
-d '{"tenure_months": 3, "monthly_charges": 89.5, "support_tickets": 4}'
Response

Click Run to see what this code prints.

Free Interactive Docs

Visiting /docs on a running FastAPI app gives you a full interactive Swagger UI for every endpoint, generated automatically from your type hints — no extra work required.

gradio: Shareable Demo UIs

Gradio solves a different problem: getting a model in front of a non-technical stakeholder — a product manager, a client, a teammate — without building any frontend at all. A few lines of Gradio wrap a Python function in a working web UI, complete with input fields and a public shareable link.

pip install gradio
demo.py
import gradio as gr
import joblib
import numpy as np
model = joblib.load("churn_model.pkl")
def predict_churn(tenure_months, monthly_charges, support_tickets):
X = np.array([[tenure_months, monthly_charges, support_tickets]])
probability = model.predict_proba(X)[0][1]
return f"Churn probability: {probability:.1%}"
demo = gr.Interface(
fn=predict_churn,
inputs=[
gr.Number(label="Tenure (months)"),
gr.Number(label="Monthly Charges"),
gr.Number(label="Support Tickets"),
],
outputs=gr.Textbox(label="Prediction"),
title="Churn Predictor Demo",
)
demo.launch()
Terminal Output

Click Run to see what this code prints.

FastAPI vs Gradio, in One Line

FastAPI builds a production endpoint other software calls programmatically. Gradio builds a demo UI a human clicks through in a browser. Many projects use both — FastAPI for the real service, Gradio for a quick internal demo of the same model.

Common Mistakes

Avoid These Mistakes
  • Loading the model inside the /predict function instead of once at startup, which reloads it from disk on every single request.
  • Skipping input validation and trusting raw request bodies, instead of using Pydantic models like FastAPI encourages.
  • Treating a Gradio demo link as production infrastructure — the free public gradio.live link is temporary and not meant for real traffic.
  • Not pinning the scikit-learn version used to save a model versus the version used to load it, which can silently break predictions.

Best Practices

  • Load models once at application startup, not per-request.
  • Use Pydantic models to define and validate the exact shape of incoming request data.
  • Return prediction probabilities alongside the class label, not just a bare true/false.
  • Use Gradio for internal demos and stakeholder buy-in; use FastAPI behind proper infrastructure for real production traffic.

Frequently Asked Questions

It is possible, and Gradio apps can be hosted persistently (for example on Hugging Face Spaces), but it is designed for demos and light interactive tools rather than high-throughput production APIs — FastAPI is the better fit there.

Not for the API itself — FastAPI only serves JSON. If you need a full interactive frontend, you would build one separately (in React, or with Gradio/Streamlit for something lighter) and have it call the FastAPI endpoints.

Pydantic validates types automatically and returns a clear error response if a field is missing or the wrong type, before your prediction code ever runs — this catches malformed input early instead of causing a confusing crash deeper in the code.

Key Takeaways

  • FastAPI wraps a trained model in a production REST API with automatic validation and docs.
  • Gradio builds a shareable demo UI for a model in just a few lines of Python.
  • Load models once at startup in both, never on every request.
  • Use FastAPI for real services, Gradio for demos and stakeholder-facing prototypes.

Summary

FastAPI and Gradio are the two most common ways a trained model reaches an actual user — one as a production endpoint, one as a demo. Next, we look at the tools that orchestrate entire pipelines of steps like this one, running on a schedule.

Lesson 23 Completed
  • You can serve a scikit-learn model behind a FastAPI /predict endpoint.
  • You can build a shareable Gradio demo for the same model.
  • You know when to reach for each tool.
Next Lesson →

Workflow Orchestration (Airflow & Prefect)