Model Serving & APIs (FastAPI & Gradio)
Learn how to wrap a trained model in a production REST API with FastAPI, and how to build a shareable demo UI in minutes with Gradio.
Introduction
A trained model that only exists inside a Jupyter notebook has not delivered any value yet — someone still has to be able to use it. This lesson covers the two most common ways Python data scientists turn a trained model into something other people or systems can actually call: a production REST API with FastAPI, or a quick shareable demo with Gradio.
- How to wrap a trained scikit-learn model in a FastAPI prediction endpoint.
- How to build a shareable web demo for a model in a few lines with Gradio.
- When to reach for a production API versus a quick demo UI.
fastapi: Production Model APIs
FastAPI is a modern web framework for building REST APIs in Python. It is the standard choice for serving a trained model in production because it is fast, generates interactive API documentation automatically, and uses type hints to validate incoming request data before your code ever sees it — catching malformed requests before they reach your model.
pip install fastapi uvicorn scikit-learn joblibAssume a churn-prediction model has already been trained and saved to disk with joblib. The API below loads it once at startup and exposes a /predict endpoint.
from fastapi import FastAPIfrom pydantic import BaseModelimport joblibimport numpy as np
app = FastAPI()model = joblib.load("churn_model.pkl")
class CustomerFeatures(BaseModel): tenure_months: float monthly_charges: float support_tickets: int
@app.post("/predict")def predict(features: CustomerFeatures): X = np.array([[ features.tenure_months, features.monthly_charges, features.support_tickets, ]]) prediction = model.predict(X)[0] probability = model.predict_proba(X)[0][1] return { "will_churn": bool(prediction), "churn_probability": round(float(probability), 3), }uvicorn main:app --reloadcurl -X POST "http://127.0.0.1:8000/predict" \ -H "Content-Type: application/json" \ -d '{"tenure_months": 3, "monthly_charges": 89.5, "support_tickets": 4}'Click Run to see what this code prints.
Visiting /docs on a running FastAPI app gives you a full interactive Swagger UI for every endpoint, generated automatically from your type hints — no extra work required.
gradio: Shareable Demo UIs
Gradio solves a different problem: getting a model in front of a non-technical stakeholder — a product manager, a client, a teammate — without building any frontend at all. A few lines of Gradio wrap a Python function in a working web UI, complete with input fields and a public shareable link.
pip install gradioimport gradio as grimport joblibimport numpy as np
model = joblib.load("churn_model.pkl")
def predict_churn(tenure_months, monthly_charges, support_tickets): X = np.array([[tenure_months, monthly_charges, support_tickets]]) probability = model.predict_proba(X)[0][1] return f"Churn probability: {probability:.1%}"
demo = gr.Interface( fn=predict_churn, inputs=[ gr.Number(label="Tenure (months)"), gr.Number(label="Monthly Charges"), gr.Number(label="Support Tickets"), ], outputs=gr.Textbox(label="Prediction"), title="Churn Predictor Demo",)
demo.launch()Click Run to see what this code prints.
FastAPI builds a production endpoint other software calls programmatically. Gradio builds a demo UI a human clicks through in a browser. Many projects use both — FastAPI for the real service, Gradio for a quick internal demo of the same model.
Common Mistakes
- Loading the model inside the /predict function instead of once at startup, which reloads it from disk on every single request.
- Skipping input validation and trusting raw request bodies, instead of using Pydantic models like FastAPI encourages.
- Treating a Gradio demo link as production infrastructure — the free public gradio.live link is temporary and not meant for real traffic.
- Not pinning the scikit-learn version used to save a model versus the version used to load it, which can silently break predictions.
Best Practices
- Load models once at application startup, not per-request.
- Use Pydantic models to define and validate the exact shape of incoming request data.
- Return prediction probabilities alongside the class label, not just a bare true/false.
- Use Gradio for internal demos and stakeholder buy-in; use FastAPI behind proper infrastructure for real production traffic.
Frequently Asked Questions
It is possible, and Gradio apps can be hosted persistently (for example on Hugging Face Spaces), but it is designed for demos and light interactive tools rather than high-throughput production APIs — FastAPI is the better fit there.
Not for the API itself — FastAPI only serves JSON. If you need a full interactive frontend, you would build one separately (in React, or with Gradio/Streamlit for something lighter) and have it call the FastAPI endpoints.
Pydantic validates types automatically and returns a clear error response if a field is missing or the wrong type, before your prediction code ever runs — this catches malformed input early instead of causing a confusing crash deeper in the code.
Key Takeaways
- FastAPI wraps a trained model in a production REST API with automatic validation and docs.
- Gradio builds a shareable demo UI for a model in just a few lines of Python.
- Load models once at startup in both, never on every request.
- Use FastAPI for real services, Gradio for demos and stakeholder-facing prototypes.
Summary
FastAPI and Gradio are the two most common ways a trained model reaches an actual user — one as a production endpoint, one as a demo. Next, we look at the tools that orchestrate entire pipelines of steps like this one, running on a schedule.
- You can serve a scikit-learn model behind a FastAPI /predict endpoint.
- You can build a shareable Gradio demo for the same model.
- You know when to reach for each tool.