LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1922 min read

High-Performance Serving (vLLM & TGI)

Learn vLLM and Text Generation Inference (TGI), the tools that serve open-weight models at production scale with high throughput.

Introduction

Ollama (lesson 18) is ideal for one developer running a model locally, but it is not built to serve hundreds of concurrent users efficiently. vLLM and Text Generation Inference (TGI) exist for exactly that gap — serving an open-weight model in production, at scale, with high throughput.

What You Will Learn
  • What changes between running a model for yourself and serving it to many concurrent users.
  • How to serve a model with vLLM behind an OpenAI-compatible API.
  • What TGI offers as Hugging Face's answer to the same problem.

A Real-Life Analogy First

Cooking dinner for yourself (lesson 18's Ollama) and running a restaurant kitchen that serves two hundred customers on a Friday night are both "cooking," but they require completely different setups. A home kitchen has one stove and cooks one dish at a time. A commercial kitchen is engineered around batching — starting several orders together, staggering steps so nothing sits idle, and organizing the line so the whole kitchen's throughput is far higher than one home cook working alone ever could be. vLLM and TGI are that commercial-kitchen engineering, applied to serving a model to many users at once.

From "It Runs" to "It Scales"

When many requests arrive at once, a naive serving setup processes them one at a time, wasting most of a GPU's capacity. vLLM and TGI both implement techniques like continuous batching (grouping multiple requests' computation together) and efficient memory management for the model's attention cache, turning the same hardware into a much higher-throughput server — the software equivalent of a head chef timing multiple tables' orders to hit the stove in the same batch instead of cooking each ticket in strict isolation.

vLLM: High-Throughput Serving

Use case: vLLM is an open-source inference and serving engine built around PagedAttention, a memory management technique that dramatically increases how many concurrent requests a single GPU can serve — and it exposes an OpenAI-compatible API out of the box.

pip install vllm

vLLM in Action

vllm serve meta-llama/Llama-3.2-3B-Instruct
Terminal Output

Click Run to see what this code prints.

from openai import OpenAI
# Same openai SDK from lesson 5 -- just pointed at your own vLLM server
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
model="meta-llama/Llama-3.2-3B-Instruct",
messages=[{"role": "user", "content": "What is horizontal scaling?"}],
)
print(response.choices[0].message.content)
Terminal Output

Click Run to see what this code prints.

The Same SDK, a Different base_url

This is the same pattern seen with Groq and Together AI in lesson 6 — an OpenAI-compatible server means your application code barely changes whether it is talking to OpenAI's cloud or a vLLM server you deployed yourself.

TGI: Hugging Face's Serving Toolkit

Use case: Text Generation Inference (TGI) is Hugging Face's own production serving toolkit, solving a similar problem to vLLM — high-throughput serving with continuous batching — with tight integration into the wider Hugging Face ecosystem (lesson 8) and first-class Docker deployment support.

docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \
--model-id meta-llama/Llama-3.2-3B-Instruct

vLLM vs TGI

AspectvLLMTGI
OriginUC Berkeley research project, now widely adoptedHugging Face
API compatibilityOpenAI-compatible out of the boxIts own API, plus OpenAI-compatible mode available
Deploymentpip install or DockerDocker-first
Best forMaximum throughput, broad model supportTeams already standardized on the Hugging Face ecosystem

Common Mistakes

Avoid These Mistakes
  • Reaching for vLLM or TGI for a single-user local project — Ollama (lesson 18) is simpler and sufficient at that scale.
  • Underestimating GPU memory requirements for the model size being served, causing out-of-memory failures under load.
  • Assuming these tools eliminate the need for standard infrastructure practices like load balancing and monitoring at real production scale.

Best Practices

  • Reach for vLLM or TGI only once you are serving a model to multiple concurrent users, not for single-user local development.
  • Benchmark throughput and latency with your actual expected traffic pattern before committing to a hardware size.
  • Take advantage of the OpenAI-compatible API to reuse the same client code across hosted and self-served models.

Frequently Asked Questions

No — these tools are specifically for serving open-weight models yourself. If you only ever call hosted provider APIs, this lesson is background knowledge rather than something you need to set up.

vLLM is built around GPU-accelerated serving and is not intended for CPU-only production use — for CPU-only local use, Ollama or llama.cpp (lesson 18) are the better fit.

Both are open-source, though license terms can change between versions — always check each project's current license before production use.

Probably not for a while — these tools matter once a company needs to serve a self-hosted model to real users at scale. Knowing they exist (and why Ollama alone isn't the answer at that scale) is far more useful to a beginner than hands-on practice with them right now.

Key Takeaways

  • Serving many concurrent users efficiently requires different tooling than running a model for yourself.
  • vLLM uses PagedAttention for high-throughput serving with an OpenAI-compatible API.
  • TGI is Hugging Face's equivalent, with tighter ecosystem integration and Docker-first deployment.
  • Both are reserved for production-scale serving, not single-user local development.

Summary

vLLM and TGI are what turn an open-weight model from something that runs on your laptop into something that can serve real production traffic — the natural next step after local inference once a project needs to scale.

Lesson 19 Completed
  • You understand what changes between running and serving a model at scale.
  • You can serve a model with vLLM behind an OpenAI-compatible API.
  • You know how TGI compares as an alternative.
Next Lesson →

The JavaScript/TypeScript AI Ecosystem