LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 824 min read

Hugging Face for LLMs: Transformers, Tokenizers & the Hub

Learn the Hugging Face transformers pipeline API, how tokenizers work, and how to discover and download open-weight models from the Hub.

Introduction

Hugging Face is where the open-weight side of generative AI lives: a hub hosting hundreds of thousands of models, a transformers library that loads nearly any of them with a few lines of code, and tokenizers that turn text into the numeric form a model actually understands. Where lessons 5 and 6 called hosted APIs, this lesson runs an open model's weights directly.

What You Will Learn
  • How to run text generation with the transformers pipeline API.
  • What a tokenizer actually does to a piece of text.
  • How to find and download a model from the Hugging Face Hub.

A Real-Life Analogy First

The Hugging Face Hub is a lot like a public library, except every book (model) was written by a different author (research lab or individual) and the library also lends out the exact tool needed to read each book. A tokenizer, meanwhile, is like a translator standing between you and a book written in a language you don't speak: it breaks your sentence into small pieces the "reader" (the model) can understand, and later translates the reader's response back into your language.

The transformers Pipeline API

Use case: transformers provides both the model architectures behind most LLMs and a high-level pipeline() function that bundles a model, its tokenizer, and pre/post-processing into one callable — the fastest way to try a model with minimal code.

pip install transformers torch

Text Generation in Action

from transformers import pipeline
generator = pipeline("text-generation", model="distilgpt2")
result = generator(
"The best way to learn a new programming language is",
max_new_tokens=25,
num_return_sequences=1,
)
print(result[0]["generated_text"])
Terminal Output

Click Run to see what this code prints.

Small Models, Local Runs

distilgpt2 is a small, older model chosen here specifically because it downloads and runs quickly on a CPU with no API key. Lesson 18 covers running larger, more capable open-weight models (like Llama) locally with dedicated inference tools.

Understanding Tokenizers

Use case: a tokenizer converts text into a sequence of integer IDs a model can process, and converts generated IDs back into text — every model has its own tokenizer, trained alongside it, which is why token counts (and cost) differ between models for the same input text.

from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("distilgpt2")
text = "Tokenization splits text into model-readable pieces."
tokens = tokenizer.encode(text)
print("Token count:", len(tokens))
print("Token IDs:", tokens)
print("Decoded back:", tokenizer.decode(tokens))
Terminal Output

Click Run to see what this code prints.

Notice the word "Tokenization" itself did not become one token — it was split into sub-word pieces. This is exactly why LLM providers bill "per token" rather than per word or character, and why the same sentence can cost a different number of tokens on different models — the same way a taxi meter charges by distance traveled, not by "number of streets," even though the two are related.

The Hugging Face Hub

The Hub (huggingface.co) is a public registry of models, datasets, and demo "Spaces" — similar in spirit to GitHub, but for machine learning artifacts. Any model name string passed to pipeline() or AutoTokenizer.from_pretrained() above is actually a Hub repository ID; transformers downloads and caches it locally the first time you use it.

pip install huggingface_hub
huggingface-cli login # needed for gated models, not for public ones

Common Mistakes

Avoid These Mistakes
  • Loading a large model on a first run without expecting a multi-gigabyte download — check a model's size on its Hub page first.
  • Assuming every model on the Hub is free to use commercially — always check the model card's license before shipping it in a product.
  • Confusing a model's tokenizer with another model's — always load the tokenizer that was published alongside that exact model.

Best Practices

  • Start with pipeline() for quick experimentation before dropping down to the lower-level model and tokenizer classes.
  • Check a model's Hub page for its license, size, and intended use before adopting it in a real project.
  • Use small models (like distilgpt2 above) while developing, and only move to larger ones once your code path works end-to-end.

Frequently Asked Questions

No — that course's lesson covers transfer learning for general ML tasks (classification, vision). This lesson focuses specifically on text generation and the tooling around LLMs.

Not for public models like the one used here. An account (and huggingface-cli login) is only required for gated models that require accepting a license first.

No — transformers runs open-weight models locally or on your own infrastructure. GPT and Claude are closed models only accessible through their provider's hosted API (lesson 5).

Because each model has its own tokenizer, the same sentence can split into a different number of tokens depending on which model reads it — and since providers bill per token, this directly changes the cost, exactly as shown in the tokenizer example above.

Key Takeaways

  • transformers' pipeline() function loads a model, tokenizer, and pre/post-processing together with minimal code.
  • A tokenizer converts text to integer IDs and back — different models use different tokenizers, which is why token counts vary between them.
  • The Hugging Face Hub is a public registry of open models, datasets, and demos that transformers downloads models from.
  • Always check a model's license on its Hub page before using it commercially.

Summary

Hugging Face is the entry point to the open-weight side of generative AI — its pipeline API, tokenizers, and Hub together make running someone else's trained model as simple as calling a hosted API.

Lesson 8 Completed
  • You can run text generation with the transformers pipeline API.
  • You understand what a tokenizer does and why token counts vary between models.
  • You know how to find and download a model from the Hub.
Next Lesson →

Prompt Orchestration Frameworks: LangChain