Natural Language Processing Libraries (NLTK & spaCy)
Learn NLTK, the classic educational NLP toolkit, and spaCy, the fast production-grade NLP library, with tokenization, stopword removal, and named entity recognition examples.
Introduction
Before transformer models like the ones in the previous lesson existed, NLP was already a mature field with its own dedicated tooling for breaking text apart, cleaning it, and extracting structure from it. Two libraries still anchor this work today: NLTK, the classic teaching and research toolkit, and spaCy, built for speed and production use.
This lesson covers both, with a tokenization and stopword-removal example in NLTK, and a named entity recognition example in spaCy.
- What NLTK is and why it is often used for teaching and research.
- How to tokenize text and remove stopwords with NLTK.
- What spaCy is and why it is preferred in production.
- How to extract named entities from text with spaCy.
- When to choose NLTK over spaCy, or the reverse.
What is NLTK?
NLTK (Natural Language Toolkit) is one of the oldest Python NLP libraries, originally built for teaching computational linguistics. It gives you direct, low-level access to a huge range of algorithms and corpora (collections of real text), which makes it excellent for learning how NLP techniques actually work, though it is slower and more manual than newer libraries for production use.
pip install nltkExample: Tokenization and Stopwords with NLTK
Tokenization splits raw text into individual words or sentences, and removing stopwords filters out common, low-information words like 'the' or 'is' that usually do not help downstream analysis.
import nltknltk.download('punkt')nltk.download('stopwords')
from nltk.tokenize import word_tokenizefrom nltk.corpus import stopwords
text = "PrograMinds teaches data science through short, practical lessons."
tokens = word_tokenize(text)stop_words = set(stopwords.words('english'))filtered_tokens = [word for word in tokens if word.lower() not in stop_words and word.isalpha()]
print("All tokens:", tokens)print("Without stopwords:", filtered_tokens)Click Run to see what this code prints.
nltk.download() fetches the specific corpora and models a function needs (like punkt for tokenization) and only needs to run once per machine — NLTK does not bundle all of its data by default, to keep the base install small.
What is spaCy?
spaCy is a newer NLP library designed from the ground up for speed and production use rather than teaching. It ships with pretrained pipelines that handle tokenization, part-of-speech tagging, and named entity recognition together in a single, fast pass over the text, using a clean, object-oriented API.
pip install spacypython -m spacy download en_core_web_smThe second command downloads a small pretrained English pipeline — spaCy separates the library itself from its language models, so you download only the languages you actually need.
Example: Named Entity Recognition with spaCy
Named entity recognition (NER) identifies real-world things mentioned in text, such as people, organizations, dates, and locations.
import spacy
nlp = spacy.load("en_core_web_sm")
text = "PrograMinds was founded to teach practical software skills, and it now reaches learners worldwide by 2026."doc = nlp(text)
for ent in doc.ents: print(f"{ent.text:20} -> {ent.label_}")Click Run to see what this code prints.
spaCy correctly identified 'PrograMinds' as an organization and '2026' as a date, entirely from a pretrained pipeline — no manual rules or training required.
NLTK vs spaCy
| Aspect | NLTK | spaCy |
|---|---|---|
| Primary purpose | Teaching, research, algorithm experimentation | Fast, production-grade NLP pipelines |
| Speed | Slower, more manual pipeline construction | Optimized, built for processing large volumes of text quickly |
| API style | Function-based, many independent modules | Object-oriented, one nlp() call returns a rich Doc object |
| Best for | Learning how NLP techniques work internally | Shipping NLP features into a real application |
Common Mistakes
- Forgetting to run nltk.download() for the specific corpus a function needs, resulting in a LookupError.
- Forgetting the separate python -m spacy download step — installing the spacy package alone does not include any language pipeline.
- Using NLTK's slower, manual pipeline in a production system where spaCy's optimized pipeline would be a better fit.
Best Practices
- Use NLTK when you are learning NLP concepts or need access to its wide range of classic corpora and algorithms.
- Use spaCy when you are building a real application and need speed and a clean API.
- Process text in batches with spaCy's nlp.pipe() rather than calling nlp() in a loop, for much better performance on large volumes of text.
- Combine spaCy's fast preprocessing with a Hugging Face transformer model for tasks that need deeper language understanding.
Frequently Asked Questions
Yes, though it is uncommon in a single pipeline. Some projects use NLTK for exploratory analysis on a small dataset, then switch to spaCy once they are building the production pipeline.
In production settings, largely yes, but NLTK remains popular in classrooms and research because its lower-level design makes it easier to see exactly how each NLP technique works.
spaCy's larger pipelines include transformer-based models for higher accuracy, while its small pipelines (like en_core_web_sm used above) use faster, non-transformer statistical models. NLTK is largely rule-based and classical, predating the deep learning era of NLP.
Key Takeaways
- NLTK is the classic, educational NLP toolkit, ideal for learning how techniques work.
- spaCy is built for speed and production, with pretrained pipelines covering tokenization through named entity recognition.
- word_tokenize() and stopwords.words() are core NLTK building blocks for cleaning text.
- spacy.load() plus nlp(text).ents extracts named entities in just a few lines.
- Choose NLTK for learning and research, spaCy for shipping real NLP features.
Summary
NLTK and spaCy represent two different eras and philosophies of NLP tooling: one built to teach the concepts, the other built to ship them fast. Knowing both gives you the flexibility to learn deeply and still build production-ready systems.
- You tokenized text and removed stopwords with NLTK.
- You extracted named entities from real text with spaCy.
- You are ready to move from text into computer vision libraries.