LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1619 min read

Natural Language Processing Libraries (NLTK & spaCy)

Learn NLTK, the classic educational NLP toolkit, and spaCy, the fast production-grade NLP library, with tokenization, stopword removal, and named entity recognition examples.

Introduction

Before transformer models like the ones in the previous lesson existed, NLP was already a mature field with its own dedicated tooling for breaking text apart, cleaning it, and extracting structure from it. Two libraries still anchor this work today: NLTK, the classic teaching and research toolkit, and spaCy, built for speed and production use.

This lesson covers both, with a tokenization and stopword-removal example in NLTK, and a named entity recognition example in spaCy.

What You Will Learn
  • What NLTK is and why it is often used for teaching and research.
  • How to tokenize text and remove stopwords with NLTK.
  • What spaCy is and why it is preferred in production.
  • How to extract named entities from text with spaCy.
  • When to choose NLTK over spaCy, or the reverse.

What is NLTK?

NLTK (Natural Language Toolkit) is one of the oldest Python NLP libraries, originally built for teaching computational linguistics. It gives you direct, low-level access to a huge range of algorithms and corpora (collections of real text), which makes it excellent for learning how NLP techniques actually work, though it is slower and more manual than newer libraries for production use.

pip install nltk

Example: Tokenization and Stopwords with NLTK

Tokenization splits raw text into individual words or sentences, and removing stopwords filters out common, low-information words like 'the' or 'is' that usually do not help downstream analysis.

import nltk
nltk.download('punkt')
nltk.download('stopwords')
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
text = "PrograMinds teaches data science through short, practical lessons."
tokens = word_tokenize(text)
stop_words = set(stopwords.words('english'))
filtered_tokens = [word for word in tokens if word.lower() not in stop_words and word.isalpha()]
print("All tokens:", tokens)
print("Without stopwords:", filtered_tokens)
Output

Click Run to see what this code prints.

One-Time Downloads

nltk.download() fetches the specific corpora and models a function needs (like punkt for tokenization) and only needs to run once per machine — NLTK does not bundle all of its data by default, to keep the base install small.

What is spaCy?

spaCy is a newer NLP library designed from the ground up for speed and production use rather than teaching. It ships with pretrained pipelines that handle tokenization, part-of-speech tagging, and named entity recognition together in a single, fast pass over the text, using a clean, object-oriented API.

pip install spacy
python -m spacy download en_core_web_sm

The second command downloads a small pretrained English pipeline — spaCy separates the library itself from its language models, so you download only the languages you actually need.

Example: Named Entity Recognition with spaCy

Named entity recognition (NER) identifies real-world things mentioned in text, such as people, organizations, dates, and locations.

import spacy
nlp = spacy.load("en_core_web_sm")
text = "PrograMinds was founded to teach practical software skills, and it now reaches learners worldwide by 2026."
doc = nlp(text)
for ent in doc.ents:
print(f"{ent.text:20} -> {ent.label_}")
Output

Click Run to see what this code prints.

spaCy correctly identified 'PrograMinds' as an organization and '2026' as a date, entirely from a pretrained pipeline — no manual rules or training required.

NLTK vs spaCy

AspectNLTKspaCy
Primary purposeTeaching, research, algorithm experimentationFast, production-grade NLP pipelines
SpeedSlower, more manual pipeline constructionOptimized, built for processing large volumes of text quickly
API styleFunction-based, many independent modulesObject-oriented, one nlp() call returns a rich Doc object
Best forLearning how NLP techniques work internallyShipping NLP features into a real application

Common Mistakes

Avoid These Mistakes
  • Forgetting to run nltk.download() for the specific corpus a function needs, resulting in a LookupError.
  • Forgetting the separate python -m spacy download step — installing the spacy package alone does not include any language pipeline.
  • Using NLTK's slower, manual pipeline in a production system where spaCy's optimized pipeline would be a better fit.

Best Practices

  • Use NLTK when you are learning NLP concepts or need access to its wide range of classic corpora and algorithms.
  • Use spaCy when you are building a real application and need speed and a clean API.
  • Process text in batches with spaCy's nlp.pipe() rather than calling nlp() in a loop, for much better performance on large volumes of text.
  • Combine spaCy's fast preprocessing with a Hugging Face transformer model for tasks that need deeper language understanding.

Frequently Asked Questions

Yes, though it is uncommon in a single pipeline. Some projects use NLTK for exploratory analysis on a small dataset, then switch to spaCy once they are building the production pipeline.

In production settings, largely yes, but NLTK remains popular in classrooms and research because its lower-level design makes it easier to see exactly how each NLP technique works.

spaCy's larger pipelines include transformer-based models for higher accuracy, while its small pipelines (like en_core_web_sm used above) use faster, non-transformer statistical models. NLTK is largely rule-based and classical, predating the deep learning era of NLP.

Key Takeaways

  • NLTK is the classic, educational NLP toolkit, ideal for learning how techniques work.
  • spaCy is built for speed and production, with pretrained pipelines covering tokenization through named entity recognition.
  • word_tokenize() and stopwords.words() are core NLTK building blocks for cleaning text.
  • spacy.load() plus nlp(text).ents extracts named entities in just a few lines.
  • Choose NLTK for learning and research, spaCy for shipping real NLP features.

Summary

NLTK and spaCy represent two different eras and philosophies of NLP tooling: one built to teach the concepts, the other built to ship them fast. Knowing both gives you the flexibility to learn deeply and still build production-ready systems.

Lesson 16 Completed
  • You tokenized text and removed stopwords with NLTK.
  • You extracted named entities from real text with spaCy.
  • You are ready to move from text into computer vision libraries.
Next Lesson →

Computer Vision Libraries (OpenCV & Pillow)