Web Scraping Libraries (BeautifulSoup & Scrapy)
Learn when to reach for BeautifulSoup for quick one-off HTML parsing versus Scrapy for large, multi-page crawling projects.
Introduction
Not every dataset comes from a clean CSV, a database, or an API. Sometimes the data you need only exists rendered inside HTML on a web page. Web scraping is the practice of programmatically extracting that data, and Python has two dominant tools for it, aimed at very different scales of problem.
- How BeautifulSoup parses HTML and extracts data for small, one-off scraping tasks.
- How Scrapy works as a full framework for larger, multi-page crawling projects.
- A clear rule of thumb for choosing between the two.
beautifulsoup4: Parsing HTML
BeautifulSoup (imported as bs4) takes a blob of raw HTML and turns it into a navigable tree you can search and extract from with simple Python. It does not fetch pages itself — you typically pair it with requests to download the HTML first, then hand that HTML to BeautifulSoup to parse.
pip install beautifulsoup4 requestsimport requestsfrom bs4 import BeautifulSoup
response = requests.get("https://quotes.toscrape.com/")soup = BeautifulSoup(response.text, "html.parser")
quotes = soup.find_all("span", class_="text")authors = soup.find_all("small", class_="author")
for quote, author in zip(quotes, authors): print(f"{quote.text} — {author.text}")Click Run to see what this code prints.
BeautifulSoup shines for small, one-off jobs: pull a table off a single Wikipedia page, extract prices from one product page, grab headlines off one news site. It is a parsing library, not a crawler — it has no built-in concept of following links across pages or running requests in parallel.
scrapy: A Full Crawling Framework
Scrapy is a complete scraping framework, not just a parsing library. It handles making requests, following links across many pages, running requests concurrently for speed, retrying failures, respecting robots.txt, and exporting results to CSV, JSON, or a database — all as built-in features. You structure a scraping project as a Spider class.
pip install scrapyimport scrapy
class QuotesSpider(scrapy.Spider): name = "quotes" start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response): for quote in response.css("div.quote"): yield { "text": quote.css("span.text::text").get(), "author": quote.css("small.author::text").get(), }
next_page = response.css("li.next a::attr(href)").get() if next_page is not None: yield response.follow(next_page, self.parse)scrapy runspider quotes_spider.py -o quotes.jsonClick Run to see what this code prints.
Because the spider yields response.follow() for the "next" link, Scrapy automatically crawled every page of quotes, not just the first — something you would have to build by hand with requests and BeautifulSoup.
BeautifulSoup vs Scrapy: When to Use Which
| Situation | Use |
|---|---|
| One page, extract a table or a few fields | BeautifulSoup + requests |
| Hundreds or thousands of pages, following links | Scrapy |
| You need results fast for a quick analysis | BeautifulSoup |
| You need retries, concurrency, and structured export (CSV/JSON/DB) | Scrapy |
| You are already inside a larger script and just need one extraction | BeautifulSoup |
| You are building a recurring, production-grade crawler | Scrapy |
Common Mistakes
- Scraping a site without checking its robots.txt or terms of service — some sites explicitly disallow scraping.
- Sending requests as fast as possible with no delay, which can get your IP address blocked and puts unnecessary load on the target site.
- Reaching for Scrapy on a tiny, single-page task where BeautifulSoup would be simpler and faster to write.
- Assuming a page's HTML structure will never change — scrapers break silently when a site redesigns its markup.
Best Practices
- Check whether the site offers an official API before scraping — an API is almost always more stable and more polite.
- Respect robots.txt and add delays between requests (Scrapy's DOWNLOAD_DELAY setting does this automatically).
- Start with BeautifulSoup for prototyping; move to Scrapy once you need to scale across many pages.
- Wrap scraped-data extraction in try/except, since missing elements (a None from .find()) are extremely common on real-world pages.
Frequently Asked Questions
No — BeautifulSoup only parses the raw HTML it is given. For pages that render content with JavaScript, you need a browser automation tool like Selenium or Playwright to render the page first.
It depends on the site's terms of service, the jurisdiction, and what you do with the data. Always check a site's robots.txt and terms before scraping, and prefer an official API when one exists.
Yes — Scrapy has its own built-in CSS and XPath selectors that do the same parsing job, so you typically do not need BeautifulSoup at all inside a Scrapy project.
Key Takeaways
- BeautifulSoup parses HTML for quick, small-scale extraction tasks, usually paired with requests.
- Scrapy is a full framework for larger crawling jobs: following links, concurrency, retries, and structured export.
- Choose based on scale — one page versus many, one-off versus recurring production crawler.
- Always check robots.txt and prefer an official API when one is available.
Summary
BeautifulSoup and Scrapy cover two very different scales of the same problem: getting structured data out of web pages. Next, we move to the opposite end of the data spectrum — libraries for processing datasets too large to fit on a single machine.
- You can extract data from a single page with BeautifulSoup.
- You can write a multi-page Scrapy Spider.
- You know how to choose between them based on project scale.