Overview
Web scraping is where Python's third-party library ecosystem, not just its standard library, starts to matter: `requests` handles the HTTP conversation with a server, and `BeautifulSoup` (from the `bs4` package) turns the raw HTML that comes back into something you can actually search and extract data from. Neither of these ships with Python itself, which is also what makes this project a natural place to practice exception handling — unlike a local file, a network request can fail in ways entirely outside your control: a timeout, a dropped connection, a server returning an error page instead of the page you expected.
Because this project makes real HTTP requests, it deliberately targets **books.toscrape.com**, a site built specifically for scraping practice, with a stable structure and no rate limits or terms of service to worry about. The techniques you learn here — fetching a page defensively, parsing it with CSS-style selectors, looping across pagination, and saving structured results to CSV — apply the same way to any other site, but you should always check a site's terms of service and `robots.txt` before pointing a scraper at a real production website; do not adapt this code to scrape a live e-commerce or social media site without doing that first.
- A `fetch_page()` function that wraps `requests.get()` in `try`/`except` for timeouts and connection errors.
- A `parse_books()` function that uses BeautifulSoup CSS selectors to pull title, price, and availability.
- A `scrape_pages()` function that loops across a chosen number of catalogue pages.
- A `save_to_csv()` function that writes every extracted book to a CSV file.
- A small command-line entry point that asks how many pages to scrape and reports progress as it runs.
Prerequisites
- Libraries — installing third-party packages with `pip install requests beautifulsoup4` and importing them.
- Loops — iterating over a page number range and over the elements BeautifulSoup finds on a page.
- Exception handling — `try`/`except` with specific exception types, not a bare `except:`.
- Dictionaries and lists — building a list of dicts, one per scraped item, the same shape used across this course's other projects.
- Basic HTML/CSS selector familiarity — knowing what a CSS class selector like `.price_color` refers to.
Project Structure
The program lives in one file, `scraper.py`, and needs two packages installed first: `pip install requests beautifulsoup4`. `requests` is responsible for everything network-related — sending the HTTP request and receiving the response — while `BeautifulSoup` is responsible for everything HTML-related, taking the raw text `requests` returns and letting you query it with CSS-style selectors like `article.product_pod` instead of manually searching through text with string methods or regular expressions.
Each catalogue page on the target site lists twenty books, and the site paginates as `catalogue/page-1.html`, `catalogue/page-2.html`, and so on, which is why `BASE_URL` below is written as a format string with a `{}` placeholder for the page number. Every book extracted, from any page, ends up as one dict in a single flat list — the same "list of dicts" shape used throughout this course's other Python projects — which is what `save_to_csv()` in Step 5 writes out at the end.
Step 1: Import Libraries and Set Up Constants
A custom `User-Agent` header and an explicit `REQUEST_TIMEOUT` are both defensive habits worth building early: some servers reject requests that do not look like they came from a real browser, and without a timeout, a single unresponsive server could make `requests.get()` hang indefinitely instead of failing and letting your script move on.
import csvimport requests # third-party library for making HTTP requests (pip install requests)from bs4 import BeautifulSoup # third-party library for parsing HTML (pip install beautifulsoup4)
# books.toscrape.com is built specifically for scraping practice, so it is safe to hit repeatedlyBASE_URL = "https://books.toscrape.com/catalogue/page-{}.html"HEADERS = {"User-Agent": "Mozilla/5.0 (educational scraping tutorial)"} # some servers reject requests with no User-Agent at allREQUEST_TIMEOUT = 10 # seconds; stops the script from hanging forever if the server never respondsStep 2: Fetch a Page Safely
`response.raise_for_status()` is what turns a bad HTTP status code, like a 404 or a 500, into a raised exception — without it, `requests` would happily hand back the server's error page as if it were normal content, and `parse_books()` in the next step would silently extract nothing from it instead of failing loudly. The two `except` blocks below catch a timeout specifically before falling back to the broader `requests.exceptions.RequestException`, which covers connection failures and the status-code error `raise_for_status()` raises.
def fetch_page(url): """Fetch a URL and return its raw HTML text, or None if the request failed for any reason.""" try: response = requests.get(url, headers=HEADERS, timeout=REQUEST_TIMEOUT) response.raise_for_status() # turns a 404/500 response into a raised exception instead of continuing silently except requests.exceptions.Timeout: print(f"Timed out fetching {url}") return None except requests.exceptions.RequestException as e: # catches connection errors and the status-code error above print(f"Failed to fetch {url}: {e}") return None return response.text # the raw HTML, ready for BeautifulSoup to parseStep 3: Parse Book Data With BeautifulSoup
`soup.select(...)` takes a CSS selector and returns every matching element, exactly like `document.querySelectorAll()` would in JavaScript — `"article.product_pod"` matches every `<article>` tag with the class `product_pod`, which is one book on the page. From each one, `select_one()` grabs the single element matching a more specific selector, and `.get_text(strip=True)` pulls out just the visible text with surrounding whitespace removed.
def parse_books(html): """Parse one catalogue page's HTML and return a list of dicts, one per book found.""" soup = BeautifulSoup(html, "html.parser") # "html.parser" is built into Python; no extra install needed books = [] for article in soup.select("article.product_pod"): # each book on the page is one <article class="product_pod"> title_tag = article.h3.a # the book's title lives in the <a> inside the <h3> price_tag = article.select_one("p.price_color") availability_tag = article.select_one("p.instock.availability")
books.append({ # the title attribute holds the full title text; the link text itself can be visually truncated with "..." "title": title_tag["title"].strip() if title_tag else "Unknown", "price": price_tag.get_text(strip=True) if price_tag else "N/A", "availability": availability_tag.get_text(strip=True) if availability_tag else "N/A", }) return booksClick Run to see what this code prints.
Step 4: Loop Across Multiple Pages
`scrape_pages()` ties `fetch_page()` and `parse_books()` together across as many pages as the caller asks for. If `fetch_page()` returns `None` for a given page — because of a timeout or an error `fetch_page()` already printed a message about — the loop uses `continue` to skip straight to the next page number instead of letting the whole scrape crash over one bad page.
def scrape_pages(num_pages): """Scrape the given number of catalogue pages and return one combined list of book dicts.""" all_books = [] for page_num in range(1, num_pages + 1): # range(1, n+1) so page numbers start at 1, not 0 url = BASE_URL.format(page_num) print(f"Fetching page {page_num}...") html = fetch_page(url) if html is None: # fetch_page() already printed why; just skip this page and keep going continue books = parse_books(html) all_books.extend(books) # extend(), not append(), because books is itself already a list print(f" found {len(books)} books") return all_booksStep 5: Write Results to CSV
`save_to_csv()` follows the same `csv.DictWriter` pattern used in the Expense Tracker project: declare the column order once with `fieldnames`, write the header row, then write every dict as one row. Checking `if not books:` first avoids creating an empty (header-only) file when every page in the scrape happened to fail.
def save_to_csv(books, filename="scraped_books.csv"): """Write the scraped book list to a CSV file.""" if not books: print("Nothing to save.") return with open(filename, "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["title", "price", "availability"]) writer.writeheader() writer.writerows(books) print(f"Saved {len(books)} books to {filename}")Step 6: Build the Main Entry Point
Unlike the menu-driven loops in the other Python projects, a scraper is naturally a run-once script: ask how many pages to scrape, do the work, save the results, and exit. The `try`/`except ValueError` around the page-count input guards against non-numeric input the same way the Todo List and Expense Tracker projects guard their numeric prompts.
def main(): try: num_pages = int(input("How many catalogue pages do you want to scrape? ")) except ValueError: print("Please enter a whole number.") return if num_pages < 1: print("Number of pages must be at least 1.") return
books = scrape_pages(num_pages) save_to_csv(books)
if __name__ == "__main__": main()Complete Code
Here is the full script. Install dependencies once with `pip install requests beautifulsoup4`, then run it with `python scraper.py`.
import csvimport requestsfrom bs4 import BeautifulSoup
BASE_URL = "https://books.toscrape.com/catalogue/page-{}.html"HEADERS = {"User-Agent": "Mozilla/5.0 (educational scraping tutorial)"}REQUEST_TIMEOUT = 10
def fetch_page(url): """Fetch a URL and return its raw HTML text, or None if the request failed for any reason.""" try: response = requests.get(url, headers=HEADERS, timeout=REQUEST_TIMEOUT) response.raise_for_status() except requests.exceptions.Timeout: print(f"Timed out fetching {url}") return None except requests.exceptions.RequestException as e: print(f"Failed to fetch {url}: {e}") return None return response.text
def parse_books(html): """Parse one catalogue page's HTML and return a list of dicts, one per book found.""" soup = BeautifulSoup(html, "html.parser") books = [] for article in soup.select("article.product_pod"): title_tag = article.h3.a price_tag = article.select_one("p.price_color") availability_tag = article.select_one("p.instock.availability")
books.append({ "title": title_tag["title"].strip() if title_tag else "Unknown", "price": price_tag.get_text(strip=True) if price_tag else "N/A", "availability": availability_tag.get_text(strip=True) if availability_tag else "N/A", }) return books
def scrape_pages(num_pages): """Scrape the given number of catalogue pages and return one combined list of book dicts.""" all_books = [] for page_num in range(1, num_pages + 1): url = BASE_URL.format(page_num) print(f"Fetching page {page_num}...") html = fetch_page(url) if html is None: continue books = parse_books(html) all_books.extend(books) print(f" found {len(books)} books") return all_books
def save_to_csv(books, filename="scraped_books.csv"): """Write the scraped book list to a CSV file.""" if not books: print("Nothing to save.") return with open(filename, "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["title", "price", "availability"]) writer.writeheader() writer.writerows(books) print(f"Saved {len(books)} books to {filename}")
def main(): try: num_pages = int(input("How many catalogue pages do you want to scrape? ")) except ValueError: print("Please enter a whole number.") return if num_pages < 1: print("Number of pages must be at least 1.") return
books = scrape_pages(num_pages) save_to_csv(books)
if __name__ == "__main__": main()Sample Run
Click Run to see what this code prints.
Extend This Project
- Follow each book's detail page link and extract its star rating and full description too.
- Add a `time.sleep()` delay between requests to be polite to the server when scraping many pages.
- Cache already-fetched pages to a local folder so re-running the script does not re-download unchanged pages.
- Filter the results by price range or availability before saving, using a list comprehension over `books`.
- Swap `save_to_csv()` for a call that inserts each book into a SQLite database using the `sqlite3` module.
Summary
You built a web scraper that separates its concerns the way a production scraper should: `fetch_page()` owns the network conversation and every way it can fail, `parse_books()` owns turning HTML into structured dicts, and `save_to_csv()` owns persistence — none of the three functions know or care about the other two's internals. The defensive habits you practiced here, timeouts, `raise_for_status()`, and specific `except` clauses instead of a bare `except:`, are exactly what separates a scraper that works once in a demo from one that keeps working when a real network hiccups.