Demo

Scraper: a website's product list to a spreadsheet

A small Python script that reads a whole product list into a CSV for Excel or Sheets.

Run on books.toscrape.com, a practice site for scrapers. A demo, not client work. The output is real: 1,000 books from 50 pages.

What you get

Not included: logins, captchas, getting around blocks, personal contact details. A scraper can break if a site changes.

First 30 rows of the output

TitlePrice (GBP)RatingIn stock
A Light in the Attic51.773yes
Tipping the Velvet53.741yes
Soumission50.101yes
Sharp Objects47.824yes
Sapiens: A Brief History of Humankind54.235yes
The Requiem Red22.651yes
The Dirty Little Secrets of Getting Your Dream Job33.344yes
The Coming Woman: A Novel Based on the Life of the Infamous Feminist, Victoria Woodhull17.933yes
The Boys in the Boat: Nine Americans and Their Epic Quest for Gold at the 1936 Berlin Olympics22.604yes
The Black Maria52.151yes
Starving Hearts (Triangular Trade Trilogy, #1)13.992yes
Shakespeare's Sonnets20.664yes
Set Me Free17.465yes
Scott Pilgrim's Precious Little Life (Scott Pilgrim #1)52.295yes
Rip it Up and Start Again35.025yes
Our Band Could Be Your Life: Scenes from the American Indie Underground, 1981-199157.253yes
Olio23.881yes
Mesaerion: The Best Science Fiction Stories 1800-184937.591yes
Libertarianism for Beginners51.332yes
It's Only the Himalayas45.172yes
In Her Wake12.841yes
How Music Works37.322yes
Foolproof Preserving: A Guide to Small Batch Jams, Jellies, Pickles, Condiments, and More: A Foolproof Guide to Making Small Batch Jams, Jellies, Pickles, Condiments, and More30.523yes
Chase Me (Paris Nights #2)25.275yes
Black Dust34.535yes
Birdsong: A Story in Pictures54.643yes
America's Cradle of Quarterbacks: Western Pennsylvania's Football Factory from Johnny Unitas to Joe Montana22.503yes
Aladdin and His Wonderful Lamp53.133yes
Worlds Elsewhere: Journeys Around Shakespeare’s Globe40.305yes
Wall and Piece44.184yes

Download all 1,000 rows (books.csv)

The code

"""Scrape the book catalogue of https://books.toscrape.com (a practice site built for scraping) into a CSV.

python scrape_books.py [--pages N] [--out books.csv]

One row per book: title, price, rating (1-5), in stock, product link, image link.
Polite by default: one request at a time, a short pause between pages, retries on errors.
"""
import argparse
import csv
import sys
import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

BASE = "https://books.toscrape.com/catalogue/"
FIRST = BASE + "page-1.html"
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
FIELDS = ["title", "price_gbp", "rating", "in_stock", "url", "image"]


def get(session, url, tries=3):
    for n in range(tries):
        try:
            r = session.get(url, timeout=20)
            if r.status_code == 404:
                return None
            r.raise_for_status()
            r.encoding = "utf-8"
            return r.text
        except requests.RequestException:
            if n == tries - 1:
                raise
            time.sleep(2 * (n + 1))


def parse(html, page_url):
    soup = BeautifulSoup(html, "html.parser")
    rows = []
    for b in soup.select("article.product_pod"):
        link = b.h3.a
        rows.append({
            "title": link["title"],
            "price_gbp": f'{float(b.select_one(".price_color").text.strip().lstrip("£")):.2f}',
            "rating": RATINGS[b.select_one("p.star-rating")["class"][1]],
            "in_stock": "In stock" in b.select_one(".availability").text,
            "url": urljoin(page_url, link["href"]),
            "image": urljoin(page_url, b.img["src"]),
        })
    nxt = soup.select_one("li.next a")
    return rows, urljoin(page_url, nxt["href"]) if nxt else None


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--pages", type=int, default=0, help="stop after N pages (0 = all)")
    ap.add_argument("--out", default="books.csv")
    ap.add_argument("--pause", type=float, default=0.3, help="seconds between pages")
    a = ap.parse_args()

    session = requests.Session()
    session.headers["User-Agent"] = "small-scraper-demo/1.0"
    url, page, rows, seen = FIRST, 0, [], set()
    while url and (not a.pages or page < a.pages):
        html = get(session, url)
        if html is None:
            break
        got, url = parse(html, url)
        rows += [r for r in got if r["url"] not in seen]
        seen.update(r["url"] for r in got)
        page += 1
        print(f"page {page}: {len(got)} books", file=sys.stderr)
        time.sleep(a.pause)
    with open(a.out, "w", newline="", encoding="utf-8-sig") as f:  # the BOM makes Excel read accents correctly
        w = csv.DictWriter(f, FIELDS)
        w.writeheader()
        w.writerows(rows)
    print(f"{len(rows)} books from {page} pages -> {a.out}")


if __name__ == "__main__":
    main()