Scraper: a website's product list to a spreadsheet
A small Python script that reads a whole product list into a CSV for Excel or Sheets.
Run on books.toscrape.com, a practice site for scrapers. A demo, not client work. The output is real: 1,000 books from 50 pages.
What you get
- One command to run it, and a readme.
- One row per item, with the fields you pick.
- Polite: one request at a time, pauses, retries.
Not included: logins, captchas, getting around blocks, personal contact details. A scraper can break if a site changes.
First 30 rows of the output
| Title | Price (GBP) | Rating | In stock |
|---|---|---|---|
| A Light in the Attic | 51.77 | 3 | yes |
| Tipping the Velvet | 53.74 | 1 | yes |
| Soumission | 50.10 | 1 | yes |
| Sharp Objects | 47.82 | 4 | yes |
| Sapiens: A Brief History of Humankind | 54.23 | 5 | yes |
| The Requiem Red | 22.65 | 1 | yes |
| The Dirty Little Secrets of Getting Your Dream Job | 33.34 | 4 | yes |
| The Coming Woman: A Novel Based on the Life of the Infamous Feminist, Victoria Woodhull | 17.93 | 3 | yes |
| The Boys in the Boat: Nine Americans and Their Epic Quest for Gold at the 1936 Berlin Olympics | 22.60 | 4 | yes |
| The Black Maria | 52.15 | 1 | yes |
| Starving Hearts (Triangular Trade Trilogy, #1) | 13.99 | 2 | yes |
| Shakespeare's Sonnets | 20.66 | 4 | yes |
| Set Me Free | 17.46 | 5 | yes |
| Scott Pilgrim's Precious Little Life (Scott Pilgrim #1) | 52.29 | 5 | yes |
| Rip it Up and Start Again | 35.02 | 5 | yes |
| Our Band Could Be Your Life: Scenes from the American Indie Underground, 1981-1991 | 57.25 | 3 | yes |
| Olio | 23.88 | 1 | yes |
| Mesaerion: The Best Science Fiction Stories 1800-1849 | 37.59 | 1 | yes |
| Libertarianism for Beginners | 51.33 | 2 | yes |
| It's Only the Himalayas | 45.17 | 2 | yes |
| In Her Wake | 12.84 | 1 | yes |
| How Music Works | 37.32 | 2 | yes |
| Foolproof Preserving: A Guide to Small Batch Jams, Jellies, Pickles, Condiments, and More: A Foolproof Guide to Making Small Batch Jams, Jellies, Pickles, Condiments, and More | 30.52 | 3 | yes |
| Chase Me (Paris Nights #2) | 25.27 | 5 | yes |
| Black Dust | 34.53 | 5 | yes |
| Birdsong: A Story in Pictures | 54.64 | 3 | yes |
| America's Cradle of Quarterbacks: Western Pennsylvania's Football Factory from Johnny Unitas to Joe Montana | 22.50 | 3 | yes |
| Aladdin and His Wonderful Lamp | 53.13 | 3 | yes |
| Worlds Elsewhere: Journeys Around Shakespeare’s Globe | 40.30 | 5 | yes |
| Wall and Piece | 44.18 | 4 | yes |
Download all 1,000 rows (books.csv)
The code
"""Scrape the book catalogue of https://books.toscrape.com (a practice site built for scraping) into a CSV.
python scrape_books.py [--pages N] [--out books.csv]
One row per book: title, price, rating (1-5), in stock, product link, image link.
Polite by default: one request at a time, a short pause between pages, retries on errors.
"""
import argparse
import csv
import sys
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
BASE = "https://books.toscrape.com/catalogue/"
FIRST = BASE + "page-1.html"
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
FIELDS = ["title", "price_gbp", "rating", "in_stock", "url", "image"]
def get(session, url, tries=3):
for n in range(tries):
try:
r = session.get(url, timeout=20)
if r.status_code == 404:
return None
r.raise_for_status()
r.encoding = "utf-8"
return r.text
except requests.RequestException:
if n == tries - 1:
raise
time.sleep(2 * (n + 1))
def parse(html, page_url):
soup = BeautifulSoup(html, "html.parser")
rows = []
for b in soup.select("article.product_pod"):
link = b.h3.a
rows.append({
"title": link["title"],
"price_gbp": f'{float(b.select_one(".price_color").text.strip().lstrip("£")):.2f}',
"rating": RATINGS[b.select_one("p.star-rating")["class"][1]],
"in_stock": "In stock" in b.select_one(".availability").text,
"url": urljoin(page_url, link["href"]),
"image": urljoin(page_url, b.img["src"]),
})
nxt = soup.select_one("li.next a")
return rows, urljoin(page_url, nxt["href"]) if nxt else None
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--pages", type=int, default=0, help="stop after N pages (0 = all)")
ap.add_argument("--out", default="books.csv")
ap.add_argument("--pause", type=float, default=0.3, help="seconds between pages")
a = ap.parse_args()
session = requests.Session()
session.headers["User-Agent"] = "small-scraper-demo/1.0"
url, page, rows, seen = FIRST, 0, [], set()
while url and (not a.pages or page < a.pages):
html = get(session, url)
if html is None:
break
got, url = parse(html, url)
rows += [r for r in got if r["url"] not in seen]
seen.update(r["url"] for r in got)
page += 1
print(f"page {page}: {len(got)} books", file=sys.stderr)
time.sleep(a.pause)
with open(a.out, "w", newline="", encoding="utf-8-sig") as f: # the BOM makes Excel read accents correctly
w = csv.DictWriter(f, FIELDS)
w.writeheader()
w.writerows(rows)
print(f"{len(rows)} books from {page} pages -> {a.out}")
if __name__ == "__main__":
main()