This project is a Python-based web scraper that extracts book data from http://books.toscrape.com using BeautifulSoup.
- Scrapes book titles, prices, and ratings
- Handles pagination (multiple pages)
- Saves data into CSV format
- Clean and modular Python code
- Python
- BeautifulSoup
- Requests
Create and activate a virtual environment:
python -m venv venv
venv\Scripts\activatepython3 -m venv venv
source venv/bin/activateThen install dependencies:
pip install -r requirements.txtpython scraper.pyData is saved in:
data/books.csv
| Title | Price | Rating | Availability |
|---|---|---|---|
| A Light in the Attic | £51.77 | Three | In stock |
| Tipping the Velvet | £53.74 | One | In stock |
| Soumission | £50.10 | One | In stock |
| Sharp Objects | £47.82 | Four | In stock |
| Sapiens: A Brief History of Humankind | £54.23 | Five | In stock |
This project demonstrates skills in:
- Web scraping
- Data extraction
- Automation using Python
- Understanding HTML structure and DOM parsing
- Extracting data using CSS selectors
- Handling pagination in web scraping
- Automating data collection workflows
- Add support for scraping all available pages dynamically
- Export data to JSON format
- Add logging and error handling
- Use Selenium for dynamic websites
- Schedule scraping tasks using cron jobs
⭐ This project is part of my learning journey in Python automation and RPA.