Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

📚 Python Web Scraper - Books to Scrape

This project is a Python-based web scraper that extracts book data from http://books.toscrape.com using BeautifulSoup.

🚀 Features

  • Scrapes book titles, prices, and ratings
  • Handles pagination (multiple pages)
  • Saves data into CSV format
  • Clean and modular Python code

🛠️ Tech Stack

  • Python
  • BeautifulSoup
  • Requests

🐍 Virtual Environment Setup

Create and activate a virtual environment:

Windows

python -m venv venv
venv\Scripts\activate

Mac/Linux

python3 -m venv venv
source venv/bin/activate

Then install dependencies:

pip install -r requirements.txt

▶️ How to Run

python scraper.py

📊 Output

Data is saved in:

data/books.csv

Sample Result

Title Price Rating Availability
A Light in the Attic £51.77 Three In stock
Tipping the Velvet £53.74 One In stock
Soumission £50.10 One In stock
Sharp Objects £47.82 Four In stock
Sapiens: A Brief History of Humankind £54.23 Five In stock

🎯 Purpose

This project demonstrates skills in:

  • Web scraping
  • Data extraction
  • Automation using Python

🧠 Key Learnings

  • Understanding HTML structure and DOM parsing
  • Extracting data using CSS selectors
  • Handling pagination in web scraping
  • Automating data collection workflows

🔮 Future Improvements

  • Add support for scraping all available pages dynamically
  • Export data to JSON format
  • Add logging and error handling
  • Use Selenium for dynamic websites
  • Schedule scraping tasks using cron jobs

⭐ This project is part of my learning journey in Python automation and RPA.

About

Python web scraping project using BeautifulSoup to extract book data with pagination and CSV export

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages