Skip to content

Repository files navigation

Here is the complete README.md formatted in the same structure, style, and tone as your Metropolitan Environment project, adapted for your Web Scraper repository:

Product Web Scraper

A Flask-based web application for extracting, viewing, and exporting product data from e-commerce platforms. The application retrieves product titles, prices, stock availability, star ratings, and direct product links using BeautifulSoup and Requests, processes the information, and displays it through an interactive web interface.

Features

  • 🛍️ Location-based product data extraction

  • 🏷️ Live product price parsing

  • 📦 Live stock availability checking

  • ⭐ Product star rating extraction

  • 🎯 Configurable scraping URL and item limit

  • 💾 Stores collected scraped data in local JSON and CSV files

  • 🔍 Dynamic client-side table search and filtering

  • 📥 One-click CSV and JSON data exports

  • 🌐 Flask-based web application with dark-mode Glassmorphism UI

Technologies Used

  • Python
  • Flask
  • BeautifulSoup4
  • Requests
  • HTML
  • CSS
  • JavaScript
  • JSON & CSV
  • Git & GitHub

Project Structure

Web-Scraper-/
│
├── app.py
├── scraper.py
├── README.md
├── .gitignore
├── requirements.txt
│
├── static/
│   ├── css/
│   │   └── style.css
│   └── js/
│       └── main.js
│
├── templates/
│   └── dashboard.html
│
├── data/
│   ├── scraped_data.json
│   └── scraped_data.csv
│
└── venv/

Note: venv/ and local data/ cache files should not be uploaded to GitHub.

How It Works

The application follows these basic steps:

  1. The user opens the web application.

  2. A target product URL and item limit are provided.

  3. Flask receives the request and triggers the scraper module.

  4. The application requests the webpage using custom HTTP headers.

  5. BeautifulSoup parses the HTML structure and extracts product metadata.

  6. The collected information is returned to the frontend as JSON.

  7. The product data is cached locally in data/scraped_data.json and data/scraped_data.csv.

  8. The frontend renders the data in an interactive table with instant filtering options.

Product Data Collected

The application collects product details such as:

  • Product Title

  • Product Price

  • Stock Availability Status

  • Star Rating

  • Direct Product URL

Security & Git Configuration

The .gitignore file should contain:

venv/
__pycache__/
*.pyc
data/

Installation

1. Clone the Repository

git clone https://github.com/ayasreebiswas-cmd/Web-Scraper-.git

2. Open the Project

cd Web-Scraper-

3. Create a Virtual Environment (Optional)

On Windows:

python -m venv venv

4. Activate the Virtual Environment

On Windows:

venv\Scripts\activate

On macOS/Linux:

source venv/bin/activate

5. Install Dependencies

pip install -r requirements.txt

If requirements.txt has not been created yet, install the required packages:

pip install Flask requests beautifulsoup4

Then create the requirements file:

pip freeze > requirements.txt

Running the Application

Activate the virtual environment:

venv\Scripts\activate

Then run:

python app.py

The Flask application runs on: [http://127.0.0.1:5000](http://127.0.0.1:5000)

Open the URL in your web browser.

API Endpoints

Home Page

GET /

Displays the main dashboard pre-populated with cached data.

Scrape Live Data

POST /api/scrape

Triggers the scraping process.

Request Example:

{
  "url": "https://books.toscrape.com/",
  "limit": 20
}

Example Response:

{
  "status": "success",
  "count": 1,
  "data": [
    {
      "title": "A Light in the Attic",
      "price": "£51.77",
      "availability": "In stock",
      "rating": "3 ★",
      "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
    }
  ],
  "source": "https://books.toscrape.com/"
}

Export Data

GET /download/<fmt>

Downloads the current cached dataset (csv or json).

Data Storage

The application stores collected product details in:

  • data/scraped_data.json

  • data/scraped_data.csv

The stored information includes:

  • Title

  • Price

  • Stock Status

  • Rating

  • Product Link

GitHub Workflow

After making changes to the project:

git status
git add .
git commit -m "Update project"
git push

Future Improvements

Possible future improvements include:

  • Support for pagination across multiple pages
  • Selenium / Playwright integration for dynamic JavaScript websites
  • Database integration (SQLite or PostgreSQL)
  • Automated scheduled scraping jobs
  • Price tracking and historical charts
  • Email alerts for price drops

Author

Ayasree Biswas

License

This project is intended for educational and project-development purposes.

About

A Flask-based web application that extracts product catalog data (titles, prices, stock, ratings) using BeautifulSoup4 and Requests. Features a modern dark-mode Glassmorphism UI, client-side live filtering, data persistence, and one-click CSV/JSON exports.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages