Here is the complete README.md formatted in the same structure, style, and tone as your Metropolitan Environment project, adapted for your Web Scraper repository:
A Flask-based web application for extracting, viewing, and exporting product data from e-commerce platforms. The application retrieves product titles, prices, stock availability, star ratings, and direct product links using BeautifulSoup and Requests, processes the information, and displays it through an interactive web interface.
-
🛍️ Location-based product data extraction
-
🏷️ Live product price parsing
-
📦 Live stock availability checking
-
⭐ Product star rating extraction
-
🎯 Configurable scraping URL and item limit
-
💾 Stores collected scraped data in local JSON and CSV files
-
🔍 Dynamic client-side table search and filtering
-
📥 One-click CSV and JSON data exports
-
🌐 Flask-based web application with dark-mode Glassmorphism UI
- Python
- Flask
- BeautifulSoup4
- Requests
- HTML
- CSS
- JavaScript
- JSON & CSV
- Git & GitHub
Web-Scraper-/
│
├── app.py
├── scraper.py
├── README.md
├── .gitignore
├── requirements.txt
│
├── static/
│ ├── css/
│ │ └── style.css
│ └── js/
│ └── main.js
│
├── templates/
│ └── dashboard.html
│
├── data/
│ ├── scraped_data.json
│ └── scraped_data.csv
│
└── venv/
Note:
venv/and localdata/cache files should not be uploaded to GitHub.
The application follows these basic steps:
-
The user opens the web application.
-
A target product URL and item limit are provided.
-
Flask receives the request and triggers the scraper module.
-
The application requests the webpage using custom HTTP headers.
-
BeautifulSoup parses the HTML structure and extracts product metadata.
-
The collected information is returned to the frontend as JSON.
-
The product data is cached locally in
data/scraped_data.jsonanddata/scraped_data.csv. -
The frontend renders the data in an interactive table with instant filtering options.
The application collects product details such as:
-
Product Title
-
Product Price
-
Stock Availability Status
-
Star Rating
-
Direct Product URL
The .gitignore file should contain:
venv/
__pycache__/
*.pyc
data/
git clone https://github.com/ayasreebiswas-cmd/Web-Scraper-.git
cd Web-Scraper-
On Windows:
python -m venv venv
On Windows:
venv\Scripts\activate
On macOS/Linux:
source venv/bin/activate
pip install -r requirements.txt
If requirements.txt has not been created yet, install the required packages:
pip install Flask requests beautifulsoup4
Then create the requirements file:
pip freeze > requirements.txt
Activate the virtual environment:
venv\Scripts\activate
Then run:
python app.py
The Flask application runs on:
[http://127.0.0.1:5000](http://127.0.0.1:5000)
Open the URL in your web browser.
GET /
Displays the main dashboard pre-populated with cached data.
POST /api/scrape
Triggers the scraping process.
Request Example:
{
"url": "https://books.toscrape.com/",
"limit": 20
}
Example Response:
{
"status": "success",
"count": 1,
"data": [
{
"title": "A Light in the Attic",
"price": "£51.77",
"availability": "In stock",
"rating": "3 ★",
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}
],
"source": "https://books.toscrape.com/"
}
GET /download/<fmt>
Downloads the current cached dataset (csv or json).
The application stores collected product details in:
-
data/scraped_data.json -
data/scraped_data.csv
The stored information includes:
-
Title
-
Price
-
Stock Status
-
Rating
-
Product Link
After making changes to the project:
git status
git add .
git commit -m "Update project"
git push
Possible future improvements include:
- Support for pagination across multiple pages
- Selenium / Playwright integration for dynamic JavaScript websites
- Database integration (SQLite or PostgreSQL)
- Automated scheduled scraping jobs
- Price tracking and historical charts
- Email alerts for price drops
Ayasree Biswas
This project is intended for educational and project-development purposes.