End-to-End Data Engineering & Analytics Platform
A modern e-commerce data platform for data ingestion, lake storage, distributed processing, data warehousing, APIs, and analytics dashboards.
The platform implements an end-to-end data engineering pipeline:
Data Sources → Python → S3 Raw → PySpark → S3 Processed → Data Warehouse → ASP.NET Core API → Blazor Dashboard
- Python — Data extraction, ingestion, and warehouse loading
- Apache Spark / PySpark — Distributed data processing and transformation
- PostgreSQL — Operational database and data warehouse
- AWS S3 — Cloud-based data lake storage
- Docker Compose — Local development environment
- Pandas
- PyArrow
- Boto3
- Psycopg2
- ASP.NET Core Web API
- .NET 10
- Npgsql
- Blazor Server
- SQL
- Staging Tables
- Star Schema
- Fact & Dimension Modeling
The platform works with multiple e-commerce data sources.
The operational database contains:
customers
products
categories
orders
order_items
inventory
Synthetic product data is generated using a Python data generator and exported to:
products.csv
A Flask-based REST API provides inventory-related data in JSON format.
The project uses an AWS S3 bucket as the central data lake:
ecommerce-data-lake-ferit-2026-a7x92
Region: us-east-1
The data lake is organized into two main zones.
Raw source data is stored with minimal transformation:
RAW/
├── customers/
│ └── *.parquet
├── orders/
│ └── *.parquet
├── products/
│ └── *.parquet
├── order_items/
│ └── *.parquet
└── inventory/
└── *.json
Cleaned and transformed datasets are stored in optimized formats:
PROCESSED/
├── customers/
│ └── *.parquet
├── orders/
│ └── *.parquet
├── products/
│ └── *.parquet
└── inventory/
└── *.parquet
The main processing flow is:
S3 Raw → PySpark → S3 Processed
Python is responsible for extracting data from the source systems and loading it into the data lake.
The ingestion layer supports:
- PostgreSQL → Parquet
- Inventory API → JSON
- Incremental Orders → Parquet
- Watermark-based extraction
Incremental extraction allows the pipeline to process only newly available order records instead of reprocessing the complete dataset.
Apache Spark / PySpark is used for data cleaning and transformation.
The processing layer handles:
Customers
Orders
Products
Inventory
The simplified processing flow is:
S3 Raw
│
▼
PySpark
│
▼
S3 Processed
Transformations are performed locally during development using a Windows environment.
The analytical data warehouse is implemented using PostgreSQL.
Database:
ecommerce_dw
The warehouse is divided into two layers.
stg_customers
stg_orders
stg_products
stg_order_items
stg_inventory
The analytical model follows a Star Schema:
┌───────────────┐
│ dim_customer │
└───────┬───────┘
│
│
┌──────────────┐ ▼ ┌──────────────┐
│ dim_product │ ──► FACT ◄── │ dim_date │
└──────────────┘ │ └──────────────┘
│
▼
fact_sales
dim_customerdim_productdim_date
fact_sales
The fact_sales table supports analytical queries for:
- Total sales
- Discounts
- Profit
- Profit margin
- Order counts
- Sales trends
The data warehouse is exposed through an ASP.NET Core Web API.
PostgreSQL Data Warehouse
│
│ SQL / Npgsql
▼
ASP.NET Core Web API
The API provides data for the dashboard, including:
- KPI metrics
- Sales trends
- Inventory summaries
The frontend is implemented using Blazor Server and provides an interactive e-commerce analytics dashboard.
Blazor Server Dashboard
│
│ HTTP / JSON
▼
ASP.NET Core Web API
│
▼
PostgreSQL Data Warehouse
The dashboard includes:
- KPI Cards
- Sales Metrics
- Inventory Overview
- Sales Charts
- Monthly Sales Trends
- Monthly Profit Trends
The dashboard displays metrics such as:
- Total Orders
- Total Order Items
- Units Sold
- Gross Sales
- Net Sales
- Discounts
- Total Cost
- Profit
- Profit Margin
- Synthetic e-commerce data generation
- Multi-source data ingestion
- PostgreSQL operational database integration
- Inventory REST API integration
- Incremental order extraction
- Watermark-based data ingestion
- AWS S3 Data Lake
- Raw and Processed data zones
- Parquet and JSON data formats
- PySpark data processing
- PostgreSQL staging layer
- Star-schema data warehouse
- ASP.NET Core REST API
- Blazor Server analytics dashboard
- Docker Compose development environment
The goal of this project is to demonstrate a practical, modern end-to-end data engineering architecture by combining:
Data Sources
↓
Data Ingestion
↓
Cloud Data Lake
↓
Distributed Processing
↓
Data Warehouse
↓
API Layer
↓
Analytics Dashboard
The project focuses on building a complete data pipeline from raw source systems to business-facing analytics.
ecommerce-data-platform/
│
├── data-generator/
│
├── extraction/
│
├── spark/
│
├── warehouse-loader/
│
├── inventory-api/
│
├── api/
│
├── dashboard/
│
├── docker-compose.yml
│
└── README.md
| Layer | Technology |
|---|---|
| Data Generation | Python |
| Extraction | Python |
| Data Lake | AWS S3 |
| Data Format | Parquet / JSON |
| Processing | Apache Spark / PySpark |
| Operational DB | PostgreSQL |
| Data Warehouse | PostgreSQL |
| Warehouse Loading | Pandas / PyArrow / Psycopg2 |
| Backend | ASP.NET Core Web API |
| Database Access | Npgsql |
| Dashboard | Blazor Server |
| Environment | Docker Compose |
E-Commerce Data Platform
Built as an end-to-end Data Engineering project