An end-to-end Azure Databricks data engineering project built with PySpark, Delta Lake, Auto Loader and Unity Catalog.
Spark eKart processes e-commerce data from Online and Offline sources and transforms it into clean, unified and analytics-ready datasets using the Medallion Architecture.
Online / Offline CSV Files
│
▼
🥉 BRONZE
Raw Delta Data
│
▼
🥈 SILVER
Clean + Unified Data
│
▼
🥇 GOLD
Analytics-Ready Tables
Spark-eKart-End-to-End-Project/
│
├── 📁 Ingestion Bronze/
│ ├── Online Customers
│ ├── Offline Customers
│ ├── Online Products
│ ├── Offline Products
│ ├── Online Orders
│ └── Offline Orders
│
├── 📁 Transformation Silver/
│ └── Silver Transformations
│
└── 📁 Transform Gold/
└── Gold Transformations
Raw data from Online and Offline sources is ingested using Databricks Auto Loader and stored as Delta tables.
- ⚡ Auto Loader
- 🔄 Structured Streaming
- 🗃️ Delta Lake
- 📍 Checkpointing
- 📝 Audit columns
The Silver layer cleans and standardizes data from both sources.
- 🧹 Data cleansing
- 🔄 Column standardization
- 🔗 Online + Offline data unification
- 📥 Incremental loading using a watermark
- 📊 Creation of unified Customer, Product and Order tables
Online Data ─────┐
├──────► Silver Unified Tables
Offline Data ────┘
The Gold layer contains business-ready datasets for analytics.
Product dimension implementing SCD Type 2 to maintain product history.
Order-level fact table containing:
- Quantity
- Price
- Discount
- Revenue
- Customer
- Product
Provides a complete view of customers including:
- Customer details
- Total orders
- Revenue
- Purchase behavior
- Customer segment
A watermark table keeps track of the last processed file_date.
New Run
│
▼
Check Last Watermark
│
▼
Process Only New Data
│
▼
Update Watermark
This ensures that already processed data is not unnecessarily processed again.
🐍 Python Transformation logic ⚡ PySpark Distributed processing ☁️ Azure Databricks Data engineering platform 🗃️ Delta Lake Reliable data storage 🚀 Auto Loader Incremental file ingestion 🔄 Structured Streaming Streaming ingestion 🔐 Unity Catalog Data governance
This project demonstrates:
- ✅ Medallion Architecture
- ✅ Incremental Data Processing
- ✅ Auto Loader
- ✅ Delta Tables
- ✅ Watermarking
- ✅ Data Transformation
- ✅ SCD Type 2
- ✅ Fact & Dimension Modeling
- ✅ Customer 360
CSV Files
│
▼
Auto Loader
│
▼
🥉 Bronze
│
▼
Incremental Processing
│
▼
🥈 Silver
│
▼
Business Transformations
│
▼
🥇 Gold
│
▼
Analytics
I built an end-to-end e-commerce data pipeline using Azure Databricks and PySpark. Data from Online and Offline sources is ingested into Bronze using Auto Loader, cleaned and unified in Silver with incremental processing, and transformed into Gold tables including a Product SCD Type 2 dimension, Order Fact table and Customer 360 dataset.
The goal is simple:
Turn raw e-commerce data into reliable, clean and analytics-ready data.
Raw Data
↓
Clean Data
↓
Unified Data
↓
Historical Data
↓
Business-Ready Data
Chethan Prabhas
Azure Data Engineer | Databricks | PySpark | Azure Data Factory | SQL
If you found this repository useful, consider giving it a ⭐ to support the project.