Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

🛒 Spark eKart

An end-to-end Azure Databricks data engineering project built with PySpark, Delta Lake, Auto Loader and Unity Catalog.

Spark eKart processes e-commerce data from Online and Offline sources and transforms it into clean, unified and analytics-ready datasets using the Medallion Architecture.


🔄 Architecture

        Online / Offline CSV Files
                    │
                    ▼
              🥉 BRONZE
             Raw Delta Data
                    │
                    ▼
              🥈 SILVER
        Clean + Unified Data
                    │
                    ▼
               🥇 GOLD
        Analytics-Ready Tables

📂 Project Structure

Spark-eKart-End-to-End-Project/
│
├── 📁 Ingestion Bronze/
│   ├── Online Customers
│   ├── Offline Customers
│   ├── Online Products
│   ├── Offline Products
│   ├── Online Orders
│   └── Offline Orders
│
├── 📁 Transformation Silver/
│   └── Silver Transformations
│
└── 📁 Transform Gold/
    └── Gold Transformations

🥉 Bronze Layer

Raw data from Online and Offline sources is ingested using Databricks Auto Loader and stored as Delta tables.

Key Concepts

  • ⚡ Auto Loader
  • 🔄 Structured Streaming
  • 🗃️ Delta Lake
  • 📍 Checkpointing
  • 📝 Audit columns

🥈 Silver Layer

The Silver layer cleans and standardizes data from both sources.

What happens here?

  • 🧹 Data cleansing
  • 🔄 Column standardization
  • 🔗 Online + Offline data unification
  • 📥 Incremental loading using a watermark
  • 📊 Creation of unified Customer, Product and Order tables
 Online Data ─────┐
                  ├──────► Silver Unified Tables
 Offline Data ────┘

🥇 Gold Layer

The Gold layer contains business-ready datasets for analytics.

📦 dim_product

Product dimension implementing SCD Type 2 to maintain product history.

🧾 fact_orders

Order-level fact table containing:

  • Quantity
  • Price
  • Discount
  • Revenue
  • Customer
  • Product

👤 customer_360

Provides a complete view of customers including:

  • Customer details
  • Total orders
  • Revenue
  • Purchase behavior
  • Customer segment

🔄 Incremental Loading

A watermark table keeps track of the last processed file_date.

             New Run
                │
                ▼
       Check Last Watermark
                │
                ▼
       Process Only New Data
                │
                ▼
        Update Watermark

This ensures that already processed data is not unnecessarily processed again.


🛠️ Technologies


🐍 Python Transformation logic ⚡ PySpark Distributed processing ☁️ Azure Databricks Data engineering platform 🗃️ Delta Lake Reliable data storage 🚀 Auto Loader Incremental file ingestion 🔄 Structured Streaming Streaming ingestion 🔐 Unity Catalog Data governance


🎯 Key Data Engineering Concepts

This project demonstrates:

  • ✅ Medallion Architecture
  • ✅ Incremental Data Processing
  • ✅ Auto Loader
  • ✅ Delta Tables
  • ✅ Watermarking
  • ✅ Data Transformation
  • ✅ SCD Type 2
  • ✅ Fact & Dimension Modeling
  • ✅ Customer 360

🚀 End-to-End Flow

CSV Files
   │
   ▼
Auto Loader
   │
   ▼
🥉 Bronze
   │
   ▼
Incremental Processing
   │
   ▼
🥈 Silver
   │
   ▼
Business Transformations
   │
   ▼
🥇 Gold
   │
   ▼
Analytics

🎤 Interview Explanation

I built an end-to-end e-commerce data pipeline using Azure Databricks and PySpark. Data from Online and Offline sources is ingested into Bronze using Auto Loader, cleaned and unified in Silver with incremental processing, and transformed into Gold tables including a Product SCD Type 2 dimension, Order Fact table and Customer 360 dataset.


⭐ Project Goal

The goal is simple:

Turn raw e-commerce data into reliable, clean and analytics-ready data.

Raw Data
   ↓
Clean Data
   ↓
Unified Data
   ↓
Historical Data
   ↓
Business-Ready Data

👨‍💻 Author

Chethan Prabhas

Azure Data Engineer | Databricks | PySpark | Azure Data Factory | SQL

If you found this repository useful, consider giving it a ⭐ to support the project.

About

park eKart is an end-to-end Azure Databricks data engineering project that processes online and offline e-commerce data using PySpark, Auto Loader, Delta Lake, and Unity Catalog. It follows the Medallion Architecture (Bronze → Silver → Gold) with incremental loading, data standardization, SCD Type 2 and Customer 360 fact and dimension tables.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages