Skip to content
#

big-data-processing

Here are 99 public repositories matching this topic...

End-to-end data pipeline transforming Olist e-commerce data through Azure cloud services. Implements medallion architecture (Bronze-Silver-Gold) with multi-source ingestion, Spark-based processing, and OLTP-to-OLAP optimization for analytics-ready datasets.

  • Updated Aug 27, 2026
  • Jupyter Notebook

Production-oriented TypeScript framework for governed, retrieval-augmented data agents with semantic-layer querying, bounded SQL, sampling, profiling, verification, and audit traces.

  • Updated Sep 10, 2026
  • TypeScript

Advanced network analytics pipeline for RIPE Atlas traceroute data. Features streaming JSONL processing, IP-to-ASN enrichment with local caching, and interactive geolocation mapping of 88M+ measurements

  • Updated Apr 17, 2026
  • HTML

A practical coursework-style project from my Master's studies in Big Data Analytics (at University of East London), showcasing hands-on use of big data tools and techniques on a real-world cyber-security dataset.

  • Updated Nov 18, 2025
  • Python

Traitement distribué d’images sur AWS (EMR, EC2, S3) avec PySpark et MobileNetV2 : extraction de features, PCA Spark et pipeline Big Data scalable.

  • Updated Nov 16, 2025
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the big-data-processing topic, visit your repo's landing page and select "manage topics."

Learn more