Workload-aware ML framework for PySpark performance prediction and configuration recommendation, validated through real benchmark execution and out-of-distribution evaluation.
-
Updated
Sep 12, 2026 - Jupyter Notebook
Workload-aware ML framework for PySpark performance prediction and configuration recommendation, validated through real benchmark execution and out-of-distribution evaluation.
End-to-end data pipeline transforming Olist e-commerce data through Azure cloud services. Implements medallion architecture (Bronze-Silver-Gold) with multi-source ingestion, Spark-based processing, and OLTP-to-OLAP optimization for analytics-ready datasets.
Hybrid time-series and block-column storage database engine written in Java
Production-oriented TypeScript framework for governed, retrieval-augmented data agents with semantic-layer querying, bounded SQL, sampling, profiling, verification, and audit traces.
🛠️ Python library to import OCR data in various formats into the canonical JSON format defined by the Impresso project.
High-performance desktop app for removing duplicate lines from massive text datasets (100GB+)
Apache Pig analysis of 464k UK road accidents (2012–2014) with pandas feature engineering
End-to-end big data pipeline for delivery operations analytics using distributed storage, batch & stream processing, and live dashboards to support operational monitoring and decision-making.
Big data analytics platform combining batch processing, real-time anomaly detection, Kafka streaming, and multi db serving.
Advanced network analytics pipeline for RIPE Atlas traceroute data. Features streaming JSONL processing, IP-to-ASN enrichment with local caching, and interactive geolocation mapping of 88M+ measurements
Sistema de streaming para predecir la probabilidad de wipe usando LSTM en GPU, demostrando la arquitectura event-driven escalable a entornos de producción distribuidos.
Repository for the CENG544 course that I have taken at IZTECH
Hands-on project demos covering infrastructure automation (Ansible, Docker), big-data processing & streaming (Hive, Spark, Kafka), and network experiments (MitM, TCP-over-UDP).
Project repository for university course in real time big data processing
All sorts of Interview Questions
This project uses Apache NiFi to construct a straightforward yet comprehensive data input and transformation pipeline. The objective is to extract a clean, deduplicated list of users from raw e-commerce transaction data that is contained in a CSV file.
A practical coursework-style project from my Master's studies in Big Data Analytics (at University of East London), showcasing hands-on use of big data tools and techniques on a real-world cyber-security dataset.
Traitement distribué d’images sur AWS (EMR, EC2, S3) avec PySpark et MobileNetV2 : extraction de features, PCA Spark et pipeline Big Data scalable.
Course covers big data fundamentals, processes, technologies, platform ecosystem, and management for practical application development.
From traffic sensors to smarter cities: real-time congestion prediction with Kafka, Spark, LSTM, XGBoost, and dynamic routing powered by graph algorithms.
To associate your repository with the big-data-processing topic, visit your repo's landing page and select "manage topics."