You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Learn PySpark through practical examples. This repository includes Spark fundamentals, DataFrame operations, Spark SQL, transformations, actions, and real-world practice for aspiring Data Engineers.
The Forex Data Pipeline is a comprehensive solution designed to collect, process, and prepare currency exchange rate data for downstream machine-learning pipelines. This repository showcases the creation of a data pipeline that fetches currency rates from an external API and performs data transformation using PySpark.
The IPL Data Analysis project aims to explore and analyze the Indian Premier League (IPL) data using PySpark for data processing and Matplotlib and Seaborn for data visualization. The goal is to derive actionable insights into player performances, match trends, and overall league dynamics.
A comparative study to understand the computing efficiencies of Pyspark architectures vs python based distributed programming methodologies such as MPI, multi-threading or multi-processing on the Yelp kaggle dataset.
An end-to-end Batch ETL Pipeline implemented on Azure Databricks using PySpark and SQL, processing over 11 million rows of real-world urban data. The project uncovers hidden correlations between New York City transit patterns and historical weather conditions.
Real-time oil well production monitoring pipeline with Kafka, PySpark, Grafana, and AWS (S3, SNS, SQS, ECR). Infrastructure managed by Terraform with CI/CD via GitHub Actions.