Building Data Lake and ETL pipelines using Amazon EMR, S3, and Apache Spark
-
Updated
Dec 23, 2022 - Python
Building Data Lake and ETL pipelines using Amazon EMR, S3, and Apache Spark
BigQuery data pipeline with dbt, Spark, Docker, Airflow, Terraform, GCP
Setting up a Spark cluster in a Docker environment for improved repeatability and reliability. This project includes a simple transformation on a dataset containing approximately 31 million rows.
End-to-end big data pipeline for delivery operations analytics using distributed storage, batch & stream processing, and live dashboards to support operational monitoring and decision-making.
Hands-on project demos covering infrastructure automation (Ansible, Docker), big-data processing & streaming (Hive, Spark, Kafka), and network experiments (MitM, TCP-over-UDP).
Kappa Architecture Based Sentiment Analysis System for User Comments
A practical coursework-style project from my Master's studies in Big Data Analytics (at University of East London), showcasing hands-on use of big data tools and techniques on a real-world cyber-security dataset.
Exploring and Implementing Scalable Data Processing Techniques
This work is from my master thesis: Condition Monitoring with Machine Learning: A Data-Driven Framework for Quantifying Wind Turbine Energy Loss.
Solved tasks of the master's degree courses of speciality "Algorithms and Systems for Big Data Processing".
"Provides tools for parallel pipeline processing of large data structures
Software basati su metodi di intelligenza artificiale per l'automazione dell'analisi di big data.
Electrical Consumption Monitoring - Big Data Pipeline using Lambda Architecture in Python
Sistema de streaming para predecir la probabilidad de wipe usando LSTM en GPU, demostrando la arquitectura event-driven escalable a entornos de producción distribuidos.
Project using Python, Hive and MapReduce to compare various techniques to find the top K words in a very large file i.e. different techniques to process Big Data.
Simple CSV parser for huge volumes of data with the use of the library Pandas for Python for getting specific columns of a CSV file and putting the extracted data into one or more files (each column in a separated file or all of them in the same output) in a short amount of time.
Sentiment-Analysis-API
The following readme file, assume that before running the Spark analytic job, you have already installed the correct versions of **Java**, **Hadoop**, **Spark** and that you are inside **Ubuntu**.
excel, markdown, csv, sql 数据源批量/单独格式互相转换
Implementation of algorithms for big data using python, numpy, pandas.
To associate your repository with the big-data-processing topic, visit your repo's landing page and select "manage topics."