Big Data serverless application
-
Updated
Feb 15, 2020 - Python
Big Data serverless application
Canari v3 - next gen Maltego framework for rapid remote and local transform development
Secondo progetto per il corso di Big Data durante A.A. 2020/2021. Progetto che si incentra sulla creazione di una architettura lambda per effettuare analisi streaming e batch su dati estratti dalla piattaforma Twitch.
A project which includes simulation of real time queries by kafka and performing stream and batch processing of the simulated queries by spark. Also, this follows lambda architecture, in which kafka is publisher and spark helps in subscribing
Create data pipeline using Lambda architecture with Spark, Kafka, Airflow and Snowflake
This project aims to predict smartphone prices using a combination of batch and stream processing techniques in a Big Data environment. The architecture follows the Lambda Architecture pattern, providing both real-time and batch processing capabilities to users.
Electrical Consumption Monitoring - Big Data Pipeline using Lambda Architecture in Python
A data and analytics engineering platform designed for real-time sports betting analytics.
This project implements an end-to-end techstack for a data platform, for local development.
Architecture Big Data temps réel (Lambda, HDFS-first) pour l'analyse de signaux planétaires — Kafka, Hadoop, Spark, MongoDB
Lambda Architecture for real-time traffic speed (TomTom) and air quality (OpenAQ) data in Ho Chi Minh City. Streams to Kafka/JSON, processes with PySpark, persists in Lakehouse (MinIO + Delta Lake), and highlights critical periods where PM2.5 spikes align with congestion, supporting smart city traffic and health decisions.
Python-based streaming analytics platform combining real-time metrics with batch historical analysis. Uses a Kafka-compatible event stream, Redis speed layer, and AWS S3 + Athena pipeline — a practical implementation of the Lambda Architecture.
Big Data project integrating Polymarket prediction data and Binance cryptocurrency rates to analyse relationships between market expectations and real prices.
Modern büyük veri araçlarıyla (Spark, Flink, Kafka) Twitter duygu analizi mimarisi. Veri işleme, anlık alarm mekanizması ve çoklu veri saklama (Parquet/Avro) formatlarını içeren full-stack data engineering projesi.
GCP Lambda architecture with Batch (Parquet ETL) and Streaming (Pub/Sub → Dataflow → BigQuery).
Spark-Streaming-Crypto is a high-performance Big Data pipeline designed for the real-time ingestion, processing, and analysis of cryptocurrency market data and its correlation with media events. Leveraging a Lambda Architecture, this project demonstrates how to handle high-velocity data streams to provide actionable insights into market volatility.
Data Pipeline for Electronic Health Records
To associate your repository with the lambda-architecture topic, visit your repo's landing page and select "manage topics."