db(列存储-大数据-实时分析数据库)ClickHouse® is a free analytics DBMS for big data
-
Updated
Feb 18, 2021 - C++
db(列存储-大数据-实时分析数据库)ClickHouse® is a free analytics DBMS for big data
Este repositório abriga o projeto acadêmico da disciplina de Tópicos de Big Data em Python. O projeto analisa os dados da pesquisa anual "State of Data Brazil", realizada pela comunidade Data Hackers em parceria com a Bain & Company.
Automated Cloud Data Engineering pipeline using Google Cloud Platform (GCP)
A search engine implementation for large-scale datasets Key features include large-scale data processing, fast indexing, advanced search capabilities, and scalable architecture
Databricks and PySpark project for airline delay analysis, visualization, machine learning prediction, and MongoDB NoSQL storage using U.S. flight data from 2016 to 2018, including more than 18.5 million flight records in the dataset.
Understanding Big Data Analytics by using Map Reduce for performing various tasks like Blooms Filter, Frequent Itemset, KMeans, Matrix Multiplication, Finding Maximum Temperature, Finding Word Count, and Analyzing Electricity Consumption
Big Data & Cloud Computing project for recommendation, cluster analysis, data visualization with Hadoop and Spark deployed in auto- scaling cloud environment, youtube link:
📁 Multi-label classification of printed media articles to topics
Predict the number of retweets that a tweet about a specific museum will have.
In this jupyter notebook file, fictional data of football players was used to perform big data analytics in python. It involves using librarires such as pandas and matplotlib.
Real-Time Analysis of Linguistic Media
→ B.Tech big data analytics laboratory implementations utilizing Scala and distributed computing frameworks.
This is data engineering pipeline that ingests Canada's Express Entry immigration data (admissions and invited candidates) from IRCC's Open Government Portal, transforms it with dbt, and lands it in BigQuery.
Big Data Exploratory Data Analysis with Spark on Scala
A lossy counting algorithm implemented to determine the top trending hashtags using the Twitter API to get a continuous stream of tweets.
Big Data using PySpark, AWS, and pgAdmin
Sentiment Analysis and Topic Modeling on the Steam Game Reviews using Hadoop and Mahout
Ben Gurion University "The Art of Analyzing Big Data - The Data Scientist’s Toolbox (372.2.5401)" course assignments & solutions
ETH analysis using big data for the QMUL Big Data Processing module. Intended to promote analysis of data retrieved via big data processing
10alytics_air_realtor ,a dynamic repository, hosts an AWS-driven data pipeline. Utilizing Apache Airflow, AWS S3, and EC2, it performs efficient ETL operations, extracting comprehensive real estate data from the Realty Mole Property API via RapidAPI. This tool empowers real estate professionals with timely insights for strategic decision-making.
To associate your repository with the big-data-analytics topic, visit your repo's landing page and select "manage topics."