This is a repo with links to everything you'd ever want to learn about data engineering
-
Updated
Aug 3, 2026 - Jupyter Notebook
This is a repo with links to everything you'd ever want to learn about data engineering
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
High-performance, scalable time-series database designed for Industrial IoT (IIoT) scenarios
GridDB is a next-generation open source database that makes time series IoT and big data fast,and easy.
A curated list of awesome big data frameworks, ressources and other awesomeness.
Upserts, Deletes And Incremental Processing on Big Data.
RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.
A Cloud Native Batch System (Project under CNCF)
100+套大数据可视化炫酷大屏Html5模板;包含行业:社区、物业、政务、交通、金融银行等,全网最新、最多,最全、最酷、最炫大数据可视化模板。陆续更新中
JuiceFS is a distributed POSIX file system built on top of Redis and S3.
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
🔨 用 JSON 来生成结构化的 SQL 语句,基于 Vue3 + TypeScript + Vite + Ant Design + MonacoEditor 实现,项目简单(重逻辑轻页面)、适合练手~
Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.
To associate your repository with the bigdata topic, visit your repo's landing page and select "manage topics."