Build a MongoDB NoSQL analytics platform for Airbnb data with disaster recovery, high availability, and geo-distributed sharding
-
Updated
Oct 6, 2026 - Python
Build a MongoDB NoSQL analytics platform for Airbnb data with disaster recovery, high availability, and geo-distributed sharding
Build Python data transformation pipelines with a visual IDE and AI assistant
Scores a five-day forecast into a crop risk rating from deterministic per-crop thresholds. FastAPI, SQLAlchemy, React, Tailwind.
End-to-end European income distribution analytics pipeline using Eurostat data. Extracts and transforms data with Python and Pandas, models it in a star schema, loads it into PostgreSQL for SQL analysis, and delivers an interactive Power BI dashboard for exploring income trends across countries and demographics.
介護制度の公式資料を構造化し、実務の疑問から根拠へ辿る公開DB。制度理解・データ設計・独立照合・現行性と人手確認を分ける品質管理の個人プロジェクト。
Multi-tenant data and AI platform for ~40 SaaS customers: event ingestion, dbt daily marts, a Gemini AI analyst with fixed tools, and Postgres RLS tenant isolation proven by CI tests. Includes CSV self-serve upload, org deletion with receipt, cost controls and kill switches. FastAPI, Supabase, Render, GitHub Actions.
自治体のAI調達を公式一次資料から比較・構造化する個人研究。仕様・質疑・訂正・評価・公開結果を接続し、有効要件と公開証拠の限界を管理。
Linked Open Data Modeling Language
🛢️ Visualize and forecast oil & gas well production with a user-friendly dashboard for data cleaning, engineering, and machine learning training.
Agentic AI for Data Vault 2.0 — requirements to contract-backed dbt code
[In progress] A comprehensive wrapper around Winnipeg Transit's API that integrates GTFS data to provide customized real-time bus location and service information, while maintaining historical records of buses and runs. The plan is to eventually integrate this application into my other transit-related projects.
Type annotations for specifying, validating, and serializing arrays with arbitrary backends in Pydantic (and beyond)
Point-in-time SEC fundamentals warehouse (dbt + DuckDB) that models financial restatements to eliminate lookahead bias
End-to-end data engineering project using Medallion Architecture, synthetic data and a Power BI Semantic Model.
SaaS growth analytics: activation, retention, churn and experimentation with SQL models, Python and Streamlit
dbt warehouse and long-stay prediction model on 173,000 Austin Animal Center records. Austin's live release rate held steady while dog adoption waits tripled; the model flags at intake which dogs and cats are likely to stay 30+ days. Tested end to end, with backtest monitoring and leakage checks in CI.
Hospital-operations analytics on 12,000 synthetic visits (FY 2024-25) — MySQL schema + KPI queries, seeded Python data generator, Power BI star schema with DAX. Revenue mix, admissions, LOS & discharge outcomes.
Analytics engineering reference project: 40 dbt models, 128 tests, SCD2 snapshots and margin-aware channel ROAS. Clone and run - builds on DuckDB with no credentials, verified end-to-end on Snowflake and Redshift. Optional Dagster orchestration, or integrate with your own Airflow.
To associate your repository with the data-modeling topic, visit your repo's landing page and select "manage topics."