Declarative data-quality + load probes for Python. Rust + tokio under the hood.
-
Updated
Jul 17, 2026 - Rust
Declarative data-quality + load probes for Python. Rust + tokio under the hood.
Orchestrated a Databricks ingestion and dbt analytics pipeline with Apache Airflow 3.3, using deferrable remote-job execution, parallel model branches, data tests, and a containerized CeleryExecutor stack.
Translating between two sets of notation for Kalman filters
Local-first CLI agent for reviewable data test suggestions, with safe profiling, optional bounded LLM generation, deterministic validation, local execution, and human review reports.
Credit Risk Classification
End-to-end data engineering pipeline: Open-Meteo APIs → Python → Snowflake star schema → Airflow orchestration → Groq AI summaries
Data-centric QA automation framework for validating ETL/ELT pipelines in Big Data environments — YAML-driven, Great Expectations validations, AWS S3 and Delta Lake ready.
Automated ETL pipeline (Kestra + DuckDB) for a wine merchant case study - multi-source reconciliation, revenue reporting, and z-score premium detection.
ETL data pipeline (API → SQLite → Tableau) with a pytest/pandas data-quality test suite and CI.
A professional ELT Data Warehouse built with Python and dbt-core, transforming fragmented CRM & ERP data into actionable sales intelligence for SMEs.
Framework d'observabilité de données léger, conteneurisé et 100% Open-Source.
End-to-end financial data quality platform enforcing data contracts, validation gates, and observability across ingestion, transformation, and analytics layers using Airflow, dbt, Snowflake, and Soda.
Dynamic data testing engine based on pySpark
Top Data Diff & Regression Testing (Opensource) 🌟 Star if you like it! 🌟
Superconductive — independent third-party profile of a public API surface, by API Evangelist. Superconductive Inc. is the company behind Great Expectations (GX), an open-source data quality framework used to validate, document, and profile data across modern data stacks. Teams declare "Expectations" — assertions about how data should behave — and G
Reusable dbt data-quality tests for healthcare: NHS number checksums, ICD-10/CPT format, date sanity, staging-to-mart reconciliation, duplicate claims and null-rate checks, proven on DuckDB by a labeled synthetic harness (1.0 recall over 160 violations, 0 false positives) with a severity-weighted scorecard from dbt artifacts.
Soda — independent third-party profile of a public API surface, by API Evangelist. Soda is an AI-native, fully automated data quality platform that helps data engineering and analytics teams define data quality checks, scan datasets, monitor data freshness, and manage data quality incidents.
Synthetic CSV test-data generator for QA, demos, ingestion tests, and repeatable edge-case validation.
Data Migration Testing | Data Reconcilation
Develop a data science project using historical sales data to build a regression model that accurately predicts future sales. Preprocess the dataset, conduct exploratory analysis, select relevant features, and employ regression algorithms for model development. Evaluate model performance, optimize hyperparameters, and provide actionable insights.
To associate your repository with the data-testing topic, visit your repo's landing page and select "manage topics."