Sandbox to test out ideas for profiling document data
-
Updated
Feb 27, 2023 - JavaScript
Sandbox to test out ideas for profiling document data
This C++ project profiles and cleanses student data, identifying anomalies and generating statistics, developed for an OOP course using data profiling and validation techniques.
DataBridge Quality Control
End-to-end data preprocessing pipeline using the IBM HR Employee Attrition dataset with Python, Pandas, and Google Colab.
VQ8 Data Profiler - High-performance CSV/TXT file profiler with exact metrics for PHI data
A simple widget for interactive EDA / QA. Works on top of Pandas [in Jupyter Notebook] using IPyWidgets with a sprinkle of Regex.
Power BI Data Analysis & Visualization Projects: A comprehensive collection of interactive Power BI dashboards showcasing advanced data modeling, DAX scripting, and storytelling techniques. This repository demonstrates practical business intelligence applications for reporting automation, data cleansing (ETL), and strategic decision-making support.
A Python toolkit for imputing, synthesizing, and validating tabular data using AI-driven profiling, GAN-based generation, and automated quality assessment.
A limited, dependency-free CSV quality preview for row counts, duplicate IDs, and missing cells.
Data quality for Spark that runs where your data already is. SODA-style metric checks, SQL rules, schema contracts — plus a valid/invalid split that tells you which rule rejected each row. Pure PySpark, no extra services.
Data Preprocessing & Feature Engineering — Customer Churn Prediction | THE PARTH SHAH
AI-powered data quality agent — profile, score and fix any dataset in one command. 📦 PyPI: pip install parseiq | 🌐 Web UI
🚀 Modern, intuitive web app for cleaning, profiling, and exporting Excel & CSV datasets. Built with Streamlit, Pandas & OpenPyXL.
🚚 Agile Data Science Workflows made easy with Pyspark
This course will teach students to use popular tools for sourcing data, transforming it, building and optimizing models, communicating these as visual stories, and deploying them in production.
Generates automated EDA reports using YData Profiling for quick data understanding.
Rule-based data quality analysis and risk assessment tool with optional LLM-powered insights.
Go library for real-time data profiling, dynamic JSON schema inference, type probability estimation, and field optionality tracking.
project in process
To associate your repository with the data-profiling topic, visit your repo's landing page and select "manage topics."