Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
-
Updated
Jul 2, 2026 - Python
Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
An elegant, opinionated framework for deploying BrightHive Data Resources with zero coding.
Databricks-native intelligent data transformation engine — coherence-scored Bronze/Silver/Gold with entity resolution and temporal reconciliation in a single deployable product.
Open standard and Python toolkit for AI-ready, traceable, and auditable documents. GalloDoc turns raw data into structured, verifiable decision artifacts.
A Databricks control pattern that certifies every record before downstream consumption. 7 contract checks, replay detection, schema drift handling, and quarantine with explicit reasons. 56 passing tests. Databricks Free Edition validated. Enterprise Data Trust, Chapter 1.
A reproducible benchmark that scores data controls against known failure scenarios with precision, recall, and ground truth. Custom approach achieved perfect recall; industry baselines missed injected drift. 37 passing tests, 10/10 gates. Enterprise Data Trust, Chapter 3.
A release control that detects when business columns collapse despite healthy schema and row counts. Distribution stability scoring, 6 publication gates, and blocked Gold refresh when the health score dropped from 1.0 to 0.20. 50 passing tests. Databricks Free Edition validated. Enterprise Data Trust, Chapter 2.
A fruit fly's brain beat Claude, GPT-5, Grok and Gemini at judging health records. Code, data and paper for SuperTruth's connectome study: the MaleCNS v1.0 fly nervous system held fixed and trained to reproduce the Data Trust Index on synthetic health records, against five controls. Pre-registered, published win or lose. DOI 10.5281/zenodo.22865215
To associate your repository with the data-trust topic, visit your repo's landing page and select "manage topics."