borghei/data-quality-auditor
> Audit data quality across pipelines, warehouses, and stores. Use when designing a DQ program, defining DQ dimensions, building rule-based checks, detecting schema drift, monitoring freshness SLAs, or responding to a DQ incident.
npx skills add https://github.com/borghei/Claude-Skills --skill data-quality-auditor
End-to-end data quality (DQ) practice: define DQ dimensions, write rule-based checks, detect schema drift, monitor freshness SLAs, respond to DQ incidents, build a maturity-graded program. Tool-agnostic — works whether you use Great Expectations, dbt tests, Soda Core, Monte Carlo, custom SQL, or hand-rolled scripts.
This skill is audit-focused, not pipeline-focused. For pipeline design, ETL, Spark/dbt, see engineering/senior-data-engineer.
| Situation | Skill applies |
|-----------|---------------|
| Setting up DQ from scratch on a new pipeline | Yes — start with DQ dimensions + check catalog |
| Auditing existing pipelines for missing DQ | Yes — dq_check_runner.py |
| Detecting schema drift in upstream sources | Yes — schema_drift_detector.py |
| Monitoring freshness / SLA on data assets | Yes — freshness_monitor.py |
| Responding to a DQ incident (bad data in prod) | Yes — incident response playbook |
| Designing a DQ governance model | Yes — DQ maturity model |
| Compliance evidence (SOC 2 PI1, GDPR, ISO 27001) | Yes — checks produce auditable artifacts |
| Building data pipelines for the first time | Use engineering/senior-data-engineer first |
Before running the audit, confirm these inputs. If any is unknown or vague, ASK — do not assume:
--data)dq_check_runner.py vs schema_drift_detector.py vs freshness_monitor.py)--max-age-min)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
Industry-standard taxonomy. Every dataset should have at least one check per dimension when at production stage.
| Dimension | Question | Example check |
|-----------|----------|---------------|
| Completeness | Are required fields populated? | users.email IS NOT NULL — fail if > 0.1% nulls |
| Accuracy | Do values match reality? | Reconciliation against source-of-truth system; sample-based human review |
| Consistency | Do values agree across systems / time? | users.email in DB matches Salesforce; row count today within 5% of yesterday |
| Timeliness / Freshness | Is data current to expectation? | events_table.max(event_time) is < 1h old; pipeline runs SLA |
| Validity | Do values conform to format / schema / business rules? | Email regex matches; country code in ISO 3166-1; status in known enum |
| Uniqueness | Are entities not duplicated? | users.user_id is unique; no two rows with same (user_id, day) |
Some teams add: Integrity (referential — FKs resolve), Conformity (matches a published standard), Reasonableness (passes basic sanity checks beyond strict validity).
Checks group into five categories applied per dataset — Volume, Freshness, Schema, Values, and Distribution. See the category summary and the full ~50-pattern catalog in references/dq-check-catalog.md.
| Tool | Purpose | Command |
|------|---------|---------|
| dq_check_runner.py | Run/profile DQ checks against tabular data; per-table pass/fail/warning with value vs threshold | python scripts/dq_check_runner.py --data t.json --checks checks.json --format json |
| schema_drift_detector.py | Diff a current schema against a baseline snapshot (added/removed/changed columns, types, ordinals) | python scripts/schema_drift_detector.py --baseline base.json --current cur.json |
| freshness_monitor.py | Check a freshness SLA: current age vs max-age budget, alerting-ready output | python scripts/freshness_monitor.py --data t.json --column updated_at --max-age-min 60 |
All scripts: stdlib only, argparse CLI, JSON or human-readable output (see Scope re: live DB integration).
Load the reference that matches the task — keep this file lean and pull detail on demand:
This skill covers:
This skill does NOT cover:
engineering/senior-data-engineer.engineering/senior-data-engineer — pipeline design, ETL, dbt, Sparkengineering/observability-designer — observability for data infrastructure (adjacent to DQ)engineering/chaos-engineering — DQ checks benefit from chaos testingra-qm-team/gdpr-dsgvo-expert — DQ underpins GDPR Art. 5(1)(d) "accuracy"ra-qm-team/soc2-compliance-expert — SOC 2 PI1 (Processing Integrity) requires DQ controlsTake borghei/data-quality-auditor from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.