borghei/senior-data-engineer
> Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or troubleshooting pipeline failures.
npx skills add https://github.com/borghei/Claude-Skills --skill senior-data-engineer
Generate pipeline configurations (Airflow, Prefect, Dagster), validate data quality with profiling and anomaly detection, and optimize SQL/Spark performance with actionable recommendations.
Before generating pipelines, confirm these inputs. If any is unknown or vague, ASK — do not assume:
--type; changes the generated DAG code)--source/--destination/--mode; shapes the pipeline)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
# Generate an Airflow DAG for incremental PostgreSQL -> Snowflake
python scripts/pipeline_orchestrator.py generate \
--type airflow --source postgres --destination snowflake \
--tables orders,customers --mode incremental --schedule "0 5 * * *"
# Validate data quality against a schema
python scripts/data_quality_validator.py validate data.csv \
--schema schema.json --detect-anomalies --json
# Profile a dataset
python scripts/data_quality_validator.py profile data.csv --json
# Optimize a slow SQL query
python scripts/etl_performance_optimizer.py analyze-sql query.sql \
--warehouse snowflake --json
# Estimate query cost
python scripts/etl_performance_optimizer.py estimate-cost query.sql \
--warehouse bigquery --stats data_stats.json --json
| Tool | Subcommands | Purpose |
|------|-------------|---------|
| pipeline_orchestrator.py | generate, validate, template | Generate Airflow/Prefect/Dagster pipeline code, validate DAGs |
| data_quality_validator.py | validate, profile, generate-suite, contract, schema | Schema validation, profiling, anomaly detection, Great Expectations |
| etl_performance_optimizer.py | analyze-sql, analyze-spark, optimize-partition, estimate-cost, template | SQL/Spark optimization, partition strategy, cost estimation |
All subcommands support --json for machine-readable output and --output for file writing.
Load the reference that matches the task — keep this file lean and pull detail on demand:
| Skill | Integration |
|-------|-------------|
| senior-data-scientist | Feature engineering consumes curated mart data |
| senior-ml-engineer | ML pipelines depend on feature store tables |
| senior-devops | CI/CD for dbt, Airflow deployment, container orchestration |
| senior-architect | Architecture reviews for lakehouse vs warehouse decisions |
| code-reviewer | Pipeline code reviews for DAGs, dbt models, Spark jobs |
Take borghei/senior-data-engineer from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.