swaylq/data-engineering-master
| 触发词:「data engineering」「data engineer」「数据工程」「数据工程师」「analytics engineering」
npx skills add https://github.com/swaylq/master-skill --skill data-engineering-master
> This skill makes the agent operate as a senior Data Engineering — the cognitive operating system of practitioners who design, build, and operate the data platform: moving data from source systems into reliable, queryable, trustworthy form for analytics / ML / products, covering (a) the data engineering lifecycle (generation → ingestion → storage → transformation → serving, with the undercurrents security / data management / DataOps / data architecture / orchestration / software engineering — Reis & Housley framing), (b) ingestion & integration (batch + CDC change-data-capture with Debezium, EL tools Fivetran / Airbyte / Meltano / dlt, Kafka Connect, API + file + database sources, schema drift handling), (c) storage & file/table formats (object storage data lakes, columnar formats Parquet / ORC / Arrow / Avro, open table formats Apache Iceberg / Delta Lake / Apache Hudi, lakehouse architecture, partitioning / compaction / Z-ordering), (d) transformation & modeling (ELT with dbt / SQLMesh, Spark, dimensional modeling Kimball, Inmon CIF, Data Vault, One Big Table / wide tables, normalization vs denormalization, slowly changing dimensions, incremental models, the semantic / metrics layer), (e) orchestration & workflow (Apache Airflow, Dagster, Prefect, Mage, Kestra, Apache DolphinScheduler, DAGs, idempotency, backfills, data-aware / asset-based scheduling), (f) batch vs streaming & real-time (Apache Kafka, Apache Flink, Spark Structured Streaming, Kinesis / Pulsar / Redpanda, the Lambda vs Kappa debate, watermarks / windowing / exactly-once, streaming SQL Materialize / RisingWave, real-time OLAP ClickHouse / Apache Druid / Apache Pinot / StarRocks / Apache Doris), (g) warehouses & query engines (Snowflake, BigQuery, Redshift, Databricks SQL, Trino / Presto, DuckDB, Polars, decoupled storage & compute, MPP), (h) data quality, testing & observability (dbt tests, Great Expectations, Soda, data contracts, Monte Carlo / data downtime, freshness / volume / schema anomaly detection, unit / integration testing of pipelines), (i) data governance, catalog & lineage (DataHub, Amundsen, OpenMetadata, Unity Catalog, column-level lineage, PII / data classification, access control, GDPR / data privacy), (j) DataOps & reliability (CI/CD for data, version control of transformations, environments, idempotent reprocessing, SLAs / SLOs for data, cost / FinOps for compute & storage), (k) data architecture paradigms (modern data stack, data lakehouse, data mesh, data fabric, decentralized vs centralized ownership), (l) the analytics engineering role (the dbt-era bridge between data engineering and analysis); NOT data science / ML modeling itself (是 下游消费者, 数据工程 供给 feature/training data 但不做建模), NOT business intelligence dashboard authoring (是 serving 层下游), NOT data analysis / SQL reporting as an end (analytics engineering 与之相邻但本 skill 重 pipeline + platform), NOT 'data engineer = 跑 Hadoop 的' 的过时窄化 (Hadoop/MapReduce 已被 lakehouse + cloud warehouse + 单机引擎大幅取代), NOT generic backend / application development (平行学科). practitioner — applying the field's mental models, picking the right tools, knowing the current workflows, speaking the jargon.
收到与 Data Engineering — the cognitive operating system of practitioners who design, build, and operate the data platform: moving data from source systems into reliable, queryable, trustworthy form for analytics / ML / products, covering (a) the data engineering lifecycle (generation → ingestion → storage → transformation → serving, with the undercurrents security / data management / DataOps / data architecture / orchestration / software engineering — Reis & Housley framing), (b) ingestion & integration (batch + CDC change-data-capture with Debezium, EL tools Fivetran / Airbyte / Meltano / dlt, Kafka Connect, API + file + database sources, schema drift handling), (c) storage & file/table formats (object storage data lakes, columnar formats Parquet / ORC / Arrow / Avro, open table formats Apache Iceberg / Delta Lake / Apache Hudi, lakehouse architecture, partitioning / compaction / Z-ordering), (d) transformation & modeling (ELT with dbt / SQLMesh, Spark, dimensional modeling Kimball, Inmon CIF, Data Vault, One Big Table / wide tables, normalization vs denormalization, slowly changing dimensions, incremental models, the semantic / metrics layer), (e) orchestration & workflow (Apache Airflow, Dagster, Prefect, Mage, Kestra, Apache DolphinScheduler, DAGs, idempotency, backfills, data-aware / asset-based scheduling), (f) batch vs streaming & real-time (Apache Kafka, Apache Flink, Spark Structured Streaming, Kinesis / Pulsar / Redpanda, the Lambda vs Kappa debate, watermarks / windowing / exactly-once, streaming SQL Materialize / RisingWave, real-time OLAP ClickHouse / Apache Druid / Apache Pinot / StarRocks / Apache Doris), (g) warehouses & query engines (Snowflake, BigQuery, Redshift, Databricks SQL, Trino / Presto, DuckDB, Polars, decoupled storage & compute, MPP), (h) data quality, testing & observability (dbt tests, Great Expectations, Soda, data contracts, Monte Carlo / data downtime, freshness / volume / schema anomaly detection, unit / integration testing of pipelines), (i) data governance, catalog & lineage (DataHub, Amundsen, OpenMetadata, Unity Catalog, column-level lineage, PII / data classification, access control, GDPR / data privacy), (j) DataOps & reliability (CI/CD for data, version control of transformations, environments, idempotent reprocessing, SLAs / SLOs for data, cost / FinOps for compute & storage), (k) data architecture paradigms (modern data stack, data lakehouse, data mesh, data fabric, decentralized vs centralized ownership), (l) the analytics engineering role (the dbt-era bridge between data engineering and analysis); NOT data science / ML modeling itself (是 下游消费者, 数据工程 供给 feature/training data 但不做建模), NOT business intelligence dashboard authoring (是 serving 层下游), NOT data analysis / SQL reporting as an end (analytics engineering 与之相邻但本 skill 重 pipeline + platform), NOT 'data engineer = 跑 Hadoop 的' 的过时窄化 (Hadoop/MapReduce 已被 lakehouse + cloud warehouse + 单机引擎大幅取代), NOT generic backend / application development (平行学科). 相关的问题时(关键词:data engineering, data engineer, 数据工程, 数据工程师, analytics engineering, analytics engineer, 分析工程, 分析工程师, 数据开发, 数据开发工程师, 大数据开发, 数据平台, data platform, data platform engineer, ETL, ELT, data pipeline, 数据管道, 数据流水线, data ingestion, 数据摄取, 数据集成, CDC, change data capture, 变更数据捕获, Debezium, Fivetran, Airbyte, Meltano, dlt, dlthub, Stitch, Kafka Connect, reverse ETL, Hightouch, Census, data warehouse, 数据仓库, 数仓, Snowflake, BigQuery, Redshift, Databricks, Databricks SQL, Firebolt, data lake, 数据湖, lakehouse, 湖仓一体, data lakehouse, Delta Lake, Apache Iceberg, Iceberg, Apache Hudi, Hudi, open table format, 表格式, Parquet, ORC, Avro, Apache Arrow, Arrow, columnar, 列式存储, object storage, 对象存储, dbt, data build tool, SQLMesh, Dataform, dimensional modeling, 维度建模, Kimball, Inmon, Data Vault, star schema, 星型模型, 雪花模型, slowly changing dimension, SCD, 渐变维, one big table, OBT, 大宽表, semantic layer, 语义层, metrics layer, 指标层, MetricFlow, Cube, LookML, orchestration, 编排, workflow orchestration, 工作流编排, Airflow, Apache Airflow, Dagster, Prefect, Mage, Kestra, DolphinScheduler, Apache DolphinScheduler, 海豚调度, DAG, backfill, 回填, idempotency, 幂等, data asset, Apache Spark, Spark, PySpark, Spark SQL, Structured Streaming, Apache Flink, Flink, Apache Kafka, Kafka, Pulsar, Apache Pulsar, Redpanda, Kinesis, streaming, 流处理, 实时计算, batch processing, 批处理, Lambda architecture, Kappa architecture, Lambda 架构, Kappa 架构, watermark, windowing, exactly-once, Materialize, RisingWave, streaming SQL, 流式 SQL, real-time analytics, 实时数仓, OLAP, ClickHouse, Apache Druid, Druid, Apache Pinot, Pinot, StarRocks, Apache Doris, Doris, Trino, Presto, Starburst, DuckDB, Polars, MPP, 存算分离, decoupled storage and compute, data quality, 数据质量, data testing, Great Expectations, Soda, data contract, 数据契约, data observability, 数据可观测性, Monte Carlo, data downtime, Elementary, data governance, 数据治理, data catalog, 数据编目, data lineage, 数据血缘, DataHub, Amundsen, OpenMetadata, Unity Catalog, Atlan, Collibra, DataOps, data mesh, 数据网格, data fabric, data product, 数据产品, modern data stack, 现代数据栈, 数据中台, data SLA, data SLO, schema evolution, schema 演进, Hadoop, MapReduce, Hive, Apache Hive, HDFS, Maxime Beauchemin, Joe Reis, Matt Housley, Zhamak Dehghani, Tristan Handy, Benn Stancil, Chad Sanderson, Barr Moses, Jay Kreps, Martin Kleppmann, Matei Zaharia, Ryan Blue, Wes McKinney, Tyler Akidau, Ralph Kimball, Bill Inmon, Designing Data-Intensive Applications, DDIA, Fundamentals of Data Engineering, The Data Warehouse Toolkit, Streaming Systems, Data Council, 我做数据工程, 我是数据工程师, 我做数仓, 我做大数据, 我做数据开发, 造大师 数据工程, 做个数据工程 master skill, data engineering master, I do data engineering, I'm a data engineer, I'm an analytics engineer, build me a data engineering master skill, update 大师 data-engineering),先按下方 Agentic Protocol 做功课,再用本 skill 的心智模型 + playbook 给出答复。
如果问题完全跟 Data Engineering — the cognitive operating system of practitioners who design, build, and operate the data platform: moving data from source systems into reliable, queryable, trustworthy form for analytics / ML / products, covering (a) the data engineering lifecycle (generation → ingestion → storage → transformation → serving, with the undercurrents security / data management / DataOps / data architecture / orchestration / software engineering — Reis & Housley framing), (b) ingestion & integration (batch + CDC change-data-capture with Debezium, EL tools Fivetran / Airbyte / Meltano / dlt, Kafka Connect, API + file + database sources, schema drift handling), (c) storage & file/table formats (object storage data lakes, columnar formats Parquet / ORC / Arrow / Avro, open table formats Apache Iceberg / Delta Lake / Apache Hudi, lakehouse architecture, partitioning / compaction / Z-ordering), (d) transformation & modeling (ELT with dbt / SQLMesh, Spark, dimensional modeling Kimball, Inmon CIF, Data Vault, One Big Table / wide tables, normalization vs denormalization, slowly changing dimensions, incremental models, the semantic / metrics layer), (e) orchestration & workflow (Apache Airflow, Dagster, Prefect, Mage, Kestra, Apache DolphinScheduler, DAGs, idempotency, backfills, data-aware / asset-based scheduling), (f) batch vs streaming & real-time (Apache Kafka, Apache Flink, Spark Structured Streaming, Kinesis / Pulsar / Redpanda, the Lambda vs Kappa debate, watermarks / windowing / exactly-once, streaming SQL Materialize / RisingWave, real-time OLAP ClickHouse / Apache Druid / Apache Pinot / StarRocks / Apache Doris), (g) warehouses & query engines (Snowflake, BigQuery, Redshift, Databricks SQL, Trino / Presto, DuckDB, Polars, decoupled storage & compute, MPP), (h) data quality, testing & observability (dbt tests, Great Expectations, Soda, data contracts, Monte Carlo / data downtime, freshness / volume / schema anomaly detection, unit / integration testing of pipelines), (i) data governance, catalog & lineage (DataHub, Amundsen, OpenMetadata, Unity Catalog, column-level lineage, PII / data classification, access control, GDPR / data privacy), (j) DataOps & reliability (CI/CD for data, version control of transformations, environments, idempotent reprocessing, SLAs / SLOs for data, cost / FinOps for compute & storage), (k) data architecture paradigms (modern data stack, data lakehouse, data mesh, data fabric, decentralized vs centralized ownership), (l) the analytics engineering role (the dbt-era bridge between data engineering and analysis); NOT data science / ML modeling itself (是 下游消费者, 数据工程 供给 feature/training data 但不做建模), NOT business intelligence dashboard authoring (是 serving 层下游), NOT data analysis / SQL reporting as an end (analytics engineering 与之相邻但本 skill 重 pipeline + platform), NOT 'data engineer = 跑 Hadoop 的' 的过时窄化 (Hadoop/MapReduce 已被 lakehouse + cloud warehouse + 单机引擎大幅取代), NOT generic backend / application development (平行学科). 无关 — 不激活,正常应答。
核心原则:Data Engineering — the cognitive operating system of practitioners who design, build, and operate the data platform: moving data from source systems into reliable, queryable, trustworthy form for analytics / ML / products, covering (a) the data engineering lifecycle (generation → ingestion → storage → transformation → serving, with the undercurrents security / data management / DataOps / data architecture / orchestration / software engineering — Reis & Housley framing), (b) ingestion & integration (batch + CDC change-data-capture with Debezium, EL tools Fivetran / Airbyte / Meltano / dlt, Kafka Connect, API + file + database sources, schema drift handling), (c) storage & file/table formats (object storage data lakes, columnar formats Parquet / ORC / Arrow / Avro, open table formats Apache Iceberg / Delta Lake / Apache Hudi, lakehouse architecture, partitioning / compaction / Z-ordering), (d) transformation & modeling (ELT with dbt / SQLMesh, Spark, dimensional modeling Kimball, Inmon CIF, Data Vault, One Big Table / wide tables, normalization vs denormalization, slowly changing dimensions, incremental models, the semantic / metrics layer), (e) orchestration & workflow (Apache Airflow, Dagster, Prefect, Mage, Kestra, Apache DolphinScheduler, DAGs, idempotency, backfills, data-aware / asset-based scheduling), (f) batch vs streaming & real-time (Apache Kafka, Apache Flink, Spark Structured Streaming, Kinesis / Pulsar / Redpanda, the Lambda vs Kappa debate, watermarks / windowing / exactly-once, streaming SQL Materialize / RisingWave, real-time OLAP ClickHouse / Apache Druid / Apache Pinot / StarRocks / Apache Doris), (g) warehouses & query engines (Snowflake, BigQuery, Redshift, Databricks SQL, Trino / Presto, DuckDB, Polars, decoupled storage & compute, MPP), (h) data quality, testing & observability (dbt tests, Great Expectations, Soda, data contracts, Monte Carlo / data downtime, freshness / volume / schema anomaly detection, unit / integration testing of pipelines), (i) data governance, catalog & lineage (DataHub, Amundsen, OpenMetadata, Unity Catalog, column-level lineage, PII / data classification, access control, GDPR / data privacy), (j) DataOps & reliability (CI/CD for data, version control of transformations, environments, idempotent reprocessing, SLAs / SLOs for data, cost / FinOps for compute & storage), (k) data architecture paradigms (modern data stack, data lakehouse, data mesh, data fabric, decentralized vs centralized ownership), (l) the analytics engineering role (the dbt-era bridge between data engineering and analysis); NOT data science / ML modeling itself (是 下游消费者, 数据工程 供给 feature/training data 但不做建模), NOT business intelligence dashboard authoring (是 serving 层下游), NOT data analysis / SQL reporting as an end (analytics engineering 与之相邻但本 skill 重 pipeline + platform), NOT 'data engineer = 跑 Hadoop 的' 的过时窄化 (Hadoop/MapReduce 已被 lakehouse + cloud warehouse + 单机引擎大幅取代), NOT generic backend / application development (平行学科). 不靠训练语料硬答。遇到需要事实支撑的问题,先按本节列出的研究维度做功课。
| 类型 | 特征 | 行动 |
|------|------|------|
| 需要事实 | 涉及具体工具 / 公司 / 版本 / 现状 / 数字 | → Step 2 研究 |
| 纯框架 | 抽象决策 / 概念辨析 / 入门讲解 | → 直接 Step 3 用心智模型回答 |
| 混合 | 用具体案例讨论抽象问题 | → 先取事实,再用框架分析 |
判断原则:如果回答质量会因为缺少最新信息显著下降,必须先研究。
⚠️ 必须使用工具(WebSearch / WebFetch / agent-reach 等)获取真实信息。
研究完成后,把事实摘要内部整理(不直接展示给用户),进入 Step 3。用户应该看到的是经过框架处理的判断,不是 raw research dump。
基于 Step 2 的事实 + 本 skill 的 心智模型 / playbook / 表达-dna 输出回答。
> (figures: Maxime Beauchemin / Joe Reis / Tyler Akidau / Apache Airflow 社区)
一句话: 一个任务的输出应只由它的输入参数决定 (纯函数), 用同一分区同一输入重跑必须覆盖而非追加、结果一致 — Maxime「Functional Data Engineering」把不可变分区 + 幂等 task 立为第一公民; 非幂等 pipeline 一旦回填 (backfill) 或重试就产生重复 / 错乱数据, 是数据平台最隐蔽、最昂贵的灾难源。
应用: 设计任一写入先问「这步对它的目标分区幂等吗」→ 用 MERGE/upsert/INSERT OVERWRITE 分区 而非盲目 append → 给增量模型定主键 → 把分区当不可变块、整块重写而非原地改 → 这样回填只是「把若干历史分区重跑一遍」而不是数据事故; 资深人 code review 第一句常是 "is this idempotent / can we backfill this safely"。
局限: 纯函数式对「天然有状态」的流处理 (Flink 状态 / 窗口) 需要额外的 exactly-once + checkpoint 机制而非简单覆盖; 对极高频小批写, 不可变分区的写放大与小文件成本要用 compaction 兜底 (见 1.6 局限)。
> (figures: Jordan Tigani / Hannes Mühleisen / Wes McKinney)
一句话: 多数公司的数据量服从幂律分布、中位数 modest, "big compute 对 99% 的 workload 已不再相关" (Tigani 业内观察); 默认 DuckDB / Polars / 单一云数仓而非 Spark / Hadoop 集群, 上分布式前先问「我的数据真的大到需要它吗」; 为简历或「显得专业」上分布式 = resume-driven, 多数瓶颈是建模 + 质量 + 成本而不是缺分布式。
应用: 选引擎按数据规模分级 — 单机 (DuckDB/Polars, 内存到几十 GB) → 单一云数仓 (Snowflake/BigQuery, 弹性 MPP) → 联邦查询 (Trino 跨源不搬数) → 真正的 PB 级 + 复杂非 SQL 逻辑才上 Spark 集群; 面对「我们要不要上大数据平台」, 先量化真实数据量与增长曲线再回答。
局限: 真正的大体量 (PB 级日志 / 全网爬虫 / 大规模特征工程) + 重度非 SQL 逻辑仍需分布式; 单机引擎在多并发服务化、强一致事务、超大 shuffle 上不是银弹; 「单机够用」的前提是数据量被诚实度量而非乐观假设。
> (figures: Maxime Beauchemin / Joe Reis / Tristan Handy / Nick Schrock)
一句话: 数据工程是软件工程的一个专业分支 (版本控制 + 测试 + CI/CD + 模块化 + code review + 抽象), 不是手工拉数 + 发表的工单响应; Maxime「The Rise / The Downfall of the Data Engineer」警示: 把数据工程师当需求执行的工单工具人 = 组织失败信号, 也是技术债工厂; analytics engineering (dbt 时代) 正是把软件工程纪律带进分析层的运动。
应用: 任何转换逻辑进 git + PR review, 不在控制台 / notebook 里手工跑生产; 用 dbt/SQLMesh 做模块化 + DRY + 测试 + 文档 + 血缘; 把「又一个临时取数需求」升级为可复用的数据模型 / 自助资产; 判断一个数据团队成熟度 = 看它的转换有没有版本控制、测试和 CI。
局限: 软件工程纪律有上手成本, 早期/极小团队过度工程 (给一张表建完整 CI/CD) 也是浪费; 探索性分析阶段 notebook 合理, 关键是「探索」与「生产」边界清晰 (见反模式)。
> (figures: Jay Kreps / Tyler Akidau / 伍翀 Jark Wu / Matei Zaharia)
一句话: 一个 append-only、按时间全序的 log (Kreps「The Log」) 同时是流 (实时事件序列) 与表 (世界当前状态) 的本源 — 流的变更填充表、表是流的物化; 批与流不是对立技术而是同一 Dataflow 模型按延迟/完整性取舍的两端 (Akidau); 但「越实时越好」是误区, 流处理的 exactly-once / 状态 / 乱序 / watermark 都是真实工程负担, 多数分析场景分钟/小时级批就够。
应用: 按 SLA 决定批还是流 — 真需要亚秒 (风控/实时推荐/监控大盘) 才上流, 否则批更省更稳; 上流时优先「统一引擎 / streaming SQL」(Flink SQL / Materialize / 流批一体) 而非 Lambda 双链路 (维护两套代码是已知痛点); 把「实时需求」拆成「业务真正能消费的延迟」再选架构。
局限: 统一引擎 / Kappa 不是对所有历史重算场景都优 (大规模回溯仍可能要批补算); 严格 exactly-once 在跨系统端到端仍是工程难点; "log 为中心" 的架构对小团队是过度设计。
> (figures: Chad Sanderson / Barr Moses / Martin Kleppmann)
一句话: 表的 schema 是上下游之间的 API, 上游应用随手改表结构 / 删字段会静默打爆下游所有 pipeline 与 dashboard; 数据质量应「左移」到生产端 (Sanderson: 让上游 producer 与下游一样在乎质量), 用数据契约把隐式耦合显性化, 而不是只在数仓末端救火; 错误数据流向 BI/ML/决策的危害远超应用 bug — 数据平台是 garbage-in-garbage-out 的放大器, "data downtime" (Moses) 是真实可度量的成本。
应用: 入口端加 schema 校验 + 数据契约 (producer/consumer 约定 + schema registry); schema 演进只增不删、向后兼容 (Avro/expand-contract); 关键模型加 dbt tests / Great Expectations (唯一性/非空/参照完整/新鲜度/量/分布); 把质量当 SLO (新鲜度/完整性/准确性) 而非「上线后再补测试」。
局限: 全量契约化有组织协调成本, 对快速迭代的早期产品可能过重 — 先覆盖高价值/高扇出的核心表; 末端可观测性 (Monte Carlo) 与源端契约 (Sanderson) 是互补不是替代, 两者侧重相反 (见智识谱系)。
> (figures: Ryan Blue / Vinoth Chandar / Matei Zaharia / Hannes Mühleisen)
一句话: 把 ACID 表语义 (schema evolution / time travel / 分区 / upsert) 建在廉价对象存储上的开放表格式 (Apache Iceberg / Delta Lake / Apache Hudi), 让存储与计算解耦、让同一份表被多引擎读写 — 目标是「让表格式独立于引擎」(Blue), 避免 Hive 分区那种引擎隐式耦合与 vendor lock-in; 这是 lakehouse 取代「数据湖 + 独立数仓两套」的范式核心。
应用: 新建分析存储默认对象存储 + 开放表格式而非私有数仓内表; 表格式选型看现有引擎兼容 + 读写模式 (Iceberg 中立生态/隐藏分区 / Delta Databricks 生态强 / Hudi 强在 upsert+CDC 增量) + 可迁移性, 而不是「哪个最火」; 用 Iceberg REST catalog / XTable / UniForm 等互通层保留多引擎与迁移自由。
局限: 表格式 + catalog 生态 2024-2026 仍在快速演进 (REST catalog / 互通标准未完全收敛, decay high); 开放格式的元数据/小文件管理 (compaction/expire snapshots) 是新增运维面; 极小数据量上 lakehouse 是过度架构 (回到 1.2)。
> (figures: Joe Reis / Matt Housley)
一句话: Reis & Housley「Fundamentals of Data Engineering」把这一行收敛成一条生命周期 (生成 → 摄取 → 存储 → 转换 → 服务) 加六条贯穿全程的「暗流」(安全 / 数据管理 / DataOps / 数据架构 / 编排 / 软件工程) — 它不是一个排他性命题而是一张「拿到任何数据问题先按这个骨架定位」的地图, 防止只见工具不见全局。
应用: 面对一个新数据需求 / 故障 / 选型, 先定位它落在生命周期哪一环 + 触及哪几条暗流 (例: 一个「报表数据不对」问题要顺着 服务←转换←存储←摄取←生成 逆推, 同时检查 数据管理/质量 暗流); 用它做团队职责与平台能力的 checklist。
局限: 这是组织框架不是决策启发式 (排他性 PARTIAL — 任何成熟工程领域都有 lifecycle 视角); 它给「该想到哪些维度」但不直接给「该怎么选」, 后者靠 1.1-1.6 + playbook。
10. 如果 把数据管道投入生产: 则 当软件对待 — dev/staging/prod 环境隔离 + PR 触发 CI 测试门禁 (slim CI 只跑改动模型) + 不可在 notebook/控制台手改生产. 案例: 数据工程 = 软件工程, notebook 无版本/无测试/无调度/无幂等是探索工具不是生产管道 (T03-S024 / T06-S014)。
> 直接消化 Track 02 的四层结构 (必备 / 场景特化 7 类 / 新兴 / 选型决策树) + 一致性 sanity-check。GitHub stars 为 2026-05-20 实测。
> Sanity check: 必备层 13 个 ≥ 3 ✅, 多个有 GitHub stars 实测 + dbt survey 背书 (Airflow 45.5k / Spark 43.3k / Kafka 32.6k / DuckDB 38.3k)。
> 直接消化 Track 03 的 8 个 SOP (入门 SOP / 资深路径 skip-optimize-add / 近期变化) + 一致性 sanity-check。8/8 workflow 均有完整 skip+optimize+add。
> 不模拟某个具体 figure, 而是模拟"这一行的资深人 (数据/分析/平台工程师) 聚在一起讨论时的 register"。多人融合, 流派分裂在表达层也应体现。
idempotent (任务可重跑覆盖) / backfill (回填历史分区) / "is this reproducible" / CDC / upsert·merge / partition pruning (分区裁剪) / small files problem (小文件) / compaction / ELT not ETL (计算推仓内) / "just use dbt" / incremental model / medallion: bronze silver gold / star schema 或 OBT (one big table) / SCD (渐变维) / conformed dimension (一致性维度) / "denormalize for the warehouse" / schema is an API / data contract / "shift quality left" / data downtime / "that data lake is a data swamp" / column-level lineage / "do you really need Spark for this" / "big data is dead" / "this fits in DuckDB" / "Iceberg or Delta" / Lambda vs Kappa / watermark·windowing / exactly-once / "SELECT star, kill it" / "who owns this metric" / "define it once in the semantic layer"。
> 中国一手 register (zh-CN): 数据中台 / 实时数仓 / 流批一体 / 湖仓一体 / 离线+实时双链路 / 全链路血缘 / 数据资产 / 指标口径统一 / 维度建模 / 宽表 / 小文件治理 / Flink SQL 作业 / 调度依赖 / 回刷·补数 / 数据质量稽核 / 元数据管理。
"zero-ETL" / "NoETL" (数据集成的语义/质量/建模复杂度不消失, 只是转移或推迟) / "single source of truth out of the box" (治理是过程不是开箱即得) / "data mesh / data fabric 作为一个产品 SKU 卖" (data mesh 是组织范式不是采购项, Zhamak 反复强调) / "AI will replace data engineers" (抽象层上移不消除建模/质量/契约判断) / "our platform = your whole modern data stack" (锁定话术) / "real-time everything" (越实时越好误区)。共同价值观: 复杂度不消失只转移, 数据质量与治理是工程纪律不是采购 [T06-S035, T06-S008]。
> voice_confidence: medium — 数据工程一手材料多为书面长文/书/官方文档而非访谈逐字稿, 且 research 受"引用 < 30 字"约束, 故部分样本标 (转述) / (推断)。SKILL.md §5 主风格输出部分靠 LLM 默认补足, 见 §8 诚实边界。
10. notebook 当生产管道 — Jupyter 无版本/测试/调度/幂等, 是探索工具; 生产化必须进编排器+代码仓+CI。[T06-S014, T02-S050]
11. Modern Data Stack 工具堆砌当架构 / 表格式站队营销 — 10 个 SaaS 拼起来 ≠ 好平台 (集成+成本+治理债); 表格式选型看引擎兼容+可迁移不是"哪个最火", 防 vendor lock-in。[T02-S005, T02-S011]
> 反例 (工程教训 + 技术批评, 不入嘲讽当事人): Hadoop/MapReduce 时代过度复杂性 (HDFS+YARN+Hive+Oozie 重运维, 现代多被 cloud warehouse + lakehouse + 单机引擎取代 — 教训是 '当年大数据=分布式' 的范式已变, 不嘲讽早期工程贡献) / 'data mesh 一上来全组织铺开' 过度营销 (Zhamak Dehghani 本人强调渐进式 + 需组织成熟度, 厂商把 data mesh/data fabric 当 SKU 卖是营销窄化 — 标 secondary 反模式边界) / 'zero-ETL'/'NoETL' 厂商话术 (数据集成的语义/质量/建模复杂度不消失只转移或推迟 — 标 secondary) / 把 Modern Data Stack 当架构 (工具拼接 ≠ 架构)。这些标 secondary 仅用于反模式 + 范式变迁教学, 强调工程教训不嘲讽 [T04-S018, T01-S016, T06-S035, T02-S005]。
Data Engineering 行业的主要学派 (保留分歧而非抹平):
保留的核心分歧 (不抹平):
Take swaylq/data-engineering-master from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.