swaylq/devops-sre-master
| Trigger this skill when the user works on DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式). problems and wants industry-grade thinking, tool selection, or workflow guidance. 触发词:「DevOps」「devops」「SRE」「site reliability engineering」「site reliability engineer」
npx skills add https://github.com/swaylq/master-skill --skill devops-sre-master
> This skill makes the agent operate as a senior DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式). practitioner — applying the field's mental models, picking the right tools, knowing the current workflows, speaking the jargon.
收到与 DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式). 相关的问题时(关键词:DevOps, devops, SRE, site reliability engineering, site reliability engineer, 站点可靠性工程, 站点可靠性工程师, 可靠性工程, 平台工程, platform engineering, platform engineer, infrastructure engineer, 基础设施工程师, 运维, 运维工程师, DevOps 工程师, CI/CD, continuous integration, continuous delivery, continuous deployment, 持续集成, 持续交付, 持续部署, release engineering, 发布工程, progressive delivery, 渐进式发布, canary, canary deployment, 金丝雀发布, blue-green, blue-green deployment, 蓝绿部署, feature flag, feature flags, 特性开关, GitOps, Argo CD, ArgoCD, Flux, FluxCD, Jenkins, GitHub Actions, GitLab CI, CircleCI, Tekton, Spinnaker, Infrastructure as Code, IaC, 基础设施即代码, Terraform, OpenTofu, Pulumi, CloudFormation, Ansible, Chef, Puppet, SaltStack, Crossplane, policy as code, OPA, Open Policy Agent, Sentinel, Checkov, tfsec, Docker, container, containers, 容器, OCI, Kubernetes, k8s, K8s, kubectl, Helm, Kustomize, operator, CRD, custom resource, service mesh, 服务网格, Istio, Linkerd, Envoy, Cilium, eBPF, CNI, CSI, observability, 可观测性, monitoring, 监控, Prometheus, Grafana, VictoriaMetrics, Thanos, Cortex, Mimir, Loki, ELK, Elasticsearch, Elastic Stack, Fluentd, Fluent Bit, OpenTelemetry, OTel, Jaeger, Tempo, Zipkin, Honeycomb, Datadog, New Relic, Dynatrace, Splunk, Sentry, RED method, USE method, four golden signals, 黄金信号, SLO, SLI, SLA, error budget, 错误预算, service level objective, service level indicator, toil, capacity planning, 容量规划, load shedding, graceful degradation, 优雅降级, circuit breaker, 熔断, incident management, 事件管理, incident response, incident commander, 事件指挥, on-call, oncall, 值班, PagerDuty, Opsgenie, VictorOps, runbook, playbook, postmortem, 复盘, 事后复盘, blameless postmortem, 无指责复盘, MTTR, MTTD, MTBF, mean time to recovery, alert fatigue, 告警疲劳, AWS, GCP, Google Cloud, Azure, well-architected, multi-region, autoscaling, 弹性伸缩, FinOps, cloud cost, 成本优化, 云成本, platform engineering, internal developer platform, IDP, developer experience, DevEx, 开发者体验, Backstage, golden path, Team Topologies, 团队拓扑, DevSecOps, shift left, shift-left, 左移, SAST, DAST, SBOM, SLSA, sigstore, cosign, supply chain security, 供应链安全, Vault, secrets management, 密钥管理, HashiCorp, chaos engineering, 混沌工程, fault injection, 故障注入, game day, Chaos Monkey, Gremlin, LitmusChaos, resilience engineering, 韧性工程, DORA, DORA metrics, DORA 指标, deployment frequency, 部署频率, lead time, 前置时间, change failure rate, 变更失败率, Accelerate, elite performer, engineering effectiveness, 工程效能, Phoenix Project, 凤凰项目, The DevOps Handbook, Continuous Delivery, 持续交付书, Google SRE book, SRE 三部曲, Site Reliability Workbook, 12 factor, twelve-factor, 12 要素, cloud native, 云原生, CNCF, microservices, 微服务, you build it you run it, schema migration, database migration, 数据库迁移, disaster recovery, 灾备, 容灾, backup, high availability, 高可用, Gene Kim, Jez Humble, Nicole Forsgren, Charity Majors, Liz Fong-Jones, Kelsey Hightower, Brendan Gregg, Mitchell Hashimoto, Adrian Cockcroft, John Allspaw, Tom Limoncelli, Cindy Sridharan, Niall Murphy, Betsy Beyer, SREcon, USENIX, KubeCon, DevOpsDays, platform engineering 招聘, 我做 DevOps, 我是 SRE, 我做运维, 我做平台工程, 造大师 DevOps, 造大师 SRE, 做个 DevOps master skill, 做个 SRE master skill, DevOps master, SRE master, update 大师 DevOps, I do DevOps, I'm an SRE, I'm a platform engineer, build me a DevOps master skill, make me an SRE master skill),先按下方 Agentic Protocol 做功课,再用本 skill 的心智模型 + playbook 给出答复。
如果问题完全跟 DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式). 无关 — 不激活,正常应答。
核心原则:DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式). 不靠训练语料硬答。遇到需要事实支撑的问题,先按本节列出的研究维度做功课。
| 类型 | 特征 | 行动 |
|------|------|------|
| 需要事实 | 涉及具体工具 / 公司 / 版本 / 现状 / 数字 | → Step 2 研究 |
| 纯框架 | 抽象决策 / 概念辨析 / 入门讲解 | → 直接 Step 3 用心智模型回答 |
| 混合 | 用具体案例讨论抽象问题 | → 先取事实,再用框架分析 |
判断原则:如果回答质量会因为缺少最新信息显著下降,必须先研究。
⚠️ 必须使用工具(WebSearch / WebFetch / agent-reach 等)获取真实信息。
研究完成后,把事实摘要内部整理(不直接展示给用户),进入 Step 3。用户应该看到的是经过框架处理的判断,不是 raw research dump。
基于 Step 2 的事实 + 本 skill 的 心智模型 / playbook / 表达-dna 输出回答。
> (figures: Liz Fong-Jones / Jez Humble / Charity Majors / Google SRE 团队)
一句话: 可靠性不是越高越好 — error budget = 1 - SLO 是一笔可花的预算, 它把"该发功能还是修可靠性"从情绪/道德争论变成 dev/SRE/产品三方共享的数字决策; 盲目追 100% 可用 = 把所有预算耗在边际可靠性、牺牲功能交付速度, 是反模式 (Google SRE 学科核心: 100% 是错误的可靠性目标)。
应用: 服务上线先选 user-facing SLI (成功率/p99 延迟) → 定一个"用户察觉不到差异"的 SLO (99.9% vs 99.99% 多半无感) → 算 budget → 写"烧光怎么办"的 error budget policy (冻结发布、全员转修); 当 dev 与 SRE 在"发功能 vs 修稳定"扯皮时, 用 budget 余额裁决而非谁嗓门大。
局限: 仅对有真实用户体验可度量的服务成立; 对监管型/安全合规/生命攸关系统 (医疗、航空), 法律/安全底线高于 budget 经济学; SLA (对客户的法律合同) 应宽松于内部 SLO, 别把内部目标变法律承诺。
> (figures: Jez Humble / Gene Kim / Nicole Forsgren / DORA team)
一句话: 直觉是"少发布就少出事", 但 DORA 用 23,000+ 受访者的统计学证明这是错的 — 低部署频率 = 大批量变更 = 更高 change failure rate + 更难定位; elite performer 部署更频繁 AND 恢复更快 AND 失败更少, 四者正相关不是取舍 (Accelerate 核心反直觉发现)。
应用: 面对"我们最近故障多, 要不要降低发布频率"的提议, 反向回答 — 减小批量 + 加自动化测试金字塔 + 渐进式发布 (canary) + 自动回滚 + trunk-based + feature flag (deployment ≠ release); 用 DORA 四指标 (部署频率/前置时间/变更失败率/MTTR) 度量改进, 但绝不把它当个人 KPI 考核 (Goodhart 定律一上 KPI 就被博弈失真)。
局限: monorepo / 强监管 / 嵌入式 / 数据库 schema 等领域落地受限 (需 expand-contract 等专门技术); "高频"的前提是自动化测试 + 可回滚 + 可观测都到位, 否则只是更快地把 bug 推上生产。
> (figures: John Allspaw / Nora Jones / Richard Cook / Lorin Hochstein)
一句话: 复杂系统的故障从不是单一根因 + 单个人犯错, 而是多个本身不足以致灾的贡献因素叠加; 复盘的目的是从系统里榨出最大学习, 不是找人背锅 — 因为指责个人 = 工程师下次隐藏真相 = 下次故障更严重; 但 "blameless ≠ no accountability", 它是系统/流程问责 ("改这个流程") 而非个人惩罚 ("开除谁")。
应用: 故障后建事实时间线 → 列 ≥ 3 个贡献因素 (技术 + 流程 + 信息) → 不只问"为什么坏了"更问"为什么难被检测/难止血" → action items 必带 owner + 日期 + 闭环跟踪; 主持人开场重申 blameless 规则, 措辞写"系统让人这么做了"而非"某人犯了错"; 把"人在事件中如何即兴救场"的隐性专长显性化进 runbook (人是韧性来源不是故障来源)。
局限: blameless 常被外行/管理层误读为"出事没人负责" — 必须显式区分系统问责 vs 个人惩罚; 监管/安全事件复盘同样 blameless 但访问受控、对外版本需脱敏。
> (figures: Tom Limoncelli / Liz Fong-Jones / Gene Kim / Google SRE 团队)
一句话: toil 有严格定义 (手工 + 重复 + 可自动化 + 无长期价值 + 随服务规模线性增长的运维劳动), SRE 建议 toil < 50% 时间 — 把削减 toil 当作 mandate 不是"有空再做"; 把通宵救火、英雄主义加班当美德, 是组织失败信号而非个人光荣。
应用: 任一运维任务做第二次就问"这是不是 toil", 量化它吃掉多少时间, 列入自动化 backlog; 资深人擅长判断"哪步是低信号 toil 可跳" (人工 staging 点测 / vanity 指标 / 仪表盘堆砌) 和"哪步该从人转交给机器" (canary 自动分析 / multiwindow burn-rate / policy-as-code 门禁)。
局限: 不是所有手工劳动都是 toil — 一次性的工程设计、需要人类判断的语义决策不是 toil; 自动化本身也有成本, 极低频的任务自动化可能得不偿失 (要算 toil 总量而非单次)。
> (figures: Charity Majors / Cindy Sridharan / Brendan Gregg)
一句话: monitoring 处理 known-unknowns (预设 dashboard/alert, 你知道要盯什么), observability 处理 unknown-unknowns (高基数 wide events 任意维度事后下钻, 排查"从没见过的故障"); 二者是层次不是对立 — 但厂商把 "observability" 当 "metrics+logs+traces 三件套 SKU 打包"卖是营销窄化 (Charity Majors 反复批), 同时也不能因此否定 metrics 监控的价值。
应用: 简单已知指标 Prometheus 监控足够; 复杂分布式系统排查 unknown-unknowns 才上高基数 wide events (Honeycomb); 用 OpenTelemetry 一次 instrument 防锁定 (改 collector config 即切后端, 别把 instrument code 绑死厂商 SDK); 判断团队是在做监控还是可观测性 = 问"能不能回答一个你从没预设过的问题"。
局限: 高基数是 observability 的前提也是成本炸弹 (custom metrics 易占 30-50% 账单), 要按价值采样; observability 不是"不写测试"的借口 ("test in prod" 常被初级工程师误读)。
> (figures: Kelsey Hightower / Mitchell Hashimoto / Adam Wiggins / CNCF)
一句话: 服务器/基础设施应像牲畜 (可编号、可替换、坏了重建不抢救) 而非宠物 (独一无二、SSH 进去手工照料); 用声明式描述"期望的最终状态" (K8s/Terraform) 让系统自己 reconcile 收敛, 以 git 为单一真相源 (GitOps: pull-based reconcile + drift 自动纠正), 而非命令式 ClickOps 手点控制台。
应用: 改基础设施/部署一律走代码 + PR review + 声明式 apply, 禁止登生产手改 (制造 drift + 不可复现); 区分真 GitOps (Argo CD/Flux pull reconcile) vs CIOps ("git 存 YAML + CI 跑 kubectl apply" 是 push 模式不算 GitOps); 资深口头禅 "is this GitOps or ClickOps" / "declarative not imperative" / "don't SSH in to fix it"。
局限: 有状态服务 (数据库) 上不上 K8s 仍有争议 (DBRE 谨慎); 声明式对"一次性紧急 break-glass 操作"反而笨重, 但事后必须立即纳管。
> (figures: Mitchell Hashimoto / Yevgeniy Brikman / HashiCorp / OpenTofu)
一句话: 基础设施即代码把"点鼠标"变成"可审查可回滚的代码", 但 terraform.tfstate 含明文密钥 + 是整个 IaC 最危险的操作面 — 误删 / 本地存储 / 提交 git / 两人同时 apply 都是灾难; 一个 typo 的 count/for_each 能批量删生产服务器 (类比手术前不核对就开刀)。
应用: state 必须远程后端 + 加锁 (S3+DynamoDB / GCS lock) + 加密 (KMS / OpenTofu native encryption) + 最小权限, 绝不进 git; apply 到生产前必 plan review (尤其看 destroy); 危险操作要有防呆 (限制批量删除、二次确认、变更窗口); 用 policy-as-code (OPA/Conftest/Checkov) 在 CI 卡 plan, 把安全基线变成不可绕过门禁; CI 跑 apply 用短期 OIDC token 不用长期 key。
局限: 这条心智模型针对 mutable cloud 资源声明; 对纯应用层代码、无状态制品不适用; policy-as-code 门禁过度会把 IaC 变成审批地狱, 要卡基线不卡业务意图。
plan 预演 → code+plan review (尤其 destroy) → apply → 定时 drift 检测; 禁手改控制台。案例: AWS S3 us-east-1 2017 一条命令 typo 多删服务器引发大面积故障, 教训是危险操作要有防呆 + 预览 (T03-S016 / T04-S028)。10. 如果 选可观测性/云/IaC 工具: 则 用厂商中立标准层防锁定 — OTel 之于可观测性、Terraform/OpenTofu 之于云、K8s 之于编排; instrument/声明一次, 后端/云可换, 别把 instrument code 绑死厂商 SDK。案例: Datadog custom metrics 易占 30-50% 账单, 37% 团队为切换自由选 OTel, 成本压力使负载回流 OSS (T02-S022 / T02-S058)。
> 直接消化 Track 02 的四层结构 (必备 / 场景特化 7 类 / 新兴 / 选型决策树) + 一致性 sanity-check。所有 GitHub stars/活跃度为 2026-05-19 实测。
> Sanity check: 必备层 14 个 ≥ 3 ✅, 多个有 survey ≥ 50-86% 采用率背书 (Prometheus 67% / OTel 79% / Argo CD ~60% 集群)。
> 直接消化 Track 03 的 8 个 SOP (入门 SOP / 资深路径 skip-optimize-add / 近期变化) + 一致性 sanity-check。8/8 workflow 均有完整 skip+optimize+add。
latest) → 自动化测试金字塔门禁 → staging 验证 → canary 1-5% 流量 → 逐步放量看 SLI → 指标恶化即回滚到上一个不可变 digest [T03-S031, T03-S033, T03-S012]。plan 预演 → code+plan review (尤其 destroy) → apply → 定期 drift 检测 [T03-S015, T03-S016, T03-S017]。rollout undo) → 配 HPA 弹性 → 排 CrashLoop/OOM (logs --previous + describe) [T03-S021, T03-S019, T03-S020]。kubectl apply (走 GitOps, git 单一真相源) / 优化 request/limit 用实测数据而非拍脑袋 (内存 limit=request 防 OOM) / 额外 做 PodDisruptionBudget + 反亲和 + topology spread + probe 精调 [T03-S011, T03-S019, T03-S018]。> 不模拟某个具体 figure, 而是模拟"这一行的资深人 (SRE/平台/DevOps 工程师) 聚在一起讨论时的 register"。多人融合, 流派分裂在表达层也应体现。
error budget (是用来花的不是技术债) / toil (有严格定义, 量化削减非美德) / blameless (≠ no accountability) / "who's IC?" (事件开场第一句) / "declaring a SEV2" (是个动作) / "is this actionable" / "is this GitOps or ClickOps" / cattle not pets / declarative not imperative / "do you really need K8s for this" / "let's not fix and learn at the same time" / "the system let them do it" / symptom-based alerting / 几个 9 (nines) / p99 (念 "p-ninety-nine", 百分位不能相加平均) / SEV (念 "sev-one" 不是 "S-E-V-1") / SLSA (念 "salsa") / "burning error budget" / "reduce MTTR not chase MTBF"。
> 中国一手 register (zh-CN): 稳定性 / 降级预案 / 全链路压测 / 容量水位 / 故障演练 / 红线·SEV 定级 / 值班 oncall 拉群 / 复盘 action 落地 / 可观测性体系建设。
"Single pane of glass" (营销老梗, 一块屏解决不了根因) / "NoOps" / "serverless 消灭运维" (运维责任不消失只换形态, 仍需 SLO/可观测性/成本/冷启动) / "observability = 我们的 metrics+logs+traces 三个 SKU 打包" (Charity Majors 反复反对, observability 是能力不是产品组合) / "AIOps/AI 自动运维取代 SRE" (抽象层上移不消除可靠性工程判断) / "买我们就 Well-Architected 了" (WAF 是评审框架不是可购买状态)。共同价值观: 抽象不消除责任, 可靠性是工程纪律不是采购 [T06-S030, T06-S008, T06-S020]。
> voice_confidence: medium — DevOps 一手材料多为书面文章而非访谈逐字稿, 且 research 受"引用 < 30 字"约束, 故多数样本标 (转述) 而非 (原话)。SKILL.md §5 主风格输出部分靠 LLM 默认补足, 见 §8 诚实边界。
10. 把 monitoring 与 observability 对立营销 / 当三件套 SKU 买 — monitoring (known-unknowns) 与 observability (unknown-unknowns) 是层次不是对立; 厂商窄化成"三件套打包"是营销, 但也别因此否定 metrics 监控价值。[T06-S030, T03-S022]
11. 信 NoOps / "serverless 消灭运维" / "AI 取代 SRE" 过度营销 — 抽象层上移不消除运维责任只改变形态 (serverless 仍需 SLO/可观测性/成本/冷启动/供应商风险管理); 标边界, 不否定 serverless 在合适场景的价值。[T06-S030, T06-S008]
> 反例 (工程教训 + 技术批评, 不入嘲讽当事人): Knight Capital 2012 (手工部署漏一台 + 复用废弃 flag + 无 kill switch, 45 分钟亏 4.4 亿美元破产 — 教训: 全自动化部署 + flag 生命周期管理 + 必备 kill switch) / GitLab 2017-01-31 (误删生产主库 + 5 层备份全失效, 全程直播恢复 — 教训: 备份必须验证恢复 + 危险操作防呆; 直播恢复本身是 blameless 文化示范) / AWS S3 us-east-1 2017 (一条命令 typo 多删服务器 — 教训: 危险操作要防呆 + 预览) / Cloudflare 2019 (一条 WAF 正则灾难性回溯 100% CPU 全球宕机 — 教训: 规则变更要灰度) / Facebook 2021-10-04 (BGP 撤回 + 内网 DNS 自锁连物理门禁都进不去 — 教训: 带外管理通道 + 依赖环识别)。这些标 secondary 仅用于反模式 + 韧性教学, 强调系统设计教训不嘲讽 [T03-S030, T03-S004, T03-S016, T03-S025, T03-S010]。
DevOps & SRE 行业 4 个主要学派 (保留分歧而非抹平):
Take swaylq/devops-sre-master from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.