Managing cloud infrastructure using declarative and imperative IaC tools. Use when provisioning cloud resources (Terraform/OpenTofu for multi-cloud, Pulumi for developer-centric workflows, AWS CDK for AWS-native infrastructure), designing reusable modules, implementing state management patterns, or establishing infrastructure deployment workflows.
npx skills add https://github.com/ancoleman/ai-design-components --skill writing-infrastructure-code
Provision and manage cloud infrastructure using code-based automation tools. This skill covers tool selection, state management, module design, and operational patterns across Terraform/OpenTofu, Pulumi, and AWS CDK.
Use this skill when:
Common requests:
Key Principles:
Benefits:
Choose IaC tools based on team composition and cloud strategy:
Terraform/OpenTofu - Declarative, HCL-based
Pulumi - Imperative, programming language-based
AWS CDK - AWS-native, programming language-based
Decision Tree:
Multi-cloud required?
├─ YES → Team composition?
│ ├─ Ops/SRE focused → Terraform/OpenTofu
│ └─ Developer focused → Pulumi
└─ NO → AWS only?
├─ YES → Language preference?
│ ├─ HCL/declarative → Terraform
│ ├─ TypeScript/Python → AWS CDK
│ └─ YAML/simple → CloudFormation
└─ NO → GCP/Azure only?
└─ Terraform or Pulumi
Remote state with locking enables team collaboration:
Backend Selection:
| Cloud Provider | Recommended Backend | Locking Mechanism |
|----------------|---------------------|-------------------|
| AWS | S3 + DynamoDB | DynamoDB table |
| GCP | Google Cloud Storage | Native |
| Azure | Azure Blob Storage | Lease-based |
| Multi-cloud | Terraform Cloud/Enterprise | Built-in |
| Pulumi | Pulumi Service | Built-in |
State Isolation Strategies:
prod/, staging/, dev/)Critical State Management Rules:
sensitive = trueComposable Module Structure:
modules/
├── vpc/ # Network foundation
├── security-group/ # Reusable security group patterns
├── rds/ # Database with backups, encryption
├── ecs-cluster/ # Container orchestration base
├── ecs-service/ # Individual microservice
└── alb/ # Application load balancer
Module Versioning:
version = "5.1.0")Module Design Principles:
When to Create a Module:
When to Keep Monolithic:
# Initialize providers and backend
terraform init
# Plan changes (preview)
terraform plan
# Apply changes
terraform apply
# Destroy infrastructure
terraform destroy
# Format HCL files
terraform fmt
# Validate syntax
terraform validate
# Show state
terraform state list
terraform state show <resource>
# Import existing resources
terraform import <resource.name> <id>
# Workspace management
terraform workspace list
terraform workspace new staging
terraform workspace select prod
# Initialize new project
pulumi new aws-typescript
# Preview changes
pulumi preview
# Apply changes
pulumi up
# Destroy infrastructure
pulumi destroy
# Show stack outputs
pulumi stack output
# Manage stacks
pulumi stack ls
pulumi stack select prod
# Import existing resources
pulumi import <type> <name> <id>
# Export/import state
pulumi stack export > state.json
pulumi stack import < state.json
# Initialize new app
cdk init app --language typescript
# Synthesize CloudFormation
cdk synth
# Preview changes
cdk diff
# Deploy stack
cdk deploy
# Destroy stack
cdk destroy
# Bootstrap account/region
cdk bootstrap
# List stacks
cdk list
Infrastructure Provisioning:
Module Development:
Operational Readiness:
For comprehensive patterns and implementation details:
Tool-Specific Patterns:
references/terraform-patterns.md - Terraform/OpenTofu best practices, HCL patternsreferences/pulumi-patterns.md - Pulumi across TypeScript/Python/GoArchitecture and Design:
references/state-management.md - Remote state, locking, isolation strategiesreferences/module-design.md - Composable modules, versioning, registriesOperations:
references/drift-detection.md - Detecting and remediating infrastructure driftPractical implementations demonstrating IaC patterns:
Terraform Examples:
examples/terraform/vpc-module/ - Multi-AZ VPC with public/private subnetsexamples/terraform/ecs-service/ - ECS service with ALB, autoscalingexamples/terraform/rds-cluster/ - Aurora cluster with backups, encryptionexamples/terraform/state-backend/ - S3 + DynamoDB backend setupPulumi Examples:
examples/pulumi/typescript/vpc/ - TypeScript VPC componentexamples/pulumi/python/ecs-service/ - Python ECS serviceexamples/pulumi/go/rds-cluster/ - Go RDS clusterexamples/pulumi/testing/ - Unit tests for Pulumi programsAWS CDK Examples:
examples/cdk/typescript/vpc-stack/ - VPC using L2 constructsexamples/cdk/typescript/ecs-fargate/ - Fargate service with ALBexamples/cdk/typescript/pipeline-stack/ - Self-mutating CDK pipelineexamples/cdk/testing/ - CDK assertions and snapshot testsAutomated validation and operational tools:
scripts/validate-terraform.sh - Terraform fmt, validate, tflintscripts/cost-estimate.sh - Infracost wrapper for cost analysisscripts/drift-check.sh - Scheduled drift detectionscripts/security-scan.sh - Checkov/tfsec security scanningscripts/state-backup.sh - State file backup automationscripts/module-release.sh - Module versioning and publishingDeployment Pipeline:
building-ci-pipelines - Automate terraform plan/apply in CI/CDgitops-workflows - GitOps-based infrastructure deploymentPlatform Engineering:
kubernetes-operations - Provision EKS, GKE, AKS clustersplatform-engineering - Internal developer platform infrastructureSecurity:
secret-management - Provision Vault, External Secrets Operatorsecurity-hardening - Implement infrastructure security controlscompliance-frameworks - Policy-as-code for complianceOperations:
observability - Provision monitoring infrastructure (Prometheus, Grafana)disaster-recovery - Infrastructure rebuild procedurescost-optimization - Implement cost controls via IaCData Platform:
data-architecture - Provision data lakes, warehousesstreaming-data - Provision Kafka, Kinesis infrastructureDevelopment Workflow:
terraform plan / pulumi preview locallyState Management:
Module Development:
examples/ directorySecurity:
sensitive = trueCost Management:
Operational Excellence:
State File Issues:
Module Design:
Operations:
Security:
State Lock Issues:
terraform force-unlock <lock-id> # Use only if certain no other process running
Import Existing Resources:
terraform import aws_vpc.main vpc-12345678
pulumi import aws:ec2/vpc:Vpc main vpc-12345678
Drift Detection:
terraform plan -detailed-exitcode # Exit 2 = drift detected
pulumi preview --diff
For detailed drift remediation, see references/drift-detection.md.
State Recovery:
# Terraform: Restore from S3 versioning
aws s3 cp s3://bucket/backup/terraform.tfstate terraform.tfstate
# Pulumi: Restore from checkpoint
pulumi stack export --version <timestamp> | pulumi stack import
For cloud-specific implementations:
aws-patterns - AWS-specific resource patternsgcp-patterns - GCP-specific resource patternsazure-patterns - Azure-specific resource patternsFor infrastructure operations:
kubernetes-operations - Manage Kubernetes clusters provisioned via IaCgitops-workflows - GitOps-based infrastructure deploymentplatform-engineering - Internal developer platformsFor security and compliance:
security-hardening - Infrastructure security controlssecret-management - Secret injection and rotationcompliance-frameworks - Policy-as-code for complianceFor deployment automation:
building-ci-pipelines - CI/CD for infrastructure codedeploying-applications - Application deployment to provisioned infrastructureFor cost and observability:
cost-optimization - FinOps practices for infrastructureobservability - Monitoring infrastructure healthAssess Kubernetes workloads and cluster configuration for AKS Automatic compatibility. Identifies incompatibilities, generates fixes, and guides migration from AKS Standard to AKS Automatic. WHEN: migrate to AKS Automatic, check AKS Automatic readiness, validate manifests for Automatic, assess cluster for Automatic compatibility, fix deployment for Automatic compatibility, identify AKS Automatic migration blockers, is my cluster ready for AKS Automatic.
Discovers available Azure OpenAI model capacity across regions and projects. Analyzes quota limits, compares availability, and recommends optimal deployment locations based on capacity requirements. USE FOR: find capacity, check quota, where can I deploy, capacity discovery, best region for capacity, multi-project capacity search, quota analysis, model availability, region comparison, check TPM availability. DO NOT USE FOR: actual deployment (hand off to preset or customize after discovery), quota increase requests (direct user to Azure Portal), listing existing deployments.
Interactive guided deployment flow for Azure OpenAI models with full customization control. Step-by-step selection of model version, SKU (GlobalStandard/Standard/ProvisionedManaged), capacity, RAI policy (content filter), and advanced options (dynamic quota, priority processing, spillover). USE FOR: custom deployment, customize model deployment, choose version, select SKU, set capacity, configure content filter, RAI policy, deployment options, detailed deployment, advanced deployment, PTU deployment, provisioned throughput. DO NOT USE FOR: quick deployment to optimal region (use preset).
Unified Azure OpenAI model deployment skill with intelligent intent-based routing. Handles quick preset deployments, fully customized deployments (version/SKU/capacity/RAI policy), and capacity discovery across regions and projects. USE FOR: deploy model, deploy gpt, create deployment, model deployment, deploy openai model, set up model, provision model, find capacity, check model availability, where can I deploy, best region for model, capacity analysis. DO NOT USE FOR: listing existing deployments (use foundry_models_deployments_list MCP tool), deleting deployments, agent creation (use agent/create), project creation (use project/create).
Intelligently deploys Azure OpenAI models to optimal regions by analyzing capacity across all available regions. Automatically checks current region first and shows alternatives if needed. USE FOR: quick deployment, optimal region, best region, automatic region selection, fast setup, multi-region capacity check, high availability deployment, deploy to best location. DO NOT USE FOR: custom SKU selection (use customize), specific version selection (use customize), custom capacity configuration (use customize), PTU deployments (use customize).
This skill should be used when working with LaminDB, an open-source data framework for biology that makes data queryable, traceable, reproducible, and FAIR. Use when managing biological datasets (scRNA-seq, spatial, flow cytometry, etc.), tracking computational workflows, curating and validating data with biological ontologies, building data lakehouses, or ensuring data lineage and reproducibility in biological research. Covers data management, annotation, ontologies (genes, cell types, diseases, tissues), schema validation, integrations with workflow managers (Nextflow, Snakemake) and MLOps platforms (W&B, MLflow), and deployment strategies.
Latch platform for bioinformatics workflows. Build pipelines with Latch SDK, @workflow/@task decorators, deploy serverless workflows, LatchFile/LatchDir, Nextflow/Snakemake integration.
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
Take ancoleman/writing-infrastructure-code from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.