google/gke-batch-hpc
>- Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).
npx skills add https://github.com/google/skills --skill gke-batch-hpc
This reference covers running batch processing and high-performance computing
(HPC) workloads on GKE.
> MCP Tools: apply_k8s_manifest, get_k8s_resource,
> describe_k8s_resource, get_k8s_logs, delete_k8s_resource,
> list_k8s_events
apiVersion: batch/v1
kind: Job
metadata:
name: batch-job
spec:
parallelism: 10
completions: 100
backoffLimit: 3
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "1"
memory: "2Gi"
restartPolicy: Never
The golden path enables JobSet monitoring (JOBSET in monitoringConfig).
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: training-job
spec:
replicatedJobs:
- name: workers
replicas: 4
template:
spec:
parallelism: 1
completions: 1
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"
Kueue manages job scheduling and resource allocation for batch workloads:
# Install Kueue
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
# Define a ClusterQueue
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: batch-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["cpu", "memory"]
flavors:
- name: default
resources:
- name: "cpu"
nominalQuota: 100
- name: "memory"
nominalQuota: "200Gi"
---
# Allow a namespace to use the queue
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: batch-local
namespace: batch-jobs
spec:
clusterQueue: batch-queue
For tightly-coupled HPC workloads that need low-latency inter-node
communication:
# Standard clusters: create node pool with compact placement
gcloud container node-pools create hpc-pool \
--cluster <CLUSTER_NAME> --region <REGION> \
--machine-type c3-standard-44 \
--placement-type COMPACT \
--num-nodes 8 \
--enable-autoscaling --min-nodes 0 --max-nodes 16 \
--quiet
Use the MPI Operator for MPI-based HPC applications:
# Install MPI Operator
kubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
apiVersion: kubeflow.org/v2beta1
kind: MPIJob
metadata:
name: hpc-simulation
spec:
slotsPerWorker: 4
mpiReplicaSpecs:
Launcher:
replicas: 1
template:
spec:
containers:
- name: launcher
image: <MPI_IMAGE>
command: ["mpirun", "-np", "32", "./simulation"]
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
Worker:
replicas: 8
template:
spec:
containers:
- name: worker
image: <MPI_IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"
limits:
cpu: "8"
memory: "16Gi"
Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint).
Use a ComputeClass with Spot-first priority and activeMigration to return to
Spot when available. See the gke-compute-classes skill for the
Spot-with-fallback pattern.
For batch clusters, allow node pools to scale to zero when no jobs are running:
scheduled
--min-nodes 0 on batch node poolsmemory, and optionally GPU/TPU) for all batch/HPC manifests. This is
critical for Kueue admission, autoscaling, and preventing resource
starvation in the cluster.
VMs/TPUs, advise using GKE maintenance exclusions to block automatic
cluster upgrades/reboots during the active training window to minimize
unnecessary preemption.
distributed MPI applications via the MPIJob custom resource.
sharing; use JobSet for multi-component tightly coupled workloads.
backoffLimit on Jobs, and implementapplication-level checkpointing (e.g., using Orbax or PyTorch checkpointing)
to survive Spot VM preemption.
Take google/gke-batch-hpc from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.