nvidia/deploy-slurm-cluster
Deploy a Slurm GPU cluster with DeepOps and prove it works. Use when asked to deploy, install, or rebuild Slurm on one or more GPU servers with this repository.
npx skills add https://github.com/NVIDIA/deepops --skill deploy-slurm-cluster
installs may reboot them; no active users or workloads).
user.
git submodule update --init --recursive
./scripts/setup.sh
cp -r config.example config
config/inventory: put the controller under [slurm-master] andcompute nodes under [slurm-node] (a single machine can be both). Set
the connection user in [all:vars] if not root.
python3 scripts/validation/deepops_doctor.py --remote --json
Fix anything in failures (each check's detail says how) and rerun.
ansible-playbook -l slurm-cluster playbooks/slurm-cluster.yml
This installs NVIDIA drivers, builds and configures Slurm, and sets up
munge, NFS, and node health checks. Expect roughly 30–60 minutes on a
first run.
recap:
python3 scripts/validation/validate_slurm.py --json
Require "ok": true with gpu_job_ok: true and
nodes_unavailable: 0.
network blip): rerun the same playbook; it is idempotent. A converged
rerun ends with changed=0.
nvidia-smi works in the validator's srun job but "fails" over SSH:that is the login GPU-hide behavior, not an error (see AGENTS.md
gotchas).
gpu_job_ok: false with driver errors: followskills/diagnose-driver-install/.
down or drained in node_states: checkscontrol show node <name> for the reason; after fixing, resume with
scontrol update nodename=<name> state=resume.
--flush-cache.
Take nvidia/deploy-slurm-cluster from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.