mcpbeat

Deploy Slurm Cluster

nvidia/deploy-slurm-cluster

Deploy a Slurm GPU cluster with DeepOps and prove it works. Use when asked to deploy, install, or rebuild Slurm on one or more GPU servers with this repository.

576 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
1464
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/NVIDIA/deepops --skill deploy-slurm-cluster

The instruction itself

4 sections, as written by the author

Deploy a Slurm GPU cluster

Preconditions

  • Ubuntu 22.04/24.04 or RHEL/Rocky 8/9 hosts you may fully manage (driver

installs may reboot them; no active users or workloads).

  • SSH access from the provisioning machine to every host as a sudo-capable

user.

  • Run everything from the repository root.

Procedure

  • Prepare the environment and verify it:
   git submodule update --init --recursive
   ./scripts/setup.sh
   cp -r config.example config
  • Edit config/inventory: put the controller under [slurm-master] and

compute nodes under [slurm-node] (a single machine can be both). Set

the connection user in [all:vars] if not root.

  • Preflight — must pass before deploying:
   python3 scripts/validation/deepops_doctor.py --remote --json

Fix anything in failures (each check's detail says how) and rerun.

  • Deploy:
   ansible-playbook -l slurm-cluster playbooks/slurm-cluster.yml

This installs NVIDIA drivers, builds and configures Slurm, and sets up

munge, NFS, and node health checks. Expect roughly 30–60 minutes on a

first run.

  • Validate on a cluster node — the success signal is this, not the play

recap:

   python3 scripts/validation/validate_slurm.py --json

Require "ok": true with gpu_job_ok: true and

nodes_unavailable: 0.

Failure branches

  • Playbook fails on a transient error (mirror timeout, apt lock,

network blip): rerun the same playbook; it is idempotent. A converged

rerun ends with changed=0.

  • nvidia-smi works in the validator's srun job but "fails" over SSH:

that is the login GPU-hide behavior, not an error (see AGENTS.md

gotchas).

  • gpu_job_ok: false with driver errors: follow

skills/diagnose-driver-install/.

  • Node shows down or drained in node_states: check

scontrol show node <name> for the reason; after fixing, resume with

scontrol update nodename=<name> state=resume.

  • Wrong or stale host facts after reprovisioning a node: rerun with

--flush-cache.

How to use it

Copy the folder

Take nvidia/deploy-slurm-cluster from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.