mcpbeat

Dgx Station Diagnose

nvidia/dgx-station-diagnose

Run and interpret the complete read-only dgx-assist diagnostic suite for NVIDIA DGX Station GB300, correlate findings with pinned NVIDIA playbooks, export a redacted support bundle, and apply one separately approved allowlisted fix. Use when the user reports a Station, CUDA, GPU health, coherency, vsloshd, Docker, CDI, MIG, cache, port, or owned inference-service failure.

1k tokens
context cost
the whole folder, loaded on every use
5
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
1211
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/NVIDIA/dgx-spark-playbooks --skill dgx-station-diagnose

What comes with it

3 439 bytes besides the instruction
agents/openai.yaml
references/bringup.md
references/findings.md
scripts/dgx-assist

The instruction itself

3 sections, as written by the author

DGX Station diagnostics

Diagnose first. Do not mutate as part of diagnosis.

Workflow

  • Run scripts/dgx-assist diagnose run --json.
  • Report the detected compatibility profile, then findings in severity order with their stable IDs and evidence. Preserve unknown states. On Software 1.0, do not reinterpret intentionally skipped Software 2.0 service checks as faults.
  • Search the pinned playbooks for each high or critical finding; cite the relevant URL, heading, lines, and commit.
  • If no passage overlaps, say so and avoid inventing a platform fix.
  • Offer diagnose bundle --report-id "<id>" when escalation is appropriate.
  • Offer at most one automatic fix at a time, and only when fix_id is present.
  • Preview with diagnose fix --report-id "<id>" --finding "<id>" --dry-run.
  • Explain exact actions, impact, privilege, and reboot state. Obtain explicit approval.
  • Repeat with --yes only after approval and report the action receipt.

Safety requirements

  • Keep diagnose run read-only.
  • Never install packages, rewrite Docker configuration, change power caps or driver parameters, kill workloads, or modify MIG through a diagnostic fix.
  • Never stop a service without current dgx-assist ownership evidence.
  • Never use Fabric Manager as a routine Station check or remediation.
  • Re-run diagnostics when finding evidence is stale.
  • Never reveal secrets or unredacted home paths in a bundle.
  • Do not execute an unregistered remediation.
  • Treat --yes only as approval already obtained.

Read references/findings.md before proposing a fix or support bundle. Read references/bringup.md when the problem concerns physical deployment, BMC or firmware verification, driver bring-up, power braking, or support escalation.

How to use it

Copy the folder

Take nvidia/dgx-station-diagnose from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.