nvidia/nvrx-attr
> Orchestration layer over nvidia_resiliency_ext attribution modules. Provides log-analysis, fr-analysis, and a Megatron-LM-oriented fault-injection feedback loop for benchmarking attribution quality on SLURM workloads.
npx skills add https://github.com/NVIDIA/nvidia-resiliency-ext --skill nvrx-attr
High-level orchestration layer over the nvidia_resiliency_ext.attribution modules.
Each subdirectory is a self-contained skill with its own SKILL.md and helper scripts.
| Directory | Purpose | Entry point |
|-----------|---------|------------|
| log-analysis/ | Analyze SLURM job logs for failure root-cause and restart decisions | NVRxLogAnalyzer (nvrx_logsage.py) |
| fr-analysis/ | Analyze NCCL flight-recorder dumps for collective-hang root-cause | CollectiveAnalyzer (fr_attribution.py) |
| fault-injection-loop/ | Run a batched SLURM fault-injection feedback loop and score attribution accuracy | prepare_node_alloc.sh / watch_and_analyze.sh |
src/nvidia_resiliency_ext/
├── attribution/
│ ├── log_analyzer/nvrx_logsage.py ← log-analysis implementation
│ ├── trace_analyzer/fr_attribution.py ← fr-analysis implementation
│ ├── analyzer/engine.py ← combined orchestration entry point
│ └── combined_log_fr/ ← optional log + FR fusion
└── skills/
└── nvrx-attr/ ← this skill bundle
├── log-analysis/
├── fr-analysis/
└── fault-injection-loop/
The Analyzer (analyzer/engine.py) is the recommended entry point when you need
request coalescing, result caching, or the combined LOG_AND_TRACE pipeline.
Use the individual skills when you want to run one analysis type directly without the
full coalescing stack.
LLM_API_KEY environment variable, LLM_API_KEY_FILE, or ~/.llm_api_keylangchain-openai installedlogsage package installed (required by log_analysis)pip install nvidia-resiliency-ext or pip install -e . from repo rootBefore using fault-injection-loop/, create the local config file from the tracked
template and fill in your site-specific values:
cp scripts/user.env.example scripts/user.env
The feedback-loop scripts require src/nvidia_resiliency_ext/skills/nvrx-attr/scripts/user.env
to exist at runtime. Keep user.env local and untracked.
Take nvidia/nvrx-attr from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.