nvidia/fr-analysis
> Analyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available. Detects the wavefront process group where collectives diverge and returns the root-cause suspect ranks.
npx skills add https://github.com/NVIDIA/nvidia-resiliency-ext --skill fr-analysis
Analyze PyTorch NCCL flight-recorder (FR) dumps to identify the collective operation hang
and isolate the ranks responsible, using CollectiveAnalyzer.
Script: scripts/fr_attribution.py → attribution/trace_analyzer/fr_attribution.py
--fr-path.Collective records (op type, ranks, process group, timing, state).returns the missing ranks at that boundary as the root-cause suspects.
--llm-analyze) over the structured findings for ahuman-readable summary.
python scripts/fr_attribution.py \
--fr-path /path/to/fr_dumps/ \
[-p "_dump_*"] \
[--verbose] \
[--health-check] \
[--llm-analyze] \
[--model MODEL] \
[--debug]
| Flag | Default | Description |
|------|---------|-------------|
| --fr-path | required | Path to a directory (or single file) containing FR dump files |
| --pattern, -p | _dump_* | Glob pattern for dump files within --fr-path |
| --verbose, -v | off | Print detailed per-rank collective tables |
| --health-check, -c | off | Include node health check results in output |
| --llm-analyze, -l | off | Pass structured findings to the LLM for a narrative summary |
| --model, -m | nvidia/nemotron-3-super-120b-a12b | LLM model (only used with --llm-analyze) |
| --debug | off | Convert binary trace files to JSON for inspection |
from nvidia_resiliency_ext.attribution.trace_analyzer.fr_attribution import CollectiveAnalyzer
analyzer = CollectiveAnalyzer({
"fr_path": "/path/to/fr_dumps/",
"pattern": "_dump_*",
"verbose": False,
"health_check": False,
"llm_analyze": False,
"model": "nvidia/nemotron-3-super-120b-a12b",
})
results = analyzer.run_sync({
"fr_path": "/path/to/fr_dumps/",
})
# results: tuple[FRAnalysisResult | str, AttributionState]
Returns (result, AttributionState) where result is the FR analysis table and describes:
--verbose)--health-check)--llm-analyze)AttributionState.STOP indicates the hang is unrecoverable; CONTINUE indicates the job
may be restartable after isolating the identified ranks.
| Format | Notes |
|--------|-------|
| _dump_* files | PyTorch FR dump prefix pattern used by the feedback loop |
| Binary pickle / JSON payloads | Detected automatically; use --debug to convert binary traces to JSON |
FR dumps are typically written to the directory specified by TORCH_NCCL_DEBUG_INFO_TEMP_FILE
or triggered automatically on NCCL timeout.
TORCH_NCCL_TRACE_BUFFER_SIZE > 0)LLM_API_KEY required only when using --llm-analyzelangchain-openai required only when using --llm-analyzeFR_DEBUG=1 env var enables verbose debug logging in the scriptTake nvidia/fr-analysis from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.