nvidia/ad-graph-dump
> Enable and interpret TensorRT-LLM AutoDeploy FX graph text dumps via AD_DUMP_GRAPHS_DIR. Use when you need before/after graphs per transform, to locate subgraphs, or to confirm a rewrite ran. Paths and behavior are grounded in tensorrt_llm/_torch/auto_deploy (GraphWriter, BaseTransform). Complements ad-add-fusion-transformation.
npx skills add https://github.com/NVIDIA/TensorRT-LLM --skill ad-graph-dump
AD_DUMP_GRAPHS_DIR)This file is part of trtllm-agent-toolkit. Commands and paths such as examples/auto_deploy/ and tensorrt_llm/ are relative to a TensorRT-LLM source checkout, not the plugin repository.
getitem, view, reshape) appeared or disappeared between dumps.| Skill | Use it for |
|-------|------------|
| ad-layer-visualizer | Extracting and visualizing a single decoder layer from a dump as a DOT/PNG diagram. |
| ad-add-fusion-transformation | Implementing or reviewing fusion passes once you know what the graphs show. |
| trtllm-codebase-exploration | Searching the TRT-LLM tree for transforms, custom ops, and patterns. |
| trtllm-code-contribution | Tests and contribution hygiene after you change TRT-LLM. |
Set:
export AD_DUMP_GRAPHS_DIR=/path/to/output/dir
Implementation: GraphWriter.DUMP_GRAPHS_ENV == "AD_DUMP_GRAPHS_DIR" in tensorrt_llm/_torch/auto_deploy/utils/graph_writer.py.
If unset, no graph files are written.
After each transform application, BaseTransform calls graph_writer.dump_graph(mod, t_name, self.config.stage.value) from tensorrt_llm/_torch/auto_deploy/transform/interface.py (immediately after _visualize_graph). So the dump reflects the module after that transform has run.
From GraphWriter.dump_graph:
AD_DUMP_GRAPHS_DIR is set.ADLogger.rank is set and is not 0, dumping is skipped (non–rank-0 processes do not write files).On the first dump on rank 0, GraphWriter removes the target directory if it already exists, then recreates it. Do not point AD_DUMP_GRAPHS_DIR at a directory that must be preserved without copying it first.
Files are named:
{NNN}_{<stage.value>}_{<transform_key>}.txt
NNN is a monotonically increasing three-digit counter (001, 002, …) in run order across all dumps in that process.config.stage value (same enum/string used in default.yaml under each transform’s stage: field).transform_name passed into dump_graph).So lexicographic sort by filename matches pipeline order for that run.
Each file is text and starts with headers similar to:
# Transform: <transform_key>
# Stage: <stage.value>
# GraphModules found: <count>
Then, for every torch.fx.GraphModule found under mod.named_modules() (including the root), the writer emits a section title and an SSA-style listing with shape/dtype metadata via dump_ssa_with_meta() in the same module.
Use this to compare operator chains, consumers, and node.meta shape/dtype hints across consecutive files.
From the root of the TensorRT-LLM clone (adjust the script and flags to your workflow):
AD_DUMP_GRAPHS_DIR=/tmp/ad-graphs \
python examples/auto_deploy/build_and_run_ad.py --model <hf-model-id> --use-registry
Pick any AutoDeploy entrypoint you already use; the requirement is only that the code path runs the transform pipeline with AD_DUMP_GRAPHS_DIR set in the environment.
While a transform runs, logging is patched so messages can be prefixed with [stage=<stage.value>, transform=<transform_key>] (see with_transform_logging in transform/interface.py). Transform summaries log [SUMMARY] with matches=<n> or skipped / disabled (_log_transform_summary). Use those lines together with the numbered dump files to tie match counts to graph shape before and after a specific transform.
AD_DUMP_GRAPHS_DIR overwrites prior output.GraphModule children, dump_graph returns without creating a new file for that step (see early return in graph_writer.py).tensorrt_llm/_torch/auto_deploy/utils/graph_writer.py — env var, filenames, SSA dump.tensorrt_llm/_torch/auto_deploy/transform/interface.py — call site after each transform; log prefix decorator.Take nvidia/ad-graph-dump from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.