> Visualize a specific transformer decoder layer from an AutoDeploy FX graph text dump as a hierarchical DOT/PNG diagram. Optionally annotate nodes with actual GPU kernel names and durations from an nsys trace. Use when the user wants to visualize, inspect, "show layer", "graph of layer", "layer visualization", "dump graph layer". Assumes graph dumps already exist in a directory (produced by AD_DUMP_GRAPHS_DIR).
npx skills add https://github.com/NVIDIA/TensorRT-LLM --skill ad-layer-visualizer
Visualize a single transformer decoder layer from an AutoDeploy SSA graph dump.
Optionally overlay actual GPU kernel names and durations from an nsys trace.
> Prerequisite knowledge: This skill assumes familiarity with the ad-graph-dump skill, which covers how to enable dumps via AD_DUMP_GRAPHS_DIR, file naming conventions, SSA format basics, and GraphModule section structure. Refer to ad-graph-dump for that context.
.txt files.txt file. If not given, pick the file with the highest numeric prefix (final transform stage)..nsys-rep or .sqlite trace file. When provided, GPU kernel names and durations are extracted and annotated onto each node in the visualization.Before starting, ask the user two questions (skip any already answered in their request):
.nsys-rep or .sqlite) — if yes, kernel names and durations from the trace will be annotated directly on each node in the diagram.Proceed once both are answered (or the user says no trace).
If the user didn't specify a file, pick the file with the highest numeric prefix (the final transform stage). See ad-graph-dump for the file naming convention and how lexicographic sort matches pipeline order.
Read the selected .txt file. If it contains multiple GraphModule sections (delimited by ======== headers), pick the one labeled monolithic.model or the first/largest one with real operation nodes.
This is the core step — you do this yourself by understanding the graph structure. Do NOT delegate to a script.
The dump is an SSA-form graph where each line is one of:
%name : shape : dtype (model inputs like input_ids, kv_cache, etc.)%name = namespace.op_name(%input1, %input2, ..., const_args...) : shape : dtypeoutput(%name1, %name2, ...)How to identify which nodes belong to layer N:
A transformer decoder layer typically contains these blocks in order:
model_layers_N_input_layernorm_weightmodel_layers_N_self_attn_* weights (q_proj, k_proj, v_proj, o_proj, kv_a_proj, kv_b_proj, q_a_proj, q_b_proj, etc.)model_layers_N_post_attention_layernorm_weightmodel_layers_N_mlp_* weights (gate_weight, shared_experts, etc.) and fused_*_N_* fused weight referencesExtraction rules:
layers_N_ or layers.N. or a fused weight pattern like fused_*_N_* (where these are non-% references, i.e. weight parameters, not activation outputs from other ops).layers_M_ where M ≠ N).sub_1, floordiv_1, mul_1, eq_1 that sit between two layers' MoE routing logic can be tricky. Trace their inputs backward — if they ultimately derive from layer M's weights/operations (not layer N's), they belong to layer M, not layer N. The suffix number on these generic ops does NOT indicate which layer they belong to; you must trace the dataflow._ad_rotary_cos_sin_N, batch metadata)After extraction, output a JSON file at <dump_dir>/<dump_stem>_layer<N>.json with this structure:
{
"layer": 5,
"source_file": "085_compile_compile_model.txt",
"nodes": [
{
"id": "noaux_tc_op_default_2",
"op": "trtllm.noaux_tc_op.default",
"shape": "(8x8, 8x8)",
"dtype": "(torch.bfloat16, torch.int32)",
"group": "moe",
"sub_group": "moe_router",
"inputs": ["dsv3_router_gemm_op_default_2"],
"weight_inputs": [
{"name": "model_layers_5_mlp_gate_e_score_correction_bias", "shape": "256", "dtype": "torch.bfloat16"}
]
}
],
"edges": [
{"from": "dsv3_router_gemm_op_default_2", "to": "noaux_tc_op_default_2"}
],
"external_inputs": [
{
"id": "trtllm_fused_allreduce_residual_rmsnorm_default_4",
"label": "Layer 4 residual output",
"shape": "(2x4x7168, 2x4x7168)"
}
]
}
Node fields:
id: the SSA name (without % prefix)op: the full operation target stringshape, dtype: output shape and dtypegroup: one of norm, attention/mla, moe, mlp, mamba, gdn, othersub_group (optional): finer classification like q_branch, kv_branch, rope, moe_router, moe_experts, shared_expertsinputs: list of node IDs that this node consumes (only nodes within the layer or external inputs)weight_inputs: list of weight parameters consumed (name, shape, dtype)Edge fields:
from, to: node IDs (both must be in nodes or external_inputs)Group assignment heuristic:
norm: ops consuming input_layernorm or post_attention_layernorm weightsmla/attention: ops consuming self_attn_* weights or named *mla*, *rope*, *sdpa*moe: ops consuming MoE weights (*moe*, *experts*), router ops (noaux_tc_op, topk), and the arithmetic ops that process router outputs (sub, floordiv, eq, mul between router and MoE fused op)mlp: ops consuming mlp_* weights that aren't MoE (e.g., shared_experts, gate_proj, up_proj, down_proj)other: everything else (getitem, view, reshape, etc.) — assign to the same group as neighborsIf the user provided a trace file, extract per-layer kernel sequences:
python <skill_dir>/scripts/extract_trace_kernels.py <trace_file> --layer <N> --output <dump_dir>/<dump_stem>_layer<N>_kernels.json
This script uses graphNodeId to extract a single CUDA graph replay, groups kernels by stream, and identifies which streams belong to which layer. It outputs a JSON with per-layer kernel sequences including short names, full names, durations, and stream IDs.
When trace kernel data is available, you must map GPU kernels to individual FX graph nodes. This is the key step — the render script will display kernel names and durations directly on each node in the diagram.
How to map kernels to nodes:
Read the trace kernel JSON and the layer JSON side by side. For each FX graph node, identify which GPU kernel(s) it corresponds to based on the op type and the kernel execution order. Common mappings:
| FX graph op | Trace kernel(s) |
|---|---|
| flashinfer_mla_with_cache | fmhaSm100... |
| finegrained_fp8_linear | fp8_blockscale + pack_fp32_to_ue8m0 + deep_gemm_fp8 |
| torch_linear_simple | nvjet_gemm + splitKreduce |
| flashinfer_rms_norm | rms_norm_reduce_fusion |
| flashinfer_fused_add_rms_norm | fused_kernel_a5fe... |
| mlir_fused (input_layernorm) | fused_kernel_1984... |
| mlir_fused (post_attn) | fused_kernel_2e06... |
| trtllm_moe_fused | bmm_E4m3... + moe_activation_deepseek + bmm_Bfloat16... + moe_finalize |
| noaux_tc_op (top-k routing) | deepseek_v3_topk |
| fused_swiglu_mlp / fused_finegrained_fp8_swiglu_mlp | deep_gemm_fp8 (×2) + act_and_mul |
| trtllm_dist_all_reduce | nccl_allreduce_symk or symm_mem_allreduce |
| symm_mem_all_gather | symm_mem_allgather or nccl_allgather |
| triton_rope_on_interleaved_qk_inputs | mla_rope_assign_qkv |
Add a "trace_kernels" list to each node that has corresponding GPU kernels:
{
"id": "flashinfer_mla_with_cache_default_5",
"op": "auto_deploy.flashinfer_mla_with_cache.default",
"group": "mla",
"trace_kernels": [
{"kernel": "fmhaSm100...", "duration_us": 50.1}
],
...
}
Also add top-level trace_summary for the whole layer:
{
"layer": 5,
"trace_summary": {
"total_duration_us": 650.9,
"kernel_count": 66,
"num_streams": 3
},
"nodes": [...]
}
The render script will display each node's kernels as ⚡ kernel_name (Xµs) lines below the node label. Nodes without trace_kernels are left unchanged.
Run the bundled visualization script:
python <skill_dir>/scripts/render_layer.py <dump_dir>/<dump_stem>_layer<N>.json --output <dump_dir>/<dump_stem>_layer<N>
This reads the JSON and produces .dot and .png files. If nodes have trace_kernels fields, the script renders kernel names and durations directly on each node's label (prefixed with ⚡).
Present the output paths to the user.
graphviz (dot command) to be installed for PNG renderingmonolithic.model module or auto-select the largest.nsys-rep to .sqlite if needed (requires nsys on PATH)Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas
Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.
Picks random winners from lists, spreadsheets, or Google Sheets for giveaways, raffles, and contests. Ensures fair, unbiased selection with transparency.
Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.
MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.
Access AlphaFold 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
Take nvidia/ad-layer-visualizer from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.