Profiler-driven hill-climbing to close the inference throughput gap between TorchTitan's unified model (running inside vLLM) and vLLM's native model. Benchmark with generate.py --benchmark, climb optimization rungs (compile / cudagraph / fused kernels), profile torchtitan vs the native target, then patch the single biggest gap at a time and re-measure. Use when the user wants to benchmark or optimize RL inference generation speed, reproduce previous hill climbing study, or invokes /inference_perf_hillclimb.
npx skills add https://github.com/pytorch/torchtitan --skill inference_perf_hillclimb
Close the gap between the TorchTitan unified model (in vLLM) and vLLM's native
model on inference throughput, one profiler-found optimization at a time. Harness:
torchtitan/experiments/rl/generate.py --benchmark. Append-only running log +
full per-rung numbers live in torchtitan/experiments/rl/docs/inference_gap_ablation.md.
Goal: match vLLM native 100%.
This skill makes torchtitan's model run as fast as vLLM's NATIVE model at a FIXED
model, workload, precision, and topology -- i.e. it closes the framework/overhead
gap (target = native at the *same* config; ratio -> 1.0). It is explicitly NOT
about the orthogonal levers that change the setup itself:
Those move ABSOLUTE throughput but apply equally to native -- they don't change
the torchtitan-vs-native ratio this skill chases. Hold them FIXED and IDENTICAL
between torchtitan and native when measuring; otherwise you're no longer measuring
the implementation gap.
TP=8, SP off). Sets the sharding, all-reduce sizes, cudagraph capture sizes, and
what "native" even is -- the dominant gap and right knobs change with it.
kernels, cudagraph behavior, custom AR, and APIs. Record all three -- numbers
across versions are not comparable.
(priority); W2 bs32/in4096/out1024 (needs --max-seq-len 8192). Sets prefill-vs-
decode weighting, AR/capture sizes, memory -- the dominant gap depends on it
(e.g. prefill capture helps short-gen W1 more than long-gen W2).
aot_eager (validated torchtitan default, per-layer torch.compile) |inductor (heavier) | off. The fair target is native(eager), NOT
native(inductor): inductor fuses the TP all-reduce into gemm/RMSNorm
(fuse_allreduce_rms), eager does not -- it's ~+5% and not in the aot_eager
path, so comparing to it invents a fake gap.
"cudagraph modes" under Harness).
Every rung is justified by a PROFILE, never intuition, and KEPT only if a
re-benchmark moves throughput. The flag list is the residue of past profiles; the
method is: profile -> single biggest gap -> patch exactly that -> re-measure ->
re-profile. Two rules: (a) trace time misleads -- a huge all-reduce duration is
usually cross-rank SPIN (sync wait), not reducible compute; confirm with tok/s.
(b) the biggest gap is often NOT a kernel -- prefill cudagraph capture and
arrival-spin were the top levers; profile broadly.
--native --compile aot_eager --cudagraph on (same workload).
--profile). Decompose for the gap:/tmp/ablation_logs/analyze_trace.py (single-rank diff: kernel us,extra kernels present in one but not the other, CPU-op deltas).
/tmp/ablation_logs/analyze_ar_spread.py over the 8 per-ranktraces -- min(dur)=pure comm, spread(dur)=ARRIVAL SPIN. Single-rank kernel
sums mislead for collectives.
decode). Low busy% on the driver rank => host-launch-bound.
kernel_ablation.py, a--model-2d variant, or a cudagraph/config flag); often "make torchtitan do
what native does here".
bitwise-identical with vs without the patch.
(trace promise != throughput).
within ~3% of target, or the remainder is structural (document it).
an EAGER region the host can't feed the GPU -> it starves. Spot: per-rank GPU
busy% < ~90% with idle gaps before compute kernels. Levers: cudagraph capture
takes the CPU out of the per-kernel path (capturing PREFILL too was the single
biggest lever, ~+20%); OR remove DTensor from the forward (--model-2d
pure-local); OR cut per-launch overhead. Hidden under a captured graph -- only
bites in eager prefill. NOTE: the structural lever is removing DTensor, not
"2D" -- compile+cudagraph kill DTensor's CPU cost but its boundaries
(from_local / subclass dispatch / placement views) fragment the captured graph
-> launch jitter. Tensor rank is a red herring (3D==2D); register_sharding
REGRESSES (-26%, more DTensor dispatch).
analyze_trace.py per-kernel us. Found: NCCL all-reduce 23us vs vLLM custom
one-shot AR 6.3us for small decode messages (a real algorithm difference);
faster RoPE/attn variants. Lever: --allreduce-vllm, --rope-kernel helion,
--attn-backend flash. Caveat: GEMM and RMSNorm are already at native parity
(torch/Quack rms_norm even BEATS vLLM's) -- don't chase them.
fused add+RMSNorm, fused QKV / gate-up GEMM, SiluAndMul). Spot: "extra kernels"
in the torchtitan trace + higher CPU-op count. Lever: fuse_qkv=True +
fused_swiglu, --fused-addnorm, --silu-vllm. Caveat: mostly a launch-count /
host win -> matters in eager regions, shrinks under cudagraph.
SPIN). Spot: analyze_ar_spread.py (min(dur)=comm==native, spread=arrival spin).
Find WHAT BOUNDS the straggler: *compute?* per-rank compute time is usually
UNIFORM, so no; *host-launch?* the DRIVER rank (rank 0) runs scheduler/sample/
output on its Python thread and falls behind in the eager prefill -> arrives
last (busy 54% vs 94%, spin ~600us) -- the usual cause; fix = capture prefill
(spin ~600us -> ~3us, the +20% lever) or lighten the per-step host path.
*hardware/topology?* if the SAME rank straggles in NATIVE too, it's a
GPU/NVLink-position effect, not fixable in the model. Rule: the AR kernel is
innocent -- fix uneven ARRIVAL, not the collective.
Operational gotchas: check GPU contention first (nvidia-smi; a co-tenant job
spikes variance to +/-150 vs clean +/-1-7 -- re-run clean); NO Python-side
logging/prints inside a patched forward (forces a torch.compile graph break);
launch long runs detached (setsid nohup ... &) and clean GPU stragglers between.
Build/restore: stock generate.py on main is an EXAMPLE (single prompt). Bring
the benchmark harness over from branch ablation-inference:
git checkout ablation-inference -- torchtitan/experiments/rl/generate.py + the
patch modules (models/{kernel_ablation,qwen3_vllm_2d,helion_rope,vllm_fused_ops}.py),
fix the import (`from torchtitan.experiments.rl.examples.alphabet_sort import
config_registry), and add rl_grpo_qwen3_32b (mirror rl_grpo_qwen3_14b`,
model_registry("32B"), TP=8). --benchmark builds the engine OUTSIDE the timed
region (excludes startup/compile/capture), runs --warmup-runs then times
--num-runs, reports tok/s = batch*gen/wall; feeds synthetic token-ids (skips the
tokenizer); sets enable_prefix_caching=False. Build the cudagraph
CompilationConfig EXPLICITLY (the stock helper only emits cudagraph_mode="full").
Launch: torchrun --nproc_per_node=<TP> generate.py --benchmark .... Env:
conda titan-rl (see torchtitan/experiments/rl/README.md to build it).
Knobs are the residue of past profiles -- NOT a required checklist, and NOT
guaranteed to exist on a fresh checkout (the harness is rebuilt each time; re-add
what you need, add NEW knobs for new gaps):
`--rope-kernel {helion,vllm} --silu-vllm --rmsnorm-vllm --allreduce-vllm
--fused-addnorm --attn-backend {custom,flash}` -- custom_op monkeypatches in
models/kernel_ablation.py / vllm_fused_ops.py (survive compile+cudagraph).
FusedQKV+gate_up: use config fuse_qkv=True + overrides/fused_swiglu.py
@override, not the superseded --merged-gemm.
--model-2d {local,localfused,spmd,local3d,dtensor}. localfused = pure-local
(no DTensor) + fused add+RMSNorm (best); spmd = localfused but the TP AR routes
through spmd_types.redistribute(P->R). spmd_types is a pre-run CHECK (validates
SPMD sharding via typechecking) -- keep it OFF for perf runs.
--compile {off,aot_eager,inductor} `--cudagraph{on,off} --cudagraph-mode {full_decode_only,full,full_and_piecewise}`. Capture
PREFILL with `--cudagraph-mode full --max-num-batched-tokens <P> --max-capture-size
<P>`, where P >= the prefill CHUNK size (= max_num_batched_tokens), NOT input_len.
Decode capture sizes default to powers of 2 up to max_num_seqs (= batch).
--nccl-algo, --native (target), --profile (per-rank chrome trace),--max-seq-len (long prompts).
These change WHO owns torch.compile:
CompilationConfig(mode=NONE) -- vLLM does NOTcompile; torchtitan's per-layer aot_eager is the only compile. full_decode_only
= decode captured, eager prefill; full = also captures prefill (needs
--max-capture-size >= chunk). KEEPS torchtitan compile.
mode=VLLM_COMPILE, backend="eager" AND `config.compile=off` -- vLLM compiles the whole model (to split the graph around collectives),
so per-layer aot_eager is turned off. DROPS torchtitan compile for vLLM's.
Mixing FULL and FAP across rungs mixes two compile strategies (a bug we hit --
re-run the WHOLE ladder on ONE mode). native FULL ~= native FAP (878.9 vs 879.1 W1).
ONE example trajectory, NOT a recipe -- each step was the biggest profiler-found
gap at that point: baseline (eager DTensor, SP off) -> +compile(aot_eager) ->
+cudagraph -> +tree all-reduce -> +Helion RoPE -> +FusedQKV/gate_up -> +SiluAndMul
-> +FA3 -> +fused add+RMSNorm -> +vLLM custom AR -> +vLLM RMSNorm -> pure-local.
Big jumps: cudagraph (0.03 -> 0.48x), vLLM custom AR (biggest KERNEL lever),
dropping DTensor (pure-local). SiluAndMul / vLLM-RMSNorm / vLLM-RoPE were
neutral-to-negative (torchtitan already fast). The biggest lever overall -- PREFILL
cudagraph capture (~+20%) -- isn't a kernel; it was found later by profiling the
residual arrival-spin straggler.
Ratios vs native eager (32B TP=8, full_decode_only; full per-rung numbers and the
FULL+prefill ladder are in the doc):
| path | W1 (bs8/in1024/gen128) | W2 (bs32/in4096/gen1024) |
|---|---|---|
| DTensor ceiling (vLLM-AR + vLLM-RMSNorm) | 0.742 | 0.759 |
| spmd_types, NCCL AR (TT_SPMD_NCCL_AR=1) | 0.698 | 0.796 |
| spmd_types, vLLM AR (set_dist shim) | 0.846 | 0.877 |
| pure-local (localfused) | 0.891 | 0.926 |
| native | 1.000 | 1.000 |
Capturing PREFILL too (--cudagraph-mode full + the max-batched-tokens/capture
knobs) lifts every row ~+0.07-0.18x: DTensor ceiling -> ~0.90x, pure-local ->
~0.975x, spmd+vLLM-AR -> ~0.93x, at BOTH W1 and W2; native barely moves (its
prefill is already lean).
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and econometric analysis.
Fine-tune models on Azure AI Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).
Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
Work with Data Commons, a platform providing programmatic access to public statistical data from global sources. Use this skill when working with demographic data, economic indicators, health statistics, environmental data, or any public datasets available through Data Commons. Applicable for querying population statistics, GDP figures, unemployment rates, disease prevalence, geographic entity resolution, and exploring relationships between statistical entities.
Take pytorch/inference_perf_hillclimb from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.