mcpbeat Sign in

Cuda Cutlass Fmha Incremental Rebuild Agent Skill

> Use when rebuilding ONNX Runtime CUDA after editing CUTLASS fused-MHA headers (onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/*.h such as kernel_forward.h or fmha_launch_template.h), or when a header edit "passed" an incremental build but test behavior did not change. Explains the nvcc depfile gotcha that produces stale Memory-Efficient-Attention (MEA) kernels and binaries, and how to force a correct recompile. Also covers disk-space frugality on shared GPU dev boxes.

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
21266
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/microsoft/onnxruntime --skill cuda-cutlass-fmha-incremental-rebuild

The instruction itself

6 sections, as written by the author

Incremental rebuilds silently use STALE CUTLASS fused-MHA kernels

> The general false-green principles (stale binary, wrong-artifact mtime) are summarised

> in the ort-test skill's "False-green taxonomy". This skill is the CUDA/CUTLASS-specific

> detail.

The gotcha (verification-integrity bug)

nvcc-generated depfiles do not track the CUTLASS fused-MHA headers under

onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/ (e.g. kernel_forward.h,

fmha_launch_template.h). These headers are #included by the fmha_sm*.cu

translation units, but the build system does not record that dependency.

Consequence: after you edit one of those headers, an incremental build.sh:

  • does not recompile fmha_sm*.cu,
  • reports [100%] Built target ... and exits 0,
  • leaves the recompiled artifacts — the fmha_sm*.cu.o objects and the

libonnxruntime_providers_cuda.so they link into — unchanged (same mtime as

the pre-edit build).

(Do not use the gtest test-exe mtime as the stale symptom: in the shared-provider

build the exe dlopens the .so and is not relinked, so its mtime stays old even

after a *correct* rebuild — see "How to confirm" below. The reliable diagnostic signal

is the fmha_sm*.cu.o / .so mtime.)

So your "successful" rebuild is running the old kernel. Tests that should now

pass (or fail) reflect the previous code, not your edit. This silently invalidates

any FAIL→PASS / PASS→FAIL verification.

The fix — force recompile the .cu units

Before rebuilding after editing any cutlass_fmha/*.h header:

touch onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/*.cu

Then run the normal build command. This forces the fmha_sm*.cu translation units

(and downstream binaries) to recompile against your header change.

How to confirm the rebuild was real (don't trust "[100%] Built")

Confirm that the artifact which actually links the recompiled fmha_sm*.cu.o

is newer than your header edit.

⚠️ **Do NOT just check the test EXE mtime — it can falsely flag a good build as

stale.** In the shared-provider build configuration (the default here), the CUDA

execution provider is a shared module: the recompiled fmha_sm*.cu.o link into

libonnxruntime_providers_cuda.so, and the onnxruntime_provider_test executable

dlopens that .so — it is not relinked. So after a *correct* rebuild the

test exe mtime stays old while the .so advances. Checking the exe alone

would wrongly conclude the build was stale.

Check the right artifact for your link mode:

  • Shared-provider build (default): the .so that links the recompiled .o

build/<dir>/<cfg>/libonnxruntime_providers_cuda.so

  • Statically-linked provider: the test exe itself (onnxruntime_provider_test)

Safest check — stat both the recompiled object and the .so, and confirm BOTH

are newer than the header edit:

stat -c '%y %n' onnxruntime/contrib_ops/cuda/bert/cutlass_fmha/kernel_forward.h
# in your build dir, e.g. build/Debug_quickbuild/Debug/:
stat -c '%y %n' libonnxruntime_providers_cuda.so
# and the actual recompiled object (path varies by build dir):
find . -name 'fmha_sm80.cu.o' -exec stat -c '%y %n' {} +

If the .so (and the fmha_sm*.cu.o) timestamps are older than (or equal to) the

header edit, the build was stale — touch the .cu files and rebuild. The most

reliable signal of all is behavioral: a test that was failing now passes (a stale

binary cannot flip its result).

This is the CUDA/CUTLASS instance of false-green mode 1 (zero-match / wrong binary) —

see the ort-test skill's "False-green taxonomy" for the general principle. In short:

attention/MEA/Flash boundary gtests (e.g. FlashStructuralEmptyRows*,

Attention_Causal_NonPadKVSeqLen_MEA_*) live in onnxruntime_provider_test, which CI

runs; onnxruntime_test_all does not contain them and gives a false green. Verify the

MEA/Flash boundary fix against onnxruntime_provider_test.

Full ORT CUDA builds are large (test binaries ~1 GB each; a build dir can reach

tens of GB). On a shared box, /home filling to 100% makes builds fail in

non-obvious places — e.g. git submodule sync reporting No space left on device

or a config.lock error, not an obvious "disk full" at the compile step.

Before a big rebuild, check free space and clean only clearly-stale, regenerable

build directories (old dated experiment dirs). Never delete another agent's active

build dir or anything ambiguous:

df -h /home
du -sh build/* | sort -h

Other skills for the same job

different authors, same section of the catalogue
Doc Coauthoring
by anthropics
vendor ×10

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

4k tokens
Changelog Generator
by frostant
×9

Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.

774 tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Writing Plans
by ZhanlinCui
×4

Use when you have a spec or requirements for a multi-step task, before touching code

816 tokens
Writing Skills
by ZhanlinCui
×4

Use when creating new skills, editing existing skills, or verifying skills work before deployment

26k tokens scripts
Crafting Effective Readmes
by softaworks
×3

Use when writing or improving README files. Not all READMEs are the same — provides templates and guidance matched to your audience and project type.

15k tokens
Humanizer
by softaworks
×3

| Remove signs of AI-generated writing from text. Use when editing or reviewing text to make it sound more natural and human-written. Based on Wikipedia's inflated symbolism, promotional language, superficial -ing analyses, vague attributions, em dash overuse, rule of three, AI vocabulary words, negative parallelisms, and excessive conjunctive phrases.

6k tokens
Opentrons Integration
by christophacham
×3

Official Opentrons Protocol API for OT-2 and Flex robots. Use when writing protocols specifically for Opentrons hardware with full access to Protocol API v2 features. Best for production Opentrons protocols, official API compatibility. For multi-vendor automation or broader equipment control use pylabrobot.

9k tokens scripts

How to use it

Copy the folder

Take microsoft/cuda-cutlass-fmha-incremental-rebuild from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.