Run ONNX Runtime tests. Use this skill when asked to run tests, debug test failures, or find and execute specific test cases in ONNX Runtime.
npx skills add https://github.com/microsoft/onnxruntime --skill ort-test
ONNX Runtime uses Google Test for C++ and unittest (preferred) / pytest for Python.
| Executable | What it tests |
|---|---|
| onnxruntime_test_all | Core framework, graph, optimizer, session tests |
| onnxruntime_provider_test | Operator/kernel tests (Conv, MatMul, etc.) across execution providers |
attention_op_test.cc files — don't confuse themThere are two same-named files testing different operators. Both build into
onnxruntime_provider_test:
| Path | Operator | gtest suite |
|---|---|---|
| test/providers/cpu/llm/attention_op_test.cc | ONNX-domain Attention (opset 23/24) | AttentionTest.* |
| test/contrib_ops/attention_op_test.cc | contrib MultiHeadAttention / GroupQueryAttention | ContribOpAttentionTest.* |
The MEA negative-offset regression tests (Attention_Causal_NonPadKVSeqLen_MEA_*,
e.g. ..._MEA_NegOffset_ForceFlashDisabled_FP16_CUDA) live in the providers/cpu/llm file —
the ONNX-domain op.
Use --gtest_filter to select specific tests:
./onnxruntime_provider_test --gtest_filter="*Conv3D*"
Always run from the build output directory — tests may fail to find dependencies otherwise.
# Linux
cd build/Linux/Release
./onnxruntime_provider_test --gtest_filter="*TestName*"
# macOS
cd build/MacOS/Release
./onnxruntime_provider_test --gtest_filter="*TestName*"
# Windows
cd build\Windows\Release
.\onnxruntime_provider_test.exe --gtest_filter="*TestName*"
You can also run all tests via the build script (assumes a prior successful build):
./build.sh --config Release --test
.\build.bat --config Release --test # Windows
The default path follows the pattern build/<Platform>/<Config>/ where Platform is Linux, MacOS, or Windows. With Visual Studio multi-config generators on Windows, the config may appear twice (e.g., build/Windows/Release/Release/). The path can also be customized via --build_dir.
If you can't find a test binary, search for it:
# Windows
Get-ChildItem -Path build -Recurse -Filter "onnxruntime_provider_test.exe" | Select-Object -ExpandProperty FullName
# Linux/macOS
find build -name "onnxruntime_provider_test" -type f
Use pytest as the test runner:
pytest onnxruntime/test/python/test_specific.py # entire file
pytest onnxruntime/test/python/test_specific.py::TestClass::test_method # specific test
pytest -k "test_keyword" onnxruntime/test/python/ # by keyword
Python test naming convention: test_<method>_<expected_behavior>_[when_<condition>]
AGENTS.md."False-green taxonomy" section below for the four ways a test can pass without testing
your change.
> test_output.txt 2>&1) — output can be large.--gtest_filter to run a targeted subset when the full suite takes too long.onnxruntime_provider_test and can run against a software Vulkan adapter (Mesa lavapipe). See the webgpu-local-testing skill.A green result is not always a real pass. Watch for all five modes:
--gtest_filter that matches no tests still exits 0 (green).Confirm the [==========] N tests ran line is non-zero — a zero-match run prints
0 tests from 0 test suites. Many operator/kernel gtests run only in
onnxruntime_provider_test (CI runs this), NOT onnxruntime_test_all; the wrong
binary matches nothing and looks green.
change (e.g. a header not tracked by the compiler's depfile), the "passing" run executes
the OLD code. A test that was failing cannot truly flip to passing without a real
rebuild — treat an unexpected FAIL→PASS with suspicion and confirm the linked artifact's
mtime advanced. CUDA/CUTLASS instance (nvcc depfiles don't track cutlass_fmha/*.h): see
the cuda-cutlass-fmha-incremental-rebuild skill.
libonnxruntime_providers_cuda.so), the test executable is NOT relinked when the provider
recompiles — its mtime stays old while the .so advances. Verify the artifact that
actually links your change, not the test exe. Detail: cuda-cutlass-fmha-incremental-rebuild
skill.
*different, correct* code path without ever exercising the one you meant to test (e.g. a
test meant for MEA silently handled by the unfused fallback). Assert/verify **which path
ran**, not just the output value — see "Verify which path/kernel actually executed" below.
launches on a large-dynamic-smem arch (e.g. sm90/H100, ~227KB) can fail to launch on a
smaller opt-in cap (sm86/89 ~99KB, sm80 ~163KB) with CUDA failure 1: invalid argument —
and a path with no fallback (e.g. ORT's MEA) turns that into a hard error, not a silent
degrade. So a green run on your local GPU can mask a launch failure on CI's arch. Verify
arch-portability, or pick a config whose shared-memory footprint fits every target arch
(e.g. a small head_size). Concrete instance: CUTLASS MEA head_size=512 FP16 exceeds
sm86's smem opt-in cap and dies at launch — live bug #28388 (the
cuda-attention-kernel-patterns skill §1 has the dispatch detail).
Value equality alone does not prove the intended code path ran — a correct fallback can
produce the right answer (false-green mode 4 above). When a test targets a specific
kernel/path, confirm it actually dispatched there instead of trusting the output:
exact strings (core/providers/cuda/llm/attention.cc):
ONNX Attention: using Flash Attention (:1400)ONNX Attention: using Memory Efficient Attention (:1451)Attention: using unified unfused path (:1482) — note: no ONNX prefix and itreads "unified unfused path", not "Unfused".
the test SKIPs (not silently passes) when the target path is unavailable — e.g.
SKIP_IF_MEA_NOT_COMPILED.
Operator-specific routing/forcing details: cuda-attention-kernel-patterns skill §1/§7.
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Use when implementing any feature or bugfix, before writing implementation code
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Take microsoft/ort-test from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.