Build and run ONNX Runtime WebGPU provider tests on Linux WITHOUT a real GPU, using a software Vulkan adapter (Mesa lavapipe). Use when you need to exercise WebGPU EP kernels off-Mac — the Linux webgpu CI leg is build-only, so software Vulkan is how you actually run WebGPU correctness tests locally. SCOPE - lavapipe only validates host-side enforce/shape bugs and MatMul-free kernels; any graph containing MatMul (including the expanded-Attention node tests) crashes lavapipe and runs ONLY on macOS-arm64 Metal, which is the source of truth for those. Covers install (dnf on Azure Linux), the --use_webgpu build flag, the onnxruntime_provider_test target, VK_ICD_FILENAMES, and the lavapipe MatMul crash gotcha.
npx skills add https://github.com/microsoft/onnxruntime --skill webgpu-local-testing
Reusable knowledge for exercising the WebGPU execution provider on a Linux box
with no physical GPU.
> Scope: Linux ORT WebGPU only. macOS uses the Metal backend and a real GPU;
> this skill is the off-Mac story for running WebGPU EP kernels in CI-less dev loops.
On Linux, ORT's WebGPU EP runs on Dawn with the Vulkan backend. Vulkan does
not require a hardware GPU — Mesa lavapipe is a software (CPU) Vulkan adapter that
Dawn enumerates like any other device. For EP correctness tests this is sufficient
because:
ORT_ENFORCEs) firebefore any shader is dispatched. E.g. the WebGPU broadcast ORT_ENFORCE runs
host-side, so the failure is observable on a software adapter without ever touching
the GPU.
lavapipe, so their numeric output can be validated against the CPU reference.
You are trading speed for not needing hardware. It is not a substitute for a real
GPU on perf-sensitive or driver-specific paths — but for kernel correctness it is the
practical local loop.
> Scope — what lavapipe can and cannot validate. Software Vulkan covers (a)
> host-side failures (shape/broadcast ORT_ENFORCEs that fire *before* any shader
> dispatch) and (b) MatMul-free kernels that dispatch. It does NOT cover any
> graph that contains a MatMul — the MatMul family crashes lavapipe's LLVM JIT (see
> §5). This explicitly includes the motivating expanded-Attention node tests
> (test_attention_4d_softcap_neginf_mask_expanded): they decompose to
> softmax(Q·Kᵀ + bias)·V, which contains MatMuls, so they **cannot run on
> lavapipe. For those, macOS-arm64 Metal is the source of truth**. Concretely, the
> #28969 WebGPU broadcast-underflow fix was validated on lavapipe via a standalone
> Add-broadcast OpTester proxy (a host-side enforce/shape path) — NOT via the
> expanded-Attention node test. Never run lavapipe green and conclude an
> Attention/MatMul fix is validated off-Mac.
dnf, NOT apt)dnf install -y mesa-vulkan-drivers vulkan-loader
# optional, for sanity-checking the adapter:
dnf install -y vulkan-tools && vulkaninfo | head
mesa-vulkan-drivers provides lavapipe; vulkan-loader provides the ICD loader.
The lavapipe ICD manifest lands at /usr/share/vulkan/icd.d/lvp_icd.<arch>.json —
lvp_icd.x86_64.json on x86_64, lvp_icd.aarch64.json on arm64. The examples below
use the x86_64 name; substitute your arch, or glob it:
VK_ICD_FILENAMES=$(echo /usr/share/vulkan/icd.d/lvp_icd.*.json).
--use_webgpu./build.sh --config Release --parallel --use_webgpu
See the ort-build skill for general build phases and flags.
WebGPU operator/kernel tests are provider op tests — they build into the
onnxruntime_provider_test target (NOT onnxruntime_test_all; see the ort-test
skill for the executable taxonomy). Point the Vulkan loader at the lavapipe ICD and
select a subset with --gtest_filter:
cd build/Linux/Release
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/lvp_icd.x86_64.json \
./onnxruntime_provider_test --gtest_filter="*WebGPU*"
VK_ICD_FILENAMES forces Vulkan to load only lavapipe, so the run is
deterministic regardless of what else is installed.
MathOpTest.MatMulFloatType (and other MatMul-family tests) crash lavapipe with:
LLVM ERROR: Instruction Combining did not reach a fixpoint after 1 iterations
This is a pre-existing limitation of software Vulkan (Mesa lavapipe's LLVM JIT),
not an ORT bug. Exclude the MatMul family from broad lavapipe runs:
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/lvp_icd.x86_64.json \
./onnxruntime_provider_test --gtest_filter="*WebGPU*:-*MatMul*"
The Linux webgpu CI leg (py-linux-webgpu-stage.yml) only builds — it does not run
WebGPU kernels. A green Linux webgpu leg therefore does not mean any WebGPU test
actually executed. The macOS-arm64 webgpu leg is the only CI leg that runs WebGPU
backend node tests. So a local lavapipe run is the practical way to **actually exercise
WebGPU kernels off-Mac** before you push.
But mind the §1 scope: lavapipe covers host-side enforce/shape paths and MatMul-free
kernels only. Any MatMul-containing graph — including the expanded-Attention node
tests (test_attention_4d_softcap_neginf_mask_expanded) — crashes lavapipe and runs
only on the macOS-arm64 Metal leg, which is the source of truth for those. A green
lavapipe run never validates a MatMul/Attention fix off-Mac.
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Use when implementing any feature or bugfix, before writing implementation code
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Take microsoft/webgpu-local-testing from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.