Run Windows E2E tests (MSI install tests or Fleet Automation/installer tests) locally against AWS-provisioned VMs
npx skills add https://github.com/DataDog/datadog-agent --skill run-windows-e2e
Run Windows E2E tests from test/new-e2e/tests/windows/ or test/new-e2e/tests/installer/windows/.
Detailed reference material lives in references/ next to this file — read the
relevant one when a step calls for it rather than duplicating it here:
references/setup.md — prerequisites, ~/.test_infra_config.yaml, dev modereferences/running.md — setup-env, local builds, go test flagsreferences/vm-access.md — connecting to a dev-mode VM (RDP/SSH)references/troubleshooting.md — test outputs, Pulumi locks, AWS profile$ARGUMENTSDetermine:
install-test, service-test, agent-package, install-script). If not provided, ask the user.TestXxx function. Most suites expect exactly one test per run — ask the user which one if not specified.--build pipeline (default), --build local, or --build release; pass --pipeline-id <id> through if given.setup-env run with --prefix STABLE_AGENT in Step 3 (see references/running.md "Upgrade tests").--branch <name> to setup-env. The default is the current git branch, which may have no pipelines if it's a local feature branch.Map suite names to Go package paths:
| Suite | Package path |
|-------|-------------|
| install-test | ./test/new-e2e/tests/windows/install-test |
| service-test | ./test/new-e2e/tests/windows/service-test |
| fips-test | ./test/new-e2e/tests/windows/fips-test |
| domain-test | ./test/new-e2e/tests/windows/domain-test |
| installer / Fleet Automation (agent-package, install-script, install-exe, ddot, apm-inject, …) | ./test/new-e2e/tests/installer/windows |
The installer / Fleet Automation tests are all one flat package (the
suites/<package>/ subdirectories were flattened in #47161) — pick the area by
test function with -run (e.g. TestAgentUpgrades, TestInstallScript,
TestDDOTExtensionViaMSI, TestAPMInjectInstalls).
If the user gives a partial name or test function, search with Glob/Grep under test/new-e2e/tests/windows/ and test/new-e2e/tests/installer/windows/ to resolve it.
test -f ~/.test_infra_config.yaml && echo "EXISTS" || echo "MISSING"
pulumi version 2>/dev/null || echo "MISSING"
If either is missing, offer to run dda inv e2e.setup and wait for the user to
complete its interactive prompts. If prerequisites exist but devMode is not
set, mention that devMode: true reuses VMs across runs (much faster for
iterative development). Full detail in references/setup.md.
Run setup-env with --fmt json to capture the required env vars (no shell
eval needed — prepend the parsed pairs inline to go test in Step 5).
# From a pipeline (most common)
dda inv new-e2e-tests.setup-env --build pipeline --fmt json [--branch <branch>] [--pipeline-id <id>]
# From a local build (run `dda inv msi.build` first, + `msi.package-oci` for installer/OCI tests)
dda inv new-e2e-tests.setup-env --build local --fmt json
For upgrade tests, run a second time with --prefix STABLE_AGENT and merge the
vars in. GitLab token handling, local-build details, and the STABLE_AGENT
flow are in references/running.md.
If devMode: true and the user is rerunning, the previous VM may still have the
agent installed. Ask whether they've cleaned up (MSI tests: uninstall the agent;
installer tests: datadog-installer.exe purge). See
references/running.md "Clean state between runs".
go test commandgo test -v -timeout 30m -tags test <package-path> -run <TestFunction>$
Two rules to apply (rationale in references/running.md):
anchor the -run regex with $ at both suite and subtest level, and use the
exact package path with no trailing /... (or output won't stream).
Show the full command to the user and confirm before running.
Warn the user that AWS SSO auth may open a browser window when the test starts
(the test pauses until login completes), and that a non-sandbox AWS_PROFILE
will cause auth errors — advise unset AWS_PROFILE.
Run with run_in_background: true since tests provision real AWS VMs.
Provisioning takes a few minutes; once the VM is up, SSH becomes available
within ~60s (Linux) / ~180s (Windows). If SSH is not available within those
windows, troubleshoot before assuming the test is still running normally
(see references/vm-access.md).
When the test completes:
~/e2e-output/latest/ (crash dumps, agent/installer logs, event logs). If devMode is on, the VM is still up — offer to help RDP/SSH in via references/vm-access.md.references/troubleshooting.md.Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Use when implementing any feature or bugfix, before writing implementation code
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Take datadog/run-windows-e2e from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.