maddhruv/absolute-deflake
> Triggers on "absolute deflake", "fix flaky tests", "CI is flaky", "this test fails randomly/intermittently".
npx skills add https://github.com/maddhruv/absolute --skill absolute-deflake
> Start your first response with the 🧪 emoji.
Find tests that pass and fail nondeterministically, diagnose the root cause of each,
and fix it — not by retrying or skipping, but by removing the source of nondeterminism.
Output is evidence (failure rate per test) → cause → fix, verified by repeated runs.
Runs the shared engine in references/health-engine.md — read it for the
DETECT → SCAN → TRIAGE → FIX → VERIFY → REPORT loop and the safety contract. This file
covers only what's specific to flaky tests.
retry/skip-marked tests that mask real flakiness.Not for tests that fail *deterministically* — that's a real bug or a real regression
(/absolute work for a fix, or just fix it). deflake targets *nondeterministic* failures.
Establish flakiness empirically — a test isn't flaky because someone said so. Use
preferences.health.deflakeRuns from config as the default N for repeat-runs (else 20):
| Ecosystem | Repeat-run / detect |
|---|---|
| Jest/Vitest | run suite N× (--run loop), randomize order (--shuffle / testSequencer) |
| pytest | pytest-randomly + pytest --count=N (pytest-repeat); -p no:randomly to A/B |
| Go | go test -count=N -shuffle=on ./..., -race |
Also mine signals: existing retry/flaky/skip annotations, CI history if reachable, and
run the suite both in isolation and in full/parallel — order- and concurrency-
dependent failures only show one way. Record a failure rate per suspect test.
| Cause | Tell | Fix |
|---|---|---|
| Test-order / shared state | passes alone, fails in suite (or vice versa) | isolate state; reset/teardown between tests |
| Time / clock | fails near midnight, DST, or under load | fake timers / inject clock; no real sleep |
| Async race / missing await | fails under parallelism or slow CI | await the actual condition; no fixed timeouts |
| Randomness | fails ~X% with no pattern | seed the RNG; fix the seed in tests |
| Network / external I/O | fails offline or on slow links | mock/stub the boundary |
| Unordered collections | fails on map/set iteration order | sort before asserting |
| Resource leak / port reuse | fails on repeat or parallel runs | unique resources; clean up |
| Wave | Class | Default |
|---|---|---|
| 1 | clear, isolated cause (seed, await, fake clock, sort) | fix now |
| 2 | shared-state / ordering — needs fixture refactor | fix this pass, per test |
| 3 | flakiness pointing at a real product race, not just the test | gated — surface; may be a genuine bug to fix in code |
A flaky test sometimes means the *code* has a race, not the test. Don't "stabilize" the test
into hiding a real concurrency bug — flag wave-3 cases for a real fix.
with -race) — green once is not deflaked; green across N randomized runs is.
retry/skip/flaky annotation that was masking it once the cause is fixed.sleep, or skipping the test —that hides flakiness, doesn't remove it.
sleep to dodge a race. Slows the suite and still flakes under load. Await the condition./absolute upgrade — a flaky suite makes upgrade verification unreliable; deflake first./absolute debt — flaky-test annotations are test debt; this clears them at the root./absolute work — when the flake is a genuine product-code race needing real design.Take maddhruv/absolute-deflake from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.