microsoft/verification-discipline
> Use when verifying that completed work actually works. Auto-surface during /verify mode, post-implementation review, or before claiming a task is done. Teaches the discipline of testing outcomes vs implementation, the unit/integration/smoke gradient, and what "done" actually means.
npx skills add https://github.com/microsoft/amplifier-bundle-skills --skill verification-discipline
Unit tests verify that code-as-written behaves as-written. Smoke and
integration tests verify that the system achieves the intended outcome.
Those are different questions. You need both.
"All unit tests pass" is necessary. It is rarely sufficient. A finding from
the field: four consecutive integration-blocking bugs, all of which passed
unit tests, all of which would have been caught by a five-minute smoke test
on a fresh environment. The bugs were not exotic — they were the cost of
declaring "done" too early.
doesn't anticipate.** The engineer writes code, then writes tests that
exercise the code as written. The tests ask "does this code do what I
wrote it to do?" They don't ask "what scenarios does the system need to
handle?"
value you told it to return. That tells you nothing about whether the
real dependency would have behaved that way.
Component B passes. Their interaction at the seam fails. The seam was
never tested.
golden-file test verifies that the default (un-activated) configuration
renders correctly. The activated configuration — the one production
actually uses — was never exercised.
Treat verification as a ladder. Skip a rung and you discover its bugs in
production.
| Tier | What it verifies | Example |
|---|---|---|
| 1. Unit | Code does what I wrote it to do | pytest tests/unit/ |
| 2. Integration | Component pairs interact correctly | pytest tests/integration/, real DB |
| 3. Smoke / E2E | System achieves the user-visible outcome | Fresh DTU launch, run real pipeline, observe artifacts |
| 4. Production-equivalent | Real environment, real load, real data | Staging deployment, canary, replay traces |
Each tier catches bugs the tier below it cannot. Each tier costs more time
than the tier below it. The economic choice is not "skip the expensive
tiers." The economic choice is "spend five minutes on tier 3 to avoid five
hours of rollback."
Before claiming a task is done, satisfy this checklist:
pass, or — if no integration tests exist for this code path — a
manual integration check is documented).
passed, ideally on a fresh environment).
AGENTS.md and .github/PULL_REQUEST_TEMPLATE.mdare satisfied.
events.jsonl analysis. Not "tests pass." Not "looks right."
If any box is unchecked, the work is not done. Say so, explicitly.
Different from classic TDD. TDD writes unit tests first. Tests-from-outcomes
writes the outcome assertion first.
1. Before writing implementation, write down the user-observable outcome.
"After running this pipeline, events.jsonl contains a `branch_completed`
event for each branch and no `contract_violation` events."
2. Write a test asserting that outcome. The test runs the real pipeline,
inspects the real events.jsonl, checks the real conditions.
3. Implement code until the test passes.
Both patterns are valuable. Unit-level TDD verifies internal correctness.
Outcome-level testing verifies that the system behaves as the user expects.
Use both.
verified the code you wrote. They did not verify the system you shipped.
question is "does the database client retry correctly?" and you mock the
database client's retry method, you have tested nothing.
is cheaper than a five-hour rollback.
≠ "it runs." "It runs" ≠ "it does the right thing."
screenshot, an artifact diff, an events.jsonl excerpt. "I checked" is
not evidence.
If CI has no integration tier, CI green tells you only that the unit
tier passed.
Four integration-blocking bugs in four consecutive shippings. All would have
been caught by a smoke test on a fresh environment. None were caught by the
unit tests that did exist, because the unit tests asked the wrong question.
The fix:
evidence linked next to each box.
Cultural change is hard. Changing the form is easy. Change the form first.
skills/per-repo-conventions/ — how to discover the specific gates agiven repo requires (AGENTS.md, PR template, CONTRIBUTING.md).
foundation:docs/PER_REPO_CONVENTIONS.md — canonical principle forper-repo discovery.
skills/integration-testing-discipline/ — concrete tactics for runningthe integration tier (observe first, fix in batches, expect long
durations).
Take microsoft/verification-discipline from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.