mcpbeat Sign in

E2e Testing Agent Skill

Use when writing or stabilizing Playwright tests that drive a real browser through multi-step journeys — durable locators, web-first assertions, storageState auth, trace/retries, and flakes that only bite in CI. NOT in-process component tests (that is testing-web), NOT WCAG auditing (that is accessibility), NOT the pre-merge gate (that is verify).

8k tokens
context cost
the whole folder, loaded on every use
6
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
105
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/ericrisco/rsc-harness --skill e2e-testing

What comes with it

20 518 bytes besides the instruction
evals/README.md
evals/cases.yaml
references/config-and-ci.md
references/flakiness-playbook.md
scripts/verify.sh

The instruction itself

12 sections, as written by the author

e2e-testing — drive a real browser, keep it deterministic

You write Playwright tests that walk a real browser through real user journeys — log in, fill a

form, check out, navigate across pages — and you keep those tests deterministic enough to gate a

merge. The whole game is one tension: e2e tests catch integration bugs nothing else can, and they

are the slowest, flakiest layer you own. Every rule below exists to buy back determinism.

Pin @playwright/test and provision browsers with npx playwright install --with-deps. Current

line is Playwright v1.60.x (v1.60.0 shipped 2026-05-11). The _react / _vue selector engines

and the :light Shadow-DOM suffix were removed in v1.58.0 — at any version you should be pinning

they are long gone, so do not reach for them.

Is this even an e2e test?

E2e is the most expensive layer. Spend it only on journeys that cross pages or services. Route the rest out.

| The goal is… | Layer | Why |

|---|---|---|

| A multi-step journey across pages/auth/services in a real browser | e2e (here) | Only a real browser proves the pieces integrate. |

| One component or pure function, rendered in-process (Vitest/Jest, Testing Library) | ../testing-web/SKILL.md | A browser round-trip to test render logic is slow and flaky for no gain. |

| "Is this page accessible?" — WCAG/ARIA as the deliverable | ../accessibility/SKILL.md | E2e may *call* axe inside a test, but auditing a11y is its own skill. |

| "Is this page fast?" — LCP/CWV budgets | ../performance/SKILL.md | Perf budgets are a different signal than journey correctness. |

| The runner matrix, caching, the pipeline itself | ../github-actions/SKILL.md | E2e contributes a *job*; owning the pipeline is theirs. |

| "Is the change done?" — run the gate, collect evidence, then merge | ../verify/SKILL.md | Running an existing suite as a pre-merge gate is not authoring or stabilizing one. |

Rule: if you can prove it without launching a browser, you should. Push logic down to testing-web.

Locators: the priority ladder

Pick the highest rung that uniquely identifies the element. Higher rungs track what the user

perceives, so they survive markup churn.

  • getByRole('button', { name: 'Buy' }) — role + accessible name. Default choice; doubles as an a11y signal.
  • getByLabel('Email') / getByPlaceholder(...) — form fields.
  • getByText('Order confirmed') — visible copy that uniquely identifies content.
  • getByTestId('cart-total') — when nothing user-facing is stable; requires a deliberate data-testid.
  • CSS as a last resort, scoped and shallow.

Never XPath, never nth-child chains, never the removed _react/_vue/:light engines.

// Bad — couples the test to DOM structure; one wrapper div breaks it.
await page.locator('div.card > button:nth-child(2)').click();

// Good — finds the button the way the user reads it.
await page.getByRole('button', { name: 'Buy' }).click();

Strict mode. A locator that matches two nodes *throws* — that is the framework catching an

ambiguous selector for you. Tighten the locator (getByRole(...).and(...), scope with

page.getByRole('listitem').filter({ hasText: 'Pro' })). Reaching for .first() to silence the

error hides the ambiguity and is the next flake.

Assertions: web-first only

// Bad — reads the DOM once, before the async update lands; races the render.
expect(await page.locator('#status').textContent()).toBe('Submitted');

// Good — re-polls the element until it says 'Submitted' or the timeout fires.
await expect(page.getByTestId('status')).toHaveText('Submitted');

expect(locator) assertions (toBeVisible, toHaveText, toHaveURL, toHaveCount) retry until

the condition holds. A read-once value (await locator.textContent() then compare) captures a single

frame and loses every race against a re-render. If you find expect(await in a test, it is a bug.

Auto-wait and the no-sleep rule

Locator actions (click, fill, check) already auto-wait: they block until the element is

visible, stable, enabled, and receiving events. So waitForTimeout(2000) is never the right wait —

it is either too short (flake) or too long (slow), and it waits for wall-clock time instead of the

thing you actually care about.

| Instead of guessing with a sleep | Wait on the real signal |

|---|---|

| "give the button time to appear" | await expect(locator).toBeVisible() |

| "wait for navigation" | await page.waitForURL('**/checkout') |

| "wait for the XHR/fetch" | const r = page.waitForResponse('**/api/order'); …action…; await r; |

| "wait for the list to fill" | await expect(page.getByRole('row')).toHaveCount(5) |

The ordering trap: subscribe to a response (or register a page.route mock) before the action

that triggers it, or you miss the event.

// Bad — handler registered after goto; the initial request already fired unmocked.
await page.goto('/orders');
await page.route('**/api/orders', route => route.fulfill({ json: [] }));

// Good — mock in place before navigation, so the first request is intercepted.
await page.route('**/api/orders', route => route.fulfill({ json: [] }));
await page.goto('/orders');

Fixtures and page objects

Fixtures give every test a fresh, isolated setup and kill copy-pasted boilerplate. Extend the base

test with your own; the code before use(value) is setup, after it is teardown.

import { test as base } from '@playwright/test';
import { CheckoutPage } from './pages/checkout';

type Fixtures = { checkout: CheckoutPage };

export const test = base.extend<Fixtures>({
  checkout: async ({ page }, use) => {
    const checkout = new CheckoutPage(page); // setup: depends on the built-in `page`
    await use(checkout);                      // hand it to the test
    // teardown after the test goes here, if any
  },
});

Option fixtures (['default', { option: true }]) let a project or test.use() flip behavior

without new fixtures. Keep page objects thin — locators and intent-named actions

(checkout.placeOrder()), no assertions buried inside them. Full page-object recipe lives in

references/config-and-ci.md.

Auth and storageState

Logging in through the UI on every test is slow and a flake surface. Log in once in a setup

project, save the authenticated session to JSON, and load it via storageState in the projects that

depend on it.

  • A setup project runs the login spec and writes playwright/.auth/<role>.json.
  • Real test projects declare dependencies: ['setup'] and use: { storageState: '…/<role>.json' }.
  • One file per role (admin, member, anon) — never share one mutated session across roles.
  • Regenerate every CI run; gitignore the .auth/ dir. Committed session state leaks secrets and

goes stale.

The full multi-role setup-project wiring is in references/config-and-ci.md.

playwright.config.ts (condensed)

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './e2e',
  fullyParallel: true,
  forbidOnly: !!process.env.CI,        // a stray test.only fails CI instead of skipping the suite
  retries: process.env.CI ? 2 : 0,     // retry only in CI; locally a flake should hurt
  workers: process.env.CI ? 1 : undefined,
  reporter: process.env.CI ? [['github'], ['html']] : 'list',
  use: {
    baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
    trace: 'on-first-retry',           // full trace captured the first time a test retries
  },
  projects: [
    { name: 'setup', testMatch: /.*\.setup\.ts/ },
    { name: 'chromium', use: { ...devices['Desktop Chrome'] }, dependencies: ['setup'] },
    { name: 'webkit',   use: { ...devices['Desktop Safari'] }, dependencies: ['setup'] },
  ],
  webServer: {
    command: 'npm run start',
    url: 'http://localhost:3000',
    reuseExistingServer: !process.env.CI,
  },
});

Open a captured trace with npx playwright show-trace. The full annotated config (per-role

storageState, firefox, blob reporter for sharding) is in

references/config-and-ci.md.

CI (GitHub Actions)

The shape: install browsers with OS deps, run, shard when one box can't finish inside the ~5–10 min

budget, upload the trace and HTML report as artifacts.

- run: npx playwright install --with-deps
- run: npx playwright test --shard=${{ matrix.shard }}/4
- uses: actions/upload-artifact@v4
  if: ${{ !cancelled() }}
  with: { name: report-${{ matrix.shard }}, path: playwright-report/, retention-days: 7 }

Scale: bump workers to use a single machine; add --shard=i/N across machines only once a single

box overruns the budget. Sharded runs emit blob reports you merge with npx playwright merge-reports.

Full workflow (matrix, blob report, merge job) is in references/config-and-ci.md.

Flakiness playbook

A 3% flake rate on a 40-minute pipeline burns roughly an engineer-day a week on reruns, so treat

flakes as bugs with named causes. Open the trace first (show-trace) — it replays the exact

failing run with DOM, network, and console; guessing from a one-line CI log is how flakes survive.

| Symptom in CI | Cause | Fix |

|---|---|---|

| Assertion races a re-render | Read-once value, not web-first | await expect(locator).toHaveText(...) |

| Mock/intercept never fires | page.route registered after goto | Register the route before the navigation |

| Test hangs / times out in an SPA | networkidle never settles (polling, websockets) | Wait on a locator/URL, not networkidle |

| waitForResponse misses the call | Subscribed after the action fired | const r = page.waitForResponse(...) before the action |

| "strict mode: resolved to 2 elements" | Ambiguous locator | Tighten with role+name/filter, not .first() |

| Passes alone, fails in the suite | storageState leak / shared mutable state | Per-role state file; fresh context per test |

| Wrong fixture/state in one file | test.use() scope confusion | Scope test.use to the right describe block |

| Green headed, red headless (or vice-versa) | Viewport/animation/timing drift | Pin viewport; reduce motion; debug in the failing mode |

Per-pattern reproduction and corrected code is in references/flakiness-playbook.md.

Anti-patterns

| Anti-pattern | Why it bites | Do instead |

|---|---|---|

| await page.waitForTimeout(2000) | Couples the test to wall-clock; too short flakes, too long drags | Wait on locator/URL/response |

| XPath or :nth-child locators | Breaks on any markup refactor the user never sees | getByRole/getByTestId ladder |

| page.$(...) / page.$$(...) element handles | No auto-wait, no retry — pre-locator API | page.locator(...) / getBy* |

| expect(await locator.textContent()).toBe(...) | Reads one frame; races the async update | await expect(locator).toHaveText(...) |

| Committing storageState JSON | Leaks session secrets, goes stale, false green | Gitignore .auth/; regenerate per run |

| trace: 'on' always | Heavy artifacts, slows every run | trace: 'on-first-retry' |

| Tests that depend on run order | One reorder cascades failures | Each test self-contained; fresh context |

| Driving pure logic through the browser | Slow + flaky for a unit-level check | Push it to ../testing-web/SKILL.md |

| .first() to silence strict mode | Hides ambiguity → the next flake | Make the locator unique |

When a flake resists the table, hand the trace to ../debug/SKILL.md — reproduce as a rate (k/N

runs), isolate one variable, fix the cause, not the symptom.

Verify

Run scripts/verify.sh [dir] over the test/config files you emit. It is a read-only static lint

(no browser, no network) that fails on the skill's own banlist: waitForTimeout(, XPath///

locators, page.$(/page.$$( handles, expect(await read-once assertions, and any

playwright.config.* missing both trace and retries. Clean or empty target exits 0.

How to use it

Copy the folder

Take ericrisco/e2e-testing from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference npx. Without those the skill loads but fails at the first command.