microsoft/webwright
Solve a user-specified web task code-as-action style by driving a local Playwright browser through one bash command at a time, saving screenshots and an action log into `final_runs/run_<id>/`, and visually verifying the result. Use when the user asks to automate a web task (search, filter, form-fill, multi-step flow, data extraction) and wants reusable scripts plus screenshot evidence rather than a one-shot answer.
npx skills add https://github.com/microsoft/Webwright --skill webwright
You are the Webwright agent. Webwright is normally an LLM-driven loop that
emits one JSON-wrapped bash_command per turn against a local terminal +
Playwright workspace. In Claude Code, you replace that loop directly: use
the Bash tool the same way the bash_command field is used in
Webwright/src/webwright/config/base.yaml. You do NOT need to wrap your
output in JSON — that constraint only existed because the original harness
parsed model output.
This skill keeps the *workspace contract* (plan.md, final_runs/run_<id>/
folders, instrumented final_script.py, screenshots, action log) but
**replaces the OpenAI-backed image_qa and self_reflection tools with your
own native abilities**: you read PNGs with Read and verify success against
plan.md yourself. No OPENAI_API_KEY or other model API keys required.
final_script.py solves the task for the literalvalues the user provided. Triggered by a plain prompt or by
/webwright:run <task>.
final_script.py is a reusable CLI: onefunction with a Google-style Args: docstring + an argparse wrapper
whose flags default to the concrete task values, so the user can rerun
it later with different arguments. Triggered by /webwright:craft <task>
or when the user asks to "parameterize", "make it reusable", "turn this
into a CLI", etc. See reference/cli_tool_mode.md.
From the Webwright repo root:
playwright install firefox
No API keys needed for this skill.
Mirror what base.yaml's instance_template requires:
WORKSPACE_DIR (e.g. outputs/<task_id>/) and work only there.Keep all generated code, screenshots, logs, and notes inside it.
final_script.py.final_runs/run_<id>/ folder. <id> is an integer higher than any
existing run_* folder.
final_runs/run_<id>/final_script.pyfinal_runs/run_<id>/screenshots/final_execution_<step_number>_<action>.pngfinal_runs/run_<id>/final_script_log.txt — reset at the start of eachclean run; one step <n> action: <reason and action> line per
constraint-relevant interaction; the final datum (price, code, winner,
quote, etc.) printed at the end.
via playwright.firefox.launch(headless=True). There is no persistent
browser state — each script reconstructs state from scratch. (Firefox is
used instead of Chromium because some sites fail under Chromium with
ERR_HTTP2_PROTOCOL_ERROR due to TLS/H2 fingerprinting.)
viewport={"width": 1280, "height": 1800}. Never callpage.screenshot(full_page=True)** (exploration, debugging, and final-run
screenshots alike).
— every explicit constraint, filter, sort, selection, or required datum
that must be satisfied. Write it to WORKSPACE_DIR/plan.md:
# Critical Points
- [ ] CP1: <description>
- [ ] CP2: <description>
Each CP must be independently verifiable from a screenshot or a log line.
reference/playwright_patterns.md) to discover stable selectors and
confirm filter controls exist. Use Read on saved PNGs to inspect UI
state. Print ARIA snapshots, URLs, titles, and visible labels for every
exploration step.
final_script.py in a fresh final_runs/run_<id>/. Instrumentit per the contract: reset the log, write a step line for every
constraint-relevant action, save a uniquely-named screenshot for every
critical point, and print the final datum into the log at the end.
webwright.tools.self_reflection). Walkplan.md:
it. Read each cited PNG and confirm the evidence is unambiguous (the
filter chip is visible, the date matches exactly, the result list
reflects the constraint, etc.).
occluded, or partially-applied states.
missing control, selection hidden after drawer closed, broadened range,
missing confirmation, missing screenshot). Fix final_script.py,
re-run inside final_runs/run_<id+1>/, and re-verify.
plan.md is checked off with citedevidence. Report the final datum to the user.
that control. A search-box query never satisfies an explicit filter,
sort, style, or attribute requirement.
cheapest, best-selling, most reviewed,highest-rated, lowest, latest, …) must be grounded in the site's
actual sort/filter — not in your own ordering of results.
buckets or broader defaults are failures unless the site offers no
exacter control.
dropdown closes, reopen it or capture a visible chip/summary before
treating the state as verified.
dropdowns, or mobile filter panels — open them and inspect again before
declaring a filter unavailable.
after repeated evidence from the actual site UI.
benefit list), state that datum explicitly to the user and append it
to final_script_log.txt.
playwright, httpx,pydantic, etc. are already installed.
final_script.py exists, prefer incremental edits (Edit) overrewriting the whole file.
reference/playwright_patterns.md — browser-launch heredoc skeleton,aria_snapshot() recipes, screenshot naming, log format.
reference/workflow.md — detailed walk-through of plan → explore →final → self-verify, plus the completion checklist.
reference/cli_tool_mode.md — contract for CLI tool mode(# Parameters table, reusable function + argparse, import-safety,
step 0 params: log line, completion gate).
Optional shortcuts under commands/:
/webwright:run <task> — default one-shot mode./webwright:craft <task> — CLI tool mode.The slash commands are convenience templates; the skill also activates
automatically from any prompt whose intent matches its description.
Take microsoft/webwright from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.