posthog/exploring-the-wizard
Run, drive, and explore the PostHog wizard headlessly against an app — boot it on the app and decide each screen yourself over the wizard-ci MCP tools (open_app / read_state / perform_action / run_agent), snapshotting the TUI to see what happened. Use to test or explore the wizard end-to-end.
npx skills add https://github.com/PostHog/wizard --skill exploring-the-wizard
Drive a real wizard run yourself: boot it on an app, read each screen, decide,
act, snapshot.
Everything goes through the wizard-ci MCP server, registered in this
repo's .mcp.json (npx tsx scripts/wizard-ci-mcp.no-jest.ts) — it runs the
wizard from this checkout's source, so whatever branch you're on is what
you're testing. If the tools (open_app, read_state, …) aren't available,
the server isn't approved yet — ask the user to approve wizard-ci, then
retry. For how the harness works underneath, read
e2e-harness/ARCHITECTURE.md.
framework detection, gatherContext, feature discovery, warehouse scan,
intro, setup questions, health check, auth.
run_agent): OAuth-equivalent bootstrap, thefull agent run, outro, MCP install, skills — creating **real PostHog
resources** (dashboard + insights) in the target project.
render_screen) exactly as a user sees it.open_app replaces the activewizard), or advance auth/run without credentials.
Two modes — pick by how far the run must go:
open_app needs just{ appDir, projectId }. Everything up to and including the auth screen
runs credential-free — enough to regression-test detection, setup
questions, and screen flow. End the run at auth.
auth, i.e.run_agent. Prompt the user for three things before starting:
keyFile (preferred: keeps the key out of logs). If they hold the key
in an environment variable instead, have them write it to a file first
(printenv THEIR_VAR > /tmp/phx-key — you never echo it) or pass
apiKey inline as a last resort. Never print or commit the key.
projectId).us (default) or eu (region).Always copy the target app to a throwaway /tmp copy (never a real
fixture) — the run edits files.
open_app({ appDir, projectId, keyFile?, region? }) — boots a livewizard on the app and returns the first screen.
read_state — current screen, run phase, secret-free session, tasks,and the actions legal right now. Call after every move.
perform_action({ action, params? }) — commit a decision:confirm_setup, dismiss_outage, choose (a setup question, e.g.
{ key, value }), set_mcp_outcome, dismiss_slack, keep_skills.
render_screen — render the current TUI to ANSI so you can _see_ it.run_agent — kicks off the real integration in the background andreturns immediately; it bootstraps credentials, so it's what advances
auth and run. Then poll read_state — runPhase goes
running → completed and the screen advances to outro.
A typical full walk:
open_app → intro → perform_action confirm_setup
read_state → health-check → perform_action dismiss_outage
read_state → auth → run_agent (returns at once; integration runs in background)
read_state (poll) → runPhase running → completed, screen → outro
outro → perform_action dismiss_outro → … → keep_skills
A detection-only walk (no key):
open_app → read_state (poll until detectionComplete) →
check integration / detectedFrameworkLabel → confirm_setup → auth → done, next app
Snapshot with render_screen at each key moment and save each frame to a
numbered file — /tmp/wz-explore-snaps/NN-<screen>.txt, incrementing NN in
visit order — so the run leaves a readable, ordered record you and the user
can review afterward (the same shape the CI route's .txt frames take).
Capture the run screen as it progresses, not just on screen changes.
The fixture library lives at
wizard-workbench/apps/basic-integration/<framework>/<app> (sibling repo).
To regression-test detection across every framework, loop the detection-only
walk over each app. Learned the hard way:
rsync -a --exclude node_modules --exclude .git (plusvendor/venv/Pods/build/dist). A plain cp -R of the workbench fills the
disk, and a copy that dies mid-write leaves a truncated app that detects as
null — a false regression. If detection returns null unexpectedly, check
the fixture (package.json present?) before blaming the code.
open_app returns the first paint, sometimes before detection lands(detectionComplete: false, integration: null). Poll read_state —
slower detectors (laravel, rails) need a beat.
/tmp/posthog-wizard.log. Recordwc -c < /tmp/posthog-wizard.log before the sweep and slice with
tail -c +OFFSET after — that's the run's own log, greppable for
detection lines and [bounded-fs] cap warnings.
gatherContext results don't appear inread_state or the log.** An empty setupQuestions implies the mode
resolved (ambiguity would raise a question), but for positive proof call
the util directly with npx tsx against the same fixture (e.g.
getNextJsRouter, getTanStackRouterMode).
router re-derives the active screen. Name actions, not keys.
auth and run advance only via run_agent. They expose no action anddon't self-advance. run_agent returns immediately and runs the integration in
the background — poll read_state for runPhase (running → completed).
Everything else is an instant commit.
run_agent creates real PostHog resources (a dashboard + insights) in theproject; each run duplicates them.
runPhase=completed means the flowfinished, not that the wizard understood the framework (e.g. it'll treat a Wasp
app as react-router). Read what it actually changed.
e2e-harness/, out of src/.Take posthog/exploring-the-wizard from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npx.
Without those the skill loads but fails at the first command.