do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.
npx skills add https://github.com/Devin-AXIS/iPolloWork --skill run-evals
Launch a real iPolloWork app and run coded eval flows against it. This skill owns
launch + run; the prove/repair/verdict loop and evidence standard live in
the fraimz skill — load that too for anything that ends in a verdict.
daytona CLI installed and logged in (daytona login), right org selected(daytona organization use "<org-name>")
.devcontainer/ files present in the repobash .devcontainer/setup-daytona-secrets-volume.sh .newtoken (never print
keys; sandboxes source every /daytona-secrets/*.env before Electron starts)
daytona organization use "<org-name>"
bash .devcontainer/test-on-daytona.sh <branch-or-commit> --artifacts-volume
The helper creates a fresh VNC-capable sandbox from the ipollowork-eval-vnc
snapshot, mounts the secrets + pnpm-store volumes, starts XFCE/noVNC, Vite, and
Electron with Daytona-safe flags, waits for CDP, then prints the CDP and noVNC
URLs. --artifacts-volume mounts /daytona-artifacts served on port 8090 for
published frame proof. Refresh the snapshot when dependencies change:
bash .devcontainer/create-daytona-ipollowork-snapshot.sh.
Verify the endpoint before running flows:
browser_list({ browser_url: "<CDP_URL>" }) # must show an "iPolloWork" target
If it fails, inspect /tmp/electron.log — the real success marker is
Chromium's DevTools listening on ws://127.0.0.1:9825/....
If the app shows the Welcome page, create a workspace first (see
evals/daytona-flows.md Flow 1: create /workspace/hello on the sandbox,
"Get started" → "Local workspace" → inject the path → "Create Workspace").
pnpm evals --list
pnpm evals --flow <flow-id> --cdp-url <printed-electron-cdp-url>
pnpm evals --all --stack den # brings up MySQL + den-api + seed for cloud flows
The runner produces machine-checkable assertions, validated screenshots, and
writes fraimz.html + report.md / report.json under
evals/results/<run-id>/. If no coded flow exists for the behavior, add one in
evals/flows/<id>.flow.mjs (see the fraimz skill and evals/README.md for
the ctx.* API); use manual browser tools only to debug or prototype — a coded
flow is the PR evidence.
Frame proof is the default deliverable; record video only when motion matters
(streaming, animations). Start with
bash .devcontainer/test-on-daytona.sh <branch> --record-video --recording-name <name>,
stop with daytona exec "$SANDBOX" -- 'bash .devcontainer/stop-daytona-recording.sh',
download via the port-8090 artifacts URL. Details: daytona-recording-artifacts.
When Daytona is down or quota-limited:
pnpm install
pnpm --filter @ipollowork/app typecheck
IPOLLOWORK_ELECTRON_REMOTE_DEBUG_PORT=9826 pnpm dev # then:
pnpm evals --flow <flow-id> --cdp-url http://127.0.0.1:9826
Report clearly whether the result came from Daytona or the local fallback — a
local run is not a Daytona validation.
daytona delete "$SANDBOX"
Take devin-axis/run-evals from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npm.
Without those the skill loads but fails at the first command.