mcpbeat Sign in

Agent First Screenshots Agent Skill

Agent-first screenshots — an agent drives the real app via CDP and produces clean, defect-free product screenshots (newsletters, landing pages, social, decks, PR). Dual-channel verification (DOM + pixels + vision) in a capture loop. Use for any "take/redo screenshots of the app" task.

4k tokens
context cost
the whole folder, loaded on every use
3
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
3593
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/Devin-AXIS/iPolloWork --skill agent-first-screenshots

What comes with it

9 453 bytes besides the instruction
scripts/capture-verify.mjs
scripts/screenshot-verify.mjs

The instruction itself

7 sections, as written by the author

Agent-First Screenshots

An agent-first screenshot is produced by an agent driving the real app via

CDP, gated by structural + pixel + vision verification before it ships. It

turns screenshots from hand-crafted artifacts into a regenerable pipeline.

Use this for "take/redo screenshots / make marketing images" tasks. It is NOT

e2e evidence — for pass/fail proof use the fraimz skill.

The Zero-Defect Bar

A shippable screenshot has ZERO obvious visual defects — the bar is not

"the content is present", it is: *would a human glancing at it immediately spot

something broken?* Overlapping or clipped text, misaligned/collapsed layout,

garbled or mid-transition content, blank regions, stray modals/tooltips — any

of these disqualifies the frame. Full stop.

The One Method

**Operate the app like a power user preparing a demo, not like a developer

hacking the DOM.** The app already looks great: get it into a real, settled

state through the UI and capture it cleanly. Never rebuild or fake it.

Golden Rules

  • Only remove, never add. Hiding a leaf distraction (notification badge,

"Sign in" footer link, status text) via display:none is safe. Never

inject fake HTML, override flex/height/width on structural containers, add

fixed-position overlays, or mutate the DOM tree — the layout engine will

collapse (scroll areas shrink to 0, content disappears).

  • If a feature isn't available, don't fake it. Enable it through the

settings UI like a real user, or skip the shot and say so.

  • Each shot starts from a clean reload. location.reload(), wait for

full render, navigate through the UI, minimal leaf cleanup, verify, capture.

Never carry CSS hacks across shots.

  • Get state naturally. Click the real tabs/buttons/pickers; close panels

via their close buttons (not CSS); run real multi-turn tasks so content is

impressive — no toy data, no mid-stream captures.

  • Match the target aspect ratio via CDP metrics override (e.g. 1440x900

at deviceScaleFactor: 2), then wait ~1s and re-verify — the override can

trigger re-layout.

Verify with three channels, in a loop

innerText existing does not mean visible; a healthy DOM rect does not mean it

rendered. Every frame must pass all three before it ships:

  • DOM pre-check (cheap, before capture): scroll area height > 100px (not

collapsed), hero text inside the viewport rect, no unexpected modal/overlay.

If it fails: reload and redo — never fix with more CSS.

  • Pixel post-check (truth, after capture): decode the PNG and sample the

DOM-derived hero rects. Calibration: for text-on-white UI, background ratio

is a bad signal (85–92% background is normal). **Variance is the reliable

signal**: content-filled region variance > ~200 (often 1000+); blank/flat

region < 50. Whole-image bgRatio > 0.97 → blank frame.

  • Vision check (mandatory gate): deterministic stats CANNOT catch the

defects that matter most — overlap, clipping, misalignment,

double-rendering. Hand the PNG to a vision-capable model with a zero-defect

rubric returning JSON:

   { any_obvious_defect: bool,   // the gate — if true, REJECT
     overlapping_text_or_elements: bool, clipped_or_cutoff: bool,
     misaligned_or_broken_layout: bool, blank_or_empty_regions: bool,
     stray_modal_tooltip_or_panel: bool, legible: bool,
     polish_score: 1-5, defects: ["..."] }

If any_obvious_defect is true the frame is rejected regardless of layers

1–2. Diagnose, fix the root cause (settle the state, close the picker,

reload), recapture.

Reusable scripts in this skill's directory: screenshot-verify.mjs (verify an

existing PNG) and capture-verify.mjs (capture at 2x + verify regions + save

only on pass); both use sharp.

Common failure modes

| Symptom | Fix |

|---------|-----|

| Overlapping/clipped text (layers 1–2 pass) | Unsettled/edit-mode/wrong-width state — settle, use preview mode, gate on vision |

| Blank screenshot / 32px scroll area | Structural element was CSS-hidden — reload, close panels via UI |

| Overlay (picker/modal) on every shot | It was opened and never closed — Escape before capturing |

| In DOM but not in pixels | Clipped/blank render — trust variance + vision, not innerText |

Anti-patterns

Injecting fake components; overriding flex properties; fixed-position

overlays; carrying state across shots; verifying via innerText only;

proof-frame mindset (the goal is "I want that", not "it didn't crash").

How to use it

Copy the folder

Take devin-axis/agent-first-screenshots from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.