calesthio/product-ad-production
Diagnose, design, script, produce, adapt, and quality-control product video advertisements from evidence-backed briefs. Use for physical-product, software, app, service, ecommerce, direct-response, brand-response, demonstration, testimonial, launch, and variant ad work where product truth, legibility, claims, platform delivery, and measurable iteration matter.
npx skills add https://github.com/calesthio/generative-media-skills --skill product-ad-production
Produce an advertisement that makes the right person recognize a relevant situation, understand what the product changes, believe the proof, and know what to do next. Visual polish is not a substitute for product comprehension or claim support.
This skill is provider-neutral. Select image, video, capture, voice, music, editing, and composition tools only after the advertising problem is defined. Do not let a generation model's favorite visual language determine the concept.
Platform behavior and ad products change. Platform facts in this skill were verified 2026-07-09. Re-check the exact placement specification, policy, interface preview, and destination requirements immediately before export and trafficking.
Do not write a script from a slogan and a product URL. Diagnose the commercial problem first.
Capture or infer each field, marking assumptions explicitly:
| Field | Required decision |
|---|---|
| Business outcome | Awareness, consideration, lead, trial, purchase, upsell, retention, or launch learning |
| Primary success measure | One decision metric; supporting and guardrail metrics are separate |
| Audience | Purchase situation and constraints, not demographics alone |
| Job to be done | Progress sought, present struggle, emotional/social stakes, and competing solutions |
| Awareness state | Unaware, problem-aware, solution-aware, product-aware, or returning customer |
| Product truth | What the product is, how it works, prerequisites, limitations, price/offer, and current availability |
| Proposition | One useful change the ad will make credible |
| Evidence | Approved substantiation, demonstrations, customer data, certifications, testimonials, and provenance |
| Objection | The most consequential reason a qualified viewer may not act |
| Placement | Platform, ad product, buying objective, duration, aspect ratio, UI overlays, and sound context |
| Destination | Exact landing page or store state; product, offer, naming, and price must agree with the ad |
| Brand system | Correct product/packaging, marks, type, colors, sonic assets, and prohibited treatments |
| Rights and policy | Talent, location, music, stock, creator, user content, AI/replica consent, category restrictions |
| Production constraints | Deadline, budget, available real product, UI build, languages, render and review path |
Return one of three diagnoses:
Jobs to Be Done theory treats products as choices people make to achieve functional, emotional, and social progress, not merely as bundles of attributes [S1]. Translate the job into an ad-specific record:
When [situation/trigger], this person is trying to [progress],
but [friction/anxiety/current alternative].
They will believe a better option when they see [proof],
and the smallest credible next action is [CTA].
Separate buyer, user, approver, and beneficiary when they differ. For B2B, map the operational user, economic buyer, security/legal reviewer, and implementation owner. An ad normally speaks to one primary role and acknowledges one decisive objection; it should not average all roles into vague corporate language.
Documented fact—United States: advertising must be truthful and non-deceptive, advertisers must possess evidence for claims, and ads cannot be unfair [S2][S3]. The FTC evaluates the ad's overall or “net” impression, including words, images, sounds, omissions, and format; a small disclosure does not neutralize a misleading main message [S4]. Other jurisdictions and regulated categories may impose additional rules. Obtain qualified legal review when stakes warrant it.
Create the ledger before scripting. Every spoken, written, visual, sonic, comparative, and implied claim gets an ID.
| Field | Record |
|---|---|
| Claim ID | Stable identifier used in script, shot list, and approvals |
| Exact expression | Wording or visual implication the viewer receives |
| Type | Existence/feature, performance, quantified, comparative, superiority, price/offer, testimonial, environmental, health/safety, or illustrative |
| Audience interpretation | Reasonable takeaway, including likely implied meaning |
| Evidence | Document, dataset, test report, product record, or approved customer source |
| Scope | Model/SKU, population, conditions, geography, dates, and exclusions |
| Qualification | Material information needed beside the claim |
| Demonstration link | Shot IDs that prove or illustrate it |
| Owner and approval | Product/legal/brand approver and status |
| Expiry trigger | Product revision, offer end, evidence age, or policy change |
Documented fact—United States: endorsements must reflect the endorser's honest experience; material connections should be clearly disclosed, and atypical results may require disclosure of generally expected results [S3][S5]. The FTC advises placing a social disclosure with the endorsement itself, making it hard to miss, and including it in the video rather than relying only on a description [S5]. The U.S. Consumer Reviews and Testimonials Rule prohibits specified fake or false reviews and testimonials, including certain AI-generated fakes [S6].
Maintain a testimonial packet: original statement, permission, identity, actual-use confirmation, compensation/material connection, editing approval, validity date, and support for any product claim repeated by the endorser. An actor may portray a dramatized user only when the ad cannot reasonably be read as that actor's genuine testimonial. Label dramatizations where needed.
Compress the strategy into:
For [audience in situation], [product] changes [friction] into [outcome]
because [credible mechanism/proof]. After viewing, do [one action].
Generate multiple concepts that differ in mechanism, not merely art direction. Useful concept families include:
Score each concept from 0–3 on audience relevance, claim support, product centrality, demonstration clarity, brand distinctiveness, placement fit, variant potential, feasibility, and rights/policy risk. Reject any concept whose appeal depends on a behavior the product cannot perform.
Research finding—platform-specific: Google's data-backed ABCD guidance recommends earning attention, branding early and throughout, creating a focused human connection, and giving clear direction; for consideration it specifically recommends making the product prominent and showing how it works [S7]. Google's reference guide also says its evidence is correlational and not a one-size-fits-all guarantee [S8]. Use this as informed prior evidence, then test it for the brand and placement.
The viewer should be able to answer, at the appropriate moment:
Design explicit recognition events:
Do not assume that a tiny logo, generic color, or beauty silhouette identifies an unfamiliar product. The Ehrenberg-Bass Institute describes a strong distinctive asset as both famous and unique to the brand; an unmeasured visual element should not be treated as a brand-name substitute [S9].
Practitioner heuristic: perform a “blurred-thumbnail test” for silhouette/color recognition, a one-second freeze-frame test for product/category recognition, and a five-second recall check with people who did not see the brief. These are diagnostic screens, not campaign-effectiveness measures.
A demonstration is a small experiment presented to viewers. Define it before shooting or generating:
| Element | Question |
|---|---|
| Claim | What exact viewer conclusion may this demo support? |
| Unit | Which real SKU, version, plan, build, and accessories are used? |
| Conditions | What environment, starting state, preparation, network, lighting, load, or operator skill matters? |
| Comparator/control | Is the baseline fair, current, and equivalently prepared? |
| Procedure | What actions happen, in what order, and how are they timed or measured? |
| Capture | Which uninterrupted wide proof shot and which detail inserts are needed? |
| Repetitions | Is one event enough, or must variability/typicality be shown? |
| Qualification | Which material conditions must appear with the result? |
| Failure rule | What result stops the advertised claim or requires revised language? |
For physical products, capture a continuity master that shows setup, product, hands, comparator, timer or measuring device, and outcome without a claim-altering cut. Detail shots can improve comprehension but cannot replace the proof master.
For software and apps:
For comparisons, match conditions and show the basis of comparison. Avoid degrading a competitor, choosing an obsolete baseline without disclosure, changing camera treatment to favor one outcome, or saying “better” without defining the dimension.
Practitioner heuristic: if the proof cannot survive a locked-off wide shot, it is not yet a demonstration; it is an illustration. Illustration can still be useful, but label and script it accordingly.
There is no mandatory hook-problem-solution template. Build beats around the viewer's decision and the placement's interruption pattern.
Possible beat functions:
For every beat, record:
time range | beat purpose | viewer question answered | visual action
product/brand cue | dialogue/VO | on-screen text | claim IDs
audio event | transition logic | source asset | disclosure/qualification
placement risks | cutdown dependency
Specify shots by advertising function, not by decorative camera vocabulary:
| Shot role | Acceptance criterion |
|---|---|
| Product identity | Exact product is recognizable; logo/label is not malformed or obscured |
| Situation | Friction is clear without stereotyping or humiliating the user |
| Operation | Contact points, sequence, orientation, and causality are readable |
| Proof master | Claim-relevant action remains continuous and conditions are visible |
| Detail | Reveals a mechanism or result the master cannot show |
| Human consequence | Performance and result match the intended audience/job |
| Offer/CTA | Product, offer, destination, terms, and next action agree |
AI-generated shots require additional continuity controls: approved pack/reference images, geometry and label verification, state tracking across shots, hand/product interaction review, and a ban on hallucinated ports, controls, accessories, ingredients, results, or interface elements. When accuracy is essential, use real product capture, composited pack shots, deterministic 3D, or real UI rather than asking a generative model to reproduce critical details.
Do not call alternate color grades “three concepts.” Build a variant matrix where each version has one named learning goal.
| Variable | Example hypothesis |
|---|---|
| Situation/hook | Recognized morning friction will improve qualified hold versus abstract intrigue |
| Audience/job | Team lead outcome language will outperform agent-level feature language for lead quality |
| Proposition | Time saved will outperform compactness for commuters |
| Proof | Continuous demonstration will outperform testimonial-only proof |
| Objection | “No setup required” will increase trials among switchers |
| Offer | Trial framing will outperform discount framing on qualified conversion |
| CTA | “Build your first board” will outperform generic “Learn more” |
| Length/placement | A native 12-second proof cut will outperform a cropped 30-second master on vertical feed |
Create a version manifest with parent version, changed variable, preserved variables, claim-ledger version, assets, platform, language, duration, export checksum, approval state, and live dates.
Documented platform method: TikTok's split-testing guidance describes isolated audience groups and holding other variables constant; Google Video Experiments can compare different video ads with the same audience [S10][S11]. Platform systems and eligibility differ, so plan enough traffic/time and use the platform's experiment design rather than judging early noise.
Practitioner heuristic: begin with materially different propositions or proof approaches, then optimize execution within the winning strategic territory. Tiny changes too early produce precise answers to unimportant questions.
Plan four layers:
Maintain a cue sheet with track/source, composer/performer, license, territories, media, term, paid-social rights, edits, stems, and expiry. “Royalty-free” does not mean unrestricted. Confirm whether a platform music library permits paid advertising and off-platform versions.
Caption all meaningful speech and relevant non-speech information when accessibility or destination context calls for it. Documented standard: WCAG 2.2 requires captions for prerecorded synchronized media in covered web contexts; captions should not obscure relevant visual information [S12]. For critical text, target WCAG's 4.5:1 contrast for normal text and 3:1 for large text where applicable [S13], while recognizing that platform video itself may be governed by different formal requirements.
Mix against the destination's current specification. Check dialogue intelligibility, true peak/clipping, phase, mono compatibility, mobile-speaker translation, and head/tail silence. For U.S. digital television delivery, ATSC A/85 is the relevant loudness recommended practice; do not apply a broadcast target blindly to social exports [S14].
Practitioner heuristic: the visual track should communicate product, change, and CTA with sound unavailable; the audio track should still identify product and proposition when the viewer glances away. This is a robustness test, not an instruction to weaken either channel.
A strong CTA names one feasible action and the immediate value of taking it:
Verify product/SKU, brand name, price, discount math, dates, inventory, shipping territory, eligibility, subscription renewal, button text, app-store state, language, tracking parameters, and page speed/function. TikTok's current policy explicitly requires consistency among ad, caption, CTA, product, and landing page [S15]; treat this as good production discipline everywhere.
Build a clean, high-resolution master with separated dialogue, music, effects, captions, graphics, textless plates, and product/pack layers. Then author placement versions.
For each placement:
These are planning anchors, not a substitute for the live spec:
Create a delivery manifest containing placement, ad product, objective, duration, canvas, pixel dimensions, frame rate, codec/profile, bitrate or quality mode, audio codec/sample rate/channels, loudness target if specified, captions mode/file, safe-zone template version, filename, checksum, source master, claims version, and approval.
Tie evaluation to the business decision:
The MRC viewability convention counts a viewable video impression when at least 50% of pixels are in view for at least two continuous seconds [S22]. That is a measurement threshold, not proof that the ad was noticed, understood, remembered, or persuasive.
Use a learning log:
hypothesis | treatment/control | single changed variable | audience/placement
primary metric | guardrails | planned sample/time | result/uncertainty
claim or product changes during test | decision | next test
Do not optimize a brand objective solely to click-through rate, or a sales objective solely to completion rate. Do not declare a winner from a few conversions, overlapping audiences, unequal offers, mid-test edits, or a platform label you do not understand. Where feasible, use randomized experiments or lift studies; Google distinguishes A/B experiments from exposed-versus-control lift studies for measuring incremental outcomes [S23].
| Failure | Why it fails | Repair |
|---|---|---|
| Beautiful category film; product appears at end | Attention is not linked to the advertised choice | Put product/mechanism into the premise and create early recognition |
| Feature inventory | No prioritized job or consequence | Choose one situation, one change, one proof |
| Generic “game-changing” language | Unsupported and non-diagnostic | State the exact outcome and show evidence |
| Demonstration built from inserts only | Cuts can conceal setup or causality | Add an uninterrupted proof master and disclose material conditions |
| Fake interface speed | Implies performance not delivered | Capture real latency or label a concept; edit around non-material pauses only |
| Testimonial by synthetic person | Can be read as a real experience | Use a documented real endorser or unmistakable dramatization |
| Tiny disclaimer contradicts headline | Net impression remains misleading | Narrow or remove the headline claim |
| One master auto-cropped everywhere | Product/text/UI are occluded; pacing mismatches placement | Re-author framing, type, beats, and CTA per placement |
| Thirty variants with no hypothesis | No interpretable learning | Create a version matrix and isolate consequential variables |
| Optimizing to cheap views | Delivery proxy displaces business goal | Predeclare primary and guardrail metrics |
| AI-mutated pack/UI between shots | Product becomes fictional | Lock approved references; composite or use real/3D/UI capture |
| Music chosen after edit | Rights, rhythm, and intelligibility break late | Decide source/license and audio role before picture lock |
| CTA and landing page disagree | Viewer trust and platform compliance suffer | QA the live destination with every trafficking version |
The following are fictional examples, not mandatory formulas. Their products, evidence, and results are invented for training and must not be reused as real claims.
Intent: direct-response ad for hybrid workers who carry a laptop stand but avoid using it because setup is annoying.
Product: fictional SnapFold Stand S2.
Constraints and evidence: 9:16 paid social; real production unit and packaging available. Approved internal dimensional inspection says the stand folds to 14 mm at its thickest point (n=20 production units, range 13.7–14.2 mm). Approved load test says it held a centered static 5 kg load for 60 minutes without collapse (n=10); this is not permission to call it “indestructible.” No ergonomic or health claims. Offer is 15% off through 2026-08-31 in the U.S.; landing page is live.
JTBD: “When I move between home and coworking spaces, I want a usable laptop setup without carrying bulky gear or performing a fiddly assembly.”
Proposition: full stand function from a slim folded object, with the setup shown in one continuous action.
Claim ledger excerpt:
| ID | Expression | Support/scope |
|---|---|---|
| C1 | “14 mm folded” | Production-unit inspection; show “at thickest point” qualification |
| C2 | “Opens in one motion” | Demonstration language; defined as one continuous unfolding action, not a timed claim |
| C3 | “15% off through Aug 31” | Approved U.S. offer and matching landing page |
Demonstration protocol: begin with stand folded beside a ruler/depth gauge; same unit remains in frame while one operator unfolds it and places the laptop. No cut from folded state through stable placement. Detail inserts may follow. Failure: if the unit requires a second adjustment hidden by framing, remove C2 or show the adjustment.
Complete beat/shot plan:
| Time | Visual | Audio | Text/claim |
|---|---|---|---|
| 0.0–2.0 | Commuter pulls a bulky old stand from a bag; it catches on the zipper. Snap cut to the slim S2 sliding out cleanly. Both are truthful props; no competitor marks. | Zipper catch, then clean slide; VO: “Your portable setup shouldn't fight the bag.” | Product name appears by 1.2 s; no comparative performance claim |
| 2.0–5.0 | Locked proof master: S2 beside depth gauge, reading visible. Hand picks up the same unit. | Music pulse begins; VO: “SnapFold S2 packs down to 14 millimeters…” | 14 mm folded* C1; *at thickest point readable beside it |
| 5.0–9.0 | Still in the proof master, hand completes one uninterrupted unfolding action and places laptop. | Real hinge click; VO: “…then opens in one motion.” | One continuous setup C2 |
| 9.0–13.5 | Detail: hinge seats; hands type; wider view keeps whole stand identifiable. | Typing under VO: “Desk height when you need it. Flat when you don't.” | “Desk height” is descriptive, not an ergonomic benefit claim |
| 13.5–16.5 | Folded S2 returns to side pocket; packaging beside it confirms exact SKU. | Hinge click becomes sonic mnemonic. | SnapFold Stand S2 |
| 16.5–20.0 | Product and three available colors on a clean end frame. Offer and CTA remain inside placement safe zone. | VO: “Choose your S2. Save 15% through August 31.” | See colors CTA; C3; U.S. only. Terms apply. |
Production inputs: real S2 unit/pack; approved depth gauge; laptop weight within product instructions; signed talent/location releases; owned hinge recording; licensed paid-social music; safe-zone template; destination screenshot.
Expected result: the viewer sees the exact object, understands the compact-to-usable change, witnesses the setup, and receives a destination-specific CTA.
Likely failures: ruler unreadable on phone; first prop accidentally implies a named-competitor comparison; generative insert mutates hinge or logo; qualification collides with caption; “desk height” read as an ergonomic health claim. Repair with macro-but-continuous evidence, generic baseline, real/composited product, safe-zone rebuild, and revised wording.
Meaningful variants: A changes only opening situation from zipper snag to crowded café table; B changes only proof from dimensional gauge to side-by-side bag-pocket fit; C keeps creative fixed and tests See colors versus Shop S2 if destinations are equivalent.
Intent: generate qualified demo bookings from customer-support operations leads, not maximize generic clicks.
Product: fictional QueueWise Routing, available only on the Enterprise plan.
Evidence and constraints: a named customer's approved case-study dataset covers eight weeks before and eight weeks after deployment. Median first-response time changed from 3 h 42 m to 1 h 58 m for email tickets; ticket mix and staffing also changed, so the company approved only observational wording, not “QueueWise cut response time by 47%.” Customer logo, quote, and spokesperson release are approved for North American paid media through 2026-12-31. Real UI capture uses synthetic ticket data. 16:9 YouTube in-stream master plus separately authored 9:16 feed version.
JTBD: “When ticket volume shifts, I need urgent work to reach the right specialist without asking agents to monitor multiple queues, while retaining control and auditability.”
Proposition: show the routing rule, then scope the observed customer result honestly.
Claim ledger excerpt:
| ID | Expression | Support/scope |
|---|---|---|
| Q1 | “Route by issue, account tier, and language” | Current Enterprise production build and documentation |
| Q2 | “At Northstar, median email first response moved from 3h42m to 1h58m” | Approved 8-week before/after case dataset; observational wording and period required |
| Q3 | Customer quote: “We stopped babysitting three inboxes.” | Signed exact quote; speaker/title current at approval |
| Q4 | “Book a routing review” | Live booking page with correct event and region |
Complete beat/shot plan:
| Time | Visual | Audio | Text/claim |
|---|---|---|---|
| 0–4 | Realistic queue view fills with three differently tagged tickets. Cursor does not move yet. | Alert texture; VO: “When every queue looks urgent, who watches the queues?” | Three queues. One team. |
| 4–9 | Real production UI: rule builder selects issue, account tier, language; cursor path and labels remain readable. | VO: “QueueWise routes by issue, account tier, and language.” | Q1; Enterprise plan adjacent |
| 9–13 | Uninterrupted capture: one synthetic ticket arrives and is routed to the matching specialist; audit entry opens. | Subtle confirmation sound; no time compression. | Rule matched → assigned → logged |
| 13–20 | Approved customer spokesperson in their real workspace, disclosure/identity lower third. Cutaway shows approved Northstar dashboard export, not recreated vanity data. | Speaker: “We stopped babysitting three inboxes.” | Q3; Northstar customer; results vary where appropriate |
| 20–25 | Simple before/after data card tied to exact scope. | VO: “Across the approved case period, Northstar's median email first response moved from three forty-two to one fifty-eight.” | Q2; 8 weeks before vs. 8 weeks after deployment. Observational result; staffing and ticket mix also changed. |
| 25–30 | Product UI and logo remain visible; booking page preview shows the exact next step. | VO: “See whether your routing is ready. Book a workflow review.” | Q4; Book a routing review |
9:16 adaptation: re-record the UI at a narrow approved responsive layout; use a guided crop only where labels remain readable; move customer identity and qualifications above the platform caption/CTA area; shorten the first beat to 2.5 seconds without cutting the rule-to-audit proof; keep the case result fully scoped.
Expected result: operations leads understand the workflow and can distinguish a product capability from one customer's observed outcome.
Likely failures: turning the case result into a causal percentage; using a generic fake dashboard; implying the feature exists on lower plans; tiny qualification; booking to the homepage; customer-rights expiration during the flight. These are release blockers.
Meaningful variants: proposition variant A leads with control/auditability; B leads with reduced queue monitoring; proof variant C removes the testimonial but preserves the UI demonstration and scoped case data. Do not change audience, offer, and CTA in the same test if the goal is to learn which proof structure works.
All web sources below were checked 2026-07-09.
Take calesthio/product-ad-production from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.