calesthio/avatar-spokesperson-production
Provider-independent production workflow for AI avatar spokesperson videos, including presenter briefs, consent and likeness rights, disclosure and platform labeling, casting, localization, performance direction, brand fit, claims review, lip-sync/audio QA, accessibility, approvals, and delivery packages. Use when creating synthetic presenter videos, corporate explainers, training modules, sales/support avatars, localization spokespeople, product announcements, compliance videos, or social avatar clips.
npx skills add https://github.com/calesthio/generative-media-skills --skill avatar-spokesperson-production
Use this skill to plan, direct, review, and package synthetic presenter videos where a humanlike avatar or digital spokesperson speaks to the viewer. It is provider-independent: adapt the workflow to the avatar, TTS, dubbing, lip-sync, compositing, or platform tools available in the current environment.
Do not treat avatar production as only a prompt-writing task. The core production risks are identity rights, audience transparency, claims accuracy, performance credibility, brand trust, localization, and QA.
This skill is not legal advice. When rights, advertising claims, union/talent terms, regulated industries, minors, political/civic content, health/financial claims, public figures, or deceased personalities are involved, escalate to the client, counsel, brand owner, performer/voice owner, or compliance reviewer before generation or publication.
Separate documented facts, project requirements, and production heuristics:
Before generation, confirm:
Capture these fields before selecting tools or writing prompts:
project:
goal: "What business or learning outcome must the avatar achieve?"
audience: "Who is watching, where, and what do they already believe?"
format: "corporate explainer | training | sales | support | localization | announcement | compliance | social"
channels: ["landing page", "YouTube", "TikTok", "LinkedIn", "LMS", "internal portal"]
duration_targets: ["primary 60s", "vertical 30s cutdown", "6s hook"]
risk_level: "low | moderate | high"
persona:
role: "host | instructor | support agent | product manager | executive proxy | narrator"
identity_type: "licensed real person | synthetic non-identifiable avatar | employee likeness | public figure likeness"
voice_type: "licensed human voice | synthetic voice | cloned voice | recorded actor | localized dub"
consent_artifacts: ["contract", "release", "usage limits", "revocation/expiry terms"]
content:
script_source: "new | client-provided | translated | regulated/compliance"
claims_inventory: ["all objective claims, endorsements, statistics, comparisons"]
forbidden_implications: ["what the avatar must not appear to say, endorse, or represent"]
localization:
languages: ["en-US"]
pronunciation_list: ["brand names", "people names", "acronyms", "technical terms"]
cultural_notes: ["gestures", "formality", "dress", "examples to avoid"]
approvals:
required_reviewers: ["brand", "legal", "compliance", "talent/performer", "regional reviewer"]
Set risk_level to high if any of the following apply: public figure or employee likeness; cloned voice; medical, legal, financial, political, employment, housing, education, or safety claims; children or vulnerable audiences; synthetic expert/customer testimonials; regulated training; crisis communications; international release; or a platform where synthetic-media policy is material to distribution.
Treat avatar identity and voice as separate rights surfaces. A realistic face, body, voice, name, title, uniform, accent, and mannerism can each create audience assumptions.
Minimum consent packet for a real or cloned person:
SAG-AFTRA summarizes its AI guardrails as consent, fair compensation, and control over performances. NAVA recommends performer contracts cover consent, use limits, opt-out or term limits, payment, exclusivity, and safe storage/tracking of voice, likeness, performance, and resulting products. Use these as production governance principles even when the project is non-union, then defer to the actual contract and counsel for binding terms.
For a synthetic non-identifiable avatar, still check whether the avatar resembles a real person, implies a protected role, or uses a voice trained on identifiable speakers. Do not describe a synthetic avatar as an actual employee, doctor, customer, founder, or public official unless the identity and approval are real.
Plan disclosure as part of the creative, not as an afterthought. The disclosure should be clear enough that an ordinary viewer understands that the presenter is synthetic or AI-generated when that fact is material to trust, identity, or platform policy.
Good disclosure locations:
Avoid disclosures that are tiny, fleeting, hidden behind a click, contradicted by the script, or placed only after a misleading impression has already formed.
For advertising and endorsements, the FTC endorsement guides address endorsements and material connections under Section 5, and FTC advertising substantiation policy requires a reasonable basis for objective claims. If an avatar states a product claim, expert opinion, customer experience, ranking, statistic, savings number, health/safety outcome, or comparative superiority claim, the client must provide support before publication. Do not let the realism of an avatar create a fake testimonial or expert endorsement.
Platform and legal rules change. Re-check platform synthetic-media requirements at production time, especially for YouTube, TikTok, Instagram/Facebook, paid ads, app stores, LMS platforms, and regional variants. As verified on 2026-07-11: YouTube requires disclosure for realistic altered/synthetic content viewers could mistake for a real person, place, scene, or event; TikTok requires labels for AI-generated content containing realistic images, audio, or video; Meta describes "AI info" labels based on industry-standard signals and self-disclosure for AI-generated or AI-modified video, audio, and image content. Treat these facts as volatile.
Write for an avatar like a camera-facing presenter, not a generic narrator. The script should be easy to perform in one breath group at a time and should not require facial nuance the avatar cannot deliver.
Create a persona brief:
persona_brief:
speaker_role: "Calm onboarding coach for new enterprise users"
audience_relationship: "helpful peer, not executive authority"
credibility_basis: "company-approved training host; not a real customer or lawyer"
warmth: "medium-high"
formality: "plain professional"
pace: "145 words per minute target; slower for compliance terms"
energy: "confident but not salesy"
eye_line: "direct-to-camera for explanations; glance down only for quoted checklist moments"
gestures: "small open-hand emphasis on transitions; no pointing at viewer"
wardrobe: "solid navy blazer, no lab coat or official uniform"
background: "brand-neutral office gradient, no fake newsroom/clinic/courtroom"
forbidden_read: "must not appear to be a real employee giving legal advice"
Script rules:
OpenMontage (OH-pun mon-TAHZH), GDPR (G D P R), SaaS (sass).[beat], [slower], [smile], [firm].For compliance or training modules, separate "must-say" approved copy from optional connective tissue. Do not paraphrase regulated language unless the compliance reviewer approves.
Casting should support comprehension and trust without tokenism, stereotype, deception, or false authority.
Check:
Prefer role clarity over realism. "Synthetic product guide" is often safer and more honest than "AI customer testimonial."
Avatar providers vary in how much control they expose. Even when controls are limited, write direction in production language so the prompt, script markup, and review are aligned.
Direct:
Break long scripts into takes. Generate and review 10-20 second segments for high-risk work before committing to a full-length render. For localization, generate a short proof clip per language that includes the hardest names, numbers, and claims.
Make the avatar fit the brand without implying unauthorized affiliation or credentials.
Wardrobe:
Background:
Brand:
Create a claims inventory before production:
claims_inventory:
- script_line: "Cut onboarding time by 40%."
claim_type: "objective performance claim"
substantiation: "client study Q2 2026, approved by marketing/legal"
risk: "moderate"
reviewer: "legal"
- script_line: "I recommend this to every sales team."
claim_type: "endorsement/testimonial"
substantiation: "none"
action: "rewrite; avatar cannot imply personal recommendation"
Escalate if the avatar:
Rewrite synthetic endorsements into transparent brand narration:
Localization quality is a trust issue. A visually polished avatar with wrong pronunciation or culturally odd gestures feels deceptive or careless.
For each language or market:
Keep source text, translated script, back-translation notes, pronunciation dictionary, voice/avatar choice, and reviewer signoff in the delivery record.
Review avatar video at normal speed, half speed, and with audio-only playback.
Lip-sync and face:
p, b, m; teeth/tongue do not flicker unnaturally.Audio:
Edit/composite:
If a defect affects consent, disclosure, claims, or identity, do not "fix in post" without re-approval. Regenerate or escalate.
For prerecorded synchronized media, WCAG 2.2 Success Criterion 1.2.2 requires captions for audio content, with exceptions for media alternatives clearly labeled as such. Use this as a baseline for avatar videos even when the distribution channel is not a website.
Caption package:
[music fades], [alert tone].Also consider audio description or an equivalent text transcript if visual information not spoken by the avatar is necessary to understand the video.
Use gates. Do not wait until the final render to discover a rights or compliance problem.
Recommended gates:
Record approvals with dates and scope. If the script, avatar, voice, language, channel, claim, edit, or disclosure changes after approval, route back to the affected reviewer.
Deliver more than one MP4 when the project will be distributed across channels.
Include:
Production intent: a 90-second internal onboarding video introducing a new expense workflow.
Approach:
Example direction:
Create a waist-up AI avatar presenter in a neutral modern office setting with soft brand-blue accents. The presenter is a synthetic training host, not a real employee. Tone: clear, calm, reassuring, professional. Eye-line: direct to camera. Gesture density: low; small open-hand gestures only at section transitions. Wardrobe: solid navy blazer, no badges, no logo shirt. Pace: about 140 words per minute, with a short pause after each numbered step. Keep the lower third and bottom 20% of frame free for captions.
Script excerpt:
"Welcome to the new expense workflow. [beat] This training presenter is AI-generated, and the process details have been approved by the People Operations team. Starting August first, submit travel receipts through the Expenses tab. The system will ask three questions: the trip name, the business purpose, and whether a client attended. [slower] If a receipt is missing, add a note before you submit."
Expected QA focus: pronunciation of policy terms, no fake employee implication, captions clear on mobile, product UI screenshots current, no promise of faster reimbursement unless substantiated.
Production intent: product support videos in English, Spanish, and Japanese.
Approach:
Example localization record:
termbase:
"OpenMontage": "OH-pun mon-TAHZH; do not translate"
"workspace": "approved localized product term per market glossary"
"render": "use product glossary term, not literal machine translation"
proof_clip_lines:
en-US: "OpenMontage keeps your workspace renders organized, even when a project has multiple language versions."
es-MX: "OpenMontage organiza los renders del espacio de trabajo, incluso cuando un proyecto tiene varias versiones de idioma."
ja-JP: "OpenMontage wa, fukusu no gengo versions ga aru project demo, workspace renders o seiri shimasu. (Example romanized placeholder; use approved Japanese copy in production.)"
reviewers:
required: ["regional support lead", "native-language reviewer", "brand"]
Expected QA focus: natural localized phrasing, mouth-shape acceptability, no mismatched captions, correct UI language, correct formality.
Production intent: 30-second vertical launch clip for LinkedIn, Instagram Reels, TikTok, and YouTube Shorts.
Approach:
Example opening:
[On-screen lower third for first three seconds: "AI-generated product guide"]
"Meet the new release in thirty seconds. [beat] OpenMontage now helps teams review avatar videos before they publish: script, captions, claims, and platform labels in one approval flow."
Claims review:
Verified on 2026-07-11 unless noted. Re-check volatile platform/provider rules at production time.
Take calesthio/avatar-spokesperson-production from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.