calesthio/performance-direction
Provider-independent performance direction for generated video, avatar/spokesperson clips, animation, ads, film scenes, social content, and voice-led media. Use when an agent must cast or direct synthetic performers, avatars, animated characters, talking heads, AI video characters, or voice performances; plan acting beats, objective/obstacle/action, blocking, gesture, eye-line, facial expression, body language, lip-sync, voice alignment, rehearsal prompts, iteration, continuity, consent, and performance QA.
npx skills add https://github.com/calesthio/generative-media-skills --skill performance-direction
Use performance direction to make a generated person or character appear to want something, react to pressure, and communicate through voice, face, gaze, and body. Do not ask a model for a generic emotion and hope it supplies acting. Give playable circumstances, beat-level actions, physical behavior, vocal behavior, continuity constraints, and a way to judge the take.
Documented facts used by this skill:
Empirical observations to collect on each job:
Production heuristics, not universal facts:
For every speaking or acting subject, write a compact brief before prompting generation:
10. Continuity: where the eyes, hands, body, prop, emotional state, and voice start/end so adjacent shots match.
11. Safety/rights: consent status, likeness/voice rights, permitted use, disclosure/provenance requirement, and prohibited resemblance.
Keep the brief short enough to fit into the generation prompt or handoff notes. If a provider has weak control over a dimension, move that dimension into QA and iteration rather than pretending it is guaranteed.
Beat direction should answer:
Use playable action verbs instead of result adjectives:
| Weak direction | Stronger performance direction |
|---|---|
| "She is worried." | "She tries to sound calm while checking whether the viewer understands; her eyes hold steady, then flick down for half a beat before she recovers." |
| "Make him confident." | "He invites the viewer in as if sharing a shortcut he has tested; relaxed shoulders, direct lens contact, small smile on the first proof point." |
| "Very emotional." | "He starts controlled, swallows the next word, then lets one breath land before saying the final sentence more quietly." |
| "Friendly avatar." | "Warm onboarding host: open chest, hands quiet at waist, nod only at the user's likely concern, pace 10% slower on setup steps." |
For short generated clips, limit each shot to one or two beat turns. If the scene needs several tactical shifts, split it into multiple shots or takes.
Name the underlying emotion only after defining the action. Then specify screen intensity:
Add contradiction when needed: "voice stays professional, face reveals concern." This prevents flat delivery and avoids asking the model to make every channel express the same emotion.
For faces, prefer observable cues:
Avoid stacking too many expression cues. A face cannot convincingly do "wide smile, clenched jaw, whispered grief, aggressive stare, relaxed warmth" in one beat unless the contradiction is intentional and physically coherent.
Blocking is performance meaning in space. State where the performer is, how they orient to the viewer/partner/camera, and why they move.
Use these defaults:
Gesture rules:
Eye-line continuity matters. For multi-shot scenes, record who/what each character is looking at at the start and end of every shot. In AI video, inconsistent eye-line is often a continuity failure, not just a face failure.
Voice performance drives perceived acting in avatar and voice-led media. Direct:
For synthetic voices, write performance markup in the syntax the provider supports; if syntax is unknown, use plain text notes adjacent to the line. Verified 2026-07-10, ElevenLabs documents pacing as a major part of voice design and advises clear pacing language for voice outputs: ElevenLabs Voice Design.
For lip-sync/avatar generation:
For dubbing/localization, maintain meaning and natural speech while respecting original timing and visible mouth/gesture openings. If perfect internal mouth-shape match makes language unnatural, choose the sync type explicitly: lip-sync, simil-sync, voice-over, narration, or subtitle-supported dub.
Do not direct every medium like live action.
Avatar presenter:
Talking photo:
AI cinematic character:
Rigged 2D/3D animation:
Voice-only:
Before creating or directing a synthetic human, decide whether the performer is:
Require written scope for custom or recognizable likeness/voice work:
If consent is missing or ambiguous, redirect to a fictional/non-identifiable performer, licensed stock avatar, or human-recorded performance. Do not create a "soundalike" or "lookalike" workaround for a real person when the intent is recognizability.
Use prompt rehearsal before expensive final renders:
Iteration notes should be specific:
Review the output with the sound on, sound off, and audio-only.
Sound on:
Sound off:
Audio-only:
Continuity:
For each performer or character, hand off:
Performer: <name or asset id>
Consent/rights: <fictional/licensed/custom consent status; allowed uses; disclosure>
Casting read: <age/presence/vocal/body energy; prohibited resemblance>
Scene objective: <what they want>
Obstacle: <what resists them>
Relationship/listener: <who they believe they address>
Beat map:
00:00-00:03 <action, expression intensity, gaze/body/voice>
00:03-00:07 <beat turn and new tactic>
Voice notes: <pace, pauses, emphasis, breath, pronunciation>
Blocking/gesture: <position, movement, hands, props, eye-line>
Continuity in/out: <start and end states>
QA priorities: <highest-risk performance details>
Production intent: 25-second SaaS onboarding video for a warm but credible avatar presenter. Medium: script-driven avatar with direct-to-camera framing.
Performance brief:
Performer: licensed stock avatar, mid-30s, calm technical host.
Objective: reassure a new user that setup is quick and safe.
Obstacle: the user expects integration setup to be annoying.
Action: demystify, then invite.
Emotion scale: warmth 2, urgency 1, confidence 2.
Eye-line: lens as user.
Gesture density: low; one small open-hand gesture on "two minutes."
Voice: conversational, 145-155 wpm, slight smile in tone, pause after first sentence.
Consent/rights: stock avatar within platform license; disclose synthetic presenter if required by channel policy.
Script/performance direction:
[steady lens contact; relaxed shoulders]
"You do not need to rebuild your workflow."
[small pause; softer smile, warmth 2]
"Connect your calendar, choose the approval rule, and we will show you the first safe draft before anything goes live."
[one open-hand gesture on "first safe draft"; pace slows 8%]
"Most teams finish setup in about two minutes."
[return hands still; invite, not pressure]
"Let us do the first pass together."
Why it is structured this way: the avatar gets one objective, one obstacle, a restrained gesture plan, and voice timing that supports lip-sync and trust. Likely failures: over-smiling, random nodding, too-fast line two. Repair by reducing expression intensity, removing extra gesture requests, or splitting the long sentence.
Production intent: 8-second dramatic shot for a short film teaser. Medium: AI video, over-the-shoulder close-up, no celebrity likeness.
Prompt-ready direction:
Shot: close over Mira's shoulder toward Jonas at a rain-streaked bus shelter at night, 50mm, shallow depth of field. Fictional characters, no resemblance to real actors.
Given circumstances: Jonas has just been caught hiding the train ticket; Mira is waiting for an explanation off camera.
Jonas objective: persuade Mira not to leave.
Obstacle: he knows his excuse is weak and she has stopped trusting him.
Action: confess just enough to keep her there.
Beat 1 (00:00-00:03): Jonas holds Mira's eye-line slightly left of lens, jaw tight, breath held, face controlled at regret 2.
Beat turn (00:03): he sees she is about to walk away.
Beat 2 (00:03-00:08): his shoulders drop, eyes flick to the ticket then back to Mira; he says quietly, "I bought two." No smile. Voice low, slower on "two."
Blocking: Jonas remains still except the ticket hand lowers into frame; Mira is a blurred shoulder foreground, no speaking.
Continuity out: Jonas looking at Mira, ticket visible at chest height, rain and neon consistent.
QA: eye-line to Mira, hand artifact on ticket, no melodramatic crying, mouth closes after "two."
Why it is structured this way: the scene is directed through objective and obstacle, not "sad man." It gives the model one physical reveal and one emotional shift. Likely failures: ticket hand distortion, eyes to camera, overacted tears. Repair with a closer crop, clearer prop silhouette, or reduced emotion scale.
Production intent: 15-second cartoon character explains why password managers help. Medium: rigged 2D character with TTS and simple viseme lip-sync.
Direction:
Character: small raccoon office helper, curious and practical, no real-person likeness.
Objective: convince the viewer to stop reusing passwords.
Obstacle: the viewer thinks password managers are extra work.
Action: make the danger concrete, then offer relief.
Animation constraints: use three poses only: lean-in warning, count-on-fingers, relaxed thumbs-up. Keep paws below chin to avoid mouth occlusion.
Emotion scale: concern 3 for first beat, relief 3 for final beat; animation style allows exaggeration.
Voice: bright, quick but clear; pause after "same key."
Line map:
00:00-00:05 lean-in warning, eyes wide, paws still:
"Using one password everywhere is like using the same key for your house, car, and diary."
00:05-00:10 count-on-fingers, one paw gesture only:
"A password manager makes different keys and remembers them for you."
00:10-00:15 relaxed thumbs-up, smile 3, slower final phrase:
"You remember one master password. That is the deal."
Lip-sync QA: check open vowels in "everywhere" and final mouth closure on "deal"; if weak, slow the final sentence or regenerate visemes from final audio.
Why it is structured this way: the acting is simple enough for a rig, gestures do not fight lip-sync, and each pose has a persuasive job.
Take calesthio/performance-direction from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.