0xsline/talking-head-guide
| Guide for editing videos where the primary content is people talking — talking-head / 口播, interview / 访谈, lecture, tutorial, podcast, course content, and similar talking-driven formats. Use when the user wants speech editing on a talking video (剪口播 / 口播剪辑 / 去口癖 / clean up fillers / smooth speech), motion graphics layered onto talking video (口播加 MG / 加动画), or B-roll on a talking video (加 B-roll / add B-roll). For motion graphics specifically, use this together with the active Motion Graphics skill/workflow available in the current OpenChatCut environment — this skill adds talking-specific guidance (speech-rhythm timing, frame-aware placement, subject/caption protection, placement verification).
npx skills add https://github.com/0xsline/OpenChatCut --skill talking-head-guide
Required input: an existing talking-head / 口播 video uploaded to the project. If the user wants to start without one (e.g., generate a fresh talking-head from scratch), this skill doesn't apply.
When the user enters this workflow without a source video uploaded yet, ask via a widget surface — bundle the file upload with the treatment selection in one flow, not two separate turns or a markdown "drag your file in" instruction. Load widget-forms for the host-specific route. Never tell the user to "拖进编辑器" / "点击素材库的上传按钮"; that's friction with no upside.
When the task creates or targets a OpenChatCut project for the user, surface the editor link early so they can watch progress, and re-confirm the visible editor matches the project before final delivery.
Independent treatments that can be applied to talking-head videos. Pick the ones that match what the user wants — not all are needed every time.
voice-isolation skill.> 用户语言为中文时,在 widget options / choices options / 对话文案里严格使用上面括号里的产品术语——别自己再翻译一遍,会跟产品其它地方对不上。
Beyond picking treatments, a talking-head edit is shaped by several orthogonal variables. When the user's ask is vague, these are what's worth clarifying first:
When more than one of these variables is missing, ask with one form after loading widget-forms. Do not ask markdown numbered questions and then append <choices/> for only one part of the same intake.
When multiple treatments have been aligned with the user, they depend on each other and must be finalized in dependency order. This section is only relevant after alignment — it doesn't tell you what to start with on a fresh request.
The speech timing (set by A-roll editing) anchors everything downstream — MG placement, B-roll cut-covers, music duration, and caption sync all reference the final speech timeline.
So: finalize A-roll editing before committing any visual, audio, or text layer. Don't write captions against pre-edit speech, don't cut music to pre-edit length, don't place MG against timing that will shift.
You must confirm the result with the user after each major step before starting the next, unless the user has explicitly asked to run end-to-end without stopping. Key checkpoints when multiple treatments apply: after A-roll editing finalizes the speech timing; before MG generation (confirm style and direction, and, when it isn't obvious, whether it sits over the video as an overlay or takes the whole frame); after MG generation; same pattern for B-roll, music, and captions. Don't bundle multiple checkpoints into one response — confirm each step separately. An upstream mistake forces redoing everything downstream (e.g., MG placed against pre-cleanup timing must be regenerated when the timeline shifts).
In a talking-head workflow, the first step is usually A-roll editing: editing the original spoken footage.
A-roll edits are ultimately applied to the timeline and change what the viewer actually hears and sees. However, the editing decisions should usually start from the transcript, because the core question is: what spoken content should the viewer hear, and what should be removed, compressed, or reordered?
A-roll editing is not only cleanup. First decide what spoken-content task the user is asking for, then choose the editing strategy and tools.
Common tasks:
Cleanup is the most common task and the one most likely to fail from bad boundary decisions. It is described in detail below. Other tasks get shorter rules, but still follow the shared A-roll principles: complete meaning, clear boundaries, and natural listening flow.
These principles apply to all A-roll tasks, not only cleanup.
[sN] segment it sits in; keeping a whole segment for one requested sentence is over-keeping that drags in unrequested speech. This applies only when the task names what to keep — never to open-ended cleanup, where you keep complete units (above).[sN], [cN], [gap], word indices, clip ids, or segment ids. The user cannot see those addresses and will not understand what they mean. Use the actual spoken content, a short quote, or a plain-language description of the edit.Good cleanup does not mean making the video as short as possible, and it does not mean rewriting the speaker into a different script.
Good cleanup means:
Bad cleanup usually falls into two failure modes:
Default principle: remove defects without changing meaning; make speech smoother, not harder; prefer small local cuts over whole-sentence or whole-segment deletion; when unsure whether a cut harms meaning, keep it.
Below are the common cleanup categories and how to make editing decisions for each.
Fillers fall into two categories.
The first category is clearly meaningless hesitation sounds. These are usually safe to remove:
umuherah呃额When they do not carry special meaning, use clean_script first for bulk cleanup.
The second category depends on context and must not be removed by word list alone:
solike然后就是嗯啊那个那对所以但是How to decide:
Examples:
um, I think this solves the main problem -> remove um.It works like a checklist -> keep like; it is a comparison.The upload failed, so we retried it -> keep so; it carries cause/result.right after the call, send the recap -> keep right; it modifies timing.然后我们再看第二点 -> keep 然后; it marks sequence.A retake is when the speaker retries the same intended idea because they misspoke, got stuck, forgot words, or restarted. Retake cleanup is not "delete repeated text." The goal is to keep one complete, natural, logically coherent version of the intended idea.
Use this decision path:
Treat it as a retake only when multiple attempts are trying to say the same intended idea. Do not treat it as a normal retake when the repetition is intentional emphasis, a rhetorical beat, a structural marker, or a second pass that adds new information or tone.
A complete version may include more than the main content sentence. It may need a lead-in, connector, section marker, topic setup, contrast, qualifier, subject, object, or conclusion. These are not filler when the kept content depends on them.
Remove only words that are wrong, dangling, abandoned, or fully covered by the kept version. The cut boundary starts at the repeated or failed idea, not automatically at the earlier transition, setup, or continuous speech. If earlier speech contains useful context that the kept version does not repeat, keep it.
If several attempts are complete, usually prefer the later one because it is often closer to the speaker's intended take. But do not choose the last attempt mechanically. If the later attempt is missing needed context, structure, subject, object, or conclusion, keep the more complete version or preserve the missing lead-in from the earlier attempt.
A repeated lead-in is redundant only when another equivalent lead-in remains naturally connected to the kept content. If removing every copy makes the result lose structure or sound abrupt, keep one natural copy and remove only the extra restarts. Do not stitch unfinished fragments from different attempts into one artificial sentence.
Examples are patterns, not a closed list:
There, there's no After Effects, no Premiere, no DaVinci Resolve learning.
Keep the complete sentence, but remove the abandoned restart:
There's no After Effects, no Premiere, no DaVinci Resolve learning.
Do not keep the stray first word just because the full sentence is otherwise useful.
And secondly, ... and secondly, we're introducing a brand new UI.
Remove the extra restart, but keep one natural lead-in attached to the kept content:
And secondly, we're introducing a brand new UI.
Do not delete every structural marker and leave only:
We're introducing a brand new UI.
Then the next one is different from comedy. It is popular on Disney Plus. It is called...
Later retake:
It is a popular Disney Plus show called Love Story.
Keep useful setup that the later retake does not repeat, and cut from the failure point:
Then the next one is different from comedy. It is a popular Disney Plus show called Love Story.
Use false starts / unfinished fragments for this category. False start is the more natural editing/transcription term for a speaker beginning a phrase and then restarting or abandoning it; unfinished fragment makes the dangling half-sentence case explicit.
Only remove a fragment when it clearly does not form useful information.
Safe to remove:
Do not remove:
If only part of a sentence or segment is wrong, do not delete the useful content around it. Remove only the bad word, phrase, or pause; if a local cut cannot sound natural, keep the segment.
Pause cleanup should default to compression, not zeroing out. Spoken video needs natural breathing room.
Default rules:
How to operate on pauses:
clean_script. This is the default path for compressing many long pauses.clean_script pause rules:silence: "compress:300" (or the requested cap).silence: "restore:500" (or the requested minimum).silence: "normalize:500".silence: "range:300-800".Any rule that makes a pause longer — restore, normalize, or the lower bound in range — never invents new silence. It only recovers pause time that already existed at that exact spot in the original recording. If the original pause was shorter than the requested value, it stops at the original pause length.
read_script({ showSilence: true }) before batch pause cleanup. By default, timeline.md hides silence markers, but clean_script can still detect and rewrite silences internally.read_script({ showSilence: true }) only when you need to inspect or manually adjust a specific pause. Then edit the visible marker: ~~[silence=0.8s]~~ to fully cut it, [silence=0.8s→0.2s] to compress it, or leave it untouched to keep it.timeline.md. If the final pacing still has many long pauses, run clean_script only="silence"; if only one or two pauses feel wrong, use showSilence: true and adjust those manually.Script gap primitive note:
[gap] on the primary video track as a pacing pause. A Script [gap] means no source is playing; on the only visible video track it renders as black. If pacing needs breathing room, preserve or restore source silence with clean_script / [silence=...], cover the moment with B-roll/MG/a full-frame visual beat, or intentionally declare the black beat in the plan.Highlight extraction is not about making the content as short as possible. It is about selecting the most valuable spoken content according to the user's criteria.
Rules:
Restructure means changing the order of spoken content. It does not mean freely breaking sentences apart.
Rules:
Hook / short version work aims to make the opening more compelling or compress long content into a shorter but still complete version.
Rules:
Target-script / script alignment means cutting the final spoken content according to a user-provided script, target paragraph, or desired content.
Rules:
Highlight, short version, excerpt, hook, restructure, and making several versions are all transcript-content tasks: drive them through Script (read_script → edit timeline.md → apply_script), never by looking up timestamps and placing source clips manually.
timeline.md and apply_script. A version on its own timeline (the user asked for separate timelines, or wants each version independently editable/exportable): manage_timelines action=duplicate — the copy carries the content and its script, so you immediately read_script → trim → apply_script on it. Building fresh from library assets: manage_timelines action=create, add the source asset, then drive it through Script.library/<filename>.md, copy the needed [sN] line(s) into timeline.md where they belong, and apply_script. This is how you pull source content onto the timeline — through Script.[sN] segments in timeline.md in version order, one version after another, then apply_script once. Reuse is just repetition — the same [sN] segment may appear in more than one version, and repeating the line replays that source range again.find_transcript and place spoken content with edit_item / split_item. If you are converting transcript segments into source frame or second ranges, you are off the editing surface — return to Script. edit_item / find_transcript are only for non-transcript placement such as MG overlays and B-roll visual timing.Check each version against its request. After assembling a version, highlight, or excerpt, re-read the result end to end and confirm every requested sentence is present, in the requested order, with no extra source carried in. Fix any dropped, duplicated, or out-of-order content before finishing.
Use this flow for any A-roll task driven by transcript meaning.
read_script, then read timeline.md once to understand the user's goal, the content structure, and whether fixed fillers or long pauses are present. If you will run clean_script, do not build the full semantic edit from this pre-clean read.clean_script for fixed hesitation sounds (um, uh, er, ah, 呃, 额) and batch pause compression. If both are present, use the default clean_script pass so both are handled together. Do not use this step for context-dependent fillers, retakes, repeated sentences, or anything that needs meaning.clean_script, always read the refreshed clean timeline.md before semantic editing. Use this refreshed file as the source of truth; clean_script changes the canonical timeline and rematerializes the script, so previously read text may be stale. Do not edit from memory based on the pre-clean script. Then edit timeline.md with semantic judgment: choose the best retake, clean false starts, remove repeated or failed attempts, preserve useful setup and context, reorder content when needed, and keep the speech natural. For long transcripts, work one clear section at a time if that improves judgment accuracy.apply_script. If apply fails, fix the markdown error or stale state, re-read the current timeline.md if needed, and apply again.apply_script, read the regenerated clean timeline.md and check what the viewer will actually hear: broken logic, missing context, over-deletion, missed cleanup, wrong order, or pauses that feel too tight or too long. Fix clear problems only. If the final result still needs batch pause adjustment, use clean_script only="silence". Use read_script({ showSilence: true }) only for manual adjustment of specific pauses.Editing timeline.md is not just changing displayed text. It describes which source media ranges should play on the timeline.
[sN] rows are ASR segments, not semantic units. A complete sentence, idea, retake, or transition may span several [sN] rows, and one [sN] row may contain only part of a sentence. Before deciding what to delete or keep, mentally reconstruct the complete spoken sentence or idea across adjacent rows.
~~...~~ removes the corresponding audible audio range.apply_script applies the result back to the timeline.Choose the editing goal and content boundaries first, then choose the tool. Do not let tool availability change the editing strategy.
clean_script: use for mechanical first-pass cleanup: bulk removal of fixed meaningless fillers and batch silence compression/adjustment. It can process silence even when timeline.md is currently rendered without silence markers. Do not use it for context-dependent fillers, retakes, repeated sentences, or semantic decisions.read_script + apply_script: the main transcript-based editing surface. Use it for real semantic editing: deleting words, sentences, pauses, reordering, or pulling library content onto the timeline.manage_transcript action fix: only fixes ASR mistakes or speaker attribution. It does not cut audio and does not change what the viewer hears.display_text forcePageBreak:true on the word that should START the new card; to MERGE a card up into the previous one, set display_text keepWithPrevious:true on that card's FIRST word (works for any break — no box resizing, no wordsPerPage fiddling). To drop a repeated/false-start word, use display_text hidden:true. Box width / fontSize / wordsPerPage are style & density knobs, NOT per-boundary segmentation levers — do not widen the box or raise wordsPerPage to merge or split a specific card. NEVER edit the transcript to fix a caption line break — manage_transcript fix is only for an ASR-misheard WORD (content), not layout. read_captions shows each page's break= reason and per-word keys for these edits.find_transcript: only locates when a phrase is spoken. It does not edit. If the next step is cutting spoken content, return to Script.Edit / Write: use these to modify timeline.md. The edit only reaches the timeline after apply_script.Script details to preserve:
read_script materializes timeline.md (current cut) and library/<filename>.md (full read-only source transcripts) in the workspace.[s1] 过去~~呢~~一个月.clean_script for batch pause cleanup. Use read_script({ showSilence: true }) only to expose [silence=Ns] markers for precise manual edits such as ~~[silence=0.8s]~~ or [silence=0.8s→0.2s].find_transcript can locate a phrase for visual timing; it is not the editing surface. Do not use find_transcript + split_item / edit_item to cut, place, or assemble transcript-based clips — this includes highlights, hooks, excerpts, and multi-version cuts. All spoken-content selection, placement, and reuse happens in Script (read_script → edit timeline.md → apply_script).Motion graphics layered into A-roll reinforce what the speaker is conveying — deepening the audience's impression of the key points and helping them grasp content that's hard to land through speech alone. Complete A-roll editing first; MG timing is based on the post-edit timeline.
This section only adds talking-head timing, frame-composition, subject/caption protection, and review constraints. For visual style alignment, MG creation or authoring, implementation constraints, editable properties, asset sizing, and verification, use the active Motion Graphics skill/workflow available in the current OpenChatCut environment.
For talking-head MG work, treat the video as one edited piece, not as isolated graphics.
Design Style is the video's confirmed visual language. It gives MGs a shared tone, color logic, typography logic, visual density, and motion language. It keeps different MGs in one family without forcing them into the same shape. It does not decide which MGs are useful, when they appear, where they sit, or whether they are transparent / opaque; those remain per-MG editing decisions.
Resolve the visual language before planning MG moments. Use the active MG workflow for the actual style-alignment interaction and implementation details:
manage_design_style action="get" before planning MG moments.Picker is a visual Design Style selector. It shows preset thumbnails so the user can choose a visual direction by sight, instead of describing style in words.
manage_design_style with action: "list". The catalog returns presetId, name, and a style summary (no scenario filter or thumbnails in this build); shortlist reasonable options by name/summary and the actual video context.manage_design_style with action: "apply" and the selected presetId, then inspect the applied Design Style with action: "get" before authoring.Persist only confirmed visual language:
manage_design_style action="apply".After applying a preset or confirming a custom direction as the project style, tell the user in one or two natural sentences that this is now the video's visual style, future MGs in this video will follow it by default, and it can be changed or adjusted later.
MG meaningfully helps comprehension or orientation when the content has:
One video should usually have one visual language, but not one universal MG shape.
Reuse a Motion Graphic asset only for intentionally recurring instances of the same component: same viewer task, same information structure, same visual form, and content changed through properties. Repeated chapter markers, recurring section labels, or a repeated status badge can share one asset. Different jobs such as an opening title, chapter marker, quote, list, diagram, and CTA should usually be separate assets that share palette, typography, motion tone, spacing, and material treatment.
An accepted first MG proves the visual language works in frame. It is not automatically a template for unrelated MGs.
For talking-head videos, do not start MG creation from transcript timing alone. Inspect the target frame first: transcript tells you what and when; the frame tells you form, placement, and background.
Before creating the MG, make four linked editor decisions. They prepare the active MG workflow and the later timeline placement.
| Decision | Question | Output |
| ---------------------- | --------------------------------------------------------------- | --------------------------------------------------------------- |
| Content | What idea deserves a visual layer? | Message or visual fact expressed by the MG. |
| Timing | When should it land with the speech? | Timeline start, duration, read time, and internal motion beats. |
| Form and placement | What kind of MG is it, and where can it live safely? | MG form / size, then timeline placement after asset creation. |
| Background | Is this an overlay on the talking-head shot, or its own moment? | Transparent overlay or opaque / full-screen beat. |
Use this subsection only when the active Motion Graphics workflow explicitly asks you to write a generation brief or request for another model or generator, such as Gemini / motion-graphic-gen.
Skip this subsection for direct-authoring workflows. If you are creating or editing JSX yourself with create_motion_graphic_from_code / edit_asset, do not use referenceAssetIds, :template, :style, role anchors, or Gemini brief language.
For generator workflows, carry the visual language into the tool call. For now, templates are generation references, not direct-apply targets. Template refs from a Design Style are no different from any other template ID. For new MG assets, pass one code reference source: same-role role anchor with referenceAssetIds: ["<roleAnchorAssetId>:template"] only when the visual job, structure, and canvas role are the same; otherwise use the matched template ID directly, for example referenceAssetIds: ["<templateId>:style"]. If no template matches, write the confirmed Direction in the brief and use any accepted role anchor only for the same role. Template slot counts are not user constraints: if the user asks for more/fewer bars, rows, items, or data points than the template shows, generate a new structure instead of asking them to fit the slots.
When a template or role anchor is passed, keep the Gemini brief focused on content, role / broad form, background, and frame constraints. Let the reference carry detailed style and motion language.
Map the four shared decisions into a generator brief like this:
Content in the brief.Timing only when the MG has its own beats. Internal Timing values say _when_ each element appears, not _how_ it moves; leave the motion style to Gemini.Size & shape in the brief, not final canvas placement. Do not write final left, top, right, bottom, coordinates, or placement anchors such as "lower-left" / "top-right" into the Gemini brief.Background: transparent or Background: opaque.Choose what the MG expresses, not just what text it repeats. The content may be a speaker identity, distilled quote, key term, statistic, list, comparison, relationship diagram, chapter marker, or another visual representation of the point.
Choose the timeline anchor first. The MG should land with the relevant speech beat or section boundary, not trail after the speaker has already made the point. Use find_transcript; pass includeWordTimestamps: true when the MG has internal rhythm such as list items appearing one by one or multi-step reveals.
Write internal timing values relative to the MG's own start time. The timeline item start is the absolute video position; internal timing is the MG-internal rhythm after that start. Exit when the point is fully made.
Choose the MG form and likely placement region before creating the asset. The active MG workflow creates the graphic; place the finished asset on the video canvas afterward.
Placement principles:
Common forms and areas:
| Content type | Common form | Common area |
| ---------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| Identity / context | Name tag or small context label | Lower-third first; lower-left or lower-right depending on the shot. |
| Key information / quotes | Typographic quote, pull quote, or emphasis treatment | Lower-center / lower-third; side area if the bottom is crowded; full-screen for a major punchline, conclusion, or pause. |
| Structured information | List, step stack, comparison layout, or compact diagram | Left/right side areas or bottom horizontal area; full-screen if the information is too dense for an overlay. |
| Chapter / topic markers | Full-frame title, title overlay, or side title panel | Full-screen for a strong intro or section break; lower-third for a light cue; side panel when one side has obvious open space. |
| Abstract concepts | Concept visual, relationship map, cycle, framework, chart | Lower-third if light and readable above captions; side area or full-screen if denser. |
| Tiny auxiliary labels | Badge, status label, logo-like mark, section marker | Top corners can work here only. Do not use top-left/top-right as the default home for primary opening titles or chapter titles. |
Use the MG's intrinsic form constraints, not final canvas placement, when deciding asset shape. Good examples: "lower-third-style name tag", "compact side treatment", "bottom horizontal strip", "full-screen title beat". Do not bake final canvas coordinates into the asset unless the MG is intentionally full-frame.
For familiar forms like speaker name tags, give the form and content without forcing dimensions early. For constrained overlays, describe the intended rough form or usable area. For full-screen MGs, make the form explicit in Size & shape and choose Background: opaque.
From the target screenshot, include canvas tone only when it affects legibility: for example, Other context: dark interior scene — keep the design bright/light enough to read clearly.
Choose background from the form:
Background: transparent for talking-head overlays: lower-thirds, side treatments, quote treatments, compact diagrams, and other graphics that sit over A-roll. A transparent root may still contain internal semi-transparent or solid panels.Background: opaque when the MG is its own visual surface: full-screen opening titles, strong chapter beats, full-screen information layouts, and full-screen emphasis moments.solid item underneath as a color matte. The MG owns the frame; change its bgColor / transparentBackground properties instead. Do not create temporary solid fallbacks; if you encounter an old transparent-MG-plus-solid fallback while replacing it with an opaque generated MG, delete both fallback pieces, not only the old MG.Default to a transparent overlay unless a full-screen beat is intended — guessing full-screen/opaque silently is what covers the speaker's face or blanks the frame.
edit_item (adds/updates). Prefer an explicit rectangle once you know the frame: left/top/width/height for direct placement, or right/bottom/width/height when right/bottom margins are clearer.left — explicit x position. right — margin from the canvas right edge. Do not pass both.top — explicit y position. bottom — margin from the canvas bottom edge, symmetric with right, e.g. { right: 80, bottom: 150, width: 500, height: 350 } for a bottom-right overlay. Caption-safe defaults: bottom: 162 (landscape 1080p) or bottom: 576 (portrait 1080×1920). Do not pass both top and bottom.width / height should tightly bound the local visible composition, not the project canvas. Place and scale that local asset on the timeline. Use timeline-sized assets only when the visible design intentionally spans the whole frame.track_progress / project state are practical aids for resizing and placement, not the final judge.Enrich visual layers and cover jump cuts left by A-roll editing.
B-roll depends on having suitable footage and adds production effort — treat it as an optional enhancement, not a default. Apply only when the user opts in or there's a clear visual problem to solve.
Footage can come from three places: clips already in the project library, stock via search_stock_media then push_asset with the returned import args, or AI generation via the video-gen skill (Seedance / Kling / Hailuo when configured). Pick based on the user's need; if unclear, align with the user upfront.
Don't cut away in the first or last 3 seconds. For dense jump cuts (<3s apart), use one long cutaway covering multiple. Don't overlap with MG by default.
First decide the B-roll mode:
If the user only says "add B-roll" and the mode is not implied by the existing edit, ask once: "Should these be full-screen cutaways or small rounded-corner PiP overlays?"
For PiP / small-window overlay:
borderRadius to 24-36 by default unless the requested style is square/sharp. Do not add a mask/effect solely for ordinary rounded PiP corners; use effects only for special shapes or item types that cannot use native borderRadius.For full-screen cutaway:
fit:"cover" first-pass so the B-roll owns the visual beat.view_asset_frames before choosing fit if you have not already viewed it. Use read_script to choose representative video moments and sample more frames with view_asset_frames when the protected region is not obvious.fit:"contain" for the foreground and add a deliberate full-screen background such as an opaque MG background/matte that matches the edit or a blurred/enlarged duplicate/background layer. If a cover attempt only trims a compact subject/action that can be recovered without hiding other protected information, try a safer reframe/crop that moves the source protection frame fully into the canvas and closer to the intended center of attention.fit:"cover" without checking each source's aspect ratio and protected content first.After editing, read back the exact item ids you changed. An asset appearing in the library is not proof it is on the timeline; read_project must show the new/updated B-roll items. If the result involved crop, fit, scale, overlay placement, or a full-screen composition trade-off, verify the affected frame with a screenshot or visual analysis before reporting success, then fix failed source/destination protection or state the unavoidable trade-off. Do not report success if the target items are unchanged.
When the user has two or more cameras recording the same moment — cues like "both angles", "the same interview", "multi angles", "alternate angle", "cut to the other angle", "angle switch", 换角度, 两个机位 — switching to another angle means the picture changes but the audio and lip-sync must stay matched to the take.
Do not hand-compute source offsets with edit_item to line angles up. Manual offsets drift wherever the underlying reference angle was cut, and the drift only shows up later as out-of-sync lips. Use the multicam_sync tool instead: it runs the editor's audio-based alignment engine and repositions each angle clip so its picture matches the reference angle's audio. Pass the angle clips' itemIds (the reference plus the follower angle(s)); optionally name the referenceItemId.
Key constraint: a single cutaway clip that spans a cut in the reference angle can't be aligned as one piece — split it at that cut with split_item first, then pass both pieces to multicam_sync so each maps to the reference segment beneath it.
multicam_sync runs in the user's editor (no backend path): if it reports the editor isn't open, ask the user to open the project, then retry. After it applies, read the project back to confirm the alignment.
A track's role is the single declaration that drives the audio mix. Set it with edit_track and the engine derives a seamless duck — followers dip under speech, then rise back in the gaps — without you hand-adjusting any volume. There are only two roles, plus off:
role to anchor. This is the track everything else ducks under — set it, or nothing ducks.role: follower (auto-ducks under every anchor).role unset (none).A track with no role behaves exactly as today — roles are additive and safe, so you only set them where the content makes the job obvious.
Read the existing layout first. Before creating tracks or placing new clips, read the current track names and roles — if a track is already tagged for this content (a follower named "Music", an anchor named "VO"), put the new clip there and match its role; only make a new track when nothing fits. Organize before you assign — roles are per-track, so aim for one role per track. If the same kind of content is scattered across several tracks (e.g. the voice on A1 _and_ A3), consolidate it onto one track _first_: move the clips with edit_item (updates[].trackId), then delete the emptied track by id with edit_track. Then assign the role once. While you're laying tracks out, stack them the way a mixer reads a session — voice/VO on A1, the top audio lane; music below it — and give each a short name like "VO" or "Music" so the spoken word stays easy to find. A sensible default, not a rule; follow the user's intent when the layout should differ. Keep deliberate separation, though — two _different_ speakers, clips that overlap in time, or intentional layering each stay on their own track (and each still gets its role). After assigning, read the project back to confirm every track that should anchor/duck does, and that you left the music's base volume alone.
Set the mood and smooth over micro-gaps in speech.
role to follower with edit_track (and the talking track's role to anchor). That single pair turns on auto-ducking — the engine dips the music under speech and lifts it back in the gaps.edit_track initialize audioRouting.duckDepthDb from the current timeline loudness when it can. Pass audioRouting.duckDepthDb yourself only when the user explicitly wants the music louder or softer under speech.decibelAdjustment natural by default. Do not pre-duck music with a large negative clip gain, then also set manual duckDepthDb; only do both when the user explicitly asks for a lower overall bed and a stronger / weaker speech duck.audioFadeIn / audioFadeOut in seconds, usually 1-2 seconds. Do not pass frame counts to these fields.Fit BGM to the final video extent after A-roll timing is finalized. The target duration runs from the BGM start to the real content end (video / visual / speech items), excluding the BGM itself so music never extends the render.
audio item at the BGM start, set its duration to the target duration, and add a fade out. Do not let the full music asset run past the last visual item.audio items until the target is covered.Ducking is automatic once the music track's role is follower: the engine dips the track under audible anchor tracks (the speech / voice), and lifts it back to full level in pauses and the outro. This needs both halves — the music track has role: follower and the talking track has role: anchor (see Track roles above). If nothing is set to anchor, nothing ducks and the music stays at full level.
Set music to a normal, audible base level where there is no speech. To tune the dip under speech, update the follower track's audioRouting.duckDepthDb; otherwise leave it unset and let edit_track auto-initialize from timeline loudness when available. Do not solve speech clarity by heavily lowering the clip and also manually deepening the duck. To keep a track out of ducking entirely (for example a stinger that should punch through), leave its role unset (none).
Improve accessibility and engagement with on-screen text.
Captions are transcribed from the source audio — they always match the speaker's language. Translation between languages is not supported. Don't ask the user what language they want captions in; pick the preset by source audio language + target aspect ratio.
Take 0xsline/talking-head-guide from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.