Localization and dubbing production
Use this skill to turn a source video into market-ready localized variants without treating translation as a word swap. Direct the localization strategy, script adaptation, voice approach, timing, accessibility, rights checks, review loop, and quality control before choosing any provider or model.
This is production guidance, not legal advice. When the work involves talent contracts, union-covered performers, biometric/voice data, likeness rights, political content, regulated claims, minors, health/finance/legal claims, or paid media in multiple jurisdictions, require qualified legal/local market review before release.
Evidence posture
Documented facts in this skill are grounded in the sources listed at the end and were verified on 2026-07-10 unless otherwise stated. Empirical observations are recurring production patterns seen in localization/dubbing work and should be checked against the actual source media. Production heuristics are decision rules for agents; override them when a client brief, platform specification, law, accessibility requirement, or native reviewer says otherwise.
Start by classifying the job
Identify the content type before making localization choices:
- Entertainment, documentary, testimonial, news, or creator content: preserve voice, identity, cultural specificity, and trust. Subtitles may be more honest than dubbing for testimony or journalism.
- Advertisement, product launch, app demo, social clip, avatar video, or explainer: optimize comprehension, retention, brand tone, and platform behavior. Dubbing or native re-voicing is often worth the cost.
- Training, compliance, medical, finance, legal, or public-sector media: prioritize accuracy, terminology control, accessibility, traceable review, and auditable approvals over fluency shortcuts.
- Children/family, literacy-limited, or sound-off mobile audiences: do not assume subtitles alone are sufficient; consider dubbing, captions, visual reinforcement, and simpler reading load.
- High-emotion drama, close-up presenter, avatar, or face-heavy ad: decide early whether lip-sync is required, because it changes the script, casting, edit, and QC burden.
Ask for or infer these essentials:
- Source language, target locales, audience, platform, release date, and region.
- Existing contracts/permissions for voice, likeness, music, on-screen people, and synthetic processing.
- Source assets: final video, clean dialogue/music/effects stems if available, transcript, captions, brand guide, approved product terms, campaign claims, legal disclaimers, and edit project or shot list.
- Delivery mode per locale: subtitles, SDH/captions, voice-over, phrase-sync dub, lip-sync dub, localized avatar performance, or fully re-edited market cut.
- Reviewers: native-language reviewer, brand/legal reviewer, subject-matter reviewer, accessibility reviewer, and final approver.
Separate facts, observations, and heuristics during planning
Documented facts to carry into production:
- WCAG 2.1 requires captions for prerecorded audio in synchronized media at Success Criterion 1.2.2 Level A and live captions at 1.2.4 Level AA. Audio description is separately required at 1.2.5 Level AA when visual information is not otherwise available.
- The FCC describes TV closed-caption quality around accuracy, synchronicity, completeness, and placement. Even when FCC rules do not apply to a given web/social deliverable, these dimensions are useful caption QC categories.
- Netflix timed text general requirements specify subtitle event duration minimum 5/6 second and maximum 7 seconds for Netflix deliveries; do not apply Netflix-specific limits blindly to non-Netflix platforms, but use them as a reality check for readability.
- ISO 17100:2015 defines requirements for core translation processes and resources for translation services; its public abstract states raw machine-translation output plus post-editing is outside its scope.
- YouTube multi-language audio lets eligible creators upload additional language audio tracks to one video or Short; uploaded tracks must be audio-only and roughly the same length as the video. YouTube automatic dubbing is separate and may contain errors in pronunciation, dialect, idioms, proper nouns, jargon, speech recognition, or voice matching.
- YouTube requires creators to disclose realistic AI-generated or meaningfully AI-altered content in Studio; YouTube says disclosure is not required for caption creation and lists cloning one's own voice for voice overs or dubs among examples creators do not need to disclose, but this is a platform policy, not a blanket rights clearance.
- WebVTT is a W3C text-track format for time-aligned captions/subtitles, descriptions, chapters, and metadata.
- NAVA defines active consent for synthetic/AI voices as an affirmative waiver/agreement by the performer or representative, excluding all other use. SAG-AFTRA public AI resources emphasize consent, disclosure, and digital replica protections in union contexts.
Empirical observations to verify on each project:
- Literal translations often break timing, humor, on-screen UI correspondence, mouth movement, legal nuance, and brand tone.
- AI dubbing systems are most likely to fail on overlapping speakers, noisy audio, regional accents, names, acronyms, jokes, sarcasm, code-switching, rapid speech, songs, and domain jargon.
- Viewers tolerate looser sync in narration, training, explainers, and off-camera voice-over than in close-up human performance, testimonials, drama, or avatar spokespeople.
- Some locales expect dubbing as a default for entertainment; others prefer subtitles for authenticity. Treat this as market research, not a universal rule.
Production heuristics:
- Localize intent first, wording second, timing third. A fluent line that misses the claim, joke, or call to action is wrong.
- Lock the source edit before localization unless the localization plan explicitly includes per-market recuts.
- Build a termbase before translating more than a sample. Fixing terminology after recording is expensive.
- Use native reviewers for market fitness, not just bilingual reviewers for dictionary accuracy.
- Prefer phrase-sync over lip-sync when the mouth is not central to trust or immersion; spend lip-sync budget only where the face sells the performance.
- Treat every synthetic voice, cloned voice, face reenactment, and avatar-lip-sync request as a consent-and-disclosure gate before production, not as a post-delivery paperwork issue.
Choose subtitles, captions, voice-over, phrase-sync, or lip-sync
Use the lightest mode that meets comprehension, trust, accessibility, and platform goals.
Subtitles are usually strongest when:
- The original performance, testimony, accent, music, or documentary authenticity matters.
- Budget or timeline cannot support native casting, direction, mix, and QC.
- The content is watched with sound on and reading load is reasonable.
- Legal or technical constraints make replacing the original voice risky.
- The source has many speakers but visible lip-sync is not required.
Captions/SDH are required when accessibility is in scope:
- Include dialogue, speaker identification where needed, and meaningful non-speech audio such as music cues, sound effects, offscreen voices, laughter, alarms, or tone-setting audio.
- Captions are not merely translated subtitles; they represent audio information for viewers who cannot hear it.
- For localized accessibility, create captions for the localized audio, not only translated subtitles from the source.
Voice-over is usually strongest when:
- The original speakers may remain audible under a translated narration bed.
- The content is documentary, news, training, webinar, interview, or educational, and face-perfect sync would feel artificial or unnecessary.
- Speed and clarity matter more than the illusion that the person is speaking the target language.
Phrase-sync dubbing is usually strongest when:
- The localized voice should begin and end near the source speaker's phrases, but exact mouth-shape matching is not needed.
- The source contains presenter shots, demos, explainers, corporate training, or ads with moderate face visibility.
- The script can be adapted to match phrase duration while preserving a natural target-language performance.
Lip-sync dubbing is usually strongest when:
- The audience sees close-up mouths, emotional performance, characters, avatars, or direct-to-camera talent.
- The localized version must create the illusion that the on-screen person is speaking the target language.
- The budget supports adaptation, casting, voice direction, retakes, alignment, and visual QC.
Avoid lip-sync when:
- The source is mostly off-camera narration, screen capture, product demo, or motion graphics.
- The target language expansion would force rushed, unnatural speech.
- Consent for face/voice manipulation is missing.
- The production cannot afford native linguistic and performance QC.
Build the localization brief
Produce or request a brief with:
- Source and target locales, not just languages: e.g. Spanish for Mexico, Spanish for Spain, French for Canada, Arabic MSA versus dialect, Portuguese Brazil versus Portugal.
- Audience: age, expertise, literacy, cultural context, accessibility needs, and likely viewing mode.
- Purpose: inform, persuade, train, entertain, convert, comply, support, or build trust.
- Brand voice: formal/informal, humor tolerance, pronoun/register policy, forbidden tones, approved slogans.
- Market constraints: legal claims, medical/financial disclaimers, regulated terminology, political sensitivity, religious/cultural restrictions, units/currency/date formats.
- Platform/delivery: YouTube multi-language audio, social burned-in captions, OTT timed-text package, LMS training module, broadcast, paid ad upload, in-app video, or sales enablement file.
- Sync target: no sync, time-sync, phrase-sync, lip-sync, avatar lip-sync, or re-edited native cut.
- Review path: who signs off terminology, claims, performance, accessibility, and final delivery.
Prepare source assets before translation
Do not start from an auto-transcript alone unless no better source exists. Create a source localization packet:
- Locked source video with timecode reference.
- Dialogue transcript with speaker IDs, timecodes, overlapping speech notes, offscreen/onscreen status, and inaudible flags.
- Intent notes per section: joke, warning, emotional beat, product claim, CTA, legal disclaimer, safety instruction, character relationship, or plot reveal.
- On-screen text, UI strings, graphics, lower thirds, captions, subtitles, title cards, end cards, supers, and legal slates.
- Pronunciation list for names, products, acronyms, places, invented terms, and domain terminology.
- Music/effects/dialogue stems when dubbing or re-mixing.
- Rights notes for talent, music, archive footage, user-generated clips, public figures, minors, and synthetic processing.
If the source is not locked, mark every downstream artifact as provisional. A changed source edit can invalidate subtitles, dubbing timing, localized graphics, and approvals.
Translate for purpose, then adapt for time
Use a two-pass script process:
- Meaning pass: translate all claims, instructions, story beats, emotional intent, terminology, names, numbers, safety/legal statements, and calls to action accurately for the target locale.
- Adaptation pass: reshape lines for reading speed, phrase duration, mouth visibility, performer breath, market idiom, brand voice, and on-screen context.
Maintain a line table for dubbed work:
| Field | Purpose |
| --- | --- |
| Source timecode in/out | Preserve timing and make retakes traceable. |
| Speaker | Support casting, subtitle labels, and mix decisions. |
| Source line | Keep reviewers anchored. |
| Literal meaning | Prevent adaptation from drifting away from intent. |
| Localized performance line | The actual recorded/synthesized line. |
| Sync constraint | No sync, time-sync, phrase-sync, lip-sync, off-camera, or can recut. |
| Pronunciation | Names, product terms, acronyms, regional terms. |
| Reviewer notes | Brand/legal/cultural/accessibility notes and status. |
For subtitles, preserve meaning and readability rather than every source word. For captions, preserve audio information. For dubbing, preserve performance intent and timing while remaining natural in the target language.
Build glossary and style controls
Create a project termbase before batch localization:
- Product names: translate, transliterate, or leave unchanged.
- UI strings: match shipped product UI exactly; do not invent local labels.
- Brand slogans and campaign lines: mark approved translations or "transcreate, do not literalize."
- Legal, safety, medical, financial, and compliance terms: require subject-matter approval.
- Units, currency, dates, phone numbers, addresses, measurements, honorifics, and formality/register.
- Competitor names, trademarks, cultural references, idioms, banned phrases, and sensitive words.
- Pronunciation: IPA, plain-language pronunciation, stress, and acceptable variants.
For multi-market campaigns, create a global glossary plus locale overrides. Do not force one Spanish/French/Arabic/etc. line across markets when usage, regulation, or audience expectation differs.
Cast and direct voices
Match communicative function before surface similarity:
- Role: narrator, expert, customer, founder, instructor, character, child, elder, support agent, avatar, announcer.
- Performance: warm, urgent, credible, playful, restrained, premium, conversational, technical, reassuring, comedic.
- Vocal attributes: range, pacing, energy, articulation, age impression where relevant, accent/dialect, breathiness, authority, texture.
- Cultural fit: native or near-native target locale performance, idiom comfort, correct register and pronunciation.
- Continuity: recurring brand voice, character continuity, multi-episode consistency.
Do not cast by stereotype. If a brief asks for a protected-class-coded voice without a production reason, reframe toward performance, locale, and audience comprehension.
For synthetic voices:
- Confirm the voice license allows the intended language, territory, media, duration, paid usage, client, derivative works, and future edits.
- Do not clone or imitate a real person, client employee, performer, celebrity, public figure, private individual, child, deceased person, or previous actor unless explicit, informed, written permission covers the exact use.
- Keep a consent record, source of voice, model/provider, approved use, expiration, revocation terms, and disclosure requirements.
- Avoid "sound like [living actor/celebrity/creator]" directions. Use production descriptors instead.
- Disclose AI use where platform policy, client policy, law, or audience trust requires it. When unsure, escalate for policy/legal review.
For phrase-sync:
- Align sentence starts, ends, pauses, and emotional beats to the source more than individual mouth shapes.
- Rewrite instead of speeding target audio into an unnatural delivery.
- Use edit points, reaction shots, B-roll, screen recordings, or graphics to hide unavoidable expansion.
- Keep breaths and pauses believable; audiences hear unnatural compression even when they cannot read the language.
For lip-sync:
- Identify "hero sync" shots: close-ups, front-facing mouths, high-emotion lines, product claims delivered to camera, and avatar monologues.
- Mark "low sync" shots: off-camera, profile, wide shot, masked mouth, fast montage, B-roll, graphics, or cutaway.
- Adapt lines to fit visible mouth closures, phrase length, and emotional beat; do not sacrifice legal/medical/safety accuracy.
- Prioritize visible plosives and mouth closures at line ends and obvious close-ups, but do not over-optimize phonetics at the expense of natural speech.
- Use retakes for performance and timing; do not rely only on time-stretching.
For avatar or AI face/lip generation:
- Confirm permission for the avatar, face likeness, voice, script, language, territory, and disclosure.
- Check that mouth motion, eye behavior, head movement, and prosody do not create uncanny or misleading effects.
- Keep translated line length within the avatar system's practical pacing; if the avatar rushes, rewrite or recut.
Handle cultural adaptation without erasing meaning
Classify each adaptation:
- Direct translation: acceptable when terminology, facts, or instructions carry cleanly.
- Localization: adjust units, currency, date formats, examples, idioms, register, UI labels, and market-specific references.
- Transcreation: rebuild a slogan, joke, CTA, rhyme, metaphor, or emotional appeal to achieve the same effect.
- Market cut: edit shots, claims, graphics, music, or examples for local regulation, cultural sensitivity, or platform behavior.
Do not change:
- Safety instructions, legal disclaimers, medical/financial claims, rights notices, source quotations, or testimony meaning without explicit approval.
- Speaker identity, consent context, or documentary truth.
- Product capabilities, price, availability, or guarantees unless approved for that market and date.
Subtitle and caption production rules
Before finalizing timed text, verify the target platform's current spec. Use these general rules when no stricter spec exists:
- Keep captions readable, synchronized, complete, and away from essential visual information.
- Use one or two lines where possible; segment at natural syntax breaks.
- Avoid flashing one-word captions unless the style intentionally requires kinetic captions and accessibility has been addressed.
- Do not cover faces, mouths, lower thirds, product UI, legal text, or calls to action.
- For social burned-in captions, design inside platform safe areas and preview on a phone; UI overlays change and vary by placement.
- For player-selectable captions/subtitles, provide sidecar files when the platform supports them. For social feeds where users may watch without caption settings, burned-in captions can be useful, but also preserve editable/source captions for accessibility and future localization.
- For SDH/captions, include speaker IDs and meaningful non-speech sound. For translation subtitles, avoid cluttering with non-speech cues unless the deliverable is explicitly SDH.
- When delivering WebVTT, SRT, TTML, SCC, or platform-native captions, validate timecodes, encoding, language tags, speaker labels, and line breaks before handoff.
Mix localized audio like a real soundtrack
For dubbing, voice-over, and synthetic audio:
- Match loudness, room tone, perspective, and emotional energy to the source. A perfect translation with mismatched acoustics feels pasted on.
- Keep music and effects under dialogue intelligible in each locale; music with lyrics may conflict with translated speech.
- Preserve meaningful source sound effects unless rights or market edits require replacement.
- When retaining low source dialogue under voice-over, keep it low enough not to compete with the target voice.
- QC with speakers and headphones; mobile speakers reveal harsh sibilance, clipped plosives, and buried consonants.
- Document any changed music, effects, or alternate mix stems for rights and client review.
Manage review without creating chaos
Use staged review:
- Strategy approval: mode per locale, rights assumptions, glossary ownership, schedule, QC standard.
- Sample approval: 30-90 seconds per target locale, including hardest lines, faces, product claims, and caption style.
- Script approval: translation/adaptation line table before recording or synthesis.
- Voice approval: casting/voice sample with pronunciation and tone notes.
- Rough dub/subtitle review: timing, performance, readability, and cultural fit.
- Final QC review: encoded deliverables, sidecars, metadata, rights/disclosure notes, and platform preview.
Force reviewers to label comments:
- Meaning error
- Terminology/glossary error
- Cultural/market issue
- Brand voice issue
- Legal/compliance issue
- Accessibility issue
- Timing/sync issue
- Performance/voice issue
- Audio mix issue
- Personal preference
Do not accept vague "make it better" notes. Ask for the locale, timecode, current line, proposed change, reason, and whether it changes approved terminology or claims.
QA checklist
Run QA per target locale and per output format.
Linguistic QA:
- Meaning, omissions, additions, numbers, names, product claims, disclaimers, units, dates, formality, register, idioms, and glossary compliance.
- Native fluency and audience appropriateness.
- No hallucinated features, prices, claims, locations, or endorsements.
Timing and sync QA:
- Subtitle/caption readability and cue timing.
- Dub phrase starts/ends, pauses, breath, line endings, and rush.
- Lip-sync on hero close-ups and avatar shots.
- No drift after edits, trims, or platform transcodes.
Audio QA:
- Loudness consistency, clipping, noise, room tone, sibilance, plosives, music/FX balance, and intelligibility on mobile speakers.
- Correct language track mapped to correct file, metadata, thumbnail, title, and caption track.
Visual/platform QA:
- Captions and localized graphics inside safe areas.
- On-screen text localized or intentionally left unchanged.
- Right-to-left layout, text expansion, fonts/glyphs, diacritics, line breaks, and encoding.
- Platform upload support for sidecar captions, multi-language audio, audio descriptions, localized thumbnails, and metadata as applicable.
Rights, safety, and disclosure QA:
- Voice/likeness/music/footage permissions cover every locale, platform, media buy, and synthetic transformation.
- AI/synthetic labels or client disclosures handled.
- No unauthorized clone, imitation, public figure misuse, or deceptive testimonial.
- Sensitive topics reviewed by qualified human reviewers.
Client handoff:
- Final video/audio files, caption/subtitle sidecars, transcript, localized script, glossary, pronunciation list, QC report, known limitations, disclosure notes, and reviewer approval record.
Failure patterns and repairs
- Dub sounds rushed: rewrite the target line, split it across a cutaway, reduce filler, or recut the shot. Do not simply speed up speech.
- Subtitle is unreadable: condense meaning, split cues at natural breaks, reduce on-screen competition, or move to a dub/voice-over mode for that segment.
- Voice sounds wrong: revisit casting specs and performance direction; do not ask for celebrity imitation.
- AI dub mispronounces terms: add pronunciation controls if the provider supports them, split/respell acronyms carefully, or record human pickup lines.
- Translation is accurate but culturally flat: shift from translation to transcreation for jokes, CTAs, metaphors, and campaign lines.
- Native reviewer conflicts with brand reviewer: use the glossary/style guide as the source of truth; escalate unresolved conflicts by timecode and business impact.
- Lip-sync fails on close-ups: rewrite for visible mouth beats, get retakes, switch to phrase-sync only with client approval, or cover with approved cutaways.
- Captions cover product UI: reposition captions, reframe, redesign lower thirds, or produce separate social-safe versions.
- Consent is unclear: stop synthetic voice/likeness work until permissions are documented.
Example: multi-market SaaS training rollout
Production intent: localize a 12-minute English onboarding video into Spanish (Mexico), French (Canada), German, and Japanese for an LMS.
Recommended approach: phrase-sync voice-over or native narration, not lip-sync. Preserve screen recordings and localize UI labels only if the product UI exists in that locale. Produce captions for each localized audio track plus a transcript for accessibility.
Workflow:
- Build glossary from shipped UI strings, product names, support terms, and compliance language.
- Translate and adapt script by timecoded segments; flag UI strings that do not exist in target product builds.
- Cast clear training voices by locale; prioritize comprehension over dramatic performance.
- Record or synthesize lines by segment, checking that each segment fits the screen action.
- Mix voices over original music/effects; lower music if consonants are masked.
- Deliver MP4 per locale for LMS, WebVTT/SRT captions per locale, localized transcript, glossary, and QC report.
Likely failure modes: translators invent UI strings; Japanese expansion/contraction affects screen timing; captions overlap UI controls; source edit changes after localization.
Example: face-heavy product ad for paid social
Production intent: localize a 30-second founder-led launch ad from English into Brazilian Portuguese and Spanish for Mexico, with close-up direct-to-camera shots and on-screen claims.
Recommended approach: native transcreated scripts plus lip-sync only for close-up hero lines; phrase-sync or cutaway coverage elsewhere. Validate claims and CTA in each market. Burn in social-safe captions for sound-off feeds and keep sidecar captions for accessible uploads where supported.
Production directions:
- Mark the first 3 seconds, product claim, and final CTA as hero-sync lines.
- Transcreate the hook and CTA; do not literalize if it weakens persuasion.
- Keep legal/disclaimer text approved by market reviewer.
- Cast voices for founder credibility and warmth, not imitation of the founder's private voice unless explicit consent covers it.
- Preview on a phone in the actual placement safe area before approval.
Likely failure modes: translated CTA too long for mouth movement; captions hidden by UI; voice resembles the founder without permission; market reviewer changes claims after recording.
Example: documentary testimony
Production intent: localize an interview-driven documentary clip featuring survivors describing personal experiences.
Recommended approach: subtitles or gentle voice-over, not full lip-sync dub, unless the commissioner has a strong editorial reason and participant consent. Preserve original voices enough to maintain authenticity. Include captions/SDH for accessibility.
Production directions:
- Preserve names, places, pauses, uncertainty, and emotional cadence.
- Do not smooth testimony into marketing copy.
- Use translator notes for idioms, trauma-sensitive language, and cultural context.
- Keep music low; do not let the localized voice flatten emotion.
- Require human native review and editorial approval.
Likely failure modes: over-adaptation changes meaning; AI voice makes testimony feel performed; captions omit emotional non-speech audio; reviewer tries to "improve" the speaker's words.
Sources verified on 2026-07-10
- Netflix Partner Help Center, "Timed Text Style Guide: General Requirements" - https://partnerhelp.netflixstudios.com/hc/en-us/articles/215758617-Timed-Text-Style-Guide-General-Requirements
- Netflix Partner Help Center, "English (USA) Timed Text Style Guide" - https://partnerhelp.netflixstudios.com/hc/en-us/articles/217350977-English-USA-Timed-Text-Style-Guide
- W3C, "Web Content Accessibility Guidelines (WCAG) 2.1" - https://www.w3.org/TR/WCAG21/
- W3C, "WebVTT: The Web Video Text Tracks Format" - https://www.w3.org/TR/webvtt1/
- W3C WAI, "Captions/Subtitles" - https://www.w3.org/WAI/media/av/captions/
- FCC, "Closed Captioning on Television" and "Closed Captioning of Video Programming on Television" - https://www.fcc.gov/consumers/guides/closed-captioning-television and https://www.fcc.gov/general/closed-captioning-video-programming-television
- Described and Captioned Media Program, "Captioning Key" - https://dcmp.org/captioningkey/print
- ISO, "ISO 17100:2015 Translation services - Requirements for translation services" - https://www.iso.org/standard/59149.html
- YouTube Help, "Add Multi-language features to your videos" - https://support.google.com/youtube/answer/13338784
- YouTube Help, "Use automatic dubbing" - https://support.google.com/youtube/answer/15569972
- YouTube Help, "Disclosing use of GenAI content" - https://support.google.com/youtube/answer/14328491
- YouTube Help, "Add subtitles & captions" - https://support.google.com/youtube/answer/2734796
- NAVA, "AI Voice Actor Resources" - https://navavoices.org/synth-ai-info/ai-voice-actor-resources/
- SAG-AFTRA, "Artificial Intelligence" - https://www.sagaftra.org/contracts-industry-resources/member-resources/artificial-intelligence