osidemedia/higgsfield-audio
> Use when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance 1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok Imagine Video). Also use when the user's prompt would benefit from audio direction but they haven't mentioned it. Also use when the user wants standalone audio — a soundtrack, ambience bed, multi-speaker scene audio (Seed Audio 1.0), or text-to-speech voiceover.
npx skills add https://github.com/OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-audio
*Routing aids — read the linked sections for the full rules.*
@Audio1 is a conditioning INPUT — beat sync, the AUDIO: Xs] script block, and the first-15s extraction trap [→seed_audio, standalone) = whole-scene audio in ONE pass — multi-speaker dialogue + music + SFX + ambience mixed →seed_audio, qwen_audio_tts (NEW — Qwen 3.0 TTS Flash, expressive instructions + cloned voices), text2speech_v2 (5 engines incl. cozy_voice), plus 3 game-pipeline-only tools — distinct from in-video joint audio →| Model | Audio type | Dialogue | SFX | Ambient | BGM | Lip-sync |
|-------|-----------|----------|-----|---------|-----|----------|
| Kling 3.0 / Omni | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language |
| Seedance 2.0 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Multi-language |
| Seedance 1.5 Pro | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ Best lip-sync |
| Veo 3 / 3.1 | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ English best |
| Grok Imagine Video | Native joint | ✅ | ✅ | ✅ | ✅ | ✅ |
| All other models | ❌ | — | — | — | — | — |
"Native joint" means audio and video are generated simultaneously in one pass —
not layered on after. This produces natural synchronization without post-production.
Models without native audio: add audio in post with Lipsync Studio or external tools.
Every audio-capable prompt should consider four layers. You don't need all four
in every prompt, but knowing which to include gives the model clear direction.
Put dialogue in quotes. Be explicit about who speaks, their tone, and language.
She says: "We need to leave. Now."
He whispers: "Not yet."
Best practices:
She speaks in Cantonese: "走啦"Taiwanese Mandarin, Shanghainese), Japanese, Korean, Spanish, Indonesian
Describe SFX at the point they happen. Tie them to visible actions.
The glass shatters on the floor — sharp crack, then settling tinkle.
Footsteps on wet concrete — splashing, rhythmic.
A door slams shut — heavy metal, echoing.
Best practices:
Set the acoustic environment. This is the continuous sound bed.
Ambient: quiet café murmur, espresso machine, rain against windows.
Ambient: forest at night — crickets, distant owl, gentle wind through leaves.
Ambient: busy intersection — traffic, horns, construction in the distance.
Best practices:
Don't name songs or artists (content filter). Describe the musical texture.
BGM: slow piano, minor key, melancholic.
BGM: tense orchestral build — low strings, rising.
BGM: lo-fi hip-hop beat, warm vinyl crackle, relaxed.
Best practices:
Add audio cues naturally within your prompt or as a dedicated block at the end.
A woman walks into a quiet library. Her heels click on the marble floor — each step
echoing. She whispers to the librarian: "Do you have the Collected Letters?"
Distant page turns. A clock ticks somewhere above.
[Scene description — visual content, action, camera]
Audio:
Dialogue: She says "We leave at dawn." He replies: "I'll be ready."
SFX: coffee cup set down, chair scraping back
Ambient: early morning kitchen — birds outside, kettle just boiled
BGM: none — silence emphasizes the tension
Lip-sync is the most failure-prone audio feature. Follow these rules strictly:
> Expressive facial acting *around* the words — forced smiles, leaking fear,
> mixed emotions during a spoken line — is driven separately by FACS Action Unit
> codes per beat. Let lip-sync shape the phonemes; schedule the brow/eye/cheek
> AUs for the performance. See ../higgsfield-facs/SKILL.md § Dialogue &
> Monologue Facial Acting.
locked-off static camera or slow Dolly In onlynodding, turning head, looking aroundcompete with the lip engine and cause desync
the generative audio engine to override your dialogue
Multi-person lip-sync matching is an unresolved limitation across all models.
The production workaround:
Field-observed word budgets for reliable lip-sync in a ~15s in-video Seedance
dialogue clip — not official limits, and not the same as how many words the model
can *voice*. The acoustic budget ≠ reliable-sync budget: the model will happily
speak more words than it can keep synced to the mouth.
| Language | Reliable-sync budget (~15s clip) | Notes |
|----------|----------------------------------|-------|
| English | ~16–20 words (5–10 per line) | Strongest Western language |
| Mandarin | — | Strongest sync overall |
| Russian | ~10–15 words | Weak — budget conservatively |
| Japanese / Korean | Under-tested | No reliable field numbers yet |
Cross-language sizing unit: "one short sentence ≈ one breath." Write dialogue
in breath-sized sentences and count breaths, not seconds.
On surfaces that accept a spoken-voice reference, an attached **rights-cleared
voice recording drives lip-sync directly** — the model syncs the mouth to your
recording instead of synthesizing a voice first. This is the most reliable
field-reported path for non-English dialogue (it sidesteps the weak-language
sync budgets above). Rights-sensitive: only use recordings you have clear
rights to — cloned or scraped voices are out.
@Audio1)The most under-used Seedance 2.0 capability: an uploaded audio file is a
conditioning input, not just an output track. The model spec lists audio
as a reference media role alongside image / video, and generate_audio
(native sound output) is documented as *independent of the audio reference
medias* — i.e. the uploaded file conditions the generation, and whether the
clip also gets generated sound is a separate switch.
This means @Audio1 has two distinct jobs, and you pick one per shot:
| Use | What @Audio1 does | Prompt discipline |
|-----|---------------------|-------------------|
| Audio-as-output | Plays the uploaded track unmodified as the clip's soundtrack | Timestamp-anchor it (plays exactly as uploaded from 0s to end) and remove all ambient/SFX/music tokens so the engine doesn't override it (see § Seedance 2.0 below) |
| Audio-as-driver (beat sync) | Drives the visuals — cut timing, camera acceleration, action pace, energy peaks | Write the audio→visual mapping explicitly (below). The clip can still get generated sound, or set generate_audio false for visuals-only. |
> Why it works (author's model — empirical, not in the official spec): the
> temporal branch that reasons about motion and pacing reads the sound's
> structure — beat positions, dynamic contour, timbral texture, song-structure
> sections — and maps it to visual rhythm. Treat the mechanism as a working
> model; treat the capability (audio reference role) as confirmed.
Upload an MP3 as @Audio1, then map audio characteristics to visual elements.
The minimum is three sentences, each handling one thing — **rhythm source /
which visual responds / how energy maps to the arc**:
Use @Audio1 as the rhythmic foundation. Sync camera transitions to the beat
positions. Visual energy builds with the audio crescendo and peaks at the drop.
You can assign different visual elements to different audio characteristics —
mixing audio-to-visual the way you'd mix a track:
@Audio1 drives the visual rhythm. Camera cuts land on the downbeats. Subject
movement accelerates into the build, holds at the peak, releases on the drop.
Colour temperature shifts warmer with the crescendo.
Camera ← beat position. Movement ← dynamic contour. Colour ← overall energy arc.
It stacks with other references — character from @Image1, camera style from
@Video1, rhythm from @Audio1, processed together:
@Image1 as character reference. Follow @Video1 camera-movement style. @Audio1 as
rhythmic foundation — sync all camera transitions to the beat positions.
Character movement should pulse with the music.
The one constraint: @Video1 camera style and @Audio1 rhythm have to be
temporally compatible. A slow continuous dolly pulled from a video reference
fighting an EDM track sends the temporal branch conflicting instructions — same
failure class as mixing reference *images* of clashing styles. Pick references
that can coexist. (Sibling of ../higgsfield-seedance/SKILL.md § Reference Roles
→ Load-Bearing Rule: references stay in their lanes.)
[AUDIO: Xs] script block — dialogue + SFX + lip-sync from text aloneNo microphone, no recording. A timestamped script inside the prompt text
generates voices, SFX, and lip-sync. Quoted text → speech with automatic
lip-sync; physical descriptions → sound effects. Each marker is a timestamp in
the clip:
[AUDIO: 0s] heavy footsteps on concrete, echoing in a corridor
[AUDIO: 2s] door bursting open, impact bang
[AUDIO: 3s] character says "Nobody move"
[AUDIO: 5s] tense silence, distant traffic
[AUDIO: 7s] character says "Put it down. Slowly."
[AUDIO: 9s] object placed on table, soft thud
The model generates the voice first, then maps facial movement to the
waveform — so lip-sync quality is mostly set by how precisely you wrote the
dialogue. Exact quoted text outperforms paraphrase. It works across
languages (write the line in Spanish/Japanese/French → speech with
phoneme-level lip-sync in that language).
This obeys the same physical rules as § Lip-Sync Rules above: a strong @Image1
character reference gives a consistent mouth structure to animate, and **close-up
framing beats wide** (a small face has too few pixels to sync). Keep individual
dialogue beats inside the 3–8s accuracy window.
It combines with beat sync in one generation — uploaded music as the
rhythmic foundation, the script block as foreground dialogue/SFX, cuts synced to
the beat:
@Audio1 as background music. Sync camera transitions to the beats.
[AUDIO: 0s] music from @Audio1 begins
[AUDIO: 3s] character says "This changes everything"
[AUDIO: 5s] sharp breath — beat drop hits simultaneously
[AUDIO: 8s] character says "Let's go"
The audio reference limit is 15s, and the model takes the first 15s of
whatever you upload. Drop in a full 3-minute track and you almost always feed it
the intro — low energy, often ambient, no rhythmic drive. Nothing for the
temporal branch to map.
The right 15s follow a build → drop arc: rising tension into a peak. That
dynamic gradient is what becomes visual energy structure. A segment with uniform
energy gives the model beats to detect but no arc — output is rhythmically
synced but dramatically flat.
Where the window lives:
Extract exactly that segment before uploading. MP3 at ≥256kbps — lower
bitrate degrades beat detection. Don't upload the full track and hope; pick the
window, cut it, upload that. (Flipping the workflow — audio in first, visuals
built around it — changes the output at a structural level, not subtly.)
Audio Speaker Attribution Format (V3/O3):
[Speaker: Character Name] "dialogue" in a [warm/confident/excited] [male/female] voice with [accent].
Add [sound: footsteps / rain / door closing] when [action].
Background ambient: [environment description].
"Audio @Audio1 plays exactly as uploaded from 0s to end. Do not modify."Then remove all ambient/SFX/music tokens to prevent the generative engine from overriding.
@Audio1 is also a visual driver — beat sync, the [AUDIO: Xs] script block,and the first-15s extraction trap are all in § Audio as a Conditioning Input above.
> Diegetic-only convention for the prompt body — a
> prompt-authoring discipline that sits on top of Seedance 2.0's
> audio capability. BGM is a valid audio *layer* (see § The Four
> Audio Layers above) — that's what Seedance can generate. The
> diegetic-only convention is what you should *write* in the
> prompt body: only sounds that physically exist in the scene
> (footsteps on wet pavement, fabric whip on motion, breath, room
> tone, weather, weapon fire, crowd reaction, stage haze) rather
> than naming songs, lyrics, or score cues. If music is intended
> for the final cut, layer it in post rather than in the prompt
> body.
>
> Two reasons the discipline matters even though BGM is
> supported: (i) score descriptors ("dramatic strings",
> "orchestral swell") underdetermine the generated audio and
> routinely produce generic music beds at odds with the scene;
> (ii) the *timestamp-anchoring* + *remove-all-music-tokens*
> pattern in the bullets above already enforces this discipline
> when an MP3 audio reference is uploaded — the diegetic-only
> convention generalizes that pattern to all Seedance prompts
> whether or not an audio reference is attached.
"This must be it," he murmured.tires screeching loudly| Problem | Cause | Fix |
|---------|-------|-----|
| Lip-sync completely off | Audio > 8s, or head motion tokens present | Trim to 5s, remove nodding/turning tokens |
| Model replaces uploaded audio | Ambient/music tokens in prompt invite generative override | Add timestamp anchoring phrase, remove all ambient/music tokens |
| Dialogue missing entirely | Non-MP3 format used (Seedance 2.0) | Convert to MP3 128-320kbps |
| SFX drowns out dialogue | Too many SFX cues competing | Reduce to 1-2 SFX per shot, prioritize dialogue |
| Audio sounds robotic | Flat emotional cues | Add emotional direction: "says warmly", "whispers with urgency" |
| Background music too loud | BGM description too prominent in prompt | Move BGM to end of prompt, reduce detail, or say "subtle BGM" |
Not every prompt needs audio direction. Skip audio cues when:
> Negative constraints: For audio-specific artifacts (lip-sync desync, background music
> overriding dialogue, SFX drowning dialogue) and their prevention phrases, see
> ../shared/negative-constraints.md — Temporal/Consistency Artifacts section.
Cinema Studio 3.0 introduces native audio-video joint generation — a fundamental shift from models that treat audio as a post-processing step.
Audio is generated simultaneously with video via a unified multimodal architecture. This means:
Always describe audio as a separate section in your prompts. The generation engine handles three parallel audio tracks:
A chef slices vegetables rapidly on a wooden cutting board.
Camera: tight close-up tracking the knife.
Style: warm kitchen lighting, shallow depth of field.
Audio: rhythmic chopping on wood, oil sizzling in a nearby pan,
soft clinking of ceramic bowls. Light acoustic guitar BGM.
| Parameter | Limit |
|-----------|-------|
| Accepted formats | MP3, WAV |
| Max audio clips | 3 per generation |
| Combined duration | ≤15s total |
| Single file size | <15MB |
| MP3 bitrate | 128–320 kbps |
Available but experimental in Cinema Studio 3.0:
Control speaking style, accent, and language by referencing a video with the desired voice:
Voiceover tone references @Video1. The narrator describes the product
in a warm, conversational tone. "This changes everything."
Audio @Audio1 plays exactly as uploaded from 0s to end.
Do not modify or replace the audio content.
Dialects written directly in the prompt work — the model understands regional speech patterns. Write dialogue in the target dialect for authentic delivery.
When uploading reference audio that must play unmodified:
Audio @Audio1 plays exactly as uploaded from 0s to end.
Do not modify or replace the audio content.
Then remove all ambient/SFX/music tokens from the prompt to prevent the generation engine from overriding the uploaded audio with generated sound.
Describe specific foley, not generic moods:
Wrong: nice ambient sounds, pleasant background noise
Right: the scratch of frosted glass, rustling of plush fabric, gentle tapping on acrylic, popping of bubble wrap, wooden floor creaking under bare feet
Specific sound descriptions directly influence the generated audio output. The more precise the foley description, the more accurate the result.
Separate from in-video joint audio above: Seed Audio 1.0 (ByteDance, released
2026-06-23 at the FORCE conference) is a standalone **one-pass whole-scene audio
generator**. One generation produces multi-speaker dialogue + music + SFX +
ambience, already mixed — a radio-drama scene, not a single voice track. Use it
to build a soundtrack for footage you'll assemble in post, or scene audio that
has no video at all.
| You need | Use | Why |
|----------|-----|-----|
| A whole scene's soundtrack: several speakers + music + SFX + ambience, mixed in one pass | Seed Audio 1.0 (seed_audio) | One-pass scene audio; script-style prompt drives the whole mix |
| One clean voice track (narration, single-speaker VO) | text2speech_v2 (pick an engine) | Single-voice TTS — simpler, engine-selectable |
| Sound baked into the generated video, synced to on-screen action and lips | Seedance generate_audio (in-video) | Native joint generation — audio and visuals in the same pass (see § Audio as a Conditioning Input) |
Model id seed_audio (output_type audio). Parameters:
| Param | Range / options | Default |
|-------|-----------------|---------|
| format | wav / mp3 / pcm / ogg_opus | wav |
| sample_rate | 8000–48000 Hz | 24000 |
| speech_rate | −50..100 | 0 |
| loudness_rate | −50..100 | 0 |
| pitch_rate | −12..+12 | 0 |
| voice_type + voice_id | preset \| element — must travel together | none |
Media roles: image_references + audio_references. Per the fal schema: up to
3 reference audio clips (each ≤30s, ≤10MB) XOR one image reference —
image and audio refs cannot combine. Reference audio inputs in the prompt by
load order: @Audio1, @Audio2, @Audio3.
Everything in this subsection is community-converged practice, not spec — treat
as a starting point, not a guarantee. Write the prompt as a **radio-drama
script**:
[Scene: busy coffee shop, morning]Host (warm, upbeat): "…"[sound: espresso machine, soft jazz fades in]format: wav) when the audio is headed for postCompact worked example:
[Scene: rain-soaked night market, closing time]
Vendor (tired, warm): "Last skewers — half price, take them."
Girl (excited): "Two! No — three!"
[sound: rain drumming on tarp canopy, a scooter passing in the distance]
Vendor (chuckling): "Three it is. Careful, they're hot."
[sound: coins dropped on a metal tray, charcoal hiss]
Music: a lonely muted trumpet fades in under the rain, wistful but hopeful.
The live standalone-audio catalog, reconciled against the models_explore
snapshot of 2026-08-01 (../../specs/models_explore_snapshot_audio_2026-08-01.json;
generated table: ../../specs/AUDIO-MODEL-SPECS.md, machine twin
../../specs/audio-model-specs.json — regenerate with python3 scripts/sync_specs.py --type audio).
The Audio tab's UI tools — Voiceover (text → speech), Change Voice (swap a
voice in any video), Translation (translate speech in any video) — sit on top
of these models:
| Model id | Name | What it does | Availability |
|----------|------|--------------|--------------|
| seed_audio | Seed Audio 1.0 (ByteDance) | One-pass whole-scene audio: dialogue + music + SFX + ambience (§ above) | General |
| qwen_audio_tts | Qwen Audio 3.0 TTS Flash (Alibaba) | Expressive TTS: natural-language instruction for emotion/dialect/speed, preset or cloned reference-element voices, 13 language hints | General *(NEW 2026-08-01)* |
| text2speech_v2 | Text to Speech V2 | Single-voice TTS; engine via variant: elevenlabs, minimax, seed_speech, vibe_voice, cozy_voice *(NEW)*; preset or reference-element voices (voice_type + voice_id) | General |
| sonilo_music | Sonilo Music (FAL) | Text-to-music with controllable duration | Game pipeline only |
| mirelo_text_to_audio | Mirelo Text to Audio (FAL) | Text-to-audio SFX with controllable duration | Game pipeline only |
| inworld_text_to_speech | Inworld TTS (FAL) | Preset-voice TTS, ~110 voices across en/zh/ja/ko/es/fr/de/ru/… | Game pipeline only |
Engine picks within text2speech_v2: seed_speech when the deliverable is
multilingual voiceover/narration; elevenlabs (Eleven v3) when fine
emotional/tone control matters; vibe_voice for long-form narration. These
are standalone audio generators — distinct from the native joint audio baked
into Kling 3.0 / Seedance 2.0 / Veo during video generation. (Catalog reflects
the 2026-08-01 snapshot; verify live before quoting pricing or availability.)
A post-generation alternative to prompting audio at all: upload the **finished
clip** to Supercomputer and ask for an analyzed voice-over (e.g. "Analyze the
video and create a voiceover for it in the style of wildlife documentaries") —
the agent analyzes the footage, writes a script, and offers voices to pick from.
Shown working in Higgsfield's Seedance-4K tutorial; useful when the visuals are
already locked and only narration is missing.
higgsfield-seedance-vfx — Footage transforms whose payoff is a camera move synced to a spoken line (crash-zoom / push-in), or preserving the source talk track through a transform (SFX and source dialogue only); see ../higgsfield-seedance-vfx/references/dialogue-timing.mdhiggsfield-models — Which models support native audiohiggsfield-troubleshoot — Audio failure diagnosishiggsfield-cinema — Cinema Studio audio workflow with Kling 3.0higgsfield-vibe-motion — Motion graphics with audio (different from AI-generated audio)Take osidemedia/higgsfield-audio from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.