mcpbeat Sign in

Flux 3 Audio Dialogue Agent Skill

Use when directing FLUX 3 audio, dialogue, or voiceover. Covers ambience, effects, music, silence, and timing.

691 tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
106
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/black-forest-labs/skills --skill flux-3-audio-dialogue

What comes with it

430 bytes besides the instruction
metadata.json

The instruction itself

1 sections, as written by the author

FLUX 3 Audio and Dialogue

Name each layer separately: speech, voiceover, ambience, effects, music, or deliberate

silence. One blurred description gives up control of all of them; every sound needs a

physical source or narrative role.

Speech. Quote the exact line, name the visible speaker (or label the line

voiceover/narration so it is not searching for a mouth to belong to), and add

no on-screen text, no subtitles when text is unwanted:

A weather presenter on camera in front of a stylized storm map, speaking directly to
the lens: "Storm season is here, and this time, we're ready." Confident delivery,
clean studio lighting. No on-screen text, no subtitles.

Voice anchors: age range, accent when relevant, register, energy, recording distance.

Reusing the same direction preserves a kind of voice, not the same performer across

generations.

Speakability. Write for the clip's real duration: short sentences, one thought per

line, room before and after the payoff; spell unusual names phonetically; shorten the

line before speeding the delivery. A line that cannot finish comfortably needs a

shorter script or a longer clip.

Effects are causal, not a detached list:

As the cup hits the tile, it cracks with one sharp ceramic snap.

Mix. Say what leads and what stays under it. Keep background voices out when one

line matters:

Her line is foreground and fully intelligible. Café chatter and espresso hiss remain
low and diffuse. A restrained piano pulse enters beneath the final words without
masking them.

Silence and post. Set generate_audio: false for a deliberately silent source

clip. Reserve for deterministic post: final loudness, EQ, ducking, and fades;

guaranteed wording or speaker identity; frame-accurate sync; subtitles and captions;

continuity across separately generated clips.

When a take misses, change one dimension at a time: speaker ownership, line length,

delivery anchors, competing layers, action-to-effect causality, or generation versus

post.

How to use it

Copy the folder

Take black-forest-labs/flux-3-audio-dialogue from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.