calesthio/elevenlabs-music
Generate and iterate music with ElevenLabs Eleven Music for production deliverables. Use when planning, prompting, API-calling, editing, inpainting, reviewing, or licensing AI-generated songs, instrumental beds, ad music, soundtrack cues, video-to-music scores, stems, or music assets using the ElevenLabs Music API or ElevenCreative Music product.
npx skills add https://github.com/calesthio/generative-media-skills --skill elevenlabs-music
Use this skill to turn a content brief into a usable music asset with ElevenLabs Eleven Music. Treat it as a production workflow: choose the right generation route, write legally safe musical direction, preserve artifact custody, and review the result against the edit before publishing.
Source-backed facts in this skill were verified on 2026-07-10 from official ElevenLabs documentation and terms. Re-check live docs before relying on pricing, limits, model defaults, licensing terms, or endpoint schemas in a paid or public production.
Do not submit prompts, lyrics, references, uploads, or metadata that name or clearly target a protected work or identifiable artist. ElevenLabs Music Terms prohibit Music inputs that include any artist real/stage name, songwriter real/stage name, song title, album title, music publisher name, label name, or substantial/distinct song lyric excerpt intended to reference a song. They also prohibit prompting likely infringement and misleading mimicry of an identifiable recording artist.
Use genre, era, region, instrumentation, production language, tempo, mood, arrangement, mix references, and functional intent instead:
For commercial release, check the active plan and Music Model-Specific Terms. As verified 2026-07-10, self-serve plans differ by generation/download limits, attribution, streaming rights, API features, concurrency, and permitted media. Self-serve media rights were documented as allowing online/offline commercial use except film, TV, radio, and Studio Games; Enterprise Music was documented as allowing all online/offline commercial use. Do not imply a client has broadcast, film, TV, Studio Games, reseller, or music-library/repository rights unless their plan/contract explicitly grants them.
Documented Eleven Music capabilities include text-to-music generation, optional vocals or instrumental output, multilingual lyrics, section/lyric editing, composition plans, inpainting with music_v2, video-to-music, upload for inpainting, and stem separation. API access is for paid users.
Key API facts verified 2026-07-10:
POST /v1/music.POST /v1/music/detailed, returning audio plus composition plan and song metadata.composition_plan are mutually exclusive in compose calls.music_length_ms with prompt generation accepts 3,000-600,000 ms.music_length_ms, while the overview and pricing pages also described a 5-minute generation/duration limit. Verify live account/API behavior before promising a cue longer than 5 minutes.prompt accepts up to 4,100 characters.model_id values documented for music include music_v1 and music_v2.music_v2 was documented as the UI default while API default remained music_v1 during a transition period. Set model_id explicitly instead of depending on defaults.output_format="auto" selects MP3 format according to the model: documented as mp3_44100_128 for v1 and mp3_48000_192 for v2.seed can improve consistency with the same parameters, but exact reproducibility is not guaranteed and outputs may change across system updates.force_instrumental=true guarantees an instrumental result with prompt generation and cannot be used with composition plans.For fast creative exploration, use a prompt with explicit duration and model_id="music_v2" unless a project requires v1 compatibility. Ask for 2-4 variants when using the web UI; with API, create separate generations and track seeds, prompt versions, and selected take.
For an edit-locked cue, use a composition plan. Plans give explicit chunk durations, section labels, lyrics placement, positive styles, negative styles, and context adherence. Use plans when the cue must hit voiceover gaps, chapter changes, trailer beats, looping regions, or a precise chorus/outro.
For scoring an existing cut, use video-to-music when the footage itself should guide the background music. Still provide a concise description and up to 10 style tags so the model understands brand tone and edit purpose. Combine videos beforehand when possible to reduce request duration and avoid ordering mistakes.
For revision after a good take, use music_v2 inpainting instead of regenerating the entire song. Generate with store_for_inpainting=true or upload an owned/non-infringing file, keep good slices as audio-reference chunks, and regenerate only weak sections.
For mix control, request or derive stems when you need to duck vocals under narration, remove drums during dialogue, create a social cutdown, or hand off to a sound mixer. Expect high latency on stem separation for longer files.
A strong Eleven Music prompt answers five questions in musical language:
ElevenLabs' best-practices docs state that high-level intent prompts can work, but production prompts should still include constraints that protect the edit: duration, vocal policy, density under dialogue, timing cues, and deliverable context.
Use "instrumental only" in the prompt and force_instrumental=true for prompt-based API calls when vocals would conflict with narration. If lyrics are acceptable, provide original lyrics and timing cues such as "lyrics begin at 15 seconds" or "instrumental only after 1:45." When lyrics are not provided, the model may generate structured lyrics that match the prompt and duration.
For isolated elements, use targeted phrasing: "solo piano," "solo nylon-string guitar," or "a cappella female vocal texture." For a full mix, specify that those elements should sit inside a full arrangement.
Use music_v2 chunk-style plans for most production work. A plan can contain up to 30 chunks. Generation chunks include:
text: section label in square brackets, original lyrics, or inline directions.duration_ms: 3,000-120,000 ms for each generation chunk.positive_styles: up to 50 style/direction terms.negative_styles: up to 50 terms to avoid.context_adherence: high, medium, or low.conditioning_ref and condition_strength for reference-conditioned regeneration.The first chunk strongly sets the overall tone. Give early chunks at least 6-7 concrete style terms until the direction is established. Use context_adherence="high" for seamless edits and medium or low only when a section should noticeably change.
Audio-reference chunks preserve slices of a stored song unchanged:
{
"song_id": "stored_song_id",
"range": { "start_ms": 0, "end_ms": 30000 }
}
Conditioning references can guide a generation chunk; the documented maximum conditioning reference length is 30 seconds. Use condition_strength low, medium, high, or xhigh depending on how tightly the new chunk should follow the stored source.
Start with a music brief:
Then run this loop:
Before handing off an ElevenLabs Music asset, verify:
force_instrumental=true for prompt mode, add "instrumental only," add "no vocals, no lyrics, no choir, no vocal chops," or use stems to remove/duck vocal content if the track is otherwise strong.music_v2, section durations are enforced in composition plans.Production intent: Create a paid-social music bed that supports a crisp voiceover and lands a product reveal at 24 seconds.
Route: POST /v1/music or SDK music.compose; prompt mode is enough because the structure is simple. Use force_instrumental=true.
Example prompt:
30-second instrumental-only music bed for a polished SaaS launch ad. Confident, modern, optimistic, and premium without sounding corporate-stale. Start with muted pulsing synth bass and soft tick percussion under voiceover. Add warm analog pads at 8 seconds, subtle handclap groove at 14 seconds, and a clean uplifting product-reveal lift at 24 seconds. 112-118 BPM, tight low end, bright but not harsh, no vocals, no lyrics, no vocal chops, no busy lead melody during narration, no artist imitation, no copyrighted references. End with a short resolved logo sting.
Example API parameters:
{
"model_id": "music_v2",
"music_length_ms": 30000,
"force_instrumental": true,
"output_format": "auto"
}
Expected review: The track should be sparse in the midrange until the reveal, feel polished at mobile-speaker volume, and have a usable final sting. If the model creates an intrusive hook, regenerate with stronger negative styles around lead melody and vocal-like synths.
Production intent: Score a 60-second cinematic product trailer with hard beats at 0:18, 0:38, and 0:55.
Route: composition plan with model_id="music_v2" so timing is explicit.
Example composition plan:
{
"chunks": [
{
"text": "[Cold Open]",
"duration_ms": 18000,
"positive_styles": [
"cinematic",
"minimal low strings",
"subtle clock-like percussion",
"tense",
"premium technology",
"80 BPM",
"dark spacious mix",
"slow riser"
],
"negative_styles": ["vocals", "lyrics", "rock drums", "comic", "bright pop"],
"context_adherence": "high"
},
{
"text": "[First Reveal]",
"duration_ms": 20000,
"positive_styles": [
"larger hybrid orchestra",
"deep braam accent at the start",
"pulsing synth bass",
"controlled intensity",
"wide percussion",
"rising brass swells"
],
"negative_styles": ["busy melody", "sung vocals", "cheerful"],
"context_adherence": "high"
},
{
"text": "[Final Build]",
"duration_ms": 17000,
"positive_styles": [
"accelerating percussion",
"heroic but restrained",
"layered strings",
"sub-bass pulses",
"big hit at final second"
],
"negative_styles": ["fade out too early", "lyrics", "soft ending"],
"context_adherence": "high"
},
{
"text": "[Logo Sting]",
"duration_ms": 5000,
"positive_styles": ["short resolved logo sting", "clean tail", "premium"],
"negative_styles": ["new melody", "vocals", "long unresolved reverb"],
"context_adherence": "high"
}
]
}
Expected review: Confirm the edit points hit, then inpaint only the weak chunk if the overall sonic identity works.
Production intent: Generate background music that follows a 75-second travel montage without manually timecoding every cut.
Route: video-to-music with the final locked reference video. Combine clips in order before uploading when possible.
Example request fields:
{
"description": "Instrumental background score for a warm, cinematic travel montage. Starts intimate and curious, grows into open-road optimism, then resolves gently for an end card. Leave room for natural location sound and occasional captions.",
"tags": ["cinematic", "warm", "instrumental", "travel", "optimistic", "acoustic", "soft percussion"],
"model_id": "music_v2",
"output_format": "auto"
}
Expected review: The score should follow broad visual energy but may not perfectly hit every cut. If exact hit points matter, move to a composition plan using the edit timecode map.
Production intent: Keep the first 45 seconds of an approved track, replace a weak ending with a cleaner CTA sting.
Route: generate or upload a stored song, then pass a mixed audio-reference/generation plan to music.compose with model_id="music_v2".
Example plan:
{
"chunks": [
{
"song_id": "stored_song_id",
"range": { "start_ms": 0, "end_ms": 45000 }
},
{
"text": "[CTA Ending]",
"duration_ms": 12000,
"positive_styles": [
"same premium electronic palette",
"clean rising transition",
"confident final logo sting",
"short reverb tail",
"instrumental"
],
"negative_styles": ["new genre", "vocals", "lyrics", "messy cymbals", "fade out"],
"context_adherence": "high"
}
]
}
Expected review: The splice should feel intentional. If the regenerated ending ignores the original palette, add a conditioning_ref from the last 10-20 seconds of the stored source and use condition_strength="high".
Official sources verified 2026-07-10:
Take calesthio/elevenlabs-music from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.