calesthio/stable-audio
Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs, and reviewing generated audio for media projects.
npx skills add https://github.com/calesthio/generative-media-skills --skill stable-audio
Stable Audio is Stability AI's generative-audio family for music and sound-design outputs, not speech, voice cloning, dialogue, or final mix/mastering. Use it when a project needs instrumental music, loops, stems/ideas, sound effects, foley-like assets, ambience, stingers, transitions, or style/continuation edits from rights-cleared audio. Route speech, dubbing, voiceover, transcript alignment, and loudness mastering to the appropriate speech/audio-post tools instead.
All volatile facts in this guide were verified on 2026-07-10 from Stability AI official pages, the Stability API OpenAPI specification, official model cards/repositories, and Stability policy pages.
Documented facts:
/v2beta/audio/stable-audio-2/..., with model choices stable-audio-2 and stable-audio-2.5./v2beta/audio/stable-audio/..., with model=stable-audio-3.mp3 or wav output format.GET /v2beta/audio/results/{id} until HTTP 200 returns the audio.mp3 or wav, and requests are capped at 50 MB.mp3 or wav, and requests are capped at 100 MB.seed accepts 0 or omitted for random generation; explicit seeds may be recorded for reproducibility.steps differs by model: Stable Audio 2 accepts 30-100 steps and defaults to 50; Stable Audio 2.5 and Stable Audio 3 accept 4-8 steps and default to 8.cfg_scale ranges 1-25. API docs describe it as prompt-adherence strength; defaults are 7 for Stable Audio 2, 1 for Stable Audio 2.5, and 1 for Stable Audio 3. Treat it as a secondary control after prompt, duration, model, steps, and source-audio strategy.strength in audio-to-audio ranges 0-1 and controls source-audio influence. Stability describes 0 as identical to the input and 1 as equivalent to no audio influence.mask_start and mask_end in seconds to choose the replacement/continuation segment.credits = 17 + 0.06 * steps per successful generation, usually 20 credits at the default 50 steps and 23 credits at 100 steps; Stable Audio 2.5 is 20 credits per successful result; Stable Audio 3.0 is 26 credits per successful result. Stability states failed generations are not charged. Stability pricing uses 1 credit = $0.01 and is subject to change.small-music, small-sfx, and medium; the official repository lists Small music/SFX as CPU-capable 433M-parameter models with 120s max length and Medium as a 1.4B CUDA model with 380s max length. Stable Audio 3 Large is API-only in the official repo.Sources: Stability API reference / OpenAPI spec, Stable Audio 3 page, Stable Audio 2.0 announcement, Stable Audio 2.5 announcement, Stable Audio 3 repository, Stable Audio Open 1.0 model card, Stable Audio Open paper, Stable Audio 3 paper, Stability pricing page.
Use hosted Stable Audio 3 when the project needs current highest-capability hosted Stable Audio, up to ~6 minute beds, async generation is acceptable, or inpainting/continuation is part of the plan. It is the cleanest choice for longer music beds, adaptive game/music cues, sonic-branding variations, and complex ambience where 190 seconds is not enough.
Use hosted Stable Audio 2.5 when you need synchronous API behavior, 190 seconds is enough, and the brief is commercial music/sound production where Stable Audio 2.5's documented improvements in speed, musical structure, mood/genre prompt adherence, and inpainting matter.
Use hosted Stable Audio 2 only when a pipeline already depends on its behavior or when you intentionally need its 30-100 step range and cfg_scale behavior. Do not default to it just because the endpoint default is stable-audio-2; explicitly set the model.
Use Stable Audio 3 open weights when local or self-hosted generation, fine-tuning/LoRA experimentation, data-control review, or offline iteration matters. Choose:
small-music for quick CPU music drafts up to 120s;small-sfx for quick CPU sound-effect drafts up to 120s;medium for higher-quality local generation up to 380s when CUDA and dependency constraints are acceptable.Use Stable Audio Open 1.0 only when the 47s limit is acceptable and the workflow needs the older open model, diffusers/stable-audio-tools compatibility, or research comparison.
Do not use Stable Audio for:
Before any audio-to-audio or inpaint request, require the user or project brief to establish that every uploaded sound is rights-cleared for transformation. Stability's API docs and Stable Audio announcements state that copyrighted content is not allowed to be uploaded and describe content recognition/compliance scanning for uploads. Treat "I found this song online" as blocked until replaced with licensed, public-domain, commissioned, self-recorded, or otherwise authorized material.
For commercial media, record:
cfg_scale if used, strength, mask times, and output format;Stability's Terms of Service state that users are responsible for inputs and must have rights, licenses, and permissions for inputs; as between the user and Stability, Stability assigns any right it has in outputs subject to compliance and applicable law. The same terms also say outputs may be similar across users, users must verify legality/appropriateness before use, and Stability may use inputs/outputs to improve services unless the user opts out where available. Privacy/security policy claims should not be overstated: Stability states it uses organizational and technical safeguards, but no internet transmission/storage is guaranteed to be fully secure.
For confidential brand libraries, unreleased music, celebrity voices, minors, medical/legal content, or contractual IP, prefer enterprise-approved settings or local/open-weight routes only after the project's data-handling requirements are known. Do not upload sensitive stems just because the API is convenient.
Sources: Stability Terms of Service, Stability Acceptable Use Policy, Stability Privacy Policy, Stable Audio 2.0 announcement, Stable Audio 2.5 announcement, Stability SOC 2/SOC 3 announcement.
Stable Audio responds best when the prompt describes what the audio should sound like, not what the video should show. Translate visual intent into sonic parameters.
Documented prompt elements from Stability guidance include genre/subgenre, style, tempo/BPM, mood, and instrument type. Production heuristics below are not documented guarantees; use them because they make review and iteration easier:
Keep prompts internally consistent. A single prompt that asks for "minimal ambient corporate piano" and "aggressive drum & bass festival drop" will produce less controllable review targets.
Duration:
Model:
model explicitly. Do not rely on endpoint defaults.stable-audio-3 for 191-380s hosted work or async pipelines.stable-audio-2.5 for synchronous 1-190s hosted work unless project history requires 2.0.Steps:
Seed:
Output format:
wav for assets that will be edited, looped, layered, mixed, or mastered.mp3 for quick previews, low-risk temp tracks, or delivery systems that require it.Audio-to-audio strength:
Inpaint masks:
mask_start slightly before the flawed region and mask_end slightly after it so the model has room to blend.Text-to-audio music bed:
wav with documented seed/parameters.Text-to-audio SFX/foley:
Audio-to-audio transformation:
Inpainting/continuation:
Intent: an optimistic bed under narration for a SaaS launch reel, no vocals, easy to duck.
Route: hosted Stable Audio 2.5 text-to-audio, synchronous, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio-2/text-to-audio
model: stable-audio-2.5
duration: 40
steps: 8
output_format: wav
seed: 0
Prompt:
40-second modern product-launch instrumental bed at 104 BPM, optimistic and confident, clean electronic pop with warm analog synth pulses, soft piano accents, light brushed percussion, subtle bass, wide polished stereo mix. Structure: 4-second gentle intro, steady lift through 25 seconds, restrained final button ending. Designed to sit under spoken narration. No vocals, no lead guitar solo, no aggressive drums, no recognizable melody.
Why structured this way: The prompt names function, duration, BPM, mood, arrangement, section behavior, mix priority, and exclusions relevant to narration.
Expected review: Check the first 4 seconds for usable intro space, verify the 25-30s region can support the hero claim, and cut/fade the 40s render to the final 30s edit.
Likely failures: too much lead melody, drums masking speech, unresolved ending, or tempo not matching the edit. Iterate by reducing instrumentation or asking for "more sparse, more sidechain space for narration."
Intent: transform a client-owned two-note chime into a family of softer onboarding sounds.
Precondition: project log confirms the client owns the chime and permits transformation.
Route: hosted Stable Audio 3 audio-to-audio, async, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio/audio-to-audio
model: stable-audio-3
duration: 8
steps: 8
strength: 0.42
output_format: wav
seed: 0
Prompt:
8-second soft premium onboarding chime derived from the source gesture, calm and reassuring, two-note identity preserved as a subtle motif, warm glass mallet tone layered with quiet felt piano resonance, gentle airy tail, close clean studio sound, no melody beyond the two-note motif, no voice, no percussion, no harsh digital sparkle.
Workflow: Submit the job, store the returned generation id, poll /v2beta/audio/results/{id}, save the returned seed and request id, then create a short candidate sheet with waveform, loudness, and subjective notes.
Likely failures: source motif disappears at high strength; source is too literal at low strength; tail is too long for UI. Adjust strength first, not prompt complexity.
Intent: replace a 12-second section with accidental foreground chatter in an otherwise useful 90-second sci-fi hallway ambience.
Route: hosted Stable Audio 2.5 inpaint, synchronous, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio-2/inpaint
model: stable-audio-2.5
duration: 90
steps: 8
mask_start: 31.5
mask_end: 45.0
output_format: wav
seed: keep original asset seed if known; otherwise 0
Prompt:
90-second seamless sci-fi hallway ambience, low ventilation rumble, distant electrical hum, subtle metallic room tone, occasional very soft servo movement far away, tense but not musical. The replacement section should match the surrounding ambience and contain no speech, no footsteps, no alarms, no rhythmic music.
Review: Listen across 28-48s on headphones for seam clicks, room-tone jump, new foreground events, and stereo image shifts. If the new section draws attention, reduce event density in the prompt and widen the mask slightly.
Before accepting a Stable Audio asset:
wav for production; derive compressed preview/delivery formats from the approved master.Take calesthio/stable-audio from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.