calesthio/seedance-2-0
Direct ByteDance Dreamina Seedance 2.0 Standard, Fast, and Mini video production across BytePlus ModelArk and verified gateways. Use for text/image/reference-to-video, native synchronized audio and dialogue, multimodal image-video-audio reference, first/last-frame animation, video editing or extension, model/gateway selection, prompt construction, async task handling, failure recovery, QA, and rights-safe production.
npx skills add https://github.com/calesthio/generative-media-skills --skill seedance-2-0
Treat this as a provider- and model-family-specific operating guide, not a generic video-prompt recipe. Establish the exact gateway contract before writing a request. “Seedance 2.0” does not imply one portable model ID or schema.
Evidence labels used here:
Prefer this family when the shot benefits from one or more of:
Choose another workflow, or plan post-production, when the requirement is:
Do not call Seedance 2.0 “best” in the abstract. ByteDance's published comparisons include internal evaluation, and public leaderboards change. Select it against the shot's references, audio, control, delivery, cost, latency, and compliance needs.
The official API is asynchronous:
POST https://ark.ap-southeast.bytepluses.com/api/v3/contents/generations/tasksGET https://ark.ap-southeast.bytepluses.com/api/v3/contents/generations/tasks/{id}Use the current model catalog, not an invented alias:
| Variant | ModelArk ID verified 2026-07-09 | Choose when | Documented output ceiling |
|---|---|---|---|
| Seedance 2.0 | dreamina-seedance-2-0-260128 | maximum family quality; 1080p or 4K required | 480p, 720p, 1080p, 4K; 4K is 10-bit HEVC |
| Seedance 2.0 Fast | dreamina-seedance-2-0-fast-260128 | lower latency/cost matters more than the last quality increment | 480p, 720p |
| Seedance 2.0 Mini | dreamina-seedance-2-0-mini-260615 | best family cost efficiency for drafts or scaled workloads | 480p, 720p |
Standard, Fast, and Mini are documented as largely sharing the core creation modes, but output and platform controls differ. Confirm the live model card and activation status before spending.
Do not translate calls by changing only the model name.
| Contract | Identifier/endpoint pattern | Important differences verified 2026-07-09 |
|---|---|---|
| BytePlus ModelArk | model ID in one content-generation endpoint | heterogeneous content[] items with role; ratio; integer duration or -1; direct Seedance 2.0 does not support seed, frames, or camera_fixed; asynchronous task ID |
| fal | bytedance/seedance-2.0/{text-to-video,image-to-video,reference-to-video} and /fast/... | separate mode endpoints; aspect_ratio; duration is an enum such as "auto" or "4"–"15"; first/last fields and modality URL arrays; exposes seed; exact resolution options differ by endpoint/tier, so read that endpoint's live schema |
| Replicate | bytedance/seedance-2.0 | unified wrapper; image, last_frame_image, reference_images, reference_videos, reference_audios; integer duration with -1; aspect_ratio; exposes seed; returns a file/URI |
Gateway-exposed seed fields do not prove that direct ModelArk has seed control. Treat reproducibility as gateway-specific and verify by repeated tests; even a fixed supported seed need not yield byte-identical output.
Gateway prompt labels also differ. ModelArk documentation refers to media by ordered type as Image 1, Video 1, and Audio 1. fal pages use forms such as @Image1 and, on other pages, [Image1]; Replicate documents [Image1]. Follow the active endpoint's schema and example exactly. Never reference an asset ID in prose when the contract expects an ordinal media label.
Do not hardcode prices, rate limits, or route availability into a production plan. Fetch them at preflight; record retrieval time, model/tier, resolution, duration, audio setting, input-video surcharge rules, and currency.
Use top-level JSON parameters, not legacy prompt suffixes.
| Field | Direct behavior verified 2026-07-09 |
|---|---|
| model | current model ID or an activated endpoint ID |
| content | ordered text/image/video/audio items; roles determine first/last/reference behavior |
| resolution | default 720p; Standard alone supports 1080p and 4K under the current contract |
| ratio | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16; default adaptive for the 2.0 series |
| duration | integer 4–15, default 5; -1 delegates integer duration selection to the model and affects billing |
| generate_audio | default true; false produces silent output; generated audio is mono |
| watermark | default false; true adds an “AI Generated” mark in the lower right; do not remove required provenance signals |
| return_last_frame | default false; request a PNG last frame for continuity workflows |
| callback_url | optional status callback; authenticate and make callback handling idempotent |
| execution_expires_after | 3,600–259,200 seconds; default 172,800 (48 h) |
| safety_identifier | stable, unique end-user identifier, at most 64 characters; hash an internal ID rather than send personal data |
| priority | Standard Seedance 2.0 only; integer 0–9 within the same endpoint queue; not an execution-speed guarantee |
Do not send frames, seed, camera_fixed, draft, or service_tier: "flex" for the direct 2.0 family merely because another Seedance version or gateway accepts them. The official API documents frames, seed, and camera_fixed as unsupported for Seedance 2.0, draft mode as a 1.5 Pro feature, and the 2.0 series as online-inference only.
content[] correctlyText item:
{"type":"text","text":"<complete prompt>"}
First-frame image:
{"type":"image_url","image_url":{"url":"<public URL, data URI, or asset:// ID>"},"role":"first_frame"}
Strict first-and-last frame mode requires exactly two image items with first_frame and last_frame. Reference mode uses reference_image, reference_video, and reference_audio roles.
First-frame, first-and-last-frame, and multimodal-reference modes are mutually exclusive at the role level. In reference mode, prompt language can ask an image to influence the opening or ending, but only first/last-frame roles strictly anchor those frames.
Direct ModelArk documents these limits:
Use ffprobe or equivalent before paying for generation. Verify codec, duration, fps, dimensions, aspect ratio, audio stream, and file size. Normalize orientation metadata. Strip accidental captions, watermarks, unrelated faces, and conflicting backgrounds when they are not intended references.
Use the smallest reference set that can express the shot. The official prompt guide warns against automatically filling every available slot: too many references obscure feature priority and can cause style conflicts, weak subject identification, and drift.
Create an asset ledger before writing the prompt:
| Ordinal | Asset | Rights/consent | Intended dimension | Must preserve | May change |
|---|---|---|---|---|---|
| Image 1 | product hero | owned | exact geometry, label colors | silhouette, materials | environment |
| Video 1 | motion plate | licensed | camera path and hand motion | timing, interaction | original product |
| Audio 1 | approved voice sample | performer release | timbre and pace | voice character | exact words |
Order high-priority assets first. Reference only the dimension needed: subject identity, product geometry, wardrobe, environment, composition, camera movement, action, effect, voice timbre, melody, or rhythm. “Use everything from every reference” creates conflicts.
For characters, the official guide recommends a clean headshot plus one full-body/styling image rather than a multi-view collage. A small face inside a busy sheet receives weak identity weight; multiple views may be interpreted as different people and increase identity drift or duplicates. On ModelArk, use its trusted-output, preset-character, or authorized-real-person asset paths rather than trying to bypass face moderation.
Start with the shot's invariant, then its temporal behavior:
Shot 1, Shot 2, Shot 3 in event order;Use the same subject label every time. Re-bind it when ambiguity is possible: the amber bottle from Image 1, not later just it when several props exist.
Prefer causal verbs and observable performance:
She is nervous and the shot is cinematic.Her shoulders tighten; she glances twice toward the locked door, exhales through parted lips, then slowly reaches for the handle. A restrained handheld medium close-up follows her hand.For multi-shot generation, describe order rather than demanding frame-accurate timecodes. BytePlus explicitly documents unstable compliance with exact segment timing such as “0–3 seconds.” Use approximate beats, generate shots separately when a cut must land exactly, and finish timing in an editor.
Do not stack push-in, orbit, pan, tilt, zoom, crane, and shake into one shot. The official guide recommends one camera movement per shot because combined movement increases instability.
For dialogue:
For on-screen text, keep phrases short and common. Even though the model supports slogans, subtitles, and speech bubbles, do not trust generated typography for legal copy or final brand delivery. Prefer clean footage plus deterministic text and logo compositing in post.
Describe subject, environment, action sequence, camera, lighting, visual finish, audio, and exclusions. Use this for original shots when reference fidelity is not required.
Describe motion away from the first frame and, for two-frame mode, the causal transformation that arrives at the last. Match input and output aspect ratios; mismatched dimensions can cause center crop, stretching, compression, or abrupt frame jumps. Use ratio: "adaptive" when exact delivery framing is not yet locked, or pre-crop both frames to the documented target raster.
Use explicit bindings such as:
Define the matte amber pump bottle in Image 1 as PRODUCT.
Use only the clockwise hand rotation and camera path from Video 1.
Use the warm, low female timbre from Audio 1 for the quoted line.
PRODUCT keeps its exact silhouette, pump geometry, amber glass, and cream label throughout.
...
Do not submit reference audio alone. When voice matching matters, describe voice traits in words as well as attaching the approved sample, and keep new lines close to the sample's register and delivery.
Name the source and scope, then state what changes and what stays:
Strictly edit Video 1. Replace only the silver can with PRODUCT from Image 1. Preserve the actor's hands, grip, finger occlusion, camera movement, lighting direction, reflections, background, timing, and all other content.
For removal, name both the object to remove and the background/occlusion behavior that must be reconstructed. For combined tasks, distinguish the reference source from the video being edited.
State forward or backward extension, the new action, and continuity constraints for subject, lighting, camera, motion, ambience, and narrative. With 2–3 source clips, describe the transition sequence explicitly.
Expect iterative extension to degrade detail and faces. Limit continuation depth, keep high-quality anchors, and regenerate from a clean state rather than extending an already degraded result repeatedly. Official guidance notes that joins can jump or roll back; align and trim in post and hide difficult joins on motivated cuts.
These are examples, not mandatory formulas. Translate field names and media labels to the active gateway.
Intent: create one 8-second, audio-on, 9:16 social shot without external references.
Request:
{
"model": "dreamina-seedance-2-0-260128",
"content": [{
"type": "text",
"text": "Photorealistic vertical lifestyle scene. A fictional woman in her early thirties stands at a sunlit kitchen island holding an unbranded ceramic travel mug. She has a short black bob, a mustard linen shirt, and a calm, candid manner. Single shot: medium close-up with a gentle handheld drift only. She turns the mug once in both hands, looks toward the hallway, smiles slightly, and says in a warm conversational voice: \"Train in ten. Let's go.\" A soft ceramic tap lands as she sets it down. Natural room tone, distant morning traffic, no music. Soft window light, realistic skin texture, neutral color grade. No subtitles, no logos, no watermark-like graphics, no duplicate fingers, no extra people."
}],
"resolution": "1080p",
"ratio": "9:16",
"duration": 8,
"generate_audio": true,
"watermark": true,
"return_last_frame": false,
"safety_identifier": "<stable-end-user-hash>"
}
Why structured this way: one speaker, one action arc, one camera motion, quoted speech, and explicit sound priorities fit the duration. The actor is fictional, the object is unbranded, and the line makes no testimonial or product-performance claim.
Expected result: a usable hero take with synchronized speech and a motivated object sound.
Likely failures and repair: if the line truncates, shorten it or increase duration; if unwanted captions appear, reinforce no subtitles and plan deterministic captioning in post; if fingers deform during the turn, simplify to a slower quarter-turn or split the set-down into another shot.
Variation: set generate_audio: false for a controlled external voiceover workflow.
Intent: make an 11-second landscape launch insert using two product images, one motion reference, and one consented voice sample.
Inputs in request order: Image 1 front product hero; Image 2 side/detail view; Video 1 licensed hand-and-camera motion plate; Audio 1 consented performer reference.
Complete prompt:
Define the matte navy portable speaker shown in Image 1 and Image 2 as SPEAKER. Preserve SPEAKER's exact rounded-rectangle silhouette, grille pattern, two coral buttons, seam placement, and logo-free front.
Use only the slow hand placement and 30-degree clockwise camera arc from Video 1. Use the clear, low, lightly textured female timbre and measured pace from Audio 1 for the quoted line; do not reuse words from Audio 1.
Shot 1: On a dark walnut desk at blue hour, a hand places SPEAKER beside a rain-speckled window. Medium close-up; the camera performs the same restrained clockwise arc from Video 1. SPEAKER remains rigid and correctly proportioned.
Shot 2: Cut to a macro close-up of one coral button depressing once. A soft tactile click is synchronized with the press; a narrow rim light travels across the grille.
Shot 3: Return to the three-quarter hero angle. Off-camera, the woman says: "Small room. Full sound." A restrained low-frequency music pulse rises after the line, then ends cleanly.
Premium product cinematography, realistic materials, controlled reflections, navy-and-coral palette, shallow depth of field. No subtitles, no added labels, no altered button count, no warped grille, no hands touching the grille, no duplicate product, no third-party logo.
Direct parameters: Standard model; resolution: "1080p"; ratio: "16:9"; duration: 11; generate_audio: true; every media item uses the relevant reference_* role.
Expected result: product identity comes from Images 1–2, camera/action from Video 1, and voice character from Audio 1 without copying its content.
Likely failures and repair: if product geometry drifts, remove the lower-priority side image or reduce camera rotation; if the model copies source-video styling, state reference only motion and camera path; if voice match is weak, describe timbre/register more precisely and make the new line closer in cadence to the sample.
Variation: use Fast at 720p for motion-layout tests, then switch to Standard only after approval; record the model switch as a new paid generation decision.
Intent: create a silent 6-second transition from a closed paper package to the same package opened with contents arranged.
Inputs: two owned images pre-cropped to the same 1:1 raster; first_frame then last_frame.
Complete prompt:
Single locked overhead shot. Beginning exactly from the closed kraft-paper package in the first frame, two gloved hands enter from the lower edge, slowly untie the black cotton cord, unfold the top flap, and place the three tea sachets into a neat fan. Every movement is continuous and physically causal, arriving exactly at the supplied last frame. Preserve the package dimensions, paper texture, cord color, sachet count, tabletop grain, lighting direction, and all printed artwork from the supplied frames. Silent video. No additional objects, no camera motion, no text changes, no hand crossing over the printed mark.
Direct parameters: resolution: "1080p"; ratio: "1:1"; duration: 6; generate_audio: false.
Expected result: exact boundary frames with model-generated causal motion between them.
Likely failures and repair: if the pack jumps or stretches, verify both images have identical raster/aspect and match the chosen output table; simplify the knot action; try adaptive only if delivery framing is flexible. Do not replace strict roles with reference_image if exact end-frame arrival is required.
Intent: replace one prop in a licensed 7-second source plate while preserving performance and lighting.
Inputs: Image 1 owned cream jar; Video 1 licensed source plate with a generic jar and no unconsented face.
Complete prompt:
Define the frosted pale-green cream jar with the flat white lid in Image 1 as JAR.
Strictly edit Video 1. Replace only the generic jar in the actor's right hand with JAR. Keep JAR's exact cylindrical proportions, frosted glass, flat lid, and blank label panel. Preserve the actor's hands and finger placement, natural occlusion around the jar, wrist motion, original camera path, rack focus, scene duration, background, wardrobe, lighting direction, shadows, reflections, and all audio from Video 1. Do not change the face, body, hand count, grip, or background. Do not add text, logos, subtitles, or a second jar.
Direct parameters: reference mode; resolution: "1080p"; ratio: "adaptive"; duration: 7; generate_audio: true if source audio continuity is required.
Expected result: local prop replacement with preserved plate motion.
Likely failures and repair: if the hand changes, repeat replace only the jar and slow/shorten the source plate; if lighting mismatches, add a product reference rendered under similar light; if exact brand graphics matter, composite the final label deterministically after generation.
Intent: build a 24-second sequence as three separately approved 8-second clips.
Workflow:
return_last_frame: true; download both video and PNG immediately.first_frame; describe only the next action and preserve subject, lens height, light, weather, screen direction, wardrobe, and ambience.Clip B prompt:
Continue exactly from the supplied first frame. The same red delivery bicycle moves left-to-right through the wet alley; the rider pedals twice, then brakes beside the blue doorway. Preserve the bicycle frame geometry, rider's yellow raincoat and black helmet, wet cobblestone reflections, overcast light, 35 mm eye-level perspective, left-to-right screen direction, and light rain ambience. One lateral tracking movement only. The brake squeak lands as the rear wheel stops; no speech or music. No new riders, no duplicated wheels, no change of weather, no subtitles or logos.
Direct parameters: Standard or Fast consistently across the chain; ratio: "16:9"; duration: 8; generate_audio: true; return_last_frame: true.
Likely failures and repair: repeated continuation may compound degradation. Regenerate from the last clean approved anchor rather than extend a degraded clip; use a cutaway to mask discontinuity; do not assume last-frame handoff guarantees identity or motion continuity.
For callbacks, validate origin/signature if the gateway provides one, accept duplicate delivery, transition task state idempotently, and fall back to bounded polling with exponential backoff and jitter. Treat queued, running, succeeded, failed, expired, and cancelled explicitly. Do not endlessly resubmit a moderation failure.
| Symptom | Likely cause | First repair |
|---|---|---|
| subject identity drifts | face/subject too small, collage/multi-view ambiguity, weak binding | use a clean close-up plus one styling image; define and repeat one subject label; order identity first |
| duplicate people/products | ambiguous multiple views or crowded prompt | use single-subject references; bind every role; prohibit duplicates; split crowded action into shots |
| wrong product geometry | motion/style references overpower product reference | reduce references; state exact invariant features; use a slower, smaller rotation; composite fine label detail later |
| style drifts realistic | realistic source conflicts with requested stylization | name the target style globally and state which source dimensions not to inherit |
| unexpected subtitles/logo | speech or source text triggers graphic priors | remove irrelevant source text; state no generated text; try landscape; finish text/logo in post; accept that prompts cannot guarantee absence |
| first/last transition jumps | mismatched aspect/raster or over-complex transformation | pre-crop both frames to the same supported raster; simplify action; use adaptive ratio only when acceptable |
| exact beat ignored | model timing is approximate | use ordered shots without tight timestamps; generate separate shots and edit timing deterministically |
| camera unstable | too many simultaneous camera commands | keep one principal move per shot; reduce subject dynamics |
| dialogue truncates/mispronounces | too many words, uncommon/polyphonic terms, duration pressure | shorten/rephrase phonetically, simplify diction, extend duration, or dub in post |
| audio clicks at tail | generated phrase/music ends on clip boundary | regenerate with a clean ending or apply a short post-production fade/envelope |
| extension join jumps | rollback or discontinuity at generated boundary | trim/align around the join, cut on action, insert a cutaway, or regenerate from a cleaner transition |
| extension quality decays | repeated re-encoding/generation from degraded outputs | limit chain depth; reuse high-quality identity anchors; regenerate from the last clean state |
| moderation rejects a real face | direct upload is not a trusted/authorized portrait path | use ModelArk's authorized-real-person asset flow, approved preset character, or eligible same-account trusted output; never evade filters |
| request rejected | wrong schema, role combination, format, duration, region, quota, or unsupported parameter | validate against the exact live endpoint; remove cross-gateway fields; inspect structured error before changing creative intent |
Regenerate only after classifying the failure as contract, asset, prompt, model variance, moderation, or post-production. Preserve successful components rather than rewriting the whole prompt.
Inspect the complete file, not only a thumbnail or first frame.
Picture and motion
Audio
Technical
Editorial and safety
Require human review for public release, advertising claims, realistic people, political/news contexts, health/finance/safety claims, minors, or any high-impact use.
Before uploading, record source, owner, license, permitted transformations, territory, term, model-training restrictions, and release status for every asset. A technically accepted upload is not evidence of permission.
Obtain explicit, scoped permission for a person's face, body, performance, and voice. ModelArk's authorized-real-person library uses verification and authorization; use that path when applicable. Do not depict a living or dead person without appropriate rights, impersonate them deceptively, create non-consensual intimate content, or bypass face/public-figure filters.
Do not request copyrighted characters, franchise worlds, celebrity likenesses, signature performances, protected logos, or living-artist imitation merely because the model can approximate them. Use owned/public-domain/original direction, or document a license. BytePlus's terms make the customer responsible for input and output legality and say outputs may be non-unique; BytePlus does not warrant non-infringement.
Do not remove or conceal required AI indicators, watermarks, credentials, identifiers, metadata, or other provenance signals. Disclose synthetic media where context or law requires it. Keep raw generations, request records, consent records, and final-edit provenance.
Never work around moderation. If a benign request is blocked, inspect the provider error, remove ambiguity, use the documented authorized-asset route, or escalate to provider support. A gateway accepting content does not make the use lawful or safe.
Volatile facts above were re-verified 2026-07-09. Re-check before production because model IDs, limits, access, schemas, and prices can change.
Primary and official:
Verified gateway documentation, not the direct BytePlus contract:
Independent evidence, use narrowly:
Take calesthio/seedance-2-0 from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.