Default downstream MiniMax H3 specialist for every image-based request unless the user explicitly declares boundary-only first/last frames with no reusable reference role. Use the official six-section full-reference format for character/person/object consistency, scene/style/action/camera/storyboard/voice/audio reference, source-video editing or continuation, ambiguous image roles, ordinary image animation, and prompt-level first/key/last-frame anchoring.
npx skills add https://github.com/unknowlei/minimax-h3-opencode-skills --skill minimax-h3-reference-video-prompt
Convert mixed reference assets and user intent into a traceable six-section full-reference prompt. Make every asset role explicit and write the target video as a detailed audiovisual timeline.
New MiniMax H3 requests should enter through minimax-h3-creative-director. This is the default specialist for any request containing images. Route away only when the user explicitly declares pure first/last boundary roles and no image also carries identity, character, person, object, costume, scene, style, action, camera, composition, or another reusable reference responsibility. Ambiguity stays here. Full-reference may use prompt-level first/key/last-frame anchoring while maintaining consistency.
Before drafting, read ../h3-prompt-writing/SKILL.md and ../h3-prompt-writing/references/ref-en.txt. Consult ../h3-prompt-writing/references/base-en.txt for the shared shot, camera, dialogue, visible-text, and sound rules when needed. Treat these official MiniMax files as the canonical prompt-format specification. If this skill or its local references conflict with them, follow the official files.
Read references/reference-rules.md before drafting.
When the director supplies a confirmed multishot_plan, treat its shot count, timing, content, framing, performance, camera, transitions, sound, active references, and continuity decisions as already answered. Do not ask those questions again. Map every confirmed shot into detailed_description, keep reference labels and responsibilities stable, and carry the continuity ledger into retention and preservation instructions. Reopen a decision only when the plan is physically impossible, conflicts with a source asset, or exceeds the effective duration.
For the official online product/API, keep output within 4-15 seconds at 24 FPS and the prompt within 7000 characters. Accept at most 9 images, 3 videos, 3 audio files, and 12 mixed files total. Each video or audio file must be 2-15 seconds; combined video duration and combined audio duration must each be at most 15 seconds. Audio cannot be the only reference type. Respect the documented per-file and API-body limits. Treat local ComfyUI constraints separately when they are narrower.
Always return the official six sections in their required order. Never return a free-form natural-language paragraph, keyword list, abbreviated prompt, base three-field prompt, or alternate schema as the final English prompt for a full-reference task.
Before assigning labels or drafting, use the host's structured-choice tool whenever the brief is sparse; gives only an image plus a generic motion request; omits two or more of action progression, scene treatment, preservation priorities, visual style, camera/editing, dialogue/voice, sound/music, or endpoint; leaves a decisive reference responsibility unclear; or asks the AI to improvise.
In OpenCode, call the built-in question tool rather than displaying Markdown options. In Codex, use the available structured user-input tool; if unavailable, ask one concise plain-text question.
Ask as many related questions as the current decision stage requires; normally 1-5, but three is not a per-call, per-turn, or per-session ceiling. A strict yes/no question may contain exactly two options; every other choice question must contain at least five materially different, feasible options. Put the context-specific recommendation first and append (Recommended) to its label. Explain the consequence of each choice, omit Other, and use multiple selection only when roles can truly be combined. Before every question tool call, count the options in the actual payload. If any non-binary choice has fewer than five, do not submit it; expand it with meaningful alternatives or make it open-ended. Do not create near-duplicates merely to reach five.
Prioritize:
For sparse or delegated briefs, ask enough high-impact questions in one batch to establish the current stage, normally 3-5, even when the asset itself is visually clear. Do not ask the user to classify every asset; default images to character/object/scene consistency and ask about the desired creative outcome. Skip questions only when the user explicitly prohibits them. A request to improvise still requires questions.
<Subject N>, <Picture N>, <Video N>, and <Audio N> labels only where their defined roles apply.keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference.Allow a full-reference task to designate a referenced image as a concrete first frame, keyframe, edited keyframe, last frame, or composition anchor. Define it as a standalone <Picture N> and state its exact role, for example:
<Picture 2> is the last frame of [Shot 3], defining the final pose, object placement, camera angle, lighting, and composition.
Combine keyframe completion with other task types in summary when appropriate. In detailed_description, describe a continuous path into or out of the anchored picture rather than repeating static image descriptions.
When a picture both anchors a boundary frame and preserves a person, character, object, costume, scene, style, or composition, keep the task in full-reference. Define both responsibilities explicitly instead of moving to the pure keyframe specialist.
Observe the current API distinction: prompt-level keyframe semantics are supported by the full-reference format, but an API request must not mix first_frame/last_frame roles with any reference_* role. When mixed reference assets are required, pass images as reference_image and express the concrete frame relationship in the prompt. Do not claim that the API role itself is a hard keyframe role.
Produce exactly these six sections in order:
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:
Write all six sections in English, preserving the original language only for dialogue, lyrics, and text visibly present in the scene.
detailed_description to a plot summary or a list of reference links.[unclear] for unintelligible spans.Return:
### English Prompttext code block containing only the complete six-section English prompt### 中文翻译Write all six sections in English except exact literal content that must remain in its source form:
<d>[Language] ...</d>Write subject definitions, reference relationships, visual description, actions, camera, editing, sound design, and music in English. Do not retain non-English prose merely because the user's request was written in that language.
After completing all six English sections, translate every prompt component into Chinese, including the meaning of dialogue, lyrics, visible text, task types, retention relationships, and sound descriptions. Keep field names, task/relationship markers, <Subject N>, <Picture N>, <Video N>, <Audio N>, [Shot N], timestamps, (Sx), and control tags visible alongside Chinese explanations. Preserve exact proper nouns and source literals, adding Chinese meaning where helpful. The Chinese section is explanatory and not an executable prompt.
Add ### 采用的假设 before the English prompt only if materially useful.
<Subject N> denotes reusable visible content; <Picture N> denotes a concrete frame or planning anchor; <Video N> denotes a whole-video source or temporal structure; <Audio N> denotes an audio signal or audio reference.(Sx) IDs; audio-only verbal cues do not create fictitious speakers.N/A.Take unknowlei/minimax-h3-reference-video-prompt from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.