mcpbeat Sign in

Happy Video Gen Skill for Claude

Universal AI video generation supporting OpenAI Sora, Google Veo 2/3, Runway Gen-3/Gen-4, Pika 2.2, Luma Dream Machine (Ray 2), FAL (Kling / Wan / Veo / Sora wrappers), Ark Seedance 1.5 Pro/Lite, Bailian Wanx (i2v), MiniMax Hailuo-02, and Vidu Q3. Use this skill whenever the user asks to generate, create, make, or synthesize a video from a text prompt or from a first-frame image. Covers text-to-video and image-to-video, with optional last-frame control on providers that support it. Typical phrases include "generate a video of ...", "make a 5-second clip of ...", "animate this image", "生成一段视频", "做个短片", or any mention of video-generation model families like Sora, Veo, Runway Gen, Kling, Wan, Seedance, Hailuo, Pika, Dream Machine, Vidu. Always use this skill even if the user does not name a specific model — pick a provider from their EXTEND.md defaults or available API keys. Do NOT use this skill when the user explicitly mentions 即梦 / Dreamina / Jimeng — those go to happy-dreamina instead.

18k tokens
context cost
the whole folder, loaded on every use
24
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
303
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/iamzhihuix/happy-claude-skills --skill happy-video-gen

The instruction itself

9 sections, as written by the author

happy-video-gen

Generates short videos (text-to-video or image-to-video) across 10 providers through one CLI: bun scripts/main.ts .... All providers are async — the CLI submits a job, polls until the provider finishes, then downloads the MP4 / WebM.

Quick usage

# Text-to-video
bun scripts/main.ts --prompt "camera slowly pushes into a calico cat on grass" --ar 16:9 --duration 5 --video ./out.mp4

# Image-to-video (first frame)
bun scripts/main.ts --prompt "subtle zoom, leaves swaying" --image ./keyframe.png --duration 5 --video ./out.mp4

# Image-to-video with last-frame control (provider-dependent)
bun scripts/main.ts --prompt "seamless morph" --image ./a.png --last-frame ./b.png --video ./out.mp4

When to invoke this skill

  • User asks to generate / create / make / synthesize a video from text.
  • User asks to animate a still image, or provides a first-frame path.
  • User names any video model family (Sora, Veo, Runway, Kling, Wan, Seedance, Hailuo, Pika, Dream Machine, Vidu).

Route to happy-dreamina if the user explicitly mentions 即梦 / Jimeng / dreamina CLI. Route to happy-image-gen if the user actually wants a still image.

Step 0: Preflight (BLOCKING)

  • Locate EXTEND.md (same resolution order as happy-image-gen):
  • ./.happy-skills/happy-video-gen/EXTEND.md
  • $XDG_CONFIG_HOME/happy-skills/happy-video-gen/EXTEND.md
  • ~/.happy-skills/happy-video-gen/EXTEND.md

If none, run bun scripts/main.ts --setup and walk the user through references/config/first-time-setup.md.

  • Verify one provider has credentials. Check env vars in the order the CLI auto-detects (see providers.md). Do not proceed without one usable provider.
  • Verify Bun. Fall back to npx -y bun if missing.
  • Warn about cost. Video generation is 10–100× more expensive per call than images. If the user asks for HD 1080P / 10-second clips, confirm before firing — show them the expected provider cost bracket from references/providers.md.

Step 1: Choose provider

Preference order:

  • --provider <id> explicitly passed.
  • EXTEND.md default_provider.
  • Auto-detect from env vars: fal > ark > minimax > runway > luma > pika > vidu > google > bailian > openai.

Pick by strength of the actual task:

  • Chinese prompts / Chinese text in frameark (Seedance) or bailian (Wanx).
  • Photorealistic portraitsgoogle (Veo 3) or runway (Gen-4).
  • Anime / stylizedfal (Kling) or luma.
  • Cheap draftark Seedance Lite, fal Kling v2.5 turbo, vidu Q1.
  • Voice-synced dialogue video (if applicable) → google Veo 3, openai Sora 2.

Step 2: Fill parameters

  • --prompt: always double-quote.
  • --image <path> / --last-frame <path>: local paths, will be base64-encoded as data URIs automatically. --last-frame only accepted by Luma and a few FAL endpoints.
  • --duration <seconds>: 5 is universal default. Caps: Sora-2 / Kling up to 10; Seedance up to 10; Luma up to 9.
  • --ar <ratio>: 16:9 / 9:16 / 1:1 / 4:3 / 3:4. See references/aspect_ratio_map.md for provider-specific quirks.
  • --resolution: 480p / 720p / 1080p. Not all providers honour it; most cap at 720p on cheap tiers.
  • --poll-timeout: default 600s (10 min). Increase for 1080P or >5s clips.

Step 3: Submit and wait

bun scripts/main.ts \
  --prompt "..." \
  --video ./out.mp4 \
  --provider ark \
  --duration 5 \
  --ar 16:9 \
  --resolution 720p

While waiting, do not fire another job on the same provider — concurrency caps on cheap tiers are strict (often 1). On success the CLI writes the MP4 and reports size + path. JSON output:

{ "success": true, "provider": "ark", "model": "doubao-seedance-1-0-lite-t2v-250408", "video": "/abs/out.mp4", "size_bytes": 4823456, "format": "mp4" }

Step 4: Timeouts and recovery

If polling exceeds --poll-timeout, the CLI throws with the provider-specific external id (task id / job id / operation name / request id). Capture it from stderr and resume later with provider-specific tooling. See references/async-protocol.md for the per-provider id format and manual resume commands.

References

  • references/providers.md — all 10 providers with env vars, defaults, cost notes, feature matrix.
  • references/async-protocol.md — external id format per provider + how to resume a stuck task.
  • references/aspect_ratio_map.md--ar mapping per provider.
  • references/error_codes.md — common errors and fixes.
  • references/config/first-time-setup.md — setup walkthrough.
  • references/config/extend-schema.md — EXTEND.md schema.
  • assets/EXTEND.template.md — config template.

Other skills for the same job

different authors, same section of the catalogue
Canvas Design
by anthropics
vendor ×13

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

1388k tokens
Algorithmic Art
by anthropics
vendor ×10

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

15k tokens scripts
Image Enhancer
by frostant
×6

Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.

635 tokens
Video Downloader
by CommandCodeAI
×4

Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.

671 tokens
Histolab
by christophacham
×3

Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.

18k tokens
Omero Integration
by christophacham
×3

Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.

32k tokens
Pydicom
by christophacham
×3

Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.

13k tokens scripts
Transformers
by christophacham
×3

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

13k tokens

How to use it

Copy the folder

Take iamzhihuix/happy-video-gen from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference npx. Without those the skill loads but fails at the first command.