mcpbeat

calesthio Skills

153 skills published by calesthio across 1 repository. Together they weigh 1 661 433 tokens — that is what loading all of them at once would cost you in context.

153 skills 1 661 433 tokens total

2d Character Rig Animation
generative-media-skills

Production guidance for reusable layered 2D character rigs in browser-rendered and video media. Use for SVG hierarchies, pivots and constraints, FK/IK planning, pose and facial libraries, acting, cycles, deterministic handoff, crop variants, and rig QA; not for character concept continuity, 3D armatures, or motion-capture solving.

6k tokens
3d Asset Production
generative-media-skills

Use this skill to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery. It covers requirements, topology repair, scale and axes, UVs, PBR baking, LODs, static and skeletal readiness, rig handoff, optimization, glTF/USD/FBX delivery, validation, QA, and rights/provenance. Do not use it for provider-specific generation prompts, full scene/world building, mocap operation, or exact CAD/manufacturing work.

22k tokens scripts
Ace Step
generative-media-skills

Use ACE-Step and ACE-Step 1.5 for local or hosted AI music generation, including text-to-music, lyrics-to-song, instrumental beds, covers, repainting, stem/track extraction, track completion, LoRA personalization, REST/Python/Gradio workflows, rights review, and music integration for video, ads, games, and social content.

7k tokens
Adobe Firefly Image
generative-media-skills

Use Adobe Firefly Services to generate, edit, expand, fill, match, composite, and upscale still images through the current REST APIs. Apply when selecting Firefly Image 5 versus Image 3/4 or custom models, implementing authenticated asynchronous image workflows, using style or structure references, building product composites, handling provenance and enterprise rights, or diagnosing Firefly image jobs in production. Exclude Firefly video, audio, and unrelated Creative Cloud APIs.

15k tokens
Alibaba Image Models
generative-media-skills

Plan, prompt, call, edit, iterate, and productionize Alibaba Cloud Model Studio image generation with current Wan 2.7 Image, Qwen-Image 2.0, and Z-Image Turbo models. Use for Alibaba/DashScope text-to-image, multi-reference generation, instruction editing, character-consistent image sets, typography and layout work, regional authentication, model selection, API integration, cost/rate-limit planning, troubleshooting, safety, rights, privacy, and output QA.

10k tokens
Alibaba Wan Video
generative-media-skills

Plan, implement, and review Alibaba Wan video generation using Alibaba Cloud Model Studio/DashScope hosted APIs or official Wan 2.1/2.2 open-weight checkpoints. Use for Wan text-to-video, image-to-video, first/last-frame or continuation, reference-to-video, video editing, speech-driven video, character animation, regional deployment, billing, licensing, and production safety.

14k tokens
Amazon Nova Canvas
generative-media-skills

Produce and review production image-generation and image-editing workflows with Amazon Nova Canvas on Amazon Bedrock, including native InvokeModel payloads, safe authentication, validation, retries, cost controls, provenance, and lifecycle migration checks. Use when a task names Nova Canvas, amazon.nova-canvas-v1:0, Bedrock image generation, Canvas inpainting/outpainting/background removal, image conditioning, color guidance, image variation, virtual try-on, or Canvas fine-tuning.

11k tokens
Amazon Nova Reel
generative-media-skills

Produce and operate Amazon Nova Reel video-generation jobs through Amazon Bedrock. Use for Nova Reel text-to-video, image-conditioned animation, automated or manual multi-shot storyboards, async S3 delivery, cost and approval gates, production continuity, output custody, and AWS-specific safety, privacy, IAM, and provenance decisions.

9k tokens
Amazon Polly
generative-media-skills

Use Amazon Polly for production text-to-speech work: selecting Standard, Neural, Long-form, or Generative engines and compatible voices; authoring SSML; creating speech marks for captions, word highlighting, or lip-sync; managing pronunciation lexicons; running synchronous, streaming, or asynchronous S3-backed synthesis; planning quotas, pricing, IAM, privacy, and QA for narration, audiobooks, accessibility audio, avatars, and multilingual media.

9k tokens
Amazon Rekognition
generative-media-skills

Use this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR, moderation, image properties, Custom Labels or moderation adapters, stored-video analysis, conditional streaming-video workflows for existing eligible accounts, searchable media libraries, confidence evaluation, S3/IAM/event architecture, privacy, biometric consent, cost control, lifecycle management, and QA.

11k tokens
Amazon Transcribe
generative-media-skills

Use Amazon Transcribe for AWS-based speech-to-text production: batch S3 transcription, real-time streaming, captions/subtitles, diarization, channel identification, custom vocabularies, vocabulary filters, language identification, PII/PHI handling, toxicity detection, Call Analytics, Medical, and secure S3/IAM/KMS workflows.

9k tokens
Anime Animation Production
generative-media-skills

>- Produce anime-style and stylized 2D animation with generative image and video tools as a craft discipline, not a filter. Use when a request asks for anime, manga-style, cel-shaded, sakuga, shonen/shojo/seinen/slice-of-life, mecha, or retro-90s-OVA animation; when planning shots, motion, and cutting rhythm to anime convention; when holding a character on-model and preventing style drift across shots or episodes; when choosing prompting vocabulary that actually controls anime look; when pairing voice, music, and impact SFX to stylized action; and when navigating the cultural, IP, and platform-policy sensitivities of AI anime (studio/artist style imitation, fan-art and IP boundaries, disclosure and monetization). This is provider-neutral; models are named only as illustrative options.

13k tokens
Assemblyai Transcription
generative-media-skills

Use this skill when an agent needs AssemblyAI for speech-to-text or speech-understanding work in media production, including pre-recorded, synchronous short-file, and real-time streaming transcription; speaker diarization or speaker identification; captions and subtitles; timestamps; language detection, code-switching, or transcript translation; audio intelligence such as summaries, chapters, topics, entities, key phrases, sentiment, content moderation, profanity filtering, and PII redaction; webhooks, scaling, rate limits, retention, security, consent, and QA for podcasts, interviews, captions, call recordings, and edit workflows.

10k tokens
Asset Continuity Management
generative-media-skills

Provider-independent asset continuity and version management for generated-media production. Use when an agent must track generated or reference assets, prompts, seeds, model/tool versions, manifests, approvals, edit decisions, derivatives, localization variants, metadata/provenance, and continuity QA across image, video, audio, avatar, product-ad, social, explainer, or film-style pipelines.

10k tokens
Audiobook Production
generative-media-skills

>- Produce full-length audiobooks and long-form narration with generative voice tools. Use when the task is to turn a manuscript or long text into hours of footnotes, tables, dialogue), casting a single narrator or full cast, controlling pronunciation and voice consistency across a whole book, running a proofing/QC listen, meeting a retailer's technical delivery specs (RMS, peak, noise floor, room tone, chapterized files, metadata), and choosing a distribution route under each platform's current AI-narration policy. Not for short TTS clips, single-line voiceover, podcast production, or captioning existing video.

11k tokens
Audio Mixing Mastering
generative-media-skills

Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social content. Use when planning, mixing, repairing, mastering, QCing, or delivering dialogue, music, ambience, and sound effects, including loudness/true-peak targets, intelligibility, accessibility, stems, stereo/immersive decisions, platform/client specs, and final audio QA.

14k tokens scripts
Audio Reactive Video Composition
generative-media-skills

Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions; not for music-video concept direction, audio generation, or a runtime-specific recipe.

4k tokens
Audioshake Stem Separation
generative-media-skills

>- Operate AudioShake's cloud source-separation service (developer.audioshake.ai) to split a recording into stems. Use when a task involves isolating vocals, drums, bass, guitar, piano, keys, strings, or winds from music; separating dialogue / music / effects (DME) for film-TV post-production, localization, and dubbing; splitting a mixed recording into one stem per speaker; lyric transcription or word-level alignment; music detection/identification; or speech denoise/dereverb. Covers the Tasks API job lifecycle (assets, targets, formats, polling vs webhooks, credits, limits), choosing which stem targets a production actually needs, reviewing separated stems for bleed / artifacts / transient smearing / phase problems and repairing them, the rights and consent questions raised by separating copyrighted or third-party recordings, and the decision boundary for when a local open model such as Demucs is the better route than the API. Not for generating or synthesizing new audio, mixing, or mastering.

9k tokens
Avatar Spokesperson Production
generative-media-skills

Provider-independent production workflow for AI avatar spokesperson videos, including presenter briefs, consent and likeness rights, disclosure and platform labeling, casting, localization, performance direction, brand fit, claims review, lip-sync/audio QA, accessibility, approvals, and delivery packages. Use when creating synthetic presenter videos, corporate explainers, training modules, sales/support avatars, localization spokespeople, product announcements, compliance videos, or social avatar clips.

10k tokens
Azure Speech
generative-media-skills

Use Microsoft Azure Speech in Foundry Tools for media-production speech workflows: speech-to-text, fast and batch transcription, diarization, captions/subtitles, real-time transcription, text-to-speech neural and HD voices, SSML, custom/personal voices, text-to-speech avatars, speech/video translation, localization, quotas, regions, privacy, consent, rights, and production QA.

14k tokens
Black Forest Labs Flux
generative-media-skills

Plan, prompt, execute, troubleshoot, and quality-control Black Forest Labs FLUX image generation and editing across the BFL direct API and licensed local/open-weight deployments. Use for FLUX.2 model selection, text-to-image, single- or multi-reference editing, typography, exact-color work, mask-based erase/inpainting, outpainting, API integration, reproducibility, deployment licensing, and production rights or safety decisions; also use when migrating or maintaining legacy FLUX.1/FLUX1.1 workflows.

12k tokens
Brand Launch Film Production
generative-media-skills

Provider-independent production workflow for AI agents creating brand launch films, product reveal videos, manifesto films, campaign hero films, website hero videos, investor/customer launch assets, and social cutdowns. Use when planning, scripting, generating, editing, reviewing, or delivering launch films that must align strategy, claims, brand voice, visual direction, product proof, accessibility, compliance, and multi-platform delivery.

11k tokens
Bria Fibo Image
generative-media-skills

Build and operate rights-aware Bria FIBO and FIBO Lite image-generation workflows with structured prompts, reference images, asynchronous status handling, webhooks, cost gates, and safe artifact downloads. Use when a user asks for Bria/FIBO generation, refinement, inspiration, reproducibility, hosted API integration, or a licensed-data image workflow; do not use for FIBO Edit, video, generic image editing, or third-party FIBO gateways.

16k tokens
Bytedance Seedream
generative-media-skills

Build and operate production image generation and natural-language image editing with ByteDance Seedream through first-party Volcengine Ark (China) or BytePlus ModelArk (global), including model and region selection, multi-reference and grouped outputs, streaming, secure artifact handling, retries, cost controls, prompting, safety, and rights review.

11k tokens
Byteplus Seed Speech Tts
generative-media-skills

Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling, and output QA. Do not use for mainland-China Volcengine Doubao Speech endpoints.

4k tokens
Captions Media Accessibility
generative-media-skills

Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content. Use when planning, authoring, reviewing, localizing, burning in, exporting, or QAing captions, subtitles, SDH, transcripts, audio description, flashing/motion safety, caption readability, or accessible media handoff files.

17k tokens scripts
Cartesia Sonic
generative-media-skills

Use Cartesia Sonic and related Cartesia voice APIs for production speech: text-to-speech, realtime WebSocket TTS, voice selection, instant and professional voice cloning, pronunciation/language/emotion controls, voice localization, voice changer, pricing/concurrency planning, privacy/security review, and QA for narration, ads, localization, dubbing, avatars, and interactive voice agents.

11k tokens
Character Design Continuity
generative-media-skills

Provider-independent character design and continuity direction for generated media. Use when an agent must create, preserve, revise, or QA a recurring character across AI-generated images, video clips, avatars, animation, comics, ads, product mascots, serialized social posts, or multi-shot campaigns; especially when building character bibles, reference packs, prompts, expression/pose ranges, costume/hair/makeup/prop continuity, identity and likeness safeguards, iteration plans, handoff notes, or continuity reviews.

10k tokens
Cinematic Shot Direction
generative-media-skills

Direct cinematic shots and coverage for AI-assisted content production. Use when an agent must translate story intent into framing, shot size, lens and perspective, camera position, movement, blocking, screen direction, lighting-aware coverage, continuity, shot specifications, or diagnose why generated or filmed shots do not cut together.

12k tokens
Cinematic Trailer Production
generative-media-skills

Provider-independent cinematic trailer production for AI agents creating movie-style trailers, teaser trailers, launch trailers, game trailers, documentary trailers, series promos, and high-impact generated-media previews. Use for trailer strategy, structure, shot planning, voiceover/title-card copy, music and sound-design direction, generated-video prompt translation, edit pacing, platform/runtime choices, delivery handoff, rights/claims caveats, and trailer QA.

12k tokens
Color Grading Finishing
generative-media-skills

Provider-independent color grading and finishing direction for generated video, ads, trailers, product films, social clips, explainers, and mixed-source edits. Use when planning, directing, reviewing, or QAing color correction, shot matching, exposure, contrast, saturation, skin/product color protection, look development, LUT/reference use, SDR/HDR color management, delivery constraints, generated-media artifacts, editor/compositor handoff, and finishing QC.

9k tokens
Comfyui Media Workflows
generative-media-skills

Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media generation; use when a task involves ComfyUI workflow JSON/API graphs, model and custom-node inventories, reproducibility, provenance, safety/rights review, local or cloud execution hygiene, or production QA.

16k tokens scripts
D3 Animated Data Visualization
generative-media-skills

Production guidance for converting sourced data and approved claims into truthful, accessible, deterministic animated visualizations with D3. Use for data-driven charts, maps, networks, hierarchies, transitions, annotations, responsive video variants, frame-by-frame browser rendering, and visualization QA; not for generic dashboards, unsourced decorative charts, or ordinary motion graphics without a data contract.

6k tokens
Deepgram Speech
generative-media-skills

Use for Deepgram speech and voice production workflows: speech-to-text transcription, live captions, diarization, audio intelligence, Aura text-to-speech, Flux and Voice Agent live audio, model selection, cost/limits/privacy checks, artifact custody, and production QA.

9k tokens
Deepmotion Animate 3d
generative-media-skills

Use DeepMotion Animate 3D for markerless video-to-3D human motion capture through the cloud portal, sales-gated API, or real-time SDK; covers capture planning, body/hand/face/multi-person limits, custom characters, retargeting, refinement, exports, DCC/game-engine handoff, QA, pricing, privacy, consent, and rights checks.

12k tokens
Dialogue Editing Adr
generative-media-skills

Provider-independent dialogue editing and ADR direction for AI agents producing generated videos, films, ads, avatar clips, explainers, localization/dubbing, podcasts, recuts, and social content. Use for dialogue prep, repair, room tone, take comping, sync, ADR cueing, dubbing direction, pronunciation, de-essing, plosives, mouth clicks, voice continuity, synthetic voice consent boundaries, captions/transcripts handoff, mix handoff, loudness, intelligibility, accessibility, and QA.

9k tokens
D ID Avatar Video
generative-media-skills

Use D-ID to plan, generate, stream, localize, and QA avatar/talking-head videos, including V2 Photo Avatar Talks, V3 Pro/Instant Avatar Clips, V4 Expressive Avatar Scenes, D-ID Agents, voices, consent, lifecycle, safety, and artifact custody.

9k tokens
Documentary Montage Production
generative-media-skills

Provider-independent workflow for AI agents producing documentary-style montages, archive-driven timelines, interview-supported sequences, nonprofit or advocacy shorts, historical explainers, and hybrid generated/archive edits. Use for editorial thesis development, research/source logs, fact-checking, archive rights, synthetic reenactment disclosure, chronology, interview selects, lower thirds, generated b-roll direction, music restraint, captions, sensitive subjects, review escalation, delivery variants, and documentary QA.

10k tokens
Ecommerce Product Imagery
generative-media-skills

Provider-independent ecommerce product imagery production for marketplace listings, product pages, hero shots, lifestyle scenes, comparison charts, infographics, variant/packaging images, virtual try-on or placement mockups, ads, and AI-assisted product visuals. Use when an agent must plan, generate, edit, review, localize, package, or QA truthful commercial product images while managing SKU continuity, marketplace requirements, claims, rights, releases, accessibility, and approval.

11k tokens
Editing Montage
generative-media-skills

Provider-independent editing and montage direction for AI agents producing generated videos, ads, trailers, social clips, explainers, product videos, documentary-style pieces, and music- or beat-driven cuts. Use when planning, revising, QAing, or handing off an edit: story structure, shot selection, continuity, montage theory, rhythm and pacing, J/L cuts, match cuts, cutaways, transitions, beat sync, platform timing, source/generated asset management, revision workflow, and composition handoff.

11k tokens
Educational Animation Production
generative-media-skills

Provider-independent production workflow for AI agents creating educational animated lessons, classroom explainers, STEM visualizations, history/social-science animations, training modules, microlearning clips, whiteboard-style explainers, diagram-driven videos, and generated or assembled instructional animation assets. Use when planning, scripting, storyboarding, generating, reviewing, localizing, or QAing animation whose primary purpose is learning.

10k tokens
Elevenlabs Agents
generative-media-skills

Build, configure, and ship production voice agents on the ElevenLabs Agents platform (branded "ElevenAgents," formerly "Conversational AI"). Use when an agent must design a spoken-conversation system prompt, pick an LLM and TTS voice/model for a real-time voice bot, wire client/server/system tools and knowledge-base RAG, connect a channel (WebRTC/WebSocket SDK, embeddable widget, or telephony via Twilio/SIP), tune turn-taking and latency, set up simulation testing and evaluation, or reason about pricing, concurrency, and the consent/disclosure obligations of a synthetic voice talking to real people. Not for plain non-conversational text-to-speech, dubbing, or music generation.

11k tokens
Elevenlabs Dubbing Voice Conversion
generative-media-skills

Use for ElevenLabs provider-specific dubbing, localization, voice changer, speech-to-speech, and voice isolation/enhancement workflows. Applies when an agent must prepare source media, choose ElevenLabs dubbing versus voice conversion, manage language/speaker/timing decisions, handle voice identity and consent, call or plan against documented ElevenLabs APIs, export dubs/transcripts, estimate cost/limits, or quality-control localized speech while avoiding generic text-to-speech guidance.

9k tokens
Elevenlabs Music
generative-media-skills

Generate and iterate music with ElevenLabs Eleven Music for production deliverables. Use when planning, prompting, API-calling, editing, inpainting, reviewing, or licensing AI-generated songs, instrumental beds, ad music, soundtrack cues, video-to-music scores, stems, or music assets using the ElevenLabs Music API or ElevenCreative Music product.

7k tokens
Elevenlabs Scribe
generative-media-skills

Use ElevenLabs Scribe for speech-to-text production workflows: transcribing audio or video files, diarization and speaker/channel handling, word/character timestamps, captions/subtitles, keyterm prompting, entity detection/redaction, webhooks, realtime STT boundaries, pricing/limits, privacy/retention, consent, artifact custody, and QA for podcasts, interviews, edits, accessibility captions, and localization prep.

9k tokens
Elevenlabs Sound Effects
generative-media-skills

Produce, direct, generate, QA, and integrate non-speech audio with ElevenLabs Text to Sound Effects / Sound Effects API. Use when an agent needs custom SFX, foley, ambience, loops, stingers, impacts, UI sounds, musical one-shots, trailer braams, or sound-design layers for video, ads, social edits, games, apps, podcasts, audiobooks, or interactive media using ElevenLabs.

9k tokens
Elevenlabs Tts
generative-media-skills

Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken text; controlling pronunciation, pacing, emotion, and code-switching; using TTS APIs; cloning or designing a voice with consent; repairing artifacts; mastering deliverables; or evaluating generated speech. Excludes music, sound-effect generation, conversational-agent design, and general speech recognition except transcription used to verify TTS.

14k tokens
Explainer Video Production
generative-media-skills

Provider-independent explainer video production for AI agents creating educational, product, concept, nonprofit/public-service, onboarding, training, animated, mixed-media, and short social explainers. Use when an agent must turn a topic, brief, research corpus, product, process, policy, dataset, or complex idea into a factual, accessible, narrated, captioned, storyboarded, generated-media-ready explainer video with claim review, localization, pacing, variants, and QA.

11k tokens
Fashion Campaign Production
generative-media-skills

Provider-independent production workflow for AI-assisted fashion campaigns, lookbooks, editorial fashion films, ecommerce/editorial hybrids, social cutdowns, virtual try-on assets, generated model/lifestyle ads, and retouched fashion visuals; use when an agent must plan, prompt, produce, review, localize, or deliver fashion imagery or video while preserving garment truthfulness, model consent, fit representation, claims substantiation, platform compliance, accessibility, and fashion QA.

12k tokens
Ffmpeg Media Finishing
generative-media-skills

Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables. Use when finalizing images, image sequences, video, audio, captions, overlays, social/export variants, checksums, manifests, delivery specs, and QA with ffmpeg/ffprobe, including transcode versus stream-copy decisions, scaling, frame-rate handling, color metadata caution, loudness normalization, subtitle burn-in or sidecars, concat/trim, GIFs, batch variants, and reproducible command reporting.

18k tokens scripts
Fish Audio Tts
generative-media-skills

>- Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically.

10k tokens
Food Beverage Content Production
generative-media-skills

Provider-independent production guidance for AI agents creating or retouching appetizing food and beverage images, recipe videos, restaurant/menu visuals, CPG ads, ecommerce product images, beverage pours, packaging/lifestyle shots, and social food clips, including truthfulness, labeling/claims discipline, styling, prompt translation, platform variants, accessibility, rights, alcohol, and QA.

13k tokens
Game Trailer Production
generative-media-skills

Provider-independent production workflow for AI agents creating game trailers and promos, including announcement, gameplay, launch, DLC/update, store-page videos, social cutdowns, and generated/engine-capture hybrid trailers. Use when planning, scripting, capturing, generating, editing, localizing, reviewing, or QAing game marketing videos where gameplay truthfulness, platform/store constraints, ratings, accessibility, claims, CTA metadata, and rights provenance matter.

10k tokens
Gemini Live Audio
generative-media-skills

Build and evaluate low-latency spoken, multimodal, and translation experiences with Google Gemini Live API and Gemini Enterprise Agent Platform Live API. Use when a media-production or voice-agent workflow needs real-time audio input/output, barge-in, voice configuration, live transcription, live translation, tool/function calling during a spoken session, WebSocket session design, latency QA, quotas/cost review, or Google/Vertex data-governance tradeoffs.

11k tokens
Generated Media Qa
generative-media-skills

Provider-independent quality assurance for AI-generated and AI-assisted media. Use when reviewing, accepting, revising, or reporting on images, video, audio, avatars, ads, product content, social clips, explainers, localization, mixed-source edits, captions, accessibility, provenance, model metadata, safety/policy, and delivery readiness.

18k tokens scripts
Google Cloud Speech
generative-media-skills

Use this skill when a media-production agent needs Google Cloud speech and voice services for transcription, captions/subtitles, long-form audio analysis, live caption planning, Text-to-Speech narration, Chirp 3 / Gemini-TTS voice selection, consented custom voice workflows, dubbing/localization planning, pricing/quotas/region checks, or speech-related safety and data-governance decisions.

9k tokens
Google Cloud Vision
generative-media-skills

Use this skill when an agent needs Google Cloud Vision API for still-image understanding: labels, object localization, OCR, document text, SafeSearch, image properties, crop hints, web detection, batch annotation, Cloud Storage based pipelines, confidence evaluation, privacy, quotas, cost, and QA. Do not use it for Gemini multimodal reasoning, video analysis, image generation, custom model training, product catalog search design, or human reference/authenticity review.

11k tokens
Google Gemini Image
generative-media-skills

Plan, generate, edit, and quality-check still images with Google's Gemini native image models (Nano Banana), including Gemini 3.1 Flash Image, Gemini 3.1 Flash-Lite Image, Gemini 3 Pro Image, and legacy Gemini 2.5 Flash Image. Use for model and API-route selection, conversational editing, multi-reference composition, thinking and Google Search grounding, resolution/aspect-ratio control, exact-text and brand workflows, SynthID/provenance, safety, privacy, rights, and migration from legacy Gemini or Imagen image endpoints. Do not use for Veo/video generation or Gemini Live/audio work.

15k tokens
Google Gemini Omni Video
generative-media-skills

Generate, reference, and conversationally edit short videos with Google's Gemini Omni Flash through the Gemini Developer API or Gemini Enterprise Agent Platform. Use when a task specifically needs Gemini Omni video, multimodal reference roles, time-directed prompting, audio-aware generation, Interactions API state, or a precise comparison with Veo, Nano Banana, consumer Gemini, or gateway routes.

15k tokens
Google Lyria
generative-media-skills

Use Google Lyria music generation models through Google Cloud / Vertex AI / Gemini Enterprise Agent Platform for production music beds, songs, vocal tracks, image-conditioned music, short-form social audio, ad music, and video scoring. Covers Lyria 2, Lyria 3 Clip, and Lyria 3 Pro model selection, prompt construction, rights/privacy checks, SynthID/C2PA provenance, artifact custody, and QA.

8k tokens
Google Veo
generative-media-skills

Direct production with Google DeepMind's Veo video-generation family across the Gemini Developer API and Google Cloud Gemini Enterprise Agent Platform (formerly Vertex AI). Use for Veo route and model selection, text/image/reference/first-last-frame/extension workflows, native-audio prompting, camera direction, async API operation, troubleshooting, QA, and responsible commercial content production.

15k tokens
Gsap Animation Composition
generative-media-skills

Production guidance for authoring, integrating, and reviewing GSAP animation in browser-rendered media. Use for deterministic kinetic typography, SVG drawing and morphing, motion paths, FLIP transitions, responsive motion systems, or frame-seekable GSAP timelines inside HTML video composition and React-based renderers; not for ordinary CSS transitions or generic frontend development.

8k tokens
Hedra Character Video
generative-media-skills

Produce Hedra character, talking-avatar, lip-sync, motion-avatar, and live avatar work. Use when an AI agent needs to choose Hedra models, prepare image/audio/script inputs, direct character performance, call or plan around the Hedra API or Studio, evaluate avatar/lip-sync results, handle consent/rights/privacy, or troubleshoot Hedra character video production.

7k tokens
Heygen Avatar Video
generative-media-skills

Produce HeyGen avatar videos with Direct Video, Video Agent, Digital Twin, Avatar Realtime, and related avatar/voice/asset APIs. Use for provider routing, avatar and voice selection, script-to-video workflows, lip-sync from audio, backgrounds/scenes, callbacks and polling, consent and rights checks, localization handoffs, artifact custody, QA, and troubleshooting HeyGen avatar-video productions.

8k tokens
Higgsfield Video
generative-media-skills

Produce video (and its supporting stills) through Higgsfield (higgsfield.ai), a multi-model creative platform that fronts third-party video models (Kling, Veo, Seedance, Wan, MiniMax, and, until retirement, Sora) plus Higgsfield's own control layer — Soul / Soul ID image generation, the DoP camera-preset engine, and the Cinema Studio filmmaking suite. Use this skill when a request routes generation through Higgsfield specifically, when the user wants aggregator-level model switching and cinematic camera/effect presets in one workspace, when planning character consistency via Soul ID, or when deciding whether to go through Higgsfield versus direct to an underlying model provider. Covers model selection, per-model prompt strategy, credit economics, the (partial) API surface, capability limits, and rights/consent policy for generated and uploaded real-person footage.

10k tokens
Hume Evi
generative-media-skills

Build production realtime voice agents with Hume's Empathic Voice Interface (EVI): speech-to-speech sessions, empathic/prosodic response design, EVI 3 versus EVI 4-mini selection, WebSocket and SDK integration, tools/function calling, interruption/barge-in, turn-taking, transcripts/audio artifacts, pricing/limits, privacy, safety, consent, and QA.

11k tokens
Hume Octave
generative-media-skills

Use Hume Octave for emotionally expressive speech and voice production: text-to-speech, voice design, voice cloning, voice conversion, streaming/realtime TTS, multilingual narration, dialogue continuity, timestamps/lip-sync, safety/rights review, and production QA.

9k tokens
Hyperframes Video Composition
generative-media-skills

Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos. Use for HyperFrames composition planning, scene architecture, media custody, animation/timing, captions/audio, deterministic preview/render QA, accessibility/flashing checks, provenance ledgers, and delivery handoff.

15k tokens scripts
Ideogram Image
generative-media-skills

Generate, remix, edit, inpaint, reframe, background-process, describe, layerize, and upscale images with the Ideogram Developer API. Use when Codex must integrate Ideogram 4.0 or 3.0, render reliable typography or structured layouts, use style or character references, build synchronous or webhook workflows, migrate legacy Ideogram endpoints, or reason about Ideogram API pricing, safety, privacy, and production QA.

14k tokens
Image Generation Gateways
generative-media-skills

Select, integrate, and operate multi-model image-generation gateways with model-specific schema discovery, version policy, asynchronous jobs, webhooks, spend approval, safe inputs and artifacts, data-governance review, and billing observability. Use for comparing or building against fal.ai, Replicate, or Together AI image APIs, including controlled failover; do not use for direct model-provider APIs, local inference, training, dedicated endpoint provisioning, video generation, or general image editing.

31k tokens scripts
Immersive Spatial Video Production
generative-media-skills

Use this skill to plan, direct, finish, and QA provider-independent mono or stereo 180-degree, 360-degree, spherical, spatial, wide-FOV, and immersive video. It covers format choice, capture rigs, stitching, nadir and seam management, stereoscopic comfort, attention direction, stabilization, ambisonic and spatial audio, titles, captions, accessibility, edit, color, VFX, projection metadata, Vision Pro, YouTube, headset delivery, device QA, and privacy, location, and likeness rights. Do not use it for interactive XR apps, volumetric/world-model generation, game-engine experiences, or provider-specific video generation prompts.

12k tokens
Kling Advanced Lip Sync
generative-media-skills

Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.

3k tokens
Kling Kolors Image
generative-media-skills

Plan, generate, edit, reference, and quality-control still images with Kling AI's hosted IMAGE surfaces or Kuaishou's open Kolors checkpoints. Use for Kling IMAGE 3.0/3.0 Omni/O1/2.1 web, official Kling CLI/MCP or Open Platform integration, and local Kolors text-to-image, image-to-image, IP-Adapter, ControlNet, inpainting, or LoRA work. Do not use for Kling video, avatars, lip-sync, unofficial gateway APIs, or for assuming that current hosted Kling Image models are identical to the 2024 open Kolors weights.

17k tokens
Kling Video
generative-media-skills

Plan, prompt, generate, reference, edit, motion-control, and quality-review videos with Kuaishou's direct Kling AI video platform and API, especially Kling VIDEO 3.0, 3.0 Omni, 3.0 Turbo, native audio, multi-shot, elements, start/end frames, and image/video references. Use for first-party Kling video production or API integration, not Kolors image generation, avatars, standalone audio, consumer-only membership advice, or third-party Kling gateways.

14k tokens
Kokoro Tts
generative-media-skills

>- Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad. Use when synthesizing speech offline, on-device, in the browser, or on your own server without per-character API cost — for high-volume narration, audiobooks, privacy-constrained pipelines, and prototyping. Covers running Kokoro via the Python `kokoro` package, ONNX / kokoro-js in the browser, and the OpenAI-compatible Kokoro-FastAPI wrapper; picking voices and language codes; blending voices; chunking long text; controlling pronunciation via misaki/espeak-ng; and judging when Kokoro is the right tool versus when its lack of voice cloning, narrow emotional range, and weaker non-English quality mean you should reach for a hosted or larger model instead. Not for voice cloning, expressive/emotional character performance, or high-fidelity multilingual work — say so and route elsewhere.

8k tokens
Leonardo Image
generative-media-skills

Create, edit, guide, upscale, and quality-control still images with Leonardo.Ai's official Production API, including native Lucid and Phoenix models, uploaded or generated references, image-to-image, ControlNet guidance, realtime-canvas inpainting, Pro and Universal upscalers, and custom Elements or models. Use for Leonardo image API implementation, async polling or webhooks, dry-run and cost governance, secure artifact handling, prompt iteration, or production troubleshooting. Do not use for Leonardo video, 3D generation, unofficial wrappers, or third-party gateway APIs.

14k tokens
Lighting Direction
generative-media-skills

Provider-independent lighting direction for AI-generated images, video clips, ads, film scenes, avatars, product shots, interviews, beauty work, and social content. Use when an agent must specify, translate, maintain, troubleshoot, or QA motivated lighting, key/fill/back/rim/practical lights, hard or soft light, color temperature, exposure, contrast, time-of-day, continuity, reference strategy, or safety/accessibility risks in generated media.

9k tokens
Livestream Event Production
generative-media-skills

Use this skill to plan, direct, troubleshoot, and hand off provider-independent live or hybrid livestream event productions, including run of show, crew, capture, audio, graphics, switching, guests, encoding, ingest, transport choices, latency, redundancy, accessibility, moderation, monitoring, incidents, recording, and delivery.

11k tokens
Localization Dubbing Production
generative-media-skills

Provider-independent localization and dubbing production direction for AI agents producing translated videos, dubbed ads, social clips, explainers, avatar videos, documentaries, training media, and multi-market campaigns. Use when planning or reviewing subtitle-versus-dub strategy, translation adaptation, glossary/style-guide setup, casting or voice matching, lip-sync or phrase-sync direction, synthetic dubbing boundaries, captions/subtitles handoff, accessibility, platform delivery, client review, or localization QA.

11k tokens
Lottie Animation Delivery
generative-media-skills

Production guidance for assessing, exporting, packaging, integrating, capturing, validating, and handing off Lottie vector animations. Use for Bodymovin/Lottie JSON, dotLottie archives, renderer and player compatibility, fonts and glyphs, image assets, markers and segments, deterministic video capture, responsive sizing, optimization, accessibility, and cross-platform QA; not for general After Effects animation craft or live-action/video delivery.

7k tokens
Ltx 2 Video
generative-media-skills

Generate and edit synchronized audio-video with LTX-2.3 using the official hosted LTX API or official local/open-weight LTX-2 repository. Use for text-to-video, image-to-video, audio-to-video, retake, extend, HDR conversion, local inference, checkpoint selection, quantization, and LTX-specific production planning.

15k tokens
Luma Photon
generative-media-skills

Use Luma AI's first-party Photon and Photon Flash image API for text-to-image, image/style/character references, image modification, asynchronous generation, secure artifact handling, production prompting, cost control, and policy-aware delivery. Trigger for Luma Photon API integration or production work; do not use for Ray video, UNI image models, the Luma consumer app, or third-party Photon gateways.

14k tokens
Luma Ray Video
generative-media-skills

Direct and operate Luma Ray video generation through the current Luma Agents API and distinguish it from the consumer Luma App and legacy Dream Machine API. Use for Ray 3.2 text/image/multi-keyframe video, edits, extends, reframes, HDR/EXR, pricing and approvals, async delivery, migration, rights, privacy, and production QA.

7k tokens
Manim Explainer Animation
generative-media-skills

Provider-independent production workflow for creating Manim-based explainer animations. Use when an agent must plan, code, render, QA, or hand off precise math, science, data, diagram, algorithm, or concept animations with Manim, including storyboard, scene architecture, formulas, coordinate systems, voiceover timing, accessibility, reproducibility, provenance, and delivery decisions.

11k tokens
Media Provenance Rights
generative-media-skills

Provider-independent provenance, rights, consent, disclosure, and release governance for AI-generated media. Use when producing or reviewing generated images, video, audio, ads, avatars, product content, social clips, documentaries, localization, music, sound, or campaign assets that need source-rights checks, model/provider term checks, likeness or voice consent, trademark/logo review, copyrighted-character/style risk triage, music licensing, C2PA/content credentials, provenance ledgers, release notes, platform disclosures, jurisdiction/client caveats, or rights QA.

9k tokens
Media Qc Delivery
generative-media-skills

Provider-independent media quality control and delivery direction for generated images, video, audio, ads, product content, social clips, broadcast and streaming assets, localization packages, and mixed-source edits. Use when an agent must define final export specs, inspect technical and perceptual quality, check captions/audio/color/accessibility/provenance/rights, create delivery sheets, version filenames, checksums, manifests, package deliverables, set acceptance or rejection criteria, and perform final delivery QA before handoff to a platform, broadcaster, streamer, client, or ad trafficker.

15k tokens scripts
Meshy 3d
generative-media-skills

>- Produce 3D assets with Meshy (meshy.ai) through its REST API and web platform — text-to-3D, image-to-3D, multi-image-to-3D, AI texturing/PBR, remesh/topology control, UV, and auto-rigging/animation. Use this skill when an agent must generate a mesh from a prompt or reference image, control polygon count and topology, apply or restyle textures, rig a humanoid for animation, export to GLB/FBX/OBJ/USDZ/STL for Unity/Unreal/Blender, drive Meshy's asynchronous job API (auth, polling, streaming, webhooks, credits, rate limits), review generated meshes for game/film/print readiness, or decide whether Meshy is the right tool at all. Not for editing an existing hand-modeled asset's geometry, hard-surface CAD with exact dimensions, or guaranteed-watertight manufacturing parts.

10k tokens
Midjourney Image
generative-media-skills

Plan, prompt, iterate, edit, migrate, and quality-check still-image work made with Midjourney's documented website and Discord interfaces. Use for Midjourney model and parameter selection, image/style/Omni references, Draft or Conversational workflows, typography-aware production, privacy and rights review, or troubleshooting a Midjourney image pipeline; do not use it to invent or automate an unsupported Midjourney API.

15k tokens
Midjourney Video
generative-media-skills

Plan and direct Midjourney Video V1 image-to-video work through the official website or Discord, with human operator handoff, motion prompts, start/end frames, loops, extensions, resolution and batch budgeting, privacy, rights, safety, and delivery QA. Use when creating Midjourney videos or when determining whether a requested API, automation, plan, or production workflow is supported.

9k tokens
Minimax Hailuo Video
generative-media-skills

Use MiniMax's first-party Hailuo video API safely and reproducibly across global and mainland-China platforms. Covers current text-to-video, image-to-video, first/last-frame, and subject-reference models; prompt and camera craft; region/account routing; exact cost approval; asynchronous jobs and callbacks; artifact provenance; and rights, likeness, disclosure, and data-governance review. Use when designing, implementing, debugging, or evaluating direct MiniMax/Hailuo video-generation API workflows. Do not use for the consumer Hailuo site, Video Agent templates, gateways, or unrelated MiniMax modalities.

17k tokens
Minimax Music
generative-media-skills

Produce music with MiniMax Music 2.6 and MiniMax cover/lyrics APIs for songs, instrumentals, AI-generated lyrics, reference-audio covers, video/social/ad soundtracks, artifact custody, rights checks, and production QA.

7k tokens
Minimax Speech
generative-media-skills

Use this skill when producing speech, narration, dubbing, localization, advertising voice, voice-clone previews, or interactive voice with MiniMax speech/audio APIs. It covers MiniMax T2A HTTP, WebSocket streaming, async long-form TTS, Speech 2.8/2.6/02 model selection, system and custom voices, rapid voice cloning, voice design, pronunciation/language/emotion controls, pricing and limits, artifact custody, consent and rights checks, and production QA.

9k tokens
Moonvalley Marey
generative-media-skills

>- Produce video with Moonvalley's Marey model family (Marey Realism v1.5) — a filmmaker-oriented, 1080p/24fps generative video model marketed as trained exclusively on licensed data. Use this skill when a task asks for Marey or Moonvalley specifically, when a brief demands brand-safe / legally reviewed AI video for commercial or studio work, or when the request needs Marey's director-style controls (camera trajectory, motion transfer, pose transfer, keyframing, reference conditioning). It covers what "commercially safe" really means versus the actual terms of service, access routes (Moonvalley web / Voyager, the fal API, ComfyUI), pricing, capability limits, prompt and reference strategy, where Marey fits production versus where other models win, iteration, and quality review. Do not use it for audio generation, for photoreal talking-head dialogue, or as a generic "best video model" default.

8k tokens
Motion Graphics Direction
generative-media-skills

Provider-independent motion graphics direction for AI agents producing title sequences, lower thirds, explainer graphics, product UI overlays, social ads, kinetic typography, diagrams, data callouts, brand films, and video post. Use when planning, prompting, storyboarding, specifying, implementing, reviewing, or QAing motion graphics across live footage, generated video, UI capture, data visualization, brand systems, typography, accessibility, claim integrity, runtime handoff, and iteration.

10k tokens
Move AI Motion Capture
generative-media-skills

Use this skill when planning, capturing, processing, validating, troubleshooting, or exporting markerless human motion capture with Move AI products, including Move One single-camera workflows and Genesis multi-camera studio systems. It covers provider selection, setup, calibration, wardrobe and lighting constraints, documented multi-person operation, processing, cleanup, skeleton and retargeting choices, DCC/game-engine export, data management, consent, and privacy. Do not use it for general keyframe animation, non-Move mocap systems, 3D asset generation, or generic animation cleanup unrelated to Move AI outputs.

11k tokens
Music Supervision Scoring
generative-media-skills

Provider-independent music supervision and scoring direction for AI agents producing generated videos, ads, trailers, explainers, product films, social clips, documentaries, game trailers, and branded content. Use when planning, prompting, selecting, licensing, mixing, repairing, or quality-reviewing music, score, temp tracks, needle drops, AI-generated tracks, cue sheets, stems, loudness, loopability, platform music constraints, and rights-aware audio handoff for visual media.

9k tokens
Music Video Production
generative-media-skills

Provider-independent production workflow for AI agents creating music videos, lyric videos, visualizers, performance/narrative/dance videos, social cutdowns, album or single launch assets, and generated-media clips synchronized to songs. Use when planning, directing, prompting, editing, reviewing, or delivering video content built around a recorded song, artist performance, lyrics, beat timing, choreography, platform variants, music rights, likeness consent, accessibility, or synthetic-media disclosure.

10k tokens
Neural Reality Capture
generative-media-skills

Use this skill for provider-independent reality capture that turns photographed or video-derived real places, objects, aerial sites, interiors, or people into photogrammetric meshes, NeRFs, or 3D Gaussian splats. It guides representation choice, capture planning, calibration and scale, reconstruction or training, cleanup, compression and LOD, coordinate and interchange handoff, visual/geometric QA, archival records, and property, location, people, and cultural rights checks. Do not use it for text-to-3D generation, synthetic asset invention, guaranteed metrology, or broad downstream finishing after handoff.

12k tokens
Nvidia Cosmos Video
generative-media-skills

Select, run, and govern NVIDIA Cosmos world-video generation across Cosmos 3 Generator, Predict2.5, Transfer2.5, downloadable checkpoints, self-hosted NIMs, and hosted preview surfaces. Use for text/image/video-to-world, controlled world transfer, multiview or action-conditioned physical-AI video, local checkpoint pinning, hardware planning, inference approval, and output validation; do not use for Cosmos Reason/Embed-only tasks.

9k tokens
Nvidia Maxine Audio Effects
generative-media-skills

Use NVIDIA Maxine / NVIDIA AI for Media audio effects for speech cleanup and enhancement in live or offline media pipelines, including Background Noise Removal, denoise+dereverb, room echo removal, acoustic echo cancellation, audio super-resolution, Studio Voice, Speaker Focus, and Voice Font. Use when selecting, integrating, prompting around, QAing, or privacy-reviewing NVIDIA AFX SDK or BNR NIM workflows for conferencing, broadcast, podcast, avatar, ad, social, and video post-production audio.

8k tokens
Nvidia Speech Nim
generative-media-skills

Use NVIDIA Speech NIM microservices for speech production and voice workflows, including self-hosted ASR/STT, TTS, text translation, speech-to-speech pipeline design, Riva Python client integration, deployment planning, GPU/runtime sizing, licensing/privacy review, and production QA for transcription, captioning, voice agents, localization, and synthetic voice output.

9k tokens
Odyssey Interactive Video
generative-media-skills

>- Plan and scope production work with Odyssey's real-time interactive video world models (odyssey.ml — Odyssey-1/Odyssey-2/Starchild-1/Agora-1). Use when a request involves generating video that a user steers live with keyboard, text, speech, or controller input (streamed frames that respond in the moment), or when someone must decide between Odyssey-style interactive generation and offline clip generators (Sora/Veo/Kling) or downloadable 3D / game engines. Covers what is verified versus announced, current access routes, latency and coherence limits, feasible-now versus speculative use cases, pricing/access terms, and content/rights considerations. This is a fast-moving research-stage product; the skill teaches honest scoping, not a promise of production reliability.

9k tokens
Openai Audio
generative-media-skills

Produce and understand audio with OpenAI request-based audio APIs and audio-capable chat models, including text-to-speech, transcription, translation, multimodal audio input/output, model routing, prompt and performance direction, artifact custody, approval gates, and safety/rights review. Use for non-realtime OpenAI audio production; route continuous live voice agents to the separate realtime voice skill.

11k tokens
Openai Gpt Image
generative-media-skills

Generate, edit, composite, stream, and production-review still images with OpenAI GPT Image models. Use when an agent must choose between GPT Image 2 and legacy GPT Image/DALL-E integrations, select the Image API or Responses API, build prompts and reference-image workflows, use masks or multiple inputs, preserve identity or brand details, render in-image text, handle output files and partial images, estimate cost and rate limits, migrate deprecated image code, or apply OpenAI image safety, consent, privacy, provenance, and rights requirements. Do not use for OpenAI video or audio generation.

12k tokens
Openai Realtime Voice
generative-media-skills

Build production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket transport, voice activity detection, tool/function calling, prompt design, logging, consent, privacy, safety, latency, and cost controls. Use for speech-to-speech agents and live voice UX; do not use for separate request-based transcription, text-to-speech, or offline audio generation work.

10k tokens
Performance Direction
generative-media-skills

Provider-independent performance direction for generated video, avatar/spokesperson clips, animation, ads, film scenes, social content, and voice-led media. Use when an agent must cast or direct synthetic performers, avatars, animated characters, talking heads, AI video characters, or voice performances; plan acting beats, objective/obstacle/action, blocking, gesture, eye-line, facial expression, body language, lip-sync, voice alignment, rehearsal prompts, iteration, continuity, consent, and performance QA.

8k tokens
Pika Video
generative-media-skills

>- Produce short-form video with Pika (pika.art) — its effect-driven tools (Pikaffects, Pikadditions, Pikaswaps, Pikaframes, Pikascenes, Pikatwists) and audio-driven lip sync (Pikaformance), plus text-to-video and image-to-video. Use when a task calls for playful, stylized, meme-able social clips, or for applying generative effects and object swaps to real uploaded footage, and when deciding whether Pika is the right engine versus a higher-fidelity model. Covers model versions, credit/plan economics, the fal.ai API route, prompt construction, iteration/repair, capability limits, and consent/rights rules for uploaded real-person media. Not a general "best video model" chooser.

8k tokens
Podcast Production
generative-media-skills

>- Produce audio-first podcast episodes with generative tools — design the show and episode format (interview, narrative, news brief, two-host conversational), write scripts for the ear, decide between fully-synthetic and hybrid (recorded human + synthetic) production and meet the disclosure duty each triggers, cast and direct multi-voice TTS for consistency and chemistry, build episode structure (cold open, intro/outro, segments, ad slots), assemble and edit, clear music and SFX rights, hit podcast loudness and delivery standards, ship metadata/chapters/transcripts through RSS, and pass platform AI-content policy and QA before publishing. Use this skill whenever the deliverable is a podcast episode or a podcast show bible, or when an agent must make production, rights, disclosure, or delivery decisions for spoken-word audio. Not for music tracks, single-voice notification prompts, or video where picture leads (route audio-for-video to a video skill).

12k tokens
Precise Video Description
generative-media-skills

Provider-independent production guidance for converting observed video into precise, objective, temporally ordered language. Use for shot descriptions, searchable metadata, dataset captions, reference logs, generation-prompt handoffs, and analysis across subject, scene, motion, spatial, and camera aspects; not for deciding what to shoot, accessibility captions, creative interpretation, or provider-specific video analysis APIs.

4k tokens
Procedural Canvas Animation
generative-media-skills

Provider-independent production guidance for deterministic Canvas 2D and p5.js animation. Use for particles, fields, trails, weather, procedural textures, generative geometry, and lightweight 2D simulations that need fixed media dimensions, seeded repeatability, transparent compositing, aspect variants, performance QA, or frame-addressable rendering.

4k tokens
Product Ad Production
generative-media-skills

Diagnose, design, script, produce, adapt, and quality-control product video advertisements from evidence-backed briefs. Use for physical-product, software, app, service, ecommerce, direct-response, brand-response, demonstration, testimonial, launch, and variant ad work where product truth, legibility, claims, platform delivery, and measurable iteration matter.

16k tokens
Production Design Direction
generative-media-skills

Provider-independent production design direction for AI-generated media. Use when translating narrative, brand, advertising, product, social, film, or worldbuilding intent into sets, locations, props, color and material palettes, era research, graphic and environmental design, continuity bibles, image/video prompt planning, reference strategy, production constraints, iteration notes, handoff specs, and QA.

10k tokens
Qwen3 Tts
generative-media-skills

Produce text-to-speech with Alibaba/Qwen Qwen3-TTS through DashScope/Model Studio or open-weight Qwen3-TTS checkpoints. Use when an agent must choose Qwen3-TTS models, voices, realtime versus non-realtime synthesis, voice cloning, voice design, instruction/prosody controls, multilingual or dialect speech, audio artifact handling, regional routing, pricing/privacy/licensing constraints, or production QA for narration, characters, assistants, audiobooks, ads, localization, or voice-enabled media.

9k tokens
Real Estate Content Production
generative-media-skills

Provider-independent real estate media production for listing photos, virtual staging, property videos, walkthrough reels, floor-plan explainers, neighborhood clips, renovation concepts, agent ads, and generated or retouched property visuals. Use when creating, editing, reviewing, prompting, packaging, or approving real estate marketing content where property truthfulness, listing/MLS compliance, fair-housing risk, staging disclosure, privacy, accessibility, platform variants, or real-estate QA matter.

14k tokens
Recraft Image Design
generative-media-skills

Design and produce raster images, native SVG artwork, brand-controlled visuals, and documented image edits with Recraft's API. Use for Recraft model and route selection, exact request construction, image-to-image, inpainting, outpainting, background work, vectorization, remix/exploration, palette and typography control, API migration, safety and rights review, and production QA.

14k tokens
Reference Media Analysis
generative-media-skills

Provider-independent reference media analysis for generated-media production. Use when an agent must analyze reference images, videos, audio, style boards, product shots, brand assets, mood boards, storyboards, performances, prior cuts, or client examples and translate them into safe, non-copying direction, prompts, QA criteria, provenance records, and handoff notes for image, video, audio, avatar, or post-production agents.

11k tokens
Remotion Video Composition
generative-media-skills

Provider-independent production workflow for assembling generated or sourced media into React/Remotion videos. Use when an agent must plan, build, render, review, variant-render, or hand off Remotion compositions with media custody, captions, audio, animation, accessibility, provenance, QA, and delivery requirements.

10k tokens
Resemble Chatterbox
generative-media-skills

Use Resemble AI Chatterbox for text-to-speech and voice-cloning workflows, including local open-weight Chatterbox, Chatterbox Multilingual, Chatterbox Turbo, and Resemble-hosted Chatterbox API routes. Apply when an agent must choose a Chatterbox variant, prepare reference-voice inputs, control emotion/paralinguistic delivery, plan multilingual TTS, handle PerTh watermarking, or make production decisions about consent, privacy, licensing, latency, and audio QA.

7k tokens
Runway Image
generative-media-skills

Generate, edit, and iterate still images with Runway's official API, especially native Gen-4 Image and Gen-4 Image Turbo reference workflows. Use when implementing Runway text-to-image, reference-driven image generation or natural-language image edits, task polling, secure artifact handling, production retries, cost controls, or Runway image API QA. Do not use for Runway video generation or third-party gateway APIs.

11k tokens
Runway Video
generative-media-skills

Build and operate production-safe Runway API video generation, video editing, and character-performance workflows with native Runway models, explicit paid-call approval, duplicate-create protection, asynchronous task handling, secure media transfer, and evidence-aware governance.

23k tokens
Saas Product Demo Production
generative-media-skills

Provider-independent workflow for producing SaaS and software product demo videos, including app walkthroughs, feature launches, sales demos, onboarding clips, help-center explainers, website hero demos, and PLG assets. Use when an agent must plan, script, capture, synthesize, edit, localize, review, or QA privacy-safe software demo videos with substantiated claims, accessible captions, screen/cursor direction, CTA variants, and stakeholder approval.

13k tokens
Screen Demo Production
generative-media-skills

Provider-independent workflow for producing terminal, IDE, documentation, browser, desktop, and application demonstration videos. Use when choosing authentic capture, browser automation, synthetic UI, or hybrid treatment; scripting actions and narration; resetting demo state; protecting secrets and personal data; directing cursor/callouts; creating crop variants; and reviewing workflow truth, readability, accessibility, and provenance.

4k tokens
Seedance 2 0
generative-media-skills

Direct ByteDance Dreamina Seedance 2.0 Standard, Fast, and Mini video production across BytePlus ModelArk and verified gateways. Use for text/image/reference-to-video, native synchronized audio and dialogue, multimodal image-video-audio reference, first/last-frame animation, video editing or extension, model/gateway selection, prompt construction, async task handling, failure recovery, QA, and rights-safe production.

15k tokens
Social Short Production
generative-media-skills

Plan, produce, adapt, finish, deliver, and improve platform-ready short vertical social videos from a brief, script, transcript, or source media. Use for TikTok videos, Instagram or Facebook Reels, YouTube Shorts, LinkedIn vertical clips, and cross-platform social cutdowns when the work requires audience and objective definition, an honest hook and promise, short-form beat or shot planning, UI-safe layout, captions and accessibility, speech/music/SFX mixing, platform-specific exports, rights and AI or sponsorship disclosure, analytics-led retention iteration, or final QA and repair. Do not use as a provider-specific image, video, voice, or music generation guide.

13k tokens
Sound Design Foley
generative-media-skills

Provider-independent sound design and Foley direction for generated media. Use when an agent must plan, prompt, source, generate, edit, mix, hand off, repair, or QA sound effects, Foley, ambience, impacts, transitions, UI/product audio, emotional sound design, or accessibility-safe captions for videos, ads, trailers, animation, avatar clips, explainers, games/trailers, product films, and social edits.

9k tokens
Stability AI Image
generative-media-skills

Operate Stability AI image generation, image-to-image, edit, control, background, and upscale APIs safely and reproducibly. Use when selecting or calling Stable Image Ultra, Core, Stable Diffusion 3.5, v2beta image edit/control services, or Stability open-weight image models; when debugging schemas, moderation, rate limits, cost, lifecycle, licenses, privacy, provenance, or migrations; and when planning production QA for Stability-generated images.

13k tokens
Stable Audio
generative-media-skills

Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs, and reviewing generated audio for media projects.

7k tokens
Still Image Retouching Finishing
generative-media-skills

Provider-independent still-image post-production skill for RAW/rendered intake, nondestructive development, exposure, white balance, tone, color, ICC proofing, masking, repair, compositing, truthful retouching by genre, generative disclosure, sharpening, noise, resampling, variants, metadata, export, QA, rights, and ethical finishing. Use when an agent must plan, perform, direct, or review still photography retouching and finishing for print, web, social, archive, publication, or client delivery; do not use for image-generation APIs, video grading, or campaign-specific production strategy.

13k tokens
Storyboard Previsualization
generative-media-skills

Provider-independent storyboard and previsualization direction for AI agents turning briefs, scripts, treatments, ads, films, social videos, and generated-media concepts into shot-by-shot boards, animatics, previs plans, camera/blocking/movement direction, reference packs, AI image/video prompt plans, review workflows, continuity checks, and production handoffs.

9k tokens
Suno Music
generative-media-skills

>- Generate music with Suno (v5 / v5.5 era, 2026) and advise a production team on what they may legally do with the output. Use this skill when a user wants to create, extend, remaster, or stem-separate a track in Suno; craft Suno prompts (style field, lyrics/metatags, exclude-styles, personas, custom voices); choose a subscription tier for a given use; or get accurate answers about ownership, commercial rights, distribution/monetization on DSPs, the Warner Music Group settlement, the ongoing label/publisher litigation, and Suno's API situation. Do not use it for other music generators (Udio, ElevenLabs Music, Google Lyria) except by explicit contrast, and do not use it as a substitute for a lawyer on a specific commercial deal.

9k tokens
Sync Labs Lipsync
generative-media-skills

>- Generate AI lip sync and visual dubbing with Sync Labs (sync.so) — choose the right model (sync-3, lipsync-2, lipsync-2-pro, lipsync-1.9.0-beta, react-1), build the POST /v2/generate request, host inputs, handle async jobs via polling or webhooks, run batch dubbing/localization, and review output for sync accuracy and identity preservation. Use when an agent must re-voice, dub, translate, personalize, or re-time the mouth of a talking-head video (or animate a still face from audio), or must debug a failed/rejected Sync generation. Also covers consent, likeness, and rights obligations for editing a real person's face.

9k tokens
Synthesia Avatar Video
generative-media-skills

Produce presenter-led AI avatar videos with Synthesia Studio and Synthesia API. Use for Synthesia-specific avatar video planning, avatar/voice/language selection, scripted scenes, template personalization, API job lifecycle, localization/dubbing, brand/training/product workflows, consent/rights/safety checks, artifact custody, troubleshooting, and final QA.

9k tokens
Talking Head Podcast Recut
generative-media-skills

Provider-independent production workflow for recutting long talking-head, podcast, interview, webinar, panel, lecture, livestream, founder call, or customer-call footage into short clips, highlight reels, trailers, audiograms, captioned vertical videos, and social cutdowns. Use when an agent must ingest source recordings, transcribe and diarize speakers, select quotes without misrepresentation, preserve context, handle consent/rights/disclosures, clean jump cuts and audio, add captions, graphics, lower thirds, b-roll, platform variants, synthetic-media boundaries, review packets, delivery specs, and recut QA.

11k tokens
Tavus Replica Video
generative-media-skills

Use for producing Tavus AI-human videos and real-time avatar conversations with Tavus Faces/Replicas, PALs/Personas, async Video Generation, CVI conversations, consent-safe likeness workflows, webhooks, backgrounds, localization, and avatar QA.

9k tokens
Tencent Hunyuan3d
generative-media-skills

>- Generate 3D assets with Tencent's Hunyuan3D family — open-weight self-hosted models (Hunyuan3D-2.0/2.1 and HunyuanWorld / HY-World for scenes) and the hosted, closed API tiers (2.5, PolyGen, 3.0/3.1 via Tencent Cloud and third-party hosts). Use when a task involves image-to-3D or text-to-3D mesh generation, the shape-then-texture (DiT + Paint / PBR) workflow, choosing between self-hosting and a hosted API, checking the Community License's Territory (EU/UK/South Korea) and 1M-MAU commercial gate, sizing GPU/VRAM for local inference, ComfyUI integration, or reviewing and post-processing (retopology, UVs, decimation) the meshes these models produce. Not for image, video, or general LLM tasks.

9k tokens
Tencent Hunyuanvideo
generative-media-skills

Generate and operate Tencent Hunyuan video through the managed TokenHub HY-Video-1.5 API or official local HunyuanVideo repositories. Use for text-to-video, image-to-video, hosted job lifecycle, pricing and region planning, local checkpoint selection, hardware and acceleration, prompt design, licensing, safety, provenance, and delivery QA.

12k tokens
Threejs Scene Composition
generative-media-skills

Production guidance for planning, building, animating, capturing, and reviewing complete Three.js scenes for rendered media. Use for browser-native 3D product shots, title sequences, procedural worlds, glTF scene assembly, camera and lighting animation, shaders, post-processing, deterministic frame export, and Three.js render QA; not for modeling standalone assets, WebXR interaction, or generic website decoration.

6k tokens
Title Kinetic Typography
generative-media-skills

Provider-independent title design and kinetic typography direction for video, animation, ads, explainers, social hooks, brand films, lyric/text videos, title cards, main titles, lower thirds, captions-as-design, and other motion-led text. Use when an agent must plan, prompt, implement, hand off, iterate, or QA typography hierarchy, readability, timing, reveal/hold/exit, motion semantics, line breaks, safe areas, localization, brand constraints, accessibility, flashing/motion safety, or production specs for animated text.

9k tokens
Topaz Video Enhancement
generative-media-skills

>- Use when enhancing, upscaling, restoring, denoising, sharpening, deinterlacing, stabilizing, motion-deblurring, colorizing, SDR-to-HDR converting, or frame-rate/slow-motion converting video (and secondarily images) with Topaz Labs — whether cleaning up AI-generated video, compressed UGC, or old archival footage, or building a generate-low-res-then-upscale delivery pipeline. Covers choosing the right named Topaz model for a specific footage problem, setting parameters, driving the Topaz Platform (Video/Image) REST API's async job lifecycle, judging when enhancement helps versus when it introduces artifacts (over-smoothing, temporal shimmer, face hallucination), and QA of enhanced output. Do not use for originating new video content, non-Topaz upscalers, or purely editorial cuts.

9k tokens
Tripo 3d
generative-media-skills

>- Produce 3D assets with Tripo (tripo3d.ai / Tripo AI by VAST) through its OpenAPI and Studio platform — text-to-3D, image-to-3D, and multiview-to-3D generation; PBR texturing; auto-rigging and preset animation; retopology/low-poly and quad remesh; format conversion (GLB/FBX/OBJ/USDZ/STL/3MF) and engine import (Unity/Unreal/Blender/3D printing). Use when an agent must drive the async task-based Tripo API (create task, poll or webhook, download), choose a model version (v3.0 / v3.1 / P1), write prompts or prepare input images, control mesh/texture parameters, estimate credits and respect rate limits, review generated mesh quality (topology, UVs, poly count, texture fidelity), run iteration and repair workflows, or reason about the rights/licensing of Tripo-generated assets. Trigger on requests to generate, texture, rig, animate, retopologize, convert, or evaluate a 3D model with Tripo, or to integrate the Tripo API into a pipeline. Not for choosing a 3D renderer, editing meshes by hand, or non-Tripo generators.

9k tokens
Twelvelabs Video Understanding
generative-media-skills

>- indexing footage, semantic/visual search across an archive, generating descriptions, summaries, chapters, highlights, and tags from video, producing multimodal embeddings, or wiring TwelveLabs into a media-production pipeline (NLE panels, logging, compliance review, metadata). Covers the current model families (Marengo for search/embeddings, Pegasus for video-to-text analysis), the v1.3 Video Understanding API (indexes, assets, tasks, search, analyze, embed), prompt construction for analysis, capability and format limits, pricing/quota math, output quality review (hallucination and timestamp accuracy), and privacy/rights obligations when footage shows real people. Not for generating or editing video pixels — this is analysis and retrieval.

9k tokens
Ugc Ad Production
generative-media-skills

Provider-independent workflow for producing UGC-style paid and social ads, creator testimonials, product demos, hooks, before/after concepts, app walkthroughs, TikTok/Reels/Shorts variants, synthetic-persona or avatar UGC, ad-policy checks, disclosure planning, rights/consent review, delivery specs, and QA.

10k tokens
Vfx Compositing
generative-media-skills

Provider-independent VFX compositing direction for generated video, ads, trailers, product films, social clips, explainers, mixed-source edits, and image/video composites. Use when an agent must plan, brief, generate, integrate, repair, or quality-check visual effects composites involving plates, alpha/mattes, keying, roto, tracking, stabilization, matchmove, lighting/color integration, grain/noise/sharpness, shadows/reflections/contact, depth/atmosphere, AI-generated elements, clean plates, editor/colorist handoff, rights/safety, or compositing QA.

11k tokens
Video Description Oversight
generative-media-skills

Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Use for pre-caption critique and post-caption revision, aspect-by-aspect factual review, critique precision/recall/constructiveness, second-stage peer review, calibration, appeals, versioned triplets, provenance, and acceptance reporting; not for pixel/audio QA, accessibility captioning, model training, or creative shot direction.

4k tokens
Video Generation Gateways
generative-media-skills

>- Select, integrate, and operate multi-model video-generation gateways — hosted inference aggregators such as fal.ai, Replicate, and WaveSpeed that expose many third-party video (and video-adjacent audio) models behind one account, one API surface, and one bill. Use when deciding whether to route video generation through a gateway versus a direct model-provider API; when comparing gateways on catalog, async/queue job design, webhooks, pricing, schema discovery, and version pinning; and when operating gateway workloads in production — spend controls, retries, cross-gateway failover, content- safety differences, data retention, and commercial-rights passthrough of the underlying model. Not for direct model-provider APIs, local/self-hosted inference, model training/fine-tuning, or image-only gateway use.

12k tokens
Video To Audio Foley
generative-media-skills

Use for video-conditioned audio and automated Foley production: adding synchronized sound effects, ambience, impacts, footsteps, cloth, prop, and environmental audio to silent or under-sounded video using video-to-audio models, text-to-sound tools, manual spot lists, editing, rights checks, and QA.

8k tokens
Vidu Video
generative-media-skills

Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame, multi-subject reference, media-reference, or Q2 multi-frame work; do not conflate the API with vidu.com consumer plans or Vidu-S1 streaming digital humans.

11k tokens
Virtual Production Icvfx
generative-media-skills

Plan, troubleshoot, and hand off provider-independent in-camera VFX and LED volume virtual-production work, with Unreal Engine/nDisplay/Live Link awareness. Use for ICVFX suitability, LED stage geometry, frustums, tracking, lens calibration, sync, render nodes, color, lighting/reflections, artifact mitigation, rehearsal, live-composite fallback, monitoring, records, QA, and handoff. Do not use for ordinary offline post compositing, generic storyboards, or full game production.

12k tokens
Visual Style Direction
generative-media-skills

Translate a creative brief, brand system, or reference set into an original, portable visual-direction system for image and video production. Use for art direction, reference analysis without imitation, design tokens, palette and contrast, typography, composition, texture and material, lighting, lens and camera language, motion language, cross-shot continuity, accessibility, production handoff, visual review, and repair; do not use for provider-specific API syntax or model-specific prompt recipes.

15k tokens
Volcengine Doubao Speech Tts
generative-media-skills

Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats, expiring results, quota/error handling, authorized cloned voices, privacy, and audio QA. Do not use for international BytePlus endpoints.

3k tokens
World Labs Marble
generative-media-skills

>- Generate persistent, explorable 3D worlds/environments with Marble by World Labs — from text, a single image, multiple images, video, or a 360 panorama, plus coarse-structure blocking with Chisel. Use this skill when a task needs a navigable 3D scene, VR/immersive backdrop, virtual-production environment, previz set, game-blockout, or web/engine-ready Gaussian splat or mesh, and you must choose an input mode, a Marble model, an export format, and a plan/API route, then review world quality. Covers the Marble web app and the World API, editing/expansion, output formats and what each supports downstream, import into Unity/Unreal/Blender/Houdini/web, pricing/credits, capability limits, and rights/licensing. Not for flat text-to-image, text-to-video clips, or real-time on-the-fly world models (Marble outputs are downloadable, static 3D environments).

8k tokens
Xai Grok Imagine Image
generative-media-skills

Generate and edit production images with xAI's first-party Grok Imagine API, including model selection, multiple references, synchronous and Batch API workflows, durable output handling, prompting, iteration, QA, cost and rate controls, privacy, safety, and rights review. Use for direct xAI image API integrations, not Grok video or third-party gateways.

12k tokens
Xai Grok Imagine Video
generative-media-skills

Produce, edit, extend, and govern short videos with xAI's direct Grok Imagine Video API. Use for text-to-video, image-to-video, multi-reference video, natural-language video edits, video continuation, exact media-cost approval, asynchronous request recovery, Files/Batch integration, moderation review, and API privacy or residency decisions. Covers `grok-imagine-video` and image-only `grok-imagine-video-1.5`; excludes Grok consumer subscriptions, Grok on X, still-image generation, voice APIs, and third-party gateways.

20k tokens