The open format is called Agent Skills and works in Claude Code, Codex, Cursor and other agents — most people know it as Claude Skills.
Every Agent Skill we could find on GitHub, deduplicated by content. 79 870 files from 1 769 authors, of which 62 217 are unique — the rest is the same skill repackaged into someone else's repository. For each one: what it weighs in tokens, whether it ships runnable scripts, and which MCP servers it needs.
Plan, troubleshoot, and hand off provider-independent in-camera VFX and LED volume virtual-production work, with Unreal Engine/nDisplay/Live Link awareness. Use for ICVFX suitability, LED stage geometry, frustums, tracking, lens calibration, sync, render nodes, color, lighting/reflections, artifact mitigation, rehearsal, live-composite fallback, monitoring, records, QA, and handoff. Do not use for ordinary offline post compositing, generic storyboards, or full game production.
>- Produce 3D assets with Meshy (meshy.ai) through its REST API and web platform — text-to-3D, image-to-3D, multi-image-to-3D, AI texturing/PBR, remesh/topology control, UV, and auto-rigging/animation. Use this skill when an agent must generate a mesh from a prompt or reference image, control polygon count and topology, apply or restyle textures, rig a humanoid for animation, export to GLB/FBX/OBJ/USDZ/STL for Unity/Unreal/Blender, drive Meshy's asynchronous job API (auth, polling, streaming, webhooks, credits, rate limits), review generated meshes for game/film/print readiness, or decide whether Meshy is the right tool at all. Not for editing an existing hand-modeled asset's geometry, hard-surface CAD with exact dimensions, or guaranteed-watertight manufacturing parts.
>- Generate 3D assets with Tencent's Hunyuan3D family — open-weight self-hosted models (Hunyuan3D-2.0/2.1 and HunyuanWorld / HY-World for scenes) and the hosted, closed API tiers (2.5, PolyGen, 3.0/3.1 via Tencent Cloud and third-party hosts). Use when a task involves image-to-3D or text-to-3D mesh generation, the shape-then-texture (DiT + Paint / PBR) workflow, choosing between self-hosting and a hosted API, checking the Community License's Territory (EU/UK/South Korea) and 1M-MAU commercial gate, sizing GPU/VRAM for local inference, ComfyUI integration, or reviewing and post-processing (retopology, UVs, decimation) the meshes these models produce. Not for image, video, or general LLM tasks.
>- Produce 3D assets with Tripo (tripo3d.ai / Tripo AI by VAST) through its OpenAPI and Studio platform — text-to-3D, image-to-3D, and multiview-to-3D generation; PBR texturing; auto-rigging and preset animation; retopology/low-poly and quad remesh; format conversion (GLB/FBX/OBJ/USDZ/STL/3MF) and engine import (Unity/Unreal/Blender/3D printing). Use when an agent must drive the async task-based Tripo API (create task, poll or webhook, download), choose a model version (v3.0 / v3.1 / P1), write prompts or prepare input images, control mesh/texture parameters, estimate credits and respect rate limits, review generated mesh quality (topology, UVs, poly count, texture fidelity), run iteration and repair workflows, or reason about the rights/licensing of Tripo-generated assets. Trigger on requests to generate, texture, rig, animate, retopologize, convert, or evaluate a 3D model with Tripo, or to integrate the Tripo API into a pipeline. Not for choosing a 3D renderer, editing meshes by hand, or non-Tripo generators.
Use NVIDIA Maxine / NVIDIA AI for Media audio effects for speech cleanup and enhancement in live or offline media pipelines, including Background Noise Removal, denoise+dereverb, room echo removal, acoustic echo cancellation, audio super-resolution, Studio Voice, Speaker Focus, and Voice Font. Use when selecting, integrating, prompting around, QAing, or privacy-reviewing NVIDIA AFX SDK or BNR NIM workflows for conferencing, broadcast, podcast, avatar, ad, social, and video post-production audio.
Use D-ID to plan, generate, stream, localize, and QA avatar/talking-head videos, including V2 Photo Avatar Talks, V3 Pro/Instant Avatar Clips, V4 Expressive Avatar Scenes, D-ID Agents, voices, consent, lifecycle, safety, and artifact custody.
Produce Hedra character, talking-avatar, lip-sync, motion-avatar, and live avatar work. Use when an AI agent needs to choose Hedra models, prepare image/audio/script inputs, direct character performance, call or plan around the Hedra API or Studio, evaluate avatar/lip-sync results, handle consent/rights/privacy, or troubleshoot Hedra character video production.
Produce HeyGen avatar videos with Direct Video, Video Agent, Digital Twin, Avatar Realtime, and related avatar/voice/asset APIs. Use for provider routing, avatar and voice selection, script-to-video workflows, lip-sync from audio, backgrounds/scenes, callbacks and polling, consent and rights checks, localization handoffs, artifact custody, QA, and troubleshooting HeyGen avatar-video productions.
Produce presenter-led AI avatar videos with Synthesia Studio and Synthesia API. Use for Synthesia-specific avatar video planning, avatar/voice/language selection, scripted scenes, template personalization, API job lifecycle, localization/dubbing, brand/training/product workflows, consent/rights/safety checks, artifact custody, troubleshooting, and final QA.
Plan, prompt, call, edit, iterate, and productionize Alibaba Cloud Model Studio image generation with current Wan 2.7 Image, Qwen-Image 2.0, and Z-Image Turbo models. Use for Alibaba/DashScope text-to-image, multi-reference generation, instruction editing, character-consistent image sets, typography and layout work, regional authentication, model selection, API integration, cost/rate-limit planning, troubleshooting, safety, rights, privacy, and output QA.
Use for producing Tavus AI-human videos and real-time avatar conversations with Tavus Faces/Replicas, PALs/Personas, async Video Generation, CVI conversations, consent-safe likeness workflows, webhooks, backgrounds, localization, and avatar QA.
Use Adobe Firefly Services to generate, edit, expand, fill, match, composite, and upscale still images through the current REST APIs. Apply when selecting Firefly Image 5 versus Image 3/4 or custom models, implementing authenticated asynchronous image workflows, using style or structure references, building product composites, handling provenance and enterprise rights, or diagnosing Firefly image jobs in production. Exclude Firefly video, audio, and unrelated Creative Cloud APIs.
Plan, prompt, execute, troubleshoot, and quality-control Black Forest Labs FLUX image generation and editing across the BFL direct API and licensed local/open-weight deployments. Use for FLUX.2 model selection, text-to-image, single- or multi-reference editing, typography, exact-color work, mask-based erase/inpainting, outpainting, API integration, reproducibility, deployment licensing, and production rights or safety decisions; also use when migrating or maintaining legacy FLUX.1/FLUX1.1 workflows.
Produce and review production image-generation and image-editing workflows with Amazon Nova Canvas on Amazon Bedrock, including native InvokeModel payloads, safe authentication, validation, retries, cost controls, provenance, and lifecycle migration checks. Use when a task names Nova Canvas, amazon.nova-canvas-v1:0, Bedrock image generation, Canvas inpainting/outpainting/background removal, image conditioning, color guidance, image variation, virtual try-on, or Canvas fine-tuning.
Build and operate rights-aware Bria FIBO and FIBO Lite image-generation workflows with structured prompts, reference images, asynchronous status handling, webhooks, cost gates, and safe artifact downloads. Use when a user asks for Bria/FIBO generation, refinement, inspiration, reproducibility, hosted API integration, or a licensed-data image workflow; do not use for FIBO Edit, video, generic image editing, or third-party FIBO gateways.
Build and operate production image generation and natural-language image editing with ByteDance Seedream through first-party Volcengine Ark (China) or BytePlus ModelArk (global), including model and region selection, multi-reference and grouped outputs, streaming, secure artifact handling, retries, cost controls, prompting, safety, and rights review.
Plan, generate, edit, and quality-check still images with Google's Gemini native image models (Nano Banana), including Gemini 3.1 Flash Image, Gemini 3.1 Flash-Lite Image, Gemini 3 Pro Image, and legacy Gemini 2.5 Flash Image. Use for model and API-route selection, conversational editing, multi-reference composition, thinking and Google Search grounding, resolution/aspect-ratio control, exact-text and brand workflows, SynthID/provenance, safety, privacy, rights, and migration from legacy Gemini or Imagen image endpoints. Do not use for Veo/video generation or Gemini Live/audio work.
Generate, remix, edit, inpaint, reframe, background-process, describe, layerize, and upscale images with the Ideogram Developer API. Use when Codex must integrate Ideogram 4.0 or 3.0, render reliable typography or structured layouts, use style or character references, build synchronous or webhook workflows, migrate legacy Ideogram endpoints, or reason about Ideogram API pricing, safety, privacy, and production QA.
Select, integrate, and operate multi-model image-generation gateways with model-specific schema discovery, version policy, asynchronous jobs, webhooks, spend approval, safe inputs and artifacts, data-governance review, and billing observability. Use for comparing or building against fal.ai, Replicate, or Together AI image APIs, including controlled failover; do not use for direct model-provider APIs, local inference, training, dedicated endpoint provisioning, video generation, or general image editing.
Create, edit, guide, upscale, and quality-control still images with Leonardo.Ai's official Production API, including native Lucid and Phoenix models, uploaded or generated references, image-to-image, ControlNet guidance, realtime-canvas inpainting, Pro and Universal upscalers, and custom Elements or models. Use for Leonardo image API implementation, async polling or webhooks, dry-run and cost governance, secure artifact handling, prompt iteration, or production troubleshooting. Do not use for Leonardo video, 3D generation, unofficial wrappers, or third-party gateway APIs.
Use Luma AI's first-party Photon and Photon Flash image API for text-to-image, image/style/character references, image modification, asynchronous generation, secure artifact handling, production prompting, cost control, and policy-aware delivery. Trigger for Luma Photon API integration or production work; do not use for Ray video, UNI image models, the Luma consumer app, or third-party Photon gateways.
Plan, generate, edit, reference, and quality-control still images with Kling AI's hosted IMAGE surfaces or Kuaishou's open Kolors checkpoints. Use for Kling IMAGE 3.0/3.0 Omni/O1/2.1 web, official Kling CLI/MCP or Open Platform integration, and local Kolors text-to-image, image-to-image, IP-Adapter, ControlNet, inpainting, or LoRA work. Do not use for Kling video, avatars, lip-sync, unofficial gateway APIs, or for assuming that current hosted Kling Image models are identical to the 2024 open Kolors weights.
Plan, prompt, iterate, edit, migrate, and quality-check still-image work made with Midjourney's documented website and Discord interfaces. Use for Midjourney model and parameter selection, image/style/Omni references, Draft or Conversational workflows, typography-aware production, privacy and rights review, or troubleshooting a Midjourney image pipeline; do not use it to invent or automate an unsupported Midjourney API.
Generate, edit, composite, stream, and production-review still images with OpenAI GPT Image models. Use when an agent must choose between GPT Image 2 and legacy GPT Image/DALL-E integrations, select the Image API or Responses API, build prompts and reference-image workflows, use masks or multiple inputs, preserve identity or brand details, render in-image text, handle output files and partial images, estimate cost and rate limits, migrate deprecated image code, or apply OpenAI image safety, consent, privacy, provenance, and rights requirements. Do not use for OpenAI video or audio generation.
Generate, edit, and iterate still images with Runway's official API, especially native Gen-4 Image and Gen-4 Image Turbo reference workflows. Use when implementing Runway text-to-image, reference-driven image generation or natural-language image edits, task polling, secure artifact handling, production retries, cost controls, or Runway image API QA. Do not use for Runway video generation or third-party gateway APIs.
Generate and edit production images with xAI's first-party Grok Imagine API, including model selection, multiple references, synchronous and Batch API workflows, durable output handling, prompting, iteration, QA, cost and rate controls, privacy, safety, and rights review. Use for direct xAI image API integrations, not Grok video or third-party gateways.
Design and produce raster images, native SVG artwork, brand-controlled visuals, and documented image edits with Recraft's API. Use for Recraft model and route selection, exact request construction, image-to-image, inpainting, outpainting, background work, vectorization, remix/exploration, palette and typography control, API migration, safety and rights review, and production QA.
Operate Stability AI image generation, image-to-image, edit, control, background, and upscale APIs safely and reproducibly. Use when selecting or calling Stable Image Ultra, Core, Stable Diffusion 3.5, v2beta image edit/control services, or Stability open-weight image models; when debugging schemas, moderation, rate limits, cost, lifecycle, licenses, privacy, provenance, or migrations; and when planning production QA for Stability-generated images.
Use this skill when an agent needs Google Cloud Vision API for still-image understanding: labels, object localization, OCR, document text, SafeSearch, image properties, crop hints, web detection, batch annotation, Cloud Storage based pipelines, confidence evaluation, privacy, quotas, cost, and QA. Do not use it for Gemini multimodal reasoning, video analysis, image generation, custom model training, product catalog search design, or human reference/authenticity review.
Use this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR, moderation, image properties, Custom Labels or moderation adapters, stored-video analysis, conditional streaming-video workflows for existing eligible accounts, searchable media libraries, confidence evaluation, S3/IAM/event architecture, privacy, biometric consent, cost control, lifecycle management, and QA.
Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.
>- Generate AI lip sync and visual dubbing with Sync Labs (sync.so) — choose the right model (sync-3, lipsync-2, lipsync-2-pro, lipsync-1.9.0-beta, react-1), build the POST /v2/generate request, host inputs, handle async jobs via polling or webhooks, run batch dubbing/localization, and review output for sync accuracy and identity preservation. Use when an agent must re-voice, dub, translate, personalize, or re-time the mouth of a talking-head video (or animate a still face from audio), or must debug a failed/rejected Sync generation. Also covers consent, likeness, and rights obligations for editing a real person's face.
Use this skill when planning, capturing, processing, validating, troubleshooting, or exporting markerless human motion capture with Move AI products, including Move One single-camera workflows and Genesis multi-camera studio systems. It covers provider selection, setup, calibration, wardrobe and lighting constraints, documented multi-person operation, processing, cleanup, skeleton and retargeting choices, DCC/game-engine export, data management, consent, and privacy. Do not use it for general keyframe animation, non-Move mocap systems, 3D asset generation, or generic animation cleanup unrelated to Move AI outputs.
Use ACE-Step and ACE-Step 1.5 for local or hosted AI music generation, including text-to-music, lyrics-to-song, instrumental beds, covers, repainting, stem/track extraction, track completion, LoRA personalization, REST/Python/Gradio workflows, rights review, and music integration for video, ads, games, and social content.
Use Google Lyria music generation models through Google Cloud / Vertex AI / Gemini Enterprise Agent Platform for production music beds, songs, vocal tracks, image-conditioned music, short-form social audio, ad music, and video scoring. Covers Lyria 2, Lyria 3 Clip, and Lyria 3 Pro model selection, prompt construction, rights/privacy checks, SynthID/C2PA provenance, artifact custody, and QA.
Generate and iterate music with ElevenLabs Eleven Music for production deliverables. Use when planning, prompting, API-calling, editing, inpainting, reviewing, or licensing AI-generated songs, instrumental beds, ad music, soundtrack cues, video-to-music scores, stems, or music assets using the ElevenLabs Music API or ElevenCreative Music product.
Use DeepMotion Animate 3D for markerless video-to-3D human motion capture through the cloud portal, sales-gated API, or real-time SDK; covers capture planning, body/hand/face/multi-person limits, custom characters, retargeting, refinement, exports, DCC/game-engine handoff, QA, pricing, privacy, consent, and rights checks.
>- Generate music with Suno (v5 / v5.5 era, 2026) and advise a production team on what they may legally do with the output. Use this skill when a user wants to create, extend, remaster, or stem-separate a track in Suno; craft Suno prompts (style field, lyrics/metatags, exclude-styles, personas, custom voices); choose a subscription tier for a given use; or get accurate answers about ownership, commercial rights, distribution/monetization on DSPs, the Warner Music Group settlement, the ongoing label/publisher litigation, and Suno's API situation. Do not use it for other music generators (Udio, ElevenLabs Music, Google Lyria) except by explicit contrast, and do not use it as a substitute for a lawyer on a specific commercial deal.
Produce music with MiniMax Music 2.6 and MiniMax cover/lyrics APIs for songs, instrumentals, AI-generated lyrics, reference-audio covers, video/social/ad soundtracks, artifact custody, rights checks, and production QA.
Produce, direct, generate, QA, and integrate non-speech audio with ElevenLabs Text to Sound Effects / Sound Effects API. Use when an agent needs custom SFX, foley, ambience, loops, stingers, impacts, UI sounds, musical one-shots, trailer braams, or sound-design layers for video, ads, social edits, games, apps, podcasts, audiobooks, or interactive media using ElevenLabs.
Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs, and reviewing generated audio for media projects.
Use Microsoft Azure Speech in Foundry Tools for media-production speech workflows: speech-to-text, fast and batch transcription, diarization, captions/subtitles, real-time transcription, text-to-speech neural and HD voices, SSML, custom/personal voices, text-to-speech avatars, speech/video translation, localization, quotas, regions, privacy, consent, rights, and production QA.
Use for ElevenLabs provider-specific dubbing, localization, voice changer, speech-to-speech, and voice isolation/enhancement workflows. Applies when an agent must prepare source media, choose ElevenLabs dubbing versus voice conversion, manage language/speaker/timing decisions, handle voice identity and consent, call or plan against documented ElevenLabs APIs, export dubs/transcripts, estimate cost/limits, or quality-control localized speech while avoiding generic text-to-speech guidance.
Use for Deepgram speech and voice production workflows: speech-to-text transcription, live captions, diarization, audio intelligence, Aura text-to-speech, Flux and Voice Agent live audio, model selection, cost/limits/privacy checks, artifact custody, and production QA.
>- Operate AudioShake's cloud source-separation service (developer.audioshake.ai) to split a recording into stems. Use when a task involves isolating vocals, drums, bass, guitar, piano, keys, strings, or winds from music; separating dialogue / music / effects (DME) for film-TV post-production, localization, and dubbing; splitting a mixed recording into one stem per speaker; lyric transcription or word-level alignment; music detection/identification; or speech denoise/dereverb. Covers the Tasks API job lifecycle (assets, targets, formats, polling vs webhooks, credits, limits), choosing which stem targets a production actually needs, reviewing separated stems for bleed / artifacts / transient smearing / phase problems and repairing them, the rights and consent questions raised by separating copyrighted or third-party recordings, and the decision boundary for when a local open model such as Demucs is the better route than the API. Not for generating or synthesizing new audio, mixing, or mastering.
Use NVIDIA Speech NIM microservices for speech production and voice workflows, including self-hosted ASR/STT, TTS, text translation, speech-to-speech pipeline design, Riva Python client integration, deployment planning, GPU/runtime sizing, licensing/privacy review, and production QA for transcription, captioning, voice agents, localization, and synthetic voice output.
Use this skill when a media-production agent needs Google Cloud speech and voice services for transcription, captions/subtitles, long-form audio analysis, live caption planning, Text-to-Speech narration, Chirp 3 / Gemini-TTS voice selection, consented custom voice workflows, dubbing/localization planning, pricing/quotas/region checks, or speech-related safety and data-governance decisions.
Produce and understand audio with OpenAI request-based audio APIs and audio-capable chat models, including text-to-speech, transcription, translation, multimodal audio input/output, model routing, prompt and performance direction, artifact custody, approval gates, and safety/rights review. Use for non-realtime OpenAI audio production; route continuous live voice agents to the separate realtime voice skill.
Use Amazon Transcribe for AWS-based speech-to-text production: batch S3 transcription, real-time streaming, captions/subtitles, diarization, channel identification, custom vocabularies, vocabulary filters, language identification, PII/PHI handling, toxicity detection, Call Analytics, Medical, and secure S3/IAM/KMS workflows.
Use this skill when an agent needs AssemblyAI for speech-to-text or speech-understanding work in media production, including pre-recorded, synchronous short-file, and real-time streaming transcription; speaker diarization or speaker identification; captions and subtitles; timestamps; language detection, code-switching, or transcript translation; audio intelligence such as summaries, chapters, topics, entities, key phrases, sentiment, content moderation, profanity filtering, and PII redaction; webhooks, scaling, rate limits, retention, security, consent, and QA for podcasts, interviews, captions, call recordings, and edit workflows.
Use Amazon Polly for production text-to-speech work: selecting Standard, Neural, Long-form, or Generative engines and compatible voices; authoring SSML; creating speech marks for captions, word highlighting, or lip-sync; managing pronunciation lexicons; running synchronous, streaming, or asynchronous S3-backed synthesis; planning quotas, pricing, IAM, privacy, and QA for narration, audiobooks, accessibility audio, avatars, and multilingual media.
Use ElevenLabs Scribe for speech-to-text production workflows: transcribing audio or video files, diarization and speaker/channel handling, word/character timestamps, captions/subtitles, keyterm prompting, entity detection/redaction, webhooks, realtime STT boundaries, pricing/limits, privacy/retention, consent, artifact custody, and QA for podcasts, interviews, edits, accessibility captions, and localization prep.
Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling, and output QA. Do not use for mainland-China Volcengine Doubao Speech endpoints.
Use Cartesia Sonic and related Cartesia voice APIs for production speech: text-to-speech, realtime WebSocket TTS, voice selection, instant and professional voice cloning, pronunciation/language/emotion controls, voice localization, voice changer, pricing/concurrency planning, privacy/security review, and QA for narration, ads, localization, dubbing, avatars, and interactive voice agents.
Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken text; controlling pronunciation, pacing, emotion, and code-switching; using TTS APIs; cloning or designing a voice with consent; repairing artifacts; mastering deliverables; or evaluating generated speech. Excludes music, sound-effect generation, conversational-agent design, and general speech recognition except transcription used to verify TTS.
>- Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically.
Use Hume Octave for emotionally expressive speech and voice production: text-to-speech, voice design, voice cloning, voice conversion, streaming/realtime TTS, multilingual narration, dialogue continuity, timestamps/lip-sync, safety/rights review, and production QA.
>- Local and self-hosted text-to-speech with Kokoro, the ~82M-parameter open-weight (Apache 2.0) model by hexgrad. Use when synthesizing speech offline, on-device, in the browser, or on your own server without per-character API cost — for high-volume narration, audiobooks, privacy-constrained pipelines, and prototyping. Covers running Kokoro via the Python `kokoro` package, ONNX / kokoro-js in the browser, and the OpenAI-compatible Kokoro-FastAPI wrapper; picking voices and language codes; blending voices; chunking long text; controlling pronunciation via misaki/espeak-ng; and judging when Kokoro is the right tool versus when its lack of voice cloning, narrow emotional range, and weaker non-English quality mean you should reach for a hosted or larger model instead. Not for voice cloning, expressive/emotional character performance, or high-fidelity multilingual work — say so and route elsewhere.
Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats, expiring results, quota/error handling, authorized cloned voices, privacy, and audio QA. Do not use for international BytePlus endpoints.
Produce text-to-speech with Alibaba/Qwen Qwen3-TTS through DashScope/Model Studio or open-weight Qwen3-TTS checkpoints. Use when an agent must choose Qwen3-TTS models, voices, realtime versus non-realtime synthesis, voice cloning, voice design, instruction/prosody controls, multilingual or dialect speech, audio artifact handling, regional routing, pricing/privacy/licensing constraints, or production QA for narration, characters, assistants, audiobooks, ads, localization, or voice-enabled media.
Answers built from the skills we actually parsed.