3 019 media skills from 496 authors. They create and process images, video and sound. Half of them fit into 1 966 tokens or less — that is what one costs your context window when the agent loads it. 768 ship runnable scripts rather than instructions alone. 30 of them cannot work without an MCP server, most often rube. We also found 321 copies of these same skills sitting in other people's repositories — counted once here, not 321 times.
3 019 unique 496 authors 1 925 updated this month 160 from vendors
Elite voice actor with 10+ years in commercial, animation, games, and audiobooks. Specializes in character voice design, emotional delivery, and studio performance. Use when: voice acting, voice-over, character voice, dubbing, audiobook narration.
Master video editor with 12+ years in commercial, documentary, and social media post-production. Specializes in narrative pacing, color science, sound design integration, and efficient NLE workflows
Expert-level Art Instructor with 15+ years of experience in drawing, painting, illustration, digital art, and art history
Expert-level Art Teacher with deep knowledge of drawing, painting, illustration, design principles, color theory, and visual arts education
Expert-level Music Instructor with 20+ years of experience in piano, guitar, violin, drums, vocals, music theory, composition, and audio production
Expert-level Music Teacher with deep knowledge of instrument pedagogy, music theory, sight-reading, ear training, practice methodology, and performance psychology
Expert Online Course Creator specializing in instructional design, multimedia content production, learning experience design, and course monetization. Expert in course platforms, video production, engagement strategies, and learner success optimization. Use when: online-course-creation, instructional-design, course-monetization, video-production, learning-experience-design, course-platforms.
Expert Speech-Language Pathologist (SLP) with 15+ years of experience in diagnosing and treating speech, language, and communication disorders
Generate images, video, audio, and 3D with Comfy Cloud — search hundreds of models and workflow templates, run custom ComfyUI workflows, and manage generation jobs through the hosted Comfy Cloud MCP server. Cloud-only — connects to the hosted service, not a local ComfyUI install. Use for any "generate/edit an image", "make a video", "text to 3D", or ComfyUI workflow task.
Bitwarden's product content style guide for end-user-facing GUI copy — voice, tone, AP-style-with-exceptions grammar, sentence case in UI, and accessibility-first language at a U.S. 7th-grade reading level.
Learn reusable 2D motion-generation knowledge from user-specified action resources with `/img2mo-learn <resource>`. Use when the user provides videos, extracted frame sequences, spritesheets, Spine assets, generated outputs, failed attempts, or reference motion folders and wants to summarize animation timing, pose beats, style traits, prompt patterns, extraction/cropping rules, or failure lessons into the project `img2mo-knowledge/` folder for later `/img2motion` generation.
Standardize a 2D animation baseline image with `/img2mo-std xxx.png/pos`. Use when preparing a character, creature, prop, or weapon baseline before motion generation, especially if the source image is too large, tightly cropped, lacks transparent action margin, has a white background, or later walk/attack/idle frames show clipping, overlap, scale popping, or inconsistent character size.
Use when a user provides one baseline image and requests game animation frames, sprite sequences, attack/walk/idle/hit/death/casting motion, transparent PNG frames, a spritesheet, keyframe prompts, or consistent whole-character pose animation.
Call vision models (Doubao, Qwen, OpenAI) to analyze images. Use when you need to understand screenshots, UI layouts, diagrams, or any image content. Supports png/jpg/webp/gif.
Diagnose flat dialogue, same-voice characters, and lack of subtext. Use when conversations feel wooden, characters sound alike, or dialogue only does one thing at a time.
Generate phonologically consistent constructed languages for fiction. Use when you need naming languages, alien speech, or fantasy tongues without deep linguistics knowledge.
Extract descriptive musical characteristics from any artist or band without using their name, building a vocabulary of sonic qualities for AI music generation, music description, or creative recombination.
Convert written documents to narrated video scripts with TTS audio and word-level timing. Use when preparing essays, blog posts, or articles for video narration. Outputs scene files, audio, and VTT with precise word timestamps. Keywords: narration, voiceover, TTS, scenes, audio, timing, video script, spoken.
Transform comprehensive written content into purposeful spoken guidance. Use when adapting for speech, converting to spoken format, optimizing for listening, or creating audio content from written material. Keywords: speech, audio, spoken, listening, adaptation, podcast.
Extract and document a writer's distinctive voice patterns for consistent reproduction. Use when you need to capture writing voice, analyze writing style, create a voice guide, or write in someone's established style. Keywords: voice, tone, style, writing analysis, fingerprint.
Generate game assets using AI image generation APIs (DALL-E, Replicate, fal.ai) and prepare them for Godot. Covers the full art pipeline from concept art and style guides to final sprites, sprite sheets, and import configuration. This skill should be used when creating game art, generating sprites, making tilesets, creating UI elements, or preparing assets for Godot import. Keywords: game assets, AI art, DALL-E, Replicate, fal.ai, sprite sheet, tileset, Godot, pixel art, character sprite, game art, texture, animation frames.
Download YouTube video or audio with yt-dlp and ffmpeg at highest available quality.
> Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all automated. photo of a person/product anchored into the collage (C-roll mode).
(project) Use when editing any file under skills/hal-skills/ or plugins/hal-voice/ to bump the plugin version before committing
Reverse-engineer any motion reference (a video from X/Twitter, Dribbble, a screen recording, a GIF) into production animation code through frame-level dissection. Use when the user shares a video/URL and says "implement this animation", "recreate this motion", "port this interaction", "how does this animate", "clone this effect", or wants to study how a reference moves before building it. Covers both timeline choreography (entrances, text sweeps, staggers) and interaction-driven motion (scrubbers, sliders, drag-driven scenes). Also fires when the user asks where to find good animation references or inspiration — it suggests curated sites and X accounts to hunt, then reverse-engineers whatever they bring back. Ports default to React/TypeScript with framer-motion; the analysis phases are framework-agnostic.
> Tạo một video MỚI cho series "so sánh / phân biệt kiến thức" của repo này — clip dọc TikTok/Reels/Shorts 30-40s, layout 3-zone cố định theo DESIGN.md, voiceover tiếng Việt sinh bằng Vbee TTS, dựng bằng HyperFrames. Dùng skill này khi người dùng nói "làm video so sánh X vs Y", "phân biệt X và Y", "thêm video mới vào series", "tạo video so sánh kiến thức", hoặc yêu cầu bất kỳ video nào theo đúng format sẵn có của repo (thư mục videos/<slug>/). KHÔNG dùng cho video ngoài format này (promo sản phẩm, video từ URL, slideshow, thêm phụ đề cho footage có sẵn).
> Tạo một video MỚI cho series "so sánh/phân biệt kiến thức" của repo này (TikTok/Reels 30-40s, layout 3-zone cố định trong DESIGN.md). Dùng khi được yêu cầu "làm video so sánh X vs Y", "thêm video mới vào series", hoặc bất kỳ video nào theo format sẵn có của repo. Không dùng cho video ngoài format này — route qua /hyperframes như bình thường.
Use for DSPy adapter selection, JSONAdapter, XMLAdapter, ChatAdapter, native function calling, structured outputs, and multimodal inputs like dspy.Image or dspy.Audio.
Turn Chinese classical poems and ci into coherent vertical Chinese-art videos with poem-driven scene grouping, GPT ImageGen stills, Docker-only Gemini I2V, retained model-generated ambience, Gemini sparkle-watermark cleanup, brush-calligraphy captions revealed character by character, optional local BGM mixing, stitching, and final-frame QA. Supports verse-driven variation across ink landscape, gongbi bird-and-flower, colored figure-and-horse painting, blue-green landscape, xuan paper, silk, and related Chinese visual languages. Use when users ask for 古诗词动态视频、诗词逐句或两句一景、国风视频、毛笔字逐字出现、整首诗拼接成片,or want the established 月落乌啼霜满天 workflow applied to another poem.
Use this skill whenever the user asks to create, improve, audit, or split prompts for AI video generators (Seedance, Kling, Veo, Runway, Luma, Pika, Sora, any image-to-video system). The skill also covers storyboards, shot lists, director treatments, dynamic montage, multi-clip story structure, camera direction, lighting, blocking, pacing, character continuity, dialogue, and sound design. Trigger even when the user says things like "придумай сцену для видео", "разбей на склейки", "сделай раскадровку", "улучши промпт для Kling", "переведи сценарий в промпты", "как снять X в AI-видео", or shares a prompt and asks to fix it.
> Image prompting skill for Nano Banana (NBP/NB2) and GPT Image 2. Writes ready-to-use картинку", "image prompt", "промпт для картинки", blog covers, slides, posters, product shots, UI mockups, storyboards, character sheets, edit/colorize, style transfer, vision analysis, image-to-prompt, nb, NBP, NB2, gpt-image-2, multi-panel grids, ecommerce product photography, fashion editorial, food/beverage ads, cinematic portraits.
Prepare ASE NEB workflow tasks with backend-agnostic controls. Use when the user needs reaction-path optimization between initial/final states with explicit image construction, spring settings, and convergence controls.
> A versatile CLI tool for converting molecular file formats, generating 3D atomic coordinates from SMILES, rendering 2D chemical structure images, and preparing or extracting structures for computational workflows. USE WHEN you need to convert between chemical file formats (e.g., xyz, pdb, mol, smi, gjf), generate 3D structures from SMILES using `--gen3d`, render molecule images (PNG/SVG), or extract geometries from simulation logs to build new inputs.
> USE WHEN requesting core chemical structural data (SMILES, formula, mass, 2D images) via IUPAC, common, or multilingual names. You MUST actively retrieve the data using this skill; DO NOT hallucinate or generate structures yourself. DO NOT USE WHEN asking for physical properties (melting point, solubility), safety/toxicity data (MSDS), or synthesis pathways.
Judges whether a user's input suits the Vibe Creating style of video-prompt writing, and when it does, distills single-scene prompts, multi-shot descriptions, emotional imagery, or mixed input into prompts that are easier for a video model to generate from — while preserving any user-specified dialogue, voiceover, music, sound effects, and other hard constraints. Use when a user wants to turn an idea, story, feeling, or rough/over-specified prompt into a strong text-to-video prompt (Seedance, Sora, Kling, Veo, Runway, etc.), or asks to "rewrite", "improve", "clean up", or "vibe-ify" a video prompt. Do NOT use for long narrative films that need precise word-for-word dialogue sync, industrial shot lists meant to be executed verbatim, or functional/UI demos and step-by-step tutorials.
Match UI/UX interaction needs to proven SwiftUI animation patterns from curated open-source catalogs. Use when planning or building iOS/macOS screens where an interaction should feel alive - loaders, likes, toggles, card decks, reveals, shaders - or when asked which animation fits a moment. Recommends system-first restraint before custom motion.
Run animation jank QA. Use for GSAP, Motion, CSS transitions, scroll animation, pinned scenes, parallax, hover states, page transitions, mobile motion, reduced-motion fallbacks, performance-heavy animations, and final motion polish before delivery.
Define animation easing language for premium websites and apps. Use for motion timing systems, easing curves, spring feel, duration scales, stagger rhythm, delay strategy, entrance/exit timing, tactile controls, cinematic sections, and translating luxury, modern, fast, calm, or playful motion words into concrete values.
Use when planning or implementing website motion, scroll choreography, microinteractions, page transitions, text reveals, pinned sections, smooth scrolling, Lottie, Three.js/R3F, WebGL, or animation-heavy premium sites. Chooses the lightest appropriate tool and defines reduced-motion fallbacks.
Translate brand voice into UI copy for websites and apps. Use for buttons, labels, empty states, loading states, errors, tooltips, form help, nav labels, product copy, onboarding microcopy, premium tone systems, and making brand voice usable in interface text instead of only marketing prose.
Improve website accessibility and performance polish. Use for semantic HTML, keyboard navigation, focus states, contrast, alt text, reduced motion, image/video loading, canvas fallbacks, bundle caution, animation performance, and final production readiness for polished websites.
Motion and interaction design skill using interpretive lenses informed by publicly available work from Emil Kowalski, Jakub Krehel, and Jhey Tompkins. Two modes: build interactive components with purposeful motion, or audit existing animations to catch AI-generated motion anti-patterns. Use when creating, adding, animating, or reviewing UI motion: transitions, hover states, micro-interactions, enter/exit animations, or motion design work in React, Framer Motion, CSS, or HTML. This skill is not authored, reviewed, or endorsed by the referenced designers.
GSAP performance guidance — prefer transforms, avoid layout thrashing, use will-change selectively, and batch animation work. Use when optimizing GSAP animations, reducing jank, or when the user asks about animation performance, FPS, or smooth 60fps.
Art-direct website hero imagery. Use for landing pages, homepages, product pages, portfolios, venue sites, editorial pages, and premium first viewports that need real or generated media, text-safe composition, responsive focal points, visual hierarchy, contrast, aspect ratios, and anti-generic image selection.
Run responsive image crop QA for websites. Use after adding or changing images, hero media, galleries, product grids, portraits, venue photos, screenshots, background images, object-fit/object-position rules, art-directed picture sources, image dimensions, layout stability, and mobile crop polish.
Run image and video loading QA. Use after adding or changing images, videos, posters, responsive sources, lazy loading, galleries, hero media, product media, background videos, image sequences, or CMS media to verify loading, dimensions, crops, network errors, alt/captions, and performance.
Define website motion direction before implementation. Use for cinematic websites, Viktor-style motion, GSAP/ScrollTrigger pages, Motion for React microinteractions, Framer-like transitions, scroll reveals, product storytelling, luxury pacing, and any site where animation should feel intentional rather than generic.
Audit website performance. Use for Core Web Vitals, LCP, INP, CLS, Lighthouse or browser performance checks, bundle and asset weight, image/video loading, font loading, third-party scripts, animation cost, hydration cost, and production-readiness performance review.