mcpbeat

Media Claude Skills

3 032 media skills from 496 authors. They create and process images, video and sound. Half of them fit into 1 962 tokens or less — that is what one costs your context window when the agent loads it. 771 ship runnable scripts rather than instructions alone. 30 of them cannot work without an MCP server, most often rube. We also found 321 copies of these same skills sitting in other people's repositories — counted once here, not 321 times.

3 032 unique 496 authors 1 920 updated this month 173 from vendors

1 962
tokens, median
what a typical one costs in context
771
ship scripts
code that runs, not instructions alone
30
need a server
most often rube
321
copies elsewhere
counted once here, not once per repository

1 537–1 584 of 3 032

page 33 of 64
Progress Monitoring Cv
by datadrivenconstruction

Monitor construction progress using computer vision. Analyze site photos and drone imagery to track work completion, detect safety issues, and compare against BIM models.

4k tokens
Director
by hoodini

Award-winning director's brain for making films/videos with AI. Runs a gated, interrogation-first process: lock the IDEA and the SCRIPT (the story) in words before generating a single pixel, then plan shots/camera/light/sound, then generate and edit. Use whenever creating a video, film, teaser, trailer, promo, ad, short, reel, documentary, or narrative piece AND the storytelling must actually work — or any time the user says the story/script keeps failing. Hebrew triggers: סרטון, טיזר, פרומו, פרסומת, ריל, סרט, תסריט, סטוריבורד. Handles development, story, and shot-planning; composes with cinematic-ai-video and yuv-fomo-teaser (style/manipulation), hyperframes (render), nano-banana-2/ElevenLabs (assets). Backed by a 28-chapter Director's Bible in references/.

178k tokens
Image Master
by hoodini

Master prompt-engineer for photoreal, artifact-free AI still images on ANY tool (Reve, Midjourney, Flux, GPT-image, Imagen, Nano Banana, Stable Diffusion). Builds prompts that hit National-Geographic-grade realism — true skin/fur texture (no plastic), correct anatomy/hands/faces, physically coherent light/shadows/reflections, clean legible text — while engineering deliberate visual IMPACT (color contrast, composition, awe/adrenaline, a striking point of view). Runs a gated process: lock the point-of-view and build the 8-block Capture Stack BEFORE generating, then run a forensic pre-submit inspection against the known artifact list. Use whenever the goal is a single still image that must look REAL and hold up under close inspection — contest entries, hero shots, product/character/wildlife/architecture/concept art, or any time AI images come out plastic, distorted, or fake. Composes with nano-banana-2 (execution) and director (motion). Hebrew triggers: תמונה, תמונות, פוטוריאליזם, ריאליזם, לייצר תמונה, פרומפט לתמונה, בלי עיוותים, עור פלסטיק, ידיים מעוותות, פרצוף מעוות, חדות, נשיונל ג'יאוגרפיק.

27k tokens
Parallax Landing Page
by hoodini

Build a scroll-driven cinematic landing page from a short video. The user provides a 5–15 second video (often AI-generated); this skill extracts every frame at HD JPEG quality, then produces a single-hero HTML page where the user's scroll gesture scrubs the frames in place (the page itself never scrolls) and 5 dramatic text overlays crossfade in/out — Google Anton headlines, Caveat handwritten accents, locked body, virtual scroll. Use this skill whenever the user wants to "turn this video into a landing page", "make a scroll-scrub landing page", "build a parallax hero from this clip", "add a new landing page to the parasites showcase", "do the same as github/lion/hope for this new video", or any variant that pairs a short clip with dramatic scroll-triggered storytelling. Trigger even if the user only says "use my video for a landing page" — that is this skill.

16k tokens scripts
Video Edit
by hoodini

Edit any video into a captioned showcase — transcribe (any language, defaults to large-v3), present a transcript_review.txt for the user to fix mishears BEFORE rendering, then build a HyperFrames composition with liquid-glass caption pills, liquid blob background, liquid morph wipes, optional behind-subject text via background removal, and render the final video. Use whenever the user provides a video file and asks to edit it, caption it, add subtitles, fix existing captions, make a reel/promo/captioned tutorial, or "do the same" pattern as a prior captioned video. Supports English, Hebrew, and any Whisper-supported language. **Renders both 16:9 (YouTube / horizontal) and 9:16 (TikTok / Instagram Reels / YouTube Shorts) from the SAME 16:9 source** — vertical mode uses a centered footage strip with a blurred backdrop + liquid blobs and a vertical-tuned caption pill, no need to re-shoot. THE PIPELINE PAUSES FOR USER APPROVAL on the transcript before final render — this is the support mechanism for getting captions perfect (especially Hebrew). Pairs with hyperframes, hyperframes-cli, hyperframes-registry, and yuv-design-system skills.

260k tokens scripts
Video To Landing Page
by hoodini

Turn any video into a cinematic scroll-driven landing page — Apple-style hero where scrolling progresses the visible frame through the video. Use when the user provides a video file and asks for "a landing page from this video", "scroll-frame website", "Apple-style scroll site", "hero that scrubs the video", "like the GitHub Copilot landing", or any equivalent. Extracts N evenly-spaced frames via ffmpeg, builds a self-contained HTML page with a sticky hero + JS scroll listener that swaps the visible frame as you scroll, plus headline, sections and CTA below. For YUV.AI projects, applies the yuv-design-system skill in Neon mode (pink/cyan/white, default for YUV.AI web) — Decks (purple/yellow) is reserved for slides only. For generic / non-YUV.AI projects, picks an appropriate palette per the source video. Output is one folder with `index.html` and a `frames/` directory — drop on any static host.

5k tokens scripts
Yuv Decks
by hoodini

Build cinematic, narrative-driven presentation decks in Yuval Avidani's signature style using @open-slide/core. The user describes a topic and audience; this skill scaffolds an open-slide project, drafts the 4-act narrative arc (Boarding → Ascent → Cruise → Descent), writes every slide in the Yuval voice (plain-language, no jargon, story-driven), applies the brand visual language from yuv-design-system, and orchestrates companion skills for hero images and video moments. Triggers on "make a deck", "create slides", "build a presentation", "build a deck", "new deck", "presentation about", "talk deck", "hackathon deck", "open-slide deck", "yuv-decks", "yuv deck", "deck like Yuval", "מצגת", "שקפים", "דק", "מצגת על", "להכין מצגת". Use proactively whenever the user asks for ANY slide-based talk; the skill self-selects the right scope.

6k tokens
Yuv Reel Covers
by hoodini

> Generate unified, on-brand Instagram Reel covers for Yuval (YUV.AI Neon Phoenix system) — the signature look is a giant Hebrew headline BEHIND the subject cutout + a punch line IN FRONT (depth effect), on rich black with a dim neural-net field, series-colored chip + phoenix mark. Use whenever Yuval asks for an Instagram/Reel/TikTok/YouTube cover, thumbnail, עטיפה, קאבר, תמבנייל, cover image for a video, or wants his Instagram grid to look consistent. One command per cover; background removal included in the pipeline.

1197k tokens scripts
Yuv Pilot
by hoodini

Top-of-pyramid orchestrator for Yuval Avidani's YUV.AI brand work. Apply when (a) the user wants YUV.AI output and the medium is ambiguous or multi-medium, (b) the user is planning a launch or cross-channel campaign for YUV.AI, or (c) explicitly invokes /yuv-pilot or asks "what should I build for YUV.AI / my brand". Triggers: "for YUV.AI", "for my brand", "YUV.AI launch", "ship something for me", "orchestrate", "cross-channel", "multi-platform for me", "yuv-pilot", Hebrew השקה, מולטי-פלטפורמה ל-יובל. Does NOT do the work — it identifies the right downstream YUV.AI skills (yuv-design-system across 3 modes, yuv-decks for slides, yuv-viral-video for short MP4s, hyperframes for HTML→MP4, nano-banana for in-brand imagery, gsap for animation), explains the composition, hands off. Does NOT apply to non-YUV.AI requests. When a request is clearly single-medium (just a deck, just a viral short), the specific skill wins — yuv-pilot is the front door for ambiguous or multi-output requests.

5k tokens
Yuv Video Director
by hoodini

> Yuval's all-in-one AI video pipeline. Turns an idea/script into a finished, on-brand MP4 by orchestrating HyperFrames (HTML→deterministic video render), Lottie (branded motion graphics), ManimCE (math / neural-network / concept animations), and a transcribe→approve caption flow — all wrapped in the YUV.AI Neon Phoenix brand via a frame.md. Use whenever Yuval wants to make, about X", "explain X as a video", "neural network animation", "turn this into a video", hyperframes, animation, "make a video", "explain ... as a video", מצגת וידאו, סרטון, הסבר וידאו. Routes each beat to the right engine, wraps in brand, self-verifies, and renders.

213k tokens scripts
Yuv Viral Video
by hoodini

Edit any selfie or screen-share footage into a viral short-form video in YUV.AI's signature style — Apple-style liquid-glass cards (real CSS backdrop-filter), dark-mode polish, MrBeast-paced cuts, video-title karaoke captions, premium GSAP motion graphics, no fake content, never covering the speaker's face. Hebrew is rendered in Rubik Black, English in Anton uppercase. Always renders BOTH 9:16 and 16:9 and always saves with _V<N> suffix for backups. Trigger when the user drops a path to an .mp4/.mov/.mkv and says "edit this", "make it viral", "turn this into a short", or any Hebrew equivalent (ערוך סרטון, סרטון ויראלי, להפוך לוויראלי, ריל, שורט). The pipeline is the COMBINATION of two skills: video-use (transcription + word-snapped cuts + base extraction) and hyperframes (HTML/CSS/GSAP visual composition + render). Do NOT use for podcast-only audio edits.

107k tokens scripts
Excalidraw Skill
by lingzhi227

Programmatic canvas toolkit for creating, editing, and refining Excalidraw diagrams via MCP tools with real-time canvas sync. Use when an agent needs to (1) draw or lay out diagrams on a live canvas, (2) iteratively refine diagrams using describe_scene and get_canvas_screenshot to see its own work, (3) export/import .excalidraw files or PNG/SVG images, (4) save/restore canvas snapshots, (5) convert Mermaid to Excalidraw, or (6) perform element-level CRUD, alignment, distribution, grouping, duplication, and locking. Requires a running canvas server (EXPRESS_SERVER_URL, default http://localhost:3000).

8k tokens scripts
Unity Vrc World SDK 3
by niaka3dayo

> VRChat World SDK 3 guide for scene and Inspector setup, component placement, optimization, and upload. Use for VRChat world scene configuration, VRC SDK components, layers, baked lighting, Quest/Android performance, Dynamics for Worlds, Build Panel warning triage, validation, and upload. Covers VRC_SceneDescriptor, VRC_Pickup, VRC_Station, VRC_Mirror, VRC_ObjectSync, VRC_CameraDolly, spawn points, collision matrices, PhysBone and Contact component placement, Box Contacts, Global Avatar PhysBone Colliders, and VRCPhysBoneCollider component setup. component placement, optimization, Quest support, light baking, upload, SDK validation, Build Panel warning, Auto Fix, red warning, yellow warning, or white warning. Do not use for UdonSharp C# or VRCTween calls; use unity-vrc-udon-sharp for runtime scripting.

49k tokens
Azure Speech Service
by membranedev

| Azure Speech Service integration. Manage data, records, and automate workflows. Use when the user wants to interact with Azure Speech Service data.

2k tokens
Canvas
by membranedev

| Canvas integration. Manage Canvases. Use when the user wants to interact with Canvas data.

2k tokens
Chatsonic
by membranedev

| Chatsonic integration. Manage Users, Chats, Images, Workspaces, Prompts. Use when the user wants to interact with Chatsonic data.

2k tokens
Dacast
by membranedev

| Dacast integration. Manage Videos, Playlists, Channels. Use when the user wants to interact with Dacast data.

2k tokens
Diffbot
by membranedev

| Diffbot integration. Manage Articles, Products, Images, Discussions, Videos. Use when the user wants to interact with Diffbot data.

2k tokens
Display Video 360
by membranedev

| Display & Video 360 integration. Manage data, records, and automate workflows. Use when the user wants to interact with Display & Video 360 data.

2k tokens
Dynapictures
by membranedev

| DynaPictures integration. Manage Images, Users, Albums, Tags. Use when the user wants to interact with DynaPictures data.

2k tokens
Generated Photos
by membranedev

| Generated Photos integration. Manage Persons. Use when the user wants to interact with Generated Photos data.

2k tokens
Google Cloud Vision
by membranedev

| Google Cloud Vision integration. Manage Images. Use when the user wants to interact with Google Cloud Vision data.

2k tokens
Heygen
by membranedev

| HeyGen integration. Manage Videos, Avatars, Templates. Use when the user wants to interact with HeyGen data.

2k tokens
Hippo Video
by membranedev

| Hippo Video integration. Manage Persons, Organizations, Deals, Leads, Activities, Notes and more. Use when the user wants to interact with Hippo Video data.

2k tokens
HTML To Image
by membranedev

| HTML to Image integration. Manage Images. Use when the user wants to interact with HTML to Image data.

2k tokens
Htmlcss To Image
by membranedev

| HTML/CSS to Image integration. Manage Images. Use when the user wants to interact with HTML/CSS to Image data.

2k tokens
Imagekitio
by membranedev

| ImageKit.io integration. Manage Images, Folders, Users. Use when the user wants to interact with ImageKit.io data.

2k tokens
Imejisio
by membranedev

| Imejis.io integration. Manage Images, Users, Projects. Use when the user wants to interact with Imejis.io data.

2k tokens
Removebg
by membranedev

| Remove.bg integration. Manage Images. Use when the user wants to interact with Remove.bg data.

2k tokens
Ritekit
by membranedev

| Ritekit integration. Manage Hashtags, Images, Texts. Use when the user wants to interact with Ritekit data.

2k tokens
Sitecreatorio
by membranedev

| Sitecreator.io integration. Manage Sites, Pages, Templates, Images, Domains, Users and more. Use when the user wants to interact with Sitecreator.io data.

2k tokens
Video SDK
by membranedev

| Video sdk integration. Manage data, records, and automate workflows. Use when the user wants to interact with Video sdk data.

2k tokens
Vimeo
by membranedev

| Vimeo integration. Manage Videos. Use when the user wants to interact with Vimeo data.

2k tokens
Agentsop Streaming Output
by agentsope

| Enhancement-overlay decision protocol for STREAMING the output of long-running LLM / agent runs from the *backend*, not just wiring a typing animation in the UI. Activates when a coder agent must stream final tokens to a chat client, surface intermediate agent steps (which tool, which node, partial reasoning), emit custom tool-progress events, choose a transport (SSE vs WebSocket), or decide what to do when the client disconnects mid-stream. The langchain / langgraph skills mention stream modes but stop at "you can stream"; this skill encodes *what to stream, over what transport, and how to fail safely*.

11k tokens
Long Audio Transcript Processor
by cafe3310

对大量语音转写稿进行校对、整理、分段处理,支持断点续传和恢复

4k tokens scripts zh
Online Content Collector
by cafe3310

对 Obsidian 仓库进行自动素材媒体剪藏,本地化特定 tag 标注的网页、视频及附件

7k tokens scripts zh
Wx Emoji Maker
by cafe3310

处理 PNG 图片目录,将其转换为适合微信使用的表情包

2k tokens scripts zh
Paste Image
by cafe3310

此技能提供一个脚本,用于将 macOS 剪贴板中的图片直接粘贴到指定的 PNG 文件。

533 tokens scripts zh
Gemini Omni Video To Sticker Gif
by cafe3310

从视频片段裁剪缩放变速以制作GIF动态表情

3k tokens scripts zh
Oneshot Website
by cafe3310

Generate immersive, one-shot single-file HTML websites with embedded CSS and JS. No external images. Hostable on CodePen or Vercel. Use for writeup showcases, AI capability demos, and portfolio pieces.

4k tokens
Slides
by majiayu000

生成口播视频背景 PPT 幻灯片(16:9 横版 PNG 序列)。当用户需要做 PPT、生成幻灯片、做演示背景图时使用

8k tokens zh
Web Asset Generator
by majiayu000

Generate web assets including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Use when users need icons, favicons, social sharing images, or Open Graph images from logos or text slogans. Handles image resizing, text-to-image generation, and provides proper HTML meta tags.

12k tokens scripts
Xiaohongshu Netfeel Guardian
by majiayu000

| 当用户用 Claude 写小红书、公众号、视频脚本等内容时,输出出现翻译腔、英文思考、西式框架,导致“像外国人假装写中文”时使用。 典型触发语句: 始终保护用户原有的 CLAUDE.md / AGENTS.md 中的人设与语气要求,只做最小必要干预。优先与 xiaohongshu 技能组合使用。

4k tokens zh
Higgsfield
by OSideMedia

> Use this skill whenever the user asks anything about Higgsfield AI — writing or refining video/image prompts, choosing a model (Kling, Sora 2, Veo, Wan, Seedance, Minimax Hailuo, DoP, Soul, Nano Banana, Seedream, Flux, GPT Image, etc.), camera controls, named motion presets, Soul ID character consistency, Cinema Studio 2.5/3.0, Vibe Motion, troubleshooting failed generations, credit optimization, Photodump, or any mention of higgsfield.ai. Also trigger on generic "write me a video prompt" or "make me an AI video prompt" requests when Higgsfield is the user's configured platform.

1274k tokens scripts
Higgsfield Canvas
by OSideMedia

Use when the user mentions Higgsfield Canvas, a node-based or node graph workspace, an infinite board/canvas, chaining generations into a pipeline, or wants to wire prompts → images → videos across models on one surface. Covers what Canvas is, the node categories, the seven models that run inside Canvas, the named canvas patterns (Simple Seedance, Extend Video, Image Edit, StoryBoard With Elements, Long Video fan-out), the build-free / generate-paid cost model, reusable templates, assets-as-nodes, and Shared Canvas live collaboration. Also trigger on 'Higgsfield ComfyUI alternative', 'node workflow', or 'connect nodes to build a scene/campaign'.

2k tokens
Higgsfield Cinema
by OSideMedia

Guides users through professional filmmaking workflows in Higgsfield Cinema Studio, including creating multi-shot sequences, configuring optical stacks, applying color grading, managing Soul Cast AI actors, and structuring per-scene prompts with Director Panel camera movements. Use when the user mentions Cinema Studio, Cinema Studio 2.5, Cinema Studio 3.0, Soul Cast, color grading, multi-shot video, shot sequences, storyboard workflow, Hero Frame, optical stack, keyframe interpolation, Elements system (@Characters/@Locations/@Props), Speed Ramp, Director Panel, Higgsfield Popcorn, Single Shot / Multi-Shot Auto / Multi-Shot Manual modes, Reference Anchor, Smart shot control, or any professional filmmaking workflow inside Higgsfield.

29k tokens
Higgsfield Audio
by OSideMedia

> Use when the user asks about audio in Higgsfield videos, needs to add dialogue or lip-sync, wants sound effects or ambient sound in generated video, asks about music or BGM in output, or is using any audio-capable model (Kling 3.0, Seedance 1.5 Pro, Seedance 2.0, Veo 3/3.1, Grok Imagine Video). Also use when the user's prompt would benefit from audio direction but they haven't mentioned it. Also use when the user wants standalone audio — a soundtrack, ambience bed, multi-speaker scene audio (Seed Audio 1.0), or text-to-speech voiceover.

7k tokens
Higgsfield Content Factory
by OSideMedia

Use when the user wants to run a full ad-campaign pipeline on top of Higgsfield Marketing Studio — 'create a campaign', 'build a content plan', 'run the content pipeline', 'generate 100 UGC videos', 'plan and schedule a launch', 'make a batch of ads from my product', or 'how much did this save vs traditional production'. Covers the 5-stage orchestration (Research → Plan → Generate → Publish → Report), the UGC-first 5-format campaign mix (UGC Entertainment, Street Interview, Unboxing, Product Review, ASMR), the even-split allocation math, button-driven onboarding, the per-batch generation gate, and the publish + cost-report tail (satellite). Defers all Marketing Studio API ground-truth (presets, params, hooks, avatars) to higgsfield-marketing-studio.

5k tokens