3 019 media skills from 496 authors. They create and process images, video and sound. Half of them fit into 1 966 tokens or less — that is what one costs your context window when the agent loads it. 768 ship runnable scripts rather than instructions alone. 30 of them cannot work without an MCP server, most often rube. We also found 321 copies of these same skills sitting in other people's repositories — counted once here, not 321 times.
3 019 unique 496 authors 1 925 updated this month 160 from vendors
口播视频转录和口误识别。生成审查稿和删除任务清单。触发词:剪口播、处理视频、识别口误
Virtual singer MV script generator — audio analysis to video storyboard pipeline
微信 (WeChat) 与 OpenClaw 的双向集成通道。基于 Wechaty + PadLocal 实现微信消息的接收和发送,支持私聊、群聊、@提及检测、图片/文件传输。当需要通过微信与 AI 助手交互、接收微信消息触发 AI 响应、或从 OpenClaw 发送消息到微信时使用此技能。
> 公众号文章全自动流水线:写作→封面→发布。三段式编排 wechat-article-writer / code-to-image / md2wechat,
微信视频号助手网页版视频发布全流程。通过浏览器自动化操控 channels.weixin.qq.com 完成登录检测、扫码登录、上传视频、填写描述和短标题、截图确认后发布或保存草稿。触发场景:用户需要发布视频到视频号、视频号发布、视频号上传视频、发视频号。
易经占卜系统。支持铜钱法、蓍草法起卦,生成本卦、互卦、变卦,提供Oracle Voice诠释。当用户请求占卜、问卦、易经解读、或寻求决策指引时使用。
Generate complete YouTube videos from a single prompt - script, voiceover, stock footage, captions, thumbnail. Self-contained, no external modules. 100% free tools.
This skill performs deep analysis of YouTube videos through **both information channels** Multimodal YouTube video analysis through both audio (transcript) and visual (frame extraction + image analysis) channels. Especially powerful for HowTo videos, tutorials, demos, and explainer videos where what is SHOWN (screenshots, UI demos, diagrams, code, physical actions) is just as important as what is SAID. Use this skill whenever a user wants to analyze, summarize, or create step-by-step guides from YouTube videos, or when they share a YouTube URL and want to understand what happens in the video. Triggers on requests like "Analyze this YouTube video", "Create a step-by-step guide from this video", "What does this video show?", "Summarize this tutorial", or any YouTube URL shared with analysis intent.
Publish and manage content on 知识星球 (zsxq.com). Supports talk posts, Q&A, long articles, file sharing, digest/bookmark, homework tasks, and tag management. Use when publishing content to 知识星球, creating/editing posts, uploading files/images/audio, managing digests, batch publishing, or formatting content for 知识星球.
> Convert text posts into visual carousel images or presentations for Threads, Instagram, LinkedIn, TikTok, YouTube. 12 slide types (incl. image/emoji/number), 6 format presets (incl. 1920×1080 wide), 8 background styles (incl. ruled paper), 3-axis style system (font × color × purpose), highlighted keywords with optional italic-box style. Generates PNG or single-file PDF via Next.js preview + browser export. RU/EN toolbar.
> Receive and verify Bunny Stream webhooks. Use when setting up Bunny Stream webhook handlers, debugging X-BunnyStream-Signature verification, or handling video encoding events like Status 3 (Finished / encoding done), Status 5 (Failed), or captions and title/description generation.
> Receive and verify Google Gemini API webhooks. Use when setting up Gemini webhook handlers for batch jobs, video generation, or Interactions API function-calling LROs, debugging signature verification, or handling events like batch.succeeded, batch.failed, video.generated, or interaction.completed.
> Receive and verify Retell AI webhooks. Use when setting up Retell webhook handlers, debugging X-Retell-Signature verification, or handling voice call events like call_started, call_ended, call_analyzed, and transcript_updated.
> Receive and verify TikTok for Developers webhooks (Login Kit, Content Posting, Data Portability). Use when setting up TikTok webhook handlers, debugging TikTok-Signature verification, or handling events like authorization.removed, video.upload.failed, video.publish.completed, or portability.download.ready. Not for TikTok Shop — see tiktok-shop-webhooks.
> Receive and verify Twilio webhooks. Use when setting up Twilio webhook handlers, debugging X-Twilio-Signature verification, or handling communications events like incoming SMS, voice calls, message status callbacks (delivered, failed), or recording status callbacks.
>- Add an SBOM (Software Bill of Materials) generation step to an existing Harness pipeline using Harness SCS SscaOrchestration (SBOM Orchestration). Supports container images and code Syft or cdxgen, SPDX or CycloneDX, optional SBOM attestation. Only works with existing pipelines. Use when asked to create an SBOM, generate a bill of materials, add SBOM to a pipeline, scan a container image or repo for components, or set up SBOM Orchestration. scan for dependencies, SBOM Orchestration, add Generate SBOM step.
Use when the user attended one or more sessions of a conference (with audio recordings + slide photos) and needs help building (1) faithful per-session reconstructions in markdown (slide visuals + speaker transcript with Whisper hallucination annotations), and (2) a downstream report deliverable whose scope and format are decided interactively with the user. Pipeline phase (raw → mlx_whisper Chinese SRT → per-slide multimodal reconstruction → official agenda cross-check via Playwright if available) is deterministic. Report phase is interactive — always quiz user on scope (single-session / single-day / multi-day synthesis), format (existing template / free-form fallback), recipient (formal / informal), and any business workstream mapping before drafting.
Use when the user asks to generate an image via GPT/Codex (e.g. 「叫 gpt 生圖」「幫我用 gpt 生圖」「gpt 畫一個 X」). The skill drafts a Chinese + English prompt pair, iterates with the user until they explicitly approve, then dispatches Codex CLI ($imagegen skill, codex built-in image_gen) in the background, monitors progress, converts the result to a jpg in the current working directory, and writes a sidecar prompt log. Does text-to-image AND img2img — drop a reference image (on-disk file) and it runs Codex `-i` to lock a face/character across scenes.
Use when the user wants to build a consistent-identity LoRA for an original character — defining the character, generating a face/body-consistent multi-angle dataset (via the gpt-image-gen skill for codex image generation), captioning it, doing base-specific homework, training on a chosen base (Pony / Z-Image / others) on a local GPU, and producing a usable LoRA. This skill ORCHESTRATES the end-to-end pipeline and gates every expensive/irreversible step; it delegates actual image generation to gpt-image-gen and never improvises training settings from memory.
Use when the user wants a Traditional Chinese (Taiwan) language review of Markdown or plain-text files — flagging genuinely comprehension-breaking grammar faults (missing subject, broken predicate, wrong measure word, misused connectives, inconsistent naming) and non-Taiwanese wording (mainland-Chinese terms, translationese) with replacement suggestions. High bar on grammar: only reports what actually misleads the reader; pure style preferences, punctuation nits and unavoidable loanwords are let through. Phase 1 is a read-only report; Phase 2 (applying edits) only runs after the user explicitly approves. NOT for changing tone or voice (that is rewrite-tone), NOT for rewriting content, adding facts, or translating between languages.
Use when the user wants privacy-respecting web, image, news, or video search through a configured local SearXNG instance instead of external search APIs. Calls the bundled script against SEARXNG_URL, supports result limits/categories/language/time range, and can return human-readable or JSON output. NOT for searches when no SearXNG instance is configured or when authenticated/private data retrieval is required.
Source b-roll for a video edit — classify each moment, scope the search, return vetted candidates, place on the word. Use when the user says "find b-roll", "/find-broll", "source clips for this edit", or wants footage/memes/screenshots to lay over a talking-head video.
Design principles and techniques for creating technical exploded-view diagrams and engineering documentation posters.
Extract key frames (I-frames) from video files using FFmpeg to identify scene changes and important moments.
Use OpenCV for image processing, grayscale conversion, and template matching to detect objects in images.
Creating technical diagrams and exploded-view illustrations using Python Pillow with geometric primitives and text annotations.
Extract key frames (I-frames) from video files using FFmpeg for analysis and processing.
Convert images from RGB/BGR to grayscale using OpenCV and save in-place.
Count occurrences of a template object in an image using OpenCV template matching with non-maximum suppression.
Techniques for drawing isometric or orthographic exploded-view hardware diagrams programmatically using Pillow, with annotation leader lines and layer separation.
Create technical poster and diagram images using Pillow (PIL), including shapes, text, lines, and layered composition.
Extract I-frame keyframes from video files using FFmpeg command line, saving them as numbered PNG images.
Count occurrences of an object in an image using OpenCV template matching (cv2.matchTemplate). Use this skill whenever the user needs to detect and count how many times a small reference image (template) appears in a larger image, such as counting coins, enemies, or other game sprites. Works on both grayscale and color images.
Convert images to grayscale in-place using Pillow (PIL), overwriting original RGB files.
Technical image generation using Python's Pillow library. Use this when you need to programmatically create diagrams, posters, or technical drawings.
Techniques for image preprocessing and template matching to count objects in images.
Tools and techniques for extracting keyframes and processing video files using FFmpeg.
Guidelines for applying strict brand color palettes and minimalist technical design standards to generated imagery.
Use ffmpeg to extract key frames (I-frames) from a video file into a directory.
Use ImageMagick to convert an image to grayscale inplace.
Detect and count occurrences of template images within a larger target image.
Convert images to grayscale in-place.
Extract key frames (I-frames) from video files using FFmpeg.
Compute and aggregate F1 score and delta from clustering evaluation across images
Advanced OpenCV template matching with dynamic threshold tuning, robust NMS filtering, and multi-scale detection for object counting in images.
Robust FFmpeg video frame extraction with error handling, alternative methods, and comprehensive frame selection strategies.
Robust image processing with format conversion, in-place modification, and comprehensive verification of image properties.
Extract I-frame keyframes from video files using FFmpeg, with output verification and naming conventions.