3 019 media skills from 496 authors. They create and process images, video and sound. Half of them fit into 1 966 tokens or less — that is what one costs your context window when the agent loads it. 768 ship runnable scripts rather than instructions alone. 30 of them cannot work without an MCP server, most often rube. We also found 321 copies of these same skills sitting in other people's repositories — counted once here, not 321 times.
3 019 unique 496 authors 1 925 updated this month 160 from vendors
Web Audio API for JARVIS audio feedback and voice processing
使用 Scrapling + html2text 从 URL 抓取可读正文(含图片),按优先级选择器提取并按字符数截断;随后自动写入飞书文档并返回文档链接。适用于用户发送文章/博客/新闻链接(尤其是微信公众号 mp.weixin.qq.com)并希望快速验收正文内容的场景。
Generate images using Qwen Image API (Alibaba Cloud DashScope). Use when users request image generation with Chinese prompts or need high-quality AI-generated images from text descriptions.
使用通义千问·万相(wan2.6-t2i)生成漫画或动漫风格的图片。当用户说"生成漫画""用万相画漫画""生成漫画风格图片""用千问画一张二次元角色"等与漫画风格图像生成相关的请求时,执行本技能。
Manage containers using Podman, the daemonless container engine. Run rootless containers, create pods, manage images, and use Docker-compatible commands. Use when working with Podman or requiring rootless container operations.
Scan container images for vulnerabilities using Trivy, Grype, and cloud-native tools. Identify security issues in base images, packages, and configurations. Use when implementing container security, building secure images, or meeting compliance requirements.
Structure complex business problems with hypothesis-driven MECE decomposition and named strategic frameworks (Five Forces, PESTLE, SWOT, 7S, VRIO, Balanced Scorecard, Ansoff, Growth-Share Matrix, Nine-Box, Three Horizons, Value Chain, Business Model Canvas, Strategy Canvas, Platform Strategy). Use when framing an ambiguous question, building issue trees, writing testable hypotheses, designing an analytical workplan, or running framework-based analysis for competitive positioning, market entry, growth strategy, organizational alignment, or portfolio allocation. For broad or high-stakes questions, apply two or more frameworks and synthesize across them. For narrow questions, one framework applied rigorously beats two applied superficially.
Transcribe audio/video files via MLX Whisper (Apple Silicon). Usage: /transcribe path/to/file.mp3
Use first for creative work—technical or everyday—when a brief is incomplete or intent is implicit and success depends on taste, voice, feeling, or human experience. Also trigger when literal compliance may lose the point. Do not wait for explicit creative wording. Skip factual lookup, mechanical or exact tasks, and fully specified work.
> Use the `gumroad` CLI to look up and manage Gumroad data from the terminal. Trigger when the user asks about Gumroad products, files, file uploads, attachments, sales, subscribers, licenses, payouts, audience emails, broadcasts, offer codes, webhooks, refund policies, or any Gumroad data lookup. Also trigger on "check my Gumroad", "look up a sale", "verify a license", "list my products", "how much have I made", "who bought", "recent sales", "refund a sale", "create a product", "upload a file", "attach a file to a product", "add a cover image", "set a product thumbnail", "get product content", "set product content", "upload product media", "publish a product landing page", "publish custom HTML", "clear custom HTML", "customize my profile page", "publish a profile landing page", "set profile custom HTML", "attach a file to a variant", "finish a failed upload", "abort an upload", "manage webhooks", "draft an email", "preview a broadcast", "send an audience email", "list drafts", "set refund policy", "check my refund policy", "check my earnings", "see my revenue", "who subscribed", "manage my store", "discount code", "coupon", "shipping status", "payout schedule", or any request to query or act on Gumroad data — even if the user doesn't say "Gumroad" explicitly but is clearly referring to their creator store or digital product sales. Do NOT trigger for Gumroad web UI, Rails, or codebase questions.
Create cinematic HTML presentations with AI video backgrounds, deployed to GitHub Pages. Use for: slides, presentation, deck, cinematic slides, video presentation, animated slides, live presentation.
Burn subtitles onto videos using FFmpeg. Use for: hardcode subtitles, embed captions, video subtitling.
Generate images with Gemini (default) or fal.ai FLUX.2 klein 4B (--cheap for fast/low-cost). Generate videos with Grok Imagine (default) or fal.ai LTX-2 (--cheap). Use for: create image, generate visual, AI image generation, poster, video generation.
Start real-time microphone transcription using ElevenLabs Scribe v2 Realtime. Use when user wants to start live transcription, dictation, or real-time speech capture. Triggers on: 'תתחיל תמלול', 'תמלל בזמן אמת', 'start transcribing', 'live transcribe', 'הקלט מה שאני אומר'. After starting, tell user they can say 'אוקי זה מספיק בוא נעצור את התמלול' to stop, or use /live-transcribe-stop.
Create professional kinetic typography videos from scratch. Includes speech writing, TTS with emotional dynamics, music generation, and animated text. Use for: promo videos, explainers, social content, inspirational speeches, product launches.
Generate AI music with ElevenLabs Music API. Use for: background music, soundtracks, jingles, theme songs, instrumental tracks, AI music composition.
Spin up an instant browser voice session (OpenAI Realtime gpt-realtime-2) to close a topic in a short conversation instead of working through documents. Generic & white-label - works for any process. Supports live data work (read/update files, JSON, run commands), and distill mode (no tools, ends with a structured deliverable). Has a generic canvas that can display images, markdown, code, html, json, video, audio - perfect for "let's go over X" flows where the agent shows you items one by one and you react in real time. Use when user says "let's close this in a voice call", "run a quick voice session about X", "תפעיל שיחה קולית", "let's go over the [images/leads/PRs/files/notes]", or when a task is faster as a 3-minute conversation than as a document edit.
Translate video subtitles to any language with native-quality refinement. Full pipeline: transcribe → translate → refine → embed RTL-safe subtitles. Use for: translate video, תרגם סרטון, video translation, foreign subtitles, Hebrew subtitles, translated captions.
Never teach a principle without an immediate rep. When designing ANY teaching content — a workshop, a lesson plan, a deck, a course module, a webinar, an explainer video, an educational post — pair every theoretical unit with an embodiment the learner does right now. Use whenever building or reviewing teaching material, lesson plans, workshop flows, course outlines, or training content, OR when the user says 'teach-by-doing', 'add an exercise', 'make it hands-on', 'don't leave it abstract', 'תרגיל לכל עיקרון', 'הטמעה מיידית'.
Upload videos to YouTube with title, description, tags. Use for: youtube upload, publish video, share on youtube.
Fetch X (Twitter) bookmarks via the official X API v2. Downloads recent bookmarks with text, images, and videos into a local folder. Use whenever user asks to grab/download/export their X bookmarks, save bookmarked tweets, or pull recent saved posts from X/Twitter. Uses OAuth 2.0 user-context auth (one-time browser consent, then refresh-token forever).
Download YouTube videos with quality presets. Use for: download youtube, yt download, video download, youtube to whatsapp, youtube mp3.
Extract a founder's or expert's point of view — core beliefs, contrarian takes, origin stories, taste — and synthesize a recommended "one big idea" (OBI) that anchors thought leadership and brand voice. Writes to marketing/expert-pov/expert-pov.md. Triggers - "expert pov", "founder point of view", "thought leadership angle", "one big idea", "OBI", "founder beliefs", "founder narrative
Extract voice patterns from existing content (website, blog, social, sales-call transcripts) and codify into voice rules. Produces voice analysis + voice guidelines (rules per pattern + violation pattern + fix template). Writes to marketing/brand/brand-voice.md as the canonical voice rules every content skill reads. Triggers - "tone of voice", "brand voice", "voice guidelines", "TOV audit", "writing rules", "extract voice
Generate or edit images with gpt-image-2 through Codex/ChatGPT subscription authentication instead of OPENAI_API_KEY. Use for text-to-image, reference-image editing, or visual assets when the user wants local Codex auth, especially when no native image tool is available. Do not use for official OpenAI API-key billing or OpenAI-compatible gateways.
All-platform pre-publish audit and adaptation for creators publishing to 抖音/Douyin、小红书/Xiaohongshu、微信视频号/Weixin Video Accounts、快手/Kuaishou. Review videos, images, graphic posts, audio, articles, transcripts, subtitles, covers, livestream scripts, product claims, reposted material, and overseas content. Use when users need one Skill to check visible content, spoken claims, AI disclosure, sources, rights, commercial context, platform mentions, livestream commerce, or platform-specific release versions. Trigger on requests such as “能不能发”“发布前审核”“查违禁词/敏感词”“会不会限流”“检查口播/字幕/封面/导流/带货”“标注 AI/转载/演绎/广告/个人观点”, “按我的经验库审核”“保存这次审核”“复盘这条内容”.
Deepfake detection and media safety — detect AI-generated audio, images, and video, trace synthesis sources, and analyze media intelligence using direct Resemble AI API calls
Publish content directly to WordPress sites via REST API with full Gutenberg block support. Create and publish posts/pages, auto-load and select categories from website, generate SEO-optimized tags, preview articles before publishing, and generate Gutenberg blocks for tables, images, lists, and rich formatting. Use when user wants to publish to WordPress, post to blog, create WordPress article, update WordPress post, or convert markdown to Gutenberg blocks.
p5.js ile generative art, flow fields ve interactive visuals oluşturma rehberi.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Visual hierarchy, z-index, shadows, animations ve white space kuralları.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Create and revise HyperFrames HTML video compositions with CSS, GSAP timelines, composition IDs, reusable components, render commands, and proof-frame validation. Use when an agent needs to build a scene, edit a composition, render HyperFrames output, or debug HyperFrames-specific scene behavior.
Preserve visual continuity across video scenes with previous-final-frame starts, camera swipes, carryover frames, transition QA, and boundary proofing. Use when a scene must begin from the exact last frame of a previous scene, when adding camera-move transitions, when replacing scene endings, or when debugging stale starting frames.
Generate, style, preview, and burn captions, and build music/percussion beds with ducking, loop/tail selection, fade-outs, and final silence. Use when captions are requested, captions are out of sync, subtitle style needs approval, background music competes with voiceover, or music must start/end at exact moments.
Integrate raw product demo recordings into agent-created videos, including crop decisions, zoom preservation, camera swipe transitions, UI readability checks, and timing around narration. Use when a user provides screen recordings, product walkthrough clips, SaaS demos, app footage, browser recordings, or asks to transition into/out of product demo footage.
Build and revise FFmpeg-based video assemblies from scene renders and narration, including audio probing, scene retiming, pauses, trim/concat, setpts, atempo, phrase-to-visual cue alignment, and final timeline manifests. Use when narration length, visual beats, product moments, or scene boundaries must sync precisely.
End-to-end video production workflow for coding agents using HyperFrames, FFmpeg, captions, audio sync, music beds, and render QA. Use when a user asks for a complete video, multi-scene edit, launch video, product/demo video, UGC-style edit, enterprise explainer, or final assembled MP4 from prompt/script/assets.
Convert rough video prompts, scripts, reference images, and edit notes into decision-complete scene specs and storyboards for agent-led video production. Use when a user describes a scene, gives timestamped narration, asks to preserve previous scenes, references a concept image, or requests a multi-scene video plan.
Processes expense receipts and creates expense report entries following company policies with approval thresholds and validation rules. Use when user says "log this expense", "process this receipt", "create expense entry", "submit expense", "add to expense report", uploads a receipt image, or provides purchase documentation to expense.
| Call local Grok CLI, Kimi Code CLI, and Claude Code CLI with the strongest current default models and native tools. Use this qiaomu skill workflow when the user wants to run Grok 4.5 for X/web search, image/video generation, research, or multi-tool agent work; run Kimi K3 (1M) for frontend UI/CSS/React/Vue work; run Claude Fable 5, Opus 4.8, or Sonnet 5 for coding/review/refactors; run independent multi-model comparisons concurrently with native stream progress, artifact verification, private logs, cancellation, and failed-job retry; choose between the three CLIs; or wrap non-interactive CLI invocations. Trigger on phrases like "用 grok cli", "用 kimi cli", "grok 搜 X", "kimi 写前端", "qiaomu-model-cli", "调用 grok4.5", "调用 k3 1m", "用 claude code", "用 fable 5", "opus 4.8 改代码", "分别调用三个模型", "多模型横评". Not for generic cloud API registry management (use qiaomu-llm) or OpenCLI site adapters.
> Create and manage articles on the PANews platform. All operations require a valid user session. upload images, search tags, apply for a column, polish or review article content.
This skill should be used when the user asks "what does this mediainfo mean", "explain this video format", "what is BT.709", "what is H.264", "container vs codec", "why is this 10bit", "what does limited range mean", or wants educational explanations of video technical concepts.
This skill should be used when the user asks "check frame rate", "is this CFR or VFR", "video has duplicate frames", "video stutters", "frame rate issues", "why does video judder", or wants to analyze frame rate characteristics and detect timing problems.
This skill should be used when the user asks "what's wrong with this video", "why does this video look bad", "detect video artifacts", "find quality issues", "video has artifacts", "identify compression artifacts", or wants to diagnose specific quality problems in a video file.
This skill should be used when the user asks "is this interlaced", "detect telecine", "video has combing", "should I deinterlace", "3:2 pulldown", "inverse telecine", or sees horizontal lines/combing artifacts in video and wants to understand if the content is truly interlaced or telecined.
This skill should be used when the user asks "compare these videos", "which source is better", "compare blu-ray vs web", "which release should I use", "compare video quality", or needs to evaluate multiple versions of the same content to determine which has better quality.
This skill should be used when the user asks to "audit this video", "analyze video quality", "check this video file", "is this video good quality", "should I reencode this", "what format is this video", or wants to understand a video file's technical properties and quality before working with it.