mcpbeat Sign in

Claude Skills

The open format is called Agent Skills and works in Claude Code, Codex, Cursor and other agents — most people know it as Claude Skills.

Every Agent Skill we could find on GitHub, deduplicated by content. 79 870 files from 1 769 authors, of which 62 217 are unique — the rest is the same skill repackaged into someone else's repository. For each one: what it weighs in tokens, whether it ships runnable scripts, and which MCP servers it needs.

62 217
unique skills
out of 79 870 files found on GitHub
17 653
are copies
same content, someone else's repository
1 743
tokens, median
what a typical skill costs you in context
7 935
name collisions
two skills with one name cannot sit side by side

38 041–38 100 of 62 217

page 635 of 1 037
Resemble Chatterbox
by calesthio

Use Resemble AI Chatterbox for text-to-speech and voice-cloning workflows, including local open-weight Chatterbox, Chatterbox Multilingual, Chatterbox Turbo, and Resemble-hosted Chatterbox API routes. Apply when an agent must choose a Chatterbox variant, prepare reference-voice inputs, control emotion/paralinguistic delivery, plan multilingual TTS, handle PerTh watermarking, or make production decisions about consent, privacy, licensing, latency, and audio QA.

7k tokens
Minimax Speech
by calesthio

Use this skill when producing speech, narration, dubbing, localization, advertising voice, voice-clone previews, or interactive voice with MiniMax speech/audio APIs. It covers MiniMax T2A HTTP, WebSocket streaming, async long-form TTS, Speech 2.8/2.6/02 model selection, system and custom voices, rapid voice cloning, voice design, pronunciation/language/emotion controls, pricing and limits, artifact custody, consent and rights checks, and production QA.

9k tokens
Topaz Video Enhancement
by calesthio

>- Use when enhancing, upscaling, restoring, denoising, sharpening, deinterlacing, stabilizing, motion-deblurring, colorizing, SDR-to-HDR converting, or frame-rate/slow-motion converting video (and secondarily images) with Topaz Labs — whether cleaning up AI-generated video, compressed UGC, or old archival footage, or building a generate-low-res-then-upscale delivery pipeline. Covers choosing the right named Topaz model for a specific footage problem, setting parameters, driving the Topaz Platform (Video/Image) REST API's async job lifecycle, judging when enhancement helps versus when it introduces artifacts (over-smoothing, temporal shimmer, face hallucination), and QA of enhanced output. Do not use for originating new video content, non-Topaz upscalers, or purely editorial cuts.

9k tokens
Alibaba Wan Video
by calesthio

Plan, implement, and review Alibaba Wan video generation using Alibaba Cloud Model Studio/DashScope hosted APIs or official Wan 2.1/2.2 open-weight checkpoints. Use for Wan text-to-video, image-to-video, first/last-frame or continuation, reference-to-video, video editing, speech-driven video, character animation, regional deployment, billing, licensing, and production safety.

14k tokens
Midjourney Video
by calesthio

Plan and direct Midjourney Video V1 image-to-video work through the official website or Discord, with human operator handoff, motion prompts, start/end frames, loops, extensions, resolution and batch budgeting, privacy, rights, safety, and delivery QA. Use when creating Midjourney videos or when determining whether a requested API, automation, plan, or production workflow is supported.

9k tokens
Higgsfield Video
by calesthio

Produce video (and its supporting stills) through Higgsfield (higgsfield.ai), a multi-model creative platform that fronts third-party video models (Kling, Veo, Seedance, Wan, MiniMax, and, until retirement, Sora) plus Higgsfield's own control layer — Soul / Soul ID image generation, the DoP camera-preset engine, and the Cinema Studio filmmaking suite. Use this skill when a request routes generation through Higgsfield specifically, when the user wants aggregator-level model switching and cinematic camera/effect presets in one workspace, when planning character consistency via Soul ID, or when deciding whether to go through Higgsfield versus direct to an underlying model provider. Covers model selection, per-model prompt strategy, credit economics, the (partial) API surface, capability limits, and rights/consent policy for generated and uploaded real-person footage.

10k tokens
Kling Video
by calesthio

Plan, prompt, generate, reference, edit, motion-control, and quality-review videos with Kuaishou's direct Kling AI video platform and API, especially Kling VIDEO 3.0, 3.0 Omni, 3.0 Turbo, native audio, multi-shot, elements, start/end frames, and image/video references. Use for first-party Kling video production or API integration, not Kolors image generation, avatars, standalone audio, consumer-only membership advice, or third-party Kling gateways.

14k tokens
Luma Ray Video
by calesthio

Direct and operate Luma Ray video generation through the current Luma Agents API and distinguish it from the consumer Luma App and legacy Dream Machine API. Use for Ray 3.2 text/image/multi-keyframe video, edits, extends, reframes, HDR/EXR, pricing and approvals, async delivery, migration, rights, privacy, and production QA.

7k tokens
Amazon Nova Reel
by calesthio

Produce and operate Amazon Nova Reel video-generation jobs through Amazon Bedrock. Use for Nova Reel text-to-video, image-conditioned animation, automated or manual multi-shot storyboards, async S3 delivery, cost and approval gates, production continuity, output custody, and AWS-specific safety, privacy, IAM, and provenance decisions.

9k tokens
Google Veo
by calesthio

Direct production with Google DeepMind's Veo video-generation family across the Gemini Developer API and Google Cloud Gemini Enterprise Agent Platform (formerly Vertex AI). Use for Veo route and model selection, text/image/reference/first-last-frame/extension workflows, native-audio prompting, camera direction, async API operation, troubleshooting, QA, and responsible commercial content production.

15k tokens
Google Gemini Omni Video
by calesthio

Generate, reference, and conversationally edit short videos with Google's Gemini Omni Flash through the Gemini Developer API or Gemini Enterprise Agent Platform. Use when a task specifically needs Gemini Omni video, multimodal reference roles, time-directed prompting, audio-aware generation, Interactions API state, or a precise comparison with Veo, Nano Banana, consumer Gemini, or gateway routes.

15k tokens
Ltx 2 Video
by calesthio

Generate and edit synchronized audio-video with LTX-2.3 using the official hosted LTX API or official local/open-weight LTX-2 repository. Use for text-to-video, image-to-video, audio-to-video, retake, extend, HDR conversion, local inference, checkpoint selection, quantization, and LTX-specific production planning.

15k tokens
Minimax Hailuo Video
by calesthio

Use MiniMax's first-party Hailuo video API safely and reproducibly across global and mainland-China platforms. Covers current text-to-video, image-to-video, first/last-frame, and subject-reference models; prompt and camera craft; region/account routing; exact cost approval; asynchronous jobs and callbacks; artifact provenance; and rights, likeness, disclosure, and data-governance review. Use when designing, implementing, debugging, or evaluating direct MiniMax/Hailuo video-generation API workflows. Do not use for the consumer Hailuo site, Video Agent templates, gateways, or unrelated MiniMax modalities.

17k tokens
Moonvalley Marey
by calesthio

>- Produce video with Moonvalley's Marey model family (Marey Realism v1.5) — a filmmaker-oriented, 1080p/24fps generative video model marketed as trained exclusively on licensed data. Use this skill when a task asks for Marey or Moonvalley specifically, when a brief demands brand-safe / legally reviewed AI video for commercial or studio work, or when the request needs Marey's director-style controls (camera trajectory, motion transfer, pose transfer, keyframing, reference conditioning). It covers what "commercially safe" really means versus the actual terms of service, access routes (Moonvalley web / Voyager, the fal API, ComfyUI), pricing, capability limits, prompt and reference strategy, where Marey fits production versus where other models win, iteration, and quality review. Do not use it for audio generation, for photoreal talking-head dialogue, or as a generic "best video model" default.

8k tokens
Pika Video
by calesthio

>- Produce short-form video with Pika (pika.art) — its effect-driven tools (Pikaffects, Pikadditions, Pikaswaps, Pikaframes, Pikascenes, Pikatwists) and audio-driven lip sync (Pikaformance), plus text-to-video and image-to-video. Use when a task calls for playful, stylized, meme-able social clips, or for applying generative effects and object swaps to real uploaded footage, and when deciding whether Pika is the right engine versus a higher-fidelity model. Covers model versions, credit/plan economics, the fal.ai API route, prompt construction, iteration/repair, capability limits, and consent/rights rules for uploaded real-person media. Not a general "best video model" chooser.

8k tokens
Tencent Hunyuanvideo
by calesthio

Generate and operate Tencent Hunyuan video through the managed TokenHub HY-Video-1.5 API or official local HunyuanVideo repositories. Use for text-to-video, image-to-video, hosted job lifecycle, pricing and region planning, local checkpoint selection, hardware and acceleration, prompt design, licensing, safety, provenance, and delivery QA.

12k tokens
Runway Video
by calesthio

Build and operate production-safe Runway API video generation, video editing, and character-performance workflows with native Runway models, explicit paid-call approval, duplicate-create protection, asynchronous task handling, secure media transfer, and evidence-aware governance.

23k tokens
Video Generation Gateways
by calesthio

>- Select, integrate, and operate multi-model video-generation gateways — hosted inference aggregators such as fal.ai, Replicate, and WaveSpeed that expose many third-party video (and video-adjacent audio) models behind one account, one API surface, and one bill. Use when deciding whether to route video generation through a gateway versus a direct model-provider API; when comparing gateways on catalog, async/queue job design, webhooks, pricing, schema discovery, and version pinning; and when operating gateway workloads in production — spend controls, retries, cross-gateway failover, content- safety differences, data retention, and commercial-rights passthrough of the underlying model. Not for direct model-provider APIs, local/self-hosted inference, model training/fine-tuning, or image-only gateway use.

12k tokens
Nvidia Cosmos Video
by calesthio

Select, run, and govern NVIDIA Cosmos world-video generation across Cosmos 3 Generator, Predict2.5, Transfer2.5, downloadable checkpoints, self-hosted NIMs, and hosted preview surfaces. Use for text/image/video-to-world, controlled world transfer, multiview or action-conditioned physical-AI video, local checkpoint pinning, hardware planning, inference approval, and output validation; do not use for Cosmos Reason/Embed-only tasks.

9k tokens
Seedance 2 0
by calesthio

Direct ByteDance Dreamina Seedance 2.0 Standard, Fast, and Mini video production across BytePlus ModelArk and verified gateways. Use for text/image/reference-to-video, native synchronized audio and dialogue, multimodal image-video-audio reference, first/last-frame animation, video editing or extension, model/gateway selection, prompt construction, async task handling, failure recovery, QA, and rights-safe production.

15k tokens
Vidu Video
by calesthio

Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame, multi-subject reference, media-reference, or Q2 multi-frame work; do not conflate the API with vidu.com consumer plans or Vidu-S1 streaming digital humans.

11k tokens
Twelvelabs Video Understanding
by calesthio

>- indexing footage, semantic/visual search across an archive, generating descriptions, summaries, chapters, highlights, and tags from video, producing multimodal embeddings, or wiring TwelveLabs into a media-production pipeline (NLE panels, logging, compliance review, metadata). Covers the current model families (Marengo for search/embeddings, Pegasus for video-to-text analysis), the v1.3 Video Understanding API (indexes, assets, tasks, search, analyze, embed), prompt construction for analysis, capability and format limits, pricing/quota math, output quality review (hallucination and timestamp accuracy), and privacy/rights obligations when footage shows real people. Not for generating or editing video pixels — this is analysis and retrieval.

9k tokens
Elevenlabs Agents
by calesthio

Build, configure, and ship production voice agents on the ElevenLabs Agents platform (branded "ElevenAgents," formerly "Conversational AI"). Use when an agent must design a spoken-conversation system prompt, pick an LLM and TTS voice/model for a real-time voice bot, wire client/server/system tools and knowledge-base RAG, connect a channel (WebRTC/WebSocket SDK, embeddable widget, or telephony via Twilio/SIP), tune turn-taking and latency, set up simulation testing and evaluation, or reason about pricing, concurrency, and the consent/disclosure obligations of a synthetic voice talking to real people. Not for plain non-conversational text-to-speech, dubbing, or music generation.

11k tokens
Gemini Live Audio
by calesthio

Build and evaluate low-latency spoken, multimodal, and translation experiences with Google Gemini Live API and Gemini Enterprise Agent Platform Live API. Use when a media-production or voice-agent workflow needs real-time audio input/output, barge-in, voice configuration, live transcription, live translation, tool/function calling during a spoken session, WebSocket session design, latency QA, quotas/cost review, or Google/Vertex data-governance tradeoffs.

11k tokens
Xai Grok Imagine Video
by calesthio

Produce, edit, extend, and govern short videos with xAI's direct Grok Imagine Video API. Use for text-to-video, image-to-video, multi-reference video, natural-language video edits, video continuation, exact media-cost approval, asynchronous request recovery, Files/Batch integration, moderation review, and API privacy or residency decisions. Covers `grok-imagine-video` and image-only `grok-imagine-video-1.5`; excludes Grok consumer subscriptions, Grok on X, still-image generation, voice APIs, and third-party gateways.

20k tokens
Hume Evi
by calesthio

Build production realtime voice agents with Hume's Empathic Voice Interface (EVI): speech-to-speech sessions, empathic/prosodic response design, EVI 3 versus EVI 4-mini selection, WebSocket and SDK integration, tools/function calling, interruption/barge-in, turn-taking, transcripts/audio artifacts, pricing/limits, privacy, safety, consent, and QA.

11k tokens
Openai Realtime Voice
by calesthio

Build production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket transport, voice activity detection, tool/function calling, prompt design, logging, consent, privacy, safety, latency, and cost controls. Use for speech-to-speech agents and live voice UX; do not use for separate request-based transcription, text-to-speech, or offline audio generation work.

10k tokens
Odyssey Interactive Video
by calesthio

>- Plan and scope production work with Odyssey's real-time interactive video world models (odyssey.ml — Odyssey-1/Odyssey-2/Starchild-1/Agora-1). Use when a request involves generating video that a user steers live with keyboard, text, speech, or controller input (streamed frames that respond in the moment), or when someone must decide between Odyssey-style interactive generation and offline clip generators (Sora/Veo/Kling) or downloadable 3D / game engines. Covers what is verified versus announced, current access routes, latency and coherence limits, feasible-now versus speculative use cases, pricing/access terms, and content/rights considerations. This is a fast-moving research-stage product; the skill teaches honest scoping, not a promise of production reliability.

9k tokens
World Labs Marble
by calesthio

>- Generate persistent, explorable 3D worlds/environments with Marble by World Labs — from text, a single image, multiple images, video, or a 360 panorama, plus coarse-structure blocking with Chisel. Use this skill when a task needs a navigable 3D scene, VR/immersive backdrop, virtual-production environment, previz set, game-blockout, or web/engine-ready Gaussian splat or mesh, and you must choose an input mode, a Marble model, an export format, and a plan/API route, then review world quality. Covers the Marble web app and the World API, editing/expansion, output formats and what each supports downstream, import into Unity/Unreal/Blender/Houdini/web, pricing/credits, capability limits, and rights/licensing. Not for flat text-to-image, text-to-video clips, or real-time on-the-fly world models (Marble outputs are downloadable, static 3D environments).

8k tokens
Write Changelog
by Automattic
2k tokens
Verify Local Changes
by googleapis
vendor

Verifies local Java SDK changes.

844 tokens
Sv Number
by sv-number

>- Order a private phone number in any country over the API, receive the SMS verification code, and hand the number back. Use when an agent has to pass a phone check during signup, login or account recovery, or needs a TOTP second factor computed locally.

6k tokens scripts
Audio Mix
by ZJU-REAL

音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audio_mix.py。与 audio-editing concat 区别:concat 是前后顺序拼接,本 SKILL 是同时叠加混音;与 video-editing bgm 区别:那个给视频配乐,本 SKILL 输出纯音频。

824 tokens zh
Audio Denoise
by ZJU-REAL

>- 音频降噪:去除录音中的背景噪声、电流声、风噪、嗡嗡声,基于 ffmpeg 滤镜链(afftdn/highpass/lowpass)。 当用户说"降噪""去噪""去杂音""消除背景噪声""电流声""风噪""录音有杂音""音频降噪"时使用。 和 audio-editing 的区别:audio-editing 做剪辑/转码/音量等通用音频操作(内置 denoise 兜底),本 SKILL 专做降噪调参。

1k tokens zh
AI Music
by ZJU-REAL

AI 音乐 / BGM 生成:给短视频、社媒内容生成原创背景音乐 / 配乐 / 纯音乐。通过可插拔 provider(阿里 DashScope / Suno 类第三方 API)文生音乐,异步提交→轮询→下载,产物可再裁剪/归一化或加到视频。当用户说“AI 音乐”“AI 配乐”“生成 BGM”“背景音乐”“原创音乐”“AI 作曲”“纯音乐”“给视频配乐”“做首曲子”时使用。与 tts-voiceover 的区别:tts-voiceover 生成人声口播/旁白,ai-music 生成背景音乐/配乐(无人声或带演唱)。

2k tokens zh
Asset Manager
by ZJU-REAL

> outputs/ 目录下的产物管理:按日期/平台/类型归档、打标签、搜索历史内容、生成素材清单。 当用户说"整理素材"、"归档"、"找之前的内容"、"搜索历史"、"素材管理"、 "outputs 整理"、"之前做的"时使用。 与其他 produce 层 SKILL 的区别:produce 层负责生成内容, asset-manager 负责生成后的产物管理(归档、检索、标签)。

6k tokens scripts zh
Audio Editing
by ZJU-REAL

通用音频处理:音频剪辑/裁剪、格式转码(mp3/wav/m4a/aac)、音量归一化、从视频提取音轨、多段拼接、淡入淡出、变速(保音高)。当用户说“剪音频”“裁一段”“转成 mp3”“调音量/响度”“提取音轨/扒音频”“拼接音频”“淡入淡出”“加速/减速音频”“变速不变调”时使用。与 audio-denoise 的区别:audio-denoise 专做降噪,audio-editing 做除降噪外的通用音频操作(也内置 denoise 作为兜底)。

1k tokens zh
Audio Visualizer
by ZJU-REAL

音频可视化视频:把纯音频(播客片段、音乐、口播金句、电台)渲染成带动态波形/频谱的视频,配封面和标题,好发到抖音/B站/视频号等只收视频的平台。当用户说 音频可视化、音频转视频、播客做成视频、音频波形视频、音乐可视化、给音频配画面、声波视频、频谱视频、把音频发到视频平台、电台切片视频 时使用。基于 shared/scripts/audio_viz.py(ffmpeg showwaves/showcqt/showspectrum)。与 audio-mix 区别:那个输出音频,本 SKILL 输出视频;与 slideshow-video 区别:那个用图片,本 SKILL 用音频驱动画面。

939 tokens zh
AI Video Gen
by ZJU-REAL

AI 视频生成:文生视频 / 图生视频 / 数字人首帧驱动。通过可插拔 provider(通义万相 Wan / 火山 Seedance / 快手可灵 / OpenAI 兼容)异步生成视频,用户自备 API key。当用户说 AI 视频生成、文生视频、图生视频、AI 生成视频、AI 短视频、让图片动起来、数字人视频、生成一段视频 时使用。与 video-strategy(选型/策略)、video-editing(剪辑处理)、clipify(切片)区别:本 SKILL 是从 0 用 AI 生成新视频。

2k tokens zh
AI Image Gen
by ZJU-REAL

通用 AI 生图:文生图 / 图生图 / 图像变体。当用户说 AI 生图、AI 画图、文生图、图生图、生成图片、生成配图、图像生成、AI 出图、AI 作图、换图、改图、图像编辑、给我画一张、生成一张图 时使用。支持 OpenAI 兼容 API 与 apimart 异步 API,用户自备 API key。

2k tokens zh
Auto Short Video
by ZJU-REAL

一句话主题 → 成品短视频:自动串联 文案→配图/AI视频→配音→字幕→BGM→合成,把 Easel 制作层零件编排成一条'一键出片'流水线。**单条视频、口播/资讯向,画面默认逐句配图 + Ken Burns 缓动,需要动态时才逐段图生视频**。当用户说 一键生成视频、自动做短视频、主题生成视频、帮我做条视频、口播视频一条龙、自动出片、短视频一键生成 时使用。**有剧情/角色/对白/反转/多集的短剧改用 short-drama(每镜强制图生视频、不用静态图冒充);只写脚本用 video-script;只生成单个片段用 ai-video-gen。**

10k tokens scripts zh
Card Quote
by ZJU-REAL

生成适合微博、知乎、公众号或 X/Twitter 分享的 16:9 横版金句卡和数据卡。当用户说“做金句卡、语录卡、数据卡、横版分享卡”时使用。小红书竖版知识卡用 card-xiaohongshu,竖版营销海报用 poster-hero。

2k tokens zh
Batch Process
by ZJU-REAL

批量处理:对一个目录里的一批图片/视频/音频统一套用同一操作——批量压缩、加水印、转格式、缩放、转比例、音量归一化等。当用户说 批量处理、批量压缩、批量加水印、批量转格式、一批图片/视频、给这个文件夹、全部转成、批量缩放、批量转竖版、整个目录 时使用。基于 shared/scripts/batch_process.py(委派 image_ops/video_ops/audio_ops)。与 image-editing/video-editing/audio-editing 区别:那些处理单文件,本 SKILL 批量套用到整个目录。

833 tokens zh
Auto Subtitle
by ZJU-REAL

自动字幕 / 语音转字幕:把音频或视频里的语音识别成字幕文件(SRT/ASS/TXT/JSON),可选把字幕烧录进视频。当用户说自动字幕、语音转字幕、视频加字幕、上字幕、转录、听写、字幕文件、生成字幕、烧字幕时使用。

1k tokens zh
Beat Sync Video
by ZJU-REAL

音乐卡点视频 / 踩点视频:检测背景音乐的节拍,让图片或片段在节拍点上切换,配推进/白闪特效,做出燃系'卡点'短视频。当用户说 卡点视频、踩点视频、音乐卡点、节奏卡点、按音乐切换、鼓点视频、踩节奏、beat 卡点、图片卡点、卡点混剪 时使用。基于 shared/scripts/beatsync.py(librosa 节拍检测)。与 slideshow-video 区别:slideshow 每图固定时长、柔和转场,本 SKILL 切换点由音乐节拍决定、硬切踩点带特效。

930 tokens zh
Card Design
by ZJU-REAL

> 社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。 当用户说“卡片难看、优化视觉、封面/海报设计、排版配色、卡片填不满、AI 味重”时使用;所有 card-*、poster-hero 等视觉产出应把它作为设计规范,而非最终渲染器。

10k tokens scripts zh
Chart Visualization
by ZJU-REAL

将数据可视化为图表。当用户需要生成柱状图、折线图、饼图、散点图、雷达图、桑基图、思维导图、流程图等图表时调用此技能,通过 curl 工具调用 AntV API 生成图表图片。产出静态图片 URL(25+ 类型);要本地渲染的信息图/GIF 动画图表用 infographic,要 CSV/JSON→整页报告用 data-report

2k tokens zh
Card Xiaohongshu
by ZJU-REAL

把已有卡片文案渲染为 1080×1440 小红书竖版知识卡片组,并按 card-design 选择视觉风格。当用户说“渲染/制作小红书卡片、知识卡、滑动卡片组”时使用。整套笔记策划与文案用 xhs-note-creator;横版金句卡用 card-quote。

2k tokens zh
Clipify
by ZJU-REAL

>- 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。

74k tokens scripts
Copywriting
by ZJU-REAL

> 国内带货转化营销文案:提炼卖点并产出种草、信息流广告、活动促销、电商详情页或落地页的标题、正文和 CTA。 当用户说“写种草/广告/活动/促销/详情页/落地页文案、提炼卖点、广告语”时使用。 整套小红书笔记用 xhs-note-creator;涨粉互动内容用 social-content;严格套 PAS/AIDA 等框架用 post-formatter。

4k tokens zh
Comparison Card
by ZJU-REAL

>- 对比图/一图流:生成 A vs B 参数对比图、优劣势对比表、产品参数一图流。 用 HTML+CSS 渲染成可截图的视觉卡片,适合小红书/微博等平台分享。 使用时机:用户说"做个对比图"、"A vs B"、"参数对比"、"优劣对比"、 "一图流"、"对比表"、"哪个好"、"对比一下"。 和 chart-visualization 的区别:chart 做数据图表(柱状图/折线图),comparison-card 做对比表/一图流。 和 infographic 的区别:infographic 做多维信息图,comparison-card 专注 A vs B 对比。

4k tokens zh
Doc Convert
by ZJU-REAL

把 Markdown 文稿排版并转换为 HTML、可打印 PDF 或长图 PNG。当用户说“Markdown/MD 转 HTML/PDF/图片、文章导出长图、MD 排版/渲染”时使用。仅处理 Markdown;DOCX/PPT 不在范围内,层级大纲转脑图用 mindmap。

532 tokens zh
Data Report
by ZJU-REAL

把 CSV、Excel 或 JSON 数据生成包含 KPI、图表和洞察的完整可视化报告页。当用户说“数据报告、CSV/Excel 转报告、做 KPI 看板、生成可视化报告页”时使用。单张图表用 chart-visualization;信息图或 GIF 动画图表用 infographic。

7k tokens scripts zh
Ecom Details Image
by ZJU-REAL

>- 生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt。 当用户说"商品主图""详情页视觉""电商配图方案""商品场景图""带货视觉""产品视觉方向""详情页设计"时使用。 本 SKILL 出视觉方案+生图 Prompt(策划);实际抠白底图用 remove-bg,实际生成图片用 ai-image-gen。

30k tokens scripts zh
Green Screen
by ZJU-REAL

绿幕抠像 / 换背景 / 合成:把绿幕(或蓝幕/指定色)拍摄的前景人物抠出来,合成到新背景——图片、视频、纯色或前景自身模糊。当用户说 绿幕、抠像、抠图换背景、去绿幕、chromakey、绿幕合成、换背景、蓝幕、把绿幕背景换掉、人物抠出来、绿布 时使用。基于 shared/scripts/chromakey.py(ffmpeg chromakey + despill)。与 video-reframe 区别:reframe 只改画幅不换背景;与 ai-image-gen 区别:那个生成新图,本 SKILL 处理已拍的绿幕素材。

816 tokens zh
Image Editing
by ZJU-REAL

>- 通用图像处理加工:改尺寸/缩放、裁剪、补边适配平台尺寸、格式转换(png/jpg/webp)、 压缩到目标大小、加文字或图片水印、圆角、多图拼接、生成缩略图、读图片信息。 基于 image_ops.py 确定性处理。 使用时机:用户说"改尺寸"、"缩放图片"、"裁剪"、"压缩图片"、"加水印"、"转格式"、 和 card-*(卡片类)的区别:card-* 是"HTML 设计→渲染出图", image-editing 只加工已有图片,不负责视觉设计。

2k tokens zh
Image Enhance
by ZJU-REAL

图片增强 / 放大 / 变清晰:高质量放大(Lanczos 2x/4x)+ 去噪 + 锐化 + 自动对比度/饱和度,改善偏糊、偏暗、噪点多的图片。当用户说 图片放大、图片变清晰、提高清晰度、图片增强、去噪点、锐化、图片太糊了、放大到高清、提升画质、优化图片、图片调亮调色 时使用。基于 shared/scripts/img_enhance.py(Pillow+OpenCV)。注意:这是传统增强非 AI 超分,凭空生成细节请用 ai-image-gen 图生图。与 image-editing 区别:那个做缩放/裁切/水印等常规操作,本 SKILL 专做画质提升。

749 tokens zh
Infographic
by ZJU-REAL

将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

9k tokens scripts zh
Meme Generator
by ZJU-REAL

表情包 / Meme 生成:给图片加经典上下大字(白字黑边)做梗图,或在图上/下加配文条做反应图('当…的时候'格式)。中英文都支持,自动换行和字号自适应。当用户说 表情包、做表情包、meme、梗图、reaction 图、配图加字、给这张图加字、做个梗、反应图、当xx的时候 时使用。基于 shared/scripts/meme_ops.py(Pillow)。与 card-quote 区别:card-quote 做精致金句卡片,meme-generator 做梗图/表情包;与 image-editing watermark 区别:那个加水印,本 SKILL 加梗字。

791 tokens zh
Mindmap
by ZJU-REAL

思维导图:把 Markdown 大纲(标题层级 + 列表)渲染成可交互思维导图 HTML,可选导出 PNG。适合知识结构、内容框架、SWOT、脑图梳理。当用户说 思维导图、脑图、mindmap、知识导图、大纲图、把要点做成脑图、内容结构图、SWOT图、树状图 时使用。基于 shared/scripts/mindmap.py(markmap 自包含 HTML + Chromium 渲染 PNG)。与 chart-visualization/infographic 区别:那些做数据图表/信息图,本 SKILL 专做层级大纲思维导图。

642 tokens zh

Claude Skills — questions

Answers built from the skills we actually parsed.

What is a Claude Skill?
A folder with a SKILL.md file: instructions that teach an agent to do one thing well, optionally with scripts and reference files alongside. The format is open and called Agent Skills — Claude Code, Codex and other agents read the same files. It is not a program you run; it is knowledge the agent loads when the task calls for it.
How is a skill different from an MCP server?
A server gives the agent new abilities — it connects to something and exposes tools. A skill gives the agent knowledge: how to use what it already has. They combine, and often literally: 11 541 of the skills here declare which MCP servers they need to work.
Why are there fewer skills here than in other catalogues?
Because we deduplicate by content. Of 79 870 files found on GitHub, 62 217 are unique — the rest is the same skill copied into someone else's repository, word for word. Catalogues that count files rather than skills show every copy as a separate entry.
What does the token count mean?
A skill is loaded into the model's context when it is used, so its size is a running cost on every request that touches it. We measure the whole folder, not just SKILL.md: one official skill is 377 tokens, another drags 83 files of fonts behind it.
How do I install a skill?
Copy the skill folder into ~/.claude/skills for personal use, or into .claude/skills inside a project. The agent picks it up by the name in the SKILL.md header — which is worth checking: 7 935 skills here share a name with another skill, and two of them cannot sit side by side.