| ListenHub CLI skills router. Routes to the correct skill based on user intent. "generate image", "generate video", "做播客", "解说视频", "朗读", "生成图片", "生成视频", "幻灯片", "slides", "音乐", "music", "generate music", "翻唱", "cover song", "pixverse", "口型", "lipsync", "对口型", "parse URL", "解析链接", "提取内容".
npx skills add https://github.com/joeseesun/qiaomu-cut-skill --skill listenhub-cli
This is a router skill. When users trigger a general ListenHub action, this skill identifies the intent and delegates to the appropriate specialized skill.
| User intent | Keywords | Route to |
|-------------|----------|----------|
| ListenHub Voice end-to-end audio | "端到端音频", "语音生成", "图片转音频", "图片生成音频", "多音色对白", "参考音频克隆", "克隆音色", "音效", "生成音效" | /listenhub-voice |
| Podcast | "podcast", "播客", "debate", "dialogue" | /podcast |
| Explainer video | "explainer", "解说视频", "tutorial video" | /explainer |
| Slides / PPT | "slides", "幻灯片", "PPT", "presentation" | /slides |
| TTS / Read aloud | "TTS", "read aloud", "朗读", "配音", "语音合成" | /tts |
| Image generation | "generate image", "画一张", "生成图片", "AI图" | /image-gen |
| Video generation | "video", "视频", "seedance", "pixverse", "生成视频", "text to video", "做视频", "口型", "lipsync", "对口型" | /video-gen |
| Music | "music", "音乐", "生成音乐", "翻唱", "cover" | /music |
| Content extraction | "parse URL", "extract content", "解析链接" | /content-parser |
| Audio transcription | "transcribe", "ASR", "语音转文字" | /asr |
| Creator workflow | "创作", "写公众号", "小红书", "口播" | /creator |
If the intent is ambiguous, ask the user to clarify:
Question: "What would you like to create?"
Options:
- "ListenHub Voice" — End-to-end audio: sound effects, multi-voice dialogue, reference-audio cloning, image→audio
- "Podcast" — Audio discussion on a topic
- "Explainer Video" — Narrated video with AI visuals
- "Slides" — Slide deck / presentation
- "Music" — AI-generated music or cover
- "Video" — AI video generation from text or reference materials
Most skills require the ListenHub CLI. Check:
listenhub auth status --json
If not installed or not logged in, handle automatically — never ask the user to run install commands:
npm install -g @marswave/listenhub-cli silentlylistenhub auth login (opens browser)Exception: /asr runs locally and needs no CLI or API key.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take joeseesun/listenhub-cli from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npm.
Without those the skill loads but fails at the first command.