Text-to-speech voice-over generation from YAML speaker notes using Azure Speech SDK with SSML pronunciation control
npx skills add https://github.com/microsoft/hve-core --skill tts-voiceover
Generates per-slide WAV voice-over files from YAML speaker_notes using Azure Speech SDK with SSML pronunciation control.
This skill reads content.yaml files from a PowerPoint skill content directory, extracts speaker_notes fields, applies SSML acronym aliases for correct pronunciation of technical terms, and produces one WAV file per slide. Supports dry-run mode for SSML template verification without Azure credentials.
SPEECH_KEY) or Microsoft Entra ID (SPEECH_RESOURCE_ID).uv for virtual environment management.SPEECH_REGION for synthesis. Operators must pin an approved region and avoid sending regulated or confidential narration.export SPEECH_KEY="your-speech-key"
export SPEECH_REGION="eastus"
Requires a custom domain on the Speech resource and Cognitive Services Speech User role.
export SPEECH_RESOURCE_ID="/subscriptions/.../Microsoft.CognitiveServices/accounts/your-resource"
export SPEECH_REGION="eastus"
Install dependencies:
# run from this skill folder
uv sync
Verify SSML templates without generating audio:
uv run scripts/generate_voiceover.py --dry-run --content-dir path/to/content
Generate voice-over WAV files:
uv run scripts/generate_voiceover.py --content-dir path/to/content --output-dir voice-over
Embed audio into a PPTX deck:
uv run scripts/embed_audio.py --input deck.pptx --audio-dir voice-over --output deck-narrated.pptx
| Parameter | Type | Default | Description |
|:----------------------|:-------|:------------------------------------|:-------------------------------------------------------------------------------------------|
| --dry-run | flag | false | Print SSML templates without generating audio |
| --voice | string | en-US-Andrew:DragonHDLatestNeural | Azure TTS voice name |
| --rate | string | +10% | Speech prosody rate |
| --content-dir | path | content | Path to slide content directory |
| --output-dir | path | voice-over | Path to WAV output directory |
| --lexicon | path | *(auto-detect)* | Custom acronyms.yaml path |
| --collapse-newlines | flag | false | Collapse newlines and whitespace runs in speaker notes into single spaces before synthesis |
| --verbose / -v | flag | false | Enable verbose (DEBUG) logging output |
Embeds WAV files into corresponding PPTX slides and adds narration timing
XML so PowerPoint recognizes the audio for video export via
File > Export > Create a Video > Use Recorded Timings and Narrations.
| Parameter | Type | Default | Description |
|:-------------------|:-----|:------------------|:--------------------------------------|
| --input | path | *(required)* | Source PPTX file path |
| --audio-dir | path | voice-over | Directory with slide-NNN.wav |
| --output | path | *-narrated.pptx | Output PPTX file path |
| --verbose / -v | flag | false | Enable verbose (DEBUG) logging output |
Generate with custom voice and rate:
uv run scripts/generate_voiceover.py \
--content-dir content \
--output-dir voice-over \
--voice "en-US-Jenny:DragonHDLatestNeural" \
--rate "+5%"
Use a custom lexicon:
uv run scripts/generate_voiceover.py \
--content-dir content \
--lexicon custom-acronyms.yaml
Collapse newlines in speaker notes (recommended for block-scalar | notes,
whose line breaks are otherwise spoken as pauses):
uv run scripts/generate_voiceover.py \
--content-dir content \
--collapse-newlines
Embed generated audio:
uv run scripts/embed_audio.py \
--input slide-deck/presentation.pptx \
--audio-dir voice-over \
--output slide-deck/presentation-narrated.pptx
The lexicon controls SSML <sub alias> replacements for acronyms and technical terms. Create an acronyms.yaml file:
acronyms:
HVE-Core: "H V E Core"
OWASP: "Oh wasp"
SBOM: "S Bomb"
SLSA: "Salsa"
CI/CD: "C I C D"
Lexicon resolution order:
--lexicon argument.acronyms.yaml in the content directory.Each slide produces an SSML document:
<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis"
xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
<voice name="en-US-Andrew:DragonHDLatestNeural">
<prosody rate="+10%">
Text with <sub alias="Oh wasp">OWASP</sub> aliases applied.
</prosody>
</voice>
</speak>
This skill reads from the PowerPoint skill's content directory structure:
content/
├── slide-001/
│ └── content.yaml # Must include speaker_notes: field
├── slide-002/
│ └── content.yaml
└── ...
Each content.yaml should contain a speaker_notes: field with the narration text. The generated WAV files are named slide-NNN.wav matching the directory names.
| Issue | Solution |
|:-----------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| Set SPEECH_KEY ... or SPEECH_RESOURCE_ID | Export SPEECH_KEY (key auth) or SPEECH_RESOURCE_ID (Entra ID) with SPEECH_REGION. |
| 401 with Entra ID auth | Verify custom domain on the Speech resource and Cognitive Services Speech User role. RBAC propagation takes up to 5 minutes. |
| Empty WAV files or skipped slides | Verify speaker_notes: is present and non-empty in content.yaml. |
| Mispronounced acronyms | Add entries to acronyms.yaml with phonetic aliases. |
| azure-cognitiveservices-speech package is required | Run uv sync in the skill directory. |
| Audio icon visible in PPTX | Reposition or resize the audio object in PowerPoint after embedding. |
| Authored slide animations missing after embedding | embed_audio.py replaces existing p:timing with narration timing; re-apply animations in PowerPoint after embedding audio. |
| Slides no longer advance on click after embedding | embed_audio.py sets advClick="0" for auto-advance. To re-enable, select all slides in PowerPoint and check Advance Slide > On Mouse Click in the Transitions tab. |
| Video export shows "No timings recorded" | Re-embed audio with the updated embed_audio.py which adds narration timing XML automatically. |
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take microsoft/tts-voiceover from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.