Canonical short-form-ad audio mix in one FFmpeg pass. VO loudnorm + 3.0× per-clip + 2.0× mix, music 0.13 base + apad+afade, sidechain compress 20:1 @ 0.01, climax line +20%, optional video-to-music duration sync. Replaces the reactive multi-round tuning that cost v03 6+ passes.
npx skills add https://github.com/gooseworks-ai/goose-skills --skill mix-master
The v03 Ironman LEARNINGS describe 6+ tuning rounds because no default mix protocol existed: round 1 music too loud, round 2 VO too quiet, round 3 Whisper mis-transcribed, round 4 sidechain ratio finally right, round 5 music ended too early, round 6 climax sat at the same loudness as the rest.
This atom encodes the protocol that landed and ships it as the default. Operators tune by flag; they don't redesign the chain.
--video <path> — video file with no audio (or whose audio gets dropped) (required)--vo <path,path,...> — comma-separated list of per-scene VO mp3s in scene order (required)--vo-starts <ms,ms,...> — start time in milliseconds for each VO clip on the timeline (required, same length as --vo)--music <path> — music bed file (required)--sfx <path,path,...> — optional comma-separated SFX list--sfx-starts <ms,ms,...> — optional SFX start times--output <path> — output mp4 (required)--total-duration <s> — target total duration in seconds (required — video AND music are sync-fit to this)--vo-boost-line <N> — 1-indexed VO clip that is the climax; gets +20% additional volume (default unset)--vo-volume <float> — per-clip VO volume multiplier (default 3.0)--vo-mix-volume <float> — extra mix-bus volume on VO (default 2.0)--music-volume <float> — base music volume (default 0.13)--music-swell-volume <float> — peak music volume during apad/swell (default 0.21)--sidechain-ratio <float> — sidechain ratio, capped at 20 by FFmpeg (default 20)--sidechain-threshold <float> — sidechain threshold (default 0.01)--target-i <lufs> — VO loudnorm integrated target (default -16)--target-tp <dbtp> — VO loudnorm true-peak (default -1.5)--target-lra <lu> — VO loudnorm LRA (default 11)--vo and --vo-starts must be same length; --sfx and --sfx-starts must be same length; ffmpeg + ffprobe must be on PATH.ffprobe). If music shorter than --total-duration, append apad=pad_dur=2.0 then trim back to target. If music longer, trim with afade=out over last 1.5s.filter_complex_script file (long chains exceed shell arg limits):[N:a]loudnorm=I=...:TP=...:LRA=...,adelay=<start>|<start>,volume=<vo_volume>[voN];--vo-boost-line N): volume becomes vo_volume * 1.2.amix=inputs=K:duration=longest,volume=<vo_mix_volume>[vo_pre];[vo_pre]asplit=2[vo_final][vo_sc];apad=pad_dur=2.0,atrim=0:<TOTAL>,afade=t=out:st=<TOTAL-1.5>:d=1.5,volume=<music_volume>[music_base];[music_base][vo_sc]sidechaincompress=threshold=<thresh>:ratio=<ratio>:attack=20:release=600[music_ducked];adelay + volume then mix into [sfx_bus].amix=inputs=<2 or 3>:duration=longest[a_out].-c:v copy).[a_out], write manifest.json and verification.md.<output> — mixed mp4 with VO + music + SFX.manifest.json — chain parameters, vo timing, climax line, integrated LUFS after.verification.md — pass/fail report (integrated LUFS within ±0.5 of expected target, no clipping).<target-i> (post-protocol; should be approximately -14 social-ready after the boosts).<target-tp> (no clipping).<total-duration> ± 100ms.afade=out — last 100ms RMS < -30 dB.--vo-boost-line was set, the corresponding clip's volume= value is vo_volume * 1.2.apad insufficient — emit a warning; video may end before music. Operator should re-source music or shorten --total-duration.create-clips Phase 4 catches it.-c:v copy to preserve quality.scripts/mix.sh \
--video peloton/ads/video-04-new-job/edits/master-no-audio.mp4 \
--vo "peloton/ads/video-04-new-job/audio/vo-scene-01.mp3,...,peloton/ads/video-04-new-job/audio/vo-scene-14.mp3" \
--vo-starts "0,2000,3500,...,28000" \
--music peloton/ads/video-04-new-job/audio/music-cue-01.wav \
--output peloton/ads/video-04-new-job/edits/master-final.mp4 \
--total-duration 30.0 \
--vo-boost-line 14
Consumed by: video-orchestrator/edit-video Phase 3–5, auto-fix-from-review-notes dispatcher (when vo-intelligibility or music-vo-balance polish-notes items reference mix-master).
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take gooseworks-ai/mix-master from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.