mcpbeat Sign in

Render Narrated Ugc Wardrobe Stitch Agent Skill

Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product B-roll) are each trimmed to their EDL window built from the VO's Whisper word boundaries and hard-concatenated via filter_complex concat (never the demuxer, which drops audio on a duration mismatch), the VO mixed over an optional sidechain-ducked instrumental bed (−20dB, 20 to 1) so the VO stays on top, karaoke-pop captions burned on every word throughout (VEED Whisper preset, re-spelled against the locked script), a landing-page scroll rendered as FFmpeg zoompan over a Playwright PNG (not i2v), and closed on the brand's real end-card PNG — never AI-rendered text. This is the FREE deterministic assembly stage (trim-to-EDL + filter_complex concat + VO and music mix + karaoke captions + landing-page zoompan + end-card append); the VO, creator, start-frames, and clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-image-fal / create-video-fal. Use for the narrated-ugc-wardrobe-stitch format.

7k tokens
context cost
the whole folder, loaded on every use
6
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
1086
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/gooseworks-ai/goose-skills --skill render-narrated-ugc-wardrobe-stitch

What comes with it

23 937 bytes besides the instruction
scripts/PIPELINE.md
scripts/README.md
scripts/config.example.json
skill.meta.json
tests/smoke-test.md

The instruction itself

3 sections, as written by the author

render-narrated-ugc-wardrobe-stitch

Assemble a narrated-UGC "stitch reply" ad from a config: a fast-cut vertical testimonial

where a single spoken VO carries a verbatim ~13-sentence reversal-hook monologue over ONE creator

across ~5 wardrobe changes in ~3 micro-worlds, interspersed with product B-roll (capsule macro,

unboxing, a landing-page scroll), ~30 hard cuts on the VO cadence, closing on a brand end card.

This capability is the FREE, deterministic assembly — trim-to-EDL, hard-concat, the VO+music

mix, the karaoke-pop caption burn, the landing-page zoompan, and the end-card append.

scripts/config.example.json is the worked example (Bioma "Do NOT buy Bioma Probiotics", ~37s

1080×1920 9:16, ~30 body cuts + a ~2s end card); scripts/PIPELINE.md maps every config block to

its source step and scripts/README.md documents the free assembly.

Run

This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are

separate capabilities — the spoken VO (create-vo-elevenlabs) Whisper-aligned so the WORD

BOUNDARIES set the cut grid; one locked creator (create-image-gpt-image-fal anchor + ~5 wardrobe

edits chained off the anchor) + 3 world wides + per-cut start-frames (create-image-fal product

composites); and one Veo/Seedance i2v clip per cut (create-video-fal). Given the VO +

vo-final.words.json + edl.json + one clip per cut + a Playwright landing-page PNG + the brand

end-card PNG, render-narrated-ugc-wardrobe-stitch trims each clip to its EDL window, hard-concats

on the VO cadence, mixes the VO over the ducked bed, burns the karaoke-pop captions, appends the

end card → the master. Re-cuts reuse the existing VO / start-frames / clips and cost $0.

Contract (the free assembly)

  • The spoken VO carries the narrative — lock it FIRST. The VO IS the narration bed; the whole

ad is cut to it. Never plan the cut grid before the VO is locked and Whisper-aligned.

  • Build the EDL from the VO's Whisper word boundaries. ~30 role-tagged cuts (hook, feature,

reaction-insert, payoff-hold, b-roll-insert, landing-page); snap every cut window to the

word boundaries. The payoff line gets a HELD payoff-hold beat (~3× mean shot length).

  • Hard cuts via filter_complex concat, not the demuxer. Trim each clip to its EDL window and

hard-concat with filter_complex concat — the -f concat demuxer drops the audio when a

drawtext/scale step shaves a clip a few ms below its window. No dissolves.

  • Karaoke-pop captions on every word, throughout. From the VO's vo-final.words.json (VEED

Whisper preset, bold yellow), on every word; re-spell brand tokens Whisper mishears against the

locked script ("synbiotic" over "symbiotic"; keep "I'ma" verbatim) — never edit the script to

match Whisper. Captions are suppressed over the end card. If VEED mis-captions a brand token,

hand-patch that sentence with local ASS karaoke.

  • Product B-roll breaks up the talking head. Capsule macro, unboxing, and a landing-page

scroll are interspersed with the creator cuts. The landing-page scroll is FFmpeg zoompan over

a Playwright-rendered PNG — not an i2v clip (i2v hallucinates the UI).

  • VO over a ducked bed. Mix the optional instrumental bed sidechain-ducked UNDER the VO

(−20dB, 20:1) so the VO stays clearly on top; the bed can drop in on the payoff beat.

  • End card via the brand's real PNG — never AI-render brand text. Append the brand's real

end-card PNG (~2s) on the tail, captions suppressed. A diffusion model garbles a wordmark.

  • FFmpeg composite, deterministic, FREE. Trim-to-EDL, filter_complex concat, VO+music mix,

caption burn, landing-page zoompan, end-card append, loudnorm I=-14 → a 1080×1920 h264+aac

master (~37s). No paid calls, no keys.

Other skills for the same job

different authors, same section of the catalogue
Canvas Design
by anthropics
vendor ×13

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

1388k tokens
Algorithmic Art
by anthropics
vendor ×10

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

15k tokens scripts
Image Enhancer
by frostant
×6

Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.

635 tokens
Video Downloader
by CommandCodeAI
×4

Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.

671 tokens
Histolab
by christophacham
×3

Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.

18k tokens
Omero Integration
by christophacham
×3

Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.

32k tokens
Pydicom
by christophacham
×3

Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.

13k tokens scripts
Transformers
by christophacham
×3

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

13k tokens

How to use it

Copy the folder

Take gooseworks-ai/render-narrated-ugc-wardrobe-stitch from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.