Assemble a two-host fake-podcast skit ad from a config — per-line lipsync clips hard-concatenated in script order, scaled/padded to 1080×1920, WHITE bottom-center captions (up to 5 words per cue, broken on sentence punctuation, word-wrapped to stay in-frame, held at least 0.9s) built from each line's OWN ElevenLabs char-level timestamps (offset by cumulative clip start, never Whisper), and closed on a Playwright/PIL brand end card composited from the real wordmark — never AI-rendered text. This is the FREE deterministic assembly stage (concat + white captions + end card + crf28 encode); the per-line VOs, photoreal gpt-image-2 base stills, expression variants, and lipsync clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the podcast-skit format.
npx skills add https://github.com/gooseworks-ai/goose-skills --skill render-podcast-skit
Assemble a two-host fake-podcast skit ad from a config: a skeptic and a believer at an
absurd themed podcast desk do a snappy back-and-forth about the product (the set is
deliberately unrelated — that is the joke). Each line is its own lipsync clip so the edit can
cut on the dialogue beat (~1.8s avg); this capability is the FREE, deterministic assembly
that concatenates those clips, renders the WHITE captions, and appends the brand end card.
scripts/config.example.json is the worked example (Ladder run-02 "Laundromat 2am", ~49s
1080×1920 9:16, ~22 lines); scripts/PIPELINE.md maps every config block to its source step
and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities: one ElevenLabs with-timestamps VO per line (one voice per host) via
create-vo-elevenlabs; two photoreal base stills at the themed desk plus ~10 expression variants
(mouths NEUTRAL/CLOSED, gpt-image-2 quality=high, not nano-banana) via
create-image-gpt-image-fal; and one lipsync clip per (still, VO) pair via create-video-fal.
Given the per-line clips + their VO timestamps + the brand wordmark SVG, render-podcast-skit
walks the scenes in script order, builds the global caption timeline, renders the WHITE captions,
hard-concats the clips, auto-appends the end card, and final-encodes crf28 → the master. Re-cuts
reuse the existing VOs / stills / clips and cost $0.
needs no music (an optional low ambience is a taste call, off by default).
order (scale/pad to 1080×1920, re-encode) — no dissolves.
global words.json by offsetting each line's char-level word timings by the cumulative clip
start, group into ≤5-word cues broken on sentence-final punctuation, and render **WHITE
#FFFFFF bottom-center captions (black outline), word-wrapped to stay in-frame** and held
≥0.9s — PIL PNG overlays when the host ffmpeg lacks libass (common), else ASS. (Yellow 3-word
karaoke was the old style, rejected in testing.) Whisper on the rendered clips mistimes; the VO
timestamps are ground truth.
lockup is a deterministic HTML → PNG → 2.5s mp4 from the brand's real wordmark SVG (black bg,
brand wordmark, CTA pill, URL), auto-appended after the last line. A diffusion model garbles
a wordmark.
burn ASS via libass), append the end-card mp4, and **final-encode -preset slow -crf 28 + aac
96k** → a 1080×1920 h264+aac master (~6MB for ~28s; the old -crf 20 produced ~16MB). No paid
calls, no keys.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take gooseworks-ai/render-podcast-skit from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.