Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the TOP ~52%, each top clip contain-fit with a darkened blurred cover-scale fill of the same clip (never black bars), a 3px brand-color divider between the zones, the creator slice cover-fit per the per-scene VO timing, scenes hard-concatenated with the body audio being the concatenated creator VO slices, an end card held on the last sharp frame ~3s, then the ASSEMBLED cut transcribed with local Whisper (not the raw VO — concat drops inter-scene silence) and word-level captions burned in the chosen style. This is the FREE deterministic assembly + caption stage (two-zone composite + blurred fill + divider + hard-concat + end card + captions); the VO comes from create-vo-elevenlabs, the anchor from create-image-gpt-image-fal, and the whole-VO lip-sync from a paid VEED Fabric 1.0 take (a no-atom upstream input). Use for the split-screen-creator format.
npx skills add https://github.com/gooseworks-ai/goose-skills --skill render-split-screen-creator
Assemble a split-screen creator ad from a config: a two-zone vertical
(1080×1920, 9:16, ~40s) format where an AI creator talking-head anchors the
BOTTOM ~48% of the frame and real 16:9 product/demo clips run uncropped in
the TOP ~52%, a 3px brand-color divider between the zones. The creator delivers
the whole VO cold-to-camera and each top clip proves the claim its VO line makes.
This capability is the FREE, deterministic assembly + captions — the two-zone
composite (contain-fit + blurred-cover fill + divider + creator slice), the
hard-concat, the end card, and the word-level caption burn from the assembled cut.
scripts/config.example.json is the worked example (Perplexity concept-10
"Bloomberg terminal", ~40s 1080×1920 9:16, 6 scenes + an end card);
scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly + captions stage — it spends
nothing. The paid inputs are separate steps — the VO (create-vo-elevenlabs,
ElevenLabs eleven_v3 with-timestamps, sliced into per-scene windows), the
photoreal MEDIUM chest-up AI-creator anchor (create-image-gpt-image-fal,
gpt-image-2) — shot at a natural webcam distance (headroom + shoulders, real
room), not a plain-background close-up headshot (see the anchor note below), and
the whole-VO lip-sync (a paid VEED Fabric 1.0 @ 720p take — a no-atom step,
image_url = the anchor, audio_url = the vo mp3, ~$0.15/sec, ~$5.90 for a 39s
VO; run its calls sequentially, veed/fabric-1.0 storage-auths 403 under
parallel load). Given the creator lip-sync take + the per-scene VO timing + one
16:9 top clip per scene + the scene-1 hook graphic + the end-card clip,
render-split-screen-creator composites the two zones, hard-concats the scenes,
appends the end card, transcribes the assembled cut, and burns the captions → the
master. Re-cuts reuse the existing VO / lip-sync / clips and cost $0.
product/demo clip contain-fit (uncropped), the BOTTOM zone is the creator
lip-sync framed head-to-shoulders: scale-to-width × a small ZOOM (~1.15–1.2)
then crop the zone with a downward offset so the face sits upper-middle and the
shoulders enter the bottom. (A plain cover + crop-toward-top shows only the head and
cuts the shoulders — and it can't rescue an anchor that was shot too close; fix the
anchor distance first.) Tune zoom/offset visually against the source video's creator
framing — it's a FREE re-assemble, no VEED re-run. A 3px brand-color divider separates
the zones. Canvas 1080×1920, top_height ~998. Keep every stacked height EVEN
(998 + 4 divider + 918 = 1920) — libx264 rejects odd dimensions.
good as the anchor. It must be photoreal/candid (real lived-in room, natural skin
cues), framed chest-up at a natural webcam distance (headroom + shoulders), not a
plain-background close-up and not a phone-selfie pose copied from another format
(e.g. ugc-walk-and-talk). VEED Fabric handles photoreal fine (unlike Seedance).
filled with a darkened blurred cover-scale of the same clip — a flat
charcoal/black bar reads cheap.
(top_start/top_end) to the on-message segment that proves its VO line.
Never loop a short clip — set the window and the assembler speed-fits it to
the scene (looping replays into a sparse/black tail).
the concatenated creator VO slices, timed per the per-scene timing.json; the
lip-sync drives the mouth.
(no dissolves); append the end card holding the last sharp frame ~3s. If the
end-card clip fades to black, hold the last sharp second (endcard.clip_end),
not the black tail.
silence, so the ad timeline ≠ the VO timeline; only the final cut's audio
yields correct caption timing. Transcribe the assembled cut with local Whisper,
build word-level cues (sentence-aware chunking), burn the ASS in the chosen
style (serif-accent, kinetic-pop, …). Keep the -precaption cut + the
.ass sidecar so captions restyle without re-rendering the composite. If the
host ffmpeg lacks libass, render the cues as timed PIL PNG overlays (ffmpeg
overlay=…:enable='between(t,st,en)') at the same placement.
hard-concat, append the end card, mux the creator VO, loudnorm I=-14 → a
1080×1920 h264+aac master. No paid calls, no keys — the VEED Fabric lip-sync is
a supplied input, produced upstream.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take gooseworks-ai/render-split-screen-creator from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.