> Yuval's all-in-one AI video pipeline. Turns an idea/script into a finished, on-brand MP4 by orchestrating HyperFrames (HTML→deterministic video render), Lottie (branded motion graphics), ManimCE (math / neural-network / concept animations), and a transcribe→approve caption flow — all wrapped in the YUV.AI Neon Phoenix brand via a frame.md. Use whenever Yuval wants to make, about X", "explain X as a video", "neural network animation", "turn this into a video", hyperframes, animation, "make a video", "explain ... as a video", מצגת וידאו, סרטון, הסבר וידאו. Routes each beat to the right engine, wraps in brand, self-verifies, and renders.
npx skills add https://github.com/hoodini/ai-agents-skills --skill yuv-video-director
The conductor for YUV.AI video. You (the agent) decide what each beat needs, route it to the
right engine, compose everything into one HyperFrames composition, wrap it in the **YUV.AI
Neon Phoenix brand, self-verify, and render** to MP4. This skill is the router + the
working reference implementations; load a reference file only when that engine is in play.
> Design source of truth: the yuv-design-system skill (Neon mode — pink #FF1464, cyan
> #00E5FF, rich-black/white, Anton+Inter+JetBrains Mono, neural-net phoenix motif). The video
> form of it is frame.md — see references/frame-md.md. Bundled
> template: assets/FRAME.md. Drop it in the project root; HyperFrames reads it.
HyperFrames renders by seeking each frame in headless Chrome → FFmpeg (frameIndex = floor(t·fps),
same input → same output). So every visual is one of two kinds:
| Pattern | Runs… | Engines | Rule |
|---|---|---|---|
| Live seekable adapter | *inside* the render, driven to time t per frame | GSAP, Lottie (window.__hfLottie), Three.js, a canvas driven by a GSAP proxy onUpdate | must be clock-driven — no Date.now(), Math.random() (seed a mulberry32), setTimeout, or .play() |
| Pre-rendered asset | *offline*, outputs a file, imported as a clip | ManimCE (Python→MP4/alpha), TTS audio, background-removal | render first, then drop in as a <video>/asset clip |
If it can be seeked, it's an adapter. If it can't, pre-render it. Manim is *always* a pre-rendered clip — it has its own renderer and runs in Python; it can never be a live adapter.
"explain X as a TEASER/promo (FOMO, cliffhanger, fast)" → references/teaser-explainer.md (the formula)
"explain a concept / math / neural network / algorithm / training" → ManimCE (pre-rendered clip — cut into BURSTS for teasers)
"branded motion: logo sting · stat reveal · icon pop · pulse" → Lottie (live, lottie-web)
"kinetic captions · titles · reveals · transitions · data callouts" → GSAP (live) ← default
"3D / spatial" → Three.js (live) (Babylon NOT used)
speech → captions → transcribe + approve webapp (see video-edit skill)
no voiceover provided → TTS (Kokoro: npx hyperframes tts)
brand colors / fonts / motifs → frame.md (picked up front)
GSAP is the reliable default for text/motion. Reach for Lottie for *designed* branded graphics, Manim for *explaining* an idea.
FRAME.md is in the project root (copy assets/FRAME.md). All colors/fonts/motifs come from it — never invent. Three brand must-haves on every video (see references/brand-kit.md + references/cinematic.md): the real phoenix logo (assets/logo-phoenix.png) at the reveal + end card, a real featured Lottie (generate one — assets/lottie-burst-generator.py), and the link end-card (logo + "LET'S FLY HIGH" + the full link set + CTA). For teaser/social pacing see references/editing.md; for psychological/cliffhanger/FOMO cuts see references/cinematic.md.py -m manim render -qh --fps 30 scene.py SceneName, copy the MP4 into assets/.video-edit skill (transcribe → approve webapp → sync) and npx hyperframes tts.npx hyperframes init <slug> --non-interactive. Author index.html (the hyperframes skill is the contract). Use:<video class="clip" muted playsinline> body clipnpx hyperframes lint (0 errors) → validate (0 console errors, WCAG AA) → render → spot-check 5 frames across the timeline. Fix → re-run. Lottie MUST be screenshot-verified (the Skottie-vs-lottie-web trap).npx hyperframes render --fps 30 --output renders/<name>_FINAL.mp4. For vertical, clone with a 1080×1920 layout (see video-edit).See references/prereqs.md. Need Node 22+, FFmpeg, Python 3.11+ with pip (Manim/captions). On this machine: real Python is py → C:\Python313 (the bare python is Hermes' venv with NO pip — don't use it). ManimCE installs via py -m pip install manim (no LaTeX needed if you author with Text()/MarkupText, not Tex/MathTex). If Manim isn't available → skip math beats or offer to install; never hard-fail the whole video.
| File | What it is |
|---|---|
| assets/FRAME.md | YUV.AI Neon Phoenix video frame spec (rebranded from HeyGen's Coral pack) |
| assets/neural-net-field.js | Deterministic, seekable neural-net phoenix canvas background |
| assets/neural-pulse.json | Hand-authored Bodymovin Lottie (renders in lottie-web, not just Skottie) |
| assets/gen_content_lotties.py | Generator for 6 production-grade CONTENT lotties (WhatsApp-collapse, decode-beam, eye-read, ghost-line hero, orb-extract, brain-fire) — transparent, persistent+continuous, lottie-web verified |
| assets/{wa-collapse,decode-beam,eye-read,ghost-line,orb-extract,brain-fire}.json | The 6 generated content lotties, ready to drop in |
| assets/manim-scene-template.py | ManimCE scene template, neon brand styling, Text-only (no LaTeX) |
| assets/what_is_nn.py | "What is a neural network" ManimCE scene (neuron → network → training, 34.5s) |
| assets/Anton-Regular.woff2 | Local Anton (renderer doesn't auto-resolve it; declare @font-face) |
| references/composition-pattern.md | The full multi-engine index.html pattern (field + Lottie + Manim clip + GSAP + flash) |
| references/teaser-explainer.md | The cinematic teaser-explainer formula: cold-open slams → manim BURSTS → face-off → FOMO montage → cliffhanger; content-synced transparent lottie beats; the seek-modulo fix; teaser music synth |
hyperframes (composition contract — always invoke when authoring), hyperframes-cli, lottie,
video-edit (transcribe + approve webapp), yuv-design-system (brand). This skill orchestrates them.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.
Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.
Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
Take hoodini/yuv-video-director from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip, npx.
Without those the skill loads but fails at the first command.