calesthio/anime-animation-production
>- Produce anime-style and stylized 2D animation with generative image and video tools as a craft discipline, not a filter. Use when a request asks for anime, manga-style, cel-shaded, sakuga, shonen/shojo/seinen/slice-of-life, mecha, or retro-90s-OVA animation; when planning shots, motion, and cutting rhythm to anime convention; when holding a character on-model and preventing style drift across shots or episodes; when choosing prompting vocabulary that actually controls anime look; when pairing voice, music, and impact SFX to stylized action; and when navigating the cultural, IP, and platform-policy sensitivities of AI anime (studio/artist style imitation, fan-art and IP boundaries, disclosure and monetization). This is provider-neutral; models are named only as illustrative options.
npx skills add https://github.com/calesthio/generative-media-skills --skill anime-animation-production
Anime is not a preset. It is a set of production conventions born from hand-drawn
cel animation under tight schedules: flat color, hard-edged shading, deliberately
uneven frame timing, and a shot grammar that spends motion where it matters and
freezes everywhere else. Generative tools can reproduce the *surface* of that look
easily, and the *grammar* almost never by default. The value an agent adds is the
grammar — knowing which convention a given genre and beat calls for, encoding it into
prompts and shot plans, and reviewing output against the things that actually break
the illusion.
This skill covers the craft as production knowledge. It is provider-independent. Any
model named (Sora, Kling, Vidu, Wan, PixVerse, Midjourney, etc.) is an illustrative
option, dated where its behavior is volatile — never the method itself.
To keep evidence honest, claims here are tagged:
of record, with a source and (for volatile items) a verification date.
independent of any tool.
model-dependent; verify against your actual model and re-test when models change.
Volatile facts in this document were verified on 2026-07-10. Model capabilities
and platform policies change fast; re-verify before relying on any dated claim.
Use this skill when the deliverable should *read as anime or stylized 2D animation* —
cel-shaded characters, limited-animation feel, anime shot grammar — regardless of the
underlying model.
Do not reach for this skill for photoreal or 3D-CG-realistic pieces, for Western
cartoon styles (Disney full animation, Cartoon Network flat-vector, cutout/Flash), or
for generic "illustrated" video that has no anime intent. Those are different craft
traditions with different timing and shading logic; applying anime conventions to them
produces an uncanny hybrid. If the user says "animated" without specifying anime,
confirm the tradition before committing to this grammar.
The defining surface of anime comes from the celluloid process: an ink outline drawn
on a clear cel, backed by flat painted color, over a separately painted background.
[Fact — the "cel" in cel-shading names the celluloid sheet used in traditional
production; painters laid flat color behind ink outlines. Sources: Wave Motion Cannon;
cel-shading references, 2026-07-10.] Its production-relevant properties:
darker than a pure black (often a brown-black or color-holds on hair/skin). Line
weight varies by era: heavier, chunkier lines in Akira-era/OVA work; thinner, more
uniform lines in modern digital TV anime. [Convention]
character. Shading is one or two discrete tones with a *crisp* edge, not a blur.
A second shadow tone ("2-shadow") appears in higher-budget or dramatic work. This
hard-edge cel shadow is the single strongest tell of "anime" versus generic
digital painting. [Convention]
sheen). [Convention]
is deliberate: characters are flat cel, backgrounds are lush painted (watercolor,
gouache, or digital matte in the Kanmanagement/Kobayashi tradition). This flat-on-painted
contrast reads as "anime" even in a still. [Convention]
Production implication: if your output shows soft, blurred, airbrushed shading on
faces, or lineless "soft render," it has drifted toward generic digital illustration
and away from cel. That is a QC failure (see Part 8), not a style choice, unless the
brief explicitly wants a painterly/"Makoto Shinkai" soft look.
Anime evolved *limited* animation — fewer unique drawings per second than full Western
animation — as an economic necessity that became an expressive language. This is the
most important and most-ignored convention in AI anime.
Frame timing vocabulary [Fact — Anime News Network lexicon; Wave Motion Cannon,
2026-07-10]:
| Term | Unique drawings/sec (at 24fps base) | Feel |
|------|-------------------------------------|------|
| On ones | 24 | Fully fluid, expensive; sakuga, fast action |
| On twos | 12 | Standard "full" animation cadence |
| On threes | 8 | TV-anime default; economical, slightly stepped |
| On fours | 6 | Very sparse; slow/held or background motion |
as a production standard adopted after the 1960s; Osamu Tezuka's *Astro Boy* (1963)
pioneered limited techniques under budget pressure.
snaps on ones, a lumbering giant crawls on threes or fours, an impact holds a single
drawing for several frames. Different rates carry different weight and speed; the
brain fills the gaps. [Convention — popularized through the Otsuka/Kanada lineages.]
large slow object runs on twos, steering the eye. [Convention]
Production implication for generative video: most AI video models interpolate to
smooth, high-frame-rate, "on-ones-everywhere" motion — which reads as *too fluid*,
uncanny, "AI-smooth," and un-anime. Countermeasures [Heuristic, 2026-07-10]:
"held frames," "8fps anime cadence," "TV anime motion." Effectiveness varies by model;
many still smooth it out.
native rate, then step/decimate frames (hold every drawing 2–3 output frames) or
render fewer keyframes and hold. This gives the stepped cadence that prompts alone
usually cannot. [Heuristic]
else more limited so the fluid beat lands.
detailed, virtuoso animation — the money shot. [Fact — term/definition per animation
references, 2026-07-10.] Producing it means spending your fluidity budget on a short
beat while keeping surrounding shots limited, so it contrasts.
high-contrast, inverted, monochrome, or abstract "flash" drawing — inserted for
1–3 frames to punch the collision. [Convention] In generation, produce as an inserted
still or a hard cut, not as smooth motion.
exist for a frame or two to sell speed (a hand becomes a blur of multiple fingers).
[Convention]
velocity or focus (radial "focus lines" for emphasis/reaction). Structural to the
medium, not decoration. [Convention]
"Kanada style" (angular, geometric flame/lightning shapes; "wakame"/seaweed-shaped
shadows). [Convention — Kanada lineage.]
and long holds on a single drawing with only a camera move or a mouth flap. [Convention]
The mouth-flap-over-a-hold and the slow pan/zoom over a still are the workhorses of TV
anime economy and are *easy* to reproduce with a still + a camera move.
From manga heritage: screen tone (halftone dot patterns and hatching for gray values
and texture) and halftone impact bursts. In color anime these appear as retro/manga-panel
stylings, flashback textures, and comedic cutaways. [Convention] Useful as a deliberate
stylistic tag ("halftone screentone shading, manga panel"), not a default.
Anime "style" is really several styles keyed to demographic and genre. Getting this
right is the difference between generic-anime and on-target. Characteristics below are
[Convention] for the traditions and [Heuristic] for the prompt tags.
linework; saturated high-contrast palette; kinetic FX (speed lines, energy, impact
frames); exaggerated musculature and expressions. Tags: *cel-shaded, dynamic action
pose, bold outlines, saturated colors, speed lines, impact frame.*
catchlights; delicate, flowing linework; soft pastel palette; emotive environmental
effects (floating petals, sparkles, screen-tone bloom, backgrounds that mirror mood).
Tags: *soft pastel palette, delicate linework, large sparkling eyes, floral bokeh.*
detailed linework and denser shading; muted/desaturated palette with selective accent
color; psychologically complex, restrained expressions. Bridges anime and Western
illustration. Tags: *realistic proportions, muted desaturated palette, detailed
shading, subdued expression.*
realism than seinen; naturalistic palette. Tags: *naturalistic palette, refined
realistic features.*
realistic everyday backgrounds (Japanese suburbia, classrooms); muted naturalistic
light; warmth over spectacle. Tags: *soft natural lighting, detailed realistic
background, gentle muted palette, everyday setting.*
highlights, often combined with dramatic FX. Tags: *mechanical detail, hard-surface
shading, panel lines.*
comedy inserts. Tags: *chibi, super deformed, simplified.*
A distinct and popular target. The look comes from a real production chain: hand-painted
cels in a small fixed palette, shot on 16/35mm film, broadcast through CRT. [Fact — 90s
production-stack references, 2026-07-10.] Reproduce it with an explicit *stack* of tags
rather than a generic "retro" word [Heuristic, 2026-07-10]:
(Akira-era), hand-painted background.*
80s), *slight gate weave, warm white point, faded/limited palette (muted red, smoky
teal, faded mustard).*
the "meme VHS filter" look rather than authentic OVA.
Reference touchstones for the era's feel (as descriptors, not to imitate a specific
protected work): *Akira*, *Cowboy Bebop*, *Evangelion*, *Sailor Moon*-era palettes.
[Heuristic, model-dependent, 2026-07-10.] Vague "anime style" produces a generic
soft-modern-digital look. Control comes from stacking four categories:
(or "cel shading," "anime cel") over "anime style." Add "flat colors, hard-edge
shadows, clean lineart."
palette," "warm sunset rim light," "high-key soft daylight").
cel shadow," "chunky specular hair highlights."
Additional levers:
gradients, soft airbrush, lineless" to hold the cel look.
"shallow depth of field bokeh background") and composes with style tags.
— pick one coherent era.
anime-tuned checkpoints/LoRAs that respond better to *booru-style tag lists* (comma-
separated attributes) than to prose; others (large general video/image models) respond
better to natural-language scene descriptions. Match prompt shape to the model.
[Heuristic, 2026-07-10.]
This is the hardest technical problem and where most AI anime fails. Two distinct
consistency targets:
scheme across every shot. [Fact — character/"identity" drift is documented across
essentially all current major video models; subtle changes to face, hair, and
proportions appear between shots and worsen over longer clips and scene changes.
Sources: multiple 2025–2026 model comparisons, 2026-07-10.]
and rendering across shots and episodes, so nothing drifts between "flat cel" and
"soft painterly" or between two different color temperaturas.
Note [Heuristic]: stylized anime characters tend to drift *less* than photoreal humans,
because exaggerated proportions and high-contrast features are easier for models to
reproduce — a genuine advantage of the medium. It is a smaller problem, not a solved one.
front / three-quarter / profile / back, a neutral and one expressive face, and a
full-body turnaround, all in the *final* style. This is the anime-production
"settei"/model-sheet concept adapted to generative work. [Convention + Heuristic]
few *distinct* angles/expressions rather than many similar ones. [Heuristic — 2–4
well-chosen references typically outperform 5–6 redundant ones; overlap gives the
model no new information; 2026-07-10.]
image reference / character reference (some models accept multiple simultaneous
references), first-frame image, or a trained character LoRA/embedding for repeated
use. [Fact — as of 2026-07-10, reference-image and first/last-frame conditioning are
available in several video models, e.g. Kling's reference and start/end-frame modes,
Sora's image reference / storyboard keyframes, Vidu's multi-reference mode; treat
exact feature names/limits as volatile and verify.]
toward it so the whole sequence doesn't drift between photoreal and illustration or
between lighting regimes. [Heuristic]
shot N+1 where continuity is tight; for episodic consistency, keep the same model
sheet, seed strategy, and LoRA across the whole batch.
regenerate off-model shots rather than "fixing in the edit."
edges, mute colors, or re-render fabric/hair — changing the character's signature.
Re-assert the core style/character tags in every prompt; don't assume they carry over.
eyes, and signature costume in every prompt, and QC color-pick key frames.
edit over one long generation when identity matters.
Anime editing rhythm is distinctive and generative shots should be planned to it, not
to a generic "smooth continuous" video model default.
still) drawing with a slow pan, zoom, or rack, plus a mouth flap or hair sway. This is
economical *and* the model reproduces it reliably: generate a strong still, then apply
a slow Ken-Burns/parallax move. Use it for dialogue and establishing beats. [Convention
+ Heuristic]
sweat drop, focus lines) — often 0.5–1s. Build these as short beats; they carry emotion
and cover the fact that most shots are held. [Convention]
in-between/smear (on ones) → impact frame + hold on the connect. Plan action as this
spend-fluidity-then-hold cadence rather than uniform continuous motion. [Convention]
background to set place/mood between scenes (the "pillow shot" / kire). [Convention]
sakuga beat, to create contrast (Part 1.2).
then hold on impact" reads more anime than "character punches smoothly."
expecting the base render to include them convincingly.
angles for tension, and long slow pushes for emotion.
Anime sound is stylized and expressive, and pairing it correctly does as much for the
"anime" read as the visuals. [Convention, with sourced anime-sound references, 2026-07-10.]
delivered; dialogue sits forward in the mix and stays intelligible over music and FX.
For AI/synth VO, direct expressive, character-distinct performances; keep dialogue
the priority channel.
synth kick/tom, an energy blast built from synth + whoosh. Footsteps, cloth, and small
motions are often over-emphasized. Layer designed hits rather than using dry realistic
thumps.
or percussive stinger on the connect sells action.
for emotional or tense beats (breath, room tone only). Don't wall-to-wall the music;
plan silence as a tool.
and drops out for ma. Genre cues matter (city-pop/jazz for a Bebop feel, orchestral for
epic shonen, gentle piano/acoustic for slice-of-life).
Mixing order of priority: dialogue intelligible first, then impact hits, then music bed,
with silence used intentionally.
This is a live, contested area. Handle it explicitly; do not treat AI anime as
consequence-free.
images "in the aesthetic of" a studio is not automatically infringement. But this is a
legal *characterization of style*, not permission to copy characters, and it is disputed
across jurisdictions.
produced viral Studio-Ghibli-style images. OpenAI stated it **refuses the style of
individual living artists but permits broader studio styles** (a policy critics note is
in tension, since Ghibli's style is inseparable from living artist Hayao Miyazaki, who
has publicly expressed strong opposition to AI animation). Sources: TechCrunch,
Fast Company, 2025.
work. In late October 2025, CODA** (Content Overseas Distribution Association),
representing ~18+ major Japanese companies including **Studio Ghibli, Bandai Namco,
Square Enix, Kodansha, Kadokawa, and Shogakukan**, publicly demanded OpenAI stop using
their content to train Sora 2, arguing Japanese copyright law implies an opt-in
standard and that reproduction during training can itself infringe. The Japanese
government/Cabinet Office also formally raised concerns. Sources: Variety, TechCrunch,
Game Developer, 2025.
Practical guidance for an agent:
("cel-shaded shonen action," "90s-OVA look") over "in the style of [named living
studio/artist]." Naming a living studio/artist as a style target is legally gray,
ethically contested, and may be blocked or filtered by the tool.
characters, logos, exact designs) for anything beyond clearly-permitted, private,
non-commercial fan use — and even then flag the IP boundary to the user. Do not present
fan-art of a franchise as ownable or commercially safe output.
franchise's characters, surface the concern (living-artist policies, active
rightsholder objections, commercial risk) and offer an original-design alternative
that hits the same genre feel. Follow the user's decision for lawful private use, but
do not help pass off imitation as original or infringe for commercial distribution.
genuine original direction, not with jailbreak phrasing.
be "significantly original and authentic"; mass-produced or repetitious near-
duplicate AI content is demonetized under the "inauthentic content" rules. A separate
AI-disclosure toggle is required for *realistic synthetic media that could mislead*
viewers into thinking a real identifiable person said/did something — a clearly
stylized anime narrator/character generally does not trigger the realistic-person
disclosure, but originality/authenticity rules still apply. Sources: YouTube policy
coverage, 2025.
and avoid template-farmed batches, both to satisfy monetization/originality rules and
because it is what makes the work worth making.
requirements; check the destination platform's current rules before publishing.
[Volatile — verify at publish time.]
and never for deceptive or defamatory use. Cameo/likeness features in some tools require
the person's opt-in; respect it.
Review every shot against these before accepting. Generic video QC misses the failures
that specifically break *anime*.
On-model / character checks (compare to the model sheet):
accessories (hands and small props are common failure points).
Style / line / shading checks:
(jittering line edges frame to frame), not breaking up or doubling. Line-boil is the
airbrushed gradients or lineless "soft render" unless intended.
the sequence and across episodes.
Frame-rate / motion feel:
If so, it fails the limited-animation feel — step/decimate or re-plan (Part 1.2).
drifting or popping, and interpolation ghosting are rejects.
Continuity:
continuous drifting AI clip.
Rule of thumb: if a shot is off-model or line-boiling, regenerate, don't paint over
it — the error usually recurs and compounds across the sequence.
Match encode and framing to destination.
container even if the *animation cadence* is limited (container fps ≠ animation rate),
captions safe-area aware, hook in first ~2s. Keep the anime cadence but ensure the
container frame rate is standard so playback isn't judder-flagged.
or 23.976, H.264/H.265, stereo or 5.1 for premium. Add the required AI/originality
handling from Part 7.
deliver at the platform's spec (often 16:9 4K/1080p, 24–30fps). Lyric/typography overlays
should respect the anime palette.
still encode into a standard container frame rate (24/25/30). Decimate/hold at the
animation layer, then encode normally, so players don't misread the file as low-fps.
These are examples, not mandatory formulas. Adapt to the model, brief, and rights
posture. Prompt wording is [Heuristic] and model-dependent; the *structure and decisions*
are the transferable part.
spike, for a 9:16 Short. Original character (no franchise IP).
final cel-shaded style); fixed palette (hair #2b2b3a, eyes #d64545, jacket #e8e2d0).
(limited, on threes feel), speed lines starting.
then hold.
cut to a 0.7s reaction (eyes narrow, dust settles), slow zoom out.
bold clean lineart, saturated high-contrast palette, [character from reference], dynamic
low-angle action pose, speed lines, hard-edge cel shadows, dramatic rim light; [shot-
specific action]; NOT photorealistic, NOT 3d, NOT soft airbrush."
fluid so it pops; add impact-frame insert; layer SFX (synth kick + whoosh + cloth) with
the loudest hit synced to the impact frame; short silence (ma) right before the strike.
drifts between shots (fix: re-anchor to sheet, shorten shots); impact reads as mush
(fix: insert real impact frame + hold); palette drift (fix: restate hex colors).
drop the flash for a grounded thud.
controllable here than full text→video, and reproduces the anime "hold + camera move"
convention directly.
painted contrast), then a slow parallax push and mouth flaps; cut to a 1s reaction insert.
hard-edge cel shadow, chunky hair highlights, heavy line weight, warm golden-hour rim
light, muted limited palette (faded mustard, smoky teal), film grain, slight gate weave;
two students on a rooftop; NOT clean modern digital, NOT 3d."
gentle piano BGM under dialogue, dialogue forward in the mix, a beat of ma (room tone
only) on a pause.
"hand-painted, painterly background"); characters too modern-clean (add "heavy line
weight, cel paint").
to a clean daylight palette.
and across a later episode.
a character LoRA/embedding on the sheet for reuse across episodes, else keep the sheet as
a persistent multi-image reference; (3) generate shots, each anchored to the sheet AND a
single style-anchor hero frame; (4) chain continuity by feeding shot N's last good frame as
shot N+1's first-frame reference where cuts are tight; (5) QC every shot against the sheet
(Part 8) and regenerate off-model shots; (6) reuse the exact sheet/LoRA/palette in the next
episode so the style doesn't drift between episodes.
adding scene context silently alters the character (re-assert core tags every prompt);
cross-episode palette drift (lock exact colors and reuse the same reference assets).
Craft and animation references:
https://wavemotioncannon.com/2016/12/31/an-introduction-to-framerate-modulation/
https://www.animenewsnetwork.com/encyclopedia/lexicon.php?id=61
https://animationobsessive.substack.com/p/a-subtle-art-with-visceral-power
https://animetudes.com/2021/03/06/the-kanada-style-in-context/
https://www.344audio.com/post/article-secrets-of-anime-sound-design ;
https://www.344sfx.com/blog-posts/inside-anime-sound-design-techniques-tricks-and-subtle-brilliance
https://www.crunchyroll.com/news/deep-dives/2023/3/23/feature-how-anime-voice-over-and-music-get-made
IP, policy, and platform (dated, volatile):
https://techcrunch.com/2025/03/26/openais-viral-studio-ghibli-moment-highlights-ai-copyright-concerns/
https://www.fastcompany.com/91308222/
https://variety.com/2025/digital/news/studio-ghibli-openai-sora2-japanese-trade-group-coda-letter-1236568751/
https://techcrunch.com/2025/11/03/
https://www.gamedeveloper.com/business/japanese-game-studios-demand-openai-stop-pilfering-their-work
https://onewrk.com/youtubes-ai-disclosure-requirements-the-complete-2025-guide/
Generative-tool capability and consistency (dated, volatile — verify against the model in use):
https://magichour.ai/blog/how-to-keep-characters-consistent-in-ai-video ; longstories.ai
https://longstories.ai/blog/maintaining-style-consistency-ai-animation
Sora 2 storyboard/keyframe & image-reference (OpenAI cookbook/API), Vidu multi-reference —
https://cookbook.openai.com/examples/sora/sora2_prompting_guide ;
https://www.runcomfy.com/models/kling/kling-video-o1/image-to-video/reference
https://nanoimagine.art/blog/best-80s/90s-retro-anime-prompts ;
https://autoweeb.com/blog/how-to-choose-the-best-anime-art-style-for-ai-anime-generations
Take calesthio/anime-animation-production from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.