mcpbeat

Social Short Production

calesthio/social-short-production

Plan, produce, adapt, finish, deliver, and improve platform-ready short vertical social videos from a brief, script, transcript, or source media. Use for TikTok videos, Instagram or Facebook Reels, YouTube Shorts, LinkedIn vertical clips, and cross-platform social cutdowns when the work requires audience and objective definition, an honest hook and promise, short-form beat or shot planning, UI-safe layout, captions and accessibility, speech/music/SFX mixing, platform-specific exports, rights and AI or sponsorship disclosure, analytics-led retention iteration, or final QA and repair. Do not use as a provider-specific image, video, voice, or music generation guide.

13k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
112
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/calesthio/generative-media-skills --skill social-short-production

What comes with it

18 663 bytes besides the instruction
EVAL.md

The instruction itself

27 sections, as written by the author

Social Short Production

Produce an intentional short video, not merely a vertical crop. Treat each platform as a distinct playback, interface, policy, and measurement environment.

Evidence labels

Interpret guidance by its label:

  • Documented fact — stated by a platform, regulator, standard, or other primary authority. Recheck volatile facts before delivery.
  • Research finding — supported in a cited study but bounded by that study's population and context.
  • Production heuristic — a starting practice to test against the brief, audience, platform preview, and analytics; never present it as an algorithmic rule.

Do not turn a heuristic such as “hook in three seconds,” “cut every two seconds,” or “master at -14 LUFS” into a universal requirement. Platform metrics and recommendation systems change, and audience intent can outweigh tempo.

Define the production contract

Before scripting or cutting, resolve the fields that can change the result:

  • Objective: awareness, qualified attention, education, community response, lead/action, or conversion. Select one primary outcome and at most two guardrails.
  • Audience state: who is watching, what they already know, what they doubt, and why this clip is relevant in this viewing context.
  • Promise: the specific value delivered by watching. Make it provable inside the clip.
  • Source of truth: approved brief, transcript, footage, product build, data, citations, or spokesperson. Mark unverified claims.
  • Distribution surfaces: exact organic, paid, boosted, or partner placements. “Social” is not a platform specification.
  • Action after viewing: remember, comment, save, share, follow, click, install, purchase, or watch a longer item.
  • Constraints: duration ceiling, languages, brand rules, accessibility, age or regulated category, review deadline, budget, and available analytics.
  • Rights and representation: owners and license scope for footage, music, graphics, fonts, trademarks, likenesses, voices, testimonials, and synthetic media.

If a critical item is unknown, state an assumption and make the plan reversible. Do not invent product claims, testimonials, measurements, or permissions.

Brief normalization output

Return a compact production contract before detailed creative work:

Primary objective:
Primary audience and viewing context:
Viewer tension or need:
Promise and proof available:
Primary platform/placement:
Secondary platforms:
Target duration range:
Desired viewer action:
Source material and gaps:
Rights/disclosure status:
Primary metric + guardrails:
Approval owner:

Build a dated platform contract

Recheck official documentation on the day of export. Record the verification date and placement; organic, ad, boosted, and in-app-created posts can have different rules.

Current reference points — verified 2026-07-09

| Surface | Documented facts useful to production | Production consequence |

|---|---|---|

| YouTube Shorts | Square or vertical uploads up to three minutes are categorized as Shorts under YouTube's current upload rules. Since 2025-03-31, a public Shorts “view” counts starts or replays with no minimum watch time; “engaged views” remains available for viewers who chose to continue watching. Shorts over one minute with an active Content ID claim are blocked globally. [S1][S2] | Do not compare public views with old Shorts cohorts without noting the definition change. Check claimed music before using a 1–3 minute cut. Export without pillarbox or letterbox bars. |

| Instagram Reels | Instagram accepts Reels from 1.91:1 through 9:16, with at least 30 fps and 720-pixel resolution. Boosted Reels have a narrower contract: currently less than 90 seconds and full-screen 9:16. [S3][S4] | Prefer a 9:16 organic master when the intent is immersive vertical viewing, but create a placement-specific boost version rather than assuming the organic upload is eligible. |

| TikTok in-feed ads | For non-Spark in-feed ads, TikTok currently recommends 9:16 at at least 540×960, accepts common video containers, and publishes placement-specific rules. TikTok says its safe zone changes with dimensions, ad-caption length, and add-ons and supplies downloadable overlays. Organic and Spark contracts differ. [S5] | Never hard-code one TikTok safe rectangle for all deliveries. Use the current overlay for the exact placement and caption/add-on configuration. Recheck organic upload behavior in the target account/app. |

| LinkedIn video / immersive feed | Uploaded video can be eligible for LinkedIn's immersive vertical experience, but creators cannot post directly into that feed. LinkedIn warns that its UI can overlap edges. Current upload limits include 3 seconds minimum, up to 15 minutes on desktop or 10 on mobile, and aspect ratios from 1:2.4 to 2.4:1. [S6][S7] | Frame a professional-feed version with edge clearance, useful context, and a non-feed-dependent CTA. Do not promise immersive-feed placement. |

Documented fact: TikTok, Meta, YouTube, and LinkedIn can change specs, labels, metrics, or eligibility independently. A local render passing yesterday's table is not final delivery approval.

Production heuristic: Use a clean 1080×1920, 9:16 working master when all target surfaces support it, then derive and preview separate platform renditions. This is a workflow convenience, not proof that one edit is optimal everywhere.

For every target, capture:

platform + placement | organic/paid/boosted | max eligible duration
aspect/resolution/fps | accepted codec/container | max file size
UI overlay/safe-zone source + date | caption mode | cover crop
music license path | sponsorship/AI disclosure controls
upload metadata | analytics definitions | test-upload result

Make the opening earn attention honestly

A hook is the first coherent reason to continue, not necessarily the first spoken sentence. It can be visual, verbal, sonic, or a combination.

  • Name the audience tension or show the result.
  • State or imply a specific promise.
  • Supply enough context to interpret the promise.
  • Begin proof quickly.
  • Pay off the promise before asking for action.

Useful opening modes include result-first demonstration, a precise contradiction, a consequential before/after, an unresolved but answerable question, or an expert claim immediately paired with evidence.

Production heuristic: Remove greetings, logos, and throat-clearing unless recognition or trust is itself the value. Keep an intro only when it improves comprehension, authority, or narrative tension.

Reject an opening when it:

  • withholds essential context merely to manufacture confusion;
  • promises an outcome the clip does not deliver;
  • uses an alarming visual unrelated to the subject;
  • presents a correlation as causation or a simulation as evidence;
  • makes a health, finance, safety, or performance claim without adequate support;
  • depends on tiny text or audio alone to be intelligible.

Choose a structure that matches the job

Select a reasoning shape instead of forcing every clip into the same formula:

  • Demonstration: result → setup → critical steps → verified result → next action.
  • Explanation: surprising claim → necessary context → mechanism → example/evidence → implication.
  • Problem/repair: recognizable failure → cause → correction → proof → prevention.
  • Story: consequential moment → minimal orientation → escalation/turn → resolution → meaning.
  • Expert excerpt: strong self-contained claim → context needed for fairness → supporting detail → conclusion/attribution.
  • Announcement: what changed → why it matters → evidence/demo → who it affects → date/action.

Research finding: Guo, Kim, and Rubin found shorter educational videos associated with more engagement in a large MOOC dataset, but that population, task, and video length range do not establish an optimal duration or cut cadence for social feeds. Use the finding as support for removing avoidable material, not as a universal duration threshold. [R1]

End when the promise is complete. A 19-second complete explanation is better than padding to 30 seconds; a careful 70-second demonstration is better than compressing away the evidence to fit folklore.

Convert the structure into a beat-and-shot ledger

Plan in beats. A beat is a change in question, claim, proof, emotion, location, or action—not an arbitrary number of seconds.

Use this schema:

| Time range | Narrative job | Viewer question | Spoken content | Visual evidence/action | On-screen text/caption | Audio event | Source/rights | Safe-zone risk | Exit condition |

|---|---|---|---|---|---|---|---|---|---|

Require every beat to do at least one job: advance, prove, orient, contrast, or resolve. Remove beats that only decorate a finished thought.

For source-led edits:

  • transcribe before selecting when practical;
  • preserve enough lead-in and follow-through to avoid changing meaning;
  • distinguish jump-cut repair from deceptive reordering;
  • never splice separate phrases into a claim the speaker did not make;
  • retain source timecodes and filenames in the ledger;
  • identify b-roll as evidence, illustration, reconstruction, or atmosphere;
  • request approval for sensitive context changes.

For capture or animation planning:

  • vary shot scale or composition when it clarifies new information;
  • reserve close detail for proof, not random energy;
  • give demonstrations readable screen time;
  • motivate transitions with action, shape, sound, subject, or idea;
  • leave short handles around source takes for editorial repair;
  • plan the first and last frames deliberately because feeds may loop.

Production heuristic: Introduce a visual reset when attention or comprehension needs one, not on a timer. A cut that adds no information can create fatigue; a held shot with evolving evidence can retain attention.

Protect content from interface occlusion

Safe zones are delivery data, not taste.

Documented fact: TikTok's published safe zone varies with caption length and add-ons, Meta provides Reels safe-zone checking resources, and LinkedIn instructs creators to keep key elements away from all edges. [S5][S6][S8]

For each rendition:

  • Obtain the current official template, preview, or placement checker.
  • Overlay it in the edit/compositor at the actual output resolution.
  • Keep faces, product proof, disclosures, captions, logos, data labels, and CTA meaning inside the unobstructed region.
  • Check the right-side action rail, bottom caption/CTA area, top account/title region, gesture/home indicators, and cover/profile crops.
  • Preview with both a short and maximum expected post caption when the UI changes with caption length.
  • Test at small-phone scale and at least one tall and one less-tall viewport.
  • Save a UI-preview screenshot with the delivery record.

Do not solve occlusion by shrinking everything. Recompose the shot, move noncritical decoration outward, shorten copy, or create a platform variant.

Caption for access and comprehension

Documented fact: WCAG 1.2.2 requires captions for prerecorded audio content in synchronized web media at Level A; captions include meaningful non-speech audio and speaker identification, not dialogue alone. [S9]

Decide between:

  • Closed/platform captions: toggleable and accessible to the platform/player; upload and correct a sidecar or review native transcription.
  • Open/burned captions: guaranteed visual styling and visibility but cannot be turned off.
  • Both: only when platform behavior will not show duplicate captions. Keep a clean master and caption file even if distributing an open-caption version.

Caption production rules:

  • transcribe meaning accurately; do not “improve” claims while captioning;
  • synchronize to speech and preserve phrase boundaries;
  • identify a changed or off-screen speaker when needed;
  • include meaningful sounds such as [alarm] or [glass breaks];
  • use high-contrast text with a stable backing treatment over variable imagery;
  • avoid covering a speaker's mouth, hands, product controls, charts, or evidence;
  • use a readable size at actual phone scale;
  • reduce copy or extend exposure when the viewer cannot read it naturally;
  • proof names, numbers, units, URLs, technical terms, and translated lines manually;
  • test without audio and with captions disabled.

Professional subtitle standards emphasize accuracy, synchronization, readable presentation, and tradeoffs between speech density and reading speed; do not treat a single characters-per-second limit as universal across languages or audiences. [P1]

If essential visual information is not conveyed by narration, provide an accessible description in the spoken track, post text, or platform-supported alternative as the delivery context permits.

Mix speech, music, and effects around meaning

Build the mix in this order:

  • Repair dialogue only as needed: remove distracting noise, clicks, and excessive room tone without creating artifacts.
  • Match dialogue tone and level across cuts; automate individual words before applying blunt compression.
  • Add music for pace, emotion, or continuity, not as mandatory wallpaper.
  • Duck or arrange music around consonants and key claims; check small-phone speakers and mono compatibility.
  • Use SFX to explain an action, establish space, punctuate a verified change, or support a transition. Remove effects that compete with language.
  • Meter the full program and inspect true peaks; then audition the encoded upload.

Documented fact: ITU-R BS.1770-5 defines algorithms for measuring programme loudness and true-peak level; it does not prescribe a universal social-platform delivery loudness. [S10]

Production heuristic: For a conventional stereo social master, begin tests around -16 to -14 LUFS integrated with true peak no higher than -1 dBTP, then adjust to the content and the platform's current behavior. This range is not a platform guarantee. Preserve a higher-quality mix master, compare speech intelligibility after platform transcode, and avoid winning loudness at the cost of distortion or fatigue.

Run four listening passes: normal device, quiet volume, phone speaker, and headphones. Run a mute pass. A music-only hook fails users watching silently; text-only proof fails users who are not watching continuously.

Run these checks before lock:

  • Sponsorship/material connection: place a clear disclosure with the endorsement, in the same language, where viewers will notice it. The FTC says a platform disclosure tool may not by itself be adequate and defines clear-and-conspicuous disclosure contextually. [S11]
  • Synthetic or altered media: use the current platform control and visible context when required. YouTube requires disclosure for realistic, meaningful altered/synthetic content; TikTok requires labels on realistic AI-generated images, audio, or video and prohibits some harmful or nonconsensual likeness uses even if labeled; Meta may apply “AI info” based on detected signals or self-disclosure. [S12][S13][S14]
  • Provenance: preserve source IDs, edit decisions, releases, licenses, model/tool involvement, and export hashes. C2PA Content Credentials can carry cryptographically bound provenance, but provenance is context—not proof that a claim is true. [S15]
  • Music: verify commercial scope, platform, territory, duration, paid/organic use, attribution, and expiration. TikTok recommends its Commercial Music Library for content promoting a brand; Meta says its general licensed library is intended for personal, non-commercial use and points commercial users toward Sound Collection; YouTube limits its copyright-safety assurance to its own Audio Library and warns that longer Shorts with active claims can be blocked. [S16][S17][S1]
  • People: obtain appropriate permission for likeness, voice, testimonial, private location, and sensitive personal data. Do not infer consent from public availability.
  • Claims: keep evidence snapshots for prices, results, rankings, quotes, and before/after comparisons. Show simulations and reconstructions as such.

Do not strip authenticity metadata unnecessarily. If a platform transcode removes it, retain the signed/source asset and manifest in the production archive and still use human-visible disclosure where required.

Export platform renditions, not one mystery file

Production heuristic — cross-platform working master: 1080×1920, square pixels, progressive H.264 High Profile in MP4, original/native frame rate, AAC-LC stereo at 48 kHz, BT.709 for SDR, and fast-start metadata. These choices align with broad support and YouTube's documented recommendations, but the exact placement contract wins. [S18]

Avoid frame-rate conversion unless required. Avoid baked black bars. Do not upscale poor source merely to satisfy a number; replace or intelligently reframe it.

Create a delivery package per platform:

video rendition
clean master
caption sidecar + caption transcript
cover/thumbnail with profile-grid crop preview
post copy, title, hashtags/mentions, CTA and destination
alt/accessibility text where supported
sponsorship and AI-disclosure settings
music/asset license ledger and attribution
spec verification date and official URL
upload/test status, platform-generated captions checked, UI screenshot

Before public release, upload privately, unlisted, as a draft, or to a test account where the surface permits. Inspect the platform transcode, crop, color, caption timing, UI collisions, cover, sound, and disclosure—not only the local file.

Technical preflight

Verify with a media inspector rather than trusting the export dialog:

  • duration, dimensions, sample aspect ratio, frame rate/time base;
  • codec/profile/pixel format, scan type, color tags, bitrate and file size;
  • audio codec, sample rate, channels, integrated loudness and true peak;
  • audio/video sync and caption end time;
  • first/last frames, unexpected black/frozen frames, offline media, flash frames;
  • spelling, supers, legal lines, URLs, QR codes, and safe-zone overlay;
  • no hidden draft audio, guide tracks, watermarks, or personal metadata.

Create variants that answer a question

Separate platform adaptation from experimentation.

  • Platform adaptation may change duration, pacing, caption position, CTA, cover, metadata, music license, disclosure, and even the amount of context.
  • Experimentation should change one meaningful variable when possible: first frame, hook wording, evidence order, duration, CTA, or caption treatment.

Name variants by hypothesis, not final-v7:

ytshort-hook-resultfirst-r1.mp4
reels-hook-contradiction-r1.mp4
tiktok-proof-demoearly-r2.mp4

Keep the body identical when testing a hook. Otherwise the result cannot isolate the hook. When the platform does not provide randomized testing, label the comparison observational and note time, audience, distribution, paid support, and sample-size differences.

Measure retention instead of repeating folklore

Define a funnel appropriate to the platform:

eligible exposure → start → chose/continued to view → time held
→ completion/rewatch → meaningful engagement → intended action

Documented fact: platform definitions differ. YouTube exposes “how many chose to view” versus swiped away and audience-retention reports; Instagram reports views, watch time, average watch time, reach, interactions, and follows; TikTok ad reporting defines starts, 2-second views, 6-second views, quartiles, completion, and average play time. Replays are handled differently by metric. [S2][S19][S20]

Therefore:

  • compare clips within the same platform, placement, metric definition, length band, and similar audience context;
  • prefer seconds watched and time-indexed retention alongside average percentage viewed;
  • report uncertainty and sample size; do not declare a winner from a handful of views;
  • separate paid from organic distribution and new from returning audiences;
  • use the primary business or communication outcome as the decision metric, not raw views alone;
  • never compare post-definition-change YouTube public Shorts views directly with old cohorts without qualification.

Retention-curve diagnosis

Treat these as hypotheses to investigate:

| Signal | Possible causes | Repair test |

|---|---|---|

| Immediate cliff | first frame unreadable, promise irrelevant, context missing, wrong audience | test result-first frame, clearer audience cue, or shorter first line |

| Drop during setup | context debt, duplicated narration/text, slow proof | move evidence earlier; remove explanation already visible |

| Dip at a term or chart | jargon, small labels, cognitive overload | define one term, simplify chart, extend evidence beat |

| Spike/replay | interest, surprise, confusion, illegible detail, loop | inspect comments and frame readability before calling it success |

| Strong completion, weak action | satisfying ending but weak/irrelevant CTA, wrong audience, destination friction | test a more congruent action and verify destination |

| Strong early hold, weak completion | hook overpromises, middle repeats, payoff delayed | align promise and proof; cut redundant middle |

Set a minimum observation window before reading results. Compare against the account's relevant baseline, not a generic “good retention” percentage.

Final editorial and platform QA

Do not ship until each item passes or has an explicit waiver:

Meaning

  • Promise is clear and paid off.
  • Claims match approved evidence; dates and numbers are current.
  • Source edits preserve context and speaker intent.
  • CTA follows naturally from the delivered value.

Picture and interface

  • Composition works at phone size; text is readable without pausing.
  • Current placement overlay shows no critical occlusion.
  • Faces, products, charts, disclosures, and captions remain visible.
  • Reframes do not crop gestures, demonstrations, or key context.
  • First frame, cover, and loop boundary are intentional.

Audio and access

  • Speech remains intelligible at quiet playback on a phone.
  • Music/SFX do not mask claims; peaks do not clip after encode.
  • Captions are accurate, synchronized, complete, and non-duplicated.
  • Silent viewing still communicates the main promise and proof.
  • Essential visuals have an accessible equivalent appropriate to the surface.

Integrity and delivery

  • Likeness, music, footage, font, and trademark rights are documented.
  • Sponsorship and synthetic-media disclosures are visible and set in-platform.
  • Technical inspection matches the dated platform contract.
  • Test upload has been reviewed after platform processing.
  • Clean master, rendition, captions, metadata, licenses, and analytics hypothesis are archived.

Repair order

Repair in this order: factual/deceptive issue → consent/rights/disclosure → inaccessible meaning → broken technical delivery → unintelligible audio → UI occlusion → weak structure/retention → cosmetic polish. Never polish around a critical integrity failure.

Complete example — 27-second product demonstration

Example, not a mandatory formula.

Intent: Produce an organic Instagram Reel and YouTube Short for a calendar app's verified “find a mutual time” feature. Audience: project leads coordinating three or more people. Objective: qualified product interest. Primary metric: saves; guardrails: average seconds watched and profile/site actions.

Inputs and constraints: approved screen recording, product claim verified in current build, no customer data, 27–32 seconds, 9:16, spoken English, open captions plus clean/caption-file master, licensed original music bed.

Promise: “Find the first time everyone can attend without opening four calendars.”

Beat ledger:

| Time | Job | Voice/text | Visual | Audio/access | Risk/control |

|---|---|---|---|---|---|

| 0:00–0:03 | Show result and audience problem | “Four calendars. One meeting. Here is the first time that works.” | Messy availability collapses into one highlighted slot | UI click + caption | Crop contains only fake names; result remains above platform caption area |

| 0:03–0:08 | Orient | “Add the people, then set the window.” | Cursor selects three demo participants and date range | Music low under speech | Cursor and fields enlarged for phone; no speed ramp through necessary action |

| 0:08–0:16 | Prove mechanism | “The overlap view removes every conflict.” | Conflicts fade; overlap remains | Subtle confirmation SFX | Label animation as product UI, not simulated result |

| 0:16–0:22 | Show verified result | “Tuesday at 2:30 is the first shared opening.” | Slot selected; confirmation panel | Captions spell time consistently | Keep time and button inside current safe overlay |

| 0:22–0:27 | Resolve and act | “Save this for the next scheduling spiral.” | Before/after split, app name, save cue | Music resolves; no extra claim | CTA is save, matching objective; no unrelated “link in bio” |

Platform adaptations: YouTube version uses a searchable title and clean last frame; Instagram version uses a profile-grid-safe cover and verifies Reel UI/caption position. The core proof remains identical. The agent does not claim equal views are comparable because YouTube and Instagram count them differently.

Variant test: Change only the opening from result-first to “Still comparing calendars?” Keep beats 0:03 onward identical. Compare within each platform to the account's similar product-demo baseline.

Expected result: A viewer understands the problem, watches the real workflow, sees the exact result, and receives one congruent action.

Likely failures and repairs: If retention dips during participant selection, shorten pointer travel but retain enough input context to prove the result. If the product UI is illegible, create a guided crop with magnified callout rather than accelerating it. If native and burned captions duplicate, upload the clean version with corrected platform captions or disable one path.

Complete example — 43-second expert excerpt from source footage

Example, not a mandatory formula.

Intent: Cut a vertical LinkedIn clip and Reel from a 35-minute interview. The expert explains why a sample-size increase reduced a survey's apparent conversion rate. Objective: correct a common analytics misconception. Primary metric: saves; guardrails: completion and substantive comments.

Selection: Use one continuous answer from source interview-a.cam1, 18:42–19:27, with two pauses tightened. Do not join words across separate answers. Retain the sentence “the earlier sample was only 42 users,” because removing it would make the conclusion misleading.

Structure and shot plan:

  • 0:00–0:05: cold-open conclusion, “The conversion rate fell because the sample got less lucky, not because the product got worse.” Show source speaker, name/role, and concise captions.
  • 0:05–0:13: necessary context, including 42-user sample. Use a simple on-screen 42 → 840 respondents comparison labeled from the approved transcript/data.
  • 0:13–0:31: mechanism. Alternate speaker with a restrained diagram only when it illustrates sampling variability; do not use unrelated attention footage.
  • 0:31–0:39: implication, “Compare intervals and cohorts, not one headline percentage.”
  • 0:39–0:43: attribution and “Save this before the next dashboard review.”

Audio/caption plan: Repair room tone around pause trims; keep natural cadence. Music is omitted because it adds no communicative value and the source room tone is clean. Upload corrected closed captions and an open-caption variant only where needed. Identify the speaker once; caption [laughs] only if it changes the tone.

Rights/integrity: Confirm interview release, chart-data approval, font license, and corporate endorsement policy. Preserve the transcript, source timecodes, and edit decision list. No synthetic-media disclosure is required for ordinary reframing, color correction, captions, or diagram assistance under the cited YouTube examples, but each target platform's current rule must still be checked. [S12]

Platform adaptations: LinkedIn includes more professional context in post copy, keeps all edges clear, and does not promise immersive-feed placement. Reels uses a 9:16 cover crop and checks current UI safe zones. Both retain the sample-size context.

Analytics repair: A replay spike on the diagram could mean interest or illegibility. Inspect the frame and comments; do not automatically add more cuts. If saves are high but completion is modest, the explanatory core may still meet the objective; test a shorter setup without deleting the 42-user fact.

Sources

Platform and policy facts below are volatile and were verified 2026-07-09 unless a source carries its own later update date.

How to use it

Copy the folder

Take calesthio/social-short-production from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.