vercel-labs/ai-cli
Generate text, images, video, and audio from the terminal using AI models.
npx skills add https://github.com/vercel-labs/ai-cli --skill ai-cli
Generate text, images, video, and audio from the terminal using AI models.
Use when you need to:
Requires AI_GATEWAY_API_KEY or a provider-specific key (e.g. OPENAI_API_KEY) in the environment.
ai text "explain this code" # generate text
ai image "a sunset over mountains" # generate an image
ai video "a spinning triangle" # generate a video
ai audio speak "hello" # generate speech
ai audio transcribe recording.mp3 # transcribe audio
ai models --type audio # list speech and transcription models
-m, --model <id> Model ID (provider/name or short name), comma-separated for multi-model
-o, --output <path> Output file or directory
-n, --count <n> Number of generations per model
-q, --quiet Suppress progress output
--json Output structured metadata as JSON (paths, timing, success/failure)
--timeout <seconds> Request timeout in seconds (see Timeouts for per-command defaults)
Chain commands for agent workflows:
# Pipe content in for summarization
cat file.txt | ai text "summarize this"
git diff | ai text "write a commit message"
# Image-to-video pipeline
ai image "a dragon" | ai video "animate this"
# Image editing via stdin
cat photo.png | ai image "make it a watercolor"
# Audio workflows
echo "Ship the changelog" | ai audio speak -o changelog.mp3
cat recording.mp3 | ai audio transcribe -o transcript.txt
Use --json to get machine-readable results:
ai image "a sunset" --json
Returns:
{
"elapsed_ms": 3420,
"count": 1,
"results": [
{
"index": 1,
"model": "openai/gpt-image-2",
"elapsed_ms": 3420,
"success": true,
"file": "/path/to/resp_abc123.png"
}
]
}
ai image "a sunset" -m "openai/gpt-image-1,bfl/flux-2-pro,xai/grok-imagine-image"
-o <dir>: saves inside directory with auto-generated namesWhen the CLI chooses a filename, it uses a response ID when available and falls back to a random 8-character ID, such as resp_abc123.png or 7f3a9c1d.mp3.
Important for agents: Always use -o to save to a file when generating images, video, or speech audio. Without -o in a non-TTY context, raw binary data is written to stdout, which wastes context and is not useful for agents. Use -o output.png, -o speech.mp3, or an output directory and read the file path from --json output instead.
Override with --timeout <seconds> when a prompt legitimately needs longer, instead of dropping to a faster model variant that changes the output:
ai image "a 72-cell sprite atlas, detailed" --timeout 600
The value is in seconds, not milliseconds, and is capped at 2147483.
0 — success1 — all generations failed2 — partial failure (some succeeded)Take vercel-labs/ai-cli from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.