The user asks a question about a video that was already watched or indexed — "what did they say about X", "what error code appears", "what happens at 2:30", "does the video show Y". Use this to answer from the persistent index with timestamped evidence and a confidence score instead of re-watching or guessing.
npx skills add https://github.com/oxbshw/watch-skill --skill asking-with-evidence
Every watched video sits in a persistent index. Questions about it are
answered from that index — text first, frames only when needed — with
timestamps, a confidence score, and an honest refusal when the video
does not show the answer. Never re-run a watch for a follow-up.
watch-skill ask <video_id-or-original-url> "<question>"
Any language works; the answer comes back in the language of the
question. The engine escalates on its own when unsure (dense re-sampling,
zoom-crop re-OCR, stronger model) and prints a ~N tokens saved line.
Three rules for reading the result:
decoration.
the answer, that is the answer. Do not invent one past it.
yourself — Read them then (or force with --frames).
Moment questions get a dense window, not a whole-video ask:
watch-skill ask <video_id> "what is on screen around 2:30?"
The answer engine pulls frames, transcript and OCR around the moment it
resolves. Agents on MCP have a dedicated get_moment tool that takes an
explicit timestamp and window; the CLI answers the same question through
ask.
watch-skill search "<phrase>"
Hybrid keyword + semantic search across every video ever watched, with
per-script normalization (Arabic folding, CJK segmentation, Thai
segmentation). Follow a hit with ask or moment on that video.
Report it so the next answer is better — see the
learning-from-mistakes skill.
Take oxbshw/asking-with-evidence from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.