mcpbeat Sign in

Asking With Evidence Agent Skill

The user asks a question about a video that was already watched or indexed — "what did they say about X", "what error code appears", "what happens at 2:30", "does the video show Y". Use this to answer from the persistent index with timestamped evidence and a confidence score instead of re-watching or guessing.

552 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
336
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/oxbshw/watch-skill --skill asking-with-evidence

What it tells the agent to use

found in the instruction text
Bash runs shell commands — read the instruction before connecting

The instruction itself

5 sections, as written by the author

Asking with evidence

Every watched video sits in a persistent index. Questions about it are

answered from that index — text first, frames only when needed — with

timestamps, a confidence score, and an honest refusal when the video

does not show the answer. Never re-run a watch for a follow-up.

Answer a question

watch-skill ask <video_id-or-original-url> "<question>"

Any language works; the answer comes back in the language of the

question. The engine escalates on its own when unsure (dense re-sampling,

zoom-crop re-OCR, stronger model) and prints a ~N tokens saved line.

Three rules for reading the result:

  • Cite the timestamps it gives you; they are real evidence, not

decoration.

  • Trust the refusal. When it says the video does not clearly show

the answer, that is the answer. Do not invent one past it.

  • Frame paths are listed only when the engine wants you to look

yourself — Read them then (or force with --frames).

"What happens at 2:30?"

Moment questions get a dense window, not a whole-video ask:

watch-skill ask <video_id> "what is on screen around 2:30?"

The answer engine pulls frames, transcript and OCR around the moment it

resolves. Agents on MCP have a dedicated get_moment tool that takes an

explicit timestamp and window; the CLI answers the same question through

ask.

Don't know which video? Search them all

watch-skill search "<phrase>"

Hybrid keyword + semantic search across every video ever watched, with

per-script normalization (Arabic folding, CJK segmentation, Thai

segmentation). Follow a hit with ask or moment on that video.

When the user corrects you

Report it so the next answer is better — see the

learning-from-mistakes skill.

How to use it

Copy the folder

Take oxbshw/asking-with-evidence from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.