mcpbeat Sign in

Evals Run Agent Skill

Workflow skill for evals run.

926 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
536
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/HoangNguyen0403/agent-skills-standard --skill evals-run

The instruction itself

13 sections, as written by the author

Evals Run Skill

> [!IMPORTANT]

> Workflow skill for evals run.

Optional args: slug=<feature>, ticket=<id/url>, mode=interactive|autonomous|channel, channel=<id>, auto_continue=true|false, profile=business|hybrid|technical.

Instructions

When the user asks to perform this workflow, execute the following steps:

description: Run blinded live skill evals and publish reproducible v2 results.

Goal

Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.

Steps

1. Choose or resume a run

  • For ordinary maintenance after a complete catalog baseline exists, run pnpm evals:baseline first. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers.
  • Review that plan before spending quota. Start workers only with pnpm evals:baseline -- --execute; the default is gpt-5.6-luna with high reasoning and one worker. Override intentionally with EVALS_MODEL, EVALS_REASONING_EFFORT, or EVALS_CONCURRENCY (maximum four workers).
  • If usage is exhausted, keep the run directory and rerun the identical --execute command after access resumes; completed answers are reused automatically.
  • Use pnpm evals:manifest -- --category <category> for one category or pnpm evals:manifest -- --all for the complete catalog.
  • Use pnpm evals:manifest -- --resume <runId> only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID.
  • Record the printed run ID. The manifest records source hashes, the v2 schema, and the generation protocol.

2. Answer each blinded case

  • Run each baseline and with-skill arm in a separate worker/context.
  • Baseline receives only the prompt. With-skill receives the same prompt plus that skill's SKILL.md.
  • Trigger cases receive only the skill name and one-line description; never open the full skill body or expose the expected label.
  • Trigger prompt filenames use opaque case IDs; never infer the expected label from filenames or ordering.
  • For all runs, write answers under answers/<category>/<skill>/<case>; category runs use answers/<skill>/<case>.
  • Mark known compromised baselines in the manifest and do not use them for delta calculations until clean reruns replace them.

3. Complete and score

  • Fill metadata.agent, metadata.model, and metadata.completedAt after every required answer exists.
  • Run pnpm evals:score -- --run <runId>.
  • Scoring refuses to write results.json while any arm is pending, verifies source hashes, and writes one immutable inputs.json snapshot before publishing v2 results.

4. Report and verify

  • Run pnpm evals:report to project aggregate runs into the newest complete category partitions and update physical history/archive records.
  • Run pnpm evals:verify -- --run <runId> and, before handoff, pnpm evals:verify -- --all.
  • Confirm case pass rate, assertion pass rate, trigger recall, trigger specificity, and balanced trigger accuracy. Treat baseline and delta as n/a for compromised arms.
  • Never hand-edit results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.

Output

Run Summary

Evidence

Known Risks

Outcome Report

feature_status: implemented | partially_implemented | blocked

requirement_trace: manifest -> inputs -> results -> report -> verification

completed_evidence: []

missing_evidence: []

decision_needed: []

recommended_next_workflow: verify-work

Other skills for the same job

different authors, same section of the catalogue
Skill Creator
by anthropics
vendor ×10

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

56k tokens scripts
Skill Creator
by vercel-labs
vendor ×10

Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.

12k tokens scripts
Skill Creator
by JayZeeDesign
×9

Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.

10k tokens scripts
Template Skill
by JayZeeDesign
×7

Replace with description of the skill and when Claude should use it.

35 tokens
Dispatching Parallel Agents
by ZhanlinCui
×5

Use when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

2k tokens
Skill Development
by anthropics
vendor ×4

This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.

9k tokens
Find Skills
by sanity-io
vendor ×4

Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.

1k tokens
Writing Skills
by ZhanlinCui
×4

Use when creating new skills, editing existing skills, or verifying skills work before deployment

26k tokens scripts

How to use it

Copy the folder

Take hoangnguyen0403/evals-run from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.