Workflow skill for evals run.
npx skills add https://github.com/HoangNguyen0403/agent-skills-standard --skill evals-run
> [!IMPORTANT]
> Workflow skill for evals run.
Optional args: slug=<feature>, ticket=<id/url>, mode=interactive|autonomous|channel, channel=<id>, auto_continue=true|false, profile=business|hybrid|technical.
When the user asks to perform this workflow, execute the following steps:
description: Run blinded live skill evals and publish reproducible v2 results.
Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.
pnpm evals:baseline first. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers.pnpm evals:baseline -- --execute; the default is gpt-5.6-luna with high reasoning and one worker. Override intentionally with EVALS_MODEL, EVALS_REASONING_EFFORT, or EVALS_CONCURRENCY (maximum four workers).--execute command after access resumes; completed answers are reused automatically.pnpm evals:manifest -- --category <category> for one category or pnpm evals:manifest -- --all for the complete catalog.pnpm evals:manifest -- --resume <runId> only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID.SKILL.md.all runs, write answers under answers/<category>/<skill>/<case>; category runs use answers/<skill>/<case>.metadata.agent, metadata.model, and metadata.completedAt after every required answer exists.pnpm evals:score -- --run <runId>.results.json while any arm is pending, verifies source hashes, and writes one immutable inputs.json snapshot before publishing v2 results.pnpm evals:report to project aggregate runs into the newest complete category partitions and update physical history/archive records.pnpm evals:verify -- --run <runId> and, before handoff, pnpm evals:verify -- --all.n/a for compromised arms.results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.feature_status: implemented | partially_implemented | blocked
requirement_trace: manifest -> inputs -> results -> report -> verification
completed_evidence: []
missing_evidence: []
decision_needed: []
recommended_next_workflow: verify-work
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Replace with description of the skill and when Claude should use it.
Use when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies
This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.
Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.
Use when creating new skills, editing existing skills, or verifying skills work before deployment
Take hoangnguyen0403/evals-run from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.