hoangnguyen0403/evals-run
Workflow skill for evals run.
npx skills add https://github.com/HoangNguyen0403/agent-skills-standard --skill evals-run
> [!IMPORTANT]
> Workflow skill for evals run.
Optional args: slug=<feature>, ticket=<id/url>, mode=interactive|autonomous|channel, channel=<id>, auto_continue=true|false, profile=business|hybrid|technical.
When the user asks to perform this workflow, execute the following steps:
description: Run blinded live skill evals and publish reproducible v2 results.
Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.
pnpm evals:baseline first. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers.pnpm evals:baseline -- --execute; the default is gpt-5.6-luna with high reasoning and one worker. Override intentionally with EVALS_MODEL, EVALS_REASONING_EFFORT, or EVALS_CONCURRENCY (maximum four workers).--execute command after access resumes; completed answers are reused automatically.pnpm evals:manifest -- --category <category> for one category or pnpm evals:manifest -- --all for the complete catalog.pnpm evals:manifest -- --resume <runId> only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID.SKILL.md.all runs, write answers under answers/<category>/<skill>/<case>; category runs use answers/<skill>/<case>.metadata.agent, metadata.model, and metadata.completedAt after every required answer exists.pnpm evals:score -- --run <runId>.results.json while any arm is pending, verifies source hashes, and writes one immutable inputs.json snapshot before publishing v2 results.pnpm evals:report to project aggregate runs into the newest complete category partitions and update physical history/archive records.pnpm evals:verify -- --run <runId> and, before handoff, pnpm evals:verify -- --all.n/a for compromised arms.results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.feature_status: implemented | partially_implemented | blocked
requirement_trace: manifest -> inputs -> results -> report -> verification
completed_evidence: []
missing_evidence: []
decision_needed: []
recommended_next_workflow: verify-work
Take hoangnguyen0403/evals-run from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.