azure/eval-authoring-skill
Author and validate Vally evals for Agent Skills under .github/skills. WHEN: "write a skill eval", "add trigger eval", "test skill routing", "add anti-trigger tests", "create skill capability eval", "harden skill graders". DO NOT USE FOR: single MCP tool prompt-to-tool coverage (use eval-authoring-tool), multi-tool or multi-turn scenarios (use eval-authoring-workflow), authoring SKILL.md itself (use skill-authoring).
npx skills add https://github.com/Azure/azure-sdk-tools --skill eval-authoring-skill
Author Vally evals that verify one Agent Skill's routing and behavior in isolation, under .github/skills/<skill>/evals/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with skill-eval context already loaded, not to duplicate it.
USE FOR: write a skill eval, add trigger eval, test skill routing, add anti-trigger tests, create skill capability eval, harden skill graders
WHEN: "write a skill eval", "add trigger eval", "test skill routing", "add anti-trigger tests", "create skill capability eval", "harden skill graders"
DO NOT USE FOR: single MCP tool prompt-to-tool coverage (use eval-authoring-tool), multi-tool or multi-turn scenarios (use eval-authoring-workflow), authoring SKILL.md itself (use skill-authoring)
.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the skill tier, then the guide's "Skill" column throughout (naming, requirements, worked example).SKILL.md — its WHEN/DO NOT USE FOR boundaries, invoked tools — and its existing evals/.eval.yaml with routing stimuli (trigger + anti-trigger, skill-invocation grader) and capability stimuli together in the same file per the guide's naming convention and four-layer pattern. Split into additional <behavior>.eval.yaml files only once a skill's coverage genuinely grows large. For a boundary prompt, mount and require the intended competing skill while disallowing the skill under test — an anti-trigger with no competing skill mounted trivially "passes" (guide anti-pattern A7).tool-calls with zero recorded calls, which looks like a routing bug but is really a missing-context prompt bug.eval-results.md and results.jsonl on failure.turn for conversation-wide assertions; pin a turn only when that turn owns the behavior unambiguously..github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-tool/eval-authoring-workflow — update it there, not per-skill.Take azure/eval-authoring-skill from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.