mcpbeat Sign in

Eval Authoring Tool Agent Skill

Author and validate hermetic single-tool Vally evals under evals/tools. WHEN: "write a tool eval", "add prompt-to-tool coverage", "test MCP tool selection", "add tool catalog eval", "create single-tool scenario", "harden tool-call grader". DO NOT USE FOR: skill routing or capability evals (use eval-authoring-skill), multi-tool, multi-turn, or live scenarios (use eval-authoring-workflow).

2k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
136
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/Azure/azure-sdk-tools --skill eval-authoring-tool

What comes with it

3 079 bytes besides the instruction
evals/eval.yaml

The instruction itself

5 sections, as written by the author

Tool Eval Authoring

Author Vally evals that verify one MCP tool is selected correctly for a given prompt, under evals/tools/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with tool-eval context already loaded, not to duplicate it.

Triggers

USE FOR: write a tool eval, add prompt-to-tool coverage, test MCP tool selection, add tool catalog eval, create single-tool scenario, harden tool-call grader

WHEN: "write a tool eval", "add prompt-to-tool coverage", "test MCP tool selection", "add tool catalog eval", "create single-tool scenario", "harden tool-call grader"

DO NOT USE FOR: skill routing or capability evals (use eval-authoring-skill), multi-tool, multi-turn, or live scenarios (use eval-authoring-workflow)

Steps

  • Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the tool tier, then the guide's "Tool" column throughout (naming, requirements, worked example).
  • Confirm the scenario expects one primary tool. If success requires orchestration, conversation state, or live services, use eval-authoring-workflow instead.
  • Add stimuli to the matching prompt-to-tool-<area>.eval.yaml; create a separate file only when it needs fixtures or outcome grading.
  • Use realistic, concrete prompts with multiple natural phrasings and collision cases that disallow the nearest competing tool. Make tool-calls the primary signal; avoid the anti-patterns in the guide (vacuous keyword-only grading, missing scoring.threshold).
  • Keep the unit tier hermetic (environment: azsdk-mcp-mock, tags.tier: unit, established area tag). Avoid git worktrees and production writes.
  • Validate locally per the guide's "Running evals locally" section, building the mock MCP first. Do not finish or open a PR until it passes; inspect recorded tool names and arguments on failure.

Rules

  • Use exact MCP tool names and assert forbidden alternatives where ambiguity exists.
  • Do not force a tool through unnatural prompt instructions unless direct invocation is the contract being tested.
  • Keep fixture paths relative to the eval file and fixture data minimal and non-secret.
  • Preserve existing namespace coverage and add new stimuli instead of duplicating files.

References

  • Eval authoring guide: .github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-skill/eval-authoring-workflow — update it there, not per-skill.
  • Repository-local eval README/configuration discovered in Step 0 (if present)

Other skills for the same job

different authors, same section of the catalogue
Skill Creator
by anthropics
vendor ×10

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

56k tokens scripts
Skill Creator
by vercel-labs
vendor ×10

Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.

12k tokens scripts
Skill Creator
by JayZeeDesign
×9

Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.

10k tokens scripts
Template Skill
by JayZeeDesign
×7

Replace with description of the skill and when Claude should use it.

35 tokens
Dispatching Parallel Agents
by ZhanlinCui
×5

Use when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

2k tokens
Skill Development
by anthropics
vendor ×4

This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.

9k tokens
Find Skills
by sanity-io
vendor ×4

Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.

1k tokens
Writing Skills
by ZhanlinCui
×4

Use when creating new skills, editing existing skills, or verifying skills work before deployment

26k tokens scripts

How to use it

Copy the folder

Take azure/eval-authoring-tool from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.