mcpbeat

Eval Authoring Workflow

azure/eval-authoring-workflow

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval". DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool).

2k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
136
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/Azure/azure-sdk-tools --skill eval-authoring-workflow

What comes with it

3 132 bytes besides the instruction
evals/eval.yaml

The instruction itself

5 sections, as written by the author

Workflow Eval Authoring

Author Vally evals that verify multi-skill, multi-tool, or multi-turn orchestration, state, and outcomes across an entire agent conversation, under evals/workflows/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with workflow-eval context already loaded, not to duplicate it.

Triggers

USE FOR: write a workflow eval, add end-to-end eval, test tool sequence, create multi-turn eval, add mock workflow scenario, add live eval

WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval"

DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool)

Steps

  • Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the workflow tier, then the guide's "Workflow" column throughout (naming, requirements, worked example).
  • Use evals/workflows/mock/ by default. Choose live/ only when the behavior cannot be represented by the mock MCP; document writes, authentication, cleanup, and nightly-only execution.
  • Model one user goal per stimulus. Use turns only when conversation state matters; otherwise keep a single prompt. Mount every candidate skill explicitly, provide minimal file/git fixtures, and bound turns, tokens, workers, and timeout.
  • Combine process and outcome graders per the guide's four-layer pattern and grader catalog: required/disallowed skills and tools, ordering where essential, files/commands, response quality. Avoid overfitting to incidental call sequences and the guide's anti-patterns (e.g. boundary anti-triggers with no competing skill mounted).
  • Scope graders to turn only when that turn owns the assertion unambiguously; otherwise grade the full conversation.
  • Validate locally per the guide's "Running evals locally" section, building the matching MCP and priming git fixtures when required. Do not finish or open a PR until it passes; inspect every turn and tool call on failure.

Rules

  • Keep mock workflows hermetic and repeatable; use fake IDs and canned responses.
  • Never run a live write scenario without the documented test area and safety environment.
  • Preserve relative path invariants for environment.skills, files, and git.source.
  • Use weights summing to 1.0 and a meaningful threshold when combining grader types.

References

  • Eval authoring guide: .github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-skill/eval-authoring-tool — update it there, not per-skill.
  • Repository-local eval README/configuration discovered in Step 0 (if present)

How to use it

Copy the folder

Take azure/eval-authoring-workflow from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.