azure/eval-authoring-workflow
Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval". DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool).
npx skills add https://github.com/Azure/azure-sdk-tools --skill eval-authoring-workflow
Author Vally evals that verify multi-skill, multi-tool, or multi-turn orchestration, state, and outcomes across an entire agent conversation, under evals/workflows/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with workflow-eval context already loaded, not to duplicate it.
USE FOR: write a workflow eval, add end-to-end eval, test tool sequence, create multi-turn eval, add mock workflow scenario, add live eval
WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval"
DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool)
.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the workflow tier, then the guide's "Workflow" column throughout (naming, requirements, worked example).evals/workflows/mock/ by default. Choose live/ only when the behavior cannot be represented by the mock MCP; document writes, authentication, cleanup, and nightly-only execution.turns only when conversation state matters; otherwise keep a single prompt. Mount every candidate skill explicitly, provide minimal file/git fixtures, and bound turns, tokens, workers, and timeout.turn only when that turn owns the assertion unambiguously; otherwise grade the full conversation.environment.skills, files, and git.source.1.0 and a meaningful threshold when combining grader types..github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-skill/eval-authoring-tool — update it there, not per-skill.Take azure/eval-authoring-workflow from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.