microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. No built agent required — an idea or description is enough. Promotes eval-first development: write evals before building. Use when anyone mentions agent evaluation, eval planning, "what should we test", "how do we know if the agent is good", test case generation, or interpreting eval results.
npx skills add https://github.com/microsoft/eval-guide --skill eval-guide
<!-- VERSION CHECK INSTRUCTIONS
When the preamble outputs text, handle it as follows:
If the output contains "UPGRADE_AVAILABLE <old> <new>":
Use AskUserQuestion to ask the user:
"eval-guide v<new> is available (you're on v<old>). Upgrade now?"
With these options:
microsoft/eval-guide must already be added via claude plugin marketplace add microsoft/eval-guide)Then run: claude plugin install eval-guide@eval-guide
Then continue with the skill normally.
Then continue with the skill normally.
The config and snooze scripts are in the same bin/ directory as the update check script.
After upgrade completes, tell the user to restart the session for the new version to take effect.
If the output contains "JUST_UPGRADED <old> <new>":
Tell the user: "Running eval-guide v<new> (just updated from v<old>)!" and continue normally.
If the output is empty:
Continue normally — the user is up to date (or check was snoozed/disabled).
-->
Help customers go from "I don't know where to start with eval" to "I have a plan, test cases, and know how to interpret results" — in one session. The customer becomes self-sufficient for future eval cycles.
You do NOT need a built agent to start. All you need is an idea, a description, or even a vague goal. This skill is designed around the eval-first approach: define what "good" looks like and write your evals before you build the agent or feature.
Why eval-first?
Start here whether you:
Stages 0 (Discover), 1 (Plan), and 2 (Generate) all work without a running agent. They help you think through your agent's purpose, design a structured eval plan, and generate test cases — all before writing a single line of agent configuration. Stage 3 (Run) is the only stage that requires a live agent, and it's optional.
This skill is grounded in Microsoft's Practical Guidance on Agent Evaluation (the 10-step playbook) — see playbook.md for the canonical methodology — together with the Eval Scenario Library, Triage & Improvement Playbook, and MS Learn agent evaluation documentation.
Important: You are an enablement accelerator, not a replacement. Each stage generates artifacts the customer can use immediately AND explains the reasoning so they internalize the methodology. After one session, they should be able to do the next eval without us.
Plan produces a populated Eval Suite Template workbook plus a companion interactive HTML review page. Generate and Interpret produce interactive HTML dashboards that open directly in the browser. Dashboard stages run against a tiny localhost HTTP server (serve.py --serve); the customer never sees, downloads, or moves a JSON file. Feedback flows from the browser → server → the AI's bash stdout, in one step.
Flow at each dashboard review stage (Generate, Interpret):
stage-1-data.json).--serve mode. The AI's bash blocks until the customer clicks Approve or Regenerate:python "$(ls ~/.claude/skills/eval-guide/dashboard/serve.py 2>/dev/null || ls ~/.claude/plugins/cache/*/eval-guide/*/skills/eval-guide/dashboard/serve.py 2>/dev/null | head -1)" --stage <name> --serve --data <file>.json
http://localhost:3118: edits fields inline, updates eval-set/case/root-cause details, and adds comments. Edits auto-save to the localhost server./api/feedback. The server captures it, prints the feedback JSON to stdout between marker lines, and shuts down. No file is downloaded; the customer never moves anything. ===EVAL_GUIDE_FEEDBACK_BEGIN===
{ "stage": "...", "status": "confirmed" | "changes_requested", "edits": {...}, "comments": "..." }
===EVAL_GUIDE_FEEDBACK_END===
Decode the JSON between those markers — that's the customer's feedback. (<stage>-feedback.json is also written next to the data file as a debugging backup, but stdout is the primary channel — read from there.)
status: "confirmed" → apply the edits, generate final deliverables (docx, CSV), proceed to next stage.status: "changes_requested" → apply the edits, regenerate the stage data file, re-launch the dashboard. Same loop.The orient stage is a pre-built static HTML (dashboard/orient-dashboard.html) — agent-agnostic, no serve.py, no JSON write, no feedback file. The skill simply opens the file in the customer's browser and continues the conversation. See *Session Start: Orient* below.
Review checkpoints: Plan uses workbook + HTML review. Generate and Interpret use dashboards. Stage 3 (Run) executes tests directly.
Key principle: No final docx or CSV files are generated until the customer confirms the relevant checkpoint. The checkpoint replaces the "does this look right?" chat-based confirmation with a structured review.
Start from wherever the customer is. Most customers come to eval guidance early — they have an idea or a description, not a finished agent. That's exactly right. The eval-first approach means defining "what good looks like" before building.
Ask: "Tell me about the agent you're building or planning to build. It could be a detailed spec, a rough idea, or even just 'we want a bot that helps with X.' We'll use that to build your eval plan — you don't need a running agent to get started."
/clone-agent to import the agent's topics, knowledge sources, and configuration. Use this to pre-fill the Agent Vision in Stage 0.The key message: Writing evals early makes the agent better. The eval plan becomes the spec, and the test cases become the acceptance criteria. Customers who define evals first build more focused agents and catch problems before they reach production.
Once the customer has described their agent in one or two sentences, give them a visual snapshot of the Per-Agent Eval Maturity Model — where their agent stands today and where this session takes it. This is the orientation moment, and it sets the frame for everything that follows.
The orient dashboard is pre-built and shipped with the skill — dashboard/orient-dashboard.html. It is identical for every agent (the maturity model and "what you walk away with" are agent-agnostic), so there is no per-session JSON write and no Python launch. Don't ask for the agent name yet — Stage 0 captures it where it's actually needed for deliverable filenames.
ORIENT_HTML="$(ls ~/.claude/skills/eval-guide/dashboard/orient-dashboard.html 2>/dev/null || ls ~/.claude/plugins/cache/*/eval-guide/*/skills/eval-guide/dashboard/orient-dashboard.html 2>/dev/null | head -1)"
case "$(uname -s 2>/dev/null)" in
Darwin) open "$ORIENT_HTML" ;;
Linux) xdg-open "$ORIENT_HTML" ;;
*) cmd.exe /C start "" "$ORIENT_HTML" ;; # Windows / Git Bash
esac
The ls ... | head -1 fallback resolves the file regardless of install location — user-global skills first (~/.claude/skills/eval-guide/), plugin-cache second.
For dev installs (skill checked out at an arbitrary path, not in ~/.claude/), the AI should know the absolute path of the SKILL.md it's reading and substitute <SKILL.md-dir>/dashboard/orient-dashboard.html.
This is a read-only stage. There is no feedback file, no confirmation gate, and no serve.py involvement. The customer reviews the snapshot in the browser while the conversation continues in chat.
When to rebuild the static HTML: if templates/orient.html, templates/base.html, or examples/stage-orient-data.json change, run python dashboard/build-orient.py once and check in the regenerated orient-dashboard.html. The build script reuses serve.py's generate_html, so the rendering stays consistent with the live dashboards.
Why this matters for the customer: The maturity model is the value moment. Without it, the customer sees a series of stages with no map. With it, they understand exactly what they're getting and what comes next — the eval-first message lands because they can see the full journey.
Skip orient when: the customer has already done a session with the toolkit and is returning for a Stage 1 / Stage 2 / Stage 4 jump-in. Don't re-orient someone who already has the map.
| Customer says... | Start at |
|---|---|
| "We're planning to build an agent for..." | Stage 0: Discover — eval-first: define evals before building |
| "We have an idea for an agent, what should we test?" | Stage 0: Discover — perfect, evals start from an idea |
| "Help us think through what good looks like" | Stage 0: Discover |
| "I want to add a new feature to my agent" | Stage 0: Discover — write evals for the feature before building it |
| "Here's our agent description, plan the eval" | Stage 1: Plan |
| "I already have a plan, generate test cases" | Stage 2: Generate |
| "I have eval results, what do they mean?" | Stage 4: Interpret |
When running the full pipeline, complete each stage, show the output, explain your reasoning, then ask: "Ready for the next stage?"
Use the Per-Agent Eval Maturity Model as an outcome scorecard to orient customers on where they are today and where this session takes them. It is the progress-framing layer over the 10-step playbook (the canonical methodology lives in playbook.md). Five pillars of eval practice, five levels each — from L100 Initial (no practice in place) to L500 Optimized (continuous improvement built into operations). Assume the agent starts at L100 Initial on all pillars. This session targets L300 Systematic on Pillars 1, 2, and 4 (in-session deliverables) and L200 Defined on Pillars 3 and 5 (via reference protocols delivered alongside the session).
The full 5×5 definitions live in maturity-model.md — that file is the canonical scorecard reference. Each pillar maps to playbook steps (P1=Step 1, P2=Steps 2–5, P3=Steps 6+8, P4=Steps 7+9, P5=Step 8). Update maturity-model.md first when level definitions change.
| Pillar | What it measures | After this session | Mechanism |
|---|---|---|---|
| 1 — Define what "good" means | Acceptance criteria quality | L300 Systematic ✓ | Stage 0 (Discover) + Stage 1 (Plan) |
| 2 — Build your eval sets | Coverage and versioning | L300 Systematic ✓ | Stage 2 (Generate) |
| 3 — Run evals across the lifecycle | Where and when evals execute (offline, pre-deploy, production) | L200 Defined ✓ | rerun-protocol-<agent>-<date>.docx (starter artifact) |
| 4 — Improve and iterate | How improvements are validated | L300 Systematic ✓ | Stage 4 (Interpret) — only if eval results are available |
| 5 — Handle changes with confidence | How changes (prompts, tools, models, architecture) get tested before shipping | L200 Defined ✓ | baseline-comparison-<agent>-<date>.xlsx (starter artifact) |
Pillars 3 and 5 stop at L200 Defined this session. L300 Systematic on those pillars requires operating practice — a release cadence with codified triggers (Pillar 3) and version-tagged baselines accumulated over multiple changes (Pillar 5). The starter artifacts get the customer to L200 in one session: a documented protocol and a fill-in workbook they can execute when triggered. Generate rerun-protocol-<agent>-<date>.docx and baseline-comparison-<agent>-<date>.xlsx at the end of Stage 2 (see deliverables C and D in Stage 2's "After confirmation" block).
Each stage below includes a maturity callout naming which pillar and level it advances.
The toolkit's canonical methodology is Microsoft's *Practical Guidance on Agent Evaluation* — a 10-step playbook (full definition in playbook.md). The operational stages below are the session UX; each one delivers specific playbook steps. Share this crosswalk with customers so they see how the accelerator maps to the guidance. Prefer stage names over numbers when talking to customers — the playbook owns the numbering.
| Operational stage | Playbook steps delivered | What it means |
|---|---|---|
| Discover | Step 1 — Plan the eval effort | Name the eval objective, classify the agent's risk tier (5 factors), name an owner. Articulate purpose/users/boundaries/success — the eval spec. |
| Plan | Steps 2–5 (plan side) | Decompose into capability eval sets and trust & safety eval sets, set pass-rate targets + hard/soft gates, specify human inputs (rubrics, ground truths, source→ground-truth map). |
| Generate | Steps 2, 3, 5 (build) + Step 8 (design) | Produce the capability + trust & safety eval sets (CSVs + manifest); tag each set gate-only | regression | exploratory for the regression suite. |
| Run | Step 6 — Run the baseline | Execute the suite vs the current build; record per-set results with version + timestamp. |
| Interpret | Step 7 — Iterate to diagnose (+ Step 9 design) | Classify each failure as eval-setup vs agent-quality; SHIP/ITERATE/BLOCK on gates; design the production optimization loop. |
| Closeout _(folded into the Interpret report)_ | Step 10 — Reusable assets | Flag reusable rubrics / trust & safety sets for the shared library (Required / Recommended / Opt-in). |
Steps 8–10 (regression suite, optimization loop, reusable assets) are designed in-session and run over time — the session leaves the customer reference artifacts to execute them.
When to share this: After Discover, show the customer the crosswalk and say: *"Today covers Steps 1–5 of Microsoft's playbook — planning the effort and building your capability and trust & safety eval sets. Once you have a running agent you'll run the baseline (Step 6), iterate (Step 7), then stand up the regression suite, optimization loop, and shared-asset library (Steps 8–10)."*
Downloadable reference: Point customers to Microsoft's Eval Guidance Kit to track progress through all ten steps independently.
Help the customer articulate what their agent is supposed to do and what "good" looks like. This is the most important stage — it shapes everything downstream.
Don't ask Q1–Q7 in chat. This was the old flow; it tested as an interrogation and customers tuned out. The new flow: extract everything you can from the customer's kickoff description, fill the gaps with domain-keyed safe defaults, summarize in 5–6 lines, and proceed straight to the Plan dashboard. The customer corrects in chat ("actually, peer comp comparison isn't a boundary for us") or via the dashboard's General Comments box. Nothing is locked until they confirm in the dashboard.
From the customer's 1–4 sentence description, extract:
If the kickoff is too thin (one sentence with no domain hint), ask one clarifying question — *"Two more sentences on what it does and who uses it would help me draft a Vision faster"* — then resume.
Domain detection runs on keywords in the kickoff description. Pick the matching default set:
| Domain trigger keywords | Default boundaries (what NOT to do) | Default risk tier |
|---|---|---|
| HR / ESS / employee / benefits / policy / leave / payroll | Legal advice; medical advice; salary negotiation; performance review interpretation; HR investigation details; peer compensation comparison; PII about other employees | HIGH (data sensitivity + regulatory exposure) |
| Customer support / refunds / billing / accounts | Refunds beyond policy; account-specific data outside this user's scope; legal-binding promises; competitor product recommendations | HIGH (reach + criticality: customer trust + financial) |
| Knowledge / documentation / FAQ / wiki | Content beyond the named knowledge sources; opinions framed as facts; regulated advice (legal/medical/financial) | MEDIUM (defaults higher if regulated content domain) |
| IT / helpdesk / troubleshooting | Remote-execute actions on user systems; reset credentials without verification; security advice that bypasses policy | MEDIUM (HIGH if security/privacy adjacent) |
| Agentic / tool-using / "submits" / "schedules" / "books" | Irreversible actions without confirmation; actions outside user's authorization scope; anything requiring approval the agent can't get | HIGH (autonomy / blast radius: writes to systems) |
| No domain detected | "Outside the named knowledge sources" + "anything the user-cohort isn't authorized for" + 1 generic safety guardrail | MEDIUM (default cautious) |
The default tier is a starting point keyed on domain. Confirm it against all five risk factors — reach (who and how many use it), criticality of error (financial/legal/safety/reputational consequence), autonomy / blast radius (does it only draft text a human reviews, or take irreversible actions?), regulatory exposure (HIPAA/GDPR/SOX/fiduciary/attorney-client), data sensitivity (PII/PHI/confidential/source code). In enterprise contexts autonomy and regulatory exposure often dominate, so bump the tier up when either is present even if reach is small.
Default success criteria (always include unless customer overrides):
Default knowledge sources when only categorized:
Multiple SharePoint sites (TBD — name in Plan dashboard) so the customer can fill names without us blocking on it.Auto-detect role-based access: if the customer's description contains "your," "personalized," "based on your," "role-specific," "tailored to," set role_based_access: true and infer 2–3 likely personalization axes from the agent's domain (HR/ESS → location, tenure, plan; customer support → account tier, region; etc.). Customer corrects if wrong.
Marketing-language capabilities like *"empower employees," "explore opportunities," "streamline X"* don't survive the concreteness check. Drop them from Core Capabilities and add a one-line note in the Vision summary: *"Note: dropped 'explore opportunities' as aspirational — not a testable feature. Tell me if it's actually a concrete capability and I'll add it back."*
This is silent removal with a flagged note, not a question. Customer can flag if they disagree.
Display the pre-extracted Vision compactly:
Agent Vision: [Name]
Eval objective: [one sentence — what "good" means + what decision the evals inform]
Purpose: [one sentence from kickoff]
Users: [extracted or default]
Knowledge: [named sources, or "TBD — confirm in Plan dashboard"]
Capabilities: [3–5 from kickoff, aspirational dropped]
Boundaries: [domain default set, listed]
Success: [default 3 criteria]
Role-based: [auto-detected: yes/no, with axes]
Risk tier: [domain default: HIGH/MEDIUM/LOW] — driven by 5 factors (reach, criticality, autonomy, regulatory, data sensitivity)
Owner: [named accountable owner, or "TBD — name before deploy"]
Then: *"This is what I extracted from your description, with safe defaults for [HR/ESS/etc.] domain agents filling the gaps. Speak up now if any of this is wrong — boundaries, risk tier, eval objective, or capabilities especially. I'm proceeding to draft the eval plan; you'll review the full criteria + matrix in the Plan dashboard."*
Don't gate on customer confirmation. Write stage-0-data.json and proceed to Stage 1 immediately. The customer either replies with corrections (which you incorporate before launching the dashboard) or stays silent (proceed). The Plan dashboard is the real review surface.
Using the Agent Vision, produce a structured eval suite plan. This works whether the agent exists or not — the plan defines what the agent SHOULD do.
Use the attached Eval Suite Template workbook when available. Populate a copy of it only. Do not rename sheets, add sheets, add columns, change headers, rewrite README text, edit Dropdown Lists, change styles, or change data validation. If the template is missing, ask for it instead of creating a different workbook.
Fill the existing 1 . Planning input cells:
Populate 2 . Eval Suite Registry with one row per eval set:
Do not generate legacy planning-artifact rows in the workbook. The registry is one row per eval set only.
Use the existing registry columns:
Target pass rate, Target rationale, Gate type, Intended use, Run cadence, and Notes; do not add a new column.Use the registry columns for human input type/author, grounding source dependency, and source-change review. Use TBD - confirm before baseline where owners or sources are unknown.
The template has no grader-validation columns. Do not add them. Record grader type and validation expectations in each registry row's Notes, e.g. programmatic check to confirm, human-review agreement, or LLM-as-judge validation against human-labeled hard and borderline cases.
Add optional baseline placeholder rows in 3 . Run Log: Run type = Baseline, result fields blank, Actionable next step = Validate grader, then run baseline, Status = Open.
Use Intended use and Run cadence in the registry. Capability sets usually become Both or Regression; most T&S sets are Gate, with a slim regression subset for model/tool/policy changes.
Populate 4 . Reusable Library only with candidates that could help other agents: reusable T&S sets, rubrics, failure-pattern templates, or production-derived edge-case categories.
Do not display a long eval-set summary in chat. Put Step 1 objective/risk/owner, capability eval sets, trust & safety eval sets, Step 4 governance, Step 5 human inputs, Step 6 grader-validation notes, Step 8 cadence, and Step 10 reusable candidates into the interactive HTML review page described below.
The customer payoff: *"You now have a workbook your PM, builder, risk owner, and source owners can review. It preserves your template and shows which eval sets exist, how each is governed, who owns human inputs, what must happen before baseline, and which assets may be reusable."*
Maturity callout — Pillar 1 / playbook Step 1 (L100 Initial → L300 Systematic): Discover + Plan advance Pillar 1 from "good lives in the builder's head" to a written objective, five-factor risk tier, accountable owner, and workbook-backed eval-set governance. Pillar 2 advances in Generate (Steps 2, 3, 5); Pillar 3 now starts with Step 6 grader validation before any baseline is trusted.
The legacy Plan dashboard is pre-v5 criteria-based and must not be used for the v5 workbook workflow. Instead, generate a draft workbook copy plus a companion HTML review page and have the customer review those artifacts.
eval-suite-<agent-name>-<YYYY-MM-DD>.xlsx as a populated copy of the user's template.eval-suite-<agent-name>-<YYYY-MM-DD>-review.html next to it using skills/eval-guide/plan-review-page.md.1 . Planning: objective, risk tier, owners, sign-off criteria.2 . Eval Suite Registry: eval-set rows, Step 4 governance, cadence, human inputs, source dependencies, grader-validation notes, reusable flags.3 . Run Log: baseline placeholders, if added.4 . Reusable Library: reusable candidates.Dropdown Lists.After confirmation, the eval plan deliverable is:
Customer-ready .xlsx eval-suite planning workbook using the /xlsx skill, named eval-suite-<agent-name>-<YYYY-MM-DD>.xlsx. It must be a populated copy of the user's template, not a recreated or redesigned spreadsheet.
Interactive HTML review page named eval-suite-<agent-name>-<YYYY-MM-DD>-review.html. The page carries the summary, eval-set explorer, TBD action list, and human review checklist so the chat response stays concise.
The workbook is the review checkpoint and the primary Plan artifact. A separate .docx narrative is optional only if the user asks for it.
Tell the customer only where to open the workbook and HTML review page. Do not duplicate the HTML page content in chat.
Generate test cases as separate CSV files per eval set from the workbook registry. These are the customer's deliverable — they can import them into Copilot Studio or use them as acceptance criteria during development.
| Artifact | Use it for |
|---|---|
| eval-<set-type>-<set-slug>-<date>.csv (per eval set — 2 columns: Question, Expected response; one row per case; the Testing method is assigned per row in Copilot Studio's Evaluate tab after import) | Paste directly into Copilot Studio Evaluation tab |
| eval-test-cases-<agent>-<date>.docx | PM / stakeholder review |
| eval-setup-guide-<agent>-<date>.docx | Step-by-step walkthrough for setting up + running the eval in Copilot Studio's Evaluate tab |
| rerun-protocol-<agent>-<date>.docx | Pillar 3 L200 — when to re-run the eval as the agent changes |
| baseline-comparison-<agent>-<date>.xlsx | Pillar 5 L200 — your version-comparison workbook |
The kit is one deliverable. CSVs go to Copilot Studio. The test-case .docx goes to your PM. The setup guide, rerun protocol, and baseline-comparison workbook go to your eval-process docs.
Default to Single Response. ~80% of agents are single-response Q&A. Conversation (multi-turn) only fits agents that do real multi-step workflows — troubleshooting flows, form-filling, slot-extracting conversations. If you're not sure, you don't need Conversation mode.
| Mode | Best for | Limits | Supported test methods |
|---|---|---|---|
| Single response *(default — fits ~80% of agents)* | Factual Q&A, tool routing, specific answers, safety tests | Up to 100 test cases per set | All 7 methods (General quality, Compare meaning, Keyword match, Capability use, Text similarity, Exact match, Custom) |
| Conversation (multi-turn) | Multi-step workflows, context retention, clarification flows, process navigation | Up to 20 test cases, max 12 messages (6 Q&A pairs) per case | General quality, Keyword match, Capability use, Custom (Classification) |
When to switch to conversation eval:
When to stay with single response (the default):
Explain the choice: "I'm recommending single response eval for your knowledge-lookup criteria because each question is independent — the agent doesn't need previous context to answer. For your troubleshooting criterion, I'm recommending conversation eval because the agent needs to gather information across multiple turns before resolving the issue."
Note for CSV generation: Single response test sets use the 2-column import CSV (Question, Expected response); the testing method is assigned per row in Copilot Studio's Evaluate tab after import (see the manifest note below). Conversation test sets can be imported via spreadsheet or generated in the Copilot Studio UI — each test case contains a sequence of user messages that simulate a multi-turn interaction.
If the Agent Vision has role_based_access: true (set in Discover), the test cases for personalization criteria need user profiles in Copilot Studio. Without profiles, the agent has no context to personalize from — and the test results are misleading.
Walk the customer through this BEFORE generating cases:
Boston-2yr-PPO)Seattle-7yr-HMO)Remote-FirstYear-HDHP)If role_based_access: false, skip this branch entirely — no profile setup needed.
When generating expected responses, the AI wraps factual content it can't independently confirm in [VERIFY: ...] markers. These are the failures-in-waiting. A wrong [VERIFY] becomes an eval test case that "passes" while hiding a production failure — the agent matches the bogus expected response and gets a green check.
The dashboard highlights every [VERIFY] span in yellow. Read every one before approving. This is the customer's most important responsibility in Stage 2; the LLM that drafted the test cases cannot do this work — only the human who knows the actual knowledge sources can.
When narrating to the customer, say: *"I've wrapped factual claims I'm guessing at in [VERIFY] markers. Please check each one against your real knowledge source — these are the most likely places the eval will lie to you about agent quality."*
eval-capability-accuracy-correctness.csveval-capability-faithfulness-groundedness.csveval-capability-reasoning-tool-use.csveval-trust-safety-sensitive-data-handling.csveval-trust-safety-prompt-injection-jailbreak.csveval-trust-safety-compliance-specific.csv (if applicable)Only create files for categories that apply.
Versioning: Name each file with a date stamp or agent version (e.g., eval-knowledge-accuracy-2026-04-22.csv) so successive sessions produce a version history rather than overwriting the baseline. Versioning is a requirement of L300 Systematic Pillar 2.
"Question","Expected response"
"How many PTO days do LA employees get?","LA employees receive 18 PTO days per year."
The Testing method is NOT a CSV column — it is assigned per row in Copilot Studio's Evaluate tab after import (see the manifest note below). The method chosen for each criterion travels in the companion .docx manifest and the eval-setup-guide-<agent>-<date>.docx, which walk the customer through the manual per-row assignment. Valid Testing method values (assigned in the UI): General quality, Compare meaning, Text similarity, Exact match, Keyword match (core five), plus Capability use and Custom (extensions). *(A 3-column -with-methods variant may be emitted as a human-readable reference only — never import it.)*
| Criterion style | Method | Why |
|---|---|---|
| Factual with known answer | Compare meaning | Semantic equivalence |
| Open-ended quality | General quality | LLM judge |
| Must-include terms (URL, email) | Keyword match | Exact presence |
| Agent should refuse | Compare meaning | Refusal matches expected |
| Domain-specific criteria (compliance, tone, policy) | Custom | Define your own rubric and pass/fail labels |
Display a summary table of test cases per eval set.
The customer payoff: *"You now have a test suite that imports directly into Copilot Studio, plus the .docx report your PM can sign off on, plus the Pillar 3 and Pillar 5 starter artifacts you'll keep for ongoing operations. That's the eval kit a new team member would need to evaluate this agent — questions, expected responses, methods, re-run protocol, comparison template."*
Maturity callout — Pillar 2 / playbook Steps 2, 3, 5 (L100 Initial → L300 Systematic): Generate advances Pillar 2 from "no established eval set" to versioned capability eval sets and separate trust & safety eval sets, coverage mapped to risk and value, each tagged gate-only | regression | exploratory for the regression suite (Step 8). Pillar 4 advances in Interpret. Pillars 3 and 5 reach L200 Defined via the rerun-protocol-<agent>-<date>.docx and baseline-comparison-<agent>-<date>.xlsx starter artifacts generated at session close — surface these to the customer when delivering them.
Before generating final CSV and report files, launch the test cases dashboard for review:
stage-2-data.json. Methods and governance metadata live at the eval-set (test_set) level, inherited from the workbook registry. {
"agent_name": "...",
"test_sets": [
{
"eval_set_id": "CAP-ACC-001",
"display_name": "Policy answer correctness",
"set_type": "capability",
"capability_dimension": "Accuracy / correctness",
"methods": ["Compare meaning", "Keyword match"],
"gate_type": "Hard floor + soft target",
"target_pass_rate": "Launch floor 90%; regression/direction after baseline",
"run_cadence": "Weekly",
"cases": [
{
"id": 1,
"question": "...",
"expected_responses": {
"Compare meaning": "Canonical answer, with [VERIFY: factual content to check] markers",
"Keyword match": "PTO, Time Off Policy, accrual"
},
"custom_rubric": ""
}
]
}
]
}
Key requirements:
methods: [] array — the methods for this eval set's CSV. Choose one method when one fits; choose multiple only when the eval set genuinely needs them. Default to one method.gate_type, target_pass_rate, target_rationale, run_cadence, owner/source/grader notes where available).expected_responses: { method → value } — one entry per method in the eval set's methods array that needs a per-case reference (Compare meaning, Text similarity, Exact match, Keyword match). Methods that grade against a set-level rubric (General quality, Capability use, Custom) do NOT need entries.[VERIFY: ...] markers inside the Compare meaning / Text similarity entries so the dashboard highlights them for review.Custom method in the eval set: also write a custom_rubric field on each set or case — a short LLM-judge rubric drafted from the eval-set purpose and expected behavior ("Rate the response Pass / Fail. Pass = …. Fail = …. Output PASS or FAIL with a one-sentence reason."). The dashboard shows this as an editable textarea. Don't leave Custom sets without a rubric.Keyword match method: the per-case expected_responses["Keyword match"] value is a comma-separated keyword list (not a reference answer). The dashboard renders this as a "Keywords" column. python "$(ls ~/.claude/skills/eval-guide/dashboard/serve.py 2>/dev/null || ls ~/.claude/plugins/cache/*/eval-guide/*/skills/eval-guide/dashboard/serve.py 2>/dev/null | head -1)" --stage generate --serve --data stage-2-data.json
===EVAL_GUIDE_FEEDBACK_BEGIN=== / ===EVAL_GUIDE_FEEDBACK_END=== markers. Apply every edit it contains, faithfully and without question. The customer's choices are final — do NOT re-litigate, do NOT suggest reverting, do NOT ask for confirmation again, do NOT partially apply. (generate-feedback.json is also on disk as a backup, but stdout is the primary channel.)This applies to ALL edit types:
[VERIFY: …] wrapper: [VERIFY: <content>] → <content>. By the time the customer has confirmed, every span is either edited (already clean) or accepted (marker is now noise).test_sets[i].cases[k].expected_responses["Compare meaning"], test_sets[i].cases[k].expected_responses["Keyword match"], etc. Each method's value updates that method's column for that case.test_sets[i].custom_rubric or test_sets[i].cases[k].custom_rubric) — the customer's refined rubric is final; use it as the LLM judge prompt verbatim.test_sets[i].methods) — adding/removing a method changes which columns and rubric blocks render for that eval set.Then narrate the edits back so the customer sees their changes were captured — count [VERIFY] corrections, count test case additions/deletions, list significant edits, restate updated total case count. Example: *"Got it — 8 [VERIFY] corrections captured, 2 new cases for CAP-ACC-001, total now 56 cases across 7 eval sets."* Don't just say "applied." The narration confirms you parsed correctly; it is NOT an invitation to re-decide.
If changes requested instead of confirmed, regenerate and re-launch.
.docx report, the eval-setup-guide .docx, the rerun-protocol .docx, and the baseline-comparison .xlsx are one delivery, produced together. The customer should see the artifact list in chat ("five files generated") and find the files on disk before they say anything more.A. CSV files — One CSV per eval set: eval-<set-type>-<set-slug>-<date>.csv. Exactly two columns:
"Question","Expected response"
No Testing method column. Copilot Studio's Evaluation tab requires the customer to set the testing method manually per row in the UI after import — it is not pre-encoded in the CSV. The companion eval-setup-guide-<agent>-<date>.docx (deliverable E below) walks the customer through that manual step in detail.
Row generation rule. One row per active case per eval set (no case × method explosion). Per row:
Question = the case's question.Expected response = whichever of the case's expected_responses is most informational, picked by this priority order against the eval set's method set:Compare meaning → case.expected_responses["Compare meaning"].Text similarity → case.expected_responses["Text similarity"].Exact match → case.expected_responses["Exact match"].Keyword match → case.expected_responses["Keyword match"] (comma-separated keyword list).General quality / Custom / Capability use) → leave the cell empty.Strip every [VERIFY: …] marker from the cell value before writing the row. Replace [VERIFY: <content>] → <content>. The markers exist only as a review aid in the dashboard — by the time the customer has clicked Approve, every span has either been confirmed or edited. The CSV is the eval set the customer is importing into Copilot Studio; it must contain clean expected responses with no review-tooling syntax. Apply the regex \[VERIFY:\s*([^\]]*)\] → $1 (or equivalent) to every Expected response cell before emitting the row.
The customer can still edit any cell in CPS or in the CSV before import — for example, switching a row from canonical-answer to keyword-list when they decide that row should use Keyword match. The eval-setup-guide.docx makes this explicit.
An eval set with 12 cases produces exactly 12 rows. (No multiplication by methods.)
Tell the customer: "One CSV per eval set — two columns: Question and Expected response. Import each into Copilot Studio's Evaluation tab. Then in the CPS UI, set the Testing method for every row — this is a manual step. The eval-setup-guide.docx walks you through which method to pick per eval set and what threshold or regression rule to use."
B. .docx report — Generate a customer-ready report using the /docx skill. The report must be:
Report structure:
[VERIFY: …] markers the same way as in the CSV — [VERIFY: <content>] → <content>. The dashboard's review markers don't belong in the customer-facing report.eval-setup-guide-<agent>-<date>.docx (step-by-step Copilot Studio setup), rerun-protocol-<agent>-<date>.docx (Pillar 3 L200), and baseline-comparison-<agent>-<date>.xlsx (Pillar 5 L200). They walk you through how to set up the run today and advance Pillars 3 and 5 from L100 Initial to L200 Defined."*| Pillar | Baseline | After this session | Next-session target |
|---|---|---|---|
| 1 — Define what "good" means | L100 Initial | L300 Systematic ✓ | — |
| 2 — Build your eval sets | L100 Initial | L300 Systematic ✓ | — |
| 3 — Run evals across the lifecycle | L100 Initial | L200 Defined ✓ (via rerun-protocol-<agent>-<date>.docx) | L300 Systematic |
| 4 — Improve and iterate | L100 Initial | L100 Initial | L300 Systematic (Stage 4) |
| 5 — Handle changes with confidence | L100 Initial | L200 Defined ✓ (via baseline-comparison-<agent>-<date>.xlsx) | L300 Systematic |
C. Pillar 3 starter — rerun-protocol-<agent>-<date>.docx — Generate using the /docx skill, sourcing structure and content from skills/eval-guide/rerun-protocol.md. This is the customer's takeaway reference for Pillar 3 L200 Defined: when to re-run evals, what scope to run, how to log the result. The docx is portable, printable, and shareable with the team.
Render the markdown sections as docx sections with the same headings (Purpose, Prerequisites, When to re-run, Run order rule, Logging discipline, Interpreting re-run results, You've reached L200 Defined when…, Path to L300 Systematic, References). Format the trigger table as a styled docx table, color-code the priority column, and put the "You've reached L200 Defined when…" exit criteria in a callout box.
D. Pillar 5 starter — baseline-comparison-<agent>-<date>.xlsx — Generate using the /xlsx skill, sourcing structure and content from skills/eval-guide/baseline-comparison-template.md. This is the customer's fill-in workbook for Pillar 5 L200 Defined: a structured template they fill in each time they compare two eval runs.
Workbook structure (auto-size columns; freeze header rows; protect instruction sheets):
| Sheet | Contents |
|---|---|
| Instructions | Purpose, when to use, prerequisites. Read-first sheet — protected. |
| Comparison | 5-metric comparison table with empty Run 1 / Run 2 / Delta cells (Overall, Capability eval-set pass rate, Trust & Safety gate status, Regression eval-set pass rate, Hard gate failures). Above the table: editable cells for Run 1 name/version, Run 2 name/version, Eval set version, Change description. |
| Case-level delta | 4-row bucket table (Pass-Pass / Fail-Pass / Pass-Fail / Fail-Fail) with empty Count and Notable cases columns. Conditional formatting highlights Pass-Fail row in red. |
| Decision rules | Variance rules, ship/hold logic. Read-only reference sheet. |
| Capability vs. regression | Cheat sheet on the two run types, when to use each. Read-only reference sheet. |
E. Eval setup guide — eval-setup-guide-<agent>-<date>.docx — Always generate this alongside the CSVs (A). It is not optional and not on-request. Without it, the customer is staring at CSVs with no instructions for the manual method-assignment step in CPS. Generate using the /docx skill, sourcing structure and content from skills/eval-guide/eval-setup-guide.md. This is the customer's step-by-step walkthrough for setting up and running the CSVs in Copilot Studio's Evaluate tab — the operational companion to the eval set.
Render the markdown sections as docx sections with the same headings (What you should have before you start, Step 1–8, Per-method setup table, How to choose a threshold, Common setup issues, You've finished setup successfully when…, Related artifacts, References). Format the per-method setup section as styled docx tables; pull the eval-set method decision tree into a callout box; preserve the troubleshooting symptom/cause/fix table verbatim.
Tell the customer: "Five artifacts: the CSVs go straight into Copilot Studio, the test case .docx is for sharing, the new eval-setup-guide-<agent>-<date>.docx walks you through the Evaluate tab step by step (open it the first time you set up the run), and rerun-protocol-<agent>-<date>.docx + baseline-comparison-<agent>-<date>.xlsx are your Pillar 3 and Pillar 5 starter kits — keep them with your eval set."
Stage 3 turns the eval set into evidence. Run your CSVs against the live agent and record the results. 10–30 minutes (depends on test count and auth setup).
eval-results-<agent>-<date>.csv — pass/fail per case, score per LLM method, judge rationale.eval-results-<agent>-<date>.json — same data, programmatic-friendly.First-run pass rate is usually 40–70%, not 80%+. Customers who get 50% on the first run sometimes spiral; they shouldn't. The valuable signal is *which categories* pass and fail, not the headline number. Stage 4 turns the failures into ranked action.
LLM-judge methods are non-deterministic — Compare meaning and General quality show ±5% variance between runs. If a result lands borderline, run it again and take the median.
| Path | When it's right | Setup cost |
|---|---|---|
| Copilot Studio UI Evaluation tab *(default — start here)* | Most customers, especially incidental users. Import eval-<set-type>-<set-slug>-<date>.csv, run, view results in the UI. Use this unless you need automation. | Agent auth only. |
| eval-runner.js (CLI) | You need to automate, run from CI, or use LLM-judge methods the UI doesn't expose. | Node, DirectLine token endpoint, ANTHROPIC_API_KEY (real $ — Claude API costs apply). |
node eval-runner.js --token-endpoint "<URL>" --csv-dir .
Or use /chat-with-agent for individual questions via the Copilot Studio SDK.
Scoring methods:
Compare meaning → semantic equivalence (0.0–1.0, LLM judge)General quality → relevance / groundedness / completeness / abstention (0.0–1.0, LLM judge)Keyword match → code-based string matching (free, deterministic)Exact match → code-based string equality (free, deterministic)Required: ANTHROPIC_API_KEY for LLM-judge methods. Code-based methods run free.
Take microsoft/eval-guide from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.