Create plugin development eval scenarios (JSON files with natural prompts and deterministic checks for testing plugin skills). NOT for Copilot Studio in-product evaluation — use /copilot-studio:create-eval-set for that.
npx skills add https://github.com/microsoft/skills-for-copilot-studio --skill create-eval
Guide the user through creating eval test cases for a Copilot Studio plugin scenario. Evals test end-to-end scenarios with natural prompts — the request routes through sub-agents (e.g., Author agent) which invoke skills internally.
The eval harness (evals/evaluate.py) works by:
claude -p "<prompt>" with a PreToolUse hook that traces skill invocations inside sub-agentsAuthoring scenarios that produce YAML files (topics, agents, knowledge sources, etc.) are the best candidates. The harness supports these check types:
| Check | What it validates | Use for |
|-------|------------------|---------|
| agent_invoked | Expected sub-agent was dispatched (e.g., Author agent) | Routing verification |
| agent_not_invoked | Unwanted sub-agents were NOT dispatched | Routing verification |
| skill_invoked | Expected skill was invoked (traced inside sub-agents via hook) | Skill routing |
| skill_not_invoked | Unwanted skills were NOT invoked | Skill routing |
| files_created | Expected files were created/modified (glob pattern) | All authoring scenarios |
| schema_validate | Full Copilot Studio schema validation (kind, required fields, IDs, Power Fx, scopes) | All YAML-producing scenarios |
| yaml_structure | Specific YAML path has expected value, min array length, or contains string | Structural assertions |
| content_contains | Keywords from prompt appear in output files | Domain relevance |
| no_placeholders | No _REPLACE, TODO, or FIXME markers left | Template completion |
| stdout_contains | CLI response text contains expected strings | Reference/info scenarios |
| stdout_not_contains | CLI response does NOT contain error strings | Error absence |
| exit_code | CLI exited with expected code | All scenarios |
| yaml_unchanged | Specific file or YAML path was NOT modified | Preservation testing |
Note: no_placeholders runs automatically when any .mcs.yml file is changed, unless explicitly set to false.
Not yet testable: Integration scenarios that call external APIs (chat-directline, manage-agent) — these need script mocking which isn't implemented yet.
Fixtures are pre-built agent directories in evals/fixtures/:
GenerativeActionsEnabled: false, one Greeting topic. Use for most authoring evals.If the scenario needs a richer agent (e.g., existing topics to modify, knowledge sources, actions), note that the fixture would need to be created first.
$ARGUMENTS is provided, use it as the scenario name. Otherwise ask the user what scenario they want to test (e.g., "topic creation", "agent settings", "knowledge sources"). Glob: skills/*/SKILL.md
Understand: What skills are involved? What YAML kinds? What files get created/modified?
Glob: evals/scenarios/<scenario-name>.json
If yes, read them and offer to add more test cases. Note the highest existing eval ID.
basic-agent)For topic-creation scenarios:
{
"agent_invoked": "copilot-studio:Copilot Studio Author",
"skill_invoked": "copilot-studio:new-topic",
"files_created": [{"pattern": "topics/*.topic.mcs.yml", "min_count": 1}],
"schema_validate": true,
"yaml_structure": [
{"path": "kind", "equals": "AdaptiveDialog"},
{"path": "beginDialog.kind", "equals": "<trigger-type>"}
],
"content_contains": ["<domain keywords>"],
"no_placeholders": true
}
For agent-settings scenarios:
{
"agent_invoked": "copilot-studio:Copilot Studio Author",
"skill_invoked": "copilot-studio:edit-agent",
"files_created": [{"pattern": "agent.mcs.yml", "min_count": 1}],
"schema_validate": true,
"yaml_structure": [
{"path": "kind", "equals": "GptComponentMetadata"}
],
"content_contains": ["<expected content>"],
"no_placeholders": true
}
For knowledge-source scenarios:
{
"agent_invoked": "copilot-studio:Copilot Studio Author",
"skill_invoked": "copilot-studio:add-knowledge",
"files_created": [{"pattern": "knowledge/*.knowledge.mcs.yml", "min_count": 1}],
"schema_validate": true,
"no_placeholders": true
}
For reference/query scenarios:
{
"stdout_contains": ["<expected content in response>"],
"exit_code": 0
}
Write: evals/scenarios/<scenario-name>.json
Format:
{
"scenario_name": "<scenario-name>",
"evals": [
{
"id": 1,
"name": "<short descriptive title>",
"prompt": "<natural language request — what a user would say>",
"fixture": "basic-agent",
"mock_scripts": [],
"checks": { ... }
}
]
}
python3 evals/evaluate.py --scenario <scenario-name> --verbose
Or for all scenarios: node evals/run.js
To generate the HTML report: python3 evals/report.py evals/results/<timestamp>/
agent_invoked and skill_invoked checks to verify correct routingschema_validate: true for ALL scenarios that produce YAML — it's the most powerful checkcontent_contains keywords should come directly from the prompt to verify domain relevanceMulti-agent autonomous startup system for Claude Code. Triggers on "Loki Mode". Orchestrates 100+ specialized agents across engineering, QA, DevOps, security, data/ML, business operations, marketing, HR, and customer success. Takes PRD to fully deployed, revenue-generating product with zero human intervention. Features Task tool for subagent dispatch, parallel code review with 3 specialized reviewers, severity-based issue triage, distributed task queue with dead letter handling, automatic deployment to cloud providers, A/B testing, customer feedback loops, incident response, circuit breakers, and self-healing. Handles rate limits via distributed state checkpoints and auto-resume with exponential backoff. Requires --dangerously-skip-permissions flag.
Use when working with error debugging multi agent review
Build evaluation frameworks for agent systems. Use when testing agent performance systematically, validating context engineering choices, or measuring improvements over time.
Diagnoses and debugs A2A agent communication issues including agent status, message routing, transport connectivity, and log analysis. Use when agents aren't responding, messages aren't being delivered, routing is incorrect, or when debugging orchestrator, coder-agent, tester-agent communication problems.
Use when working with error debugging multi agent review
Rapidly creates atomic, focused skills optimized with evidence-based prompting, specialist agents, and systematic testing. Each micro-skill does one thing exceptionally well using self-consistency, program-of-thought, and plan-and-solve patterns. Enhanced with agent-creator principles and functionality-audit validation. Perfect for building composable workflow components.
Ultimate multi-agent framework for Google Antigravity. Orchestrates specialized domain agents (PM, Frontend, Backend, Mobile, QA, Debug) via Serena Memory.
This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines.
Take microsoft/create-eval from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.