mcpbeat

Waza Runner

microsoft/waza-runner

| Run evaluations on Agent Skills to measure their effectiveness. "check skill triggers", "skill compliance check", "measure skill performance", "run evals on [skill-name]", "grade skill execution". (use sensei), or general testing unrelated to skills.

1k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
1144
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/microsoft/waza --skill waza-runner

What comes with it

3 171 bytes besides the instruction
references/EVAL-SPEC.md

The instruction itself

13 sections, as written by the author

Skill Eval Runner

> Evaluate Agent Skills like you evaluate AI Agents

This skill runs evaluations on other skills to measure their effectiveness using the same patterns that power AI agent evaluations.

When to Use

  • Running quality evaluations on a skill
  • Testing if a skill triggers on correct prompts
  • Measuring skill behavior quality
  • Generating eval reports for CI/CD

Commands

Run Evals

Run evals on <skill-name>

Initialize Eval Suite

Create evals for <skill-name>

Generate Report

Generate eval report for <skill-name>

Workflow

  • Check for Eval Suite: Look for eval.yaml in the skill directory
  • Load Tasks: Parse task definitions from tasks/*.yaml
  • Execute: Run each task through the configured graders
  • Report: Output results in JSON or Markdown format

Metrics Measured

| Metric | Description | Default Threshold |

|--------|-------------|-------------------|

| Task Completion | Did the skill accomplish the goal? | 80% |

| Trigger Accuracy | Was skill invoked on correct prompts? | 90% |

| Behavior Quality | Tool calls, efficiency, reasoning | 70% |

Grader Types

  • Code Graders: Deterministic assertions, regex matching
  • LLM Graders: Model-as-judge with configurable rubrics
  • Human Graders: Manual review workflow

Example Usage

Running Evals

# From CLI
waza run ./my-skill/eval.yaml

# Output to file
waza run ./my-skill/eval.yaml -o results.json

Interpreting Results

{
  "summary": {
    "pass_rate": 0.85,
    "composite_score": 0.82
  },
  "metrics": {
    "task_completion": { "score": 0.9, "passed": true },
    "trigger_accuracy": { "score": 0.95, "passed": true }
  }
}

References

  • Eval Specification - Full eval.yaml schema
  • Writing Tasks - Task definition guide
  • Grader Reference - Available graders

How to use it

Copy the folder

Take microsoft/waza-runner from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.