mcpbeat

Vally Eval

microsoft/vally-eval

Author, validate, and run Vally evaluation suites for agent skills. TRIGGERS: create eval, write eval, add eval, run eval, validate eval, vally eval, eval.yaml, add stimulus, map test to eval, migrate test to eval, eval graders, eval scoring, add eval to CI.

2k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
242
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/microsoft/GitHub-Copilot-for-Azure --skill vally-eval

What comes with it

2 630 bytes besides the instruction
references/ci-test.md

The instruction itself

11 sections, as written by the author

Vally eval suites

Skills in the azure-skills plugin are required to have integration tests that run prompts against an LLM agent to evaluate whether they help the agent accomplish goals in target scenarios. Such integration tests are written as vally eval suites, using vally as the underlying tool for running tests and grading the agent outcome.

Write vally eval suites

Vally eval suites are written as yaml documents. All eval suites share eval spec.

Refer to the official documentation on the schema of the spec and the schema of the eval suites writing-eval-specs.

Vally eval suites for azure-skills plugin have the following file layout. The shared eval spec is located at <repo-root>/.vally.yaml. The eval suites are categorized by plugin and skills. The eval suites for each skill are located at <repo-root>/evals/<plugin-dirname>/<skill-name>/*.yaml, e.g. <repo-root>/evals/azure-skills/azure-ai/eval.yaml.

Use meaningful file names to categorize tests. If a skill needs fixture files for its eval suites, it should organize such fixture files in a fixture directory under its directory, e.g. <repo-root>/evals/azure-skills/azure-ai/fixture/. The vally test runner and stimulus validation script will load all *.yaml files except for those under a fixture/ directory. Make sure to put all fixture files under the fixture/ directory.

Migrate integration tests

azure-skills plugin have implemented JavaScript integration test using Jest as the underlying test runner. All such integration tests are under tests/**/integration.test.ts files.

To migrate integration test for a skill to vally suites, create its eval suite spec at <repo-root>/evals/<plugin-dirname>/<skill-name>/eval.yaml, add a suite that runs the same prompt and uses vally's built-in graders to grade the trajectory of the agent run. If the integration test grades the agent run in a way that vally's built-in graders don't support, refer to the official documentation on how to create a custom grader writing-custom-grader.

Why is there a custom executor

The legacy Jest based integration test framework implemented features that vally doesn't support yet, such as early termination, follow up, system prompt modification, screenshot taking, etc. Besides, test-all-integration runs automated integration tests, collects its exported data and feeds the data to a dashboard web app under <repo-root>/dashboard/ to monitor skill integration test results.

If you intend to have your vally suites use any of the extended features or have their results be consumed by the dashboard, you MUST use the custom executor in your vally suites.

Use tags to control the custom executor

The custom executor uses special tag values to control the behavior of the custom executor. See tag-helpers.ts to learn what special tags are supported.

> Note: If an eval suite specifies an earlyTerminate condition, the suite MUST NOT use the completed grader because early terminated runs will always fail the completed grader by design.

Validate vally eval suites

Vally eval suites in this repo follow certain conventions. For example, all eval suites must have a type, tier, cost and area tag so they can be run for a corresponding target group. To ensure all eval suites follow the conventions, a script is added to validate the eval suites and report errors when it sees any violation. To run the script, execute this command from the scripts/ directory.

# cwd as <repo-root>/scripts/
npm run vally validate-stimulus

Extended features such as early termination are implemented using tags and many of them use serialized JSON objects as input. This validation script also validates the values of these special tags.

Run vally eval suites locally

Use vally-cli to run vally eval suites. In most cases, you would like to use a command like this.

# In tests/
npm run test:vally -- --plugin $PLUGIN_DIR --skill $SKILL

See vally test runner on how it composes the vally commands under the hood.

Run vally eval suites in CI

Vally eval suites implemented in this repo can be added to the CI test workflow to be run nightly and publish results for reviewing. Refer to ci-test on how to add the Vally eval suites to the CI test workflow.

Extend with custom grader

Custom graders can be added to grade trajectories in ways built-in graders don't support. To add a custom grader, follow the examples in the official vally documentation to create a tests/vally/<custom-name>-grader.ts module and register the new custom grader in tests/vally/vally-graders.ts. The npm run test:vally command internally loads all the custom graders when testing skills.

Re-grade an existing trajectory

Test authors commonly need to fine-tune grader configurations to reduce result flakiness. Vally supports re-grading an existing trajectory using a command like this:

# in tests/
npx @microsoft/vally-cli grade --eval-spec ../evals/<plugin-dirname>/<skill-name>/eval.yaml --verbose < results/<test-run-name>/results.jsonl

You can keep tuning the grader config in eval.yaml and re-grade the trajectory until the results meet your expectations.

If your skill uses a custom grader, add --grader-plugin to load the custom graders.

# in tests/
npx @microsoft/vally-cli grade --eval-spec ../evals/<plugin-dirname>/<skill-name>/eval.yaml --grader-plugin
../../../tests/vally/vally-graders.ts --verbose < results/<test-run-name>/results.jsonl

Note that the grader plugin's path is relative to the parent directory of the eval spec to run. For example, if the eval spec to run is <repo-root>/evals/azure-skills/azure-ai/eval.yaml, resolving this relative path ends at <repo-root>/tests/vally/vally-executor.ts.

Collect test results

When running locally, the test results can be found at the following directories:

  • tests/reports/<test-run-name>/
  • tests/results/<test-run-name>/

When running in CI, the test results can be found in the GitHub Action artifacts or at a storage account that the workflow publishes to. You can also use the integration tests dashboard to view the test results from nightly test runs.

How to use it

Copy the folder

Take microsoft/vally-eval from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference npx. Without those the skill loads but fails at the first command.