mcpbeat Sign in

SDK AI Bot Run Evaluation Agent Skill

Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case. WHEN: "run evaluation", "run eval", "evaluate the bot", "run perf evaluation", "run basic evaluation", "run a single test case", "evaluate one question", "run all scenarios", "score the bot", "run evals locally". DO NOT USE FOR: preparing or uploading datasets, pipeline troubleshooting, knowledge-graph indexing.

2k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
136
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/Azure/azure-sdk-tools --skill sdk-ai-bot-run-evaluation

What comes with it

4 436 bytes besides the instruction
references/running-evaluations.md

The instruction itself

6 sections, as written by the author

Run QA Bot Evaluation

Run the Azure SDK QA bot evaluation locally with evals_run.py in the package

tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation. It calls the bot /completion endpoint

concurrently to collect answers, then grades them with the Foundry builtin LLM evaluators.

Run all commands from tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation with the package

.venv active and az login done. See running evaluations for flags, the bot endpoint options, single-case and all-scenario recipes, and how to read results.

Triggers

USE FOR: run an evaluation on a curated dataset (basic or perf); run a single scenario; run all scenarios; run a single test case; grade against the local or deployed bot; inspect results and the pass/fail gate

WHEN: "run evaluation", "run eval", "evaluate the bot", "run perf evaluation", "run basic evaluation", "run a single test case", "evaluate one question", "run all scenarios", "score the bot", "run evals locally"

DO NOT USE FOR: preparing or uploading datasets, pipeline troubleshooting, knowledge-graph indexing

Rules

  • --dataset accepts a local evaluation_datasets/<target>/<scenario>.jsonl path or a Foundry asset name qa-bot-<target>-<scenario>[:version] (the asset name resolves to the local file; the version is informational).
  • There is no single-testcase flag. To run one case, make a one-row JSONL file and pass it as --dataset (see the reference).
  • Use --is_ci False locally (uses az login); CI uses pipeline identity.
  • The bot endpoint comes from BOT_SERVICE_ENDPOINT; if unset it defaults to the local http://localhost:8089 — start the agent server.py first for local runs.
  • Default evaluators are all seven; pass --evaluators to subset.

Environment

Confirm and remind the user to set these before running (loads .env; copy and fill

tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation/env-variables):

| Purpose | Variables |

| --------------- | -------------------------------------------------------------------------------------------- |

| Foundry grading | AZURE_AI_PROJECT_ENDPOINT, AZURE_EVALUATION_MODEL_NAME, EVALUATE_THRESHOLD (default 3) |

| Deployed bot | BOT_SERVICE_ENDPOINT + (BOT_AGENT_TOKEN_RESOURCE or BOT_AGENT_ACCESS_TOKEN) |

| Local bot | none — run agent server.py (defaults to http://localhost:8089) |

| Tenant routing | STORAGE_BLOB_ACCOUNT, BOT_CONFIG_CONTAINER, BOT_CONFIG_CHANNEL_BLOB |

| Auth | az login (with --is_ci False) |

Common commands

# One scenario (all evaluators), against the local server, no baseline gate:
python evals_run.py --dataset evaluation_datasets/perf/typespec.jsonl \
  --is_ci False --baseline_check False --cache_result full

# Same via the asset name:
python evals_run.py --dataset "qa-bot-perf-typespec:latest" --is_ci False --baseline_check False

# Subset of evaluators:
python evals_run.py --dataset "qa-bot-basic-python:latest" \
  --evaluators "bot_evals,groundedness" --is_ci False --baseline_check False

For a single test case, all scenarios, the deployed bot, and reading

results / the gate, see running evaluations.

Steps

  • Ensure the required env vars are set; for a local run, start the agent server.py.
  • Choose the dataset: a scenario file, an asset name, or a one-row file for a single case.
  • Run python evals_run.py --dataset <...> --is_ci False (add --baseline_check False

for ad-hoc runs, --evaluators to subset).

  • Open the printed Foundry Report URL and review the per-case score table.
  • With --cache_result full, inspect cache/<scenario>-*.json for per-case and

failed-case detail.

Other skills for the same job

different authors, same section of the catalogue
Modal
by christophacham
×3

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

17k tokens
Github Workflow Automation
by ComeOnOliver
×3

Advanced GitHub Actions workflow automation with AI swarm coordination, intelligent CI/CD pipelines, and comprehensive repository management

9k tokens
Gcloud
by Dicklesworthstone
×2

Google Cloud Platform CLI - manage GCP resources including Compute Engine, Cloud Run, GKE, Cloud Functions, Storage, BigQuery, and more.

2k tokens
Backend Architect
by ComeOnOliver
×2

Expert backend architect specializing in scalable API design, microservices architecture, and distributed systems. Masters REST/GraphQL/gRPC APIs, event-driven architectures, service mesh patterns, and modern backend frameworks. Handles service boundary definition, inter-service communication, resilience patterns, and observability. Use PROACTIVELY when creating new backend services or APIs.

7k tokens
Modal
by ComeOnOliver
×2

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

37k tokens
Aspire
by github
vendor ×1

Aspire skill covering the Aspire CLI, AppHost orchestration, service discovery, integrations, MCP server, VS Code extension, Dev Containers, GitHub Codespaces, templates, dashboard, and deployment. Use when the user asks to create, run, debug, configure, deploy, or troubleshoot an Aspire distributed application.

21k tokens
Bigquery Pipeline Audit
by github
vendor ×1

Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness. Returns a structured report with exact patch locations.

1k tokens
Msstore CLI
by github
vendor ×1

Microsoft Store Developer CLI (msstore) for publishing Windows applications to the Microsoft Store. Use when asked to configure Store credentials, list Store apps, check submission status, publish submissions, manage package flights, set up CI/CD for Store publishing, or integrate with Partner Center. Supports Windows App SDK/WinUI, UWP, .NET MAUI, Flutter, Electron, React Native, and PWA applications.

4k tokens

How to use it

Copy the folder

Take azure/sdk-ai-bot-run-evaluation from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.