Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case. WHEN: "run evaluation", "run eval", "evaluate the bot", "run perf evaluation", "run basic evaluation", "run a single test case", "evaluate one question", "run all scenarios", "score the bot", "run evals locally". DO NOT USE FOR: preparing or uploading datasets, pipeline troubleshooting, knowledge-graph indexing.
npx skills add https://github.com/Azure/azure-sdk-tools --skill sdk-ai-bot-run-evaluation
Run the Azure SDK QA bot evaluation locally with evals_run.py in the package
tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation. It calls the bot /completion endpoint
concurrently to collect answers, then grades them with the Foundry builtin LLM evaluators.
Run all commands from tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation with the package
.venv active and az login done. See running evaluations for flags, the bot endpoint options, single-case and all-scenario recipes, and how to read results.
USE FOR: run an evaluation on a curated dataset (basic or perf); run a single scenario; run all scenarios; run a single test case; grade against the local or deployed bot; inspect results and the pass/fail gate
WHEN: "run evaluation", "run eval", "evaluate the bot", "run perf evaluation", "run basic evaluation", "run a single test case", "evaluate one question", "run all scenarios", "score the bot", "run evals locally"
DO NOT USE FOR: preparing or uploading datasets, pipeline troubleshooting, knowledge-graph indexing
--dataset accepts a local evaluation_datasets/<target>/<scenario>.jsonl path or a Foundry asset name qa-bot-<target>-<scenario>[:version] (the asset name resolves to the local file; the version is informational).--dataset (see the reference).--is_ci False locally (uses az login); CI uses pipeline identity.BOT_SERVICE_ENDPOINT; if unset it defaults to the local http://localhost:8089 — start the agent server.py first for local runs.--evaluators to subset.Confirm and remind the user to set these before running (loads .env; copy and fill
tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation/env-variables):
| Purpose | Variables |
| --------------- | -------------------------------------------------------------------------------------------- |
| Foundry grading | AZURE_AI_PROJECT_ENDPOINT, AZURE_EVALUATION_MODEL_NAME, EVALUATE_THRESHOLD (default 3) |
| Deployed bot | BOT_SERVICE_ENDPOINT + (BOT_AGENT_TOKEN_RESOURCE or BOT_AGENT_ACCESS_TOKEN) |
| Local bot | none — run agent server.py (defaults to http://localhost:8089) |
| Tenant routing | STORAGE_BLOB_ACCOUNT, BOT_CONFIG_CONTAINER, BOT_CONFIG_CHANNEL_BLOB |
| Auth | az login (with --is_ci False) |
# One scenario (all evaluators), against the local server, no baseline gate:
python evals_run.py --dataset evaluation_datasets/perf/typespec.jsonl \
--is_ci False --baseline_check False --cache_result full
# Same via the asset name:
python evals_run.py --dataset "qa-bot-perf-typespec:latest" --is_ci False --baseline_check False
# Subset of evaluators:
python evals_run.py --dataset "qa-bot-basic-python:latest" \
--evaluators "bot_evals,groundedness" --is_ci False --baseline_check False
For a single test case, all scenarios, the deployed bot, and reading
results / the gate, see running evaluations.
server.py.python evals_run.py --dataset <...> --is_ci False (add --baseline_check Falsefor ad-hoc runs, --evaluators to subset).
--cache_result full, inspect cache/<scenario>-*.json for per-case andfailed-case detail.
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
Advanced GitHub Actions workflow automation with AI swarm coordination, intelligent CI/CD pipelines, and comprehensive repository management
Google Cloud Platform CLI - manage GCP resources including Compute Engine, Cloud Run, GKE, Cloud Functions, Storage, BigQuery, and more.
Expert backend architect specializing in scalable API design, microservices architecture, and distributed systems. Masters REST/GraphQL/gRPC APIs, event-driven architectures, service mesh patterns, and modern backend frameworks. Handles service boundary definition, inter-service communication, resilience patterns, and observability. Use PROACTIVELY when creating new backend services or APIs.
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
Aspire skill covering the Aspire CLI, AppHost orchestration, service discovery, integrations, MCP server, VS Code extension, Dev Containers, GitHub Codespaces, templates, dashboard, and deployment. Use when the user asks to create, run, debug, configure, deploy, or troubleshoot an Aspire distributed application.
Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness. Returns a structured report with exact patch locations.
Microsoft Store Developer CLI (msstore) for publishing Windows applications to the Microsoft Store. Use when asked to configure Store credentials, list Store apps, check submission status, publish submissions, manage package flights, set up CI/CD for Store publishing, or integrate with Partner Center. Supports Windows App SDK/WinUI, UWP, .NET MAUI, Flutter, Electron, React Native, and PWA applications.
Take azure/sdk-ai-bot-run-evaluation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.