mcpbeat Sign in

Deploying Scalable Agents Agent Skill

>- Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry. Covers deployment patterns (client-hosted, hosted agents, agent workflows), the agent lifecycle, model routing, response caching, evaluation gates, human-in-the-loop approval, observability with OpenTelemetry, cost optimisation, and smoke-testing deployed agents with the AI Smoke Test action. Based on Lesson 16 of AI Agents for Beginners. Foundry Agent Service, model routing, response caching, evaluation gate, release gate, human approval workflow, agent observability, agent tracing, agent cost optimisation, smoke test a hosted agent, production customer support agent. on-device (use local-ai-agents / Lesson 17), Azure infrastructure provisioning unrelated to agents, non-Foundry deployment targets.

2k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
71207
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/microsoft/ai-agents-for-beginners --skill deploying-scalable-agents

The instruction itself

10 sections, as written by the author

Deploying Scalable Agents with Microsoft Foundry

> Companion skill for Lesson 16 – Deploying Scalable Agents.

> Use it to help a learner move an agent from prototype to a scalable, observable

> production deployment. Ground every recommendation in the lesson content and

> the runnable notebook; do not invent Foundry APIs.

Triggers

Activate this skill when a learner wants to:

  • Deploy an agent to Microsoft Foundry as a hosted agent and make it versioned/observable.
  • Choose between client-hosted, hosted-agent, and agent-workflow deployment patterns.
  • Add model routing, response caching, or bounded concurrency to control latency and cost.
  • Add an evaluation gate so a bad agent version cannot ship.
  • Add a human-in-the-loop approval step for high-risk actions.
  • Instrument an agent with OpenTelemetry tracing for production observability.
  • Smoke-test a deployed agent as a fast post-deploy gate.

Core mental model

A production agent is mostly the operational skeleton *around* the model (~80%),

not the model itself. Map every recommendation to one of these concerns:

| Concern | Prototype → Production |

|---------|------------------------|

| Hosting | notebook → versioned hosted service |

| Identity | your az login → managed identity + scoped RBAC |

| State | in-memory → externalised thread/memory store |

| Failure | traceback → retries, fallbacks, alerts |

| Cost | "a few cents" → tracked, routed, cached, budgeted |

| Quality | eyeballing → automated evaluation gate |

| Trust | you approve → policy + human-in-the-loop |

Deployment patterns (pick one, or combine)

  • Client-hosted — the reasoning loop runs in your process. Max control; you own scaling/state.
  • Hosted agent (Foundry Agent Service) — Foundry hosts the loop, stores threads, enforces RBAC/content safety, shows the agent in the portal. Less control, far less operational surface.
  • Agent workflow — multiple agents/tools composed into a graph with branching, approval nodes, and durable checkpoints.

Lifecycle (the loop that ships an agent)

create → version → evaluate (gate) → deploy hosted → observe online → collect failures → repeat.

Offline evaluation is a gate, not an afterthought — a version does not ship

unless it clears the threshold. Online observability feeds real failures back

into the offline test set.

Scaling and cost levers (in priority order)

  • Right-size the model — use the smallest model that passes the evaluation gate.
  • Route by complexity — small/fast model for simple requests, large model for real reasoning (DIY classifier or Foundry Model Router).
  • Cache — serve near-duplicate requests without a model call.
  • Stateless design + bounded concurrency — externalise state; retry with backoff.

Key patterns to reproduce

Point the learner at these from the notebook

16-python-agent-framework.ipynb:

  • Request handler: cache → route by complexity → trace span → run → cache.
  • Evaluation gate: score an offline test set; return pass_rate >= threshold and only deploy if true.
  • Human approval: @tool(approval_mode="always_require") for actions like large refunds.
  • Tracing: wrap each request in tracer.start_as_current_span(...) and set attributes like routed.model, customer.id.

Smoke-testing a deployed agent

After deploy, verify the endpoint actually answers (a green deploy can still be

silent). Use the AI Smoke Test

action via .github/workflows/smoke-test.yml

with the catalog in tests/. The runner POSTs each

prompt to POST {project_endpoint}/agents/{agent_name}/endpoint/protocols/openai/responses

and asserts on the reply text. The identity needs the Azure AI User role at

Foundry project scope; the token audience must be https://ai.azure.com/.

Layer the gates: smoke test (reachable/responding, every deploy) → **offline

evaluation (good enough to ship, before promotion) → online evaluation** (how

is it doing in the wild, continuous).

Enterprise controls

  • RBAC: give each hosted agent a managed identity with least privilege.
  • MCP in production: treat every MCP server as an untrusted boundary — pin the version, scope its identity, validate outputs, rate-limit, never expose secrets.

Guardrails for the assistant

  • Prefer the canonical FoundryChatClient(...) + provider.as_agent(...) pattern used across the course.
  • Do not promise live-Azure results you have not verified; recommend the smoke-test workflow to confirm a deployment.
  • Keep evaluation and cost advice tied together: evaluation sets the quality floor, routing/caching keep cost near that floor.

Other skills for the same job

different authors, same section of the catalogue
Lamindb
by christophacham
×3

This skill should be used when working with LaminDB, an open-source data framework for biology that makes data queryable, traceable, reproducible, and FAIR. Use when managing biological datasets (scRNA-seq, spatial, flow cytometry, etc.), tracking computational workflows, curating and validating data with biological ontologies, building data lakehouses, or ensuring data lineage and reproducibility in biological research. Covers data management, annotation, ontologies (genes, cell types, diseases, tissues), schema validation, integrations with workflow managers (Nextflow, Snakemake) and MLOps platforms (W&B, MLflow), and deployment strategies.

22k tokens
Latchbio Integration
by christophacham
×3

Latch platform for bioinformatics workflows. Build pipelines with Latch SDK, @workflow/@task decorators, deploy serverless workflows, LatchFile/LatchDir, Nextflow/Snakemake integration.

12k tokens
Modal
by christophacham
×3

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

17k tokens
Pyhealth
by christophacham
×3

Comprehensive healthcare AI toolkit for developing, testing, and deploying machine learning models with clinical data. This skill should be used when working with electronic health records (EHR), clinical prediction tasks (mortality, readmission, drug recommendation), medical coding systems (ICD, NDC, ATC), physiological signals (EEG, ECG), healthcare datasets (MIMIC-III/IV, eICU, OMOP), or implementing deep learning models for healthcare applications (RETAIN, SafeDrug, Transformer, GNN).

22k tokens
Github Workflow Automation
by ComeOnOliver
×3

Advanced GitHub Actions workflow automation with AI swarm coordination, intelligent CI/CD pipelines, and comprehensive repository management

9k tokens
Lamindb
by ComeOnOliver
×3

This skill should be used when working with LaminDB, an open-source data framework for biology that makes data queryable, traceable, reproducible, and FAIR. Use when managing biological datasets (scRNA-seq, spatial, flow cytometry, etc.), tracking computational workflows, curating and validating data with biological ontologies, building data lakehouses, or ensuring data lineage and reproducibility in biological research. Covers data management, annotation, ontologies (genes, cell types, diseases, tissues), schema validation, integrations with workflow managers (Nextflow, Snakemake) and MLOps platforms (W&B, MLflow), and deployment strategies.

41k tokens
Ml Pipeline Workflow
by ComeOnOliver
×3

Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment. Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows.

5k tokens
Pyhealth
by ComeOnOliver
×3

Comprehensive healthcare AI toolkit for developing, testing, and deploying machine learning models with clinical data. This skill should be used when working with electronic health records (EHR), clinical prediction tasks (mortality, readmission, drug recommendation), medical coding systems (ICD, NDC, ATC), physiological signals (EEG, ECG), healthcare datasets (MIMIC-III/IV, eICU, OMOP), or implementing deep learning models for healthcare applications (RETAIN, SafeDrug, Transformer, GNN).

39k tokens

How to use it

Copy the folder

Take microsoft/deploying-scalable-agents from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.