mcpbeat Sign in

Dynamo Recipe Runner Agent Skill

Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes. Use for model/backend/GPU/deployment-mode recipe bring-up; use router-starter for router-only mode work and troubleshoot for broken deployments.

9k tokens
context cost
the whole folder, loaded on every use
7
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
2778
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/NVIDIA/skills --skill dynamo-recipe-runner

What comes with it

28 032 bytes besides the instruction
BENCHMARK.md
evals/evals.json
references/k8s-recipe-workflow.md
scripts/recipe_tool.py
skill-card.md
skill.oms.sig

The instruction itself

18 sections, as written by the author

Dynamo Recipe Runner

<!--

SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.

SPDX-License-Identifier: CC-BY-4.0

-->

Purpose

Get from user intent to a working Dynamo recipe endpoint with minimal back and

forth. Do not create new guide content. Operate on the existing recipes/

tree, patch the smallest necessary set of manifests, deploy when the user has

cluster access, and prove success with an OpenAI-compatible smoke request.

Prerequisites

  • Python 3.10+ on the operator machine.
  • kubectl configured with a working cluster context.
  • Cluster has a default storage class for model-cache PVCs.
  • Hugging Face token stored in a Kubernetes secret named hf-token-secret

(or equivalent) in the target namespace.

  • Read access to the recipes/ tree in the ai-dynamo/dynamo repository.

Required Inputs

Collect or infer these before changing manifests:

  • recipe target: model, framework (vllm, sglang, trtllm, tokenspeed), deployment mode, and GPU type/count
  • Kubernetes context and namespace
  • Hugging Face secret name, usually hf-token-secret
  • storage class for model cache PVCs
  • runtime image tag if the recipe uses a placeholder or stale test image
  • whether to run commands or only produce exact commands

If a required value is missing and cannot be inferred from the selected recipe,

ask for only that value.

Instructions

1. Preflight

Run read-only checks first:

git status --short
python3 scripts/recipe_tool.py list --format table
kubectl config current-context
kubectl get storageclass
kubectl get nodes -o wide
kubectl get namespace "${NAMESPACE}"
kubectl get secret hf-token-secret -n "${NAMESPACE}"

If kubectl is unavailable or the cluster is unreachable, continue by

selecting and validating the recipe, then return exact commands instead of

pretending the deployment ran.

2. Select The Recipe

Use the recipe matrix from recipes/README.md and the scanner:

python3 scripts/recipe_tool.py list \
  --query qwen --framework vllm --mode disagg --format table

Prefer an exact existing recipe. Do not invent new manifests unless the user

explicitly asks to author a new recipe.

3. Inspect And Validate

Read the selected recipe README, model-cache manifests, deploy.yaml, and

perf.yaml if present. Then run:

python3 scripts/recipe_tool.py validate \
  recipes/<model>/<framework>/<mode>

Resolve reported blockers before applying manifests: storage class, model cache

PVC, image tag, HF token secret, GPU count, frontend service name, and router

mode.

4. Patch Minimal Values

Patch only recipe-specific values needed for this run. Do not reformat whole

YAML files. Common patches:

  • storageClassName
  • image repository/tag
  • model path or model cache mount path
  • GPU resource requests/limits
  • frontend DYN_ROUTER_MODE
  • namespace only when a manifest hardcodes it

Never write Hugging Face tokens into files or logs. Use Kubernetes secrets.

5. Deploy

Follow the selected recipe README when it differs from the default sequence.

The default sequence is:

kubectl apply -f recipes/<model>/model-cache/ -n "${NAMESPACE}"
kubectl wait --for=condition=Complete job/model-download -n "${NAMESPACE}" --timeout=6000s
kubectl apply -f recipes/<model>/<framework>/<mode>/deploy.yaml -n "${NAMESPACE}"
kubectl get dynamographdeployment -n "${NAMESPACE}"
kubectl get pods -n "${NAMESPACE}" -o wide

Wait for the frontend and workers to be ready before testing.

6. Smoke Test

Port-forward the frontend service, then verify /v1/models and one chat

completion:

kubectl port-forward svc/<deployment-name>-frontend 8000:8000 -n "${NAMESPACE}"
curl http://127.0.0.1:8000/v1/models

If dynamo-router-starter is also installed, prefer its scripts/check_router_health.py

for the full OpenAI-compatible smoke test. If this fails, switch to

dynamo-troubleshoot.

Available Scripts

| Script | Purpose | Arguments |

|---|---|---|

| scripts/recipe_tool.py list | Enumerate available recipes, optionally filtered | --query, --framework, --mode, --format |

| scripts/recipe_tool.py validate | Validate a recipe directory before apply | positional recipe path |

Invoke via the agentskills.io run_script() protocol:

run_script("scripts/recipe_tool.py", args=["list", "--framework", "sglang", "--format", "table"])
run_script("scripts/recipe_tool.py", args=["validate", "recipes/nemotron-3-super-fp8/sglang/agg"])

Examples

List sglang recipes that fit a single 8xB200 node:

python3 scripts/recipe_tool.py list --framework sglang --format table

Validate a specific recipe and resolve blockers before applying:

python3 scripts/recipe_tool.py validate recipes/nemotron-3-super-fp8/sglang/agg

Equivalent through the agent protocol:

run_script("scripts/recipe_tool.py", args=["validate", "recipes/nemotron-3-super-fp8/sglang/agg"])

Output Contract

Return:

  • selected recipe path and why it was selected
  • exact values patched
  • commands run or commands to run
  • endpoint and smoke-test result
  • unresolved blockers, if any
  • next troubleshooting step when deployment does not become healthy

Limitations

  • Operates on the existing recipes/ tree only. Does not author new manifests.
  • Cluster-mutating apply steps require kubectl permission to the target namespace.
  • Smoke-test depth is intentionally minimal; for full router/endpoint coverage use dynamo-router-starter.
  • Multi-node disagg transport correctness is out of scope; use dynamo-interconnect-check after deploy.

Troubleshooting

| Symptom | Likely cause | Next step |

|---|---|---|

| kubectl cluster unreachable | Context not set or VPN down | Return exact commands instead of running them; resume when cluster is reachable |

| validate reports missing storage class | Cluster has no default StorageClass | Patch storageClassName on the model-cache manifest before applying |

| Model-cache job stuck Pending | PVC unbound or HF secret missing | Inspect PVC events; create or rename the HF secret to match the recipe |

| Worker pods ImagePullBackOff | Stale image tag or missing pull secret | Patch the image tag; verify image pull secret in the namespace |

| /v1/models 4xx/5xx after deploy | Frontend not ready or wrong service port | Wait for pods Ready; re-run port-forward; switch to dynamo-troubleshoot if it persists |

Benchmark

See BENCHMARK.md for the NVCARPS-EVAL performance report (auto-generated by the NVSkills CI pipeline). To refresh, re-run /nvskills-ci on an upstream PR touching this skill.

References

  • Read references/k8s-recipe-workflow.md for command templates and readiness checks.
  • Use scripts/recipe_tool.py for recipe discovery and lightweight validation.

Other skills for the same job

different authors, same section of the catalogue
Azure Kubernetes Automatic Readiness
by microsoft
vendor ×3

Assess Kubernetes workloads and cluster configuration for AKS Automatic compatibility. Identifies incompatibilities, generates fixes, and guides migration from AKS Standard to AKS Automatic. WHEN: migrate to AKS Automatic, check AKS Automatic readiness, validate manifests for Automatic, assess cluster for Automatic compatibility, fix deployment for Automatic compatibility, identify AKS Automatic migration blockers, is my cluster ready for AKS Automatic.

13k tokens
Capacity
by microsoft
vendor ×3

Discovers available Azure OpenAI model capacity across regions and projects. Analyzes quota limits, compares availability, and recommends optimal deployment locations based on capacity requirements. USE FOR: find capacity, check quota, where can I deploy, capacity discovery, best region for capacity, multi-project capacity search, quota analysis, model availability, region comparison, check TPM availability. DO NOT USE FOR: actual deployment (hand off to preset or customize after discovery), quota increase requests (direct user to Azure Portal), listing existing deployments.

6k tokens scripts
Customize
by microsoft
vendor ×3

Interactive guided deployment flow for Azure OpenAI models with full customization control. Step-by-step selection of model version, SKU (GlobalStandard/Standard/ProvisionedManaged), capacity, RAI policy (content filter), and advanced options (dynamic quota, priority processing, spillover). USE FOR: custom deployment, customize model deployment, choose version, select SKU, set capacity, configure content filter, RAI policy, deployment options, detailed deployment, advanced deployment, PTU deployment, provisioned throughput. DO NOT USE FOR: quick deployment to optimal region (use preset).

8k tokens
Deploy Model
by microsoft
vendor ×3

Unified Azure OpenAI model deployment skill with intelligent intent-based routing. Handles quick preset deployments, fully customized deployments (version/SKU/capacity/RAI policy), and capacity discovery across regions and projects. USE FOR: deploy model, deploy gpt, create deployment, model deployment, deploy openai model, set up model, provision model, find capacity, check model availability, where can I deploy, best region for model, capacity analysis. DO NOT USE FOR: listing existing deployments (use foundry_models_deployments_list MCP tool), deleting deployments, agent creation (use agent/create), project creation (use project/create).

26k tokens scripts
Preset
by microsoft
vendor ×3

Intelligently deploys Azure OpenAI models to optimal regions by analyzing capacity across all available regions. Automatically checks current region first and shows alternatives if needed. USE FOR: quick deployment, optimal region, best region, automatic region selection, fast setup, multi-region capacity check, high availability deployment, deploy to best location. DO NOT USE FOR: custom SKU selection (use customize), specific version selection (use customize), custom capacity configuration (use customize), PTU deployments (use customize).

9k tokens
Lamindb
by christophacham
×3

This skill should be used when working with LaminDB, an open-source data framework for biology that makes data queryable, traceable, reproducible, and FAIR. Use when managing biological datasets (scRNA-seq, spatial, flow cytometry, etc.), tracking computational workflows, curating and validating data with biological ontologies, building data lakehouses, or ensuring data lineage and reproducibility in biological research. Covers data management, annotation, ontologies (genes, cell types, diseases, tissues), schema validation, integrations with workflow managers (Nextflow, Snakemake) and MLOps platforms (W&B, MLflow), and deployment strategies.

22k tokens
Latchbio Integration
by christophacham
×3

Latch platform for bioinformatics workflows. Build pipelines with Latch SDK, @workflow/@task decorators, deploy serverless workflows, LatchFile/LatchDir, Nextflow/Snakemake integration.

12k tokens
Modal
by christophacham
×3

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

17k tokens

How to use it

Copy the folder

Take nvidia/dynamo-recipe-runner from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.