> Use this skill when the user is invoking `doca_telemetry_utils` on a host with DOCA installed — discovering the diagnostic-counter schema, translating counter names to binary Data IDs, validating per-device counter support before committing a DOCA Telemetry exporter config, or reverse-resolving a captured Data ID. Trigger even when the user does not explicitly mention "doca_telemetry_utils" or "Data ID" — typical implicit phrasings include "my exporter ships but the collector sees nothing", "this metric silently drops downstream", "which counters does this BlueField expose", "translate this 0x... back to a counter name", "what do node / pcie_index / depth mean here", or "is this counter supported on this device before I commit it". Refuse and route elsewhere for developer-side collector / exporter library programming, DTS deployment, or DOCA install / repair — those belong to doca-telemetry, doca-public-knowledge-map, and doca-setup.
npx skills add https://github.com/NVIDIA/skills --skill doca-telemetry-utils
Where to start: This is a tool skill for invoking
doca_telemetry_utils — the documented host-side CLI that
supports a DOCA Telemetry exporter / collector pipeline by
discovering the counter schema, translating counter names ↔
Data IDs, and probing per-device counter support. Open
TASKS.md and start at
## install for the host-side
prerequisites and ## run for the
three documented invocation classes
(enumerate / name→ID / ID→name). Open
CAPABILITIES.md when the question is
*what does this tool actually discover about the telemetry
schema*, *how does it pair with the developer-side
doca-telemetry and
exporter libraries*, *how do I confirm a device supports a
counter before committing an exporter config to it*, or
*why does my exporter pipeline silently drop a metric*.
This skill is the operator-side support tool for a
DOCA Telemetry deployment. It is NOT the developer-side
collector library (that is
doca-telemetry),
NOT the developer-side publisher library (that is
doca-telemetry-exporter — see
doca-telemetry ## Related skills),
and NOT a DOCA Telemetry Service (DTS) deployment guide
(route via
doca-public-knowledge-map).
Three separate surfaces; conflating them is the most
common telemetry first-touch error.
The CLASSES of doca_telemetry_utils questions this skill
is built to answer, each with one worked example. The
class is the load-bearing piece; the worked example is one
instance.
port_rx_bytes butnothing shows up downstream — what did I get wrong?"** —
worked example: *"my exporter config has a counter name
string and my collector sees no events with that
name"*. Answered by the name ↔ Data ID translation
step in
CAPABILITIES.md ## Capabilities and modes
+ the per-device-support probe in
TASKS.md ## test: the exporter
ships a Data ID, not a name; a name in the config that
resolves to a Data ID the device does not support is
silently dropped.
actually expose?"** — worked example: *"enumerate the
full counter schema for my BlueField-3 before I write
the exporter config"*. Answered by the schema-discovery
invocation class in
CAPABILITIES.md ## Capabilities and modes
(doca_telemetry_utils get-counters lists every
counter name the diagnostic-data surface knows about;
pair with a per-device probe to confirm support).
was that?"** — worked example: *"a downstream consumer
emitted Data ID: 0x1160000600030201 — translate it
back so I can correlate against the public guide"*.
Answered by the reverse-resolve invocation class in
CAPABILITIES.md ## Capabilities and modes
+ the
ID-encodes-properties rule in
TASKS.md ## use (a Data ID carries
the counter's property dimensions; the reverse-resolve
reports them).
what values are valid?"** — worked example: *"I know
the counter is pcie_link_write_stalled_time_* — what
do node / pcie_index / depth mean and what
values does the device accept?"*. Answered by the
property-dimension table in
CAPABILITIES.md ## Capabilities and modes
+ the <name> invocation without arguments which
prints the documented property options + units +
unit-specific axes.
commit it to the exporter config?"** — worked example:
*"validate that port_rx_bytes with node=1 is
exposed on the BlueField at PCIe address X before I
write the exporter config"*. Answered by the
per-device-support probe (`<device PCI> <name>
[properties]`) in
CAPABILITIES.md ## Capabilities and modes
+ the gate-before-commit rule in
TASKS.md ## use.
doca_telemetry_utils on my install, and is itpaired with the matching doca-telemetry library
version?"** — worked example: *"is the diagnostic-
data counter set my exporter targets on this DOCA
version?"*. Answered by the version-overlay in
CAPABILITIES.md ## Version compatibility,
which redirects to the canonical
doca-version chain
and adds the *tool ↔ doca-telemetry library
schema-version* match rule.
This skill serves **external operators, developers, and
AI agents standing up or debugging a DOCA Telemetry
exporter / collector pipeline who need to confirm the
counter schema, validate per-device support, or
translate between human-readable counter names and the
binary Data IDs the exporter actually ships**.
Concretely:
exporter on a BlueField fleet who needs to confirm
which counters the target devices actually expose
before committing the exporter config.
app linking doca-telemetry,
or a third-party aggregator consuming via DTS) who
has a captured Data ID stream and needs to translate
IDs back to counter names + properties.
downstream"* / *"this metric is silently missing"*
report against a deployed exporter — the
schema-discovery + per-device-support probes are the
canonical *"is the counter even supposed to work on
this device"* first step before suspecting the
collector or the network path.
for this BlueField + this DOCA version"* answer
honestly — with each counter resolved to its Data ID
and each Data ID confirmed against the per-device
capability probe.
It is not for users debugging the doca_telemetry_utils
binary itself, not a substitute for the live public
DOCA Telemetry guides, and not the right place for
learning how to *write* an exporter application (that
audience belongs in
doca-telemetry-exporter skill
when present, or via
doca-public-knowledge-map).
The tool is shipped as a CLI binary under
/opt/mellanox/doca/tools/, not a library you link
against. The skill uses the same kind: tool
three-file shape as the rest of the bundle so the
agent's task-verb contract is uniform across libraries,
services, and tools.
doca_telemetry_utils is a C host-side CLI that uses
the documented doca_telemetry_diag surface to list
counter types and probe per-device support. Its inputs
are command-line arguments (counter name, property
values, Data ID, optional device PCI address); its
outputs are human-readable Data ID + properties
mappings. The skill keeps workflow guidance language-
neutral; downstream consumers in any language can use
the resolved Data IDs against their own collector code.
Load this skill when the user is — or the agent needs
to — invoke doca_telemetry_utils on a real host with
DOCA installed, alongside an exporter / collector
pipeline that needs schema discovery, per-device
support validation, or Data ID translation.
Concretely:
a fresh DOCA Telemetry exporter config (get-counters).
resolves to a Data ID the target device actually
supports, before committing the config.
downstream consumer back to a counter name +
properties for correlation against the public guide.
nothing"* report (the canonical schema-mismatch /
unsupported-counter failure mode).
BlueField + this DOCA version"* artifact as part of
a structured exporter-config baseline.
and confirming each previously-supported counter
still resolves cleanly on the new version.
Do not load this skill for general DOCA orientation,
collector / exporter library programming, DTS
deployment, or DOCA install. For those, route to
doca-public-knowledge-map,
doca-telemetry,
or doca-setup.
This is a thin loader. Substantive material lives
in two companion files:
CAPABILITIES.md — what doca_telemetry_utilsdiscovers + how it pairs with the developer-side
surfaces: the three documented invocation classes
(enumerate / name→ID / ID→name), the property-
dimension model (counters carry property axes such
as node, pcie_index, depth, plus per-unit
axes), the optional per-device capability probe
(<device PCI> <name> runs the resolved counter
against the device), the operator-side support
role (this is NOT the developer-side library; it
exists to make exporter / collector setups
honest), the version overlay (tool ↔
doca-telemetry library schema version pairing),
the layered error taxonomy (install / parse /
unknown-counter / unknown-data-id / property-out-
of-range / device-not-supported / version /
cross-cutting), the observability surface
(stdout-only), and the safety policy (the tool is
read-only; mistakes appear downstream as silent
metric drops, not crashes).
TASKS.md — step-by-step workflows for thein-scope task verbs: install (host-side DOCA +
telemetry component prerequisites), configure
(axis decisions: which invocation class, which
device, which counter), build (route to install
— the binary is shipped), modify (refuse — do
not patch the binary; modify the invocation and
the exporter / collector config that consumes the
resolved Data IDs), run (the three documented
invocations), test (round-trip a chosen
counter through name → Data ID → per-device
probe → exporter config → collector receipt),
debug (walk the error taxonomy), use (the
hand-off into the developer-side doca-telemetry
collector / exporter pipeline), plus a `Deferred
task verbs` block.
The skill assumes a host where DOCA is already
installed with the telemetry component, a BlueField
the operator is targeting is visible to DOCA, and the
exporter / collector pipeline that consumes the
resolved Data IDs is either being authored or has
already been authored against the matching
doca-telemetry
library version.
This skill is agent guidance, not a samples or
scripts bundle. To keep the boundary clean, it
deliberately does not contain — and pull requests
should not add:
table.** The counter set evolves across DOCA
versions and BlueField generations; an
inventory pinned in this skill would silently rot.
doca_telemetry_utils get-counters on the
installed binary is the authoritative source.
are install-, device-, and use-case-specific; a
packaged config in this skill would mislead
operators on a different setup.
service with its own public guide; route via
doca-public-knowledge-map.
This skill *resolves the counter names + IDs* a
DTS pipeline consumes; it does not configure DTS.
that consume the tool's output. The output is a
simple text mapping; users who want to script
against it should read the live guide and write
the parser against their installed version.
samples/ or reference/ subtree. This isa thin loader for a shipped CLI; substantive
material lives on the public page, in --help,
and in doca-telemetry.
SKILL.md first to confirm the user'squestion is in scope (operator-side support
tooling for a telemetry pipeline, not the
developer-side library programming).
property-dimension model, the per-device support
probe, the version overlay, the error taxonomy,
the observability surface, and the safety
policy, see CAPABILITIES.md.**
discover → resolve → validate → consume
workflow — install, configure, build,
modify, run, test, debug, use — see
TASKS.md.**
doca-telemetry— the developer-side collector library whose
schema this tool helps the operator discover. Pair
them in every exporter-pipeline triage session.
The collector library skill teaches the
collector-vs-exporter rule, the schema-must-match
contract with the publisher, and the consumer-
queue-full back-pressure rule; this tool's role
is to make the *schema* half of that contract
inspectable from the operator side. Conflating
the library and the tool is the most common
telemetry first-touch error.
doca-public-knowledge-map— routing to the public DOCA Telemetry guide and
the DOCA Telemetry Service (DTS) page on
docs.nvidia.com, plus the on-disk install
layout for the tool.
doca-version —canonical DOCA version-handling rules. The
## Version compatibility section in
CAPABILITIES.md is a concise
overlay that redirects here for the body and
adds the *tool ↔ doca-telemetry library
schema-version pairing* rule.
doca-setup — envpreparation, install verification, and the *I
have no install yet* path with the public NGC
DOCA container. This skill assumes its
preconditions are satisfied (DOCA installed
with the telemetry component, BlueField visible
to DOCA).
doca-debug — thecross-cutting debug ladder. Telemetry-utils
feeds the cross-cutting ladder by surfacing the
counter schema + per-device support truth at the
runtime layer; exporter-pipeline regressions
often resolve at the schema layer before
touching the collector / network path.
doca-structured-tools-contract— the bundle's detect → prefer → fall back →
report contract for structured helper tools. The
command appendix in TASKS.md
honors this contract.
doca-programming-guide— general DOCA programming patterns shared by
every library / tool surface, including the
cross-library DOCA_ERROR_* taxonomy this
tool's host-side error layer overlays on top of
when downstream collector / exporter code
surfaces a related error.
Assess Kubernetes workloads and cluster configuration for AKS Automatic compatibility. Identifies incompatibilities, generates fixes, and guides migration from AKS Standard to AKS Automatic. WHEN: migrate to AKS Automatic, check AKS Automatic readiness, validate manifests for Automatic, assess cluster for Automatic compatibility, fix deployment for Automatic compatibility, identify AKS Automatic migration blockers, is my cluster ready for AKS Automatic.
Discovers available Azure OpenAI model capacity across regions and projects. Analyzes quota limits, compares availability, and recommends optimal deployment locations based on capacity requirements. USE FOR: find capacity, check quota, where can I deploy, capacity discovery, best region for capacity, multi-project capacity search, quota analysis, model availability, region comparison, check TPM availability. DO NOT USE FOR: actual deployment (hand off to preset or customize after discovery), quota increase requests (direct user to Azure Portal), listing existing deployments.
Interactive guided deployment flow for Azure OpenAI models with full customization control. Step-by-step selection of model version, SKU (GlobalStandard/Standard/ProvisionedManaged), capacity, RAI policy (content filter), and advanced options (dynamic quota, priority processing, spillover). USE FOR: custom deployment, customize model deployment, choose version, select SKU, set capacity, configure content filter, RAI policy, deployment options, detailed deployment, advanced deployment, PTU deployment, provisioned throughput. DO NOT USE FOR: quick deployment to optimal region (use preset).
Unified Azure OpenAI model deployment skill with intelligent intent-based routing. Handles quick preset deployments, fully customized deployments (version/SKU/capacity/RAI policy), and capacity discovery across regions and projects. USE FOR: deploy model, deploy gpt, create deployment, model deployment, deploy openai model, set up model, provision model, find capacity, check model availability, where can I deploy, best region for model, capacity analysis. DO NOT USE FOR: listing existing deployments (use foundry_models_deployments_list MCP tool), deleting deployments, agent creation (use agent/create), project creation (use project/create).
Intelligently deploys Azure OpenAI models to optimal regions by analyzing capacity across all available regions. Automatically checks current region first and shows alternatives if needed. USE FOR: quick deployment, optimal region, best region, automatic region selection, fast setup, multi-region capacity check, high availability deployment, deploy to best location. DO NOT USE FOR: custom SKU selection (use customize), specific version selection (use customize), custom capacity configuration (use customize), PTU deployments (use customize).
This skill should be used when working with LaminDB, an open-source data framework for biology that makes data queryable, traceable, reproducible, and FAIR. Use when managing biological datasets (scRNA-seq, spatial, flow cytometry, etc.), tracking computational workflows, curating and validating data with biological ontologies, building data lakehouses, or ensuring data lineage and reproducibility in biological research. Covers data management, annotation, ontologies (genes, cell types, diseases, tissues), schema validation, integrations with workflow managers (Nextflow, Snakemake) and MLOps platforms (W&B, MLflow), and deployment strategies.
Latch platform for bioinformatics workflows. Build pipelines with Latch SDK, @workflow/@task decorators, deploy serverless workflows, LatchFile/LatchDir, Nextflow/Snakemake integration.
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
Take nvidia/doca-telemetry-utils from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.