> Run `doca_bench` (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm. Use it to discover enabled benchmark libraries, capture a reproducible command/version/device/environment baseline, compare stable runs against a declared tolerance, or diagnose configuration, device-binding, workload-precondition, and measurement failures. Trigger for requests such as measuring BlueField compression speed, NIC RDMA throughput, crypto latency, or a pre-upgrade baseline. Do not use for application end-to-end timing, custom benchmark code, DOCA installation, or binary patches.
npx skills add https://github.com/NVIDIA/skills --skill doca-bench
doca_bench)Where to start: This is a tool skill for invoking doca_bench,
the cross-library micro-benchmark harness. Open
TASKS.md and start at
## configure for the three-axis decision
(target library × workload shape × measurement axis), then
## run for the smoke-before-bulk flow. Open
CAPABILITIES.md when the question is *what
doca_bench can measure*, *which DOCA libraries it can drive*, or
*how to interpret throughput / latency / op-rate output without
fooling yourself on warm-up or steady-state*. If DOCA is not
installed yet, route to
doca-setup first; if the install
version is < 2.7.0, doca_bench is not shipped on this host.
The CLASSES of doca_bench questions this skill is built to answer,
each with one worked example. The class is the load-bearing piece;
the worked example is one instance.
worked example: *"throughput of DOCA Compress on my BlueField-3"*.
Answered by the three-axis configuration in
CAPABILITIES.md ## Capabilities and modes
+ the smoke-before-bulk flow in
TASKS.md ## run. The *same* shape answers
*"send-side throughput of DOCA RDMA"* — doca_bench is
cross-library, not single-library.
doca_bench actually drive on thisinstall?"** — worked example: *"is doca_sha enumerable on a
granular-build install"*. Answered by the built-in query system
surfaced in
CAPABILITIES.md ## Capabilities and modes
+ TASKS.md ## configure step 2
(probe-before-bench). Empty enumeration = library not installed,
not bench failure.
worked example: *"why does my first-second number differ from my
steady-state number"*. Answered by the measurement-soundness
overlay in
CAPABILITIES.md ## Error taxonomy
layer 5 + TASKS.md ## test (the eval-loop
overlay treats warm-up / steady-state / outliers as
re-iteration triggers, not one-shot facts).
with the public docs."** — worked example: *"doca_bench shows
zero ops for AES-GCM but doca_caps says the device supports
it"*. Answered by the layered error taxonomy in
CAPABILITIES.md ## Error taxonomy
(config-syntax → device-binding → library-precondition →
workload-precondition → measurement-soundness → version →
cross-cutting) + TASKS.md ## debug.
against?"** — worked example: *"snapshot decompress throughput
on this BlueField + DOCA version before a firmware update"*.
Answered by the CSV output + version-overlay rule in
TASKS.md ## test (capture command line +
version + device + as-deployed environment alongside the
numbers; quoting numbers without the four-tuple is the
cross-version regression-hunt failure mode).
doca_bench returns nothing for library X — what does thatmean?"** — worked example: *"empty output for DOCA SHA"*.
Answered by the empty-output interpretation rules in
TASKS.md ## debug +
CAPABILITIES.md ## Error taxonomy.
Re-route through
doca-caps for the coarse
per-device per-library capability ground truth, then back
into bench once the capability is confirmed present.
This skill serves **external operators, developers, and AI agents
who need a reproducible, vendor-supported way to measure DOCA
library performance on the user's actual install and device**.
Concretely:
COMPRESS vs SHA vs DMA throughput) before committing an
application design.
driver upgrade, firmware burn) by re-running a captured
doca_bench baseline against the new state.
device delivers today"* artifact that downstream consumers
(capacity planning, regression bisection) can cite.
expect from DOCA library X on device Y?"* honestly — with a
measured number, the command line that produced it, and the
version + device + environment that scopes it — instead of
guessing from datasheet headlines.
It is not for users debugging the doca_bench source code,
and not a substitute for the live public DOCA Bench guide on
docs.nvidia.com.
doca_bench is shipped as a tool (a single CLI binary plus a
companion app for the remote half of remote-memory / RDMA / Eth
scenarios), not a library you link against. The skill uses the
same kind: tool three-file shape as the rest of the bundle so
the agent's task-verb contract
(configure / build / modify / run / test / debug) is uniform
across libraries, services, and tools — even when individual
verbs collapse to a routing stub for a shipped binary.
Load this skill when the user is — or the agent needs to — invoke
doca_bench on a real host with DOCA ≥ 2.7.0 installed (or
inside the public NGC DOCA container with the equivalent version)
to measure performance of a DOCA library. Concretely:
workload (RDMA vs COMPRESS vs DMA, etc.).
latency vs precision latency vs max-bandwidth) — the four modes
defined in tools/bench/doca_bench/configuration.hpp are not
interchangeable.
honestly report *"this library is not exposed on this install"*
instead of inventing a workload.
+ as-deployed environment + numbers) for later regression hunts.
and obtaining two consecutive runs within that tolerance before
reporting a stable result; otherwise escalating the variance.
results (the error-taxonomy walk in
TASKS.md ## debug).
Do not load this skill for general DOCA orientation, library
API work, or installation. For those, use
doca-public-knowledge-map,
the matching libs/<library> skill, or
doca-setup. Do not load it for
*application-level* end-to-end benchmarking either — doca_bench
measures the DOCA library surface, not the user's application
above it.
This is a thin loader. Substantive material lives in two
companion files:
CAPABILITIES.md — what doca_bench can measure (thecross-library scope, the three-axis configuration model, the
documented operating modes, the warm-up / pipeline / multi-core
concepts that constrain measurement soundness), the version
overlay (doca-bench-specific facts on top of the canonical
doca-version rules), the layered error taxonomy
(config-syntax / device-binding / library-precondition /
workload-precondition / measurement-soundness / version /
cross-cutting), the observability surface (screen + CSV
output, real-time stats, query system), and the safety
posture (the public guide's *"not for production"* warning,
the host vs BlueField execution rule, the companion-app
attack surface).
TASKS.md — step-by-step workflows for the in-scope taskverbs: configure (the three-axis decision + the
probe-before-bench step), build (route to install — the
binary is shipped, the companion app is shipped), modify
(refuse — do not patch the bench binary; modify the bench
*invocation* instead), run (the smoke-before-bulk flow),
test (the eval loop — warm-up, steady-state, outliers,
cross-version), debug (walk the error taxonomy layer by
layer), plus a Deferred task verbs block routing
out-of-scope questions and a Command appendix of
doca_bench-specific invocation classes.
The skill assumes a host where DOCA ≥ 2.7.0 is already installed
(or the public NGC DOCA container is running at an equivalent
version) and the operator has whatever permissions the public
guide requires for doca_bench to bind devices and allocate
resources on their platform.
This skill is agent guidance, not a samples or scripts
bundle. To keep the boundary clean, it deliberately does not
contain — and pull requests should not add:
beyond what the public DOCA Bench guide documents.** The flag
surface evolves and is install-specific; the documented
invocations + --help on the installed version are the
authoritative answer. Inventing a flag is the most common
hallucination failure for this skill.
Bench output is device-, version-, firmware-, NUMA-, and
tuning-specific. A captured number pinned to one platform and
one DOCA version misleads operators on a different
platform / version.
doca_bench CSV or stdout. The output formats are documented;
if a user wants to script against them, the right answer is
"read the live guide, write the parser against your installed
version".
samples/ or reference/ subtree. This is a thinloader for a documented CLI; substantive material lives on
the public page and in --help.
SKILL.md first to confirm the user's question isin scope (the user actually wants to invoke doca_bench for
measurement, not learn about a DOCA library in general).
doca_bench measures, the three-axis model, theversion overlay, the error taxonomy, observability surface,
and safety posture, see CAPABILITIES.md.**
workflow — configure, build, modify, run, test,
debug — see TASKS.md.**
doca-public-knowledge-map— routing to the public DOCA Bench page on docs.nvidia.com
and the rest of the public DOCA documentation set.
doca-version — the canonicalversion-detection chain, four-way match rule, NGC container
semantics, and headers-win-over-docs rule. The
## Version compatibility section in this skill is a thin
overlay on top of doca-version; the body lives there.
doca-structured-tools-contract— the bundle-wide contract for structured-output helper tools.
Bench-runner / bench-snapshot executables that satisfy the
detect-prefer-fallback-report loop are deferred to PR2; the
contract is consumed here in advance so the
## Command appendix in TASKS.md is infra-aware
from PR1.
doca-setup — env preparation,install verification, hugepages, NUMA awareness, and the *I
have no install yet* path with the public NGC DOCA container.
doca-debug — the cross-cuttingdebug ladder. Bench surfaces *its own* error taxonomy in
CAPABILITIES.md ## Error taxonomy;
when the cause turns out to be below DOCA (driver, firmware,
NUMA), the bench taxonomy hands off to doca-debug.
doca-caps — the sibling DOCA toolfor the coarse per-device per-library capability snapshot.
Bench probes capability at finer grain via its own query
system; doca_caps is the cheaper first step to confirm the
device is even visible to DOCA.
libs/<library> skill — e.g.doca-comch,
doca-compress — for
the workload-side preconditions, capability-query rules, and
error-taxonomy overlays of the library under test. Bench
drives the library; the library skill explains *what
"healthy" means for it*.
> Look up Y Combinator companies, batches, and startup ecosystem data using the yc-oss API (read-only). Use this skill whenever the user wants to research YC-backed startups, find companies in a specific batch or industry, check which YC companies are hiring, explore top YC companies, or analyze startup trends by sector or tag. "top Y Combinator companies", "find YC companies tagged AI", "W25 batch", "S24 companies", "YC stats", "Y Combinator portfolio", "startup research", "which YC companies do X", "venture research on YC", any mention of Y Combinator, YC batch, or YC-backed companies in the context of startup research, venture analysis, or market intelligence. This is a read-only data source — the API is a static JSON dataset updated daily.
Shopify/DTC profit margin calculator for sellers. Calculate cost breakdowns including ad spend, CAC, payment processing fees, and 3PL costs. Includes LTV/CAC analysis and DTC-specific benchmarks. No API key required.
创建和更新宣传文稿、一页简介、投资者备忘录、加速器申请、财务模型和融资材料。当用户需要面向投资者的文件、预测、资金用途表、里程碑计划或必须在多个融资资产中保持内部一致性的材料时使用。
Decompose portfolio returns into explainable components to identify where value was added or lost. Use when the user asks about Brinson attribution, allocation vs selection effects, factor-based attribution, fixed-income attribution, or currency attribution. Also trigger when users mention 'what drove my returns', 'was it stock picking or sector bets', 'alpha decomposition', 'multi-period linking', 'interaction effect', 'active return breakdown', or ask why their portfolio outperformed or underperformed the benchmark.
创建和更新宣传文稿、一页简介、投资者备忘录、加速器申请、财务模型和融资材料。当用户需要面向投资者的文件、预测、资金用途表、里程碑计划或必须在多个融资资产中保持内部一致性的材料时使用。
Legal contract analysis using CUAD dataset (41 risk categories). Supports NDA, SaaS, M&A, employment, payment/merchant, and finder/broker agreements. Identifies red flags, suggests redlines, compares to market standards.
Inspect the actual holdings dataset to understand its schema, field names, and data patterns before applying filters. Use this to identify the correct field name that distinguishes equity securities from bonds, options, warrants, and other non-stock instruments.
Retrieve and load the complete holdings dataset for a specific fund using its accession_number. Use this to access position-level details needed for AUM, stock count, and comparative analysis.
Take nvidia/doca-bench from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.