Use when working on the Evaluator plugin CLI, jobs, SDK-backed specs, metric types, or plugin-owned Evaluator skills.
npx skills add https://github.com/NVIDIA/skills --skill nemo-evaluator-plugin
Use this skill for evaluation tasks against a running NeMo Platform server. The plugin-backed CLI interface is nemo evaluator; the legacy generated nemo evaluation API command group is not the target surface for new guidance.
nemo CLI: source .venv/bin/activateCheck plugin status from the CLI:
nemo evaluator info
To view available metric names, run:
nemo evaluator metric-types
To view a specific metric schema, pass a metric name from the metric_types list above:
nemo evaluator metric-types <metric-name>
Inspect all the registered metric schema contracts:
nemo evaluator evaluate explain
> Note: use nemo evaluator evaluate explain as the source of truth for the current plugin input schema. It will return a large json schema response, so strongly prefer nemo evaluator metric-types when you only need metric names and corresponding schemas.
Evaluation spec is a payload that is provided to CLI as an input to execute evaluation.
At a high level, a spec describes:
metrics: bundled Evaluator SDK metric configurationsdataset: inline rows to evaluate or platform FilesetRef that contains the datasetparams: optional Evaluator SDK execution parameterstarget: optional model or agent target for online evaluationSee the LLM-judge spec example at assets/specs/llm_as_judge.json.
The checked-in spec examples use bundled SDK metrics. The fields under metrics[*].payload are generated by bundle_metric(metric, CloudpickleMetricBundlePackager()).
To see the pattern for configuring a pre-defined SDK metric, for example ExactMatchMetric, and converting it into bundled metric JSON, inspect build_metric_bundle_example() in generate_example_specs.py and run:
uv run --frozen python skills/nemo-evaluator-plugin/scripts/generate_example_specs.py
When using the nemo evaluator evaluate run command, results are saved into local temporary directories and the link is printed to stdout.
Prefer the --spec-file named argument over inline shell JSON because metric bundles include serialized payloads.
Examples of various specs are provided in the assets/specs directory.
exact-match metricSee the spec example at assets/specs/exact_match_metric.json.
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_benchmark.json
LLM-Judge metricUses an LLM to score responses. See the spec example at assets/specs/llm_as_judge.json.
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/llm_as_judge.json
Use the nemo evaluator evaluate submit command to create a durable evaluation job. The response of this command returns a job handler object instead of the evaluation result.
nemo evaluator evaluate submit \
--spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
The submit response includes the generated job's name field, for example nemo-evaluator-zlhn1ecd. Wait for the job to complete, then list and download the job results.
nemo jobs get-status <job-name>
nemo jobs get <job-name>
nemo jobs results list <job-name>
nemo jobs results download aggregate-scores --job <job-name> --output-file aggregate-scores.json
nemo jobs results download row-scores --job <job-name> --output-file row-scores.jsonl
Evaluator Python SDK client is exposed as evaluator variable on NeMoPlatform instance:
from nemo_platform import NeMoPlatform
platform_client = NeMoPlatform(base_url="http://localhost:8080")
status = platform_client.evaluator.plugin_status()
See examples of using the plugin SDK interface in plugin_sdk_examples.py.
Make sure not to print any secrets to stdout since this can be collected as logs
For LLM-judge setup notes, see LLM Judge Notes.
For evaluator API key auth, see Evaluator API Auth.
For local and cluster troubleshooting, see Evaluation Troubleshooting.
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.
React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.
Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling
Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification
Take nvidia/nemo-evaluator-plugin from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.