nvidia/nemo-evaluator-plugin
Use when working on the Evaluator plugin CLI, jobs, SDK-backed specs, metric types, or plugin-owned Evaluator skills.
npx skills add https://github.com/NVIDIA/skills --skill nemo-evaluator-plugin
Use this skill for evaluation tasks against a running NeMo Platform server. The plugin-backed CLI interface is nemo evaluator; the legacy generated nemo evaluation API command group is not the target surface for new guidance.
nemo CLI: source .venv/bin/activateCheck plugin status from the CLI:
nemo evaluator info
To view available metric names, run:
nemo evaluator metric-types
To view a specific metric schema, pass a metric name from the metric_types list above:
nemo evaluator metric-types <metric-name>
Inspect all the registered metric schema contracts:
nemo evaluator evaluate explain
> Note: use nemo evaluator evaluate explain as the source of truth for the current plugin input schema. It will return a large json schema response, so strongly prefer nemo evaluator metric-types when you only need metric names and corresponding schemas.
Evaluation spec is a payload that is provided to CLI as an input to execute evaluation.
At a high level, a spec describes:
metrics: bundled Evaluator SDK metric configurationsdataset: inline rows to evaluate or platform FilesetRef that contains the datasetparams: optional Evaluator SDK execution parameterstarget: optional model or agent target for online evaluationSee the LLM-judge spec example at assets/specs/llm_as_judge.json.
The checked-in spec examples use bundled SDK metrics. The fields under metrics[*].payload are generated by bundle_metric(metric, CloudpickleMetricBundlePackager()).
To see the pattern for configuring a pre-defined SDK metric, for example ExactMatchMetric, and converting it into bundled metric JSON, inspect build_metric_bundle_example() in generate_example_specs.py and run:
uv run --frozen python skills/nemo-evaluator-plugin/scripts/generate_example_specs.py
When using the nemo evaluator evaluate run command, results are saved into local temporary directories and the link is printed to stdout.
Prefer the --spec-file named argument over inline shell JSON because metric bundles include serialized payloads.
Examples of various specs are provided in the assets/specs directory.
exact-match metricSee the spec example at assets/specs/exact_match_metric.json.
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_benchmark.json
LLM-Judge metricUses an LLM to score responses. See the spec example at assets/specs/llm_as_judge.json.
nemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/llm_as_judge.json
Use the nemo evaluator evaluate submit command to create a durable evaluation job. The response of this command returns a job handler object instead of the evaluation result.
nemo evaluator evaluate submit \
--spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.json
The submit response includes the generated job's name field, for example nemo-evaluator-zlhn1ecd. Wait for the job to complete, then list and download the job results.
nemo jobs get-status <job-name>
nemo jobs get <job-name>
nemo jobs results list <job-name>
nemo jobs results download aggregate-scores --job <job-name> --output-file aggregate-scores.json
nemo jobs results download row-scores --job <job-name> --output-file row-scores.jsonl
Evaluator Python SDK client is exposed as evaluator variable on NeMoPlatform instance:
from nemo_platform import NeMoPlatform
platform_client = NeMoPlatform(base_url="http://localhost:8080")
status = platform_client.evaluator.plugin_status()
See examples of using the plugin SDK interface in plugin_sdk_examples.py.
Make sure not to print any secrets to stdout since this can be collected as logs
For LLM-judge setup notes, see LLM Judge Notes.
For evaluator API key auth, see Evaluator API Auth.
For local and cluster troubleshooting, see Evaluation Troubleshooting.
Take nvidia/nemo-evaluator-plugin from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.