mcpbeat

Config Evals

n8n-io/config-evals

>- Builds and maintains configuration-based evaluations on a workflow with the eval-config tool. Use when the user asks to set up, add, view, change, or remove an evaluation, score, grade, or judge a workflow's output, or measure answer quality against a test dataset. This is the only eval form Instance AI handles — it does not touch on-canvas evaluation nodes.

3k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
199283
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/n8n-io/n8n --skill config-evals

What comes with it

5 095 bytes besides the instruction
references/config-eval-playbook.md

The instruction itself

7 sections, as written by the author

Config-based Evaluations

Use this skill to attach a configuration-based evaluation to a workflow with the

eval-config tool. A config eval pairs a workflow with a name, a start node, an

end node, one or more judged metrics, and a Data Table dataset. Nothing is added

to the canvas — the config lives off-canvas via the evaluation-config API.

Config evals are the only evaluation form you work with. Do not add, read,

rewire, or reason about on-canvas evaluation nodes (EvaluationTrigger,

Evaluation/checkIfEvaluating/setOutputs/setMetrics). If the user asks for those,

build a config eval instead and briefly say that is how you set up evaluations.

What a Config Eval Needs

  • name — a human-readable evaluation name.
  • startNodeName — the node where a run begins; it is fed one test-input row.

Must be a node with an incoming connection — never a trigger (see step 2).

  • endNodeName — the node whose output is judged.
  • dataTableId — a Data Table holding the test dataset. Create and populate it

with the data-tables tool first, then link it here by id.

  • metrics — one or more judged metrics (see below).

Default Procedure

  • Identify the target workflow and read it. Trace the main path from trigger to

the node that produces the answer.

  • Pick the nodes:
  • startNodeName is the first node after the trigger — the node that

receives the input the dataset varies. Never use the trigger itself: an

eval run swaps the trigger for a dataset-driven one, so the start node must

have an incoming connection or the run fails to compile. For a chat/agent

workflow this is usually the agent node (often the same as endNodeName).

  • endNodeName is the node whose output you want scored (usually the AI agent

or the final response node).

  • Resolve the dataset. Call data-tables(action="list") to find an existing

dataset, or create and seed one with data-tables before creating the config.

Never invent a dataTableId; use one returned by data-tables.

  • Choose metrics and build the actualAnswer / expectedAnswer / userQuery

expressions (see Metrics).

  • Call eval-config (action="create"), or update when changing an existing

config. The tool shows an approval card automatically — call it and respect

the result; do not ask for chat approval first.

  • Close with facts: evaluation name, workflow, start/end nodes, dataset name and

id, and the metrics configured.

Metrics

Each metric is LLM-judged and needs a judge model: a credentialId, a model,

and an outputType (numeric, the default, or boolean). Reuse an LLM

credential the workflow already uses when one fits.

Do not set provider unless you know the exact chat-model node type — it is

derived automatically from the credential you pass (each credential type maps to

one provider). Just pick the credential and the model.

Two presets are available:

  • correctness — compares the produced answer to a ground-truth answer.

Requires expectedAnswer (an n8n expression resolving to the ground-truth

value, typically a dataset column, e.g. ={{ $json.expected_output }}).

  • helpfulness — judges the produced answer against the user's query.

Requires userQuery (an n8n expression for the input the user asked, e.g.

={{ $json.input }}).

Every metric also needs actualAnswer: an n8n expression resolving to the

workflow's produced answer at the end node, e.g. ={{ $json.output }}.

userQuery and expectedAnswer name dataset columns (the input the user

asked; the ground-truth answer). actualAnswer names a field of the workflow's

produced output. Write all of them as ={{ $json.<name> }} — the evaluation

reads dataset columns from the dataset row and actualAnswer from the end node

automatically. Do not reference the trigger or any node by name.

Expression fields must begin with =

actualAnswer, userQuery, and expectedAnswer are n8n expressions — they

read a value out of each test row at runtime. The leading = is what tells n8n

to evaluate the {{ … }} template. **Without it the string is stored as literal

text**: the field shows {{ $json.output }} verbatim and the judge scores that

raw string instead of the resolved value.

  • Correct: ={{ $json.output }}, ={{ $json.expected_output }}
  • Wrong: {{ $json.output }} (no = → treated as fixed text)

Only add = when the value references workflow data via {{ … }}. A genuinely

fixed constant (rare for these fields) is written as plain text without =.

Pick correctness when the dataset has a known right answer to compare against;

pick helpfulness when there is no single ground truth and quality is judged

relative to the request. Use prompt only to override the default judge prompt.

Dataset Boundary

  • Build the dataset with the data-tables tool: one column for each input the

evaluation varies, plus a ground-truth column when using correctness.

  • The config only references the dataset by dataTableId; the eval-config tool

does not create or populate rows. If no suitable dataset exists, create one

first, then create the config.

  • Do not weaken the evaluation to fit a thin dataset — seed the dataset to match

the metrics, or ask the user for the expected answers.

More Detail

Use references/config-eval-playbook.md for

tool-call recipes, worked examples, and output shapes.

How to use it

Copy the folder

Take n8n-io/config-evals from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.