get-convex/analyze-eval
Investigate a single failing eval from the convex-evals system. Use when the user shares a visualizer URL pointing to a specific eval, asks about a specific failing eval, or references a specific eval ID.
npx skills add https://github.com/get-convex/convex-evals --skill analyze-eval
https://convex-evals.netlify.app/experiment/.../run/$runId/$category/$evalIdThe visualizer URL pattern is:
/experiment/$experimentId/run/$runId/$category/$evalId?tab=steps
$runId — the Convex document ID for the run (e.g. jn7922j1w29pdxm76bj9ps0enx80mg9e)$evalId — the Convex document ID for the specific eval (e.g. jh73jvjz2n00gfeve1dt5h963s80mbc6)You need the evalId to query.
Run the internal action from the evalScores/ directory. Always use --prod to query the production database (where CI writes results):
npx convex run --prod debug:getEvalDebugInfo '{"evalId": "<evalId>"}'
This returns a JSON object with:
| Field | Contents |
|-------|----------|
| eval | Name, category, evalPath, status (pass/fail + failure reason), task text |
| run | Model name, provider, experiment name, run status |
| steps | Array of step results: filesystem, install, deploy, tsc, eslint, tests — each with pass/fail/skipped and failure reason |
| outputFiles | Map of file path -> file content from the model's generated output (unzipped) |
| evalSourceFiles | Map of file path -> file content from the eval source (answer dir, grader, TASK.txt, etc.) |
With the data returned, compare:
steps for the first entry with status.kind === "failed". The failureReason field has the error message.outputFiles for the model's code.evalSourceFiles for the answer directory and grader test files.eval.task for the TASK.txt content.Common failure patterns:
outputFiles against evalSourceFiles (look for files like grader.test.ts or answer/) to understand what the tests expected.Classify the failure as one of:
Summarize:
Take get-convex/analyze-eval from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npx.
Without those the skill loads but fails at the first command.