> Run and interpret skill evaluations. Use when you need to evaluate a skill, run probe/test/PR-val, check if a PR regresses quality, compare two versions, or diagnose why the score dropped. Handles the full eval lifecycle including solver trajectory diagnosis for tool-level error detection.
npx skills add https://github.com/Tencent/SkillHone --skill skillhone-evaluation
Run evaluations and diagnose results to inform improvement decisions.
The evaluator is not just a scorer. It is a harness that runs the skill against
private tasks and leaves a structured evidence trail:
logic, task contract, and optional compiler/audit helpers.
may create files, call scripts, and write the required artifact.
call task-local validators, compilers, parsers, renderers, or audit helpers.
trajectory.jsonl, produced artifacts, and stderr-likesignals that explain failures the score cannot explain.
This separation matters. A low score may come from weak skill instructions, but
it may also come from missing files, tool crashes, invalid compiled artifacts,
over-strict verifier rules, or infrastructure errors. Evaluation work is about
mapping the failure to the correct harness layer.
On Forgejo-backed repos, status.py is a read-only context check before PR
validation or merge decisions:
python3 ~/.skillhone/skills/skillhone/scripts/status.py
# Probe — fast iteration signal
python3 ~/.skillhone/skills/skillhone/scripts/eval.py \
--skill-dir /path/to/skill --eval-dir /path/to/eval-repo \
--split probe --output _data/probe_result.json
# Test — final benchmark (NEVER during iteration)
python3 ~/.skillhone/skills/skillhone/scripts/eval.py \
--skill-dir /path/to/skill --eval-dir /path/to/eval-repo \
--split test --output test_result.json
The output JSON includes a "workdir" field pointing to solver working
directories (e.g. /data/tmp/eval_agent_xyz/) containing trajectory.jsonl
files and produced artifacts for deeper diagnosis.
| Subagent | What it does |
|----------|-------------|
| trajectory-analyzer | Reads workdir/work_<uid>/trajectory.jsonl files to diagnose tool errors (rate limits, wrong tool calls, script crashes). Outputs redacted _data/trajectory_diagnosis.json safe to share with improver. |
| pr-quality-reviewer | Merge gate for skill PRs — runs static check + rubric scoring, posts PR comment, returns APPROVE/REQUEST_CHANGES. |
probe_result.json captures scores and some error categories, but it is only
the top of the evidence trail. It cannot fully explain runtime behavior such as:
web_search directly instead of Bash("python3 scripts/web_search.py ...")The trajectory-analyzer subagent fills this gap by reading raw solver logs. Its output distinguishes infrastructure failures (fix scripts/config) from skill failures (fix SKILL.md). This distinction is critical for avoiding wasted iterations.
Some artifact tasks are compiler-like: Mermaid, LaTeX, TypeScript, Python tests,
SQL parsers, JSON/YAML schema validators, browser renderers, and similar tools
produce actionable stderr or diagnostics. Do not reduce these failures to
wrong_answer.
When a failed trace produced an artifact, inspect the solver workdir from the
workdir field and run the task-local compiler, validator, renderer, or audit
helper on that artifact. Prefer commands and helpers shipped by the eval repo or
described in its README/contract; if none exist, use the standard local compiler
for that artifact type. Capture only concise, non-gold diagnostic summaries:
answer file"
Write this as _data/compiler_diagnosis.json or include it in the existing
diagnosis file. It is safe to share with the improver when it contains only
error messages, failed score names, and artifact snippets needed to identify the
syntax class; do not include gold answers or full eval questions.
If the task-local verifier already exposes detailed failed score keys, preserve
those names. They are usually better improvement signals than a rewritten
natural-language summary.
| Field | Meaning |
|-------|---------|
| score | pass rate (0.0–1.0) |
| avg_duration_s | efficiency; rising duration with flat score = looping/waste |
| traces[].error | "hard timeout", "agent_process_error", or empty |
| workdir | path to solver trajectories for deeper analysis |
Decision thresholds:
≥ +0.02 → real improvement±0.02 → noise, don't claim improvement≤ −0.04 → regression, revertAlways label which harness run produced a score. In a full SkillHone run there
may be several valid scores: baseline probe, iteration probe, PR validation, and
a final driver re-score after the master agent exits. These can differ because
they may use different skill checkouts, regenerated eval data, or output paths.
When writing issues, PR comments, wiki observations, or user summaries, cite the
score source in words: split, output JSON path if available, workdir if useful,
and whether it is an internal iteration score or final harness score. Do not
collapse multiple scores into one number without naming the source.
| Split | Purpose | Who sees |
|-------|---------|----------|
| probe | Iteration signal | orchestrator (redacted traces) |
| pr_val | PR merge gate | orchestrator (aggregate only) |
| test | Final benchmark | orchestrator only, NEVER during iteration |
test during iteration — it contaminates the final benchmark.Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
Replace with description of the skill and when Claude should use it.
Use when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies
This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.
Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.
Use when creating new skills, editing existing skills, or verifying skills work before deployment
Take tencent/skillhone-evaluation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.