agent-skills-hub/agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
This is a copy. The original lives at comeonoliver/agent-evaluation.
npx skills add https://github.com/agent-skills-hub/agent-skills-hub --skill agent-evaluation
Take agent-skills-hub/agent-evaluation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.