microck/ordinary-claude-machine-learning-llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
This is a copy. The original lives at comeonoliver/skillshub-llm-evaluation.
npx skills add https://github.com/Microck/ordinary-claude-skills --skill llm-evaluation
Take microck/ordinary-claude-machine-learning-llm-evaluation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.