kennethkhoocy/llm-gold-bound-failure-check
| Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections).
npx skills add https://github.com/kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Take kennethkhoocy/llm-gold-bound-failure-check from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.