borghei/prompt-engineer-toolkit
> chain-of-thought, few-shot, regression testing, and rubrics. Use when designing production prompts, running A/B tests, or building prompt libraries.
npx skills add https://github.com/borghei/Claude-Skills --skill prompt-engineer-toolkit
The complete lifecycle for production prompts: design patterns that work, testing frameworks that catch regressions, versioning systems that track changes, and evaluation rubrics that replace subjective "looks good" with measurable quality. This treats prompts as production code with the same rigor — not clever tricks.
Tags: prompt engineering, chain-of-thought, few-shot, evaluation, testing, prompt versioning
Before designing or testing the prompt, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
| Tool | Purpose | Command |
|------|---------|---------|
| eval_scorer.py | Score evaluation results from JSON test cases (exact/contains/regex) | python scripts/eval_scorer.py suite.json --fail-under 0.80 --json |
| prompt_analyzer.py | Analyze prompt files for clarity, instruction density, few-shot coverage, tokens | python scripts/prompt_analyzer.py my_prompt.txt --json |
| prompt_diff.py | Compare two prompt versions for structural changes, instruction deltas, risk | python scripts/prompt_diff.py v2.txt v3.txt --show-diff --json |
Load the reference that matches the task — keep this file lean and pull detail on demand:
This skill covers:
This skill does NOT cover:
engineering/model-training-pipeline for training workflowsengineering/context-engine for context retrieval architectureengineering/agent-designer for agent system designengineering/llm-gateway-design for inference infrastructure patterns| Skill | Integration | Data Flow |
|-------|-------------|-----------|
| agent-designer | Agent system prompts are the highest-stakes prompts; use this toolkit to test and version them | Agent specs → prompt layers → tested system prompts |
| self-improving-agent | Prompt degradation signals feed into self-improvement loops for automatic correction | Test suite results → regression alerts → prompt iteration |
| context-engine | Retrieved context quality directly impacts prompt effectiveness; coordinate retrieval tuning with prompt testing | Retrieved chunks → prompt context layer → evaluation scores |
| ab-test-setup | A/B test prompt variants in production with statistical rigor before full rollout | Prompt candidates → traffic split → scoring comparison → winner promotion |
| llm-gateway-design | Gateway handles prompt routing, versioning, and model fallback at the infrastructure layer | Versioned prompts → gateway config → model routing → response logging |
| code-review-automation | Code review prompts are high-frequency production prompts that benefit from this toolkit's testing framework | Review criteria → prompt design → test suite → deployed reviewer prompt |
Take borghei/prompt-engineer-toolkit from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.