oliver-kriska/lab:autoresearch
> Self-improving loop for plugin skills. Reads program.md, proposes one mutation per iteration, evaluates against deterministic scorer, keeps improvements via git, reverts failures. Targets weakest skill+dimension. Use with /loop for overnight runs.
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix --skill lab:autoresearch
Iteratively improve plugin skills via the autoresearch pattern:
propose one mutation -> eval -> keep/revert -> repeat.
/lab:autoresearch # Targeted: attack weakest skill+dimension
/lab:autoresearch --skill review # Focus on one skill
/lab:autoresearch --strategy sweep # Process all skills alphabetically
/lab:autoresearch --dry-run # Show what would change, don't commit
For overnight runs:
/loop 5m /lab:autoresearch --strategy sweep --max-iterations 200
keep or revert command (never skip)All eval/git/journal operations go through ONE script. Do NOT run these manually.
# Find the weakest skill+dimension
python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
# Score a skill (before mutation, to get baseline)
python3 lab/autoresearch/scripts/run-iteration.py score <skill-name>
# After mutation: score + checks + compare → verdict (KEEP or REVERT)
python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>
# Act on verdict:
python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
--desc "what changed" --asi '{"hypothesis": "why", "mechanism": "how"}'
python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
--desc "what was attempted" --asi '{"hypothesis": "why", "regression": "what broke", "avoid": "do not retry this"}'
# Check overall progress
python3 lab/autoresearch/scripts/run-iteration.py status
lab/autoresearch/program.md (goals, mutable surface, rules)lab/autoresearch/ideas.md if it exists (deferred optimizations)python3 lab/autoresearch/scripts/run-iteration.py statusRun: python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
Parse the JSON: skill, dimension, failing_checks. If all_perfect → STOP.
lab/eval/evals/{skill}.jsonideas.md for deferred ideas about this skill${CLAUDE_SKILL_DIR}/references/mutation-strategies.mdpython3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>verdict fieldIf verdict is KEEP:
python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
--desc "..." --asi '{"hypothesis": "...", "mechanism": "..."}'
If verdict is REVERT:
python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
--desc "..." --asi '{"hypothesis": "...", "regression": "...", "avoid": "..."}'
If during analysis you discovered a promising optimization you can't act on now:
lab/autoresearch/ideas.md as a bullet${CLAUDE_SKILL_DIR}/references/mutation-strategies.md — mutation type catalog${CLAUDE_SKILL_DIR}/references/state-management.md — git protocol, journalinglab/autoresearch/program.md — research agenda (read every iteration)Take oliver-kriska/lab:autoresearch from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.