ai4s-research/experiment-suite
Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.
npx skills add https://github.com/ai4s-research/ai4s-skills --skill experiment-suite
End-to-end experiment package builder. Single stage, full quality from the start. The agent (Claude Code / Cursor / Aider / Codex / …) writes everything directly using its own tools (Write, Bash, WebFetch, …). This skill contains procedure + reference playbooks + figure-example scripts — no Python runtime, no LLM SDK.
The substantive work is decomposed into reference playbooks under references/:
| Reference | Topic |
|---|---|
| references/00-incremental-execution.md | how to do this without losing work: batches, persistence, resume — read first |
| references/01-design-depth.md | what a real experiment design contains (motivation → hypothesis → datasets → baselines → metrics → ablations → budget) |
| references/01a-data-contract.md | runtime dataset binding: source, access route, version, split, and reuse boundary |
| references/02-code-quality.md | code-skeleton standards — runnable model.py, data.py, train.py, evaluate.py |
| references/03-results-protocol.md | results.json schema; measured / simulated / illustrative provenance |
| references/04-publication-figures.md | publication-grade charts, multi-panel layouts, taste rules |
| references/04a-figure-contract.md | figure logic before plotting: conclusion, panel map, reviewer risk |
| references/04b-figure-qa.md | export bundle, editable text, statistics and image-integrity QA |
| references/05-report-structure.md | structured experiment_report.md (problem → design → method → results → analysis → limitations) |
| references/06-quality-gate.md | self-check before delivery |
Also: figure_examples/ — publication-style matplotlib scripts plus a shared style kit the agent can use as starting points.
Read the relevant reference _before_ writing, not after. The full pass does not fit in a single turn — references/00-incremental-execution.md is the only execution mode that completes.
paper-writer.literature-survey.Confirm with the user:
results.json or run train.py against real data later.results.json as a placeholder. Every figure/table caption must say "simulated".If the user has data and time, push toward measured mode. If not, simulated is acceptable provided disclosures are honest in every artefact.
QUESTION="<research_question>"
SLUG=$(python3 -c "import re,hashlib,sys; t=sys.argv[1]; n=re.sub(r'[\\s_]+','-',re.sub(r'[^\\w\\s-]','',t.lower().strip())).strip('-')[:40].rstrip('-'); h=hashlib.sha1(t.encode()).hexdigest()[:8]; print(f'{n}-{h}')" "$QUESTION")
TS=$(date +%Y-%m-%d_%H%M%S)
RUN=output/experiment-suite/$SLUG/$TS
mkdir -p "$RUN/experiment" "$RUN/figures"
ln -sfn "$TS" "output/experiment-suite/$SLUG/latest"
In commands below $RUN = output/experiment-suite/<slug>/latest.
The agent will create five top-level files inside $RUN/:
experiment_design.mddata_contract.mdexperiment/{model.py,data.py,train.py,evaluate.py,config.yaml,requirements.txt,README.md}results.jsonfigures/*.pdf plus their make_*.py source and a manifest.jsonexperiment_report.mdOpen references/00-incremental-execution.md first. Then carry out the six tracks below across many turns, persisting state to $RUN/ after every batch.
Open: references/01-design-depth.md and references/01a-data-contract.md. First write $RUN/data_contract.md as the dataset contract for this run. It must say whether the data are user-supplied, agent-discovered, reused public, controlled, or synthetic fallback. Then write $RUN/experiment_design.md as a real design (≥ 700 words): motivation → hypothesis → datasets → baselines → metrics → ablations → compute budget. Justify every choice.
Open: references/02-code-quality.md. Fill $RUN/experiment/ with code that an engineer could launch with python train.py --config config.yaml. Real (if minimal) model class, real data loader, real train loop, real eval. The generated data.py and config.yaml are runtime products of this run and should bind to $RUN/data_contract.md, not to a repository-wide hard-coded benchmark. Add a README.md with run instructions.
Open: references/03-results-protocol.md. Produce $RUN/results.json with a well-formed schema: per-seed entries, per-method per-metric mean & std, ablation block, and a provenance field that names the source.
experiment/train.py (or supplies a results JSON) and the agent loads it into $RUN/results.json, setting "simulated": false and "provenance": "loaded from <path>".$RUN/data_contract.md; results.json provenance must point back to that binding."simulated": true.Open: references/04-publication-figures.md, references/04a-figure-contract.md, references/04b-figure-qa.md, and figure_examples/. Before writing plotting code, define the figure contract in a small working note under $RUN/figures/figure_contract.md:
Then plan and generate at minimum:
Save each figure into $RUN/figures/<basename>.pdf with its make_*.py source alongside. Prefer saving an editable .svg and print-grade .tiff alongside the PDF when the environment supports it. Append entries to $RUN/figures/manifest.json storing basenames only (never absolute paths) so paper-writer can copy them in directly. Apply the shared publication style (embedded fonts, explicit palette, panel labels, simulated watermark when applicable).
If simulated, watermark the figures or always note "simulated" in their captions in the report.
Open: references/05-report-structure.md. Write $RUN/experiment_report.md with sections: problem statement → design rationale → method → setup → results → analysis → limitations. Reference figures by filename. This report is the primary deliverable for users who want only the experiment package (no paper-writer follow-up).
Open: references/06-quality-gate.md. Targets: design ≥ 700 words, code imports cleanly (python -c "import experiment.model" from inside $RUN), results.json passes schema check, ≥ 3 figures, report ≥ 6 sections.
Report:
output/experiment-suite/<slug>/latest/experiment_design.mdoutput/experiment-suite/<slug>/latest/experiment/ — runnable code package.output/experiment-suite/<slug>/latest/results.json — with provenance.output/experiment-suite/<slug>/latest/figures/ — publication-grade charts + manifest.json.output/experiment-suite/<slug>/latest/experiment_report.md — structured report.references/06-quality-gate.md.The paper-writer skill computing the same slug for the same topic will look here:
output/experiment-suite/<slug>/latest/results.json — source of the numbers and the "simulated" flag (drives the disclosure clause in the paper).output/experiment-suite/<slug>/latest/figures/*.pdf (+ manifest.json) — figures to reuse rather than redraw.Always store basenames in manifest.json. Absolute paths in the manifest break paper-writer's \includegraphics{figures/<basename>}.
import anthropic / import openai. The skill is SKILL.md + references + figure examples only.results.json ("simulated": true), in figure captions, in the report's top-of-page disclosure, and in any downstream paper's \thanks footnote.experiment/README.md.Take ai4s-research/experiment-suite from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.