Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
npx skills add https://github.com/wshobson/agents --skill checkpoint-promotion
The Phase 5 gate for the whole
plugin: a checkpoint that trains
cleanly and beats its task metric
still doesn't ship without
clearing all four stages below.
eval-harness-first built the
suite re-run here — this skill is
where that suite's baseline
decides something.
Input: a trained checkpoint,
eval/baseline-<model>.json from
eval-harness-first, and the
frozen eval/drift-suite.yaml.
Output format:
promotion-report.md — the
four-stage evidence plus a
terminal PROMOTE or REJECT
verdict that /finetune Phase 5
and /promote-checkpoint consume
directly.
Each stage gates the next — a
failure at stage 2 means stage 3
doesn't run. Stages 2 and 3 share
one expensive inference pass, so
running them concurrently and
applying gate order at verdict
time is licensed on a
deterministic arena (nothing
saved by serializing); a
judge-based arena should still
wait for stage 2 first — that's
where the real savings are.
any eval touches the
checkpoint: dedup the training
set, check for eval-goldens
leakage (the exact failure
trace-to-training-data's
Hygiene section exists to
prevent), and scan for label
noise. A checkpoint trained on
leaked goldens invalidates
every later stage.
capability-drift suite.**
Re-run eval-harness-first's
eval/drift-suite.yaml —
MMLU/GSM8K/IFEval plus 200–500
domain-adjacent items — against
the checkpoint and diff against
baseline-<model>.json per
benchmark against the Drift
Budget table below.
Position-randomized judge,
checkpoint vs. base model, same
prompts — or the deterministic
paired-comparison variant in
references/gate-templates.md
when every grader in the
harness is deterministic (no
LLM-judge; position
randomization N/A there).
**A holdout win that
loses the live arena does not
ship** — stage-2 numbers and
stage-3 judgments must agree; a
win on frozen goldens and a
loss in paired comparison is a
real signal, not a discrepancy
to explain away.
rollout with auto-rollback for
any checkpoint reaching
production traffic. **Local-only
users stop at stage 3** —
skipping stage 4 for a local
deployment is the correct
stopping point, not a shortcut.
| Drift (pts) | Verdict |
|---|---|
| ≤1 | Noise — proceed |
| 2–5 | Rerun with seed variation before deciding |
| >5 | HARD FAIL — no exception for task gains |
The >5pt row governs regardless
of the others: a checkpoint that
gained 8 points on the target
task and lost 6 points of general
capability still fails here —
task improvement never buys back
a drift-budget breach.
**Item count derives from the
budget, not convenience:** the
strict n for a half-width under
half the 5pt hard-fail threshold
is ~1,300 at typical accuracy
(p≈0.7); n=200 is a pragmatic
floor (±6pt half-width at that
same p, n=50 ±13pt) — report the
half-width with every verdict,
and treat a margin smaller than
it as REJECT (uncertain), not
PASS/HARD FAIL. Full math and a
5-run cautionary example:
references/gate-templates.md.
RERUN is not a verdict. A
2–5pt drift only ever produces a
PROMOTE or REJECT after the
seed-variation rerun completes —
PROMOTE requires landing back
at ≤1pt (noise); any rerun still
>1pt — 2–5pt band or >5pt breach
alike — resolves stage 2 to a
hard REJECT. No report may
reach the Verdict section with
stage 2 still showing RERUN.
Unmanaged LoRA fine-tuning loses
real general capability, and
stage 2 is what catches it:
unmanaged** — no replay, no
regularization.
— some replay or a conservative
LR.
disciplined case.
mix is the standard
mitigation** — blend general-
domain data into training
rather than target-task data
alone.
If a checkpoint hits the >5pt
hard fail in stage 2, work this
escalation ladder in order — the
one canonical order this skill
and references/gate-templates.md
both point to:
fraction — swap rows, don't
add them** (adding confounds
fraction with total optimizer
steps). Dose is not monotonic
at small-run scale (<~100
steps) — re-check drift after
any swap.
same rank/LR levers
lora-qlora-recipes and
preference-optimization tune
for the training run, applied
here in reverse.
This order is a default, not a
law: **remediation guidance from
a single before/after run pair
is a hypothesis** — label it
low-confidence once any lever
produces a reversal, and prefer
a seed-variation repeat over
trusting the next rung blindly.
A lever that clears the drift
breach but drops a
success-criterion metric below
target is a two-sided tradeoff
for a human, not a reason to
keep descending the ladder. Full
reasoning and the 5-run
trajectory behind both caveats:
references/gate-templates.md.
**Disclose drift-suite
instruction reuse.** A replay row
copying the drift harness's exact
instruction phrasing (not just
disjoint source items) makes that
benchmark's post-replay score an
upper bound — flag it
instruction-familiar, or re-probe
with a paraphrase, before
treating a near-budget pass as
clean.
promotion-report.md covers all
four stages as sections and
**must end with a terminal
verdict: PROMOTE or REJECT**,
the evidence that produced it,
and exactly one top remediation
when the verdict is REJECT.
Template: references/gate-templates.md.
The terminal contract other
skills parse:
## Verdict
REJECT
Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.
Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.
error.** A checkpoint that
fails stage 2's drift budget or
stage 3's arena comparison did
its job. Don't treat a REJECT
as a failed run needing a rerun
of this skill; it's the correct
output of a working gate.
menu.** Evidence sections may
list everything observed; the
verdict section names the
single highest-leverage fix per
the escalation ladder above. A
report that hedges across three
possible fixes hasn't done the
prioritization this skill
exists to do.
skill produces a verdict and a
report, not a re-triggered
training run. A REJECT hands
the remediation back to a human
decision at
finetuning-method-selection or
the relevant training skill.
eval-harness-first — owns thedrift suite and baseline this
skill re-runs and diffs
against; no baseline-<model>.json
means nothing to gate against.
quantized-export — the onlyvalid next step after a
PROMOTE verdict.
preference-optimization andlora-qlora-recipes — own the
LR and rank levers in the
Catastrophic Forgetting
escalation path; this skill
diagnoses the breach, those
skills own the config that
caused it.
dataset-curation — owns thereplay-mix construction recipe
the escalation ladder's first
rung applies.
Complete promotion-report.md
template with all four stages,
the drift-suite scoring table,
the paired-arena protocol (item
count, position randomization,
win-rate threshold), and a
replay-mix configuration example:
references/gate-templates.md.
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and econometric analysis.
Add unsigned integer (uint) type support to PyTorch operators by updating AT_DISPATCH macros. Use when adding support for uint16, uint32, uint64 types to operators, kernels, or when user mentions enabling unsigned types, barebones unsigned types, or uint support.
Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.
Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.
Take wshobson/checkpoint-promotion from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.