wshobson/dataset-curation
Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.
npx skills add https://github.com/wshobson/agents --skill dataset-curation
This skill assumes finetuning-method-selection
already routed here — the next step is preparing
data, not choosing a method. What follows: format
selection by target method, the template/packing
mechanics behind the most common silent training
failures, rules for mixing in synthetic data
without collapse, and the dataset card that closes
out Phase 2 before a run starts.
Input: raw examples (demonstrations, preference
judgments, or task prompts) plus a routing decision
from finetuning-method-selection.
Output format: a formatted, packed, validated
JSONL dataset plus a completed dataset card — the
Phase 2 artifact /finetune checks before launching
training.
| Method | Shape | Rows |
|---|---|---|
| SFT, single-turn | Instruct (instruction/response or prompt/completion) | ~1,000+ floor |
| SFT, multi-turn | Conversation / ChatML messages list | ~1,000+ floor |
| DPO / ORPO | Preference pair (prompt, chosen, rejected) | Method-dependent, see preference-optimization |
| KTO | Unpaired (prompt, completion, label) | Method-dependent, see preference-optimization |
| GRPO / RLVR | Prompt-only (prompt + verifier metadata) | Method-dependent, see grpo-rlvr-training |
not a target. Below it, a handful of low-quality
or duplicate examples can dominate the gradient;
above it, quality over quantity — a smaller
verified, deduplicated set beats a larger noisy one.
formats plus a ShareGPT conversion note live in
references/formats-and-templates.md:
{"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]}
Apply the target model's chat template before
any concatenation or packing, never after — packing
raw text and templating the packed blob afterward
corrupts turn boundaries, landing role markers in
the wrong place relative to each example.
loss (-100 in the labels tensor) over system/user
turns and the template's own role markers — only
assistant-turn content tokens contribute to loss.
failure mode.** A model trained against one chat
template but served or evaluated with a different
one degrades without erroring. Verify the same
template string used in training is applied at
inference and eval time.
messages shape and letthe trainer template and mask it
(assistant_only_loss=True in current TRL) —
pre-rendering to a flat text field destroys the
turn boundaries masking needs. Full code sketch:
references/formats-and-templates.md. Sanity-check
before training — decode only unmasked positions;
expect only assistant text:
keep = batch["labels"][0] != -100
print(tokenizer.decode(batch["input_ids"][0][keep]))
**Without packing, 40–70% of compute is spent on
padding** — variable-length examples batched at a
fixed sequence length waste the gap between each
example's length and the batch's max. Packing
concatenates multiple examples into one sequence
up to the max length, cutting most of that waste.
sequence can contain several original examples, so
"steps per epoch" and any LR schedule keyed to
example count shift once packing is on — recompute
schedule milestones against packed-sequence count.
packed sequences before scaling to a full run.**
Confirm example boundaries land where expected,
template markers are intact per sub-example, and
the loss mask is still assistant-only within each
packed sequence. Not optional — packing bugs are
silent (the loss curve looks normal) and only
surface in eval quality, hours later:
for seq in packed_dataset.select(range(10)):
print(tokenizer.decode(seq["input_ids"]))
Training on a growing share of model-generated
data without a real-data floor drives measurable
quality collapse over successive generations —
25% real is the minimum that holds the line.
**General-domain replay rows
count toward this floor** —
"real" means "not generated
for this task from this
student," not "human-authored."
An all-synthetic-by-construction
dataset can meet the ≥25% floor
through replay alone (see
references/synthetic-data.md's
Replay-Mix Construction recipe);
state which rows count as "real"
in the dataset card rather than
leaving the floor structurally
unmeetable.
workhorses.** Magpie extracts prompts from the
model's own template prior; rejection sampling
generates several candidates per prompt and keeps
only the ones a filter passes. Both beat naive
single-shot generation.
generation by 1.3–2x sample efficiency** — aiming
at the student's actual failure modes hits a
quality bar with fewer filtered examples.
10–30%.** Plan volume accordingly — a 10,000-row
target at 15% accept needs ~65,000+ raw generations.
mix construction, and distillation pattern:
references/synthetic-data.md.
Every dataset that reaches training gets a card —
the required Phase 2 artifact /finetune checks
before launching. The card is not free-form
documentation; it MUST carry these fields:
source(s), synthetic method(s), or both),
traceable to trace-to-training-data output.
(train/eval/held-out) if split.
checked against the ≥25% real floor above.
(embedding threshold), or both; see the filter
funnel in references/synthetic-data.md.
string/identifier, kept consistent through
inference and eval — this is what ties an
eval-harness-first run back to the checkpoint.
max sequence length, and confirmation the
5–10-sequence manual inspection above was done.
A dataset missing any of these six fields isn't
ready for /finetune — the card is a gate, not a
summary written after the fact.
Before handing off to /finetune, confirm:
references/formats-and-templates.md — JSONLexamples per format, current-TRL masking code,
and the ShareGPT conversion note.
references/synthetic-data.md — generation-methodranking, filter funnel, replay-mix construction,
and teacher→student distillation pattern.
Related skills: finetuning-method-selection routes
here; lora-qlora-recipes, vision-sft, and
preference-optimization consume the datasets this
skill produces; trace-to-training-data is the
provenance source for graded-trajectory datasets;
eval-harness-first grades the resulting checkpoint.
Take wshobson/dataset-curation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.