willoscar/schema-normalizer
| Normalize cross-skill JSONL interfaces (ids + titles + citation key formats) so downstream skills do not rely on best-effort joins.
npx skills add https://github.com/WILLOSCAR/research-units-pipeline-skills --skill schema-normalizer
Purpose: close a common failure mode in skills-first pipelines: schema drift across JSONL artifacts.
When fields are inconsistent (missing ids/titles, mixed citation-key formats), downstream skills start doing best-effort joins and fragile parsing.
This skill makes the interface explicit and deterministic.
outline/outline.yml (source of truth for section/subsection ids + titles)citations/ref.biboutline/subsection_briefs.jsonloutline/chapter_briefs.jsonloutline/evidence_bindings.jsonloutline/evidence_drafts.jsonloutline/anchor_sheet.jsonloutline/writer_context_packs.jsonloutput/SCHEMA_NORMALIZATION_REPORT.md (always written; PASS/FAIL + what changed).bak.* is created if changes are applied).For any record with sub_id: "<H2>.<H3>":
section_id exists (derived from the prefix before the dot)title, section_title exist (filled from outline/outline.yml)For any record with section_id: "<H2>":
section_title exists (filled from outline/outline.yml)Within these C2-C4 JSONL artifacts, normalize citation keys so they are raw BibTeX keys (no @ prefix):
"citations": ["smith2023", "jones2024"]Notes:
[@smith2023].Recommended placement in arxiv-survey(-latex):
evidence-draft + anchor-sheet and before writer-context-pack + evidence-selfloop.outline/evidence_drafts.jsonl and outline/anchor_sheet.jsonl are schema-stable before drafting packs are built.outline/outline.yml is missing or cannot be parsed, the skill FAILs.uv run python .codex/skills/schema-normalizer/scripts/run.py --helpuv run python .codex/skills/schema-normalizer/scripts/run.py --workspace <workspace>--workspace <dir>--unit-id <U###>--inputs <semicolon-separated>--outputs <semicolon-separated>--checkpoint <C#>uv run python .codex/skills/schema-normalizer/scripts/run.py --workspace <workspace> --inputs outline/outline.yml;citations/ref.bib;outline/subsection_briefs.jsonl;outline/chapter_briefs.jsonl;outline/evidence_bindings.jsonl;outline/evidence_drafts.jsonl;outline/anchor_sheet.jsonl --outputs output/SCHEMA_NORMALIZATION_REPORT.mdwriter-context-pack):uv run python .codex/skills/schema-normalizer/scripts/run.py --workspace <workspace> --inputs outline/outline.yml;citations/ref.bib;outline/writer_context_packs.jsonl --outputs output/SCHEMA_NORMALIZATION_REPORT.mdTake willoscar/schema-normalizer from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.