willoscar/literature-engineer
| Multi-route literature expansion + metadata normalization for evidence-first surveys.
npx skills add https://github.com/WILLOSCAR/research-units-pipeline-skills --skill literature-engineer
Goal: build a large, verifiable candidate pool for downstream dedupe/rank, mapping, notes, citations, and drafting.
This skill is intentionally evidence-first: if you can't reach the target size with verifiable IDs/provenance, the correct behavior is to block and ask for more exports / enable network, not to fabricate.
Always read:
references/domain_pack_overview.md — how domain packs drive topic-specific behaviorDomain packs (loaded by topic match):
assets/domain_packs/llm_agents.json — pinned classic/survey arXiv IDs for LLM agent topicsUse scripts/run.py only for:
Do not treat run.py as the place for:
queries.mdkeywords, exclude, max_results, time windowpapers/import.(csv|json|jsonl|bib)papers/arxiv_export.(csv|json|jsonl|bib)papers/imports/*.(csv|json|jsonl|bib)papers/snowball/*.(csv|json|jsonl|bib)papers/papers_raw.jsonltitle (str), authors (list[str]), year (int|""), url (str)arxiv_id and/or doiabstract (str; may be empty in offline mode)source (str) + provenance (list[dict])papers/papers_raw.csv (human scan)papers/retrieval_report.md (route counts, missing-meta stats, next actions)provenance.retrieval_policy.minimum_records, use that value; survey profiles may instead derive a stricter pool target from core_size.arxiv_id or doi, plus url).uv run python .codex/skills/literature-engineer/scripts/run.py --helpuv run python .codex/skills/literature-engineer/scripts/run.py --help.queries.md.papers/import.(csv|json|jsonl|bib), papers/arxiv_export.(csv|json|jsonl|bib), papers/imports/*.(csv|json|jsonl|bib).papers/snowball/*.(csv|json|jsonl|bib).--online and/or --snowball.ref.bib can include must-cite anchors even when keyword search misses them.r.jina.ai proxy so the pipeline can still self-boot without manual exports.0 records due to transient network errors, a simple rerun is often sufficient (the pipeline should not fabricate).papers/imports/ then run:uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace>uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --input path/to/a.bib --input path/to/b.jsonluv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --onlineuv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --snowballSymptom:
papers/papers_raw.jsonl is below the explicit or profile-derived minimum declared by the locked Workflow.Causes:
Solutions:
papers/imports/ (multiple routes/queries).papers/snowball/.--online --snowball.Symptom:
arxiv_id and doi.Solutions:
--online to backfill arXiv IDs.Take willoscar/literature-engineer from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.