mcpbeat

Agent Survey Corpus

willoscar/agent-survey-corpus

| Download a small corpus of open-access arXiv survey/review PDFs about agentic systems and extract text for style learning.

3k tokens
context cost
the whole folder, loaded on every use
2
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
496
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/WILLOSCAR/research-units-pipeline-skills --skill agent-survey-corpus

What comes with it

11 139 bytes besides the instruction
scripts/run.py

The instruction itself

9 sections, as written by the author

Agent Survey Corpus (arXiv PDFs → text extracts)

Goal: create a small, local reference library so you can learn from real agent surveys when refining:

  • C2 outline structure (paper-like sectioning)
  • C4 tables/claims organization
  • C5 writing style and density

This is intentionally *not* part of the pipeline; it is an optional, repo-level toolkit.

Inputs

  • ref/agent-surveys/arxiv_ids.txt

Outputs

  • ref/agent-surveys/pdfs/
  • ref/agent-surveys/text/
  • ref/agent-surveys/STYLE_REPORT.md (tracked; auto-generated summary)

Workflow

1) Edit ref/agent-surveys/arxiv_ids.txt (one arXiv id per line).

2) Run the downloader to fetch PDFs and extract the first N pages to text.

3) Skim the extracted text under ref/agent-surveys/text/:

  • look at section counts (H2), subsection granularity (H3), and how they transition between chapters.
  • identify repeated rhetorical patterns you want the pipeline writer to imitate.

Script

Quick Start

  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --help
  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace . --max-pages 20

All Options

  • --workspace <dir> (use . to write into repo root)
  • --inputs <semicolon-separated> (default: ref/agent-surveys/arxiv_ids.txt)
  • --max-pages <N> (default: 20)
  • --sleep <seconds> (default: 1.0)
  • --overwrite (re-download + re-extract)

Examples

  • Download/extract into repo root ref/:
  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace . --max-pages 20
  • Download/extract into a specific folder (treated as workspace root):
  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace /tmp/surveys --max-pages 30

Troubleshooting

  • Download fails / timeout: rerun with a larger --sleep, or try fewer ids.
  • Text extract is empty: the PDF may be scanned; try another survey or increase --max-pages.
  • Files showing up in git status: PDFs/text are ignored via .gitignore (ref//pdfs/, ref//text/).

How to use it

Copy the folder

Take willoscar/agent-survey-corpus from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.