mcpbeat Sign in

Taxonomy Builder Skill for Codex

| Build a 2+ level taxonomy (`outline/taxonomy.yml`) from a core paper set and scope constraints, with short descriptions per node.

10k tokens
context cost
the whole folder, loaded on every use
15
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
496
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/WILLOSCAR/research-units-pipeline-skills --skill taxonomy-builder

What comes with it

36 531 bytes besides the instruction
assets/domain_packs/embodied_ai.yaml
assets/domain_packs/gen_image.yaml
assets/domain_packs/llm_agents.yaml
assets/domain_packs/rag_evaluation.yaml
assets/taxonomy_schema.json
references/archetypes_generic.md
references/domain_pack_embodied_ai.md
references/domain_pack_gen_image.md
references/domain_pack_llm_agents.md
references/examples_bad.md
references/examples_good.md
references/overview.md
references/taxonomy_principles.md
scripts/run.py

The instruction itself

15 sections, as written by the author

Taxonomy Builder (router, compatibility mode)

Build outline/taxonomy.yml from papers/core_set.csv.

P0 compatibility note:

  • The output contract stays the same (outline/taxonomy.yml, YAML list, >=2 levels, concrete descriptions).
  • Curated domain taxonomies now live in assets/domain_packs/*.yaml instead of Python prose.
  • scripts/run.py stays a deterministic scaffold/helper: detect domain pack -> load pack when available -> otherwise fall back to the generic builder.

Load Order

  • references/overview.md
  • references/taxonomy_principles.md
  • If a domain pack applies, read its references/domain_pack_<domain>.md and assets/domain_packs/<domain>.yaml
  • Otherwise read references/archetypes_generic.md
  • Calibrate naming/description quality with references/examples_good.md and references/examples_bad.md

Current compatibility packs:

  • llm_agents
  • gen_image
  • embodied_ai
  • rag_evaluation

Explicit refinement marker

Create outline/taxonomy.refined.ok only after reviewing a manually refined taxonomy. The marker is honored only while it is newer than the taxonomy, its upstream evidence, and the generator; stale markers are removed and the prior taxonomy is backed up before regeneration.

Inputs

  • papers/core_set.csv (required)
  • Optional: papers/papers_dedup.jsonl
  • Optional: DECISIONS.md, GOAL.md, queries.md

Outputs

  • outline/taxonomy.yml

Asset contract

  • assets/taxonomy_schema.json: machine-readable shape for domain packs / output expectations
  • assets/domain_packs/*.yaml: compatibility domain packs for supported domains

Script role

Use scripts/run.py only for deterministic help:

  • never overwrite non-placeholder user taxonomy
  • preserve current CLI flags / output path
  • load a supported domain taxonomy only when GOAL.md / queries.md explicitly match its detection contract
  • keep the generic fallback builder for non-packed domains

When to refine manually

Refine the generated taxonomy before marking the unit DONE if:

  • top-level buckets feel like keyword clusters instead of chapter-level questions
  • leaf names are generic (Overview, Benchmarks, Open Problems, Misc)
  • descriptions lack scope cues or representative paper anchors
  • domain detection chose the wrong pack

Quick start

  • uv run python .codex/skills/taxonomy-builder/scripts/run.py --help
  • uv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>

Execution notes

When running in compatibility mode, scripts/run.py currently reads:

  • papers/core_set.csv as the required corpus input
  • papers/papers_dedup.jsonl when present for generic corpus signals
  • GOAL.md and queries.md as the authoritative domain-pack selection intent; corpus term co-occurrence cannot override them

Script

Quick Start

  • uv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>

All Options

  • --workspace <dir>
  • --top-k <int>
  • --min-freq <int>
  • --unit-id <id>
  • --inputs <a;b;...>
  • --outputs <a;b;...>
  • --checkpoint <C*>

Examples

  • uv run python .codex/skills/taxonomy-builder/scripts/run.py --workspace <workspace>

Troubleshooting

  • If the wrong domain pack is chosen, inspect GOAL.md, queries.md, and the pack detect rules before changing Python.
  • If outline/taxonomy.yml already contains a real non-placeholder taxonomy, the script intentionally returns without overwriting it.
  • If no pack matches, the script falls back to the generic builder.

Other skills for the same job

different authors, same section of the catalogue
Content Research Writer
by frostant
×10

Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section. Transforms your writing process from solo effort to collaborative partnership.

4k tokens
Lead Research Assistant
by frostant
×8

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

2k tokens
Notebooklm
by ZhanlinCui
×6

Use this skill to query your Google NotebookLM notebooks directly from Claude Code for source-grounded, citation-backed answers from Gemini. Browser automation, library management, persistent auth. Drastically reduced hallucinations through document-only responses.

26k tokens scripts
Biorxiv Database
by christophacham
×4

Efficient database search tool for bioRxiv preprint server. Use this skill when searching for life sciences preprints by keywords, authors, date ranges, or categories, retrieving paper metadata, downloading PDFs, or conducting literature reviews.

9k tokens scripts
Openalex Database
by christophacham
×4

Query and analyze scholarly literature using the OpenAlex database. This skill should be used when searching for academic papers, analyzing research trends, finding works by authors or institutions, tracking citations, discovering open access publications, or conducting bibliometric analysis across 240M+ scholarly works. Use for literature searches, research output analysis, citation analysis, and academic database queries.

13k tokens scripts
Uspto Database
by christophacham
×4

Access USPTO APIs for patent/trademark searches, examination history (PEDS), assignments, citations, office actions, TSDR, for IP analysis and prior art searches.

21k tokens scripts
Denario
by christophacham
×3

Multiagent AI system for scientific research assistance that automates research workflows from data analysis to publication. This skill should be used when generating research ideas from datasets, developing research methodologies, executing computational experiments, performing literature searches, or generating publication-ready papers in LaTeX format. Supports end-to-end research pipelines with customizable agent orchestration.

11k tokens
Hypogenic
by christophacham
×3

Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.

7k tokens

How to use it

Copy the folder

Take willoscar/taxonomy-builder from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.