oaustegard/building-github-index-v2
DEPRECATED - Use building-github-index instead. Superseded per-file GitHub API implementation of progressive disclosure repository indexes.
npx skills add https://github.com/oaustegard/claude-skills --skill building-github-index-v2
⚠️ DEPRECATED: Use the building-github-index skill instead.
Despite the "-v2" directory name, this is the *older* implementation. It fetches
the repo tree and then every file individually through api.github.com, which is
slow and burns per-file rate limit. building-github-index supersedes it with a
single-request tarball download that processes files locally.
This directory is retained only so existing references resolve. It receives no
further updates.
Create markdown indexes of GitHub repositories optimized for Claude project knowledge. Indexes enable retrieval via GitHub API with semantic descriptions for effective matching.
# Documentation repos (markdown/notebooks)
python scripts/github_index.py owner/repo -o index.md
# Code repos (extract symbols via tree-sitter)
python scripts/github_index.py owner/repo --code-symbols -o index.md
# Multiple repos combined
python scripts/github_index.py owner/repo1 owner/repo2 -o combined.md
| Flag | Description |
|------|-------------|
| -o, --output | Output file (default: github_index.md) |
| --token | GitHub PAT; also reads GITHUB_TOKEN env |
| --include-patterns | Only index matching globs: "docs/" "src/" |
| --exclude-patterns | Skip matching globs: "test/**" |
| --max-files | Cap files per repo (default: 200) |
| --skip-fetch | Tree only, no content fetch (fast, filename-only descriptions) |
| --code-symbols | Include code files, extract function/class names via tree-sitter |
title: and description: fields--code-symbols)Some repos have stub files (links to external docs, empty readmes). In these cases:
Manual curation recommended. Use the tree output and domain knowledge:
# Get tree structure only (fast)
python scripts/github_index.py owner/repo --skip-fetch -o skeleton.md
# Then manually enhance descriptions based on domain knowledge
For code-heavy repos with embedded apps:
acc_wav_gen → "ACC waveform generation"# {Repo} - Content Index
**Repository:** {url}
**Branch:** `{branch}`
## Retrieval Method
{API curl commands}
---
## {Category}
| Description | Path |
|-------------|------|
| {What this covers} | `{path/file.md}` |
Description column leads (relevance matching), path follows (retrieval key).
Enumerate files:
curl -sL "https://api.github.com/repos/OWNER/REPO/git/trees/BRANCH?recursive=1"
Fetch content:
curl -s "https://api.github.com/repos/OWNER/REPO/contents/PATH?ref=BRANCH" \
-H "Accept: application/vnd.github+json" | \
python3 -c "import sys,json,base64; print(base64.b64decode(json.load(sys.stdin)['content']).decode())"
Allowlist: api.github.com, raw.githubusercontent.com
accessing-github-repos - Private repos, PAT setup, tarball downloadtree-sitting - Detailed code structure (methods, imports, line numbers)For token-constrained project knowledge, use the condensed script:
python scripts/pk_index.py owner/repo -o repo_pk.md
Produces ~80% smaller output:
path — descriptionIdeal when adding multiple repo indexes to project knowledge.
Take oaustegard/building-github-index-v2 from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.