mcpbeat Sign in

Codebase Inspection Agent Skill

Inspect codebases: LOC, languages, ratios via pygount.

969 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
117
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/HezaoHezao/poirot --skill codebase-inspection

The instruction itself

11 sections, as written by the author

Codebase Inspection

Analyze repositories for lines of code, language breakdown, file counts, and

code-vs-comment ratios using pygount.

When to Use

  • User asks for LOC (lines of code) count
  • User wants a language breakdown of a repo
  • User asks about codebase size or composition
  • User wants code-vs-comment ratios
  • General "how big is this repo" questions

Prerequisites

pip install pygount

1. Basic Summary (Most Common)

cd /path/to/repo
pygount --format=summary \
  --folders-to-skip=".git,node_modules,venv,.venv,__pycache__,.cache,dist,build,.next,.tox,.eggs,*.egg-info" \
  .

IMPORTANT: Always use --folders-to-skip to exclude dependency/build

directories, otherwise pygount will crawl them and take a very long time.

2. Common Folder Exclusions

# Python project
--folders-to-skip=".git,__pycache__,venv,.venv,.tox,.eggs,*.egg-info,.pytest_cache,.mypy_cache,.ruff_cache"

# Node.js project
--folders-to-skip=".git,node_modules,dist,build,.next,.cache,coverage"

# General (safe default)
--folders-to-skip=".git,node_modules,venv,.venv,__pycache__,.cache,dist,build,.tox,.eggs,*.egg-info,.pytest_cache,.mypy_cache"

3. Detailed Per-File Output

pygount --format=summary \
  --folders-to-skip=".git,node_modules,.venv,__pycache__" \
  --names-to-skip="*.pyc,*.pyo,*.so,*.dylib" \
  /path/to/repo

4. Language Breakdown Only

pygount --format=summary /path/to/repo 2>/dev/null | grep -E "^\s+\w" | sort -t$'\t' -k2 -rn

5. JSON Output (for further processing)

pygount --format=json \
  --folders-to-skip=".git,node_modules,.venv,__pycache__" \
  /path/to/repo > codebase_stats.json

python3 -c "
import json
with open('codebase_stats.json') as f:
    data = json.load(f)
# Aggregate by language
from collections import defaultdict
by_lang = defaultdict(lambda: {'code': 0, 'files': 0})
for entry in data:
    lang = entry.get('language', 'unknown')
    by_lang[lang]['code'] += entry.get('code', 0)
    by_lang[lang]['files'] += 1
for lang, stats in sorted(by_lang.items(), key=lambda x: -x[1]['code']):
    print(f'{lang:20s} {stats[\"code\"]:8d} lines  {stats[\"files\"]:4d} files')
"

6. Quick LOC Count (without pygount)

If pygount isn't available, use find + wc:

# Count lines in all Python files (excluding venvs)
find . -name "*.py" -not -path "*/venv/*" -not -path "*/.venv/*" -not -path "*/__pycache__/*" | xargs wc -l | tail -1

# Count by language
echo "Python: $(find . -name '*.py' -not -path '*/.venv/*' | xargs wc -l 2>/dev/null | tail -1)"
echo "JavaScript: $(find . -name '*.js' -not -path '*/node_modules/*' | xargs wc -l 2>/dev/null | tail -1)"
echo "TypeScript: $(find . -name '*.ts' -not -path '*/node_modules/*' | xargs wc -l 2>/dev/null | tail -1)"

7. File Count by Type

# Count files by extension
find . -type f -not -path "*/.git/*" -not -path "*/node_modules/*" -not -path "*/.venv/*" | \
  sed 's/.*\.//' | sort | uniq -c | sort -rn | head -20

Pitfalls

  • Always exclude dependency dirs: node_modules, venv, .venv,

__pycache__, dist, build — otherwise pygount hangs or counts millions

of irrelevant lines.

  • Binary files: pygount skips them, but find + wc doesn't. Use

--names-to-skip for pygount, or filter with grep -I for find.

  • Generated files: *.min.js, *_pb2.py, auto-generated code inflates

counts. Exclude with --names-to-skip.

  • Encoding: pygount may fail on non-UTF-8 files. Use --encoding=utf-8

or --encoding=chardet for mixed-encoding repos.

Other skills for the same job

different authors, same section of the catalogue
MCP Builder
by anthropics
vendor ×13

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

30k tokens scripts
Changelog Generator
by frostant
×9

Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.

774 tokens
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
MCP Builder
by JayZeeDesign
×7

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

37k tokens scripts
Vercel React Native Skills
by vercel-labs
vendor ×6

React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.

39k tokens
Vercel React Best Practices
by ratacat
×5

React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.

34k tokens
Next Best Practices
by vercel-labs
vendor ×4

Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling

20k tokens
Using Git Worktrees
by ZhanlinCui
×4

Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification

1k tokens

How to use it

Copy the folder

Take hezaohezao/codebase-inspection from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip. Without those the skill loads but fails at the first command.