mcpbeat Sign in

Code Audit Skill for Claude

> Two-phase code audit workflow for empirical research scripts. Use this skill when the user asks to review, audit, check, validate, or verify R or Python research code — especially code that processes licensed or sensitive data (WRDS, CRSP, WellDatabase, PLIDA). Also use when the user says "check my code", "review this script", "does this look right", or "audit my analysis."

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
146
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/aspi6246/Claude-Code-Skills-for-Academics --skill code-audit

The instruction itself

9 sections, as written by the author

Research Code Audit Protocol

Overview

This is a two-phase audit. Phase 1 is strictly read-only — do not

execute any code. Phase 2 only proceeds with explicit user approval.

This separation matters because research code often touches licensed

datasets (WRDS, CRSP, WellDatabase, PLIDA) that have access restrictions,

and because running code on large datasets without review risks producing

results from buggy pipelines that then get embedded in manuscripts.


Phase 1: Static Review (read-only)

Read the code without executing anything. Produce a structured report

covering the sections below.

1. Data integrity

  • Are there checks for duplicates in the panel identifier?
  • Are missing values / NAs handled explicitly (not silently dropped)?
  • Is the panel balanced or unbalanced, and is this acknowledged?
  • Are merges validated? Check:
  • Join type (left, inner, full) — is it intentional?
  • Many-to-one vs one-to-one — is this documented?
  • Post-merge row count check (did the merge inflate observations?)
  • Are data filters documented and justified in comments?
  • Are there sample selection steps that could introduce bias?

2. Econometric concerns

  • Is the identification strategy clearly implemented in the code?
  • Are standard errors clustered at the appropriate level?
  • Panel data: at minimum, cluster at the unit level
  • If iid SEs are used, flag this as likely incorrect
  • Are fixed effects consistent with the stated research design?
  • For Difference-in-Differences:
  • Is the parallel trends assumption testable with this data?
  • Is it tested (e.g., event study plot, pre-treatment coefficients)?
  • For staggered adoption: is a robust estimator used (Sun & Abraham,

Callaway & Sant'Anna, etc.) or is vanilla TWFE flagged as problematic?

  • For Instrumental Variables:
  • Is the first stage reported?
  • Is the first-stage F-statistic checked (rule of thumb: F > 10)?
  • Is the exclusion restriction discussed in comments?
  • Are there potential bad controls (Angrist & Pischke)?
  • Variables that are themselves affected by treatment
  • Post-treatment outcomes included as controls
  • Is winsorisation or trimming applied? At what level? Is it justified?

3. Reproducibility

  • Is a random seed set where randomisation is used?
  • Are file paths relative to the project root (not absolute)?
  • Are package versions recorded (renv.lock, sessionInfo(), or equivalent)?
  • Can the pipeline run end-to-end from raw data to final output?
  • Is there a clear execution order (numbered scripts, makefile, or README)?
  • Are intermediate outputs saved so individual scripts can be re-run?

4. Code quality

  • Redundant computations or copy-paste patterns that should be functions
  • Hardcoded values that should be parameters (e.g., winsorisation cutoffs,

sample date ranges, variable names repeated as strings)

  • Missing comments on non-obvious transformations
  • Variable names that are ambiguous or misleading
  • Scripts that are excessively long (>300 lines) and should be split

Phase 2: Execution checks (requires explicit user approval)

Do NOT proceed to this phase unless the user explicitly says something like

"go ahead and run it", "you can execute", or "test it."

When approved:

  • Run on a small sample or synthetic data first — never on the full

dataset as a first step

  • Verify output dimensions match expectations
  • Check that coefficient magnitudes are reasonable (not astronomically

large or exactly zero when they shouldn't be)

  • Compare against any baseline results the user provides
  • Check for convergence warnings or estimation issues in the log

Output format

Present findings as a structured report with severity levels:

  • 🔴 Critical: Results may be wrong. Must fix before interpreting output.

Examples: wrong clustering, bad merge inflating observations, post-treatment

controls, missing fixed effects.

  • 🟡 Warning: Potential issue worth investigating. May or may not affect

conclusions. Examples: no duplicate check, missing parallel trends test,

hardcoded sample restrictions.

  • 🟢 Note: Suggestion for improvement. Won't affect results but improves

code quality or reproducibility. Examples: absolute file paths, missing

comments, inefficient code patterns.

End every audit with a Recommendations section listing concrete,

actionable fixes ordered by severity. Each recommendation should reference

the specific line(s) of code involved.

Other skills for the same job

different authors, same section of the catalogue
Networkx
by christophacham
×3

Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python. Use when working with network/graph data structures, analyzing relationships between entities, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks, or visualizing network topologies. Applicable to social networks, biological networks, transportation systems, citation networks, and any domain involving pairwise relationships.

15k tokens
Pyzotero
by christophacham
×3

Interact with Zotero reference management libraries using the pyzotero Python client. Retrieve, create, update, and delete items, collections, tags, and attachments via the Zotero Web API v3. Use this skill when working with Zotero libraries programmatically, managing bibliographic references, exporting citations, searching library contents, uploading PDF attachments, or building research automation workflows that integrate with Zotero.

9k tokens
Networkx
by ComeOnOliver
×3

Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python. Use when working with network/graph data structures, analyzing relationships between entities, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks, or visualizing network topologies. Applicable to social networks, biological networks, transportation systems, citation networks, and any domain involving pairwise relationships.

17k tokens
Networkx
by K-Dense-AI
×1

Create, analyze, and visualize complex networks and graphs in Python with NetworkX. Use when working with network/graph data structures, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks (random, scale-free, small-world), reading/writing graph file formats, or drawing network topologies. Common applications include social, biological, transportation, and citation networks.

15k tokens
Neurokit2
by K-Dense-AI
×1

Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.

45k tokens scripts
Pyzotero
by K-Dense-AI
×1

Interact with Zotero reference management libraries using the pyzotero Python client. Retrieve, create, update, and delete items, collections, tags, and attachments via the Zotero Web API v3. Use this skill when working with Zotero libraries programmatically, managing bibliographic references, exporting citations, searching library contents, uploading PDF attachments, or building research automation workflows that integrate with Zotero.

10k tokens
Networkx
by zLanqing
×1

Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python. Use when working with network/graph data structures, analyzing relationships between entities, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks, or visualizing network topologies. Applicable to social networks, biological networks, transportation systems, citation networks, and any domain involving pairwise relationships.

15k tokens
Fp Ts Pragmatic
by lingxling
×1

A practical, jargon-free guide to fp-ts functional programming - the 80/20 approach that gets results without the academic overhead. Use when writing TypeScript with fp-ts library.

4k tokens

How to use it

Copy the folder

Take aspi6246/code-audit from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.