aperivue/deidentify
> De-identify clinical research data before LLM-assisted analysis. Standalone Python CLI detects PHI via regex + heuristics with 10 country locale packs (kr, us, jp, cn, de, uk, fr, ca, au, in). Interactive terminal review. No LLM touches raw data — the script runs locally without any network or AI calls.
npx skills add https://github.com/Aperivue/medsci-skills --skill deidentify
You are guiding a medical researcher through data de-identification. The actual
de-identification is performed by a standalone Python script that runs WITHOUT
any LLM. Your role is to explain, guide, and verify — not to see or process raw
PHI data.
The script processes data locally. You never need to see patient-level data.
(SHA-256 hashes only), and de-identified output (PHI already removed).
English for technical terms (PHI, HIPAA, Safe Harbor, etc.).
${CLAUDE_SKILL_DIR}/references/hipaa_18_identifiers.md — HIPAA Safe Harbor checklist${CLAUDE_SKILL_DIR}/references/korean_phi_patterns.md — Korean-specific regex patterns${CLAUDE_SKILL_DIR}/references/date_shift_guide.md — Date shifting best practicesRead relevant references before advising the researcher.
openpyxl (for .xlsx files): pip install openpyxlAsk the researcher:
Based on answers, recommend the appropriate command:
python deidentify.py full <file> --locale <code>python deidentify.py scan <file> --locale <code> firstAvailable locale codes: kr (Korea), us (USA), jp (Japan), cn (China), de (Germany),
uk (United Kingdom), fr (France), ca (Canada), au (Australia), in (India).
If --locale is omitted, the script shows an interactive country selection menu.
Users can provide a custom locale file via --locale-file custom.json.
Guide the researcher to run the script. The script is located at:
${CLAUDE_SKILL_DIR}/deidentify.py
Full pipeline (recommended for most users):
python ${CLAUDE_SKILL_DIR}/deidentify.py full data.xlsx \
--locale kr \
--output-dir ./deidentified/ \
--auto-accept-safe
Step-by-step (for careful review):
# Step 1: Scan
python ${CLAUDE_SKILL_DIR}/deidentify.py scan data.xlsx --locale kr --output-dir ./deidentified/
# Step 2: Review (interactive)
python ${CLAUDE_SKILL_DIR}/deidentify.py review ./deidentified/scan_report.json
# Step 3: Apply
python ${CLAUDE_SKILL_DIR}/deidentify.py apply ./deidentified/reviewed_report.json
Options:
--locale CODE: Country locale for PHI patterns (kr, us, jp, cn, de, uk, fr, ca, au, in)--locale-file PATH: Custom locale JSON file (copy locales/_template.json to create one)--auto-accept-safe: Skip confirmation for columns classified as SAFE (faster for large datasets)--hash-mapping: Store SHA-256 hashes instead of original values in mapping file (one-way, more secure)--output-dir: Where to save de-identified file, mapping, and audit log-v/--verbose: Enable debug loggingThe script's terminal review has three passes:
The researcher confirms or overrides each classification.
with more sample values displayed.
individual decisions before confirming.
Coach the researcher. Deliver these prompts in the researcher's preferred language:
After the script completes, help the researcher verify:
cat ./deidentified/audit_log.csv | head -20
Verify the number of changes, affected columns, and PHI types.
Read a few rows to confirm pseudonyms (P0001, etc.), date shifts, and [REDACTED] markers
appear where expected.
Verify no original names, phone numbers, or RRN values remain.
Generate a de-identification methods paragraph for the manuscript or IRB:
Template:
> Protected health information was removed from the dataset prior to analysis using
> a rule-based de-identification tool (deidentify.py, medsci-skills) with the [COUNTRY]
> locale pattern pack. The tool scanned column names and cell values using regex patterns
> for country-specific identifiers (e.g., national ID numbers, phone numbers), email
> addresses, dates, and addresses. Each column classification was reviewed by the
> researcher in an interactive terminal session. Names were replaced with pseudonyms
> (P0001, P0002, ...), dates were shifted by a random per-patient offset (±365 days)
> preserving relative temporal intervals, and direct identifiers (phone numbers, email
> addresses, national ID numbers) were suppressed. A total of [N] cells across [M]
> columns were de-identified. The de-identification mapping file was stored separately
> under restricted access (file permissions 0600).
Customize based on the actual audit log statistics.
clean-data in the research pipeline/clean-data for data quality profiling/analyze-stats can safely process the de-identified output/write-paper Methods section should reference the de-identification process/write-protocol can use the HIPAA/PIPA reference files for protocol documentation| File | Contains PHI? | Safe for Claude? | Purpose |
|------|:------------:|:----------------:|---------|
| *_deidentified.xlsx/csv | No | Yes | De-identified data for analysis |
| mapping.json | YES | No | Original ↔ pseudonym mapping |
| audit_log.csv | No (hashes only) | Yes | What was changed and where |
| scan_report.json | No | Yes | Column classification results |
| reviewed_report.json | No | Yes | Researcher-reviewed classifications |
Supported (v1):
--locale-file with templateNOT supported (planned for v2):
Take aperivue/deidentify from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.