>-
npx skills add https://github.com/daymade/claude-code-skills --skill asr-transcribe-to-text
Transcribe audio/video to speaker-labeled text. Default pipeline (decoupled,
WhisperX-style): Qwen3-ASR transcribes the full audio with context intact,
mlx-whisper supplies a word-level timing lattice, pyannote supplies speaker
segments, and an aligner merges the three — the audio is never cut before ASR,
so transcription quality stays at full-audio fidelity.
| Mode | When | Speed | Cost |
|------|------|-------|------|
| Local MLX | macOS Apple Silicon | 15-27x realtime | Free |
| Remote API | Any platform, or when local unavailable | Depends on GPU | API/self-hosted |
Configuration persists in ${CLAUDE_PLUGIN_DATA}/config.json.
> Speaker labels are the default. Every run produces [start-end] SPEAKER_xx: text
> + CSV. Plain-text-only output is the opt-out (--no-diarization) for monologues,
> podcasts, or when you just want a summary — see Step 3.
>
> One-time setup for diarization: pyannote is a gated HuggingFace model — it
> needs a token once (## Speaker Diarization & Identification below). First run
> without it FAILS with setup steps; after setup, full capability is permanent
> and auto-detected.
cat "${CLAUDE_PLUGIN_DATA}/config.json" 2>/dev/null
If config exists, read values and proceed to Step 1.
If config does not exist, auto-detect platform first:
python3 -c "
import sys, platform
is_mac_arm = sys.platform == 'darwin' and platform.machine() in ('arm64', 'aarch64')
print(f'Platform: {sys.platform} {platform.machine()}')
print(f'Apple Silicon: {is_mac_arm}')
if is_mac_arm:
print('RECOMMEND: local-mlx')
else:
print('RECOMMEND: remote-api')
"
Then use AskUserQuestion with platform-aware defaults:
For macOS Apple Silicon (recommended: local):
ASR setup — your Mac has Apple Silicon, so local transcription is recommended.
Q1: Transcription mode?
A) Local MLX — runs on your Mac's GPU, no API key needed, 15-27x realtime (Recommended)
B) Remote API — send audio to a server (vLLM, Tailscale workstation, etc.)
Q2: Does your network have an HTTP proxy that might intercept traffic?
A) Yes — bypass proxy for ASR traffic (Recommended if using Shadowrocket/Clash)
B) No — direct connection
For other platforms (recommended: remote):
ASR setup — local MLX requires macOS Apple Silicon. Using remote API mode.
Q1: ASR Endpoint URL?
A) https://asr.example.com/v1/audio/transcriptions (Self-hosted remote ASR)
B) http://localhost:8002/v1/audio/transcriptions (Local ASR server)
C) Custom URL
Q2: Proxy bypass needed?
A) Yes (Recommended for Shadowrocket/Clash/corporate proxy)
B) No
Save config:
mkdir -p "${CLAUDE_PLUGIN_DATA}"
python3 -c "
import json
config = {
'mode': 'MODE', # 'local-mlx' or 'remote-api'
'model': 'MODEL_ID', # local: 'mlx-community/Qwen3-ASR-1.7B-8bit', remote: 'Qwen/Qwen3-ASR-1.7B'
'max_tokens': 200000, # local only, critical for long audio
'endpoint': 'URL', # remote only
'noproxy': True,
'max_timeout': 900 # remote only
# 'diarization_declined': True # set only after the user explicitly declines
# the pyannote setup in Step 3 — every run then warns + goes plain-text
# until an HF token appears (auto-detected)
}
with open('${CLAUDE_PLUGIN_DATA}/config.json', 'w') as f:
json.dump(config, f, indent=2)
print('Config saved.')
"
Accept local files, direct media URLs, or web/podcast episode pages.
og:audio, media tags, JSON-LD, RSS-style enclosure links), downloads URLs with atomic temp-file replacement, verifies remote Content-Length when present, computes SHA-256, and validates the result with ffprobe.uv run ${CLAUDE_SKILL_DIR}/scripts/resolve_media_input.py \
INPUT_FILE_OR_URL [INPUT_FILE_OR_URL2 ...] \
--output-dir OUTPUT_DIR \
--manifest OUTPUT_DIR/media_manifest.json
For suspicious or high-value downloads, add --decode-check to make ffmpeg decode the whole file before transcription:
uv run ${CLAUDE_SKILL_DIR}/scripts/resolve_media_input.py \
"https://www.xiaoyuzhoufm.com/episode/EPISODE_ID" \
--output-dir OUTPUT_DIR \
--manifest OUTPUT_DIR/media_manifest.json \
--decode-check
Expected output:
Downloaded ... bytes in ...s -> OUTPUT_DIR/episode-title.m4a
OUTPUT_DIR/episode-title.m4a
Use the printed local path as INPUT_AUDIO in later steps. If your runtime shows the literal ${CLAUDE_SKILL_DIR} instead of a substituted path, resolve the skill directory per the Troubleshooting entry at the bottom of this document.
For third-party public podcasts or copyrighted media, save the transcript as a local file for the user's personal analysis. Do not paste a full long transcript into chat; provide a path, previews, summaries, or short excerpts instead.
For video files (mp4, mov, mkv, avi, webm), extract as 16kHz mono WAV:
ffmpeg -i INPUT_VIDEO -vn -acodec pcm_s16le -ar 16000 -ac 1 OUTPUT.wav -y
Audio files (wav, mp3, m4a, flac, ogg) can be used directly. Get duration:
ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 INPUT_FILE
Cleanup: After transcription succeeds, delete extracted WAV files to save disk space.
Run this BEFORE transcription when either applies:
sessions into fixed-length files (e.g. TX02_MIC024_....wav, TX02_MIC025_....wav;
TX01/TX02 = DJI MIC MINI 2S internal recording — device roster and the
recorder→Feishu-Minutes paths: the meeting-ingest skill's meeting-ingest/references/architecture.md §①-L0).
Merge them and transcribe the merged file: full-audio context is the quality basis
of the decoupled pipeline (Step 3), so transcribing segments separately throws away
exactly what the architecture buys.
pitch-PRESERVED speedup cuts billed duration directly, and modern ASR does not care:
1.3x was user-verified on Feishu Minutes (2026-07-16) with no perceptible recognition
difference, and public Whisper benchmarks show no sharp WER drop until 2.0x
(≤1.5x = safe zone, ~3% WER increase at 1.5x; >2x unusable).
Use the bundled script — it merges, normalizes to 16 kHz mono, optionally speeds up,
and verifies its own output instead of trusting the ffmpeg exit code:
uv run ${CLAUDE_SKILL_DIR}/scripts/prepare_asr_input.py SEG1.wav SEG2.wav -o merged.wav # merge only
uv run ${CLAUDE_SKILL_DIR}/scripts/prepare_asr_input.py SEG*.wav -o upload.m4a --speed 1.3 # merge + quota-saving speedup
Expected output:
Merge order:
1. SEG1.wav [pcm_s24le 48000Hz ch=1 1800.14s]
2. SEG2.wav [pcm_s24le 48000Hz ch=1 1800.15s]
[OK] duration: 4946.19s vs expected 4946.18s (delta +0.00s)
[OK] boundary 1 @ 1384.7s: max_volume -15.5 dB
[info] overall: mean_volume -38.3 dB, max_volume 0.0 dB
Wrote upload.m4a
YYYYMMDD_HHMMSS timestamp embedded in their filenames whenevery file has one (recorder dumps do); otherwise the given order is kept with a note —
eyeball the printed merge order before transcribing.
otherwise); each splice gets a 10 s volume spot-check (dead air at a boundary = wrong
order or a missing segment); overall loudness prints for comparison with the source.
atempo-style pitch-preserved stretch — never sample-rate trickery,which shifts pitch and breaks both ASR accuracy and diarization voiceprints.
| Destination | Format | Why |
|---|---|---|
| Local MLX pipeline (Path A) | .wav or .m4a | Both feed the pipeline directly (m4a verified 2026-07-18: a 3-min slice transcribed cleanly). M4A is ~5x smaller — 324 MB WAV → 63 MB M4A on a 2h49m merge, duration identical to the second |
| Metered upload (Feishu Minutes, per-minute quota) | .m4a + --speed 1.3 | AAC 48k is speech-transparent for ASR, ~30% smaller than mp3 at equal speech quality; speedup cuts billed duration ~23% |
| Lossless archive | .flac | ~50% of WAV, bit-perfect |
| Only when the target rejects the above | .mp3 | Compatibility fallback |
Run the decoupled speaker pipeline — it handles dependency pins, model loading,
and the critical max_tokens parameter internally.
uv run ${CLAUDE_SKILL_DIR}/scripts/speaker_transcribe.py \
INPUT_AUDIO [INPUT_AUDIO2 ...] OUTPUT_DIR
Expected output (per file):
Device: mps
+ uv run .../transcribe_local_mlx.py ... (leg 1: full-audio text)
+ uv run .../word_timestamps_whisper.py ... (leg 2: timing lattice)
... diarization ... (leg 3: pyannote segments)
STEM: 42 turns, speakers=['SPEAKER_00', 'SPEAKER_01'], anchored_ratio=0.93
Wrote STEM.txt, STEM.csv, STEM.alignment.json
Outputs per input: <stem>.txt ([MM:SS - MM:SS] SPEAKER_xx + text),
<stem>.csv (file,start,end,duration,speaker,text — feeds review UIs and
voiceprint ID), <stem>.diarization.json, <stem>.alignment.json (provenance
+ anchored_ratio trust signal; < 0.5 prints a loud warning — verify labels
against the audio before trusting them). Intermediate legs are cached in
OUTPUT_DIR/_align/ so re-runs are cheap (--force redoes them).
Before a long first run, smoke-test the Qwen3 leg once:
uv run ${CLAUDE_SKILL_DIR}/scripts/transcribe_local_mlx.py --smoke-test
Expected output includes `Dependency stack: mlx-audio 0.3.1, mlx-lm 0.30.5,
transformers 5.0.0rc3 and Smoke test OK`. For performance details and the
max_tokens truncation issue, see references/local_mlx_guide.md.
How it works (and why): full-audio Qwen3-ASR text + mlx-whisper word
timestamps + pyannote speaker segments, aligned after the fact — the audio is
never cut before transcription, so ASR keeps full context. Architecture,
alignment algorithm, and failure modes: references/decoupled_speaker_alignment.md.
First run: pyannote needs a one-time HuggingFace token. If the script exits
with the setup hint (exit code 3), STOP and use AskUserQuestion:
Speaker diarization needs a one-time setup (gated model, free):
1. Accept terms at https://hf.co/pyannote/speaker-diarization-3.1
2. Run `huggingface-cli login` (or set HF_TOKEN)
Options:
A) Set it up now — I'll wait, then rerun with full speaker labels (Recommended)
B) Continue without speakers this time — plain text only
auto-detected every run; full capability is permanent from then on.
diarization_declined: true in config.json) andrerun the SAME command. The script detects the flag, prints a one-line warning
with the two setup steps, and auto-falls back to plain text for that run —
no need to pass --no-diarization (the fallback is automatic now, enforced in
the script not just the doc). The same warn-and-continue happens on every
later run while the token is still missing. When a token later appears,
diarization resumes automatically (the flag is ignored once a token is
present) — mention this so the user knows setup is all that's needed.
Plain-text fast path (monologue, podcast, "just summarize it"):
uv run ${CLAUDE_SKILL_DIR}/scripts/speaker_transcribe.py \
INPUT_AUDIO OUTPUT_DIR --no-diarization
Remote/pre-made ASR text (e.g. from Path B, or another ASR service): skip
the Qwen3 leg and align that text instead. --text-file pairs ONE transcript
with ONE input wav — passing multiple inputs is rejected (one transcript can't
be aligned to several files):
uv run ${CLAUDE_SKILL_DIR}/scripts/speaker_transcribe.py \
INPUT_AUDIO OUTPUT_DIR --text-file TRANSCRIPT.txt
Non-Apple-Silicon machines: the whisper timing leg is MLX-only. Without it
there is no timing lattice to align speakers onto — run with --no-diarization
and tell the user speaker mode currently requires Apple Silicon (cloud ASR with
built-in diarization, e.g. Feishu Minutes, is the no-local-GPU alternative).
Before batching many short files (promo clips, montage cuts — anything that
may contain music-only audio), read ## Batch Transcription (many short files)
below: one music-only clip can stall the whole batch for 10+ minutes.
The remote endpoint returns plain text only — speakers are added locally by
aligning that text (leg 1) with the local timing + diarization legs. So Path B
= fetch text remotely, then run Path A's pipeline with --text-file.
Health check first (skip if already verified this session):
python3 -c "
import json, subprocess, sys
with open('${CLAUDE_PLUGIN_DATA}/config.json') as f:
cfg = json.load(f)
base = cfg['endpoint'].rsplit('/audio/', 1)[0]
noproxy = ['--noproxy', '*'] if cfg.get('noproxy', True) else []
result = subprocess.run(
['curl', '-s', '--max-time', '10'] + noproxy + [f'{base}/models'],
capture_output=True, text=True
)
if result.returncode != 0 or not result.stdout.strip():
print(f'HEALTH CHECK FAILED: {base}/models', file=sys.stderr)
sys.exit(1)
print(f'Service healthy: {base}')
"
Read config and send via curl:
python3 -c "
import json, subprocess, sys, os, tempfile
with open('${CLAUDE_PLUGIN_DATA}/config.json') as f:
cfg = json.load(f)
noproxy = ['--noproxy', '*'] if cfg.get('noproxy', True) else []
timeout = str(cfg.get('max_timeout', 900))
audio_file = 'AUDIO_FILE_PATH'
output_json = tempfile.mktemp(suffix='.json', prefix='asr_')
result = subprocess.run(
['curl', '-s', '--max-time', timeout] + noproxy + [
cfg['endpoint'],
'-F', f'file=@{audio_file}',
'-F', f'model={cfg[\"model\"]}',
'-o', output_json
], capture_output=True, text=True
)
with open(output_json) as f:
data = json.load(f)
if 'text' not in data:
print(f'ERROR: {json.dumps(data)[:300]}', file=sys.stderr)
sys.exit(1)
text = data['text']
print(f'Transcribed: {len(text)} chars', file=sys.stderr)
print(text)
os.unlink(output_json)
" > OUTPUT.txt
Then attach speakers locally (Apple Silicon + pyannote token required):
uv run ${CLAUDE_SKILL_DIR}/scripts/speaker_transcribe.py \
INPUT_AUDIO OUTPUT_DIR --text-file OUTPUT.txt
If remote health check fails, diagnose in order:
ping -c 1 HOST or tailscale status | grep HOSTtailscale ssh USER@HOST "curl -s localhost:PORT/v1/models"--noproxy '*' toggledAfter transcription, check for truncation — the most common failure mode:
max_tokens was exhaustedanchored_ratio should be ≥ 0.5 (the script warns when lower), the speaker count should be plausible for the recording (a two-person interview showing 5 speakers, or a monologue split into 2+, means diarization over-segmented — see references/speaker_diarization.md for when to distrust labels)If truncated or wrong, use AskUserQuestion:
Transcription may be truncated:
- Expected: ~[N] chars for [M] minutes of audio
- Got: [actual] chars ([pct]% of expected)
- Last line: "[last 100 chars...]"
Options:
A) Retry with higher max_tokens (current: [N], try: [N*2])
B) Switch mode — try [local/remote] instead
C) Save as-is — the output looks complete to me
D) Abort
If single remote request fails (timeout, OOM), fall back to chunked transcription:
python3 ${CLAUDE_SKILL_DIR}/scripts/overlap_merge_transcribe.py \
--config "${CLAUDE_PLUGIN_DATA}/config.json" \
INPUT_AUDIO OUTPUT.txt
Splits into 18-minute chunks with 2-minute overlap, merges using punctuation-stripped fuzzy matching. See references/overlap_merge_strategy.md for algorithm details.
For local MLX mode, overlap-merge is unnecessary — the bundled script handles chunking internally with max_tokens=200000.
ASR output always contains recognition errors — homophones, garbled technical terms, broken sentences. After successful transcription, proactively suggest running the transcript-fixer skill on the output:
Transcription complete: [N] chars saved to [output_path].
ASR output typically contains recognition errors (homophones, garbled terms, broken sentences).
Would you like me to run /daymade-audio:transcript-fixer to clean up the text?
Options:
A) Yes — run daymade-audio:transcript-fixer on the output now (Recommended)
B) No — the raw transcription is good enough for my needs
C) Later — I'll run it myself when ready
If the user chooses A, invoke the transcript-fixer skill with the output file path. The two skills form a natural pipeline: transcribe → correct → review.
rm "${CLAUDE_PLUGIN_DATA}/config.json"
Then re-run Step 0.
Passing many files to one transcribe_local_mlx.py invocation is efficient (model loads once) — but only when every file contains actual speech. If the batch may include music-only / BGM-only clips (short promo videos, montage clips with subtitles instead of voiceover), do NOT batch them in one process:
max_tokens=200000 — one such file can stall for 10+ minutes and starve the whole batch.timeout 240 / perl -e 'alarm 240; exec @ARGV' around each invocation, skip on timeout, second pass for failures). A stuck file then costs 4 minutes, not the batch.--max-tokens 3000: real speech in a short clip fits comfortably; a looping file gets truncated output you can classify.len(set(words))/len(words) < 0.06 on a 40+ char output), the clip almost certainly has no voiceover — label it as such rather than delivering the loop text. (Downstream OCR of on-screen captions is the actual fix for subtitle-only videos.)mlx-whisper's word timing is now the timing leg of the default speaker pipeline (leg 2 — scripts/word_timestamps_whisper.py runs it automatically). This section is for using word timestamps STANDALONE: subtitle generation, aligning narration to shot boundaries, per-clip captioning.
Qwen3-ASR is an LLM-decoder ASR: it emits plain text with no alignment information, on both local and remote paths. When the task needs to know *when* each word is spoken, use mlx-whisper with word_timestamps=True. Whisper's cross-attention word alignment is the de-facto local solution for this class of task.
Key facts (full recipe in references/whisper_word_timestamps.md):
mlx-community/whisper-large-v3-turbo (~1.6GB). Its Chinese WER is higher than Qwen3-ASR for pure transcription, but for alignment tasks Qwen3-ASR is not an option at all; prime domain terms via initial_prompt.select='gt(scene,0.3)') for the visual side; avoid PySceneDetect on non-ASCII paths.Speaker labels are the DEFAULT output of Step 3 (decoupled architecture:
full-audio Qwen3-ASR text + whisper timing lattice + pyannote segments,
aligned — never cut-then-transcribe). This section covers the pieces.
scripts/speaker_transcribe.py runs all three legs +alignment in one command and writes the speaker-labeled transcript + CSV.
Architecture, alignment algorithm, trust signals (anchored_ratio), and
failure modes: references/decoupled_speaker_alignment.md. Production
pitfalls (over-segmentation, mic-domain effects, when to distrust labels):
references/speaker_diarization.md.
scripts/diarize_speakers.py emits just thespeaker × time segments (no transcription).
scripts/speaker_transcribe_cascade.py is the oldcut-then-transcribe variant (diarize → slice audio per turn → ASR each
slice). It breaks ASR context at every cut and lowers text quality; kept
only for extremely noisy / heavy-overlap audio where per-slice isolation of
a dominant near-field speaker beats full-audio ASR. Everything else uses
the decoupled default.
(SPEAKER_00…) and per-file. To map them to real names, unify a speaker
across files, or collapse diarization's over-segmentation, use CAM++
voiceprints via scripts/voiceprint_id.py. Recipe **and the critical
acoustic-domain caveat** — a voiceprint built from one mic type matches the
same person on a different mic far less well:
references/voiceprint_speaker_id.md.
One-time pyannote setup (gated model): accept terms at
hf.co/pyannote/speaker-diarization-3.1, then huggingface-cli login once
(or set HF_TOKEN). Auto-detected on every run afterward.
After diarization you get a CSV per file (file,start,end,duration,speaker,text). The bundled audit HTML generator turns those CSVs into a single, reader-first review page with audio playback, per-turn flags/notes, speaker aliasing, and export.
Generate it from a speaker-transcribe output directory:
uv run ${CLAUDE_SKILL_DIR}/scripts/generate_audit_html.py \
OUTPUT_DIR \
--output OUTPUT_DIR/audit/index.html \
--audio-dir /path/to/original/audio
Defaults assume a flat layout under PROJECT_DIR: PROJECT_DIR/*.csv transcripts, PROJECT_DIR/*.diarization.json, and the original audio files placed next to the outputs. speaker_transcribe.py itself writes the CSV, TXT, and diarization files flat under its OUTPUT_DIR. If your project uses a different structure, override any of those paths:
uv run ${CLAUDE_SKILL_DIR}/scripts/generate_audit_html.py \
/path/to/project \
--output /path/to/project/audit/index.html \
--csv-dir /path/to/project/csv \
--txt-dir /path/to/project/txt \
--diarization-dir /path/to/project/diarization \
--audio-dir /path/to/project/audio \
--original-dir /path/to/project/original \
--manifest /path/to/project/manifest.json \
--title "Project Audit" \
--subtitle "Speaker-labeled transcript review" \
--storage-key "project-audit" \
--known-speaker "Speaker A" \
--known-speaker "Speaker B"
Key CLI options:
| Option | Meaning |
|--------|---------|
| project_dir | Base project directory (required) |
| --output | Where to write index.html |
| --csv-dir | Directory containing *.csv transcript files |
| --txt-dir | Directory containing *.txt plain-text transcripts (optional) |
| --diarization-dir | Directory containing *.diarization.json files |
| --audio-dir | Directory containing playback audio files |
| --original-dir | Directory containing original source media (optional) |
| --manifest | JSON manifest mapping file IDs to metadata (optional) |
| --title / --subtitle | Page title and subtitle |
| --storage-key | localStorage namespace for state persistence |
| --known-speaker | Repeatable; "Name" auto-assigns a color, "Name=#hex" sets one explicitly |
| --material-final / --material-rough | Repeatable material classification labels used for filtering |
The output is a single self-contained HTML file with no external dependencies. Open it in a browser to review, flag, and annotate turns; the export button produces a report of all flagged rows with reasons and notes.
If model loading fails with an error like:
AttributeError: 'str' object has no attribute '__module__'
the agent is probably using an unpinned or stale copy of the local MLX script. The known-good stack is:
mlx-audio 0.3.1
mlx-lm 0.30.5
transformers 5.0.0rc3
Run the bundled --smoke-test command and confirm the dependency stack line matches. Do not start a long transcription until the smoke test succeeds.
${CLAUDE_SKILL_DIR} is not substitutedScript paths in this skill use ${CLAUDE_SKILL_DIR} — the skill's own directory, which Claude Code substitutes when the skill loads. If a command reaches you with the literal ${CLAUDE_SKILL_DIR} (some runtimes don't substitute), resolve the skill directory in this order:
Base directory for this skill: <path> → <path> is the skill directory.find ~/.claude ~/.claude-profiles ~/.codex ~/workspace -maxdepth 7 -type d -name asr-transcribe-to-text 2>/dev/null | head -5
Substitute the resolved absolute path for ${CLAUDE_SKILL_DIR} everywhere in this document.
Scripts:
resolve_media_input.py — Resolve local paths, direct media URLs, and podcast/web pages into validated local media filesprepare_asr_input.py — Merge multi-segment recordings + normalize for ASR (16 kHz mono), optional pitch-preserved speedup for metered uploads; self-verifies duration math and splice boundariestranscribe_local_mlx.py — Local MLX transcription (macOS ARM64, PEP 723 deps)speaker_transcribe.py — DEFAULT pipeline: decoupled multi-speaker transcription (full-audio Qwen3-ASR + whisper word timing + pyannote diarization, aligned) → speaker-labeled transcript + CSV; --no-diarization plain-text fast path; --text-file for remote/pre-made ASR textalign_speakers.py — Decoupled alignment core (stdlib): maps full transcript onto whisper word lattice + pyannote segments; usable standalone for debuggingword_timestamps_whisper.py — mlx-whisper word-level timestamps → JSON timing lattice (Apple Silicon)speaker_transcribe_cascade.py — LEGACY cut-then-transcribe variant (extremely noisy / heavy-overlap audio only)diarize_speakers.py — Speaker diarization alone (pyannote 3.1 @ MPS) → per-segment JSONvoiceprint_id.py — CAM++ voiceprint enroll/match: map anonymous SPEAKER_xx to real namesoverlap_merge_transcribe.py — Chunked transcription with overlap merge (remote API fallback)generate_audit_html.py — Build a self-contained HTML audit/review page from speaker-transcribe CSV outputsReferences:
decoupled_speaker_alignment.md — The default architecture: why decouple, alignment algorithm, trust signals, failure modesspeaker_diarization.md — Production pitfalls: over-segmentation, mic-domain effects, when to distrust labels; legacy cascade notesvoiceprint_speaker_id.md — CAM++ speaker ID: enroll/match, threshold+margin gates, the acoustic-domain caveat, bootstraplocal_mlx_guide.md — Performance benchmarks, max_tokens truncation, model compatibilitywhisper_word_timestamps.md — mlx-whisper word timing: the timing leg of the default pipeline; standalone subtitle/AV-alignment recipeoverlap_merge_strategy.md — Why naive chunking fails, fuzzy merge algorithmIntegration with protocols.io API for managing scientific protocols. This skill should be used when working with protocols.io to search, create, update, or publish protocols; manage protocol steps and materials; handle discussions and comments; organize workspaces; upload and manage files; or integrate protocols.io functionality into workflows. Applicable for protocol discovery, collaborative protocol development, experiment tracking, lab protocol management, and scientific documentation.
Analyzes job descriptions and generates tailored resumes that highlight relevant experience, skills, and achievements to maximize interview chances
Generate Excalidraw diagrams from natural language descriptions. Use when asked to "create a diagram", "make a flowchart", "visualize a process", "draw a system architecture", "create a mind map", or "generate an Excalidraw file". Supports flowcharts, relationship diagrams, mind maps, and system architecture diagrams. Outputs .excalidraw JSON files that can be opened directly in Excalidraw.
Build and distribute Expo development clients locally or via TestFlight
Use when you have a written implementation plan to execute in a separate session with review checkpoints
Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
Benchling R&D platform integration. Access registry (DNA, proteins), inventory, ELN entries, workflows via API, build Benchling Apps, query Data Warehouse, for lab data management automation.
Comprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use bioservices.
Take daymade/asr-transcribe-to-text from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.