telagod/processing-pdfs
Processes PDF files. Extracts text and tables, fills forms, merges and splits documents, batch-processes files, converts to images, and generates PDFs programmatically. Use when working with .pdf files. Do NOT use for Word documents, spreadsheets, or presentations.
npx skills add https://github.com/telagod/code-abyss --skill processing-pdfs
Essential PDF operations using Python libraries and CLI tools.
| Task | Best Tool | Reference |
|------|-----------|-----------|
| Merge / split / metadata / rotate | pypdf | recipes.md |
| Extract text (layout preserved) | pdfplumber | recipes.md |
| Extract tables | pdfplumber | recipes.md |
| Create new PDF | reportlab | recipes.md |
| Batch CLI ops | qpdf / pdftk | recipes.md |
| OCR scanned PDFs | pytesseract + pdf2image | advanced.md |
| Add watermark / extract images / encrypt | pypdf / pdfimages | advanced.md |
| Fill PDF forms | pdf-lib / pypdf | FORMS.md |
| Advanced pypdfium2 / pdf-lib JS | — | REFERENCE.md |
from pypdf import PdfReader
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
text = "".join(page.extract_text() for page in reader.pages)
| Library | Use for |
|---------|---------|
| pypdf | Merge, split, metadata, encryption, rotation |
| pdfplumber | Text extraction with layout, tables |
| reportlab | Generate PDFs programmatically |
| pdf2image + pytesseract | OCR scanned documents |
| qpdf / pdftk (CLI) | Batch ops, no Python needed |
Take telagod/processing-pdfs from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.