guia-matthieu/web-scraper
Extract structured data from websites. Use when: collecting competitor pricing; scraping product listings; extracting contact information; gathering research data; monitoring website changes
npx skills add https://github.com/guia-matthieu/clawfu-skills --skill web-scraper
> Extract structured data from websites using BeautifulSoup and requests - turn any webpage into usable data.
| Claude Does | You Decide |
|-------------|------------|
| Structures analysis frameworks | Strategic priorities |
| Synthesizes market data | Competitive positioning |
| Identifies opportunities | Resource allocation |
| Creates strategic options | Final strategy selection |
| Suggests implementation approaches | Execution decisions |
pip install beautifulsoup4 requests pandas click lxml
python scripts/main.py scrape https://example.com --selector "h1,h2,p"
python scripts/main.py scrape https://example.com --selector ".product-price"
python scripts/main.py links https://example.com
python scripts/main.py links https://example.com --internal-only
python scripts/main.py emails https://example.com
python scripts/main.py emails https://example.com --depth 2
python scripts/main.py structured https://example.com/article --schema article
python scripts/main.py structured https://example.com/product --schema product
python scripts/main.py scrape https://competitor.com/pricing --selector ".price,.plan-name"
# Output:
# Extracted 6 elements
# 1. Starter - $29/mo
# 2. Pro - $99/mo
# 3. Enterprise - Contact us
python scripts/main.py structured https://blog.example.com/post --schema article
# Output: article_data.json
# {
# "title": "How to Scale Your Startup",
# "author": "Jane Doe",
# "date": "2024-01-15",
# "content": "...",
# "word_count": 1523
# }
| Selector | Description | Example |
|----------|-------------|---------|
| tag | Element type | h1, p, div |
| .class | Class name | .price, .title |
| #id | Element ID | #main-content |
| tag.class | Tag with class | div.product |
| tag[attr] | Has attribute | a[href] |
| parent > child | Direct child | ul > li |
| tag1, tag2 | Multiple | h1, h2, h3 |
category: automation
subcategory: data-extraction
dependencies: [beautifulsoup4, requests, pandas]
difficulty: intermediate
time_saved: 5+ hours/week
Take guia-matthieu/web-scraper from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.