asgard-ai-platform/algo-seo-crawl
Implement a web crawler pipeline covering URL discovery, fetching, parsing, and storage. Use this skill when the user needs to build a site crawler, audit website structure, or collect web data systematically — even if they say 'scrape a website', 'crawl all pages', or 'site audit spider'.
npx skills add https://github.com/asgard-ai-platform/skills --skill algo-seo-crawl
A web crawler systematically traverses web pages by discovering URLs, fetching content, parsing HTML, and storing results. Uses BFS or priority-based frontier management. Performance is I/O-bound, typically limited by politeness constraints rather than compute.
Trigger conditions:
When NOT to use:
IRON LAW: Respect robots.txt and Rate Limits
A crawler MUST:
1. Parse and obey robots.txt before crawling any path
2. Enforce crawl-delay (default 1s if unspecified)
3. Identify itself with a descriptive User-Agent
Ignoring these is unethical and will get your IP blocked.
Parse seed URLs, fetch and parse robots.txt for each domain, set crawl scope (same-domain, subdomain, or cross-domain).
Gate: Valid seed URLs, robots.txt rules loaded, scope defined.
Check: no robots.txt violations in crawl log, no duplicate pages stored, all discovered URLs accounted for.
Gate: Crawl completed within scope, politeness maintained.
Return site map with pages, link graph, and extracted metadata.
{
"pages": [{"url": "...", "status": 200, "title": "...", "links_out": 15, "depth": 2}],
"metadata": {"pages_crawled": 500, "errors": 12, "duration_seconds": 300, "domain": "example.com"}
}
Input: Seed: "https://example.com", max_depth: 2, max_pages: 100
Expected: Crawl tree with homepage at depth 0, linked pages at depth 1-2, respecting robots.txt
| Input | Expected | Why |
|-------|----------|-----|
| robots.txt disallows / | Zero pages crawled | Must respect full disallow |
| Redirect loop | Stop after 5 redirects | Prevent infinite loop |
| Soft 404 (200 with error page) | Flag as soft 404 | Status code alone is insufficient |
http://Example.COM/path/ and http://example.com/path are the same URL. Normalize: lowercase host, remove default port, remove trailing slash, sort query params.references/url-normalization.mdreferences/distributed-crawl.mdTake asgard-ai-platform/algo-seo-crawl from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.