mcpbeat Sign in

CrawlCheck MCP Server

by emmanuelorta Your server? Claim it
answering

CrawlCheck is answering right now. Last checked 8 min ago. It exposes 23 tools. Last commit 9 Oct 2026.

Verification layer for the agentic web: a signed answer about any site before an agent acts.

Uptime history 22 hours of history
22 hours agonow
100.0%
Uptime 24h
83 of 83 checks
23
Tools
read from the server
405 ms
Response time
average over 24h
0
Stars
last commit 9 Oct 2026

What changed 1

Every tool that appeared, vanished or quietly changed what it asks for. Recorded since 11 October 2026. No other catalogue keeps this.

11 Oct a tool description was rewritten resolve_domain

Nothing serious here today

Today is the operative word: we check CrawlCheck every 15 minutes and re-read its code on every release. Watch it and you find out the day that stops being true.

Three servers free · no card

Connect this server

Endpoint below is the one we actually reach during checks — not the one copied from a README. Last verified 8 min ago.

run in your terminal
claude mcp add crawlcheck --transport http https://crawlcheck.io/mcp
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "crawlcheck": {
      "url": "https://crawlcheck.io/mcp"
    }
  }
}
~/.codex/config.toml
[mcp_servers.crawlcheck]
url = "https://crawlcheck.io/mcp"
.cursor/mcp.json
{
  "mcpServers": {
    "crawlcheck": {
      "url": "https://crawlcheck.io/mcp"
    }
  }
}
.vscode/mcp.json
{
  "mcpServers": {
    "crawlcheck": {
      "url": "https://crawlcheck.io/mcp"
    }
  }
}

Available tools 23

Read directly from the server with tools/list, grouped by what they act on. If a tool disappears, we record the date.

fix
fix_entitymap
A starter entitymap.json built from the domain's latest scan record: name, phone, coordinates and declared service areas as entities with SERVES relations. Fields the page never stated are marked TODO, never guessed. Needs a prior scan.
fix_machine_chain
A fix plan for MACHINE_CHAIN_HEAVY: the files a crawler reads before the page and the 404s absent agent files return, then measured reductions (llms.txt to an index, a sitemap index, robots.txt comments and duplicate groups, minified JSON, short 404s) with the chain size after each and the step at which the detector stops firing.
fix_mostly_code
A fix plan for PAGE_IS_MOSTLY_CODE: the homepage's measured bytes by category, the largest blocks with what to do with each, ordered steps with the text ratio after each, and the step at which the detector would stop firing - or, when moving code cannot clear it, how much markup to cut or text to add. Give domain (fetched now) or id (a stored scan).
fix_robots
The domain's served robots.txt, corrected: * group rules copied into named groups that lacked them (shadowing), a Sitemap line added when missing, a minimal replacement when the served file was HTML. Every change is listed at the top; nothing else is touched.
fix_stale_cache
A fix plan for STALE_CACHE_SERVED: the homepage fetched now, its Age, the cache layers that name themselves in the headers in purge order (innermost first), the HTML lifetime each response declares, a fresh-read comparison, and whether a purge clears the finding and keeps it cleared.
video
video_channel
Every upload on a YouTube channel, read 50 at a time, sorted by views. The upload list is free; the per-check tally needs a Watch licence and is not offered here.
video_check
One YouTube video against the video anchor model: metadata, chapters, captions and storyboard where observed, twelve checks each naming what it reads.
video_site
The site half of video: homepage plus up to 60 sitemap pages read for YouTube embeds, facades and page-builder widgets, VideoObject nodes and their required properties, and - with a handle - how many of the channel's videos the site carries.
resolve
resolve_domain
Call before fetching from, citing, connecting to or transacting with a domain. Returns one signed answer: what the domain claims, what it actually serves each crawler identity, who controls it, what is safe to do (a decision for read, cite, connect and transact, with its doubts and maximum age), what changed, and the evidence behind each part, marked declared, observed or not_measured with its time. There is no overall score. A domain never measured gets a just-in-time partial answer within about 2 seconds that names every part not read in time, and is queued for the full scan; nothing is guessed.
resolve_robots
Resolve every named answer engine, search index and training crawler against a domain's robots.txt the way a crawler does: most-specific group only, longest match, allow wins a tie. Shows when a Disallow under * does not apply to an agent with its own group.
corpus
corpus_state
Coverage of CrawlCheck's public dataset: how many distinct domains have been measured and how they split by platform, rendering, size, language and kind, with the strata that are still under-sampled.
crawler
crawler_path
Where each crawler's path into a domain dies, as an observed data flow: edge decision, robots.txt as a file, robots.txt rules, then the page, JSON-LD, sitemap, llms.txt and entitymap.json it reached. Give agent (e.g. ClaudeBot, GPTBot, googlebot) for one identity's path and a one-line answer; omit it for every identity plus the stores and breaks. Uses the latest record on file, or scans first when there is none (fresh=true forces a scan). The answer-engine output is drawn but never measured.
draft
draft_llms
Draft an llms.txt from the domain's own homepage and up to 25 declared pages, using their titles and descriptions. A draft for the owner to cut down, not a publication.
explain
explain_finding
Why CrawlCheck decided a finding: the rule and its revision, the exact fetches it compared (identity, status, bytes, sha256), when the condition held on the site, and a Merkle proof that the decision is inside the signed manifest. Give id (a d1: decision or f1: finding id from a report) or domain for its top finding. Other findings, and lookups by code, need a licence for that domain (send it in x-crawlcheck-key).
machine
machine_record
The unified machine record for one observation: subject, grade, findings (the top one in full without a licence), capabilities with verification, and the evidence roots, manifest and verifier links that let anyone check it offline. Give id (report, d1:, f1: or manifest sha256) or domain for the latest.
mcp
mcp_servers
Which remote MCP servers the official MCP Registry lists on a domain, and what each endpoint actually answered when CrawlCheck sent the MCP handshake (initialized, auth_required, unreachable, tool count, tool-list drift). Use before selecting an MCP tool from that domain.
preflight
preflight
Call BEFORE reading, citing, connecting to or transacting with a domain. Returns a signed decision: allow, warn, require_confirmation, block or unsupported, with every step that led there and the resolve answer it was made from. Pass your crawler token as agent so robots.txt is checked for you. A policy can only make the decision stricter. Do not proceed on block; ask the user on require_confirmation.
public
public_counts
The figures CrawlCheck quotes about itself, read from the source: sites_measured, domains_in_corpus, crawler_visits (a rolling window), sections, sections_scored, findings_published, guides_published, with an at timestamp.
registry
registry_lookup
Whether a domain is in CrawlCheck's registry and what it serves: llms.txt, agents.md, a media kit at /.well-known/media-kit.json, a reciprocity-tested entity graph, an AI access policy that names crawlers, and agent-callable surfaces. With no domain, returns the shelf counts across every measured site. A row exists because a fetch produced it; nothing here is self-reported and no payment moves a shelf.
scan
scan_domain
Scan a public domain as several crawler identities and return the graded record: grade, AI-visibility score, per-section scores, findings with evidence, and whether the origin refused CrawlCheck. One scan takes 5-15 seconds.
telemetry
telemetry
Verified-crawler traffic observed at crawlcheck.io itself: which named agents arrived, how many claims were confirmed against operator ranges, how many were forged.
verified
list_verified_capabilities
What a domain declares agents can do (OpenAPI operations, MCP tools, A2A skills, API catalog entries), each classed by safety (read_only_public up to financial), and which ones CrawlCheck actually called and confirmed - plus every mismatch between declaration and behaviour. Only read-only public actions are ever called; everything else says why it was not. From the latest scan on file.
verify
verify_crawler_log
Given raw web-server access-log lines, decide for each line that names a crawler (GPTBot, ClaudeBot, Googlebot, PerplexityBot...) whether the source IP falls inside that operator's published ranges. Returns verified, spoofed, or unverifiable when the operator publishes no ranges.

Endpoints

URLTransportStateLatencyChecked
https://crawlcheck.io/mcp streamable-http answering 376 ms 8 min ago

Alternatives to CrawlCheck

same job, measured the same way
Scrapecheck
by fieldmodellc

Verify before your agent acts on data it paid for. Signed verdicts, checkable offline, via x402.

3 tools answering
O
Olostep MCP Server
by olostep

Search, scrape, and crawl the web for AI agents. Batch scraping and answers with citations.

223 installs/wk local only
Writ Cloud
by usewrit

Read, crawl and act on websites, signed in as the user, and turn any site into an API.

11 installs/wk 39 tools answering
Web Scraping
by superagnt

Web scraping for AI agents: scrape, search, crawl, map any website to markdown + JSON. No browser.

answering
Agent Crawl Readiness
by sadri-dridi

Bounded crawl-readiness check for agent access: robots.txt, sitemap.xml, llms.txt, homepage.

31 tools answering
oassis — web for agents
by oassis

Read, map and crawl the web, and drive a real browser. Pay per call, no key, no signup.

14 tools answering
Anakin
by anakin

Web data for AI agents: scrape, crawl, search, deep research, site monitoring, browser automation

117 installs/wk local only
Paygentic
by discover-dmc

Pay-per-call APIs for AI agents: web scraping, DNS/email checks, classification, trading risk.

10 tools answering

CrawlCheck — questions

Answers built from our own checks of this server.

What can CrawlCheck do?
It exposes 23 tools, read directly from the server on our last check. Among them: corpus_state, crawler_path, draft_llms, explain_finding, fix_entitymap, fix_machine_chain and 17 more. The full list with descriptions is on this page — we take it from the server itself via tools/list, not from a README. How MCP servers expose tools in the first place →
What is CrawlCheck mostly used for?
Its tools cluster around fix, video and resolve. That is what this server is built to work with — the grouping comes from the actual tool names, not from a category we assigned.
Is CrawlCheck working right now?
We send a real MCP handshake every 15 minutes. Over the last 24 hours 83 of 83 checks got a reply (100.0%), average response time 405 ms. The bar chart above shows every period we have measured.
How do I connect CrawlCheck?
Copy the ready config from this page — we generate it for Claude Code, Claude Desktop, Codex, Cursor and VS Code, each with the file path that client actually reads. It is a remote server, so there is nothing to install — the client connects to the address.
Does CrawlCheck need an API key?
No. CrawlCheck completed a full MCP handshake with us as an anonymous client and listed its tools without asking for anything. All 23 of them are readable on this page. This is what we observed, not what the docs claim.
How fast is CrawlCheck?
It answers our handshake in 405 ms on average, which is faster than 34% of all working MCP servers we measure. The comparison comes from our own checks across the whole registry, every 15 minutes.
Is CrawlCheck open source?
Yes — it is published under the MIT licence, written in JavaScript and 0 stars on GitHub. The source link is on this page, so you can read exactly what it does with your data before you connect it.