Web Content Extractor runs on your own machine — the client starts it, so there is no endpoint to ping. 81 installs a week from npm. Last commit 2 Apr 2026.
Extract and process web content into clean, structured formats optimized for LLMs.
Today is the operative word: we check Web Content Extractor every 15 minutes and re-read its code on every release. Watch it and you find out the day that stops being true.
This server runs on your own machine — install it with the package manager and the client starts it for you. Package name taken from the official registry entry.
claude mcp add web-content-extractor -- npx -y @agenson-horrowitz/web-content-extractor-mcp
{
"mcpServers": {
"web-content-extractor": {
"args": [
"-y",
"@agenson-horrowitz/web-content-extractor-mcp"
],
"command": "npx"
}
}
}
[mcp_servers.web-content-extractor]
command = "npx"
args = ["-y", "@agenson-horrowitz/web-content-extractor-mcp"]
{
"mcpServers": {
"web-content-extractor": {
"args": [
"-y",
"@agenson-horrowitz/web-content-extractor-mcp"
],
"command": "npx"
}
}
}
{
"mcpServers": {
"web-content-extractor": {
"args": [
"-y",
"@agenson-horrowitz/web-content-extractor-mcp"
],
"command": "npx"
}
}
}
Extract and structure any web page into clean JSON. Free tier: 10/day.
Turn documents into structured data: parse, extract, classify, split, and fill PDF forms.
Extract structured web data (product, article, company) for AI agents via Agent API Gateway
Scrape any website, extract structured data, and collect web content at scale with AI agents
Extract structured data from PDFs and documents with Suparse AI OCR.
JSON contract registry and validator for structured LLM and agent outputs.
Operating contracts for AI agents. Structured instructions compiled into system prompts.
Give AI agents clean, LLM-ready web data — scrape any URL to markdown or extract structured JSON.
Answers built from our own checks of this server.