trackmcp
Back to directory

Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.

2,319 stars RustOthers Updated Sep 4, 2026
ai-agentsclillmmarkdownmcprustself-hostedweb-extractionweb-scrapingweb-crawlerai-scrapinghtml-to-markdownmcp-servertls-fingerprintingfirecrawl-alternativeapify-alternativescraperapi-alternativescrapingbee-alternativecrawl4ai-alternativejina-alternative

Documentation

English |

webclaw

Turn websites into clean markdown, JSON, and LLM-ready context.

CLI, MCP server, REST API, and SDKs for AI agents and RAG pipelines.


Most web scraping tools give your agent one of two bad outputs:

  • a blocked page, login wall, or empty app shell
  • raw HTML full of nav, scripts, styling, ads, and duplicated boilerplate

webclaw.io is the hosted web extraction API for webclaw. This repo contains the open-source CLI, MCP server, extraction engine, and self-hostable server.

webclaw turns a URL into clean content your tools can actually use.

bash
webclaw https://example.com --format markdown
md
# Example Domain

This domain is for use in illustrative examples in documents.

You may use this domain in literature without prior coordination or asking for permission.

Use it from the terminal, wire it into Claude/Cursor through MCP, call the hosted API from your app, or self-host the OSS server.


Install

Agent setup

The fastest way to connect webclaw to Claude Code, Claude Desktop, Cursor, Windsurf, OpenCode, Codex CLI, and other MCP-compatible tools:

bash
npx create-webclaw

The installer detects supported clients and configures the MCP server for you.

Homebrew

bash
brew tap 0xMassi/webclaw
brew install webclaw

Prebuilt binaries

Download macOS, Linux, and Windows binaries from GitHub Releases.

Docker

bash
docker run --rm ghcr.io/0xmassi/webclaw https://example.com

Cargo

bash
cargo install --git https://github.com/0xMassi/webclaw.git webclaw-cli
cargo install --git https://github.com/0xMassi/webclaw.git webclaw-mcp

If building from source fails because native build tools are missing, install the platform prerequisites:

OSCommand
Debian / Ubuntu`sudo apt install -y pkg-config libssl-dev cmake clang git build-essential`
Fedora / RHEL`sudo dnf install -y pkg-config openssl-devel cmake clang git make gcc`
Arch`sudo pacman -S pkg-config openssl cmake clang git base-devel`
macOS`xcode-select --install`

Quick Start

Scrape one page

bash
webclaw https://stripe.com --format markdown

Return LLM-optimized text

bash
webclaw https://docs.anthropic.com --format llm

Keep only the main content

bash
webclaw https://example.com/blog/post --only-main-content

Include or exclude selectors

bash
webclaw https://example.com \
  --include "article, main, .content" \
  --exclude "nav, footer, .sidebar, .ad"

Crawl a documentation site

bash
webclaw https://docs.rust-lang.org --crawl --depth 2 --max-pages 50

Workflow examples

Extract brand assets

bash
webclaw https://github.com --brand

Compare a page over time

bash
webclaw https://example.com/pricing --format json > pricing-old.json
webclaw https://example.com/pricing --diff-with pricing-old.json

MCP Server

webclaw ships with an MCP server for AI agents.

Zero-install — point any MCP client at the npx launcher:

json
{
  "mcpServers": {
    "webclaw": {
      "command": "npx",
      "args": ["-y", "@webclaw/mcp"]
    }
  }
}

Or run `npx create-webclaw` to auto-detect your AI tools and write their configs for you.

Then ask your agent things like:

text
Scrape these competitor pricing pages and summarize the differences.
text
Crawl this documentation site and prepare clean context for a RAG index.
text
Extract the brand colors, fonts, and logos from this company website.

Use as an agent skill

Add webclaw to Claude Code, Cursor, Windsurf, and other MCP agents in one command:

bash
npx skills add 0xMassi/webclaw-skill

Your agent gets scrape, crawl, map, extract, summarize, diff, brand, and search

as native tools. Most sites extract locally with no API key. Set `WEBCLAW_API_KEY`

to handle bot-protected and JavaScript-rendered pages.

Find it on skills.sh.


Tools

ToolWhat it doesLocal
`scrape`Extract one URL as markdown, text, JSON, LLM format, or HTMLYes
`crawl`Follow same-origin links and extract discovered pagesYes
`map`Discover URLs without extracting every pageYes
`batch`Scrape multiple URLs in parallelYes
`extract`Convert page content into structured dataYes, with local or configured LLM
`summarize`Summarize a pageYes, with local or configured LLM
`diff`Compare page content snapshotsYes
`brand`Extract colors, fonts, logos, and metadataYes
`search`Search the web and scrape resultsHosted API
`research`Multi-source research workflowHosted API

SDKs

bash
npm install @webclaw/sdk
pip install webclaw
go get github.com/0xMassi/webclaw-go

TypeScript

ts
import { Webclaw } from "@webclaw/sdk";

const client = new Webclaw({ apiKey: process.env.WEBCLAW_API_KEY! });

const page = await client.scrape({
  url: "https://example.com",
  formats: ["markdown"],
  only_main_content: true,
});

console.log(page.markdown);

Python

python
from webclaw import Webclaw

client = Webclaw(api_key="wc_your_key")

page = client.scrape(
    "https://example.com",
    formats=["markdown"],
    only_main_content=True,
)

print(page.markdown)

cURL

bash
curl -X POST https://api.webclaw.io/v1/scrape \
  -H "Authorization: Bearer $WEBCLAW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "formats": ["markdown"],
    "only_main_content": true
  }'

Output Formats

FormatUse it when you need
`markdown`Clean page content with structure preserved
`llm`Compact context for agents and RAG pipelines
`text`Plain text with minimal formatting
`json`Structured metadata, links, images, and extracted fields
`html`Cleaned HTML for custom processing

Local First, Hosted When Needed

The CLI and MCP server work locally without an account for the core extraction path.

Use the hosted API at webclaw.io when you need:

  • protected-site access without managing infrastructure
  • JavaScript rendering
  • async crawl and research jobs
  • web search
  • watches and production usage tracking
  • SDKs for application code
bash
export WEBCLAW_API_KEY=wc_your_key

webclaw https://example.com --cloud

What You Can Build

Use caseExample
AI agent web accessGive Claude, Cursor, or another MCP client clean page context
RAG ingestionCrawl docs, help centers, blogs, and knowledge bases
Competitor monitoringTrack pricing pages, changelogs, docs, and product pages
Structured extractionTurn messy pages into typed JSON for automations
Research workflowsSearch, scrape, summarize, and cite multiple sources
Brand intelligenceExtract logos, colors, fonts, and social metadata

Architecture

text
webclaw/
  crates/
    webclaw-core     HTML to markdown, text, JSON, and LLM-ready output
    webclaw-fetch    Fetching, crawling, batching, and mapping
    webclaw-llm      Local and hosted LLM provider support
    webclaw-pdf      PDF text extraction
    webclaw-mcp      MCP server for AI agents
    webclaw-cli      Command-line interface

`webclaw-core` is pure extraction logic: no network I/O, small surface area, and usable independently from the fetching layer.


Configuration

VariableDescription
`WEBCLAW_API_KEY`Hosted API key
`OLLAMA_HOST`Ollama URL for local LLM features
`OPENAI_API_KEY`OpenAI-compatible LLM provider key
`OPENAI_BASE_URL`OpenAI-compatible base URL
`ANTHROPIC_API_KEY`Anthropic-compatible LLM provider key
`ANTHROPIC_BASE_URL`Anthropic-compatible base URL
`ORCAROUTER_API_KEY`OrcaRouter LLM provider key
`ORCAROUTER_BASE_URL`OrcaRouter base URL (defaults to https://api.orcarouter.ai/v1)
`WEBCLAW_PROXY`Single proxy URL
`WEBCLAW_PROXY_FILE`Proxy pool file

Contributing

The most useful contributions right now are practical and small:

  • add examples for real agent and RAG workflows
  • improve SDK snippets
  • report pages that extract poorly
  • add failing fixtures for messy HTML
  • improve docs for MCP clients and local setup
  • test the CLI on more Linux/macOS environments

Good first places to start:

If a page extracts badly, include:

text
URL:
Command or API request:
Expected output:
Actual output:
Format used: markdown / llm / text / json / html
CLI, MCP, SDK, or API:

Please remove secrets, cookies, private tokens, and customer data from logs before posting.


Strategic Partner

picks the right one instead of leaving a white

wordmark invisible on light mode. Colours unmodified, per their

media kit. -->

.

Give real-time data to your AI agents and enhance their responses with SerpApi’s

structured search engine results. SerpApi supports webclaw as a Strategic Partner.


Infrastructure Partner

ColdProxy supports webclaw as an Infrastructure Partner, providing residential IPv4,

residential IPv6, and datacenter IPv6 proxy infrastructure across 195+ countries for public data

collection, regional testing, monitoring, and web scraping workflows. Explore

's latest plans and available offers directly on the website.

Use code webclaw8Off for 8% off your first payment.

See the

for a hands-on walkthrough of wiring ColdProxy into webclaw.


Studio Partners

NodeMaven is the most reliable proxy provider with the highest-quality IPs on the market.

Best solution for automation, web scraping, SEO research, and social media management: 99.9% uptime,

sticky sessions up to 7 days, IP filtering (all proxies under a 97% fraud score), no KYC, and cashback up

to 10% on traffic. Use WEBCLAW35 for 35% off Mobile and Residential proxies, or

WEBCLAW40 for 40% off ISP (Static) proxies at

.

MangoProxy provides residential, ISP, datacenter, and mobile proxies across 200+ locations, backed by a 90M+ IP pool with HTTP and SOCKS5 support and high stability for web scraping and data collection at scale.

Use code 0XMASSI for 8% off ISP (Static) proxies at

.


Community Plugins

Third-party plugins that integrate webclaw with AI agent platforms:

PluginPlatformWhat it does
openclaw-webclawOpenClawNative webclaw v1 API plugin with 9 tools: scrape, search, crawl, extract, summarize, diff, map, batch, brand
hermes-webclawHermes AgentWeb search provider and 9 dedicated tools for the full v1 API surface. Install with `hermes plugins install jal-co/hermes-webclaw`

Built a webclaw integration? Open a PR to add it here.


Contributors

Thanks to everyone improving webclaw through issues, examples, docs, bug reports, and pull requests.


Star History


License

AGPL-3.0

Frequently asked questions

What is webclaw?

webclaw is Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.

How do I install webclaw?

Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

Is webclaw open source?

Yes — it is hosted on GitHub at https://github.com/0xMassi/webclaw and has 2,319 stars.

Related MCP tools

Run your own MCP server? See who uses it and what to fix.

Measure it with TrackMCP