webclaw
A fast, locally-focused web content extraction tool supporting CLI, REST API, and MCP server.
- Type
- MCP
- Transport
- stdio
- Open source
- Yes
- GitHub Stars
- ★ 1.5k
- Source
- mcp-github
- Repository
- github.com/0xMassi/webclaw
Overview
webclaw is a fast, locally-first web content extraction tool designed for LLMs. It can scrape, crawl, and extract structured data, with support for Rust implementation. Offers multiple access methods including CLI, REST API, and MCP server. With webclaw, AI can convert web pages into clean Markdown, JSON, or LLM-ready context. Ideal for scenarios requiring useful information extraction from web pages, such as document crawling and competitor analysis.
Capabilities
- ▪Scrape individual web pages
- ▪Crawl entire websites
- ▪Extract structured data
- ▪Generate LLM-optimized text
- ▪Retain only main content
- ▪Include or exclude specific selectors
Use cases
Setup
npx create-webclaw
This information was compiled by AI from public sources and may contain inaccuracies — please refer to the source.
FAQ
How to install webclaw?
Install quickly using `npx create-webclaw`.
What output formats does webclaw support?
Supports Markdown, JSON, LLM-optimized text, and more.
Related skills
Scrapling
An adaptive web scraping framework capable of handling everything from single requests to large-scale scraping.
TrendRadar
An AI-driven tool for sentiment monitoring and trend analysis, supporting multi-platform aggregation, RSS subscription, and intelligent notifications.
gpt-researcher
An autonomous agent for deep research using any LLM provider.
video-search-and-summarization
NVIDIA Video Search and Summarization Blueprint enables real-time video analytics, visual question answering, and automated reporting.
SurfSense
SurfSense is an open-source AI agent network research platform providing real-time data connectivity via REST API or MCP server.
google-maps-scraper
Scrape business data from Google Maps, such as name, address, phone number, etc.