Sign inSign up

superbizon007/crawl-web-mcp

By superbizon007

•Updated 6 months ago

server that exposes web crawling capabilities as tools for AI models.

Image
0

303

superbizon007/crawl-web-mcp repository overview

⁠Crawl Web MCP

An MCP (Model Context Protocol) server that exposes web crawling capabilities as tools for AI models. Tracks HTTP requests, manages cookies, and stores page snapshots via SQLite.

⁠Features

  • MCP Tools — 7 tools accessible via stdio transport (Claude Desktop, MCP clients)
  • Cookie Management — Persistent cookie storage across requests (SQLite)
  • Network Tracking — In-memory log of the last 100 HTTP requests
  • Page Snapshots — Store and retrieve last-visited HTML content
  • HTML Parsing — CSS selector-based element extraction (BeautifulSoup)
  • Realistic Browser Headers — Chrome 120 User-Agent and Sec-Fetch headers
  • Proxy Support — Automatic HTTP_PROXY / HTTPS_PROXY / NO_PROXY support

⁠Quick Start

⁠Install
# Using uv (recommended)
uv pip install -e .

# Or pip
pip install -e .
⁠Run
# Via entry point (stdio transport)
crawl-web-mcp

# Or via module
python -m app.server

# Or with uv
uv run crawl-web-mcp

The server communicates over stdio — it does not listen on a port. Connect it to an MCP client (e.g., Claude Desktop).

⁠Claude Desktop Configuration

Add to your Claude Desktop claude_desktop_config.json:

{
  "mcpServers": {
    "crawl-web": {
      "command": "crawl-web-mcp"
    }
  }
}

Or with uv:

{
  "mcpServers": {
    "crawl-web": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/crawl_web_mcp", "crawl-web-mcp"]
    }
  }
}

⁠MCP Tools

ToolParametersDescription
open_urlurl, headers?Fetch a URL, store HTML/cookies, track the request
get_snapshot—Return the last stored HTML page source and URL
find_elementselector, limit?Find elements in last HTML via CSS selector
get_cookies—List all stored cookies
set_cookiename, value, domain, path?, expires?, secure?, http_only?Add or update a cookie
delete_cookies—Remove all stored cookies
get_network_requests—List tracked network requests (last 100, in-memory)
⁠Tool Response Examples

open_url

{
  "success": true,
  "url": "https://example.com",
  "status_code": 200,
  "content_length": 1234,
  "cookies_count": 2,
  "message": "Page opened successfully"
}

find_element

{
  "success": true,
  "selector": "h1",
  "elements": [
    {"html": "<h1>Title</h1>", "text": "Title", "attributes": {}}
  ],
  "count": 1
}

⁠Project Structure

crawl_web_mcp/
├── app/
│   ├── server.py              # MCP server & tool definitions (entry point)
│   ├── config.py              # Settings (debug flag via pydantic-settings)
│   ├── models/
│   │   └── schemas.py         # CookieSchema, NetworkRequestSchema
│   └── services/
│       ├── crawler.py         # CrawlerService — HTTP requests + browser headers
│       ├── storage.py         # StorageService — SQLite persistence
│       └── network_tracker.py # NetworkTracker — in-memory request log
├── tests/
│   └── test_storage.py        # StorageService unit tests
├── storage/                   # Created at runtime
│   └── crawl_agent.db         # SQLite database
├── pyproject.toml
├── ARCHITECTURE.md
└── README.md

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
DEBUGfalseEnable debug logging
HTTP_PROXY—HTTP proxy server
HTTPS_PROXY—HTTPS proxy server
NO_PROXY—Comma-separated hosts to bypass proxy

Create a .env file in the project root or export the variables directly.

⁠Storage

SQLite database at storage/crawl_agent.db (auto-created):

  • cookies — Persistent cookies, unique by (name, domain, path)
  • page_snapshot — Last visited page (single row: URL + HTML)
  • request_headers — Last request headers as JSON (single row)

Network requests are tracked in-memory only (last 100, not persisted).

⁠Development

⁠Prerequisites
  • Python 3.13+
  • uv or pip
⁠Setup
# Install with dev dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Run with coverage
pytest --cov=app --cov-report=html

⁠Tech Stack

  • mcp[cli]⁠ — FastMCP server (stdio transport)
  • requests — HTTP client with realistic browser headers
  • beautifulsoup4 — HTML parsing and CSS selector queries
  • pydantic / pydantic-settings — Data validation and configuration
  • SQLite — Persistent storage (cookies, snapshots, headers)

Tag summary

Content type

Image

Digest

sha256:d623abece…

Size

186.3 MB

Last updated

6 months ago

docker pull superbizon007/crawl-web-mcp