# Web Crawler Data Bridge

> Use this tool when you need to bridge web crawler data with AI language models for analysis and querying. It solves problems in AI-powered website analysis, content auditing, SEO optimization, and research by providing full-text search, filtering, and pagination of crawled web content. The tool accepts various crawler formats as input and outputs queried data through a JSON configuration interface, making it ideal for workflows that require fast indexing and search of previously crawled web content.

Canonical page: https://skillsregistry.net/skills/pragmar-webcrawl  
JSON: https://api.skillsregistry.net/v1/skills/pragmar-webcrawl

## Description

This MCP server by Ben Caulfield bridges web crawler data with AI language models, supporting five major crawler formats: WARC files, wget archives, InterroBot databases, Katana HTTP text files, and SiteOne captures. Built with Python and featuring full-text search with boolean queries, field-specific filtering by HTTP status and content type, resource pagination, and advanced extras like thumbnail generation for images, markdown conversion, contextual snippets, and XPath extraction. The implementation uses in-memory SQLite databases for fast indexing and search, supports multiple simultaneous crawler configurations, and integrates with Claude Desktop through JSON configuration, making it valuable for AI-powered website analysis, content auditing, SEO optimization, and research workflows that need to query and analyze previously crawled web content.

## Trust

- **Trust score (0–1):** 0.27
- **Verification tier:** scanned
- **Last scanned:** 2026-09-19

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** database
- **Updated:** 2026-09-19

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/pragmar-webcrawl)
- **Repository:** <https://github.com/pragmar/mcp-server-webcrawl>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "pragmar-webcrawl"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/pragmar-webcrawl` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/pragmar-webcrawl/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
