Web Crawler Data Bridge
This MCP server by Ben Caulfield bridges web crawler data with AI language models, supporting five major crawler formats: WARC files, wget archives, InterroBot databases, Katana HTTP text files, and SiteOne captures. Built with Python and featuring full-text search with boolean queries, field-specific filtering by HTTP status and content type, resource pagination, and advanced extras like thumbnail generation for images, markdown conversion, contextual snippets, and XPath extraction. The implementation uses in-memory SQLite databases for fast indexing and search, supports multiple simultaneous crawler configurations, and integrates with Claude Desktop through JSON configuration, making it valuable for AI-powered website analysis, content auditing, SEO optimization, and research workflows that need to query and analyze previously crawled web content.
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately. Last scanned 2026-09-19.
Scan details: Circle-IR · 2026-09-19 · Appeal
View full trust & usage report →Metadata
- Version
- 1.0.0
- Skill type
- atomic
- Execution layer
- mcp-remote
- Category
- database
- Source
- PulseMCP
- Repository
- github.com/pragmar/mcp-server-webcrawl
- Author type
- human
- Last scanned
- 2026-09-19
- Updated
- 2026-09-19
Use via MCP
Resolve Web Crawler Data Bridge from your agent
Streamable HTTP transport at https://api.skillsregistry.net/mcp. No auth for read tools. Discovery: .well-known/mcp.json.
One command in your shell — Claude Code wires it up and verifies the connection. Run /mcp in any session to confirm.
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp --scope user for --scope project to commit it to .mcp.json.