# The Crawler

> Use this tool when you need to extract specific data from web pages or documents, such as text, images, or metadata, and require output in a structured format like markdown. The Crawler solves problems related to web scraping, data extraction, and document processing, supporting various file types including PDF and DOCX. It is ideal for use cases where large-scale data extraction is necessary, with a cost-effective pricing model of $0.003 per page scraped.

Canonical page: https://skillsregistry.net/skills/manchittlab-the-crawler  
JSON: https://api.skillsregistry.net/v1/skills/manchittlab-the-crawler

## Description

The Crawler is a web scraping tool that extracts text, links, images, metadata, tables, and structured data from web pages and documents, returning output in LLM-ready markdown with RAG chunking. It supports PDF and DOCX ingestion and uses CheerioCrawler for pure HTTP scraping without a headless browser. Charges $0.003 per page scraped.

## Trust

- **Trust score (0–1):** 0.54
- **Verification tier:** scanned
- **Last scanned:** 2026-09-02

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** browser-automation
- **Updated:** 2026-09-02

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/manchittlab-the-crawler)
- **Repository:** <https://github.com/manchittlab/thecrawler>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "manchittlab-the-crawler"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/manchittlab-the-crawler` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/manchittlab-the-crawler/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
