# TheCrawler

> Use this tool when you need to extract data from websites, documents, or files in various formats, such as PDF and DOCX. TheCrawler solves problems of data collection and information retrieval by scraping web pages and converting content into LLM-ready markdown with RAG chunking. It takes in URLs, file paths, or raw text as input and outputs formatted markdown, making it ideal for use cases involving data mining, research, and content analysis.

Canonical page: https://skillsregistry.net/skills/io-github-manchittlab-thecrawler  
JSON: https://api.skillsregistry.net/v1/skills/io-github-manchittlab-thecrawler

## Description

Universal web scraper with LLM-ready markdown, RAG chunking, PDF/DOCX support.

## Trust

- **Trust score (0–1):** 0.58
- **Verification tier:** scanned
- **Last scanned:** 2026-09-19

## Facts

- **Version:** 0.1.1
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** browser-automation
- **Updated:** 2026-09-19

## Source

- **Source listing:** [MCP Registry](https://registry.modelcontextprotocol.io/v0/servers/io.github.manchittlab%2Fthecrawler)
- **Repository:** <https://github.com/manchittlab/TheCrawler>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "io-github-manchittlab-thecrawler"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/io-github-manchittlab-thecrawler` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/io-github-manchittlab-thecrawler/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
