# doc-scraper

> doc-scraper — sriram-pr-doc-scraper. Use this tool when you need to extract and convert documentation from websites into clean Markdown format for Large Language Model (LLM) ingestion, such as training data for RAG. It takes in website URLs and outputs formatted Markdown files, making it ideal for data preparation and ingestion pipelines. The doc-scraper tool simplifies the process of collecting and preprocessing documentation for AI model training.

Canonical page: https://skillsregistry.net/skills/sriram-pr-doc-scraper  
JSON: https://api.skillsregistry.net/v1/skills/sriram-pr-doc-scraper

## Description

Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).

## Trust

- **Trust score (0–1):** 0.65
- **Verification tier:** scanned
- **Last scanned:** 2026-09-28

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** container
- **Runtime environment:** vm
- **License:** Apache-2.0
- **Updated:** 2026-09-28

## Source

- **Source listing:** [GitHub](https://github.com/Sriram-PR/doc-scraper)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "sriram-pr-doc-scraper"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/sriram-pr-doc-scraper` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/sriram-pr-doc-scraper/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
