# Markdown Web Crawl

> Use this tool when you need to extract and save website content as markdown files, solving problems like content aggregation, site archiving, or building knowledge bases from web sources. It takes in website URLs and configuration settings, and outputs markdown files with optional content indexes. Ideal for small-scale personal projects or larger data collection tasks, this tool streamlines web scraping with features like batch URL processing and concurrent requests.

Canonical page: https://skillsregistry.net/skills/jmh108-md-webcrawl  
JSON: https://api.skillsregistry.net/v1/skills/jmh108-md-webcrawl

## Description

This MCP implementation, developed by JMH, is a Python-based web crawler designed for extracting and saving website content as markdown files. It offers features like website structure mapping, batch URL processing, and configurable output settings. The project integrates with FastMCP for easy installation and deployment, and leverages libraries such as BeautifulSoup and requests for efficient web scraping. Its focus on markdown output and straightforward configuration makes it particularly suitable for content aggregation, site archiving, or building knowledge bases from web sources. The crawler's ability to create content indexes and its support for concurrent requests set it apart as a tool for both small-scale personal projects and larger data collection tasks.

## Trust

- **Trust score (0–1):** 0.65
- **Verification tier:** scanned
- **Last scanned:** 2026-09-02

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** devops-ci
- **Updated:** 2026-09-02

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/jmh108-md-webcrawl)
- **Repository:** <https://github.com/jmh108/md-webcrawl-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "jmh108-md-webcrawl"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/jmh108-md-webcrawl` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/jmh108-md-webcrawl/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
