# Markdown Web Extractor

> Use this tool when you need to extract clean and readable markdown content from web pages, solving problems like content analysis and documentation processing by filtering out unnecessary elements. It takes in web page URLs and outputs markdown text, allowing for customizable options like image and link inclusion. It is particularly useful for web scraping workflows where accurate text extraction is required, handling JavaScript-heavy sites and supporting platforms like Confluence.

Canonical page: https://skillsregistry.net/skills/vishwajeetdabholkar-markdown-web-extractor  
JSON: https://api.skillsregistry.net/v1/skills/vishwajeetdabholkar-markdown-web-extractor

## Description

Built by Vishwajeet Dabholkar, this server extracts clean markdown content from web pages using Playwright's headless Chrome browser. It intelligently identifies main content areas while filtering out navigation, headers, footers, and advertisements, with special optimizations for platforms like Confluence. The implementation supports configurable options including image and link inclusion, custom CSS selectors for waiting, timeout controls, and handles JavaScript-heavy sites by waiting for content to load. This makes it valuable for content analysis, documentation processing, and web scraping workflows where clean, readable text extraction is needed.

## Trust

- **Trust score (0–1):** 0.98
- **Verification tier:** verified
- **Last scanned:** 2026-09-19

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** browser-automation
- **Updated:** 2026-09-19

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/vishwajeetdabholkar-markdown-web-extractor)
- **Repository:** <https://github.com/vishwajeetdabholkar/markdown-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "vishwajeetdabholkar-markdown-web-extractor"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/vishwajeetdabholkar-markdown-web-extractor` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/vishwajeetdabholkar-markdown-web-extractor/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
