# mcp-trafilatura-server

> Use this tool when you need to extract clean web content from URLs or HTML, and want to customize the output format and extraction options. It solves problems of web scraping and data mining by providing a simple interface to input URLs or HTML and output extracted content in various formats. Ideal for use cases where structured data is required from unstructured web pages.

Canonical page: https://skillsregistry.net/skills/achieveai-mcp-web-extractor  
JSON: https://api.skillsregistry.net/v1/skills/achieveai-mcp-web-extractor

## Description

This MCP server enables clean web content extraction from URLs or HTML using Trafilatura, supporting multiple output formats and configurable extraction options.

## Trust

- **Trust score (0–1):** 0.69
- **Verification tier:** scanned
- **Last scanned:** 2026-08-31

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** other
- **Updated:** 2026-08-31

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/grg9dxa7m3)
- **Repository:** <https://github.com/achieveai/mcp-web-extractor>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "achieveai-mcp-web-extractor"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/achieveai-mcp-web-extractor` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/achieveai-mcp-web-extractor/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
