# PDF Extraction

> Use this tool when you need to extract content from PDF files, such as text and images, to solve problems like document analysis, text mining, or content indexing. It takes PDF files as input and outputs extracted text, supporting specific page ranges and various PDF formats. Use it in contexts where AI assistants or applications require programmatic access to PDF content without manual parsing and OCR complexity.

Canonical page: https://skillsregistry.net/skills/xraywu-pdf-extraction  
JSON: https://api.skillsregistry.net/v1/skills/xraywu-pdf-extraction

## Description

This PDF extraction MCP server, developed by an unnamed author, provides tools for extracting content from PDF files. Built with Python 3.11+ and leveraging libraries like PyPDF2, pytesseract, and PyMuPDF, it offers both text extraction and OCR capabilities. The implementation focuses on flexibility, allowing extraction from specific page ranges and supporting various PDF formats. It's particularly useful for tasks like document analysis, text mining, or content indexing, enabling AI assistants or applications to access PDF content programmatically without dealing with the complexities of PDF parsing and OCR directly.

## Trust

- **Trust score (0–1):** 0.65
- **Verification tier:** scanned
- **Last scanned:** 2026-09-01

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** productivity
- **Updated:** 2026-09-01

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/xraywu-pdf-extraction)
- **Repository:** <https://github.com/xraywu/mcp-pdf-extraction-server>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "xraywu-pdf-extraction"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/xraywu-pdf-extraction` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/xraywu-pdf-extraction/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
