# MCP-PDF-Extractor-server

> MCP-PDF-Extractor-server — rayenmalouche-mcp-pdf-extractor-server. Use this tool when you need to extract content and metadata from various file types, such as PDF, DOCX, and TXT, and require a server-based solution with REST API access. It solves problems related to file analysis, data retrieval, and document processing, providing outputs in HTML, text, and metadata formats. Ideal for use cases involving automated document processing, data mining, and content analysis, where a scalable and compliant server-based extraction solution is necessary.

Canonical page: https://skillsregistry.net/skills/rayenmalouche-mcp-pdf-extractor-server  
JSON: https://api.skillsregistry.net/v1/skills/rayenmalouche-mcp-pdf-extractor-server

## Description

A Java-based server leveraging Apache Tika to extract content and metadata from files (PDF, DOCX, TXT, etc.) in a local files-to-extract directory. Supports HTML (with CSS styling) and text extraction, file listing, and metadata retrieval via MCP-compliant tools and REST APIs. Built with Spring Boot, Jetty, and MCP SDK.

## Trust

- **Trust score (0–1):** 0.82
- **Verification tier:** verified
- **Last scanned:** 2026-09-28

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** container
- **Runtime environment:** vm
- **Updated:** 2026-09-28

## Source

- **Source listing:** [GitHub](https://github.com/RayenMalouche/MCP-PDF-Extractor-server)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "rayenmalouche-mcp-pdf-extractor-server"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/rayenmalouche-mcp-pdf-extractor-server` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/rayenmalouche-mcp-pdf-extractor-server/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
