# doc-ingestor-mcp

> Use this tool when you need to convert unstructured documents into machine-readable format for AI processing and RAG pipelines. The doc-ingestor-mcp takes in various file types such as PDFs, Office documents, images, and audio, and outputs clean Markdown. It solves the problem of ingesting and preprocessing diverse document formats, making it ideal for use cases where structured data is required for downstream AI applications.

Canonical page: https://skillsregistry.net/skills/saleemh-doc-ingestor  
JSON: https://api.skillsregistry.net/v1/skills/saleemh-doc-ingestor

## Description

An MCP server that uses Docling to convert PDFs, Office documents, images, audio, and more into clean Markdown for AI processing and RAG pipelines.

## Trust

- **Trust score (0–1):** 0.68
- **Verification tier:** scanned
- **Last scanned:** 2026-09-03

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** devops-ci
- **Updated:** 2026-09-03

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/r5l658vrt1)
- **Repository:** <https://github.com/saleemh/doc-ingestor>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "saleemh-doc-ingestor"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/saleemh-doc-ingestor` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/saleemh-doc-ingestor/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
