# Gemini OCR MCP Server

> Use this tool when you need to extract text from images with high accuracy, such as processing CAPTCHAs or recognizing text in scanned documents. It takes file paths or base64 strings as input and outputs extracted text, providing a simple interface for text recognition tasks. Ideal for use cases requiring automated text extraction from visual data, enabling efficient processing and analysis of image-based content.

Canonical page: https://skillsregistry.net/skills/windoc-gemini-ocr-mcp  
JSON: https://api.skillsregistry.net/v1/skills/windoc-gemini-ocr-mcp

## Description

Provides OCR services powered by Google's Gemini API to extract text from images via file paths or base64 strings. It enables high-accuracy text recognition and CAPTCHA processing through simple MCP tools.

## Trust

- **Trust score (0–1):** 0.60
- **Verification tier:** unverified

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-04-21

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/avpos8sswa)
- **Repository:** <https://github.com/WindoC/gemini-ocr-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "windoc-gemini-ocr-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/windoc-gemini-ocr-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/windoc-gemini-ocr-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
