# vision-mcp

> vision-mcp — miaomiaozii-vision-mcp. Use this tool when you need to enhance text-only LLMs with image recognition and UI grounding capabilities, solving problems like visual context understanding and object detection. It accepts image inputs and outputs recognized objects or UI elements, supporting both local and cloud vision backends. Ideal for applications requiring multimodal understanding, such as visual question answering or image-based dialogue systems.

Canonical page: https://skillsregistry.net/skills/miaomiaozii-vision-mcp  
JSON: https://api.skillsregistry.net/v1/skills/miaomiaozii-vision-mcp

## Description

Adds image recognition and UI grounding capabilities to text-only LLMs through MCP tools, supporting local and cloud vision backends.

## Trust

- **Trust score (0–1):** 0.69
- **Verification tier:** scanned
- **Last scanned:** 2026-09-03

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-09-03

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/t5yyd4zuos)
- **Repository:** <https://github.com/miaomiaozii/vision-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "miaomiaozii-vision-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/miaomiaozii-vision-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/miaomiaozii-vision-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
