# Multimodal MCP

> Multimodal MCP — erickpxd-multimodal-mcp. Use this tool when you need to interact with images using natural language, enabling semantic search and visual question answering. It solves problems such as finding specific images based on text descriptions and answering questions about image content. The Multimodal MCP tool takes in text queries and image inputs, outputting relevant images and answers to user questions.

Canonical page: https://skillsregistry.net/skills/erickpxd-multimodal-mcp  
JSON: https://api.skillsregistry.net/v1/skills/erickpxd-multimodal-mcp

## Description

Enables semantic image search using CLIP embeddings and visual question answering through MCP tools, allowing natural language interaction with images.

## Trust

- **Trust score (0–1):** 0.70
- **Verification tier:** verified
- **Last scanned:** 2026-08-30

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-08-30

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/asyvxooc9q)
- **Repository:** <https://github.com/erickpxd/Multimodal-MCP>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "erickpxd-multimodal-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/erickpxd-multimodal-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/erickpxd-multimodal-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
