# multimodal-mcp

> multimodal-mcp — linjiankun-multimodal-mcp. Use this tool when you need to enhance text-only models with multimodal capabilities, such as image description, audio transcription, and video analysis, to solve problems like multimedia data processing and generation. It takes in various media inputs, including images, audio, and video, and outputs corresponding text descriptions, transcriptions, or generated media. Ideal for use cases where text-only models are insufficient, such as multimedia content creation and analysis.

Canonical page: https://skillsregistry.net/skills/linjiankun-multimodal-mcp  
JSON: https://api.skillsregistry.net/v1/skills/linjiankun-multimodal-mcp

## Description

Local MCP server that adds multimodal capabilities to text-only models like Codex/DeepSeek, offering tools for image description, audio transcription, video analysis, image/video generation, and speech synthesis.

## Trust

- **Trust score (0–1):** 0.69
- **Verification tier:** scanned
- **Last scanned:** 2026-09-03

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-09-03

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/qvylj8rni9)
- **Repository:** <https://github.com/LinJianKun/multimodal-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "linjiankun-multimodal-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/linjiankun-multimodal-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/linjiankun-multimodal-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
