# multimodal-mcp

> Use this tool when you need to generate multimedia content from text prompts, solving the problem of fragmented media creation across multiple providers. The multimodal-mcp server takes text inputs and produces images, videos, audio, and transcriptions as outputs, unifying the capabilities of OpenAI, xAI, Gemini, ElevenLabs, and BFL. Ideal for use cases requiring diverse media generation from a single interface, streamlining content creation workflows.

Canonical page: https://skillsregistry.net/skills/rsmdt-multimodal-mcp  
JSON: https://api.skillsregistry.net/v1/skills/rsmdt-multimodal-mcp

## Description

Multi-provider media generation MCP server that generates images, videos, audio, and transcriptions from text prompts using OpenAI, xAI, Gemini, ElevenLabs, and BFL through a single unified interface.

## Trust

- **Trust score (0–1):** 0.68
- **Verification tier:** scanned
- **Last scanned:** 2026-08-31

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-08-31

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/x2h2jvplhb)
- **Repository:** <https://github.com/rsmdt/multimodal-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "rsmdt-multimodal-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/rsmdt-multimodal-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/rsmdt-multimodal-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
