Video & Audio Text Extraction
An MCP server that provides text extraction capabilities from various video platforms and audio files using OpenAI's Whisper for high-quality speech recognition. Built by Seazhang, it supports downloading and transcribing content from YouTube, Bilibili, TikTok, Instagram, Twitter/X, Facebook, Vimeo, and other platforms supported by yt-dlp, offering four main tools: video download, audio extraction, video-to-text transcription, and audio file transcription. The server handles multi-language recognition, supports various audio formats, and includes configurable Whisper model sizes for balancing speed versus accuracy, making it valuable for content analysis, accessibility improvements, meeting transcription, and extracting insights from video content across diverse platforms.
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
View full trust & usage report →Metadata
- Version
- 1.0.0
- Skill type
- atomic
- Execution layer
- mcp-remote
- Category
- social-media
- Source
- PulseMCP
- Repository
- github.com/sealingp/mcp-video-extraction
- Author type
- human
- Updated
- 2026-04-25
Use via MCP
Resolve Video & Audio Text Extraction from your agent
Streamable HTTP transport at https://api.skillsregistry.net/mcp. No auth for read tools. Discovery: .well-known/mcp.json.
One command in your shell — Claude Code wires it up and verifies the connection. Run /mcp in any session to confirm.
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp --scope user for --scope project to commit it to .mcp.json.