Local Speech-to-Text
This MCP server provides local speech-to-text transcription using whisper.cpp optimized for M1 MacBook Pro performance, built by SmartLittleApps with TypeScript and the Model Context Protocol SDK. It offers six core tools: basic audio transcription with automatic format conversion via ffmpeg, long audio file processing with intelligent chunking and overlap handling, speaker diarization using PyTorch and pyannote.audio for multi-speaker identification, model management for whisper model discovery and recommendations, health checking for installation validation, and version reporting. The implementation features automatic audio format conversion, M1-specific optimizations, Python integration for advanced diarization capabilities, and support for multiple output formats including VTT, SRT, JSON, and plain text, making it valuable for podcast transcription, meeting notes, interview processing, and any workflow requiring private, local speech-to-text processing without cloud dependencies.
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
View full trust & usage report →Metadata
- Version
- 1.0.0
- Skill type
- atomic
- Execution layer
- mcp-remote
- Category
- media
- Source
- PulseMCP
- Author type
- human
- Updated
- 2026-04-25
Use via MCP
Resolve Local Speech-to-Text from your agent
Streamable HTTP transport at https://api.skillsregistry.net/mcp. No auth for read tools. Discovery: .well-known/mcp.json.
One command in your shell — Claude Code wires it up and verifies the connection. Run /mcp in any session to confirm.
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp --scope user for --scope project to commit it to .mcp.json.