# Local Speech-to-Text

> Use this tool when you need to transcribe spoken audio into text locally, without relying on cloud services, and require features like speaker diarization, automatic format conversion, and support for multiple output formats. It solves problems like podcast transcription, meeting notes, and interview processing, offering a private and efficient solution for speech-to-text tasks. It accepts audio files as input and produces transcribed text in various formats, including VTT, SRT, JSON, and plain text.

Canonical page: https://skillsregistry.net/skills/smartlittleapps-local-speech-to-text  
JSON: https://api.skillsregistry.net/v1/skills/smartlittleapps-local-speech-to-text

## Description

This MCP server provides local speech-to-text transcription using whisper.cpp optimized for M1 MacBook Pro performance, built by SmartLittleApps with TypeScript and the Model Context Protocol SDK. It offers six core tools: basic audio transcription with automatic format conversion via ffmpeg, long audio file processing with intelligent chunking and overlap handling, speaker diarization using PyTorch and pyannote.audio for multi-speaker identification, model management for whisper model discovery and recommendations, health checking for installation validation, and version reporting. The implementation features automatic audio format conversion, M1-specific optimizations, Python integration for advanced diarization capabilities, and support for multiple output formats including VTT, SRT, JSON, and plain text, making it valuable for podcast transcription, meeting notes, interview processing, and any workflow requiring private, local speech-to-text processing without cloud dependencies.

## Trust

- **Trust score (0–1):** 0.50
- **Verification tier:** unverified

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** media
- **Updated:** 2026-04-25

## Source

- **Source listing:** [PulseMCP](https://www.pulsemcp.com/servers/smartlittleapps-local-speech-to-text)
- **Repository:** <https://github.com/smartlittleapps/local-stt-mcp/tree/HEAD/mcp-server>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "smartlittleapps-local-speech-to-text"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/smartlittleapps-local-speech-to-text` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/smartlittleapps-local-speech-to-text/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
