APIs, MCP servers and models for speech to text
12 tools in the catalog do this, by name.
- AssemblyAIAPISpeech-to-text API for recorded files and live audio streams
- CartesiaModelReal-time voice models for text to speech and speech to text
- DeepgramAPISpeech-to-text, text-to-speech and voice agent APIs
- ElevenLabsModelAI voice models for text to speech, transcription and voice agents
- faster-whisperPackageWhisper speech recognition on CTranslate2, faster and lighter on CPU and GPU
- GladiaAPISpeech-to-text API for recorded and real-time audio in 100+ languages
- GroqProviderLow-latency inference API for open LLMs and Whisper
- OpenAIModelAPI for GPT language, image and audio models
- SpeechmaticsAPISpeech-to-text and text-to-speech API in 55+ languages, cloud or on-prem
- SupadataAPIAPI for YouTube and social video transcripts, with AI transcription fallback
- VoskPackageSpeech recognition toolkit with small models that work without a network
- WhisperModelOpenAI's open-source speech recognition model that runs on your machine