Vosk
PackageSpeech recognition toolkit with small models that work without a network
- Price
- Free, open source
- Access
- None, runs locally
About
Kaldi-based Apache-2.0 speech recognition toolkit from Alpha Cephei. Its 50 MB models cover 20+ languages, Turkish included, and run on a Raspberry Pi, Android or servers (iOS on request), with streaming results and bindings for Python, Java, C#, Node.js and Go. Vosk Server adds WebSocket and gRPC.
What you can do with it
- Transcribe microphone speech in real time with streaming partial results
- Add voice commands to an Android app or a Raspberry Pi device
- Generate SRT subtitles from video files with vosk-transcriber
Get started
- Install the Python package with pip3 install vosk
- Use 16 kHz 16-bit mono PCM WAV, or install ffmpeg for vosk-transcriber
Example
import wave
import sys
from vosk import Model, KaldiRecognizer
wf = wave.open(sys.argv[1], "rb") # WAV format, mono PCM
model = Model(lang="en-us")
rec = KaldiRecognizer(model, wf.getframerate())
rec.SetWords(True)
while True:
data = wf.readframes(4000)
if len(data) == 0:
break
if rec.AcceptWaveform(data):
print(rec.Result())
print(rec.FinalResult())Details
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python, Java, C#, JavaScript/TypeScript, Go
- MCP server
- None