Vosk

Package

Speech recognition toolkit with small models that work without a network

Price
Free, open source
Access
None, runs locally

About

Kaldi-based Apache-2.0 speech recognition toolkit from Alpha Cephei. Its 50 MB models cover 20+ languages, Turkish included, and run on a Raspberry Pi, Android or servers (iOS on request), with streaming results and bindings for Python, Java, C#, Node.js and Go. Vosk Server adds WebSocket and gRPC.

What you can do with it

  • Transcribe microphone speech in real time with streaming partial results
  • Add voice commands to an Android app or a Raspberry Pi device
  • Generate SRT subtitles from video files with vosk-transcriber

Get started

  1. Install the Python package with pip3 install vosk
  2. Use 16 kHz 16-bit mono PCM WAV, or install ffmpeg for vosk-transcriber

Example

import wave
import sys

from vosk import Model, KaldiRecognizer

wf = wave.open(sys.argv[1], "rb")  # WAV format, mono PCM
model = Model(lang="en-us")

rec = KaldiRecognizer(model, wf.getframerate())
rec.SetWords(True)

while True:
    data = wf.readframes(4000)
    if len(data) == 0:
        break
    if rec.AcceptWaveform(data):
        print(rec.Result())

print(rec.FinalResult())

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python, Java, C#, JavaScript/TypeScript, Go
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .