Whisper
ModelOpenAI's open-source speech recognition model that runs on your machine
- Price
- Free, open source
- Access
- None, runs locally
About
Whisper is OpenAI's MIT-licensed speech recognition model and Python package. It transcribes speech in many languages, translates it into English and detects the language, fully offline. It needs ffmpeg, and the large model wants about 10 GB of GPU memory.
What you can do with it
- Transcribe audio files offline on your own machine or server
- Translate speech in other languages into English text
- Generate SRT or VTT subtitles from video or podcast audio
Get started
- Install ffmpeg, e.g. brew install ffmpeg
- Install the package with pip install -U openai-whisper
Example
import whisper
model = whisper.load_model("turbo")
result = model.transcribe("audio.mp3")
print(result["text"])Details
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python
- MCP server
- None