Whisper

Model

OpenAI's open-source speech recognition model that runs on your machine

Price
Free, open source
Access
None, runs locally

About

Whisper is OpenAI's MIT-licensed speech recognition model and Python package. It transcribes speech in many languages, translates it into English and detects the language, fully offline. It needs ffmpeg, and the large model wants about 10 GB of GPU memory.

What you can do with it

  • Transcribe audio files offline on your own machine or server
  • Translate speech in other languages into English text
  • Generate SRT or VTT subtitles from video or podcast audio

Get started

  1. Install ffmpeg, e.g. brew install ffmpeg
  2. Install the package with pip install -U openai-whisper

Example

import whisper

model = whisper.load_model("turbo")
result = model.transcribe("audio.mp3")
print(result["text"])

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .