faster-whisper

Package

Whisper speech recognition on CTranslate2, faster and lighter on CPU and GPU

Price
Free, open source
Access
None, runs locally

About

Python library from SYSTRAN that runs OpenAI's Whisper models on the CTranslate2 inference engine. The README reports up to 4 times the speed of openai-whisper at the same accuracy with less memory, plus 8-bit quantization on CPU and GPU. Audio is decoded with PyAV, so no system FFmpeg is needed.

What you can do with it

  • Transcribe audio files on a CPU with int8 quantization
  • Batch-transcribe long recordings quickly on an NVIDIA GPU
  • Get word-level timestamps and skip silence with the Silero VAD filter

Get started

  1. Install Python 3.9 or newer, then pip install faster-whisper
  2. For GPU, install cuBLAS for CUDA 12 and cuDNN 9

Example

from faster_whisper import WhisperModel

model_size = "large-v3"

# Run on GPU with FP16
model = WhisperModel(model_size, device="cuda", compute_type="float16")
# or run on CPU with INT8
# model = WhisperModel(model_size, device="cpu", compute_type="int8")

segments, info = model.transcribe("audio.mp3", beam_size=5)

print("Detected language '%s' with probability %f" % (info.language, info.language_probability))

for segment in segments:
    print("[%.2fs -> %.2fs] %s" % (segment.start, segment.end, segment.text))

Details

Hosting
Runs locally
Available in
Worldwide
Official SDKs
Python
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .