faster-whisper
PackageWhisper speech recognition on CTranslate2, faster and lighter on CPU and GPU
- Price
- Free, open source
- Access
- None, runs locally
About
Python library from SYSTRAN that runs OpenAI's Whisper models on the CTranslate2 inference engine. The README reports up to 4 times the speed of openai-whisper at the same accuracy with less memory, plus 8-bit quantization on CPU and GPU. Audio is decoded with PyAV, so no system FFmpeg is needed.
What you can do with it
- Transcribe audio files on a CPU with int8 quantization
- Batch-transcribe long recordings quickly on an NVIDIA GPU
- Get word-level timestamps and skip silence with the Silero VAD filter
Get started
- Install Python 3.9 or newer, then pip install faster-whisper
- For GPU, install cuBLAS for CUDA 12 and cuDNN 9
Example
from faster_whisper import WhisperModel
model_size = "large-v3"
# Run on GPU with FP16
model = WhisperModel(model_size, device="cuda", compute_type="float16")
# or run on CPU with INT8
# model = WhisperModel(model_size, device="cpu", compute_type="int8")
segments, info = model.transcribe("audio.mp3", beam_size=5)
print("Detected language '%s' with probability %f" % (info.language, info.language_probability))
for segment in segments:
print("[%.2fs -> %.2fs] %s" % (segment.start, segment.end, segment.text))Details
- Hosting
- Runs locally
- Available in
- Worldwide
- Official SDKs
- Python
- MCP server
- None