Faster-Whisper Engine
Faster-Whisper is a reimplementation of OpenAI’s Whisper model using CTranslate2, a fast inference engine for Transformer models. It is up to 4x faster than openai/whisper with significantly reduced memory footprint.
Key Advantages
- Universal Multilingual Support: 99+ languages with automatic language detection.
- Word-Level Timestamps: Precise alignment for subtitles (SRT/VTT).
- Quantization Support: Native
int8,int8_float16, andfloat16execution.
Configuration Example
yaml
ASR-Providers:
faster-whisper:
enable: true
engine: faster-whisper
load:
model_name: medium # tiny, base, small, medium, large-v3
model_path: models/faster-whisper-medium
device: cuda # or cpu
compute_type: float16 # or int8Performance Benchmark
| Model Size | Parameters | VRAM (fp16) | RTF (NVIDIA RTX 4090) |
|---|---|---|---|
| base | 74M | 1.0 GB | 0.03x |
| small | 244M | 2.0 GB | 0.06x |
| medium | 769M | 4.5 GB | 0.12x |
| large-v3 | 1550M | 8.0 GB | 0.22x |
