Skip to content

Faster-Whisper Engine ​

Faster-Whisper is a reimplementation of OpenAI’s Whisper model using CTranslate2, a fast inference engine for Transformer models. It is up to 4x faster than openai/whisper with significantly reduced memory footprint.


Key Advantages ​

  • Universal Multilingual Support: 99+ languages with automatic language detection.
  • Word-Level Timestamps: Precise alignment for subtitles (SRT/VTT).
  • Quantization Support: Native int8, int8_float16, and float16 execution.

Configuration Example ​

yaml
ASR-Providers:
  faster-whisper:
    enable: true
    engine: faster-whisper
    load:
      model_name: medium # tiny, base, small, medium, large-v3
      model_path: models/faster-whisper-medium
      device: cuda # or cpu
      compute_type: float16 # or int8

Performance Benchmark ​

Model SizeParametersVRAM (fp16)RTF (NVIDIA RTX 4090)
base74M1.0 GB0.03x
small244M2.0 GB0.06x
medium769M4.5 GB0.12x
large-v31550M8.0 GB0.22x