Skip to content

Hardware Acceleration & Optimization ​

OneASR is designed to extract maximum performance from whatever hardware is available, from multi-GPU server clusters to lightweight edge laptops.


1. NVIDIA CUDA & TensorRT ​

For high-throughput server workloads, NVIDIA GPUs offer the lowest latency and highest concurrency.

yaml
ASR-Providers:
  faster-whisper:
    enable: true
    load:
      device: "cuda"
      compute_type: "float16" # Or "int8_float16" for lower VRAM

VRAM Requirements Table: ​

Model / EngineMin VRAM (int8)Recommended VRAM (fp16)Concurrency (Concurrent Streams)
Faster-Whisper Tiny/Base1.0 GB2.0 GB20+
Faster-Whisper Small/Medium2.5 GB5.0 GB8 - 12
Faster-Whisper Large-v34.5 GB8.0 GB4 - 6
X-ASR (Zipformer ONNX)< 500 MB< 1.0 GB50+
Qwen3-ASR (1.7B)3.5 GB6.0 GB4 - 8

2. Apple Silicon (M1 / M2 / M3 / M4) ​

OneASR leverages unified memory and Metal Performance Shaders (MPS) or optimized CPU SIMD on macOS.

yaml
ASR-Providers:
  faster-whisper:
    enable: true
    load:
      device: "cpu"
      compute_type: "int8" # Extremely fast on Apple Silicon with 8+ CPU cores

TIP

On Apple Silicon Macs, running Faster-Whisper with compute_type: int8 and 4 CPU worker threads delivers over 4x real-time speed (RTF < 0.25) with near-zero heat and power consumption.


3. CPU Multithreading Optimization ​

When running without dedicated GPUs (e.g., standard VPS or Docker container):

  • Set OMP_NUM_THREADS and CT2_USE_EXPERIMENTAL_PACKED_GEMM=1 for optimal matrix math.
  • Use int8 quantization to cut memory bandwidth usage by 50%.