Hardware Acceleration & Optimization
OneASR is designed to extract maximum performance from whatever hardware is available, from multi-GPU server clusters to lightweight edge laptops.
1. NVIDIA CUDA & TensorRT
For high-throughput server workloads, NVIDIA GPUs offer the lowest latency and highest concurrency.
Recommended Configuration (config.yaml):
yaml
ASR-Providers:
faster-whisper:
enable: true
load:
device: "cuda"
compute_type: "float16" # Or "int8_float16" for lower VRAMVRAM Requirements Table:
| Model / Engine | Min VRAM (int8) | Recommended VRAM (fp16) | Concurrency (Concurrent Streams) |
|---|---|---|---|
| Faster-Whisper Tiny/Base | 1.0 GB | 2.0 GB | 20+ |
| Faster-Whisper Small/Medium | 2.5 GB | 5.0 GB | 8 - 12 |
| Faster-Whisper Large-v3 | 4.5 GB | 8.0 GB | 4 - 6 |
| X-ASR (Zipformer ONNX) | < 500 MB | < 1.0 GB | 50+ |
| Qwen3-ASR (1.7B) | 3.5 GB | 6.0 GB | 4 - 8 |
2. Apple Silicon (M1 / M2 / M3 / M4)
OneASR leverages unified memory and Metal Performance Shaders (MPS) or optimized CPU SIMD on macOS.
yaml
ASR-Providers:
faster-whisper:
enable: true
load:
device: "cpu"
compute_type: "int8" # Extremely fast on Apple Silicon with 8+ CPU coresTIP
On Apple Silicon Macs, running Faster-Whisper with compute_type: int8 and 4 CPU worker threads delivers over 4x real-time speed (RTF < 0.25) with near-zero heat and power consumption.
3. CPU Multithreading Optimization
When running without dedicated GPUs (e.g., standard VPS or Docker container):
- Set
OMP_NUM_THREADSandCT2_USE_EXPERIMENTAL_PACKED_GEMM=1for optimal matrix math. - Use
int8quantization to cut memory bandwidth usage by 50%.
