Deploy industrial-grade ASR infrastructure on your own servers or Mac in seconds. Aggregate Faster-Whisper, SenseVoice, FireRedASR, and Zipformer with full OpenAI API compatibility.
Everything you need to run high-throughput, private voice transcription pipelines.
Drop-in replacement for /v1/audio/transcriptions. Connect seamlessly with Dify, FastGPT, Open WebUI, DuRT, and standard OpenAI SDKs.
Dynamically route requests across Faster-Whisper, Zipformer (X-ASR), Qwen-ASR, and FireRedASR via simple YAML configuration.
WebSocket duplex streaming audio protocol for sub-second realtime captions, live meetings, and interactive voice agents.
Silero-VAD powered intelligent voice activity detection, natural pause splitting, and duplicate hallucination suppression.
Interactive web UI for testing microphone inputs, uploading long audio files, monitoring active workers, and benchmarking latency.
Keep every byte of voice data inside your infrastructure. Zero third-party telemetry, cloud proxying, or vendor lock-in.
Choose the best engine suited for your language, latency, and hardware profile.
| Engine | Category | Supported Languages | Latency / RTF | Hardware Acceleration | Best Use Case |
|---|---|---|---|---|---|
| Faster-WhisperRecommended for Files | File & Batch ASR | 99+ Languages (Multilingual) | 0.12x - 0.25x RTF | CUDA, Metal, CPU (int8/fp16) | Multilingual file transcription, podcast processing, subtitle generation |
| X-ASR (Zipformer2)Lowest Latency | Realtime Streaming | Chinese, English, Mixed | < 160ms chunk delay | CPU, ONNX, Metal, CUDA | Ultra-low latency live subtitles, speech-to-text live assistants |
| Qwen3-ASR (1.7B) | File & Batch ASR | 30+ Languages, Dialects | 0.15x - 0.35x RTF | CUDA (vLLM / PyTorch), Metal | Complex Chinese dialects, noisy environments, domain adaptation |
| FireRedASR | Industrial ASR | Chinese, English | 0.18x - 0.30x RTF | CUDA, PyTorch | Enterprise-grade Chinese meeting minutes and contact centers |
From audio ingestion to structured transcriptions in a standardized pipeline.
Audio streams or files are ingested via REST or WebSocket, normalized to 16kHz mono, and chunked with Silero-VAD.
Requests are routed to the optimal preloaded engine (Whisper, X-ASR, Qwen) based on language and latency requirements.
Acoustic models generate phonetic tokens and timestamped words with GPU TensorRT, CUDA, or Metal acceleration.
Punctuation restoration, repetition filtering, and formatting into OpenAI JSON, SRT, or VTT subtitles.
Integrate OneASR into your application with familiar tools and standard protocols.
from openai import OpenAI
# Initialize client pointing to your local OneASR gateway
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="sk-oneasr-v1-p_L3kXm9QZ8sT2vA4wE7rY1u"
)
# Transcribe audio file with word-level timestamps
with open("podcast.mp3", "rb") as audio_file:
transcript = client.audio.transcriptions.create(
model="faster-whisper",
file=audio_file,
response_format="verbose_json",
timestamp_granularities=["word", "segment"]
)
print(transcript.text)