Skip to content
Self-Hosted · Multi-Engine · OpenAI Compatible

Unified Speech RecognitionAPI Gateway for AI Systems

Deploy industrial-grade ASR infrastructure on your own servers or Mac in seconds. Aggregate Faster-Whisper, SenseVoice, FireRedASR, and Zipformer with full OpenAI API compatibility.

# 1. Pull and run OneASR with GPU acceleration
$ docker run -d -p 8000:8000 --gpus all codeandxv/oneasr:latest
INFO: Preloading ASR Engines [whisper-medium, xasr-zh-en]...
INFO: Application startup complete. Uvicorn running on http://0.0.0.0:8000
OpenAI API ready at /v1/audio/transcriptions

Engineered for Performance & Privacy

Everything you need to run high-throughput, private voice transcription pipelines.

🔄Drop-in

OpenAI Compatible API

Drop-in replacement for /v1/audio/transcriptions. Connect seamlessly with Dify, FastGPT, Open WebUI, DuRT, and standard OpenAI SDKs.

⚡Flexible

Multi-Engine Aggregation

Dynamically route requests across Faster-Whisper, Zipformer (X-ASR), Qwen-ASR, and FireRedASR via simple YAML configuration.

🌊< 200ms

Low-Latency Streaming

WebSocket duplex streaming audio protocol for sub-second realtime captions, live meetings, and interactive voice agents.

🎯Smart VAD

Built-in VAD & Chunking

Silero-VAD powered intelligent voice activity detection, natural pause splitting, and duplicate hallucination suppression.

🖥️Built-in UI

Vue 3 Web Console

Interactive web UI for testing microphone inputs, uploading long audio files, monitoring active workers, and benchmarking latency.

🔒Zero Leak

100% Private & Self-Hosted

Keep every byte of voice data inside your infrastructure. Zero third-party telemetry, cloud proxying, or vendor lock-in.

Supported ASR Engines

Choose the best engine suited for your language, latency, and hardware profile.

EngineCategorySupported LanguagesLatency / RTFHardware AccelerationBest Use Case
Faster-WhisperRecommended for FilesFile & Batch ASR99+ Languages (Multilingual)0.12x - 0.25x RTFCUDA, Metal, CPU (int8/fp16)Multilingual file transcription, podcast processing, subtitle generation
X-ASR (Zipformer2)Lowest LatencyRealtime StreamingChinese, English, Mixed< 160ms chunk delayCPU, ONNX, Metal, CUDAUltra-low latency live subtitles, speech-to-text live assistants
Qwen3-ASR (1.7B)File & Batch ASR30+ Languages, Dialects0.15x - 0.35x RTFCUDA (vLLM / PyTorch), MetalComplex Chinese dialects, noisy environments, domain adaptation
FireRedASRIndustrial ASRChinese, English0.18x - 0.30x RTFCUDA, PyTorchEnterprise-grade Chinese meeting minutes and contact centers

How OneASR Works

From audio ingestion to structured transcriptions in a standardized pipeline.

Step 01

1. Ingestion & Pre-processing

Audio streams or files are ingested via REST or WebSocket, normalized to 16kHz mono, and chunked with Silero-VAD.

Step 02

2. Dynamic Engine Dispatcher

Requests are routed to the optimal preloaded engine (Whisper, X-ASR, Qwen) based on language and latency requirements.

Step 03

3. Neural Inference & Alignment

Acoustic models generate phonetic tokens and timestamped words with GPU TensorRT, CUDA, or Metal acceleration.

Step 04

4. Post-processing & Output

Punctuation restoration, repetition filtering, and formatting into OpenAI JSON, SRT, or VTT subtitles.

Developer Experience First

Integrate OneASR into your application with familiar tools and standard protocols.

from openai import OpenAI

# Initialize client pointing to your local OneASR gateway
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="sk-oneasr-v1-p_L3kXm9QZ8sT2vA4wE7rY1u"
)

# Transcribe audio file with word-level timestamps
with open("podcast.mp3", "rb") as audio_file:
    transcript = client.audio.transcriptions.create(
        model="faster-whisper",
        file=audio_file,
        response_format="verbose_json",
        timestamp_granularities=["word", "segment"]
    )

print(transcript.text)