Skip to content

OneASR Overview & Architecture ​

OneASR is an open-source, high-throughput speech recognition (ASR) aggregation gateway and self-hosted inference server. It standardizes diverse speech models into unified, OpenAI-compatible APIs and ultra-low latency WebSocket streams.


Why OneASR? ​

Modern AI applications and voice agents require reliable, low-latency, and private speech recognition. However, running state-of-the-art ASR models often introduces fragmented interfaces, missing VAD/chunking pipelines, and high cloud costs.

OneASR solves this by providing:

  1. Unified OpenAI Compatibility: Direct drop-in replacement for POST /v1/audio/transcriptions.
  2. Multi-Engine Multiplexing: Switch between Faster-Whisper, X-ASR (Zipformer), Qwen3-ASR, and FireRedASR with zero client code modifications.
  3. Integrated VAD & Chunking: Silero-VAD pipeline handles audio normalization, pause segmentation, and repetition filtering automatically.
  4. 100% Data Sovereignty: Self-hosted on your Mac, local GPU workstation, or Kubernetes cluster with zero external telemetry.

System Architecture ​

[ Clients & AI Apps ] (DuRT, Dify, Open WebUI, Python SDK, cURL)
        │
        ▼ (HTTP / WebSocket)
┌────────────────────────────────────────────────────────┐
│                   OneASR Gateway (FastAPI)             │
│                                                        │
│  ┌───────────────────────┐  ┌───────────────────────┐  │
│  │   Auth & API Layer    │  │  ASR-Toolkit (VAD)    │  │
│  │  • /v1/audio/*        │  │  • Silero-VAD         │  │
│  │  • /api/v1/stream     │  │  • Dynamic Chunking   │  │
│  └───────────┬───────────┘  └───────────┬───────────┘  │
│              │                          │              │
│              ▼                          ▼              │
│  ┌──────────────────────────────────────────────────┐  │
│  │             Engine Registry & Dispatcher         │  │
│  └──────────────────────────┬───────────────────────┘  │
└─────────────────────────────┼──────────────────────────┘
                              │
         ┌────────────────────┼────────────────────┐
         ▼                    ▼                    ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│  Faster-Whisper  │ │  X-ASR Zipformer │ │    Qwen3-ASR     │
│ (CUDA/Metal/CPU) │ │ (ONNX / Stream)  │ │ (vLLM / PyTorch) │
└──────────────────┘ └──────────────────┘ └──────────────────┘

Core Capabilities ​

  • Realtime Streaming: Sub-200ms chunk processing with WebSocket full-duplex protocol.
  • Async Batch Jobs: Upload large multi-gigabyte audio/video files with background worker execution and progress polling.
  • Millisecond Timestamps: Generate word-level and sentence-level timestamps for subtitle alignment and speaker diarization.
  • Embedded Web Console: Interactive Vue 3 dashboard for microphone testing, model metrics, and task management.

Next Steps ​