음성 전사 (POST /v1/audio/transcriptions)
완전한 OpenAI API 호환성을 갖추고 오디오를 지정된 형식의 텍스트로 전사합니다.
HTTP 요청
http
POST /v1/audio/transcriptions
Authorization: Bearer sk-oneasr-...
Content-Type: multipart/form-data요청 본문 (Multipart Form)
| 필드 | 타입 | 필수 여부 | 설명 |
|---|---|---|---|
file | binary | 필수 | 오디오 파일 바이너리 객체 (mp3, mp4, wav, m4a, ogg, webm, flac 등). |
model | string | 필수 | 대상 엔진/모델 이름 (예: faster-whisper, qwen, whisper-1). |
language | string | 선택 | ISO-639-1 언어 코드 (예: ko, en, zh, ja, de, es). |
prompt | string | 선택 | 철자나 전문 용어를 유도하기 위한 선택적 가이드 프롬프트. |
response_format | string | 선택 | 출력 형식: json (기본값), text, srt, verbose_json, vtt. |
temperature | number | 선택 | 샘플링 온도 (0.0 ~ 1.0, 기본값: 0.0). |
timestamp_granularities | array | 선택 | verbose_json 타임스탬프 단위: ["word"], ["segment"]. |
코드 예제
cURL
bash
curl -X POST http://localhost:8000/v1/audio/transcriptions -H "Authorization: Bearer sk-oneasr-v1-xxx" -F file="@podcast.mp3" -F model="faster-whisper" -F response_format="verbose_json"Python (공식 OpenAI SDK)
python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="sk-oneasr-v1-xxx"
)
with open("speech.mp3", "rb") as f:
result = client.audio.transcriptions.create(
model="faster-whisper",
file=f,
response_format="verbose_json"
)
print(result.text)응답 예시 (verbose_json)
json
{
"task": "transcribe",
"language": "english",
"duration": 5.24,
"text": "Welcome to OneASR private voice gateway.",
"segments": [
{
"id": 0,
"seek": 0,
"start": 0.0,
"end": 5.2,
"text": "Welcome to OneASR private voice gateway.",
"tokens": [50364, 4843, 284, 1530, 50589],
"temperature": 0.0,
"avg_logprob": -0.12,
"compression_ratio": 1.1,
"no_speech_prob": 0.001
}
]
}