Audio Transcription
Convert audio to text using OpenAI-compatible Whisper endpoints —through the same gateway.
Overview
The transcription API follows the OpenAI audio/transcriptions format exactly. If you're already using the OpenAI SDK for transcription, just change the base_url and you're done.
What you need
- An audio file (supported:
flac,mp3,mp4,mpeg,mpga,m4a,ogg,wav,webm) - A configured provider that supports Whisper (OpenAI or Azure OpenAI)
- Max file size: 25 MB
Transcribe Audio
POST /v1/audio/transcriptions
Send your audio file as multipart/form-data:
Request
curl -X POST https://api.2kw.ai/v1/audio/transcriptions \
-H "Authorization: Bearer sk_your_api_key" \
-F "file=@meeting.mp3" \
-F "model=whisper-1" \
-F "language=en"
Request Parameters
| Field | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Audio file to transcribe |
model | string | Yes | Platform model name (e.g., whisper-1) or provider/model for BYOK |
language | string | No | ISO-639-1 code (en, de, fr, etc.) |
prompt | string | No | Guide the transcription style |
response_format | string | No | json, text, srt, verbose_json, vtt (default: json) |
temperature | string | No | Sampling temperature from 0 to 1 as a decimal, e.g. 0.2. A value that is not a number is ignored |
timestamp_granularities[] | array | No | word and/or segment |
Improve accuracy with language
Setting language explicitly improves accuracy and speed —Whisper doesn't have to auto-detect.
Response Formats
Choose the output format that fits your use case:
JSON (default)
{ "text": "Welcome to today's meeting. We'll be discussing the Q4 roadmap..." }
Verbose JSON
Get word-level timestamps for precise alignment:
curl ... -F "response_format=verbose_json" \
-F "timestamp_granularities[]=word" \
-F "timestamp_granularities[]=segment"
{
"task": "transcribe",
"language": "en",
"duration": 45.2,
"text": "Hello and welcome...",
"words": [
{"word": "Hello", "start": 0.0, "end": 0.4},
{"word": "and", "start": 0.5, "end": 0.6},
{"word": "welcome", "start": 0.7, "end": 1.1}
]
}
SRT / VTT Subtitles
For video captioning, use response_format=srt for SRT or response_format=vtt for WebVTT.
Generate Subtitle Files
Generate subtitles
from openai import OpenAI
client = OpenAI(
api_key="sk_your_api_key",
base_url="https://api.2kw.ai/v1"
)
srt = client.audio.transcriptions.create(
model="whisper-1",
file=open("video.mp3", "rb"),
response_format="srt"
)
with open("subtitles.srt", "w") as f:
f.write(srt)
Supported Models
The platform provides whisper-1 out of the box on all tiers.
With BYOK (Team, Business and Enterprise plans), you can route through your own provider:
| Provider | Model | Notes |
|---|---|---|
| OpenAI | openai/whisper-1 | Your own OpenAI Whisper API key |
| Azure OpenAI | azure-openai/{deployment} | Your Azure Whisper deployment |
Errors
Errors use the OpenAI error format with the type invalid_request_error, so the OpenAI SDK raises them as usual:
{
"error": {
"message": "Unsupported audio format: pdf. Supported formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm",
"type": "invalid_request_error"
}
}
| Status | code | Cause |
|---|---|---|
400 | — | The file extension is not one of the supported formats, or response_format is not one of the supported values |
400 | audio_duration_unknown | The built-in whisper-1 could not read the audio's duration before transcribing |
400 | missing_required_part | The request has no file or no model part |
413 | — | The file is larger than 25 MB |
502 | — | The transcription provider failed; retry later |
Unsupported format
The format is checked by file extension, so make sure the file name ends in one of flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm. To convert other audio or video, use a tool such as ffmpeg:
ffmpeg -i input.m4v output.mp3
Built-in model: audio duration unknown
The built-in whisper-1 needs the audio's duration before it transcribes, and reads it from mp3, mp4, m4a, and wav files. Files in other formats, such as ogg, webm, or flac, can be refused with 400 and the code audio_duration_unknown. Convert them to one of those four formats first. A BYOK model accepts every supported format.
File too large
Files over 25 MB are refused with 413. Make the file smaller, or split it and transcribe the parts one by one:
# Lower the bitrate and mix down to mono
ffmpeg -i input.wav -ac 1 -b:a 64k output.mp3
# Split into 10-minute parts
ffmpeg -i input.mp3 -f segment -segment_time 600 output_%03d.mp3