OpenAI-compatible
Speech to text
Use the Privatemode speech-to-text API to generate text from audio files. The API is compatible with the OpenAI transcriptions API. To generate text from audio, send your requests to the Privatemode proxy. Audio requests and responses are encrypted, both in transit and during processing.
Request body
modelstringrequired
The name of the model to use for transcription, e.g., whisper-large-v3.
filefilerequired
The audio file to transcribe. Supported formats are flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, and webm.
Privatemode's serving backend only supports files up to 50 MB in size. For larger files, consider splitting the audio into smaller segments, or try compressing the file to reduce its size.
languagestring
The language of the audio in ISO-639-1 (e.g., en) format. Not setting the correct language can lead to poor accuracy and performance.
promptstring
An optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
response_formatstringdefault: json
The format of the response. Options are json (default) or verbose_json. The verbose_json format additionally returns duration and segments with timestamps and metadata for each transcribed segment. Note: verbose_json requires setting language.
streambooleandefault: false
Whether to stream the response as server-sent events, following the OpenAI transcription streaming format.
For additional parameters see the vLLM transcriptions API documentation.
Returns
The response is a transcription object or a stream of transcription events.
textstring
The transcribed text from the audio.
usageobject
Duration-based usage statistics.
Show properties
typestring
Always "duration".
secondsinteger
The billed duration of the audio in seconds.
durationfloat
The duration of the audio in seconds. Only returned with verbose_json.
segmentslist
Metadata for each transcribed segment. Only returned with verbose_json.
Show properties
idinteger
The segment ID.
seekinteger
The seek offset of the segment.
startfloat
The start time of the segment in seconds.
endfloat
The end time of the segment in seconds.
textstring
The transcribed text of the segment.
tokenslist
The token IDs of the segment.
temperaturefloat
The temperature used for the segment.
avg_logprobfloat
The average log probability of the segment.
compression_ratiofloat
The compression ratio of the segment.
no_speech_probfloat
The probability that the segment contains no speech.
Other fields are consistent with the OpenAI API specifications.