Skip to main content
Both endpoints take a multipart form and return {"text": "..."}.
For English text from speech in another language, call client.audio.translations.create with whisper-large-v3-turbo or whisper-large-v3, or POST /v1/audio/translations with the same form. The answer is always JSON with text. NagaAI ignores response_format, timestamp_granularities and temperature, so there are no timestamps, SRT or VTT output.