A model’s
supported_endpoints in /v1/models says which of the three it serves.
Chat models that accept audio, such as gemini-3.7-flash, can also take audio inside a chat request. See Multimodal content.Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Text to speech, transcription and translation to English, in the OpenAI Audio format.
| Task | Endpoint | Example models |
|---|---|---|
| Text to speech | POST /v1/audio/speech | gpt-4o-mini-tts, eleven-v3 |
| Speech to text | POST /v1/audio/transcriptions | gpt-4o-transcribe, whisper-large-v3-turbo |
| Speech to English text | POST /v1/audio/translations | whisper-large-v3-turbo, whisper-large-v3 |
supported_endpoints in /v1/models says which of the three it serves.
Chat models that accept audio, such as gemini-3.7-flash, can also take audio inside a chat request. See Multimodal content.