Skip to main content
POST
curl

Authorizations

Authorization
string
header
required

An account API key. Send it as Authorization: Bearer <key>.

Headers

x-naga-routing-strategy
enum<string>

The routing strategy when the body names none. An unknown name is ignored. A NagaAI extension. The order in which NagaAI tries the providers that serve a model. A NagaAI extension.

Available options:
reliability,
balanced,
latency,
throughput,
cache

Body

application/json

The OpenAI-compatible Responses request. The gateway takes an unrecognized key and drops it, and keeps nothing between requests.

model
string
required

The model to run, named by an id or alias from NagaAI's catalog. An entry without the chat.completions capability draws a refusal.

input
required

The prompt: one string, or the conversation as typed items, oldest first. Required.

One entry of the input list, chosen by its type. Six of twenty-six tags do something.

routing
object | null

How NagaAI orders the providers for this request. It reorders them and never removes one. A NagaAI extension.

instructions
string | null

A system instruction put ahead of the conversation.

max_output_tokens
integer<int64> | null

The ceiling on generated tokens, reasoning included. At least 1, and the gateway raises it where the reasoning budget would not fit.

Required range: x >= 1
temperature
number<double> | null

How much randomness goes into the reply, between 0 and 2. A provider may still refuse a value this range allows.

Required range: 0 <= x <= 2
top_p
number<double> | null

The sampling cut-off, between 0 and 1, and an alternative to temperature.

Required range: 0 <= x <= 1
tools
object[] | null

The tools the model may call. web_search and web_search_preview are the only hosted ones the gateway runs; other types travel no further.

One entry of tools[], chosen by its type. Five of the eighteen types reach the model; the gateway drops the other thirteen.

tool_choice

Which tool, if any, the model is to call. Naming a tool the gateway does not run reads as if the key were absent.

Available options:
auto,
none,
required
parallel_tool_calls
boolean | null

Whether the model may return several tool calls in one turn. It reaches only some providers.

truncation
enum<string> | null

What to do when the conversation outgrows the context window. It reaches only some providers.

Available options:
auto,
disabled
reasoning
object | null

How much reasoning to do, and what summary of it comes back.

text
object | null

The shape of the reply, and how long it should run.

stream
boolean

Whether the reply arrives incrementally as the model writes it. An explicit null draws a refusal.

Response

The response. application/json carries one response object; text/event-stream carries the event-typed lifecycle frames.

A response: the non-streaming body, and the snapshot every lifecycle frame carries.

instructions
string | null
required

The system instructions the turn ran with, or null.

tools
any[]
required

The tools the model was offered, in the shape the gateway took them.

tool_choice
any
required

How the model was told to choose among them.

temperature
any
required

The sampling temperature the turn ran with; 1 when the request set none.

top_p
any
required

The sampling cut-off the turn ran with; 1 when the request set none.

parallel_tool_calls
boolean
required

Whether the model could call several tools at once.

metadata
any
required

The metadata the request sent, echoed back unchanged. The gateway drops it upstream.

text
any
required

The output-format settings the turn ran with.

truncation
any
required

What was to happen if the conversation outgrew the context window.

max_output_tokens
integer<int64> | null
required

The ceiling on generated tokens the turn ran with, or null.

id
string
required

This response's identifier, repeated on every frame. No handle to resume from.

object
enum<string>
required

What kind of object this is. Always response.

Available options:
response
created_at
integer<int64>
required

When the response was created, in seconds since the Unix epoch.

status
string
required

Where the turn stands, as the provider named it.

output
(Message item · object | Reasoning item · object | Synthesized reasoning item · object | Function call item · object | Function call output item · object | Web search call item · object | Image generation call item · object | Image generation call, added view · object)[]
required

What the model produced, item by item, in order.

One item the model produced, addressed by output_index.

error
object | null
required

What went wrong, or null when nothing did.

incomplete_details
object | null
required

Why the turn stopped short, or null.

usage
object | null
required

What the request cost in tokens, or null while the turn runs.

model
string | null
required

The model that answered, by its id in the NagaAI catalog.

reasoning
any

The reasoning settings the turn ran with.