Skip to main content
POST
curl

Authorizations

Authorization
string
header
required

An account API key. Send it as Authorization: Bearer <key>.

Headers

x-naga-routing-strategy
enum<string>

The routing strategy when the body names none. An unknown name is ignored. A NagaAI extension. The order in which NagaAI tries the providers that serve a model. A NagaAI extension.

Available options:
reliability,
balanced,
latency,
throughput,
cache

Body

application/json

The OpenAI-compatible Chat Completions request. The gateway takes an unrecognized key and drops it.

model
string
required

The model to run, named by an id or alias from NagaAI's catalog. An entry without the chat.completions capability draws a refusal.

messages
object[]
required

The conversation so far, oldest first; one entry at least.

Minimum array length: 1
routing
object | null

How NagaAI orders the providers for this request. It reorders them and never removes one. A NagaAI extension.

tools
object[] | null

The tools the model is allowed to call.

One entry of tools[], told apart by type.

functions
any[] | null

The pre-tools function-calling surface. A request carrying it draws a refusal.

tool_choice

Which tool, if any, the model calls. An explicit null draws a refusal.

response_format
object

The shape the model's output must take.

temperature
number<double> | null

How much randomness goes into the reply, between 0 and 2, and an alternative to top_p.

Required range: 0 <= x <= 2
top_p
number<double> | null

The sampling cut-off, between 0 and 1, and an alternative to temperature.

Required range: 0 <= x <= 1
stream
boolean | null

Whether the reply arrives incrementally as the model writes it.

stream_options
object | null

Extra streaming behavior. include_usage adds one trailing chunk carrying the totals.

stop

Strings that end generation the moment the model produces one.

max_completion_tokens
integer<int64> | null

The ceiling on generated tokens, reasoning included. At least 1.

Required range: x >= 1
max_tokens
integer<int64> | null

An alias of max_completion_tokens, read only when that key is absent.

Required range: x >= 1
reasoning_effort
enum<string> | null

How much reasoning to spend before answering, from none to xhigh.

Available options:
none,
minimal,
low,
medium,
high,
xhigh
presence_penalty
number<double> | null

How hard a token is penalized for having appeared, from -2 to 2.

Required range: -2 <= x <= 2
frequency_penalty
number<double> | null

How hard repetition is penalized, scaled by how often, from -2 to 2.

Required range: -2 <= x <= 2
logit_bias
object | null

A map from token id to sampling bias. The gateway drops it.

parallel_tool_calls
boolean | null

Whether the model may return several tool calls in one turn.

prediction
object | null

Content the reply is expected to reproduce, which the gateway drops.

web_search_options
object | null

Settings for the hosted web-search tool. The gateway drops them.

image_config
object | null

An aspect ratio and a size for image output. A NagaAI extension the gateway drops.

Response

The completion. application/json carries one chat.completion object when stream is false; text/event-stream carries one chat.completion.chunk per event, then the data: [DONE] sentinel, when it is true.

The completion. The gateway answers a request it does not forward with the second shape, which omits logprobs and refusal.

id
string
required

This answer's identifier, repeated on every frame of a stream.

object
string
required

What kind of object this is. Always chat.completion.

Example:

"chat.completion"

created
integer<int64>
required

When the answer was created, in seconds since the Unix epoch.

model
string
required

The model that answered, by its id in the NagaAI catalog.

choices
object[]
required

The answers the model produced — one entry, see index.

usage
object | null

What the request cost in tokens, absent when the provider reported none.