Code for each endpoint: Chat Completions, Responses, Messages.
How NagaAI handles effort
- One scale for every model. NagaAI converts the level for each model. If a model does not support the level you asked for, NagaAI uses the nearest level it does support.
- Budgets become levels. In Messages,
thinking.budget_tokensmaps to a level: 1024 to 8191 islow, up to 32767 ismedium, up to 65535 ishigh, and 65536 or more isxhigh.output_config.effortwins overthinkingwhen you send both.maxruns asxhigh. - Models without reasoning ignore it. If
reasoning_effortis missing from a model’ssupported_parameters, NagaAI removes it from the request. - Reasoning costs the output price. NagaAI bills reasoning tokens as output, even when the provider hides the reasoning text.
Multi-turn conversations
Some models need their earlier reasoning back to keep it across tool calls. Send the previous assistant turn back unchanged: the whole Chat Completions message withreasoning_details, the Responses reasoning items, or the Messages thinking blocks with their signature. NagaAI checks where a signature came from and removes signatures that another provider issued, so a conversation can move between providers without errors.