Core concepts
Reasoning
Give a model room to reason before it answers, and set how hard it works.
Reasoning helps on hard tasks: complex logic, math and multi-step agent work. How you control it depends on the model family.
GPT models
Set reasoning_effort on a Chat Completions request: low, medium, high, xhigh or max. Higher effort reasons longer and uses more output tokens; a value the model does not support returns 400.
Claude models
Set the same reasoning_effort on a Claude model: low, medium, high, xhigh or max, with minimal read as low. Each model takes the nearest level it supports, so a valid value never fails. Leave it out, or send none, and Claude answers without reasoning.
Reasoning is off for a request that continues a tool loop (it ends on a tool result) or whose tool_choice forces a tool. The first turn of a tool loop still reasons.
In Claude Code, on the Messages API, the client sets thinking and effort itself. Newer Claude models use an effort control instead of a fixed token budget; Roteo adapts a legacy thinking budget automatically and reports anything it adjusts in the x-roteo-repaired-params response header.
Reading the reasoning
Reasoning comes back as a summary in reasoning_content, on each streamed chunk's delta. A non-streamed Claude response carries it on the message.
Open-weight models
Open-weight models run with reasoning off.
Billing
Reasoning tokens are output tokens, billed at the model's output rate. They are spent from max_tokens, so leave room for the answer.