Core concepts

Reasoning

Give a model room to reason before it answers, and set how hard it works.

Reasoning helps on hard tasks: complex logic, math and multi-step agent work. How you control it depends on the model family.

GPT models

Set reasoning_effort on a Chat Completions request: low, medium, high, xhigh or max. Higher effort reasons longer and uses more output tokens; a value the model does not support returns 400.

curl https://api.roteo.ai/v1/chat/completions \
-H "Authorization: Bearer $ROTEO_API_KEY" \
-H "content-type: application/json" \
-d '{
  "model": "gpt-6.1-sol",
  "reasoning_effort": "high",
  "messages": [{ "role": "user", "content": "Prove that sqrt(2) is irrational." }]
}'
bash · 8 lines

Claude models

Set the same reasoning_effort on a Claude model: low, medium, high, xhigh or max, with minimal read as low. Each model takes the nearest level it supports, so a valid value never fails. Leave it out, or send none, and Claude answers without reasoning.

Reasoning is off for a request that continues a tool loop (it ends on a tool result) or whose tool_choice forces a tool. The first turn of a tool loop still reasons.

In Claude Code, on the Messages API, the client sets thinking and effort itself. Newer Claude models use an effort control instead of a fixed token budget; Roteo adapts a legacy thinking budget automatically and reports anything it adjusts in the x-roteo-repaired-params response header.

Reading the reasoning

Reasoning comes back as a summary in reasoning_content, on each streamed chunk's delta. A non-streamed Claude response carries it on the message.

Open-weight models

Open-weight models run with reasoning off.

Billing

Reasoning tokens are output tokens, billed at the model's output rate. They are spent from max_tokens, so leave room for the answer.

Reasoning | Roteo