Endpoints
Chat Completions
POST /v1/chat/completions: the OpenAI-compatible endpoint. Serves Claude, GPT and open-weight models.
A drop-in OpenAI Chat Completions API: the OpenAI SDK, chat apps and anything that speaks /v1/chat/completions work when pointed at the Roteo base URL. Routing is by the model field.
https://api.roteo.ai/v1/chat/completions- Base URL:
https://api.roteo.ai/v1 - Auth:
Authorization: Bearer sk_your_key
Request parameters
modelstringrequiredThe model ID to call: a claude-*, gpt-* or open-weight id, for example claude-sonnet-4-6, gpt-5.6-sol or kimi-k3. See Models. IDs are served verbatim, so copy them exactly.
messagesarrayrequiredThe conversation, as OpenAI { role, content } objects. role is system, developer, user, assistant, or tool. content is a string or an array of parts (text, image_url).
max_tokensintegerMaximum output tokens. Also accepts max_completion_tokens. Capped by a server ceiling; when clamped, the response carries an x-roteo-clamped-max-tokens header.
streambooleandefault: falseWhen true, tokens stream back as chat.completion.chunk server-sent events, ending with data: [DONE]. See Streaming.
stream_optionsobjectSet { "include_usage": true } to receive a final usage-only chunk (choices: [], usage: {...}) before [DONE].
toolsarrayFunction tools the model may call, in OpenAI { type: "function", function: {...} } shape. See Tool use.
tool_choicestring | objectauto, none, required, or a specific { type: "function", function: { name } }.
reasoning_effortstringlow, medium, high, xhigh or max on GPT and Claude models; Claude also takes none and minimal, and uses the nearest level it supports. Not used by open-weight models. See Reasoning.
temperaturenumberdefault: 1.0Sampling randomness. Stripped automatically for models that no longer accept it.
top_pnumberNucleus sampling. Use either temperature or top_p.
stopstring | arrayUp to four sequences where generation stops.
Unknown fields are passed through to the model.
Example
Model families on this endpoint
- GPT models work with everything, including agent frameworks. See Agents & OpenAI-compatible clients.
- Open-weight models work for direct requests and with agent frameworks; they are the recommended choice for opencode.
- Claude models work for direct requests (chat, single calls, SDK use). Agent frameworks are refused on them; for a Claude agent, use Claude Code.
The full matrix is in the model support table.
Response
Streaming responses arrive as chat.completion.chunk events; the usage chunk appears last when you set stream_options.include_usage.
Response headers
Errors
See Errors. Common cases: 400 (invalid request or unknown model), 401 (bad key), 402 (wallet underfunded for this call), 429 (rate limited).