Endpoints

Chat Completions

POST /v1/chat/completions: the OpenAI-compatible endpoint. Serves Claude, GPT and open-weight models.

A drop-in OpenAI Chat Completions API: the OpenAI SDK, chat apps and anything that speaks /v1/chat/completions work when pointed at the Roteo base URL. Routing is by the model field.

POSThttps://api.roteo.ai/v1/chat/completions
  • Base URL: https://api.roteo.ai/v1
  • Auth: Authorization: Bearer sk_your_key

Request parameters

modelstringrequired

The model ID to call: a claude-*, gpt-* or open-weight id, for example claude-sonnet-4-6, gpt-5.6-sol or kimi-k3. See Models. IDs are served verbatim, so copy them exactly.

messagesarrayrequired

The conversation, as OpenAI { role, content } objects. role is system, developer, user, assistant, or tool. content is a string or an array of parts (text, image_url).

max_tokensinteger

Maximum output tokens. Also accepts max_completion_tokens. Capped by a server ceiling; when clamped, the response carries an x-roteo-clamped-max-tokens header.

streambooleandefault: false

When true, tokens stream back as chat.completion.chunk server-sent events, ending with data: [DONE]. See Streaming.

stream_optionsobject

Set { "include_usage": true } to receive a final usage-only chunk (choices: [], usage: {...}) before [DONE].

toolsarray

Function tools the model may call, in OpenAI { type: "function", function: {...} } shape. See Tool use.

tool_choicestring | object

auto, none, required, or a specific { type: "function", function: { name } }.

reasoning_effortstring

low, medium, high, xhigh or max on GPT and Claude models; Claude also takes none and minimal, and uses the nearest level it supports. Not used by open-weight models. See Reasoning.

temperaturenumberdefault: 1.0

Sampling randomness. Stripped automatically for models that no longer accept it.

top_pnumber

Nucleus sampling. Use either temperature or top_p.

stopstring | array

Up to four sequences where generation stops.

Unknown fields are passed through to the model.

Example

curl https://api.roteo.ai/v1/chat/completions \
-H "Authorization: Bearer sk_your_key" \
-H "content-type: application/json" \
-d '{
  "model": "gpt-5.6-sol",
  "messages": [{ "role": "user", "content": "Hello, Roteo!" }]
}'
bash · 7 lines

Model families on this endpoint

  • GPT models work with everything, including agent frameworks. See Agents & OpenAI-compatible clients.
  • Open-weight models work for direct requests and with agent frameworks; they are the recommended choice for opencode.
  • Claude models work for direct requests (chat, single calls, SDK use). Agent frameworks are refused on them; for a Claude agent, use Claude Code.

The full matrix is in the model support table.

Response

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1700000000,
  "model": "gpt-5.6-sol",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello! How can I help?" },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 12, "completion_tokens": 8, "total_tokens": 20 }
}
json · 14 lines

Streaming responses arrive as chat.completion.chunk events; the usage chunk appears last when you set stream_options.include_usage.

Response headers

HeaderMeaning
x-request-idTrace ID for the request.
x-roteo-billing-chainThe chain this call settled on.
x-roteo-clamped-max-tokensPresent when max_tokens was reduced to the server ceiling.

Errors

See Errors. Common cases: 400 (invalid request or unknown model), 401 (bad key), 402 (wallet underfunded for this call), 429 (rate limited).

Chat Completions | Roteo