Core concepts

Streaming

Receive the answer token by token as it is generated, over server-sent events.

Set "stream": true on a Chat Completions request. The events match the OpenAI Chat Completions API, so any compatible client parses them unchanged, and the stream ends with data: [DONE].

curl -N https://api.roteo.ai/v1/chat/completions \
-H "Authorization: Bearer $ROTEO_API_KEY" \
-H "content-type: application/json" \
-d '{
  "model": "gpt-6.1-sol",
  "stream": true,
  "messages": [{ "role": "user", "content": "Write a haiku about the ocean." }]
}'
bash · 8 lines

The same request streams for Claude, GPT and open-weight models. Claude Code streams over the Messages API on its own, with nothing to configure.

If you ask for more max_tokens than the model allows, Roteo lowers it and reports your original value in the x-roteo-clamped-max-tokens response header.

Streaming | Roteo