Endpoints
Messages
POST /v1/messages: the Anthropic Messages API, currently serving Claude Code.
This is the endpoint Claude Code talks to: point Claude Code at Roteo and it calls /v1/messages for you. Requests from other clients currently get 403; for direct Claude requests from your own code, use Chat Completions.
https://api.roteo.ai/v1/messagesRequest parameters
modelstringrequiredThe model ID to call, for example claude-sonnet-4-6. See Models.
max_tokensintegerrequiredThe maximum number of tokens to generate before stopping. Large values are capped by a server ceiling; when that happens the response carries an x-roteo-clamped-max-tokens header with your original value.
messagesarrayrequiredThe conversation so far, as an array of { role, content } objects. role is user or assistant; content is a string or an array of content blocks (text, images, tool results).
systemstring | arrayA system prompt: instructions and context applied to the whole conversation.
temperaturenumberdefault: 1.0Amount of randomness, between 0.0 and 1.0. Lower is more deterministic. Stripped where a model no longer accepts it; see Parameter repair below.
top_pnumberNucleus sampling. Use either temperature or top_p, not both.
top_kintegerOnly sample from the top K options for each token.
stop_sequencesarrayCustom text sequences that will stop generation.
streambooleandefault: falseWhen true, tokens are streamed back as server-sent events. See Streaming.
toolsarrayTool definitions the model may call. See Tool use.
tool_choiceobjectControls whether and which tool the model must use (auto, any, or a specific tool).
thinkingobjectEnable extended thinking with a token budget. On newer models this is adapted automatically to effort. See Reasoning.
metadataobjectAn object with a user_id and other opaque metadata about the request.
Any other field from the Anthropic Messages API is accepted and passed through.
Parameter repair
Newer models reject some legacy sampling controls. Roteo adjusts the request instead of failing it: temperature, top_p and top_k are stripped where a model no longer accepts them, and a legacy fixed-budget thinking is rewritten to the model's effort control. Anything adjusted is listed in the x-roteo-repaired-params response header.
Response
A successful call returns a message object:
The usage block is what you are billed on.
Prompt caching
A stable prefix of the request can be cached, so repeat calls that share it are billed at a fraction of the input price. The last block to cache is marked with cache_control: a system prompt, the tools array or a message block. Claude Code does this for you.
A cache write is billed at 2x the input rate and a cache read at roughly a tenth of it, over a 1-hour cache window. Both appear in the response usage as cache_creation_input_tokens and cache_read_input_tokens.
Context editing
For long agent runs, context_management has the model trim older tool results and turns automatically as the context window fills, instead of resending a hand-pruned history each turn. Roteo forwards it unchanged and adds the beta header it requires, so it behaves the same as calling Anthropic directly.
Response headers
Errors
See Errors for status codes. Common cases: 401 (bad key), 402 (wallet underfunded for this call), 403 (a client other than Claude Code), 429 (rate limited).