Get started

Models

Every model ID you can call, and how to choose between them.

Pass one of the model IDs below in the model field of your request. New models land as they ship, and live pricing for each is in the dashboard.

Choosing a model

  • Claude models (claude-*) are best used with Claude Code. For direct requests, call them through Chat Completions; the Messages API currently serves Claude Code.
  • GPT models (gpt-*) work with any OpenAI-compatible client: the Chat Completions and Responses APIs, the OpenAI SDK, Codex CLI, chat apps and coding agents.
  • Open-weight models work through Chat Completions, for direct requests and agent frameworks.

For an agent framework, use an open-weight or GPT model. See the model support matrix.

Claude models

Call these from Claude Code, or through /v1/chat/completions for direct requests.

Model IDContextBest for
claude-fable-5-11MThe newest Fable, the premium flagship.
claude-fable-51MThe previous Fable, at the top of the lineup.
claude-opus-5-51MThe newest Opus, for the hardest reasoning and agentic work.
claude-opus-51MThe previous Opus, for demanding reasoning and agentic work.
claude-opus-4-81MTop-tier reasoning and coding.
claude-opus-4-71MA pinned earlier Opus snapshot.
claude-opus-4-61MA pinned earlier Opus snapshot.
claude-sonnet-5-51MThe newest Sonnet, a strong, fast default.
claude-sonnet-51MThe previous Sonnet, fast and capable.
claude-sonnet-4-61MA pinned Sonnet snapshot, fast and balanced.
claude-haiku-4-5200KThe fastest and most economical option.

All text models support tool use. See Tool use and Reasoning.

GPT models

Call these through the OpenAI-compatible /v1/chat/completions or /v1/responses endpoints.

Model IDContextBest for
gpt-6-astra1MThe flagship, for the most demanding agentic work.
gpt-6.1-sol1MThe newest Sol, fast and economical for everyday agentic work.
gpt-6-luna1MThe lightest and most economical option.
gpt-5.6-sol1MThe previous Sol, a pinned agentic model.
gpt-5.6-terra1MA balanced mid-tier option.

GPT models support tool calling and reasoning. See GPT (OpenAI-compatible) to connect the OpenAI SDK, Cursor, or curl.

Open-weight models

Call these through the OpenAI-compatible /v1/chat/completions endpoint; they aren't reachable through /v1/messages or /v1/responses. They are the recommended choice for opencode.

Model IDContextBest for
deepseek-v4-flash-07311MA fast, economical open-weight model from DeepSeek.
qwen3.8-max1MAlibaba's flagship Qwen model.
kimi-k31MMoonshot AI's Kimi model.
glm-5.3256KZ.ai's agent-oriented GLM, built for reasoning and tool use.
glm-5.3-flash1MThe cheapest model on this endpoint. Natively multimodal, with the larger context of the two GLM tiers.

Image models

Call these through /v1/images/generations and /v1/images/edits.

Model IDBest forAccepts reference images
nano-banana-proHighest quality text-to-image and editing.Yes
nano-banana-2Fast, economical generation.Yes
seedream-proOpen-weight generation and editing.Yes
flux-uncensoredOpen-weight generation, lowest cost per image.No

Every image model supports text-to-image, a range of sizes, and all ten aspect ratios.

Editing is per model. /v1/images/edits needs a model from the "accepts reference images" column above; sending one to a model without that support returns 400. flux-uncensored is generation-only.

seedream-pro takes up to 10 references for multi-image blends and subject consistency; cite each in the prompt as @image1, @image2, and so on. The other editors take a single reference image. See the Seedream guide.

Video models

Call these through the async /v1/videos endpoint.

Model IDBest forInputDurations
seedance-2.5Long clips with synchronized audio, and multi-image references.Text, image, or references5, 8, 12, 16, 20, 24, 30s
pixverse-v6Fast general video with synchronized audio.Text or image4, 8, 12s
wan-spicyOpen-weight animation, lowest cost per clip.Image only5, 8s

Durations are per model: a value one model accepts, another rejects. pixverse-v6 renders landscape or portrait at 720p and every clip includes synchronized audio.

seedance-2.5 is the only video model that takes more than one reference. How many input_reference files you send picks the mode:

  • None: text-to-video from the prompt alone.
  • One: animates that image as the first frame.
  • Two or more (up to 30): uses them as subject and style references rather than frames. Cite them in the prompt as @image1, @image2, and so on.

It renders at 480p or 720p, landscape (854x480, 1280x720) or portrait (480x854, 720x1280), and generates synchronized audio at no extra cost. Clips run up to 30 seconds, and price scales with both duration and resolution.

wan-spicy is image-to-video only. It animates a starting image, so a request without one returns 400. It is landscape only, and the clip takes its aspect ratio from the frame you supply. It renders at 480p, 720p or 1080p.

A prompt is required in every mode. When you supply a starting frame it should describe the motion, not the scene: the scene is already the frame you uploaded.

Pricing

Live per-token and per-image prices for every model are on the pricing page.

Listing models in code

Call GET /v1/models for the OpenAI-style JSON list, or point an agent at /v1/llms.txt for the live models and endpoints in plain text.

Models | Roteo