Endpoints
Videos
Async, OpenAI-Sora-compatible text-to-video and image-to-video generation.
Generate video from a prompt, or animate a reference image. Video is asynchronous: submit a job, poll until it finishes, then download the MP4.
https://api.roteo.ai/v1/videoshttps://api.roteo.ai/v1/videos/{id}https://api.roteo.ai/v1/videos/{id}/contentLifecycle
- Submit to
/v1/videos. It returns immediately with anidandstatus: "in_progress". - Poll
GET /v1/videos/{id}every few seconds.statusmovesin_progress→completed(orfailed). Rendering takes about 1 to 5 minutes. - Download from
GET /v1/videos/{id}/contentoncestatusiscompleted.
You are charged per delivered clip, by duration; a failed job is free.
Submit
modelstringrequiredAny video model, e.g. pixverse-v6. See Models: sizes, durations and whether an image is required all vary per model.
promptstringrequiredRequired in every mode. When you supply a starting frame, describe the motion rather than the scene. Up to 8,000 characters.
sizestringdefault: 1280x720Pixverse: 1280x720 (landscape) or 720x1280 (portrait). Seedance: 854x480, 480x854, 1280x720 or 720x1280. Wan is landscape only, at 640x480, 1280x720 or 1920x1080.
secondsintegerdefault: 4Pixverse: 4, 8 or 12. Wan: 5 or 8. Seedance: 5, 8, 12, 16, 20, 24 or 30. A value one model accepts another rejects with 400.
input_referencefileReference image (multipart requests only). PNG, JPEG, or WebP, up to 3.5MB each and 4MB in total.
Optional on Pixverse and Seedance, required on Wan, which animates a starting frame and returns 400 without one.
Repeat the part to send several (-F input_reference=@a.jpg -F input_reference=@b.jpg). One reference is a starting frame; two or more are subject and style references, which only seedance-2.5 accepts (up to 30). Sending more than a model takes returns 400 rather than quietly dropping the extras.
A single Seedance frame is cover-cropped to the video size for you. Pixverse and Wan take their aspect ratio from the frame itself, so nothing is cropped, and multi-image references are never cropped.
last_referencefileOptional LAST frame for first+last-frame animation (seedance-2.5 only). Send it alongside exactly one input_reference (the first frame) and Seedance animates the transition between them: a controlled morph, or continuation from a previous clip. Sending it with no input_reference, with several, or on another model returns 400.
Text-to-video
Image-to-video
Send multipart/form-data with an input_reference image plus the same scalar fields.
Reference-to-video
seedance-2.5 only. Send two or more input_reference images (up to 30) and cite them in the prompt as @image1, @image2, and so on. They are treated as subject and style references rather than a starting frame, so none of them is cropped.
Choosing a reference image
Use a detailed photo or rich scene: the model animates from the reference, so it needs real structure to work with. A flat or near-blank image (a logo, icon or solid color) gives it nothing to animate and can stall or produce poor motion. Transparency is flattened to a solid background.
Pixverse and Wan take the clip's aspect ratio from the frame, so send it already framed the way you want the video. Seedance cover-crops a single frame for you, and multi-image references are never cropped.