Images and Video
Content parts use the OpenAI shape. A user message’s content can be a list mixing text, image_url and video_url parts. Check a model’s input_modalities on GET /v1/models to see whether it accepts images or video.
Images
Local files as data URLs
To send a local file, base64-encode it into a data URL and pass that as the url.
Multiple images
Add one image_url part per image to the same message. Some models cap the number of images per prompt: reka-edge-2603 accepts at most three and returns a 400 error beyond that.
Video
Models that list video in input_modalities, such as reka-edge-2603 and qwen3.8-27b, accept a video_url part. The URL must be publicly downloadable; sites such as YouTube block direct downloads.
Sampling frames yourself
To control frame sampling or keep uploads small, extract frames and send them as image_url parts. This command extracts one frame per second:
Then send the frames in order, followed by the question. Pick a model whose image limit fits your frame count; reka-edge-2603 accepts at most three images per prompt.
For long videos you want to index, search and query repeatedly, the Vision API uploads and processes the video once.