Skip to navigation

Images and Video

Content parts use the OpenAI shape. A user message’s content can be a list mixing text, image_url and video_url parts. Check a model’s input_modalities on GET /v1/models to see whether it accepts images or video.

Images

Image of a cat on a keyboard
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.reka.ai/v1",
api_key=os.environ["REKA_API_KEY"],
)
response = client.chat.completions.create(
model="reka-edge-2603",
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://v0.docs.reka.ai/_images/000000245576.jpg"}},
{"type": "text", "text": "What animal is this? Answer briefly."},
],
}
],
)
print(response.choices[0].message.content)

Local files as data URLs

To send a local file, base64-encode it into a data URL and pass that as the url.

import base64
with open("cat.jpg", "rb") as f:
data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="reka-edge-2603",
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": data_url}},
{"type": "text", "text": "Describe this image."},
],
}
],
)

Multiple images

Add one image_url part per image to the same message. Some models cap the number of images per prompt: reka-edge-2603 accepts at most three and returns a 400 error beyond that.

response = client.chat.completions.create(
model="reka-edge-2603",
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/image_1.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/image_2.jpg"}},
{"type": "text", "text": "What colours and shapes are present in both images?"},
],
}
],
)

Video

Models that list video in input_modalities, such as reka-edge-2603 and qwen3.8-27b, accept a video_url part. The URL must be publicly downloadable; sites such as YouTube block direct downloads.

response = client.chat.completions.create(
model="reka-edge-2603",
messages=[
{
"role": "user",
"content": [
{"type": "video_url", "video_url": {"url": "https://demo-videos-bucket-reka.s3.eu-west-2.amazonaws.com/evYsMdBrFXc.mp4"}},
{"type": "text", "text": "Describe this video in one sentence."},
],
}
],
)
print(response.choices[0].message.content)

Sampling frames yourself

To control frame sampling or keep uploads small, extract frames and send them as image_url parts. This command extracts one frame per second:

ffmpeg -i input.mp4 -vf "fps=1" -q:v 2 frame_%03d.jpg

Then send the frames in order, followed by the question. Pick a model whose image limit fits your frame count; reka-edge-2603 accepts at most three images per prompt.

import base64
import glob
content = []
for path in sorted(glob.glob("frame_*.jpg")):
with open(path, "rb") as f:
content.append({
"type": "image_url",
"image_url": {"url": "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()},
})
content.append({"type": "text", "text": "Describe what happens in this video."})
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": content}],
)
print(response.choices[0].message.content)

For long videos you want to index, search and query repeatedly, the Vision API uploads and processes the video once.