Skip to navigation

Video Chat

For chat over text, images and short videos, use the OpenAI-compatible Chat Completions API.

POST /v2/chat answers natural language questions about up to five of your videos at once. Answers draw on each video’s transcript and captions, and can also look at the frames of a specific time range.

To ask about a short clip without uploading it first, send the video inline with the Chat Completions API with a video input.

Prerequisites

Every video in the request needs the captions feature to be ready. If it is not, the API returns an error listing the missing features. To prepare a video, call Plan Features with {"desired": ["captions"]} and trigger what it returns.

Ask a question

Bash

curl -X POST https://vision-agent.api.reka.ai/v2/chat \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "What is happening in this video?"}
],
"context": [
{"video_id": "550e8400-e29b-41d4-a716-446655440000"}
]
}'

Python

import requests
BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
payload = {
"messages": [
{"role": "user", "content": "What is happening in this video?"}
],
"context": [
{"video_id": "550e8400-e29b-41d4-a716-446655440000"}
],
}
response = requests.post(f"{BASE_URL}/v2/chat", json=payload, headers=headers)
response.raise_for_status()
print(response.json()["response"])

Response

{
"response": "A presenter walks through a slide deck, pausing on a bar chart of quarterly revenue.",
"model": "MODEL_NAME"
}
  • response: The answer to the question.
  • model: The model that generated the answer.

Request parameters

  • messages (required): The conversation so far, 1 to 20 messages. Each message has a role (user or assistant) and a non-empty content string. The last message must be from the user.
  • context (required): The videos to analyze, 1 to 5 entries. Each entry has:
    • video_id (required): The video to analyze.
    • start (optional): Start of a time range in seconds. Setting it turns on visual analysis, which extracts and looks at frames from that range. The video’s upload must be complete.
    • end (optional): End of the time range in seconds. Defaults to start plus 10 seconds, clamped to the video’s duration.

Multi-turn conversations

Send earlier answers back as assistant messages to ask follow-up questions.

curl -X POST https://vision-agent.api.reka.ai/v2/chat \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "What is happening in this video?"},
{"role": "assistant", "content": "The video shows a person walking on the beach during sunset."},
{"role": "user", "content": "What color is the car in the background?"}
],
"context": [
{"video_id": "550e8400-e29b-41d4-a716-446655440000"}
]
}'

Ask about a specific moment

Pair chat with video search: take the video_id, start, and end of a search result and pass them as context, so the answer looks at the frames of that moment.

search = requests.post(
f"{BASE_URL}/v2/search",
json={"query": "person unboxing a laptop", "page_limit": 3},
headers=headers,
).json()
context = [
{"video_id": r["video_id"], "start": r["start"], "end": r["end"]}
for r in search["data"]
]
answer = requests.post(
f"{BASE_URL}/v2/chat",
json={
"messages": [{"role": "user", "content": "Which laptop brand is being unboxed in each clip?"}],
"context": context,
},
headers=headers,
).json()
print(answer["response"])

Question examples

  • General: “What is happening in this video?”
  • Specific: “What color is the car in the video?”
  • Temporal: “What happens at the beginning of the video?”
  • Comparative: “How do the two product demos differ?”
  • Descriptive: “Describe the setting and atmosphere.”

Error handling

  • Video not found: Check that each video_id in context exists and belongs to you.
  • Video not ready: Trigger the captions feature and wait for it to be ready.
  • Upload not complete: Wait for the upload to finish before sending a start time.