Skip to navigation

Video Insights

For chat over text, images and short videos, use the OpenAI-compatible Chat Completions API.

Once a feature is ready, you can read its output directly. These endpoints share the same query parameters and pagination:

  • start, end (optional): Restrict results to a time range, in seconds.
  • page_limit (optional): Items per page, minimum 1. Defaults to 50.
  • page_token (optional): The next_page_token from a previous response.

Each returns {"data": [...], "next_page_token": ...}. When next_page_token is null, you have the last page.

EndpointReturnsNeeds feature
GET /v2/videos/{video_id}/transcriptSpoken wordstranscript
GET /v2/videos/{video_id}/captionsVisual descriptionscaptions
GET /v2/videos/{video_id}/scenesScene boundariestranscript
GET /v2/videos/{video_id}/objectsDetected objectsobjects

List transcript

GET /v2/videos/{video_id}/transcript returns the transcript. The format query parameter picks the shape: segments (the default) and words return paginated entries, while text returns the whole transcript as a single string.

Bash

curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/transcript?start=0&end=60" \
-H "X-Api-Key: YOUR_API_KEY"

Python

import requests
BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
video_id = "550e8400-e29b-41d4-a716-446655440000"
response = requests.get(
f"{BASE_URL}/v2/videos/{video_id}/transcript",
params={"start": 0, "end": 60},
headers=headers,
)
response.raise_for_status()
for segment in response.json()["data"]:
print(segment["start"], segment["end"], segment["text"])

Response

{
"data": [
{"start": 0.0, "end": 14.2, "text": "Welcome back to the quarterly review."},
{"start": 14.2, "end": 31.8, "text": "As you can see, revenue grew in every region."}
],
"next_page_token": null
}

With format=text:

{
"text": "Welcome back to the quarterly review. As you can see, revenue grew in every region."
}

List captions

GET /v2/videos/{video_id}/captions returns the AI-generated visual description of each chunk.

curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/captions" \
-H "X-Api-Key: YOUR_API_KEY"
{
"data": [
{"start": 0.0, "end": 14.2, "caption": "A presenter stands beside a projected title slide."},
{"start": 14.2, "end": 31.8, "caption": "The presenter points at a bar chart of revenue by region."}
],
"next_page_token": null
}

List scenes

GET /v2/videos/{video_id}/scenes returns detected scene boundaries. Scenes are produced by the transcript feature.

curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/scenes" \
-H "X-Api-Key: YOUR_API_KEY"
{
"data": [
{"index": 0, "start": 0.0, "end": 14.2},
{"index": 1, "start": 14.2, "end": 31.8}
],
"next_page_token": null
}

List objects

GET /v2/videos/{video_id}/objects returns detections from the objects feature, grouped into time segments. The optional type query parameter filters by detection type. Object detection currently tracks people only.

curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/objects?start=0&end=30" \
-H "X-Api-Key: YOUR_API_KEY"
{
"data": [
{
"start": 0.0,
"end": 14.2,
"detections": [
{
"type": "person",
"person_id": null,
"bbox": [412, 188, 655, 710],
"frame_idx": 24,
"frame_timestamp_start": 1.0,
"frame_timestamp_end": 1.04
}
]
}
],
"next_page_token": null
}

Each detection carries its type, a bbox as a list of integers, the frame_idx it came from, and that frame’s start and end timestamps. person_id may be null.

Segment a video

POST /v2/videos/{video_id}/segment detects the objects you describe in text, frame by frame, within a time range of up to 15 seconds. It needs only a finished upload, not any processed feature.

Bash

curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/segment \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompts": [
{"type": "text", "text": "laptop"},
{"type": "text", "text": "red car"}
],
"start": 12.0,
"end": 20.0,
"threshold": 0.4
}'

Python

payload = {
"prompts": [
{"type": "text", "text": "laptop"},
{"type": "text", "text": "red car"},
],
"start": 12.0,
"end": 20.0,
"threshold": 0.4,
}
response = requests.post(f"{BASE_URL}/v2/videos/{video_id}/segment", json=payload, headers=headers)
response.raise_for_status()
for frame in response.json()["frames"]:
for det in frame["detections"]:
print(frame["timestamp"], det["label"], det["score"], det["bbox"])

Parameters

  • prompts (required): 1 to 10 objects to detect. Each prompt is {"type": "text", "text": "..."}, where text is 1 to 200 characters. Every prompt runs against every sampled frame.
  • start (optional): Start of the time range in seconds. Defaults to 0.
  • end (optional): End of the time range in seconds. Defaults to start plus 15 seconds, clamped to the video’s duration. The range can be at most 15 seconds.
  • threshold (optional): Confidence threshold between 0 and 1. Detections scoring below it are dropped. Defaults to 0.3.

Response

{
"frames": [
{
"timestamp": 12.0,
"detections": [
{
"label": "laptop",
"prompt_index": 0,
"score": 0.87,
"bbox": {"x_min": 320.0, "y_min": 210.0, "x_max": 610.0, "y_max": 430.0}
}
]
}
],
"frame_size": {"width": 1920, "height": 1080},
"frame_count": 1
}
  • frames: One entry per sampled frame, with its timestamp and detections. Each detection has the matched label, the prompt_index of the prompt that matched, a score, and a bbox with x_min, y_min, x_max, and y_max.
  • frame_size: Width and height of the analyzed frames.
  • frame_count: Number of frames returned.