Skip to navigation

Video Features

For chat over text, images and short videos, use the OpenAI-compatible Chat Completions API.

Uploading a video does not process it. You choose which features to compute, and each feature enables a different part of the API:

FeatureWhat it producesEnables
transcriptSpeech-to-text transcription, video chunking, and scene boundariesTranscript and scenes
captionsVisual descriptions for each video chunkVideo chat, captions, text search over captions
embeddingsVisual and text vector embeddingsSemantic search
objectsObject detection and tracking on sampled framesObjects

Features depend on each other. captions and objects need transcript, and embeddings needs both transcript and captions. Triggering a feature before its dependencies are ready returns an error that names the missing dependency.

Feature statuses

Every feature on a video has one of these statuses, visible in the features map of Get Video:

  • none: Not requested yet.
  • pending: Requested and queued.
  • processing: Currently running.
  • ready: Finished and available.
  • failed: Processing failed.
  • blocked: Cannot run because a dependency failed.

Get the feature catalog

GET /v2/features lists the features you can trigger, with their dependencies.

curl https://vision-agent.api.reka.ai/v2/features \
-H "X-Api-Key: YOUR_API_KEY"
{
"features": [
{
"name": "transcript",
"description": "Speech-to-text transcription and video chunking",
"depends_on": [],
"produces": ["transcript", "scenes"],
"note": "Trigger via POST /v2/videos/{id}/features/transcript"
},
{
"name": "captions",
"description": "VLM visual descriptions per video chunk",
"depends_on": ["transcript"],
"produces": [],
"note": null
},
{
"name": "embeddings",
"description": "Vector embeddings (visual + text) for semantic search",
"depends_on": ["transcript", "captions"],
"produces": [],
"note": null
},
{
"name": "objects",
"description": "Object detection and tracking on sampled frames",
"depends_on": ["transcript"],
"produces": [],
"note": null
}
]
}

Plan features

POST /v2/videos/{video_id}/features/plan takes the features you want and tells you what to trigger next. Call it after upload, trigger everything in actionable, then call it again until done is true.

Bash

curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/plan \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"desired": ["embeddings"]}'

Python

import requests
BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
video_id = "550e8400-e29b-41d4-a716-446655440000"
response = requests.post(
f"{BASE_URL}/v2/videos/{video_id}/features/plan",
json={"desired": ["embeddings"]},
headers=headers,
)
response.raise_for_status()
plan = response.json()
print(plan["actionable"], plan["done"])

Response

{
"video_id": "550e8400-e29b-41d4-a716-446655440000",
"all_required": ["captions", "embeddings", "scenes", "transcript"],
"statuses": {
"captions": "none",
"embeddings": "none",
"scenes": "none",
"transcript": "none"
},
"actionable": ["transcript"],
"processing": [],
"blocked": [],
"reconfigure": {},
"errors": {},
"done": false
}
  • all_required: The desired features plus everything they depend on.
  • statuses: Current status of each required feature.
  • actionable: Features whose dependencies are ready, so you can trigger them now.
  • processing: Features currently running.
  • blocked: Features that cannot run because a dependency failed.
  • reconfigure: Features that need a parent feature re-triggered with a different configuration, with the trigger, config, and a message explaining why.
  • errors: Error messages for failed features.
  • done: true once every desired feature is ready.

Trigger a feature

Each feature has its own trigger endpoint. All of them accept force (default false) to re-run a feature that is already ready, and return the feature’s current status.

FeatureEndpoint
transcriptPOST /v2/videos/{video_id}/features/transcript
captionsPOST /v2/videos/{video_id}/features/captions
embeddingsPOST /v2/videos/{video_id}/features/embeddings
objectsPOST /v2/videos/{video_id}/features/objects

Bash

curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/transcript \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{}'

Python

response = requests.post(
f"{BASE_URL}/v2/videos/{video_id}/features/transcript",
json={},
headers=headers,
)
response.raise_for_status()
print(response.json())

Response

{
"video_id": "550e8400-e29b-41d4-a716-446655440000",
"feature": "transcript",
"status": "processing"
}

A trigger for a feature that is already ready returns ready without re-running it, unless you pass force. A trigger for a feature that is already processing returns processing without starting a second run.

Feature options

  • Transcript: chunking_config sets min_len (default 15.0) and max_len (default 25.0), the bounds on chunk length in seconds.
  • Captions: caption_prompt customizes the prompt used to describe each chunk.
  • Objects: person_localization tunes detection with max_fps (default 15), max_failed_frames (default 10), num_photos_per_person (default 3), max_objects_per_chunk (default 10), conf (default 0.9), and iou (default 0.7). Object detection currently tracks people only.
  • Embeddings: No options beyond force.
curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/captions \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"caption_prompt": "Describe the products shown and any on-screen text."}'

End-to-end example

This script uploads a video, processes it for semantic search, and waits until it is ready.

import time
import requests
BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
upload = requests.post(
f"{BASE_URL}/v2/videos",
headers=headers,
data={"video_url": "https://example.com/video.mp4", "video_name": "demo.mp4"},
)
upload.raise_for_status()
video_id = upload.json()["video_id"]
while requests.get(f"{BASE_URL}/v2/videos/{video_id}", headers=headers).json()["status"] == "uploading":
time.sleep(5)
while True:
plan = requests.post(
f"{BASE_URL}/v2/videos/{video_id}/features/plan",
json={"desired": ["embeddings"]},
headers=headers,
).json()
if plan["done"]:
break
if plan["blocked"] or plan["errors"]:
raise RuntimeError(plan)
for feature in plan["actionable"]:
requests.post(f"{BASE_URL}/v2/videos/{video_id}/features/{feature}", json={}, headers=headers)
time.sleep(10)
print(f"{video_id} is ready for search")