Video Insights
For chat over text, images and short videos, use the OpenAI-compatible Chat Completions API.
Once a feature is ready, you can read its output directly. These endpoints share the same query parameters and pagination:
start,end(optional): Restrict results to a time range, in seconds.page_limit(optional): Items per page, minimum 1. Defaults to 50.page_token(optional): Thenext_page_tokenfrom a previous response.
Each returns {"data": [...], "next_page_token": ...}. When next_page_token is null, you have the last page.
List transcript
GET /v2/videos/{video_id}/transcript returns the transcript. The format query parameter picks the shape: segments (the default) and words return paginated entries, while text returns the whole transcript as a single string.
Bash
Python
Response
With format=text:
List captions
GET /v2/videos/{video_id}/captions returns the AI-generated visual description of each chunk.
List scenes
GET /v2/videos/{video_id}/scenes returns detected scene boundaries. Scenes are produced by the transcript feature.
List objects
GET /v2/videos/{video_id}/objects returns detections from the objects feature, grouped into time segments. The optional type query parameter filters by detection type. Object detection currently tracks people only.
Each detection carries its type, a bbox as a list of integers, the frame_idx it came from, and that frame’s start and end timestamps. person_id may be null.
Segment a video
POST /v2/videos/{video_id}/segment detects the objects you describe in text, frame by frame, within a time range of up to 15 seconds. It needs only a finished upload, not any processed feature.
Bash
Python
Parameters
prompts(required): 1 to 10 objects to detect. Each prompt is{"type": "text", "text": "..."}, wheretextis 1 to 200 characters. Every prompt runs against every sampled frame.start(optional): Start of the time range in seconds. Defaults to0.end(optional): End of the time range in seconds. Defaults tostartplus 15 seconds, clamped to the video’s duration. The range can be at most 15 seconds.threshold(optional): Confidence threshold between 0 and 1. Detections scoring below it are dropped. Defaults to0.3.
Response
frames: One entry per sampled frame, with itstimestampanddetections. Each detection has the matchedlabel, theprompt_indexof the prompt that matched, ascore, and abboxwithx_min,y_min,x_max, andy_max.frame_size: Width and height of the analyzed frames.frame_count: Number of frames returned.