> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.reka.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server.

# Video Insights

> Read transcripts, captions, scenes, and objects, and detect prompted objects in a video segment

> **Info**
>
> For chat over text, images and short videos, use the OpenAI-compatible [Chat Completions API](/chat/overview).

Once a [feature](/vision/video-features) is `ready`, you can read its output directly. These endpoints share the same query parameters and pagination:

* **`start`**, **`end`** (optional): Restrict results to a time range, in seconds.
* **`page_limit`** (optional): Items per page, minimum 1. Defaults to 50.
* **`page_token`** (optional): The `next_page_token` from a previous response.

Each returns `{"data": [...], "next_page_token": ...}`. When `next_page_token` is `null`, you have the last page.

| Endpoint                               | Returns             | Needs feature |
| -------------------------------------- | ------------------- | ------------- |
| `GET /v2/videos/{video_id}/transcript` | Spoken words        | `transcript`  |
| `GET /v2/videos/{video_id}/captions`   | Visual descriptions | `captions`    |
| `GET /v2/videos/{video_id}/scenes`     | Scene boundaries    | `transcript`  |
| `GET /v2/videos/{video_id}/objects`    | Detected objects    | `objects`     |

## List transcript

`GET /v2/videos/{video_id}/transcript` returns the transcript. The `format` query parameter picks the shape: `segments` (the default) and `words` return paginated entries, while `text` returns the whole transcript as a single string.

#### Bash

```bash
curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/transcript?start=0&end=60" \
  -H "X-Api-Key: YOUR_API_KEY"
```

#### Python

```python
import requests

BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
video_id = "550e8400-e29b-41d4-a716-446655440000"

response = requests.get(
    f"{BASE_URL}/v2/videos/{video_id}/transcript",
    params={"start": 0, "end": 60},
    headers=headers,
)
response.raise_for_status()
for segment in response.json()["data"]:
    print(segment["start"], segment["end"], segment["text"])
```

### Response

```json
{
  "data": [
    {"start": 0.0, "end": 14.2, "text": "Welcome back to the quarterly review."},
    {"start": 14.2, "end": 31.8, "text": "As you can see, revenue grew in every region."}
  ],
  "next_page_token": null
}
```

With `format=text`:

```json
{
  "text": "Welcome back to the quarterly review. As you can see, revenue grew in every region."
}
```

## List captions

`GET /v2/videos/{video_id}/captions` returns the AI-generated visual description of each chunk.

```bash
curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/captions" \
  -H "X-Api-Key: YOUR_API_KEY"
```

```json
{
  "data": [
    {"start": 0.0, "end": 14.2, "caption": "A presenter stands beside a projected title slide."},
    {"start": 14.2, "end": 31.8, "caption": "The presenter points at a bar chart of revenue by region."}
  ],
  "next_page_token": null
}
```

## List scenes

`GET /v2/videos/{video_id}/scenes` returns detected scene boundaries. Scenes are produced by the `transcript` feature.

```bash
curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/scenes" \
  -H "X-Api-Key: YOUR_API_KEY"
```

```json
{
  "data": [
    {"index": 0, "start": 0.0, "end": 14.2},
    {"index": 1, "start": 14.2, "end": 31.8}
  ],
  "next_page_token": null
}
```

## List objects

`GET /v2/videos/{video_id}/objects` returns detections from the `objects` feature, grouped into time segments. The optional `type` query parameter filters by detection type. Object detection currently tracks people only.

```bash
curl "https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/objects?start=0&end=30" \
  -H "X-Api-Key: YOUR_API_KEY"
```

```json
{
  "data": [
    {
      "start": 0.0,
      "end": 14.2,
      "detections": [
        {
          "type": "person",
          "person_id": null,
          "bbox": [412, 188, 655, 710],
          "frame_idx": 24,
          "frame_timestamp_start": 1.0,
          "frame_timestamp_end": 1.04
        }
      ]
    }
  ],
  "next_page_token": null
}
```

Each detection carries its `type`, a `bbox` as a list of integers, the `frame_idx` it came from, and that frame's start and end timestamps. `person_id` may be `null`.

## Segment a video

`POST /v2/videos/{video_id}/segment` detects the objects you describe in text, frame by frame, within a time range of up to 15 seconds. It needs only a finished upload, not any processed feature.

#### Bash

```bash
curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/segment \
  -H "X-Api-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompts": [
      {"type": "text", "text": "laptop"},
      {"type": "text", "text": "red car"}
    ],
    "start": 12.0,
    "end": 20.0,
    "threshold": 0.4
  }'
```

#### Python

```python
payload = {
    "prompts": [
        {"type": "text", "text": "laptop"},
        {"type": "text", "text": "red car"},
    ],
    "start": 12.0,
    "end": 20.0,
    "threshold": 0.4,
}
response = requests.post(f"{BASE_URL}/v2/videos/{video_id}/segment", json=payload, headers=headers)
response.raise_for_status()
for frame in response.json()["frames"]:
    for det in frame["detections"]:
        print(frame["timestamp"], det["label"], det["score"], det["bbox"])
```

### Parameters

* **`prompts`** (required): 1 to 10 objects to detect. Each prompt is `{"type": "text", "text": "..."}`, where `text` is 1 to 200 characters. Every prompt runs against every sampled frame.
* **`start`** (optional): Start of the time range in seconds. Defaults to `0`.
* **`end`** (optional): End of the time range in seconds. Defaults to `start` plus 15 seconds, clamped to the video's duration. The range can be at most 15 seconds.
* **`threshold`** (optional): Confidence threshold between 0 and 1. Detections scoring below it are dropped. Defaults to `0.3`.

### Response

```json
{
  "frames": [
    {
      "timestamp": 12.0,
      "detections": [
        {
          "label": "laptop",
          "prompt_index": 0,
          "score": 0.87,
          "bbox": {"x_min": 320.0, "y_min": 210.0, "x_max": 610.0, "y_max": 430.0}
        }
      ]
    }
  ],
  "frame_size": {"width": 1920, "height": 1080},
  "frame_count": 1
}
```

* **`frames`**: One entry per sampled frame, with its `timestamp` and `detections`. Each detection has the matched `label`, the `prompt_index` of the prompt that matched, a `score`, and a `bbox` with `x_min`, `y_min`, `x_max`, and `y_max`.
* **`frame_size`**: Width and height of the analyzed frames.
* **`frame_count`**: Number of frames returned.