> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.reka.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server.

# Video Features

> Process uploaded videos on demand: transcript, captions, embeddings, and objects

> **Info**
>
> For chat over text, images and short videos, use the OpenAI-compatible [Chat Completions API](/chat/overview).

Uploading a video does not process it. You choose which features to compute, and each feature enables a different part of the API:

| Feature      | What it produces                                                   | Enables                                                                                         |
| ------------ | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- |
| `transcript` | Speech-to-text transcription, video chunking, and scene boundaries | [Transcript and scenes](/vision/video-insights)                                                 |
| `captions`   | Visual descriptions for each video chunk                           | [Video chat](/vision/video-chat), [captions](/vision/video-insights), text search over captions |
| `embeddings` | Visual and text vector embeddings                                  | [Semantic search](/vision/video-search)                                                         |
| `objects`    | Object detection and tracking on sampled frames                    | [Objects](/vision/video-insights#list-objects)                                                  |

Features depend on each other. `captions` and `objects` need `transcript`, and `embeddings` needs both `transcript` and `captions`. Triggering a feature before its dependencies are `ready` returns an error that names the missing dependency.

## Feature statuses

Every feature on a video has one of these statuses, visible in the `features` map of [Get Video](/vision/video-management#get-a-video):

* **`none`**: Not requested yet.
* **`pending`**: Requested and queued.
* **`processing`**: Currently running.
* **`ready`**: Finished and available.
* **`failed`**: Processing failed.
* **`blocked`**: Cannot run because a dependency failed.

## Get the feature catalog

`GET /v2/features` lists the features you can trigger, with their dependencies.

```bash
curl https://vision-agent.api.reka.ai/v2/features \
  -H "X-Api-Key: YOUR_API_KEY"
```

```json
{
  "features": [
    {
      "name": "transcript",
      "description": "Speech-to-text transcription and video chunking",
      "depends_on": [],
      "produces": ["transcript", "scenes"],
      "note": "Trigger via POST /v2/videos/{id}/features/transcript"
    },
    {
      "name": "captions",
      "description": "VLM visual descriptions per video chunk",
      "depends_on": ["transcript"],
      "produces": [],
      "note": null
    },
    {
      "name": "embeddings",
      "description": "Vector embeddings (visual + text) for semantic search",
      "depends_on": ["transcript", "captions"],
      "produces": [],
      "note": null
    },
    {
      "name": "objects",
      "description": "Object detection and tracking on sampled frames",
      "depends_on": ["transcript"],
      "produces": [],
      "note": null
    }
  ]
}
```

## Plan features

`POST /v2/videos/{video_id}/features/plan` takes the features you want and tells you what to trigger next. Call it after upload, trigger everything in `actionable`, then call it again until `done` is `true`.

#### Bash

```bash
curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/plan \
  -H "X-Api-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"desired": ["embeddings"]}'
```

#### Python

```python
import requests

BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}
video_id = "550e8400-e29b-41d4-a716-446655440000"

response = requests.post(
    f"{BASE_URL}/v2/videos/{video_id}/features/plan",
    json={"desired": ["embeddings"]},
    headers=headers,
)
response.raise_for_status()
plan = response.json()
print(plan["actionable"], plan["done"])
```

### Response

```json
{
  "video_id": "550e8400-e29b-41d4-a716-446655440000",
  "all_required": ["captions", "embeddings", "scenes", "transcript"],
  "statuses": {
    "captions": "none",
    "embeddings": "none",
    "scenes": "none",
    "transcript": "none"
  },
  "actionable": ["transcript"],
  "processing": [],
  "blocked": [],
  "reconfigure": {},
  "errors": {},
  "done": false
}
```

* **`all_required`**: The desired features plus everything they depend on.
* **`statuses`**: Current status of each required feature.
* **`actionable`**: Features whose dependencies are ready, so you can trigger them now.
* **`processing`**: Features currently running.
* **`blocked`**: Features that cannot run because a dependency failed.
* **`reconfigure`**: Features that need a parent feature re-triggered with a different configuration, with the trigger, config, and a message explaining why.
* **`errors`**: Error messages for failed features.
* **`done`**: `true` once every desired feature is `ready`.

## Trigger a feature

Each feature has its own trigger endpoint. All of them accept `force` (default `false`) to re-run a feature that is already `ready`, and return the feature's current status.

| Feature      | Endpoint                                         |
| ------------ | ------------------------------------------------ |
| `transcript` | `POST /v2/videos/{video_id}/features/transcript` |
| `captions`   | `POST /v2/videos/{video_id}/features/captions`   |
| `embeddings` | `POST /v2/videos/{video_id}/features/embeddings` |
| `objects`    | `POST /v2/videos/{video_id}/features/objects`    |

#### Bash

```bash
curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/transcript \
  -H "X-Api-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{}'
```

#### Python

```python
response = requests.post(
    f"{BASE_URL}/v2/videos/{video_id}/features/transcript",
    json={},
    headers=headers,
)
response.raise_for_status()
print(response.json())
```

### Response

```json
{
  "video_id": "550e8400-e29b-41d4-a716-446655440000",
  "feature": "transcript",
  "status": "processing"
}
```

A trigger for a feature that is already `ready` returns `ready` without re-running it, unless you pass `force`. A trigger for a feature that is already `processing` returns `processing` without starting a second run.

### Feature options

* **Transcript**: `chunking_config` sets `min_len` (default `15.0`) and `max_len` (default `25.0`), the bounds on chunk length in seconds.
* **Captions**: `caption_prompt` customizes the prompt used to describe each chunk.
* **Objects**: `person_localization` tunes detection with `max_fps` (default `15`), `max_failed_frames` (default `10`), `num_photos_per_person` (default `3`), `max_objects_per_chunk` (default `10`), `conf` (default `0.9`), and `iou` (default `0.7`). Object detection currently tracks people only.
* **Embeddings**: No options beyond `force`.

```bash
curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/captions \
  -H "X-Api-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"caption_prompt": "Describe the products shown and any on-screen text."}'
```

## End-to-end example

This script uploads a video, processes it for semantic search, and waits until it is ready.

```python
import time
import requests

BASE_URL = "https://vision-agent.api.reka.ai"
headers = {"X-Api-Key": REKA_API_KEY}

upload = requests.post(
    f"{BASE_URL}/v2/videos",
    headers=headers,
    data={"video_url": "https://example.com/video.mp4", "video_name": "demo.mp4"},
)
upload.raise_for_status()
video_id = upload.json()["video_id"]

while requests.get(f"{BASE_URL}/v2/videos/{video_id}", headers=headers).json()["status"] == "uploading":
    time.sleep(5)

while True:
    plan = requests.post(
        f"{BASE_URL}/v2/videos/{video_id}/features/plan",
        json={"desired": ["embeddings"]},
        headers=headers,
    ).json()
    if plan["done"]:
        break
    if plan["blocked"] or plan["errors"]:
        raise RuntimeError(plan)
    for feature in plan["actionable"]:
        requests.post(f"{BASE_URL}/v2/videos/{video_id}/features/{feature}", json={}, headers=headers)
    time.sleep(10)

print(f"{video_id} is ready for search")
```