> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.reka.ai/vision/video-features/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server. # Video Features > Process uploaded videos on demand: transcript, captions, embeddings, and objects > **Info** > > For chat over text, images and short videos, use the OpenAI-compatible [Chat Completions API](/chat/overview). Uploading a video does not process it. You choose which features to compute, and each feature enables a different part of the API: | Feature | What it produces | Enables | | ------------ | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | | `transcript` | Speech-to-text transcription, video chunking, and scene boundaries | [Transcript and scenes](/vision/video-insights) | | `captions` | Visual descriptions for each video chunk | [Video chat](/vision/video-chat), [captions](/vision/video-insights), text search over captions | | `embeddings` | Visual and text vector embeddings | [Semantic search](/vision/video-search) | | `objects` | Object detection and tracking on sampled frames | [Objects](/vision/video-insights#list-objects) | Features depend on each other. `captions` and `objects` need `transcript`, and `embeddings` needs both `transcript` and `captions`. Triggering a feature before its dependencies are `ready` returns an error that names the missing dependency. ## Feature statuses Every feature on a video has one of these statuses, visible in the `features` map of [Get Video](/vision/video-management#get-a-video): * **`none`**: Not requested yet. * **`pending`**: Requested and queued. * **`processing`**: Currently running. * **`ready`**: Finished and available. * **`failed`**: Processing failed. * **`blocked`**: Cannot run because a dependency failed. ## Get the feature catalog `GET /v2/features` lists the features you can trigger, with their dependencies. ```bash curl https://vision-agent.api.reka.ai/v2/features \ -H "X-Api-Key: YOUR_API_KEY" ``` ```json { "features": [ { "name": "transcript", "description": "Speech-to-text transcription and video chunking", "depends_on": [], "produces": ["transcript", "scenes"], "note": "Trigger via POST /v2/videos/{id}/features/transcript" }, { "name": "captions", "description": "VLM visual descriptions per video chunk", "depends_on": ["transcript"], "produces": [], "note": null }, { "name": "embeddings", "description": "Vector embeddings (visual + text) for semantic search", "depends_on": ["transcript", "captions"], "produces": [], "note": null }, { "name": "objects", "description": "Object detection and tracking on sampled frames", "depends_on": ["transcript"], "produces": [], "note": null } ] } ``` ## Plan features `POST /v2/videos/{video_id}/features/plan` takes the features you want and tells you what to trigger next. Call it after upload, trigger everything in `actionable`, then call it again until `done` is `true`. #### Bash ```bash curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/plan \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"desired": ["embeddings"]}' ``` #### Python ```python import requests BASE_URL = "https://vision-agent.api.reka.ai" headers = {"X-Api-Key": REKA_API_KEY} video_id = "550e8400-e29b-41d4-a716-446655440000" response = requests.post( f"{BASE_URL}/v2/videos/{video_id}/features/plan", json={"desired": ["embeddings"]}, headers=headers, ) response.raise_for_status() plan = response.json() print(plan["actionable"], plan["done"]) ``` ### Response ```json { "video_id": "550e8400-e29b-41d4-a716-446655440000", "all_required": ["captions", "embeddings", "scenes", "transcript"], "statuses": { "captions": "none", "embeddings": "none", "scenes": "none", "transcript": "none" }, "actionable": ["transcript"], "processing": [], "blocked": [], "reconfigure": {}, "errors": {}, "done": false } ``` * **`all_required`**: The desired features plus everything they depend on. * **`statuses`**: Current status of each required feature. * **`actionable`**: Features whose dependencies are ready, so you can trigger them now. * **`processing`**: Features currently running. * **`blocked`**: Features that cannot run because a dependency failed. * **`reconfigure`**: Features that need a parent feature re-triggered with a different configuration, with the trigger, config, and a message explaining why. * **`errors`**: Error messages for failed features. * **`done`**: `true` once every desired feature is `ready`. ## Trigger a feature Each feature has its own trigger endpoint. All of them accept `force` (default `false`) to re-run a feature that is already `ready`, and return the feature's current status. | Feature | Endpoint | | ------------ | ------------------------------------------------ | | `transcript` | `POST /v2/videos/{video_id}/features/transcript` | | `captions` | `POST /v2/videos/{video_id}/features/captions` | | `embeddings` | `POST /v2/videos/{video_id}/features/embeddings` | | `objects` | `POST /v2/videos/{video_id}/features/objects` | #### Bash ```bash curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/transcript \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{}' ``` #### Python ```python response = requests.post( f"{BASE_URL}/v2/videos/{video_id}/features/transcript", json={}, headers=headers, ) response.raise_for_status() print(response.json()) ``` ### Response ```json { "video_id": "550e8400-e29b-41d4-a716-446655440000", "feature": "transcript", "status": "processing" } ``` A trigger for a feature that is already `ready` returns `ready` without re-running it, unless you pass `force`. A trigger for a feature that is already `processing` returns `processing` without starting a second run. ### Feature options * **Transcript**: `chunking_config` sets `min_len` (default `15.0`) and `max_len` (default `25.0`), the bounds on chunk length in seconds. * **Captions**: `caption_prompt` customizes the prompt used to describe each chunk. * **Objects**: `person_localization` tunes detection with `max_fps` (default `15`), `max_failed_frames` (default `10`), `num_photos_per_person` (default `3`), `max_objects_per_chunk` (default `10`), `conf` (default `0.9`), and `iou` (default `0.7`). Object detection currently tracks people only. * **Embeddings**: No options beyond `force`. ```bash curl -X POST https://vision-agent.api.reka.ai/v2/videos/550e8400-e29b-41d4-a716-446655440000/features/captions \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"caption_prompt": "Describe the products shown and any on-screen text."}' ``` ## End-to-end example This script uploads a video, processes it for semantic search, and waits until it is ready. ```python import time import requests BASE_URL = "https://vision-agent.api.reka.ai" headers = {"X-Api-Key": REKA_API_KEY} upload = requests.post( f"{BASE_URL}/v2/videos", headers=headers, data={"video_url": "https://example.com/video.mp4", "video_name": "demo.mp4"}, ) upload.raise_for_status() video_id = upload.json()["video_id"] while requests.get(f"{BASE_URL}/v2/videos/{video_id}", headers=headers).json()["status"] == "uploading": time.sleep(5) while True: plan = requests.post( f"{BASE_URL}/v2/videos/{video_id}/features/plan", json={"desired": ["embeddings"]}, headers=headers, ).json() if plan["done"]: break if plan["blocked"] or plan["errors"]: raise RuntimeError(plan) for feature in plan["actionable"]: requests.post(f"{BASE_URL}/v2/videos/{video_id}/features/{feature}", json={}, headers=headers) time.sleep(10) print(f"{video_id} is ready for search") ``` > One OpenAI-compatible API for Reka models and a curated selection of open models.