Video Features
For chat over text, images and short videos, use the OpenAI-compatible Chat Completions API.
Uploading a video does not process it. You choose which features to compute, and each feature enables a different part of the API:
Features depend on each other. captions and objects need transcript, and embeddings needs both transcript and captions. Triggering a feature before its dependencies are ready returns an error that names the missing dependency.
Feature statuses
Every feature on a video has one of these statuses, visible in the features map of Get Video:
none: Not requested yet.pending: Requested and queued.processing: Currently running.ready: Finished and available.failed: Processing failed.blocked: Cannot run because a dependency failed.
Get the feature catalog
GET /v2/features lists the features you can trigger, with their dependencies.
Plan features
POST /v2/videos/{video_id}/features/plan takes the features you want and tells you what to trigger next. Call it after upload, trigger everything in actionable, then call it again until done is true.
Bash
Python
Response
all_required: The desired features plus everything they depend on.statuses: Current status of each required feature.actionable: Features whose dependencies are ready, so you can trigger them now.processing: Features currently running.blocked: Features that cannot run because a dependency failed.reconfigure: Features that need a parent feature re-triggered with a different configuration, with the trigger, config, and a message explaining why.errors: Error messages for failed features.done:trueonce every desired feature isready.
Trigger a feature
Each feature has its own trigger endpoint. All of them accept force (default false) to re-run a feature that is already ready, and return the feature’s current status.
Bash
Python
Response
A trigger for a feature that is already ready returns ready without re-running it, unless you pass force. A trigger for a feature that is already processing returns processing without starting a second run.
Feature options
- Transcript:
chunking_configsetsmin_len(default15.0) andmax_len(default25.0), the bounds on chunk length in seconds. - Captions:
caption_promptcustomizes the prompt used to describe each chunk. - Objects:
person_localizationtunes detection withmax_fps(default15),max_failed_frames(default10),num_photos_per_person(default3),max_objects_per_chunk(default10),conf(default0.9), andiou(default0.7). Object detection currently tracks people only. - Embeddings: No options beyond
force.
End-to-end example
This script uploads a video, processes it for semantic search, and waits until it is ready.