Chat Completions
POST https://api.reka.ai/v1/chat/completions takes the OpenAI Chat Completions request and returns the OpenAI response shape. Authenticate with Authorization: Bearer $REKA_API_KEY, or let the OpenAI SDK set the header from api_key.
Basic request
The response follows the OpenAI shape:
metadata.weight_version names the checkpoint that served the request.
System prompts and multi-turn conversations
Send the whole conversation in messages on every request. Roles are system, user, assistant and tool.
Sampling parameters
- max_tokens: the maximum number of tokens to generate, capped by the model’s maximum output length. If
finish_reasonis"length", the output was cut off. - temperature: lower values give more predictable output, higher values more varied output.
- top_p: nucleus sampling.
- stop: a string or list of strings that end generation.
Other parameters, such as top_k, seed, frequency_penalty and presence_penalty, depend on the model. Each model lists the ones it honors in supported_sampling_parameters on GET /v1/models.
Structured outputs
Models that list structured_outputs in their supported_features accept a JSON schema in response_format, and return content that matches it.
This returns content like {"city": "Toronto", "country": "Canada"}.
Reasoning models
Models that list reasoning in their supported_features think before they answer. The answer is in message.content; the thinking trace, when the model returns one, arrives in a separate field on the message. Reasoning tokens count toward completion_tokens and are not billed twice.
Streaming
Set stream to true to receive the response as server-sent events. Each chat.completion.chunk carries the next piece of the message in choices[0].delta.
Tools
Models that list tools in their supported_features accept OpenAI-format tools and tool_choice. When the model calls a tool, finish_reason is "tool_calls". Run the function, append a tool message with the matching tool_call_id, and send the conversation again.