Skip to navigation

Chat Completions

POST https://api.reka.ai/v1/chat/completions takes the OpenAI Chat Completions request and returns the OpenAI response shape. Authenticate with Authorization: Bearer $REKA_API_KEY, or let the OpenAI SDK set the header from api_key.

Basic request

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.reka.ai/v1",
api_key=os.environ["REKA_API_KEY"],
)
response = client.chat.completions.create(
model="reka-flash-3",
messages=[{"role": "user", "content": "What is the fifth prime number?"}],
)
print(response.choices[0].message.content)

The response follows the OpenAI shape:

{
"id": "6ea73acf9bb34fceb7e2d7a513c7f4c4",
"object": "chat.completion",
"created": 1790758370,
"model": "reka-flash-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The first five prime numbers are:\n\n1. 2 \n2. 3 \n3. 5 \n4. 7 \n5. 11 \n\nSo, the fifth prime number is **11**."
},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 11, "completion_tokens": 43, "total_tokens": 54, "reasoning_tokens": 0},
"metadata": {"weight_version": "default"}
}

metadata.weight_version names the checkpoint that served the request.

System prompts and multi-turn conversations

Send the whole conversation in messages on every request. Roles are system, user, assistant and tool.

response = client.chat.completions.create(
model="reka-flash-3",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "My name is Matt."},
{"role": "assistant", "content": "Nice to meet you, Matt."},
{"role": "user", "content": "What was my name?"},
],
)
print(response.choices[0].message.content)

Sampling parameters

  • max_tokens: the maximum number of tokens to generate, capped by the model’s maximum output length. If finish_reason is "length", the output was cut off.
  • temperature: lower values give more predictable output, higher values more varied output.
  • top_p: nucleus sampling.
  • stop: a string or list of strings that end generation.

Other parameters, such as top_k, seed, frequency_penalty and presence_penalty, depend on the model. Each model lists the ones it honors in supported_sampling_parameters on GET /v1/models.

Structured outputs

Models that list structured_outputs in their supported_features accept a JSON schema in response_format, and return content that matches it.

response = client.chat.completions.create(
model="reka-flash-3",
messages=[{"role": "user", "content": "Name the largest city in Canada."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "city",
"strict": True,
"schema": {
"type": "object",
"properties": {
"city": {"type": "string"},
"country": {"type": "string"},
},
"required": ["city", "country"],
"additionalProperties": False,
},
},
},
)
print(response.choices[0].message.content)

This returns content like {"city": "Toronto", "country": "Canada"}.

Reasoning models

Models that list reasoning in their supported_features think before they answer. The answer is in message.content; the thinking trace, when the model returns one, arrives in a separate field on the message. Reasoning tokens count toward completion_tokens and are not billed twice.

Streaming

Set stream to true to receive the response as server-sent events. Each chat.completion.chunk carries the next piece of the message in choices[0].delta.

stream = client.chat.completions.create(
model="reka-flash-3",
messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)

Tools

Models that list tools in their supported_features accept OpenAI-format tools and tool_choice. When the model calls a tool, finish_reason is "tool_calls". Run the function, append a tool message with the matching tool_call_id, and send the conversation again.

tools = [{
"type": "function",
"function": {
"name": "get_product_availability",
"description": "Check whether a product is in stock.",
"parameters": {
"type": "object",
"properties": {"product_id": {"type": "string"}},
"required": ["product_id"],
},
},
}]
response = client.chat.completions.create(
model="reka-edge-2603",
messages=[{"role": "user", "content": "Is product a-12345 in stock?"}],
tools=tools,
)
print(response.choices[0].message.tool_calls)

Next steps