> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.reka.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server.

# Create Chat Completion

POST https://api.reka.ai/v1/chat/completions
Content-Type: application/json

Generate the next assistant message for a conversation. The request and response follow the OpenAI Chat Completions shape.

Reference: https://docs.reka.ai/chat/api-reference/create-chat-completion

## Authentication

- `Authorization` header (bearer token, required) — Your Reka API key, sent as `Authorization: Bearer <key>`.

## Request

### Body (application/json)

This endpoint expects a ChatCompletionRequest.

- `model` (string, required) — A model ID from `GET /v1/models`, such as `reka-flash-3`. IDs have no namespace prefix.
- `messages` (list of Message, required) — The conversation so far.
- `stream` (boolean, optional, default: false) — Return the response as server-sent events. Usage totals arrive on the final chunk.
- `max_tokens` (integer, optional) — Maximum number of tokens to generate, capped by the model's maximum output length.
- `temperature` (double, optional) — Sampling temperature. Lower values give more predictable output.
- `top_p` (double, optional) — Nucleus sampling probability mass.
- `top_k` (integer, optional) — Only sample from the `top_k` most likely tokens. Honored by models that list it in `supported_sampling_parameters`.
- `stop` (ChatCompletionRequestStop, optional) — A string or list of strings that end generation.
- `seed` (integer, optional) — Random seed. Honored by models that list it in `supported_sampling_parameters`.
- `frequency_penalty` (double, optional) — Penalize tokens by how often they have appeared. Honored by models that list it in `supported_sampling_parameters`.
- `presence_penalty` (double, optional) — Penalize tokens that have already appeared. Honored by models that list it in `supported_sampling_parameters`.
- `tools` (list of Tool, optional) — Functions the model may call. Supported by models that list `tools` in `supported_features`.
- `tool_choice` (enum, optional) — `auto` lets the model decide, `none` disables tool calls, `required` forces at least one call. Defaults to `auto` when `tools` is set.
  - Allowed values: `auto`, `none`, `required`
- `response_format` (ResponseFormat, optional) — Constrain the output to a JSON schema. Supported by models that list `structured_outputs` in `supported_features`.
- `logprobs` (boolean, optional) — Return log probabilities of the output tokens. Supported by models that list `logprobs` in `supported_features`.

## Response

### 200

The completion. With `stream` set to `true`, the body is a stream of server-sent events, each carrying a `chat.completion.chunk` object.

- `id` (string, required)
- `object` (string, required) — `chat.completion`, or `chat.completion.chunk` for streamed events.
- `created` (integer, required) — Unix timestamp in seconds.
- `model` (string, required)
- `choices` (list of Choice, required)
- `usage` (Usage, optional)
- `metadata` (ChatCompletionMetadata, optional)

## Errors

### 400 Bad Request Error

The request body is invalid.

- `error` (ErrorResponseError, required)

### 401 Unauthorized Error

Missing or revoked API key.

- `error` (ErrorResponseError, required)

### 404 Not Found Error

Unknown model ID.

- `error` (ErrorResponseError, required)

### 429 Too Many Requests Error

Rate or capacity limit reached. Retry with backoff.

- `error` (ErrorResponseError, required)

## Types

### Message

- `role` (enum, required)
  - Allowed values: `system`, `user`, `assistant`, `tool`
- `content` (MessageContent, optional, nullable) — The message text, or a list of content parts for multimodal input.
- `tool_calls` (list of ToolCall, optional) — Tool calls made by the assistant. Send them back unchanged in the conversation history.
- `tool_call_id` (string, optional) — On a `tool` message, the ID of the tool call this message answers.

### ChatCompletionRequestStop

A string or list of strings that end generation.

### Tool

- `type` (enum, required)
  - Allowed values: `function`
- `function` (ToolFunction, required)

### ResponseFormat

Constrain the output to a JSON schema. Supported by models that list `structured_outputs` in `supported_features`.

- `type` (enum, required)
  - Allowed values: `text`, `json_schema`
- `json_schema` (ResponseFormatJsonSchema, optional)

### Choice

- `index` (integer, required)
- `message` (Message, optional)
- `delta` (Message, optional) — On streamed chunks, the next piece of the message.
- `finish_reason` (string, optional, nullable) — `stop` when the model finished, `length` when it hit `max_tokens`, `tool_calls` when it called a tool.

### Usage

- `prompt_tokens` (integer, optional)
- `completion_tokens` (integer, optional) — Output tokens, including reasoning tokens.
- `total_tokens` (integer, optional)
- `reasoning_tokens` (integer, optional) — Reasoning tokens, already counted in `completion_tokens`, on models that report them here.

### ChatCompletionMetadata

- `weight_version` (string, optional) — The checkpoint that served the request.

### ErrorResponseError

- `message` (string, required)
- `type` (string, required)
- `code` (string, optional, nullable)
- `param` (string, optional, nullable)

### MessageContent

The message text, or a list of content parts for multimodal input.

### ToolCall

- `id` (string, required)
- `type` (enum, required)
  - Allowed values: `function`
- `function` (ToolCallFunction, required)

### ToolFunction

- `name` (string, required)
- `parameters` (ToolFunctionParameters, required) — A JSON schema for the function's arguments.
- `description` (string, optional)

### ResponseFormatJsonSchema

- `name` (string, required)
- `schema` (ResponseFormatJsonSchemaSchema, required)
- `strict` (boolean, optional)

### ToolCallFunction

- `name` (string, required)
- `arguments` (string, required) — The arguments as a JSON string.

### ToolFunctionParameters

A JSON schema for the function's arguments.

### ResponseFormatJsonSchemaSchema

## Examples

**Request**

```json
{
  "model": "reka-flash-3",
  "messages": [
    {
      "role": "user",
      "content": "What is the fifth prime number?"
    }
  ]
}
```

**Response**

```json
{
  "id": "6ea73acf9bb34fceb7e2d7a513c7f4c4",
  "object": "chat.completion",
  "created": 1790758370,
  "model": "reka-flash-3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The first five prime numbers are:\n\n1. 2  \n2. 3  \n3. 5  \n4. 7  \n5. 11  \n\nSo, the fifth prime number is **11**."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 11,
    "completion_tokens": 43,
    "total_tokens": 54,
    "reasoning_tokens": 0
  },
  "metadata": {
    "weight_version": "default"
  }
}
```

**SDK Code**

```python OpenAI Python SDK
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.reka.ai/v1",
    api_key=os.environ["REKA_API_KEY"],
)

response = client.chat.completions.create(
    model="reka-flash-3",
    messages=[{"role": "user", "content": "What is the fifth prime number?"}],
)
print(response.choices[0].message.content)

```

```typescript OpenAI JavaScript SDK
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.reka.ai/v1",
  apiKey: process.env.REKA_API_KEY,
});

const response = await client.chat.completions.create({
  model: "reka-flash-3",
  messages: [{ role: "user", content: "What is the fifth prime number?" }],
});
console.log(response.choices[0].message.content);

```

```go
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.reka.ai/v1/chat/completions"

	payload := strings.NewReader("{\n  \"model\": \"reka-flash-3\",\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"What is the fifth prime number?\"\n    }\n  ]\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("Authorization", "Bearer <token>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.reka.ai/v1/chat/completions")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["Authorization"] = 'Bearer <token>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"model\": \"reka-flash-3\",\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"What is the fifth prime number?\"\n    }\n  ]\n}"

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.reka.ai/v1/chat/completions")
  .header("Authorization", "Bearer <token>")
  .header("Content-Type", "application/json")
  .body("{\n  \"model\": \"reka-flash-3\",\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"What is the fifth prime number?\"\n    }\n  ]\n}")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.reka.ai/v1/chat/completions', [
  'body' => '{
  "model": "reka-flash-3",
  "messages": [
    {
      "role": "user",
      "content": "What is the fifth prime number?"
    }
  ]
}',
  'headers' => [
    'Authorization' => 'Bearer <token>',
    'Content-Type' => 'application/json',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.reka.ai/v1/chat/completions");
var request = new RestRequest(Method.POST);
request.AddHeader("Authorization", "Bearer <token>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"model\": \"reka-flash-3\",\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"What is the fifth prime number?\"\n    }\n  ]\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = [
  "Authorization": "Bearer <token>",
  "Content-Type": "application/json"
]
let parameters = [
  "model": "reka-flash-3",
  "messages": [
    [
      "role": "user",
      "content": "What is the fifth prime number?"
    ]
  ]
] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.reka.ai/v1/chat/completions")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```