> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.reka.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server.

# Run locally

> Run the open Reka Edge weights on your own hardware.

Reka Edge weights are published on Hugging Face. See the [Reka Edge repository](https://huggingface.co/RekaAI/reka-edge-2603) for instructions on running it locally.

## macOS

* **OS:** macOS 13+
* **Hardware:** Apple Silicon Mac with 32 GB+ unified memory (M1 Pro/Max or later recommended)
* **Python:** 3.12+
* [**uv**](https://docs.astral.sh/uv/) (recommended), which handles dependencies automatically

## Linux with an OpenAI-compatible server

For high-throughput serving, the [vllm-reka](https://github.com/reka-ai/vllm-reka) plugin extends [vLLM](https://github.com/vllm-project/vllm) to support Reka's custom architectures and tokenizer. Follow the [vllm-reka installation instructions](https://github.com/reka-ai/vllm-reka/blob/main/README.md) to install the plugin along with vLLM.

* **OS**: Linux with CUDA. macOS is not supported for serving.
* **Hardware**: NVIDIA GPU, ideally with ≥24 GB VRAM. This has been tested to work on GTX 3090 GPUs with 40-50 tokens/s.
* **Python**: 3.10 ≥ x > 3.14
* **vLLM**: 0.15.x (0.15.0 ≥ x > 0.16.0)