> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.reka.ai/chat/run-locally/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.reka.ai/_mcp/server. # Run locally > Run the open Reka Edge weights on your own hardware. Reka Edge weights are published on Hugging Face. See the [Reka Edge repository](https://huggingface.co/RekaAI/reka-edge-2603) for instructions on running it locally. ## macOS * **OS:** macOS 13+ * **Hardware:** Apple Silicon Mac with 32 GB+ unified memory (M1 Pro/Max or later recommended) * **Python:** 3.12+ * [**uv**](https://docs.astral.sh/uv/) (recommended), which handles dependencies automatically ## Linux with an OpenAI-compatible server For high-throughput serving, the [vllm-reka](https://github.com/reka-ai/vllm-reka) plugin extends [vLLM](https://github.com/vllm-project/vllm) to support Reka's custom architectures and tokenizer. Follow the [vllm-reka installation instructions](https://github.com/reka-ai/vllm-reka/blob/main/README.md) to install the plugin along with vLLM. * **OS**: Linux with CUDA. macOS is not supported for serving. * **Hardware**: NVIDIA GPU, ideally with ≥24 GB VRAM. This has been tested to work on GTX 3090 GPUs with 40-50 tokens/s. * **Python**: 3.10 ≥ x > 3.14 * **vLLM**: 0.15.x (0.15.0 ≥ x > 0.16.0) > One OpenAI-compatible API for Reka models and a curated selection of open models.