Skip to navigation

Run locally

Reka Edge weights are published on Hugging Face. See the Reka Edge repository for instructions on running it locally.

macOS

  • OS: macOS 13+
  • Hardware: Apple Silicon Mac with 32 GB+ unified memory (M1 Pro/Max or later recommended)
  • Python: 3.12+
  • uv (recommended), which handles dependencies automatically

Linux with an OpenAI-compatible server

For high-throughput serving, the vllm-reka plugin extends vLLM to support Reka’s custom architectures and tokenizer. Follow the vllm-reka installation instructions to install the plugin along with vLLM.

  • OS: Linux with CUDA. macOS is not supported for serving.
  • Hardware: NVIDIA GPU, ideally with ≥24 GB VRAM. This has been tested to work on GTX 3090 GPUs with 40-50 tokens/s.
  • Python: 3.10 ≥ x > 3.14
  • vLLM: 0.15.x (0.15.0 ≥ x > 0.16.0)