Run locally
Reka Edge weights are published on Hugging Face. See the Reka Edge repository for instructions on running it locally.
macOS
- OS: macOS 13+
- Hardware: Apple Silicon Mac with 32 GB+ unified memory (M1 Pro/Max or later recommended)
- Python: 3.12+
- uv (recommended), which handles dependencies automatically
Linux with an OpenAI-compatible server
For high-throughput serving, the vllm-reka plugin extends vLLM to support Reka’s custom architectures and tokenizer. Follow the vllm-reka installation instructions to install the plugin along with vLLM.
- OS: Linux with CUDA. macOS is not supported for serving.
- Hardware: NVIDIA GPU, ideally with ≥24 GB VRAM. This has been tested to work on GTX 3090 GPUs with 40-50 tokens/s.
- Python: 3.10 ≥ x > 3.14
- vLLM: 0.15.x (0.15.0 ≥ x > 0.16.0)