Run local LLMs on Apple Silicon 3x faster with native multi-token prediction and speculative decoding
Native inference engine optimized for DeepSeek V4 Flash, GLM 5.2, and PRO models with Metal, CUDA, and ROCm support
High-performance neural network inference framework optimized for mobile platforms