Local LLM inference engine - run any AI model offline on any device with zero setup
High-performance C++ inference engine for local audio AI models - TTS, STT, voice conversion & music generation
Run frontier MoE models (744B-2.8T params) on consumer hardware — pure C, zero deps, experts streamed from disk
High-performance single-GPU inference engine for Qwen models on RTX 5090
Run local LLMs on Apple Silicon 3x faster with native multi-token prediction and speculative decoding
Native inference engine optimized for DeepSeek V4 Flash, GLM 5.2, and PRO models with Metal, CUDA, and ROCm support
High-performance neural network inference framework optimized for mobile platforms