Experimental ultra-low precision quantization formats for AMD hardware, enabling faster LLM inference with 2-8 bit weights
Run multimodal AI models fully on-device for iOS, Android & HarmonyOS
Lightweight tensor library powering local LLM inference on any hardware