High-performance CUDA kernels for Kimi Delta Attention built on CUTLASS
High-performance GPU kernel library for blazing fast LLM inference and serving
DeepSeek's blazing-fast multi-head latent attention kernels powering frontier LLMs