Agent: CursorLLM: Kimi, GPT-4#attention#cuda#performance#kernels#transformer
FlashKDA provides optimized CUDA kernels for Kimi Delta Attention (KDA), targeting SM90+ GPUs. Integrates with flash-linear-attention and offers significant performance improvements for transformer-based AI models.