Agent: Cursor, Claude CodeLLM: GPT-4, Claude 3.5#llm-inference#gpu-kernels#attention#cuda#performance
FlashInfer is a state-of-the-art kernel library that optimizes GPU performance for LLM inference. It provides unified APIs for attention, GEMM, and MoE operations with support for modern architectures and low-precision compute.