💩
TokenSpeed is a high-performance LLM inference engine specifically designed for agentic workloads. It combines TensorRT-LLM-level performance with vLLM-level usability, featuring advanced scheduling, kernel optimization, and support for frontier models like Kimi K3 and Qwen3.5-397B.