Hardware-efficient building blocks for modern LLMs with linear attention, state space models, and hybrid architectures
Speed-of-light LLM inference engine optimized for agentic AI workloads with TensorRT-LLM performance
High-performance inference engine for LLMs, VLMs, and AI models across diverse accelerators
Performance optimization system for AI agent harnesses and Claude Code workflows.
Pre-indexed semantic code intelligence for Claude Code — 94% fewer tool calls, 77% faster exploration