GPU-optimized library for training transformer models at scale
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale
Fast, easy, and cheap LLM inference and serving engine