Fault-tolerant GPU orchestration for training trillion-parameter LLMs without the headaches
GPU-optimized library for training transformer models at scale
Run frontier AI models locally across multiple devices in a distributed cluster
PyTorch distributed training library for LLMs/VLMs with out-of-the-box Hugging Face support