Unified library of SOTA model optimization techniques for compressing and accelerating deep learning models
A flexible framework for cutting-edge LLM inference and fine-tuning optimizations with CPU-GPU heterogeneous computing
Tencent's toolkit for compressing LLMs and VLMs with quantization and speculative decoding