💩
A comprehensive toolkit for optimizing AI models through quantization, pruning, distillation, and neural architecture search. Seamlessly integrates with deployment frameworks like TensorRT-LLM and vLLM to accelerate inference for production AI applications.