A flexible framework for cutting-edge LLM inference and fine-tuning optimizations with CPU-GPU heterogeneous computing