Community recipes for serving modern LLMs locally on RTX 3090/4090/5090 GPUs
Fine-tune LLMs from one YAML config - train 8B models on 4GB laptop GPUs with layer streaming
Run 70B LLMs on 4GB GPUs with zero quantization using memory-optimized inference