π©
A comprehensive collection of working configurations and benchmarks for running large language models like Qwen and Gemma on consumer RTX GPUs. Supports multiple engines (vLLM, llama.cpp) and provides optimized setups for single or dual-card configurations.