💩
A lightweight inference engine that runs massive AI models (up to 2.8T parameters) on consumer hardware by treating storage, RAM, and VRAM as a unified hierarchy. Supports GLM, Kimi K3, DeepSeek, Qwen and more with chat, server, and web interfaces.