Local LLM inference engine - run any AI model offline on any device with zero setup
AI-powered penetration testing assistant that runs local LLM analysis on vulnerability scans
Community recipes for serving modern LLMs locally on RTX 3090/4090/5090 GPUs
Turn your Android phone into a local OpenAI-compatible LLM inference server
Cross-platform AI client with local on-device LLM inference and seamless cloud API fallback
Native inference engine optimized for DeepSeek V4 Flash, GLM 5.2, and PRO models with Metal, CUDA, and ROCm support
A coding agent tuned for small local models, optimized for laptops and built on pi framework
Ultra-efficient 1-bit language models with vision, tool calling, and reasoning that run locally on any device
Local OpenAI-compatible API proxy for Qwen Chat with session management
Find the best local LLM for your hardware with real benchmarks, not guesses.
The fastest local AI engine for Apple Silicon β 4.2x faster than Ollama, drop-in OpenAI replacement.
The original local LLM interface β 100% offline, private, and AI-powered.
Hot-swap between local AI models with zero dependencies
Run powerful LLMs locally on any device, no GPU or API required.