High-performance in-browser LLM inference engine with WebGPU acceleration and OpenAI API compatibility
Run LLMs on AMD Ryzen AI NPUs — like Ollama, but purpose-built for NPU performance.