High-performance in-browser LLM inference engine with WebGPU acceleration and OpenAI API compatibility