Agent: Cursor, Claude CodeLLM: Claude 3.5, GPT-4#inference-engine#apple-silicon#mlx#qwen#speculative-decoding
High-performance inference engine for Apple Silicon that accelerates local LLM execution using native multi-token prediction. Features Qwen model support, Metal optimization, and speculative decoding without external drafters.