Run local LLMs on Apple Silicon 3x faster with native multi-token prediction and speculative decoding