High-performance runtime for running generative AI models on-device with ONNX
Microsoft's generative AI runtime that powers LLM inference with ONNX models. Supports major architectures like Llama, Mistral, and Phi across multiple platforms with hardware acceleration. Used by VS Code AI Toolkit and Windows ML.
Sign in to leave a comment.
No comments yet.