High-performance GPU kernel library for blazing fast LLM inference and serving
Route, manage, and analyze LLM requests across multiple providers with unified API interface
A coding agent tuned for small local models, optimized for laptops and built on pi framework
Real-time data platform with ML/AI-driven automation for instant action on streaming data
High-performance inference engine for LLMs, VLMs, and AI models across diverse accelerators
A self-organizing Obsidian vault that gives AI coding agents persistent memory across conversations