Give text-only AI models vision capabilities - just paste an image and get structured JSON analysis
Give Claude the ability to watch and understand videos with frame extraction and multimodal audio analysis
A benchmark for evaluating AI agents in tool-agent-user interactions across real-world domains
Cross-platform AI client with local on-device LLM inference and seamless cloud API fallback
Open-source framework for building full-stack AI-powered apps in JavaScript, Go, and Python
Post-training recipes and SDK for customizing language models with distributed fine-tuning
AI-powered framework for video understanding, editing, and creative remixing
Fine-tune 600+ LLMs and 300+ MLLMs with PEFT, full-parameter training, and advanced RL algorithms.
Open-source multimodal AI agent stack connecting cutting-edge models with GUI automation and MCP tools
Open-source multimodal data stack for logging, visualizing, and querying robotics and AI data
Fast, flexible LLM inference engine in Rust with multimodal support and agentic features
All-in-one multimodal RAG framework powered by advanced AI
The original local LLM interface — 100% offline, private, and AI-powered.
AI memory for your screen — record, search, and automate everything you see and do