Local-first AI agent workspace for real work with controlled permissions and recoverable execution logs
A benchmark for evaluating AI agents in tool-agent-user interactions across real-world domains