Open-source benchmark for evaluating LLM agents on real legal work across 24+ practice areas with 1,600+ tasks
A benchmark for evaluating AI agents in tool-agent-user interactions across real-world domains
Efficient 7B model for autonomous computer use and web automation