Benchmark for measuring frontier coding agents on long-horizon software engineering tasks across 113 real-world challenges
Evaluation and evolution tool for Agent Skills with automated improvement loops
Agent skill harness for turning ideas into evaluated workflows
CLI framework for building, testing, and benchmarking AI agent skills across models