π©
DeepSWE provides comprehensive benchmarking for AI coding agents across TypeScript, Go, Python, JavaScript, and Rust. It features 113 original tasks from active open-source repositories with isolated environments and program-based verifiers to measure agent performance on real software engineering challenges.