Comprehensive benchmarking tool for measuring generative AI model performance across inference solutions
Browser automation skill that lets Claude write and execute custom Playwright scripts on-the-fly
Evaluation and evolution tool for Agent Skills with automated improvement loops
Test, evaluate, and red-team LLM apps with declarative configs and CI/CD integration.
Token-efficient CLI for AI agents to automate browser testing and web interactions.