Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.
Superpowers-evals is a behavioral evaluation laboratory for testing coding-agent command-line interfaces (Claude, Codex, Gemini, Kimi, and others) built in TypeScript and running on Bun. The project uses Quorum, a Gauntlet-backed framework, to drive agents through QA scenarios and grade them on workflow compliance against defined criteria and deterministic post-checks. The evaluation system is built around a modular architecture spanning scenario orchestration, check verification, transcript normalization, cost estimation, and a web dashboard for viewing results, with support for local and containerized (Linux amd64) execution alongside a shared remote appliance for team-scale evaluations.