← All repos

superpowers-evals

Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.

ai-agentscoding-agentsevalspythonsuperpowers
Browse cluster: Claude AI Agent Frameworks & MCP Tools
1,264commits
3contributors
6languages

Tech stack & purpose

Superpowers-evals is a behavioral evaluation laboratory for testing coding-agent command-line interfaces (Claude, Codex, Gemini, Kimi, and others) built in TypeScript and running on Bun. The project uses Quorum, a Gauntlet-backed framework, to drive agents through QA scenarios and grade them on workflow compliance against defined criteria and deterministic post-checks. The evaluation system is built around a modular architecture spanning scenario orchestration, check verification, transcript normalization, cost estimation, and a web dashboard for viewing results, with support for local and containerized (Linux amd64) execution alongside a shared remote appliance for team-scale evaluations.

Languages

TypeScript
95.1%
Shell
4.0%
JavaScript
0.5%
CSS
0.2%
Python
0.1%
Dockerfile
0.1%

Contributors