obra/superpowers: A Skills Framework That Makes Coding Agents Follow Engineering Discipline

Published · AI Daily — AI-assisted deep research, methodology & disclosure

obra/superpowers is an MIT-licensed skills framework that imposes a software development method on coding agents. It runs a seven-step workflow: brainstorming, isolated worktrees, written plans, subagent-driven or inline execution, strict red-green-refactor testing, staged code review, and branch completion. The skills trigger automatically and are described as mandatory. The repository had 294,554 GitHub stars on 3 October 2026. This analysis rests on the README and repository structure. It reports no independent benchmark, so the practical gains remain unverified.

Background and Problem Definition

Coding agents can write code, but they often skip the discipline that engineering demands. They start building before the goal is clear. They declare success without running tests. They change many files at once, and the result is hard to review. The gap is rarely raw model capability. It is the absence of a working method that the agent actually follows.

obra/superpowers targets that gap directly. The project describes itself as "a complete software development methodology for your coding agents," built from a set of composable skills and a short set of initial instructions that make sure the agent uses them. According to the README, the skills trigger automatically, so the user does not need special commands.

The project is built by Jesse Vincent and the team at Prime Radiant. It is released under the MIT License. As of 3 October 2026, the GitHub repository shows 294,554 stars. That figure signals strong community interest, but it says nothing about whether the method works. This article rests on the README and the repository's skills directory. It does not report an independent benchmark.

Architectural Core and Technical Principles

The core of Superpowers is a seven-step workflow. Each step is a skill, and each skill activates when the previous stage produces its output. Brainstorming starts before any code is written. It refines a rough idea through questions, explores alternatives, and presents the design in sections for approval, then saves a design document. After approval, using-git-worktrees creates an isolated workspace on a new branch, runs project setup, and confirms a clean test baseline. Writing-plans then breaks the work into tasks of roughly two to five minutes each. Every task names exact file paths, includes complete code, and lists verification steps.

Execution has two modes. Subagent-driven-development dispatches a fresh subagent for each task and reviews after each one. The README calls it the most thorough option. Executing-plans runs every task inline in one session and reviews the whole branch once at the end, which the README calls the cheapest option. Test-driven-development then enforces a red-green-refactor cycle: write a failing test, watch it fail, write the minimal code, watch it pass, and commit. The README states that code written before its test gets deleted. Requesting-code-review runs between tasks and reports issues by severity, and critical issues block progress. Finishing-a-development-branch verifies the tests, offers merge, pull request, keep, or discard, and removes the worktree. The design places its constraints in the right layer. The rules live in skill files, not in a single line of a prompt. The agent is told to check for relevant skills before any task. The README describes these as mandatory workflows, not suggestions. One design choice deserves attention. The two-stage review in subagent-driven-development checks spec compliance first, then code quality. Separating "did it build the right thing" from "is it built well" keeps reviewers from letting style issues crowd out functional ones.

Practical Evaluation and Applications

The skills directory contains fifteen skills, which fall into four groups. Testing centers on test-driven-development. Debugging includes systematic-debugging, a four-phase root-cause process with supporting techniques for root-cause tracing, defense in depth, and condition-based waiting, plus verification-before-completion, which asks whether a problem is actually fixed. Collaboration covers design, planning, parallel dispatch, requesting and receiving code review, and branch completion. The meta group covers writing-skills and using-superpowers. Installation depends on the host. Claude Code users can install from Anthropic's official plugin marketplace or from the Superpowers marketplace. Antigravity users run agy plugin install with the repository URL. The README lists more than a dozen supported hosts, including Cursor, Codex CLI, Gemini CLI, OpenCode, and Hermes Agent. Each host needs its own installation. Evaluation needs care in three places. First, the workflow aims to prevent familiar failures: skipped tests, large unreviewed changes, and claims of completion without verification. Whether it reduces those failures needs a controlled comparison on real projects, with and without the framework. This article contains no such data. Second, the README says agents can sometimes work autonomously for a couple of hours without drifting from the plan. That is the project's own description, and this article cannot verify it. Third, a heavier process costs more tokens. Subagent-driven-development spawns a new subagent per task, so it costs more than executing-plans. Teams should match the mode to the size of the task.

Two boundaries also matter. The project states that it does not generally accept contributions of new skills, and that any change to a skill must work across every supported coding agent. Superpowers is therefore an opinionated methodology, not an open marketplace of skills. Second, the optional visual companion in brainstorming loads a logo from the Prime Radiant website by default, and that request includes the Superpowers version in use. The README says the project does not see project details, prompts, or clicks. Setting SUPERPOWERS_DISABLE_TELEMETRY to a true value turns this off, and the README also honors Claude Code's DISABLE_TELEMETRY and CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC opt-outs. Organizations should know this before rollout.

Industry Impact and Outlook

Superpowers matters less for any one piece of code than for what it packages. Engineering discipline has long lived in senior engineers' habits and review checklists. Here, those habits ship as skills, install alongside the agent, and get checked before each task begins. The effect reaches two audiences. Individual developers get defaults that they would otherwise have to remember to enforce. Teams get a shared starting point, so that members' agents follow the same sequence and reviews align more easily. The project also raises a question worth tracking over time: will methodological constraints stay useful as models improve? A process that is essential today may become overhead on a stronger model. In the other direction, more capable agents may need firmer boundaries to keep autonomous work on plan. Answering this needs longitudinal evidence, not a single demonstration.

The community's size shows real demand. It does not prove quality. The next questions to watch are whether independent teams can reproduce the gains on production repositories, whether the review stages catch real defects, and whether behavior differences between hosts erode the promise that one methodology works everywhere. Within the scope of this article, the conclusion is clear. Superpowers packages a well-structured engineering method with explicit constraints. The design case is sound. The practical benefit still needs independent measurement.

Sources

FAQ

What problem does obra/superpowers address?

It addresses coding agents that skip requirements clarification, testing, and review. It packages a seven-step workflow as composable skills that the agent checks before each task.

Which execution mode is the most thorough?

Subagent-driven-development, which dispatches a fresh subagent per task and reviews after each one. The README calls executing-plans the cheapest option.

Is the benefit independently measured?

No. This analysis relies on the README and the skills directory. The README's claim about autonomous runs of a couple of hours is the project's own description.