gstack: 23 Role-Based Slash Commands That Turn Claude Code into a Virtual Engineering Team
gstack is an MIT-licensed Claude Code setup published by Garry Tan, President and CEO of Y Combinator. It splits software work into 23 role-based slash commands, such as CEO review, engineering review, design review, QA and release, plus eight power tools. Everything is plain Markdown. The repository has about 135,000 stars. Its real value is turning a blank prompt into a repeatable workflow. The author's claim of 810 times more output is self-reported and not independently reproduced.
Background and Problem Definition
In March 2026, Andrej Karpathy said on a podcast that he had barely typed a line of code since December. The gstack README opens with that quote. It points at a new problem. When a model can write most of the code, the bottleneck moves from "can it write" to "how is the work organized".
Most people meet Claude Code as a blank prompt box. The same model can fill in a few lines for one person and deliver a full feature for another. The gap usually comes from the structure of the request, not from the model. Without structure, output is uneven, and the habits of a real team, such as review, testing and release discipline, never appear.
gstack targets that gap. It was published by Garry Tan, President and CEO of Y Combinator, in the repository garrytan/gstack under the MIT license. It has about 135,000 GitHub stars. The author calls it his open source software factory and says he uses it every day. The README names three audiences: technical founders who still want to ship, first-time Claude Code users who want structured roles instead of a blank prompt, and tech leads who want strict review, QA and release automation on every pull request.
Architectural Core and Technical Principles
The central idea is to split one general model into several roles with clear duties. The README lists a CEO who rethinks the product, an engineering manager who locks the architecture, a designer who catches AI slop, a reviewer who finds production bugs, a QA lead who opens a real browser, a security officer who runs OWASP and STRIDE audits, and a release engineer who ships the pull request. The stated total is 23 specialists plus eight power tools. In form, gstack has no backend service and no new runtime. The README says it is all slash commands, all Markdown, free, and MIT licensed. Each role is therefore a prepared prompt file that Claude Code loads when you call the matching command. This design has three consequences. First, the entry barrier is very low. Installation means placing files where Claude Code can read them, and the README says setup takes about 30 seconds. Second, the system is auditable. Every behavior is written as text, so you can read, edit and fork it line by line. Third, and most important, it has no enforcement power. A role is only a prompt. Whether the model obeys it depends on the model. There is no type system, no compiler and no mechanism that proves the "security officer" actually completed an audit.
The quick start shows the same division of labor. You run `/office-hours` to describe what you are building, `/plan-ceo-review` to challenge the idea, `/review` on a branch with changes, and `/qa` on a staging URL or on an isolated local API, CLI, job or webhook. The order imitates a real team: clarify the need, review the plan, review the code, test, release.
Practical Evaluation and Applications
The skill list visible in this session shows that the commands cover many stages of the software lifecycle: brainstorming (office-hours), strategy, architecture and design review (plan-ceo-review, plan-eng-review, plan-design-review), design systems (design-consultation), debugging (investigate), testing (qa, qa-only), code review (review), visual audit (design-review), shipping (ship, land-and-deploy), documentation updates (document-release), retrospectives (retro), and safety guardrails (careful, freeze, guard). A headless browser, described as about 100 ms per command, supports QA and screenshots.
The QA role deserves the most attention. Many AI coding flows stop when the code is written. gstack asks the agent to open a real browser and verify page state. That replaces "I think it works" with "I saw it work", which is a practical hedge against hallucination. The guardrail commands are also useful. They warn before dangerous operations such as `rm -rf` or a forced push, or they limit edits to one directory.
One fact must be stated plainly: the most prominent numbers in the README come from the author. He reports that his 2026 rate of logical code change is about 810 times his 2013 rate, at 11,417 versus 14 logical lines per day. He also reports that the first months of 2026, through April 18, already equal 240 times the whole of 2013, measured across 40 public and private repositories. He admits that raw line counts inflate with AI, and he links a methodology document that normalizes for this. Even so, nobody has independently reproduced the figures, and the private repositories cannot be checked from outside. More code is not the same as more value, and the gain cannot be credited to gstack alone. We ran no paired controlled comparison, so we cannot claim any specific multiplier of improvement.
Industry Impact and Outlook
The importance of gstack is not an algorithmic breakthrough. It is the practice it represents: encoding team process as a shareable bundle of prompts. This is the same family of ideas as writing style rules into lint configuration or writing deployment into CI scripts, with an agent as the target. About 135,000 stars show strong demand for ready-made agent workflows.
It also exposes risks. Role prompts can become generic, and real projects differ widely in their constraints, so copying them directly may produce ritual reviews with no substance. Without enforced verification, the output of review commands still needs a human to check it. And the commands are built around Claude Code, so moving to another agent would need a rewrite.
Two directions seem likely. One is to connect role prompts to real deterministic checks, for example letting the security role call a scanner and produce a log that others can verify, instead of writing only a text report. The other is measurement: paired, per-sample data that compares defect and rework rates with and without role-based workflows. Until such evidence exists, the careful view is this. gstack is a well designed, very cheap workflow template that is worth trying. Its productivity promise remains the author's self-report, not a verified result.
Sources
FAQ
Who published gstack and under which license?
Garry Tan, President and CEO of Y Combinator, published it as garrytan/gstack under the MIT license. It has about 135,000 GitHub stars.
What is gstack technically, and does it need a backend?
It needs no backend. According to the README, it is a set of slash commands written entirely as Markdown prompt files, with no separate service or new runtime.
Can the 810 times output claim be trusted?
It is self-reported by the author, includes private repositories, and has no independent reproduction or paired controlled comparison. More code is also not the same as more value, so treat it with caution.