pstack-claude: Porting a Cursor Skill Stack to Many Coding Agents

Published · AI Daily — AI-assisted deep research, methodology & disclosure

pstack-claude ports Lauren Tan's Cursor skill stack, pstack, to Claude Code, Codex, Pi, OpenCode, Gemini and Prime Agent. You state a goal to poteto-mode, and it picks the right workflow while keeping code concise, simple and verified. The port tracks upstream and carries named policy forks declared in tools/forks.json. A setup-pstack skill sets models and reasoning effort per role. The README says there is no server and no telemetry, scripts run locally, and the license is MIT.

Background and Problem Definition

Coding agents have improved quickly, yet the same model can behave very differently from one task to the next. In many cases the weakness is not the model but the process around it. The agent starts editing before it has reproduced the failure. It declares a fix without checking evidence. Or it turns a small change into a wide refactor. pstack targets this gap in process.

According to the README, Lauren Tan's pstack is an opinionated Cursor skill stack that improves agent outcomes. pstack-claude is a port for Claude Code, Codex, Pi and other agent harnesses. The repository description also names OpenCode, Gemini and Prime Agent. The hard part of such a port is not copying text. It is translating the primitives that Cursor provides, such as subagents, questions and wake-ups, into their counterparts in other environments. The repository summarizes this as "Cursor primitives translated for other harnesses".

The problem is practical. Teams often use several agent tools, and each one has its own skill format and invocation style. If a good workflow is bound to one editor, the lessons cannot be reused. A port lets the same method behave consistently across environments. One note on method: this article is based on reading the repository README. We did not install or run the project, so every judgment below about effect is analysis of the design, not a measured result.

Architectural Core and Technical Principles

The README lets us reconstruct the project in four layers. The first layer is the entry skill, poteto-mode. You state a goal, for example "fix the search filter resetting when I change pages", and it chooses the right workflow. The README says it keeps your code concise, simple and verified. The SKILL.md file lists playbooks for planning, features, refactoring, performance issues, investigations, prototypes, PR maintenance, shipping and longer projects. The second layer is a set of composable skills. The README example mentions `how`, `why` and `architect`. For a bug, the flow is as follows. The agent reproduces the failure. It investigates with `how` and `why`. It delegates the fix. Then it reruns the failing case. If the fix crosses a function boundary, poteto-mode brings in `architect` before implementation. You receive the fix, plus the failing and passing evidence. This evidence-first delivery is the most valuable part of the design.

The third layer is the runtime adaptation. On Claude Code, you install through the plugin marketplace with `/plugin marketplace add michael-denyer/pstack-claude` and `/plugin install pstack@pstack-claude`. Codex uses `codex plugin marketplace add` and `codex plugin add`. Pi uses `pi install git:github.com/michael-denyer/pstack-claude`. On Pi, the package also loads a pstack extension. That extension adds subagent, question and wake-up tools, a `/loop` command and the routing instruction. This shows that the port does more than move skill text. It supplies implementations where a harness lacks a primitive. The fourth layer is policy and configuration. The port tracks upstream and carries named policy forks, each declared in `tools/forks.json`. The `setup-pstack` skill changes model defaults and sets a reasoning effort per role. The README example is `arena runners: opus @xhigh, fable @max`. On Claude Code, such settings are dispatched through the plugin's `pstack:effort-` or `pstack:poteto-agent-` agents. Roles without a level keep the session's effort, unless the `default effort` line in the sheet names one. Automatic routing can be turned off.

Practical Evaluation and Applications

Three features deserve attention. The first is automatic routing. On Claude Code and Codex, the plugin installs a routing hook so that an incoming task goes through poteto-mode first. Codex asks you to trust the hook through `/hooks` before it runs. This is a sound safety choice, because a hook changes agent behavior and should not be enabled silently. On Pi, the extension injects the same routing instruction. The second is per-role model and effort settings. You can spend expensive, high-effort reasoning only on the roles that need it and run the rest at a lower level. This gives a fine-grained trade-off between quality and cost. The third is the data boundary. The README states that pstack has no server and no telemetry. Anything its skills ask your agent to read, including session transcripts, goes to your model provider. Scripts run locally, and PR tools use your GitHub CLI login. This is a candid statement. It does not promise that data never leaves the machine. It says where the data goes.

As for usage, the stack suits tasks with clear acceptance criteria, such as a defect that you can express as a failing case. For vague exploratory work, the planning and investigation playbooks fit better. The limits should be stated too. A stricter process costs more calls and more time per task. The result depends on how well the underlying model follows instructions. Whether the primitive translations are fully equivalent in every environment is something you must check in your own target. At the time of writing, the repository shows about 1,194 stars. We found no reproducible benchmark, so we draw no conclusion about how much the results improve.

Industry Impact and Outlook

pstack-claude reflects a trend that is taking shape: skills and workflows are becoming portable assets, separate from any one agent product. Prompts and procedures used to live in one editor's settings. Now plugin marketplaces, skill directories and hook systems have similar forms across several agent tools, so "write once, install in many places" is becoming practical. It also shows a way to maintain an upstream and a downstream side by side. The port tracks upstream and writes its own policy differences into an explicit list in `tools/forks.json`. That is more transparent than quiet edits, and it tells users which rules they have installed. The author also publishes a separate plugin, agent-formal-verify, which adds TLA+ model checking and Lean proofs for concurrency bugs and invariants that tests cannot reach. The split is clear. pstack covers rigor in everyday workflows. Formal verification covers the corners that tests miss.

Two open questions remain. First, will the differences between harness primitives grow as each tool evolves, and will the cost of porting rise with them? Second, can the quality of a workflow be measured objectively? Only public, reproducible evaluations can move a skill stack from "looks reasonable" to "shown to work". A team that wants to try it should start in a low-risk repository, take one existing failing case, and check whether the delivered evidence really reduces rework.

Sources

FAQ

How does pstack-claude relate to Lauren Tan's pstack?

pstack is Lauren Tan's skill stack for Cursor. pstack-claude is a port of it for Claude Code, Codex, Pi and other agent harnesses. It tracks upstream and declares its own named policy forks in tools/forks.json.

How do I install it in Claude Code?

Run /plugin marketplace add michael-denyer/pstack-claude, then /plugin install pstack@pstack-claude. To change models and reasoning effort, use /pstack:setup-pstack.

Where does my data go?

The README says pstack has no server and no telemetry. Anything its skills ask your agent to read, including session transcripts, goes to your model provider. Scripts run locally, and PR tools use your GitHub CLI login.