bitcreed

Engineering Practice

GSD Core in practice: plans, reviews, and cost control

An AI coding agent can produce a convincing implementation quickly. Keeping it aligned with requirements, checking its assumptions and finishing the work takes more discipline. That is where I use GSD Core — Git. Ship. Done.

GSD is a spec-driven development and context-engineering framework for coding agents, including Claude Code and Codex. It turns a large request into bounded phases, gives specialist agents fresh context, and keeps decisions and progress in files. This guide covers the current @opengsd/gsd-core package, version 1.15.0, checked on 4 October 2026. GSD Core

My interest is practical: can the process make AI-assisted software engineering more dependable without consuming the whole day in ceremony? After hundreds of millions of recorded tokens across my projects, the answer is useful, but expensive.

What GSD adds to an AI coding session

Long conversations accumulate stale assumptions and competing instructions. GSD addresses this by sending research, planning and implementation into focused agents, while preserving project knowledge under .planning/.

A milestone describes a release or version. Its phases divide the work into outcomes small enough to plan and verify. PROJECT.md explains the project, REQUIREMENTS.md records what it must deliver, ROADMAP.md orders the phases, and STATE.md records where work stands. Each phase adds decisions, plans, implementation summaries and verification evidence.

Those files give the next session something concrete to resume from. They also let me inspect what the agent thinks it is building before paying for the implementation. GSD’s phase loop makes that inspection part of the workflow.

When one project becomes several

GSD’s project files solve continuity within a repository. When I began coordinating several GSD projects, I hit a different problem: repeatedly switching directories to check each roadmap, phase and handoff. I built GSD Meta Manager to give me one terminal view across them.

The manager reads each project’s .planning/ directory directly, so a status check does not need to start Claude or spend tokens. A filesystem watcher refreshes the view when those files change; Rust, Ratatui and Tokio keep the terminal interface responsive. It has become a practical companion to the phase workflow: I can see where each project stands and decide where my attention is needed, while the project files remain the source of truth.

GSD Meta Manager dashboard showing 14 projects, their active phases, status, progress and backlog counts The dashboard puts project status and next actions in one terminal view.

I used GSD’s own planning workflow to develop the manager. That feedback loop made the value concrete: phase plans and handoffs are useful inside one project, and reading those same artifacts across projects makes parallel work easier to coordinate. The tool is published on crates.io; install it with cargo install gsd-meta-manager. The repository has setup and usage details.

Install GSD and complete one phase

The current package requires Node.js 24 or newer and npm 10 or newer. Run the installer and choose your coding host and installation scope:

npx @opengsd/gsd-core@latest

Restart the host if needed, then use /gsd-new-project for a new project or /gsd-onboard for an existing repository. These examples use the documented hyphen spelling; invocation varies by host, and Claude plugin installations use a /gsd-core:<command> namespace. Use the commands your installation exposes. Installation guide

For phase 1, the core loop is:

/gsd-discuss-phase 1
/gsd-plan-phase 1
/gsd-execute-phase 1
/gsd-verify-work 1
/gsd-ship 1

Discuss resolves implementation choices. Plan researches the problem and writes bounded tasks with acceptance checks. Inspect those plans: this is the cheapest place to catch a wrong direction. Execute runs them, using dependency waves where work can proceed independently, and commits the implementation. Verify checks the delivered behavior and routes failures into diagnosis and follow-up work.

Ship checks repository and verification prerequisites, pushes the branch and creates a pull request using the planning evidence. It supports draft PRs and optional review integration. Deployment still belongs to your repository’s release process. Visual phases can also use /gsd-ui-phase to establish a design contract before implementation. Verify and ship

Use the full loop for work whose decisions and dependencies deserve it. /gsd-quick gives smaller tasks a tracked workflow; /gsd-fast handles trivial edits inline.

Capture ideas without interrupting the work

The current /gsd-capture command brings tasks, notes and future ideas into one entry point:

/gsd-capture Check reconnect behavior after suspend
/gsd-capture --note Firmware reports a different capability on this board
/gsd-capture --backlog Add an export workflow
/gsd-capture --seed Revisit local inference when hardware permits
/gsd-capture --list

The default creates a structured todo. Flags route the thought to a note, roadmap backlog or seed with conditions for revisiting it. That matters when an agent finds something interesting halfway through a phase: capture it, then finish the agreed scope. Capture command

Have Codex review Claude’s plans, or Claude review Codex’s

Independent review is one of GSD’s most useful features. Once the phase plans exist, an installed and authenticated external CLI can review them:

/gsd-review --phase 1 --codex
/gsd-plan-phase 1 --reviews

Use --claude for the reverse direction. GSD collects severity-ranked concerns into the phase’s REVIEWS.md; replanning with --reviews incorporates that feedback. Configure default reviewers through /gsd-config --integrations, or enable /gsd-plan-review-convergence to repeat the review and revision cycle within a configured limit. Cross-AI review guide

This reviews the plan. For implemented code, use /gsd-code-review, or configure workflow.code_review_command for an external reviewer during Ship. A second model offers another perspective, but domain evidence still matters. In my Linux USB tool project, generated tests agreed with incorrect assumptions; real hardware and the specification exposed the bugs.

What it costs in my projects

My refreshed audit covers 14 distinct project repositories, with Claude Code transcripts available for 12, from August and September 2026. It records approximately 839 million input, cache-write and output tokens, plus 33.2 billion cache-read tokens reported separately. Around 691 million of the first total are attributable to GSD agents and commands. Coordinator work is not always identifiable as GSD, so this is only a partial accounting of my usage rather than a lifetime total.

Across the recorded samples, the median phase used about 6.8 million tokens, while a quick task used about 443,000. These are different workloads, not a controlled comparison. Executors accounted for 39% of GSD-attributed tokens; the remainder includes planning, coordination, research, reviews and verification. Much of that work is the reason to use the framework.

My working rule of thumb remains roughly four times the tokens of a direct agent task. A milestone with three phases can occupy a workday, although parallelism, caching and model speed mean token usage does not translate directly into elapsed time. On my Claude Max 20× setup, two intensive implementation days consume roughly half the displayed weekly allowance. That is my workload experience, not a subscription requirement.

I prefer Opus for planning and use Sonnet for execution when the tasks are clear. Anthropic also describes this division as a way to balance quality and cost. Model orchestration guidance

Extensions: cut unnecessary work before building it

GSD supports extensions called capabilities. They can contribute skills, agents, configuration and hooks at defined points in the workflow. I am building an Elon Musk Algorithm-inspired simplification capability that asks whether a requirement or implementation step should exist before GSD expands it into more work. It defers unnecessary scope, removes speculative layers and gives small phases a leaner path. Capability model

As of this audit, five project ledgers report 82 applied cuts. The extension’s built-in pricing model estimates about 28.3 million net tokens avoided so far.

Part of my aim is to bring the roughly 4:1 overhead closer to 3:1. That remains a target but token-cuts are not the only goal. The more useful engineering question is already clear: which requirements, abstractions and process steps are earning their cost?

Keep human attention on the decisions that matter

In my setup, auto mode lets the workflow continue without launching Claude with --dangerously-skip-permissions. GSD automation and the host’s tool permissions are separate settings; the permissions flag is optional in the first-project tutorial.

I keep a coordinating session and dispatch phase work into subagents, which lets me oversee several projects. Four or five can run in parallel, but fewer leave me less drained. I also tell the agents to perform routine checks themselves and reserve my verification time for behavior, domain assumptions and decisions they cannot establish independently. Before hitting a quota limit, /gsd-pause-work preserves the handoff for the next session.

Start with one bounded phase. Read its plan, verify its behavior and inspect the diff before shipping. GSD makes the work traceable; engineering judgment determines whether the right thing was built.