* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
GSD Core documentation
Documentation is organised into four quadrants: tutorials help you learn by doing, how-to guides solve specific tasks, reference states authoritative facts, and explanation explores concepts and design decisions.
Language versions: English · Português (pt-BR) · 日本語 · 简体中文
Tutorials
- Your first project — install to first shipped phase, one guaranteed path
- Onboarding an existing codebase — bring GSD Core to a brownfield repo
- Build your first capability — author a tiny declarative capability and watch it act in the loop
- Install your first capability — install a third-party capability end-to-end: consent, verify, check for updates, remove
How-to guides
- Install on your runtime — runtime-specific install steps for all 16 supported runtimes
- Install a minimal GSD and add skills later — install only the core skills, then grow the surface with profiles and
/gsd-surface - Attach a plugin-provided skill to a GSD agent — use the
global:plugin:skillentry form to load Claude Code plugin skills into agent prompts - Discuss a phase — capture implementation decisions before planning begins
- Resolve edge-coverage findings — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- Probe edges in a non-English project — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it"
- Resolve prohibition findings — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- Resolve an unreachable-workflow finding — wire or fully sweep a shipped workflow that no command, agent, or skill references
- Change the STATE.md schema — add, change or remove a STATE.md frontmatter key and keep the template and all five reference documents in step
- Resolve verify-command path findings — fix an
<automated>verify command whose target directory does not resolve from the executor's cwd - State a failing direction — say what output constitutes failure for an
<automated>verify command, and migrate a phase planned before the rule - Resolve a contract-drift finding — bring an agent's completion contract, read-tag gate, or deleted-file test reference back into agreement with the registry
- Resolve unreachable-guard findings — fix shell guards whose fallback arm cannot run, and tell "nothing to report" apart from "could not look"
- Diagnose which gsd-tools is running — tell this package's tool apart from the predecessor's colliding binary and from a gsd-core too old to identify itself
- Resolve an ESLint glob-coverage finding — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- Read the statusline freshness marker — turn on
state ~N commits back, and tell "STATE.md is fresh" apart from "freshness could not be established" - Consume the planning snapshot — read
planning inspectfrom a dashboard or harness, and tell "nothing to report" apart from "could not look" - Consume the state contract — read
.planning/state.jsonfrom a workbench or editor extension, gate on the contract version, and tell "nothing to show" apart from "could not look" - Keep planning docs out of a shared repo — make
.planning/local-only, including untracking files git already tracks (the step.gitignorealone cannot do) - Publish PRs without planning artifacts — keep
.planning/committed locally, so worktrees and/gsd-undokeep working, whileplanning.pr_strictkeeps every planning path out of the branch you push - Plan a phase — run research, decompose work, and verify plan quality
- Verify a dependency-compatibility claim — act on a compatibility claim the researcher left
[ASSUMED], and tell "nothing declared" apart from "a constraint is declared" and "the lookup failed" - Execute a phase — run plans in parallel waves with fresh-context subagents
- Enable parallel reviewer lanes — cut a multi-reviewer
/gsd-reviewpass toward its slowest lane, and tell a rate-limited lane apart from one that was never selected - Verify and ship — walk through completed work, diagnose failures, and create the PR
- Catch complexity before it compounds — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it
- Run phases autonomously — use autonomous mode for unattended phase execution
- Handle quick and fast tasks — use
/gsd-quickand/gsd-fastfor ad-hoc work outside the phase loop - Configure model profiles — switch between quality, balanced, and budget model tiers
- Control which host runtime GSD reports — read the
agent_runtimeladder, understand what host detection looks at, and pin the runtime when detection is not what you want - Set up cross-AI review — configure a second AI to review code produced by the primary agent
- Scope code review depth by path — escalate
/gsd-code-reviewtodeepfor sensitive directories while the rest of the repo stays at the default depth - Work in parallel with workstreams — run independent lines of work simultaneously using workstreams
- Isolate work with workspaces — use workspaces to sandbox experimental or risky changes
- Debug a failed execution — diagnose and recover from broken or incomplete phase execution
- Interpret scope-conformance warnings — read the advisory the worktree-wave merge emits when a plan branch commits outside its declared scope
- Interpret install-shadow warnings — read the advisory GSD Core emits when a
/gsd-*trigger is installed at both scopes and one silently wins, and tell "nothing to report" apart from "could not look" - Interpret
state validateresults — read thescopereason codes and tell "nothing to report" apart from "could not look" - Spike and sketch — use
/gsd-spikeand/gsd-sketchfor exploratory work before committing to a plan - Design a UI phase — use the UI phase loop for frontend and visual work
- Enable live-DOM verification — opt a project into browser-backed UI acceptance checks during execution, handle the browser-profile lock, and tell "nothing to report" apart from "could not look"
- Develop a Capability for GSD 1.5+ — add feature Capabilities, hook fragments, and registry entries
- Ship a reviewer lane in your capability — declare a
reviewerbody so/gsd-reviewdiscovers, invokes, and renders your external review CLI or model endpoint - List your reviewer lane in the registry — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- Take over a capability or EoS integration — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing
- Add or update a host's integration — set a host's documentation-sourced
runtime.hostIntegrationaxes (ADR-1239 Phase A), with theundocumentedsentinel rule - Migrate an install test to the executed plan — convert an
fs.existsSync-probing install test group to a value assertion againstinstallRuntimeArtifacts's executed-plan return, and test against a fake fs adapter - Vendor a dependency — add a third-party package
gsd-core/bin/**needs at runtime as a verbatim vendored artifact, keep it out ofdependencies, and pick the right upstream bundle - Turn a capability off (and keep it off) — disable a capability via the surface, or gate individual hooks off without removing the capability
- Drive GSD from a tracker issue — start a phase from a GitHub, Linear, or Jira issue
- Migrate from GSD 2 — upgrade an existing GSD 2 project to GSD Core
- Update GSD — re-run the installer to pick up the latest release
- Clean up get-shit-done-cc — remove leftover old-package artifacts that cause a spurious
⬆ /gsd-updateindicator after migrating to@opengsd/gsd-core - Fix the worktree base-mismatch (exit 42) error — resolve the branch-divergence condition that halts parallel phase execution
- Recover and troubleshoot — fix common problems, rebuild context, and uninstall
Reference
- Commands — every command with flags and examples
- Configuration — full config schema, model profiles, git branching strategies
- CLI tools —
gsd-tools.cjsprogrammatic API for workflows and agents - JSON error mode —
gsd-toolsfailure channels: faults (stderr, exit 1) vs degraded results (stdout, exit 0), and the reason-code taxonomy - Features — complete feature index
- Inventory — installed skills and surface map
- STATE.md schema — field-by-field reference for
.planning/STATE.md - CONTEXT.md schema — field-by-field reference for
.planning/phases/<N>/CONTEXT.md - PLAN.md schema — field-by-field reference for
.planning/phases/<N>/PLAN.md - Planning artifacts — all
.planning/files and their roles - Review and verification capabilities — code review, security, and Nyquist capability ownership and hook contracts
- Gate predicates — canonical specification of the phase-gate predicate vocabulary
- Capability matrix — generated catalogue of every capability's role, tier, extension points, hook kinds, and
engines.gsd - Capability manifest — the full
capability.jsonschema and validation rules gsd capabilitycommand — install / update / remove / list reference for third-party capabilities- Workflow fragments — in-file
<!-- gsd:section -->marker grammar for fragmentizing workflow markdown at emission time - Reviewer Lane Registry — generated catalogue of third-party reviewer lanes, with their flags, transport, and install commands
Explanation
- Context engineering — how context rot forms and how GSD Core prevents it
- The phase loop — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- Multi-agent orchestration — how subagents are spawned, scoped, and coordinated
- Security model — trust boundaries, permissions, and safe automation
- The capability trust model — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- How overlay capabilities compose — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- Architecture — system architecture, agent model, and data flow
- The Embeddable Orchestration System — one public, versioned contract for embedding GSD across many hosts
- Discuss modes — assumptions mode vs interview mode for
/gsd-discuss-phase - Context monitoring — context window monitoring hook architecture
- Issue-driven orchestration — recipe for driving GSD from a tracker issue using existing primitives
Related
- What's new in 1.7.0 — curated highlights of the 1.7.0 release
- Root README — landing page, quickstart, and documentation overview
- Changelog — release history