Commit Graph

139 Commits

Author SHA1 Message Date
Tom Boucher
90eb9e0f8b Merge pull request #3607 from gsd-build/feat/3592-test-rewrite-text-existence-checks-into-
test: rewrite alias coverage as behavioral contract
2026-05-15 23:11:34 -04:00
Tom Boucher
55e1360b2c Merge pull request #3616 from gsd-build/fix/3599-bug-roadmap-get-phase-no-longer-matches-
fix(3599): preserve project-code prefix when looking up roadmap phases
2026-05-15 23:04:32 -04:00
Tom Boucher
55e50cf392 fix(3600): count project-code-prefixed phase dirs in milestone filter
`init.new-milestone` reported `phase_dir_count: 0` for projects whose
phase directories carry a project_code prefix (`.planning/phases/CK-01-name`)
when the ROADMAP used numeric `### Phase N:` headings. Verified via a
temp-project repro that mirrors the reporter's setup.

Root cause: `getMilestonePhaseFilter` builds an `isDirInMilestone(dirName)`
predicate that tries two paths:

  1) Numeric — requires the dir name to START with a digit. `CK-01-name`
     starts with `C`, so this skips.
  2) Custom-ID — captures the leading kebab token (`CK-01-name` as a
     whole) and compares it to the normalised milestone phase IDs
     (`{"1"}`). No match.

There was no path that stripped the project_code prefix before retrying
the numeric match. Added a third path that strips the same shape
`normalizePhaseName` already recognises (`^[A-Z]{1,6}-(?=\d)`) and retries
the numeric match. This runs AFTER the custom-ID path so a ROADMAP that
uses `### Phase PROJ-42:` continues to win via the custom-ID match for
a `PROJ-42` directory; the new branch only fires when the milestone is
keyed on the bare numeric form.

The fix lands in both:

  - get-shit-done/bin/lib/core.cjs:isDirInMilestone (active CJS runtime)
  - sdk/src/query/state.ts:isDirInMilestone (SDK twin)

`getMilestonePhaseFilter` is shared by multiple callers — init.new-milestone,
phase complete, verify-work, validate-health — so the fix benefits every
caller that walks `.planning/phases/` against a numeric ROADMAP.

Regression test
(tests/bug-3600-milestone-phase-filter-project-code-prefix.test.cjs):

  1. Reporter's case: CK-01-name + CK-02-build dirs against Phase 1 / 2
     headings → phase_dir_count === 2.
  2. Existing contract: 01-first dir against Phase 1 heading still counts.
  3. Custom-ID contract: PROJ-42 dir against `### Phase PROJ-42:` still
     counts via the existing custom-ID match (no regression).
  4. Counter-test: CK-99-backlog and CK-100-future dirs MUST NOT count
     against a milestone with only Phase 1 — the strip-and-retry must
     still respect the milestone's actual phase set.

All assertions go through `init new-milestone --json` (typed payload —
`phase_dir_count`). No raw text matching.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 22:28:49 -04:00
Tom Boucher
2336c27141 fix(3599): preserve project-code prefix when looking up roadmap phases
`roadmap get-phase PROJ-42` returned `{found: false}` because
phaseMarkdownRegexSource() unconditionally strips the project-code prefix
(`^[A-Z]{1,6}-(?=\d)`) before matching, building the regex `0*42` which
matches `### Phase 42:` but never `### Phase PROJ-42:`. The function's
own docstring promised a fallback to escapeRegex(phaseNum) for custom
IDs, but the line-680 regex match consumes the stripped-numeric form
before that branch is reachable.

Fix: add phaseMarkdownRegexSourceExact() that returns the exact-escaped
source for project-code-prefixed inputs (or null for un-prefixed). Update
cmdRoadmapGetPhase to do a two-pass search — try the exact-prefixed form
first, only fall back to the existing padding-tolerant numeric form if
the exact heading is not present.

Two-pass at the call site (rather than alternation inside the regex
source) is required: a roadmap containing both `### Phase 42:` and
`### Phase PROJ-42:` cannot be disambiguated by a single alternation
because regex match-position is leftmost-wins, so the bare numeric
heading at line N would always intercept the match intended for the
prefixed sibling at line M.

The #3537 contract is preserved: `roadmap get-phase CK-01` against a
roadmap that uses `### Phase 1:` prose still resolves correctly via the
numeric fallback, because the exact-prefixed pass returns null and the
existing padded-numeric pass runs unchanged.

Tests added (tests/bug-3599-roadmap-get-phase-project-code-prefix.test.cjs):
1. PROJ-42 query against `### Phase PROJ-42:` heading — found
2. Counter-test: bare `42` query against `### Phase PROJ-42:` — NOT found
3. #3537 contract preserved: CK-01 query → `### Phase 1:` heading
4. Disambiguation: both `### Phase 42:` and `### Phase PROJ-42:` in
   one roadmap; each query resolves to its specific match

All assertions go through runGsdTools + JSON parse — typed payload,
no raw text matching on stdout.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 22:08:28 -04:00
Tom Boucher
f3f088a03c test: add behavioral alias dispatch contract 2026-05-15 19:34:31 -04:00
Tom Boucher
a3ca6ff6d6 feat(3553): Project-Root Resolution Module via generator (Phase 4 of #3524)
Phase 4 of the CJS↔SDK hard-seam migration (parent #3524).
Eliminates the `findProjectRoot` duplication that lived at
bin/lib/core.cjs:74-140 and sdk/src/query/helpers.ts:497-590,
the drift carrier behind historical bugs #1362 and #2561.

- sdk/src/project-root/index.ts — source of truth (120 lines,
  pure-with-sync-fs). Exports findProjectRoot(startDir: string)
  and FIND_PROJECT_ROOT_MAX_DEPTH constant.
- sdk/src/project-root/index.test.ts — 13 vitest pinning fixtures
  covering all four heuristics, the #1362 guard, malformed
  config fallback, empty sub_repos, deep nesting, and depth-limit
  enforcement.
- sdk/scripts/gen-project-root.mjs — generator. Captures
  function body via Function.prototype.toString() from compiled
  sdk/dist/. Emits CJS preamble for destructured node:fs /
  node:path / node:os imports.
- sdk/scripts/check-project-root-fresh.mjs — freshness check.
  Imports the generator function directly (Phase 3's cleaner
  pattern).
- get-shit-done/bin/lib/project-root.generated.cjs — generator-
  emitted CJS mirror.
- tests/project-root-generator.test.cjs — 11 parity assertions
  comparing SDK source and generated CJS for every fixture.

- sdk/src/query/helpers.ts: -127 lines. The 94-line inline
  findProjectRoot plus the FIND_PROJECT_ROOT_MAX_DEPTH constant
  (originally at line 471) replaced by a single re-export:
  `export { findProjectRoot } from '../project-root/index.js';`
  Removed unused `parse as parsePath` import.
- get-shit-done/bin/lib/core.cjs: -83 lines net. The 67-line
  inline findProjectRoot replaced by a single
  `require('./project-root.generated.cjs')`. The detectSubRepos
  helper at lines 40-56 stays (used by loadConfig migration).

- sdk/package.json: gen:project-root + check:project-root-fresh
  scripts.
- package.json: proxy for the freshness check.
- .githooks/pre-commit: drift block.
- .github/workflows/test.yml: drift check step after the
  state-document drift step.
- CONTEXT.md: Project-Root Resolution Module entry.
- docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json:
  +1 module count, +1 row.

- Full suite: 9226/9226 pass (baseline 9215 + 11 new parity
  fixtures).
- SDK vitest: 1804/1804 pass.
- Reader shrink: -127 SDK + -83 CJS = 210 lines of duplication
  deleted across the two Readers. New shared Module is 120 lines.

1. Depth limit canonicalization. CJS findProjectRoot previously
   had no explicit walk-up bound (walked until dir === root or
   homedir). The new Module uses FIND_PROJECT_ROOT_MAX_DEPTH = 10,
   matching the SDK's pre-existing value. Only affects paths
   nested more than 10 levels deep from a .planning/ root — a
   pathological case in practice. None of the existing 22 CJS
   findProjectRoot tests covered this; the new parity test does.
2. platformReadSync → readFileSync. The old CJS findProjectRoot
   used the platformReadSync wrapper from
   shell-command-projection.cjs for reading .planning/config.json,
   which returns null on read failure. The Module uses raw
   readFileSync, which throws — caught by the surrounding
   try/catch that already swallowed errors. Functionally
   equivalent for the existing code path; no test exercises the
   null-return semantic.

Closes #3553.
2026-05-15 10:07:10 -04:00
Tom Boucher
2de2d185fa feat(3536): Configuration Module via shared manifests + generator (Phase 2 of #3524)
Phase 2 of the CJS↔SDK hard-seam migration (parent #3524).
Eliminates the structural drift surface that produced bug class

After this phase, neither bin/lib/ nor sdk/src/ defines
CONFIG_DEFAULTS, VALID_CONFIG_KEYS, DYNAMIC_KEY_PATTERNS, or the
four legacy-key normalizations inline. All come from one canonical
source: the Configuration Module (sdk/src/configuration/index.ts)
+ two JSON manifests (sdk/shared/config-{defaults,schema}.manifest.json).
The CJS mirror is generator-emitted (get-shit-done/bin/lib/configuration.generated.cjs)
with a CI freshness check (sdk/scripts/check-configuration-fresh.mjs).

- sdk/shared/config-defaults.manifest.json — canonical nested defaults,
  union of CJS + SDK keys (includes security_*, post_planning_gaps,
  agent_skills, mode, every git/workflow/hooks sub-section).
- sdk/shared/config-schema.manifest.json — VALID_CONFIG_KEYS array,
  RUNTIME_STATE_KEYS array, DYNAMIC_KEY_PATTERNS array with source
  strings (regex reconstructed at runtime).
- sdk/src/configuration/index.ts — source of truth. Exports
  loadConfig (pure read), normalizeLegacyKeys (pure, idempotent,
  returns Normalization[]), mergeDefaults (deep-merge), migrateOnDisk
  (explicit opt-in disk writeback), plus CONFIG_DEFAULTS,
  VALID_CONFIG_KEYS, RUNTIME_STATE_KEYS, DYNAMIC_KEY_PATTERNS.
- sdk/src/configuration/index.test.ts — 29 vitest pinning tests.
- sdk/scripts/gen-configuration.mjs — generator (Function.prototype.toString()
  inspection of compiled SDK dist, plus brace-balanced text scan for
  internal helpers, matching the Phase 1 pattern).
- sdk/scripts/check-configuration-fresh.mjs — CI freshness gate.
- tests/configuration-generator.test.cjs — 27 parity assertions
  (CJS-generated == SDK source).
- tests/configuration-migrate-config.test.cjs — 3 cases for the new
  gsd-tools migrate-config subcommand.

- bin/lib/core.cjs: CONFIG_DEFAULTS literal now sources values from
  CANONICAL_CONFIG_DEFAULTS (the manifest), with a thin flat
  projection at the load boundary to preserve the existing
  flat-shape return contract for the ~21 CJS test files and 100+
  consumers. All four legacy-key migration blocks (branching_strategy,
  sub_repos, multiRepo, depth — historically lines 351-358, 388-397,
  401-408, 416-423) collapse to a single normalizeLegacyKeys call
  in each code path. The inline platformWriteSync writeback stays
  for now to preserve sync loadConfig semantics; the new async
  migrateOnDisk is reachable via gsd-tools migrate-config.
- bin/lib/config-schema.cjs: 135 → 31 lines. Re-exports from the
  generated Module.
- bin/lib/config.cjs: adds cmdMigrateConfig handler (calls
  migrateOnDisk on the explicit user-driven path).
- bin/gsd-tools.cjs: wires migrate-config into command dispatch.

- sdk/src/config.ts: re-exports CONFIG_DEFAULTS and mergeDefaults
  from the Module. loadConfig now calls normalizeLegacyKeys before
  mergeDefaults (replaces the inline branching_strategy graft).
- sdk/src/query/config-schema.ts: 160 → 36 lines. Re-exports from
  the Module.

- tests/config-schema-sdk-parity.test.cjs: refactored from
  "CJS Set equals SDK Set" (trivially true post-migration) to
  "both sides source from the manifest" — structural plus runtime
  invariant.
- Four other tests that text-grepped source files for valid keys
  (plan-review-convergence, bug-3212, bug-2492, feat-3210) are
  updated to use runtime VALID_CONFIG_KEYS.has() or manifest JSON
  lookups.

- CONTEXT.md: new Configuration Module entry with full Interface
  contract.
- Root package.json: check:configuration-fresh proxy script.
- sdk/package.json: gen:configuration + check:configuration-fresh.
- .githooks/pre-commit: configuration drift block.
- .github/workflows/test.yml: configuration drift step after the
  alias drift check.

- 9201 CJS tests pass (baseline pre-cycle: 9195; +6 net new tests
  across migrate-config + parity refactor)
- 1872 SDK vitest tests pass
- 29 Configuration Module vitest fixtures
- 27 CJS/SDK parity fixtures
- Net diff: +388 / −519 = 131-line reduction across the seven cycles,
  despite adding the new Module, manifests, generator, freshness
  check, and two new test files.

1. SDK CONFIG_DEFAULTS now includes manifest-canonical keys
   (resolve_model_ids: false, context_window: 200000, phase_naming,
   claude_md_path, git.create_tag, workflow.security_*,
   workflow.code_review_*, planning.*, hooks.workflow_guard, ship.*).
   Consumers accessing via [key: string]: unknown index get
   the manifest default instead of undefined.
2. SDK mergeDefaults is now proper recursive deep-merge instead of
   spread-per-section. Overlay { workflow: { research: false } }
   now preserves sibling workflow keys; previously it replaced
   the entire workflow section with only research + the section's
   defaults. Semantically identical for the common case;
   strictly better for partial nested overrides.
3. New gsd-tools migrate-config CLI subcommand for the explicit,
   opt-in on-disk migration path.

Closes #3536.
2026-05-15 00:02:56 -04:00
Tom Boucher
a7f0af2ce9 fix(3537): route every phase-number ROADMAP regex through phaseMarkdownRegexSource (#3538)
* fix(3537): route every phase-number ROADMAP regex through phaseMarkdownRegexSource

v1.42.1 added the padding-tolerant `phaseMarkdownRegexSource()` helper but
wired it into only 1 of 8 call sites that build phase-number regexes against
ROADMAP/STATE prose. The other 7 used raw `escapeRegex(phaseNum)` or partial
`0*${escapeRegex(...)}` (tolerated extra padding, not missing), so when
skills passed the resolved padded form (`02.7`) against un-padded ROADMAP
prose (`### Phase 2.7:`, `- [ ] **Phase 2.7:**`), the verbs silently no-op'd
while reporting success.

This consolidates every phase-number ROADMAP/STATE regex through the
canonical helper:

- Promote `phaseMarkdownRegexSource` from `roadmap.cjs` to `core.cjs` so
  `phase.cjs` and `core.cjs` itself can consume it (no circular dep —
  both already import `core.cjs`).
- Wire the helper into the 7 remaining sites:
  - `core.cjs:getRoadmapPhaseInternal` (replaces hand-rolled `isNumeric`
    branch that only padded integers, not decimals).
  - `roadmap.cjs:cmdRoadmapGetPhase` (searchPhaseInContent escapedPhase).
  - `roadmap.cjs:cmdRoadmapAnalyze` checkbox lookup.
  - `roadmap.cjs:cmdRoadmapAnnotateDependencies` phase header lookup.
  - `phase.cjs:cmdPhaseNextDecimal` ROADMAP prose scan.
  - `phase.cjs:cmdPhaseInsert` target anchor + decimal scan + header.
  - `phase.cjs:cmdPhaseComplete` (3 regexes: checkbox, plan-count,
    REQUIREMENTS extraction).

Adds `tests/bug-3537-padded-id-against-unpadded-roadmap.test.cjs` — a
parity-style regression matching CONTEXT.md DEFECT.GENERATIVE-FIX: for
each user-facing verb, asserts that the padded form (`02.7`) and the
un-padded form (`2.7`) produce identical ROADMAP.md against an identical
fixture. Includes one control case (`update-plan-progress`, already wired
in 1.42.1) to prove the parity assertion is non-vacuous.

Closes #3537

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(3537): add changeset fragment (pr: placeholder, amended post-create)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(3537): pin changeset pr: field to #3538

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:17:38 -04:00
Tom Boucher
75dffc8069 fix: address branching strategy review findings 2026-05-14 19:44:01 -04:00
Tom Boucher
3ab24c6c56 fix(3523): self-healing migration of legacy top-level branching_strategy
Three-part fix for the false "unknown config key(s)" warning fired for
top-level `branching_strategy` in .planning/config.json:

1. On-disk migration (option 3, mirroring multiRepo → planning.sub_repos):
   When loadConfig reads a config.json with top-level `branching_strategy`
   set and `git.branching_strategy` unset, it grafts the value into
   `git.branching_strategy` and deletes the top-level key, then persists.
   If `git.branching_strategy` is already set, the nested value wins
   (matches SDK mergeDefaults precedence, PR #3116).

2. KNOWN_TOP_LEVEL safety net: 'branching_strategy' added to the deprecated-
   keys bucket so the warning never fires even on the first read of a root
   config that feeds a workstream merge (where `parsed` may still carry it).

3. Double-emission guard: a module-level `_warnedUnknownConfigKeys` Set
   deduplicates the unknown-key warning across multiple loadConfig calls
   within a single CLI invocation (init phase-op N called it twice).

Closes #3523

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-14 19:03:39 -04:00
Tom Boucher
fb6633ceda fix(workflow): detect nested git worktree in new-project bootstrap (#3491)
The `has_git` boolean returned by `init new-project` and `init ingest-docs`
was derived from a shallow `pathExists(cwd, '.git')` check, so a subdirectory
of an existing repo reported `has_git: false`. The workflow then ran
`git init`, creating a nested `.git` inside the outer worktree and silently
diverting subsequent `gsd-sdk commit` calls into the nested repo.

Replace the shallow check with `git rev-parse --is-inside-work-tree`
semantics in both CJS (`get-shit-done/bin/lib/init.cjs`) and TS
(`sdk/src/query/init.ts`, `sdk/src/query/init-complex.ts`) handlers via a new
shared `gitWorktreeInfoInternal` helper, and expose `git_worktree_root` +
`in_nested_subdir` so the workflows can refuse `git init` inside an existing
worktree and warn that planning files will track to the outer repo.

Regression test: `tests/bug-3491-nested-git-worktree.test.cjs`.

Fixes #3491

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 09:11:59 -04:00
Tom Boucher
1e091d2bcb refactor(shell-projection): remove deprecated wrappers + finalize ADRs (Phase 4, #3468) (#3484)
* refactor(shell-projection): remove deprecated wrappers + finalize ADRs (Phase 4, #3468)

Final phase of the shell-command-projection expansion. Removes the legacy
core.cjs wrappers (`atomicWriteFileSync`, `safeReadFile`, `normalizeMd`)
now that every call site lives behind the seam, plus three Phase-3
stragglers (`graphify.cjs`, `template.cjs`, dead import in
`profile-pipeline.cjs`).

Documentation:
- ADR-0009: addendum noting Phase 1–4 scope expansion (subprocess +
  file I/O ownership), supersession of "does not execute" constraint,
  and resolution of open Q4.
- ADR-0010: status changed to Superseded by ADR-0009 with explanation.
- CONTEXT.md "Shell Command Projection Module" entry already current
  from Phase 1 — no edit needed.

Tests:
- `tests/atomic-write.test.cjs` deleted — wrapper it tested is gone;
  `atomic-write-coverage.test.cjs` (Phase 3) covers platformWriteSync.
- `tests/core.test.cjs::safeReadFile` + `::normalizeMd` describes
  deleted — wrappers are gone.
- `tests/concurrency-safety.test.cjs` normalizeMd suite (behavioral /
  perf / snapshot) repointed via 2-line shim at the seam's
  `normalizeContent` — full regression coverage preserved.

Test result: 9059/9041/18 — exact pre-Phase-4 baseline. All 18
failures are pre-existing path-with-spaces local-env issues.

Closes #3468

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate remaining raw fs.writeFileSync sites (Phase 4, #3468)

Sweeps the 7 raw fs.writeFileSync call sites that bypassed the seam through Phase 3,
folding them into platformWriteSync. Net -14 lines: deletes the local writeFileAtomicSync
helper in installer-migrations.cjs and collapses surface.cjs's manual tmp+rename into a
single seam call.

Sites migrated:
- drift.cjs (1) — frontmatter write
- learnings.cjs (1) — learning record JSON write
- install-profiles.cjs (1) — profile marker write (collapsed redundant mkdir)
- gsd2-import.cjs (1) — imported file write (collapsed redundant mkdir)
- surface.cjs (1) — surface state write (replaced manual tmp+rename block)
- installer-migrations.cjs (3) — journal init/finalize + rewrite-json action;
  deleted private writeFileAtomicSync helper and its three call sites

Two sites intentionally retained outside the seam:
- planning-workspace.cjs:241 — workspace lock (wx-flag atomic-create; previously excluded by Phase 3)
- installer-migrations.cjs:220 — install migration lock (fd write into wx-opened handle)
- writeInstallState (installer-migrations.cjs) — strict atomic contract for install state;
  the seam's fallback-to-direct-write on rename failure would silently violate the
  invariant that install state must never be left half-written. Inline tmp+rename with
  rethrow keeps the original guarantee.

Tests: 9059 / 9041 / 18 — exactly the pre-Phase-4 baseline; 18 failures are the
pre-existing path-with-spaces local-env issues, identical files as before.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(installer-migrations): use strict atomic write for rollback install-state restore

The rollback path was restoring INSTALL_STATE via platformWriteSync, which falls
back to a direct write on rename failure and would silently violate the
half-written invariant that the install-state contract guarantees elsewhere.

Extracts the strict tmp+rename logic from writeInstallState into a shared
atomicWriteInstallState(configDir, content) helper and routes both
writeInstallState and rollbackAppliedMigrationResult through it. Preserves the
existing null-handling (rmSync when previousInstallStateBytes === null) and
existing failure-collection (failures.push on caught errors).

Byte-faithful restore: previousInstallStateBytes is written as-is (no JSON
parse round-trip), preserving the exact prior file contents on restore.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 20:46:02 -04:00
Tom Boucher
439d9ceacd refactor(shell-projection): migrate all fs call sites to platform* seam (Phase 3, #3467) (#3481)
* refactor(shell-projection): migrate roadmap.cjs writes to platformWriteSync (#3467)

2 atomicWriteFileSync calls → platformWriteSync. The seam owns markdown
normalization, so the explicit utf-8 encoding arg is no longer needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate config.cjs writes to platformWriteSync (#3467)

- 3 atomicWriteFileSync calls → platformWriteSync
- 1 raw fs.writeFileSync (depth→granularity migration) → platformWriteSync
- 2 fs.mkdirSync(planningBase, { recursive: true }) → platformEnsureDir

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate docs.cjs reads to platformReadSync (#3467)

6 try { fs.readFileSync } catch {} patterns → platformReadSync(path) with
explicit null guards. detectProjectType now reads package.json once and
shares it across has_cli_bin/is_monorepo/has_tests checks. JSON.parse is
still wrapped in a try (parsing is a separate failure mode from missing file).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate audit.cjs reads to platformReadSync (#3467)

8 try { fs.readFileSync(safeFilePath, 'utf-8') } catch { continue } patterns
→ const content = platformReadSync(safeFilePath); if (content === null) continue;

The single safeSum case (where catch set status='unreadable' rather than
continue) maps to an if/else that preserves the same semantics.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate planning-workspace.cjs to platform* seam (#3467)

- 2 try { fs.readFileSync } catch {} → platformReadSync (null on missing)
- 2 fs.writeFileSync (workstream pointer writes) → platformWriteSync
- 3 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir

The .lock file write at withPlanningLock is intentionally NOT migrated.
That call uses { flag: 'wx' } for atomic exclusive-create, which is the
correct lock-acquisition primitive. platformWriteSync's atomic-rename
pattern would silently overwrite an existing lock file and break the
locking guarantee.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate milestone.cjs writes to platform* seam (#3467)

- 5 atomicWriteFileSync calls → platformWriteSync (4 dropped normalizeMd
  wrapper; seam handles .md normalization automatically)
- 2 raw fs.writeFileSync (archive ROADMAP.md / REQUIREMENTS.md) → platformWriteSync
- 2 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir
- Dropped normalizeMd import (only used as write pre-call here)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate intel.cjs to platform* seam (#3467)

- 7 fs.readFileSync (existsSync+readFileSync patterns and try/catch) → platformReadSync
- 2 fs.writeFileSync → platformWriteSync
- 1 fs.mkdirSync(intelPath, { recursive: true }) → platformEnsureDir
- Consolidated dual-check (existsSync + readFileSync) into single platformReadSync
  call returning null on missing file

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate workstream.cjs to platform* seam (#3467)

- 5 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir
- 1 fs.writeFileSync (STATE.md initial scaffold) → platformWriteSync

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate init.cjs reads/writes to platform* seam (#3467)

- 11 try/readFileSync and existsSync+readFileSync patterns → platformReadSync
- 1 fs.writeFileSync (skill-manifest.json) → platformWriteSync

Three bare fs.readFileSync calls remain (ROADMAP/STATE reads in code paths
where the file is required to exist) — these are not "Done when" violations
(no try/catch wrapping, no inline existsSync guard).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate commands.cjs reads/writes to platform* seam (#3467)

- 6 try/readFileSync and existsSync+readFileSync patterns → platformReadSync
- 2 fs.writeFileSync → platformWriteSync
- 3 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir
- Removed unused safeReadFile import (zero call sites in this file)

Three bare fs.readFileSync calls remain (sourcePath at line 752, fullPath at
443, roadmapPath in cmdAuditOpen) — preceded by existsSync guards or in code
paths where file presence is required; not "Done when" violations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate profile-output.cjs to platform* seam (#3467)

- 6 safeReadFile (from core.cjs) calls preserved by aliasing platformReadSync
  as safeReadFile in the import — same semantics, zero call-site changes
- 3 try/JSON.parse(readFileSync) patterns → platformReadSync + try/JSON.parse
- 1 existsSync+readFileSync pattern (claude.md update) → platformReadSync
- 5 fs.writeFileSync → platformWriteSync
- 4 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir

Two bare fs.readFileSync calls remain (template reads where file must exist
or fail loudly) — not "Done when" violations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate state.cjs to platform* seam (#3467)

- 4 atomicWriteFileSync calls → platformWriteSync (3 dropped normalizeMd
  wrapper; seam handles .md normalization)
- 4 try/readFileSync and existsSync+readFileSync patterns → platformReadSync
- 1 fs.writeFileSync (WAITING.json) → platformWriteSync
- 1 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir
- Dropped normalizeMd and atomicWriteFileSync imports (only used as write
  pre-calls here)

Bare fs.readFileSync calls remain in code paths where STATE.md is required
to exist (statePath reads in cmd handlers, dry-run prune) — not "Done when"
violations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate core.cjs to platform* seam (#3467)

- 7 try/readFileSync and existsSync+readFileSync patterns → platformReadSync
- 3 fs.writeFileSync (config writes + large-payload temp file) → platformWriteSync
- 1 fs.mkdirSync (GSD_TEMP_DIR) → platformEnsureDir

Three fs calls remain — they are the internal implementations of the
safeReadFile and atomicWriteFileSync wrappers that core.cjs exports for
backward compatibility. The wrappers are scheduled for removal in Phase 4
(#3468) and will not be migrated here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate phase.cjs writes to platform* seam (#3467)

- 6 atomicWriteFileSync calls → platformWriteSync
- 3 fs.writeFileSync(path.join(dirPath, '.gitkeep'), '') → platformWriteSync
- 3 fs.mkdirSync(..., { recursive: true }) → platformEnsureDir

Bare fs.readFileSync calls remain for roadmapPath/planPath reads where the
file is required to exist; these are not "Done when" violations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate verify.cjs to platform* seam (#3467)

- 8 safeReadFile (from core.cjs) calls preserved by aliasing platformReadSync
  as safeReadFile in the import — same semantics, zero call-site changes
- 1 existsSync+readFileSync inline ternary → safeReadFile (returns null)
- 5 fs.writeFileSync (config writes + milestones writes) → platformWriteSync

Bare fs.readFileSync calls remain for code paths where the file is required
to exist (roadmap/state/config full reads); these are not "Done when"
violations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate frontmatter.cjs + update atomic-write test (#3467)

- frontmatter.cjs: 2 atomicWriteFileSync calls → platformWriteSync. The
  legacy normalizeMd wrapper is dropped because the seam handles markdown
  normalization. safeReadFile preserved by aliasing platformReadSync.
- atomic-write-coverage.test.cjs: update the #1972 structural invariant
  to assert on platformWriteSync. platformWriteSync uses the same
  tmp-file + atomic-rename primitive that atomicWriteFileSync did — the
  no-partial-write guarantee is preserved across the migration.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): add entry for shell-projection Phase 3 migration (#3467)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(coderabbit): disable ESLint tool (repo uses custom lint scripts)

CodeRabbit's review surface emits a "skipped: no ESLint configuration"
warning because the repo doesn't ship ESLint config. The repo
intentionally does not use ESLint — it ships its own targeted lint
scripts (scripts/lint-no-source-grep.cjs, npm run lint:tests) that
enforce repo-specific test-quality invariants. Adding ESLint config
purely to satisfy CR would add an external dependency
(CONTRIBUTING.md: "No external dependencies in core") and overlap
with the existing custom lint surface.

Disable the ESLint tool in CR's tools config so the skip warning
stops appearing on every PR.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 19:34:23 -04:00
Tom Boucher
639e4d603a refactor(shell-projection): migrate all subprocess call sites to exec*/probeTty seam (Phase 2, #3466) (#3476)
* refactor(shell-projection): migrate planning-workspace.cjs tty probe to probeTty seam (#3466)

Replaces direct execFileSync('tty') with probeTty() from the shell-projection
seam. Removes try/catch — probeTty() returns null on error/non-tty/win32.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate commands.cjs to execGit seam (#3466)

- Replaces execSync('git diff --cached --name-only') with execGit array call
- Migrates 14 existing execGit(cwd, args) callers from core.cjs's local
  wrapper to the seam's execGit(args, { cwd }) signature
- Drops execGit from the core.cjs destructure to resolve naming collision

Drops try/catch around git diff — execGit returns exitCode without throwing,
so the no-staged-files / not-a-git-repo case is detected by exitCode !== 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate check-latest-version.cjs to execNpm seam (#3466)

Routes the default-spawn path through execNpm — execNpm owns the win32
shell-flag policy. The injection point remains spawnSync-shaped for test
compatibility; an internal adapter translates { exitCode } → { status }.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate init.cjs git calls to execGit seam (#3466)

Replaces 3 execSync calls with execGit array-args:
- detectChildRepos: git status --porcelain
- cmdInitNewWorkspace: git --version (worktree availability probe)
- cmdRemoveWorkspace: git status --porcelain

Drops 3 try/catch blocks — execGit returns exitCode without throwing, so
best-effort handling becomes a clean exitCode === 0 check.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate core.cjs to execGit seam delegation (#3466)

- Removes direct require('child_process') from core.cjs
- Replaces execFileSync('git check-ignore') with seam's execGit
- Local execGit wrapper now a thin adapter delegating to seam — keeps the
  legacy (cwd, args) positional signature and derived timedOut field for
  the verify.cjs and worktree-safety.cjs consumers that are out of Phase 2
  scope (the wrapper proper would only be removed once those consumers
  migrate, tracked separately)

Extends the seam's _spawnResult to expose signal and error fields so callers
can compute timedOut without bypassing the seam. The Phase 1 test suite
asserts on required field presence only, so the extension is
backward-compatible.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): migrate graphify.cjs to execTool seam (#3466)

- execGraphify: spawnSync('graphify', ...) → execTool with env passthrough,
  preserving the ENOENT/TIMEOUT/EXIT_NONZERO typed reason mapping using the
  seam's signal/error fields
- checkGraphifyInstalled: spawnSync('graphify', ['--help']) → execTool
- checkGraphifyVersion strategy 1: graphify --version via execTool
- checkGraphifyVersion strategy 2: python3 importlib.metadata via execTool

Adds env option to execTool — graphify needs PYTHONUNBUFFERED=1 to drain
buffered stdout on long-running operations.

Changes seam internals to access spawnSync/execFileSync via the non-destructured
childProcess module reference. Destructured imports capture references at load
time and are un-mockable by mock.method(childProcess, 'spawnSync', ...) — which
breaks all the graphify subprocess tests. Non-destructured access restores
mockability without changing public behavior.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(shell-projection): execGit defaults to non-interactive git env (#3466)

Bakes GIT_TERMINAL_PROMPT=0 and GCM_INTERACTIVE=never into execGit's default
env. Without these, a credential prompt or terminal-input probe blocks the
git subprocess indefinitely until our 10s timeout kills it — surfacing as
a generic timeout instead of the actual auth-prompt cause.

These were previously set ad-hoc in worktree-safety.cjs's local execGitDefault
wrapper. Moving them to the seam makes them the consistent default for every
git call across the codebase. Callers can override via opts.env.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(shell-projection): remove core.cjs local execGit wrapper; migrate verify + worktree-safety (#3466)

Completes the Phase 2 "remove local execGit wrapper" criterion. All callers
now use the shell-projection seam's execGit(args, opts) signature directly.

- core.cjs: delete the local execGit wrapper and the execGit export. The
  isGitIgnored seam check now calls execGit from the seam. Worktree-safety
  function calls drop their execGit DI passthrough — worktree-safety's
  internal execGitDefault now delegates to the seam and adds timedOut.
- verify.cjs: 6 callers migrate from execGit(cwd, args) to
  execGit(args, { cwd }). Imports execGit from the seam directly. The
  inspectWorktreeHealth DI passes the seam's execGit (worktree-safety now
  matches that shape).
- worktree-safety.cjs: local execGitDefault becomes a thin adapter over
  the seam — no more direct spawnSync. 11 internal callers migrate to the
  new shape. DI contract for tests changes from (cwd, args) → (args, opts).
- graphify.cjs: 2 remaining execGit callers migrate from core.cjs (now
  removed) to the seam directly.
- test mocks updated in 3 worktree-safety test files to match the new
  (args, opts) DI shape — most mocks were shape-agnostic and required no
  changes; only those that destructured cwd/args needed updates.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): add entry for shell-projection Phase 2 migration (#3466)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr3476): address CodeRabbit review findings

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 18:33:32 -04:00
Tom Boucher
a60e05c714 fix(claude): restore namespaced /gsd:<command> references (#3452)
* fix(claude): restore namespaced /gsd:<command> references

* test(claude): align slash-command expectations to /gsd: form

* test(claude): align generated command references to /gsd:

* test(claude): finish /gsd: namespace expectation updates
2026-05-12 21:53:24 -04:00
Tom Boucher
a33cbe72f5 fix(worktree): bound git subprocesses with timeout + surface degraded health (#3281) (#3283)
* test: red — bounded git subprocess + structured worktree warnings (#3281)

Regression tests for #3281: worktree-related git subprocess calls have no
timeout bound, and timeout/error outcomes are not surfaced as structured signals.

Failing assertions:
- planWorktreePrune / listLinkedWorktreePaths / snapshotWorktreeInventory must
  return reason=git_timed_out (not generic git_list_failed) when execGit returns
  timedOut:true — enables callers to distinguish timeout from auth failure
- executeWorktreePrunePlan must include timedOut:true in result when the git
  prune call itself times out

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(worktree): bounded git subprocess + structured warning surfacing (#3281)

Root cause (PRED.k014): execGit / execGitDefault called spawnSync with no
timeout, so `git worktree list --porcelain` against a hung/locked repo
blocked the parent process indefinitely.  Downstream callers in core.cjs
and verify.cjs then swallowed any resulting failure silently via
catch { /* intentionally empty */ } (PRED.k302).

Fix:
- worktree-safety.cjs: execGitDefault now passes timeout:10000 to spawnSync.
  Detects SIGTERM+ETIMEDOUT and returns { timedOut:true } in the result shape.
  readWorktreeList maps timedOut:true -> reason:'git_timed_out' (distinct from
  generic git_list_failed) so callers can emit a structured warning.
  executeWorktreePrunePlan propagates timedOut:true as a first-class result field.
- core.cjs: execGit receives the same timeout+timedOut treatment (PRED.k014
  uniform-fix discipline).  pruneOrphanedWorktrees now emits a [gsd-tools]
  WARNING to stderr when the git prune call times out instead of silent-catch.
- verify.cjs: Check 11 branches on worktreeHealth.ok to surface W018 warning
  when the worktree list times out, instead of silent-catch on ok:false.

Backward-compatible: exitCode/stdout/stderr continue to work for all existing
callers; timedOut and error are additive new fields.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* changeset: pr=3283 for #3281

* fix(verify): rename W020 for worktree-timeout warning to avoid W018 collision

W018 is already used for milestone archive drift (Check 12). The new
worktree-health-degraded timeout warning was assigned W018, causing
warning-code ambiguity in triage. Rename to W020 (next available code).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-09 01:53:50 -04:00
Tom Boucher
8bc255c266 fix(workstream): normalize migration workstream names (#3269)
* fix(workstream): normalize migrate-name to valid slug

* docs(context): record workstream migrate-name slug invariant

* fix(catalog-cjs): balanced fallback for unknown profile (CR finding A)

profiles[profile] could return undefined for any profile key absent from
the catalog entry, causing downstream callers like formatAgentToModelMapAsTable
to crash on .length. Add ?? profiles.balanced fallback to match the SDK adapter.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(sdk): anchor path resolution on import.meta.url not cwd (CR finding B)

resolve(process.cwd(), '..') breaks when Vitest is invoked from the repo root
because cwd is already the repo root and '..' goes one level above. Replace
with a file-relative path using fileURLToPath(new URL('../../../', import.meta.url))
anchored at the test file's location (sdk/src/query/).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test: derive Group B runtime list from catalog (CR finding C)

Hardcoded ['kilo', 'cline', ...] throws TypeError if a runtime name is
removed from the catalog. Derive group B dynamically via
Object.keys(catalog.runtimeTierDefaults).filter(r => !r.opus) so the
test never goes stale and auto-covers future Group B additions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(workflow): add hermes to Step B runtime options (CR finding D)

hermes appears in the Group A built-in defaults table but was missing from
the AskUserQuestion options in Step B, forcing users to manually type it via
'Other (Group B or custom)'. Add explicit hermes entry for UI consistency.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(config): refresh dynamic_routing tier table; fix stale L671 (findings E+F)

Finding E: tier table was missing 6 heavy-tier agents and 15 standard/light
agents added by this PR. Updated all three rows to match catalog routingTier
assignments (33 agents total).

Finding F: removed stale '18 of 31' claim and agent enumeration; replaced
with accurate note that all 33 agents have explicit catalog entries. Updated
authoritative source pointers to model-catalog.cjs / model-catalog.ts.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(core): add profile-fallback unit tests for quality and budget (CR nitpick G)

The PR introduced quality→opus and budget→haiku unknown-agent fallbacks but
only balanced→sonnet and inherit→inherit were tested. Add two tests covering
the remaining two branches to complete coverage.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* adr: define planning workspace and worktree seam

* refactor(worktree): extract worktree safety policy module

* refactor(workstream): extract active workstream pointer store seam

* test(worktree): cover policy branch paths and persist seam guardrails

* refactor(worktree): centralize health inventory seam for W017

* fix(workspace): align SDK project path policy with CJS planningDir

* refactor(query): unify SDK planning path projection seam

* refactor(init): route workspace projection through planningPaths seam

* docs(adr): add SDK architecture and planning path ADRs

* refactor(worktree): deepen name, pointer, inventory, and config seams

* docs(config): harmonize claude-opus-4-6 to 4-7 in resolve_model_ids example (CR finding 2)

* fix(sdk): return undefined for model_profile='inherit' sentinel (CR finding 3)

* docs(adr): renumber conflicting 0003-sdk-package-seam-module to 0007, update seam-map reference (CR finding 4)

* fix(workstream): align CJS and SDK name validation to accept dots, guard path traversal via includes('..') (CR finding 5)

* fix(sdk): guard writeActiveWorkstream against non-existent workstream directory, k014/k031 parity (CR finding 6)

* chore(changeset): add #3269 changeset (CR finding 1 — proper changeset for this PR)

* docs(inventory): register 3 new CLI modules in INVENTORY.md/MANIFEST (active-workstream-store, workstream-name-policy, worktree-safety)

* fix(sdk): use relPlanningPath(workstream) in planningPaths, fix setActiveWorkstream/getActiveWorkstream name errors in workstream.ts

* fix(sdk): validate GSD_WORKSTREAM in planningPaths before use (#3269 regression)

planningPaths() called resolveWorkspaceContext() which returned GSD_WORKSTREAM
raw (no validation). An invalid value like '../evil' was used as effectiveWorkstream,
constructing a bad path; roadmapAnalyze() caught the ENOENT and returned a
no-phase_count error object instead of the root ROADMAP result.

Fix: validate envCtx.workstream with validateWorkstreamName() in planningPaths()
before accepting it as effectiveWorkstream. Invalid env → null → root .planning/
fallback, preserving the bug-2791 contract: invalid GSD_WORKSTREAM is silently
ignored and falls back to the root context (phase_count: 0 for empty root ROADMAP).

The bug-2791 regression test now passes. No other call sites read GSD_WORKSTREAM
without validation: query-runtime-context.ts already validates; cli.ts already
validates; context-engine.ts takes a caller-validated workstream parameter.

Closes #3268 (regression introduced by #3269 workstream-name-policy work).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-09 00:15:04 -04:00
Tom Boucher
96806003c5 fix(#3229): shared model catalog source of truth for agent profiles + runtime tier defaults (#3230)
* docs(adr): add ADR-0003 model catalog module

* fix(#3229): add shared model catalog as source of truth for agent profiles and runtime tier defaults

Research / design (ADR-0003):
- Existing drift came from 4 independent model truths:
  1. CJS model-profiles.cjs
  2. SDK config-query.ts stale copy (18 agents)
  3. settings-advanced.md runtime tier table
  4. session-runner Claude-only profile map
- New design: one machine-readable Model Catalog Module in sdk/shared/
  that both packages ship and consume.

Implementation:
- sdk/shared/model-catalog.json — canonical source of truth for:
  - full 33-agent registry
  - per-agent golden (quality) alias + balanced/budget aliases
  - adaptive derivation from routingTier
  - agent→phaseType map
  - agent→dynamic-routing default tier map
  - runtime tier defaults for all supported runtimes
- get-shit-done/bin/lib/model-catalog.cjs — CJS adapter over the catalog
- sdk/src/model-catalog.ts — SDK adapter over the same catalog
- CJS model-profiles.cjs now re-exports derived data from model-catalog.cjs
- SDK config-query.ts now re-exports MODEL_PROFILES/VALID_PROFILES from
  model-catalog.ts instead of maintaining its own list
- sdk/src/query/helpers.ts runtime list now comes from the catalog (fixes hermes drift)
- sdk/src/session-runner.ts Claude profile→model-id mapping now resolves via catalog
- docs/CONFIGURATION.md + settings-advanced.md runtime tables updated to match catalog

Behavior changes:
- resolve-model now covers every shipped agent file on disk (33 agents)
- unknown-agent fallback is profile-semantic, not hardcoded sonnet:
  quality→opus, budget→haiku, balanced/adaptive→sonnet, inherit→inherit
- Group B runtimes remain known runtimes but do not get built-in tier defaults

Tests (RED→GREEN):
- root tests: shipped agent files must equal MODEL_PROFILES keys
- sdk tests: shipped agent files must equal MODEL_PROFILES keys
- direct fix assertion: gsd-code-reviewer resolves to opus under quality with no unknown_agent
- runtime defaults parity test: settings-advanced.md + CONFIGURATION.md tables must match catalog
- helper tests: hermes included in SUPPORTED_RUNTIMES and getRuntimeConfigDir()

Closes #3229

* chore(changeset): update #3229 changeset pr field to 3230

* fix(ci): update inherit fallback expectations and inventory parity for model catalog
2026-05-08 21:25:37 -04:00
Tom Boucher
aa64638176 Merge pull request #3112 from gsd-build/fix/3101-plan-summary-matcher-in-core-cjs-reports
fix: canonicalize plan-summary matching for suffixless summaries
2026-05-04 23:35:34 -04:00
Tom Boucher
dbbc7f0942 Merge pull request #3117 from gsd-build/fix/3056-pruneorphanedworktrees-destroys-linked-w
fix: make orphaned worktree prune non-destructive by default
2026-05-04 23:31:13 -04:00
Tom Boucher
ecd5d11b32 fix(worktree): disable destructive orphaned-worktree removal by default 2026-05-04 23:08:13 -04:00
Tom Boucher
c7886415c3 fix(phase): canonicalize plan-summary matching for suffixless summaries 2026-05-04 22:51:15 -04:00
Tom Boucher
19e580137d fix: scope milestone complete stats to explicit version 2026-05-04 22:06:22 -04:00
Tom Boucher
e1d661ece0 feat(#3024): dynamic routing with failure-tier escalation (#3031)
* feat(#3024): dynamic routing with failure-tier escalation

Adds a `dynamic_routing` block to .planning/config.json that lets
the resolver start agents on a cheap tier and escalate one tier up
when the orchestrator detects a soft failure (verification
inconclusive, plan-check FLAG, etc.). Solves the "pay Opus rates as
insurance" anti-pattern by making escalation observed-quality-driven.

Architecture:
- AGENT_DEFAULT_TIERS map (light/standard/heavy) — every agent in
  MODEL_PROFILES declares a default tier; tests assert coverage
  so adding a new agent without updating the map fails CI.
- nextTier(currentTier) helper — light → standard → heavy → heavy
  (heavy stays at heavy; can't go further).
- resolveModelForTier(cwd, agentType, attempt) — new resolver. The
  orchestrator tracks the attempt counter and passes 0 for the
  first spawn, 1+ on escalation. The resolver caps internally at
  max_escalations so the orchestrator can blindly bump the counter.
- Schema validation: dynamic_routing.enabled / escalate_on_failure /
  max_escalations / tier_models.<light|standard|heavy>. Unknown
  tiers and unknown sub-keys rejected at config-set time.
- SDK schema mirror updated to keep CJS/SDK in lockstep (#2653).

Resolution precedence (highest → lowest):
  1. model_overrides[<agent>]              (full IDs accepted)
  2. dynamic_routing.tier_models[<tier>]   (NEW; escalation-aware)
  3. models[<phase_type>]                  (#3023 phase-type map)
  4. model_profile                         (per-agent column)
  5. Runtime default

Backward compatibility: dynamic_routing is disabled by default
(enabled: false or block omitted). resolveModelForTier short-
circuits to resolveModelInternal in that case, so callers can
adopt unconditionally without breaking existing behavior.

This PR delivers the JS-layer infrastructure: schema + tier map +
resolver. Orchestrator adoption (workflow markdown updates that
detect soft failures and call resolveModelForTier with attempt+1)
is incremental follow-up — verifier / plan-checker / integration-
checker each adopt the protocol when ready.

Tests (23 cases, all structural-IR — no stdout grep):
- Schema invariants: AGENT_DEFAULT_TIERS coverage, VALID_AGENT_TIERS
  exact match, every assignment uses a valid tier
- nextTier helper: light→standard→heavy→heavy, null on invalid input
- Disabled mode: no block + enabled:false both no-op (back-compat)
- Enabled mode: attempt=0 returns default tier model, attempt=1
  escalates, beyond max_escalations caps, heavy agents stay heavy,
  default max_escalations=1 when omitted
- Precedence: per-agent override beats dynamic_routing,
  dynamic_routing beats phase-type models
- Validation: every settings key accepted, unknown tiers/sub-keys
  rejected, bare `dynamic_routing` rejected as config-set target

Documentation:
- get-shit-done/references/model-profiles.md — full reference section
- docs/CONFIGURATION.md — full settings table + escalation flow
- docs/USER-GUIDE.md — task-oriented "Cheap-by-default" section
- docs/FEATURES.md — config row cross-link

Verification:
- 23/23 pass on regression test
- 6843/6843 full suite (23 net new from 6820)
- lint-no-source-grep clean (376 test files)
- SDK schema mirror keeps CJS/SDK in sync per #2653 parity test

Closes #3024

* fix(#3024): honor escalate_on_failure:false + 3 CR follow-ups

CodeRabbit on PR #3031 (4 findings — 1 Major + 2 Minor + 1 Nitpick):

1. **Major (inline)** — get-shit-done/bin/lib/core.cjs:1668
   resolveModelForTier ignored dynamic_routing.escalate_on_failure.
   When the user set it to false, escalation should be disabled, but
   the resolver only checked attempt/max_escalations. An orchestrator
   that always passes attempt+1 on retry would silently escalate
   despite the user opting out.
   Fix: gate effectiveAttempt on `dr.escalate_on_failure !== false`
   so false short-circuits every attempt back to the default tier.

2. **Minor (inline)** — docs/CONFIGURATION.md:123-126
   The dynamic_routing rows in the Core Settings table had 4 cells
   instead of 5 (missing the Options column), breaking the table
   structure. Added explicit Options values for enabled / escalate_on_failure
   / max_escalations rows.

3. **Minor (outside-diff)** — references/model-profiles.md:179-195
   "Resolution Logic" sketch was pre-#3024 and didn't include
   dynamic_routing in the precedence ladder. Updated to a 6-step
   block with dynamic_routing at step 3 (between override and
   phase-type).

4. **Nitpick** — tests/feat-3024-dynamic-routing.test.cjs:189+
   Tests used `if (lightAgent) { ... }` guards that silent-pass
   when AGENT_DEFAULT_TIERS drifts. Replaced all 5 conditional
   skips with `assert.ok(lightAgent, '...')` preconditions so a
   tier-mapping change surfaces as a test failure.

Plus: 2 new regression tests for the Major fix:
- escalate_on_failure:false caps every attempt at default tier
- escalate_on_failure:true (explicit) still escalates normally

Verification:
- 25/25 pass on regression test (23 prior + 2 escalate_on_failure)
- 6845/6845 full suite (2 net new)
- lint-no-source-grep clean

* docs(#3024): align precedence + add fence language tags (CR follow-up)

CodeRabbit (3 minor):

1. docs/CONFIGURATION.md:691 — "Per-Phase-Type Models → Resolution
   precedence" was a 4-step block written pre-#3024; readers got
   contradictory rules between the per-phase-type section and the
   later dynamic_routing section. Updated to the same 5-step ladder
   with dynamic_routing at step 2, and noted that dynamic_routing
   is disabled by default so this section's behavior is unchanged
   when the kill-switch is off.

2. docs/CONFIGURATION.md:770 — escalation-flow code fence missing
   language tag (MD040). Added `text`.

3. references/model-profiles.md:184 — resolution-ladder code fence
   missing language tag (MD040). Added `text`.

No code changes; docs only. Verification: regression test still 25/25.

* docs(#3024): clarify precedence prose — five layers, not four (CR nitpick)

CodeRabbit nitpick: the "Per-Phase-Type Models → Resolution
precedence" prose said "The four layers compose..." but the ladder
above lists five (including Runtime default). Also "dynamic_routing
escalates per-attempt above all of them" misreads as suggesting
dynamic_routing wins over model_overrides — actually overrides still
win at step 1.

Reworded top-down so the precedence direction is unambiguous:
  - model_profile = base
  - models = phase-level override
  - dynamic_routing = per-attempt escalation
  - model_overrides = per-agent exception (top)
  - runtime default = fallback

No code changes; docs only.

* docs(#3024): note escalate_on_failure:false in escalation-flow diagram (CR)

CodeRabbit nitpick: the escalation-flow diagram in
docs/CONFIGURATION.md described the soft-failure → respawn →
tier_models[next_tier_up] path, but didn't surface the
`dynamic_routing.escalate_on_failure: false` kill-switch right next
to it. Users reading the flow diagram (which is the canonical place
to understand attempt behavior) wouldn't see that the kill-switch
overrides the soft-failure branch.

Added a one-paragraph note immediately after the flow listing,
before the tier-sequence example, so the kill-switch is visible
exactly where users decide whether escalation will happen.

No code changes; docs only.
2026-05-02 14:26:35 -04:00
Tom Boucher
d812c66020 feat(#3023): per-phase-type model map in .planning/config.json (#3030)
* feat(#3023): per-phase-type model map in .planning/config.json

Adds a new `models` block to .planning/config.json with six phase-type
slots (planning / discuss / research / execution / verification /
completion). Lets users express coarse tuning ("Opus for planning,
Sonnet for the rest") without learning the agent taxonomy.

Resolution precedence (highest → lowest):
  1. Per-agent `model_overrides[agent]`     (full IDs; targeted exception)
  2. Phase-type `models[phase_type]`         (NEW; tier alias)
  3. Profile table (`model_profile`)          (per-agent column)
  4. Runtime default

The three layers compose: `models` defaults a phase, `model_overrides`
carves an exception. Phase-type values are tier aliases (opus/sonnet/
haiku/inherit) so the runtime-resolution chain (#2517) stays correct
end-to-end without further branching.

Implementation:
- model-profiles.cjs: new AGENT_TO_PHASE_TYPE map + VALID_PHASE_TYPES
  set. Each agent in MODEL_PROFILES gets one phase-type assignment;
  tests assert coverage so adding a new agent without updating the
  table fails CI.
- core.cjs (resolveModelInternal): inserts phase-type tier lookup
  between per-agent override and profile-derived tier. Skips runtime
  resolution when the resolved tier is 'inherit' (was previously gated
  only on profile === 'inherit'; phase-type can now produce inherit
  independently).
- core.cjs (loadConfig): pass `parsed.models` through both code paths
  so resolveModelInternal can read it.
- config-schema.cjs + sdk/src/query/config-schema.ts: dynamic-pattern
  validator accepts only the six known phase-types; unknown slots
  rejected at config-set time.

Backward compat: configs without `models` behave exactly as today.

Tests (15 cases, all structural-IR — no stdout grep):
- Schema: AGENT_TO_PHASE_TYPE coverage, VALID_PHASE_TYPES exact match
- Resolver: phase-type alone; per-agent override beats phase-type;
  phase-type beats profile; issue's full example; "inherit"; empty
  block is no-op; no block is no-op
- Validation: each of the 6 slots accepted; unknown slot rejected;
  bare `models` (no slot) rejected

Verification:
- 15/15 pass on new regression test
- 6808/6808 full suite (5 net new), 0 fail
- lint-no-source-grep clean across 375 test files

Closes #3023

* docs(#3023): document `models` per-phase-type config in user-facing docs

Adds `models` block coverage to the three user-facing docs that ship
with each release:

1. docs/CONFIGURATION.md
   - New "Per-Phase-Type Models" section between "Per-Agent Overrides"
     and "Non-Claude Runtimes" with:
       * full example mixing models + model_overrides
       * phase-type → agent mapping table
       * resolution-precedence pseudocode
       * accepted values (tier alias only)
       * "When to use which" decision matrix
       * validation behavior + example error
   - Added `"models": {}` to the Full Schema snippet
   - Added a row for `models.<phase_type>` to the config keys table
     (next to model_profile_overrides for adjacency)

2. docs/FEATURES.md
   - Added a row for models.<phase_type> in the Configurable Settings
     table (right under model_profile)
   - Cross-link to CONFIGURATION.md for the full surface

3. docs/USER-GUIDE.md
   - New task-oriented "Tuning model cost by phase" section above
     "Using Non-Claude Runtimes" — leads with the concrete config
     and shows the override pattern (one-shot phase + targeted exception)
   - Cross-link to CONFIGURATION.md

Verification:
- 29/29 pass on config-schema-docs-parity + docs-update + new feature
  test (parity-check passes, so the config-schema entry I added in the
  feature commit is now matched by a docs row)
- 6808/6808 full suite pass
- lint-no-source-grep clean

Doc style follows the same pattern used by the existing model_profile,
model_overrides, and model_profile_overrides sections — example-led,
table-backed, cross-referenced. Each doc surfaces the feature at the
right depth (reference / settings table / task guide).

* fix(#3023): mirror phase-type tier in resolveReasoningEffortInternal (CR Major)

CodeRabbit caught a real Codex correctness bug + 3 minor docs/test issues:

1. **Major (outside-diff)** — resolveReasoningEffortInternal in core.cjs
   derived its tier exclusively from the profile table, ignoring the
   models.<phase_type> override added in #3023. Failure mode on Codex:

     Config: model_profile=balanced, models.execution=opus, agent=gsd-executor
     resolveModelInternal:           tier=opus    → gpt-5.4
     resolveReasoningEffortInternal: tier=sonnet  → reasoning_effort=medium
                                                     ↑
                                                  WRONG — should be xhigh
                                                  (opus tier on Codex)

   The runtime received a mismatched (model, effort) pair. Mirrored the
   phase-type lookup from resolveModelInternal so both functions derive
   from the same tier source. 'inherit' phase-type returns null effort
   (no runtime entry maps to 'inherit'; let runtime decide).

2. Minor — .changeset/per-phase-type-models.md `pr: TBD` → `pr: 3030`.

3. Minor (outside-diff) — model-profiles.md "Resolution Logic" section
   omitted the new phase-type tier. Updated the 4-step block to a 5-step
   block including `models[phase_type]` between override and profile,
   plus a paragraph noting that `model` and `reasoning_effort` derive
   from the same tier source.

4. Nitpick — added 2 typo-safety tests:
   - models.research = "haiku3" (typo) → falls through to profile
   - models.research = "openai/gpt-5" (full ID) → falls through to profile
   Plus 5 new reasoning_effort tests covering the Major fix:
   - exported correctly
   - phase-type override flips both model AND effort to same tier
   - inherit phase-type returns null effort
   - per-agent override still bypasses phase-type for effort
   - claude runtime ignores models.* (no effort propagation)

Verification:
- 24/24 pass on regression test (15 original + 2 typo-safety + 5 effort + 2 outside-diff related)
- 6815/6815 full suite (7 net new from 6808)
- lint-no-source-grep clean

The reasoning_effort tests are written semantically (phase-type override
must produce the SAME effort as a profile-only opus config) rather than
hard-coding tier-specific effort strings, so changes to the runtime tier
map don't break them.

* fix(#3023): phase-type override beats profile=inherit (CR Major round 2)

CodeRabbit caught another precedence inversion: when
  { model_profile: 'inherit', models: { execution: 'opus' } }
both resolvers short-circuited on `profile === 'inherit'` BEFORE the
phase-type override could be honored. Result: model returned 'inherit'
and reasoning_effort returned null — both contradicting the documented
precedence where models[phase_type] wins over model_profile.

Fix in resolveModelInternal:
- Compute tier from phase-type FIRST. If phase-type is a valid alias,
  it wins. Otherwise, fall back to profile-derived tier OR 'inherit'
  (when profile === 'inherit').
- Gate the runtime-resolution branch on `tier !== 'inherit'` (was
  `profile !== 'inherit'`) so phase-type=opus can flip runtime mapping
  on even when profile=inherit.
- Gate the inherit-return on `tier === 'inherit'` (was
  `profile === 'inherit'`).

Fix in resolveReasoningEffortInternal:
- Remove the `if (profile === 'inherit') return null;` early-return.
- Compute tier from phase-type first, fall back to profile. If
  phase-type is explicitly 'inherit' OR the resolved tier is 'inherit',
  return null (no runtime entry maps to inherit).

Tests added (5 new):
- model: phase-type wins over profile=inherit (with explicit opus, with
  haiku for one phase + planner-without-slot still inheriting)
- model: profile=inherit + no models block → all agents inherit (no
  regression on existing inherit semantics)
- model: profile=inherit + models block but agent has no slot → that
  agent inherits, agents with slots get phase-type tier
- effort: phase-type opus + profile=inherit → produces opus-tier
  effort, NOT null (the original bug)

Verification:
- 27/27 pass on regression test (24 prior + 3 model + 1 effort)
- 6820/6820 full suite (5 net new)
- lint-no-source-grep clean

The effort test reads the expected value by running a profile-only opus
config and comparing — semantic check, not hard-coded effort string. So
runtime tier map changes don't break the test.
2026-05-02 13:19:15 -04:00
Tom Boucher
f55069ecbf test(#2974): migrate 8 test files to typed-IR assertions (#3016)
* test(#2974): migrate 8 test files to typed-IR assertions

Replaces raw stdout/stderr substring matching with structured-field
assertions per CONTRIBUTING.md "Prohibited: Raw Text Matching on Test
Outputs". Adds shared infrastructure for typed error emission so this
pattern is the easy path going forward.

Shared infrastructure:
- core.cjs: ERROR_REASON frozen enum + setJsonErrorMode/getJsonErrorMode
- gsd-tools.cjs: --json-errors CLI flag, parsed before subcommand dispatch
- config.cjs: typed reasons at all 7 error sites
- graphify.cjs: GRAPHIFY_REASON enum + reason/timeout_ms in execGraphify result
- bin/install.js: pure buildSdkFailFastReport() IR builder + renderer
- hooks/gsd-session-state.sh, gsd-phase-boundary.sh: emit Claude Code
  hookSpecificOutput JSON envelope with typed state_present/config_mode/
  planning_modified/file_path fields (no-op when hooks.community is off)

Test migrations (all pass, 171 tests across the 8 files):
- bug-2649-sdk-fail-fast: assert on ir.reason / ir.context / ir.fix_command
- bug-2687-config-read-warning-parity: assert.equal stderr === ''
- bug-2796-arg-parsing-regression: assert on result.json.updated/.phase
- bug-2838-summary-rescue: parse rescue footer, assert mtime invariant
- bug-2943-config-get-context-window: parse JSON, assert ERROR_REASON.CONFIG_KEY_NOT_FOUND
- graphify: assert reason === GRAPHIFY_REASON.ENOENT/TIMEOUT
- hooks-opt-in: parse hookSpecificOutput, assert typed fields
- security-scan: reclassified as source-text-is-the-product (scan label
  output and CI workflow YAML ARE the deployed contract)

Verification: lint-no-source-grep clean (0 violations), full suite
6741/6741 pass.

Closes #2974

* test(#2974): address CR feedback — typed code field, robust idempotency

Two CodeRabbit findings on #3016 addressed:

1. tests/hooks-opt-in.test.cjs:355 (Minor, inline) —
   parsed.reason.includes('Conventional Commits') was still substring
   matching after the typed-IR migration. Fixed at the source: the
   gsd-validate-commit hook now emits a typed `code` field
   ('CONVENTIONAL_COMMITS_VIOLATION', 'COMMIT_SUBJECT_TOO_LONG')
   alongside the human-readable `reason`. Test asserts strictEqual
   on the code; the prose copy is no longer part of the test contract.

2. tests/bug-2838-summary-rescue-gitignored-planning.test.cjs:224-250
   (Outside-diff) — mtimeMs alone can stay unchanged on coarse-grained
   filesystems (HFS+, FAT) when two rewrites land within the same
   timestamp tick, falsely passing the idempotency assertion.
   Replaced with a full snapshot (mtimeMs, ctimeMs, size, ino, sha256
   of contents) compared via assert.deepStrictEqual — the hash
   catches any rewrite the timestamp would miss.

Verification: 30/30 pass on the two affected files; lint-no-source-grep
clean (0 violations across 368 test files).
2026-05-02 09:27:23 -04:00
teknium1
b126c0579a feat(install): add Hermes Agent runtime support (#2841)
Adds Hermes Agent as a supported installation target. Users can run
\`npx get-shit-done-cc --hermes\` to install all 86 GSD commands as
skills under \`~/.hermes/skills/gsd-*/SKILL.md\`, following the same
open skill standard as Claude Code 2.1.88+, Qwen Code, Antigravity,
Trae, Augment, and Codebuddy.

Hermes Agent is an open-source AI agent framework by Nous Research
(NousResearch/hermes-agent, MIT). Its skill loader accepts the Claude
skill format as-is: frontmatter parsed with PyYAML SafeLoader (unknown
keys like \`allowed-tools\` / \`argument-hint\` ignored), body XML tags
(\`<objective>\`, \`<execution_context>\`, \`<process>\`) passed directly
to the model. Compatibility proven end-to-end with all 86 GSD skills
loading cleanly, \`skill_view()\` returning full bodies, and
\`build_skills_system_prompt()\` emitting them into the agent system
prompt — zero Hermes code changes required.

Changes:
- \`bin/install.js\`: --hermes flag, getDirName/getGlobalDir/getConfigDirFromHome
  support, HERMES_HOME env var (native to Hermes — used for profile
  mode / Docker deploys), install/uninstall pipelines, interactive
  picker option 10 (alphabetical: between Gemini and Kilo), .hermes
  path replacements in copyCommandsAsClaudeSkills and
  copyWithPathReplacement, legacy commands/gsd cleanup, CLAUDE.md ->
  HERMES.md and "Claude Code" -> "Hermes Agent" content rewrites in
  skills/agents/hooks, runtime-appropriate finish message.
- \`get-shit-done/bin/lib/core.cjs\`: add hermes to KNOWN_RUNTIMES;
  add RUNTIME_PROFILE_MAP.hermes with OpenRouter-slug defaults
  (Hermes is provider-agnostic; these defaults resolve across
  OpenRouter, native Anthropic, and Copilot via Hermes' aggregator-
  aware resolver, and are overridable per-tier via
  model_profile_overrides.hermes.{opus,sonnet,haiku}).
- \`README.md\`: Hermes Agent in tagline, runtime list, verification
  command, install/uninstall examples, \`--hermes\` flag reference.
- \`tests/hermes-install.test.cjs\`: new, 14 tests covering directory
  mapping, HERMES_HOME env var precedence, install/uninstall
  lifecycle, user-skill preservation, engine cleanup.
- \`tests/hermes-skills-migration.test.cjs\`: new, 11 tests covering
  frontmatter conversion, path replacement (~/.claude/ ->
  \$HERMES_HOME/skills/), CLAUDE.md -> HERMES.md, "Claude Code" ->
  "Hermes Agent", stale skill cleanup, SKILL.md format validation.
- \`tests/multi-runtime-select.test.cjs\`: updated for new option
  numbering (hermes=10, kilo=11, opencode=12, qwen=13, trae=14,
  windsurf=15, all=16).
- \`tests/kilo-install.test.cjs\`: updated assertions for Kilo having
  moved from option 10 to option 11.

Closes #2841

Implementation notes:
- Zero custom code paths: Hermes reuses copyCommandsAsClaudeSkills()
  identical to Qwen Code / Antigravity pattern.
- Path replacement: ~/.claude/, \$HOME/.claude/, ./.claude/ ->
  .hermes equivalents in skill/agent/hook content.
- Config precedence: --config-dir > HERMES_HOME > ~/.hermes (matches
  how Hermes itself resolves its home directory).
- Legacy cleanup: removes commands/gsd/ if present from a prior
  install, preserving dev-preferences.md (same as Qwen).
- No external dependencies added.

Testing: 5841 / 5841 tests pass (0 failures, 0 regressions)
- 14 new tests in hermes-install.test.cjs
- 11 new tests in hermes-skills-migration.test.cjs
- multi-runtime-select.test.cjs renumbered + 1 new test (single choice for hermes)
2026-04-30 17:24:53 -04:00
Tom Boucher
abb2cb63f6 refactor: extract planning-workspace seam from core.cjs (#2901)
* refactor: extract planning workspace seam from core

* docs: document planning-workspace module and inventory updates

* fix: harden planning lock timeout and preserve workstream set contract

---------

Co-authored-by: Tom Boucher <thomas.boucher@sas.com>
2026-04-30 11:38:13 -04:00
Tom Boucher
c3a42d66f9 Revert "feat(install): add Hermes Agent runtime support" (#2849) 2026-04-29 07:44:49 -04:00
teknium1
5a636bc90a feat(install): add Hermes Agent runtime support (#2841)
Adds Hermes Agent as a supported installation target. Users can run
\`npx get-shit-done-cc --hermes\` to install all 86 GSD commands as
skills under \`~/.hermes/skills/gsd-*/SKILL.md\`, following the same
open skill standard as Claude Code 2.1.88+, Qwen Code, Antigravity,
Trae, Augment, and Codebuddy.

Hermes Agent is an open-source AI agent framework by Nous Research
(NousResearch/hermes-agent, MIT). Its skill loader accepts the Claude
skill format as-is: frontmatter parsed with PyYAML SafeLoader (unknown
keys like \`allowed-tools\` / \`argument-hint\` ignored), body XML tags
(\`<objective>\`, \`<execution_context>\`, \`<process>\`) passed directly
to the model. Compatibility proven end-to-end with all 86 GSD skills
loading cleanly, \`skill_view()\` returning full bodies, and
\`build_skills_system_prompt()\` emitting them into the agent system
prompt — zero Hermes code changes required.

Changes:
- \`bin/install.js\`: --hermes flag, getDirName/getGlobalDir/getConfigDirFromHome
  support, HERMES_HOME env var (native to Hermes — used for profile
  mode / Docker deploys), install/uninstall pipelines, interactive
  picker option 10 (alphabetical: between Gemini and Kilo), .hermes
  path replacements in copyCommandsAsClaudeSkills and
  copyWithPathReplacement, legacy commands/gsd cleanup, CLAUDE.md ->
  HERMES.md and "Claude Code" -> "Hermes Agent" content rewrites in
  skills/agents/hooks, runtime-appropriate finish message.
- \`get-shit-done/bin/lib/core.cjs\`: add hermes to KNOWN_RUNTIMES;
  add RUNTIME_PROFILE_MAP.hermes with OpenRouter-slug defaults
  (Hermes is provider-agnostic; these defaults resolve across
  OpenRouter, native Anthropic, and Copilot via Hermes' aggregator-
  aware resolver, and are overridable per-tier via
  model_profile_overrides.hermes.{opus,sonnet,haiku}).
- \`README.md\`: Hermes Agent in tagline, runtime list, verification
  command, install/uninstall examples, \`--hermes\` flag reference.
- \`tests/hermes-install.test.cjs\`: new, 14 tests covering directory
  mapping, HERMES_HOME env var precedence, install/uninstall
  lifecycle, user-skill preservation, engine cleanup.
- \`tests/hermes-skills-migration.test.cjs\`: new, 11 tests covering
  frontmatter conversion, path replacement (~/.claude/ ->
  \$HERMES_HOME/skills/), CLAUDE.md -> HERMES.md, "Claude Code" ->
  "Hermes Agent", stale skill cleanup, SKILL.md format validation.
- \`tests/multi-runtime-select.test.cjs\`: updated for new option
  numbering (hermes=10, kilo=11, opencode=12, qwen=13, trae=14,
  windsurf=15, all=16).
- \`tests/kilo-install.test.cjs\`: updated assertions for Kilo having
  moved from option 10 to option 11.

Closes #2841

Implementation notes:
- Zero custom code paths: Hermes reuses copyCommandsAsClaudeSkills()
  identical to Qwen Code / Antigravity pattern.
- Path replacement: ~/.claude/, \$HOME/.claude/, ./.claude/ ->
  .hermes equivalents in skill/agent/hook content.
- Config precedence: --config-dir > HERMES_HOME > ~/.hermes (matches
  how Hermes itself resolves its home directory).
- Legacy cleanup: removes commands/gsd/ if present from a prior
  install, preserving dev-preferences.md (same as Qwen).
- No external dependencies added.

Testing: 5841 / 5841 tests pass (0 failures, 0 regressions)
- 14 new tests in hermes-install.test.cjs
- 11 new tests in hermes-skills-migration.test.cjs
- multi-runtime-select.test.cjs renumbered + 1 new test (single choice for hermes)
2026-04-29 04:27:46 -07:00
Tom Boucher
eeaf9c556f fix(#2787): track fenced code blocks in extractCurrentMilestone (#2812)
* fix(#2787): track fenced code blocks in extractCurrentMilestone

The milestone-end search used a multiline regex against the raw
restContent string. Lines inside fenced code blocks (``` or ~~~)
that matched the milestone-heading pattern (e.g. `# note v1.0`)
prematurely set sectionEnd, hiding all phases after the block from
roadmap analyze, roadmap get-phase, and every downstream command.

Replace the regex match with a line-by-line scan that tracks fence
state. Lines inside an open fence are skipped regardless of content.
Adds three regression tests covering backtick fences, tilde fences,
and the roadmap get-phase code path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#2787): track fence delimiter instead of toggling bare boolean

Replace the inFence boolean with fenceChar/fenceLen tracking so that
indented fences (up to 3 leading spaces) and mixed-delimiter content
(~~~ inside a backtick fence) are parsed correctly. A closing fence
is only recognised when it uses the same character as the opening
delimiter and has at least the same run length, matching the CommonMark
spec.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#2787): require fence-only closing line — reject info-string lines as closers

A closing fence delimiter must contain only optional trailing whitespace.
A line like \`\`\`js inside an open fence has an info string and must not
close it. The previous regex /^\s{0,3}([`~]{3,})/ matched the opening of any
such line, so the closing check could toggle fenceChar off on an info-string
line and expose subsequent heading-like content to the milestone-end detector.

Fix: capture the trailing portion of every fence-candidate line and only clear
fenceChar when trailing matches /^\s*$/ (per CommonMark §4.5).

Adds a regression test covering the ```text / ```js nesting scenario.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-28 20:37:47 -04:00
Tom Boucher
3246810876 feat: extend RUNTIME_PROFILE_MAP to gemini/qwen/opencode/copilot + settings-advanced UI (#2754)
Closes #2612

- Add gemini, qwen, opencode, and copilot entries to RUNTIME_PROFILE_MAP in core.cjs
- Group B runtimes (kilo, cline, cursor, windsurf, augment, trae, codebuddy, antigravity) intentionally have no built-in map and fall through to the existing unknown-runtime fallback
- Add 40 new tests to tests/issue-2517-runtime-aware-profiles.test.cjs covering each new runtime's three tiers, Group B fall-through, and partial override merge semantics
- Add Section 7 "Runtime Model Tiers" to settings-advanced.md with interactive UI to view and override built-in tier defaults per runtime
- Update docs/CONFIGURATION.md built-in tier table to include all four new runtimes

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-26 13:33:54 -04:00
Tom Boucher
2b95ccbddd Merge pull request #2743 from gsd-build/fix/2714-workstream-config-inheritance
feat(config): workstream config.json inherits from root .planning/config.json
2026-04-26 11:49:30 -04:00
Tom Boucher
cb149383c1 fix(config): bump MODEL_ALIAS_MAP to claude-opus-4-7 (#2733)
* fix(config): bump MODEL_ALIAS_MAP and RUNTIME_PROFILE_MAP to claude-opus-4-7

Opus 4.7 shipped Q1 2026 but MODEL_ALIAS_MAP and RUNTIME_PROFILE_MAP.claude.opus
were still pinned to claude-opus-4-6. Users with resolve_model_ids: true received
stale model IDs in logs and agent-tool calls.

Also adds a resolve_model_ids: true test suite — this path had zero coverage,
which is why the stale ID survived undetected.

Closes #2712

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(config): derive RUNTIME_PROFILE_MAP.claude from MODEL_ALIAS_MAP (coderabbit)

RUNTIME_PROFILE_MAP.claude was duplicating model IDs that MODEL_ALIAS_MAP
already owns. Future model bumps now only require updating MODEL_ALIAS_MAP.
Also fixes stale test assertion (claude-opus-4-6 → claude-opus-4-7).

* fix(tests): update stale claude-opus-4-6 refs to claude-opus-4-7; DRY: derive RUNTIME_PROFILE_MAP.claude from MODEL_ALIAS_MAP

- Update 3 hardcoded `claude-opus-4-6` assertions in tests/issue-2517-runtime-aware-profiles.test.cjs to `claude-opus-4-7`
- Update comment on line 128 that referenced the old model ID
- Replace manual per-tier expansion of RUNTIME_PROFILE_MAP.claude with Object.fromEntries so future alias bumps only require updating MODEL_ALIAS_MAP

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-26 11:48:47 -04:00
Tom Boucher
1cb4bebcf5 feat(config): workstream config.json inherits from root .planning/config.json
- Add _deepMergeConfig() with correct null-override semantics
- loadConfig() reads root config.json when GSD_WORKSTREAM is set, then
  deep-merges with workstream config (workstream wins on conflict)
- Workstream without config.json falls back to root config entirely
- Migrations and disk writes operate on fileData (on-disk content) only,
  never on the merged result, to prevent workstream pollution
- Fixes null-override bug from PR #2717: explicit null in workstream now
  correctly overrides root value instead of falling back to root
- Tests: inherit root model_overrides, workstream override, nested
  workflow.* deep merge, explicit null override, missing workstream config

Closes #2714
2026-04-26 11:15:16 -04:00
Tom Boucher
b7ff14fe51 fix(#2687): derive KNOWN_TOP_LEVEL from DYNAMIC_KEY_PATTERNS to eliminate read-side drift (#2706)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 12:08:59 -04:00
Tom Boucher
b1a670e662 fix(#2697): replace retired /gsd: prefix with /gsd- in all user-facing text (#2699)
All workflow, command, reference, template, and tool-output files that
surfaced /gsd:<cmd> as a user-typed slash command have been updated to
use /gsd-<cmd>, matching the Claude Code skill directory name.

Closes #2697

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 10:59:33 -04:00
Tom Boucher
c811792967 fix(#2660): capture prose after labeled bold in extractOneLinerFromBody (#2679)
* fix(#2660): capture prose after label in extractOneLinerFromBody

The regex `\*\*([^*]+)\*\*` matched the first bold span, so for the new
SUMMARY template `**One-liner:** Real prose here.` it captured the label
`One-liner:` instead of the prose. MILESTONES.md then wrote bullets like
`- One-liner:` with no content.

Handle both template forms:
- Labeled:  `**One-liner:** prose`  → prose
- Bare:     `**prose**`             → prose (legacy)

Empty prose after a label returns null so no bogus bullets are emitted.

Note: existing MILESTONES.md entries generated under the bug are not
regenerated here — that is a follow-up.

Closes #2660

* fix(#2660): normalize CRLF before one-liner extraction

Windows-authored SUMMARY files use CRLF line endings; the LF-only regex
in extractOneLinerFromBody would fail to match. Normalize \r\n and \r
to \n before stripping frontmatter and matching the one-liner pattern.

Adds test case (h) covering CRLF input.
2026-04-24 20:22:29 -04:00
Tom Boucher
06463860e4 fix(#2638): write sub_repos to canonical planning.sub_repos (#2668)
loadConfig's multiRepo migration and filesystem-sync writers targeted the
top-level parsed.sub_repos, but KNOWN_TOP_LEVEL (the unknown-key validator's
allowlist) only recognizes planning.sub_repos (canonical per #2561). Each
migration/sync therefore persisted a key the next loadConfig call warned was
unknown.

Redirect both writers to parsed.planning.sub_repos, ensuring parsed.planning
is initialized first. Also self-heal legacy/buggy installs by stripping any
stale top-level sub_repos on load, preserving its value as the
planning.sub_repos seed if that slot is empty.

Tests cover: (a) canonical planning.sub_repos emits no warning, (b) multiRepo
migration writes to planning.sub_repos with no top-level residue,
(c) filesystem sync relocates to planning.sub_repos, (d) stale top-level
sub_repos from older buggy installs is stripped on load.

Closes #2638
2026-04-24 18:05:33 -04:00
Tom Boucher
74da61fb4a fix(#2619): prevent extractCurrentMilestone from truncating on phase-vX.Y headings (#2624)
* fix(#2619): prevent extractCurrentMilestone from truncating on phase-vX.Y headings

extractCurrentMilestone sliced ROADMAP.md to the current milestone by
looking for the next milestone heading with a greedy regex:

    ^#{1,N}\s+(?:.*v\d+\.\d+|✅|📋|🚧)

Any heading that mentioned a version literal matched — including phase
headings like "### Phase 12: v1.0 Tech-Debt Closure". When the current
milestone was at the same heading level as the phases (### 🚧 v1.1 …),
the slice terminated at the first such phase, hiding every phase that
followed from phase.insert, validate.health W007, and other SDK commands.

Fix: add a `(?!Phase\s+\S)` negative lookahead so phase headings can
never be treated as milestone boundaries. Phase headings always start
with the literal `Phase `, so this is a clean exclusion.

Applied to:
- get-shit-done/bin/lib/core.cjs (extractCurrentMilestone)
- sdk/src/query/roadmap.ts (extractCurrentMilestone + extractNextMilestoneSection)

Regression tests:
- tests/roadmap-phase-fallback.test.cjs: extractCurrentMilestone does not
  truncate on phase heading containing vX.Y (#2619)
- sdk/src/query/roadmap.test.ts: extractCurrentMilestone bug-2619: does
  not truncate at a phase heading containing vX.Y

Closes #2619

* fix(#2619): make milestone-boundary Phase lookahead case-insensitive

CodeRabbit follow-up on #2619: the negative lookahead `(?!Phase\s+\S)`
in the SDK milestone-boundary regex was case-sensitive, so headings like
`### PHASE 12: v1.0 Tech-Debt` or `### phase 12: …` still truncated the
milestone slice. Add the `i` flag (now `gmi`).

The sibling CJS regex in get-shit-done/bin/lib/core.cjs already uses the
`mi` flag, so it is already case-insensitive; added a regression test to
lock that in.

- sdk/src/query/roadmap.ts: change flags from `gm` → `gmi`
- sdk/src/query/roadmap.test.ts: add PHASE/phase regression test
- tests/roadmap-phase-fallback.test.cjs: add PHASE/phase regression test

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:53:20 -04:00
Tom Boucher
1a3d953767 feat: add unified post-planning gap checker (closes #2493) (#2610)
* feat: add unified post-planning gap checker (closes #2493)

Adds a unified post-planning gap checker as Step 13e of plan-phase.md.
After all plans are generated and committed, scans REQUIREMENTS.md and
CONTEXT.md <decisions> against every PLAN.md in the phase directory and
emits a single Source | Item | Status table.

Why
- The existing Requirements Coverage Gate (§13) blocks/re-plans on REQ
  gaps but emits two separate per-source signals. Issue #2493 asks for
  one unified report after planning so that requirements AND
  discuss-phase decisions slipping through are surfaced in one place
  before execution starts.

What
- New workflow.post_planning_gaps boolean config key, default true,
  added to VALID_CONFIG_KEYS, CONFIG_DEFAULTS, hardcoded.workflow, and
  cmdConfigSet (boolean validation).
- New get-shit-done/bin/lib/decisions.cjs — shared parser for
  CONTEXT.md <decisions> blocks (D-NN entries). Designed for reuse by
  the related #2492 plan/verify decision gates.
- New get-shit-done/bin/lib/gap-checker.cjs — parses REQUIREMENTS.md
  (checkbox + traceability table forms), reads CONTEXT.md decisions,
  walks PHASE_DIR/*-PLAN.md, runs word-boundary coverage detection
  (REQ-1 must not match REQ-10), formats a sorted report.
- New gsd-tools gap-analysis CLI command wired through gsd-tools.cjs.
- workflows/plan-phase.md gains §13e between §13d (commit plans) and
  §14 (Present Final Status). Existing §13 gate preserved — §13e is
  additive and non-blocking.
- sdk/prompts/workflows/plan-phase.md gets an equivalent
  post_planning_gaps step for headless mode.
- Docs: CONFIGURATION.md, references/planning-config.md, INVENTORY.md,
  INVENTORY-MANIFEST.json all updated.

Tests
- tests/post-planning-gaps-2493.test.cjs: 30 test cases covering step
  insertion position, decisions parser, gap detector behavior
  (covered/not-covered, false-positive guard, missing-file
  resilience, malformed-input resilience, gate on/off, deterministic
  natural sort), and full config integration.
- Full suite: 5234 / 5234 pass.

Design decisions
- Numbered §13e (sub-step), not §14 — §14 already exists (Present
  Final Status); inserting before it preserves downstream auto-advance
  step numbers.
- Existing §13 gate kept, not replaced — §13 blocks/re-plans on
  REQ gaps; §13e is the unified post-hoc report. Per spec: "default
  behavior MUST be backward compatible."
- Word-boundary ID matching avoids REQ-1 matching REQ-10 and avoids
  brittle semantic/substring matching.
- Shared decisions.cjs parser so #2492 can reuse the same regex.
- Natural-sort keys (REQ-02 before REQ-10) for deterministic output.
- Boolean validation in cmdConfigSet rejects non-boolean values
  matches the precedent set by drift_threshold/drift_action.

Closes #2493

* fix(#2493): expose post_planning_gaps in loadConfig() + sync schema example

Address CodeRabbit review on PR #2610:

- core.cjs loadConfig(): return post_planning_gaps from both the
  config.json branch and the global ~/.gsd/defaults.json fallback so
  callers can rely on config.post_planning_gaps regardless of whether
  the key is present (comment 3127977404, Major).
- docs/CONFIGURATION.md: add workflow.post_planning_gaps to the Full
  Schema JSON example so copy/paste users see the new toggle alongside
  security_block_on (comment 3127977392, Minor).
- tests/post-planning-gaps-2493.test.cjs: regression coverage for
  loadConfig() — default true when key absent, honors explicit
  true/false from workflow.post_planning_gaps.
2026-04-22 23:03:59 -04:00
Tom Boucher
cc17886c51 feat: make model profiles runtime-aware for Codex/non-Claude runtimes (closes #2517) (#2609)
* feat: make model profiles runtime-aware for Codex/non-Claude runtimes (closes #2517)

Adds an optional top-level `runtime` config key plus a
`model_profile_overrides[runtime][tier]` map. When `runtime` is set,
profile tiers (opus/sonnet/haiku) resolve to runtime-native model IDs
(and reasoning_effort where supported) instead of bare Claude aliases.

Codex defaults from the spec:
  opus   -> gpt-5.4        reasoning_effort: xhigh
  sonnet -> gpt-5.3-codex  reasoning_effort: medium
  haiku  -> gpt-5.4-mini   reasoning_effort: medium

Claude defaults mirror MODEL_ALIAS_MAP. Unknown runtimes fall back to
the Claude-alias safe default rather than emit IDs the runtime cannot
accept. reasoning_effort is only emitted into Codex install paths;
never returned from resolveModelInternal and never written to Claude
agent frontmatter.

Backwards compatible: any user without `runtime` set sees identical
behavior — the new branch is gated on `config.runtime != null`.

Precedence (highest to lowest):
  1. per-agent model_overrides
  2. runtime-aware tier resolution (when `runtime` is set)
  3. resolve_model_ids: "omit"
  4. Claude-native default
  5. inherit (literal passthrough)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(#2517): address adversarial review of #2609 (findings 1-16)

Addresses all 16 findings from the adversarial review of PR #2609.
Each finding is enumerated below with its resolution.

CRITICAL
- F1: readGsdRuntimeProfileResolver(targetDir) now probes per-project
  .planning/config.json AND ~/.gsd/defaults.json with per-project winning,
  so the PR's headline claim ("set runtime in project config and Codex
  TOML emit picks it up") actually holds end-to-end.
- F2: resolveTierEntry field-merges user overrides with built-in defaults.
  The CONFIGURATION.md string-shorthand example
    `{ codex: { opus: "gpt-5-pro" } }`
  now keeps reasoning_effort from the built-in entry. Partial-object
  overrides like `{ opus: { reasoning_effort: 'low' } }` keep the
  built-in model. Both paths regression-tested.

MAJOR
- F3: resolveReasoningEffortInternal gates strictly on the
  RUNTIMES_WITH_REASONING_EFFORT allowlist regardless of override
  presence. Override + unknown-runtime no longer leaks reasoning_effort.
- F4: runtime:"claude" is now a no-op for resolution (it is the implicit
  default). It no longer hijacks resolve_model_ids:"omit". Existing
  tests for `runtime:"claude"` returning Claude IDs were rewritten to
  reflect the no-op semantics; new test asserts the omit case returns "".
- F5: _readGsdConfigFile in install.js writes a stderr warning on JSON
  parse failure instead of silently returning null. Read failure and
  parse failure are warned separately. Library require is hoisted to top
  of install.js so it is not co-mingled with config-read failure modes.
- F6: install.js requires for core.cjs / model-profiles.cjs are hoisted
  to the top of the file with __dirname-based absolute paths so global
  npm install works regardless of cwd. Test asserts both lib paths exist
  relative to install.js __dirname.
- F7: docs/CONFIGURATION.md `runtime` row no longer lists `opencode` as
  a valid runtime — install-path emission for non-Codex runtimes is
  explicitly out of scope per #2517 / #2612, and the doc now points at
  #2612 for the follow-on work. resolveModelInternal still accepts any
  runtime string (back-compat) and falls back safely for unknown values.
- F8: Tests now isolate HOME (and GSD_HOME) to a per-test tmpdir so the
  developer's real ~/.gsd/defaults.json cannot bleed into assertions.
  Same pattern CodeRabbit caught on PRs #2603 / #2604.
- F9: `runtime` and `model_profile_overrides` documented as flat-only
  in core.cjs comments — not routed through `get()` because they are
  top-level keys per docs/CONFIGURATION.md and introducing nested
  resolution for two new keys was not worth the edge-case surface.
- F10/F13: loadConfig now invokes _warnUnknownProfileOverrides on the
  raw parsed config so direct .planning/config.json edits surface
  unknown runtime values (e.g. typo `runtime: "codx"`) and unknown
  tier values (e.g. `model_profile_overrides.codex.banana`) at read
  time. Warnings only — preserves back-compat for runtimes added
  later. Per-process warning cache prevents log spam across repeated
  loadConfig calls.

MINOR / NIT
- F11: Removed dead `tier || 'sonnet'` defensive shortcut. The local
  is now `const alias = tier;` with a comment explaining why `tier`
  is guaranteed truthy at that point (every MODEL_PROFILES entry
  defines `balanced`, the fallback profile).
- F12: Extracted resolveTierEntry() in core.cjs as the single source
  of truth for runtime-aware tier resolution. core.cjs and bin/install.js
  both consume it — no duplicated lookup logic between the two files.
- F14: Added regression tests for findings #1, #2, #3, #4, #6, #10, #13
  in tests/issue-2517-runtime-aware-profiles.test.cjs. Each must-fix
  path has a corresponding test that fails against the pre-fix code
  and passes against the post-fix code.
- F15: docs/CONFIGURATION.md `model_profile` row cross-references
  #1713 / #1806 next to the `adaptive` enum value.
- F16: RUNTIME_PROFILE_MAP remains in core.cjs as the single source of
  truth; install.js imports it through the exported resolveTierEntry
  helper rather than carrying its own copy. Doc files (CONFIGURATION.md,
  USER-GUIDE.md, settings.md) intentionally still embed the IDs as text
  — code comment in core.cjs flags that those doc files must be updated
  whenever the constant changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 23:00:37 -04:00
Tom Boucher
73c1af5168 fix(#2543): replace legacy /gsd-<cmd> syntax with /gsd:<cmd> across all source files (#2595)
Commands are now installed as commands/gsd/<name>.md and invoked as
/gsd:<name> in Claude Code. The old hyphen form /gsd-<name> was still
hardcoded in hundreds of places across workflows, references, templates,
lib modules, and command files — causing "Unknown command" errors
whenever GSD suggested a command to the user.

Replace all /gsd-<cmd> occurrences where <cmd> is a known command name
(derived at runtime from commands/gsd/*.md) using a targeted Node.js
script. Agent names, tool names (gsd-sdk, gsd-tools), directory names,
and path fragments are not touched.

Adds regression test tests/bug-2543-gsd-slash-namespace.test.cjs that
enforces zero legacy occurrences going forward. Removes inverted
tests/stale-colon-refs.test.cjs (bug #1748) which enforced the now-obsolete
hyphen form; the new bug-2543 test supersedes it. Updates 5 assertion
tests that hardcoded the old hyphen form to accept the new colon form.

Closes #2543

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 12:04:25 -04:00
Tom Boucher
0d6349a6c1 fix(#2554): preserve leading zero in getMilestonePhaseFilter (#2585)
The normalization `replace(/^0+/, '')` over-stripped decimal phase IDs:
`"00.1"` collapsed to `".1"`, while the disk-side extractor yielded
`"0.1"` from `"00.1-<slug>"`. Set membership failed and inserted decimal
phases were silently excluded from every disk scan inside
`buildStateFrontmatter`, causing `state update` to rewind progress
counters.

Strip leading zeros only when followed by a digit
(`replace(/^0+(?=\d)/, '')`), preserving the zero before the decimal
point while keeping existing behavior for zero-padded integer IDs.

Closes #2554
2026-04-22 12:03:56 -04:00
Tom Boucher
2b494407e5 feat(assembly): add link mode for CLAUDE.md @-reference sections (#2484)
* feat(assembly): add link mode for CLAUDE.md @-reference sections (#2415)

Adds `claude_md_assembly.mode: "link"` config option that writes
`@.planning/<source>` instead of inlining content between GSD markers,
reducing typical CLAUDE.md size by ~65%. Per-block overrides available
via `claude_md_assembly.blocks.<section>`. Falls back to embed for
sections without a real source file (workflow, fallbacks).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): add positive assertion for embedded workflow content (CodeRabbit #2484)

The negative assertion only confirmed @GSD defaults wasn't written.
Add assert.ok(content.includes('GSD Workflow Enforcement')) to verify
the workflow section is actually embedded inline when link mode falls back.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 18:27:55 -04:00
Tom Boucher
5afcd5577e fix: zero-padded phase numbers bypass archived-phase guard; stale current_milestone (#2458)
* fix(sdk): extractCurrentMilestone Backlog leak + state.begin-phase flag parsing

Closes #2422
Closes #2420

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(sdk): skip stateVersion early-return for shipped milestones

When STATE.md has a stale `milestone: v1.0` entry but v1.0 is already
shipped (heading contains ✅ in ROADMAP.md), the stateVersion early-return
path in getMilestoneInfo was returning v1.0 instead of detecting the new
active milestone.

Two-part fix:
1. In the stateVersion block: skip the early-return when the matched
   heading line includes ✅ (shipped marker). Fall through to normal
   detection instead.
2. In the heading-format fallback regex: add a negative lookahead
   `(?!.*✅)` so the regex never matches a ✅ heading regardless of
   whether stateVersion was present. This handles the no-STATE.md case
   and ensures fallthrough from part 1 actually finds the next milestone.

Adds two regression tests covering both ✅-suffix (`## v1.0 ✅ Name`)
and ✅-prefix (`## ✅ v1.0 Name`) heading formats.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(core): allow padded-and-unpadded phase headings in getRoadmapPhaseInternal

The zero-strip normalization (01→1) fixed the archived-phase guard but
broke lookup against ROADMAP headings that still use zero-padded numbers
like "Phase 01:". Change the regex to use 0*<normalized> so both formats
match, making the fix robust regardless of ROADMAP heading style.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 10:08:40 -04:00
Logan
fbf30792f3 docs: authoritative shipped-surface inventory with filesystem-backed parity tests (#2390)
* docs: finish trust-bug fixes in user guide and commands

Correct load-bearing defects in the v1.36.0 docs corpus so readers stop
acting on wrong defaults and stale exhaustiveness claims.

- README.md: drop "Complete feature"/"Every command"/"All 18 agents"
  exhaustiveness claims; replace version-pinned "What's new in v1.32"
  bullet with a CHANGELOG pointer.
- CONFIGURATION.md: fix `claude_md_path` default (null/none -> `./CLAUDE.md`)
  in both Full Schema and core settings table; correct `workflow.tdd_mode`
  provenance from "Added in v1.37" to "Added in v1.36".
- USER-GUIDE.md: fix `workflow.discuss_mode` default (`standard` ->
  `discuss`) in the workflow-toggles table AND in the abbreviated Full
  Schema JSON block above it; align the Options cell with the shipped
  enum.
- COMMANDS.md: drop "Complete command syntax" subtitle overclaim to
  match the README posture.
- AGENTS.md: weaken "All 21 specialized agents" header to reflect that
  the `agents/` filesystem is authoritative (shipped roster is 31).

Part 1 of a stacked docs refresh series (PR 1/4).

* docs: refresh shipped surface coverage for v1.36

Close the v1.36.0 shipped-surface gaps in the docs corpus.

- COMMANDS.md: add /gsd-graphify section (build/query/status/diff) and
  its config gate; expand /gsd-quick with --validate flag and list/
  status/resume subcommands; expand /gsd-thread with list --open, list
  --resolved, close <slug>, status <slug>.
- CLI-TOOLS.md: replace the hardcoded "15 domain modules" count with a
  pointer to the Module Architecture table; add a graphify verb-family
  section (build/query/status/diff/snapshot); add Graphify and Learnings
  rows to the Module Architecture table.
- FEATURES.md: add TOC entries for #116 TDD Pipeline Mode and #117
  Knowledge Graph Integration; add the #117 body with REQ-GRAPH-01..05.
- CONFIGURATION.md: move security_enforcement / security_asvs_level /
  security_block_on from root into `workflow.*` in Full Schema to match
  templates/config.json and the gsd-sdk runtime reads; update Security
  Settings table to use the workflow.* prefix; add planning.sub_repos
  to Full Schema and description table; add a Graphify Settings section
  documenting graphify.enabled and graphify.build_timeout.

Note: VALID_CONFIG_KEYS in bin/lib/config.cjs does not yet include
workflow.security_* or planning.sub_repos, so config-set currently
rejects them. That is a pre-existing validator gap that this PR does
not attempt to fix; the docs now correctly describe where these keys
live per the shipped template and runtime reads.

Part 2 of a stacked docs refresh series (PR 2/5), based on PR 1.

* docs: make inventory authoritative and reconcile architecture

Upgrade docs/INVENTORY.md from "complete for agents, selective for others"
to authoritative across all six shipped-surface families, and reconcile
docs/ARCHITECTURE.md against the new inventory so the PR that introduces
INVENTORY does not also introduce an INVENTORY/ARCHITECTURE contradiction.

- docs/AGENTS.md: weaken "21 specialized agents" header to 21 primary +
  10 advanced (31 shipped); add new "Advanced and Specialized Agents"
  section with concise role cards for the 10 previously-omitted shipped
  agents (pattern-mapper, debug-session-manager, code-reviewer,
  code-fixer, ai-researcher, domain-researcher, eval-planner,
  eval-auditor, framework-selector, intel-updater); footnote the Agent
  Tool Permissions Summary as primary-agents-only so it no longer
  misleads.

- docs/INVENTORY.md (rewritten to be authoritative):
  * Full 31-agent roster with one-line role + spawner + primary-doc
    status per agent (unchanged from prior partial work).
  * Commands: full 75-row enumeration grouped by Core Workflow, Phase &
    Milestone Management, Session & Navigation, Codebase Intelligence,
    Review/Debug/Recovery, and Docs/Profile/Utilities — each row
    carries a one-line role derived from the command's frontmatter and
    a link to the source file.
  * Workflows: full 72-row enumeration covering every
    get-shit-done/workflows/*.md, with a one-line role per workflow and
    a column naming the user-facing command (or internal orchestrator)
    that invokes it.
  * References: full 41-row enumeration grouped by Core, Workflow,
    Thinking-Model clusters, and the Modular Planner decomposition,
    matching the groupings docs/ARCHITECTURE.md already uses; notes
    the few-shot-examples subdirectory separately.
  * CLI Modules and Hooks: unchanged — already full rosters.
  * Maintenance section rewritten to describe the drift-guard test
    suite that will land in PR4 (inventory-counts, commands-doc-parity,
    agents-doc-parity, cli-modules-doc-parity, hooks-doc-parity).

- docs/ARCHITECTURE.md reconciled against INVENTORY:
  * References block: drop the stale "(35 total)" count; point at
    INVENTORY.md#references-41-shipped for the authoritative count.
  * CLI Tools block: drop the stale "19 domain modules" count; point
    at INVENTORY.md#cli-modules-24-shipped for the authoritative roster.
  * Agent Spawn Categories: relabel as "Primary Agent Spawn Categories"
    and add a footer naming the 10 advanced agents and pointing at
    INVENTORY.md#agents-31-shipped for the full 31-agent roster.

- docs/CONFIGURATION.md: preserve the six model-profile rows added in
  the prior partial work, and tighten the fallback note so it names the
  13 shipped agents without an explicit profile row, documents
  model_overrides as the escape hatch, and points at INVENTORY.md for
  the authoritative 31-agent roster.

Part 3 of a stacked docs refresh series (PR 3/4). Remaining consistency
work (USER-GUIDE config-section delete-and-link, FEATURES.md TOC
reorder, ARCHITECTURE.md Hook-table expansion + installation-layout
collapse, CLI-TOOLS.md module-row additions, workflow-discuss-mode
invocation normalization, and the five doc-parity tests) lands in PR4.

* test(docs): add consistency guards and remove duplicate refs

Consolidates USER-GUIDE.md's command/config duplicates into pointers to
COMMANDS.md and CONFIGURATION.md (kills a ghost `resolve_model_ids` key
and a stale `discuss_mode: standard` default); reorders FEATURES.md TOC
chronologically so v1.32 precedes v1.34/1.35/1.36; expands
ARCHITECTURE.md's Hook table to the 11 shipped hooks
(gsd-read-injection-scanner, gsd-check-update-worker) and collapses
the installation-layout hook enumeration to the *.js/*.sh pattern form;
adds audit/gsd2-import/intel rows and state signal-*, audit-open,
from-gsd2 verbs to CLI-TOOLS.md; normalizes workflow-discuss-mode.md
invocations to `node gsd-tools.cjs config-set`.

Adds five drift guards anchored on docs/INVENTORY.md as the
authoritative roster: inventory-counts (all six families),
commands/agents/cli-modules/hooks parity checks that every shipped
surface has a row somewhere.

* fix(convergence): thread --ws to review agent; add stall and max-cycles behavioral tests

- Thread GSD_WS through to review agent spawn in plan-review-convergence
  workflow (step 5a) so --ws scoping is symmetric with planning step
- Add behavioral stall detection test: asserts workflow compares
  HIGH_COUNT >= prev_high_count and emits a stall warning
- Add behavioral --max-cycles 1 test: asserts workflow reaches escalation
  gate when cycle >= MAX_CYCLES with HIGH > 0 after a single cycle
- Include original PR files (commands, workflow, tests) as the branch
  predated the PR commits

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(docs,config): PR #2390 review — security_* config keys and REQ-GRAPH-02 scope

Addresses trek-e's review items that don't require rebase:

- config.cjs: add workflow.security_enforcement, workflow.security_asvs_level,
  workflow.security_block_on to VALID_CONFIG_KEYS so gsd-sdk config-set accepts
  them (closed the gap where docs/CONFIGURATION.md listed keys the validator
  rejected).
- core.cjs: add matching CONFIG_DEFAULTS entries (true / 1 / 'high') so the
  canonical defaults table matches the documented values.
- config.cjs: wire the three keys into the new-project workflow defaults so
  fresh configs inherit them.
- planning-config.md: document the three keys in the Workflow Fields table,
  keeping the CONFIG_DEFAULTS ↔ doc parity test happy.
- config-field-docs.test.cjs: extend NAMESPACE_MAP so the flat keys in
  CONFIG_DEFAULTS resolve to their workflow.* doc rows.
- FEATURES.md REQ-GRAPH-02: split the slash-command surface (build|query|
  status|diff) from the CLI surface which additionally exposes `snapshot`
  (invoked automatically at the tail of `graphify build`). The prior text
  overstated the slash-command surface.

* docs(inventory): refresh rosters and counts for post-rebase drift

origin/main accumulated surfaces since this PR was authored:

- Agents: 31 → 33 (+ gsd-doc-classifier, gsd-doc-synthesizer)
- Commands: 76 → 82 (+ ingest-docs, ultraplan-phase, spike, spike-wrap-up,
  sketch, sketch-wrap-up)
- Workflows: 73 → 79 (same 6 names)
- References: 41 → 49 (+ debugger-philosophy, doc-conflict-engine,
  mandatory-initial-read, project-skills-discovery, sketch-interactivity,
  sketch-theme-system, sketch-tooling, sketch-variant-patterns)

Adds rows in the existing sub-groupings, introduces a Sketch References
subsection, and bumps all four headline counts. Roles are pulled from
source frontmatter / purpose blocks for each file. All 5 parity tests
(inventory-counts, agents-doc-parity, commands-doc-parity,
cli-modules-doc-parity, hooks-doc-parity) pass against this state —
156 assertions, 0 failures.

Also updates the 'Coverage note' advanced-agent count 10 → 12 and the
few-shot-examples footnote "41 top-level references" → "49" to keep the
file internally consistent.

* docs(agents): add advanced stubs for gsd-doc-classifier and gsd-doc-synthesizer

Both agents ship on main (spawned by /gsd-ingest-docs) but had no
coverage in docs/AGENTS.md. Adds the "advanced stub" entries (Role,
property table, Key behaviors) following the template used by the other
10 advanced/specialized agents in the same section.

Also updates the Agent Tool Permissions Summary scope note from
"10 advanced/specialized agents" to 12 to reflect the two new stubs.

* docs(commands): add entries for ingest-docs, ultraplan-phase, plan-review-convergence

These three commands ship on main (plan-review-convergence via trek-e's
4b452d29 commit on this branch) but had no user-facing section in
docs/COMMANDS.md — they lived only in INVENTORY.md. The commands-doc-parity
test already passes via INVENTORY, but the user-facing doc was missing
canonical explanations, argument tables, and examples.

- /gsd-plan-review-convergence → Core Workflow (after /gsd-plan-phase)
- /gsd-ultraplan-phase → Core Workflow (after plan-review-convergence)
- /gsd-ingest-docs → Brownfield (after /gsd-import, since both consume
  the references/doc-conflict-engine.md contract)

Content pulled from each command's frontmatter and workflow purpose block.

* test: remove redundant ARCHITECTURE.md count tests

tests/architecture-counts.test.cjs and tests/command-count-sync.test.cjs
were added when docs/ARCHITECTURE.md carried hardcoded counts for commands/
workflows/agents. With the PR #2390 cleanup, ARCHITECTURE.md no longer
owns those numbers — docs/INVENTORY.md does, enforced by
tests/inventory-counts.test.cjs (scans the same filesystem directories
with the same readdirSync filter).

Keeping these ARCHITECTURE-specific tests would re-introduce the hardcoded
counts they guard, defeating trek-e's review point. The single-source-of-
truth parity tests already catch the same drift scenarios.

Related: #2257 (the regression this replaced).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 09:31:34 -04:00
Tom Boucher
3589f7b256 fix(worktrees): prune orphaned worktrees in code, not prose (#2367)
* feat: add /gsd-spec-phase — Socratic spec refinement with ambiguity scoring (#2213)

Introduces `/gsd-spec-phase <phase>` as an optional pre-step before discuss-phase.
Clarifies WHAT a phase delivers (requirements, boundaries, acceptance criteria) with
quantitative ambiguity scoring before discuss-phase handles HOW to implement.

- `commands/gsd/spec-phase.md` — slash command routing to workflow
- `get-shit-done/workflows/spec-phase.md` — full Socratic interview loop (up to 6
  rounds, 5 rotating perspectives: Researcher, Simplifier, Boundary Keeper, Failure
  Analyst, Seed Closer) with weighted 4-dimension ambiguity gate (≤ 0.20 to write SPEC.md)
- `get-shit-done/templates/spec.md` — SPEC.md template with falsifiable requirements
  (Current/Target/Acceptance per requirement), Boundaries, Acceptance Criteria,
  Ambiguity Report, and Interview Log; includes two full worked examples
- `get-shit-done/workflows/discuss-phase.md` — new `check_spec` step detects
  `{padded_phase}-SPEC.md` at startup; displays "Found SPEC.md — N requirements
  locked. Focusing on implementation decisions."; `analyze_phase` respects `spec_loaded`
  flag to skip "what/why" gray areas; `write_context` emits `<spec_lock>` section
  with boundary summary and canonical ref to SPEC.md
- `docs/ARCHITECTURE.md` — update command/workflow counts (74→75, 71→72)

Closes #2213

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(worktrees): auto-prune merged worktrees in code, not prose

Adds pruneOrphanedWorktrees(repoRoot) to core.cjs. It runs on every
cmdInitProgress call (the entry point for most GSD commands) and removes
linked worktrees whose branch is fully merged into main, then runs
git worktree prune to clear stale references. Guards prevent removal of
the main worktree, the current process.cwd(), or any unmerged branch.

Covered by 4 new real-git integration tests in
tests/prune-orphaned-worktrees.test.cjs (TDD red→green).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-17 10:11:08 -04:00
Tom Boucher
875b257c18 fix(init): embed auto_advance/auto_chain_active/mode in init plan-phase output (#2228)
Prevents infinite config-get loops on Kimi K2.5 and other models that
re-execute bash tool calls when they encounter config-get subshell patterns.
Values are now bundled into the init plan-phase JSON so step 15 of
plan-phase.md can read them directly without separate shell calls.

Closes #2192

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-15 14:59:15 -04:00
Andreas Brauchli
72af8cd0f7 fix: display relative time in intel status output (#2132)
* fix: display relative time instead of UTC in intel status output

The `updated_at` timestamps in `gsd-tools intel status` were displayed
as raw ISO/UTC strings, making them appear to show the wrong time in
non-UTC timezones. Replace with fuzzy relative times ("5 minutes ago",
"1 day ago") which are timezone-agnostic and more useful for freshness.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: add regression tests for timeAgo utility

Covers boundary values (seconds/minutes/hours/days/months/years),
singular vs plural formatting, and future-date edge case.

Addresses review feedback on #2132.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-12 15:54:17 -04:00