* fix(#376): rewrite /gsd: → /gsd- in Claude-installed hook .js files
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#376): preserve .sh branch + {{GSD_VERSION}} stamp in restructured hook-copy loop
Trim the .js branch comment/whitespace so the `else {` and
`entry.endsWith('.sh')` fall within the 1500/2000-char assertion windows
anchored on `configDirReplacement` in the regression tests for #1834 and
#2136. The .sh read+substitute+chmod path is intact; the new #376 hyphen-
namespace rewrite for .js/.cjs files is also preserved.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Introduces KNOWN_TEMPLATE_DEFAULTS and KNOWN_STATUS_PATTERNS in state-document.cjs
to enumerate every string a GSD handler writes. stateReplaceFieldIfTemplate
consults this table and only replaces the field when the existing value is a known
template default (or absent) — executor-authored values are left untouched.
Wire-in:
- record-session: Resume File now only overwritten when caller passes --resume-file
OR existing value is 'None'. Router no longer defaults resume_file to 'None'
before calling the handler.
- advance-plan (both branches): Status and Last Activity guarded via
stateReplaceFieldIfTemplate.
- updateCurrentPositionFields: Status and Last activity in the Current Position
section guarded; bare ISO date shape is the trigger for replacement, prose
narrative is preserved.
- planned-phase: Status and Last Activity guarded the same way.
Regression tests (7 cases) in tests/bug-397-state-preserve-executor-authored.test.cjs
cover each data-loss shape; all 106 existing state.test.cjs tests still pass.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove `DEFAULT_SMOKE_TESTS` and `WINDOWS_SMOKE_TESTS` constants (now dead after the unconditional injection block is dropped)
- Drop the `addAll(targeted, DEFAULT_SMOKE_TESTS)` / `addAll(windows, WINDOWS_SMOKE_TESTS)` block from the `codeChanged` branch
- When `codeChanged && targetedTests.length === 0`, push `'unit'` as the fallback suite token
- Two new regression tests in `tests/ci-test-scope.test.cjs` covering the no-injection and unit-fallback contracts (bug #408)
* fix(#416): return null when active milestone has no archive (no fall-through to prior milestone's archive)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#416): handle bold-formatted Milestone: field in archive dir resolver
The STATE.md regex in getActiveMilestoneArchiveDir failed to extract the
version from **Milestone:** vX.Y format (bold wraps the label+colon).
The old pattern captured '**' instead of the version, causing the
milestone→archive lookup to produce a false candidate path, then return
null (post-fix behavior) instead of falling through to the version-sort
fallback — breaking the #3164 consistency scanner tests.
Fix: extend the regex to skip optional trailing '**' after the colon so
both 'milestone: vX.Y' and '**Milestone:** vX.Y' parse correctly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
cmdWebsearch called fetch() with no timeout and no retry, so a hung
connection blocked indefinitely and transient 429/5xx/network failures
were not recovered. Add AbortSignal.timeout (configurable via
GSD_WEBSEARCH_TIMEOUT_MS, default 10s) and a bounded retry loop
(max 2 retries, exponential backoff + jitter) for 429/5xx/network
errors, honoring Retry-After on 429 (capped at 60s). Non-429 4xx fail
immediately (no wasted retries). Transient-exhausted failures report an
`attempts` count. Worst-case time is bounded by timeout*(1+retries)+backoff.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The update worker queried the unscoped 'get-shit-done-redux' via
`npm view`, which returns E404 — so `latest` stayed null and
`update_available` could never become true. Now derives the name from
package.json (`require('../package.json').name`) so it always matches
the actual published scoped name (@opengsd/get-shit-done-redux).
Fixes#378.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
config-get calls lacked --default, causing stderr noise ("Key not found") and a fragile empty-variable fallback when workflow.nyquist_validation was absent; added --default true (matching schema default). Fixes#138.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When a session dies mid-execution (hang, token exhaustion, API drop),
STATE.md's current_phase can be advanced past a phase that still has
PLAN.md files without matching SUMMARY.md files. Without this fix,
/gsd-next and /gsd-progress would route by current_phase and silently
skip the partially-executed phase, producing a data-loss-shape outcome.
Adds a Route 0 cross-phase incomplete-execution scan to both
next.md and progress.md. Before any current_phase-based routing,
the scan finds the lowest-numbered phase where plans outnumber
summaries and routes to /gsd-execute-phase <that-phase> to resume it.
Opt out with --no-resume to fall back to the prior-phase defer prompt;
--force bypasses all gates as before.
Rework (codex review):
- Route 0 now ordered AFTER Gates 1-3 (repo/state validity always run)
but BEFORE the prior-phase completeness-scan defer prompt — eliminating
the double-decision where the default path would both prompt the user
(C/S/F) AND resume the phase anyway. Prior-phase defer prompt moved to
a new prior_phase_completeness step; only reached via --no-resume.
- --force flow made coherent across all three steps: safety_gates jumps
directly to determine_next_action, skipping Gates, Route 0, AND
prior_phase_completeness. resume_incomplete_phase and prior_phase_completeness
now correctly state --force never reaches them. success_criteria entry
updated to reflect --force → determine_next_action (not prior_phase_completeness).
- Scan uses $GSD_SDK (canonical resolver form) throughout next.md, matching
the file's existing convention. progress.md uses $ROADMAP already loaded
by analyze_roadmap. Neither file uses bare gsd-sdk.
- Errors are surfaced rather than suppressed: removed 2>/dev/null on the
main roadmap.analyze call; added explicit WARNING emission when the scan
cannot run, so the invariant fails closed instead of failing open.
- Predicate aligned to plans-without-summaries (plans.length > summaries.length)
in both files, consistent with determine_next_action Route 4.
- command references use canonical /gsd: namespace form
Fixes#160
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
finishInstall called configureOpencodePermissions unconditionally, causing
fs.mkdirSync + fs.writeFileSync to run even under GSD_TEST_MODE='1', violating
the side-effect-free contract. Guarded the call with !process.env.GSD_TEST_MODE.
Fixes#130
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Static regex literal `byPhaseTablePattern` was recompiled on every call to `updatePerformanceMetricsSection`; hoisted to module scope (compiled once; stateless /i used with .match → safe to share across calls). `phaseRowPattern` uses dynamic interpolation and stays in-function. `Fixes #320`.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Lint workflows used fetch-depth:0 (full clone); switched to depth 50 +
explicit base-ref fetch so the three-dot diff (origin/${base}...HEAD) has
its merge-base; fails closed if merge-base is deeper than 50. Added
policy test asserting fetch-depth:50 on both workflows. Fixes#319.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Release jobs ran `npm install -g npm@latest` before each publish step,
adding ~30 s and version drift risk on every run; removed both occurrences,
relying on Node 24's bundled npm pinned via setup-node. Added a policy test
(tests/policy-release-no-npm-self-upgrade.test.cjs) that will fail RED if
the antipattern is re-introduced in release.yml or hotfix.yml. Fixes#318.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per-PostToolUse hot path did stat-then-read ×3 (config.json, metrics bridge,
warn sentinel); collapsed to read-with-ENOENT-catch (fewer blocking syscalls,
no TOCTOU), behavior identical. Fixes#317.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
acquireStateLock was allocating a fresh SharedArrayBuffer on every retry
iteration via Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), ...).
The buffer is never mutated and never escapes, so hoisting it before the loop
is a provably-equivalent transformation — Atomics.wait always sees value 0
whether the buffer is fresh or reused.
Fixes#316
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
loadConfig called detectSubRepos(cwd) at up to 3 sites per invocation (root-config
requiresFilesystem migration, workstream-config requiresFilesystem migration, and the
planning.sub_repos filesystem re-sync) — all with the same cwd, yielding identical
results. Introduce a per-call lazy memo (getDetectedSubRepos) so the directory scan
runs at most once per loadConfig call while preserving all conditional logic. Fixes#315.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Hot path in cmdRoadmapAnnotateDependencies called planData.find() on every
checklist line. Replaced with a first-wins Map built once before the loop so
each line resolves in O(1); first-wins preserves exact .find() semantics and
null-on-miss → wave-1 default is unchanged.
Fixes#314
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Root cause: the affected-tests runner called runAllSuites() on critical-path
changes (running every suite including install/slow on all matrix cells including
Windows), and pickAffectedTests injected DEFAULT_SMOKE_TESTS (an install test)
as the empty-selection fallback — causing install suite tests to run on PR lanes
where they are push-only per docs/TESTING-SUITES.md.
Fix: PR_EXCLUDED_SUITES filter at the pickAffectedTests chokepoint strips
install/slow from every selection path (direct-change, reverse-index, stem-match).
Empty selection now returns [] and the caller runs the unit suite as smoke.
Critical-path fallback replaces runAllSuites with PR_FULL_SUITES
(unit, integration, security). suiteOf exported from run-tests.cjs
(with require.main guard) so affected-tests-lib reuses canonical detection.
Fixes#370
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
intelExtractExports deduped export names with `if (!arr.includes(x)) arr.push(x)`
across ~8 extraction loops (one doubly-nested over an export block) — O(n^2).
Accumulate into Sets (add/has/size) and materialize to an array once at return.
Set dedups by value and preserves insertion order, so the returned export list
and its first-seen order are identical. Adds behavior-lock tests for dedup + order.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
parseNamedArgs re-scanned argv with indexOf/includes once per flag —
O(flags * argv) — on the command-dispatch hot path (24 call sites across
gsd-tools + init/state/validate routers). Build a first-index Map of argv
tokens in a single pass and use it for the flag lookups, dropping it to
O(argv + flags). Semantics are identical: firstIndex.get(t)??-1 === indexOf(t),
firstIndex.has(t) === includes(t); first-occurrence-wins and the value-token
rejection are preserved. Adds the first behavior-lock tests for the module.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
cmdCommitToSubrepo routed each changed file to a sub-repo via
subRepos.find(...) inside the file loop — O(files * repos). Extract a pure
groupFilesBySubrepo() that buckets sub-repos by first path segment and scans
only the matching bucket, dropping it to expected O(files + repos). First-
match-in-array-order semantics (incl. multi-segment sub-repos) are preserved
exactly. Adds the first behavior-lock test for the routing path.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The hook detaches a background rebuild; the tests raced it two ways on
slow/contended Docker (Mac passed): (1) a tight ~2.2s status poll budget
expired before the detached rebuild wrote its terminal status -> "1 subtest
failed"; (2) cleanupHookRepo's rmSync threw EBUSY/ENOTEMPTY while the child
was still writing -> "failed running after hook".
Replace the ad-hoc poll budgets with a shared waitForBuildStatus() that
waits for the real terminal status ('ok'/'failed') under a generous 30s
deadline (all assertions are outcome-based, so this is deterministic, not a
timing assertion), and make teardown best-effort so a residual temp dir can
never fail a passing test. No production/hook code changed; no assertion
weakened.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
release.yml (rc + finalize) and hotfix.yml (finalize) ran the full
`npm run test:coverage` suite on the release path, redundantly re-running
the integration/install/security/slow suites that already passed on the
PR lanes into next. Switch those three sites to `npm run test:coverage:unit`
(same c8 config, unit suite only) to cut release latency. Full-suite
coverage remains available via the dedicated lanes / `test:coverage:all`.
Adds a workflow-contract regression test asserting the release/hotfix gates
invoke the unit coverage command (exact-line match, not substring).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
learningsCopyFromProject called learningsWrite K times in one process,
and each call re-scanned the entire learnings store to dedupe — O(K*N).
Build the content_hash -> id index once at the start of the bulk import
and thread it through; single-write behavior and the return contract are
unchanged (no caller reads `id` on the created:false branch). O(K*N) ->
O(N+K). Adds a regression test asserting store scan count is independent
of import size.
The larger persistent on-disk index (atomic updates, corruption rebuild,
cross-process dedupe) is deferred — needs design decisions and is not
required to resolve the bulk-import scan this issue reports.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replace the per-render readdirSync().filter().map(statSync).sort() chain
with a single-pass max-by-mtime loop. Drops the O(n log n) sort and the
throwaway intermediate array; I/O and resolved-file behavior are identical.
Adds the first behavior-lock test for the todo-resolution path.
The larger disk-backed cache win from the issue is deferred: statusline is
a fresh child process per render, so any cache must be disk-backed with
invalidation/atomic-write design that needs maintainer input.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The Pass-2 topological level assignment in cmdPhasePlanIndex dequeued its
Kahn's-algorithm queue with Array.shift(), which is O(n) per call in V8, so the
BFS was O(V^2) and slowed superlinearly on deep queues (wide fan-in plan
graphs). Extract the traversal into a pure, exported computeDependencyLevels
(rawPlans, planMap, canonicalToId) and dequeue via a head index (queue[head++])
-> O(V+E). Behavior is identical: same FIFO order, same longest-path levels,
same visited-count cycle detection. A complexity-contract comment above the loop
documents why shift() must not be reintroduced.
Adds tests/phase-dependency-levels.test.cjs with deterministic behavior and
edge-case coverage (linear chain, diamond longest-path, independent set, cycle,
canonical-prefix resolution, empty, self-loop, duplicate edge, external dep). A
timing-based complexity guard was intentionally omitted: the O(V+E) Map-build
constant dilutes the O(V^2) signal until impractical N (~1e6), so an empirical
guard is inherently flaky on contended CI — the contract is enforced by the
inline comment and correctness tests instead.
Fixes#307
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#145): extractCurrentMilestone selects active sub-milestone over closed sibling
extractCurrentMilestone used a non-global regex with content.match()
(first-match) to locate the milestone section for the STATE.md version.
When the milestone is a shared semver prefix (e.g. v8.0) and ROADMAP.md
holds both a closed sub-milestone (## v8.0 ... CLOSED/FAIL) and an active
one (## v8.0-B ... STARTED), the first match was always the closed
heading, so the active section and its phases were excised from the slice
and downstream phase ops failed with "Phase N not found in current
milestone".
Switch to a global matchAll over candidate headings, skip headings
carrying a closed marker (CLOSED/ARCHIVED/ABANDONED/SHIPPED/FAILED/FAIL/
✅/🗄️), and select the first non-closed match (falling back to the first
match when every candidate is closed, preserving legacy behavior). Anchor
the preamble slice to the first heading index so a closed sibling's body
no longer leaks into the preamble when the selected section is later.
Harden version matching with a trailing word boundary (so v8.0-B does not
match v8.0-Beta), narrow the FAIL marker to FAILED, match a bare 🗄, and
add an active-marker override (STARTED/🚧/ACTIVE) so a heading carrying an
explicit active status is never treated as closed even if its name contains
a completion word. Also anchor the preamble at the first any-version
milestone heading so unmatched sibling sections do not leak into the
preamble.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(#145): add changeset fragment for milestone-selection fix
Adds the required .changeset/*.md fragment for this user-facing fix
(changeset-lint gate).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#373): replace unquoted $GSD_SDK with space-safe gsd_run launcher
Workflow bash blocks resolved the runtime as GSD_SDK="node $GSD_TOOLS" and
invoked it unquoted ($GSD_SDK query ...). On install paths containing spaces
(e.g. /Volumes/Mini Me/...) the unquoted expansion word-split into
`node /Volumes/Mini gsd-tools.cjs ...`, failing with "Cannot find module
'/Volumes/Mini'" and getting masked by `2>/dev/null || echo "{}"` into a
silent empty state.
Replace the string variable with a single-line shell launcher that defines a
gsd_run function, invokes the runtime with a fully-quoted path and "$@", and
preserves the local-cjs / installed-gsd-tools-on-PATH fallback (#3668) plus
the loud not-found error and install hint. The launcher uses _GSD_SHIM_NAME
indirection so no workflow emits the /gsd-tools substring that the do.md
dispatcher-parity scanner would misread, and is single-line to stay within
the per-file progressive-disclosure line budgets (#2551).
The canonical launcher lives in
get-shit-done/workflows/_runtime-launcher.snippet.sh, is propagated once per
file by scripts/sync-runtime-launcher.cjs, and is locked by
tests/runtime-launcher-parity.test.cjs (fails CI on drift, on a reappearing
$GSD_SDK token, or on a /gsd-tools substring). Dependent workflow-assertion
tests are updated from $GSD_SDK to gsd_run, and the runtime launcher is
registered in CONTEXT.md.
Fixes#373
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(#373): add changeset fragment and allow-test-rule for parity guard
The runtime-launcher parity test is a structural drift guard that reads
workflow markdown to assert the canonical launcher is present and the
retired $GSD_SDK / /gsd-tools tokens are absent; annotate it with
allow-test-rule per the no-source-grep lint escape hatch. Add the required
.changeset fragment for this user-facing fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(#373): make parity PATH-fallback assertion cross-platform (Windows)
Subtest (E) compared GSD_TOOLS against the Node-side absolute temp path,
but the value originates from git-bash which reports the POSIX form, so the
prefix comparison failed on windows-latest while the launcher itself worked
(the installed stub was invoked). Assert the resolved binary by normalized
suffix (/bin/gsd-tools, not .cjs) instead of the absolute prefix; the
behavioral stub-invocation assertion is unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#371): invoke installed bin via shell for Windows .cmd shims in release-tarball-smoke
`node <gsd-tools.cmd>` cannot execute a Windows batch shim as a JS script, so
runSmoke returned bin_not_callable for every check on Windows. Route .cmd/.bat
shims through shell:true (required by Node >=18.20/20.12) and keep the POSIX
node-invocation path unchanged. Add an exported binInvocation seam plus a
platform-agnostic regression test.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(#371): surface bin-invocation failure details in release-tarball-smoke
The smoke harness captured stderr/stdout in `details` but never printed
them, so Windows bin_not_callable failures gave no actionable cause in CI.
Log the resolved bin, invocation descriptor, exit status/signal/error, and
captured stderr/stdout on spawn-derived failures so the real error is visible.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(#371): include smoke result details in install-test assertion messages
The node --test TAP runner swallows in-test console.error, so Windows
bin_not_callable failures gave no cause. Embed code + details (incl. captured
stderr/stdout) into the assertion messages, which DO reach the CI log, so the
real Windows failure is diagnosable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#371): resolve installed bin from prefix root on Windows
npm install -g --prefix X writes bin shims to X\ (the prefix root) on
Windows, not X\node_modules\.bin\. The smoke harness only searched
node_modules\.bin on win32, so the installed gsd-tools/installer bin was
never found and runSmoke returned bin_not_callable before invoking anything.
Search the prefix root first (then node_modules/.bin as fallback), report the
searched candidates on miss, and drop the TAP-swallowed console.error probes.
The .cmd-via-shell binInvocation fix remains for actually running the shim.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#372): use non-login shell in bug-3668 resolver test for hermetic PATH
bash -lc re-sourced profile files (e.g. Homebrew shellenv) that prepended
real bin dirs ahead of the test's injected PATH, so the installed-gsd-tools
fallback subtest resolved a host-global gsd-tools instead of the injected
fake — failing on dev machines and CI-adjacent benches while passing on
clean CI. Use bash -c (non-login) so the injected PATH is authoritative.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#372): split workflow snippet on CRLF in bug-3668 resolver test
extractResolverSnippet split on '\n' and exact-matched the closing \`fi\`
line, so on a Windows checkout (CRLF) the line was \`fi\r\`, the end marker
was never found, and the test failed with "SDK resolution snippet must end
with fi". Split on /\r?\n/ so extraction works on LF and CRLF checkouts and
the snippet handed to bash is carriage-return-free.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(#372): assert resolver bin by normalized suffix, not exact OS path
On Windows the snippet runs under Git bash, so GSD_TOOLS is reported POSIX-
style / mixed-separator while the test built its expected regex from Node
path.join (backslashes) — the resolver was correct (installed:/runtime: output
proves the right bin ran) but the exact-path assertions failed. Normalize
separators and assert the GSD_TOOLS path suffix, keeping the behavioral
installed:/runtime: assertions as the primary checks.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: use active-workstream resolver exported by store module
* fix: wire verify codebase-drift alias and sync inventory docs
* chore: add changeset for next gate regression fixes
* fix: normalize changeset fragment metadata for docs-lint