* fix(pr-branch): handle sub_repos from config with git -C (#666)
Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:
- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
changes, pushes, and opens a companion PR via `gh pr create`
All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.
Closes#666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: update changeset pr number to 667
* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness
Resolves all three blockers and seven robustness issues raised in PR #667 review:
Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)
Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix
Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected
Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
combinations and the execGit global-trim edge case uniformly
Tests: 17/17 pass, lint: 0 errors
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false
- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
standalone bug-666-*.test.cjs into tests/commands.test.cjs under
describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
Adds allow-test-rule: source-text-is-the-product see #666 for the
workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.
* fix(pr-branch): remove obsolete regression tests for sub-repos handling
* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout
- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
(+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
call in the handle_sub_repos dirty-scan (repo convention: every git
subprocess is bounded, never hangs).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): handle rename staging and split changedFiles from filesToStage
For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): rollback on push failure in cmdPrSubrepo
If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): do not rollback after commit on push failure; add push-fail regression test
Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.
Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): validate sub-repo paths before git invocation in pr-branch.md
The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.
Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.
Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.
Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#666): make sub-repo traversal scan test genuinely fail-first
The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md
Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.
- dirty-scan: realpathSync the root once, and realpathSync each candidate before
the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
backend must still be reported). Confirmed fails-first — regressing the scan to
path.resolve makes the symlink case leak.
Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases
Pre-emptive hardening of the workflow changes from the symlink fix:
- continue-outside-loop: the "skip companion PR on seam failure" block used a
bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
(try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
creation lacks privileges) so it doesn't hard-fail on Windows CI.
Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1145): implement query user-story.validate handler
`query user-story.validate` was a phantom command invoked by mvp-phase.md
(line 102) and verify-work.md (line 170) but had no CJS handler. Every
call exited with "Unknown command: user-story". The dotted-form dispatcher
strips the `query` prefix, splits on `.`, yielding command='user-story'
which fell through to the `default:` case with no registered capability
handler.
Adds a `case 'user-story':` handler inline in gsd-tools.cjs (same pattern
as the #1140 fix in PR #1148). The handler:
- Validates "As a [role], I want to [capability], so that [outcome]."
- Uses \S anchors to require non-whitespace content in each slot
(whitespace-only slots like "As a , I want to ..." now correctly
return valid:false — found by adversarial Codex review)
- Returns { valid: boolean, errors: string[], slots: {role, capability,
outcome} | null }
- Supports --pick valid for bare boolean output (verify-work.md usage)
- Added 'user-story' to SKIP_ROOT_RESOLUTION (pure string validation)
- Added 'user-story' to TOP_LEVEL_USAGE command list
Regression tests added to tests/commands.test.cjs (per lint-regression-
test-names policy; new bug-NNNN standalone files are banned).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(changeset): backfill PR number for #1145 fix
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
groupFilesBySubrepo selected the first sub_repos entry in array order
whose prefix matched a file, so a file under a more-specific nested
sub-repo (e.g. packages/core/widget.js with sub_repos
["packages", "packages/core"]) was mis-routed to the less-specific
parent ("packages").
Select the longest (most-specific) matching prefix within each
first-segment bucket instead, making routing independent of sub_repos
array order. String()-guard the length comparison so non-string entries
still never throw (preserves the #311 tolerance). Update the stale doc
comment that claimed first-match semantics.
Closes#391.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#628): read snake_case requirements_completed in summary-extract
cmdSummaryExtract read only the kebab frontmatter key
`requirements-completed`, but the tool's own JSON output key and the
milestone-audit `--pick` both use the snake form `requirements_completed`.
A SUMMARY written in the snake form the tool itself emits was silently
read back as [] — a false negative in milestone requirement traceability
with no diagnostic.
Make the reader tolerant of both key forms (kebab takes precedence), so a
round-tripped field is no longer dropped. extractFrontmatter does no
hyphen<->underscore normalization, so distinct keys had to be read
explicitly.
Adds regression tests: snake-only fixture now returns the IDs, and a
both-forms-present fixture asserts kebab precedence.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#628): add changeset for summary-extract snake-key fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/
Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.
Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
`perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
preserves the five legitimate slug variants that are NOT the directory:
get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
(README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).
New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
ADR-0008 installer migration. On upgrade it walks the legacy
`~/.claude/get-shit-done/` tree, classifies each file via the prior install
manifest, and emits remove-managed / backup-and-remove for managed files
while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
root and symlinked entries; bounds-checks every path under configDir). The
framework rolls back on install failure. Emptied dirs may remain (framework
has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
`get-shit-done` directory token (split token to avoid self-match; case-
insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
mechanical sweep had wrongly rewritten the old-name patterns it exists to
detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
changeset + docs/installer-migrations.md row added.
Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.
Closes#604
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): unsweep pending changesets + allowlist injection-example docs
CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
like CHANGELOG); reverted those body edits so 5 pre-existing malformed
fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
prompt-injection-scan.sh: they contain intentional injection examples /
security-model prose; the path-reference rewrites are kept.
CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): resolve CodeQL alerts surfaced on this PR
The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:
- scripts/ci-test-scope.cjs: build the config-path match from string
.includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
keep the meaningful POSIX-class conversion (js/identity-replacement).
Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)
The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
reaching static regex `.test(file)` calls (not the config rule). Removed ALL
regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
`.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
loop (replace until stable) plus a final bare-opener strip.
Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL
CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): unblock security base64 scan on the large rename diff
The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.
- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
can't carry base64-obfuscated *text* and feeding NUL bytes through the
per-line scanner is pathologically slow. collect_files already filtered
binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
to accommodate very large diffs (the scan itself is unchanged).
Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): sweep get-shit-done refs introduced by merging next
The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)
Verified: guard 0 violations; build green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant
The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.
Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)
CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.
Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan
The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.
Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* enh: bump opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#466): changeset for opus-tier model-ID refresh
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#443): RED unified effort + fast_mode + resolve-execution
All 68 tests failing as expected — no implementation yet.
Covers: effort cascade (tier defaults, overrides, invalid fallthrough),
fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier
escalation, renderEffortForRuntime clamping, resolve-execution CLI,
config schema new keys, QA hostile-input matrix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query
Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max)
and fast_mode propagation knobs, with per-runtime rendering that clamps the unique
tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps
to low on Claude).
Key changes:
- config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys;
add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides,
fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment
- config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults
- model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE
- model-profiles.cjs: re-export new catalog exports
- core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier,
VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig
- commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort;
add cmdResolveExecution (superset command with effort_rendered, effort_param,
effort_propagation, fast_mode, fast_mode_supported)
- gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags
- tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix
- tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort
- docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections
- settings-advanced.md: list new effort/fast_mode keys in confirmation table
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime
- Remove resolveReasoningEffortInternal (catalog-driven effort function) from
core.cjs and its export; remove from commands.cjs destructure import
- Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort
assertions now use resolveEffortInternal + renderEffortForRuntime; Claude
effort is first-class (output_config.effort); unknown runtimes assert param===null
- Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire
resolveReasoningEffortInternal describe with unified effort assertions;
effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(#443): ADR for unified cross-provider effort + fast-mode routing
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* test(#443): architecture-level QA invariants + test-strategy doc
Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs)
covering 8 architectural invariants: cross-provider validity (never emit a value
the real API would 400 on), param/channel contract stability, resolve-execution
JSON contract (all 8 keys + correct types), totality across the full 33-agent
registry, fast-mode honesty (claude always fast_mode_supported=false), precedence
first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing
composition (effort escalation independent of model tier), and config-set round-trip
for all new effort/* and fast_mode/* key namespaces. Append test-strategy section
with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#443): add failing install-wiring tests for effort per-runtime injection (RED)
TDD RED: 10 failing tests covering:
- Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high)
- Gemini .md does NOT get effort: (already passing — Gemini-safe)
- Codex .toml gets model_reasoning_effort via unified resolver
- Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml
- Source purity: agents/*.md have no effort: key (already passing)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified)
- Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs
- Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from
.planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback),
same probe pattern as readGsdRuntimeProfileResolver
- Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching
resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high')
without loadConfig side-effects (no sub-repo detection, no migration writes)
- Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for
runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay
effort-free (Gemini-safe source contract preserved in agents/*.md)
- generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from
unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps
max → xhigh via renderEffortForRuntime('codex', ...)
- installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to
generateCodexAgentToml so per-project config wins for Codex .toml too
- Update failing tests to GREEN: 12/12 pass; all 17 install tests pass;
2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#443): source install effort defaults from manifest (kill drift) + guard test
Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in
resolveInstallTimeEffort with values read from config-defaults.manifest.json at
module init, using the same __dirname-relative path install.js already uses for
all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert
equality between install.js's runtime constants and the manifest on every CI run.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#443): reconcile Codex TOML tests with unified effort design
The #443 unified effort resolver makes generateCodexAgentToml always emit
model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides).
The test 'generated TOML omits reasoning_effort when runtime has none' had an
obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses
unified effort. Convert it to assert the new invariant: Codex TOML always carries a
valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner,
a heavy-tier agent), while model_profile_overrides model override is still respected.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity)
Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read)
and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog()
getter that initialises on first call from resolveInstallTimeEffort /
generateCodexAgentToml / Claude .md effort injection. Requiring install.js in
unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers
manifest IO or throws, eliminating the load-time side effect that changed
subprocess exit codes / stderr on the bench.
Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT)
preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift
still validates them without forcing eager load.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity
runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition
to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR,
GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses
os.homedir() directly — including the ~/.cache/gsd update-check deletion,
~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes —
never touches the real \$HOME during the test.
Without the HOME isolation the install test (which is new to this branch and
is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or
delete files under the real \$HOME, causing runtime-launcher-parity test (D)
to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is
absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback
arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists.
Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess
that the global installer spawns — irrelevant to effort-wiring assertions,
slow, and potentially writes to ~/.npm cache.
All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit
suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#443): add changeset fragment for effort + fast-mode routing
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak
Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main
install block (guarded by !GSD_TEST_MODE), performing a real global Claude
install into $HOME/.claude/. On CI ubuntu where node is on standard PATH,
the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing
runtime-launcher-parity test (D) to exit zero when it must exit non-zero.
Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs
alphabetically before runtime-launcher-parity.test.cjs in the same node
--test invocation. Each runs in a separate worker process but shares the
same HOME. The drift test's install leaks gsd-tools.cjs into that HOME,
then the launcher test's bash subprocess finds it via the $HOME/.claude arm.
Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard
test, before the require(installPath) call. This matches the pattern used
by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring
.install.test.cjs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings)
Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent.
Replace find(non-dash) with a proper flag-consuming loop that collects a single
positional; validate missing/extra positionals and malformed --attempt values.
Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra")
verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from
core.cjs) before accepting a value, mirroring resolveEffortInternal exactly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions
Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects
EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>'
before the closing '---' delimiter using the same EOL as the surrounding
frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the
LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on
Windows (git core.autocrlf=true checkout).
Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and
complex frontmatter cases. Exports injectEffortFrontmatter from module.exports.
Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs
on windows-latest runners (lines 138, 145, 152, 261, 345, 356).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
cmdWebsearch called fetch() with no timeout and no retry, so a hung
connection blocked indefinitely and transient 429/5xx/network failures
were not recovered. Add AbortSignal.timeout (configurable via
GSD_WEBSEARCH_TIMEOUT_MS, default 10s) and a bounded retry loop
(max 2 retries, exponential backoff + jitter) for 429/5xx/network
errors, honoring Retry-After on 429 (capped at 60s). Non-429 4xx fail
immediately (no wasted retries). Transient-exhausted failures report an
`attempts` count. Worst-case time is bounded by timeout*(1+retries)+backoff.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
cmdCommitToSubrepo routed each changed file to a sub-repo via
subRepos.find(...) inside the file loop — O(files * repos). Extract a pure
groupFilesBySubrepo() that buckets sub-repos by first path segment and scans
only the matching bucket, dropping it to expected O(files + repos). First-
match-in-array-order semantics (incl. multi-segment sub-repos) are preserved
exactly. Adds the first behavior-lock test for the routing path.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
On machines where os.tmpdir() returns a path with a space (e.g.
/Volumes/Mini Me/tmp), runGsdTools() string args were whitespace-split by
the helper tokeniser, truncating paths at the first space. Switch all
calls that embed a dynamic path into the argument list to the array form
of runGsdTools() so execFileSync receives each path as a single argv slot.
Fixes#3509
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(adr): add ADR-0003 model catalog module
* fix(#3229): add shared model catalog as source of truth for agent profiles and runtime tier defaults
Research / design (ADR-0003):
- Existing drift came from 4 independent model truths:
1. CJS model-profiles.cjs
2. SDK config-query.ts stale copy (18 agents)
3. settings-advanced.md runtime tier table
4. session-runner Claude-only profile map
- New design: one machine-readable Model Catalog Module in sdk/shared/
that both packages ship and consume.
Implementation:
- sdk/shared/model-catalog.json — canonical source of truth for:
- full 33-agent registry
- per-agent golden (quality) alias + balanced/budget aliases
- adaptive derivation from routingTier
- agent→phaseType map
- agent→dynamic-routing default tier map
- runtime tier defaults for all supported runtimes
- get-shit-done/bin/lib/model-catalog.cjs — CJS adapter over the catalog
- sdk/src/model-catalog.ts — SDK adapter over the same catalog
- CJS model-profiles.cjs now re-exports derived data from model-catalog.cjs
- SDK config-query.ts now re-exports MODEL_PROFILES/VALID_PROFILES from
model-catalog.ts instead of maintaining its own list
- sdk/src/query/helpers.ts runtime list now comes from the catalog (fixes hermes drift)
- sdk/src/session-runner.ts Claude profile→model-id mapping now resolves via catalog
- docs/CONFIGURATION.md + settings-advanced.md runtime tables updated to match catalog
Behavior changes:
- resolve-model now covers every shipped agent file on disk (33 agents)
- unknown-agent fallback is profile-semantic, not hardcoded sonnet:
quality→opus, budget→haiku, balanced/adaptive→sonnet, inherit→inherit
- Group B runtimes remain known runtimes but do not get built-in tier defaults
Tests (RED→GREEN):
- root tests: shipped agent files must equal MODEL_PROFILES keys
- sdk tests: shipped agent files must equal MODEL_PROFILES keys
- direct fix assertion: gsd-code-reviewer resolves to opus under quality with no unknown_agent
- runtime defaults parity test: settings-advanced.md + CONFIGURATION.md tables must match catalog
- helper tests: hermes included in SUPPORTED_RUNTIMES and getRuntimeConfigDir()
Closes#3229
* chore(changeset): update #3229 changeset pr field to 3230
* fix(ci): update inherit fallback expectations and inventory parity for model catalog
* feat(#2982): extend no-source-grep lint to catch var-binding readFileSync.includes()
The base lint (scripts/lint-no-source-grep.cjs) only catches
readFileSync(...).<text-method>() chained directly. The much more
common var-binding form escapes it:
const src = fs.readFileSync(p, 'utf8');
// 50 lines later
if (src.includes('foo')) {} // ← still grep, lint missed it
Scan of the test suite found ~141 files using this pattern.
Implementation built TDD per #2982 with structured-IR assertions:
scripts/lint-no-source-grep-extras.cjs
- detectVarBindingViolations(src) — pure detector, two passes:
pass 1 collects vars bound from readFileSync, pass 2 finds any
<var>.<includes|startsWith|endsWith|match|search>( on those vars.
- detectWrappedAssertOkMatch(src) — flags
assert.ok(<expr>.match(...)) which escapes the assert.match rule.
- VIOLATION enum exposes stable codes for tests to assert on.
scripts/lint-no-source-grep.cjs
- Wires the new detectors into the existing per-file check; one
additional violation row per file with the first 3 sample tokens.
tests/bug-2982-lint-var-binding.test.cjs
- 13 tests, all assertions on typed VIOLATION enum / structured
records. Covers all 5 text-match methods, multi-var, no-bind,
string literal (must NOT trigger), wrapped assert.ok(.match),
and assert.match (must NOT double-flag).
Migration backlog (#2974 expanded scope):
- 42 files annotated `// allow-test-rule: source-text-is-the-product`
(legitimate — they read .md/.json/.yml files whose deployed text
IS the product)
- 3 files annotated `// allow-test-rule: pending-migration-to-typed-ir [#2974]`
(read .cjs/.js source — clear migration debt)
- 95 files annotated `pending-migration-to-typed-ir [#2974]` with
`Per-file review may reclassify as source-text-is-the-product
during migration` (mixed — manual review under #2974)
After this lands the lint reports 0 violations on main; new
violations in PRs surface immediately.
Closes#2982
Refs #2974
* test(#2982): fix truncated test name per CR
The label ended with a bare '(' from a copy-paste mishap. Now reads
'does NOT flag .matchAll(...) — matchAll is not match, so
assert.ok(.matchAll(...)) is not flagged'.
* chore(#2982): add changeset fragment for PR #2985
* chore(#2982): add changeset fragment for PR #2985
When ROADMAP.md uses unpadded phase numbers (e.g. "Phase 1:") and
the phases/ directory uses zero-padded names (e.g. "01-auth"), the
phasesByNumber Map held two separate entries — one keyed "1" from the
ROADMAP heading scan and one keyed "01" from the directory scan —
doubling phases_total in /gsd-stats output.
Apply normalizePhaseName() to all Map keys in both the ROADMAP heading
scan and the directory scan so the two code paths always produce the
same canonical key and merge into a single entry.
Closes#2195
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md
- Replace require('node:assert') with require('node:assert/strict') across
all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
as that is appropriate for per-function resource guards, not lifecycle hooks
Closes#1674
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* perf(tests): add --test-concurrency=4 to test runner for parallel file execution
Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): allowlist verify.test.cjs in prompt-injection scanner
tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Add check-commit command to gsd-tools that acts as a pre-commit guard.
When commit_docs is false, rejects commits that stage .planning/ files
with an actionable error message including the unstage command.
Recreated cleanly on current main — previous version carried stale
shared fixes that are now upstream.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Phases with all summaries but no passing VERIFICATION.md now show as
"Executed" instead of "Complete", preventing false progress reporting.
Adds determinePhaseStatus() helper used by both cmdStats() and
cmdProgressRender(). Also fixes duplicate phase directory accumulation
in cmdStats() — plans/summaries from directories sharing the same
phase number are now summed instead of silently overwritten.
New statuses: Executed (summaries done, no verification), Needs Review
(verification exists with human_needed status).
Closes#1459
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The phase-extraction regex /(\d+)-/ only matched the last integer
segment before a dash, so decimal phases like 45.14 were misresolved
to phase 14 — silently switching to the wrong branch.
Closes#1402
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When branching_strategy is "phase" or "milestone", the branch was only
created during execute-phase — but discuss-phase, plan-phase, and
new-milestone all commit artifacts before that, landing them on main.
Move branch creation into cmdCommit() so the strategy branch is created
at the first commit point in any workflow. execute-phase's existing
handle_branching step becomes a harmless no-op (checkout existing branch).
Also fixes websearch test mocks broken by #1276 (fs.writeSync change).
Closes#1278
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extend resolve_model_ids to accept "omit" value: returns empty string
so non-Claude runtimes (OpenCode, Codex, Gemini, etc.) use their
configured default model instead of unresolvable Claude aliases.
- resolve_model_ids: "omit" short-circuits before alias resolution
- model_overrides still respected (checked first) for explicit IDs
- Installer sets resolve_model_ids: "omit" in ~/.gsd/defaults.json
for non-Claude runtimes during install
- 4 new tests covering omit behavior and override passthrough
- Fix websearch test mocks for fs.writeSync output change (#1276)
Closes#1156
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The summary template puts the one-liner as a `**bold**` line after the
`# Phase N` heading, but `cmdSummaryExtract` and `cmdMilestoneComplete`
only checked frontmatter `one-liner` field — which is often empty.
Adds `extractOneLinerFromBody()` to core.cjs as a fallback that parses
the first `**...**` line after the heading.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add a cross_reference_todos step to discuss-phase that surfaces relevant
backlog items before scope-setting decisions are made.
Implementation:
- New 'todo match-phase <N>' CLI command (commands.cjs) that scores
pending todos against a phase's ROADMAP goal using three heuristics:
keyword overlap, area match, and file path overlap
- New cross_reference_todos step in discuss-phase.md between
load_prior_context and scout_codebase
- CONTEXT.md template gains 'Folded Todos' subsection in <decisions>
and 'Reviewed Todos (not folded)' subsection in <deferred>
Design:
- No AI call for matching — pure keyword/area/file heuristics for speed
- Silent skip when todo_count is 0 or no matches (no workflow slowdown)
- Auto mode folds all todos with score >= 0.4 automatically
- Scoring: keywords (up to 0.6), area match (0.3), file overlap (0.4)
Tests: 5 new tests covering empty state, keyword matching, unrelated
todo exclusion, area matching, and score sorting.
Closes#1111
Add a new `/gsd:stats` command that displays comprehensive project
statistics including phase progress, plan execution metrics, requirements
completion, git history, and timeline information.
- New command definition: commands/gsd/stats.md
- New workflow: get-shit-done/workflows/stats.md
- CLI handler: `gsd-tools stats [json|table]`
- Stats function in commands.cjs with JSON and table output formats
- 5 new tests covering empty project, phase/plan counting, requirements
counting, last activity, and table format rendering
Co-authored-by: ashanuoc <ashanuoc@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Add requirements_completed field to summary-extract output, mapping
from SUMMARY frontmatter requirements-completed key. Enables
/gsd:audit-milestone cross-check to receive data from SUMMARY source.
Re-applied from #631 against refactored codebase (commands.cjs + split tests).
Co-authored-by: Colin Johnson <Solvely-Colin@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
When orphaned SUMMARY.md files cause totalSummaries > totalPlans, the
progress percentage exceeds 100%, making String.repeat() throw RangeError
on negative arguments. Clamp percent to Math.min(100, ...) at all three
computation sites (state, commands, roadmap).
Closes#633
Co-authored-by: vinicius-tersi <vinicius-tersi@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Move 81 tests (18 describe blocks) from single monolithic test file
into 7 domain-specific test files under tests/ with shared helpers.
Test parity verified: 81/81 pass before and after split.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>