Commit Graph

44 Commits

Author SHA1 Message Date
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
115433bba6 fix(#2539): anchor commit phase-token detection; drop silent wrong-branch switch (#2669)
* fix(#2539): anchor commit phase-token extraction to the phases/ segment; drop silent switch-to-existing

cmdCommit auto-detected the commit's phase from --files with an unanchored
`match(/(\d+(?:\.\d+)*)-/)`, which returns the leftmost digit-run-then-hyphen
anywhere in the joined path. A project_code ending in a digit (PROJECT_V2) made
`.planning/phases/PROJECT_V2-07-name/…` match the `2-` inside `V2-` before the
real `07-` token, resolving phase 2. findPhaseInternal also searches archived
milestones, so an existing archived phase 2 produced a real branch name and the
silent `git checkout <existing-branch>` fallback switched the whole working tree
onto the wrong branch in the same call that then committed.

The extraction now anchors to the directory segment immediately under
`.planning/phases/` (or `.planning/milestones/<v>-phases/`) and runs it through
the existing project-code-aware extractPhaseToken helper — the single owner
shared by the other 6 call sites — rather than introducing a fourth independent
copy of phase-token-matching logic. The auto-switch keeps create-if-absent only
(the #1278 intent: ensure the branch exists before the first commit on it); it
no longer force-switches an already-checked-out working branch onto a different
existing branch.

Adds two regression fixtures: a digit-suffixed project_code + an archived phase
whose number collides with the trailing digit (the silent-wrong-branch case),
and a pre-existing phase branch that must not be silently switched onto.

* test(#2539): assert non-silent warning; hoist execFileSync; normalizePhaseName guard

Address orthogonal-review findings on the #2539 fix:

- Spec AC2 ('an auto-checkout mid-commit must never happen silently'): the
  no-switch path now writes a 'Warning: resolved phase branch X already
  exists; committing on Y instead' line to stderr when checkout -b fails
  because the branch already exists. The regression test captures stderr via
  spawnSync and asserts the warning, so neither direction of the branching
  resolution is silent.
- Spec AC3 ('reuse normalizePhaseName/extractPhaseToken/stripProjectCodePrefix'):
  the token-acceptance guard now runs the candidate token through
  normalizePhaseName and accepts it only when it normalizes to a numeric phase
  form, rather than the brittle 'token !== phaseDir && /\d/.test(token)' check
  that leaned on extractPhaseToken's undocumented dirName fallback.
- Standards (Duplicated Code): hoist execFileSync/spawnSync requires to the top
  of the 'commit command' describe block instead of inlining them per test.

* fix(#2539): build phase-token shape from PHASE_NUMBER_TOKEN_SOURCE (#2128 guard)

The acceptance guard regex in detectPhaseNumberFromFiles was a hardcoded
`/^\d+[A-Z]?(?:\.\d+)*$/i` — a literal re-derivation of the canonical
phase-number grammar, which the #2128 phase-id drift guard
(tests/phase-id-drift-guard.test.cjs) rejects unless sanctioned with a
`// phase-id-owner:` marker. Build it from the single-owner
PHASE_NUMBER_TOKEN_SOURCE export instead, so this read-side acceptance check
cannot drift from every other phase-token reader. gsd-test reported this as
2 failures (linux-node22 + linux-node24) on the prior commit.

* docs(#2539): backfill changeset pr: 2669
2026-07-26 14:14:28 -04:00
Tom Boucher
a5180d96a3 fix: add regression-test-presence gate + missing tests for #2429/#2279 (#2563)
lint-fix-has-regression-test.cjs: new gate that fails if a fix(#NNNN)
or feat(#NNNN) commit has zero behavioral test files (*.test.cjs,
excluding auto-generated fixtures/baselines) in its diff. Wired into
lint:ci so it runs before PR creation.

Missing regression tests added:
- #2429: codex local scope does not set $HOME/.agents skills home;
  global scope does (tests/runtime-artifact-layout.test.cjs)
- #2279: map-codebase instructions say to overwrite existing dates,
  not just replace [YYYY-MM-DD] placeholders (tests/commands.test.cjs)
2026-07-23 12:13:12 -04:00
Tom Boucher
200daa456c fix(#2269): add --files to three unscoped query commit call sites (#2549)
* test(#2269): regression test for query commit --files scoping

Verify secure-phase.md, validate-phase.md, and next.md all pass --files
to their query commit calls.

* fix(#2269): add --files to three unscoped query commit call sites

secure-phase.md, validate-phase.md, and next.md were the only 3 of 65
query commit call sites that omitted --files, causing blanket staging
of .planning/ and committing unrelated files. Add --files with the
specific artifact path to each.

Closes #2269

* docs(#2269): backfill changeset PR number (2549)
2026-07-23 07:37:15 -04:00
Tom Boucher
be5113abfe fix(#2408): fold colliding phase statuses + add W023 collision warning (#2461)
* fix(#2408): fold colliding phase statuses + add W023 collision warning

Two coupled bugs from #2408:

1. cmdStats last-write-wins status (src/commands.cts:1610-1618): when two
   on-disk phase directories normalize to the same phase key (e.g.
   `05-real/` and `05-real-stray/`), the directory-scan merge overwrote
   `status` with whatever the *current* directory in scan order computed,
   discarding `existing?.status` entirely. fs.readdirSync order is not
   stable across platforms, so /gsd-stats could silently report `Not
   Started` for a phase that is actually `Complete`. Plan/summary counts
   were already additively merged; only `status` was wrong.

   Fix: introduced a `foldPhaseStatus(a, b)` helper that returns whichever
   status is further along the precedence ladder
   `Complete > Needs Review > Executed > In Progress > Planned > Not Started`
   (with `Pending` and unrecognized statuses ranked after). The merge site
   now calls `existing ? foldPhaseStatus(existing.status, status) : status`.
   The fold is commutative, so the result is identical regardless of read
   order.

2. cmdValidateHealth had no collision-detection pass (src/verify.cts):
   codes W001-W022 cover every condition except normalized-key collisions,
   so an operator got zero signal that anything was wrong.

   Fix: added W023 — groups phaseDirEntries by their normalized phase key
   (via the existing normalizePhaseName + extractPhaseToken helpers from
   phase-id.cjs — same normalization cmdStats uses) and emits a warning
   for any group with ≥2 dirs. The warning names the normalized key, both
   directory names (sorted by comparePhaseNum for stable output), and each
   directory's independently-computed status (via determinePhaseStatus
   imported from commands.cjs). Wording is deliberately neutral — never
   guesses which directory is the real one. The optional --repair path
   from the issue is intentionally NOT implemented in this PR (the issue
   marked it lower priority and acceptance criterion 4 is vacuously
   satisfied by omission).

Triage correction applied: the issue proposed W022, but that code is
already in use for config.json model-tier validation (src/verify.cts
:1372-1391). The next free code is W023, used here.

Tests:
- tests/commands.test.cjs: integration test that 05-real/ (Complete) +
  05-real-stray/ (empty/Not Started) collide and stats reports the merged
  phase as Complete regardless of read order; plus a direct unit test of
  foldPhaseStatus asserting commutativity + correct precedence for every
  status pair + correct handling of unrecognized statuses.
- tests/health-validation.test.cjs: integration test that W023 fires on
  the collision naming both dirs + their statuses (and uses neutral
  wording), plus a negative test that no W023 fires when only one dir
  exists per key.

References: #2408; reporter's three-layer triage + acceptance criteria;
triage correction that W022 is already in use (model-tier validation).

* chore(#2408): backfill pr:2461 in .changeset/graceful-koalas-forage.md
2026-07-20 15:49:57 -04:00
Tom Boucher
81f7ab4df1 refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4)

Cutover all 55 remaining case arms from runCommand's switch to
HOST_COMMAND_ROUTERS. runCommand now contains only its default case
(~40 lines): the three-layer dispatch (capability → overlay → host table)
plus the unknown-command diagnostic. The 73-case switch is dissolved.

Each case body was relocated verbatim to a module-scope route*Command
function via a brace-matching extractor; inner break; statements (from
_dispatchNonFamily early-exit patterns) were converted to return;
(5 arms affected); loop break; statements preserved.

Closes #2384

* chore: retrigger CI
2026-07-17 18:29:28 -04:00
Tom Boucher
b0f672f88c fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity.

Closes #2337. Admin-merged (self-review bypass) with full green CI.
2026-07-17 14:01:38 -04:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
85ed50cc4f test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file
that owns each subject-under-test, across 52 existing suites (state, config, frontmatter,
roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard,
health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe
wrappers; 881 subtests conserved 1:1. No new test files.

Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations
(intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe.

Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across
8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule
exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md +
ADR-0002/443/1235/3524 test-file references. lint:ci green.

Part of epic #1969. Closes #1972.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 08:59:23 -04:00
Enes Yağız
8748e95ed1 fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666)

Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:

- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
  changes, pushes, and opens a companion PR via `gh pr create`

All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.

Closes #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr number to 667

* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness

Resolves all three blockers and seven robustness issues raised in PR #667 review:

Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
  never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)

Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix

Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
  containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected

Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
  instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
  combinations and the execGit global-trim edge case uniformly

Tests: 17/17 pass, lint: 0 errors

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false

- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
  standalone bug-666-*.test.cjs into tests/commands.test.cjs under
  describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
  Adds allow-test-rule: source-text-is-the-product see #666 for the
  workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
  filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.

* fix(pr-branch): remove obsolete regression tests for sub-repos handling

* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout

- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
  (+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
  call in the handle_sub_repos dirty-scan (repo convention: every git
  subprocess is bounded, never hangs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): handle rename staging and split changedFiles from filesToStage

For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): rollback on push failure in cmdPrSubrepo

If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): do not rollback after commit on push failure; add push-fail regression test

Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.

Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): validate sub-repo paths before git invocation in pr-branch.md

The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.

Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.

Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.

Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#666): make sub-repo traversal scan test genuinely fail-first

The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md

Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.

- dirty-scan: realpathSync the root once, and realpathSync each candidate before
  the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
  containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
  against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
  backend must still be reported). Confirmed fails-first — regressing the scan to
  path.resolve makes the symlink case leak.

Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases

Pre-emptive hardening of the workflow changes from the symlink fix:

- continue-outside-loop: the "skip companion PR on seam failure" block used a
  bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
  loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
  prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
  (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
  creation lacks privileges) so it doesn't hard-fail on Windows CI.

Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:34:00 -04:00
Tom Boucher
3ebd13c45d fix(#1145): implement query user-story.validate handler (#1193)
* fix(#1145): implement query user-story.validate handler

`query user-story.validate` was a phantom command invoked by mvp-phase.md
(line 102) and verify-work.md (line 170) but had no CJS handler. Every
call exited with "Unknown command: user-story". The dotted-form dispatcher
strips the `query` prefix, splits on `.`, yielding command='user-story'
which fell through to the `default:` case with no registered capability
handler.

Adds a `case 'user-story':` handler inline in gsd-tools.cjs (same pattern
as the #1140 fix in PR #1148). The handler:

- Validates "As a [role], I want to [capability], so that [outcome]."
- Uses \S anchors to require non-whitespace content in each slot
  (whitespace-only slots like "As a  , I want to ..." now correctly
  return valid:false — found by adversarial Codex review)
- Returns { valid: boolean, errors: string[], slots: {role, capability,
  outcome} | null }
- Supports --pick valid for bare boolean output (verify-work.md usage)
- Added 'user-story' to SKIP_ROOT_RESOLUTION (pure string validation)
- Added 'user-story' to TOP_LEVEL_USAGE command list

Regression tests added to tests/commands.test.cjs (per lint-regression-
test-names policy; new bug-NNNN standalone files are banned).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1145 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 23:34:44 -04:00
Tom Boucher
1d90ad3c30 fix(commands): route nested sub_repos by longest prefix, not array order (#1130)
groupFilesBySubrepo selected the first sub_repos entry in array order
whose prefix matched a file, so a file under a more-specific nested
sub-repo (e.g. packages/core/widget.js with sub_repos
["packages", "packages/core"]) was mis-routed to the less-specific
parent ("packages").

Select the longest (most-specific) matching prefix within each
first-segment bucket instead, making routing independent of sub_repos
array order. String()-guard the length comparison so non-string entries
still never throw (preserves the #311 tolerance). Update the stale doc
comment that claimed first-match semantics.

Closes #391.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 17:33:26 -04:00
Tom Boucher
56a815c722 fix(#628): read snake_case requirements_completed in summary-extract (#640)
* fix(#628): read snake_case requirements_completed in summary-extract

cmdSummaryExtract read only the kebab frontmatter key
`requirements-completed`, but the tool's own JSON output key and the
milestone-audit `--pick` both use the snake form `requirements_completed`.
A SUMMARY written in the snake form the tool itself emits was silently
read back as [] — a false negative in milestone requirement traceability
with no diagnostic.

Make the reader tolerant of both key forms (kebab takes precedence), so a
round-tripped field is no longer dropped. extractFrontmatter does no
hyphen<->underscore normalization, so distinct keys had to be read
explicitly.

Adds regression tests: snake-only fixture now returns the IDs, and a
both-forms-present fixture asserts kebab precedence.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#628): add changeset for summary-extract snake-key fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 08:06:31 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
c7e5a88353 enh(#466): refresh opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5) (#467)
* enh: bump opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#466): changeset for opus-tier model-ID refresh

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 11:52:47 -04:00
Tom Boucher
5ca646f015 feat(#443): unified cross-provider effort controls + fast-mode-aware routing (#463)
* test(#443): RED unified effort + fast_mode + resolve-execution

All 68 tests failing as expected — no implementation yet.
Covers: effort cascade (tier defaults, overrides, invalid fallthrough),
fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier
escalation, renderEffortForRuntime clamping, resolve-execution CLI,
config schema new keys, QA hostile-input matrix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query

Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max)
and fast_mode propagation knobs, with per-runtime rendering that clamps the unique
tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps
to low on Claude).

Key changes:
- config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys;
  add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides,
  fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment
- config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults
- model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE
- model-profiles.cjs: re-export new catalog exports
- core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier,
  VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig
- commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort;
  add cmdResolveExecution (superset command with effort_rendered, effort_param,
  effort_propagation, fast_mode, fast_mode_supported)
- gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags
- tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix
- tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort
- docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections
- settings-advanced.md: list new effort/fast_mode keys in confirmation table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime

- Remove resolveReasoningEffortInternal (catalog-driven effort function) from
  core.cjs and its export; remove from commands.cjs destructure import
- Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort
  assertions now use resolveEffortInternal + renderEffortForRuntime; Claude
  effort is first-class (output_config.effort); unknown runtimes assert param===null
- Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire
  resolveReasoningEffortInternal describe with unified effort assertions;
  effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#443): ADR for unified cross-provider effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* test(#443): architecture-level QA invariants + test-strategy doc

Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs)
covering 8 architectural invariants: cross-provider validity (never emit a value
the real API would 400 on), param/channel contract stability, resolve-execution
JSON contract (all 8 keys + correct types), totality across the full 33-agent
registry, fast-mode honesty (claude always fast_mode_supported=false), precedence
first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing
composition (effort escalation independent of model tier), and config-set round-trip
for all new effort/* and fast_mode/* key namespaces. Append test-strategy section
with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#443): add failing install-wiring tests for effort per-runtime injection (RED)

TDD RED: 10 failing tests covering:
- Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high)
- Gemini .md does NOT get effort: (already passing — Gemini-safe)
- Codex .toml gets model_reasoning_effort via unified resolver
- Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml
- Source purity: agents/*.md have no effort: key (already passing)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified)

- Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs
- Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from
  .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback),
  same probe pattern as readGsdRuntimeProfileResolver
- Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching
  resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high')
  without loadConfig side-effects (no sub-repo detection, no migration writes)
- Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for
  runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay
  effort-free (Gemini-safe source contract preserved in agents/*.md)
- generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from
  unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps
  max → xhigh via renderEffortForRuntime('codex', ...)
- installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to
  generateCodexAgentToml so per-project config wins for Codex .toml too
- Update failing tests to GREEN: 12/12 pass; all 17 install tests pass;
  2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): source install effort defaults from manifest (kill drift) + guard test

Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in
resolveInstallTimeEffort with values read from config-defaults.manifest.json at
module init, using the same __dirname-relative path install.js already uses for
all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert
equality between install.js's runtime constants and the manifest on every CI run.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): reconcile Codex TOML tests with unified effort design

The #443 unified effort resolver makes generateCodexAgentToml always emit
model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides).
The test 'generated TOML omits reasoning_effort when runtime has none' had an
obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses
unified effort. Convert it to assert the new invariant: Codex TOML always carries a
valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner,
a heavy-tier agent), while model_profile_overrides model override is still respected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity)

Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read)
and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog()
getter that initialises on first call from resolveInstallTimeEffort /
generateCodexAgentToml / Claude .md effort injection.  Requiring install.js in
unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers
manifest IO or throws, eliminating the load-time side effect that changed
subprocess exit codes / stderr on the bench.

Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT)
preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift
still validates them without forcing eager load.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity

runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition
to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR,
GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses
os.homedir() directly — including the ~/.cache/gsd update-check deletion,
~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes —
never touches the real \$HOME during the test.

Without the HOME isolation the install test (which is new to this branch and
is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or
delete files under the real \$HOME, causing runtime-launcher-parity test (D)
to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is
absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback
arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists.

Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess
that the global installer spawns — irrelevant to effort-wiring assertions,
slow, and potentially writes to ~/.npm cache.

All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit
suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#443): add changeset fragment for effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak

Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main
install block (guarded by !GSD_TEST_MODE), performing a real global Claude
install into $HOME/.claude/. On CI ubuntu where node is on standard PATH,
the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing
runtime-launcher-parity test (D) to exit zero when it must exit non-zero.

Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs
alphabetically before runtime-launcher-parity.test.cjs in the same node
--test invocation. Each runs in a separate worker process but shares the
same HOME. The drift test's install leaks gsd-tools.cjs into that HOME,
then the launcher test's bash subprocess finds it via the $HOME/.claude arm.

Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard
test, before the require(installPath) call. This matches the pattern used
by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring
.install.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings)

Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent.
Replace find(non-dash) with a proper flag-consuming loop that collects a single
positional; validate missing/extra positionals and malformed --attempt values.

Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra")
verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from
core.cjs) before accepting a value, mirroring resolveEffortInternal exactly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions

Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects
EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>'
before the closing '---' delimiter using the same EOL as the surrounding
frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the
LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on
Windows (git core.autocrlf=true checkout).

Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and
complex frontmatter cases. Exports injectEffortFrontmatter from module.exports.

Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs
on windows-latest runners (lines 138, 145, 152, 261, 345, 356).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 10:32:58 -04:00
Tom Boucher
bd98e568f6 fix(#308): bound websearch fetch with timeout and retry (#387)
cmdWebsearch called fetch() with no timeout and no retry, so a hung
connection blocked indefinitely and transient 429/5xx/network failures
were not recovered. Add AbortSignal.timeout (configurable via
GSD_WEBSEARCH_TIMEOUT_MS, default 10s) and a bounded retry loop
(max 2 retries, exponential backoff + jitter) for 429/5xx/network
errors, honoring Retry-After on 429 (capped at 60s). Non-429 4xx fail
immediately (no wasted retries). Transient-exhausted failures report an
`attempts` count. Worst-case time is bounded by timeout*(1+retries)+backoff.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 21:32:06 -04:00
Tom Boucher
25d24219bf perf(#311): index subrepo routing by first path segment (#390)
cmdCommitToSubrepo routed each changed file to a sub-repo via
subRepos.find(...) inside the file loop — O(files * repos). Extract a pure
groupFilesBySubrepo() that buckets sub-repos by first path segment and scans
only the matching bucket, dropping it to expected O(files + repos). First-
match-in-array-order semantics (incl. multi-segment sub-repos) are preserved
exactly. Adds the first behavior-lock test for the routing path.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:29 -04:00
Tom Boucher
bd8b9eaf9d fix(#344): harden windows test assertions and pin windows-2025 labels (#349) 2026-05-26 16:14:15 -04:00
Tom Boucher
7d54416cc3 fix(ci): keep current-timestamp on CJS path to avoid Windows bridge crash 2026-05-24 22:45:46 -04:00
Tom Boucher
d4d4178603 Merge pull request #3483 from radioflyer28/feat/agent-launch-reasoning-transport-3474
feat: transport resolved reasoning effort to agent launches
2026-05-14 20:01:32 -04:00
Tom Boucher
c21bbf116e fix(tests): use array form for space-containing paths in CLI invocations
On machines where os.tmpdir() returns a path with a space (e.g.
/Volumes/Mini Me/tmp), runGsdTools() string args were whitespace-split by
the helper tokeniser, truncating paths at the first space.  Switch all
calls that embed a dynamic path into the argument list to the array form
of runGsdTools() so execFileSync receives each path as a single argv slot.

Fixes #3509

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 12:43:48 -04:00
radioflyer28
810778e137 fix(sdk): gate reasoning effort by runtime allowlist 2026-05-13 22:27:15 -04:00
radioflyer28
f70829d067 feat: transport resolved reasoning effort 2026-05-13 19:43:28 -04:00
Tom Boucher
96806003c5 fix(#3229): shared model catalog source of truth for agent profiles + runtime tier defaults (#3230)
* docs(adr): add ADR-0003 model catalog module

* fix(#3229): add shared model catalog as source of truth for agent profiles and runtime tier defaults

Research / design (ADR-0003):
- Existing drift came from 4 independent model truths:
  1. CJS model-profiles.cjs
  2. SDK config-query.ts stale copy (18 agents)
  3. settings-advanced.md runtime tier table
  4. session-runner Claude-only profile map
- New design: one machine-readable Model Catalog Module in sdk/shared/
  that both packages ship and consume.

Implementation:
- sdk/shared/model-catalog.json — canonical source of truth for:
  - full 33-agent registry
  - per-agent golden (quality) alias + balanced/budget aliases
  - adaptive derivation from routingTier
  - agent→phaseType map
  - agent→dynamic-routing default tier map
  - runtime tier defaults for all supported runtimes
- get-shit-done/bin/lib/model-catalog.cjs — CJS adapter over the catalog
- sdk/src/model-catalog.ts — SDK adapter over the same catalog
- CJS model-profiles.cjs now re-exports derived data from model-catalog.cjs
- SDK config-query.ts now re-exports MODEL_PROFILES/VALID_PROFILES from
  model-catalog.ts instead of maintaining its own list
- sdk/src/query/helpers.ts runtime list now comes from the catalog (fixes hermes drift)
- sdk/src/session-runner.ts Claude profile→model-id mapping now resolves via catalog
- docs/CONFIGURATION.md + settings-advanced.md runtime tables updated to match catalog

Behavior changes:
- resolve-model now covers every shipped agent file on disk (33 agents)
- unknown-agent fallback is profile-semantic, not hardcoded sonnet:
  quality→opus, budget→haiku, balanced/adaptive→sonnet, inherit→inherit
- Group B runtimes remain known runtimes but do not get built-in tier defaults

Tests (RED→GREEN):
- root tests: shipped agent files must equal MODEL_PROFILES keys
- sdk tests: shipped agent files must equal MODEL_PROFILES keys
- direct fix assertion: gsd-code-reviewer resolves to opus under quality with no unknown_agent
- runtime defaults parity test: settings-advanced.md + CONFIGURATION.md tables must match catalog
- helper tests: hermes included in SUPPORTED_RUNTIMES and getRuntimeConfigDir()

Closes #3229

* chore(changeset): update #3229 changeset pr field to 3230

* fix(ci): update inherit fallback expectations and inventory parity for model catalog
2026-05-08 21:25:37 -04:00
Tom Boucher
918f987a19 feat(#2982): extend no-source-grep lint to catch var-binding readFileSync.includes() (#2985)
* feat(#2982): extend no-source-grep lint to catch var-binding readFileSync.includes()

The base lint (scripts/lint-no-source-grep.cjs) only catches
readFileSync(...).<text-method>() chained directly. The much more
common var-binding form escapes it:

  const src = fs.readFileSync(p, 'utf8');
  // 50 lines later
  if (src.includes('foo')) {}        // ← still grep, lint missed it

Scan of the test suite found ~141 files using this pattern.

Implementation built TDD per #2982 with structured-IR assertions:

  scripts/lint-no-source-grep-extras.cjs
    - detectVarBindingViolations(src) — pure detector, two passes:
      pass 1 collects vars bound from readFileSync, pass 2 finds any
      <var>.<includes|startsWith|endsWith|match|search>( on those vars.
    - detectWrappedAssertOkMatch(src) — flags
      assert.ok(<expr>.match(...)) which escapes the assert.match rule.
    - VIOLATION enum exposes stable codes for tests to assert on.

  scripts/lint-no-source-grep.cjs
    - Wires the new detectors into the existing per-file check; one
      additional violation row per file with the first 3 sample tokens.

  tests/bug-2982-lint-var-binding.test.cjs
    - 13 tests, all assertions on typed VIOLATION enum / structured
      records. Covers all 5 text-match methods, multi-var, no-bind,
      string literal (must NOT trigger), wrapped assert.ok(.match),
      and assert.match (must NOT double-flag).

Migration backlog (#2974 expanded scope):

  - 42 files annotated `// allow-test-rule: source-text-is-the-product`
    (legitimate — they read .md/.json/.yml files whose deployed text
    IS the product)
  - 3 files annotated `// allow-test-rule: pending-migration-to-typed-ir [#2974]`
    (read .cjs/.js source — clear migration debt)
  - 95 files annotated `pending-migration-to-typed-ir [#2974]` with
    `Per-file review may reclassify as source-text-is-the-product
    during migration` (mixed — manual review under #2974)

After this lands the lint reports 0 violations on main; new
violations in PRs surface immediately.

Closes #2982
Refs #2974

* test(#2982): fix truncated test name per CR

The label ended with a bare '(' from a copy-paste mishap. Now reads
'does NOT flag .matchAll(...) — matchAll is not match, so
assert.ok(.matchAll(...)) is not flagged'.

* chore(#2982): add changeset fragment for PR #2985

* chore(#2982): add changeset fragment for PR #2985
2026-05-01 19:50:10 -04:00
Tom Boucher
7b0a8b6237 fix: normalize phase numbers in stats Map to prevent duplicate rows (#2220)
When ROADMAP.md uses unpadded phase numbers (e.g. "Phase 1:") and
the phases/ directory uses zero-padded names (e.g. "01-auth"), the
phasesByNumber Map held two separate entries — one keyed "1" from the
ROADMAP heading scan and one keyed "01" from the directory scan —
doubling phases_total in /gsd-stats output.

Apply normalizePhaseName() to all Map keys in both the ROADMAP heading
scan and the directory scan so the two code paths always produce the
same canonical key and merge into a single entry.

Closes #2195

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-15 14:59:35 -04:00
Tom Boucher
2703422be8 refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md (#1675)
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md

- Replace require('node:assert') with require('node:assert/strict') across
  all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
  and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
  as that is appropriate for per-function resource guards, not lifecycle hooks

Closes #1674

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(tests): add --test-concurrency=4 to test runner for parallel file execution

Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): allowlist verify.test.cjs in prompt-injection scanner

tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:29:03 -04:00
Tibsfox
9d90c7b420 feat(hooks): add check-commit guard for commit_docs enforcement (#1395)
Add check-commit command to gsd-tools that acts as a pre-commit guard.
When commit_docs is false, rejects commits that stage .planning/ files
with an actionable error message including the unstage command.

Recreated cleanly on current main — previous version carried stale
shared fixes that are now upstream.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 16:11:00 -07:00
Tom Boucher
4afc5969ef Merge pull request #1477 from Tibsfox/fix/stats-verification-status-1459
fix(stats): require verification for Complete status, add Executed state
2026-04-01 17:26:49 -04:00
Tibsfox
3321d48279 fix(stats): require verification for Complete status, add Executed state
Phases with all summaries but no passing VERIFICATION.md now show as
"Executed" instead of "Complete", preventing false progress reporting.

Adds determinePhaseStatus() helper used by both cmdStats() and
cmdProgressRender(). Also fixes duplicate phase directory accumulation
in cmdStats() — plans/summaries from directories sharing the same
phase number are now summed instead of silently overwritten.

New statuses: Executed (summaries done, no verification), Needs Review
(verification exists with human_needed status).

Closes #1459

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 12:18:03 -07:00
Tibsfox
c1fd72f81f fix(branching): capture decimal phase numbers in commit regex
The phase-extraction regex /(\d+)-/ only matched the last integer
segment before a dash, so decimal phases like 45.14 were misresolved
to phase 14 — silently switching to the wrong branch.

Closes #1402

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 19:42:19 -07:00
Tom Boucher
9ddd6c1bdc fix: create strategy branch before first commit, not at execute-phase (#1278)
When branching_strategy is "phase" or "milestone", the branch was only
created during execute-phase — but discuss-phase, plan-phase, and
new-milestone all commit artifacts before that, landing them on main.

Move branch creation into cmdCommit() so the strategy branch is created
at the first commit point in any workflow. execute-phase's existing
handle_branching step becomes a harmless no-op (checkout existing branch).

Also fixes websearch test mocks broken by #1276 (fs.writeSync change).

Closes #1278

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 01:08:06 -04:00
Tom Boucher
02254db611 fix: resolve ProviderModelNotFoundError on non-Claude runtimes (#1156)
Extend resolve_model_ids to accept "omit" value: returns empty string
so non-Claude runtimes (OpenCode, Codex, Gemini, etc.) use their
configured default model instead of unresolvable Claude aliases.

- resolve_model_ids: "omit" short-circuits before alias resolution
- model_overrides still respected (checked first) for explicit IDs
- Installer sets resolve_model_ids: "omit" in ~/.gsd/defaults.json
  for non-Claude runtimes during install
- 4 new tests covering omit behavior and override passthrough
- Fix websearch test mocks for fs.writeSync output change (#1276)

Closes #1156

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 00:57:20 -04:00
j2h4u
104a39e573 fix(lifecycle): extract one-liner from summary body when not in frontmatter
The summary template puts the one-liner as a `**bold**` line after the
`# Phase N` heading, but `cmdSummaryExtract` and `cmdMilestoneComplete`
only checked frontmatter `one-liner` field — which is often empty.

Adds `extractOneLinerFromBody()` to core.cjs as a fallback that parses
the first `**...**` line after the heading.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-19 21:29:25 +05:00
Tom Boucher
e84999663c feat(discuss-phase): cross-reference pending todos against phase scope
Add a cross_reference_todos step to discuss-phase that surfaces relevant
backlog items before scope-setting decisions are made.

Implementation:
- New 'todo match-phase <N>' CLI command (commands.cjs) that scores
  pending todos against a phase's ROADMAP goal using three heuristics:
  keyword overlap, area match, and file path overlap
- New cross_reference_todos step in discuss-phase.md between
  load_prior_context and scout_codebase
- CONTEXT.md template gains 'Folded Todos' subsection in <decisions>
  and 'Reviewed Todos (not folded)' subsection in <deferred>

Design:
- No AI call for matching — pure keyword/area/file heuristics for speed
- Silent skip when todo_count is 0 or no matches (no workflow slowdown)
- Auto mode folds all todos with score >= 0.4 automatically
- Scoring: keywords (up to 0.6), area match (0.3), file overlap (0.4)

Tests: 5 new tests covering empty state, keyword matching, unrelated
todo exclusion, area matching, and score sorting.

Closes #1111
2026-03-18 12:25:04 -04:00
Colin Johnson
0fde04f561 fix(stats): correct git and roadmap reporting (#1066) 2026-03-15 19:31:28 -06:00
ashanuoc
944df19926 feat: add /gsd:stats command for project statistics
Add a new `/gsd:stats` command that displays comprehensive project
statistics including phase progress, plan execution metrics, requirements
completion, git history, and timeline information.

- New command definition: commands/gsd/stats.md
- New workflow: get-shit-done/workflows/stats.md
- CLI handler: `gsd-tools stats [json|table]`
- Stats function in commands.cjs with JSON and table output formats
- 5 new tests covering empty project, phase/plan counting, requirements
  counting, last activity, and table format rendering

Co-authored-by: ashanuoc <ashanuoc@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 21:58:09 -06:00
Ethan Hurst
c9e73e9c94 fix(quick-1): remove HIGH severity test overfitting in 4 test files
- frontmatter.test.cjs: delete FRONTMATTER_SCHEMAS describe block (4 tests asserting on lookup table shape)
- core.test.cjs: replace 12 per-profile per-agent matrix tests with 1 structural validation test
- commands.test.cjs: URL test uses URL.searchParams.get() instead of raw string includes
- state.test.cjs: date assertion uses before/after window to handle midnight boundary safely
2026-02-26 05:49:31 +10:00
Ethan Hurst
c027fe6542 test(07-02): add cmdCommit and cmdWebsearch tests for commands.cjs
cmdCommit tests (5 cases) with real git repos via createTempGitProject:
- commit_docs=false skip, .planning gitignored skip, nothing-to-commit
- Successful commit with hash and git log verification
- Amend mode with commit count verification

cmdWebsearch tests (6 cases) with fetch mocking and stdout interception:
- No API key returns available=false, no query returns error
- Successful response with correct result shape
- URL parameter construction verification
- API error (non-200) handling, network failure handling

Covers requirements CMD-04, CMD-05.
2026-02-26 05:47:33 +10:00
Ethan Hurst
d4ec9c69a9 test(07-01): add utility function tests for commands.cjs
Add 14 new tests across 5 describe blocks:
- cmdGenerateSlug: normal text, special chars, numbers, hyphens, empty input
- cmdCurrentTimestamp: date, filename, full, default format
- cmdListTodos: empty dir, multiple, area filter match/miss, malformed
- cmdVerifyPathExists: file, directory, missing, absolute, no-arg
- cmdResolveModel: known agent, unknown agent, default profile, no-arg

Covers requirements CMD-01, CMD-02, CMD-03.
2026-02-26 05:47:33 +10:00
TÂCHES
0176a3ec5b fix: map requirements-completed frontmatter in summary-extract (#627) (#716)
Add requirements_completed field to summary-extract output, mapping
from SUMMARY frontmatter requirements-completed key. Enables
/gsd:audit-milestone cross-check to receive data from SUMMARY source.

Re-applied from #631 against refactored codebase (commands.cjs + split tests).

Co-authored-by: Colin Johnson <Solvely-Colin@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:55:22 -06:00
TÂCHES
3704829aa6 fix: clamp progress bar percent to prevent RangeError crash (#715)
When orphaned SUMMARY.md files cause totalSummaries > totalPlans, the
progress percentage exceeds 100%, making String.repeat() throw RangeError
on negative arguments. Clamp percent to Math.min(100, ...) at all three
computation sites (state, commands, roadmap).

Closes #633

Co-authored-by: vinicius-tersi <vinicius-tersi@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:51:28 -06:00
Tyler Satre
fa2e156887 refactor: split gsd-tools.test.cjs into domain test files
Move 81 tests (18 describe blocks) from single monolithic test file
into 7 domain-specific test files under tests/ with shared helpers.
Test parity verified: 81/81 pass before and after split.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 10:30:54 -05:00