Phase 1 of parent #1641. Implements the contract documented in the
Phase 0 ADR-0174 §5 amendment (#1642 / #1643).
src/command-routing-hub.cts
* InvalidArgsResult interface gains optional exitReason?: string
(carries an ERROR_REASON enum value, separate from reason which is
the explanation text).
* makeInvalidArgs(arg, reason, exitReason?) factory conditionally adds
the field only when the third arg is truthy — preserves the strict-
keys invariant tested at command-routing-hub.test.cjs:444.
* _VARIANT_SCHEMA.InvalidArgs.allowed Set extended to include
'exitReason' so the runtime validator does not coerce well-formed
extended Results to HandlerFailure.
src/cjs-command-router-adapter.cts
* Honestified the wrapper comment: the runtime check ('ok' in result)
already passes any {ok:*} object through, so the historical
{ok:true, data} return type was a lie for err Results. The lying
cast is preserved because the Hub's export = syntax doesn't expose
HubResult for import; the Hub's _validateErrResult runtime-validates
the actual shape.
* Result→error() translation branched: when InvalidArgs carries
exitReason, the adapter calls error(result.reason, result.exitReason)
so the JSON-error envelope (GSD_JSON_ERRORS=1) preserves the typed
ERROR_REASON value. When exitReason is absent, error(msg) is called
with exactly one arg — byte-identical with prior behavior.
* RouteCjsCommandFamilyOptions.error and RouteHubCommandFamilyOptions
.error callback types widened from (message) to (message, reason?)
to match io.cts's actual error() signature.
CONTEXT.md
* Command Routing Hub predicate updated to document the new field,
factory signature, and dispatcher translation contract.
Tests (TDD red→green)
* tests/command-routing-hub.test.cjs: 8 new tests covering 2-arg
(strict-keys), 3-arg (key present), undefined, empty string, frozen
result, hub.dispatch propagation, and validator acceptance.
* tests/cjs-command-router-adapter.test.cjs: 2 new tests covering
exitReason passed as second arg + byte-identical prior behavior when
absent.
Verification
* npm run test:unit: 2448 tests, 0 fail (no regressions)
* gsd-test-summary on docker: outcome=passed, 0 failures
(RULESET.PR-FLOW.docker-before-push)
Memtrace blast radius: LOW (get_impact makeInvalidArgs → 3 nodes; the
optional field is non-breaking for the 1 existing caller routePhaseCommand).
* fix(#1634): honor capability hook matcher and node-prefix command
Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.
- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
.sh and others keep the bare quoted path (unchanged).
Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.
Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.
* chore(#1634): backfill changeset pr:1638
* fix(#1634): resolve lint and windows CI failures
- validator: replace the control-character range regex with a char-code loop.
The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
codes are equally precise and lint-clean. Behavior unchanged (still rejects
matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
honor POSIX write modes (a 0o644 write reads back as 0o666), so the
precondition is meaningless there and failed the windows-latest lane. The
node-prefix assertion — the actual fix — is platform-independent and still
runs everywhere.
* docs(#1634): amend ADR-894 for optional lifecycle hook matcher
The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.
* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md
Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers
The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.
Test-only; no production code touched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1178): note AGENTS_DIR export is for future call sites
Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1619): normalize pruned mise node execPath to the stable shim
resolveNodeRunner() bakes process.execPath into managed .js hook commands.
Node realpaths execPath, so under mise it resolves to a concrete
<data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after
which every managed hook 404s — the same ephemeral-path failure #977 fixed
for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise
versioned install path to the stable sibling shim <data>/shims/node when it
exists (deriving <data> from execPath so a custom MISE_DATA_DIR works),
falling back to the raw execPath otherwise. Tests folded into
install.test.cjs per the regression test-name lint.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(changeset): set pr number to 1621
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Joe Seymour <joese@iarx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).
- planner: add a Severity column to the STRIDE threat register; assign
severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
severity enum; redefine threats_open as the count of OPEN threats whose
severity is at or above block_on (none => 0). Below-threshold opens are
reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.
No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.
- New reference gsd-core/references/security-asvs-levels.md defines L1
(opportunistic), L2 (standard), L3 (comprehensive) for both planner
threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
by extracting the goal-backward worked example to planner-guidance.md.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':
1. Missing guards: workflow.security_block_on (enum) and
workflow.security_asvs_level (integer 1-3) had no store-time validation.
2. Systemic JSON-coercion bypass: every string-enum guard used
VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
before validation, String(["member"]) === "member" let a JSON array
slip through and an array was stored in a scalar key. Reproduced on
human_verify_mode, statusline.context_position, context_guard_mode,
fallow.scope/profile, source_grounding_authority, drift_action, context.
3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
including coerced arrays/objects and out-of-enum strings like
code_review_depth=garbage — was stored silently.
Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore.
Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact.
5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts.
Refs #1629 (Finding B; Finding A addressed in #1630).
PR #1622 (issue #1615) shipped Windsurf /gsd-* workflow wrappers that delegate to command bodies at <targetDir>/.windsurf/gsd-core/commands/gsd/X.md via a hardcoded @~/.claude/gsd-core/commands/gsd/ path. The path-rewrite pipeline correctly substitutes ~/.claude/ to the install target. But the source gsd-core/ dir does not ship with commands/ — the canonical command source lives at the package root (commands/gsd/). Without this copy, every /gsd-* workflow in Cascade references a file that does not exist. The slash commands appear in the / menu but silently fail when invoked because the LLM is told to read a missing file.
None of the original reviews caught this: not the security review, not Codex's adversarial orthogonal review (gpt-5.5/high), not Memtrace's graph-backed review. It was surfaced by a #1629 regression test that verifies 'every workflow @-reference target exists on disk after install' — the test failed, revealing the bug.
Fix: for Windsurf local installs, copy commands/gsd/*.md into <targetDir>/gsd-core/commands/gsd/ via copyWithPathReplacement (applies the same path+brand rewrites as the rest of the install). Guarded on isWindsurf && !isGlobal since global Windsurf workflow install is an explicit no-op.
Documented as DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED in CONTEXT.md so the pattern is locked in: any new converter emitting a wrapper that delegates to another file MUST verify the delegation target is actually installed.
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes.
For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target).
Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode.
Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615.
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body.
Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message).
Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation.
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only.
Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX.
Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input.
Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring.
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.
- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
block (extractFrontmatter can't — its `-` items are scalars-only; this is a
focused parser, sibling of parseMustHavesBlock), validates each entry, and
classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
AND non-empty all-`pass` verification AND zero validation errors. Everything
else — judgment, empty/failing verification, malformed entry — routes to the
human (fail-safe). A malformed block falls back to legacy prose extraction and
surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
human_judgment:true); verify-work extract_tests consumes it; create_uat_file
marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
for the null-entry/comment-header/mis-indent cases found in adversarial review.
Closes#1602
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* enhance(#1549): validate PR-title issue-ref convention at open time
The release changelog is title-driven: release.yml generates "What's Changed"
from PR titles, then format-github-release-notes.cjs buckets each line by its
conventional-commit prefix and relies on a `(#<issue>)` in the title to render
the issue link. Both rules were enforced only socially, so titles like
`fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag
defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke
the changelog, landing on the maintainer as release-time cleanup.
Extract the title matcher into one shared module consumed by BOTH the changelog
classifier and a new PR-title CI gate, so a title that passes the gate cannot
mis-bucket in the changelog (single source of truth).
- scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle +
the anchored regexes. One matcher, two consumers.
- scripts/release-notes/format-github-release-notes.cjs: classifyTitle now
delegates to classifyBucket (behavior preserved; existing tests green).
- .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on
pull_request opened/edited/reopened/synchronize, for ALL authors (the drift
came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout.
- tests/conventional-title.test.cjs (new): bucket + gate cases incl. the
leading-tag mis-bucket (backfills the untested classifyTitle case) and a
cross-check that the classifier delegates to the shared matcher.
- CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag.
Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk
* fix(#1549): check out the PR in pr-title-validator so the new matcher resolves
The workflow checked out the base branch (next) as a trusted policy source, but
the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR
and does not exist on next yet — so require() failed and validate-title errored
on its own introducing PR. Check out the PR's merge ref instead: the matcher
under review is present, the check is self-consistent, and a fork pull_request
runs read-only with no secrets, so running the PR's own pure-string regex is
safe.
* fix(#1549): move conventional-title.cjs out of installed scripts/lib/
bin/install.js bundles every file under scripts/lib/ into the user-installed
payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts
that exact set. The new matcher is release/CI tooling that must NOT ship to
users, so placing it in scripts/lib/ both broke the install manifest test and
would have shipped dead code. Relocate it next to its consumer in
scripts/release-notes/ (which the installer does not copy) and update the three
require paths (classifier, workflow, test) + the CONTRIBUTING reference.
install.test.cjs now 125/125; conventional-title + release-notes suites green;
lint:ci clean.
* fix(#1549): load title matcher from trusted base ref, not PR code
Addresses review (Solvely-Colin + trek-e): the gate checked out the PR
merge ref and require()'d evaluatePrTitle from PR-controlled code, so any
future PR could edit conventional-title.cjs to return { valid: true } and
wave its own malformed title through — a self-bypassable required check.
Load the matcher from a base-branch checkout instead (ref:
github.event.pull_request.base.ref), the same trusted-policy-source pattern
pr-target-validator.yml already uses. The PR can change its title but not
the ruler that measures it. An existsSync bootstrap guard skips the check
when the matcher isn't on the base branch yet (the introducing PR); every
PR after merge is fully gated. This keeps the single shared matcher (#1549's
whole point) rather than forking the regex into the workflow.
Also per review:
- add tests/conventional-title.property.test.cjs (fast-check): any
`type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket
are total functions (never throw).
- pin the `fix(#):` zero-digit boundary as missing-issue-ref.
Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.
Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).
Closes#1592
Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
Per review (test-standard items):
- Property test (RULESET.TESTS.property-based-testing): extractRetiredPhaseNumbers
is the parsing core, so add a fast-check property — k of n checklist phases
struck → exactly the k canonical keys returned, across randomized phase counts
and numeric/zero-padded/project-code ID forms. Exposed via a `_`-prefixed test
seam (mirrors the existing _setLockProbes seams), no public API surface added.
- Boundary (RULESET.TESTS.boundary-coverage): all-retired case (k === n) →
total_phases 0, via state json.
No production behavior change; the exclusion logic is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase B-provide of epic #1258. Adds a build-generated skills/ dir +
a skills manifest field so plugin-installed GSD exposes gsd-core:<skill>
the native Claude Code way. Closes the gap where plugin-only installs
lacked the skill surface because bin/install.js never ran.
- scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md
to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill
- .claude-plugin/plugin.json: add "skills": "./skills/"
- package.json: add skills to files, gen:plugin-skills to build chain
- tests/issue-766-plugin-manifest.test.cjs: Section H conformance
(manifest field + dir + frontmatter + count parity) + C2 skills symlink
- docs/adr/766-*.md: dated amendment adding skills surface row
- .changeset/rapid-bears-hum.md: type Added
- skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed)
Closes#1596
CodeQL alert #41 (js/incomplete-sanitization) flagged the partial
escape class /[-]/g at tests/worktree-safety.test.cjs:645 — it only
escaped hyphen-minus, leaving 13 other regex metacharacters (notably
backslash) unescaped. The canonical class /[.*+?^${}()|[\]\\]/g is
what every sibling escape in the test suite already uses
(bug-2839, bug-2760, 4-phase-complete, phase6-capstone-conformance).
Today dormant: the flag array is a hardcoded [a-z-] literal, so the
expanded class is a no-op for the four existing flags and the regexes
they produce are byte-identical. The fix prevents future drift — a
contributor adding e.g. '--output=file' would have silently introduced
a regex wildcard.
All 69 tests in the file pass. No user-facing behavior change.
Fixes#1589
#1531/#1532 replaced withPlanningLock's mtime-staleness with PID-liveness
(process.kill(pid,0)). perf-407 plants pid: process.pid + 1 and relied on the
old mtime model to force the retry/sleep path; under the new model that pid's
liveness is environment-dependent, so the retry path was taken on some runners
and skipped on others (sleepCallCount: 0 precondition failure) — flaky CI red
on next that blocks the merge queue.
Pin the planted holder live via the _setLockProbes seam that #1532 added, and
_resetLockProbes() in afterEach. The retry/sleep path is now exercised
deterministically on every runner. No assertion weakened; no variable renamed.
perf-316 is unaffected (its worker writes the parent's own, always-live pid).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The back-merge of next carried #1559/#1565 (audit installer compatibility
exports), which removed convertGeminiToolName from bin/install.js's
module.exports. This PR's regression test destructured it from
require('../bin/install.js'), so after the merge it was undefined →
"TypeError: convertGeminiToolName is not a function" across all test lanes.
Rewrite the regression to assert the user-visible behavior through the
still-exported convertClaudeToGeminiAgent: a Skill/SlashCommand/AskUserQuestion
tools entry must not appear in the emitted Gemini frontmatter (the lowercase
fallback would emit invalid tool names that abort agent load — #1394/#3362),
while mapped tools (Read→read_file, WebFetch→web_fetch) survive. Folds the
dropped AskUserQuestion/ask_user coverage into the behavior test and removes
the internal-function import, so the test no longer depends on a private export
#1559 intentionally pruned. The exclusion fix itself is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The copilot path correction (.github/copilot-instructions.md) lengthened
the new-project.md instruction-file prose by 16 bytes past the prior
baseline. Regenerated; growth is justified by the more accurate path.
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet (fold into the owning module or use the fix- prefix). Matches the
fix-1445 precedent for the same total_phases subsystem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A retired/folded phase (struck through in ROADMAP, marked [x], with a
directory but no completion artifact) was counted in the total_phases
denominator via max(phaseDirs.length, roadmapPhaseCount), yet could never
satisfy the numerator (no SUMMARY → never "completed"), freezing shipped
milestones below 100% (e.g. 5/6 = 83%).
Both STATE counting paths now read the current-milestone ROADMAP scope and
exclude retired phases from BOTH the disk phase-dir set and the heading
count, so a retired phase counts toward neither denominator nor numerator:
- buildStateFrontmatter (`state json`)
- cmdStateSync (`state sync --verify` / rebuild) — previously re-derived the
inflated denominator and reported "no drift", per the issue.
Retired detection (extractRetiredPhaseNumbers) is scoped to the lines that
canonically mark a phase retired — a checklist entry (`- [x] …`) or a phase
heading — and within those, only a struck span whose SUBJECT is the phase
(`~~**Phase 04: Delta**~~`). So struck prose, a struck goal line, and the fold
target ("folded into Phase 05") are not misread as retired.
Phase matching uses the canonical phase-id helpers (normalizePhaseName +
extractPhaseToken), so numeric, decimal, and project-code IDs (PROJ-42) match
consistently across ROADMAP tokens and on-disk dir names.
Scope boundaries (separate, pre-existing concerns left unchanged):
- `roadmap analyze` (src/roadmap.cts) intentionally trusts the [x] checkbox
(incl. externally-completed phases) — a different reporting surface.
- cmdStateSync does not apply the milestone phase-dir filter (so 999.x /
other-milestone dirs can still affect its count); that is the #1445 /
milestone-filter axis, independent of retired phases.
Same counting family as #549 / #500 / #1445.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit
Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on
- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
slug, status, scope, trigger_when, planted, title. Optional case-insensitive
status filter. User-controlled content is sanitized (sanitizeForDisplay) and
every path validated (requireSafePath); read-only. Independent of
audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
renders the seed table.
Closes#441
* chore(#441): point changeset fragment at PR #722
* test(#441): allowlist list-seeds test in prompt-injection scan
The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.
* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow
Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.
* docs(#441): sync help full.md + INVENTORY for --list-seeds
Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.
* docs(#441): add --list-seeds how-to + drop phantom statuses
Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):
- USER-GUIDE.md Seeds section (how-to): extend the task to cover
auditing parked seeds on demand via --list-seeds, including the
status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
from the list-seeds filter vocabulary; the system only produces
dormant|active|triggered (src/audit.cts scanSeeds). Reference must
be factually accurate and complete.
* fix(#441): guard non-scalar status frontmatter in cmdListSeeds
A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.
Adds regression coverage for empty and array `status:` and non-scalar fields.
Refs #441
* docs(#441): align list-seeds workflow status vocabulary
The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.
Refs #441
* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds
Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.
Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).
* test(#441): add fast-check property coverage and count=1 boundary for list-seeds
Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).
Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).
* chore(#441): sync runtime launcher snippet into list-seeds workflow
Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.
* test(#441): record list-seeds.md in workflow size baseline (#1074)
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* feat(#1318): require external reviewers to verify plan claims against source
/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.
Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): add changeset for reviewer source-grounding
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt
Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1318): harden build_prompt fence extraction to be fence-run-aware
Addresses maintainer review on PR #1421 (required-before-merge).
The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.
Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.
Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1383): resolve GSD version without a top-level require of the runtime-root package.json
The extracted runtime-artifact-conversion module sits in the gsd-tools
loader chain and did a module-load `require('../../../package.json')`.
On Codex (whose runtime root has no package.json) that threw
`Cannot find module '../../../package.json'`, crashing every gsd-tools
command before it did anything. Even on Claude the synthetic
`{"type":"commonjs"}` has no `version`, so the sole consumer already
emitted `version: undefined`.
Resolve the version lazily and defensively instead: read the installed
gsd-core/VERSION, else lazily require the runtime-root package.json, else
degrade to '' so the caller omits the field. Both sources are validated
against the repo's semver-prefix convention (mirrors update-context.cts)
so a garbled VERSION is never emitted verbatim. install.js's dead
duplicate converter is intentionally left untouched (scoped to the crash).
Adds a #1383 regression block exercising resolveVersionFrom across
VERSION-only / package.json-only / neither / malformed-VERSION layouts,
asserting no-throw and the correct version string.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1383): add changeset for the Codex gsd-tools crash fix
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1383): record resolveVersionFrom export in CONTEXT.md glossary
Maintainer review gate on PR #1409: the lazy resolveVersionFrom seam added on
the Runtime Artifact Conversion Module must be recorded in CONTEXT.md so the
canonical glossary doesn't drift from the exported surface.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1383): reword changeset to drop product-name parenthetical
product-name-purity (#1777) rejects 'Codex (…)' parentheticals that render
verbatim into CHANGELOG.md. Reword to a comma clause; no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1551): match dash-separated milestone phase IDs in roadmap analyze checklist scan
The checklist scanner in cmdRoadmapAnalyze allowed only a dot separator
(?:\.\d+)* while the detail-heading scanner allows [.-], so milestone-prefixed
IDs (1-01) truncated at the dash (-> 1) and reported phantom missing detail
sections on every well-formed milestone roadmap. Widen the char class to
(?:[.-]\d+)* to match the detail scanner and the shared phaseMarkdownRegexSource
helper.
Fixes#1551
Claude-Session: https://claude.ai/code/session_01H96MxPGMJJUiJLV2NgzV16
* chore(changeset): Fixed fragment for #1552 (roadmap milestone-id checklist scan)
Claude-Session: https://claude.ai/code/session_01H96MxPGMJJUiJLV2NgzV16
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
trek-e's review found the M1 PID-liveness backport dropped two pieces of
capability-lock.cts's steal-safety machinery, reopening the #500/#905/#1230
lost-update family:
- Empty-body window (state.cts): acquireStateLock creates the lock with O_EXCL
and writes the pid in a separate writeSync; a lock observed in that gap has an
empty body, reads as not-verified-live, and was stolen at age ~0 — robbing a
holder mid-creation. Add a fresh-create floor scoped to the unverifiable-body
case: an empty/unparseable body that is fresh is treated as mid-creation and is
NOT stolen, while a COMPLETE dead-pid body is still stolen promptly (preserves
the prompt-dead-steal contract). planning-workspace writes its body atomically
(flag:'wx') so it has no empty-body window.
- Double-steal (both locks): the steal was a bare fs.unlinkSync with no identity
re-confirm, so two waiters could both reclaim a dead holder and end up holding
concurrently. Replace with an atomic renameSync (only one racer wins the inode)
guarded by a (dev,ino,body) identity re-confirm immediately before the steal;
body content is part of the identity to defeat inode reuse.
Tests (seam-driven, no wall-clock, each proven RED-before-GREEN):
- clock-seam: fresh empty-body lock is not stolen at age ~0; a racer-recreated
live lock is not double-stolen (identity re-confirm). Adds a beforeSteal seam.
- planning-workspace: racer-recreated live lock is not double-stolen.
- Updated the two #1217 unlinkSync-failure tests to the renameSync steal path
(the bounded-backoff/no-busy-spin guarantee is preserved and re-asserted).
Uncontended acquire path is byte-for-byte unchanged.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* fix(core): roadmap upgrade rollback must restore .planning regardless of git tracking (#1542)
applyMigration rolled back a failed migration with git reset --hard + git clean
-fd .planning/phases/. For a commit_docs:false project (.planning gitignored —
the default) that restores NOTHING (reset ignores untracked, clean without -x
skips ignored), yet it threw 'Migration failed (rolled back to <sha>)' — a false
claim leaving .planning half-migrated. git reset --hard is also a whole-repo op.
Replace it with a surgical, git-independent rollback: record the exact renames
performed and snapshot each file before rewriting it, then on failure reverse the
renames and restore the snapshots (deleting files that did not previously exist).
Correct whether .planning is tracked or ignored; touches only what it changed.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* chore(changeset): Fixed fragment for #1543 (roadmap upgrade surgical rollback)
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* test(core): update bug-685 execSync count floor after surgical rollback (#1542)
The #1542 surgical, git-independent rollback removed the rev-parse/reset/clean
git execSync calls from roadmap-upgrade.cts, leaving only the git status
precondition. bug-685 asserted calls.length >= 4; lower the floor to >= 1 — the
durable guard (every remaining git execSync sets windowsHide:true) is unchanged.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>