workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).
- planner: add a Severity column to the STRIDE threat register; assign
severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
severity enum; redefine threats_open as the count of OPEN threats whose
severity is at or above block_on (none => 0). Below-threshold opens are
reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.
No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.
- New reference gsd-core/references/security-asvs-levels.md defines L1
(opportunistic), L2 (standard), L3 (comprehensive) for both planner
threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
by extracting the goal-backward worked example to planner-guidance.md.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':
1. Missing guards: workflow.security_block_on (enum) and
workflow.security_asvs_level (integer 1-3) had no store-time validation.
2. Systemic JSON-coercion bypass: every string-enum guard used
VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
before validation, String(["member"]) === "member" let a JSON array
slip through and an array was stored in a scalar key. Reproduced on
human_verify_mode, statusline.context_position, context_guard_mode,
fallow.scope/profile, source_grounding_authority, drift_action, context.
3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
including coerced arrays/objects and out-of-enum strings like
code_review_depth=garbage — was stored silently.
Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore.
Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact.
5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts.
Refs #1629 (Finding B; Finding A addressed in #1630).
PR #1622 (issue #1615) shipped Windsurf /gsd-* workflow wrappers that delegate to command bodies at <targetDir>/.windsurf/gsd-core/commands/gsd/X.md via a hardcoded @~/.claude/gsd-core/commands/gsd/ path. The path-rewrite pipeline correctly substitutes ~/.claude/ to the install target. But the source gsd-core/ dir does not ship with commands/ — the canonical command source lives at the package root (commands/gsd/). Without this copy, every /gsd-* workflow in Cascade references a file that does not exist. The slash commands appear in the / menu but silently fail when invoked because the LLM is told to read a missing file.
None of the original reviews caught this: not the security review, not Codex's adversarial orthogonal review (gpt-5.5/high), not Memtrace's graph-backed review. It was surfaced by a #1629 regression test that verifies 'every workflow @-reference target exists on disk after install' — the test failed, revealing the bug.
Fix: for Windsurf local installs, copy commands/gsd/*.md into <targetDir>/gsd-core/commands/gsd/ via copyWithPathReplacement (applies the same path+brand rewrites as the rest of the install). Guarded on isWindsurf && !isGlobal since global Windsurf workflow install is an explicit no-op.
Documented as DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED in CONTEXT.md so the pattern is locked in: any new converter emitting a wrapper that delegates to another file MUST verify the delegation target is actually installed.
The commandName validation tests legitimately contain real injection payloads (newline + system-role override phrases, fake [SYSTEM] tags, jailbreak strings) to prove the validator rejects them. The scanner cannot distinguish a test fixture asserting rejection from an actual injection attempt, so CI failed on the test that adds the security control.
Added tests/windsurf-conversion.test.cjs to scripts/prompt-injection-scan.sh ALLOWLIST with a comment citing the defect class.
Also added DEFECT.PROMPT-INJECTION-SCAN-COLLISION-WITH-TESTS to CONTEXT.md so the pattern is documented. Initial draft of that predicate ITSELF triggered the scanner (it quoted the literal injection phrase as an example) — reworded to use descriptive references ('scanner-matching payload', 'instruction-override phrase') since the scanner scans CONTEXT.md too. That meta-collision is now called out in the fix-forward and prevention subkeys.
Resolved conflict in docs/reference/skill-mapping-matrix.md: kept next's Antigravity row (flat layout per #1614, doc-fixed by #1617) AND kept HEAD's Windsurf row (workflows layout per #1615). Also updated the Structural Facts section counts that the original #1615 PR left stale: '15 skill-bearing runtimes' → '14' (Windsurf no longer skill-bearing), 'Nine runtimes stay flat' → 'Eight', removed windsurf from the 'FLAT (unconfirmed)' row in the loader-verification table.
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes.
For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target).
Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode.
Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615.
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body.
Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message).
Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation.
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only.
Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX.
Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input.
Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring.
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.
- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
block (extractFrontmatter can't — its `-` items are scalars-only; this is a
focused parser, sibling of parseMustHavesBlock), validates each entry, and
classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
AND non-empty all-`pass` verification AND zero validation errors. Everything
else — judgment, empty/failing verification, malformed entry — routes to the
human (fail-safe). A malformed block falls back to legacy prose extraction and
surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
human_judgment:true); verify-work extract_tests consumes it; create_uat_file
marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
for the null-entry/comment-header/mis-indent cases found in adversarial review.
Closes#1602
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* enhance(#1549): validate PR-title issue-ref convention at open time
The release changelog is title-driven: release.yml generates "What's Changed"
from PR titles, then format-github-release-notes.cjs buckets each line by its
conventional-commit prefix and relies on a `(#<issue>)` in the title to render
the issue link. Both rules were enforced only socially, so titles like
`fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag
defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke
the changelog, landing on the maintainer as release-time cleanup.
Extract the title matcher into one shared module consumed by BOTH the changelog
classifier and a new PR-title CI gate, so a title that passes the gate cannot
mis-bucket in the changelog (single source of truth).
- scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle +
the anchored regexes. One matcher, two consumers.
- scripts/release-notes/format-github-release-notes.cjs: classifyTitle now
delegates to classifyBucket (behavior preserved; existing tests green).
- .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on
pull_request opened/edited/reopened/synchronize, for ALL authors (the drift
came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout.
- tests/conventional-title.test.cjs (new): bucket + gate cases incl. the
leading-tag mis-bucket (backfills the untested classifyTitle case) and a
cross-check that the classifier delegates to the shared matcher.
- CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag.
Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk
* fix(#1549): check out the PR in pr-title-validator so the new matcher resolves
The workflow checked out the base branch (next) as a trusted policy source, but
the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR
and does not exist on next yet — so require() failed and validate-title errored
on its own introducing PR. Check out the PR's merge ref instead: the matcher
under review is present, the check is self-consistent, and a fork pull_request
runs read-only with no secrets, so running the PR's own pure-string regex is
safe.
* fix(#1549): move conventional-title.cjs out of installed scripts/lib/
bin/install.js bundles every file under scripts/lib/ into the user-installed
payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts
that exact set. The new matcher is release/CI tooling that must NOT ship to
users, so placing it in scripts/lib/ both broke the install manifest test and
would have shipped dead code. Relocate it next to its consumer in
scripts/release-notes/ (which the installer does not copy) and update the three
require paths (classifier, workflow, test) + the CONTRIBUTING reference.
install.test.cjs now 125/125; conventional-title + release-notes suites green;
lint:ci clean.
* fix(#1549): load title matcher from trusted base ref, not PR code
Addresses review (Solvely-Colin + trek-e): the gate checked out the PR
merge ref and require()'d evaluatePrTitle from PR-controlled code, so any
future PR could edit conventional-title.cjs to return { valid: true } and
wave its own malformed title through — a self-bypassable required check.
Load the matcher from a base-branch checkout instead (ref:
github.event.pull_request.base.ref), the same trusted-policy-source pattern
pr-target-validator.yml already uses. The PR can change its title but not
the ruler that measures it. An existsSync bootstrap guard skips the check
when the matcher isn't on the base branch yet (the introducing PR); every
PR after merge is fully gated. This keeps the single shared matcher (#1549's
whole point) rather than forking the regex into the workflow.
Also per review:
- add tests/conventional-title.property.test.cjs (fast-check): any
`type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket
are total functions (never throw).
- pin the `fix(#):` zero-digit boundary as missing-issue-ref.
Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.
Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).
Closes#1592
Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
Per review (test-standard items):
- Property test (RULESET.TESTS.property-based-testing): extractRetiredPhaseNumbers
is the parsing core, so add a fast-check property — k of n checklist phases
struck → exactly the k canonical keys returned, across randomized phase counts
and numeric/zero-padded/project-code ID forms. Exposed via a `_`-prefixed test
seam (mirrors the existing _setLockProbes seams), no public API surface added.
- Boundary (RULESET.TESTS.boundary-coverage): all-retired case (k === n) →
total_phases 0, via state json.
No production behavior change; the exclusion logic is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Resolves TBD entries in the skill-mapping matrix provision/consumption
table. All 15 non-Claude platforms assessed as N/A — none have a
documented plugin skill-provision model. File-copy install path
handles skill provision for all runtimes. Updates claude row to
reflect Phase B-provide merged status. Closes epic #1258 acceptance.
Closes#1600
Phase B-provide of epic #1258. Adds a build-generated skills/ dir +
a skills manifest field so plugin-installed GSD exposes gsd-core:<skill>
the native Claude Code way. Closes the gap where plugin-only installs
lacked the skill surface because bin/install.js never ran.
- scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md
to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill
- .claude-plugin/plugin.json: add "skills": "./skills/"
- package.json: add skills to files, gen:plugin-skills to build chain
- tests/issue-766-plugin-manifest.test.cjs: Section H conformance
(manifest field + dir + frontmatter + count parity) + C2 skills symlink
- docs/adr/766-*.md: dated amendment adding skills surface row
- .changeset/rapid-bears-hum.md: type Added
- skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed)
Closes#1596
Phase A of epic #1258. Adds the single authoritative ADR documenting
the per-runtime skill mapping + converter transform-contract catalog
that ADR-3660 (layout) and ADR-1508 (module ownership) each carry a
third of. Ships a companion reference matrix projecting all 16
runtimes from their capability descriptors.
- docs/adr/1593-skill-mapping-converter-methodology.md (new ADR)
- docs/reference/skill-mapping-matrix.md (new reference page)
- docs/adr/README.md (index row + status fixes: 3660, 1016 Proposed->Accepted)
- docs/adr/1016-runtime-capability-descriptor.md (header Proposed->Accepted;
the ConverterName enum is already code-enforced per ADR-857 phase 5e)
Doc-only. No code, no behavior change. Closes#1593.
CodeQL alert #41 (js/incomplete-sanitization) flagged the partial
escape class /[-]/g at tests/worktree-safety.test.cjs:645 — it only
escaped hyphen-minus, leaving 13 other regex metacharacters (notably
backslash) unescaped. The canonical class /[.*+?^${}()|[\]\\]/g is
what every sibling escape in the test suite already uses
(bug-2839, bug-2760, 4-phase-complete, phase6-capstone-conformance).
Today dormant: the flag array is a hardcoded [a-z-] literal, so the
expanded class is a no-op for the four existing flags and the regexes
they produce are byte-identical. The fix prevents future drift — a
contributor adding e.g. '--output=file' would have silently introduced
a regex wildcard.
All 69 tests in the file pass. No user-facing behavior change.
Fixes#1589