Commit Graph

3958 Commits

Author SHA1 Message Date
Tom Boucher
e1d768dd78 fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op (#1664)
* fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op

cmdFrontmatterSet reported {updated:true} even when spliceFrontmatter returned the
content unchanged, which happened whenever the new value's extractFrontmatter projection
equalled the original's — notably for object-list fields like must_haves, whose
{path,provides} items flatten to scalar strings under the lossy parser. Detect a no-op
(newContent === content) for a dict-valued field and surface an error directing the user
to edit the file directly, instead of silently accepting a no-op set. Scalars and scalar
arrays round-trip faithfully, so idempotent sets of those are intentionally NOT flagged
(two precision regression tests lock this). Folded into frontmatter-cli.test.cjs.

* chore(#1660): backfill changeset pr ref to 1664

* refactor(#1660): extract noOpObjectListSetError as pure tested helper (Stryker coverage)

cmdFrontmatterSet is not in Stryker's property/unit test set, so the inline no-op
detection added survivors that dropped the frontmatter module below its 62% mutation
threshold. Extract the detection into a pure exported helper noOpObjectListSetError and
unit-test every branch directly (changed content, scalar, scalar-array, null, dict
no-op). cmdFrontmatterSet now calls the helper. Same pattern as the #1572 spliceFrontmatter
coverage fix.
2026-06-24 13:39:17 -04:00
Tom Boucher
f615eb9ef3 fix(#1572): preserve must_haves object-lists across frontmatter set/merge (#1656)
* fix(#1572): preserve must_haves object-lists across frontmatter set/merge

spliceFrontmatter round-tripped the WHOLE frontmatter through extractFrontmatter
(a scalar-only parser) then reconstructFrontmatter (a lossy serializer), so any
must_haves object-list — artifacts {path, provides}, prohibitions {statement,
status} — was flattened to scalar strings and re-emitted as a malformed inline
array whenever an UNRELATED field changed, silently dropping every provides:/
status: value. The write now preserves the original raw text for any top-level
key whose value is structurally unchanged between the original parse and the new
object (generalizing the existing whole-document no-op guard to per-key
fidelity), and regenerates only the key that actually changed. The key set is
still defined by newObj (the cmdSet/cmdMerge flow always passes the full merged
object). spliceFrontmatter's only callers are cmdFrontmatterSet/Merge — the
STATE.md read-modify-write family calls reconstructFrontmatter directly and is
unaffected. Regression cases folded into tests/frontmatter-cli.test.cjs:
artifacts/prohibitions object-lists survive set and merge; idempotent on repeat
sets. Asserted via parseMustHavesBlock (the structure-preserving parser).

* chore(#1572): backfill changeset pr ref to 1656

* fix(#1572): fail-closed when set/merge would emit [object Object] (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that directly setting a must_haves
object-list (a CHANGED key) still routed through the lossy reconstructFrontmatter,
emitting literal "[object Object]" and destroying the data. The reported case
(mutating an UNRELATED field) was already fixed by per-key raw-text preservation,
but the changed-object-list path was still silently lossy. Add fail-closed: when a
regenerated key's text contains the "[object Object]" sentinel, spliceFrontmatter
throws — cmdFrontmatterSet/Merge error out WITHOUT writing, directing the user to
edit the file directly. The no-frontmatter (generate-from-scratch) path is guarded
the same way. Adds a test that a refused set leaves the file unchanged and the
original object-list intact. Codex finding #2 (a contrived flattened-projection
no-op) is a deeper limitation noted in the PR — non-destructive, and the fail-closed
message already directs users to edit object-list blocks directly.

* test(#1572): add spliceFrontmatter per-key preservation + fail-closed unit coverage

Stryker mutates gsd-core/bin/lib/frontmatter.cjs against tests/frontmatter.{property,unit}.test.cjs
(MinScore 62). The #1572 regression cases live in frontmatter-cli.test.cjs, which is NOT in
Stryker's test set, so the new functions (sliceTopLevelFrontmatterSegments, the per-key
preserve/regenerate/drop/append loop, regenerateFrontmatterKey's [object Object] fail-closed)
had surviving mutants that dropped the module below threshold. Add unit-level coverage in
frontmatter.unit.test.cjs exercising every new branch directly via spliceFrontmatter:
unchanged object-list preserved (provides survives) when a scalar sibling changes; changed
scalar regenerates only that key; orphan keys dropped; new keys appended; indented nested
block stays attached to its parent key; whole-document no-op returns input verbatim; both
fail-closed paths (changed object-list + no-frontmatter) throw.
2026-06-24 13:30:46 -04:00
Tom Boucher
b205e4c2b2 fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form

bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.

* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject

The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.

* chore(#1639): backfill changeset pr ref to 1665
2026-06-24 13:26:25 -04:00
Tom Boucher
4631923982 fix(#1569): preserve explicit resolve_model_ids in non-Claude installs (#1653)
* fix(#1569): preserve explicit resolve_model_ids in non-Claude installs

The non-Claude finishInstall step keyed its resolve_model_ids:"omit" write on
!== "omit", so an explicit true opt-in (resolveModelInternal returns full model
IDs) was silently clobbered on every install/upgrade across all 14 non-Claude
runtimes, making generated agent manifests inherit the active chat model instead
of pinning the resolved model. Now only absent/falsy is defaulted to "omit"; an
explicit true (and an existing "omit") is preserved. Regression test
parameterizes across codex/opencode/gemini and covers the absent/false/idempotent/
claude/malformed boundaries.

* chore(#1569): backfill changeset pr ref to 1653

* fix(#1569): default non-canonical resolve_model_ids values to omit (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that the original allowlist-by-
enumeration condition (undefined/null/false -> omit) preserved malformed values
(0, "", "yes", {}) instead of defaulting them to omit, letting them leak Claude
aliases a non-Claude runtime cannot resolve. Switch to an allowlist condition
(existing !== true && existing !== 'omit') so only an explicit canonical true
opt-in and an existing omit are preserved; everything else defaults to the safe
non-Claude omit. Adds a parameterized test over [0, "", "yes", {}].
2026-06-24 13:21:19 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
Tom Boucher
eef8f9b8c3 Merge pull request #1648 from open-gsd/chore/sync-next-version-1.6.0-rc.3
chore: sync next package version to 1.6.0-rc.3
2026-06-23 23:52:27 -04:00
github-actions[bot]
2e406e8e05 chore: sync next package version to 1.6.0-rc.3 2026-06-24 03:52:20 +00:00
Tom Boucher
35478b615e refactor(#1646): route capability routers through Command Routing Hub per ADR-959 (#1647)
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959

Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.

src/cjs-command-router-adapter.cts
  * Imported ERROR_REASON from io.cjs.
  * UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
    as the second arg to error() — additive for host routers (their
    existing one-arg error callbacks ignore the second arg), required
    for capability routers whose tests assert reason === 'sdk_unknown_command'
    on the JSON-error envelope.

src/graphify-command-router.cts
  * Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing/invalid --budget) now
    return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
    instead of calling error() directly (Q2=C, Q4=ii from grilling).
  * Success handlers keep direct output() calls.
  * Subcommands array is alphabetical for byte-identical 'Available:'
    text in the unknown-subcommand message.
  * The unknown-subcommand path is now owned by the Hub's manifest
    check (the adapter passes SDK_UNKNOWN_COMMAND).

src/intel-command-router.cts
  * Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing filePath for patch-meta
    and extract-exports) return makeInvalidArgs Results.
  * Preserved the timeAgo mutation in the non-raw status handler.
  * Preserved the lazy require('./intel.cjs') inside the route function.

src/audit-command-router.cts
  * routeAuditUat: routes through the Hub with a synthetic 'run'
    defaultSubcommand (no real subcommands). Gives uniform observability.
  * routeAuditOpen: captures --json in a closure, strips it from args
    before Hub dispatch (so it isn't mistaken for a subcommand by the
    manifest check), then branches on wantJson inside the handler to
    preserve the formatAuditReport success-path quirk.

docs/CONFIGURATION.md
  * Observability section: noted capability commands (graphify, intel,
    audit-uat, audit-open) now emit DispatchEvent records since #1646.

.changeset/capability-routers-via-hub.md
  * Changed fragment describing the user-visible audit-trail expansion.
    pr:0 placeholder will be backfilled after gh pr create returns the
    real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).

Verification
  * graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
    error path, JSON-errors, and registry assertions)
  * intel cutover tests: 39/39 pass
  * audit cutover tests: 24/24 pass
  * bug-974-graphify-budget-missing-value regression test: pass
  * npm run test:unit (full suite): 2384 tests, 0 fail
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.

* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
2026-06-23 23:21:06 -04:00
Tom Boucher
6214039358 refactor(#1644): Hub extension — exitReason? field on InvalidArgs + adapter honestification (#1645)
Phase 1 of parent #1641. Implements the contract documented in the
Phase 0 ADR-0174 §5 amendment (#1642 / #1643).

src/command-routing-hub.cts
  * InvalidArgsResult interface gains optional exitReason?: string
    (carries an ERROR_REASON enum value, separate from reason which is
    the explanation text).
  * makeInvalidArgs(arg, reason, exitReason?) factory conditionally adds
    the field only when the third arg is truthy — preserves the strict-
    keys invariant tested at command-routing-hub.test.cjs:444.
  * _VARIANT_SCHEMA.InvalidArgs.allowed Set extended to include
    'exitReason' so the runtime validator does not coerce well-formed
    extended Results to HandlerFailure.

src/cjs-command-router-adapter.cts
  * Honestified the wrapper comment: the runtime check ('ok' in result)
    already passes any {ok:*} object through, so the historical
    {ok:true, data} return type was a lie for err Results. The lying
    cast is preserved because the Hub's export = syntax doesn't expose
    HubResult for import; the Hub's _validateErrResult runtime-validates
    the actual shape.
  * Result→error() translation branched: when InvalidArgs carries
    exitReason, the adapter calls error(result.reason, result.exitReason)
    so the JSON-error envelope (GSD_JSON_ERRORS=1) preserves the typed
    ERROR_REASON value. When exitReason is absent, error(msg) is called
    with exactly one arg — byte-identical with prior behavior.
  * RouteCjsCommandFamilyOptions.error and RouteHubCommandFamilyOptions
    .error callback types widened from (message) to (message, reason?)
    to match io.cts's actual error() signature.

CONTEXT.md
  * Command Routing Hub predicate updated to document the new field,
    factory signature, and dispatcher translation contract.

Tests (TDD red→green)
  * tests/command-routing-hub.test.cjs: 8 new tests covering 2-arg
    (strict-keys), 3-arg (key present), undefined, empty string, frozen
    result, hub.dispatch propagation, and validator acceptance.
  * tests/cjs-command-router-adapter.test.cjs: 2 new tests covering
    exitReason passed as second arg + byte-identical prior behavior when
    absent.

Verification
  * npm run test:unit: 2448 tests, 0 fail (no regressions)
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

Memtrace blast radius: LOW (get_impact makeInvalidArgs → 3 nodes; the
optional field is non-breaking for the 1 existing caller routePhaseCommand).
2026-06-23 22:47:57 -04:00
Tom Boucher
bcc5a6d1ba fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command

Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.

- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
  absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
  .sh and others keep the bare quoted path (unchanged).

Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.

Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.

* chore(#1634): backfill changeset pr:1638

* fix(#1634): resolve lint and windows CI failures

- validator: replace the control-character range regex with a char-code loop.
  The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
  codes are equally precise and lint-clean. Behavior unchanged (still rejects
  matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
  honor POSIX write modes (a 0o644 write reads back as 0o666), so the
  precondition is meaningless there and failed the windows-latest lane. The
  node-prefix assertion — the actual fix — is platform-independent and still
  runs everywhere.

* docs(#1634): amend ADR-894 for optional lifecycle hook matcher

The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.

* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md

Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
2026-06-23 21:49:43 -04:00
Tom Boucher
61be100a02 docs(#1642): amend ADR-0174 §5 — reconcile Result type + add exitReason? (#1643)
Two changes to ADR-0174 §5 (Sync dispatch with tight-typed Result<T>):

1. Reconcile the documented Result type to the as-built code. The
   original §5 text planned 'Unknown' / 'BadArgs' / 'ValidationFailed'
   / 'NotImplemented' / 'HandlerFailed'; the SDK retirement migration
   kept the ADR-0012 names (UnknownCommand / InvalidArgs / HandlerFailure)
   and never added the planned ValidationFailed or NotImplemented
   variants. HandlerRefusal was added during implementation but never
   back-filled into this ADR. The ADR now documents what consumers
   actually depend on.

2. Add the optional exitReason?: string field on the InvalidArgs variant
   (and update the makeInvalidArgs factory signature). This carries an
   ERROR_REASON enum value separately from the existing reason
   explanation text, so capability routers migrating from direct
   error(msg, ERROR_REASON.USAGE) calls to makeInvalidArgs(...) Results
   preserve ERROR_REASON granularity through the Hub Result →
   error(msg, exitReason) translation. The field is additive and
   backward-compatible.

Also tightens the amendment requirement: 'Adding a new variant OR
adding a field to an existing variant requires amending this ADR.'

Phase 0 of parent #1641. No code changes; pure ADR amendment.
2026-06-23 21:42:16 -04:00
Behruz Nassre Esfahani
fc5ca178a2 test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers

The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.

Test-only; no production code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1178): note AGENTS_DIR export is for future call sites

Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:34:08 -04:00
Joe Seymour
9d12725e4e fix(#1619): normalize pruned mise node execPath to the stable shim in normalizeNodePath (#1621)
* fix(#1619): normalize pruned mise node execPath to the stable shim

resolveNodeRunner() bakes process.execPath into managed .js hook commands.
Node realpaths execPath, so under mise it resolves to a concrete
<data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after
which every managed hook 404s — the same ephemeral-path failure #977 fixed
for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise
versioned install path to the stable sibling shim <data>/shims/node when it
exists (deriving <data> from execPath so a custom MISE_DATA_DIR works),
falling back to the raw execPath otherwise. Tests folded into
install.test.cjs per the regression test-name lint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(changeset): set pr number to 1621

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Joe Seymour <joese@iarx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:17:03 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Tom Boucher
f9d9dfb4bc fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.

- New reference gsd-core/references/security-asvs-levels.md defines L1
  (opportunistic), L2 (standard), L3 (comprehensive) for both planner
  threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
  hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
  L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
  by extracting the goal-backward worked example to planner-guidance.md.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 18:48:58 -04:00
Tom Boucher
94be6d5b60 fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) 2026-06-23 17:24:59 -04:00
Tom Boucher
01ef08ed2a Merge pull request #1632 from open-gsd/fix/1628-config-set-security-enum-validation 2026-06-23 17:24:23 -04:00
Tom Boucher
0c4d570541 fix(#1628): type-safe config-set validation — close JSON-coercion enum bypass + enforce capability schema
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':

1. Missing guards: workflow.security_block_on (enum) and
   workflow.security_asvs_level (integer 1-3) had no store-time validation.

2. Systemic JSON-coercion bypass: every string-enum guard used
   VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
   before validation, String(["member"]) === "member" let a JSON array
   slip through and an array was stored in a scalar key. Reproduced on
   human_verify_mode, statusline.context_position, context_guard_mode,
   fallow.scope/profile, source_grounding_authority, drift_action, context.

3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
   25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
   including coerced arrays/objects and out-of-enum strings like
   code_review_depth=garbage — was stored silently.

Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 16:10:34 -04:00
Tom Boucher
adef8b293a Merge pull request #1631 from open-gsd/fix/1629b-legacy-cleanup
fix(#1629): cleanup legacy .devin/skills/gsd- dirs on Windsurf reinstall
2026-06-23 16:08:38 -04:00
Tom Boucher
118aaa1098 chore(#1629): backfill changeset PR number 2026-06-23 15:54:07 -04:00
Tom Boucher
db2b4d326a fix(#1629): cleanup legacy .devin/skills/gsd- dirs on Windsurf reinstall
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore.

Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact.

5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts.

Refs #1629 (Finding B; Finding A addressed in #1630).
2026-06-23 15:54:07 -04:00
Tom Boucher
4abdfcaaef Merge pull request #1630 from open-gsd/fix/1629a-install-order
fix(#1629): copy Windsurf command bodies so workflow delegation targets resolve
2026-06-23 15:43:17 -04:00
Tom Boucher
7fdca5ecc6 chore(#1629): add changeset for Windsurf command body copy fix 2026-06-23 15:31:14 -04:00
Tom Boucher
b65939cdff fix(#1629): copy Windsurf command bodies so workflow delegation targets exist
PR #1622 (issue #1615) shipped Windsurf /gsd-* workflow wrappers that delegate to command bodies at <targetDir>/.windsurf/gsd-core/commands/gsd/X.md via a hardcoded @~/.claude/gsd-core/commands/gsd/ path. The path-rewrite pipeline correctly substitutes ~/.claude/ to the install target. But the source gsd-core/ dir does not ship with commands/ — the canonical command source lives at the package root (commands/gsd/). Without this copy, every /gsd-* workflow in Cascade references a file that does not exist. The slash commands appear in the / menu but silently fail when invoked because the LLM is told to read a missing file.

None of the original reviews caught this: not the security review, not Codex's adversarial orthogonal review (gpt-5.5/high), not Memtrace's graph-backed review. It was surfaced by a #1629 regression test that verifies 'every workflow @-reference target exists on disk after install' — the test failed, revealing the bug.

Fix: for Windsurf local installs, copy commands/gsd/*.md into <targetDir>/gsd-core/commands/gsd/ via copyWithPathReplacement (applies the same path+brand rewrites as the rest of the install). Guarded on isWindsurf && !isGlobal since global Windsurf workflow install is an explicit no-op.

Documented as DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED in CONTEXT.md so the pattern is locked in: any new converter emitting a wrapper that delegates to another file MUST verify the delegation target is actually installed.
2026-06-23 15:28:54 -04:00
Tom Boucher
c3ab4807ae Merge pull request #1622 from open-gsd/fix/1615-workflows-agents-not-installing-for-wind
fix(#1615): install Windsurf slash workflows
2026-06-23 15:15:59 -04:00
Tom Boucher
c63fa35b0a fix(#1615): allowlist windsurf-conversion.test.cjs in prompt-injection-scan
The commandName validation tests legitimately contain real injection payloads (newline + system-role override phrases, fake [SYSTEM] tags, jailbreak strings) to prove the validator rejects them. The scanner cannot distinguish a test fixture asserting rejection from an actual injection attempt, so CI failed on the test that adds the security control.

Added tests/windsurf-conversion.test.cjs to scripts/prompt-injection-scan.sh ALLOWLIST with a comment citing the defect class.

Also added DEFECT.PROMPT-INJECTION-SCAN-COLLISION-WITH-TESTS to CONTEXT.md so the pattern is documented. Initial draft of that predicate ITSELF triggered the scanner (it quoted the literal injection phrase as an example) — reworded to use descriptive references ('scanner-matching payload', 'instruction-override phrase') since the scanner scans CONTEXT.md too. That meta-collision is now called out in the fix-forward and prevention subkeys.
2026-06-23 15:05:07 -04:00
Tom Boucher
481ca1b18e Merge origin/next into fix/1615-workflows-agents-not-installing-for-wind
Resolved conflict in docs/reference/skill-mapping-matrix.md: kept next's Antigravity row (flat layout per #1614, doc-fixed by #1617) AND kept HEAD's Windsurf row (workflows layout per #1615). Also updated the Structural Facts section counts that the original #1615 PR left stale: '15 skill-bearing runtimes' → '14' (Windsurf no longer skill-bearing), 'Nine runtimes stay flat' → 'Eight', removed windsurf from the 'FLAT (unconfirmed)' row in the loader-verification table.
2026-06-23 15:01:16 -04:00
Tom Boucher
658ea33cb6 fix(#1615): applySurface rewrites commands kind, not just skills
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes.

For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target).

Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode.

Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615.
2026-06-23 14:52:32 -04:00
Tom Boucher
4ed208e74b fix(#1615): validate commandName to prevent workflow prompt injection
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body.

Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message).

Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation.
2026-06-23 14:38:38 -04:00
Tom Boucher
527142ad2e fix(#1615): normalize Windows backslash paths in workflow content
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only.

Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX.

Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input.

Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring.
2026-06-23 14:26:44 -04:00
Tom Boucher
d71694a166 Merge pull request #1624 from open-gsd/docs/1618-claude-flat-skill-docs
docs(#1618): correct Claude skill layout from nested to flat
2026-06-23 14:14:24 -04:00
Tom Boucher
e32eac56a6 docs(#1618): correct Claude skill layout from nested to flat 2026-06-23 14:11:22 -04:00
Tom Boucher
8f641bfab1 Merge pull request #1623 from open-gsd/docs/1617-antigravity-flat-skill-docs
docs(#1617): correct Antigravity skill layout from nested to flat
2026-06-23 14:09:53 -04:00
Tom Boucher
08dfcad5f1 docs(#1617): correct Antigravity skill layout from nested to flat 2026-06-23 14:05:49 -04:00
Tom Boucher
4d2f88377a chore(#1615): backfill changeset PR number 2026-06-23 13:26:59 -04:00
Tom Boucher
6d782e309d test(#1615): update Windsurf workflow expectations 2026-06-23 12:51:46 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
c28cccbf85 Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
2026-06-23 10:56:29 -04:00
Tom Boucher
1b95762661 Merge pull request #1568 from behruznassre/fix/1514-retired-phase-total-phases
fix(#1514): exclude retired/folded phases from progress.total_phases
2026-06-23 10:55:41 -04:00
Tom Boucher
afabb7c526 Merge pull request #1616 from open-gsd/fix/1614-antigravity-flat-skills
fix(#1614): install Antigravity skills flat
2026-06-23 10:54:27 -04:00
Tom Boucher
ce7fffffc4 test(#1614): classify Antigravity skills as flat 2026-06-23 10:40:55 -04:00
Tom Boucher
6db580d694 chore(#1614): backfill changeset PR number 2026-06-23 10:22:57 -04:00
Tom Boucher
da2de3a183 Merge branch 'next' into fix/1514-retired-phase-total-phases 2026-06-23 10:22:18 -04:00
Tom Boucher
cbd21092a9 fix(#1614): install Antigravity skills flat 2026-06-23 10:21:06 -04:00
Tom Boucher
610720d163 Merge pull request #1611 from open-gsd/claude/loving-cannon-745e64
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
2026-06-23 09:26:09 -04:00
Tom Boucher
50eb6176fb chore(#1602): add changeset for #1611
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Rezolv
652142521b enhance(#1549): validate PR-title issue-ref convention at open time (#1576)
* enhance(#1549): validate PR-title issue-ref convention at open time

The release changelog is title-driven: release.yml generates "What's Changed"
from PR titles, then format-github-release-notes.cjs buckets each line by its
conventional-commit prefix and relies on a `(#<issue>)` in the title to render
the issue link. Both rules were enforced only socially, so titles like
`fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag
defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke
the changelog, landing on the maintainer as release-time cleanup.

Extract the title matcher into one shared module consumed by BOTH the changelog
classifier and a new PR-title CI gate, so a title that passes the gate cannot
mis-bucket in the changelog (single source of truth).

- scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle +
  the anchored regexes. One matcher, two consumers.
- scripts/release-notes/format-github-release-notes.cjs: classifyTitle now
  delegates to classifyBucket (behavior preserved; existing tests green).
- .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on
  pull_request opened/edited/reopened/synchronize, for ALL authors (the drift
  came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout.
- tests/conventional-title.test.cjs (new): bucket + gate cases incl. the
  leading-tag mis-bucket (backfills the untested classifyTitle case) and a
  cross-check that the classifier delegates to the shared matcher.
- CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag.

Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk

* fix(#1549): check out the PR in pr-title-validator so the new matcher resolves

The workflow checked out the base branch (next) as a trusted policy source, but
the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR
and does not exist on next yet — so require() failed and validate-title errored
on its own introducing PR. Check out the PR's merge ref instead: the matcher
under review is present, the check is self-consistent, and a fork pull_request
runs read-only with no secrets, so running the PR's own pure-string regex is
safe.

* fix(#1549): move conventional-title.cjs out of installed scripts/lib/

bin/install.js bundles every file under scripts/lib/ into the user-installed
payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts
that exact set. The new matcher is release/CI tooling that must NOT ship to
users, so placing it in scripts/lib/ both broke the install manifest test and
would have shipped dead code. Relocate it next to its consumer in
scripts/release-notes/ (which the installer does not copy) and update the three
require paths (classifier, workflow, test) + the CONTRIBUTING reference.

install.test.cjs now 125/125; conventional-title + release-notes suites green;
lint:ci clean.

* fix(#1549): load title matcher from trusted base ref, not PR code

Addresses review (Solvely-Colin + trek-e): the gate checked out the PR
merge ref and require()'d evaluatePrTitle from PR-controlled code, so any
future PR could edit conventional-title.cjs to return { valid: true } and
wave its own malformed title through — a self-bypassable required check.

Load the matcher from a base-branch checkout instead (ref:
github.event.pull_request.base.ref), the same trusted-policy-source pattern
pr-target-validator.yml already uses. The PR can change its title but not
the ruler that measures it. An existsSync bootstrap guard skips the check
when the matcher isn't on the base branch yet (the introducing PR); every
PR after merge is fully gated. This keeps the single shared matcher (#1549's
whole point) rather than forking the regex into the workflow.

Also per review:
- add tests/conventional-title.property.test.cjs (fast-check): any
  `type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket
  are total functions (never throw).
- pin the `fix(#):` zero-digit boundary as missing-issue-ref.

Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 22:44:44 -04:00
Tom Boucher
fcf44ac02a Merge pull request #1604 from open-gsd/docs/1603-adr-1239-opencode-binding
docs(#1603): add opencode host-plugin binding to adr-1239
2026-06-22 22:33:49 -04:00
Tom Boucher
1985d4a4de docs(#1603): add opencode host-plugin binding to adr-1239 2026-06-22 22:29:48 -04:00