Commit Graph

294 Commits

Author SHA1 Message Date
Andreas Brauchli
cbf7c82841 feat(#323): fish-shell support in post-install PATH suggestion (#727)
* feat(#323): fish-shell support in post-install PATH suggestion

Two additive changes to the post-install PATH-suggestion seam, both scoped
to existing functions.

A. Projection: add a fish entry to the persist-mode shell-action list in
   projectPathActionProjection() (src/shell-command-projection.cts). fish has
   no `export`/`$PATH`-list syntax, so the existing zsh/bash `export PATH=...`
   commands are inert when pasted. The new entry emits the fish-native
   `fish_add_path '<dir>'` (fish 3.2+, persists via the universal-variable
   store, de-duplicating). The directory is single-quoted with the same POSIX
   literal escaping as the zsh/bash siblings; verified round-tripping through
   real fish 3.7.0 for paths containing quotes, spaces, `$`, `*`, backticks
   and unicode.

B. Detection: add homePathCoveredByFishConfig() in bin/install.js, called
   from maybeSuggestPathExport() alongside homePathCoveredByRc(). fish does
   not use sh-style `export PATH=` rc files, so a fish user whose
   fish_user_paths already covers the global bin would otherwise get a
   false-positive "not on your PATH" warning on every install. Two
   side-effect-free detection routes (no fish subprocess):

   1. The universal-variable store (~/.config/fish/fish_variables). fish
      serializes this with `full_escape`: every byte outside [A-Za-z0-9/_]
      becomes `\xHH` (space -> \x20, `-` -> \x2d, `.` -> \x2e, `$` -> \x24,
      unicode -> \uXXXX) and list elements are joined by the literal 4-char
      token `\x1e` (NOT a raw 0x1e byte). The detector splits on `\x1e`,
      decodes the escapes, then compares each as an absolute literal — a
      decoded `$` is part of the directory name, not an unexpanded variable.
      Verified against real fish 3.7.0 output.
   2. config.fish (`fish_add_path`, `set -gx PATH`, `set -Ux fish_user_paths`)
      — plain shell tokens: HOME forms ($HOME/${HOME}/~) are expanded and a
      token still holding `$` (e.g. `$PATH`, `$fish_user_paths`) is skipped.

   Honours $XDG_CONFIG_HOME and always also checks ~/.config/fish.

No behaviour change for bash/zsh/PowerShell/cmd/Git-Bash users: their entries
and command strings are unchanged; the fish entry is additive and the fish
detector only narrows the set of cases that warn.

Tests: update the projection length assertion (2 -> 3) and fish escaping in
bug-3441; add fish detection + suppression cases in install-path-detection
(uvar store with real fish escaping, dot/hyphen/space/$-literal decode
regressions, config.fish routes, commented-out, relative-segment guard,
unreadable-file fault injection, suppression and emission via
maybeSuggestPathExport).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(changeset): add Changed fragment for #323 fish PATH support (#727)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#323): address review — action-only fish docs, decoder property test, win32 guard

Addresses @trek-e's review on #727:

- docs (blocker): keep the how-to action-only (Diátaxis). Drop the
  `# fish — persists via …` comment and the internal-mechanism clause
  naming fish_variables/config.fish; leave one command + the exec-fish
  directive.
- tests (minor): extract decodeFishUniversalValue to a pure, exported
  module function and add fast-check round-trip properties
  (decode(fishEscape(p)) === p over arbitrary unicode, abs-path variant,
  totality). Consolidated into install-path-detection.test.cjs to respect
  the install test-file-count ratchet.
- tests (follow-up): port #721's win32 negative-projection test (no fish
  action on win32; persist projection is PowerShell/cmd.exe/Git Bash).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#323): address review — drop unused 'after' import, clarify escaping comment

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:20:52 -04:00
Alex V.
1a46109b97 enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb

Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb
(coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain;
gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt.
Non-breaking — additive only.

arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR).

* fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline

- C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives)
- C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws
- C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface)
- C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access)
- inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows)
- size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step)
- eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage)

* fix(#1579): register eval family in alias-drift gates

Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs
families and to familyArrayKeys in the manifest-coverage test, so the eval
family lands under the same drift guard as every sibling family
(state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review.

check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10.

* docs(#1579): use half-open verdict band ranges in CLI-TOOLS

overall_score is fractional and thresholds are >=80/>=60/>=40, so a score
in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel
bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit.

* fix(#1579): validate eval.score CLI inputs

Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior.
2026-06-24 17:16:13 -04:00
Behruz Nassre Esfahani
3870fafe74 fix(#1571): resolve schema-drift phase by token, not substring (#1640)
* fix(#1571): resolve schema-drift phase by token, not substring

verify schema-drift <phase> resolved the phase directory with a naive
entry.name.includes(phaseArg) test, so a non-existent phase could
silently match a different phase whose directory name merely contained
the requested token (e.g. "1" matched "11-expansion"), running the drift
gate against the wrong phase. Use the canonical phaseTokenMatches +
normalizePhaseName, matching find-phase, verify phase-completeness, and
this file's own unstarted-phase check.

Regression coverage folded into tests/schema-drift.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1571): add changeset for schema-drift token-match fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 17:03:01 -04:00
Tom Boucher
c583bcc02c fix(#1659): dedup By-Phase rows across padded/unpadded phase numbers (#1663)
* fix(#1659): dedup By-Phase rows across padded/unpadded phase numbers

phaseRowPattern matched the phase number literally (escapeRegex(String(phaseNum))),
so a seeded zero-padded row '| 05 |' was not matched by 'phase complete 5' (pattern '| 5 |'),
producing a duplicate row that double-counted the phase. Canonicalize a numeric phase to
its integer form (Number('05')===Number('5')===5) and match with a 0* prefix so 5/05/005
all collapse to the same row in either direction. Regression folded into state.test.cjs:
seeded '| 05 |' + 'phase complete 5' yields exactly one phase-5 row. Non-numeric phase IDs
retain the literal escapeRegex match.

* chore(#1659): backfill changeset pr ref to 1663

* fix(#1659): add verification fixture to padded-dedup test under #1522 gate
2026-06-24 14:41:45 -04:00
Tom Boucher
ff161f2281 fix(#1582): derive phase-complete velocity from By-Phase table (idempotent) (#1655)
* fix(#1582): derive phase-complete velocity from By-Phase table (idempotent)

updatePerformanceMetricsSection blind-added summaryCount onto the prior velocity
total on every phase complete, so re-running phase complete on an already-complete
phase incremented the total each time (the sibling of #4, which fixed the Completed
Phases counter the same way). The velocity total is now derived as the sum of the
By-Phase table's Plans column AFTER the row upsert — re-completing a phase upserts
the same row, so the sum is stable; a hand-edited inflated total self-heals downward
to the true sum on the next completion. When the By-Phase table is absent the total
is left unchanged (no crash). Strengthens the misnamed 'idempotent' test (its comment
explicitly declined to assert velocity idempotency — the latent gap) and adds a
self-heal regression; corrects the #320 behavior-lock velocity assertion which had
encoded the blind-add (3 = 1+2 double-count) — the derived value is 2.

* chore(#1582): backfill changeset pr ref to 1655

* fix(#1582): velocity sum tolerates indented By-Phase rows (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that byPhaseTablePattern's
data-row capture allows leading whitespace ([ \t]*\|), but the derive sum was
anchored at ^\| and would skip indented hand-edited/legacy rows — capturing them
in the table but silently undercounting. Align the sum regex (^\s*\|) with the
table capture's tolerance. Adds an indented-row regression. Two other codex
findings are pre-existing and out of scope: padded/unpadded phase dedup
(phaseRowPattern, identical in old code — derive yields the same value as the old
blind-add) and CRLF tables (the shared byPhaseTablePattern header requires bare
\n, so the upsert was already broken on CRLF; the fix changes stale-vs-
double-count, does not worsen it).

* fix(#1582): add verification fixtures to velocity tests under #1522 gate

Post-rebase onto next+#1548, the #1582 velocity tests (self-heal, indented-row) use
phase complete, which now fail-closes under #1522's canonical verification gate without a
passed *-VERIFICATION.md. Add writePassedVerification(tmpDir,'02-next','02') to both.
2026-06-24 14:36:01 -04:00
Tom Boucher
80607bec93 fix(#1658): make byPhaseTablePattern CRLF-tolerant on STATE.md tables (#1662)
* fix(#1658): make byPhaseTablePattern CRLF-tolerant on STATE.md tables

byPhaseTablePattern required a bare \n after the header and separator rows, so a
STATE.md with CRLF (\r\n) line endings (Windows, or hand-edited) had its By-Phase
table treated as absent: phase complete never upserted the row (and the velocity-from-
table derivation went stale). Make the header/separator terminators and the closing
lookahead CRLF-tolerant ([ \t]*\r?\n, (?=\r?\n|$)). Backward-compatible with LF.
Regression folded into tests/state.test.cjs: phase complete on a CRLF STATE.md upserts
the row and removes the placeholder. CONTRIBUTING's QA matrix lists Mixed CRLF/LF as a
required parser case.

* chore(#1658): backfill changeset pr ref to 1662
2026-06-24 13:52:47 -04:00
Tom Boucher
e1d768dd78 fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op (#1664)
* fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op

cmdFrontmatterSet reported {updated:true} even when spliceFrontmatter returned the
content unchanged, which happened whenever the new value's extractFrontmatter projection
equalled the original's — notably for object-list fields like must_haves, whose
{path,provides} items flatten to scalar strings under the lossy parser. Detect a no-op
(newContent === content) for a dict-valued field and surface an error directing the user
to edit the file directly, instead of silently accepting a no-op set. Scalars and scalar
arrays round-trip faithfully, so idempotent sets of those are intentionally NOT flagged
(two precision regression tests lock this). Folded into frontmatter-cli.test.cjs.

* chore(#1660): backfill changeset pr ref to 1664

* refactor(#1660): extract noOpObjectListSetError as pure tested helper (Stryker coverage)

cmdFrontmatterSet is not in Stryker's property/unit test set, so the inline no-op
detection added survivors that dropped the frontmatter module below its 62% mutation
threshold. Extract the detection into a pure exported helper noOpObjectListSetError and
unit-test every branch directly (changed content, scalar, scalar-array, null, dict
no-op). cmdFrontmatterSet now calls the helper. Same pattern as the #1572 spliceFrontmatter
coverage fix.
2026-06-24 13:39:17 -04:00
Tom Boucher
f615eb9ef3 fix(#1572): preserve must_haves object-lists across frontmatter set/merge (#1656)
* fix(#1572): preserve must_haves object-lists across frontmatter set/merge

spliceFrontmatter round-tripped the WHOLE frontmatter through extractFrontmatter
(a scalar-only parser) then reconstructFrontmatter (a lossy serializer), so any
must_haves object-list — artifacts {path, provides}, prohibitions {statement,
status} — was flattened to scalar strings and re-emitted as a malformed inline
array whenever an UNRELATED field changed, silently dropping every provides:/
status: value. The write now preserves the original raw text for any top-level
key whose value is structurally unchanged between the original parse and the new
object (generalizing the existing whole-document no-op guard to per-key
fidelity), and regenerates only the key that actually changed. The key set is
still defined by newObj (the cmdSet/cmdMerge flow always passes the full merged
object). spliceFrontmatter's only callers are cmdFrontmatterSet/Merge — the
STATE.md read-modify-write family calls reconstructFrontmatter directly and is
unaffected. Regression cases folded into tests/frontmatter-cli.test.cjs:
artifacts/prohibitions object-lists survive set and merge; idempotent on repeat
sets. Asserted via parseMustHavesBlock (the structure-preserving parser).

* chore(#1572): backfill changeset pr ref to 1656

* fix(#1572): fail-closed when set/merge would emit [object Object] (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that directly setting a must_haves
object-list (a CHANGED key) still routed through the lossy reconstructFrontmatter,
emitting literal "[object Object]" and destroying the data. The reported case
(mutating an UNRELATED field) was already fixed by per-key raw-text preservation,
but the changed-object-list path was still silently lossy. Add fail-closed: when a
regenerated key's text contains the "[object Object]" sentinel, spliceFrontmatter
throws — cmdFrontmatterSet/Merge error out WITHOUT writing, directing the user to
edit the file directly. The no-frontmatter (generate-from-scratch) path is guarded
the same way. Adds a test that a refused set leaves the file unchanged and the
original object-list intact. Codex finding #2 (a contrived flattened-projection
no-op) is a deeper limitation noted in the PR — non-destructive, and the fail-closed
message already directs users to edit object-list blocks directly.

* test(#1572): add spliceFrontmatter per-key preservation + fail-closed unit coverage

Stryker mutates gsd-core/bin/lib/frontmatter.cjs against tests/frontmatter.{property,unit}.test.cjs
(MinScore 62). The #1572 regression cases live in frontmatter-cli.test.cjs, which is NOT in
Stryker's test set, so the new functions (sliceTopLevelFrontmatterSegments, the per-key
preserve/regenerate/drop/append loop, regenerateFrontmatterKey's [object Object] fail-closed)
had surviving mutants that dropped the module below threshold. Add unit-level coverage in
frontmatter.unit.test.cjs exercising every new branch directly via spliceFrontmatter:
unchanged object-list preserved (provides survives) when a scalar sibling changes; changed
scalar regenerates only that key; orphan keys dropped; new keys appended; indented nested
block stays attached to its parent key; whole-document no-op returns input verbatim; both
fail-closed paths (changed object-list + no-frontmatter) throw.
2026-06-24 13:30:46 -04:00
Tom Boucher
b205e4c2b2 fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form

bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.

* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject

The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.

* chore(#1639): backfill changeset pr ref to 1665
2026-06-24 13:26:25 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
Tom Boucher
35478b615e refactor(#1646): route capability routers through Command Routing Hub per ADR-959 (#1647)
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959

Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.

src/cjs-command-router-adapter.cts
  * Imported ERROR_REASON from io.cjs.
  * UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
    as the second arg to error() — additive for host routers (their
    existing one-arg error callbacks ignore the second arg), required
    for capability routers whose tests assert reason === 'sdk_unknown_command'
    on the JSON-error envelope.

src/graphify-command-router.cts
  * Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing/invalid --budget) now
    return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
    instead of calling error() directly (Q2=C, Q4=ii from grilling).
  * Success handlers keep direct output() calls.
  * Subcommands array is alphabetical for byte-identical 'Available:'
    text in the unknown-subcommand message.
  * The unknown-subcommand path is now owned by the Hub's manifest
    check (the adapter passes SDK_UNKNOWN_COMMAND).

src/intel-command-router.cts
  * Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing filePath for patch-meta
    and extract-exports) return makeInvalidArgs Results.
  * Preserved the timeAgo mutation in the non-raw status handler.
  * Preserved the lazy require('./intel.cjs') inside the route function.

src/audit-command-router.cts
  * routeAuditUat: routes through the Hub with a synthetic 'run'
    defaultSubcommand (no real subcommands). Gives uniform observability.
  * routeAuditOpen: captures --json in a closure, strips it from args
    before Hub dispatch (so it isn't mistaken for a subcommand by the
    manifest check), then branches on wantJson inside the handler to
    preserve the formatAuditReport success-path quirk.

docs/CONFIGURATION.md
  * Observability section: noted capability commands (graphify, intel,
    audit-uat, audit-open) now emit DispatchEvent records since #1646.

.changeset/capability-routers-via-hub.md
  * Changed fragment describing the user-visible audit-trail expansion.
    pr:0 placeholder will be backfilled after gh pr create returns the
    real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).

Verification
  * graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
    error path, JSON-errors, and registry assertions)
  * intel cutover tests: 39/39 pass
  * audit cutover tests: 24/24 pass
  * bug-974-graphify-budget-missing-value regression test: pass
  * npm run test:unit (full suite): 2384 tests, 0 fail
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.

* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
2026-06-23 23:21:06 -04:00
Tom Boucher
6214039358 refactor(#1644): Hub extension — exitReason? field on InvalidArgs + adapter honestification (#1645)
Phase 1 of parent #1641. Implements the contract documented in the
Phase 0 ADR-0174 §5 amendment (#1642 / #1643).

src/command-routing-hub.cts
  * InvalidArgsResult interface gains optional exitReason?: string
    (carries an ERROR_REASON enum value, separate from reason which is
    the explanation text).
  * makeInvalidArgs(arg, reason, exitReason?) factory conditionally adds
    the field only when the third arg is truthy — preserves the strict-
    keys invariant tested at command-routing-hub.test.cjs:444.
  * _VARIANT_SCHEMA.InvalidArgs.allowed Set extended to include
    'exitReason' so the runtime validator does not coerce well-formed
    extended Results to HandlerFailure.

src/cjs-command-router-adapter.cts
  * Honestified the wrapper comment: the runtime check ('ok' in result)
    already passes any {ok:*} object through, so the historical
    {ok:true, data} return type was a lie for err Results. The lying
    cast is preserved because the Hub's export = syntax doesn't expose
    HubResult for import; the Hub's _validateErrResult runtime-validates
    the actual shape.
  * Result→error() translation branched: when InvalidArgs carries
    exitReason, the adapter calls error(result.reason, result.exitReason)
    so the JSON-error envelope (GSD_JSON_ERRORS=1) preserves the typed
    ERROR_REASON value. When exitReason is absent, error(msg) is called
    with exactly one arg — byte-identical with prior behavior.
  * RouteCjsCommandFamilyOptions.error and RouteHubCommandFamilyOptions
    .error callback types widened from (message) to (message, reason?)
    to match io.cts's actual error() signature.

CONTEXT.md
  * Command Routing Hub predicate updated to document the new field,
    factory signature, and dispatcher translation contract.

Tests (TDD red→green)
  * tests/command-routing-hub.test.cjs: 8 new tests covering 2-arg
    (strict-keys), 3-arg (key present), undefined, empty string, frozen
    result, hub.dispatch propagation, and validator acceptance.
  * tests/cjs-command-router-adapter.test.cjs: 2 new tests covering
    exitReason passed as second arg + byte-identical prior behavior when
    absent.

Verification
  * npm run test:unit: 2448 tests, 0 fail (no regressions)
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

Memtrace blast radius: LOW (get_impact makeInvalidArgs → 3 nodes; the
optional field is non-breaking for the 1 existing caller routePhaseCommand).
2026-06-23 22:47:57 -04:00
Tom Boucher
bcc5a6d1ba fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command

Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.

- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
  absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
  .sh and others keep the bare quoted path (unchanged).

Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.

Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.

* chore(#1634): backfill changeset pr:1638

* fix(#1634): resolve lint and windows CI failures

- validator: replace the control-character range regex with a char-code loop.
  The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
  codes are equally precise and lint-clean. Behavior unchanged (still rejects
  matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
  honor POSIX write modes (a 0o644 write reads back as 0o666), so the
  precondition is meaningless there and failed the windows-latest lane. The
  node-prefix assertion — the actual fix — is platform-independent and still
  runs everywhere.

* docs(#1634): amend ADR-894 for optional lifecycle hook matcher

The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.

* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md

Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
2026-06-23 21:49:43 -04:00
Joe Seymour
9d12725e4e fix(#1619): normalize pruned mise node execPath to the stable shim in normalizeNodePath (#1621)
* fix(#1619): normalize pruned mise node execPath to the stable shim

resolveNodeRunner() bakes process.execPath into managed .js hook commands.
Node realpaths execPath, so under mise it resolves to a concrete
<data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after
which every managed hook 404s — the same ephemeral-path failure #977 fixed
for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise
versioned install path to the stable sibling shim <data>/shims/node when it
exists (deriving <data> from execPath so a custom MISE_DATA_DIR works),
falling back to the raw execPath otherwise. Tests folded into
install.test.cjs per the regression test-name lint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(changeset): set pr number to 1621

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Joe Seymour <joese@iarx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:17:03 -04:00
Tom Boucher
0c4d570541 fix(#1628): type-safe config-set validation — close JSON-coercion enum bypass + enforce capability schema
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':

1. Missing guards: workflow.security_block_on (enum) and
   workflow.security_asvs_level (integer 1-3) had no store-time validation.

2. Systemic JSON-coercion bypass: every string-enum guard used
   VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
   before validation, String(["member"]) === "member" let a JSON array
   slip through and an array was stored in a scalar key. Reproduced on
   human_verify_mode, statusline.context_position, context_guard_mode,
   fallow.scope/profile, source_grounding_authority, drift_action, context.

3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
   25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
   including coerced arrays/objects and out-of-enum strings like
   code_review_depth=garbage — was stored silently.

Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 16:10:34 -04:00
Tom Boucher
658ea33cb6 fix(#1615): applySurface rewrites commands kind, not just skills
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes.

For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target).

Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode.

Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615.
2026-06-23 14:52:32 -04:00
Tom Boucher
4ed208e74b fix(#1615): validate commandName to prevent workflow prompt injection
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body.

Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message).

Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation.
2026-06-23 14:38:38 -04:00
Tom Boucher
527142ad2e fix(#1615): normalize Windows backslash paths in workflow content
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only.

Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX.

Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input.

Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring.
2026-06-23 14:26:44 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
1b95762661 Merge pull request #1568 from behruznassre/fix/1514-retired-phase-total-phases
fix(#1514): exclude retired/folded phases from progress.total_phases
2026-06-23 10:55:41 -04:00
Tom Boucher
da2de3a183 Merge branch 'next' into fix/1514-retired-phase-total-phases 2026-06-23 10:22:18 -04:00
Tom Boucher
cbd21092a9 fix(#1614): install Antigravity skills flat 2026-06-23 10:21:06 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Behruz Nassre Esfahani
0271231910 test(#1514): add fast-check property + all-retired boundary for retired parser
Per review (test-standard items):
- Property test (RULESET.TESTS.property-based-testing): extractRetiredPhaseNumbers
  is the parsing core, so add a fast-check property — k of n checklist phases
  struck → exactly the k canonical keys returned, across randomized phase counts
  and numeric/zero-padded/project-code ID forms. Exposed via a `_`-prefixed test
  seam (mirrors the existing _setLockProbes seams), no public API surface added.
- Boundary (RULESET.TESTS.boundary-coverage): all-retired case (k === n) →
  total_phases 0, via state json.

No production behavior change; the exclusion logic is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 17:22:29 -07:00
Behruz Nassre Esfahani
c6edeb3eb3 Merge branch 'next' into fix/1514-retired-phase-total-phases 2026-06-22 17:13:25 -07:00
Tom Boucher
3a06b4888a Merge branch 'next' into feat/1173-wire-agent-converters-descriptor 2026-06-22 13:04:46 -04:00
Tom Boucher
79d0e79c61 Merge branch 'next' into fix/1394-gemini-skill-tool-exclusion 2026-06-22 12:15:57 -04:00
Tom Boucher
1a16fa7ecc Merge branch 'next' into fix/1400-agent-skills-stdout-flush 2026-06-22 12:06:09 -04:00
Tom Boucher
b1297db50c Merge branch 'next' into fix/1531-core-lock-liveness 2026-06-22 11:56:04 -04:00
Behruz Nassre Esfahani
384b8e9d8b Merge branch 'next' into fix/1400-agent-skills-stdout-flush 2026-06-22 07:32:47 -07:00
Tom Boucher
b2c0086c1b fix(#1574): resolve review — copilot instruction file is .github/copilot-instructions.md
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
2026-06-22 09:56:15 -04:00
Tom Boucher
bf9bd1f4e0 fix(#1529): emit runtime-native instruction file from new-project 2026-06-22 09:32:29 -04:00
Tom Boucher
90f123434d Merge branch 'next' into fix/1394-gemini-skill-tool-exclusion 2026-06-22 07:56:58 -04:00
Behruz Nassre Esfahani
e3acdd89cb fix(#1514): exclude retired/folded phases from progress.total_phases
A retired/folded phase (struck through in ROADMAP, marked [x], with a
directory but no completion artifact) was counted in the total_phases
denominator via max(phaseDirs.length, roadmapPhaseCount), yet could never
satisfy the numerator (no SUMMARY → never "completed"), freezing shipped
milestones below 100% (e.g. 5/6 = 83%).

Both STATE counting paths now read the current-milestone ROADMAP scope and
exclude retired phases from BOTH the disk phase-dir set and the heading
count, so a retired phase counts toward neither denominator nor numerator:
  - buildStateFrontmatter (`state json`)
  - cmdStateSync (`state sync --verify` / rebuild) — previously re-derived the
    inflated denominator and reported "no drift", per the issue.

Retired detection (extractRetiredPhaseNumbers) is scoped to the lines that
canonically mark a phase retired — a checklist entry (`- [x] …`) or a phase
heading — and within those, only a struck span whose SUBJECT is the phase
(`~~**Phase 04: Delta**~~`). So struck prose, a struck goal line, and the fold
target ("folded into Phase 05") are not misread as retired.

Phase matching uses the canonical phase-id helpers (normalizePhaseName +
extractPhaseToken), so numeric, decimal, and project-code IDs (PROJ-42) match
consistently across ROADMAP tokens and on-disk dir names.

Scope boundaries (separate, pre-existing concerns left unchanged):
  - `roadmap analyze` (src/roadmap.cts) intentionally trusts the [x] checkbox
    (incl. externally-completed phases) — a different reporting surface.
  - cmdStateSync does not apply the milestone phase-dir filter (so 999.x /
    other-milestone dirs can still affect its count); that is the #1445 /
    milestone-filter axis, independent of retired phases.

Same counting family as #549 / #500 / #1445.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 23:29:05 -07:00
Joe
e12a2abfd8 feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit

Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on

- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
  seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
  slug, status, scope, trigger_when, planted, title. Optional case-insensitive
  status filter. User-controlled content is sanitized (sanitizeForDisplay) and
  every path validated (requireSafePath); read-only. Independent of
  audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
  renders the seed table.

Closes #441

* chore(#441): point changeset fragment at PR #722

* test(#441): allowlist list-seeds test in prompt-injection scan

The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.

* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow

Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.

* docs(#441): sync help full.md + INVENTORY for --list-seeds

Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.

* docs(#441): add --list-seeds how-to + drop phantom statuses

Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):

- USER-GUIDE.md Seeds section (how-to): extend the task to cover
  auditing parked seeds on demand via --list-seeds, including the
  status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
  from the list-seeds filter vocabulary; the system only produces
  dormant|active|triggered (src/audit.cts scanSeeds). Reference must
  be factually accurate and complete.

* fix(#441): guard non-scalar status frontmatter in cmdListSeeds

A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.

Adds regression coverage for empty and array `status:` and non-scalar fields.

Refs #441

* docs(#441): align list-seeds workflow status vocabulary

The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.

Refs #441

* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds

Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.

Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).

* test(#441): add fast-check property coverage and count=1 boundary for list-seeds

Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).

Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).

* chore(#441): sync runtime launcher snippet into list-seeds workflow

Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.

* test(#441): record list-seeds.md in workflow size baseline (#1074)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:59:41 -04:00
Tom Boucher
a570cd049c refactor(#1559): audit installer compatibility exports (#1565) 2026-06-22 00:38:55 -04:00
Tom Boucher
793fab0fd6 Merge branch 'next' into fix/1394-gemini-skill-tool-exclusion 2026-06-21 23:40:52 -04:00
Behruz Nassre Esfahani
0224f5bcf3 fix(#1383): resolve GSD version without a top-level require of the runtime-root package.json (#1409)
* fix(#1383): resolve GSD version without a top-level require of the runtime-root package.json

The extracted runtime-artifact-conversion module sits in the gsd-tools
loader chain and did a module-load `require('../../../package.json')`.
On Codex (whose runtime root has no package.json) that threw
`Cannot find module '../../../package.json'`, crashing every gsd-tools
command before it did anything. Even on Claude the synthetic
`{"type":"commonjs"}` has no `version`, so the sole consumer already
emitted `version: undefined`.

Resolve the version lazily and defensively instead: read the installed
gsd-core/VERSION, else lazily require the runtime-root package.json, else
degrade to '' so the caller omits the field. Both sources are validated
against the repo's semver-prefix convention (mirrors update-context.cts)
so a garbled VERSION is never emitted verbatim. install.js's dead
duplicate converter is intentionally left untouched (scoped to the crash).

Adds a #1383 regression block exercising resolveVersionFrom across
VERSION-only / package.json-only / neither / malformed-VERSION layouts,
asserting no-throw and the correct version string.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1383): add changeset for the Codex gsd-tools crash fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#1383): record resolveVersionFrom export in CONTEXT.md glossary

Maintainer review gate on PR #1409: the lazy resolveVersionFrom seam added on
the Runtime Artifact Conversion Module must be recorded in CONTEXT.md so the
canonical glossary doesn't drift from the exported surface.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1383): reword changeset to drop product-name parenthetical

product-name-purity (#1777) rejects 'Codex (…)' parentheticals that render
verbatim into CHANGELOG.md. Reword to a comma clause; no behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 23:40:48 -04:00
Rezolv
dee40cd392 fix(#1551): match dash-separated milestone phase IDs in roadmap analyze checklist scan (#1552)
* fix(#1551): match dash-separated milestone phase IDs in roadmap analyze checklist scan

The checklist scanner in cmdRoadmapAnalyze allowed only a dot separator
(?:\.\d+)* while the detail-heading scanner allows [.-], so milestone-prefixed
IDs (1-01) truncated at the dash (-> 1) and reported phantom missing detail
sections on every well-formed milestone roadmap. Widen the char class to
(?:[.-]\d+)* to match the detail scanner and the shared phaseMarkdownRegexSource
helper.

Fixes #1551

Claude-Session: https://claude.ai/code/session_01H96MxPGMJJUiJLV2NgzV16

* chore(changeset): Fixed fragment for #1552 (roadmap milestone-id checklist scan)

Claude-Session: https://claude.ai/code/session_01H96MxPGMJJUiJLV2NgzV16

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 23:30:27 -04:00
Dave
fc019b2688 fix(#1531): race-safe steal for the two core-path locks (PR #1532 review)
trek-e's review found the M1 PID-liveness backport dropped two pieces of
capability-lock.cts's steal-safety machinery, reopening the #500/#905/#1230
lost-update family:

- Empty-body window (state.cts): acquireStateLock creates the lock with O_EXCL
  and writes the pid in a separate writeSync; a lock observed in that gap has an
  empty body, reads as not-verified-live, and was stolen at age ~0 — robbing a
  holder mid-creation. Add a fresh-create floor scoped to the unverifiable-body
  case: an empty/unparseable body that is fresh is treated as mid-creation and is
  NOT stolen, while a COMPLETE dead-pid body is still stolen promptly (preserves
  the prompt-dead-steal contract). planning-workspace writes its body atomically
  (flag:'wx') so it has no empty-body window.

- Double-steal (both locks): the steal was a bare fs.unlinkSync with no identity
  re-confirm, so two waiters could both reclaim a dead holder and end up holding
  concurrently. Replace with an atomic renameSync (only one racer wins the inode)
  guarded by a (dev,ino,body) identity re-confirm immediately before the steal;
  body content is part of the identity to defeat inode reuse.

Tests (seam-driven, no wall-clock, each proven RED-before-GREEN):
- clock-seam: fresh empty-body lock is not stolen at age ~0; a racer-recreated
  live lock is not double-stolen (identity re-confirm). Adds a beforeSteal seam.
- planning-workspace: racer-recreated live lock is not double-stolen.
- Updated the two #1217 unlinkSync-failure tests to the renameSync steal path
  (the bounded-backoff/no-busy-spin guarantee is preserved and re-asserted).

Uncontended acquire path is byte-for-byte unchanged.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
2026-06-21 23:22:17 -04:00
Rezolv
9acd0208cd fix(#1542): roadmap upgrade rollback restores .planning regardless of git tracking (#1543)
* fix(core): roadmap upgrade rollback must restore .planning regardless of git tracking (#1542)

applyMigration rolled back a failed migration with git reset --hard + git clean
-fd .planning/phases/. For a commit_docs:false project (.planning gitignored —
the default) that restores NOTHING (reset ignores untracked, clean without -x
skips ignored), yet it threw 'Migration failed (rolled back to <sha>)' — a false
claim leaving .planning half-migrated. git reset --hard is also a whole-repo op.

Replace it with a surgical, git-independent rollback: record the exact renames
performed and snapshot each file before rewriting it, then on failure reverse the
renames and restore the snapshots (deleting files that did not previously exist).
Correct whether .planning is tracked or ignored; touches only what it changed.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1543 (roadmap upgrade surgical rollback)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* test(core): update bug-685 execSync count floor after surgical rollback (#1542)

The #1542 surgical, git-independent rollback removed the rev-parse/reset/clean
git execSync calls from roadmap-upgrade.cts, leaving only the git status
precondition. bug-685 asserted calls.length >= 4; lower the floor to >= 1 — the
durable guard (every remaining git execSync sets windowsHide:true) is unchanged.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 23:21:39 -04:00
Rezolv
3857912ff6 fix(#1540): platformWriteSync retries transient rename locks instead of truncating readers (#1541)
* fix(core): platformWriteSync must retry transient rename locks, not truncate readers (#1540)

platformWriteSync fell back to a non-atomic fs.writeFileSync(filePath) on ANY
error from the temp+rename path. On Windows, renameSync onto a target a reader
holds open throws EPERM/EBUSY/EACCES (the common case for hot files like
STATE.md), so the fallback fired and a concurrent reader saw the file
mid-truncation.

Mirror the capability-ledger rename-retry idiom: retry transient lock errnos
(EPERM/EBUSY/EACCES) with a bounded Atomics.wait backoff; on a persistent lock,
surface the error rather than do the truncating non-atomic write. Genuinely
unrenameable cases (EXDEV cross-device) and tmp-write failures still fall back.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1541 (platformWriteSync rename retry)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 23:14:43 -04:00
Rezolv
22e38f0b4a fix(#1538): roadmap upgrade must not process.exit inside the no-throw hub (#1539)
* fix(core): roadmap upgrade must not process.exit inside the no-throw hub (#1538)

The upgrade handler called process.stderr.write + process.exit(1) on an
unsupported --convention, structurally bypassing the command-routing-hub's
no-throw contract (ADR-0012). It also parsed only the space-separated
--convention <value> form, so --convention=<value> was silently dropped and
defaulted to milestone-prefixed, running the migration the user did not request.

Throw instead of exit (the hub converts to HandlerFailure and the adapter routes
it through error()); parse both --convention forms and fail closed on any
missing/unsupported value.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1539 (roadmap upgrade hub contract)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 23:04:34 -04:00
Rezolv
41bd333b30 fix(#1535): make punctuated adr-parser header synonyms reachable (#1536)
* fix: make punctuated CANONICAL_HEADERS synonyms reachable (audit M7)

classifyHeader received an already-normalized header (via normalizeAdrHeader,
which collapses [\s:._-]+ to a space and strips [^\w\s]) but compared it against
the RAW synonym strings. So any synonym carrying a hyphen/apostrophe ('trade-offs',
'non-goals', 'anti-goals', 'follow-up', 'cross-cuts', 'post-grilling', "how we'll
know", "won't do/have") could never match — its ADR section silently went unmapped.
Nine synonyms across six buckets were dead; the repo had characterization tests
pinning that quirk ('unreachable synonym').

Root-cause fix (per ADR-1372's 'compound, don't accrete' guidance): normalize BOTH
sides via a module-load-precomputed index, instead of pre-baking 9 normalized
literals into the data table. Closes the abstraction asymmetry once for all current
and future synonyms; the table stays human-readable. Also de-dupes 'trade-offs' from
considered_options so it no longer shadows risks once both normalize to 'trade offs'.

Insertion order preserved → first-match-wins + exact-then-prefix precedence byte-
identical for already-normalized synonyms; only the 9 dead synonyms gain matching.
Updates the 7 characterization tests to the corrected behavior and adds a
reachability+no-collision invariant test guarding the whole class against regression.

Note: the audit's M7 premise (trade-offs *misclassified into considered_options*)
was factually wrong — it was unmapped, and tested as such. adr-parser is CLI-only
(ADR-1372 T2), so the behavior change cannot reach in-process gates.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1536 (adr-parser punctuated synonyms)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:55:22 -04:00
Rezolv
75552f7ea0 fix(#1533): prototype-pollution guard in _deepMergeConfig (#1534)
* fix: prototype-pollution guard in _deepMergeConfig (audit M4)

The root↔workstream config merge iterated Object.keys(overlay) with no
__proto__/constructor/prototype guard, while four sibling paths in the same
file (lines ~315/319/331/341/549) guard them. A workstream/root config.json
with {"__proto__": {...}} could pollute the merged object's prototype chain
and spoof unset config flags (per-object, not global Object.prototype).

Adds the same three-key continue guard at the top of the overlay loop plus a
regression test for __proto__/constructor/prototype overlay keys.

Closes a gap missed by the closed config proto-pollution hardening
(#751/#1406/#663).

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1534 (config proto-pollution guard)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:46:34 -04:00
Enes Yağız
8748e95ed1 fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666)

Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:

- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
  changes, pushes, and opens a companion PR via `gh pr create`

All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.

Closes #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr number to 667

* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness

Resolves all three blockers and seven robustness issues raised in PR #667 review:

Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
  never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)

Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix

Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
  containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected

Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
  instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
  combinations and the execGit global-trim edge case uniformly

Tests: 17/17 pass, lint: 0 errors

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false

- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
  standalone bug-666-*.test.cjs into tests/commands.test.cjs under
  describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
  Adds allow-test-rule: source-text-is-the-product see #666 for the
  workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
  filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.

* fix(pr-branch): remove obsolete regression tests for sub-repos handling

* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout

- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
  (+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
  call in the handle_sub_repos dirty-scan (repo convention: every git
  subprocess is bounded, never hangs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): handle rename staging and split changedFiles from filesToStage

For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): rollback on push failure in cmdPrSubrepo

If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): do not rollback after commit on push failure; add push-fail regression test

Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.

Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): validate sub-repo paths before git invocation in pr-branch.md

The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.

Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.

Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.

Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#666): make sub-repo traversal scan test genuinely fail-first

The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md

Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.

- dirty-scan: realpathSync the root once, and realpathSync each candidate before
  the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
  containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
  against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
  backend must still be reported). Confirmed fails-first — regressing the scan to
  path.resolve makes the symlink case leak.

Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases

Pre-emptive hardening of the workflow changes from the symlink fix:

- continue-outside-loop: the "skip companion PR on seam failure" block used a
  bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
  loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
  prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
  (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
  creation lacks privileges) so it doesn't hard-fail on Windows CI.

Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:34:00 -04:00
Tom Boucher
94e7e3f88f refactor(#1558): plan runtime artifact uninstall removal (#1564) 2026-06-21 21:55:29 -04:00
Tom Boucher
405ae9b3b7 refactor(#1557): add runtime artifact install plan module (#1560) 2026-06-21 20:40:13 -04:00
Rezolv
195d356d7c Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio 2026-06-21 17:41:37 -04:00
Behruz Nassre Esfahani
faac9331f2 feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298
2026-06-21 15:38:44 -04:00