Commit Graph

241 Commits

Author SHA1 Message Date
Rezolv
75552f7ea0 fix(#1533): prototype-pollution guard in _deepMergeConfig (#1534)
* fix: prototype-pollution guard in _deepMergeConfig (audit M4)

The root↔workstream config merge iterated Object.keys(overlay) with no
__proto__/constructor/prototype guard, while four sibling paths in the same
file (lines ~315/319/331/341/549) guard them. A workstream/root config.json
with {"__proto__": {...}} could pollute the merged object's prototype chain
and spoof unset config flags (per-object, not global Object.prototype).

Adds the same three-key continue guard at the top of the overlay loop plus a
regression test for __proto__/constructor/prototype overlay keys.

Closes a gap missed by the closed config proto-pollution hardening
(#751/#1406/#663).

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1534 (config proto-pollution guard)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:46:34 -04:00
Enes Yağız
8748e95ed1 fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666)

Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:

- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
  changes, pushes, and opens a companion PR via `gh pr create`

All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.

Closes #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr number to 667

* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness

Resolves all three blockers and seven robustness issues raised in PR #667 review:

Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
  never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)

Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix

Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
  containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected

Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
  instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
  combinations and the execGit global-trim edge case uniformly

Tests: 17/17 pass, lint: 0 errors

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false

- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
  standalone bug-666-*.test.cjs into tests/commands.test.cjs under
  describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
  Adds allow-test-rule: source-text-is-the-product see #666 for the
  workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
  filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.

* fix(pr-branch): remove obsolete regression tests for sub-repos handling

* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout

- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
  (+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
  call in the handle_sub_repos dirty-scan (repo convention: every git
  subprocess is bounded, never hangs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): handle rename staging and split changedFiles from filesToStage

For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): rollback on push failure in cmdPrSubrepo

If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): do not rollback after commit on push failure; add push-fail regression test

Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.

Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): validate sub-repo paths before git invocation in pr-branch.md

The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.

Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.

Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.

Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#666): make sub-repo traversal scan test genuinely fail-first

The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md

Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.

- dirty-scan: realpathSync the root once, and realpathSync each candidate before
  the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
  containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
  against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
  backend must still be reported). Confirmed fails-first — regressing the scan to
  path.resolve makes the symlink case leak.

Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases

Pre-emptive hardening of the workflow changes from the symlink fix:

- continue-outside-loop: the "skip companion PR on seam failure" block used a
  bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
  loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
  prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
  (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
  creation lacks privileges) so it doesn't hard-fail on Windows CI.

Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:34:00 -04:00
Tom Boucher
94e7e3f88f refactor(#1558): plan runtime artifact uninstall removal (#1564) 2026-06-21 21:55:29 -04:00
Tom Boucher
405ae9b3b7 refactor(#1557): add runtime artifact install plan module (#1560) 2026-06-21 20:40:13 -04:00
Rezolv
195d356d7c Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio 2026-06-21 17:41:37 -04:00
Behruz Nassre Esfahani
faac9331f2 feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298
2026-06-21 15:38:44 -04:00
Tom Boucher
21fe9d3627 chore(#1544): test-quality + doc cleanups from the #1507 conversion-module review (#1547)
Follow-up hygiene from the #1507 epic review (release blocker fixed in #1537):

- path-replacement.test.cjs: delete the hand-reimplemented computePathPrefix copy
  (ADR-1508 said to); route all cases (incl. Windows + outside-home) through the
  real _computePathPrefix so the test can no longer drift from the function.
- enh-1511: add an isWindowsHost no-op tripwire characterization test.
- enh-1511: add a deterministic (fs-method monkeypatch, root/OS-independent)
  error-path test for applyRuntimeContentRewritesForCommandsInPlace — asserts the
  temp dir is rm'd on a read failure with no orphaned gsd-cmd-rewrites-* leak.
- runtime-artifact-conversion.cts: @internal note on rewriteStagedCommandBodies
  (deep-seam companion to rewriteStagedSkillBodies; no production caller today).
- ADR-1508 + CONTEXT.md: qualify the "single owner" claim with the deliberate
  opencode/kilo applyOpencodeFamilyPathPrefix pre-conversion carve-out (#784).

No user-facing behavior change (tests + internal docs + one code comment).
Refs #1507.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 15:20:46 -04:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Dave
c09c13f295 enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).

Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.

Closes #1346

Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
2026-06-21 10:44:48 -04:00
Tom Boucher
eb81faaeab refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay (#1513)
* refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay

Phase 2 of epic #1507 (ADR-1508). Behavior-preserving: makes the Runtime
Artifact Conversion Module the single owner of per-runtime content rewriting and
removes the last upward .cts -> bin/install.js dependency.

- src/runtime-artifact-conversion.cts now owns the engine (_applyRuntimeRewrites,
  5-arg with INJECTED attribution), the staged-content walkers
  (applyRuntimeContentRewritesInPlace / ...ForCommandsInPlace), computePathPrefix
  (private, exported as _computePathPrefix for tests), and the deep public seam
  rewriteStagedSkillBodies / rewriteStagedCommandBodies({runtime, configDir,
  scope, homedir?, platform?, resolveAttribution?}).
- src/surface.cts:applySurface calls rewriteStagedSkillBodies directly (no
  resolveAttribution -> undefined). Co-Authored-By is absent from ALL rewritten
  content, so processAttribution is vacuous there and undefined is provably
  behavior-identical. surface no longer imports getInstallExports.
- src/runtime-artifact-layout.cts: deleted getInstallExports / loadInstallExports
  / InstallExports + the GSD_TEST_MODE require('bin/install.js') relay.
- bin/install.js: binds computePathPrefix / the two walkers / _applyRuntimeRewrites
  from the conversion module (single implementation, exports preserved for Hyrum);
  install callsites pass getCommitAttribution(runtime) as the injected attribution.
  getCommitAttribution stays here (impure install-time config I/O).
- DEFECT.GENERATIVE-FIX guard: tests assert install.X === conversion.X reference
  identity for computePathPrefix + both walkers (no drift).

New tests/enh-1511-*.test.cjs (engine, attribution injection, deep seam, prefix,
layout-no-relay guard, reference-identity). 316 affected-suite tests green; lint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD

* test(#1511): make rewrite-engine path assertions Windows-robust

The deep seam normalizes paths as path.resolve(configDir).replace(/\\/g,'/')
and compares homedir().replace(/\\/g,'/'). Three assertions in the new test
rebuilt expected paths without that normalization, so they passed on Mac/Linux
but failed on Windows CI (PR #1513):

- two absolute-branch asserts rebuilt resolvedTarget via path.resolve(configDir)
  without the backslash→slash replace → mismatch on Windows.
- the $HOME-branch test fed a POSIX-literal /home/testuser, which Windows
  path.resolve re-roots onto the cwd drive (D:/home/...), so the
  resolvedTarget.startsWith(homeDir) check failed and the $HOME shorthand was
  never produced.

Fix is test-only (engine unchanged, still behavior-preserving): mirror the
engine's .replace(/\\/g,'/') in the two absolute-branch asserts, and use a real
absolute path (path.resolve(os.tmpdir(), ...)) + platform: process.platform for
the $HOME-branch test so the comparison holds on all platforms.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 23:59:50 -04:00
Tom Boucher
cadc944781 refactor(#1510): relocate getDirName + processAttribution out of bin/install.js (#1512)
Phase 1 of epic #1507 (ADR-1508): behavior-preserving relocation of the pure
rewrite-engine helpers out of the hand-authored installer so the conversion
module can own them without importing bin/install.js.

- getDirName -> src/runtime-name-policy.cts (pure runtime->dir-name switch).
- processAttribution -> src/runtime-artifact-conversion.cts (pure Co-Authored-By
  content transform).
- bin/install.js imports both back via destructure/binding and re-exports
  getDirName unchanged (Hyrum: install.test.cjs + runtime install tests import
  getDirName from bin/install.js).

Two refinements to ADR-1508's Phase 1 (verified against the source):
- getCommitAttribution STAYS in bin/install.js: it is impure install-time
  config I/O (reads runtime settings.json, uses install-time config-dir state +
  attributionCache), not a content transform. Phase 2 will inject the resolved
  attribution into the engine rather than move this function.
- The convertClaudeToAugmentMarkdown dedup is deferred to Phase 2's cleanup: the
  local copy is entangled with a converter cluster (convertSlashCommandsTo
  AugmentSkillMentions is used only by it; the family is partly dead-local via
  the ...runtimeArtifactConversion export spread), deduping only augment would be
  arbitrary, and it is not required to unblock Phase 2 (the engine will call the
  conversion module's own copy when it moves).

New tests/enh-1510-*.test.cjs: getDirName at its new home (all runtimes +
fallback + install.js re-export identity) and processAttribution
(null/undefined/string/$-escape/CRLF/global). 487 affected-suite tests green.


Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:29:15 -04:00
Tom Boucher
900462a59c Merge pull request #1505 from open-gsd/feat/1452-context-guard-mode
feat(#1452): add workflow.context_guard_mode — proactive context exhaustion guard for execute-phase
2026-06-20 18:53:24 -04:00
Tom Boucher
2dedbdd11c fix(#1455): resolve project-code-prefixed roadmap headings
* fix: remove hardcoded phase project-code prefix cap

Centralize project-code prefix stripping/matching and replace fixed {1,6} caps so long codes (for example MANIFOLD-117) resolve across phase, roadmap parser, roadmap upgrade, and validate flows.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: address review feedback on prefix parsing

- allow project_code prefixes with digits and underscores while preserving milestone parsing
- use shared optional project-code prefix source in phase dir parsing
- extend regression coverage for APP1/APP_1 prefixed phases

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: resolve review nit in phase-id test comment

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: resolve project-code-prefixed roadmap headings

Make getRoadmapPhaseInternal recover from drifted project-code-prefixed ROADMAP headings while preserving canonical bare-heading preference. Add init.phase-op and parser regressions for #1455 and guard gsd-roadmapper against emitting project_code in headings.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: add changeset for project-code phase fix

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* chore: update roadmapper agent size baseline

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix(#1455): tighten project-code prefix regex to [A-Z] start; add boundary test; document source order

Resolves blockers from review:
- Regex changed from [A-Z_][A-Z0-9_]* to [A-Z][A-Z0-9_]* so leading
  underscores (_FOO-7, _-7) are never misread as project-code prefixes;
  adds boundary test asserting both do NOT strip.
- Adds comment to roadmapPhaseLookupSources explaining why 3 sources are
  needed (order-dependent canonical-heading preference).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Solvely-Colin <211764741+Solvely-Colin@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:45:41 -04:00
Tom Boucher
4eca5ac96c feat(#1452): add workflow.context_guard_mode to guard execute-phase against context exhaustion
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:10:49 -04:00
Tom Boucher
6425a3cb72 fix(#1493): read workflow.drift_action/drift_threshold from nested config shape in verify.cts
loadConfig() returns a flattened object with no nested `workflow` key, so
config?.workflow was always undefined, making drift_action permanently 'warn'
and drift_threshold permanently 3 regardless of .planning/config.json. Fixes
by reading the raw config.json directly (matching the pattern in
check-command-router.cts:readWorkflowConfig). Adds two behavioral regression
tests that fail under the old code and pass under the fix.

Closes #1493

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 16:37:54 -04:00
Tom Boucher
4719b36d41 fix(#1445,#1446): exclude 999.x backlog from milestone totals; allow total_phases downward correction (#1490)
* fix(#1445,#1446): exclude 999.x backlog phases from milestone totals; allow total_phases downward correction

#1445: deriveProgressFromRoadmap (phase-lifecycle.cts), the roadmapPhaseCount
loop (state.cts), and getMilestonePhaseFilter (roadmap-parser.cts) all now
skip phase tokens matching /^999\b/ — consistent with the existing init.cts
filter. 999.x backlog dirs are consequently excluded from phaseDirs too.

#1446: shouldPreserveExistingProgress (state-document.cts) no longer includes
total_phases in its ratchet check. total_phases always takes the freshly
derived value; only completed_phases, total_plans, and completed_plans retain
ratchet behaviour.

Regression tests added for both bugs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #1445/#1446 progress-backlog-exclusion-and-ratchet

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1445,#1446): rename test files to fix-NNN convention; fix changeset pr: null

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:32 -04:00
Tom Boucher
fa1ffb4824 fix(#1437): add phase.list-plans to gsd-tools (#1485)
* fix(#1437): add phase.list-plans to gsd-tools

Register phase.list-plans in PHASE_COMMAND_ALIASES, implement
cmdPhaseListPlans in src/phase.cts (uses findPhaseInternal + scanPhasePlans
to return plan_count/has_plans/plans/phase_dir), and wire the handler in
phase-command-router. Previously every call produced "Unknown phase
subcommand".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1437): register new test file in lint-test-file-count allowlist

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1437): rename test to fix-NNN convention; update file-count allowlist

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:22 -04:00
Tom Boucher
ba3181204b fix(#1422,#1447): fix sub_repos/.git precedence; guard uncommitted data in new-milestone (#1484)
* fix(#1422): sub_repos config takes precedence over .git in findProjectRoot

When heuristic-3 (.git + parent .planning/) fires, do a lookahead walk
over ancestors above the matching parent to check if any further ancestor
has a sub_repos entry that explicitly claims the starting directory. If
found, return that ancestor instead — explicit config wins over the
implicit .git signal.

Regression tests updated: the heuristic-3 precedence test now asserts the
correct new behavior (sub_repos wins), and two additional coverage cases
added (startDir directly in sub_repo child, and startDir nested 2+ levels
inside the child).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1447): guard against uncommitted changes before deleting phase dirs in new-milestone

cmdPhasesClear now runs `git status --porcelain <phasesDir>` before
executing any rmSync. If uncommitted changes are detected it calls
error() and aborts, preventing silent data loss when new-milestone's
§6 "phases.clear --confirm" fires before the operator has archived
or committed outgoing phase work.

A new --force flag is added to bypass the guard for callers that have
already verified archival is done (or explicitly accept the loss).
When git is unavailable or the directory is not inside a git repo the
guard silently skips, preserving the existing behaviour for non-git
projects.

Five new regression tests cover: untracked files abort, staged-but-
uncommitted abort, --force bypasses, committed files pass, non-git
project passes without guard.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #1422 and #1447

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:16 -04:00
Tom Boucher
bb88a78faa fix(#1472,#1454): workstream-aware health paths; exclude active worktree from W017 (#1483)
* fix(#1472,#1454): validate health workstream-aware paths; exclude active worktree from W017

#1472: cmdValidateHealth now uses planningRoot(cwd) for shared-root files
(PROJECT.md, config.json, MILESTONES.md) and planningDir(cwd) for
workstream-scoped files (ROADMAP.md, STATE.md, phases/). Previously a
single planningDir() call was used for all paths, causing false
E002/E003/E004/W003 when GSD_WORKSTREAM is set.

#1454: W017 no longer fires for a stale worktree whose path equals or is
an ancestor of process.cwd(), preventing advice to remove the active
session's own worktree.

Regression tests added for both bugs; all 40 existing health tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:09 -04:00
Tom Boucher
7c93d9e222 feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 10:50:12 -04:00
Tom Boucher
08d1c57d6e fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 09:10:04 -04:00
Tom Boucher
f7c6ce7a1f fix(#1461): loader never crashes on a malformed overlay (skip-with-warning); bound the source fetch
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 03:24:47 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
c866ac1b24 fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) 2026-06-19 19:55:10 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
34bc096ec2 feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI

ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs
install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a
user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the
six subcommands, dispatching to the existing lifecycle/ledger:

- install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]…
- update [<id>|--all] [--scope] [--yes] [--shared-file]  (re-resolves recorded source)
- remove <id> [--purge-data] [--scope]  (first-party rejected)
- list [--json]  (first-party + overlay, both scopes, JSON array)
- disable|enable <id>  (activation-state alias of capability set --off/--on)

Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home,
project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json).
Consent is non-interactive: --yes grants; without it an executable install aborts after
printing the disclosure and writes nothing. Best-effort reconcile before each mutation.

Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs,
GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip,
remove round-trip + first-party guard, disable/enable, unknown subcommand.
Docs: docs/reference/gsd-capability-command.md reconciled to the real surface
(ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned);
docs/COMMANDS.md gains the gsd capability entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug

Adversarial-review (Codex) fixes:
- capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate
  fail-closes on it (was silently downgrading to permissive)
- installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection
  (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't
  act on a different id if the recorded source was retargeted
- capability update: prints the consent disclosure, exits non-zero on --all partial failure,
  no longer masks the resolved id
- capability remove: ledger-first ordering so an overlay is removable even if it shadows a
  first-party name; first-party guard only fires for ids not in the ledger
- gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle
  not yet wired through this path)

Silent-output bug (root cause, not waved off as pre-existing):
- captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw
  command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of
  stdout. Now it flushes the captured buffer before re-throwing (exit code preserved).
- cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws
  ExitError so the wrapper flushes — matches the repo's no-process-exit architecture.
- Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout.

Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed

- confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors
  safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it.
- mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or
  another capability's) server is skipped, so install/remove can't silently clobber user MCP config
  (hooks already append; the map-keyed mcpServers path was the gap).
- capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead
  of silently downgrading the strict_known_registries policy to permissive.
- Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry
  preserved; unparseable config blocks an external install. capability suite 83/83, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc

- install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never
  fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per
  the lifecycle contract; latent today, hardened for future status additions).
- Clarify capResolveScope comment (project scope === already-resolved cwd) and document that
  strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide
  allowlist) in gsd-capability-command.md.
- Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that
  looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1451): backfill changeset PR number → #1457

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:02:36 -04:00
Tom Boucher
1abebbf4fd feat(#1434): registry-driven dispatch for third-party capabilities (ADR-1244 Phase 5) (#1450)
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.

Closes #1434.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:37:43 -04:00
Behruz Nassre Esfahani
c26947808a feat(#1269): expand same-prefix numeric ID ranges in --phase-req-ids (#1419)
* feat(#1269): expand same-prefix numeric ID ranges in --phase-req-ids

normalizePhaseReqIds treated a range token like "SEL-01..SEL-03" as a single
literal ID, so gap-analysis reported the range string as a missing requirement
even when SEL-01/02/03 existed individually.

Add a per-token range expander (run AFTER the existing split, preserving the
string[] | null | undefined contract): a token matching <PREFIX>-<NN>..<PREFIX>-<MM>
with identical prefixes, ascending bounds, and EQUAL digit width expands to the
individual IDs preserving that width; anything ambiguous stays literal
(fail-closed). Differing-width bounds stay literal so the expander never invents
a zero-padding the author didn't type, and ranges beyond MAX_PHASE_REQ_RANGE
(1000) stay literal as a DoS guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1269): add changeset for --phase-req-ids range expansion

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1269): mark changeset docs-exempt (internal flag, no user docs surface)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1269): isolate the DoS-cap branch with a same-width range; doc nits

Review fixes: the AC4 DoS test used REQ-1..REQ-100000, whose differing digit
widths trip the width guard before the cap is reached. Use a same-width
REQ-0001..REQ-1001 (span 1001 > 1000) so the test actually exercises the cap.
Clarify the PHASE_REQ_RANGE_RE capture-group JSDoc and the property-test width
assertion comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-18 21:59:10 -04:00
Rezolv
dcceb1a004 fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing

resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first
bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also
had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy
dir (probed first). Regression from #217 — pre-#217 returned
~/.gemini/antigravity unconditionally.

Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested
descriptor: a two-pass probe prefers the candidate GSD installed into, then
bare existence, then probe[0]. Behavior is byte-identical when probeMarker is
absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for
installer/operator guidance on already-misinstalled users (auto-relocation
ruled out per ADR-0008's single-configDir migration bound).

Regression tests fail before / pass after: coexistence + marker-priority
cases, end-to-end through the registry descriptor.

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB

* chore(#1441): add changeset for antigravity resolver fix

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
2026-06-18 21:47:18 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
1abb0d3427 feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3)

ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability
install command + consent gate are Phase 4). NEW modules only; not yet wired into
bin/install.js.

- src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one
  adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/
  git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install),
  tarball (https download + sha512 integrity-before-extraction + tar), registry (stub).
  SECURITY: install never executes capability code (copy/extract only, --ignore-scripts);
  integrity verified before staging; symlink members rejected at interior, source-root,
  AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction;
  npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell);
  git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging
  (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd
  pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet.
- src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest:
  readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent,
  prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against
  hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir).
- 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex
  adversarial rounds; all 12 findings fixed + regression-tested.
- .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored
  (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen.

Closes #1432

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1432): add changeset for capability source resolver + ledger

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 16:23:20 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
d220ba6fd6 fix(#1414): resolve project root from descendant via nearest ancestor .planning/ (Resolution Provenance P1) (#1423)
Add heuristic (4) to findProjectRoot: after heuristics (1)-(3) (sub_repos,
multiRepo, .git+parent-.planning/) are exhausted without a match, perform a
second bounded walk-up within FIND_PROJECT_ROOT_MAX_DEPTH to locate the
nearest ancestor directory containing a .planning/ subdirectory. Returns that
ancestor as the project root, so loadConfig finds the correct config instead
of falling through to defaults when gsd-tools is invoked from a plain
descendant subdirectory of a single-repo project (#1366 cwd-drift gap).

Ordering is load-bearing: the new walk runs AFTER the existing loop so
sub_repos workspaces (where a child sub-repo may have its own .planning/)
still resolve correctly to the parent workspace. The existing own-.planning/
guard (heuristic 0, #1362) and the depth bound (FIND_PROJECT_ROOT_MAX_DEPTH=10)
are both preserved unchanged.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 00:50:30 -04:00
Tom Boucher
666d933e16 ci(#1401): add no-adhoc-markdown-parsing rule (fence-strip + section-collect) + grandfather burn-down (#1402)
Tightens the over-broad heading-walk detection: removes heading-walk
entirely and narrows fence-regex to require a multiline body ([\s\S]),
so single-line tests like /^```/ and /^###\s+/ are no longer flagged.
Grandfathers the 10 genuine section-collect sites across state.cts,
milestone.cts, audit.cts, and phase-lifecycle.cts with concrete reasons.
Adds 12 RuleTester tests (3 positive, 9 negative) to eslint-rules.test.cjs.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 20:14:12 -04:00
Tom Boucher
ce8bcb1b95 refactor(#1398): migrate state.cts section-collects onto markdown-sectionizer seam, byte-identical STATE.md (epic #1372 T6) (#1399)
* refactor(#1398): migrate state.cts section-collect regexes onto markdown-sectionizer seam (epic #1372 T6)

Replace ≈18 hand-rolled `/(heading)([\s\S]*?)(?=stop)/` regex splices in
state.cts with direct `tokenizeHeadings` calls that compute the exact
[bodyStart, stopOffset) span, preserving byte-identical STATE.md output.

Non-migratable site left in place: `cmdStateRecordMetric`'s metricsPattern
captures table-header rows in group 1 — not a standard heading+body shape.
Write orchestration (readModifyWriteStateMd / syncStateFrontmatter /
shouldPreserveExistingProgress / #952 no-op guard) is UNTOUCHED.

Verification: t6-headtohead.cjs head-to-head harness runs 25 ops across 6
STATE.md fixture variants (inline, trailing-blanks, CRLF, no-frontmatter,
nested-acc, post-milestone) against origin/next and reports 0 diffs.
No-op guard confirmed: record-session on recorded:false leaves STATE.md
byte-identical. All 150 state tests and 62 milestone/forensics tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1398): capture T6 state-section-splice characterization; drop throwaway harness

Remove scripts/t6-headtohead.cjs (committed throwaway HEAD-vs-origin/next
byte-compare harness). Capture its coverage as 23 behavioral characterization
tests appended to tests/state.test.cjs, exercising all migrated cmdState*
write-ops across 7 fixture variants (inline, trailing-blanks, CRLF,
no-frontmatter, nested-acc, no-current-pos, post-milestone). Includes the
#952 no-op guard (recorded:false + byte-unchanged assertion) and
CRLF/trailing-blanks edge-case coverage. Also removes the dead
spliceStateSection helper (defined but never called) that was surfacing as
an @typescript-eslint/no-unused-vars warning.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 17:56:17 -04:00
Tom Boucher
308c7be481 refactor(#1396): T5 — migrate uat-predicate + uat onto markdown-sectionizer seam (#1397)
- uat-predicate.cts: replace local _stripFencedBlocks (and its private
  FenceState/StripFencedResult types) with a call to stripFencedCode from
  markdown-sectionizer.cjs (ADR-1372 T5). Both stripFalsePositiveContexts
  step (c) and analyzeMarkdown now route through the seam. The three other
  passes in stripFalsePositiveContexts — frontmatter strip, HTML-comment
  strip, blockquote-line filter — remain caller-side (seam does not do these).
  The unterminatedFence signal consumed by analyzeMarkdown is preserved; it
  is now returned by stripFencedCode (same machine, same contract).

- uat.cts: migrate the ## Current Test, ## Tests, and ## Human Verification
  section-collect patterns onto collectSection/tokenizeHeadings from the seam.
  UAT-specific item parsing (### N. Name blocks, expected/result fields,
  categorization logic) stays caller-side. The HTML-comment strip within the
  Current Test body remains caller-side (UAT document structure, not seam scope).

- tests/markdown-sectionizer.test.cjs: remove the 18-case tautological parity
  guard (DEFECT.GENERATIVE-FIX). Once uat-predicate imports the seam the guard
  compares the seam to itself — removing it is the T5 commitment per ADR-1372.

4-space-indent behavior change (CommonMark correctness improvement): the seam
uses /^( {0,3})/ (CommonMark §4.5 ≤3-space indent); the retired
_stripFencedBlocks used /^(\s*)/ (any indent). A 4-space-indented ``` is no
longer treated as a fence opener (it is an indented code block per CommonMark).
Head-to-head over 9 corpus inputs: 0 diffs on all standard cases; 2 diffs only
on the synthetic 4-space-indent edge cases. No UAT fixture or test in the suite
exercises 4-space-indented fences. The change is a correctness improvement.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 15:44:11 -04:00
Tom Boucher
6e9f8bf50b refactor(#1393): migrate roadmap-parser onto markdown-sectionizer seam (epic #1372 T4) (#1395)
Removes all three inline copies of the fenced-code state machine from
src/roadmap-parser.cts and replaces them with calls to
tokenizeHeadings() from the canonical markdown-sectionizer seam
(ADR-1372 T4).

Changes:
- Drop stripFencedLines() function (copy 1 of 3 — standalone helper)
- Rewrite computeSectionEnd() using tokenizeHeadings() offsets into
  the original content (copy 2 — inline fence loop); the returned
  character offset is preserved exactly for all inputs
- Rewrite getMilestonePhaseFilter() versionOverride fence loop using
  tokenizeHeadings() (copy 3 — inline); sectionEnd offset preserved
- Replace stripFencedLines(roadmap) + unanchored phasePattern.exec()
  with tokenizeHeadings(roadmap) filtered by level and phase-heading
  pattern; headings in inline HTML comments (<!-- ## Phase N: -->)
  are no longer mis-counted (anchored ATX detection is correct)
- Add import { tokenizeHeadings } from ./markdown-sectionizer.cjs

Offset preservation: tokenizeHeadings() records h.offset as the
character index of '#' in the ORIGINAL content; computeSectionEnd()
and the versionOverride path both use h.offset directly as the
section-end character offset — no stripping, no shift.

Corpus head-to-head: 29/30 slots are byte-identical to origin/next.
The 1 diff (getMilestonePhaseFilter phaseCount for the HTML-comment
fixture: 3→2) is an improvement: the old unanchored regex counted
"## Phase 998:" embedded in "<!-- ## Phase 998: ... -->" mid-line;
tokenizeHeadings() correctly requires '#' at line-start (ATX rule).
No existing test asserts on that count; feat-3594 passes unchanged.

All 122 roadmap/milestone/phase tests green.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 15:13:14 -04:00
Tom Boucher
3a1961ebae refactor(#1390): migrate check-command-router + gap-checker onto markdown-sectionizer seam (epic #1372 T3) (#1392)
- check-command-router: stripCommentsAndFences delegates fenced-code
  stripping to seam's stripFencedCode; HTML-comment stripping stays
  caller-side. extractPlanDesignatedSections replaces hand-rolled
  split(/\r?\n/) + /^#{1,6}\s+/ heading walk with collectSections
  driven by DESIGNATED_HEADINGS_RE.

- gap-checker: parseRequirements checkbox-bullet detection migrates
  to iterateBullets (checkbox markers); **ID** extracted caller-side.
  Table-row path and separator-row skip stay caller-side.

- Adds refactor-1390-t3-characterization.test.cjs (43 behavioral
  tests) that were green before and remain green after.

- T1 fail-loud / could-not-parse gate semantics unchanged (verified
  by decisions.test.cjs 59/59 and direct gate invocation).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 14:32:43 -04:00
Tom Boucher
75c2e259d0 refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (epic #1372 T2) (#1388)
* refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (T2)

Replace hand-rolled parseSections (split/heading-regex walk/body accumulation)
with collectSections(content, () => true) from the canonical seam. Replace
hand-rolled splitEntries bullet-strip regex with iterateBullets from the seam,
preserving the plain-text-line fallback for byte-identical output. Removes the
last inline heading/bullet scanning from adr-parser.cts; normalizeAdrHeader and
all ADR-specific classification logic are unchanged. 218/218 tests pass before
and after.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): keep splitEntries flat (iterateBullets changed its contract); seam adoption stays in parseSections

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): drop dead preamble reconstruction; add targeted adr-parser mutation tests

parseSections' preamble reconstruction block (heading: null entry) was dead code:
both consumers (parseAdrMarkdown and parseStatusFromSections) skip heading: null
sections immediately on entry. Confirmed via analytical trace and zero-diff corpus
head-to-head across all 44 docs/adr/*.md files.

Adds 11 targeted behavioral tests to kill cheap surviving mutants:
- pushUnique intra-values dedup (kills seen.add removal mutant)
- body split/join round-trip with multi-line prose and entries
- parseStatusFromSections [0] indexing (only first line determines status)
- classifyHeader equality vs prefix-match boundary (exact match, prefix match,
  synonym+letter non-match)
- goal section prose vs entries distinction (bullet markers preserved in context)
- normalizeAdrHeader non-word char removal (parens and slash behavior)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 13:45:05 -04:00
Tom Boucher
e58b5e1721 fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof)

Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
  (these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
  for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate

#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.

#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.

parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.

Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D

FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.

FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.

FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).

FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.

Adds 14 behavioral regression tests (fail-first verified manually before fixes).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions

Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.

Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.

Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1364): backfill changeset PR number (1386)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 12:35:38 -04:00
Tom Boucher
7ca8011cd9 refactor(#1373): add canonical markdown-sectionizer seam (epic #1372 T0) (#1381)
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0)

Establishes the canonical markdown-structure parsing seam per ADR-1372.
No existing parsers are modified; this is the foundational T0 tier only.

- docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the
  seam interface, the tiered migration plan (T0-T7), and the prohibition
  enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7).
- src/markdown-sectionizer.cts: Pure module, Node built-ins only.
  Exports: stripFencedCode (CommonMark-correct state machine ported from
  uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence
  signal), tokenizeHeadings (ATX headings outside fenced blocks),
  collectSections (line-by-line predicate-driven section collection),
  collectSection (single named section, levelBounded stop, optional
  stripFences), iterateBullets (dash/checkbox/numbered + continuation).
- tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering
  the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences,
  unterminated fences, nested levels, all bullet markers, continuation
  lines, empty/non-string input) plus 4 fast-check property tests
  (idempotence, output shape, never-throws, length monotonicity).
- CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate).

Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory

- src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets;
  add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks,
  tagName regex-escaped, caller decides fence-stripping) and replaceSection(content,
  section, newBody) (pure character-offset splice for read-modify-write callers);
  update collectSections/collectSection to populate bodyStart/bodyEnd.
- tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for
  extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard
  that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a
  shared 9-item corpus; documents the known 4-space-indent divergence.
- docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and
  replaceSection in §"The seam".
- CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports.
- docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between
  loop-resolver and milestone).
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage

FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both
collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of
the raw stop-line offset, eliminating the trailing-newline overcounting that caused
replaceSection to drop separator newlines (## A\nbody## B gluing bug).

FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty
ATX headings (## / ##   ), text=''. 4-space indent correctly excluded.

FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose
level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections
that also stop at ### without abusing levelBounded.

FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark).
Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences
unaffected.

FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section
doc comments.

FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks
the non-greedy close-at-first-</tag> behavior.

FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so
tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3).

Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish)

The seam's compiled artifact must be a build-at-publish output like every other
src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed
file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 11:14:17 -04:00
Tom Boucher
66085d0080 fix(#1374): surface diagnostic when configured agent skills all fail to resolve (#1376)
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve

buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.

Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1374): backfill changeset PR number (#1376)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:16:51 -04:00
Tom Boucher
120f85164b feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams

GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.

- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
  pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
  resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
  registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
  `query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
  SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): add changeset for teams-detect guard

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)

The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): register teams-status.cjs in the inventory manifest

The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:52:23 -04:00
Tom Boucher
c03f3cc6af fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape (#1363)
* fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape

reconcileCodexHooksJsonEvent preserved whatever shape it read, so on an empty,
absent, or legacy top-level hooks.json it wrote top-level event keys
(`{ "SessionStart": [...] }`) that current Codex (deny_unknown_fields) rejects,
instead of the canonical `{ "hooks": { "SessionStart": [...] } }`.

- Lift any top-level event arrays (legacy, empty, or mixed nested+top-level)
  into the nested `hooks` table, merging same-named events so no user/legacy
  entry is dropped and no stray top-level event key survives. Mirrors
  reconcileCursorHooksJson.
- Collapse an empty hook table back to `{}` so removal on an absent file does
  not write a spurious `{ "hooks": {} }`.
- Read path still tolerates both shapes; dedup/removal unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1348): add changeset for Codex hooks.json canonicalization

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:51:11 -04:00
Tom Boucher
c53fd1f654 fix(#1324): resolve glued phase tokens (#1353) 2026-06-16 21:55:58 -04:00
Tom Boucher
aab26c7bf4 fix(#1343): parse decision bullets with text before the colon (#1358)
* fix(#1343): parse decision bullets with text before the colon

parseDecisions() silently dropped any `- **D-NN ...:**` decision bullet
whose header had freeform text (a parenthetical, em-dash, or prose) before
the `:**`, so the blocking check.decision-coverage-plan gate computed
coverage over a narrowed set and reported a false pass.

- Broaden bulletRe to tolerate a freeform run before the colon while
  preserving the optional [bracket] tag capture (drives `trackable`).
- Add a parse-miss guard: a line that looks like a D-NN bullet but still
  fails the regex flushes the current decision and warns instead of
  vanishing — the gate-integrity floor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1343): add changeset for decision-coverage false-pass fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1343): relocate decision-parser regression into owning module test file

CI's lint-regression-test-names bans new bug-NNNN-*.test.cjs files. Move the
9 regression cases from tests/bug-1343-parsedecisions-drop.test.cjs into the
owning parser test file tests/post-planning-gaps-2493.test.cjs (which already
exercises parseDecisions) and delete the banned file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:51 -04:00
Rezolv
00c05eb717 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 17:35:55 -04:00