* fix: prototype-pollution guard in _deepMergeConfig (audit M4)
The root↔workstream config merge iterated Object.keys(overlay) with no
__proto__/constructor/prototype guard, while four sibling paths in the same
file (lines ~315/319/331/341/549) guard them. A workstream/root config.json
with {"__proto__": {...}} could pollute the merged object's prototype chain
and spoof unset config flags (per-object, not global Object.prototype).
Adds the same three-key continue guard at the top of the overlay loop plus a
regression test for __proto__/constructor/prototype overlay keys.
Closes a gap missed by the closed config proto-pollution hardening
(#751/#1406/#663).
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* chore(changeset): Fixed fragment for #1534 (config proto-pollution guard)
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(pr-branch): handle sub_repos from config with git -C (#666)
Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:
- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
changes, pushes, and opens a companion PR via `gh pr create`
All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.
Closes#666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: update changeset pr number to 667
* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness
Resolves all three blockers and seven robustness issues raised in PR #667 review:
Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)
Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix
Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected
Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
combinations and the execGit global-trim edge case uniformly
Tests: 17/17 pass, lint: 0 errors
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false
- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
standalone bug-666-*.test.cjs into tests/commands.test.cjs under
describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
Adds allow-test-rule: source-text-is-the-product see #666 for the
workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.
* fix(pr-branch): remove obsolete regression tests for sub-repos handling
* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout
- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
(+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
call in the handle_sub_repos dirty-scan (repo convention: every git
subprocess is bounded, never hangs).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): handle rename staging and split changedFiles from filesToStage
For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): rollback on push failure in cmdPrSubrepo
If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): do not rollback after commit on push failure; add push-fail regression test
Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.
Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): validate sub-repo paths before git invocation in pr-branch.md
The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.
Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.
Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.
Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#666): make sub-repo traversal scan test genuinely fail-first
The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md
Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.
- dirty-scan: realpathSync the root once, and realpathSync each candidate before
the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
backend must still be reported). Confirmed fails-first — regressing the scan to
path.resolve makes the symlink case leak.
Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases
Pre-emptive hardening of the workflow changes from the symlink fix:
- continue-outside-loop: the "skip companion PR on seam failure" block used a
bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
(try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
creation lacks privileges) so it doesn't hard-fail on Windows CI.
Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Follow-up hygiene from the #1507 epic review (release blocker fixed in #1537):
- path-replacement.test.cjs: delete the hand-reimplemented computePathPrefix copy
(ADR-1508 said to); route all cases (incl. Windows + outside-home) through the
real _computePathPrefix so the test can no longer drift from the function.
- enh-1511: add an isWindowsHost no-op tripwire characterization test.
- enh-1511: add a deterministic (fs-method monkeypatch, root/OS-independent)
error-path test for applyRuntimeContentRewritesForCommandsInPlace — asserts the
temp dir is rm'd on a read failure with no orphaned gsd-cmd-rewrites-* leak.
- runtime-artifact-conversion.cts: @internal note on rewriteStagedCommandBodies
(deep-seam companion to rewriteStagedSkillBodies; no production caller today).
- ADR-1508 + CONTEXT.md: qualify the "single owner" claim with the deliberate
opencode/kilo applyOpencodeFamilyPathPrefix pre-conversion carve-out (#784).
No user-facing behavior change (tests + internal docs + one code comment).
Refs #1507.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs
Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).
Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
never calls it — so a real `--codex`/`--cursor`/etc. install emitted
`--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
workflow ran executors unisolated against the main checkout. (#1515/#1519 were
also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
Claude default.
Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
claude`; called from both `_applyRuntimeRewrites` and, crucially,
`copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
everything-else -> inline` (research: only Codex can background-nest the
pipeline's subagents; all others run inline, which they support).
Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.
Closes#1521
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1521): backfill changeset PR number (#1537)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees
A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:
1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
without `--raw`, so config-get's JSON-quoted output ("codex") was captured
verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
the Codex fail-closed guard was dead even when runtime:codex was explicit,
and Claude's own worktree degrade-check was dead too. Add `--raw` to those
reads across execute-phase, autonomous, manager, diagnose-issues, quick.
2. The conversion engine emitted `--default claude` for every runtime. Stamp
the codex-emitted workflows to `--default codex` (runtime) and
`--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
a neutral config on a Codex install resolves runtime=codex / worktrees off.
Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).
Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.
Closes#1515
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1515): backfill changeset PR number (#1519)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).
Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.
Closes#1346
Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
* refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay
Phase 2 of epic #1507 (ADR-1508). Behavior-preserving: makes the Runtime
Artifact Conversion Module the single owner of per-runtime content rewriting and
removes the last upward .cts -> bin/install.js dependency.
- src/runtime-artifact-conversion.cts now owns the engine (_applyRuntimeRewrites,
5-arg with INJECTED attribution), the staged-content walkers
(applyRuntimeContentRewritesInPlace / ...ForCommandsInPlace), computePathPrefix
(private, exported as _computePathPrefix for tests), and the deep public seam
rewriteStagedSkillBodies / rewriteStagedCommandBodies({runtime, configDir,
scope, homedir?, platform?, resolveAttribution?}).
- src/surface.cts:applySurface calls rewriteStagedSkillBodies directly (no
resolveAttribution -> undefined). Co-Authored-By is absent from ALL rewritten
content, so processAttribution is vacuous there and undefined is provably
behavior-identical. surface no longer imports getInstallExports.
- src/runtime-artifact-layout.cts: deleted getInstallExports / loadInstallExports
/ InstallExports + the GSD_TEST_MODE require('bin/install.js') relay.
- bin/install.js: binds computePathPrefix / the two walkers / _applyRuntimeRewrites
from the conversion module (single implementation, exports preserved for Hyrum);
install callsites pass getCommitAttribution(runtime) as the injected attribution.
getCommitAttribution stays here (impure install-time config I/O).
- DEFECT.GENERATIVE-FIX guard: tests assert install.X === conversion.X reference
identity for computePathPrefix + both walkers (no drift).
New tests/enh-1511-*.test.cjs (engine, attribution injection, deep seam, prefix,
layout-no-relay guard, reference-identity). 316 affected-suite tests green; lint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
* test(#1511): make rewrite-engine path assertions Windows-robust
The deep seam normalizes paths as path.resolve(configDir).replace(/\\/g,'/')
and compares homedir().replace(/\\/g,'/'). Three assertions in the new test
rebuilt expected paths without that normalization, so they passed on Mac/Linux
but failed on Windows CI (PR #1513):
- two absolute-branch asserts rebuilt resolvedTarget via path.resolve(configDir)
without the backslash→slash replace → mismatch on Windows.
- the $HOME-branch test fed a POSIX-literal /home/testuser, which Windows
path.resolve re-roots onto the cwd drive (D:/home/...), so the
resolvedTarget.startsWith(homeDir) check failed and the $HOME shorthand was
never produced.
Fix is test-only (engine unchanged, still behavior-preserving): mirror the
engine's .replace(/\\/g,'/') in the two absolute-branch asserts, and use a real
absolute path (path.resolve(os.tmpdir(), ...)) + platform: process.platform for
the $HOME-branch test so the comparison holds on all platforms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 1 of epic #1507 (ADR-1508): behavior-preserving relocation of the pure
rewrite-engine helpers out of the hand-authored installer so the conversion
module can own them without importing bin/install.js.
- getDirName -> src/runtime-name-policy.cts (pure runtime->dir-name switch).
- processAttribution -> src/runtime-artifact-conversion.cts (pure Co-Authored-By
content transform).
- bin/install.js imports both back via destructure/binding and re-exports
getDirName unchanged (Hyrum: install.test.cjs + runtime install tests import
getDirName from bin/install.js).
Two refinements to ADR-1508's Phase 1 (verified against the source):
- getCommitAttribution STAYS in bin/install.js: it is impure install-time
config I/O (reads runtime settings.json, uses install-time config-dir state +
attributionCache), not a content transform. Phase 2 will inject the resolved
attribution into the engine rather than move this function.
- The convertClaudeToAugmentMarkdown dedup is deferred to Phase 2's cleanup: the
local copy is entangled with a converter cluster (convertSlashCommandsTo
AugmentSkillMentions is used only by it; the family is partly dead-local via
the ...runtimeArtifactConversion export spread), deduping only augment would be
arbitrary, and it is not required to unblock Phase 2 (the engine will call the
conversion module's own copy when it moves).
New tests/enh-1510-*.test.cjs: getDirName at its new home (all runtimes +
fallback + install.js re-export identity) and processAttribution
(null/undefined/string/$-escape/CRLF/global). 487 affected-suite tests green.
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
loadConfig() returns a flattened object with no nested `workflow` key, so
config?.workflow was always undefined, making drift_action permanently 'warn'
and drift_threshold permanently 3 regardless of .planning/config.json. Fixes
by reading the raw config.json directly (matching the pattern in
check-command-router.cts:readWorkflowConfig). Adds two behavioral regression
tests that fail under the old code and pass under the fix.
Closes#1493
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1445,#1446): exclude 999.x backlog phases from milestone totals; allow total_phases downward correction
#1445: deriveProgressFromRoadmap (phase-lifecycle.cts), the roadmapPhaseCount
loop (state.cts), and getMilestonePhaseFilter (roadmap-parser.cts) all now
skip phase tokens matching /^999\b/ — consistent with the existing init.cts
filter. 999.x backlog dirs are consequently excluded from phaseDirs too.
#1446: shouldPreserveExistingProgress (state-document.cts) no longer includes
total_phases in its ratchet check. total_phases always takes the freshly
derived value; only completed_phases, total_plans, and completed_plans retain
ratchet behaviour.
Regression tests added for both bugs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1445/#1446 progress-backlog-exclusion-and-ratchet
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1445,#1446): rename test files to fix-NNN convention; fix changeset pr: null
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): add phase.list-plans to gsd-tools
Register phase.list-plans in PHASE_COMMAND_ALIASES, implement
cmdPhaseListPlans in src/phase.cts (uses findPhaseInternal + scanPhasePlans
to return plan_count/has_plans/plans/phase_dir), and wire the handler in
phase-command-router. Previously every call produced "Unknown phase
subcommand".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): register new test file in lint-test-file-count allowlist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): rename test to fix-NNN convention; update file-count allowlist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1422): sub_repos config takes precedence over .git in findProjectRoot
When heuristic-3 (.git + parent .planning/) fires, do a lookahead walk
over ancestors above the matching parent to check if any further ancestor
has a sub_repos entry that explicitly claims the starting directory. If
found, return that ancestor instead — explicit config wins over the
implicit .git signal.
Regression tests updated: the heuristic-3 precedence test now asserts the
correct new behavior (sub_repos wins), and two additional coverage cases
added (startDir directly in sub_repo child, and startDir nested 2+ levels
inside the child).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1447): guard against uncommitted changes before deleting phase dirs in new-milestone
cmdPhasesClear now runs `git status --porcelain <phasesDir>` before
executing any rmSync. If uncommitted changes are detected it calls
error() and aborts, preventing silent data loss when new-milestone's
§6 "phases.clear --confirm" fires before the operator has archived
or committed outgoing phase work.
A new --force flag is added to bypass the guard for callers that have
already verified archival is done (or explicitly accept the loss).
When git is unavailable or the directory is not inside a git repo the
guard silently skips, preserving the existing behaviour for non-git
projects.
Five new regression tests cover: untracked files abort, staged-but-
uncommitted abort, --force bypasses, committed files pass, non-git
project passes without guard.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1422 and #1447
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1472,#1454): validate health workstream-aware paths; exclude active worktree from W017
#1472: cmdValidateHealth now uses planningRoot(cwd) for shared-root files
(PROJECT.md, config.json, MILESTONES.md) and planningDir(cwd) for
workstream-scoped files (ROADMAP.md, STATE.md, phases/). Previously a
single planningDir() call was used for all paths, causing false
E002/E003/E004/W003 when GSD_WORKSTREAM is set.
#1454: W017 no longer fires for a stale worktree whose path equals or is
an ancestor of process.cwd(), preventing advice to remove the active
session's own worktree.
Regression tests added for both bugs; all 40 existing health tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation
ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
capabilities is 'gsd capability list', not this generated file.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1435): Added changeset for the capability matrix reference
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1435): address code-review — non-vacuous matrix test + generator polish
- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
(import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1435): backfill changeset PR number → #1458
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI
ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs
install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a
user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the
six subcommands, dispatching to the existing lifecycle/ledger:
- install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]…
- update [<id>|--all] [--scope] [--yes] [--shared-file] (re-resolves recorded source)
- remove <id> [--purge-data] [--scope] (first-party rejected)
- list [--json] (first-party + overlay, both scopes, JSON array)
- disable|enable <id> (activation-state alias of capability set --off/--on)
Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home,
project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json).
Consent is non-interactive: --yes grants; without it an executable install aborts after
printing the disclosure and writes nothing. Best-effort reconcile before each mutation.
Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs,
GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip,
remove round-trip + first-party guard, disable/enable, unknown subcommand.
Docs: docs/reference/gsd-capability-command.md reconciled to the real surface
(ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned);
docs/COMMANDS.md gains the gsd capability entry.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug
Adversarial-review (Codex) fixes:
- capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate
fail-closes on it (was silently downgrading to permissive)
- installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection
(capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't
act on a different id if the recorded source was retargeted
- capability update: prints the consent disclosure, exits non-zero on --all partial failure,
no longer masks the resolved id
- capability remove: ledger-first ordering so an overlay is removable even if it shadows a
first-party name; first-party guard only fires for ids not in the ledger
- gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle
not yet wired through this path)
Silent-output bug (root cause, not waved off as pre-existing):
- captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw
command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of
stdout. Now it flushes the captured buffer before re-throwing (exit code preserved).
- cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws
ExitError so the wrapper flushes — matches the repo's no-process-exit architecture.
- Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout.
Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed
- confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors
safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it.
- mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or
another capability's) server is skipped, so install/remove can't silently clobber user MCP config
(hooks already append; the map-keyed mcpServers path was the gap).
- capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead
of silently downgrading the strict_known_registries policy to permissive.
- Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry
preserved; unparseable config blocks an external install. capability suite 83/83, lint clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc
- install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never
fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per
the lifecycle contract; latent today, hardened for future status additions).
- Clarify capResolveScope comment (project scope === already-resolved cwd) and document that
strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide
allowlist) in gsd-capability-command.md.
- Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that
looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1451): backfill changeset PR number → #1457
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.
Closes#1434.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1269): expand same-prefix numeric ID ranges in --phase-req-ids
normalizePhaseReqIds treated a range token like "SEL-01..SEL-03" as a single
literal ID, so gap-analysis reported the range string as a missing requirement
even when SEL-01/02/03 existed individually.
Add a per-token range expander (run AFTER the existing split, preserving the
string[] | null | undefined contract): a token matching <PREFIX>-<NN>..<PREFIX>-<MM>
with identical prefixes, ascending bounds, and EQUAL digit width expands to the
individual IDs preserving that width; anything ambiguous stays literal
(fail-closed). Differing-width bounds stay literal so the expander never invents
a zero-padding the author didn't type, and ranges beyond MAX_PHASE_REQ_RANGE
(1000) stay literal as a DoS guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1269): add changeset for --phase-req-ids range expansion
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1269): mark changeset docs-exempt (internal flag, no user docs surface)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1269): isolate the DoS-cap branch with a same-width range; doc nits
Review fixes: the AC4 DoS test used REQ-1..REQ-100000, whose differing digit
widths trip the width guard before the cap is reached. Use a same-width
REQ-0001..REQ-1001 (span 1001 > 1000) so the test actually exercises the cap.
Clarify the PHASE_REQ_RANGE_RE capture-group JSDoc and the property-test width
assertion comment.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing
resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first
bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also
had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy
dir (probed first). Regression from #217 — pre-#217 returned
~/.gemini/antigravity unconditionally.
Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested
descriptor: a two-pass probe prefers the candidate GSD installed into, then
bare existence, then probe[0]. Behavior is byte-identical when probeMarker is
absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for
installer/operator guidance on already-misinstalled users (auto-relocation
ruled out per ADR-0008's single-configDir migration bound).
Regression tests fail before / pass after: coexistence + marker-priority
cases, end-to-end through the registry descriptor.
Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
* chore(#1441): add changeset for antigravity resolver fix
Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)
Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):
- Extract the conformance validator to a shared runtime-callable module
(gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
<root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
prefixes); full merged-set cross-capability validation; engines.gsd load-time
re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
+ config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
config-set call (never eager at module load, never wrong-cwd); first-party path
unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
capability-validator.cjs stays linted (#551 migration coverage).
Closes#1431
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1431): add changeset for runtime capability registry overlay
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)
The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].
Changes:
- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
{value, configured, reason, warnings}, makeResolution<T>() builder, and
AgentSkillsValue {block, skills_count}. No other src/ imports.
- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
skills_count} field (built via makeResolution). All existing flat fields
(agent_type, block, skills_count, warnings, configured, reason, source,
degraded) are retained unchanged for back-compat.
- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
naming it the canonical read-verb envelope. No JSON change.
- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
the canonical mutation-verb result (warnings=advisory, errors=operation-
not-applied). No JSON change.
- CONTEXT.md: new ### Resolution Convention glossary entry after
### Resolution Provenance.
- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.
- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
and back-compat of all flat fields.
All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.
Part of #1411
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add heuristic (4) to findProjectRoot: after heuristics (1)-(3) (sub_repos,
multiRepo, .git+parent-.planning/) are exhausted without a match, perform a
second bounded walk-up within FIND_PROJECT_ROOT_MAX_DEPTH to locate the
nearest ancestor directory containing a .planning/ subdirectory. Returns that
ancestor as the project root, so loadConfig finds the correct config instead
of falling through to defaults when gsd-tools is invoked from a plain
descendant subdirectory of a single-repo project (#1366 cwd-drift gap).
Ordering is load-bearing: the new walk runs AFTER the existing loop so
sub_repos workspaces (where a child sub-repo may have its own .planning/)
still resolve correctly to the parent workspace. The existing own-.planning/
guard (heuristic 0, #1362) and the depth bound (FIND_PROJECT_ROOT_MAX_DEPTH=10)
are both preserved unchanged.
Part of #1411
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Tightens the over-broad heading-walk detection: removes heading-walk
entirely and narrows fence-regex to require a multiline body ([\s\S]),
so single-line tests like /^```/ and /^###\s+/ are no longer flagged.
Grandfathers the 10 genuine section-collect sites across state.cts,
milestone.cts, audit.cts, and phase-lifecycle.cts with concrete reasons.
Adds 12 RuleTester tests (3 positive, 9 negative) to eslint-rules.test.cjs.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#1398): migrate state.cts section-collect regexes onto markdown-sectionizer seam (epic #1372 T6)
Replace ≈18 hand-rolled `/(heading)([\s\S]*?)(?=stop)/` regex splices in
state.cts with direct `tokenizeHeadings` calls that compute the exact
[bodyStart, stopOffset) span, preserving byte-identical STATE.md output.
Non-migratable site left in place: `cmdStateRecordMetric`'s metricsPattern
captures table-header rows in group 1 — not a standard heading+body shape.
Write orchestration (readModifyWriteStateMd / syncStateFrontmatter /
shouldPreserveExistingProgress / #952 no-op guard) is UNTOUCHED.
Verification: t6-headtohead.cjs head-to-head harness runs 25 ops across 6
STATE.md fixture variants (inline, trailing-blanks, CRLF, no-frontmatter,
nested-acc, post-milestone) against origin/next and reports 0 diffs.
No-op guard confirmed: record-session on recorded:false leaves STATE.md
byte-identical. All 150 state tests and 62 milestone/forensics tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1398): capture T6 state-section-splice characterization; drop throwaway harness
Remove scripts/t6-headtohead.cjs (committed throwaway HEAD-vs-origin/next
byte-compare harness). Capture its coverage as 23 behavioral characterization
tests appended to tests/state.test.cjs, exercising all migrated cmdState*
write-ops across 7 fixture variants (inline, trailing-blanks, CRLF,
no-frontmatter, nested-acc, no-current-pos, post-milestone). Includes the
#952 no-op guard (recorded:false + byte-unchanged assertion) and
CRLF/trailing-blanks edge-case coverage. Also removes the dead
spliceStateSection helper (defined but never called) that was surfacing as
an @typescript-eslint/no-unused-vars warning.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- uat-predicate.cts: replace local _stripFencedBlocks (and its private
FenceState/StripFencedResult types) with a call to stripFencedCode from
markdown-sectionizer.cjs (ADR-1372 T5). Both stripFalsePositiveContexts
step (c) and analyzeMarkdown now route through the seam. The three other
passes in stripFalsePositiveContexts — frontmatter strip, HTML-comment
strip, blockquote-line filter — remain caller-side (seam does not do these).
The unterminatedFence signal consumed by analyzeMarkdown is preserved; it
is now returned by stripFencedCode (same machine, same contract).
- uat.cts: migrate the ## Current Test, ## Tests, and ## Human Verification
section-collect patterns onto collectSection/tokenizeHeadings from the seam.
UAT-specific item parsing (### N. Name blocks, expected/result fields,
categorization logic) stays caller-side. The HTML-comment strip within the
Current Test body remains caller-side (UAT document structure, not seam scope).
- tests/markdown-sectionizer.test.cjs: remove the 18-case tautological parity
guard (DEFECT.GENERATIVE-FIX). Once uat-predicate imports the seam the guard
compares the seam to itself — removing it is the T5 commitment per ADR-1372.
4-space-indent behavior change (CommonMark correctness improvement): the seam
uses /^( {0,3})/ (CommonMark §4.5 ≤3-space indent); the retired
_stripFencedBlocks used /^(\s*)/ (any indent). A 4-space-indented ``` is no
longer treated as a fence opener (it is an indented code block per CommonMark).
Head-to-head over 9 corpus inputs: 0 diffs on all standard cases; 2 diffs only
on the synthetic 4-space-indent edge cases. No UAT fixture or test in the suite
exercises 4-space-indented fences. The change is a correctness improvement.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Removes all three inline copies of the fenced-code state machine from
src/roadmap-parser.cts and replaces them with calls to
tokenizeHeadings() from the canonical markdown-sectionizer seam
(ADR-1372 T4).
Changes:
- Drop stripFencedLines() function (copy 1 of 3 — standalone helper)
- Rewrite computeSectionEnd() using tokenizeHeadings() offsets into
the original content (copy 2 — inline fence loop); the returned
character offset is preserved exactly for all inputs
- Rewrite getMilestonePhaseFilter() versionOverride fence loop using
tokenizeHeadings() (copy 3 — inline); sectionEnd offset preserved
- Replace stripFencedLines(roadmap) + unanchored phasePattern.exec()
with tokenizeHeadings(roadmap) filtered by level and phase-heading
pattern; headings in inline HTML comments (<!-- ## Phase N: -->)
are no longer mis-counted (anchored ATX detection is correct)
- Add import { tokenizeHeadings } from ./markdown-sectionizer.cjs
Offset preservation: tokenizeHeadings() records h.offset as the
character index of '#' in the ORIGINAL content; computeSectionEnd()
and the versionOverride path both use h.offset directly as the
section-end character offset — no stripping, no shift.
Corpus head-to-head: 29/30 slots are byte-identical to origin/next.
The 1 diff (getMilestonePhaseFilter phaseCount for the HTML-comment
fixture: 3→2) is an improvement: the old unanchored regex counted
"## Phase 998:" embedded in "<!-- ## Phase 998: ... -->" mid-line;
tokenizeHeadings() correctly requires '#' at line-start (ATX rule).
No existing test asserts on that count; feat-3594 passes unchanged.
All 122 roadmap/milestone/phase tests green.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- check-command-router: stripCommentsAndFences delegates fenced-code
stripping to seam's stripFencedCode; HTML-comment stripping stays
caller-side. extractPlanDesignatedSections replaces hand-rolled
split(/\r?\n/) + /^#{1,6}\s+/ heading walk with collectSections
driven by DESIGNATED_HEADINGS_RE.
- gap-checker: parseRequirements checkbox-bullet detection migrates
to iterateBullets (checkbox markers); **ID** extracted caller-side.
Table-row path and separator-row skip stay caller-side.
- Adds refactor-1390-t3-characterization.test.cjs (43 behavioral
tests) that were green before and remain green after.
- T1 fail-loud / could-not-parse gate semantics unchanged (verified
by decisions.test.cjs 59/59 and direct gate invocation).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (T2)
Replace hand-rolled parseSections (split/heading-regex walk/body accumulation)
with collectSections(content, () => true) from the canonical seam. Replace
hand-rolled splitEntries bullet-strip regex with iterateBullets from the seam,
preserving the plain-text-line fallback for byte-identical output. Removes the
last inline heading/bullet scanning from adr-parser.cts; normalizeAdrHeader and
all ADR-specific classification logic are unchanged. 218/218 tests pass before
and after.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#1387): keep splitEntries flat (iterateBullets changed its contract); seam adoption stays in parseSections
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#1387): drop dead preamble reconstruction; add targeted adr-parser mutation tests
parseSections' preamble reconstruction block (heading: null entry) was dead code:
both consumers (parseAdrMarkdown and parseStatusFromSections) skip heading: null
sections immediately on entry. Confirmed via analytical trace and zero-diff corpus
head-to-head across all 44 docs/adr/*.md files.
Adds 11 targeted behavioral tests to kill cheap surviving mutants:
- pushUnique intra-values dedup (kills seen.add removal mutant)
- body split/join round-trip with multi-line prose and entries
- parseStatusFromSections [0] indexing (only first line determines status)
- classifyHeader equality vs prefix-match boundary (exact match, prefix match,
synonym+letter non-match)
- goal section prose vs entries distinction (bullet markers preserved in context)
- normalizeAdrHeader non-word char removal (parens and slash behavior)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1364,#1365): add decisions regression tests (fail-first proof)
Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
(these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate
#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.
#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.
parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.
Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D
FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.
FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.
FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).
FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.
Adds 14 behavioral regression tests (fail-first verified manually before fixes).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions
Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.
Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.
Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1364): backfill changeset PR number (1386)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0)
Establishes the canonical markdown-structure parsing seam per ADR-1372.
No existing parsers are modified; this is the foundational T0 tier only.
- docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the
seam interface, the tiered migration plan (T0-T7), and the prohibition
enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7).
- src/markdown-sectionizer.cts: Pure module, Node built-ins only.
Exports: stripFencedCode (CommonMark-correct state machine ported from
uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence
signal), tokenizeHeadings (ATX headings outside fenced blocks),
collectSections (line-by-line predicate-driven section collection),
collectSection (single named section, levelBounded stop, optional
stripFences), iterateBullets (dash/checkbox/numbered + continuation).
- tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering
the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences,
unterminated fences, nested levels, all bullet markers, continuation
lines, empty/non-string input) plus 4 fast-check property tests
(idempotence, output shape, never-throws, length monotonicity).
- CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate).
Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory
- src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets;
add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks,
tagName regex-escaped, caller decides fence-stripping) and replaceSection(content,
section, newBody) (pure character-offset splice for read-modify-write callers);
update collectSections/collectSection to populate bodyStart/bodyEnd.
- tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for
extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard
that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a
shared 9-item corpus; documents the known 4-space-indent divergence.
- docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and
replaceSection in §"The seam".
- CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports.
- docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between
loop-resolver and milestone).
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage
FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both
collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of
the raw stop-line offset, eliminating the trailing-newline overcounting that caused
replaceSection to drop separator newlines (## A\nbody## B gluing bug).
FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty
ATX headings (## / ## ), text=''. 4-space indent correctly excluded.
FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose
level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections
that also stop at ### without abusing levelBounded.
FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark).
Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences
unaffected.
FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section
doc comments.
FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks
the non-greedy close-at-first-</tag> behavior.
FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so
tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3).
Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish)
The seam's compiled artifact must be a build-at-publish output like every other
src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed
file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve
buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.
Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1374): backfill changeset PR number (#1376)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1355): detect-and-warn guard for claude-code agent-teams
GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.
- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
`query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): add changeset for teams-detect guard
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)
The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): register teams-status.cjs in the inventory manifest
The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape
reconcileCodexHooksJsonEvent preserved whatever shape it read, so on an empty,
absent, or legacy top-level hooks.json it wrote top-level event keys
(`{ "SessionStart": [...] }`) that current Codex (deny_unknown_fields) rejects,
instead of the canonical `{ "hooks": { "SessionStart": [...] } }`.
- Lift any top-level event arrays (legacy, empty, or mixed nested+top-level)
into the nested `hooks` table, merging same-named events so no user/legacy
entry is dropped and no stray top-level event key survives. Mirrors
reconcileCursorHooksJson.
- Collapse an empty hook table back to `{}` so removal on an absent file does
not write a spurious `{ "hooks": {} }`.
- Read path still tolerates both shapes; dedup/removal unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1348): add changeset for Codex hooks.json canonicalization
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1343): parse decision bullets with text before the colon
parseDecisions() silently dropped any `- **D-NN ...:**` decision bullet
whose header had freeform text (a parenthetical, em-dash, or prose) before
the `:**`, so the blocking check.decision-coverage-plan gate computed
coverage over a narrowed set and reported a false pass.
- Broaden bulletRe to tolerate a freeform run before the colon while
preserving the optional [bracket] tag capture (drives `trackable`).
- Add a parse-miss guard: a line that looks like a D-NN bullet but still
fails the regex flushes the current decision and warns instead of
vanishing — the gate-integrity floor.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1343): add changeset for decision-coverage false-pass fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1343): relocate decision-parser regression into owning module test file
CI's lint-regression-test-names bans new bug-NNNN-*.test.cjs files. Move the
9 regression cases from tests/bug-1343-parsedecisions-drop.test.cjs into the
owning parser test file tests/post-planning-gaps-2493.test.cjs (which already
exercises parseDecisions) and delete the banned file.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>