7efd9032eee62fa1905fadfeec4a591fea9c32ab
3220 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7efd9032ee |
fix(#4544): cover hooks/, scripts/, CHANGELOG and the manifest in codex rollback (#4760)
* test(#4544): failing-first rollback-coverage tests for codex install * test(#4544): scope the rollback suite to its own requires The appended describe used bare describe/test/os/cleanup and the folded block's runCodexInstall — none visible at file scope, so the whole test file failed to load. Wrap it in its own block with local requires and a local harness copy, matching the folded-block idiom. * test(#4544): import beforeEach/afterTest hooks into the suite scope * fix(#4544): restore manifest-tracked files and hooks/ on codex rollback restoreCodexSnapshot and _codexPreConfigRollback knew only the five #3245 targets (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, VERSION). A Codex install also writes hooks/, gsd-core/CHANGELOG.md, gsd-core/.gsd-runtime, scripts/changeset|lib + the standalone scripts, and rewrites gsd-file-manifest.json -- none were captured, so any rollback left the new payload in place for all of them. The pre-install capture now records every path the PRIOR install's own gsd-file-manifest.json lists (bytes when present, absence-marker when not, so a path deleted between installs is re-deleted rather than resurrected), the manifest file itself, and the whole hooks/ tree -- wholesale, because the Codex manifest deliberately omits hooks/ and hooks/ is shared space, so restore returns user files that predated the install and drops everything the failed install staged. Both rollback paths share one restore closure; malformed or missing prior state degrades to today's behavior. Known residual, documented: files the FAILED install adds under manifest-tracked dirs survive a rollback that fires before the new manifest is written (they are named by no prior state). The five original targets and hooks/ have no residual. * test(#4544): align malformed-manifest fixtures with pre-install-state semantics Row 7 seeded VERSION and then asserted its absence -- but a seeded VERSION is pre-install state the fix must restore, not remove. Row 8 asserted a pre-existing array-shaped manifest must not survive, when restoring those exact bytes IS the contract. Both were fixture bugs; the probe-verified installer behavior was correct. * docs(#4544): add Fixed changeset for codex rollback coverage * fix(#4544): harden snapshot per adversarial review — minimal mode, symlinks, clean installs Review (three independent passes) found five defects and one coverage gap in the first cut; all fixed: - BLOCKER: the capture gate is off in minimal mode but the restore call was not, so a minimal-mode rollback wholesale-deleted the user's entire hooks/ directory (empirically confirmed by the reviewer). The restore now consults a captured flag: no snapshot means do nothing. - MAJOR: the hooks/ walk followed file symlinks — a repo-shipped .codex/hooks symlink to a FIFO would hang the installer, to /dev/zero exhaust memory, or to private data copy that data into the snapshot. The walk lstats every entry and captures only true regular files; anything else marks the capture incomplete. - Incomplete captures now downgrade the restore to per-file: put back what was captured, remove only the names GSD itself stages (the hoisted CODEX_HOOKS_TO_COPY set + CommonJS marker), never wholesale- delete a tree the snapshot did not fully see. GSD-owned removal runs before the restore so a name in both sets keeps its pre-install bytes. - A pre-existing hooks FILE (not directory) is left alone instead of deleted. - readInstallManifest now rejects a manifest whose files field is a JSON array (typeof [] === 'object'), which previously produced numeric-key paths. - Clean FIRST installs: with no prior manifest nothing recorded the payload, so a failed clean install rolled back to a half-written tree. Capture now enumerates the same source directories the installer copies plus the two standalone files (CHANGELOG.md from the repo root, generated .gsd-runtime) and records absence — a failed clean install now rolls back to actually nothing. Tests: minimal-mode preservation regression, symlink never-followed regression, helpers.cjs temp dirs, the injected-failure message is asserted, residue assertions made unconditional. * test(#4544): pin symlink-downgrade semantics the final run exposed The symlink itself marks the capture incomplete, so the restore takes the per-file downgrade — which preserves uncaptured pre-install state (the link) rather than wholesale-dropping it. The probe run verified exactly this; the assertion guessed the wholesale branch. Pin the verified behavior: referent untouched, link preserved and resolving, no leak. * docs(#4544): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
aeac47b95a |
fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD (#4472)
* fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD `--phase` documented a primary path reading `.planning/.phase-manifest.json`, but nothing in the repository writes that file, so the documented fallback was the only reachable path: git log --oneline --no-merges --all | grep -E "\(0*${TARGET_PHASE}(-[0-9]+)?\):" | head -50 That selection has no milestone bound and no reachability bound, and it feeds `git revert --no-commit`. On a project that reuses a phase number it stages deletion of a previous milestone's files, under a confirmation gate that displays only `{hash} — {message}` — the one field that carries no milestone discriminator. Port the #3995 anchor already live in code-review.md: resolve the phase's own directory via `find-phase` (which resolves through planningDir, so it is workstream-correct), take PHASE_START as the first commit adding anything under it, and select over `PHASE_START^..HEAD`. Drop `--all`. Fail closed when no anchor resolves rather than widening to a repository-wide search, and report truncation instead of silently capping at 50. Also resolve the dead manifest read rather than leaving documented-but- unreachable behaviour, and hoist dependency_check onto the workstream-resolved planning root — it read a hardcoded `.planning/ROADMAP.md` and `.planning/phases/`, which are the wrong tree under an active workstream. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv * fix(#4465): close three defects the pre-create adversarial review found Root-commit off-by-one: `${PHASE_START}..HEAD` EXCLUDES PHASE_START, so when the phase's first commit is the repository root the selection silently dropped it and the undo refused legitimate work. The root branch now selects over `HEAD`. Empty-selection exit status: `grep` exits 1 on no match, and the removed `| head -50` had been masking that rc. Both selection pipelines now end in `|| true` so an empty selection reaches the workflow's own Empty check instead of aborting the block. Truncation stop was documented for MODE=phase only; MODE=plan could still cap silently. Both modes now carry it. Also documents two residuals the review surfaced rather than leaving them implicit: a revision range is ancestry and not chronology, so a pre-phase side branch merged in after PHASE_START stays selectable; and the `--diff-filter=A` anchor does not follow renames, so an archived phase directory under-selects (reverts too little or refuses, never too much). Tests: 12 assertions, negative-controlled. A-H and K-L are RED against the pre-fix workflow; J is RED against the intermediate revision that carried the off-by-one; I is green in both directions by design. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv * chore(#4465): backfill the changeset fragment's pr field The fragment could not carry `pr:` before the PR existed; both `scripts/changeset/lint.cjs` and `scripts/lint-docs-required.cjs` require it and reported `missing_pr` until now. Both are green with it filled in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv * chore(#4465): acknowledge the deliberate growth of undo.md The emitted-attribution gate flags undo.md growing 11881 -> 17435 bytes. The growth is the fix: a one-line `git log --all | grep` selection is replaced by an anchored, HEAD-bounded selection for BOTH modes, each with its own fail-closed branch and truncation stop, plus three residuals documented next to the code they qualify rather than left implicit. A workflow document is the executable contract, so the residuals belong in it. Emitted-Drift-Ack-Growth: undo.md — replaces a one-line unbounded commit-subject grep with an anchored HEAD-bounded selection in both --phase and --plan, each with a fail-closed branch and a truncation stop, plus three residuals documented in-workflow (#4465) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv * test(#4465): execute undo.md's selection fences against a git fixture The shipped regression test matched substrings in undo.md's bash fences. A substring cannot tell a live invocation from a dead one, and this is the boundary/reachability defect class this repo's own conventions want driven with limit-1/limit/limit+1 execution proof. So the fences the runtime runs — sliced out of undo.md by content anchor, never by position — are now replayed with `bash -c` inside createTempGitProject fixtures against the real `gsd-tools.cjs`, the same fence-execution shape #2308 and #2352 use. Ten executed cases: the fixture's own negative control (the retired `--all` grep over-selects the archived milestone and a dead branch), `--phase` and `--plan` selecting only the current milestone's HEAD-reachable instance, the single-milestone selection unchanged, limit-1 (a matching pre-phase commit excluded, PHASE_START itself included), the root-commit branch selecting over `HEAD`, fail-closed on an absent phase in both --phase and --plan (empty PHASE_DIR and UNDO_RANGE), the active workstream's phase directory winning over the root's, and dependency_check's `planning inspect --pick generated_from.planning_root --raw` resolving the workstream root, the project root, and the `.planning` fallback. Skipped on win32 with the #2352 precedent's reason: the fences are POSIX bash driven through `bash -c`; the shape half still runs there. Negative-controlled against the pre-fix undo.md (upstream/next): 11 of 12 shape assertions red, and the executed block fails at fence extraction. Shape test L now pins the >50 refusal message in both modes, not only its heading — the cross-AI round review removed the paragraphs under intact headings and L stayed green; it now fails on that mutation. * fix(#4465): refuse an archived-milestone PHASE_DIR in both undo modes `find-phase` searches the live `phases/` directory first, then every `milestones/v<X.Y>-phases/` directory in ascending version order, and its ambiguity check is scoped to one directory: the `matches.length > 1` test sits inside `cmdFindPhase`'s per-`searchDir` loop (`src/phase.cts`). A phase number that is not live therefore resolves silently to the OLDEST archived milestone carrying one, with no warning. Anchoring there is wrong in both directions at once. The oldest commit adding that path is the archival move, so the phase's real work predates the window and falls outside it, while the window runs forward from that archival through every later milestone -- where the subject grep matches THEIR same-numbered phase. Driven on a two-archived-milestone fixture, `--phase 03` selected v2.0's `feat(03-01): add search index` and `docs(03-01): v2.0 phase plan` and excluded v1.0's own `feat(03-01): implement auth endpoint`. On `git revert --no-commit` that is the cross-milestone contamination this PR exists to close, recurring. Both modes now blank PHASE_DIR on an archived resolution and refuse with their own message, so the existing fail-closed rule stays load-bearing even if the refusal's prose is not honored. The "archived or renamed phase directory under-selects" residual claimed this failure could revert "too little or refuse, never too much". That was wrong in the direction that matters; it is rewritten to cover renames only, and the archival case is recorded as refused rather than disclosed. Tests: a two-archived-milestone fixture, a negative control asserting the unguarded window selects v2.0's two commits and none of v1.0's, refusals in both --phase and --plan, and an inertness check on a live resolution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * docs(#4465): close the three Minor review items on undo.md and its test Purpose line: it still advertised rolling back "using the phase manifest", the exact `.planning/.phase-manifest.json` mechanism this PR removes. Test B already reads the whole file, so scope is not why it missed this -- B greps the hyphenated `phase-manifest` token, the filename, while the purpose line named the same dead mechanism in prose. Test N pins the <purpose> block itself, which is spelling-independent; widening B's pattern to /manifest/i instead would fire on any future sentence that merely mentions one. Merge-commit anchors: `git log --diff-filter=A -- "${PHASE_DIR}"` does not walk merge diffs by default. A directory added on a side branch is still found -- the side-branch commit that added it is in history -- so the uncovered case is narrower than "introduced via a merge": it is a directory first appearing in the merge RESOLUTION. The information is not absent from history, only unrequested: `git log -m` prints that add once per parent. Recorded as a residual rather than fixed -- the failure is a refusal, and the evil-merge fixture costs more than a safe-direction branch is worth. allow-test-rule marker: suppression is site-scoped, and CONTRIBUTING.md pins the window at MAX_MARKER_LOOKAHEAD_LINES = 8 with only blanks and comments between. The file-header marker sat 43 lines above the first `readFileSync` with requires and a function definition in between, so it was inert for both read sites. Moved to each site. The markers are belt-and-braces today, and NOT because the rule ignores `RegExp.test` -- it handles `regex.test(tracked)` explicitly (no-source-grep.cjs:239, :597-605). Neither read is tracked at all: `looksLikeSourcePath` (:378-390) admits only .cjs/.cts/.js/.mjs/.mts/.ts and UNDO_PATH is a .md, and the second site's reader is `readFileNormalized`, which the rule does not recognise as `readFileSync`. They become load-bearing if either scope widens. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4465): record the archived-milestone refusal in the changeset The fragment described the PHASE_START bound and the fail-closed rule but not the archived-resolution refusal added this round, which is a user-visible behaviour change: `--phase N` on a number that is no longer live now stops rather than anchoring on an archived directory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * docs(#4465): disclose the same-number-and-slug residual in undo.md The round's own cross-model body audit found this stated in the PR description and nowhere in the deployed workflow -- which is the half that survives merge, and the half this PR's whole argument says residuals belong in. The anchor is the CURRENT path and `--diff-filter=A` does not follow renames, so a later milestone that re-creates the same literal directory (`03-auth` again, not merely phase `03` again) makes the oldest add at that path the previous occupant's. The archived-milestone refusal added this round structurally cannot reach it: `find-phase` returns the LIVE directory, so nothing is under `milestones/` to refuse. Driven -- v1.0 and v2.0 both using `.planning/phases/03-auth`, v1.0 archived in between: PHASE_DIR resolves live, the guard correctly does not fire, the anchor is `docs(03-01): v1 plan`, and the selection returns all four v1+v2 phase-03 commits. `code-review.md` carries the same residual on the same anchor, where it is read-only; on `git revert --no-commit` it is not, so the note points at `/gsd:undo --last N`. Documented rather than fixed: following renames across a re-created path needs a phase identity that a directory name does not carry, which is the same wall residual 1 hits. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4465): refuse a LIVE phase directory whose name is also archived Self-found this round, from the cross-model audit of the previous commit: that commit disclosed same-number-and-slug reuse as a residual and asserted closing it needed a phase identity a directory name does not carry. The audit refuted the second half by producing a fix, and it is cheap and safe-direction, so the case is now refused rather than documented. `--diff-filter=A` does not follow renames, so a later milestone that re-creates the same literal directory (`03-auth` again, not merely phase `03` again) anchors on the EARLIER occupant's add commit and the window opens there. The archived refusal added earlier in this round structurally cannot reach it: `find-phase` returns the LIVE directory, so nothing is under `milestones/` to refuse. Driven before the guard -- v1.0 and v2.0 both at `.planning/phases/03-auth`, v1.0 archived in between: PHASE_DIR resolves live, the archive guard correctly stays inert, the anchor is `docs(03-01): v1 plan`, and all four v1+v2 phase-03 commits are selected. That is a previous milestone's work staged for `git revert --no-commit`. The guard needs no identity reconstruction: the same basename present under an archived `milestones/v*-phases/` means the path has been used before, so the anchor is untrustworthy and both modes refuse with their own message. It fails toward refusing a legitimate undo of the reusing milestone, which `--last N` covers; the alternative is reverting the earlier one's commits. Tests: a reused-slug fixture, a negative control asserting the unguarded window reaches back into v1.0 (all four commits), refusals in both modes, and an inertness check on a distinct slug. All three new checks red against the pre-fix fence. Also narrows the changeset, which still claimed selection "no longer" reaches a previous milestone -- true of the archived route, not of this one until now -- and corrects the concurrent-workstreams residual, which claimed the window removed the previous-milestone class "entirely" while this case remained open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4465): harden the reused-path refusal against its own false positives The cross-model audit of the previous commit refuted four of its claims. All four were right; this closes them. A REAL BUG: the refusal message interpolated `${PHASE_DIR}` after the fence had already blanked it, so it would have rendered "Phase 03 resolves to , but that directory name is also archived at ...". The live path is now preserved in `PHASE_DIR_LIVE` before blanking, and a test pins that it survives. THREE FALSE-REFUSAL ROUTES, each of which could block a legitimate undo: - `[ -e ]` accepted a regular FILE where an archived phase directory would sit. The evidence the refusal claims is "an earlier milestone used this path", and only a directory is that. Now `[ -d ]`. - The glob `v*-phases` accepted milestone directory names `cmdFindPhase` itself rejects -- its filter is /^v\d+.*-phases$/, so `vnondigit-phases` is not a milestone it would ever resolve. Refusing over one is refusing on evidence the producer discards. Now `v[0-9]*-phases`, in the collision check and in the archived-resolution guard alike. - `${PHASE_DIR%/phases/*}` silently left the path unchanged when it carried no `/phases/` segment, so the scan ran against the wrong root and read as clean. The strip must now have fired. Outside today's producer contract either way -- cmdFindPhase's live output always carries `/phases/` -- so this hardens a claim rather than fixing a reachable defect, and is stated as such. The commit message and the changeset both asserted the guard proves the name is "archived"; before this commit it proved only that some entry existed. Both are narrowed to what the fence now actually establishes, and the changeset headline no longer claims more than the anchor plus the two refusals deliver -- residual 2 (a range is ancestry, not chronology) is untouched by either. Tests: a file-not-directory twin, a malformed `vnondigit-phases` twin, and the preserved live path. All three red against the pre-hardening fence, along with shape test M. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * docs(#4465): disclose what the reused-path scan does not survive A fifth audit pass refuted three claims in the previous commit. Two are wording, one is a real fail-open; none is fixed by more shell, so all three are stated. The wording: `v[0-9]*-phases` was described as MIRRORING cmdFindPhase's /^v\d+.*-phases$/. It is not equivalent -- the glob's `*` matches a newline where the regex's `.` does not, so a directory named `v6<newline>-phases` is accepted here and rejected there. It tracks the filter closely enough to reject the malformed siblings that motivated it; it does not mirror it, and the comment no longer says so. The changeset likewise said the twin is "an archived phase directory", where the check establishes only that a directory of that name exists under an archived milestone -- narrowed to that. The fail-open: this collision check is the only fence in undo.md that relies on pathname expansion (every other one uses `case`, which `set -f` does not affect). Under a runtime with globbing disabled the scan is skipped silently and a genuine collision passes; under `shopt -s failglob` a NON-match aborts the fence. Both sit outside the shell state this workflow assumes throughout, and defending only this fence while the rest of the file assumes defaults would be inconsistent -- so it is a documented residual rather than a hardened one. Also disclosed: the check is conservative at two edges. `[ -d ]` follows symlinks, and an empty directory of the right name counts, so either can refuse an undo a stricter ownership test would allow. That direction is the intended one -- refusing too often costs a `--last N`, refusing too rarely reverts another milestone's work. No test is added. Asserting behaviour under `set -f` would pin a shell state the workflow does not otherwise support, and the two conservative edges are the documented intent rather than defects. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4465): register the new regression test in the conformance-tier list Round 3's only blocker. The platform-conformance-tier classifier, its generated registry and the test that asserts the registry is fresh all landed together on `next` in `bcd99696d` (#4591 / #4598) on 2026-09-10 08:36 -0400 -- a day after this branch's last push, so the gate did not exist when the branch was last green. This PR adds a unit-suite test file, the classifier's content scan selects it, and a list regenerated on `next` without it is therefore stale the moment the two meet. (The review cites `0b928fe28c` for that landing. That commit is #4592, four hours later the same day, and `git diff-tree` shows it touched only `scripts/ci-test-scope.cjs`, `scripts/gen-platform-conformance-tier.cjs` and those two files' tests -- not the generated registry at all. The date and the diagnosis hold either way.) Regenerated with `node scripts/gen-platform-conformance-tier.cjs --write` on the rebased tree: one line, 267 -> 268 entries. The `--target macos` list is unaffected -- it writes a different file, `macos-conformance-tier.generated.cjs`, and its classifier does not select this test (198, already matching) -- so `lint:generated-sync`'s second conformance-tier link needed nothing. **Two different baselines, stated so the counts are not read as one.** CI's failure on the prior head reads `546 !== 547`, and this commit's diff reads 267 -> 268. Both are correct and they are not the same tree: at the merge commit `54b0b197f` the committed list held 546 entries, so the classifier wanted 547. `4d65c248e` (#4641, "narrow the tier to 28.5%") then landed on `next` at 2026-09-11 17:00 -0400 -- after that CI run started at 13:26Z -- and collapsed the list to 266, with `a2331c01f` taking it to 267. Hence 268 here. The two 547-entry lists are not the same file set: upstream's includes `tests/execute-phase-decimal-arithmetic.test.cjs`. Order matters and is worth stating: the registry is derived from the `tests/` tree, and the base range removed 282 entries from it and added 3 (`git diff --numstat` reports `3 282`; the familiar 279 is the net shrinkage, not the count of entries changed). Regenerating before the rebase would have produced a 547-entry list against a base carrying 267, and a three-way `git merge-file` control over that pair does conflict. Rebase first, regenerate second. This clears both reds on the prior head, not one. `lint-tests` is the one the review named; `test (ubuntu-latest, 24, shard 3/3)` is the same staleness seen through `tests/platform-conformance-tier.test.cjs:278`, which asserts the committed list is fresh. It was the only failure in 1267 tests on that shard, and it completed at 13:38:33Z -- five minutes after the review was submitted at 13:33:54Z, which is why the review recorded the ubuntu matrix as green. Control: removing the added line reds `real tests/ tree classification matches the committed list` (1 of 55, `267 !== 268`); restoring it gives 55/55. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R4EJcuDjaaF1QeeT8AxYbj * test(#4465): pin what find-phase returns for the workstream archive layout Review round 4 found the archived-phase handling blind to the second archive layout, milestones/ws-<name>-<date>/phases/, that `workstream complete` writes. The suite had no fixture for it and its comments described one archive reader where there are two. - Names both readers in the resolver-contract comment: find-phase routes to cmdFindPhase, which admits only /^v\d+.*-phases$/ under milestones/, while the phase locator (listArchiveVersionDirs) also enumerates ws-*. - Adds a ws-* fixture built the way `workstream complete` builds it (the whole workstream dir moved into milestones/ws-feat-<date>/). - Negative control: a re-created workstream with the same phase slug, run without the collision guard, selects the archived generation's commits. - Pins the actual find-phase contract: a phase living only in a ws-* archive resolves to nothing, and nothing is selected. - Pins that an ordinary deleted plan file inside a live phase is not treated as a previous occupant. runFences gains an explicit env override so a test can activate a workstream after the leak scrub. Every test here passes on the pre-fix workflow; the assertions that need the fix land with it in the next commit. * fix(#4465): refuse a re-created phase path by its history, not an archive-layout glob Review round 4: the archived-phase handling encoded the archive layout as a literal `v[0-9]*-phases` glob, while the phase locator owns two layouts. The second, milestones/ws-<name>-<date>/phases/, is what `workstream complete` writes. Driven on the real fences: a workstream `feat` completed into ws-feat-<date>/ and then re-created with the same 03-auth slug resolves to the live path, the collision glob finds no v*-phases twin, the anchor opens on the archived generation's first commit, and --phase 03 selects that generation's commits. The collision check now asks git whether this exact path went EMPTY somewhere in HEAD's history and came back: `git log -m --no-renames --diff-filter=D` over PHASE_DIR, then `git ls-tree -d` at each deleting commit to tell a vacated directory from an ordinary deleted plan file. That answers for every way a path can be vacated -- the flat archive, the ws-* archive, and a phase removed and re-added under the same slug, which no layout glob could see -- without a second reader of the layout to keep in sync. The refusal message now names the commit that vacated the path. The archived-resolution refusal names both layouts. find-phase (cmdFindPhase) searches only the flat archives today, so a ws-*-only phase resolves to nothing and fails closed on the not-found rule; the ws-* arm keeps the refusal correct if find-phase is ever taught the locator's second layout, and a stubbed-resolver test pins it for both modes. The retired glob's shell-state residual (pathname expansion under set -f / failglob) is gone with it. Its replacement residual is documented in-workflow: the history check misses a single commit that both moves the directory away and re-creates it, and over-refuses when a side branch emptied it and the merge kept it. * fix(#4465): run undo's path-scoped git calls from the project root find-phase answers against Found by this round's pre-push review. find-phase prints PHASE_DIR relative to the PROJECT ROOT -- gsd-tools resolves the root before it dispatches -- while the workflow's shell stays wherever the user invoked /gsd:undo. A git pathspec is read relative to git's cwd, so from a subdirectory every PHASE_DIR-scoped call looked in the wrong place: - the anchor (`git log --diff-filter=A`) came back empty, so a legitimate --phase or --plan undo was refused. Fail-closed, but a refusal of valid work, and it dates from round 1; - the collision check's `git log --diff-filter=D` came back empty, so a re-created path was never recognised. The anchor failing too is the only reason this did not over-select. Both modes now take PROJECT_ROOT from `planning inspect --pick generated_from.cwd` -- the directory find-phase itself resolved against -- and run the three path-scoped calls as `git -C "${PROJECT_ROOT:-.}"`. The project root, not `git rev-parse --show-toplevel`: a project need not sit at the top of its repository, and a control test pins that choice. Selection, revert and rev-parse calls are SHA-only and unchanged. Tests: every behavioural test ran from the fixture root, where the two readings coincide. New ones run the fences from sub/dir -- normal selection in both modes, refusal of a re-created path for both archive layouts in both modes, and a dropped plan file NOT refused, which is how a mis-rooted ls-tree would fail (it lists nothing, and nothing reads as "vacated"). Shape test O pins every PHASE_DIR-scoped git call to the root. The two tests renamed in the previous commit are relabelled as false-positive controls: they pass with the collision check deleted, and say so. Reversion controls, all driven: the pre-change undo.md fails 7 tests; one mode's ls-tree mis-rooted fails 3 in either mode; PROJECT_ROOT from --show-toplevel fails 2. * fix(#4465): refuse a phase planned in another repository, and never widen a foreign anchor Found by this round's second pre-push review, against the previous commit. Running the anchor from the project root is right while the project root and the caller share a repository. In a `sub_repos` project they do not: .planning/ lives in a parent repository and the code in child ones, and from a child findProjectRoot returns the parent. The anchor then came from the PARENT's history -- a commit the child does not hold. `git rev-parse "${PHASE_START}^"` failed in the child, the root-commit arm read that as "no parent", set UNDO_RANGE=HEAD, and selection ran over the child's whole history. Driven: both child commits selected, including one older than the phase. Before the previous commit that case resolved no anchor and failed closed; the previous commit made it destructive. Two layers, both modes: - a same-repository refusal in the guard fence: the caller's git directory and PROJECT_ROOT's, each physically resolved, must be the same one (per-worktree, so a linked worktree compares correctly). A different one sets PHASE_DIR_FOREIGN, blanks PHASE_DIR, and the workflow stops with a --last N pointer; - in the anchor: a PHASE_START that is not a commit this repository holds is blanked before the root-commit arm. That ambiguity -- "no parent" read as "root commit" when it can also mean "not a commit here" -- dates from round 1; the previous commit made it reachable. Shape test O now also pins that PROJECT_ROOT is assigned only from planning inspect: a rooted call is only as good as the root it is handed, and a later reassignment passed every spelling check (review, claim 8). Shape test P pins both layers in both modes. Correction to the previous commit's message: it said an empty anchor was the only reason a missed collision could not over-select before that commit. The review drove a counterexample -- a tracked `.planning` path coinciding with the caller's subdirectory resolves a wrong, non-empty anchor -- so that sentence overstated it. Reversion controls, driven: the previous commit's undo.md fails 3 (P, the sub_repos refusal, defense in depth); the refusal made inert fails 2 (P and the refusal; the anchor check alone still keeps the range empty); the anchor check made inert fails 2 (P and defense in depth; the refusal alone still refuses). A verbatim negative control shows the pre-hardening root-commit arm widening the parent's anchor to all of HEAD. * fix(#4465): compare the repository, not the worktree, and anchor only inside HEAD's history Found by this round's third pre-push review, against the previous commit. Its repository gate compared per-worktree git directories. gsd-tools maps a linked worktree with no .planning/ of its own to the MAIN worktree (resolveMainWorktreeCwd), so from such a worktree PROJECT_ROOT is the main checkout: one repository, one object database, two per-worktree git dirs -- and the gate refused a legitimate undo. It now compares the COMMON git directory, physically resolved, which is the repository's identity; a sub_repos child is still a different one and is still refused. Driving that case showed the anchor check was too weak for it. The anchor comes from the main worktree's branch history, and nothing guaranteed the linked HEAD contains it: a linked branch that left main before the phase started would get a window bounded by a commit outside its own history. The check is now "PHASE_START is an ancestor of HEAD" (`git merge-base --is-ancestor`), which subsumes the previous "is a commit here" test -- a missing object is not an ancestor either -- and changes nothing in the ordinary case, where the anchor is read from HEAD's own log. Tests: a linked-worktree fixture (the linked checkout has no .planning/, which is what makes gsd-tools map it to main; the test asserts the mapping happened) is allowed in both modes and selects the phase; the same shape branched before the phase resolves no range. Shape test P pins the common-dir comparison, forbids --absolute-git-dir, and pins the ancestry check. Reversion controls, driven: the previous commit's undo.md fails 3 (P and both linked-worktree tests); per-worktree git dirs restored fails the same 3; the ancestry check replaced by the previous cat-file test fails 2 (P and the branched-before test, which then selects a linked commit whose history never held the phase); no anchor check at all fails 3 (P, defense in depth, the branched-before test). --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
26e650f312 |
fix(#4462): normalize whitespace workstream environment values (#4509)
* test(#4462): expose whitespace workstream scope * fix(#4462): normalize the workstream environment scope * docs(#4462): add changeset for #4509 * fix(#4462): normalize sibling planning scope readers * test(#4462): cover workstream normalization properties * fix(#4462): share normalized workstream resolution --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e5c4941f2a |
fix(#4485): isolate hook E2E tests from ambient capabilities (#4529)
* test(#4485): isolate hook E2E capability homes * test(#4485): isolate the remaining plan hook E2E * test(#4485): make ambient-home assertions real * test(#4485): remove stale helper imports --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e2bfc06558 |
fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase (#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until this lands. THE DEFECT, MEASURED Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking value, indistinguishable from the same call with a canonical id. Measured on |
||
|
|
85026f6a05 |
feat(#4668): add StateWriteIntent type surface and opaque-transform guard recognition (ADR-4629 C1) (#4676)
* feat(#4668): add StateWriteIntent type surface and opaque-transform guard recognition (ADR-4629 C1) Child C1 of epic #4629 — migration-order step (1) of ADR-4629: the guard/type scaffolding, with NO behavior change and no caller migrated. 1. StateWriteIntent (src/state-transition.cts) extends StateTransaction with the ADR-4629 section 8.1 concepts: field/section assertions marked required vs best-effort, plus a declared mutation scope (narrow | broad). Frozen like its base. createStateWriteIntent builds one from an existing transaction. Nothing in production constructs it yet — section 8.1's caller-side rule is Required in Phase 2; C2 (the verifying executor) and C3+ (caller migration) consume it. 2. findOpaqueStateTransforms (scripts/lint-state-write-path-drift.cjs) recognizes a residual readModifyWriteStateMd(path, (content) => ...) write whose transform is an inline anonymous arrow/function — the opaque shape section 8.1 replaces with a declared StateWriteIntent. readModifyWriteStateMd goes THROUGH the seam (it is not a raw-write bypass, Axis 2's concern), but its opaque body transform is neither verified (section 8.2) nor bounded (section 8.3). This ships recognition as a CAPABILITY: exported and unit-tested (positive control on a seeded fixture) but DELIBERATELY NOT wired into collect()'s failing scan. Wiring it now would turn the ~16 residual callers red at once, and ADR-3473 section 8.6 retired the ratchet that would otherwise absorb them. C2 wires it terminal as the verifying executor lands and callers migrate under ADR-3408 section 6 phasing. No behavior change: the guard is green on the tree (detection not wired), every state verb's output is unchanged, and the relevant suites (1763 tests) plus lint:ci pass. Regression tests are failing-first: positive/negative controls for the guard capability and a shape test for the type, plus a pin that collect() has no opaque-transform findings (C1 must not enforce; that is C2). Closes #4668 * chore(#4668): backfill changeset pr field to the real PR number (#4676) pr: 0 is rejected by parseFragment as invalid_pr (it is not a valid placeholder); the fragment must carry the real PR number, which fixes both changeset-lint and docs-lint (fail_invalid_fragment / fail_malformed_fragment). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5d4c98cde7 |
chore(#4729): guard the retired-runtime name, and finish the locale residue (#4753)
* chore(#4729): guard the retired-runtime name, and finish the locale residue Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one concern: make the tree clean, and keep it clean. The guard is inert until the tree is clean, and shipping the cleanup without the guard is the one-bug-at-a-time pattern this epic exists to end. WHY A GUARD, AND WHY LAST Nothing in CI answered "does any shipped surface still present a retired runtime as live?", and the two gates that look like they should cannot. checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it reported 0 violations while all four locale mirrors still documented --gemini as a live reviewer flag, with usage examples. And tests/gemini-runtime-removed.test.cjs is scoped by construction - its own docblock limits it to the installer CLI contract and the runtime-name-policy exports; it never reads docs/**, gsd-core/workflows/**, commands/** or agents/**. Every extension to it during this epic was a hand-added assertion for a surface somebody had already noticed. A guard written earlier would have red-flagged the very references phases 1b-4b were removing, which is why it lands last. PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE Each site was judged against its ENGLISH counterpart, not on its own: README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46 English README.md has ZERO occurrences -> substituted "Antigravity CLI, Kimi CLI" how-to/execute-a-phase.md:88 x4 locales fixed in #4728 -> substitute how-to/verify-and-ship.md:89 x4 locales fixed in #4728 -> substitute FEATURES.md cross-AI CLI list :1419 no Gemini -> DELETE FEATURES.md REQ-MULTI-RT-01 :1709 -> substitute FEATURES.md REQ-SKILLS-03 :1952 -> rewrite FEATURES.md REQ-QUOTA-02 :3256 deleted upstream -> delete VERSIONING.md:133 stale manifest -> see below The twelve README occurrences were an adversarial reviewer's BLOCKER, and the reason they survived my own sweep is structural: root-level *.md was outside the guard's scan set, so the repo's most-read runtime-advertising surface was invisible to the guard meant to police it. :46 is a live installer-runtime claim - it tells the reader the installer will offer a runtime that no longer exists. Checked for the duplicate-name trap before substituting: neither Antigravity nor Kimi appears anywhere in those four files. Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work, not a deferral. Two more show why "substitute Gemini -> Antigravity" is the wrong default: in the cross-AI list and REQ-QUOTA-02 English DELETES the name, because Antigravity was already in the list or the classifier had dropped it. Substituting would have duplicated a name - the identical trap ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot. VERSIONING.md:133 is a different and worse defect than translation lag. Under "Manifest Version Sync" it listed gemini-extension.json as a version-synced manifest. That file is ABSENT from the repo, and scripts/sync-manifest-versions.cjs says so in its own comment - "#1928: gemini-extension.json was removed with the gemini runtime ... it is no longer a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json, marketplace.json and vscode/package.json. So the doc named a manifest that does not exist AND omitted the one that replaced it. Both fixed, verified against the owning code rather than inferred from the name. The replacement bullet cites #1942, the issue that actually registered vscode/package.json, matching the convention of its neighbours. pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has no REQ-QUOTA-02 line. Skipped and recorded, never invented. PART B - THE GUARD scripts/lint-retired-runtime-name.cjs, modelled on scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem shape (forbid a retired token, allowlist frozen content, self-exempt via a split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1). Case sensitivity IS the mechanism, not an accident. The naive guard - "the string gemini must not appear" - is WRONG, not merely noisy: that string is load-bearing across Antigravity's real on-disk contract. A case-sensitive, standalone, capitalised name works because every legitimate reference is spelled differently and therefore cannot match: lowercase config homes (~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and GEMINI.md. Table-driven, so the next retired runtime costs one row. THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two rounds of isolated adversarial review reshaped it; both are recorded in .gsd/bug/chore-4729-gemini-drift-guard/60-review.json. ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can never be complete, so "no runtime word found" silently exempted every phrasing nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`, `Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0, as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and `Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime` or `运行时`. The fix was to stop enumerating exceptions and invert the evidence direction: Tier 1(a) - the hook DIALECT Antigravity inherits. Position is language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The marker must now form an ADJACENT COMPOUND with the name, not merely sit in a +/-24-character window - that window let `| Antigravity | Gemini-style hooks | Gemini support is live |` exit 0, one legitimate reference licensing a fresh live claim 21 characters later. The runtime-word veto is now LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files" in the same sentence); each is an explicit pin rather than a reason to loosen the veto for everyone. Measured: widening it surfaced exactly those ten and no others. Tier 1(b) - the provider/model axis. A version optionally followed by a qualifier, including full-width digits and CJK punctuation, AND positive model-axis evidence on the line, AND no runtime word. The positive requirement is the part that matters: all eight real model-axis lines in the repo name a model explicitly, so requiring it costs nothing on the real tree while flagging every laundering attempt. It is also the honest resolution of the agent/target tension below - rather than guess at an exhaustive veto list, stop treating an empty veto as evidence. Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match falling INSIDE an occurrence of its own snippet. Line-level containment let `Known provider menu update: Gemini CLI is once again a selectable GSD runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your GSD runtime.` both exit 0, because a short snippet elsewhere on the line pre-approved a brand-new claim. Span scoping makes short snippets safe: `Google (Gemini)` can only ever excuse the match inside those 15 characters. A LOAD-TIME validator now requires every pin to contain a retired name, and it immediately caught five of MY OWN pins whose snippets sat BESIDE the name rather than covering it - each would have shipped permanently inert and permanently reported stale. All pins were then reconciled in one pass. A pin is also marked used by PRESENCE on the line now, rather than only on the Tier-2 branch. Previously a pinned line that a general rule also matched never marked its pin used, producing a provably FALSE "no line matches pinned snippet" whose printed remedy told the maintainer to delete a pin that was still needed. Tier 3 - blanket trust, and a new occurrence inside it IS invisible. CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one surface - plus six append-only directories. All 21 `.changeset/` hits were measured to be fragments DESCRIBING the retirement or a fix to it, 464 of them under archived/; a fragment can only describe what already shipped and is deleted at release, so pinning them would be friction with no signal. The cost is stated in the guard's own header rather than hidden. THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and `references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a live claim introduced there was invisible at BOTH ends. The escape hatch must now carry a justification (`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is checked first, excuses the whole line, and the failure message advertises it, so an unexplained one is indistinguishable from a silenced defect. Plus an anti-vacuity floor counting files actually READ, not files listed - a candidate count stays healthy-looking even if every read failed. A FALSE NEGATIVE I INTRODUCED, AND CLOSED The model-display escape began as a blanket /^ \d/ - "space then a digit" - which also matched "Install for Gemini 2.5 CLI as a supported runtime.", laundering a genuine stale-runtime claim through an attached version number. That was the THIRD appearance of one failure shape in this epic: an exclusion added to suppress false positives creating a false negative. #4716's sweep excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby hid a stale review.models.gemini row whose example value was "gemini-2.5-pro" ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which is why the fix this time was to invert the rule's evidence direction rather than to enumerate more exceptions. The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI" matched inside "client" and would have vetoed legitimate model lists - and raw for CJK terms, where \b is ASCII-word-based and would never fire beside an ideograph, so anchoring them would silently disable the veto in ja/ko/zh. "agent" and "target" were deliberately left OUT: both occur throughout ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so vetoing on them would red correct content instead of catching runtime claims. The reasoning is in the guard's comment, not just the omission - and Tier 1(b)'s positive-evidence requirement is what makes that omission safe, since the rule no longer depends on the veto list being complete. COVERAGE tests/lint-retired-runtime-name.test.cjs drives the guard through its GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a guard, and this epic already shipped one that was vacuous for 2 of its 5 files, so properties are paired against BOTH failure modes - too broad silently absorbs a future defect, too narrow reds on legitimate content. Floor boundaries are covered at 149/150/151. The round-2 reviewer's sharpest point was about that claim, and it was right: the first matrix's pairing was "true of the properties chosen, not of the predicate's actual surface" - not one of its twenty properties could see the dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an over-broad pin colliding with a new line. Every one of those is now a committed regression using the reviewer's own attack line verbatim, and the local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0). That harness earned a finding of its own. Its first run reported PASS=2 FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so nothing staged, every fixture hit the empty-walk error path, and the two checks that assert an ABSENCE passed off that error path rather than off real guard logic. A staging failure is now fatal and every absence-asserting check first proves the walk ran and the expected violation was flagged. Later, one case failed because its fixture supplied only one of a pinned file's two approved lines, so the stale-pin check fired correctly - the expectation was wrong, not the guard. Telling those two apart is the whole value of running a matrix rather than reasoning about one. On the two orthogonal reviews: the isolated adversarial pass executed a great deal of code, across two rounds, against its own fixture repos. The security pass did NOT - it self-discloses that it verified by reading only, because node --test is hard-blocked here. Saying so plainly, because "two orthogonal reviews" without that caveat overstates what the second one established. It also raised, and I cleared by measurement, a concern that importing escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt clone: six other tracked scripts already require that exact path, three of them already in lint:ci, and .github/workflows/test.yml:192-193 runs `npm run build:lib` immediately before it for exactly this reason. Part A has no new test deliberately - those edits are covered by the guard itself inside lint:ci, and a separate per-locale assertion would duplicate it and then drift from it. The one exception is the root README case, which IS pinned: that residue was invisible to the guard rather than merely unasserted, so the fix is a scan-set change and needs its own regression test. No mode-bit read-failure fixture was added on purpose: the benches run as root, where chmod-based IO injection is vacuous, so such a test would assert nothing. The test's fixture helpers write throwaway docs/ paths, which trips lint-docs-guard-registration's reader-name heuristic. Resolved the way that lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry - because the fixtures only WRITE scratch data and never read shipped docs; the baseline was re-confirmed, not merely extended, each time locale and adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own --write path, since a new test file changes the count lint:generated-sync reads. Fixes #4729 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4729): backfill changeset PR number (#4753) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
39db16a653 |
fix(#4678): resolve :LINE citations against their underlying files (#4719)
* fix(#4678): check :LINE citations instead of dropping or misreporting them cmdVerifyReferences had two opposite failures on citations carrying a trailing line suffix. The backtick extractor anchored the closing backtick right after the extension, so `src/foo.ts:42` matched neither regex and was silently dropped -- an all-line-numbered document reported {valid:true, found:0, missing:[], total:0}, indistinguishable from one with no citations. The @-extractor treated ':' as a legal path character, so @src/foo.ts:42 was probed with the suffix glued on and a resolvable file was reported missing. Strip the :N / :N-M suffix (stripLineSuffix) before existsSync in both loops -- for filesystem resolution only; found/missing keep reporting the original citation text -- and let the backtick regex accept the optional suffix so those citations are counted at all. URL skip, template-placeholder skip, dedup, ~/ expansion and the output contract are unchanged. Whether a line number past EOF counts as missing is an open design question and stays out of scope. * chore(#4678): backfill changeset pr field with PR number The fragment shipped as the documented pr: 0 placeholder; the merge gate requires pr > 0 before a fragment can land, so backfill 4719 now that the PR number exists. --------- Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
48271de43f |
fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis (#4744)
* test(#4660): pin the letter-axis parity defect across all 6 shell/markdown phase-id sites Extends tests/nsegment-phase-grammar.test.cjs (#4568) one axis over: for each of the six sites, reads the live regex off disk and asserts it agrees with src/phase-id.cts's PHASE_NUMBER_TOKEN_SOURCE on the letter axis in BOTH directions — accepts `12A` / `3A` / `03A` / `23A.1.2`, still rejects `3a`, `3AB`, `A3` and the other canonical-invalid shapes — and that the two extracting sites return the full letter-suffixed token rather than its digit prefix (or nothing). Negative control against the unfixed tree: 22 failures, exactly the "(fails before the fix)" cases; every reject-parity case already green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis Adds `[A-Z]?` after the leading digit run at all six sites #4568 widened — the ERE translation of src/phase-id.cts's `\d+[A-Z]?(?:\.\d+)*` — so a documented, canonical-valid id like `12A` or `23A.1.2` is no longer refused by the four validating sites (code-review.md, code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md) or truncated to its digit prefix by the two extracting sites (execute-plan.md's plan-filename grep, plan-phase.md's --research-phase capture). Behaviour is byte-identical for every id that matched before; the adjacent comment and error-message text now names the grammar it mirrors. Driven: `init code-review 3A` on a fixture with a `03A-slug/` directory and a `### Phase 3A:` heading emits `padded_phase: "03A"`, which the old regex rejects and the widened one accepts — nothing upstream of the validator mangles the id. At execute-plan.md the trailing `-[0-9]+` is the PLAN number and stays digit-only; plan and milestone dimensions are out of scope per the brief. `CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE` derives from the canonical source by a literal `.replaceAll('A-Z', 'A-Za-z')`, so src/phase-id.cts is deliberately untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4634): extend lint-phase-id-drift to ban a letter-less phase-id mirror in workflows/ and agents/ Adds findLetterlessPhaseMirrorDrift — the letter-axis twin of the #4568 single-segment rule — flagging the unbounded-segment shape `[0-9]+(\.[0-9]+)*` (and its \d / doubled-backslash near-variants) whose digit run is NOT followed by the `[A-Z]?` class, on any phase-carrying line across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md and agents/**/*.md. Sanctioned the same way (`<!-- phase-id-owner: ... -->`), tolerates the case-flexible `[A-Za-z]?` directory-scanning variant so it cannot force that separate axis to narrow, and is wired into scanAll. Confirmed zero violations against the real tree post-#4660 fix, and one violation when a single site is reverted. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * docs(#4660): add Fixed changeset Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore: regenerate conformance-tier manifests for the extended grammar test tests/nsegment-phase-grammar.test.cjs now requires the compiled gsd-core/bin/lib/phase-id.cjs (to assert the canonical grammar agrees with each site's live regex), which moves it to a different platform-conformance tier; `gen-platform-conformance-tier.cjs --check` in lint:ci flagged the macOS manifest as stale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * test(#4660): reword a comment that tripped lint-docs-guard-registration The comment mentioned `docs/CONFIGURATION.md` between two backticked tokens, which the lint's template-literal detector read as a docs/ path expression. The test reads no docs/ file. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): refresh the compact-content benchmark baseline and acknowledge emitted growth plan-phase.md grew by 4 bytes (`[A-Z]?`), which moves the committed compact-content benchmark; refreshed with `benchmark-compact-content.cjs --write`. The six shipped files below grew by the widened regex literal plus the comment and error-message text that now names the canonical grammar. Emitted-Drift-Ack-Growth: code-review.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: code-review-fix.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: gsd-code-fixer.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the defense-in-depth comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: execute-plan.md — #4660: `[A-Z]?` in the plan-filename phase extraction (6 bytes) Emitted-Drift-Ack-Growth: plan-phase.md — #4660: `[A-Z]?` in the --research-phase capture (6 bytes) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): set changeset fragment pr to 4744 * chore: re-trigger Validate Branch Name The required check-branch context was cancelled on this head by the workflow's cancel-in-progress group when the changeset pr-field backfill push landed three seconds after the PR opened; no completed run exists for the current head, and a fork contributor cannot re-run it. Empty commit to re-run it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
06845717fe |
feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition Failing-first coverage for the Loop Host Contract role partition. At this commit crossCheckRoleFamilies does not exist, so the rows throw "crossCheckRoleFamilies is not a function" -- the RED proof they bind to behavior rather than restating it. ADR-894 section 3 assigns roles per step but parenthesises the assignment as "(illustrative roles)", and nothing enforced it. The only thing standing in the way was a single deepEqual in this same file, which is editable prose. Rows cover: each step's own family accepted; a strict subset accepted; a foreign role rejected at every step; an unknown role rejected; an unknown step failing CLOSED; capitalization not silently matched; every offending role reported rather than only the first; and purity, because buildContract puts the same array into the generated contract. Two rows exist because an earlier cut of this suite was vacuous. The purity fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot fail an in-place sort(), and the mutant was being killed by three unrelated rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the exact same role-name domain, both directions: they are parallel constants over one domain, so divergence is the generative-fix class CLAUDE.md names. Every negative row asserts the offending ROLE NAME and the STEP NAME appear in the message. A count-only assertion survives a mutant that reports the wrong role, which the 80% Stryker gate would surface only after a full CI round-trip. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#4740): reject a cross-family agent-role declaration Orchestration and execution are distinct functions of the loop and must not drift into one another. That partition was real but unenforced: ADR-894 section 3 calls its own role assignment "illustrative", and the generator accepted anything. Adding orchestrator to execute-phase.md's agent-roles line compiled, --check passed once regenerated, and capability-validator.cjs then began accepting into:"orchestrator" at every execute point. ROLE_FAMILY maps every role to one of orchestration, planning or execution. EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family. crossCheckRoleFamilies rejects a cross-family role, a role outside the vocabulary, and an unknown step. It reports every offender, not the first. It fails CLOSED on an unknown step, deliberately diverging from assertPointsCoverage's "unknown step -- caught elsewhere". For points that is true: the canonical-set and duplicate checks catch it. For roles there is no second net, so failing open would leave an unknown step as the one input that bypasses the gate. crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles to agent FILES and the orchestrator is the host, owning none -- admissibility and agent-file presence are separate concerns with separate checks. Additive to section 3's existing rule that contribution.into must be a member of the step's agentRoles, which is unchanged. That governs what a CAPABILITY may target; this governs what a WORKFLOW may declare. No capability is affected, and all five workflows already declare single-family sets, so the gate is green on the commit that introduces it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4740): make the ADR-894 role assignment normative Section 3 parenthesises its per-step role assignment as "(illustrative roles)". That word was accurate about the list's PURPOSE -- it illustrated the shape of a generated contract entry -- and wrong about its STATUS, because the assignment was load-bearing from the moment the generator consumed it. Read literally it makes the partition an example rather than a rule. Appended as a dated in-place section per docs/contributor-standards.md, which records that an accepted ADR is never rewritten and names this the default pattern. Section 3's original body is untouched. The amendment states the three disjoint families, the one family each step admits, that a step may declare a strict subset but never outside it, and why this is a clarification rather than a new decision: the contract is generated from the workflow markers "so it cannot drift into a lie", and all five workflows have always declared single-family sets. What was absent was any statement that it is required, and any check that it holds. It also pins the distinction that is easy to re-merge: contribution.into being a member of agentRoles governs what a CAPABILITY may target and is unchanged; the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary entry for the Loop Host Contract records the same, beside the agent-reference drift guard it already documented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): backfill changeset pr number Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both go from invalid_pr(0) to ok. Without that env both report success without evaluating the branch at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): stop injecting the orchestrator procedure into executors claude-orchestration declared a contribution at execute:wave:pre with into:"executor". loop-hook-dispatch.md defines a contribution as "inject fragment.inline verbatim into the context for the role named in into", so its 267 lines were injected into EXECUTOR prompts whenever the capability was enabled. Those lines are orchestration end to end -- construct a wave manifest, resolve the dispatch backend, invoke the Workflow tool to spawn executors, bridge per-agent results into the merge chain. An executor can act on none of it. Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT carries no orchestrator entry by design: the orchestrator IS the host, and the host's procedure lives in execute-phase.md. A step's agentRoles enumerates agents a capability may inject context INTO, so adding orchestrator there would model the host as an injectable agent -- the same category error pointed the other way, and it would need an exception carved into the partition the same issue just made normative. So the defect is the mechanism, not the label. A contribution injects into an agent's context; "replace step 3's inline dispatch loop" is a change to what the HOST does. The contribution channel was serving as a host-behaviour directive because it was the only channel available at an execute point. The entry is removed. plan:post into:"planner" is correct and untouched. The procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the capability -- it is the only copy in the repo -- and is no longer injected anywhere. Consequence, not softened: the Workflow backend now has no loop wiring. Detection, emission and config remain and the design is intact, but nothing dispatches it. Under the separation ADR-1143 itself asserts it never had a legitimate channel; ADR-1143's own audit already records the end-to-end path has never been exercised. Wiring it properly needs a host-level mechanism that does not exist today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): invert the stale execute:wave:pre registry assertions Removing the contribution left four surfaces asserting or describing the old state. Caught by an isolated review before a verification run was spent, which is the point of reviewing first: the first of these was a guaranteed CI red. execute-wave-post-gate-pipeline-e2e asserted against the REAL generated registry that byLoopPoint['execute:wave:pre'] held exactly one contribution with capId claude-orchestration. It now holds zero. Inverted to assert exactly 0 -- not a vague >= 0 -- and the #2285 comment above it now explains the current state rather than the one it was written for. CONTEXT.md's Claude Orchestration entry claimed two contributions at wired points. It is now one, and the entry's execute:wave:post label was already wrong before this change: the manifest said execute:wave:pre. Rewritten to one plan:post contribution, why the execute-point one was removed, and where the procedure now lives. One assertion in claude-orchestration.test.cjs could not fail. It tested for the prose "(into the executor)" while the doc says "(`into: executor`)", so no plausible wording matched it and the paired plan:post assertion was carrying the row. Replaced with a check on the structural claim, and proved RED by restoring the two-contribution wording before reverting. The moved procedure keeps section headings that speak as a live contribution -- "When this contribution is active", "Why execute:wave:pre". Preserving the body verbatim was deliberate, so the headings stay and an editor's note under the header explains why they read that way. A sweep of all 17 files referencing byLoopPoint found no further siblings: the remaining hits are a synthetic capability fixture and an empty-points test that already expected no active hooks, both correct before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eb49ff98df |
fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime
#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.
The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:
- install-on-your-runtime.md English has NO `### Gemini CLI` section -> deleted
- USER-GUIDE.md :843 "…, Antigravity CLI, Kilo)" -> substituted
- ARCHITECTURE.md English has NO Gemini CLI table row -> row deleted
- ARCHITECTURE.md :24 English holds `Kimi CLI` in that slot -> Kimi CLI
- context-monitor.md :3 "`AfterTool` for Antigravity CLI" -> substituted
- spike-and-sketch.md :93 "(Codex, Antigravity CLI, etc.)" -> substituted
- configure-model-profiles "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
- COMMANDS.md English keeps only hyphen + Codex bullets -> colon bullet deleted
- FEATURES.md source docs/features/multi-runtime-support.md:10
lists no Gemini CLI -> name removed
ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.
The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.
The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.
Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.
Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.
Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.
Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.
Fixes #4728
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate
A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.
1. I committed the exact error I claimed to have avoided. The commit message
boasted that ARCHITECTURE.md:24 proved the value of reading English rather
than substituting blind, because Antigravity already appeared later in that
list. Five hundred lines further down the SAME four files, my
`Gemini:` -> `Antigravity:` substitution produced TWO consecutive
`- Antigravity:` bullets, because an Antigravity bullet was already there.
English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
locales, reusing each locale's existing words.
2. `--gemini` survived in the runtime-detection CLI flag list in all four
locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
already lists `--antigravity` later, so this is another place where
substituting Antigravity would have duplicated it. Now `--kimi`.
3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
line I had already corrected -- the "Adaptive (Recommended)" option in
settings.md:192 and new-project/steps/auto-mode-config.md:95.
4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
`non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
without the word "runtimes", and execute-phase.md:1028 has no parenthetical
at all. The reviewer proved it by re-introducing Gemini at both lines and
watching the assertion stay GREEN. That same blind spot is what hid finding 3.
Replaced with a case-sensitive `/\bGemini\b/` walk over every
`gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
reference in that tree is spelled differently and cannot match: Antigravity's
paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
`Gemini` there means the retired RUNTIME is being named. The walk asserts it
found at least 50 files so an empty walk cannot pass vacuously, and it now
covers the nested `new-project/steps/` directory where finding 3 lived.
Two allowlist entries, both by line CONTENT and both justified:
reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
`Known provider` menu. The second was escalated by the agent rather than
decided: Section 8 of that file says model policy is defined "independently"
of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
the same axis as the lowercase model ids -- and must keep working.
Proven to fail, not just asserted: the predicate reports 0 offenders on the
real tree and exactly 2 on a /tmp copy with Gemini re-injected at
health.md:52 and execute-phase.md:1028.
Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.
The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.
Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @
|
||
|
|
b647e28313 |
chore(#4727): name the tool-conversion helpers for the runtime that uses them (#4732)
* chore(#4727): name the tool-conversion helpers for the runtime that uses them
GSD has had no Gemini runtime since #1928 removed it (Google sunset Gemini CLI
on 2026-06-18, shipped 1.8.0), yet two helpers were still named for it:
claudeToGeminiTools -> claudeToAntigravityTools
convertGeminiToolName -> convertAntigravityToolName
The sole consumer is convertClaudeAgentToAntigravityAgent, whose own comment
read "Map tools to Gemini equivalents (reuse existing convertGeminiToolName)".
Nothing named Gemini consumes them, because nothing named Gemini exists. The
new names follow the convention the file already sets with its neighbouring
Copilot pair, claudeToCopilotTools / convertCopilotToolName.
Zero behavior change. Every mapped VALUE is byte-identical, deliberately:
read_file, write_file, replace, run_shell_command, glob,
search_file_content, google_web_search, web_fetch, write_todos
Those are Gemini's built-in tool dialect and Antigravity genuinely speaks it.
This rename covers only the identifiers, which are the one part of the surface
that was GSD's choice rather than Google's contract.
Renamed in BOTH copies. CLAUDE.md labels bin/install.js "(generated)", but no
script emits it -- build:lib is tsc -p tsconfig.build.json and writes only
gsd-core/bin/lib/**. These converters are the #1099/#1173/#1182 situation: they
were extracted into src/runtime-artifact-conversion.cts while bin/install.js
kept its own working inline copies, so each symbol existed twice in two
independently hand-maintained files. Renaming one would have left two names for
one concept. Verified first that no capability descriptor resolves either by
name -- antigravity's descriptor names only convertClaudeCommandToAntigravitySkill
and convertClaudeAgentToAntigravityAgent, neither of which moved.
Comments keep their reasoning and their issue refs (#3362 AskUserQuestion,
#1394 Skill/SlashCommand); only the subject is corrected, from "Gemini CLI" to
Antigravity speaking the Gemini dialect. Those describe the dialect's behavior,
which is still Antigravity's behavior, so deleting them would destroy the record
of two real bugs.
docs/research/gemini-to-antigravity-migration.md is left unedited and carries a
dated addendum instead: it is pinned to
|
||
|
|
9b750dc00a |
fix(#4505): resolve models through the active runtime and the tier table (#4726)
* test(#4505): cover runtime-aware overrides and routing precedence Failing-first for both halves of the issue, plus the precedence layers a naive fix silently defeats. Every row drives the REAL CLI in a subprocess. That is load-bearing: the defect is WHICH function the shipped call sites reach, so a row calling the resolver in-process would pass while every real spawn stayed broken. It also makes GSD_RUNTIME hermetic -- it is ambient, and an in-process row would leak it into its neighbours. Two fixture mechanics are documented in the helper because each silently invalidates a row when got wrong, and both were found by measuring rather than by reading the loader: - the loader reads `process.env.GSD_HOME || os.homedir()`, so redirecting only HOME leaves a developer's real ~/.gsd/defaults.json in play; - the mere EXISTENCE of a .planning/ directory disables the shared-defaults layer, so a fixture that creates one stops exercising the "poisoned global" path #2297 acceptance #4 is about. Measured: .planning/ with config -> gpt-5.6-terra; .planning/ present but empty -> gpt-5.6-terra; no .planning/ at all -> "". Rows cover: both reported repros; the init payload a real spawn reads; the omit gate, the runtime tier map and model_overrides each outranking the tier table; model/tier coherence under dynamic routing; resolve-execution with the attempt absent; max_escalations at limit-1/limit/limit+1 plus a cap of 0; and four fail-safe rows pinning that only a value canonicalizing to a recognised non-Claude runtime may outrank an omit. Registers the docs-guard exemption path: the rows quote the documented first-spawn contract in comments. The file still never READS a docs/ path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4505): resolve models through the active runtime and the tier table Consolidates #4495 and #4493. One gap: the function every agent spawn goes through consulted neither mechanism that was supposed to make resolution runtime- and tier-aware. Half A -- the runtime was READ from config, not resolved. resolveActiveRuntime (GSD_RUNTIME -> config.runtime -> per-install marker -> claude) existed and worked, but was called exactly once in the file. Every other site read config['runtime'] raw, and that key is normally absent, so model_profile_overrides.<runtime>.<tier> was inert for any install that identifies its runtime through the environment or the marker. Same override both times, differing only in WHERE the runtime is declared: GSD_RUNTIME=opencode, no runtime key -> sonnet (ignored) runtime:"opencode" in the config -> TEST-OPENCODE (works) Half B -- nothing consulted dynamic_routing on the first spawn, though docs/features/dynamic-routing-with-failure-tier-escalation.md documents "the resolver picks tier_models[default_tier] for the FIRST spawn". The tier-table lookup is extracted into ONE helper both entry points call, so the first-spawn value and the escalated value cannot drift; resolveModelInternal calls it at attempt 0 and resolveModelForTier at the real attempt. Placement is the documented composition, not a convenience. The same doc says "model_overrides always wins; dynamic_routing.tier_models[<tier>] resolves above models.<phase_type> and model_profile" -- so the step sits BELOW model_overrides, the model_policy preset, the runtime tier map, the resolve_model_ids:"omit" gate and the claude tier override, and ABOVE the profile lookup. An earlier cut routed every call site through resolveModelForTier instead, which returns the tier model directly and therefore skipped three of those layers: with an omit and a non-Claude runtime it handed out a model id where the gate had returned "". Criterion 1 is applied in full, including the two value-policy reads #4192 had recorded as "NOT via resolveActiveRuntime". The tests decided it: switching them breaks nothing, so that reading was never enforced -- and the old behaviour defeated #4192's own principle that an explicit pin must not be silently unpinned (claude-opus-4-8 under GSD_RUNTIME=opencode collapsed to the Claude-only alias opus). #4192's comment is updated in place rather than left stale. The step-3 opt-in signal is CANONICALIZED. Comparing the raw config field against the literal 'claude' made runtime:"Claude", "claude-code" and even 5 count as non-Claude opt-ins and outrank an explicit omit -- failing OPEN in exactly the #2297 case the guard exists to protect. null now covers both "not a string" and "not a runtime we recognise", and both read as NOT an opt-in. Verified cell by cell against a pristine origin/next worktree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4505): backfill changeset PR number (#4726) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9c2927bff4 |
fix(#4733): derive the win32 chunk cap, isolation bar, and unknown-file weight (#4737)
* test(#4733): pin the cap, unknown-file weight, and isolation rules Failing-first coverage for the three defects that let a Windows conformance chunk be killed at the 600s per-chunk backstop with zero failing tests. The previous boundary rows were VACUOUS: they asserted literal arithmetic (21 * 18122 <= 400000) that cannot fail, and in doing so masked a shipped win32 cap of 23 -- a value that violates the very inequality they claimed to pin. These rows constrain defaultMaxFilesPerChunk itself, from both sides, so the shipped value is a derived maximum rather than a magic number. A second vacuous row was caught by review and removed: it recomputed the isolated set from the function under test using the identical predicate, so it was empty by construction. It is replaced by an exact deepEqual against the expected basenames, a cross-platform identity row, dynamism rows in both directions, an inclusive boundary triplet, and invalid-threshold throw rows. The cross-platform identity row is the regression guard for a threshold that was briefly anchored to the per-platform file-COUNT cap; it fails if isolation ever becomes platform-dependent again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4733): derive the win32 cap, isolation bar, and unknown weight A Windows conformance chunk was killed at the 600000ms per-chunk backstop with no test having failed, taking next red. Three compounding defects. The win32 cap of 40 permitted 40 * 18122 = 724880ms against a 600000ms backstop -- 121% of it -- so two rounds of budget-tuning could not hold. The cap is now derived: 22 is the largest value satisfying cap * 18122 <= 400000. The budget is 400000, not the raw backstop, because the chunk that died summed to only ~348328ms of per-file time -- a per-chunk overhead gap of at least 1.72x that no per-file table models. A file absent from the timings table was priced at medianWeight. The table is skewed 18.8x, so an unknown weighed 0.0533 -- 19x cheaper than average, and measured 17.5x under its real cost. Unknowns are now priced at the mean. ISOLATED_HEAVY_FILES was a static Set, stale by construction. Isolation is now derived from an absolute ms bar (0.3 * 400000 = 120000ms) converted to weight units via the live table's mean, so a file that gets heavy is isolated automatically instead of waiting for someone to edit a list. Review caught that an earlier cut anchored that bar to the per-platform file-COUNT cap -- a category error, count vs weight, which silently returned seven of the historical eight files to the shared pool on linux/darwin. Since macOS runs the full matrix only after merge, that would have planted a red next no PR could catch. The bar is absolute and platform-independent. Also from review: isolation no longer requires unit-suite membership, so fragment-single-edit-propagation.install.test.cjs -- 575000ms, 96% of the backstop in one file -- is eligible; partitionIsolatedFiles throws on a non-finite or non-positive threshold instead of silently isolating nothing; and stale per-shard figures no test pinned are removed rather than recomputed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4733): backfill changeset pr number --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ca8d9d4459 |
fix(#4429): stop a large commit_types config blocking or bypassing the gate (#4723)
* test(#4429): regression coverage for three defects in the commit hook Failing-first coverage. Every conforming-subject row is red against the unfixed hook, and each defect gets an explicit CONTROL row that reconstructs the pre-fix form and asserts the defect reproduces -- without those, the passing rows would pass with or without the fix. 1. SIGPIPE (the reported defect). The pre-fix first-line extraction used a `head -1` pipeline; once CONFIG_OUT exceeds the 64 KiB pipe buffer printf is killed and `set -euo pipefail` aborts the hook. That fix is already on next -- it landed incidentally in #4537, whose message never mentions #4429 -- and nothing in the tree would notice its removal. 2. regcomp. The commit-type alternation grew with the CONFIGURED list and exceeded bash's 64 KiB compiled-pattern cap. Boundary rows pin the cliff at 6051/6052, with controls on BOTH sides so limit-1 is not vacuous. 3. Ambient subprocess statuses (found by this change's security review). Defects 1 and 2 cannot be separated: each configured type adds len+1 bytes to CONFIG_OUT and len+1 to the alternation, so the smallest payload that overflows the pipe (N=6059) already puts the alternation past the ceiling. The SIGPIPE control accepts either SIGPIPE (141, Linux) or a reported write error (macOS bash 3.2's builtin printf, exit 1). Asserting only the message would go red on every CI lane, since the remote matrix is Linux-only. Named to bucket with gsd-validate-commit-crash-policy.test.cjs, which covers this same hook: lint-test-file-count derives a test's owning module from its filename prefix, and `validate-commit-*` collided with the `validate` module, already at its 2-file cap. Harness note, learned from three vacuous control runs: hooks/lib/git-cmd.js requires ../gsd-core/bin/lib/token-scanner.cjs relative to the hooks dir's parent, so a copy in a bare tmpdir fails open and returns 0 for any input. The layout symlinks gsd-core beside the copy, and every row that can prove it asserts the run was substantive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4429): bound the commit-type regex and isolate subprocess statuses Two fixes in the same file, both of the same shape: a value computed for one purpose was being read as authority about something else. 1. The commit-type alternation could not be compiled. COMMIT_TYPE_ALT joined every CONFIGURED type into one regex, so the pattern grew without bound. bash caps a compiled pattern at 64 KiB. Bisected on bash 3.2.57 (this repo's macOS target): a 65504-byte alternation compiles, 65515 fails. `[[ =~ ]]` returns 2 on a compile failure, and `if !` cannot tell that from "the subject does not conform" -- so the hook blocked a valid `feat(auth): ...` with CONVENTIONAL_COMMITS_VIOLATION while printing `feat` in its own valid_types. Match the shape with a fixed-size pattern, capture the type, then test membership against the COMMIT_TYPES array. The character class is exactly the `^[a-z][a-z0-9-]*$` safe-token filter the config loader already applies, so it captures every type that can legally reach COMMIT_TYPES and no token that cannot. Review verified equivalence over 46 handcrafted plus 6000 randomized adversarial subjects against a type list containing prefix-overlapping, digit-bearing and trailing-hyphen types: zero divergences. The loop adds no subprocess and no pipe, which is the hazard class #4429 is about. COMMIT_TYPE_ALT is now unused and removed. types pre-fix `feat(auth): ...` fixed 10 accept accept 6051 accept accept 6052 BLOCK accept 20000 BLOCK accept 2. Subprocess statuses were inherited from the environment. Each status is captured as `... || VAR=$?`, which assigns ONLY on the failure branch; on success the variable kept whatever it already held, and `${VAR:-0}` defaults only when unset or empty. So an EXPORTED CONFIG_STATUS, CMD_STATUS or CLASSIFY_STATUS -- from a CI wrapper, a .envrc, or another hook -- survived into the success path and was read as "the subprocess failed". Since the hook fails OPEN on a genuine subprocess failure by design (#3838), the result was a silent bypass. Measured: `CLASSIFY_STATUS=3 git commit -m "nope: bad"` printed "validator disabled for this call" and exited 0. The three are now initialised before use. The fail-open path is unchanged and verified byte-identical to origin/next with a failing node. hooks/dist/ is gitignored and rebuilt from hooks/ by scripts/build-hooks.js, so there is no second copy to sync. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): register the new suite with the conformance manifests Both conformance-tier manifests embed the test-file list, so adding a test file makes them stale. Regenerated with their own generators: node scripts/gen-platform-conformance-tier.cjs --write node scripts/gen-platform-conformance-tier.cjs --target macos --write Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): pin both fail-open-prone controls to their named cause Two rows in the ambient-status block asserted `status === 0`, which the hook also returns when the harness layout is broken -- so either row could have passed for entirely the wrong reason. This is the same vacuity trap the rest of the suite already guards, applied inconsistently to the rows added last. Measured, rather than reasoned about: genuine ambient bypass (pre-fix hook, CLASSIFY_STATUS=3) rc=0, no CLASSIFIER_THREW orphaned layout (no gsd-core symlink) rc=0, CLASSIFIER_THREW genuine fail-open (node shim exits 3) rc=0, no CLASSIFIER_THREW So assertSubstantive separates the intended cause from the harness failure in both rows, and each now pins its pass to the cause it names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): stop asserting a macOS-only regex cap on every platform First verification run was RED: 45636/45638 passed, both failures in this new suite on linux-node24. Cause is mine -- I measured the compiled-pattern ceiling on macOS and encoded it as a cross-platform expectation. Measured in the tester image itself: engine 6051 6052 20000 bash 3.2.57 / BSD libc (macOS) compiles rc 2 rc 2 bash 5.2.15 / glibc (Linux) compiles compiles compiles (228943 B) glibc has no reachable cap, so the regcomp defect cannot occur there and the control asserting a block at 6052 was red for a behaviour the platform cannot produce. The control now calibrates at runtime: it runs the pre-fix form and, when this engine compiled the alternation, it SKIPS with a message naming the reason rather than asserting. Skipped out loud, never silently passed -- a green row there would read as "the defect is covered" on a platform where it cannot occur. Both branches verified: the capped branch asserts (macOS 17/17, zero skipped), and the uncapped branch was exercised by forcing the payload to a size that always compiles, producing a skip and not a failure. Consequence stated rather than hidden: the remote matrix is Linux-only, so this one control is skipped in CI and really runs only on a macOS workstation. The rows that run everywhere are the ones carrying the regression weight -- the shipped hook accepting a conforming commit at every payload size, the gate still blocking unknown types, the SIGPIPE control, and all seven ambient-status rows. Note this also narrows the coupling claim: SIGPIPE and regcomp are coupled only on a capped engine. On glibc the SIGPIPE defect is directly testable without the regcomp fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4429): scope the regex-cap claim to the platform it applies to The changeset told users the validator "built a regular expression bigger than bash can compile" past ~6,000 configured types. That is false on Linux: glibc compiled a 228,943-byte alternation without complaint, so a Linux reader would have been misled about their own exposure. These are user-facing release notes, so the claim is now scoped to macOS (bash 3.2 / BSD libc) and says explicitly that glibc was never affected by this half. The hook's own comment led with the same overstatement -- "bash caps a compiled pattern at 64 KiB" -- before qualifying it. Reworded so the first clause states what is actually true: the limit is a property of the platform's regex engine. Text only; no behaviour change. Suite 17/17, eslint and lint:ci clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): backfill changeset PR number (#4723) * chore(#4429): backfill changeset PR number (#4723) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
40628d8050 |
fix(#4409): delete the fold's shadow copies of extractShellBlocks and collectWorkflowFiles (#4720)
* test(#4409): prove the fold shadows two module-scope helpers Failing-first regression coverage for #4409. tests/runtime-launcher-parity.test.cjs declares extractShellBlocks twice. The module-scope copy (line 220) splits on /\r?\n/ and pushes {index, lang, lines}; the folded copy (line 1315) splits on '\n' and pushes {lines} only. The fold opens with a bare `{` at line 1226, so the inner declaration shadows the outer one for everything inside it -- and on a CRLF checkout every extracted line keeps a trailing \r. collectWorkflowFiles is duplicated the same way and is currently byte-identical, i.e. latent rather than active. Every row walks an AST via espree, a declared devDependency. None reads a .cjs and calls .includes(): that is local/no-source-grep's exact shape, and it is also the wrong instrument -- "how many declarations exist" is a construct count, and a regex would match the name inside a comment or a string literal too. Row 2 pins the rest of the corpus as an EXACT SORTED LIST of the 26 fold-shadowed helpers that remain in 15 other files, measured with this same walker rather than guessed. A bare count was rejected: `27 !== 26` names no offender and costs a round-trip to diagnose. Row 3 keeps that baseline from rotting into strings that match nothing. Row 4 is the behavioural half -- the defect a Windows user actually hits -- and asserts the surviving extractShellBlocks splits on a regex that tolerates \r, not on the bare '\n' literal. Red round: rows 1, 2 and 4 fail; row 3 passes, since every baseline entry is real today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4409): delete the fold's shadow copies of two helpers Both folded copies are removed so the fold resolves the module-scope pair. Deletion, not repair: patching the folded split to /\r?\n/ would leave two definitions and the next edit could diverge them again, and the issue asks for one definition per helper. Verified safe before deleting, not assumed: - the fold's WORKFLOWS_DIR and SNIPPET_FILE are byte-identical to the module-scope ones, and escapeRegex resolves to the same gsd-core/bin/lib/pattern.cjs, so the surviving functions close over the same values the folded copies did; - the fold has exactly ONE call site (test (E)) and it consumes blocks only as blocks.flatMap(b => b.lines). The module-scope extractor returns {index, lang, lines}, a SUPERSET of the folded {lines}, so no caller is starved. The deepStrictEqual 11 lines below that call compares `missing`, an array of path strings -- not block objects. Also removes the fold's now-dead `escapeRegex` re-require. It was used only by the folded extractShellBlocks; lint:ci caught it as an unused binding. The module-scope require stays -- the surviving extractor needs it. Applied by AST range rather than line numbers so the deletion boundaries are exact. Not fixed here, and not a deferral: the same walker finds 28 fold-shadowed helpers across 16 test files, of which 16 DIVERGE from their module-scope twin. The 15 diverged pairs outside this file are the same defect class, but the tracker already drew this boundary -- #4337 removed three runBashFile shadows "only because that PR had to edit all four identically" and deliberately left these two to #4409. A diverged shadow also cannot be deleted mechanically the way these could: its fold's tests were written against its own copy. The new suite pins all 26 remaining as an exact list so none can be forgotten. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4409): close five holes both reviewers found in the new suite Review round. Every change is to the new suite; the fix itself is untouched. 1. The corpus walk was NON-RECURSIVE, so 64 files under tests/{dispatch, qa,observability,health-diagnostic-rules,helpers} were invisible to the baseline guard. Now recursive: 1013 files parsed, up from 949. The baseline is unchanged at 26 -- coverage widened without moving the pin. 2. `catch { continue }` silently exempted any file espree could not parse, so the guard could go quietly blind on exactly the file that broke. Parse failures now propagate, and row 2 asserts the scanned count. 3. Only FunctionDeclaration was detected, so a `const helper = () => {}` shadow was invisible. Arrow and function-expression bindings now count too. Zero such shadows exist today; the hole was future-facing. 4. The old row 3 was VACUOUS -- subsumed by row 2, which already fails on a stale baseline entry because the lists stop matching. Deleted; its message folded into row 2. 5. The CRLF row inspected only declarations[0]. On the pre-fix tree that is the MODULE-SCOPE copy, which was already correct -- so only the length===1 assertion went red and the row never actually saw the bug. It now checks EVERY declaration, and pins the split to the one applied to that declaration's own content parameter rather than to whichever `.split()` appears first in the body. Also removes the `allow-test-rule: source-text-is-the-product` marker. It was inert: the readFileSync result feeds espree.parse, never a text method, so local/no-source-grep has no candidate site to suppress. Proven by running the rule through the Linter API with suppression neutralized -- 0 messages either way. A marker naming an exemption that does not exist is misleading. The divergence figure in the header comment now states its normalization. Both reviewers were right and measured different things: over the 26 remaining shadows, whitespace-only normalization gives 16 diverged / 10 identical; stripping comments as well gives 15 / 11. Exactly one pair differs only in its comments. Control re-run against the pre-fix subject file with the new logic: 3 of 3 rows fail, 3 of 3 pass after. Every row is now load-bearing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b54c1c5848 |
fix(#4709): retire the Gemini CLI reviewer lane (#4716)
* fix(#4709): retire the Gemini CLI reviewer lane
Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.
The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(
|
||
|
|
6f99e493e7 |
fix(#4395): make the debug session manager's own gsd-debugger spawn blocking (#4718)
* test(#4395): prove the manager spawns its debugger without blocking
Failing-first regression coverage for #4395.
debug.md:209 mandates the orchestrator to session-manager spawn carry
run_in_background: false, and says why outright: "Claude Code backgrounds
subagents by default, and only that flag makes the spawn return the
compact session summary directly" (#2196).
The session-manager to debugger spawn, one level down, carries no flag.
Measured: run_in_background appears nowhere under agents/ -- only in
gsd-core/workflows/. So by the rule #2196 itself states, that spawn is
backgrounded, Step 3 ("Handle Agent Return") has no return to inspect,
the manager emits CONTINUE_REQUIRED, the orchestrator auto-resumes per
#2257/#3448, and a second detached debugger races the first on
.planning/debug/<slug>.md.
Row 4 is the load-bearing one: it closes the CLASS by requiring every
subagent spawn under agents/ to declare run_in_background explicitly, so
the next agent that spawns one has to decide rather than inherit a silent
host default. It is scoped to agents/ precisely so it cannot misfire on
the workflows that deliberately use true for parallel fan-out.
Rows 5-7 are pins, not fixes: the #2196 mandate one level up, Step 2 as
the single spawn-format source that the eight continuation sites delegate
to, and the survival of CONTINUE_REQUIRED (which has a legitimate trigger
unrelated to this defect).
Red round: 4 of 7 rows fail. Row 3 needed hardening first -- asserting
only that the two variants AGREE passed vacuously, because two missing
flags are also equal; it now asserts each is present before comparing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4395): make the manager's own debugger spawn blocking
The orchestrator-to-manager hop already requires a blocking spawn and says
why (#2196, debug.md:209): Claude Code backgrounds subagents by default,
and only run_in_background=false makes the spawn return its summary. The
manager-to-debugger hop, one level down, carried no flag -- measured,
run_in_background appeared nowhere under agents/ at all.
So that spawn was backgrounded. Step 3 ("Handle Agent Return") opens
"Inspect the return output for the structured return header" -- with
nothing to inspect, the manager correctly declined to fabricate a terminal
summary and returned CONTINUE_REQUIRED; the orchestrator correctly
auto-resumed (#2257/#3448); the resumed manager reached Step 2 and spawned
a SECOND detached debugger. Both then raced on .planning/debug/<slug>.md.
Every observable in the report follows with no further assumption,
including the count: the reporter saw exactly three collisions in one
invocation, and debug.md:251 caps auto-resumes at three per slug -- one
collision per cycle.
Fixed at the cause, in both shipped variants, kept byte-consistent. The
eight continuation sites say "see Step 2 format", so they inherit it.
The issue offered two remedies. The second -- have the auto-resume path
reconcile a still-running debugger before spawning another -- is not taken:
it treats the symptom, and needs machinery that does not exist (no portable
way to enumerate or stop another runtime's live agents, plus an in-flight
sentinel with staleness and recovery rules, or an orphaned marker deadlocks
the session permanently). With the spawn blocking, the manager cannot reach
Step 4 while a debugger is live, so such a guard would also be unreachable.
#2257, #3448, the anti-loop heuristic, the cap of three, and the
CONTINUE_REQUIRED shape are all correct and untouched. CONTINUE_REQUIRED
keeps its legitimate trigger: the manager genuinely exhausting its own turn
budget mid-investigation.
Also corrects the red-round test to the canonical CALL form. debug.md
writes run_in_background=false inside Agent(...) and run_in_background:
false in prose; the first draft asserted the prose form, which the shipped
call would never have matched.
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — the blocking spawn flag plus the note recording why an unstated flag produced colliding debuggers
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — same change as its full sibling, kept byte-consistent with it
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(#4395): add changeset fragment
pr:0 placeholder is backfilled with the real number once the PR exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(#4395): refresh the variant benchmark baseline
Two entries move.
gsd-debug-session-manager.md 4766/4477 -> 4938/4649 is this change: the
blocking-spawn flag plus its explanatory note, added to BOTH variants to
keep them byte-consistent, so the compact sibling grows by the same amount
and the pair's reduction ratio dips 6.06 -> 5.85. The compact file remains
strictly smaller than its canonical sibling, which is what the variant
guard's size check actually requires.
gsd-code-fixer.md 10741 -> 10740 is NOT from this branch -- the file is
untouched here. It has scored 10740 since
|
||
|
|
ed819aa4d6 |
fix(#4379): make the TDD RED-commit pathspec language-agnostic (#4715)
* fix(#4379): make the TDD RED pathspec language-agnostic The pathspec IS this gate's definition of "a test file", and it listed only JS/TS conventions. Go's *_test.go matches none of them, so a commit adding a failing Go test was invisible, RED_COMMIT came back empty, and every behaviour-adding task halted with TDD GATE TRIPPED. references/tdd.md already advertises `go test ./...` and `cargo test` as supported, so the gate was refusing to see tests the docs promised to support. Two corrections, both measured against a seeded repo rather than reasoned: - cover the conventions tdd.md advertises: *_test.go, test_*.py, *_test.py, *_test.exs, *_spec.rb, *_test.rb. - drop the `**/` prefix. It does NOT match a path with no directory component, so a root-level foo.test.js was invisible even in the language the gate did support -- a second defect the report did not mention. A bare glob matches at every depth. Deliberately not widened to ordinary source: a pathspec matching implementation files would make the gate pass on any in-scope commit, which is worse than tripping wrongly. Rust is a known gap for that exact reason and is now documented rather than silently broken. Driving the shipped pathspec against a seeded repo: before, 0 of 7 language/root conventions matched; after, 7 of 7, with src/impl.go and src/lib.rs correctly unmatched in both. Emitted-Drift-Ack-Growth: execute-phase.md — the widened pathspec plus the comment recording why it must not cover ordinary source Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4379): drive the shipped RED pathspec against a real repo The existing row pinned the pathspec as a literal string, which the fix makes stale. Re-point it, and add behavioural coverage that EXTRACTS the pathspec from the shipped workflow and runs git log with it against a seeded repo -- re-typing the pattern into the test would only assert that two copies of a string agree. Rows: every advertised convention is visible; a root-level test file is not invisible (the half the report missed); existing JS/TS still matches; implementation files never match, so the gate can still trip; and Rust inline #[test] stays out of reach, asserted rather than left silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4379): add changeset fragment Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4379): be honest about the widened pathspec's cost Adversarial review: the rationale comment claimed the change was safe without naming what it gives up. `*.spec.*` can match a non-test file carrying the word (api.spec.json, openapi.spec.yaml), which lets the gate pass on a commit touching only that. Not new -- `**/*.spec.*` already matched those at any nested path, so dropping `**/` extends the same class to the root -- but the comment should say so rather than imply the widening is free. Also: the tdd.md list named Ruby and Elixir as recognised while the detection step above it enumerates only Node/Python/Go/Rust. Say explicitly that the gate's pathspec is wider than the detected project types, and why that is deliberate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4379): drop a vacuous row, fold its point into a real one Adversarial review: the "rust inline #[test] remains out of reach" row asserted src/lib.rs never matches -- the identical assertion to the "implementation files never match" row directly below it. It exercised nothing about #[test] semantics and would have passed against almost any fix, so it was coverage theatre. Delete it and move its rationale into the row that already carries the assertion, where it explains WHY the Rust gap follows from that row holding: the two cannot both be satisfied by a path-based gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4379): move the pathspec rationale out of a size-capped file execute-phase.md sits under a FROZEN byte ceiling (ADR-857 Phase 6, #1168: < 93600). The 20-line rationale comment I added pushed it to 93933 and tripped seven tests, all the same ceiling. Base was 92371, so the budget was 1229 bytes and the comment spent 1481. Keep six lines at the call site -- what the pathspec is, why it is not wider, where to read more -- and move the trade-off detail to references/tdd.md, which has no ceiling. That is the right home anyway: the workflow is loaded into context on every run, the reference is read on demand. 92882 bytes, 718 under. Pathspec line byte-identical; re-proved behaviour after the trim: 0/7 conventions before, 7/7 after, no implementation files matched either way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4379): refresh the compact-content benchmark baseline execute-phase.md changed size, so the committed baseline drifted. The script's own contract makes it a report that exits 0, but the test asserts the committed baseline is up to date -- refresh via --write, which is what the drift message instructs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4379): backfill the changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ec22377a9d |
fix(#4351): run-scope preserved review evidence (#4713)
* fix(#4351): run-scope preserved review evidence The preserve block copied lane output into a flat .review-diagnostics/ using each file's source basename. A lane slug is stable across runs, so the destination was stable across runs too -- and cp over an existing file is a success, so a second review of the same phase destroyed the first run's evidence with no error and no warning, in the one directory that exists to outlive the rm -rf beside it. Copy into one subdirectory per run instead. $RUN_DIR is mktemp -d, so its basename is already unique per run by construction; the UTC stamp in front is only a sort key and is omitted if date fails. Destination-only: nothing inside $RUN_DIR is renamed, because both prepare_trimmed_prompt_for_reviewer and the lane invocation resolver depend on those exact basenames. Verified by extracting the real fence and running it twice against one phase dir: before, one report survived and it was run 2's; after, both. Emitted-Drift-Ack-Growth: review.md — per-run diagnostics subdirectory plus the comment explaining why uniqueness cannot come from the clock Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4351): cover repeated runs, resolve through the run subdir Adds the two-run rows the issue asks for: both runs' reports recoverable, and a lane failing identically twice leaving two stubs. Both assert on CONTENT, not a file count -- a clobber producing the same number of files would pass a count-only check, and the defect is that run 1's bytes were replaced. runWriteReviewsFlow gains an optional phaseDir so a caller can run the flow twice against one phase directory, which is the only arrangement that can observe the overwrite. Omitted, it mints a fresh one as before. The existing rows asserted a flat readdir of the diagnostics root, which the fix makes stale. They now resolve through preservedPath/preservedNames so they keep asserting WHICH files were preserved rather than silently becoming assertions about the layout; the layout is pinned once, explicitly, by oneRunSubdirectoryPerRun_4351. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4351): add changeset fragment Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4351): name the run, not the clock, in the changeset Adversarial review: the fragment said each run gets its own "timestamped subdirectory", which reads as though the timestamp provides the separation. It does not -- collision-safety is mktemp's random basename, and the stamp is a sort key that is dropped entirely when date fails. The code comment already said so; the user-facing text now agrees. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4351): backfill the changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1110c3b4ee |
fix(#4709): stop minting the retired gemini runtime id in shipped surfaces (#4711)
* test(#4709): assert no shipped surface mints a retired runtime id Extends the #1928 removal guard to the surfaces it structurally could not reach. Its own docblock scopes it to the installer CLI contract and the runtime-name-policy exports; it spawns the installer and inspects module exports, and never reads gsd-core/workflows/**, commands/** or skills/**. Four structural assertions, all RED on next: - every RUNTIME= assignment must name a canonical runtime - the runtime->model-tier table must name only model-catalog runtimes - runtime selection menus must offer only canonical runtimes - config-set runtime / model_profile_overrides examples must be canonical Structural, not textual: each asserts the literal is canonical or the runtime exists as a catalog key, never that the string "gemini" is absent. That string is load-bearing across Antigravity's real on-disk contract, so a fifth test pins that contract from the descriptor (not from a resolved path, which would read $ANTIGRAVITY_CONFIG_DIR and the real $HOME -- the #4312 defect class). An over-broad gemini -> antigravity replacement fails there rather than ships. The menu assertion is scoped by the nearest preceding `question:` matching /runtime/i, because the same file carries a provider menu (anthropic, openai) and a budget menu (high, medium, low) whose labels are single lowercase tokens too and name neither a runtime nor anything the policy should judge. Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4709): stop minting the retired gemini runtime id in workflow text #1928 removed the gemini runtime after Google sunset Gemini CLI on 2026-06-18, but the removal stopped at the installer boundary. Runtime-loaded workflow text kept assigning the id, and the name policy's unknown-id fallbacks then applied a default designed for a never-known FUTURE runtime to an id GSD itself retired: getRuntimeLabel('gemini') is 'Claude Code', getProjectInstructionFile('gemini') is 'AGENTS.md', getGlobalConfigDir('gemini') is ~/.claude. A stale id produced a plausible wrong answer instead of an error. Those fallbacks are DELIBERATE and are left untouched here -- four docblocks document them, src/runtime-name-policy.cts:220-222 calls the label default "the always-safe default, fail-closed", and an existing test in this very suite pins getProjectInstructionFile('gemini') === 'AGENTS.md'. This commit removes the REACHABILITY of the retired id instead: - new-project.md, ingest-docs.md: the runtime-detection cascade mapped /.gemini/ and $GEMINI_CONFIG_DIR to RUNTIME=gemini. Both now map /.gemini/antigravity{,-ide,-cli}/ and $ANTIGRAVITY_CONFIG_DIR to RUNTIME=antigravity, the documented successor. ingest-docs.md was not in the original report; the new structural test found it. - settings-advanced.md: dropped the `gemini` row from the runtime->model-tier table. The model catalog has no gemini runtime (runtimeTierDefaults has 18 keys, none of them gemini), so the row advertised built-in defaults for a runtime whose config key is ignored. Its three model IDs were copied from the `google` PROVIDER preset -- a provider axis rendered as a runtime axis. - settings-advanced.md: removed the `gemini` / "Gemini CLI." runtime menu option and its group listing, so no menu offers a runtime GSD cannot install. - settings-advanced.md: repointed the config examples from `runtime gemini` to `runtime antigravity`, which ships no built-in tier defaults and is therefore the case those overrides actually exist for. - reapply-patches.md: $GEMINI_CONFIG_DIR -> $ANTIGRAVITY_CONFIG_DIR, ~/.gemini/gsd-local-patches -> ~/.gemini/antigravity/gsd-local-patches, and the local scan's bare .gemini -> .agents (Antigravity's localConfigDir). This file is hand-written, so `npm run sync:launcher` never reached it. - update.md: bare ~/.gemini and ./.gemini as GSD config dirs -> the real ~/.gemini/antigravity and ./.agents. Antigravity's own Gemini-family surfaces are untouched by design: ~/.gemini as its configHome parent, ~/.gemini/config for global skills/agents (#3738), hookEvents "gemini", GEMINI.md as its projectInstructionFile, the ~/.gemini/antigravity{,-ide,-cli} ambiguity probes (#1441), and every gemini-* model ID. The launcher's own GEMINI_CONFIG_DIR arm is left to #4632, which absorbed #4347 for it. Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4709): record the Gemini -> Antigravity migration research Primary-source research note behind #4709: the sense taxonomy that separates a runtime-axis `gemini` (stale) from Antigravity's on-disk contract, Google's model IDs, and release history (all load-bearing); the PRESERVE table; the guard-gap analysis; and the per-file inventory with file:line citations. Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4709): keep the legacy patches probe, and stop tripping two lint gates Three review findings, fixed inline. 1. Dropping the global ~/.gemini/gsd-local-patches probe was a regression: a pre-#1928 Gemini CLI install put patches there, and a stranded patches dir is still the user's work. Restored as an explicitly-labelled legacy arm probed AFTER Antigravity, so a live install always wins. This is a directory probe, not a runtime home -- it assigns no runtime id, so it does not reintroduce the defect this PR closes. The $GEMINI_CONFIG_DIR env probe is deliberately NOT restored: that names a runtime config home, which tests/declarative-reference-antigravity.test.cjs:307 pins as ignored. 2. The comment added in (1) originally contained the literal string that the new structural test matches, so the test flagged its own fix's comment as a mint. Reworded. The test was right; a comment in shipped workflow text is as readable to a matcher as code is. 3. docs/research/gemini-to-antigravity-migration.md used the colon slash-form inside a quoted manifest description. lint-docs-command-form rejects it: docs are never passed through the install-time converters, so the colon form names a command no runtime registers. Normalised to the hyphen form. Also verified, rather than assumed: the local scan's .agents entry is unambiguous. Antigravity is the ONLY runtime declaring localConfigDir '.agents' across all 19 capability manifests; grok and codex use ~/.agents as a GLOBAL home, and this scan is local (./$dir). Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4709): retire Gemini CLI from the PR templates, and close two review gaps Adversarial review findings, all fixed inline. 1. All three .github/PULL_REQUEST_TEMPLATE/*.md still offered "Gemini CLI" under "Runtimes tested", and none offered Antigravity. #1928's follow-up dropped Gemini CLI from .github/ISSUE_TEMPLATE/*.yml but missed the PR templates, so every contributor opening a fix/feature/enhancement PR has been asked for two releases which runtime they tested and offered a retired one. Now Antigravity. Guarded by a new assertion: runtime checklist labels in the PR templates must appear in the runtime label table. Proven non-vacuous by reverting one template line and watching the probe report the offender. 2. The #4709 scanning corpus excluded agents/, which also ships runtime-loaded markdown including .compact.md variants. Widened: 318 -> 382 files (+64), zero new offenders, so the gap was coverage rather than a live defect. 3. gsd-core/workflows/sync-skills.md said "grok and gemini have no dedicated installer flag — they alias the codex and claude skills roots respectively." The gemini half is wrong twice over: the runtime is retired, and it never aliased claude -- canonicalizeRuntimeName returns null for it and the caller's fail-closed default merely happens to be claude. Describing that as designed aliasing is exactly the confusion this issue is about. Reduced to grok, which genuinely does alias the codex skills root. 4. The changeset said the runtime was removed in 1.11. It shipped in 1.8.0 (CHANGELOG.md:1023 is the enclosing release heading for the #1928 entry at :1124). Corrected. Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4709): follow the corrected sync-skills prose, and refresh the compact baseline Three GREEN-run failures, all caused by this PR's own edits. 1. tests/sync-skills-cross-runtime-refuse.test.cjs pinned the literal phrase "grok and gemini have no dedicated installer flag" — a test REQUIRING shipped text to name a runtime retired in 1.8.0, which is the exact class #4709 exists to remove. The assertion and its rationale comment now track the corrected prose ("grok has no dedicated installer flag"), and the docblock's runtime list drops gemini. The remaining assertions in that file — the guard's exit, the installer pointer, the $DEST reference, guard-before-copy ordering — are untouched, so #3025's contract is otherwise intact. 2. tests/fixtures/compact-content-benchmark-baseline.json drifted because the new-project.md edits changed its compacted size (split "new-project": off 14279 -> 14308, on 12335 -> 12364; aggregate off 107411 -> 107440). Refreshed with `node scripts/benchmark-compact-content.cjs --write`, which is that script's own documented remedy. 3. emitted-attribution reported four grown workflow files with no acknowledgment. Acked below as commit trailers per ADR-3942, which moved the acknowledgment out of tests/emitted-drift-acks/*.json fragments and into the PR's own commit range (read with three-dot base...head). Exactly the four files the gate named are acked — settings-advanced.md and sync-skills.md shrank and are deliberately absent, since a trailer no delta consumed is a staleAcks error. Refs #4709 Emitted-Drift-Ack-Growth: ingest-docs.md — the runtime-detection cascade now names Antigravity's three real directories (/.gemini/antigravity{,-ide,-cli}/) and $ANTIGRAVITY_CONFIG_DIR in place of the single retired /.gemini/ arm and $GEMINI_CONFIG_DIR; three correct paths cost more bytes than the one wrong path they replace. Emitted-Drift-Ack-Growth: new-project.md — same runtime-detection correction as ingest-docs.md, plus dropping "gemini/" from the two GEMINI.md instruction-file sentences so the prose stops contradicting getProjectInstructionFile, which returns AGENTS.md for that retired id. Emitted-Drift-Ack-Growth: reapply-patches.md — restores the legacy ~/.gemini/gsd-local-patches probe as an explicitly-labelled arm after an adversarial-review finding that dropping it stranded a pre-#1928 user's patches, and repoints the env/global probes at Antigravity; the four-line comment is load-bearing, since a bare retired-runtime path with no explanation is exactly what the next reader would delete. Emitted-Drift-Ack-Growth: update.md — bare ~/.gemini and ./.gemini as GSD config dirs are replaced by the real ~/.gemini/antigravity and ./.agents, which are longer strings; no content was added beyond the corrected paths. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4709): backfill the changeset PR number pr: 0 -> 4711, now that the PR exists. Never guessed ahead of the number. Refs #4709 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f334f277dd |
fix(#4324): stop the retired /gsd: prefix reaching users (#4712)
* test(#4324): prove colon tokens the installer cannot convert leak Failing-first regression coverage for #4324. The install rewrite (transformContentToHyphen) is gated on an exact match against the commands/gsd stem list, so any /gsd:<token> whose token is not a registered stem survives the install and reaches the user as the deprecated colon form. The gate is load-bearing -- it is the only thing protecting the workflow DSL marker family (gsd:section, gsd:protected, gsd:loop-host, gsd:guard, gsd:dispatch, gsd:plan-revision-conflicts), which workflow-fragments parses as a literal. So this suite asserts the shipped text is convertible rather than asserting the transform is broad, and pins the marker family as explicit negative space. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): stop unconvertible colon tokens reaching the user The install rewrite is gated on an exact match against the commands/gsd stem list, so a /gsd:<token> whose token is not a registered stem survives the install and reaches the user as the deprecated colon form. That gate is load-bearing -- it protects the gsd:section / gsd:protected / gsd:loop-host marker family -- so the fix is in the shipped text, and the source stays colon per CONTEXT.md's two-tier rule. - quick-batch command + skill description: close the command token at a boundary so `/gsd:quick`-shaped converts instead of being skipped. - gsd-code-fixer (both variants): execute-plan and diagnose-issues are workflows, not commands, so they never converted and rendered beside two hyphenated siblings on the same line. Name them as workflows. - help topic-mode: the extraction rule hard-coded a colon prefix that the converted full.md never ships, so --brief could never match a signature line and silently fell back on every topic. Describe the signature line without a literal prefix. - update.md: drop the prefix from prose describing a stale command. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): locate the help summary per reference variant Adversarial review finding. Restoring the signature-line match (the #4324 fix) activated a latent defect in the clause next to it: compact scope emitted "the single non-blank line immediately after" the signature, and that clause is only correct for full.md. full.compact.md puts the summary on the signature line itself, after an em-dash, and its next non-blank line is an unrelated "Usage:" line. Both variants ship and both are served, so before this commit the compact variant would have emitted the wrong line as the summary. It was masked until now only because the stale colon prefix meant no signature line ever matched at all. Name the two placements and pick per line, and say explicitly that a Usage: line is never a summary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): de-vacuum the help parity check, narrow the marker waiver Two adversarial review findings against the #4324 coverage. The help-parity assertion went vacuous the moment the fix landed: once topic.md stops spelling a literal prefix, the matched set is empty and the assertion holds for any rewording, correct or not. It now also asserts across BOTH served reference variants that each ships signature lines under the hyphen prefix, that the two genuinely disagree about where the summary sits, and that topic.md still names both placements and the Usage: guard. The marker waiver keyed on "sits inside an HTML comment", which waves through a real broken reference that happens to be commented out -- `<!-- see /gsd:typo-cmd -->` scored clean. Enumerate the six marker families instead. Verified the narrowed rule catches that probe and still passes over the tree; it also surfaced a seventh family, write-continue, that the broad rule was hiding. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): normalize the namespace in skill descriptions Both hyphen-namespace skill converters ran the hyphen transform over the body but rebuilt the frontmatter description from the raw field, so a /gsd:<cmd> mention in a command description survived into the installed SKILL.md -- the exact field the host's skill picker renders, which is the surface this issue was filed about. The local flat-command path was already correct because it rewrites the whole file; only the skills path, used by a global install, was affected. Confirmed by installing into a fake HOME before and after. Fixed in both copies: bin/install.js and the src/ source of truth that compiles into gsd-core/bin/lib. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): assert descriptions through the real converters The previous version of this check called transformContentToHyphen on the description line itself and passed, while a real install still shipped the colon form -- the converter never calls that transform on the description. It asserted a proxy for the behaviour instead of the behaviour. Drive convertClaudeCommandToClaudeSkill and convertClaudeCommandToClineSkill over every registered command and assert on the emitted description. Verified it fails against the pre-fix converters and passes against the fixed ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): regenerate skills after the description change skills/<name>/SKILL.md is generated by gen-plugin-skills, not hand-maintained, and lint:generated-sync caught the hand edit. The regenerated file emits the hyphen form, which also corrects the assumption behind the scan comment in the namespace test: skills/ is runtime-emitter output, not colon source. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): re-sanction normalizeKimiSkillName's real end line The description-normalisation fix inserted five lines above normalizeKimiSkillName in src/runtime-artifact-conversion.cts, moving its closing brace from 635 to 640. MAJOR-1 pins that line deliberately, so the planted violation landed INSIDE the exempted body and went unflagged -- 0 !== 1. Re-sanction the value rather than derive it: the array is named sanctionedRealEndLines, and a pinned line that fails loudly on drift is the design. Deriving it would remove the human check the name asks for. Verified by executing all four MAJOR-1 rows against the real tree: each planted violation is flagged at realEndLine+1 and each unmodified file stays exempt. Emitted-Drift-Ack-Growth: gsd-code-fixer.md — names execute-plan and diagnose-issues as workflows rather than as slash commands that do not exist Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — same rewording as its full sibling, kept byte-consistent with it Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): backfill the changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c0b2a05d2f |
fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser The isolation guards decided whether a run-scoped sentinel applied to a dispatch by regex-scraping model-authored prose. The scrape returned values in a different namespace from the ones the sentinel records, so the comparison could never succeed: sentinel { phase: "03", plan: "03-02-hardening" } <- $PHASE_NUMBER / $plan_id prose "Execute plan 02 of phase 03-auth." scraped { phase: "03-auth.", plan: "02" } <- greedy (\S+), both wrong #4594 reports only the phase half. Measured against a real phase-plan-index run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose carries a bare in-phase plan number — so the plan field mismatches too, and the Claude path is dead rather than latent. A fresh sentinel was therefore discarded on every executor dispatch and every legitimate ISOLATION=none degrade was denied, leaving the work unrun. hooks/lib/dispatch-identity.js is now the single owner of both halves. The two prompt-body producers emit a canonical marker carrying the same shell values the sentinel records, so producer and consumer agree by construction. The prose frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and deliberately reporting no plan — an absent identifier means "cannot compare" and is safe; a wrong one is a false mismatch and is not. The prose sentence itself is byte-identical: the executor agent reads it too, so the marker is purely additive (Hyrum's Law). An inapplicable sentinel is now named in the guards' deny reason instead of being dropped silently — the silence is why this survived three producers and two consumers unnoticed. Interpolated values come from a sentinel file and from prompt text, so both are length-bounded and stripped of control characters. ADR-4630 locks the seam and maps the epic's three phases. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4594): resolve eight review findings across the dispatch-identity seam Three orthogonal review engines ran on 43418af144 — the code-review skill's Standards and Spec axes, and an isolated adversarial security pass — plus a self-review of the committed diff. Every finding is fixed here; none deferred. F1 (major, reproduced). A keyless or unknown-key-only marker — the literal "[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and returned source:'marker' with both fields null, suppressing the prose fallback entirely. Any prompt text containing that literal silently disabled identity narrowing, so a fresh sentinel applied to a dispatch it was never scoped to, defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker that yields neither recognized key is no longer a marker: the scan continues to later markers, then later texts, then prose. Forward-compatible tolerance of unknown keys is unchanged. F2/F3 (major). The first cut duplicated sanitizeForReason, describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across both guards — the exact defect class this epic exists to delete, and with no cold-load justification, since both hooks already require hooks/lib/. They now live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in isolation-sentinel.js beside the comparison it mirrors, returning the nested {sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke four-field bag that renamed the pairs already flowing through the seam. F4 (hard violation). The visibility test asserted on the deny reason's prose. CONTRIBUTING.md prohibits raw text matching on hook output, which is why every deny carries a stable reason_code. The discard is now a structured sentinel_discarded field on each hook's stdout JSON, and the test asserts that; the sentence stays for the operator but is no longer the contract. F5 (hard violation). The 64-character truncation limit had no boundary coverage. 63/64/65 are now exercised against the single consolidated helper. F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or the bidi overrides, so a crafted value could still reflow or reverse the message. Both classes are stripped, with a test each. F7 (major). The producer/template parity test was vacuous — it rendered a marker and re-parsed its own output, and would have passed with both templates deleted. It now reads the two workflow templates, extracts each marker line, substitutes the measured values and asserts the owner's parser returns them. Proven red by deleting one template's marker line before being proven green. F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the orchestrator-worktree path because that prompt is built in shell. It is not: executor-isolation-dispatch.md:131 says plainly that those are template placeholders, not shell variables, so {plan_id} is model-substituted there too. A false guarantee in a design lock is worse than a stated limit. Both documents now say the marker is model-substituted on both paths and that the prose fallback is the real floor everywhere. The "3 workflow templates" count was also wrong — 3 prose sites across 2 files, 2 of which carry the marker. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth Refs #4630. The dispatch-identity marker and its substitution note grew gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts two real-tree guards that lint:ci does not run: - tests/benchmark-compact-content.test.cjs asserts the committed baseline is "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on 23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline regenerated with scripts/benchmark-compact-content.cjs --write. - tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer for any emitted file that grows, keyed on the bare filename. Added below. The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line inside the Agent() prompt's <objective>, and the note telling the orchestrator to substitute {plan_id} with the plan's id verbatim. Both are load-bearing -- the marker is what lets a guard hook match a dispatch to the sentinel the per-plan gate wrote, and without the note the orchestrator has no instruction telling it the value must not be paraphrased. Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): set changeset fragment pr to 4693 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a155ff4fb6 |
fix(#4055): verify a phase branch is genuinely new before create-and-switch (#4694)
* test(#4055): full-lifecycle regression for merged-branch resurrection * fix(#4055): verify a phase branch is genuinely new before create-and-switch * test(#4055): assert branch absence via the gitOrThrow throw contract * test(#4055): observe the refusal disclosure via the process seam * chore(#4055): add changeset fragment * test(#4055): drop an unused fixture variable * fix(#4055): name the milestone-arm residual and the guard degradations --------- Co-authored-by: sim <sim@local> |
||
|
|
4f487e4e75 |
fix(#3929): seed install-time capability validation with the merged registry (#4691)
* test(#3929): regression tests for singleton-map install validation * fix(#3929): seed install-time cross-capability validation with the merged registry * fix(#3929): seed install-time cross-capability validation with the merged registry * fix(#3929): drop a seed overlay whose suite run throws, mirroring load * test(#3929): match the issue repro tier so the live tier-monotone check passes it * test(#3929): give the poisoned step its required onError field * test(#3929): planted overlays must satisfy the full manifest contract * chore(#3929): backfill changeset PR number (4691) * fix(#3929): honor the generator override in central keys and skip reserved-id overlays in the seed --------- Co-authored-by: sim <sim@local> |
||
|
|
0763326ced |
fix(#3780): serialize WINDOWS.md ledger mutations on a cross-process lock (#4681)
* test(#3780): regression tests for parallel ledger-writer loss * fix(#3780): serialize WINDOWS.md mutations on a cross-process ledger lock * fix(#3780): keep the ledger-unavailable degrade contract intact under the lock wrapper * chore(#3780): backfill changeset PR number (4681) --------- Co-authored-by: sim <sim@local> |
||
|
|
ccb39aec15 |
test(#4528): migrate final seam-dispatch batch and retire the timeout-literal allowlist (#4684)
Batch 17 of 17 — the terminal batch — in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/cjs-command-router-adapter.test.cjs, tests/dispatcher.test.cjs, tests/run-tests-temp-root.test.cjs, and tests/shell-command-projection-dispatch.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. No src/bin file touched, no numeric value changed anywhere. Eslint ground truth (9 sites) matches the issue's own stated count exactly for the first time in this epic — no drift to disclose. Reuses PROBE_TIMEOUT_MS (1 site) and QUICK_SPAWN_TIMEOUT_MS (1 site). Adds four file-local constants for shapes with no existing match: RUN_TESTS_ISOLATED_PROBE_TIMEOUT_MS and RUN_TESTS_HARNESS_SPAWN_TIMEOUT_MS (run-tests-temp-root.test.cjs, distinguishing a `node -e` isolated function call from a real end-to-end spawn of the test runner itself, despite each coinciding numerically with an unrelated existing constant), and EXEC_TOOL_OPTION_PASSTHROUGH_TIMEOUT_MS and DISPATCH_FORCED_TIMEOUT_MS (shell-command-projection-dispatch.test.cjs — a mocked-spawnSync pass-through fixture and a deliberately-forced real timeout, neither a real subprocess bound in the usual sense). Terminal-batch cleanup: deletes eslint-rules/no-adhoc-timeout-literal.allowlist.json entirely, drops its require and the allowlist option from eslint.config.mjs's local/no-adhoc-timeout-literal registration (now a bare 'error', mirroring local/no-unbounded-spawn's own already-terminal configuration in the same file), and updates TESTING-STANDARDS.md's enforcement note to match — a stale pointer to the deleted file caught by review, fixed inline. A full-repo eslint run with no cache confirms zero violations anywhere in the tree under the now allowlist-free rule. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a841575037 |
test(#4527): migrate planning/review-lane batch to named timeout constants (#4680)
Batch 16 of 17 in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in 9 test files with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. No src/bin file touched, no numeric value changed anywhere. Ground truth via eslint found 48 sites, not the issue's stated 38 — tests/code-review.test.cjs alone has 11, not 1 (a 10-site undercount, the largest single-file drift in this epic). All 11 are migrated. Reuses PROBE_TIMEOUT_MS (17 sites across assumption-delta.test.cjs and code-review.test.cjs), LOOP_HOOK_POINT_CLI_TIMEOUT_MS (1 site), QUICK_SPAWN_TIMEOUT_MS (1 site). Adds a new shared constant, HTTP_REACHABLE_PROBE_TIMEOUT_FIXTURE_MS, promoted because two files (reviewer-manifest-body.test.cjs, reviewer-trust-disclosure.test.cjs) independently arrived at the same probe-fixture value across 3 sites. File-local constants elsewhere for values not shared across files: FALLOW_AUDIT_TIMEOUT_MS (code-review-pipeline-regression.test.cjs); a 5-constant set covering runBashScript's own override/forced-timeout/ boundary-triple tests (plan-phase-stall-detection.test.cjs); 9 lane-specific NATIVE_TIMEOUT_MS constants, one per shipped reviewer CLI tool, even where 4 lanes coincidentally share a value (review-lane-invocation.test.cjs); and a 6-constant set covering reviewer-manifest-body.test.cjs's own probe-kind fixtures and its separate, unrelated boundary/invalid set that coincidentally overlaps in shape (not value) with plan-phase-stall-detection.test.cjs's triple. Fixes applied inline from review: GENERATOR_SCRIPT_TIMEOUT_MS was misapplied to 3 sites in plan-review-convergence.test.cjs whose actual shape (nested bash -> node -> gsd-tools.cjs chain) doesn't match that constant's documented direct-spawn class, confirmed against its own cited precedent files — replaced with a new file-local constant naming the correct class, same value. Two stray un-renamed literals in review-lane-invocation.test.cjs (missed by eslint's object-literal-only detection) were investigated and left as disclosed literals with an explanatory comment rather than force-fit onto an unrelated lane's constant, since they belong to a distinct config-override code path. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2a5d919a03 |
test(#4526): migrate gate/predicate evaluators batch to named timeout constants (#4679)
Batch 15 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/check-predicate.test.cjs, tests/gate-predicate-evaluator.test.cjs, tests/policy-160-route0-resume.test.cjs, tests/phase6-capstone-conformance.test.cjs, and tests/prohibition-enforcement.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 5 files from the rule's allowlist. Ground truth via eslint found 17 sites, not the issue's stated 16 -- prohibition-enforcement.test.cjs has 2 sites, not 1 -- disclosed in the PR body. Reuses QUICK_SPAWN_TIMEOUT_MS and LOOP_HOOK_POINT_CLI_TIMEOUT_MS across 1 file; no new shared constants needed. Adds file-local constants for a real bounded-shell subprocess class (including one deliberately non-generous value to force a timeout), a predicate's own declarative timeout FIELD as fixture/validation data (including a deliberately-invalid zero/negative pair proving rejection), a heavier gsd-tools.cjs check CLI dispatch tier, and two distinct deliberately-short enforcement bounds forcing a fast hang-timeout path. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0cee0eee47 |
test(#4525): migrate statusline/teams batch to named timeout constants (#4677)
Batch 14 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/gsd-statusline.test.cjs and tests/teams-status.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 2 files from the rule's allowlist. Issue #4525 cautions against forcing either file onto PROBE_TIMEOUT_MS without checking, since both invoke a "long-lived status renderer." Checking each file's actual production handler separately found the caution applies to one file and not the other: gsd-statusline.test.cjs's 7 sites all spawn hooks/gsd-statusline.js, which does real rendering work (context-window, git, teams state) -- two new file-local constants (STATUSLINE_HOOK_TIMEOUT_MS=4000, 6 sites; STATUSLINE_HOOK_GIT_SHIM_TIMEOUT_MS =5000, 1 site for a heavier git-shim test). teams-status.test.cjs's 5 sites spawn `gsd-tools.cjs query teams-status`, whose handler (gsd-core/bin/lib/teams-status.cjs's cmdTeamsStatus) is a lightweight env-truthiness check with no rendering or fan-out -- genuinely matching PROBE_TIMEOUT_MS's class, confirmed by reading the handler directly rather than assumed from the matching value. No new shared-helper constant is needed this batch. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
56b706a07a |
test(#4524): migrate task-content resolution batch to named timeout constants (#4675)
Batch 13 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeoutMs object-literal property in tests/task-content-resolution.test.cjs, tests/task-command-router-resolve-content.test.cjs, and tests/task-content-resolver-grammar-parity.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 3 files from the rule's allowlist. Every one of the 13 sites describes the same field -- a task-content- resolver manifest's invoke.timeoutMs -- as fixture/validation data; none is a real subprocess spawn timeout, verified by tracing each site to a pure function, a garbage-shape rejection path, or a fully-injected fake exec function. Adds one new shared constant, TASK_RESOLVER_INVOKE_TIMEOUT_MS, used by 2 files in this batch (crossing the promotion bar). Adds file-local constants for a value used by only 1 file, plus three deliberately-invalid values (zero, negative, non-integer) inside one findResolver garbage-shapes test proving the validator rejects a malformed manifest regardless of which way its timeout is invalid. Also fixes a review-caught defect outside the mechanical rule's own scope: a bare-literal duplicate of the new shared constant inside a ResolverTimeoutError assertion (a call argument, not an object-literal property, so the lint rule never flagged it) -- renamed in this same PR per the no-deferrals rule. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bbdf7e8e84 |
chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold. THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser, not grep) found what the epic never enumerated: ADR-4650 named seven containment implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13 files, several guarding a write or an `fs.rmSync`. Two verified by reading rather than pattern-matching — `research-store.cts` comments its own as "ensure the resolved file path stays inside the store dir" immediately before a write, and `capability-lifecycle.cts` gates `fs.rmSync` with one. So the epic's Done-when "one containment predicate, used at every site" was FALSE when Phase 3 reported it satisfied. It is true now: the rule is clean across src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist. WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose first argument is a managed root and whose later arguments derive from argv. That is a taint analysis over 2046 call sites, in ESLint, without type information; "derives from argv" is not locally decidable. Any approximation either floods or is trivially evaded, and a rule that fires on hundreds of correct sites earns an allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is the COMPARISON, not the join, and that has one recognizable shape. Arm 1 X.startsWith(Y + sep) the hand-rolled containment idiom Arm 2 a containment predicate called as a bare statement, answer discarded Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper was called". Its example `validatePath(x, root).resolved` is already structurally impossible — Phase 3 un-exported `validatePath` — so the remaining expressible failure is ignoring the answer, which is the defect that recurred five times in this epic. The census found exactly one live instance (`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers stop re-deriving the path the comment above it was extracted to stop them re-deriving. The rule deliberately does NOT try to catch validate-one-path-use-another where the answer is used but a different variable flows onward. That needs flow analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the two are complementary. PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a lexical site onto the realpath family broke four tests and was caught only by the matrix. Every migrated site was triaged individually. The six installer-migrations tree-walks and the six capability-lifecycle gates take the LEXICAL family because their operands are already realpath-resolved and they deliberately treat the final component as a link; boundary sites take realpath. TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken. `installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT `target === root` by contract, while the canonical comparison ACCEPTS it. Swapped naively, a migration could `rmdir` the user's config root and a third-party descriptor could write at configHome itself. Both keep `=== root` as an explicit additional arm alongside the predicate call — the predicate decides containment, the call site keeps its own extra condition (ADR-4650 decision 6). ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was byte-identical to `isContainedIn` and said so in its own docstring. `isContainedIn` is now exported for callers that have already resolved both operands and need only the comparison, with a doc note that a caller which has NOT resolved them must use a full predicate instead. THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable reason. Two justifications: (a) not a containment decision — an ancestor-walk loop condition, sub-repo grouping, worktree identity matching, declared-path coverage; (b) it IS containment but the canonical predicate is unreachable — `capability-validator.cjs` is a committed pre-build `.cjs` and the compiled `security.cjs` is untracked build output, so requiring it would break a fresh clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure. The marker was renamed from `allow-lexical-prefix-match` mid-phase because that name asserted only (a) and would have stated something false at the (b) sites. A marker suppresses BEFORE the violation counter increments, so a file whose every occurrence is marked still reports `staleAllowlistEntry` — otherwise a drained entry lingers and silently re-permits the site later. DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/` file made `npm run lint` fail with the rule's full guidance message; removing it returned the tree to clean. Both halves recorded — red alone proves nothing, since a rule red for an unrelated reason looks identical. DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two distinct rejection messages and two manual realpath calls; routing it through `tryWithinRoot` collapses them to one message, and a missing module now surfaces as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The security property is preserved and slightly strengthened — the candidate is realpathed and containment re-checked, and the dangling-symlink oracle closure comes along with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4654): record the containment ratchet in CONTEXT.md and the security model Both entries previously described the seam without the thing that keeps it a seam. They now state what the rule bans, and — more usefully for whoever reads this next — what it deliberately does NOT attempt: deciding per path.join call whether an argument came from user input. That question is not locally decidable, and an approximation across ~2000 join sites would earn an exemption list of hundreds, which is the opposite of a ratchet. Also records the marker's two legitimate justifications and that its reason is mandatory, so the escape stays reviewable rather than becoming a mute button. Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json regenerated and confirmed byte-identical rather than assumed — which also confirms eslint-rules/ is not a shipped path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): close review findings and the two matrix failures MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my evidence for collapsing it was wrong. I searched tests/ for the literal string "resolves outside its install root", found nothing, and reported that no test asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module pointing OUTSIDE the install root is not loaded" — the test guarding the exact property I claimed was preserved. defaultRequireFromInstallRoot now does both checks again with both messages byte-identical, each routed through the canonical predicate, which is better than the original since that hand-rolled both comparisons. MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments, so a suppression marker inside a plan body drifts the baseline exactly as an edit does. Measured: with markers in place, two of the four still differed from their committed checksums. The four shipped bodies are now byte-identical to next, and the rule's config excludes those four paths BY NAME rather than by a directory wildcard, so a NEW migration is still covered. Six containment comparisons stay un-ratcheted there; that gap is recorded in the rule's Known gaps, in CONTEXT.md and in the security model rather than left implicit. Justification (c) is removed from the marker's documented reasons, because a marker was proven unable to express it. ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT shape while permitting the incorrect one: startsWith(root) with no separator is the genuinely unsafe form, since it accepts a sibling such as root-evil, and my own test blessed it as valid. Flagging every bare startsWith would swamp the rule, so that stays a STATED gap rather than a silent one. Closed for real: the template-literal spelling, which the census never saw because it only inspected plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated. A separator reached through a const alias is now resolved via scope analysis. And isContainedIn, exported in Phase 3, was missing from the discarded-result set, so a bare no-op call went unflagged on the one function the epic funnels through. SECURITY REVIEW — the marker could over-suppress two ways: a block comment worked identically to a line comment, and one marker silently covered every violation sharing its line. It now requires a Line comment positioned after the flagged node ends, so it anchors to the node it trails. Four sites had dropped an unreachable-but-deliberate equality rejection against the root; each is restored as the call site's own arm. eslint.config.mjs still documented the OLD marker token, which my rename missed — it would have sent the next author in circles. A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a stale eslint cache while twelve real violations existed. Every lint check here now clears the cache first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): anchor a suppression marker to the violation it actually trails The matrix caught this; my own test caught it, on its first execution. The case "two violations on one line: trailing marker suppresses only the one it trails" expected 1 error and got 0 — both were suppressed. ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose range started at or after the node's end. A trailing marker at the END of a line sits after EVERY node on that line, so that condition held for all of them. "After the node" does not identify WHICH node the marker trails. The fix reads as correct and is not. FIX: deferred reporting. Violations accumulate during traversal instead of being reported immediately; at Program:exit each marker claims exactly ONE pending violation — the one on its line whose end is nearest before the marker begins — and every unclaimed violation is then counted and reported. One marker, one suppression. An earlier violation sharing the line is still reported, which is the property the security review asked for and the previous attempt only appeared to deliver. The counter now increments at flush time rather than during traversal, so a suppressed occurrence still does not keep an allowlist entry alive. AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test` is hard-blocked here, so this rule's test file could only ever be executed on the remote matrix — which is why a broken anchoring shipped into a run. ESLint's programmatic Linter API is not a test runner, and exercising the rule through it verifies every case locally in seconds. All 24 now pass locally, including the two-on-one-line case that failed remotely. That loop should have been built before the rule was first sent to the matrix rather than after it failed twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs now carries both, with the how-to quadrant skipped for a stated reason rather than an empty field. The audience for this deliverable is a contributor who trips the rule, and the task-oriented guidance reaches them in the ESLint message itself — which names the correct predicate, says how to choose between the realpath and lexical families, cites the Phase 3 regression caused by choosing wrong, and gives the marker syntax. A docs/how-to page would be a second, driftable copy read by nobody at the moment of failure. enablementSequence is recorded as what it actually is: a VERIFICATION sequence, not an enablement one. The rule is never off, so there is no off-to-on transition to describe. scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not evaluate against the mandated pr:0 placeholder. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6edd506cc7 |
fix(#4558): report a byte-identical restore destination as already_present (#4599)
* fix(#4558): report a byte-identical restore destination as already_present restore-custom-files treated a destination that is byte-identical to its backup exactly like a missing one: plan reported it as `eligible`, --apply re-copied the same bytes and reported `restored`, and both counters included it. Because update.md drives its restore question off eligible_count, the workflow re-offered the same no-op restore on every update and accepting it never settled anything. Emit a distinct `already_present` outcome for that case. It is excluded from eligible_count and restored_count, --apply writes nothing for it, and the backup is left intact. A differing destination is still skipped_destination_exists and a missing one still restores normally. Regression tests cover the identical-destination plan/apply paths, the idempotence-after-success cycle (missing -> restored -> silent plan), and a mixed backup. Docs for the outcome enum are updated to match. * fix(#4558): tighten already_present wording after review update.md's RESTORE_ELIGIBLE == 0 branch now also names the already-present case, CLI-TOOLS.md no longer calls the follow-up plan run "silent" (the entry is still reported, just never offered), and a test message reads correctly. * chore(#4558): add changeset fragment for #4599 * chore(#4558): acknowledge update.md growth from the restore-outcome guidance Emitted-Drift-Ack-Growth: update.md — added already_present restore-outcome guidance for #4558 --------- Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d6788a6805 |
test(#4523): migrate config/env/locking/perf batch to named timeout constants (#4673)
Batch 12 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/check-env.test.cjs, tests/config-get-default.test.cjs, tests/federated-config.test.cjs, tests/gsd-check-update-worker-platform-gate.test.cjs, tests/gsd-mcp-server-bin.test.cjs, tests/health-validation.test.cjs, tests/locking-bugs-1909-1916-1925-1927.test.cjs, tests/perf-316-state-lock-buffer-alloc.test.cjs, tests/perf-317-context-monitor-fs.test.cjs, and tests/pi-config-dir-env-override.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 10 files from the rule's allowlist. Ground truth via eslint matched the issue's stated 26 sites across 10 files exactly. Reuses PROBE_TIMEOUT_MS, GENERATOR_SCRIPT_TIMEOUT_MS, GSD_TOOLS_CLI_MODERATE_TIMEOUT_MS, and INSTALL_TIMEOUT_MS across 4 files. No new shared constants needed -- every recurring value across files was independently verified to be a genuinely different operation class, per this migration's standing rule that numeric coincidence is never identity. Adds 11 new file-local constants, three of which are not real subprocess timeouts at all (a config-merge fixture value, and two node:test per-test timeout options bounding ReDoS/lock-retry regression backstops). Per the issue's explicit mandate, gsd-check-update-worker-platform-gate.test.cjs now imports (read-only) NPM_VIEW_TIMEOUT_MS from gsd-core/bin/check-latest-version.cjs for disclosure -- this file and that production module once independently guessed the same 15000ms value, causing the PR #4428 Windows double-SIGKILL collision. No site in this file's current bare literals actually wraps a live npm-view call needing margin arithmetic, so the import documents the historical relationship honestly rather than fabricating a computation. No src/bin file edited (only a read-only import added), no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d1d9ee85a5 |
feat(#4142): thread convention through the completion-path membership seam (#3644)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
6f0e5ccf85 |
fix(#4636,#4653): close the symlink hole, revert a wrong collapse, fix six review findings
The RED checkpoint and two orthogonal reviews found eight defects. All fixed here.
THE COLLAPSE THAT WAS WRONG — installer-migrations. Routing ensureInsideConfig's
containment decision through the realpath-based canonical predicate broke four
tests, and the failure message says it plainly: "migration path escapes
configDir: extensions/gsd.cjs". That module's entire contract is that a
symlinked managed path is snapshotted, restored and backed up AS A LINK and
never dereferenced. The canonical predicate dereferences, then rejects the
result for escaping configDir — so it destroys exactly the thing the module
exists to preserve. Reverted to lexical, with the ruling recorded above the
function so it is not collapsed a third time. normalizeRelPath is the real
pre-gate there; it throws on absolute paths and '..' before this check runs.
That makes THREE deliberately-retained implementations, not two, and they share
one shape worth naming: a realpath-based predicate is the wrong tool wherever a
symlink must be PRESERVED rather than resolved. CONTEXT.md and
docs/explanation/security-model.md are corrected — both previously described
ensureInsideConfig as collapsed.
THE MISSED CONSUMER. tests/security-prompt-injection.security.test.cjs
destructures validatePath from the compiled lib; un-exporting it turned five
tests into TypeError. It appeared in my own earlier search output and I did not
follow it up. Translated under the same rule as the rest: assertions on the
rejection REASON go through assertWithinRoot, boolean-only through
tryWithinRoot.
VALIDATE-ONE-PATH-USE-ANOTHER, FOUND TWICE MORE. This is the fourth and fifth
occurrence in this epic of the exact defect it exists to prevent.
- scripts/check-glossary-refs.cjs decided containment on `token` and then
stat'd a separately re-joined path.join(ROOT, token). The ContainedPath is
now carried through to the probe, so the validated value is the probed one.
- src/init.cts computed skillPathContained and DISCARDED it, re-joining from
the raw input for the existsSync and read. The branded type exists to make
that a type error and here it was inert.
AND THE OVER-CORRECTION OF THAT FIX, caught before it shipped. The first attempt
also substituted the validated value into the EMITTED `ref` for a global skill.
That value is a display token, not a path anything reads through — the only fs
access in that branch runs on the lexical path beforehand — so substituting it
changed emitted output two ways: it is realpath-resolved, so a symlinked global
skills directory would have emitted its resolved target instead of the user's
own path, and it came from path.join, so Windows would have emitted a backslash
where the template has a literal '/'. Restored, with the distinction recorded:
the containment check there is a GATE, not a path producer.
A TEST THAT COULD NOT FAIL. The first symlink regression planted its symlink
from inside a hooked fs.readdirSync and never asserted the planting happened —
if the hook did not fire, the "nothing was written outside" assertion passed
trivially, green against vulnerable code. It now asserts the plant, matching its
sibling. The other two were re-checked: one already asserted its equivalent, the
other plants synchronously and cannot silently no-op.
THE SYMLINK FIX ITSELF, now that the tests are proven red on the matrix.
isPathConfined is lexical by design and structurally cannot see a symlink; three
callers relied on it with no defense of their own. install-engine.cts:1608 and
install-profiles.cts:880 refuse to mkdir/write through a link — mkdirSync with
recursive:true does NOT throw on an existing symlink-to-directory, so a planted
link redirected the SKILL.md write outside the install root.
install-profiles.cts:755 refuses to read through one — statSync FOLLOWS links,
so an outside file's contents were returned and installed as a skill body. Each
mirrors the guard retired-artifact-cleanup.cts:77 already uses.
Severity stated accurately rather than dramatically: only the read at :755 needs
no race. _removeGsdEntries sweeps a pre-planted link at :1608 before the write
loop, and :880's stageDir is a fresh mkdtemp, so both of those require winning a
window. They are fixed as defense-in-depth, not as live exploits.
ALSO: the Changed changeset claimed "every command's observable behavior [is]
unchanged". Three rejection messages are reworded. It now says so, and says that
none of them reveals a host path it previously hid. A stale comment in
verify.cts still named validatePath; an init.cts warning hardcoded "resolves
outside the project directory" for a check that also rejects absolute paths, NUL
bytes and empty strings; and the rationale deleted with check-glossary-refs'
retired helper is restored, noting honestly that a rejected token is now
realpath-resolved before rejection rather than rejected by string comparison.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7086dcbbc9 |
test(#4636): failing-first coverage for symlink escape past lexical confinement
Tests only; no fix. These MUST fail. Found while draining the containment duplicates in Phase 3, not by looking for it: isPathConfined's docstring claims its callers are kept symlink-safe upstream, and checking that claim showed it does not hold for three of its four call sites. isPathConfined is lexical by design and structurally cannot see a symlink, so where the claim fails there is no defense at all. install-engine.cts:1608 mkdirSync(recursive) + writeFileSync under dest install-profiles.cts:880 the same shape under stageDir install-profiles.cts:755 statSync/readFileSync, and statSync FOLLOWS links SEVERITY, STATED HONESTLY, because the three are not equal and the first impression was wrong. :755 is the real one and needs no race. It reads <capDir>/skills/<stem>/SKILL.md. Nothing prunes that path and nothing randomizes it, so a symlink planted there is followed and its content is returned — and then written into the install tree as a skill body. :1608 and :880 are defense-in-depth against a local race, not plain write-through. _removeGsdEntries (install-engine.cts:851-854) rmSyncs every entry whose name starts with kind.prefix, and the skill names ARE prefix+stem — so a symlink planted before the call is unlinked before the write loop reaches it. :880's stageDir is a fresh mkdtempSync path, so its name cannot be guessed in advance either. Both require winning a window. That is also why the first two tests do not simply pre-plant a link and call the function: written that way they PASS today, against vulnerable code, for the wrong reason. They instead open the window deterministically by hooking a call the function is guaranteed to make — the technique CLAUDE.md section 4 prescribes for injecting filesystem conditions, and the one already used in tests/install-runtime-artifacts.test.cjs. No sleeps, no concurrency, no flakiness: the window is opened by a synchronous side effect, not raced for. Each test asserts the OUTCOME rather than the mechanism — the canary file outside the root is byte-identical afterwards, or the installed content does not contain it — so a fix is free to close the hole any way it likes. Symlink creation is guarded so the Windows lanes skip rather than fail, and no test uses chmod 0o000, which is vacuous on a bench running as root. Folded into tests/install-write-confinement.test.cjs, the owning suite, whose existing F2/M1 coverage coincidentally proves the gap: it tests the case where the destination DIRECTORY is a symlink, which is already defended, and never the per-entry case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7e7a239a65 |
refactor(#4653): replace the allowAbsolute flag with a named acceptance policy
Phase 3 of epic #4636, stage 3b. Satisfies #4653's criterion that `opts.allowAbsolute` become "a named acceptance policy on the predicate, not a per-call-site boolean". The flag was actively misleading at the call site. `{ allowAbsolute: true }` reads as "containment is relaxed here". It never was: an absolute path that resolves outside the root is rejected exactly as a traversal is. The flag only ever controlled whether an absolute candidate was CONSIDERED. On a security predicate that is the wrong thing for a reviewer to have to infer, and 31 call sites were asking them to infer it. PathAcceptance.RelativeOnly relative candidates only PathAcceptance.AbsoluteInsideRoot absolute accepted, containment unchanged The three exported wrappers take the policy and translate it inward. validatePath keeps its internal `{ allowAbsolute }` opts and its body untouched — the engine is not re-derived here either, only the exported surface is renamed. MEASURED, NOT ESTIMATED. 31 call sites across 10 files, counted by walking the AST with the repo's own @typescript-eslint/parser rather than grepping: a text match would have folded in the options-type declaration, default parameter values and comments. All 31 pass the literal `true`; none passes `false` or a dynamic value, so the migration is uniform and `RelativeOnly` is purely the existing default made nameable. audit.cts alone holds 18 of them. This migration is compiler-verified in a way the containment-value migration in the previous commit was not: the parameter type changed from an object to a string union, so any missed site is a build error rather than a silent behavioral difference. That is why a 31-site mechanical edit is acceptable in the phase whose stated risk is the width of mechanical change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
26384ca988 |
refactor(#4653): make validatePath module-internal
Phase 3 of epic #4636, stage 3a. ADR-4650 decision 2: the engine stops being a public shape. The only exported containment surface is now assertWithinRoot / tryWithinRoot / requireSafePath, none of which can hand a caller a usable path when the answer is unsafe. The src/security.cts diff is one keyword. The engine body is byte-identical — the dangling-symlink existence-oracle closure, the ancestor canonicalization and the separator-aware boundary test are untouched, which is the whole constraint this phase operates under. WHAT THE TRANSLATION COST, AND THE RULE THAT KEPT IT AT ZERO. Roughly thirty test call sites consumed validatePath directly, including the two BLOCKER regressions that are this refactor's safety net. Translating them all to `tryWithinRoot(...) === null` would have looked correct and silently destroyed one of them: BLOCKER-1 asserts the rejection reason contains "unresolvable symbolic link", which is what distinguishes a DANGLING symlink from an ordinary escape. tryWithinRoot returns a bare null and cannot tell those apart, so that assertion would have degenerated into "it failed somehow" — and the existence-oracle closure could regress with the test still green. So the rule applied throughout is: an assertion on the rejection REASON goes through assertWithinRoot, whose throw carries the engine's message verbatim; only assertions on the boolean go through tryWithinRoot. Under that rule no coverage is lost. BLOCKER-1 still pins "unresolvable symbolic link" and BLOCKER-2 still pins the exact canonicalized resolved value. Three success-path tests came out BETTER than they went in. They previously carried `expected safe:true, got error: ${result.error}` as an assertion message; routing them through assertWithinRoot means an engine regression now surfaces the real reason in the failure itself rather than as a hand-built string. The two describe blocks named after validatePath are renamed — a block named for a symbol the module no longer exports is a false signpost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
889c7efba0 |
test(#4653): failing-first coverage for the narrowed containment export
Phase 3 of epic #4636. Tests only; no implementation. These MUST fail. A refactor changes what a good test looks like: the behavior under test must be IDENTICAL before and after, so most of this phase's safety comes from invariance rather than new assertions. That safety net already exists and is untouched here — tests/security.test.cjs already pins the two engine behaviors a re-derivation would silently lose: :177 a DANGLING symlink to a non-existent OUTSIDE target stays safe:false (the existence-oracle closure) :213 a not-yet-created file in a not-yet-created subdir under a non-canonical base stays safe:true (ancestor canonicalization) plus traversal, absolute in/out, null bytes, empty, non-string, and requireSafePath's throw. Those 0 deletions are the point: if any of them had to change, the engine would have changed, and the engine is not supposed to. What is new is the export surface Phase 3 introduces: assertWithinRoot(candidate, root, label?, opts?) -> ContainedPath (throws) tryWithinRoot(candidate, root, opts?) -> ContainedPath | null Two shapes rather than one, because several call sites need a NON-throwing check — findPhaseArtifact probes a direct path, then a .planning/ path, then each readdir entry, and throwing on the first miss would break it outright. ADR-4650 names only the throwing form; this is the gap between the ADR and the call sites, recorded rather than papered over. Rows that exist because they are the ones nobody enumerates: - tryWithinRoot must return EXACTLY null on escape, and its return must not contain the escaping path's basename. The shape being replaced populates its "resolved" field with the escaping path precisely on the traversal branch, so a caller who ignores the boolean gets a usable attacker-controlled value. That is the defect the narrowing exists to remove, so it is asserted directly. - A seeded parity property: tryWithinRoot returns non-null if and only if assertWithinRoot does not throw, and the values agree. Two exported shapes over one engine is a divergence pair by construction. - The rejection text still contains the phrase "escapes allowed directory". Another suite surfaces it through a user-facing "reason" field, and a refactor is exactly where wording drifts unnoticed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f64f8a0e7b |
test(#4652): correct the absolute-filename test to match the basename guard
The test asserted that an absolute filename is folded under the pending dir and fails as "not found" rather than being rejected. That held for exactly one commit. The basename guard rejects any name containing a separator before any join happens, so an absolute path never reaches containment or the filesystem at all. Now asserts the USAGE rejection the CLI actually emits, verified by running it. All four outside-file protections are kept unchanged — the file still exists, its content is byte-identical, it never lands in completed/, and cleanup runs in finally. Those are the assertions that carry the security value; only the claim about HOW the rejection happens was stale. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9c20b7b40a |
fix(#4652): use the validated path, collapse the duplication, correct two false claims
Seven findings from the two-axis review, all fixed in place. THE ONE THAT MATTERS: cmdTodoComplete validated sourcePath and targetPath and then ran every fs call against the RAW strings — existsSync, statSync, readFileSync, platformWriteSync, unlinkSync, and the dry-run path payload — never sourceCheck.resolved / targetCheck.resolved. That is the exact "validate one path, use another" shape ADR-4650 names as the defect this epic exists to prevent, and it is the same bug this phase had just fixed in check-command-router. Committed inside the fix for it. All I/O now uses the resolved paths; user-facing messages still echo the raw filename, never a resolved absolute path. A VACUOUS TEST, and the false doc claim it was propping up. The test "[RED #4327] an absolute path outside the project is rejected" would have passed with ZERO containment logic: path.join(pendingDir, '/abs/outside/x') yields <pendingDir>/abs/outside/x — Node does not let a later absolute segment escape — so the name is FOLDED under the root, passes containment, and simply 404s. The test only ever observed "Todo not found". It now asserts what is actually true and actually valuable: an absolute name is neutralized, and the real outside file is not read, not moved, and still present afterward. docs/CLI-TOOLS.md claimed such a path "is rejected as a usage error", which was false; it now describes the fold-under-root behavior. Traversal and embedded separators ARE rejected, and those claims stand. DUPLICATION THIS EPIC EXISTS TO REMOVE. resolvePath already did isAbsolute-or-join + validatePath + reject; cmdGapAnalysisPlanPost and cmdCheckPredicate each re-inlined the identical triplet in the same file. Both now call resolvePath. Cost, stated rather than hidden: its generic message replaces the two sites' distinct "phase-dir escapes…" wording. The message still names the offending input, and one predicate with one message is the point. SYMLINK COVERAGE was required by #4652's "Done when" and was missing. Added for both the todos root and --phase-dir, skipping cleanly on EPERM so the Windows lanes do not fail where unprivileged symlink creation is disallowed. Both fast-check properties were UNSEEDED. Seeded now. The changeset named "check decision-coverage-plan" as a boundary; that is a caller of the shared resolvePath, which the body never mentioned. Corrected. DISCLOSED, not hidden: ctx.phaseDir is now always the resolved ABSOLUTE path, so ${PHASE_DIR} interpolation and the "not found in <targetDir>" message show an absolute value where a relative --phase-dir previously produced a relative one. That is an observable output change. A test pins it and docs/reference/gate-predicates.md states it. Also regenerated scripts/lib/platform-conformance-tier.generated.cjs and its macos twin — the new tests changed check-predicate.test.cjs's tier classification. Caught by npm run lint:ci locally rather than by a bench run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
374300da17 |
fix(#4652): confine every boundary that joins argv to a managed root
Phase 2 of epic #4636, absorbing #4327 and #4354. Implements ADR-4650 decision 3: containment is a boundary concern — the predicate runs where external input enters, not at whichever interior call site remembered. Four boundaries now validate against their managed root and reject with a USAGE-shaped error before touching the filesystem: todo complete <name> -> todosDir(cwd) check predicate --phase-dir <dir> -> projectDir check decision-coverage-plan <dir> -> projectDir (via resolvePath) check gap-analysis.plan-post <dir> -> projectDir #4327 understated its own severity. It reports that a traversal name "resolves outside the todos root", which reads as an information leak. Measured, it was destructive: the command exited 0, MOVED the outside file into completed/, and unlinked the original. cmdTodoComplete ends in fs.unlinkSync(sourcePath), so an unconfined name consumed across the boundary rather than merely reading across it. Validation now precedes every fs call — existsSync, readFileSync, ensureDir, writeSync, unlinkSync — and both halves of the move are confined, so neither source nor destination can land outside the root. --dry-run is rejected on the same terms; a preview must not leak a resolved outside path either. #4354 reproduces exactly: a BLOCKING gate returned block:false sourced entirely from a SECURITY.md in a caller-chosen directory outside the project. THE HARDER HALF, found by the isolated adversarial review of the first attempt: validating a path and then using a DIFFERENT one closes nothing. The first fix validated `--phase-dir` joined against `--cwd`, then passed the RAW unjoined value into the predicate context. gate-predicate-evaluator uses it as-is and findPhaseArtifact resolves a relative path against the REAL process cwd — so validation and the read used two different roots whenever process.cwd() differed from --cwd. Reproduced: running from a directory holding a plan with `secret_field: LEAKED_VALUE`, a predicate declared against an empty --cwd project exited 0 and returned "actual":"LEAKED_VALUE". The rule now applied at all three router sites: **use the validated resolved path, never the raw input.** Independently re-verified after the fix — the lookup resolves in the --cwd project and no value leaks. gate-predicate-evaluator.cts is untouched and still imports no fs. Confining in the router is what keeps that pure-leaf contract intact AND covers ${PHASE_DIR} interpolation into command-exit-zero, which an evaluator-local fix would have missed entirely. Also fixed, same review: `todo complete .` and `..` passed containment (they resolve to the pending dir, which IS inside the root) and then threw an uncaught EISDIR with an absolute-path stack trace. Now a clean USAGE rejection naming the real reason — "todo name is not a file" — rather than borrowing the escape message, which would have stated something false. Ripples discharged BEFORE the verification checkpoint rather than after, per the Phase 1 retrospective: docs/reference/gate-predicates.md and docs/CLI-TOOLS.md document the new constraints, CONTEXT.md records why containment lives at the router rather than the evaluator, the changeset is written, and the install-tree goldens were regenerated to confirm unchanged (no new shipped file) rather than assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3925839f2a |
test(#4652): failing-first coverage for the four unconfined boundaries
Phase 2 of epic #4636, absorbing #4327 and #4354. Tests only; no fix. These MUST fail. Four CLI boundaries join externally-supplied input to a managed root with no containment validation. Each was driven through the real CLI and confirmed unconfined before the assertions were written: todo complete <name> src/commands.cts cmdTodoComplete check predicate --phase-dir <dir> check-command-router cmdCheckPredicate check decision-coverage-plan <dir> check-command-router resolvePath check gap-analysis.plan-post <dir> check-command-router Boundary 1 is worse than the issue describes. #4327 reports that a traversal name "resolves outside the todos root", which reads as an information leak. Measured, it is destructive: `todo complete ../../../../b1out/leak.md` exited 0, MOVED the outside file into completed/, and unlinked the original. The file was gone. cmdTodoComplete ends in fs.unlinkSync(sourcePath), so an unconfined name does not merely read across the boundary, it consumes across it. Boundary 2 reproduces #4354 exactly: a BLOCKING gate returned {"block":false,"details":{"match":true}} sourced entirely from a SECURITY.md in a directory the caller chose, outside the project. Boundaries 3 and 4 are not named in the epic. Both accepted an outside phase dir and exited 0. Rows that exist because they are the ones nobody enumerates: - ORDERING. A real file is created outside the todos root, then the traversal name targeting it is asserted rejected AND the outside file asserted still present and unmoved. #4327 notes the existence check and the move target BOTH follow the unvalidated join, so a rejection that lands after the read has already leaked — and, per the finding above, after the unlink has already destroyed. - `a/../../b.md` — looks balanced, resolves outside. - --dry-run must reject too; a preview must not leak a resolved outside path. - ${PHASE_DIR} interpolation into a command-exit-zero predicate is the SECOND predicate kind, which a fix inside gate-predicate-evaluator.cts would miss. - An absolute path INSIDE the project must still be accepted at every boundary — absolute is not a synonym for escaping. Cross-boundary rows loop over one shared list of escaping inputs and assert all four reject with the same shape, so four sites adopting one predicate cannot drift into four rejection contracts. Property tests cover BOTH directions — outside is always rejected, inside is always accepted. A property asserting only rejection is satisfied by a predicate that rejects everything, which is the degenerate-implementation trap found in Phase 1's review. Both are seeded. Regressions fold into the owning module suites rather than a new tests/fix-NNNN-*.test.cjs, per scripts/lint-regression-test-names.cjs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c99d7bb2be |
test(#4522): migrate core CLI/domain state batch to named timeout constants (#4662)
Batch 11 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/state-document.test.cjs, tests/phase.test.cjs, tests/commands.test.cjs, tests/pattern.test.cjs, tests/adr-612-bracket-coherence.test.cjs, tests/adr-612-bracket-read-tolerance.test.cjs, tests/milestone-lock.test.cjs, tests/init.test.cjs, tests/state-todos-render.test.cjs, tests/quick-batch.test.cjs, tests/graphify.test.cjs, and tests/effort-surface-axis.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 12 files from the rule's allowlist. Ground truth via eslint found 25 sites, not the issue's stated 24 (phase.test.cjs has 5, not 4) -- disclosed in the PR body. Reuses PROBE_TIMEOUT_MS, GIT_TIMEOUT_MS, and LOOP_HOOK_POINT_CLI_TIMEOUT_MS across 8 files. Adds two new shared constants to tests/helpers/timeouts.cjs (each independently arrived at by 2 files in this batch, crossing the promotion bar): PATHOLOGICAL_INPUT_TEST_TIMEOUT_MS (node:test's own per-test timeout option, not a subprocess bound) and GSD_TOOLS_CLI_MODERATE_TIMEOUT_MS (a single gsd-tools.cjs CLI subcommand spawn, distinct tier from PROBE_TIMEOUT_MS/LOOP_HOOK_POINT_CLI_TIMEOUT_MS). Adds 3 file-local constants for values used by only 1 file in this batch. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
241646a43a |
fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail. The guard classifies a name by comparing everything after `.env.` as one token against a set whose members are FINAL EXTENSIONS. So `.env.local.example` yields suffix `local.example`, which is not a member, and a committed secret-free template is refused. That is a category error, not strictness. Two arms are covered because the same classification is hand-rolled twice in one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret` (`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a guard that allows `cat .env.local.example` while refusing `Grep --glob '.env.local.example'` — the same file, the same hook, opposite answers. A cross-arm parity loop over one shared list asserts the two cannot drift. Rows that exist because they are the ones nobody enumerates: - `.env.example.local` must stay BLOCKED. Final extension is `local`; this is dotenv's documented local-override convention and a real secret. Any fix shaped as "contains example" admits it. - `.env.local.` must stay BLOCKED — empty final extension is not a member. - `.env.` must stay ALLOWED. Note #4580's proposed patch adds `if (suffix === '') return true;`, which flips it to blocked; that breaks the existing `allows` assertion in this suite and broadens the protected set, which epic #4636's non-goals forbid. Not applied. - `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select `.env.local`, and a partial literal cannot be classified. - `*.example` and `*` must stay ALLOWED — regression protection on the arm that already works. Local behavioral repro of the current guard, confirming the tests fail for the right reason rather than by construction: .env.local.example rc=2 (blocked) <- the defect .env.example rc=0 (allowed) .env.local rc=2 (blocked) .env.example.local rc=2 (blocked) .env. rc=0 (allowed) glob .env.local.example rc=2 <- the second arm Regressions are folded into the owning module's suite rather than a new tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4651): classify by final extension so .env.<name>.example is readable Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5. The guard compared everything after `.env.` as ONE token against a set whose members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`, which is not a member, so a committed, secret-free template was refused — the guard blocked the one file that exists so nobody has to open the real `.env`. That is a category error, not strictness. The fix is not "add local.example to the set"; it is to compare the right token. hooks/lib/filename-classification.js now owns that distinction and is the only place it is expressed. Both arms are fixed, because the same classification was hand-rolled twice in this one file: - isSecretBasename (Read/Bash) now tests finalExtension(suffix). - globAltSelectsSecret (Grep --glob) split its first branch. With no wildcard the alternative IS a whole filename, so it is classified exactly via isSecretBasename. With a wildcard present the literal is only a PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and cannot be classified, so the original conservative rule stays. Fixing only the first would have shipped a self-contradicting guard: `cat .env.local.example` allowed while `Grep --glob '.env.local.example'` refused — same file, same hook, opposite answers. A cross-arm parity loop over one shared list now asserts the two cannot drift. Two deliberate departures from #4580's suggested patch, both verified: - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.` from allowed to blocked, breaking an existing assertion in this suite and broadening the protected set, which epic #4636's non-goals forbid. - `fullSuffix` was drafted alongside finalExtension and removed before commit: zero production consumers, and none planned (Phases 2-4 are containment, duplicate draining and the path-join ratchet, none of which classify filenames). A zero-caller export is dead code. The distinction is pinned instead by a test asserting finalExtension('local.example') is 'example' and explicitly NOT 'local.example'. The protected set is unchanged. `.env.example.local` stays BLOCKED — its final extension is `local`, dotenv's local-override convention and a real secret; any fix shaped as "contains example" admits it. Scoped out by measurement, not assumption: src/validate.cts:395 and src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding them in would repeat this same category error in the opposite direction. Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed, 26 failures / 45330, all 26 in the two new test files, zero pre-existing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4651): document the widened template exemption and cover the Bash arm Two findings from the isolated adversarial review, both fixed in place. 1. The header's "Stated cost" passage named only the four literal template names, but since this change the exemption keys on the FINAL EXTENSION, so the trusted set is `.env.<anything>.{example,sample,template,dist}` — an unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example` is allowed. That is the deliberate and necessary cost of fixing #4580, but it was materially larger than what the header disclosed, and a silent expansion of a security guard's trusted set is not acceptable. The passage now states the family, the concrete bypass, and that it applies across Read, Grep and Bash alike. 2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely distinct. The Bash arm was covered only by two one-off tests outside the shared table, so the table could not have caught a drift there. The loop now drives all three arms from the same TEMPLATES/SECRETS arrays. No classification logic changed. The Read-arm behavioral table is byte-identical before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for .env.local / .env.example.local / .env / .secrets. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4651): treat trailing dots and spaces as aliases of the protected file Closes a Windows path-alias bypass surfaced by the isolated adversarial review of this phase. Maintainer-approved as in scope. Win32 strips trailing dots and spaces from every path component, so `.env.`, `.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve to the real `.env` / `.secrets` on Windows. The guard allowed every one of them — a bypass of a file it already protects, reachable from Read, Grep and Bash alike. `isSecretBasename` now normalizes the basename before classifying. The whole class is fixed, not the reported name. `.env.` alone would have left `.secrets.` and the trailing-space forms open, which is the same one-cause-explains-every-failure trap this epic exists to close. Two consequences, both measured rather than assumed: - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted `.env.example` template, so this is correct; it was previously blocked only because the trailing dot broke final-extension parsing. - A Bash token that is exactly `.env` plus trailing whitespace flips allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive class: the bare `.env` token was ALREADY blocked as an operand in the same position before this change, so the alias now simply behaves like the thing it aliases. The header's "No whitespace trimming" guarantee is preserved and now stated precisely: leading and interior whitespace is still never trimmed, so prose like a commit message mentioning `.env` in a sentence stays prose and stays allowed. Only TRAILING dots and spaces are stripped. Two tests pin that. This lands at the same behavior #4580's proposed `if (suffix === '') return true;` would have produced for `.env.`, which this phase earlier rejected. The rejection was correct on its stated grounds — that line broadens the protected set, which epic #4636's non-goals forbid. The Windows framing is different: normalizing an alias of an already-protected file is not a broadening, and the fix is reached by normalization rather than by special-casing an empty suffix, so it generalizes to `.secrets.` and the space forms. Cannot be reproduced on this host — the remote matrix is Linux-only and Windows coverage arrives from CI — so this ships on the Win32 path-normalization contract plus the CI lane, and that limitation is stated rather than implied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4651): one owner for path segmentation, closing a Read/Grep divergence Four findings from the two-axis review, all fixed in place. The real one: the guard had TWO path-segmentation rules. `lastSegment` (used by Read and Bash via `namesSecret`) splits on both `/` and `\`, while `classifyGrepGlob` hand-rolled its own on `/` only. Measured: Read of `config\.env` rc=2 BLOCKED Grep --glob 'config\.env' rc=0 ALLOWED Same logical file, opposite answers — precisely the divergence this epic exists to remove, sitting inside the file this phase was already fixing. `lastSegment` now lives in hooks/lib/filename-classification.js and both arms call it. All five path-bearing cases (both separators) now agree. Note on how this was nearly missed: the first measurement of it reported "both allow", which looked like the reviewer was wrong. That reading was a measurement artifact — `config\.env` inside a printf'd JSON payload is an invalid escape, so the hook fails open at rc=0 and the test was observing JSON breakage rather than the predicate. Re-measured with correct escaping, the divergence is real. The tests added here use properly escaped literals and were verified by running, not by reasoning about the escaping. Also fixed: - Both fast-check properties were satisfied by a degenerate always-return-'' implementation: "never contains a dot / is a suffix" and "never ends with dot-or-space / is a prefix" are both trivially true of the empty string. They now additionally pin content preservation — the removed tail must match /^[. ]*$/, and a name with nothing to strip must come back unchanged. - The cross-arm parity loop used only bare basenames, so it could not have caught the divergence above. It now covers path-bearing names with both separators. - That loop's description overclaimed: Read and Bash BOTH route through `namesSecret`, so they are not independent paths; only the Grep glob arm is genuinely separate. The description now says so rather than implying three-way independence. - `normalizeWindowsBasename` runs on every platform, not only Windows. Its doc now states that explicitly: the guard must answer identically everywhere, and a name is judged by what Win32 would resolve it to. No classification logic changed; the 12-name regression sweep is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs Three things, all consequences of the fix rather than new behavior. 1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED file — package.json `files` includes `hooks` — so every per-runtime install tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11 failures one-for-one. This ripple was identified at design time and then not acted on. Fleet's impact preview named golden-install-tree.test.cjs before any code was written, and 40-design.md records it under "Ripples identified". Writing a risk down is not the same as discharging it, and a full matrix run was spent discovering something already known. 2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's protected set: it named `.env.example` / `.sample` / `.template` / `.dist` as the four exempt names. The exemption keys on the FINAL EXTENSION, so the exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`. The page now states that family, the widened residual, that order matters and only the last segment counts (`.env.example.local` is a secret), and that trailing dots and spaces are stripped because Windows resolves them to the protected file. A wrong user-facing model of what a security guard protects is worth correcting even though Fixed/Security changesets are exempt from the required-docs rule. docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as `.env.example` exempt" — non-exhaustive, still true, deliberately left alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same hedged phrasing; hand-translating a security description unreviewed is not something to do silently. 3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass closed is `Security`. Folding the second into the first would under-report it in the release notes. Both carry `pr: 0` for backfill once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4651): backfill changeset PR number to 4659 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a2331c01f1 |
fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites Manually traced against the current tree: the validating regex at code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical behavior for 1- and 2-segment ids, character class unchanged): code-review.md, code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating sites, plus their comment/error-message text), execute-plan.md's plan-filename extraction, and plan-phase.md's --research-phase flag capture. Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor, which was matching an unrelated earlier `--research-phase` occurrence (line 77's generic-value capture) instead of the targeted site (line 131). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/ Adds findSingleSegmentPhaseRegexDrift, banning the bounded `[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any phase-carrying line across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned the same way as the existing shell-arith rule. Wired into scanAll; confirmed zero violations against the real tree post-#4568 fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4568): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file The emitted-attribution gate also flags 4 files growing: code-review-fix.md (+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9 bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each site's validation regex widened from a bounded single-optional-dotted-segment shape to the unbounded form, and the accompanying comment/error-message text grew by a few characters to mention the new 3-segment example. Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568) Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4568): backfill changeset pr number to 4646 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4d65c248e5 |
fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector Tests only, committed ahead of the implementation so the RED run is real. - tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative cases proving seam calls and path-call-plus-slash-literal are not platform signals; positive pins that genuine platform content, seam-bypassing spawns, chmod and symlink still classify in; macOS signal set and generated list unchanged. - tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest rows and test-conformance still has 3 windows + 1 macOS. - tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint test does, and resolveSelection rejects the retired windows scope. Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): delete the second Windows selector and narrow the conformance tier Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped conformance tier on real Windows/macOS — was not met. Measured on PR #4640 (run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7 changed test files running on a real Windows runner twice. Two selectors, only one in the epic's scope. The test job's three scope:windows shards predate the epic (#494, sharded #3057) and gate on product_changed, not full_matrix, so they fire on every product PR whatever Phase 3's classifier decides. They are deleted; test-conformance becomes the sole Windows selector, as it already was for macOS. Non-Linux jobs 7 -> 4. Gating the lane instead was rejected as provably redundant: for a test file reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file), and that same predicate sets full_matrix, which turns test-conformance on. Every file a gated lane would run is already covered in the same run. The lane's one non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix instead of feeding a parallel lane. Two detectors matched the repo's own test idiom rather than any platform signal and carried 226 of the tier's sole-signal membership against 41 for the other eight: process-seam-subprocess (335 files, 118 unique) matches the tests/helpers.cjs entry points nearly every CLI test uses, and going through the seam is the opposite of a platform signal since shell-command-projection takes platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs only a path call anywhere plus a slash literal anywhere, and that class is already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed. Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured. Adds the size gate Phase 2 never had, as a ratio against a live denominator so it cannot stop binding as the suite grows. 292 files leave real-OS Windows execution. The drop-out set was audited: 14 have a platform-suggestive filename and all 14 are static source-text analyses or seam-mediated CLI tests. raw-child-process was investigated as a suspected false negative and left unchanged -- relaxing it adds 13 files, all false positives. macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated macos-conformance-tier.generated.cjs is byte-identical at 196 files. Fixes #4641 Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): register the new ADR path in the docs-guard exempt baseline tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md in a comment justifying the retired windows scope; lint-docs-guard-registration tracks that reference set, so the baseline needs the new path. Verified the exemption still holds: the path is prose, not a filesystem read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): make the escalation tier-backed and drop every hardcoded count Three follow-ups from measuring the first pass rather than trusting it. The windows-hint escalation now requires tier membership as well as the filename hint. Setting full_matrix runs test-conformance, which runs only the tier; escalating on a test that is NOT in the tier costs four jobs and still never runs that test on Windows. Measured over the 16 RULES entries the narrowed predicate fires on exactly the same rules today, so this is correct-by-construction rather than a behavior change. The broader variant -- escalate on any tier member a rule pulls in, ignoring the hint -- was measured at 14/16 rules and rejected as over-broad. Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as soon as the rebase pulled in one new test file from #4253, which is the whole argument against them: the ceilings are ratios against a live denominator, the committed lists are pinned by comparison against a fresh classification of the live tree, and the three named probe files now assert on their SIGNAL rather than on membership in a literal list -- asserting by filename is the exact error this PR fixes in the classifier. Regenerates both lists against the rebased tree. Same-tree figures are now 547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none added; macOS is unchanged at 197 with a zero-line diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke An isolated adversarial review found a real false negative. Removing the blanket process-seam-subprocess detector also removed the only coverage for tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook spawns options.interpreter via real spawnSync, so runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary executing a shell script extracted from workflow markdown. The seam argument holds for src/shell-command-projection.cts, which takes platform as an injected parameter; it does NOT hold for the test helpers, which spawn real binaries. Conflating the two is what made the blanket detector look purely noisy -- it was 99% noise wrapping a real signal. Adds a narrow shell-interpreter-spawn category keyed on a real interpreter option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33% ceiling. All 9 confirmed by reading the matching source line, zero comment or fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured and rejected -- each adds 9 files but misses the counterexample entirely. Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD strictly rather than carrying the Windows report-only carve-out. The carve-out now lives solely on test-conformance, whose matrix does include windows. Repairs seven pre-existing assertions the category removal invalidated, preserving each case's purpose rather than deleting coverage, and converts the last hardcoded tier bounds to live-derived ratios -- including the macOS sanity range that was still a magic [100, 350]. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): keep the confinement test on a real OS via a documented allowlist A security review found tests/external-descriptor-confinement.test.cjs had dropped out of the Windows tier. It must stay in, and no content signal can express why: it exercises isPathConfined (src/external-descriptor-trust.cts), which uses the AMBIENT path module -- path.resolve(root, target) and path.sep -- with no injection. Its win32 semantics (drive letters, UNC, separator) are only reachable by actually running on Windows, and it is a security-relevant write-confinement gate. A content classifier cannot see 'this module reads the ambient path module', so no regex belongs here. Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows tier only. A Map rather than a list so an entry without a reason is impossible by construction, and tests assert every entry names a file that exists on disk so a stale entry fails loudly instead of rotting. This is the centrally- enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's portability-vocab.cjs already models -- deliberately not a heuristic. Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is untouched and byte-identical: the win32 concern does not apply to a POSIX runner, and a test asserts the allowlist does not leak into that tier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): inject the path impl into isPathConfined and correct the ADR count Two review findings, both fixed rather than dispositioned. A security review found tests/external-descriptor-confinement.test.cjs had left real-OS execution. The allowlist pinned it back, but that only restored INCIDENTAL coverage: isPathConfined used the ambient path module, and its test carried POSIX-only literals, so a win32 confinement escape was unverified on every platform including Windows. isPathConfined now takes an optional third parameter carrying the path implementation, defaulting to the ambient module. Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change is purely additive and every existing two-argument caller is byte-identical. Tests now inject path.win32 and path.posix, covering a different drive letter, a cross-drive absolute, backslash and forward-slash traversal, UNC, and the startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators. Proved load-bearing: dropping the + p.sep from the prefix check fails exactly the two boundary cases and nothing else. Callers' suites 149/149. The spec review caught an off-by-one: the ADR narrated a 264-file tier while the committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264 -> 265 (28.5%). Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated classify() as zeroing windows_tests, a key this change removes -- kept as historical narration but labelled as such. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#131): make the unwritable-HOME test actually test something Found by sweeping for the root-bypass class after fixing commit-files-deletion. This one is the silent variant, and it was broken twice over. First, the condition: the test made a fake HOME unwritable with chmod 0o500. The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed writable and the hostile condition never existed. Replaced with a HOME whose PARENT is a regular file, so every write under it fails ENOTDIR at the VFS layer for every uid -- no permission check is involved at all. Second, and more fundamental: the probe was npm --version, which on npm 11.19.0 performs zero filesystem I/O against HOME. Proven rather than assumed -- neutralizing runNpm()'s isolation turned the sibling test red while this one stayed green, so its assertion could never detect the regression it guards, on any uid, with or without the condition fix. npm config get cache was tried next and proved vacuous the same way (it only string-resolves the path). The probe is now npm cache verify, which really does mkdir _cacache under HOME. Re-proved load-bearing after the change: with isolation neutralized the test now fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was restored and verified diff-clean; suite 13/13. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the net drop-out figure in ADR-4641 The Consequences section still said 292 files leave real-OS Windows execution. That was the count before the narrow shell-interpreter-spawn replacement restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary categories rather than only raw-child-process, and clarifies that the 14-file filename audit was against the 292 initially dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the rejected concentration ceiling and its measurement Applying Goodhart's own question to the new ceiling -- how would you make this metric look good without improving what it represents -- surfaces a real weakness: a ratio can be satisfied by inflating the denominator, so adding OS-agnostic tests loosens it without narrowing the tier. The obvious companion gate was a sole-signal concentration ceiling, since the original defect was one detector carrying half the tier. Measured and rejected: peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the original defect; any threshold below it fails on a legitimate category. The discriminator is whether a signal is platform-meaningful, which no threshold encodes. Weakness disclosed rather than covered by a gate that does not bind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4641): add the changeset fragment for the confinement-check change changeset-lint failed on PR #4643: the PR touches user-facing paths and carried no fragment. The earlier no-changeset call matched #4604's CI-only precedent and was correct then; it was not revisited once the PR grew a src/ change, which is my miss. The fragment describes the real user-visible improvement: the external-descriptor write-confinement check's Windows semantics are now verified deterministically rather than only when the suite happened to run on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the tier count in TESTING-SUITES.md Said the tier narrowed from 546 to 254. The final committed list is 265 of 931 eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored 9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in the ADR, in a live reference page rather than a dated record, so it states the current truth rather than carrying an amendment note. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the measured aggregate from real CI job lists Epic #4589's closeout asserted its reduction from a static count; #4641's acceptance criterion asks for a figure read off a real run. Recorded here: test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4. Also states the caveat that a PR's total CHECK count is not a clean before/after comparison, since many gates are path-scoped and this change touches a broader path set -- the like-for-like figure is the test.yml job count. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): compare job totals the same way on both sides The measured-aggregate table put #4640's COMPLETED run total (21) against this run's count at matrix-expansion time (15). Those are not the same measurement: the completed total includes the post-test Coverage gate and baseline-publisher jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux jobs 7 -> 4, was correct and is unchanged. Called out in the table rather than silently corrected -- comparing two differently-derived numbers is exactly the error class this ADR is about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record measured conformance wall-clock and date the stale counterfactual Adds the per-job durations from both runs. The honest read is that this is a correctness win more than a speed one: file count fell 52% but wall-clock only 9-29%, because what was removed were the cheap static tests and what remains is concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects a future narrowing to buy time proportional to file count. The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about -- pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its tier was UNCHANGED at 197 files, which fixes that as runner variance and is noted as a caution against reading a single duration as signal. Also dates the symlink-keyword counterfactual, which cited a 254-file tier from before the replacement category and allowlist took it to its final 265. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero next gained #4644 mid-flight, so every absolute count shifted. Re-measured on the tree this actually ships against (932 eligible): 548 -> 257 by detector removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed, 9 restored. macOS 198, unchanged by this PR. The percentages did not move across three rebases (58.8% -> 28.5%), which is the whole argument for expressing the ceilings as ratios rather than counts -- noted in the ADR since it is now evidence rather than assertion. Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own win32 test cases introduced the literal win32 into the pinned file, so it classifies in on content via win32-darwin-literal. The entry stays and the reason is written down, because the file's real-OS need is a property of the code under test (isPathConfined reads the ambient path module), not of the test's text -- the text that currently saves it is incidental and could be refactored away silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
db4d8a9bae |
fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic $((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is valid shell-arithmetic syntax at all, and the failed expansion aborts the rest of the snippet in a non-interactive shell. safe_resume_gate runs unconditionally before trusting STATE.md or dispatching any executor, so execute-phase failed at its own gate before the first executor on any decimal phase, regardless of workflow.tdd_mode. Regression from #4194. Fixes all 4 sites: safe_resume_gate and the TDD gate in workflows/execute-phase.md, the completion-signal spot-check fallback in workflows/execute-phase/steps/completion-reconciliation.md, and the executor gate validation example in references/tdd.md. Each now zero-strips only the leading integer segment into a *_INT variable (via %%.* / # parameter expansion — always valid shell syntax regardless of what follows) and keeps the remainder as an escaped-dot string for the anchored commit- scope regex, exactly as issue #4619 verified in both bash and zsh. A plain integer phase (12, 01) computes byte-identically to before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug Behavioral coverage via real bash execution: the old $((10#01.1)) form throws (characterizes the bug, matching the issue's own reproduction); the new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):, feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring issue #4619's own verified table exactly. Updates safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions (one per site) to the new fixed text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic With #4619's fix in place, the guard's original "ban $((10#... outright, match any occurrence" was too blunt: it flagged a comment merely mentioning the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an already-%%.*-stripped integer, and the always-safe plan-id arithmetic (plan ids are plain integers, never decimal). Refines the detector to skip full-line comments and to only flag a captured variable/placeholder name that contains "phase" and does NOT end in _INT/_int — the naming convention the #4619 fix establishes at all four sites for "already reduced to a safe integer." A plan-id variable was never phase-number arithmetic in the first place and is excluded on the same basis. This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new exemptions") and D7 ("a decimal and N-segment phase id survive an end-to-end execute-phase selection without error") for real — the guard now reports zero violations across all five .cts/.md rules. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): cover the plain-padded-integer near-miss matrix too Review found the anchored-ERE near-miss coverage only exercised the decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also verifies the plain padded-integer case (01 -> PHASE_N=1) against its own near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the missing assertion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote only 2 backslash characters in JS source, which single-quoted-string parsing collapses to 1 real backslash at runtime -- but the workflow/reference files actually contain 2 raw backslash bytes at that position (needed so bash's ${var//pattern/replacement} produces the correct single-backslash output). Write 4 backslash characters in the JS source at all 4 occurrences so the runtime string matches the files' real bytes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): refresh the committed compact-content benchmark baseline The new PHASE_INT/PHASE_FRAC arithmetic lines added to gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): note the safe_resume_gate arithmetic growth in the test header The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846 bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading integer segment via base-10 arithmetic instead of forcing the whole value through $((10#...)) and hitting a hard shell syntax error on the first dot. A blank line previously separated the Emitted-Drift-Ack-Growth trailer from the Co-Authored-By trailer below it, which splits git's trailer-block detection: only the last contiguous non-blank run of Key: Value lines at the end of a commit message is recognized as trailers, so the growth ack was silently read as ordinary body text and the differential-attribution gate failed with the growth unacknowledged. Joining the two trailers into one contiguous block fixes it. Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4208): replace chmod-based restore-failure injection with a root-proof git shim `tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated an unwritable index via a `post-index-change` hook running `chmod a-w` on the git dir. That relies on the OS enforcing the *owner's own* permission bits against itself, which uid 0 (a routine identity inside this repo's Docker-based gsd-test benches) does not: every DAC check short-circuits true for root, so the write the chmod meant to block silently succeeds, the restore comes back clean, and the disclosure/rollback behavior under test never actually gets exercised. This is CLAUDE.md's own named anti-pattern for I/O-failure injection ("Cross-platform test IO-failure injection" — chmod tricks fail under root Docker/CI). It is confirmed as the actual root cause here, not a production defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure logic (added by #4253, merged just before this run) was hand-traced and manually reproduced end to end on an unprivileged workstation against a freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces exactly the `staging_failed` + "could not be restored" / "could NOT be restored during rollback" results both tests assert. The other `post-index-change`-based tests in this file (a `sleep` to force a timeout; a real `update-index` to flip a restored entry's mode) are unaffected because neither depends on a permission check — consistent with only the two chmod-based tests failing on the real remote run. Replaces the chmod fixture with a fake `git` placed ahead of the real one on PATH that fails only `update-index --add --cacheinfo` — the one call the restore makes — unconditionally, regardless of privilege level. Every other git invocation execs straight through to the real binary, so the rest of each scenario (`rm --cached`, the restore's own `ls-files` verification, etc.) is exercised exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): backfill changeset pr number to 4644 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI Passing the script as a `-c "<script>"` argv element made it subject to Windows' CreateProcess command-line argument encoding, which silently dropped the escaped-dot backslashes before bash ever saw them (observed on PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the same script via stdin instead removes argv entirely from the transport, so there is nothing for Windows to re-encode. POSIX behavior is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5e0a7b1b56 |
fix(#4433,#4569,#4126): consolidate the phase-identity seam at name-validity, allocation, and branch-slug (#4640)
* fix(#4433): apply the name-validity guard symmetrically to every milestone-name capture extractMilestoneHeadingName already refused a punctuation-only captured name (#4134), but its two sibling capture sites in getMilestoneInfo — the STATE.md-anchored 🚧-bullet match and the no-STATE.md in-progress 🚧-bullet fallback — skipped straight to a bare truthiness check, so a malformed bullet whose only content past the version was punctuation passed through as a real milestone name. Extracts the existing inline /[\p{L}\p{N}]/u check into a single shared hasNameableContent predicate and applies it at all three capture sites, so the guard is one owner rather than a copy that happened to land at only one of them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4433): pin the name-validity guard at all three milestone-name capture sites Failing-first coverage for the hasNameableContent extraction: a punctuation-only 🚧-bullet name must not surface as a real milestone name, either on the STATE.md-anchored path or the no-STATE.md in-progress fallback, while a real name (including a digits-only one) still resolves COMPLETE exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4569): consolidate decimal-phase-number allocation into one function cmdPhaseInsert allocated its next decimal sub-phase number by scanning only on-disk phases/ directories and ### Phase N.M: headings, never the roadmap summary checklist — so a decimal that existed only as a checklist bullet (no heading yet, no on-disk directory yet) was invisible, and phase insert could silently reallocate an already-used number. It also always nested one level deeper under afterPhase, with no way to request a sibling. cmdPhaseNextDecimal had its own separate, near-identical two-source scan (missing the checklist source too) — the exact "duplicate implementations kept in sync instead of deleted" pattern this issue exists to close. Extracts scanExistingDecimalPhaseNumbers (directories + headings + checklist bullets, in one place) and migrates both cmdPhaseInsert and cmdPhaseNextDecimal onto it — deleting cmdPhaseNextDecimal's own copy rather than patching it in parallel. Adds an allocation: 'nested' | 'sibling' argument to cmdPhaseInsert (default 'nested', matching every existing caller's behavior); a top-level phase with no existing decimal segment falls back to nested since there is no sibling level to join. No CLI flag wires 'sibling' yet — that is a separate, disclosed follow-up. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4569): pin decimal-allocation coverage across phase insert and next-decimal Failing-first coverage for scanExistingDecimalPhaseNumbers: a checklist-only decimal must not be reallocated by phase insert; a decimal present in heading, checklist, and on-disk directory simultaneously must count once; an unrelated phase family's checklist bullet must not cross-pollute; and phase next-decimal (migrated onto the same shared helper) must see a checklist-only decimal too, closing the same gap in a second command. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): extend the phase-id drift guard for name-validity and shell arithmetic The epic's ratchet requirement: lint-phase-id-drift.cjs must cover the two new predicates this PR introduces, and must also scan shell inside gsd-core/workflows/**/*.md and gsd-core/references/**/*.md for integer-coercing phase-number arithmetic ($((10#...)) and friends), which neither the canonical TypeScript module nor a source-only lint can reach. Adds findNameValidityDrift (bans re-deriving /[\p{L}\p{N}]/u outside hasNameableContent's owner file) and findShellPhaseArithDrift + scanMarkdownShellArith (bans $((10#...)) in workflow/reference markdown, sanctioned via <!-- phase-id-owner: --> on the preceding line). scanRepo keeps its existing, narrower contract (src/**/*.cts only) so the already-passing "the live repo is clean" test is untouched; a new scanAll merges both for the CLI's full report. Running the guard directly against this tree correctly reports the 7 pre-existing #4619 shell sites (workflows/execute-phase.md x4, workflows/execute-phase/steps/completion-reconciliation.md x2, references/tdd.md x1) as violations — demonstrating the ratchet works, not fixing them. #4619 is a live regression tracked and fixed separately; this PR does not touch those markdown files. A characterization test pins the current count of 7 so a future change to that number is investigated rather than silently absorbed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4569): wire --sibling through phase insert's CLI so the argument is reachable cmdPhaseInsert's allocation parameter had no CLI path to 'sibling' — shipped, untested, unreachable code (code-review finding: a guaranteed surviving mutant). Adds --sibling to phase insert's argument parsing, threads it through, and documents the flag in docs/CLI-TOOLS.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4569): exercise --sibling end-to-end through the real CLI Confirms --sibling joins afterPhase's parent decimal level rather than nesting, and falls back to nested when afterPhase has no existing decimal segment (no sibling level to join). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4634): demonstrate the two new drift detectors end-to-end via a planted violation The epic asks for the guard to be "demonstrated by watching it go red" on a reintroduced copy. The two new detectors (name-validity, shell-arith) had only unit-level fixture tests; mirrors the existing bracket-rule's planted-violation-in-a-temp-tree test for both, proving they're actually wired into scanRepo/scanMarkdownShellArith end-to-end, not just correct in isolation. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): consolidate the drift guard's own owner-sanction-check logic Standards review flagged the "walk to nearest preceding non-blank line, check for a phase-id-owner comment" logic as duplicated across all four detector functions in a PR whose whole point is eliminating exactly that pattern. Extracts isSanctionedByPrecedingComment, shared by all four; behavior-preserving (verified: identical output before/after, same 7 known #4619 violations, zero token/bracket/name-validity). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): add Fixed changeset for the name-validity guard and allocation consolidation pr:0 placeholder — backfilled once the real PR number exists. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4126): consolidate branch-name slug substitution into one shared renderer cmdCommit (commands.cts) and cmdInitExecutePhase (init.cts) each independently implemented branch-name template substitution, and both substituted the literal string 'phase' when phase_slug was empty or undeliverable — producing a non-identifying branch name (gsd/phase-08-phase) that contradicted the honestly-reported phase_slug: null in the same payload. Same structural defect as the other three gaps in this epic: two consumers reimplementing one concept independently instead of sharing an owner. Adds renderPhaseBranchName (src/phase-id.cts) as the sole owner: a real slug substitutes normally; an empty/undeliverable one drops the {slug} token plus one adjacent separator (collapsing/trimming the result) rather than substituting a placeholder word, for the shipped default template and any user-configured shape alike. Both call sites now delegate to it; the old inline duplicates are deleted, not kept in sync. {project} substitution stays a separate step in init.cts, unchanged, since it is a config-level field with its own fallback contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4126): pin renderPhaseBranchName and both migrated call sites Property-based coverage for the shared renderer's degrade-path invariant (output, when non-null, never contains {slug} and never starts/ends with a separator), plus example coverage for real-slug substitution, empty/null/ non-string slug, token position at either edge, a doubled-separator template, and the only-{slug} -> null case. One regression test each in commands.test.cjs and init.test.cjs confirms a phase with no derivable slug no longer produces a branch name ending in the literal '-phase'. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: route scanExistingDecimalPhaseNumbers through the canonical enumeration owner Caught by an actual gsd-test run, not a hypothesis: the new decimal-scan helper (fix(#4569)) enumerated phases/ directories via a raw fs.readdirSync, which the pre-existing phase-enumeration drift guard (#3185/#3882) correctly flags as an unsanctioned re-derivation outside its canonical owner (listAllPhaseDirs / isSentinelPhaseId). Ironic given this epic's own thesis, and exactly why the guard exists: consolidating one seam can reintroduce drift in an adjacent one if the new code doesn't route through what's already there. Migrates the enumeration to listAllPhaseDirs; identical decimal-detection output for every existing case. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): extend the drift guard for branch-slug fallback; fix a real regex bug Adds the fourth detector the epic's ratchet section names ("both branch-name sites"): bans a `.replace('{slug}', ... || 'phase')` call outright, sanctioned via renderPhaseBranchName or a dedicated comment. Wired into scanRepo (no per-file exemption — this is a banned anti-pattern everywhere, not a grammar with one legitimate owner). Now that #4126's fix (prior commit) has landed, scanRepo reports zero violations across all four .cts-scanning rules, restoring the simple "the live repo is clean" assertion instead of a pinned-known-count characterization. Also fixes a real bug an actual gsd-test run caught: findNameValidityDrift's regex didn't tolerate the doubled-backslash template-string form its own test claimed to cover (0 !== 1) — widened to \{1,2} matching TOKEN_DRIFT_RE's existing tolerance for the same two forms. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4126): document the {slug} degrade behavior; update changeset for the full seam Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: detectPhaseNumberFromFiles wrongly rejected bare, slug-less phase directories Caught by an actual gsd-test run on the #4126 regression test, not a hypothesis: a bare phase directory with no slug remainder (e.g. .planning/phases/01/) has extractPhaseToken correctly return "01" — which is simply identical to the directory name in that case, not its no-match fallback. A stale `token !== phaseDir` check treated that equality as "no numeric token found" and rejected it regardless, leaving phaseNum null and silently skipping cmdCommit's phase-branching block entirely (the commit proceeded on whatever branch was already checked out instead of the phase branch). phaseTokenShape.test(normalized) already excludes every genuine non-phase case on its own: extractPhaseToken's real no-match fallback only fires for a dirName that doesn't start with a digit or short letter+digit prefix, and normalizePhaseName's leading-\d+ requirement rejects those regardless. The equality check was redundant for real rejections and actively wrong for bare-numeric directories. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: backfill changeset PR number to 4640 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |