6b0b92674aca56f0de97afbd92f339adcdfd0f43
19 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6b0b92674a |
refactor: remove dead descriptor-driven mechanisms (no remaining consumer)
Drop code whose only consumer was a retired runtime: hostBehaviors readers for agentFileExtension, localTargetIsProjectRoot, sharedHooksDirName, skillsManifestPrefix and skipCodexSkillsManifest; the empty NON_REGISTRY_CONFIG_HOME_DESCRIPTORS array and live-config-guard plumbing; the empty RUNTIME_NOTE_AUDIENCE_BY_HEADING filter; the unused resolveVersionFrom export; and the WINDSURF_SESSION_ID workstream session key. Delete tests that only exercised those mechanisms. |
||
|
|
6cfa0c55d2 |
refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes, cline, codebuddy and pi end to end: capability descriptors, installer branches and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters, hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi migrations, Kimi payload normalization in the hook guards, dead hostBehaviors vocabulary, launcher home probes, fixtures, runtime-specific tests and the prose that presented them as supported. Installer output for the six kept runtimes is byte-identical to before the prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and read-injection-scanner are left in place pending a decision. |
||
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
88b5775dc8 |
enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI gsd-ui-auditor is chartered to audit interaction and handed a capture driver with no interaction verb: `npx playwright screenshot` cannot click, fill, hover, press or snapshot, so a hover state, an open menu, a focus ring or a form's validation state never appears in its evidence and every Experience Design finding degrades to code reading. Implements the shape approved at triage, not a new capability: - capabilities/ui/capability.json declares `workflow.ui_interaction_capture` (boolean, default false) on the capability that already owns the auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated. - gsd-core/workflows/ui-review.md reads the key through gsd_run and hands it to the auditor as `interaction_capture:` in the spawn <config> block — the auditor carries no gsd_run resolver, so the key travels by value. - agents/gsd-ui-auditor.md gains an anchored interaction-capture section AFTER the static block. With the key on and a Chrome binary resolved it starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on an --isolated profile, opens the dev URL the static block reached, takes the a11y snapshot for element uids, captures the baseline and a Tab focus-ring state, drives the UI-SPEC's interactive components, saves console output, and stops the daemon unconditionally. Key off, no dev server, or no Chrome: one status line, and the Playwright-only static path runs exactly as before — the static fence is untouched. Needs only Bash: no MCP server, no tools: change. Chromium-only by nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so readiness is polled through evaluate_script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): bind the interaction-capture shape and containment - manifest, generated registry, config schema and config-set/loadConfig all know workflow.ui_interaction_capture as a default-off boolean, and hand-written non-booleans fall to the slice default - the orchestrator reads the key and hands it down; the auditor never grows a gsd_run dependency - the static fence stays Playwright-only and the interaction fence chrome-devtools-only, so key-off is today's path - the interaction fence runs under bash with a stub driver on PATH: key off / absent / no dev server / no Chrome invoke nothing; the happy path starts first and stops last on the [selected] pageId with the documented flags; a failed capture is removed and not counted; new_page and start failures still honour the stop-only-if-started rule; CHROME_BIN and CHROME_DEVTOOLS_MCP_VERSION overrides flow through - docs/CONFIGURATION.md row shape; registered in the docs-guard lane Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * docs(#4223): document workflow.ui_interaction_capture and its how-to - docs/CONFIGURATION.md: one row in the workflow.* table, default-off - docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the interaction-capture section adds, skips and never claims - docs/how-to/enable-ui-interaction-capture.md: turn it on, read the `**Interaction captures:**` outcomes, what it does not do, turn it off - docs/README.md: index the how-to beside live-DOM verification Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): add changeset Added-type fragment; pr: carries the issue number until the PR exists. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace; the hyphen form is retired there and the slash-command-namespace guard rejects it. docs/ keep the hyphen form by convention. Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered. Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): per-run daemon session, bounded navigation, and step failures that count Three findings from the pre-file adversarial review of the interaction fence, folded in: - `--sessionId <epoch>-<pid>` on every driver call. `start` restarts whatever daemon shares its session and --isolated isolates only the browser profile, so two concurrent audits — or an audit beside the operator's own CLI daemon — would otherwise stop each other. The CLI accepts hex and dashes only; the id is validated by the test stub. - `new_page --timeout 30000`: the one verb that takes a bound, placed before every verb that does not, so a hung page is caught first. - a failed take_snapshot or press_key now increments the failure count and is named on stdout; two clean screenshots can no longer read as `0 failed` after the step that gives the interactions their uids failed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal Second review round, both reviewers: - session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a subshell, so two audits forked from one parent in the same second shared an id and could stop each other's daemon (driven by the reviewer) - `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver under Git Bash still matches the `$` anchor, and `|| true` on the assignment so a failed new_page cannot abort the block under `set -e -o pipefail` before the unconditional stop - a failed take_snapshot removes any snapshot.txt it left or inherited from a reused directory, so stale uids never drive the interactions - `<config>` placeholder is `{interaction_capture}`, lowercase like its `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt template, not a bash heredoc - how-to: the `not captured` row no longer claims the daemon started Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges Third review round: - a new_page that prints a page line and then exits non-zero is a failed navigation, not a page id: the exit status is checked in an `if` before the output is parsed (driven by the reviewer against the previous `|| true`, which masked exactly that) - regression cases for what the last two rounds added: CRLF driver output, a stale snapshot removed on failure, partial-output new_page failure, and the whole fence under `set -e -o pipefail` (both the failed-navigation path and the happy path) - the harness whitelist gains `date`; the session-id assertion now requires all three parts, so a silently empty epoch cannot hide again Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder - the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes, over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces; the interaction section's comments are tightened to the same content in fewer bytes (23559 now). No bash changed — the fence's own tests and the real-browser run are unchanged. - .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which the post-create backfill rewrites to the PR's own number. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): set changeset fragment pr to 4477 * test(#4223): compare the fence's status path with the separator the fence uses On the windows-latest lane the happy-path case failed on `\interaction` vs `/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a literal slash, and the assertion built its expectation with path.join. Every other case in the file passed on that lane, including the CRLF and errexit/pipefail ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): drop the inert file-header allow-test-rule marker Review round 1 on #4477: the `source-text-is-the-product` marker sat at line 2, outside no-source-grep's 8-line lookahead of every readFileSync site (the first is ~60 lines down), so it suppressed nothing. It was also unnecessary: every read in this file is a .md/.json path, which the rule does not trigger on. Deleted rather than relocated — there is no site to relocate it to. Negative control: `eslint` on the file is clean without it. * chore(#4223): regenerate the platform-conformance-tier lists for the new test Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier gate after this branch opened; its two committed lists must name every file under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never been in them. Once the branch was updated against next the lists were stale and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier --check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's fragment-single-edit-propagation, which sees the same staleness as regen:derived touching files beyond the fragment edit under test. Regenerated with the repo's own generators, no hand-editing. The general tier goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry; both --check arms are clean. Verified the red is this PR's own file and not base drift: at upstream/next both generators report "list matches" (546 / 196), and our committed copies were byte-identical to next's before this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ * fix(#4223): bound, confine and trap the chrome-devtools driver fence Round 4 — three findings in one fence, interleaved on the same lines, so one commit: - Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a background job in its own process group (`set -m`) under a watchdog that kills the whole group at the ceiling — TERM, then KILL two seconds later. One pid is not enough: npm forwards SIGTERM only to its direct child, so killing `npx` alone leaves the client holding the fence's stdout and a `$(cdt … new_page)` capture blocked past the ceiling (driven against a real npx tree by the round's adversarial review; the pid-only first cut of this commit had exactly that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )` subshell: a subshell inherits bash's saved copies of the caller's stdio (the fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a watchdog outlived its kill — measured as the intermittent 30 s test run the review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 -- -pgid`, every 0.1 s) and stands down by itself once the group is empty; nothing ever signals it. The group, not the leader pid: a child can outlive the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll stood down at once and left the substitution open-ended (driven by the round's adversarial review at 6× the ceiling; a pgid cannot be reused while any member lives, which a bare pid can). The daemon `start` launches is spawned detached (its own session) and never in that group. Two platforms forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM; sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git Bash a signal to a watchdog still starting up hung the fence's `wait` for it: 18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs, while a fence slowed by xtrace, or three tests run alone, never hit it (a startup race; the mechanism is not pinned further). Polling: 0/300 slow calls and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3 full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU, msys and busybox sleep; BSD sleep documents it). A clock that cannot launch (`sleep … || exit 0`) stands the watchdog down rather than firing at once and killing a healthy call — by design that leaves a hung call unbounded, the pre-round-4 behaviour, instead of failing a healthy one. A hung call returns once its group is gone: at the ceiling, plus up to the 2 s TERM-to-KILL grace. The KILL after the grace is sent only to a group that is still alive: a pgid freed during the grace can be reused, and an unconditional KILL could hit an unrelated group (the round's review). `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog rather than either. - --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may write under the run's interaction/ directory and nowhere else. Relative, like every --filePath (unchanged from rounds 1-3): the daemon resolves both against one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and path.resolve()s both), and a relative path needs no dialect translation — an absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native daemon cannot resolve (CI's windows conformance shard caught the first cut). --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified), so the documented floor moves from ^1.8.0 to ^1.9.0, where --allowUnrestrictedPaths is deprecated. - `stop` is owed by an EXIT trap after a successful `start`, not by position (it replaces any earlier EXIT trap — none exists in this file); the explicit call keeps it in order, a flag makes the trap a no-op afterwards, and only the shell that installed the trap may act: a subshell copy of the fence state carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced locally at 3/40 under load: the second `stop` came from a subshell pid, never main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist and CI's macos conformance job showed the guard comparing empty to empty. The fence was driven under bash 3.2.57 for the injected-subshell, errexit failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy paths. A failed resize_page is a counted failed step now, not the one bare command an errexit runner could abort on. Prose in the section is tightened to pay for the mechanism: 23559 -> 24517 bytes against the 24576 DEFAULT-tier cap. Tests: the stub driver hangs as a real child tree (sh waiting on a child that holds stdout — never an exec), so a pid-only kill fails the new aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative- controlled: it blocks for the harness's whole cap on the old wrapper). A hung start and a hung capture are cut off within ceiling + grace + slack and still reach stop; an injected bare failure under errexit reaches stop through the trap, exactly once; an injected subshell call of cdt_stop issues nothing; the happy path issues exactly one stop; every driver call site names a ceiling and the only bare $CDT is the wrapper's own spawn; the start line carries --workspace with the capture directory, every --filePath lies under it, and no code line carries --allowUnrestrictedPaths. A driver whose leader exits at once while a child keeps holding the capture pipe is still cut off at the ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone (negative-controlled: the trap form kills `start` in under 20 ms). The harness EXPORTS its stub-only PATH — unexported, the exec'd watchdog fell through to bash's compiled-in default PATH and never saw the stub dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash, pid-preserving). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the accessibility tree, with entered form values) and console.txt (which can carry tokens) were committable by `git add .`. The gate now ignores `interaction/` as a directory — the next artifact type is covered by construction — and it appends whatever an existing .gitignore lacks instead of writing once. The write-once form was the same defect one step later: every project that had already run an audit would never have received the new pattern at all. Tests run the gate fence under bash: a fresh file carries every pattern; an image-only file from an earlier audit gains interaction/ and keeps its own header without duplicating present lines; a second run appends nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * test(#4223): declare the interaction-capture anchor as a comment marker The #4324 colon-token gate (slash-command-namespace) landed on next after this branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an unconvertible /gsd: command token. It is a section anchor of the same family as gsd:live-dom-families and gsd:write-continue, so it is declared in COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added gates against the merged tree; CI at ca8d2508 predates the gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
c9a5cc3e12 |
fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head. |
||
|
|
c72fb34e9a |
fix(#4658): give the ui plan gate's evidence check a native branch (#4807)
* test(#4658): add failing-first coverage for native frontend evidence hasStaticFrontendEvidence recognised only JS-ecosystem evidence, so computeUiPlanGate could never block for a SwiftUI/Compose/Flutter/XAML project. Adds evidence-level fixtures for the four suggested markers (import-matched for .swift/.kt/.dart, extension-alone for .xaml), the reporter's non-UI Swift control case, marker-exactness and SKIP_DIRS and I/O-degrade negatives, gate-level block assertions through makeProject's new native frontendEvidence modes, and a pinned-seed fast-check property. All new assertions are RED until src/ui-frontend-evidence.cts grows the native branch. * fix(#4658): give the ui plan gate's evidence check a native branch hasStaticFrontendEvidence recognised only JS-ecosystem evidence (a root package.json UI-framework dep, or a .tsx/.jsx/.vue/.svelte file), so computeUiPlanGate could never block for a SwiftUI, Jetpack Compose, Flutter, or .NET MAUI project — the #3312 gate was structurally unreachable for them. Adds a native BFS over the same bounds and skip rules: .xaml is evidence by extension alone (the .tsx analogue), while .swift/.kt/.dart count only when their content carries the ecosystem's UI import marker (import SwiftUI / import UIKit, androidx.compose, package:flutter) — matched on the import, not the extension, so a non-UI Swift package stays silent exactly as the issue's 37-file control case requires. Marker reads are bounded to a 64 KiB prefix; any I/O failure degrades to false per the module contract. The #3718 vocabulary filter, the JS evidence rules, and the weaker-extension exclusion are untouched. * chore(#4658): regenerate the macos conformance tier list The native-evidence additions to tests/check-ui-plan-gate.test.cjs move the file into the macOS conformance tier per the classifier; the committed generated list is a derived artifact and must match the live tests/ tree (the same sync the fragment-single-edit-propagation install test enforces). * fix(#4658): accept both Dart quote styles and extract the shared bounded walk The isolated reviews' remaining findings: the Dart marker carried only the single-quote anchor, missing legal double-quoted imports (a spec-narrowing deviation); the two evidence walks duplicated the subtle MAX_WALK_ENTRIES cap semantics verbatim, so they are extracted into one walkProjectFiles BFS with a visit callback; tests now use the createTempDir helper, shared fixture literals that cannot drift from NATIVE_UI_CONTENT_MARKERS, a double-quoted Flutter import case, and drop a vacuous assertion and a mid-body re-require alias. * docs(#4658): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
ad1477d659 |
enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap * enhance(260903-m7p): validate configured entrypoints before success * test(260903-m7p): require pre-success entrypoint validation * enhance(260903-m7p): gate install success on entrypoints * test(260903-m7p): cover configured entrypoints across runtimes * enhance(260903-m7p): cover emitted runtime entrypoints * fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number - finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process before the new configured-entrypoint assertion throws. Without a HOME + config-location-env sandbox that write resolved through the ambient environment and landed in the developer's live ~/.gsd (confirmed absent on origin/next baseline, present only on this branch — full-suite HERMETICITY WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the duration of the test, matching the existing in-process finishInstall/ install() pattern in tests/install.test.cjs (#2665). - .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the fork PR number until the upstream PR number is known. * fix(260903-m7p): repair cross-platform and pre-existing shape fallout - tests/configured-entrypoint-validation.test.cjs: the win32 branch of ensureCodexHooksJsonSessionStart writes a .cmd shim under <codexRoot>/hooks/; create that dir in the test (the real installer only calls this once hooks/gsd-check-update.js already exists) and assert the platform-common entrypoint shape instead of a fixed non-Windows array, since win32 legitimately emits two entries (cmd shim + script). - tests/install.test.cjs: finishInstall's shared settings-json return now carries configuredEntrypoints/rollbackInstallerMigrations for every runtime on that path (trae included, not just Claude/Cursor/Windsurf); update the trae install() exact-shape assertion to match. * fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash configuredEntrypointsForHook's shell branch dropped interpreterCandidates entirely when resolveBashExecutable returned null, unlike the sibling portableHooks runner entry a few lines below (which correctly falls back to the literal 'bash' token). Found via agy adversarial review; verified unreachable through the current call graph (buildHookCommand's own resolveBashRunner==null gate already short-circuits before recordConfiguredHookCommand runs), so this is a defensive consistency fix, not a live-bug patch — kept for the next caller that does not share that gate. * chore(260903-m7p): backfill changeset pr number to the opened upstream PR .changeset/quick-wasps-sing.md carried the fork PR number (16) as a placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now open, so record its real number per CONTRIBUTING.md's changeset pr-field convention. * fix(#4154): track already-registered hooks for entrypoint validation on update applySettingsJsonHooks registers each guard hook only if absent, so a hook already present from a prior install keeps its stale on-disk command. The new entrypoint tracker always records the freshly-computed command for it, which never matches what is actually persisted, so the exact-string filter in finishInstall silently dropped it from validation — the Blocker case this feature exists to catch (an already-installed entrypoint going stale between installs) was exactly the case it never validated. Match on the managed script's basename instead, which the persisted command carries either way, so an already-registered hook stays in the validated set. Regression test forces this path by mutating a freshly-installed hook's persisted command before a second install. * fix(#4154): distinguish an unreadable script from a missing one validateConfiguredEntrypoints folded an EACCES statSync failure into the same 'missing' reason as ENOENT, misreporting a real permission problem as an absent file. Check the error code and report 'unreadable' instead. * docs(#4154): document entrypoint validation's rollback and PATH scope CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite bin/install.js x CONTEXT.md being this repo's strongest co-change pairing. The update-gsd.md how-to overstated what a validation failure undoes: for Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/ config.toml inside install() before the aggregate validation call runs, so there is no rollback path for that write regardless of "where available" phrasing. Also note that interpreter resolution checks the installer's own PATH, not necessarily the PATH a hook fires under later (#2979 launchers). * chore(#4154): point changeset pr field at the fork PR while CI runs there Mirrors the branch's own prior backfill commit: pr: matches whichever PR number changeset-lint is currently validating against (fork PR #16 during the fork-first CI/review loop), flipped back to the upstream PR number right before the final push to open-gsd/gsd-core. * fix(#4249): address adversarial-review findings in entrypoint validation An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR found several real gaps beyond the human reviewer's Blocker, verified against source before fixing: - Codex's install() result bound rollbackInstallerMigrations to the narrow installer-migrations-only rollback instead of restoreCodexSnapshot (#3245), the full pre-install snapshot/restore Codex already owns for exactly this case — a validation failure discovered outside install() reverted nothing of the config.toml/hooks.json that call had already written. - The register-only-if-absent basename match from the prior fix used a bare substring, which an unrelated user command mentioning the same filename could false-positive into GSD's validated set — anchored on the `/hooks/<basename>` path segment instead. - nodeCandidates checked raw process.execPath (always true — we're running in that process) instead of normalizeNodePath's stable version-manager alias, the same one buildNodeRunnerChainToken bakes as its first choice — a false green regardless of whether that alias itself still resolves. - An entry with no interpreterCandidates (Cline's PreToolUse hook, or a Windows-Claude .sh hook invoked without a bash runner) runs via its own shebang; validateConfiguredEntrypoints checked only file-type, never the execute bit. Cline's writer also never reported an entrypoint at all. - Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor hook registered across several events) were validated once per duplicate. Each fix is covered by a new or extended test; the Codex one required inlining runCodexInstall's env sandboxing so the rollback closure — which re-resolves the $HOME-relative skills root live — runs before the sandbox is torn down, matching how installAllRuntimes' real aggregate gate calls it. * docs(#4249): document the round-2 entrypoint-validation fixes Runtime Hooks Surface Module and Installer Module entries now name ConfiguredEntrypoint's not-executable reason, the normalizeNodePath alignment, Cline's tracked hook, and which install() result the finishInstall/installAllRuntimes rollback path actually reverts per runtime (Codex's full snapshot vs. the others' narrow migrations-only rollback). * chore(#4249): point changeset pr field at the upstream PR now that fork CI is green * fix(#4249): address agy adversarial-review findings - validateConfiguredEntrypoints: statSync alone never detects a chmod-000 script (it only needs parent-dir search permission), so an interpreter-invoked entry with an unreadable script passed validation. Add an explicit R_OK check for the interpreterCandidates branch only — the candidate-less/shebang branch already has its own X_OK gate. - docs/how-to/update-gsd.md: the blanket "does not revert" claim was false for Codex, which reverts config.toml/hooks.json via its full pre-install snapshot; qualify it per runtime. - tests/codex-config.test.cjs: the #4249 rollback regression test asserted skills/ and VERSION were reverted but never asserted config.toml/hooks.json were too, despite the test's own stated intent. - CONTEXT.md: qualify which interpreterCandidates entries get normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's Windows .cmd shim as a third candidate-less case that relies on extension dispatch, not a shebang. * fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node', so it needs the execute bit (like any shebang-invoked entry), but its interpreter is looked up on PATH by 'env' at hook-fire time (unlike every other GSD JS hook, which bakes an absolute node path specifically to avoid that dependency). The candidate-less/interpreterCandidates fork treated these as mutually exclusive, so Cline's entry silently skipped interpreter resolution entirely — a completely missing 'node' on PATH would still validate successfully. Add an orthogonal selfExecutable flag so both checks run for entries that need them. (CodeRabbit finding on the fork rehearsal PR.) * fix(#4249): address second-round adversarial review findings (opus + agy) - validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry, not just interpreterCandidates ones — a self-executable shebang script is still opened and read by its kernel-invoked interpreter, so X_OK alone never proved it was readable. - selfExecutable is now the sole, explicit source of truth for the execute-bit check (every producer that needs it sets the flag) instead of being partly inferred from an absent interpreterCandidates, which Cline's hybrid entry also carries. - The execute-bit check now skips explicitly on win32 (matching resolveExecutableBinary's own carve-out) instead of relying on Node's accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows machine and not a test that simulates win32 on a POSIX runner. - bin/install.js: fixed a stale comment claiming no runtime's install()-time writes have a rollback path — Codex's does (restoreCodexSnapshot) — and added the omitted Cline to both that comment and CONTEXT.md's equivalent lists. - CONTEXT.md: fixed the Cline description left stale by the previous commit's selfExecutable addition, and rewrote the validation-mechanism paragraph for clarity (writing-for-agents pass). - docs/how-to/update-gsd.md: split an overloaded 4-clause sentence. - Removed a fault-injection integration test that could not reliably exercise the real installAllRuntimes -> finalize -> rollback wiring without fighting the installer's own pre-registration existence guards; the constituent pieces remain covered individually. * fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner X_OK is a POSIX-only concept, skipped entirely when an entry's platform is win32 (matching production). Two test entries omitted platform, defaulting to process.platform — on an actual windows-latest CI runner that silently skipped the very check they were meant to exercise, turning 'not-executable' into a false pass. Pin platform: 'linux' so these are deterministic regardless of which OS runs the suite. * fix(#4249): classify EPERM the same as EACCES in statSync error handling Windows raises EPERM (not EACCES) for a parent directory that couldn't be traversed into — was falling through to 'missing', misreporting a genuine permission problem as a nonexistent path. * docs(#4249): address final CodeRabbit doc-completeness findings - CONTEXT.md: install()'s documented result shape omitted configuredEntrypoints; the ConfiguredEntrypoint shape omitted selfExecutable. - docs/how-to/update-gsd.md: the failure-mode sentence omitted unreadable and lacks-execute-permission, which the installer also rejects. * fix(#4249): stop double-validating every configured entrypoint on install/update installAllRuntimes' finalize() already runs assertConfiguredEntrypoints once over the aggregate set; finishInstall then re-ran the identical check per runtime in the printSummaries loop right after, so every entrypoint paid its statSync/accessSync/interpreter-resolution cost twice on every install and update. Add entrypointsAlreadyValidated to skip the redundant pass specifically on that path, while leaving the check intact for any caller that invokes finishInstall directly. * chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there * perf(#4249): memoize interpreter candidate resolution across entrypoints resolveExecutableBinary walked PATH once per (entry, candidate) pair; a typical install has a dozen-plus entries sharing the same few candidate lists (process.execPath for JS hooks, bash for shell hooks). Cache by (platform, candidate) so each distinct pair resolves once per validation call instead of once per entry. * chore(#4249): point changeset pr field at the rebased rehearsal fork PR * fix(#4249): drop entrypoint tracking from the now-dead Codex event writer #2586 (landed on next after this branch forked) removed install.js's CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent no longer runs during install or update. The ConfiguredEntrypoint records this branch added inside it were therefore unreachable and untested. Restore the function to its upstream shape; the entrypoints it used to report were never collected by any caller. * refactor(#4249): drop the revalidation bypass flag and the candidate cache Both were this PR's own micro-optimisations over a set of roughly a dozen entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall gate off to save one statSync/accessSync pass; `resolvedCandidateCache` memoised resolveExecutableBinary across entries that are already deduped by (configPath, scriptPath). Neither is measurable, and the flag was the only way to reach finishInstall with validation disabled. finishInstall now always validates what it is given. * chore(#4249): point the changeset pr field back at the upstream PR * refactor(#4249): track settings.json entrypoints without the hooksSurface gate The install-surface writer only tracked configured entrypoints when the runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing asserts that axis agrees with `installSurface`, so a descriptor that broke the coupling would silently pass `configuredEntrypoints: undefined` and drop that runtime out of the validation this PR adds — reintroducing the exact 'reports Done! over a broken entrypoint' failure #4154 exists to close. Remove the dependence rather than test it: everything recorded on this path lands in settings.json by construction, and the registered-command filter already discards entries no persisted hook references. * chore(#4249): put the changeset body in the documented two-part format CONTRIBUTING.md and .changeset/README.md both show `**<bold change>** — <symptom-led explanation>.`; the fragment was a single unbolded sentence. * chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there * fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback #3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-* and gsd-core/VERSION. The install overwrites every other GSD-owned file too — hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest itself — before the entrypoint-validation gate runs, so a validation failure left the new payload sitting on top of the restored old config. Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims, before runInstallerMigrations so the bytes are the true pre-install state, and restore it from both Codex rollback closures ahead of the per-surface restores. Files only the failed install introduced are removed, read from the manifest now on disk. The manifest is already the authoritative record of what GSD owns, so no second hand-written list can drift out of sync, and user-owned files are never snapshotted or removed. Every path is confined through resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into an arbitrary-path write. Non-Codex runtimes are unaffected: the snapshot is gated on the same tomlConfigInstall + non-minimal condition as #3245's. * fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest Two follow-on defects in the previous commit's snapshot: - The capture was gated on `!isMinimalMode`, copied from #3245. A core/ --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the snapshot came back empty while the rollback still ran — and its removal pass would have deleted every file the new manifest lists. Gate on tomlConfigInstall alone, matching where the rollback actually reaches. - An unreadable or unparseable prior manifest was caught alongside ENOENT and treated as a fresh install. That is the same empty-snapshot state, so a failed update over a real install with a corrupt manifest could delete its prior payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means known-empty; any other read error or a parse failure means unknown, and the restore closure returns without touching anything, degrading to #3245's narrower rollback. Deliberately not fatal — a corrupt manifest has to stay repairable by reinstalling over it. Both paths are covered by red-checked regression tests. * fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too commit removed from the manifest snapshot. restoreCodexSnapshot is reachable for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-* skill dir and gsd-* agent file the snapshot does not claim — so with an empty minimal-mode snapshot a rollback deleted the whole skills/agents surface with nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via the ADR-1239 skills-kind home override, so this is also the reason manifest `skills/` keys do not resolve under configDir: that surface belongs to this snapshot, not to the manifest-driven one. Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal mode — doing nothing on an early failure is the non-destructive side. Covered by a red-checked regression test that plants bytes in an alternate-home skill file, reinstalls under the core profile marker, and asserts the rollback restores it. * fix(#4249): never remove on rollback unless a prior manifest proves what predates the install Three defects in the manifest-driven Codex rollback, all in its removal half: - ENOENT marked the snapshot usable, arming the removal pass on a FIRST install. GSD may have overwritten a user's file at a manifest-tracked path there, and no prior manifest records the difference — so rollback deleted it where before it merely left it overwritten. Absent, unreadable and malformed manifests now all leave the prior set UNKNOWN and skip removal entirely. - Membership was tested against the map of files whose pre-install read SUCCEEDED, so a tracked file that existed but was unreadable read as introduced-by-this-install and was removed. Track the prior manifest's paths in their own Set and test against that. - The unreachable "delete the manifest when there was no prior one" branch is gone: usable now implies a parsed prior manifest. Also adds the end-to-end test the aggregate gate was missing — the four Codex rollback tests drove the closure directly, proving the restore but not the wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes Cline's `env node` entry fail validation for real, and asserts Codex's payload comes back. Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper instead of six copies. Both new tests are red-checked. * test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its Windows-EBUSY retry budget. That budget is for directory trees; this removes a single file, which unlinkSync says more precisely and the rule does not flag. * chore(#4249): point the changeset pr field back at the upstream PR * fix(#4249): use an unambiguous dedup key and surface partial-restore failures trek-e's 2026-09-08 adversarial pass flagged two findings in the new entrypoint-validation/rollback code: - assertConfiguredEntrypoints' dedup key already used a raw NUL separator (introduced in ceebb65f2d), but git/Read render NUL as a space, so the key looked like a plain-space join to every reviewer that read the diff. Replace it with JSON.stringify([configPath, scriptPath]) so the separator is visible and unambiguous. - restoreManagedFileSnapshot's per-file restore catch block claimed to 'surface the original error' but only swallowed it, matching (and widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add an actual console.warn using the existing best-effort-warning convention, scoped to just this PR's new function. * fix(#4249): treat a files-less prior manifest as unknown, not known-empty agy's gemini-3.8-flash-high adversarial pass (round 5) found and I reproduced empirically: a structurally-valid manifest missing the files key (e.g. {"version":1}) parses without throwing, so Object.keys(undefined || {}) silently read as 'zero files predate this install' instead of the UNKNOWN state the malformed-manifest guard exists to produce. Rollback's removal pass then deleted every GSD-owned file the failed install's own manifest listed, including ones that predated it — the exact data loss the #4249 CodeRabbit malformed-manifest fix was supposed to prevent, reachable through a JSON.parse success instead of a failure. Route the shapeless case into the same catch-all UNKNOWN path via an explicit shape check. Regression test reproduces the deletion before the fix and confirms the file survives after it. Also extend restoreManagedFileSnapshot's removal-pass rmSync and final manifest-rewrite catches with the same real console.warn trek-e's round-4 review asked for on the per-file restore catch — same rollback function, same operator-facing-signal gap. * docs(#4249): correct which runtimes actually leave a written config on rollback agy's completeness audit (round 5, holistic pass) caught this new paragraph claiming 'for every other runtime, the configuration file(s) already written during that update are left in place' — false for Claude Code and other settings.json-based runtimes, whose write never happens on failure (assertConfiguredEntrypoints runs before finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline actually match that description, since they persist their config file inside install() ahead of the gate. Split the one sentence into the three actual outcomes; matches the PR body's own accurate Before/After wording, which this doc addition had drifted from. * fix(#4249): clean up doc/comment mismatches and dead fields from opus review Opus critical-code-reviewer + ponytail-review pass on the final diff: - assertConfiguredEntrypoints carried finishInstall's old docblock ("Apply statusline config, then print completion message") from before this function was inserted between comment and callee. finishInstall already has its own accurate #4249 comment, so the stale docblock is removed rather than moved. - checked: number on ConfiguredEntrypointValidationResult and error.configuredEntrypointValidation on the thrown error: the first had zero consumers anywhere in the repo, including its own defining file, and is removed. The second matches an existing repo convention (bin/install.js's installerMigrationRollbackFailures, #4249 predates this PR) of attaching structured diagnostic context to a re-thrown Error even before a consumer exists, so it's kept. - finishInstall's own assertConfiguredEntrypoints call is a redundant backstop on the real production path (installAllRuntimes's aggregate call already validates the superset first), but its comment read as though this call alone provided the before-the-write guarantee. Clarified rather than removed — it's the only gate for a caller that invokes finishInstall directly. * chore(#4249): split the manifest-driven rollback engine out into #4544 Issue #4154 asked the installer to consume a validation failure "through the existing rollback mechanism, without a second transaction mechanism". The manifest-driven rollback widening added during review (capture every path the prior gsd-file-manifest.json claims, restore those bytes, remove what only the failed install introduced) is that second mechanism on a plain reading. It is a real fix for a #3245-era gap, but an independent one, so it moves to its own bug report and PR. Removed here: - bin/install.js: the pre-install managed-file capture block and restoreManagedFileSnapshot, plus its call sites in _codexPreConfigRollback and restoreCodexSnapshot (99 lines). - tests/configured-entrypoint-validation.test.cjs: the five tests that exercise the manifest engine. - CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the widened restore. update-gsd.md again documents the #3245 surfaces only. Kept, because it is #4154's own scope: - the entrypoint-validation gate itself; - Codex's install() result binding rollbackInstallerMigrations to restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, gsd-core/VERSION); - the !isMinimalMode gate removal on that snapshot. Binding the closure to the result made it reachable for a core/--minimal install, where its pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot does not claim; an empty minimal-mode snapshot therefore deleted the whole surface with nothing to restore. The surviving aggregate-failure test now asserts on config.toml, a surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which only the manifest engine restored. Refs #4544 * test(#4249): cover configured entrypoints through the packed install path #4154's scope lists install smoke coverage alongside the installer gate — "assert representative configured entrypoints resolve for supported runtime profiles". The gate itself (assertConfiguredEntrypoints / validateConfiguredEntrypoints) is unit-covered by in-process install() calls; nothing proved the property survives npm pack -> npm install -g -> install.js. Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct config surfaces GSD writes launch paths into (settings.json, and hooks.json + config.toml) — run the tarball-installed installer into a throwaway HOME, then re-read that runtime's own written config and return the new ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check becomes a release gate on every matrix host without workflow changes. The scan re-derives paths from the written config instead of reusing the installer's own entrypoint list, and test I shows why that matters: a registration the installer never touched during a run is invisible to the in-process gate, so the install exits 0 and only reading the config back off disk catches the dangling launch path. * ci(#4249): pack a publish-shaped tarball in the install smoke lane `npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs after `npm ci` carries no hook scripts at all — the lane has been smoking a package that differs from the published one in exactly the artifacts the lifecycle smoke is supposed to launch. That went unnoticed because the lane's init runs `--local`, which registers no statusline and therefore registers no hook whose target is missing. A `--global` install on the same tarball exits 1 on #4249's own gate (`gsd-statusline.js (missing)`), which is what the new configured-entrypoint cycle performs, so without this step the cycle would report INIT_FAILED instead of checking anything. Build hooks before packing so the smoked tarball matches prepublishOnly. The CLI now reports 16 configured entrypoints for claude and 1 for codex instead of zero. * fix(#4249): scope Codex's full snapshot restore to entrypoint failures Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage exception un-install a Codex install that had already succeeded and already printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints` call. Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration` scopes finalize-stage rollback to installer *migrations* ("the executor uses the journal to restore modified paths"), and this PR's own operator-facing paragraph in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/ agents/VERSION revert to entrypoint-validation failures specifically ("If a script is missing, unreadable, ... For Codex, this reverts ..."). The wide behaviour is also incoherent as a transaction abort: the same doc says Cursor, Windsurf, Kimi and Cline keep the config they wrote inside install(). Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install bytes — on an update, silently downgrading a working Codex install to the previous version while the user had just been told it was Done. Select the rollback by error kind instead. `assertConfiguredEntrypoints` already tags its error with `configuredEntrypointValidation`, so the full snapshot restore runs for that error (and anything downstream of it, including finishInstall's per-runtime backstop) and the installer-migrations-only closure runs for everything else. The codex result now also exposes that narrow closure as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning the full restore, so the direct-call contract asserted by tests/codex-config.test.cjs is unchanged. Adds a regression test that installs codex+kilo together, injects EACCES on the Kilo permission write by monkeypatching node:fs (restored in a finally — never chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps the bytes the successful install wrote. Verified red against the pre-fix unconditional path. Cline cannot host this test: its plan is writesSharedSettings:false + finishPermissionWriter:null, so its finishInstall performs no write and has no non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally (unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded fs.writeFileSync. * docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing CONTEXT.md still described Codex's rollback as an unconditional bind to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint- validation failures via rollbackInstallerMigrationsOnly and the configuredEntrypointValidation error tag. Caught during the round-6 PR body pass. * fix(#4249): stop rollbackInstallerMigrations meaning its own opposite Codex's install() result bound `rollbackInstallerMigrations` to restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the actual installer-migrations-only closure behind `rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name meant the opposite of what it says, and CONTEXT.md had to concede as much in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure for every runtime, matching both its name and the meaning it already has on next, and the snapshot restore gets its own Codex-only field, `rollbackPreInstallSnapshot`. The selection in rollbackFinalizedInstallerMigrations collapses to one line and no longer needs a fallback chain. Also in this commit, all against the same rollback path: - Correct the rollbackFinalizedInstallerMigrations comment. It read as if the round-6 narrowing prevented any sibling-triggered revert of a Codex install the user has already seen "Done!" for. It does not, and is not meant to: `wide` is true for ANY entrypoint-validation error from ANY runtime, because the aggregate gate is all-or-nothing — an invalid Cline entrypoint reverts Codex's snapshot, which tests/configured-entrypoint-validation.test.cjs's 'an aggregate entrypoint validation failure rolls the Codex install back (#4249)' asserts directly. The discriminator is the error's KIND, not which runtime owns the failing path. Comment and CONTEXT.md now say that. - Name the runtime in the "Configured entrypoint validation failed" error. ConfiguredEntrypointInvalid already carries `runtime`; the message threw it away, leaving an operator of a multi-runtime install unable to tell whose entrypoint broke — which matters precisely because the failure can revert a runtime that was itself fine. - Set `configuredEntrypoints: []` explicitly on the copilot-instructions early return. Every other branch states the key; this one relied on installAllRuntimes' `(result.configuredEntrypoints || [])` defence. `[]` is correct, not a workaround: every Copilot hook is an inline printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no GSD-managed script or interpreter to resolve. No behaviour change beyond the error-message text. * docs(#4249): narrow the smoke scan's config-surface claim to what it checks RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is told to launch is registered in one of settings.json / hooks.json / config.toml, and that nothing else in a config dir is runtime configuration. Both halves are false as stated. Cline registers its hook at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native [[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi), a directory separate from Kimi's own GSD configDir — the same gap installer-migration 007 already documents as structurally unreachable. The scan is in fact correct for what it runs against: entrypointRuntimes defaults to claude + codex, whose launch paths do all live in those three top-level files. Restate the docstring at that scope, name the two known out-of-scope surfaces, and warn that adding either runtime to entrypointRuntimes without teaching scanConfiguredEntrypoints about its surface yields a scan that finds zero entrypoints and proves nothing. The entrypointRuntimes default comment carried the same overgeneralization ("every other runtime reuses one of them") and is corrected with it. Documentation only; no code change. * fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found that install()'s settings-json early return for an unparseable settings.local.json omitted configuredEntrypoints and rollbackInstallerMigrations from its result, unlike every other branch. rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations unconditionally, so this branch silently dropped its own installer-migration rollback on a later finalize-stage failure. Completed the return shape: configuredEntrypoints: [] (matching Copilot's equally-early no-entrypoints-yet return) and rollbackInstallerMigrations (already in closure scope). Red-then-green regression test added. * test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs Same internal adversarial review (round 8): this suite's before() packed the tarball directly, without the ensureHooksDist() guard every sibling install-test suite (install.test.cjs, install-minimal-hooks.test.cjs, mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run in isolation ahead of a suite that builds hooks/dist itself, this suite's pack would ship a tarball with no hook scripts and fail closed on SMOKE.INIT_FAILED instead of testing anything. * fix(#4249): refresh stale test-timings weight for the codex-config split next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the CI shard packer's weight table (tests/test-timings.json) was never updated: codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block this PR extends with its own #4249 install()-pipeline test — had no entry at all, so the packer would silently underestimate it at the table's median weight (roughly a 9x underestimate against its real cost). trek-e's most recent review flagged a Windows shard timeout in-flight on codex-config.test.cjs, plausibly aggravated by this PR's own addition to that file before the rebase moved it. Re-measured both files locally (node --test --test-reporter=tap, max of 3 runs, matching the table's own max-across-streams methodology) and patched just these two entries — not a full regeneration, which would need real multi-lane CI data this session doesn't have access to. * fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists next's platform-conformance-tier classifier (#4591/#4598) landed after this branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's gen-platform-conformance-tier --check and both the Linux and macOS conformance suites. * fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation next's #4733 (landed after this branch's last rebase) replaced the static ISOLATED_HEAVY_FILES set with a threshold derived live from tests/test-timings.json, and pins the current derived result in EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still named codex-config.test.cjs, whose own weight this PR already dropped from 127783ms to 189ms (after splitting its heavy install()-pipeline blocks into codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The live-computed set correctly no longer includes it; the pinned expectation is updated to match. * fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it trek-e's review flagged two Major gaps: the thrown error read identically regardless of which of three real outcomes a runtime hit (nothing persisted / snapshot reverted / config left broken on disk), and no test proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/ Kimi/Cline — only Codex's revert path was ever asserted. assertConfiguredEntrypoints now tags each invalid entry with its actual consequence, mirrored from docs/how-to/update-gsd.md's existing rollback-matrix disclosure. A new test drives the same aggregate failure through Cline (whose own entrypoint is the one that fails) and asserts its hook file is still on disk afterward. * fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR One review pass (gemini-3.8-flash-high via the antigravity review lane) against this PR's full diff against next, findings independently verified against source before fixing: - Copilot's install() return object was the only one of 6 runtime branches missing rollbackInstallerMigrations — reachable now that this PR's own aggregate gate runs rollback across every result on any runtime's entrypoint failure, not just Copilot's own. - buildHookCommand's unresolved-bash early return skipped track() entirely, so a win32 install with no Git Bash silently produced an unregistered .sh hook instead of the 'unresolved-interpreter' validation failure configuredEntrypointsForHook's own comment said it would. - release-tarball-smoke.cjs reported a Cycle 4 install failure under SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing SMOKE.INSTALL_FAILED. - SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's trailing args, which also truncated any configDir containing a space (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan. Anchored the match on the already-known configDir prefix instead of a generic absolute-path guess: removes the ambiguity outright rather than patching the character class, and stays a raw-text scan on purpose (it catches a writer that emits a path without registering it — a JSON.parse of the expected schema would miss exactly that case). One suggested finding (test-timings.json "missing" the new test file) was verified false — that table only holds measured CI timings, populated after a file's first real run — and one Ponytail suggestion (a JSON.stringify dedup key) was rejected as it would reintroduce a real, if narrow, key-collision risk for no benefit. * fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment changeset-lint requires pr: to match the PR it runs on (16 on the fork, not the eventual upstream number) — rehearsal-branch convention already established earlier in this PR's history. lint-docs-guard-registration's quote-pairing heuristic doesn't require the docs/ path itself to be quoted — it flags a file once ANY quote-delimited span containing "docs/" appears anywhere in it, alongside any real fs read call. A comment ending "...update-gsd.md's rollback-matrix paragraph" supplied the closing quote character (the possessive apostrophe) the heuristic paired with an unrelated single-quoted string earlier in the file. Reworded to avoid the unquoted apostrophe next to the path. * chore(#4249): point the changeset pr field back at the upstream PR Fork rehearsal (PR #16) is green; the real target for this changeset is upstream PR #4249. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
cbbde6786a |
fix(#4546): deferred UAT follow-ups no longer block completion and promote to the backlog (#4769)
* test(#4546): failing-first tests for deferred uat follow-ups * chore(#4546): regenerate derived lists for the deferred-promotion suite The new verify-work-deferred-promotion suite changes the tests/ tree the macOS conformance-tier classifier tracks and is a novel file under the verify prefix in the test-file-count ratchet; both derived lists are regenerated/registered per their own guards' instructions. * fix(#4546): deferred uat follow-ups no longer block, and get promoted Two halves of one disconnect (#1921's deferral design vs the completion predicate): - uat-predicate: the item parser now captures the block's reason: line alongside result:. A skipped item whose reason carries the verify-work writer's 'Deferred follow-up:' template is a deliberate deferral -- non-blocking, flagged deferred in the report. Quote- tolerant (the writer wraps the value) and case-insensitive. A reasonless skip, a non-deferral reason, pending/blocked/issue/ failed/missing all still block, exactly as before. - verify-work complete_session: when the Deferred Follow-Ups section is non-empty, offer to promote the items to a ROADMAP.md 999.x backlog entry reusing next.md's prior_phase_completeness entry shape, with a --files-scoped commit. Offer, not auto-mutation -- matches the workflow's interactive convention and next.md's own prompt style. * chore(#4546): refresh compact-content benchmark baseline verify-work.md grew (the #4546 deferred-follow-up promotion offer in complete_session); the registered split's token counts moved with it. Baseline recomputed with the script's own --write. Emitted-Drift-Ack-Growth: verify-work.md — complete_session gained the deferred-follow-up promotion offer (detection, [P]/[K] choice, the next.md-shaped 999.x entry template, and the --files-scoped ROADMAP.md commit); the growth is the new contract text, not duplication * fix(#4546): gate/audit agreement and review fixes for deferred follow-ups - src/uat.cts categorizeItem: a skipped item carrying the deferred follow-up template reason now categorizes as 'deferred' (the category already existed for deferred-items.md entries) instead of being misfiled into the blocked families by keyword match -- the gate/audit agreement #3078-CR expects, restored in the permissive direction the #1921 design intends. Checked BEFORE the keyword families so '... on the release build next version' is not build_needed. - verify-work.md promotion step: numbering scans for the smallest free 999.n (count races + non-contiguous history), one backlog entry per deferred follow-up, ROADMAP.md-absent behavior specified, idea text newline-flattened, Deferred at placeholder harmonized with next.md. - DEFERRED_REASON_RE: trust assumption documented (authoring contract, not a security boundary; non-matching spellings block fail-closed). - tests: the property now drives evaluateUatPassed and derives expectations from the input spec (never restates the matcher), includes the no-result-line branch, and pins its seed; the parity test drops try/finally for the approved pattern, uses createTempDir, sites its allow-test-rule marker at the suppression site, and asserts the literal [P]/[K] choices. * fix(#4546): close promotion-test docstring, drop fc replay-path misuse, refresh baseline The final matrix run caught three defects in my own review-fix commit: the parity test file's JSDoc was left unterminated (the whole file parsed as one comment -- zero tests registered, hence the file-level 'test failed' the runner reported); fast-check's replay-path parameter was misused as a label (invalid path at replay); and the workflow-text ambiguity fixes re-drifted the compact-content benchmark baseline. * docs(#4546): add Fixed changeset for deferred follow-up coverage * docs(#4546): backfill changeset PR number * fix(#4546): use the pattern seam escapeRegex for shape-marker matching The hand-rolled metacharacter escape in the shape-marker assertion tripped local/no-adhoc-regex-escape, whose named remedy this adopts. --------- Co-authored-by: sim <sim@local> |
||
|
|
e2bfc06558 |
fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase (#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until this lands. THE DEFECT, MEASURED Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking value, indistinguishable from the same call with a canonical id. Measured on |
||
|
|
5d4c98cde7 |
chore(#4729): guard the retired-runtime name, and finish the locale residue (#4753)
* chore(#4729): guard the retired-runtime name, and finish the locale residue Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one concern: make the tree clean, and keep it clean. The guard is inert until the tree is clean, and shipping the cleanup without the guard is the one-bug-at-a-time pattern this epic exists to end. WHY A GUARD, AND WHY LAST Nothing in CI answered "does any shipped surface still present a retired runtime as live?", and the two gates that look like they should cannot. checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it reported 0 violations while all four locale mirrors still documented --gemini as a live reviewer flag, with usage examples. And tests/gemini-runtime-removed.test.cjs is scoped by construction - its own docblock limits it to the installer CLI contract and the runtime-name-policy exports; it never reads docs/**, gsd-core/workflows/**, commands/** or agents/**. Every extension to it during this epic was a hand-added assertion for a surface somebody had already noticed. A guard written earlier would have red-flagged the very references phases 1b-4b were removing, which is why it lands last. PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE Each site was judged against its ENGLISH counterpart, not on its own: README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46 English README.md has ZERO occurrences -> substituted "Antigravity CLI, Kimi CLI" how-to/execute-a-phase.md:88 x4 locales fixed in #4728 -> substitute how-to/verify-and-ship.md:89 x4 locales fixed in #4728 -> substitute FEATURES.md cross-AI CLI list :1419 no Gemini -> DELETE FEATURES.md REQ-MULTI-RT-01 :1709 -> substitute FEATURES.md REQ-SKILLS-03 :1952 -> rewrite FEATURES.md REQ-QUOTA-02 :3256 deleted upstream -> delete VERSIONING.md:133 stale manifest -> see below The twelve README occurrences were an adversarial reviewer's BLOCKER, and the reason they survived my own sweep is structural: root-level *.md was outside the guard's scan set, so the repo's most-read runtime-advertising surface was invisible to the guard meant to police it. :46 is a live installer-runtime claim - it tells the reader the installer will offer a runtime that no longer exists. Checked for the duplicate-name trap before substituting: neither Antigravity nor Kimi appears anywhere in those four files. Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work, not a deferral. Two more show why "substitute Gemini -> Antigravity" is the wrong default: in the cross-AI list and REQ-QUOTA-02 English DELETES the name, because Antigravity was already in the list or the classifier had dropped it. Substituting would have duplicated a name - the identical trap ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot. VERSIONING.md:133 is a different and worse defect than translation lag. Under "Manifest Version Sync" it listed gemini-extension.json as a version-synced manifest. That file is ABSENT from the repo, and scripts/sync-manifest-versions.cjs says so in its own comment - "#1928: gemini-extension.json was removed with the gemini runtime ... it is no longer a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json, marketplace.json and vscode/package.json. So the doc named a manifest that does not exist AND omitted the one that replaced it. Both fixed, verified against the owning code rather than inferred from the name. The replacement bullet cites #1942, the issue that actually registered vscode/package.json, matching the convention of its neighbours. pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has no REQ-QUOTA-02 line. Skipped and recorded, never invented. PART B - THE GUARD scripts/lint-retired-runtime-name.cjs, modelled on scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem shape (forbid a retired token, allowlist frozen content, self-exempt via a split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1). Case sensitivity IS the mechanism, not an accident. The naive guard - "the string gemini must not appear" - is WRONG, not merely noisy: that string is load-bearing across Antigravity's real on-disk contract. A case-sensitive, standalone, capitalised name works because every legitimate reference is spelled differently and therefore cannot match: lowercase config homes (~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and GEMINI.md. Table-driven, so the next retired runtime costs one row. THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two rounds of isolated adversarial review reshaped it; both are recorded in .gsd/bug/chore-4729-gemini-drift-guard/60-review.json. ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can never be complete, so "no runtime word found" silently exempted every phrasing nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`, `Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0, as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and `Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime` or `运行时`. The fix was to stop enumerating exceptions and invert the evidence direction: Tier 1(a) - the hook DIALECT Antigravity inherits. Position is language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The marker must now form an ADJACENT COMPOUND with the name, not merely sit in a +/-24-character window - that window let `| Antigravity | Gemini-style hooks | Gemini support is live |` exit 0, one legitimate reference licensing a fresh live claim 21 characters later. The runtime-word veto is now LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files" in the same sentence); each is an explicit pin rather than a reason to loosen the veto for everyone. Measured: widening it surfaced exactly those ten and no others. Tier 1(b) - the provider/model axis. A version optionally followed by a qualifier, including full-width digits and CJK punctuation, AND positive model-axis evidence on the line, AND no runtime word. The positive requirement is the part that matters: all eight real model-axis lines in the repo name a model explicitly, so requiring it costs nothing on the real tree while flagging every laundering attempt. It is also the honest resolution of the agent/target tension below - rather than guess at an exhaustive veto list, stop treating an empty veto as evidence. Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match falling INSIDE an occurrence of its own snippet. Line-level containment let `Known provider menu update: Gemini CLI is once again a selectable GSD runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your GSD runtime.` both exit 0, because a short snippet elsewhere on the line pre-approved a brand-new claim. Span scoping makes short snippets safe: `Google (Gemini)` can only ever excuse the match inside those 15 characters. A LOAD-TIME validator now requires every pin to contain a retired name, and it immediately caught five of MY OWN pins whose snippets sat BESIDE the name rather than covering it - each would have shipped permanently inert and permanently reported stale. All pins were then reconciled in one pass. A pin is also marked used by PRESENCE on the line now, rather than only on the Tier-2 branch. Previously a pinned line that a general rule also matched never marked its pin used, producing a provably FALSE "no line matches pinned snippet" whose printed remedy told the maintainer to delete a pin that was still needed. Tier 3 - blanket trust, and a new occurrence inside it IS invisible. CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one surface - plus six append-only directories. All 21 `.changeset/` hits were measured to be fragments DESCRIBING the retirement or a fix to it, 464 of them under archived/; a fragment can only describe what already shipped and is deleted at release, so pinning them would be friction with no signal. The cost is stated in the guard's own header rather than hidden. THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and `references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a live claim introduced there was invisible at BOTH ends. The escape hatch must now carry a justification (`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is checked first, excuses the whole line, and the failure message advertises it, so an unexplained one is indistinguishable from a silenced defect. Plus an anti-vacuity floor counting files actually READ, not files listed - a candidate count stays healthy-looking even if every read failed. A FALSE NEGATIVE I INTRODUCED, AND CLOSED The model-display escape began as a blanket /^ \d/ - "space then a digit" - which also matched "Install for Gemini 2.5 CLI as a supported runtime.", laundering a genuine stale-runtime claim through an attached version number. That was the THIRD appearance of one failure shape in this epic: an exclusion added to suppress false positives creating a false negative. #4716's sweep excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby hid a stale review.models.gemini row whose example value was "gemini-2.5-pro" ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which is why the fix this time was to invert the rule's evidence direction rather than to enumerate more exceptions. The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI" matched inside "client" and would have vetoed legitimate model lists - and raw for CJK terms, where \b is ASCII-word-based and would never fire beside an ideograph, so anchoring them would silently disable the veto in ja/ko/zh. "agent" and "target" were deliberately left OUT: both occur throughout ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so vetoing on them would red correct content instead of catching runtime claims. The reasoning is in the guard's comment, not just the omission - and Tier 1(b)'s positive-evidence requirement is what makes that omission safe, since the rule no longer depends on the veto list being complete. COVERAGE tests/lint-retired-runtime-name.test.cjs drives the guard through its GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a guard, and this epic already shipped one that was vacuous for 2 of its 5 files, so properties are paired against BOTH failure modes - too broad silently absorbs a future defect, too narrow reds on legitimate content. Floor boundaries are covered at 149/150/151. The round-2 reviewer's sharpest point was about that claim, and it was right: the first matrix's pairing was "true of the properties chosen, not of the predicate's actual surface" - not one of its twenty properties could see the dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an over-broad pin colliding with a new line. Every one of those is now a committed regression using the reviewer's own attack line verbatim, and the local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0). That harness earned a finding of its own. Its first run reported PASS=2 FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so nothing staged, every fixture hit the empty-walk error path, and the two checks that assert an ABSENCE passed off that error path rather than off real guard logic. A staging failure is now fatal and every absence-asserting check first proves the walk ran and the expected violation was flagged. Later, one case failed because its fixture supplied only one of a pinned file's two approved lines, so the stale-pin check fired correctly - the expectation was wrong, not the guard. Telling those two apart is the whole value of running a matrix rather than reasoning about one. On the two orthogonal reviews: the isolated adversarial pass executed a great deal of code, across two rounds, against its own fixture repos. The security pass did NOT - it self-discloses that it verified by reading only, because node --test is hard-blocked here. Saying so plainly, because "two orthogonal reviews" without that caveat overstates what the second one established. It also raised, and I cleared by measurement, a concern that importing escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt clone: six other tracked scripts already require that exact path, three of them already in lint:ci, and .github/workflows/test.yml:192-193 runs `npm run build:lib` immediately before it for exactly this reason. Part A has no new test deliberately - those edits are covered by the guard itself inside lint:ci, and a separate per-locale assertion would duplicate it and then drift from it. The one exception is the root README case, which IS pinned: that residue was invisible to the guard rather than merely unasserted, so the fix is a scan-set change and needs its own regression test. No mode-bit read-failure fixture was added on purpose: the benches run as root, where chmod-based IO injection is vacuous, so such a test would assert nothing. The test's fixture helpers write throwaway docs/ paths, which trips lint-docs-guard-registration's reader-name heuristic. Resolved the way that lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry - because the fixtures only WRITE scratch data and never read shipped docs; the baseline was re-confirmed, not merely extended, each time locale and adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own --write path, since a new test file changes the count lint:generated-sync reads. Fixes #4729 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4729): backfill changeset PR number (#4753) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
48271de43f |
fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis (#4744)
* test(#4660): pin the letter-axis parity defect across all 6 shell/markdown phase-id sites Extends tests/nsegment-phase-grammar.test.cjs (#4568) one axis over: for each of the six sites, reads the live regex off disk and asserts it agrees with src/phase-id.cts's PHASE_NUMBER_TOKEN_SOURCE on the letter axis in BOTH directions — accepts `12A` / `3A` / `03A` / `23A.1.2`, still rejects `3a`, `3AB`, `A3` and the other canonical-invalid shapes — and that the two extracting sites return the full letter-suffixed token rather than its digit prefix (or nothing). Negative control against the unfixed tree: 22 failures, exactly the "(fails before the fix)" cases; every reject-parity case already green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis Adds `[A-Z]?` after the leading digit run at all six sites #4568 widened — the ERE translation of src/phase-id.cts's `\d+[A-Z]?(?:\.\d+)*` — so a documented, canonical-valid id like `12A` or `23A.1.2` is no longer refused by the four validating sites (code-review.md, code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md) or truncated to its digit prefix by the two extracting sites (execute-plan.md's plan-filename grep, plan-phase.md's --research-phase capture). Behaviour is byte-identical for every id that matched before; the adjacent comment and error-message text now names the grammar it mirrors. Driven: `init code-review 3A` on a fixture with a `03A-slug/` directory and a `### Phase 3A:` heading emits `padded_phase: "03A"`, which the old regex rejects and the widened one accepts — nothing upstream of the validator mangles the id. At execute-plan.md the trailing `-[0-9]+` is the PLAN number and stays digit-only; plan and milestone dimensions are out of scope per the brief. `CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE` derives from the canonical source by a literal `.replaceAll('A-Z', 'A-Za-z')`, so src/phase-id.cts is deliberately untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4634): extend lint-phase-id-drift to ban a letter-less phase-id mirror in workflows/ and agents/ Adds findLetterlessPhaseMirrorDrift — the letter-axis twin of the #4568 single-segment rule — flagging the unbounded-segment shape `[0-9]+(\.[0-9]+)*` (and its \d / doubled-backslash near-variants) whose digit run is NOT followed by the `[A-Z]?` class, on any phase-carrying line across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md and agents/**/*.md. Sanctioned the same way (`<!-- phase-id-owner: ... -->`), tolerates the case-flexible `[A-Za-z]?` directory-scanning variant so it cannot force that separate axis to narrow, and is wired into scanAll. Confirmed zero violations against the real tree post-#4660 fix, and one violation when a single site is reverted. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * docs(#4660): add Fixed changeset Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore: regenerate conformance-tier manifests for the extended grammar test tests/nsegment-phase-grammar.test.cjs now requires the compiled gsd-core/bin/lib/phase-id.cjs (to assert the canonical grammar agrees with each site's live regex), which moves it to a different platform-conformance tier; `gen-platform-conformance-tier.cjs --check` in lint:ci flagged the macOS manifest as stale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * test(#4660): reword a comment that tripped lint-docs-guard-registration The comment mentioned `docs/CONFIGURATION.md` between two backticked tokens, which the lint's template-literal detector read as a docs/ path expression. The test reads no docs/ file. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): refresh the compact-content benchmark baseline and acknowledge emitted growth plan-phase.md grew by 4 bytes (`[A-Z]?`), which moves the committed compact-content benchmark; refreshed with `benchmark-compact-content.cjs --write`. The six shipped files below grew by the widened regex literal plus the comment and error-message text that now names the canonical grammar. Emitted-Drift-Ack-Growth: code-review.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: code-review-fix.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: gsd-code-fixer.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the defense-in-depth comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: execute-plan.md — #4660: `[A-Z]?` in the plan-filename phase extraction (6 bytes) Emitted-Drift-Ack-Growth: plan-phase.md — #4660: `[A-Z]?` in the --research-phase capture (6 bytes) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): set changeset fragment pr to 4744 * chore: re-trigger Validate Branch Name The required check-branch context was cancelled on this head by the workflow's cancel-in-progress group when the changeset pr-field backfill push landed three seconds after the PR opened; no completed run exists for the current head, and a fork contributor cannot re-run it. Empty commit to re-run it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
eb49ff98df |
fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime
#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.
The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:
- install-on-your-runtime.md English has NO `### Gemini CLI` section -> deleted
- USER-GUIDE.md :843 "…, Antigravity CLI, Kilo)" -> substituted
- ARCHITECTURE.md English has NO Gemini CLI table row -> row deleted
- ARCHITECTURE.md :24 English holds `Kimi CLI` in that slot -> Kimi CLI
- context-monitor.md :3 "`AfterTool` for Antigravity CLI" -> substituted
- spike-and-sketch.md :93 "(Codex, Antigravity CLI, etc.)" -> substituted
- configure-model-profiles "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
- COMMANDS.md English keeps only hyphen + Codex bullets -> colon bullet deleted
- FEATURES.md source docs/features/multi-runtime-support.md:10
lists no Gemini CLI -> name removed
ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.
The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.
The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.
Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.
Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.
Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.
Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.
Fixes #4728
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate
A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.
1. I committed the exact error I claimed to have avoided. The commit message
boasted that ARCHITECTURE.md:24 proved the value of reading English rather
than substituting blind, because Antigravity already appeared later in that
list. Five hundred lines further down the SAME four files, my
`Gemini:` -> `Antigravity:` substitution produced TWO consecutive
`- Antigravity:` bullets, because an Antigravity bullet was already there.
English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
locales, reusing each locale's existing words.
2. `--gemini` survived in the runtime-detection CLI flag list in all four
locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
already lists `--antigravity` later, so this is another place where
substituting Antigravity would have duplicated it. Now `--kimi`.
3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
line I had already corrected -- the "Adaptive (Recommended)" option in
settings.md:192 and new-project/steps/auto-mode-config.md:95.
4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
`non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
without the word "runtimes", and execute-phase.md:1028 has no parenthetical
at all. The reviewer proved it by re-introducing Gemini at both lines and
watching the assertion stay GREEN. That same blind spot is what hid finding 3.
Replaced with a case-sensitive `/\bGemini\b/` walk over every
`gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
reference in that tree is spelled differently and cannot match: Antigravity's
paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
`Gemini` there means the retired RUNTIME is being named. The walk asserts it
found at least 50 files so an empty walk cannot pass vacuously, and it now
covers the nested `new-project/steps/` directory where finding 3 lived.
Two allowlist entries, both by line CONTENT and both justified:
reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
`Known provider` menu. The second was escalated by the agent rather than
decided: Section 8 of that file says model policy is defined "independently"
of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
the same axis as the lowercase model ids -- and must keep working.
Proven to fail, not just asserted: the predicate reports 0 offenders on the
real tree and exactly 2 on a /tmp copy with Gemini re-injected at
health.md:52 and execute-phase.md:1028.
Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.
The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.
Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @
|
||
|
|
9c2927bff4 |
fix(#4733): derive the win32 chunk cap, isolation bar, and unknown-file weight (#4737)
* test(#4733): pin the cap, unknown-file weight, and isolation rules Failing-first coverage for the three defects that let a Windows conformance chunk be killed at the 600s per-chunk backstop with zero failing tests. The previous boundary rows were VACUOUS: they asserted literal arithmetic (21 * 18122 <= 400000) that cannot fail, and in doing so masked a shipped win32 cap of 23 -- a value that violates the very inequality they claimed to pin. These rows constrain defaultMaxFilesPerChunk itself, from both sides, so the shipped value is a derived maximum rather than a magic number. A second vacuous row was caught by review and removed: it recomputed the isolated set from the function under test using the identical predicate, so it was empty by construction. It is replaced by an exact deepEqual against the expected basenames, a cross-platform identity row, dynamism rows in both directions, an inclusive boundary triplet, and invalid-threshold throw rows. The cross-platform identity row is the regression guard for a threshold that was briefly anchored to the per-platform file-COUNT cap; it fails if isolation ever becomes platform-dependent again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4733): derive the win32 cap, isolation bar, and unknown weight A Windows conformance chunk was killed at the 600000ms per-chunk backstop with no test having failed, taking next red. Three compounding defects. The win32 cap of 40 permitted 40 * 18122 = 724880ms against a 600000ms backstop -- 121% of it -- so two rounds of budget-tuning could not hold. The cap is now derived: 22 is the largest value satisfying cap * 18122 <= 400000. The budget is 400000, not the raw backstop, because the chunk that died summed to only ~348328ms of per-file time -- a per-chunk overhead gap of at least 1.72x that no per-file table models. A file absent from the timings table was priced at medianWeight. The table is skewed 18.8x, so an unknown weighed 0.0533 -- 19x cheaper than average, and measured 17.5x under its real cost. Unknowns are now priced at the mean. ISOLATED_HEAVY_FILES was a static Set, stale by construction. Isolation is now derived from an absolute ms bar (0.3 * 400000 = 120000ms) converted to weight units via the live table's mean, so a file that gets heavy is isolated automatically instead of waiting for someone to edit a list. Review caught that an earlier cut anchored that bar to the per-platform file-COUNT cap -- a category error, count vs weight, which silently returned seven of the historical eight files to the shared pool on linux/darwin. Since macOS runs the full matrix only after merge, that would have planted a red next no PR could catch. The bar is absolute and platform-independent. Also from review: isolation no longer requires unit-suite membership, so fragment-single-edit-propagation.install.test.cjs -- 575000ms, 96% of the backstop in one file -- is eligible; partitionIsolatedFiles throws on a non-finite or non-positive threshold instead of silently isolating nothing; and stale per-shard figures no test pinned are removed rather than recomputed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4733): backfill changeset pr number --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ca8d9d4459 |
fix(#4429): stop a large commit_types config blocking or bypassing the gate (#4723)
* test(#4429): regression coverage for three defects in the commit hook Failing-first coverage. Every conforming-subject row is red against the unfixed hook, and each defect gets an explicit CONTROL row that reconstructs the pre-fix form and asserts the defect reproduces -- without those, the passing rows would pass with or without the fix. 1. SIGPIPE (the reported defect). The pre-fix first-line extraction used a `head -1` pipeline; once CONFIG_OUT exceeds the 64 KiB pipe buffer printf is killed and `set -euo pipefail` aborts the hook. That fix is already on next -- it landed incidentally in #4537, whose message never mentions #4429 -- and nothing in the tree would notice its removal. 2. regcomp. The commit-type alternation grew with the CONFIGURED list and exceeded bash's 64 KiB compiled-pattern cap. Boundary rows pin the cliff at 6051/6052, with controls on BOTH sides so limit-1 is not vacuous. 3. Ambient subprocess statuses (found by this change's security review). Defects 1 and 2 cannot be separated: each configured type adds len+1 bytes to CONFIG_OUT and len+1 to the alternation, so the smallest payload that overflows the pipe (N=6059) already puts the alternation past the ceiling. The SIGPIPE control accepts either SIGPIPE (141, Linux) or a reported write error (macOS bash 3.2's builtin printf, exit 1). Asserting only the message would go red on every CI lane, since the remote matrix is Linux-only. Named to bucket with gsd-validate-commit-crash-policy.test.cjs, which covers this same hook: lint-test-file-count derives a test's owning module from its filename prefix, and `validate-commit-*` collided with the `validate` module, already at its 2-file cap. Harness note, learned from three vacuous control runs: hooks/lib/git-cmd.js requires ../gsd-core/bin/lib/token-scanner.cjs relative to the hooks dir's parent, so a copy in a bare tmpdir fails open and returns 0 for any input. The layout symlinks gsd-core beside the copy, and every row that can prove it asserts the run was substantive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4429): bound the commit-type regex and isolate subprocess statuses Two fixes in the same file, both of the same shape: a value computed for one purpose was being read as authority about something else. 1. The commit-type alternation could not be compiled. COMMIT_TYPE_ALT joined every CONFIGURED type into one regex, so the pattern grew without bound. bash caps a compiled pattern at 64 KiB. Bisected on bash 3.2.57 (this repo's macOS target): a 65504-byte alternation compiles, 65515 fails. `[[ =~ ]]` returns 2 on a compile failure, and `if !` cannot tell that from "the subject does not conform" -- so the hook blocked a valid `feat(auth): ...` with CONVENTIONAL_COMMITS_VIOLATION while printing `feat` in its own valid_types. Match the shape with a fixed-size pattern, capture the type, then test membership against the COMMIT_TYPES array. The character class is exactly the `^[a-z][a-z0-9-]*$` safe-token filter the config loader already applies, so it captures every type that can legally reach COMMIT_TYPES and no token that cannot. Review verified equivalence over 46 handcrafted plus 6000 randomized adversarial subjects against a type list containing prefix-overlapping, digit-bearing and trailing-hyphen types: zero divergences. The loop adds no subprocess and no pipe, which is the hazard class #4429 is about. COMMIT_TYPE_ALT is now unused and removed. types pre-fix `feat(auth): ...` fixed 10 accept accept 6051 accept accept 6052 BLOCK accept 20000 BLOCK accept 2. Subprocess statuses were inherited from the environment. Each status is captured as `... || VAR=$?`, which assigns ONLY on the failure branch; on success the variable kept whatever it already held, and `${VAR:-0}` defaults only when unset or empty. So an EXPORTED CONFIG_STATUS, CMD_STATUS or CLASSIFY_STATUS -- from a CI wrapper, a .envrc, or another hook -- survived into the success path and was read as "the subprocess failed". Since the hook fails OPEN on a genuine subprocess failure by design (#3838), the result was a silent bypass. Measured: `CLASSIFY_STATUS=3 git commit -m "nope: bad"` printed "validator disabled for this call" and exited 0. The three are now initialised before use. The fail-open path is unchanged and verified byte-identical to origin/next with a failing node. hooks/dist/ is gitignored and rebuilt from hooks/ by scripts/build-hooks.js, so there is no second copy to sync. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): register the new suite with the conformance manifests Both conformance-tier manifests embed the test-file list, so adding a test file makes them stale. Regenerated with their own generators: node scripts/gen-platform-conformance-tier.cjs --write node scripts/gen-platform-conformance-tier.cjs --target macos --write Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): pin both fail-open-prone controls to their named cause Two rows in the ambient-status block asserted `status === 0`, which the hook also returns when the harness layout is broken -- so either row could have passed for entirely the wrong reason. This is the same vacuity trap the rest of the suite already guards, applied inconsistently to the rows added last. Measured, rather than reasoned about: genuine ambient bypass (pre-fix hook, CLASSIFY_STATUS=3) rc=0, no CLASSIFIER_THREW orphaned layout (no gsd-core symlink) rc=0, CLASSIFIER_THREW genuine fail-open (node shim exits 3) rc=0, no CLASSIFIER_THREW So assertSubstantive separates the intended cause from the harness failure in both rows, and each now pins its pass to the cause it names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): stop asserting a macOS-only regex cap on every platform First verification run was RED: 45636/45638 passed, both failures in this new suite on linux-node24. Cause is mine -- I measured the compiled-pattern ceiling on macOS and encoded it as a cross-platform expectation. Measured in the tester image itself: engine 6051 6052 20000 bash 3.2.57 / BSD libc (macOS) compiles rc 2 rc 2 bash 5.2.15 / glibc (Linux) compiles compiles compiles (228943 B) glibc has no reachable cap, so the regcomp defect cannot occur there and the control asserting a block at 6052 was red for a behaviour the platform cannot produce. The control now calibrates at runtime: it runs the pre-fix form and, when this engine compiled the alternation, it SKIPS with a message naming the reason rather than asserting. Skipped out loud, never silently passed -- a green row there would read as "the defect is covered" on a platform where it cannot occur. Both branches verified: the capped branch asserts (macOS 17/17, zero skipped), and the uncapped branch was exercised by forcing the payload to a size that always compiles, producing a skip and not a failure. Consequence stated rather than hidden: the remote matrix is Linux-only, so this one control is skipped in CI and really runs only on a macOS workstation. The rows that run everywhere are the ones carrying the regression weight -- the shipped hook accepting a conforming commit at every payload size, the gate still blocking unknown types, the SIGPIPE control, and all seven ambient-status rows. Note this also narrows the coupling claim: SIGPIPE and regcomp are coupled only on a capped engine. On glibc the SIGPIPE defect is directly testable without the regcomp fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4429): scope the regex-cap claim to the platform it applies to The changeset told users the validator "built a regular expression bigger than bash can compile" past ~6,000 configured types. That is false on Linux: glibc compiled a 228,943-byte alternation without complaint, so a Linux reader would have been misled about their own exposure. These are user-facing release notes, so the claim is now scoped to macOS (bash 3.2 / BSD libc) and says explicitly that glibc was never affected by this half. The hook's own comment led with the same overstatement -- "bash caps a compiled pattern at 64 KiB" -- before qualifying it. Reworded so the first clause states what is actually true: the limit is a property of the platform's regex engine. Text only; no behaviour change. Suite 17/17, eslint and lint:ci clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): backfill changeset PR number (#4723) * chore(#4429): backfill changeset PR number (#4723) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d1d9ee85a5 |
feat(#4142): thread convention through the completion-path membership seam (#3644)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
9c20b7b40a |
fix(#4652): use the validated path, collapse the duplication, correct two false claims
Seven findings from the two-axis review, all fixed in place. THE ONE THAT MATTERS: cmdTodoComplete validated sourcePath and targetPath and then ran every fs call against the RAW strings — existsSync, statSync, readFileSync, platformWriteSync, unlinkSync, and the dry-run path payload — never sourceCheck.resolved / targetCheck.resolved. That is the exact "validate one path, use another" shape ADR-4650 names as the defect this epic exists to prevent, and it is the same bug this phase had just fixed in check-command-router. Committed inside the fix for it. All I/O now uses the resolved paths; user-facing messages still echo the raw filename, never a resolved absolute path. A VACUOUS TEST, and the false doc claim it was propping up. The test "[RED #4327] an absolute path outside the project is rejected" would have passed with ZERO containment logic: path.join(pendingDir, '/abs/outside/x') yields <pendingDir>/abs/outside/x — Node does not let a later absolute segment escape — so the name is FOLDED under the root, passes containment, and simply 404s. The test only ever observed "Todo not found". It now asserts what is actually true and actually valuable: an absolute name is neutralized, and the real outside file is not read, not moved, and still present afterward. docs/CLI-TOOLS.md claimed such a path "is rejected as a usage error", which was false; it now describes the fold-under-root behavior. Traversal and embedded separators ARE rejected, and those claims stand. DUPLICATION THIS EPIC EXISTS TO REMOVE. resolvePath already did isAbsolute-or-join + validatePath + reject; cmdGapAnalysisPlanPost and cmdCheckPredicate each re-inlined the identical triplet in the same file. Both now call resolvePath. Cost, stated rather than hidden: its generic message replaces the two sites' distinct "phase-dir escapes…" wording. The message still names the offending input, and one predicate with one message is the point. SYMLINK COVERAGE was required by #4652's "Done when" and was missing. Added for both the todos root and --phase-dir, skipping cleanly on EPERM so the Windows lanes do not fail where unprivileged symlink creation is disallowed. Both fast-check properties were UNSEEDED. Seeded now. The changeset named "check decision-coverage-plan" as a boundary; that is a caller of the shared resolvePath, which the body never mentioned. Corrected. DISCLOSED, not hidden: ctx.phaseDir is now always the resolved ABSOLUTE path, so ${PHASE_DIR} interpolation and the "not found in <targetDir>" message show an absolute value where a relative --phase-dir previously produced a relative one. That is an observable output change. A test pins it and docs/reference/gate-predicates.md states it. Also regenerated scripts/lib/platform-conformance-tier.generated.cjs and its macos twin — the new tests changed check-predicate.test.cjs's tier classification. Caught by npm run lint:ci locally rather than by a bench run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
db4d8a9bae |
fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic $((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is valid shell-arithmetic syntax at all, and the failed expansion aborts the rest of the snippet in a non-interactive shell. safe_resume_gate runs unconditionally before trusting STATE.md or dispatching any executor, so execute-phase failed at its own gate before the first executor on any decimal phase, regardless of workflow.tdd_mode. Regression from #4194. Fixes all 4 sites: safe_resume_gate and the TDD gate in workflows/execute-phase.md, the completion-signal spot-check fallback in workflows/execute-phase/steps/completion-reconciliation.md, and the executor gate validation example in references/tdd.md. Each now zero-strips only the leading integer segment into a *_INT variable (via %%.* / # parameter expansion — always valid shell syntax regardless of what follows) and keeps the remainder as an escaped-dot string for the anchored commit- scope regex, exactly as issue #4619 verified in both bash and zsh. A plain integer phase (12, 01) computes byte-identically to before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug Behavioral coverage via real bash execution: the old $((10#01.1)) form throws (characterizes the bug, matching the issue's own reproduction); the new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):, feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring issue #4619's own verified table exactly. Updates safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions (one per site) to the new fixed text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic With #4619's fix in place, the guard's original "ban $((10#... outright, match any occurrence" was too blunt: it flagged a comment merely mentioning the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an already-%%.*-stripped integer, and the always-safe plan-id arithmetic (plan ids are plain integers, never decimal). Refines the detector to skip full-line comments and to only flag a captured variable/placeholder name that contains "phase" and does NOT end in _INT/_int — the naming convention the #4619 fix establishes at all four sites for "already reduced to a safe integer." A plan-id variable was never phase-number arithmetic in the first place and is excluded on the same basis. This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new exemptions") and D7 ("a decimal and N-segment phase id survive an end-to-end execute-phase selection without error") for real — the guard now reports zero violations across all five .cts/.md rules. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): cover the plain-padded-integer near-miss matrix too Review found the anchored-ERE near-miss coverage only exercised the decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also verifies the plain padded-integer case (01 -> PHASE_N=1) against its own near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the missing assertion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote only 2 backslash characters in JS source, which single-quoted-string parsing collapses to 1 real backslash at runtime -- but the workflow/reference files actually contain 2 raw backslash bytes at that position (needed so bash's ${var//pattern/replacement} produces the correct single-backslash output). Write 4 backslash characters in the JS source at all 4 occurrences so the runtime string matches the files' real bytes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): refresh the committed compact-content benchmark baseline The new PHASE_INT/PHASE_FRAC arithmetic lines added to gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): note the safe_resume_gate arithmetic growth in the test header The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846 bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading integer segment via base-10 arithmetic instead of forcing the whole value through $((10#...)) and hitting a hard shell syntax error on the first dot. A blank line previously separated the Emitted-Drift-Ack-Growth trailer from the Co-Authored-By trailer below it, which splits git's trailer-block detection: only the last contiguous non-blank run of Key: Value lines at the end of a commit message is recognized as trailers, so the growth ack was silently read as ordinary body text and the differential-attribution gate failed with the growth unacknowledged. Joining the two trailers into one contiguous block fixes it. Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4208): replace chmod-based restore-failure injection with a root-proof git shim `tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated an unwritable index via a `post-index-change` hook running `chmod a-w` on the git dir. That relies on the OS enforcing the *owner's own* permission bits against itself, which uid 0 (a routine identity inside this repo's Docker-based gsd-test benches) does not: every DAC check short-circuits true for root, so the write the chmod meant to block silently succeeds, the restore comes back clean, and the disclosure/rollback behavior under test never actually gets exercised. This is CLAUDE.md's own named anti-pattern for I/O-failure injection ("Cross-platform test IO-failure injection" — chmod tricks fail under root Docker/CI). It is confirmed as the actual root cause here, not a production defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure logic (added by #4253, merged just before this run) was hand-traced and manually reproduced end to end on an unprivileged workstation against a freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces exactly the `staging_failed` + "could not be restored" / "could NOT be restored during rollback" results both tests assert. The other `post-index-change`-based tests in this file (a `sleep` to force a timeout; a real `update-index` to flip a restored entry's mode) are unaffected because neither depends on a permission check — consistent with only the two chmod-based tests failing on the real remote run. Replaces the chmod fixture with a fake `git` placed ahead of the real one on PATH that fails only `update-index --add --cacheinfo` — the one call the restore makes — unconditionally, regardless of privilege level. Every other git invocation execs straight through to the real binary, so the rest of each scenario (`rm --cached`, the restore's own `ls-files` verification, etc.) is exercised exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): backfill changeset pr number to 4644 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI Passing the script as a `-c "<script>"` argv element made it subject to Windows' CreateProcess command-line argument encoding, which silently dropped the escaped-dot backslashes before bash ever saw them (observed on PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the same script via stdin instead removes argv entirely from the transport, so there is nothing for Windows to re-encode. POSIX behavior is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4cc2a466b5 |
fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec `cmdCommit`'s `--files` list can stage an addition but never a deletion: the #2014 guard skips a missing explicit entry because the filesystem cannot tell "moved away" from "not written yet". A caller that moves a file therefore had two forms, both wrong — a directory entry records the move but also commits every unrelated file in that directory (a concurrent session's in-flight todo, in the unattended execute-phase sweep), and a file entry leaves the old path's deletion dangling with the todo tracked at both paths. `--files-removed <paths>` is the caller-declared delete intent. Each entry names a file, or a directory whose tracked-but-absent files are the removals; those paths are staged with `git rm --cached` and join the commit pathspec. `--files` keeps its skip-if-missing contract untouched. A file entry still present on disk fails the commit closed with the existing staging-failure rollback; a never-tracked path is a no-op. `--files-removed` alone is a declared scope, not the unscoped .planning/ sweep. The dispatcher previously folded every non-flag token after `--files` into that list, so a second list flag could not exist; each list now runs from its flag to the next `--` token. The execute-phase todo sweep names the moved todos on both sides from CLOSED[@], and cleanup's archive commit moves .planning/phases/ and .planning/quick/ under --files-removed. Fixes #4208 Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries * chore(#4208): set changeset fragment pr to 4253 * fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard Three CI failures, all consequences of this PR's own change. 1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The three-line rationale comment plus the four-line array-building block added 318 bytes to a file that had only 141 of headroom on next. Move the rationale to docs/CLI-TOOLS.md -- which this PR already extends with the --files-removed contract, and which is where the ADR-857 gate wants call-site detail to live rather than in the host workflow -- and fold the array build onto one line. 93,577 -> 93,372. 2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415 guarantee to its old MECHANISM: it regex-matched the literal .planning/todos/{completed,pending}/ directory pathspecs in the commit --files list. This PR deliberately replaced those with named files (a directory entry also committed an unrelated todo a concurrent session dropped in mid-close), so the guard failed on a change it should have accepted. Re-point it at the new mechanism without weakening it: assert the ADDED array reaches --files, the REMOVED array reaches --files-removed, STATE.md is still committed, and -- newly -- that the two arrays are built from $COMPLETED_DIR and $PENDING_DIR respectively. Verified by negative control: deleting --files-removed "${REMOVED[@]}" from the workflow still fails the test, so the #2415 regression remains caught. Note for the merge queue: #4233 also grows execute-phase.md (+114). The two are additive -- different regions, no textual conflict -- so with both landed the file reaches ~93,486, over the 93,400 margin though under the 93,600 hard ceiling. Whichever merges second will need to reclaim ~86 bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv * fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin Rebasing onto next surfaced the byte-gate collision flagged earlier on this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354), so this PR's +113 landed at 93,467 against the <= 93,400 margin in tests/claude-orchestration.test.cjs. Compact the close_phase_todos step this PR already edits -- drop the PHASE_NUM indirection, fold the normaliser and the match guard, print the closed list with one printf, shorten the step's prose -- without touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no longer spends any of next's 46 bytes of headroom. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback Review of #4253 found three Majors with one root cause: the removal side judged presence by fs.lstatSync alone, where the addition side already reads `git ls-files -v` state. Absence from the worktree is not removal: - a submodule gitlink (mode 160000) whose directory was deleted by hand lists like a file and was `rm --cached` with no .gitmodules cleanup; - a skip-worktree path is never materialised by a cone-mode sparse checkout, so a directory entry over a sparse-excluded tree dropped that whole tree from the index; - an assume-unchanged path's worktree state is not something git itself consults; - an intent-to-add entry (`git add -N`) renders as a plain cached entry on the empty blob, yet nothing tracked exists to remove and no rollback can restore the flag. The index listing now carries each entry's `ls-files -v -s` tag, mode and stage. Only a plain cached (H), stage-0, non-gitlink entry is a removal candidate; every other state is left alone under a directory entry (exactly like a present file) and fails closed when named directly, with the state in the error. "Named directly" is decided on RESOLVED paths, not strings -- realpath of the longest existing prefix with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a symlinked spelling of the tree (macOS `/var` -> `/private/var`, where `process.cwd()` is the real path and the caller's absolute path is not -- CI on this round's first push) all resolve to the same entry, where a string compare against git's cwd-relative output silently took the directory polarity (pre-push review, driven; the symlink case is driven with an aliased fixture directory). The enumeration's domain is what `ls-files -v -s` can emit for an index entry, stated at the classifier. The third Major -- on an unborn HEAD a successful `rm --cached` was never rolled back when a later entry failed -- is fixed differently from the review's suggestion. Pushing the path into stagedPaths would put it on the commit pathspec, which a root commit refuses ("pathspec did not match", driven), and `git reset -- <path>` cannot restore an entry with no HEAD anyway. Instead every index entry this call removes is recorded (mode, blob) before the `rm` and put back with `update-index --cacheinfo` on rollback. That also restores a caller-pre-staged blob at a removed path exactly, where a reset would have silently replaced it with HEAD's version. The rollback is best-effort, as the addition-side reset already was, and the docs say so. Eight tests: gitlink under a directory entry, named directly, and named by absolute path; skip-worktree both forms; intent-to-add both forms; assume-unchanged named; unborn-HEAD partial failure restores the removal; pre-staged blob survives the rollback. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step Review nit on #4253: inserting the --files-removed rationale between the original bash block and its closing fence left an empty ```bash``` pair before </step>. Harmless at runtime, a formatting artifact of this PR's own diff; removed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): a boolean flag inside a commit path list no longer ends the list Review minor on #4253: collectList stopped at the next `--` token, so a positional wedged between a boolean flag and the next list flag (`--files a --amend b --files-removed c`) was claimed by neither list and silently dropped -- a regression in shape against the old slice-to-end parse, which filtered `--` tokens and kept `b`. No current call site interleaves that way, but the gap was real. A list now runs to the next LIST flag (`--files` / `--files-removed`) and skips boolean flags on the way, and a REPEATED list flag merges its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse did -- a first cut stopped at the repeat and dropped `b`, the same silent-drop shape one level over (pre-post comment audit). The only change #4208 makes to parsing is that a second list flag can exist. Tests: STATE.md wedged between --no-verify and --files-removed lands in the commit; both runs of a repeated --files reach it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * test(#4208): drive the reappearance window with a post-index-change hook Review nit on #4253: the defensive re-check for a file recreated between the absence test and `git rm --cached` -- the concurrent-session race this PR's own changeset names -- had no test. git fires post-index-change the moment `rm --cached` writes the index, so a hook that copies the file back exactly then exercises the window deterministically. The call reports staging_failed / "reappeared on disk", commits nothing, and the rollback restores the removed entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): restore a staged removal when the call records nothing A `git rm --cached` that succeeds mutates the index whether or not a commit follows. Only the staging-failure rollback put those entries back, so a call that reached `nothing_to_commit` reported no state change while the removal sat staged -- riding along on the caller's next commit. The review named the unborn-HEAD, removal-only shape. Keying on `headExists` would have fixed half of it: the guard also fires with a real HEAD when the removed path is index-only (added, never committed), because `diff HEAD` reads clean with the path absent on both sides. Both shapes now restore, at both `nothing_to_commit` exits. The failure exits are deliberately left alone -- they report a failure rather than no-change, and the addition side leaves its own staged paths there too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot `cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)` now owns the index-state classification, path canonicalisation and entry recording, returning the pathspec entries and the recorded removals its caller merges. Pure motion: no branch, message or probe changed. Only the two accumulators became local names, and `restoreRemovedEntries` stays with the caller because the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the extracted helper is 277. (Figures corrected after publication: an earlier version of this message said 854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit is 581, so this is above it, not below.) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): property-test the two-list commit parser RULESET.TESTS.property-based-testing asks a parser for at least one property test asserting a domain invariant; `collectList` had only hand-picked examples, one per shape a review round had already broken. Hoisted it to module scope as `collectListFlagValues` and exported it in the file's existing exported-for-tests convention -- a parser reachable only by spawning the CLI can be tested one example at a time and no faster. Three properties over generated argv: every positional lands in exactly the run open at it whatever the flag order or count; no positional after the first list flag is dropped or double-claimed; and with `--files-removed` absent the parse equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run ending at any `--` token, and a repeated list flag that does not merge -- each of which the properties catch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): pin cleanup.md's archive commit to --files-removed execute-phase.md's rewrite is pinned by the #2415 guard in this file; cleanup.md's equivalent was not, so reverting its routing would have been caught by nothing -- the mechanism's unit tests never read this file and pass either way. Asserts the two archived directories are under --files-removed and NOT under --files (where a directory entry sweeps in a concurrent session's in-flight writes), and that the destinations and STATE.md stay on the additive half. Controlled by restoring the pre-#4208 sweep, which fails it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): pin that a symlink to a directory is one tracked path Review of #4253 read the `lstatSync(...).isDirectory()` test as a symlink-following defect. Driving it says the opposite: git tracks the link as a single blob (mode 120000) and does not traverse it, so the tracked paths "under" it live at the real directory and were never named by the caller. Following the link would stage those -- the directory sweep #4208 exists to remove -- while the named entry still sat present on disk. Pinned rather than changed, with the premise driven in the test body. Swapping `lstatSync` for `statSync` -- the prescription as written -- fails it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit The base range added `tests/benchmark-compact-content.test.cjs` and a committed token baseline over the compacted workflows. This PR edits `gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on that entry and on the aggregate. Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is those two entries and nothing else. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): report a removal the call could not put back Round review of this round found the restore itself unchecked: the helper ignored `update-index`'s exit code, so a FAILED restore still reported `nothing_to_commit` -- the same false "no state changed" the restore exists to prevent, surviving one level down on the restore-failure path. It now returns a boolean. The two no-change exits report `staging_failed` naming the paths left staged; the staging-failure rollback still ignores it, deliberately, because it is already reporting a failure and an unwritable index is usually the failure being reported. Driven with a post-index-change hook that makes the git dir unwritable the moment `rm --cached` lands, so the restore cannot take its lock. Reverting both guards fails the test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): disclose a removal the rollback could not restore Round review refuted the reasoning behind leaving the rollback path's restore unchecked. The claim was that this exit is already reporting a failure, so the restore's result adds nothing. The counterexample is the ordinary case: the reported failure is usually a DIFFERENT cause -- a contradictory declaration, a reappeared path -- so a caller reading `failures` sees only that cause and learns nothing about the removal still sitting in its index. The rollback now appends a disclosure entry per un-restored removal, naming the path. The reason and `file` still report the failure that caused the rollback; the disclosure is additive. Also moves the restore-failure test's chmod into a `finally`: `t.after` runs AFTER the parent `afterEach`, so a throw before it left the fixture undeletable. Both driven; reverting the disclosure fails the new test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): decide index state by observation, never by an exit code The restore added two commits earlier keyed both its record decision and its success verdict on git's exit code. An exit code answers "did the command succeed", never "did the index change" -- execGit collapses a spawn timeout to a non-zero exit, and a killed git can already have written the index. Round review drove four failures from that one assumption, in both directions: - a failed `rm` still contributed an entry, so the rollback disclosed a removal that was never staged (stale index.lock); - a timed-out `rm` whose write DID land contributed none, so a real mutation was neither restored nor disclosed; - a timed-out `update-index` whose write landed reported failure, publishing a "could NOT be restored" disclosure that was false; - and the read-back that replaced it omitted `-z`, so core.quotePath rendered `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not restored -- the same quoting defect this PR already fixed for `preStaged`. Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the path: gone means this call owns the removal and records it; still there means nothing was staged; a probe that cannot answer becomes its own failure entry rather than an assumption. The restore verifies the same way, comparing the WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a path-only test accepts an entry that came back as something else. The verdict is three-valued -- `restored` / `not-restored` / `unverified` -- and the unverified wording says the restore could not be VERIFIED rather than that it failed. The rm's own failure is pushed ahead of any probe diagnostic so a timed-out removal keeps `timed_out: true` and its own message as the reported cause. Five regression cases, each negative-controlled against the shape it pins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): treat a declared removal path as a path, not a pathspec An index path handed back to git is parsed as a PATHSPEC, and the removal side handed several back. Three driven harms, all of them the sweep-in this flag exists to remove, arriving through the operand rather than through a directory entry: - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it removed `peer.md` and `stays.md` too, only the declared entry was recorded, so the rollback restored one of three and the other two rode out as staged deletions the result disclosed nowhere; - the same name reached `git commit -- <paths>`, which globbed and committed an undeclared `M peer.md` alongside the declared removal; - and the intent-to-add probe (`diff --cached` over the path) matched a STAGED PEER instead of itself, so an `add -N` entry was misclassified as ordinary content, removed, and restored by `--cacheinfo` -- which cannot restore the intent flag. It came back as a real staged addition. Every operand on this path is now `:(literal)`: the `rm`, both index probes, the intent-to-add probe, the restore read-back, the entry-level `ls-files` / `ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream `ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep whatever pathspec behaviour they have today; that is not this change's to alter. `:(literal)` still resolves a directory to its descendants (driven), so the directory form is unchanged. Closes what an earlier cut of this commit declared as a residual: a filename beginning with `:` is now removable end to end, because the commit pathspec no longer reinterprets it. Also fixes a MINOR from the same review: cleanup.md's contract test checked the destinations' position relative to `--files-removed` but never that `--files` was present at all, so deleting the flag still passed. Un-literalising the seven sites fails three of the new tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): scope the rollback to the caller's own name space Round review drove a rollback that destroyed the caller's own staged work. Two causes, one of them pre-existing: - `git diff --cached` prints REPO-relative paths whatever the cwd, while `stagedPaths` holds the caller's cwd-relative names. In a project nested inside its repo (`<repo>/sub/.planning/...`) the two name spaces never intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`, and the reset unstaged a caller-staged deletion and modification that this call had never touched. `--relative` makes the two sets comparable, and is a no-op when the project IS the repo root. This governs the `--files` side too and predates this flag. - the rollback's `reset` was the last place a removal-derived name reached git as a bare pathspec; it takes `asPathspec` like every other site. Driven on a nested fixture; dropping `--relative` fails the new test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): gate six fixtures that Windows cannot construct CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two primitives the new fixtures rely on do not exist on Windows, both driven on a real Windows host rather than inferred: - a filename containing `*` or `:` cannot be created at all (`IOException` / `FileNotFoundException`), which is four of the pathspec fixtures; - `chmod` cannot make a directory unwritable — a write into a ReadOnly directory succeeds — so the two restore-failure fixtures cannot drive the failure they exist to drive. Each is skipped on win32 with its measured reason, in the repo's existing `{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The behaviours they pin are platform-independent; only the fixtures are not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): build git's index-syntax path with forward slashes The remaining Windows red was mine, not the platform's: `git rev-parse :<path>` takes a forward-slash path, and `path.join` yields backslashes there, so git rejected it as an ambiguous argument. The hook in the same test already used the slash form. Not gated — the behaviour it pins is portable; only the argument was not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4208): refresh the compact-content baseline against the rebased base `next` moved the `new-project` split and the aggregate under this PR's execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs --write` so the only leaves differing from the base's copy are the execute-phase split and the aggregate it feeds. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh * chore(#4208): regenerate the macOS conformance tier for this PR's fixtures `next` gained the macOS-specific conformance tier (#4593) after this branch was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target macos`) now selects `tests/commit-files-deletion.test.cjs` on the `chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the chmod-driven failed-restore cases and the symlink-to-directory case). Regenerated with `--target macos --write`; the platform tier was already in sync. The file was modified, not added, which is why the added-files check did not surface it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1e47560e34 |
feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the same 546-file, Windows-oriented conformance-tier list as windows-latest -- built from signals like windows-shell-token/windows-env-var that have nothing to do with macOS. Issue #4593 asked for macOS coverage sized to its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin- specific behavior) instead. Issue #4593 was filed before Phase 5 (#4603) existed and referenced updating test-full's macOS legs -- that job is gone. Corrected the issue's body before any code was touched: the "shrink from full replay" half of the original ask was already done by Phase 5; what remained was narrowing the still-Windows-oriented tier macOS was inheriting. Two design assumptions were measured and rejected before accepting a design (documented in docs/adr/4593-macos-conformance-tier-architecture.md): - Reusing the general tier's signals minus its 3 Windows-specific categories barely narrows anything (546 -> 424, 78% retained) -- most files match multiple signals and only need one to survive exclusion. - A standalone CRLF/autocrlf signal, despite the issue naming "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit 143/930 files. Root cause: CRLF is primarily a Windows checkout concern in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST- PORTABILITY), so the signal was really re-selecting Windows-relevant files already covered by the general tier, not narrowing macOS specifically. Built 5 new, genuinely macOS-specific signals instead: darwin-literal (darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch, case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim from the general tier (genuinely Unix-relevant, not Windows-motivated). Measured against the real tree: 196 of 930 eligible unit-suite files (21%), versus the general tier's 546 (59%) -- a real, evidence-backed narrowing. scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/ classifyMacosTree/renderMacosGeneratedFile and a --target windows (default, unchanged)/--target macos CLI flag, so the same generator produces two independent, gated outputs rather than needing a second script. New committed output: scripts/lib/macos-conformance-tier. generated.cjs. .github/workflows/test.yml's test-conformance job: only the macos-latest leg's file-list source changes; windows-latest is byte-for-byte untouched. New shipped-file ripples handled proactively (19 install-tree fixtures regenerated, bin/install.js registered). An isolated code-review pass found one real defect: the ADR's per- category count table had drifted by 1 (zsh-dispatch, case-sensitivity) because the new test file's own fixture strings joined the tree it classifies after the table was authored -- fixed, with the union total (196, what CI actually gates on) confirmed unaffected. An isolated security-review pass found no qualifying findings. The ADR also records an explicit requirement for any future widening proposal: check whether the motivating regression is already covered by Phase 1's no-rendered-text-length-assert lint rule (#4590) before re-proposing full macOS/Linux parity, since that is exactly what #4421's root cause was (a rendered-text-length assertion, not a real behavioral divergence). Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |