01dbda9c499bc0e8b15ab54b970c500194b7067d
5940 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
01dbda9c49 |
docs(#4910): ADR-4910 — the PlanningDoc parse → mutate → serialize seam — Phase 0 of #4906 (#4911)
* docs(#4910): ADR-4910 — the PlanningDoc parse → mutate → serialize seam — Phase 0 of #4906 Design lock for epic #4906. Docs-only; no production code lands here. Eight decisions: one PlanningDoc seam composing markdown-sectionizer, markdown-table and frontmatter as layers; node-replacement writes so a field write cannot reach past its own value; byte-stable serialization for untouched regions; escape-or-refuse shared between each artifact's writer and its reader, with an explicit accepted-superset-of-emittable split; a typed parse error scoped to the node rather than the document; a type-narrowed write boundary paired with a lint; a positive control per accepted grammar; and one implementation per shared pattern. Two corrections to the epic's stated mechanism, both load-bearing for later phases: - The epic asks for a ratchet where a reintroduced content.replace() "does not typecheck". It cannot: fs.writeFileSync(p, s.replace(...)) typechecks fine because fs has never heard of PlanningDoc. Enforcement is type-narrowing plus a lint that owns the bypass, and the ADR says so rather than shipping a guarantee one require() defeats. - The epic names scripts/lint-planning-artifact-writer-drift.cjs as the drain point. That script is a registry-completeness guard that states "No ratchet / no baseline" by design and never inspects how a write is performed. It is correct on its own axis and left untouched; the right home is local/no-adhoc-markdown-parsing. ADR-2143's Phase 4 already shipped the table-regex and replace-mutation detectors, so the ADR scopes the remaining gap precisely rather than asking for that work twice: adhocReplaceMutation keys on a table-or-section regex, and #4852's pattern is a bold-label field regex — which is why the defect sits in src/phase.cts, a file the rule does lint, with lint:ci green. Per docs/adr/README.md lifecycle rule 3, ADR-1372 and ADR-2143 are NOT given the reciprocal `Subsumed by` back-link here: a Proposed ADR's Subsumes claim is prospective, so its targets are not marked until ratification. Both back-links land in the Phase 6 ratification PR. ADR index regenerated. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4910): apply review findings — link every ADR cross-reference and state the node-scoping interpretation Three review passes ran against the ADR: an isolated adversarial fact-check of every citation, and both axes of /code-review as separate sub-agents. Standards axis (hard violation of docs/adr/README.md lifecycle rule 2): 14 bare ADR-1372 / ADR-2143 / ADR-1411 references in body prose. gen-adr-index.cjs does not catch this — its bare-id check runs only over relation-field values, never body prose — so a green lint:generated-sync did not clear it. Every bare cross-reference is now a file link; only the self-reference ADR-4910 remains bare, which is not a cross-reference. Spec axis: §5 scopes the parse error to the node, while #4906's criterion reads "an unparseable shape surfaces could-not-parse with the offending span" with no document-or-node qualifier. Both readings are faithful to that sentence and they produce materially different Phase 1 and Phase 4 work. §5 now records the distributive reading explicitly, names the document-scoped alternative, states what it would cost (one bad table failing phase list and init.progress alongside roadmap.analyze), and names what changes if the epic meant the other one. A silent narrowing became a stated one. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4910): make each phase's acceptance a structural property, not a list of fixed issues Phases 2-5 read as "fixes #4852, #4862, #4499" — a point-fix list with a seam attached. That is the failure the epic names in its own words: "an implementation that does that has not closed this epic, even with every symptom gone and CI green." New section "What makes a phase done" locks three criteria every phase carries: - Census -> zero. A phase enumerates every instance of its anti-pattern in the tree, publishes the count in its PR, and closes when it is zero. Not "the reported ones". - Deletion, not coexistence. Bespoke implementations are removed, not kept in sync beside the seam — including src/roadmap.cts:1196's three-arm planCountPattern, which is correct today and still goes, because a correct copy of a rule the seam owns is the two-copies-that-agree case. - Unrepresentable by construction. Each phase ships one property that makes its class impossible rather than currently absent: a property over generated documents, a type that does not admit the wrong shape, or a drift guard. The absorbed issues are demoted to fail-first regression evidence. A phase may not close on those tests alone. Each phase restated accordingly, with its own census / deletion / unrepresentable / evidence breakdown. The six community point-fix PRs and their issues were closed unmerged (#4762/#4736, #4897/#4837, #4848/#4661, #4610/#4605, #4609/#4606, #4530/#4499). The ADR now records that as executed rather than pending, which is what lets census-to-zero be an acceptance criterion at all — a landed point fix would make the tree look healthier than it is. Phase PRs reference those six with Refs, since CONTRIBUTING forbids a closing keyword against an already-closed issue. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eea9247c93 |
enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select Auto-mode used to auto-select a checkpoint:decision's first <option> unconditionally, making a decision checkpoint's safety depend on option presentation order rather than an authored choice. Add an optional auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode now escalates to a human exactly like gate="blocking-human" does; present, it names the option auto-mode selects; naming an id with no matching <option id> is a hard structural-validation error at plan-parse time rather than a silent fallback to the first option. gate="blocking-human" continues to win over everything, unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys An isolated adversarial review of the auto_select work found that both new attribute regexes used \b as their left anchor, which is a word boundary, not a "start of attribute name" boundary. A decoy attribute ending in the same word (e.g. data-id="...") sitting before the real id="..." on the same <option> tag matched first, silently corrupting the extracted option id. Anchor on (?:^|\s) instead so only the real attribute name can match. Adds a regression test reproducing the exact decoy-attribute shape, plus a Unicode option-id test from the same review pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane lint-docs-guard-registration failed: the new test reads docs/reference/ plan-md.md but was not registered, so a future edit to that doc could silently desync from the test without the guard catching it on the PR that changed the doc. Registered alongside its direct precedents (precondition-element.test.cjs, reversibility-tagging.test.cjs), which read the same file for the same reason. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling The remote gsd-test run caught what local checks missed: execute-phase.md carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes of headroom before this change, and the original checkpoint:decision wording pushed it to 93772 (over the ceiling). Cascaded into failures in phase6-capstone-conformance, execute-phase-completion-reconciliation, claude-orchestration, and the compact-content drift-report test. Also caught: tests/package-legitimacy-gate.test.cjs anchors a "decision is conditional, not unconditional" safety check on the literal phrase "first option" in the decision bullet — which #4095 deliberately removes, since there is no more unconditional first-option pick. The test was asserting an assumption this change intentionally makes obsolete; re-anchored on tokens that still identify the bullet ('decision', 'auto-spawn') without weakening what the test actually verifies (the bullet must still carry a blocking-human carve-out). Also fixed a word-order mismatch between my own new test's regex and the actual doc text it was asserting against (tests/auto-select-attribute.test.cjs). Regenerated the compact-content benchmark baseline (tests/fixtures/compact-content-benchmark-baseline.json) to match the new byte counts. Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap. Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4095): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
029acd9158 |
fix(#4434): win32 unmeasured test files weigh the documented ~2.2x Windows-cost floor, not the Linux-measured mean (#4903)
* fix(#4434): win32 unmeasured test files weigh the documented ~2.2x Windows-cost floor, not the Linux-measured mean test(#4434): failing-first — win32 unmeasured-file weight must use the documented Windows-cost multiplier, not the plain table mean fix(#4434): makeFileWeigher takes a platform and applies WINDOWS_UNMEASURED_COST_MULTIPLIER (2.2) to the unmeasured-file fallback on win32 only; measured files and other platforms are unaffected chore(#4434): changeset fragment Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4434): add fast-check property coverage for the win32 unmeasured-file weight multiplier CLAUDE.md requires a fast-check property test for budget-limit logic; the prior example-based #4434 tests didn't satisfy that. Adds two seeded, bounded property tests: the unmeasured-file fallback matches the platform rule for any measured table, and a measured file's weight is platform-invariant for any measured table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4434): backfill changeset PR number (4903) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4434): the Windows unmeasured-file multiplier must not apply when there is no timings table at all Real Windows CI on this PR caught a regression my own linux-only gsd-test verification couldn't see: makeFileWeigher applied WINDOWS_UNMEASURED_COST_MULTIPLIER even when `timings` is null (missing/ corrupt/empty table), breaking the pre-#2456 "no table degrades to uniform weight 1" invariant several existing tests depend on. The multiplier now only applies to a file absent from an otherwise-loaded table — the actual #4434 mechanism (a real, loaded, Linux-measured table with unmeasured entries) — never to the no-table-at-all path. Six pre-existing tests that called makeFileWeigher/pack helpers with no explicit platform (silently inheriting whatever OS runs them) now pin an explicit 'linux' platform, since they test the platform-agnostic mean-vs- median and Object.prototype-safety invariants, not #4434's Windows behavior. One subprocess-level test is isolated from the real committed tests/test-timings.json via RUN_TESTS_TIMINGS_FILE, matching this file's existing isolation convention for cost-profile-sensitive assertions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ccfed63355 |
fix(#4823): the Current Plan reset is scoped to the Current Position section (#4898)
* test(#4823): failing-first — the Current Plan reset must not rewrite prose outside Current Position * fix(#4823): the Current Plan reset is scoped to the Current Position section — the whole-body 'Plan' fallback matched hard-wrapped prose lines starting with plan: * chore(#4823): changeset fragment * chore(#4823): backfill changeset PR number (4898) --------- Co-authored-by: sim <sim@local> |
||
|
|
bd79a97df0 |
fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error, never status missing (#4896)
* test(#4806): failing-first — unparseable VERIFICATION.md frontmatter is a parse error, not status missing / Field not found * fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error — never 'missing' or 'Field not found' * test(#4806): census pin 66→67 — cmdFrontmatterGet's unparseable-frontmatter error is a new output({error}) call site * chore(#4806): backfill changeset PR number (4896) --------- Co-authored-by: sim <sim@local> |
||
|
|
e2d879f681 |
fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse (#4895)
* test(#4802): failing-first — acknowledge must refuse an unparseable-frontmatter target instead of splicing over it * fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse — never splices the marker-marked object over the file * chore(#4802): backfill changeset PR number (4895) --------- Co-authored-by: sim <sim@local> |
||
|
|
a8394713d1 |
fix(#4801): init.manager resolves archived phase directories through findPhaseInternal (#4893)
* test(#4801): failing-first — an archived phase directory must resolve and report complete in init.manager * fix(#4801): init.manager resolves archived phase directories through findPhaseInternal The private current-milestone-only scan (matchPhaseDirs over listMilestonePhaseDirs) counted archived phase directories as missing, so an archived phase with a passed verification reported no_directory/phase_complete:false. findPhaseInternal — already imported, already the shared primitive for five other init commands — searches the live directory first and falls back through the workstream-scoped archive (#2855); the private copy and its single-consumer entries list are retired. * chore(#4801): backfill changeset PR number (4893) --------- Co-authored-by: sim <sim@local> |
||
|
|
88b5775dc8 |
enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI gsd-ui-auditor is chartered to audit interaction and handed a capture driver with no interaction verb: `npx playwright screenshot` cannot click, fill, hover, press or snapshot, so a hover state, an open menu, a focus ring or a form's validation state never appears in its evidence and every Experience Design finding degrades to code reading. Implements the shape approved at triage, not a new capability: - capabilities/ui/capability.json declares `workflow.ui_interaction_capture` (boolean, default false) on the capability that already owns the auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated. - gsd-core/workflows/ui-review.md reads the key through gsd_run and hands it to the auditor as `interaction_capture:` in the spawn <config> block — the auditor carries no gsd_run resolver, so the key travels by value. - agents/gsd-ui-auditor.md gains an anchored interaction-capture section AFTER the static block. With the key on and a Chrome binary resolved it starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on an --isolated profile, opens the dev URL the static block reached, takes the a11y snapshot for element uids, captures the baseline and a Tab focus-ring state, drives the UI-SPEC's interactive components, saves console output, and stops the daemon unconditionally. Key off, no dev server, or no Chrome: one status line, and the Playwright-only static path runs exactly as before — the static fence is untouched. Needs only Bash: no MCP server, no tools: change. Chromium-only by nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so readiness is polled through evaluate_script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): bind the interaction-capture shape and containment - manifest, generated registry, config schema and config-set/loadConfig all know workflow.ui_interaction_capture as a default-off boolean, and hand-written non-booleans fall to the slice default - the orchestrator reads the key and hands it down; the auditor never grows a gsd_run dependency - the static fence stays Playwright-only and the interaction fence chrome-devtools-only, so key-off is today's path - the interaction fence runs under bash with a stub driver on PATH: key off / absent / no dev server / no Chrome invoke nothing; the happy path starts first and stops last on the [selected] pageId with the documented flags; a failed capture is removed and not counted; new_page and start failures still honour the stop-only-if-started rule; CHROME_BIN and CHROME_DEVTOOLS_MCP_VERSION overrides flow through - docs/CONFIGURATION.md row shape; registered in the docs-guard lane Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * docs(#4223): document workflow.ui_interaction_capture and its how-to - docs/CONFIGURATION.md: one row in the workflow.* table, default-off - docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the interaction-capture section adds, skips and never claims - docs/how-to/enable-ui-interaction-capture.md: turn it on, read the `**Interaction captures:**` outcomes, what it does not do, turn it off - docs/README.md: index the how-to beside live-DOM verification Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): add changeset Added-type fragment; pr: carries the issue number until the PR exists. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace; the hyphen form is retired there and the slash-command-namespace guard rejects it. docs/ keep the hyphen form by convention. Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered. Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): per-run daemon session, bounded navigation, and step failures that count Three findings from the pre-file adversarial review of the interaction fence, folded in: - `--sessionId <epoch>-<pid>` on every driver call. `start` restarts whatever daemon shares its session and --isolated isolates only the browser profile, so two concurrent audits — or an audit beside the operator's own CLI daemon — would otherwise stop each other. The CLI accepts hex and dashes only; the id is validated by the test stub. - `new_page --timeout 30000`: the one verb that takes a bound, placed before every verb that does not, so a hung page is caught first. - a failed take_snapshot or press_key now increments the failure count and is named on stdout; two clean screenshots can no longer read as `0 failed` after the step that gives the interactions their uids failed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal Second review round, both reviewers: - session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a subshell, so two audits forked from one parent in the same second shared an id and could stop each other's daemon (driven by the reviewer) - `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver under Git Bash still matches the `$` anchor, and `|| true` on the assignment so a failed new_page cannot abort the block under `set -e -o pipefail` before the unconditional stop - a failed take_snapshot removes any snapshot.txt it left or inherited from a reused directory, so stale uids never drive the interactions - `<config>` placeholder is `{interaction_capture}`, lowercase like its `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt template, not a bash heredoc - how-to: the `not captured` row no longer claims the daemon started Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges Third review round: - a new_page that prints a page line and then exits non-zero is a failed navigation, not a page id: the exit status is checked in an `if` before the output is parsed (driven by the reviewer against the previous `|| true`, which masked exactly that) - regression cases for what the last two rounds added: CRLF driver output, a stale snapshot removed on failure, partial-output new_page failure, and the whole fence under `set -e -o pipefail` (both the failed-navigation path and the happy path) - the harness whitelist gains `date`; the session-id assertion now requires all three parts, so a silently empty epoch cannot hide again Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder - the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes, over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces; the interaction section's comments are tightened to the same content in fewer bytes (23559 now). No bash changed — the fence's own tests and the real-browser run are unchanged. - .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which the post-create backfill rewrites to the PR's own number. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): set changeset fragment pr to 4477 * test(#4223): compare the fence's status path with the separator the fence uses On the windows-latest lane the happy-path case failed on `\interaction` vs `/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a literal slash, and the assertion built its expectation with path.join. Every other case in the file passed on that lane, including the CRLF and errexit/pipefail ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): drop the inert file-header allow-test-rule marker Review round 1 on #4477: the `source-text-is-the-product` marker sat at line 2, outside no-source-grep's 8-line lookahead of every readFileSync site (the first is ~60 lines down), so it suppressed nothing. It was also unnecessary: every read in this file is a .md/.json path, which the rule does not trigger on. Deleted rather than relocated — there is no site to relocate it to. Negative control: `eslint` on the file is clean without it. * chore(#4223): regenerate the platform-conformance-tier lists for the new test Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier gate after this branch opened; its two committed lists must name every file under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never been in them. Once the branch was updated against next the lists were stale and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier --check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's fragment-single-edit-propagation, which sees the same staleness as regen:derived touching files beyond the fragment edit under test. Regenerated with the repo's own generators, no hand-editing. The general tier goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry; both --check arms are clean. Verified the red is this PR's own file and not base drift: at upstream/next both generators report "list matches" (546 / 196), and our committed copies were byte-identical to next's before this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ * fix(#4223): bound, confine and trap the chrome-devtools driver fence Round 4 — three findings in one fence, interleaved on the same lines, so one commit: - Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a background job in its own process group (`set -m`) under a watchdog that kills the whole group at the ceiling — TERM, then KILL two seconds later. One pid is not enough: npm forwards SIGTERM only to its direct child, so killing `npx` alone leaves the client holding the fence's stdout and a `$(cdt … new_page)` capture blocked past the ceiling (driven against a real npx tree by the round's adversarial review; the pid-only first cut of this commit had exactly that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )` subshell: a subshell inherits bash's saved copies of the caller's stdio (the fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a watchdog outlived its kill — measured as the intermittent 30 s test run the review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 -- -pgid`, every 0.1 s) and stands down by itself once the group is empty; nothing ever signals it. The group, not the leader pid: a child can outlive the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll stood down at once and left the substitution open-ended (driven by the round's adversarial review at 6× the ceiling; a pgid cannot be reused while any member lives, which a bare pid can). The daemon `start` launches is spawned detached (its own session) and never in that group. Two platforms forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM; sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git Bash a signal to a watchdog still starting up hung the fence's `wait` for it: 18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs, while a fence slowed by xtrace, or three tests run alone, never hit it (a startup race; the mechanism is not pinned further). Polling: 0/300 slow calls and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3 full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU, msys and busybox sleep; BSD sleep documents it). A clock that cannot launch (`sleep … || exit 0`) stands the watchdog down rather than firing at once and killing a healthy call — by design that leaves a hung call unbounded, the pre-round-4 behaviour, instead of failing a healthy one. A hung call returns once its group is gone: at the ceiling, plus up to the 2 s TERM-to-KILL grace. The KILL after the grace is sent only to a group that is still alive: a pgid freed during the grace can be reused, and an unconditional KILL could hit an unrelated group (the round's review). `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog rather than either. - --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may write under the run's interaction/ directory and nowhere else. Relative, like every --filePath (unchanged from rounds 1-3): the daemon resolves both against one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and path.resolve()s both), and a relative path needs no dialect translation — an absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native daemon cannot resolve (CI's windows conformance shard caught the first cut). --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified), so the documented floor moves from ^1.8.0 to ^1.9.0, where --allowUnrestrictedPaths is deprecated. - `stop` is owed by an EXIT trap after a successful `start`, not by position (it replaces any earlier EXIT trap — none exists in this file); the explicit call keeps it in order, a flag makes the trap a no-op afterwards, and only the shell that installed the trap may act: a subshell copy of the fence state carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced locally at 3/40 under load: the second `stop` came from a subshell pid, never main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist and CI's macos conformance job showed the guard comparing empty to empty. The fence was driven under bash 3.2.57 for the injected-subshell, errexit failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy paths. A failed resize_page is a counted failed step now, not the one bare command an errexit runner could abort on. Prose in the section is tightened to pay for the mechanism: 23559 -> 24517 bytes against the 24576 DEFAULT-tier cap. Tests: the stub driver hangs as a real child tree (sh waiting on a child that holds stdout — never an exec), so a pid-only kill fails the new aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative- controlled: it blocks for the harness's whole cap on the old wrapper). A hung start and a hung capture are cut off within ceiling + grace + slack and still reach stop; an injected bare failure under errexit reaches stop through the trap, exactly once; an injected subshell call of cdt_stop issues nothing; the happy path issues exactly one stop; every driver call site names a ceiling and the only bare $CDT is the wrapper's own spawn; the start line carries --workspace with the capture directory, every --filePath lies under it, and no code line carries --allowUnrestrictedPaths. A driver whose leader exits at once while a child keeps holding the capture pipe is still cut off at the ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone (negative-controlled: the trap form kills `start` in under 20 ms). The harness EXPORTS its stub-only PATH — unexported, the exec'd watchdog fell through to bash's compiled-in default PATH and never saw the stub dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash, pid-preserving). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the accessibility tree, with entered form values) and console.txt (which can carry tokens) were committable by `git add .`. The gate now ignores `interaction/` as a directory — the next artifact type is covered by construction — and it appends whatever an existing .gitignore lacks instead of writing once. The write-once form was the same defect one step later: every project that had already run an audit would never have received the new pattern at all. Tests run the gate fence under bash: a fresh file carries every pattern; an image-only file from an earlier audit gains interaction/ and keeps its own header without duplicating present lines; a second run appends nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * test(#4223): declare the interaction-capture anchor as a comment marker The #4324 colon-token gate (slash-command-namespace) landed on next after this branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an unconvertible /gsd: command token. It is a section anchor of the same family as gsd:live-dom-families and gsd:write-continue, so it is declared in COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added gates against the merged tree; CI at ca8d2508 predates the gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
0977d0a475 |
fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it (#4892)
* test(#4797): failing-first on Windows — a .cmd/.bat shim under a spaced path must run * fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it The private /d /s /c argv-array copy quoted the shim path (any space in the path) and cmd.exe /s stripped the first-and-last quote of the whole /c string, so the pre-space fragment became the program name: exit 1, empty stdout, on every repo whose path contains a space. The declared seam wraps the whole command line in one extra quote pair with windowsVerbatimArguments — the reporter verified the shape against the repro on Windows. * chore(#4797): backfill changeset PR number (4892) --------- Co-authored-by: sim <sim@local> |
||
|
|
2e14b4df17 |
fix(#4415): treat an absent worktree as removed, not as a branch mismatch (#4612)
* fix(#4415): treat an absent worktree as removed, not as a branch mismatch Claude Code removes a subagent's worktree the moment the subagent finishes with a clean tree. A gsd-executor that committed everything — SUMMARY.md included, under `commit_docs: true` — is exactly that case, so by the time the orchestrator reaches wave cleanup the directory is routinely gone while the branch it left behind is intact and mergeable. `git -C <gone> rev-parse --abbrev-ref HEAD` fails, and nothing distinguished that filesystem failure from a real branch disagreement: both reached the same `if`, so the entry blocked `branch_mismatch`, NOTHING merged, and the branch was left dangling. When the directory instead vanished after the merge landed, `git worktree remove` failed "is not a working tree" and the entry blocked `worktree_remove_failed`, leaving the branch undeleted and the operator to run `git worktree prune` + `git branch -D` + `rm -rf` by hand every wave. Disambiguated at the point of failure rather than ahead of it. A SUCCESSFUL in-worktree read still decides identity exactly as before — a present worktree on the wrong branch blocks, unchanged — and only a FAILED read consults the filesystem. Two reads can fail, and they are not the same path: * The branch read fails with the directory absent. There is no checkout for identity to come from, so it falls back to `refs/heads/<branch>` read from repoRoot; a missing ref still blocks, so an absent worktree never becomes a silent pass. The SUMMARY rescue and the dirty check are then skipped. * The branch read succeeded and the later `status` read fails with the directory now absent — the harness removed it while the repoRoot-side base, deletion and scope checks ran. Identity was already established from the checkout and the rescue has already run; only the dirty decision is skipped. Without this, a mid-entry removal still blocked `worktree_dirty` with nothing merged: the same bug, one window later. Skipping those reads is not a claim that the worktree was clean. This code cannot tell who removed the directory, and a forced or manual `rm -rf` of a DIRTY worktree would already have destroyed an uncommitted SUMMARY before cleanup ran. The narrow thing that is true either way is that a missing source cannot be read. The two reads also fail differently: the default SUMMARY finder catches the unreadable directory and returns no files, while `git -C <gone> status` errors — and that error is what surfaced as `worktree_dirty`. A rescue that genuinely FAILS still blocks, since a copy that errored part-way can mean an uncommitted SUMMARY was really lost. Teardown prunes the stale .git/worktrees admin entry rather than removing a path that is not there, re-reading presence instead of reusing the branch-step answer since the harness can act in between. For an entry accepted as ABSENT it prunes ONLY and never issues `worktree remove --force`: that entry was merged without the rescue and dirty checks, so force-removing a checkout recreated at that path would delete contents that never passed either one — strictly worse than the bug being fixed. A genuine prune failure still reports `worktree_remove_failed`, and a blocked teardown still withholds the branch delete. `git worktree prune` is repository-wide maintenance, not an entry-scoped operation. The presence probe resolves `worktree_path` against repoRoot, the way git does. `normalizeCleanupManifestEntry` takes the path from the manifest verbatim, so it can be relative, and every git call passes it as `-C <path>` with `cwd: plan.repoRoot`; a bare `fs.existsSync` would have resolved it against the PROCESS working directory instead. Those differ whenever cleanup runs from elsewhere, reachable today through gsd-tools' `--cwd` override, and the mismatch reads both ways: a present checkout reported absent — skipping the dirty check that would have blocked it — or an absent one reported present. An earlier cut resolved presence UP FRONT, before the branch read. That broke 52 existing tests: every cleanup-wave test uses a fake path that does not exist on disk and injects no `existsSync`, so all of them re-routed down the absent branch. Disambiguating at the point of failure leaves those tests reading as they did. Three rows still needed their premise stated — each stubs a git failure against a worktree that is genuinely present — and now inject `existsSync: () => true`. No assertion in any of the three changed. Fourteen rows added. Every early row held presence CONSTANT and so could not reach the windows that matter, since the bug is caused by a directory that changes state WHILE cleanup runs: removal after the branch read, a present worktree whose status fails (which must still block), removal between the clean status read and teardown, a reappeared checkout at teardown, #2852 isolation of a blocked absent entry from the entries after it, and relative-path resolution. Verified: ran the issue's own reproduction verbatim against a build of this branch — `merged_removed`, merge commit present, branch deleted, no prunable entry in `git worktree list`. The same reproduction against a build at the merge-base returns blocked/branch_mismatch, no merge, branch present, `wt1 ... prunable`. Five of the first eight rows go red against the true merge-base file; the three that stay green are the safety-preservation rows. The rows added after each review round go red against the commit that round reviewed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * chore(#4415): add changeset Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * fix(#4415): build the probe-path expectation with path.resolve, not path.join The row asserting that the presence probe resolves a relative `worktree_path` against repoRoot failed on windows-latest while the code under test was correct. On win32 `path.resolve` prepends the current drive to a drive-less absolute path (`/repo/main` -> `D:\repo\main`) and `path.join` does not, so a join-built expectation disagrees with correct behavior: expected: '\repo\main\.claude\worktrees\agent-a1' actual: 'D:\repo\main\.claude\worktrees\agent-a1' `path.resolve` is what the fix must use — it is how git resolves `-C <path>` against `cwd: plan.repoRoot` — so the expectation moves to resolve as well. Two `notEqual` rows keep that from being circular: the probe must receive neither the raw relative path nor a process-cwd resolution. Verified by mutation — dropping the repoRoot anchoring in `worktreeExists` turns the row red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * fix(#4415): confirm absence before skipping the rescue and dirty checks `fs.existsSync` answers false for a genuinely missing path AND for one it merely cannot traverse — EACCES on a parent directory, an unreachable mount. Verified: with a parent at mode 000, `existsSync` returns false while `statSync` throws EACCES. That distinction carries weight here, because "absent" is what lets an entry skip the SUMMARY rescue and the dirty check. An unreadable-but-present worktree read as absent, so cleanup merged over uncommitted work that the dirty check exists to refuse — and it contradicted this code's own comment that a present checkout whose git read fails stays blocked. Before this PR a failed git read blocked unconditionally, so treating unreadable as present is not a new safety rule; it is the one that was already there. The default probe becomes `statSync`, which reports WHY it failed. Only ENOENT is absence; anything else reads as present and blocks. An injected probe stays authoritative, so tests state presence directly with no hidden dependency on the real filesystem, and may throw to state that a path is unreadable. Two rows added: an unreadable worktree still blocks as branch_mismatch with no merge and no teardown, and a confirmed-ENOENT probe still takes the absent path. Verified by mutation — reverting the discrimination to the permissive `return false` turns the unreadable row RED while the ENOENT row stays green, which is what distinguishes discrimination from over-blocking. The mutation was confirmed to reach the compiled artifact the test loads. Also from this round: the row named for a checkout that "reappeared" never modeled reappearance (production probes presence once, at identification), so it is renamed to the unconditional contract it does prove; the comment crediting the notEqual rows with removing circularity is narrowed to what they actually establish; and the changeset now says only a confirmed absence takes the new path. Found by Codex full-PR review (round 3) before pushing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): source identity from git's registration, removal from the errno Maintainer review rejected the premise this fix rested on. It held that once the worktree directory is gone there is no checkout to read, so identity must fall back to `refs/heads/<branch>`. Git does not lose the binding — measured, after `rm -rf`: worktree /path/to/wt branch refs/heads/feat-x prunable gitdir file points to non-existent location The ref fallback weakened identity from "the checkout registered at this path is on this branch" to "a branch by this name exists", which let a foreign sibling branch merge. Identity now comes from `git worktree list --porcelain`, so the #3677 swap control keeps its teeth on the absent path; the new swap row is what would have caught this, and dropping the branch conjunct turns only that row red. Two defects in the first cut of the porcelain rework, both measured rather than reasoned about: `prunable` is not a removal test. With a parent directory at mode 000, git prints `prunable gitdir file points to non-existent location` for a checkout that is STILL THERE — it cannot traverse the parent, so it reports the gitdir file as missing. Treating prunable as "removed" would skip the rescue and dirty checks and merge over uncommitted work in an unreadable worktree, reintroducing the review's Major finding by another route. Each source now answers only what it can prove: porcelain for identity, `statSync`'s errno for removal. Only ENOENT is removal; EACCES/EIO blocks, as it did before this PR. `git worktree prune` is repository-wide. Measured: two removed worktrees plus ONE prune leaves neither registration behind. Reading the list per entry therefore let the first absent entry's teardown erase the identity evidence of every entry after it, merging one worktree per wave and blocking the rest as branch_mismatch — worse than the bug being fixed, since a wave of parallel executors is the normal case. The identity read is now a snapshot, captured lazily on the first entry that needs it and reused for the wave, which is both pre-prune and off the happy path. The `existsSync` probe and its dep locals are deleted; the filesystem is consulted only for the errno. The comment calling repository-wide prune "Harmless" was wrong under the new identity rule and says so now. Tests: identity and removal are stated on their own axes rather than through one present/absent boolean. Added the absent-path #3677 swap row, the two-absent-entry prune row, a bare `prunable` marker row, and a fail-safe row for an unreadable worktree list. Three mutations each kill exactly the intended rows, verified against the compiled artifact the tests load. One fixture that still stated presence through the removed `existsSync` seam was passing for the wrong reason and now states both axes. Verified: lint:ci exit 0; full suite 24/24 chunks, 37,164 tests, 0 failures; tests/worktree-safety.test.cjs 422/422. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): re-confirm absence before teardown, and prove the porcelain claim against real git Maintainer review, Major. Presence was classified once, at identification, and everything between that point and teardown — the base, deletion and scope gates, and the merge itself — is a window in which a worktree can reappear. The defence was "prune only, and a live checkout would make `branch -D` fail visibly", which holds only while prune's own staleness check is not fooled by the same filesystem-visibility gap that produced the false absence one call earlier. If it is, prune clears the admin entry, `branch -D` then SUCCEEDS, and a live, unreviewed, un-rescued worktree loses its branch. That asymmetry is the argument for the fix: the bug this PR set out to repair only ever BLOCKED, while this path could DESTROY state. Absence is now re-confirmed with `confirmedGone()` immediately before teardown — no new subprocess, just the statSync already in hand — and a reappeared directory blocks as `worktree_remove_failed` instead of reaching prune or the branch delete. The review was also right that the gap was known and unverified: the existing row said so in its own comment ("it does NOT model the reappearance transition itself"). It is modelled now, by a stat that answers "gone" at identification and "present" at teardown. Mutation-verified: removing the re-confirmation turns ONLY the new row red while the old "prune, never force-remove" row stays green, which is exactly why that row could not have caught this. Minor, same review: the #4415 block was entirely mock-based, so the factual claim the identity mechanism rests on was asserted in comments and measured out of band but never proved executably. Two real-git rows now prove it — that git keeps the path -> branch binding after the checkout is deleted and marks the entry prunable, and that it ALSO reports prunable for an unreadable worktree that is still there, which is why removal is confirmed by errno rather than by prunable. The second row skips as root, where mode 000 does not deny traversal. Minor 2 (rescueSummaryArtifacts resolving worktree_path against process.cwd() while the new code resolves against plan.repoRoot) is pre-existing and not reachable through the CLI's same-cwd invocation; left for a follow-up issue rather than widened into this PR. Verified: lint:ci exit 0; full suite 27/27 chunks, 37,739 tests, 0 failures, against the true merge-base; tests/worktree-safety.test.cjs 425/425. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * test(#4415): make the real-git rows platform-correct The Windows conformance shard caught both rows on their first push, and both failures were mine, not the code's. Path separators: git reports porcelain paths with FORWARD slashes on every platform, while `path.join` yields backslashes on win32, so `includes()` compared separator styles rather than paths and the registration assertions failed. Both sides are normalised before comparison now. Premise setup: the unreadable-worktree row establishes "git cannot traverse the parent" with mode 000, which win32 does not honour for directory traversal at all — the row would have asserted `prunable` against a perfectly readable worktree and failed for a reason unrelated to the behaviour under test. It now skips on win32 for the same reason it already skipped as root, with both reasons stated together. Verified: lint:ci exit 0; tests/worktree-safety.test.cjs 425/425 locally. The Windows shard is the real check for the separator fix, since macOS cannot reproduce it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): warn when an entry is accepted as absent, giving prunable its consumer Maintainer review round 3, both Medium findings — they close together, as the review noted. The absent path reported `merged_removed`/`ok` indistinguishably from an ordinary merge. This code cannot tell "the harness cleanly removed a finished executor" from "an operator or an external process removed this path": git keeps the path -> branch registration and `statSync` reports ENOENT in both cases. Before this path existed every anomalous absence blocked loudly, so accepting the routine case silently took the operator's only signal away from the case that is not routine. The module already carries an advisory channel for a materially less risky condition — scope conformance, a few lines below — so withholding one here was inconsistent with its own pattern. `WAVE_CLEANUP_WARNING.ACCEPTED_ABSENT_WORKTREE` is now emitted at both acceptance sites, carrying git's own `prunable` reason. Advisory, never a gate: the entry still merges. That also gives `WorktreeEntry.prunable` a consumer. It was parsed, documented as "worth surfacing to an operator", and then never read — the errno rework made it unused for the predicate and the parsing stayed behind. Quoting git's reason here is what it was for. The bare-marker test was vacuous, as the review said: it asserted `merged_removed`, which is driven by `confirmedGone` and the branch match, not by the bare-marker parsing it claimed to cover, so a regression in that parsing would not have reddened it. It now asserts the parsed value reaches the warning. A bare `prunable` line normalises to the literal 'prunable' — a truthiness signal, not a reason — so the warning reports null there rather than quoting a marker back at an operator as though git had said something. `WAVE_CLEANUP_WARNING`'s locked code set is updated deliberately, with the reason recorded in the test: the lock exists so a new advisory code is a decision rather than something that appears because a branch needed one. Verified: mutation — suppressing the warning at both sites turns both new rows red; lint:ci exit 0; full suite 27/27 chunks, 38,245 tests, 0 failures; tests/worktree-safety.test.cjs 426/426. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
822934c901 |
fix(#4794): decision-coverage answers an unmeasured shape on could-not-parse; a non-file context path fails closed (#4889)
* test(#4794): failing-first — could-not-parse must answer an unmeasured shape; a directory context path fails closed * fix(#4794): could-not-parse answers an unmeasured shape (null counts, unreadable ids, no uncovered); a non-file context path fails closed * chore(#4794): backfill changeset PR number (4889) * test(#4794): skip the directory-identity probe when the platform cannot discriminate (windows runner volume collapse, measured) Two consecutive windows conformance runs failed the probe with measured identical (dev, ino) for two distinct mkdtemp directories (dev=3606225537, ino=9007199255243448 for both) — a runner-volume property, not a regression in the guard. On such a platform the guard's identity containment degrades to refuse-everything (fail-closed, documented); the probe asserts capability, so the honest response is an explicit t.skip carrying the measurement (ADR-2719 §6), not a red lane for every PR. --------- Co-authored-by: sim <sim@local> |
||
|
|
3d2cb1fb01 |
test(#4850): read freshness under the fixture git timeout in derivationIsNotMemoizedAcrossRenders (#4870)
The row asserts that deriveStateFreshness is not memoized across renders, and it did so through the hook's real git spawn bounded by the 1500 ms production timeout. On a loaded Windows runner the second spawn can exceed that bound, and readStateHeadCommits returns its designed null, which the exact-count assertion reads as a failure (null !== 10). Wrap the two freshness reads in the block's existing withSpawnSpy, forwarding to the real execFileSync with the fixture-scoped GIT_FIXTURE_TIMEOUT_MS in place of the production bound. The spawn stays real, so a genuine memoization regression (second read returning 5) still fails the row. No production file changes; STATE_FRESHNESS_GIT_TIMEOUT_MS stays 1500. Closes #4850 Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8a5166598c |
fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next (#4873)
* fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next Commit |
||
|
|
09e1110e68 |
fix(#4788): code spans in the decision bold lead-in are opaque to the separator grammar (#4883)
* test(#4788): failing-first — code spans in the decision bold lead-in are opaque to the separator grammar * fix(#4788): the decision lead-in runs are code-span-aware — a backticked span is data, never grammar * chore(#4788): backfill changeset PR number (4881) * chore(#4788): correct changeset PR number (4883, was a guessed 4881) --------- Co-authored-by: sim <sim@local> |
||
|
|
5906a24ede |
fix(#4786): plan-row detection accepts the bare planId stem — suffix-less hand-written lists tick in place (#4880)
* test(#4786): failing-first — a suffix-less hand-written plan list is ticked in place, never duplicated * fix(#4786): plan-row detection accepts the bare planId stem — a suffix-less hand-written list is ticked in place, never duplicated * chore(#4786): backfill changeset PR number (4880) --------- Co-authored-by: sim <sim@local> |
||
|
|
6dcc0428dd |
fix(#4784): Xcode gates pass -project and derive the destination from available simulators (#4879)
* test(#4784): failing-first — Xcode gates must pass -project and derive the destination from available simulators * fix(#4784): Xcode gates pass -project '$XCODEPROJ' and resolve the destination from available simulators, skipping loudly when none exist Also surfaces the -collect-test-diagnostics never / workflow.test_gate_timeout guidance on a 124 timeout (a real iOS suite measured ~600s of sysdiagnose collection after the tests passed). * test(#4784): reachability-honest test shapes — apostrophe-confused prose (the path the false positive actually reaches past the splitter) and the file's single-quote convention * fix(#4784): review fold-ins — unanchor the simulator extraction (real simctl lines end in a state suffix), escape apostrophes for the bash -c layer, assign the command variables on the skip path, behavioral extraction test The isolated reviewer's Finding 1 was critical: the first cut's $-anchored sed matched ZERO real simctl lines (every device line ends in (Shutdown)/(Booted)), so the gates would have skipped on every machine. The 46,454-test green could not see it (content assertions do not execute the pipeline); the new behavioral test runs the gate's own extraction against a real-format line. * chore(#4784): backfill changeset PR number (4879) --------- Co-authored-by: sim <sim@local> |
||
|
|
d435723c95 |
fix(#4782): claude's agents kind skips compact variants — consumed only by the non-claude persona gate (#4878)
* test(#4782): failing-first — claude install must not stage compact agent variants (and must prune stale ones) * fix(#4782): claude's agents kind skips compact variants — they are consumed only by the non-claude persona-fallback gate Emitted-Drift-Ack-Hash: agents/gsd-advisor-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ai-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-assumptions-analyzer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-code-fixer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-code-reviewer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-codebase-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-debug-session-manager.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-classifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-writer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-dom-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-domain-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-eval-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-eval-planner.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-framework-selector.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-integration-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-intel-updater.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-mempalace-curator.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-nyquist-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-pattern-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-project-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-research-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-roadmapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-security-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-user-profiler.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies * chore(#4782): changeset fragment * chore(#4782): backfill changeset PR number (4878) --------- Co-authored-by: sim <sim@local> |
||
|
|
4fe2837ab5 |
fix(#4774): plan-criteria R4 requires the pipe not to be doubled — a logical-OR fallback is handled, not swallowed (#4877)
* test(#4774): failing-first — R4 must not read a logical-OR fallback as a pipeline stage * fix(#4774): R4 requires the pipe not to be doubled — a logical-OR fallback is a handled failure, not a swallowed one Also corrects the rows-3/4 test's makeCriteriaPlan usage (second arg is the <verify> block, not a second criteria line). * chore(#4774): backfill changeset PR number (4877) --------- Co-authored-by: sim <sim@local> |
||
|
|
a87b83d485 |
fix(#4764): dep_phases extracts only Phase-prefixed references from Depends-on prose (#4876)
* test(#4764): failing-first — dep_phases must extract only Phase-prefixed references, never dates/shas/ledger ids/self * fix(#4764): dep_phases anchors phase references to their 'Phase' prose context and never emits the row's own number * fix(#4764): review fold-ins — hoist the anchored dep-reference grammar to phase-id, cover Oxford lists and hyphen ranges, repair the property test Adversarial review found: Oxford-comma lists under-extracted ('Phases 1, 2, and 3' dropped the tail member — a silent real-blocker clear, the dangerous direction); hyphen ranges ('Phases 1-3') kept only the first endpoint; the property test called fc.hexaString (absent in fast-check 4.8, threw every run) and passed the junk arbitrary unspread (vacuous guard) with no completeness assertion; planning-inspect's extractDependencyTokens carried the same whole-field scrape (generative-fix divergence). The anchored grammar now lives beside PHASE_NUMBER_TOKEN_SOURCE in phase-id.cts and both readers interpolate it. * chore(#4764): backfill changeset PR number (4876) --------- Co-authored-by: sim <sim@local> |
||
|
|
969456c46d |
fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path (#4875)
* test(#4759): failing-first — preserved-foreign hooks warning must not claim may-not-load for a commonjs package.json * fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path * test(#4759): review fold-ins — assert installer output on stdout+stderr (warn is stderr), register cleanup before the spawn, add the type-less foreign case * chore(#4759): changeset fragment * chore(#4759): backfill changeset PR number (4875) --------- Co-authored-by: sim <sim@local> |
||
|
|
d36514b816 |
fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() (#4872)
* test(#4758): failing-first — rescue must resolve a relative worktree_path against repoRoot * fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() * test(#4758): review fold-ins — t.after cleanup pattern, post-resolution reader contract comment * chore(#4758): changeset fragment * chore(#4758): backfill changeset PR number (4872) * test(#4758): windows lanes key rescue fakes on resolved path identity, not verbatim strings win32 path.resolve rewrites driveless-absolute POSIX-style fixture values to the current drive, so the rescue's (correct) resolved-path handoff stopped matching verbatim string keys: #3804/#245/#2556/B7/#2852 fakes silently skipped the rescue and my seam test compared against a POSIX literal. Fakes now key on path.resolve(repoRoot, …) identity — the same semantics the code and git -C use — so every rescue test exercises the rescue on every platform. --------- Co-authored-by: sim <sim@local> |
||
|
|
63edc777e6 |
fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading (#4868)
* test(#4588): observed fork-from-HEAD must suppress the stale-origin degrade (failing first) * fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading * fix(#4588): a throwing probe git call is an inconclusive observation, not a crash * test(#4588): name the observed reason in the inconclusive-row assertions * test(#4588): hermetic state I/O in the inconclusive-row fixtures * chore(#4588): changeset fragment * test(#4588): contained cache I/O, emit-payload assertions (review fold-ins) * fix(#4588): review fold-ins — invalidation fixture, hermetic confirm row, gate comment, cache+scope rows * chore(#4588): backfill changeset PR number (4868) --------- Co-authored-by: sim <sim@local> |
||
|
|
9a41a95212 |
fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams * fix(#4717): consult the per-install runtime marker at both identity seams resolveReportedRuntime (agent_runtime) and loadConfigResolved (config.runtime) both ignored the per-install .gsd-runtime marker that resolveRuntime and the model-resolver gate already read. On a multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing misreported every Claude Code session as codex, and a shared defaults.json stamped by the first non-Claude install leaked its runtime to every other one. Seam 1: the reported-runtime ladder becomes explicit > install marker > host detection > claude. Seam 2: loadConfigResolved fills an empty config.runtime from GSD_RUNTIME then the marker, copy-on-write (the builtin-defaults branch returns a shared object). Explicit runtimes and marker-less trees are unchanged. * fix(#4717): a marker-detected runtime opts into its tier map (decision a) * fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins * chore(#4717): backfill changeset PR number (4861) --------- Co-authored-by: sim <sim@local> |
||
|
|
58c7bbb16a |
fix(#4667): rewrite codex @ includes to the codex install root (#4858)
* test(#4667): add failing-first coverage for the codex @-include rewrite Behavioral end-to-end: a real in-process install(true,'codex') into a temp CODEX_HOME must leave zero @~/.claude includes in GSD-owned .md artifacts, rewrite the issue's own example include to @~/.codex/, keep the deliberate _GSD_RUNTIME_ROOT .claude fallback chains byte-identical, and stay idempotent across a reinstall (no doubled prefix). All four are RED until the installer grows the manifest-scoped rewrite pass. * fix(#4667): rewrite codex @ includes to the codex install root Codex-installed agents and commands kept @~/.claude/gsd-core/... (and @/Users/trekkie/.claude/gsd-core/...) include references pointing into the Claude install: silent wrong-copy reads on dual-runtime machines at divergent versions, missing files on codex-only ones. Several emitters bypass the per-runtime converters, so the per-emitter fixes since #570 rotted. Adds a manifest-scoped rewrite pass in install() beside the leak scanner: for codex, every manifest-tracked .md/.toml artifact has the @-include forms rewritten to the codex root. The pass matches the exact include literal only — the _GSD_RUNTIME_ROOT/$PREFERRED_CONFIG_DIR fallback chains, prose .claude mentions, and CHANGELOG.md are untouched — and runs before the scanner, which remains the verification backstop for anything a future emitter introduces. * test(#4667): cover the HOME-anchored include form and sync the pass comment Adds behavioral coverage for the second rewrite literal (@$HOME/.claude/gsd-core/ -> @$HOME/.codex/gsd-core/) via the plan-review-convergence command, and corrects the pass's comment: the agent .tomls are generated after it and prefix themselves, so the .toml branch of the rewrite is inert by design. * test(#4667): baseline the offline upgrade against the deployed tree (sanctioned) * chore(#4667): backfill changeset PR number (4858) * test(#4667): sandbox the install home in the include-rewrite tests (#3712 guard) * test(#4667): install via subprocess with an isolated env in the include-rewrite tests --------- Co-authored-by: sim <sim@local> |
||
|
|
e1f72cd324 |
fix(#4741): the plan checkbox tick respects the superseded exclusion (#4851)
* test(#4741): a superseded plan must not be ticked from its summary (failing first) * fix(#4741): the plan checkbox tick respects the superseded-plan exclusion * fix(#4741): review fold-ins — changeset typo, dedupe planId derivation * test(#4741): exercise a halted summary on the active plan in the #2830 pin * chore(#4741): backfill changeset PR number (4851) --------- Co-authored-by: sim <sim@local> |
||
|
|
11b3091df0 |
fix(#4738): record opencode's staged skills in the install manifest (#4847)
* test(#4738): opencode manifest must record its staged skills (failing first) * fix(#4738): record opencode's staged skills in the install manifest * test(#4738): use the centralized temp-dir helper in the manifest tests * fix(#4738): retire the dead hostBehaviors vocabulary entry, tighten detector asserts, temper changeset * chore(#4738): backfill changeset PR number (4847) --------- Co-authored-by: sim <sim@local> |
||
|
|
c5629bbe74 |
fix(#4734): degrade worktree isolation when the root has no git repository (#4843)
* test(#4734): non-git root must degrade worktree isolation (failing first) * fix(#4734): degrade worktree isolation when the root has no git repository * fix(#4734): review fold-ins — 3972 ladder fixture, parity fixture, docs row, message wording * chore(#4734): backfill changeset PR number (4843) --------- Co-authored-by: sim <sim@local> |
||
|
|
8d0b6868ae |
fix(#4725): write normalization preserves tight paragraph-list shape (#4842)
* test(#4725): write normalization must not reflow untouched prose (failing first) * fix(#4725): stop write normalization injecting a blank before a list after prose * test(#4725): repair ordered-list fixture and list-spacing snapshot * test(#4725): assert whole-file prose stability, fix heading-list comment * chore(#4725): backfill changeset PR number (4842) --------- Co-authored-by: sim <sim@local> |
||
|
|
c9a5cc3e12 |
fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head. |
||
|
|
bff99a8bb5 |
fix(#4731): read hard-wrapped Goal/Requirements fields past the line break (#4826)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; isolated adversarial review round completed (MEDIUM table-bleed finding fixed with RED/GREEN evidence) and sha-pinned bench 46331/0 on the merged head. |
||
|
|
fb3e228a0d |
fix(#4724): classify Surefire/Failsafe XML as RED evidence (#4825)
* test(#4724): add failing-first coverage for Surefire XML RED evidence * fix(#4724): classify Surefire/Failsafe XML as RED evidence check tdd-red-evidence parsed only node:test TAP, so a JVM project's genuine Maven red scored INVALID_RED while hand-written synthetic TAP scored RED_EVIDENCE_OK — the gate was passable only by fabricating its input (issue #4724's measured repro). classifyRedEvidence detects Surefire/Failsafe XML (a <testsuite> element) and parses it by TAG-BOUNDARY scanning: each <testcase> owns its own tag (self-closing) or the segment up to its </testcase> closer, so the issue's warned-about spanning trap (a lazy lazy match from a green self-closing case to the next closing tag) cannot misreport names. A <failure> or <error> child marks the case failing; the target matches at class granularity (exact classname, dotted-suffix, or method name). Any parse anomaly degrades to not-failing — the module stays fail-closed and PURE (no fs/clock; report freshness remains the workflow's run-start check per the issue's implementation notes). TAP classification is byte-identical: all existing fixtures stay green. * test(#4724): pin the scanner hardening — truncation, TAP-message flip, CDATA phantom * docs(#4724): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
7d0c6339d0 |
fix(#4705): emit Antigravity-native tool names as a YAML sequence (#4822)
* test(#4705): add failing-first coverage for native Antigravity tool sequences * fix(#4705): emit Antigravity-native tool names as a YAML sequence convertClaudeAgentToAntigravityAgent and the installer's twin emitted Gemini CLI tool names as a comma-separated scalar. Antigravity's documented subagent contract (antigravity.google/docs/subagents) wants a YAML sequence of native names — view_file, grep_search, run_command, replace_file_content are the documented examples, and wrong or malformed grants can hang the subagent per Antigravity's own warning. Map values move to the native vocabulary where documented (Read -> view_file, Edit -> replace_file_content, Bash -> run_command, Grep -> grep_search); undocumented entries keep their best-known grant rather than being dropped (dropping would silently remove a restriction). The emitter writes one '- name' item per line; an agent whose every tool was filtered emits an explicit tools: [] instead of an empty scalar. Pre-existing pins updated to the native vocabulary. * test(#4705): update the #4727 map-value pin to the Antigravity-native vocabulary The #4727-era pin held the map VALUES at the Gemini CLI dialect on the belief that Antigravity speaks it; the confirmed bug #4705 (with Antigravity's own documented subagent contract) supersedes that for the four documented names. Key/shape pinning is preserved; only the values move. * docs(#4705): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
be1b76dddd |
fix(#4700): queue the headless mempalace mine on the palace lock (#4821)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase * fix(#4699): skip already-complete phases in the next_phase cascade Both next-phase scans selected the numerically lowest phase above N without consulting completion state, so completing a reopened phase persisted an already-[x] phase as STATE.md current_phase while roadmap.analyze correctly named the outstanding one (issue repro: completing 2 with phases 1 and 3 already [x] returned next_phase 03). The cascade collects the complete phase numbers from the roadmap checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in both the disk scan and the roadmap scan; a [x] checkbox row and its heading sibling both name a phase that is never next. Heading-only and checkbox-less roadmaps behave exactly as before. * test(#4699): align the negative-control expectation with the disk spelling * test(#4699): pin the STATE.md persistence and the all-later-complete tail corner Review findings: the regression never asserted STATE.md current_phase (the issue's actual harm), and the all-later-phases-[x] corner (is_last_phase true, next_phase null) was unpinned. A changeset fragment is included. * docs(#4699): backfill changeset PR number * test(#4700): add failing-first coverage for the queued headless mine * fix(#4700): queue the headless mempalace mine and surface skipped captures The capture's mine ran in the foreground with no lock handling: MemPalace wraps every mine in a per-palace lock, so any concurrent writer (two phases finishing a stage at once, a git-hook refresh mining the same palace) made it exit 1 (MineAlreadyRunning) and the onError: skip step silently dropped the capture — unlost for CONTEXT/PLAN/SUMMARY files that can be re-filed, unrecoverable for execute:wave:post problem-fix pairs. The mine now queues via --daemon --background (MemPalace #2029: the daemon holds a job refused the lock and runs it when the holder exits), and the report step gains the queued and skipped outcomes per #4700's requirement that a skipped capture never stay silent. Option 2 (write_routing.cli) is unreleased at MemPalace 3.9.0; option 3 (retry) re-enters the same lock race — both declined in the PR body. * fix(#4700): queue the wave:post problems fragment's headless mine too The issue names the execute:wave:post problem-fix pair as the unrecoverable loss (no source file to re-file later); the capture-problems fragment's headless mine ran foreground like the capture capability's did. Same fix: --daemon --background, with the lock-deferral rationale inline. * docs(#4700): backfill changeset PR number * fix(#4682): register the stale-reverification part in the capability registry The new steps/ part is a shipped workflow file; gen-capability-registry --check requires it in the committed registry. --------- Co-authored-by: sim <sim@local> |
||
|
|
d707318e0c |
fix(#4699): skip already-complete phases in the next_phase cascade (#4820)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase * fix(#4699): skip already-complete phases in the next_phase cascade Both next-phase scans selected the numerically lowest phase above N without consulting completion state, so completing a reopened phase persisted an already-[x] phase as STATE.md current_phase while roadmap.analyze correctly named the outstanding one (issue repro: completing 2 with phases 1 and 3 already [x] returned next_phase 03). The cascade collects the complete phase numbers from the roadmap checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in both the disk scan and the roadmap scan; a [x] checkbox row and its heading sibling both name a phase that is never next. Heading-only and checkbox-less roadmaps behave exactly as before. * test(#4699): align the negative-control expectation with the disk spelling * test(#4699): pin the STATE.md persistence and the all-later-complete tail corner Review findings: the regression never asserted STATE.md current_phase (the issue's actual harm), and the all-later-phases-[x] corner (is_last_phase true, next_phase null) was unpinned. A changeset fragment is included. * docs(#4699): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
2bfff17ff8 |
fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing * fix(#4682): route stale verification to the verifier regeneration path The stale routing entry sent users to /gsd-verify-work — but verify-work never rewrites VERIFICATION.md (its only write is the human_needed canonicalization), so following the advice re-ran UAT, reached the same stale check, and looped. init's projector and execute-phase's generic next_command presentation both mirror this entry, so the dead end appeared on three surfaces. The stale entry now routes to execute-phase, and execute-phase's all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping the spine under its frozen ADR-857 ceiling) mirroring the missing route: skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier — regenerating VERIFICATION.md and its digest, marked phase or not. The non-stale fall-through. Staleness detection, the digest format (#4623), every other routing entry, and the #3684 resume arms are untouched. Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682) Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682) * test(#4682): register the stale-reverification part and align projected commands The new steps/ part must be registered in the inventory manifest and the per-runtime golden install trees (regen:derived); the projected stale next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the runtime surface), the human_needed bare-report probe keeps routing to verify-work (unchanged semantics), and init-manager's recommended action follows the new command. * test(#4682): prefix the remaining stale routing assertions with the runtime surface Nine stale next_command assertions and the human_needed bare-report probe still carried the unprefixed or flipped forms from the earlier line-number edit; all now assert the shipped /gsd-execute-phase <phase> projection, with the human_needed probe reverted to its unchanged verify-work routing. * test(#4682): align the last stale projection assertions with the execute-phase route * docs(#4682): backfill changeset PR number * test(#4682): refresh the compact-content baseline after the rebase The rebase onto the #4670 squash brought verify-work.md's bounded reconciliation text into this branch; the committed compact-content baseline now reflects the post-rebase split sizes. Local --check is clean; the previous bench drift (+243) was the baseline, not the diff. * fix(#4682): carry the response_language directive in the stale-reverification part The new steps/ part is its own coverage unit for lint-response-language-coverage; it takes the shared canonical directive line like its sibling execute-phase parts. --------- Co-authored-by: sim <sim@local> |
||
|
|
651511d1e3 |
fix(#4670): bound the commit-claim window to the plan's own history (#4813)
* test(#4670): add failing-first coverage for the bounded commit-claim window * fix(#4670): bound the commit-claim window to the plan's own history The reconciliation measured plan_head_before..HEAD — a window that grows with every later plan's task and SUMMARY commits plus execute-phase's own phase-completion commit — so an honest plan flagged commit_claim_mismatch as soon as anything landed after it (real project: claims 3/5/2 measured 20/10/5). The executor now also records plan_head_after (HEAD at its measurement moment, after the last task commit, before the SUMMARY commit), and verify-work reconciles exactly against plan_head_before..plan_head_after with a merge-base ancestry check; SUMMARYs without the anchor fall back to the legacy warning path instead of an unsound BLOCKER. Both #3968 failure modes (claimed commits never made; task commits lost) still block, driven by the issue's own fixture scenarios. Emitted-Drift-Ack-Growth: verify-work.md — reconciliation gains the bounded window and legacy fallback (#4670) Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670) * fix(#4670): name the history-rewrite case and keep the executor under its cap Review findings: the BLOCKER enumeration named only the two #3968 causes, so an honest plan whose recorded window was rewritten afterwards (rebase, amend, cherry-pick) got a mislabeled diagnosis — the clause now names that case with the manual-recount remedy. The executor's growth crossed the LARGE-tier hard cap (49152), so the plan_head_after documentation is compressed to the minimal capture + frontmatter write (verify-work.md carries the semantics), the #2751 PROSE_ALLOWLIST entry is re-pointed at the shifted line (#4670 moved it from 823 to 825), the compact-content baseline is regenerated, and the changeset records the two un-established edges (shared-base waves, subrepo ledgers). Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670) * docs(#4670): backfill changeset PR number * docs(#4670): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
652796e903 |
fix(#4665): route --fix past the empty-scope exit (#4810)
* test(#4665): add failing-first contract coverage for the --fix empty-scope recovery * fix(#4665): route --fix past the empty-scope exit check_empty_scope exited the entire workflow whenever REVIEW_FILES was empty — before dispatch-fix — so with #3661's incremental scoping, a phase whose only post-review changes were planning artifacts could never run --fix against its standing REVIEW.md findings, and the skip output did not even mention the flag. The skip is now a self-contained guarded fence (explicit REVIEW_FILES emptiness check): it fires only when --fix is absent OR the phase's REVIEW.md does not exist. Otherwise the workflow proceeds directly to dispatch-fix, which delegates to code-review-fix.md — the canonical fix implementation that already documents handling an existing REVIEW.md — while the fresh-review steps (structural pre-pass, reviewer lanes, spawn_reviewer, commit_review) are skipped: nothing new to review, nothing to commit. dispatch-fix.md's route docstring is synced. Emitted-Drift-Ack-Growth: code-review.md — check_empty_scope gains the --fix recovery branch (#4665) * docs(#4665): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
85545a77a5 |
fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate verify-work.md's complete_session step flips VERIFICATION.md to passed on 'zero issues' alone, so a session whose every UAT row is blocked (a session that observed nothing) canonicalizes the report. Pins the deployed contract the fix must satisfy: the flip runs the unflagged phase uat-passed predicate inside the human_needed branch, frontmatter.set sits inside a passed==true guard, a refusal message carries the blocker count and keeps human_needed, and an indeterminate pre-check fails closed. All four new assertions are RED until the workflow grows the guard. * fix(#4663): gate the canonicalization on the uat-passed predicate complete_session flipped VERIFICATION.md to passed whenever the session recorded zero issues and the status was human_needed — but blocked rows are not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked session (one that observed nothing) rewrote the canonical report to passed. Every later reader (transition.md's preliminary check, resume paths, validate-phase, verification.status) then inherited the unearned pass while the phase-close predicate correctly refused it. The flip now runs the phase-close predicate in a new --uat-only form before canonicalizing: UAT rows evaluated (at least one pass, no pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers skipped — they must be, because the report still reads human_needed at pre-check time and that status is itself a blocking verification entry, so the full predicate could never pass there and the flip would deadlock (found by isolated review, probed). The --require-verification call stays the transition gate; the refusal branch reports the blocker count and keeps human_needed; an indeterminate pre-check fails closed. Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * test(#4663): align the canonicalize pre-check needles with the shipped line The workflow line carries a 2>/dev/null redirect the needles did not include, so both pre-check assertions fail against the committed fix (fixed-string grep verified). Reviewer-found; needle and message aligned. * fix(#4663): reword the canonicalize prose and refresh its size baseline The rationale paragraph mentioned the flagged transition-gate call by its flag, putting a --require-verification literal before the first phase uat-passed occurrence and breaking the existing ordering pin; the prose now describes it without the literal. verify-work.md's growth also drifted the committed compact-content baseline; regenerated via benchmark-compact-content.cjs --write (derived artifact, report-not-gate contract). Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * docs(#4663): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
c72fb34e9a |
fix(#4658): give the ui plan gate's evidence check a native branch (#4807)
* test(#4658): add failing-first coverage for native frontend evidence hasStaticFrontendEvidence recognised only JS-ecosystem evidence, so computeUiPlanGate could never block for a SwiftUI/Compose/Flutter/XAML project. Adds evidence-level fixtures for the four suggested markers (import-matched for .swift/.kt/.dart, extension-alone for .xaml), the reporter's non-UI Swift control case, marker-exactness and SKIP_DIRS and I/O-degrade negatives, gate-level block assertions through makeProject's new native frontendEvidence modes, and a pinned-seed fast-check property. All new assertions are RED until src/ui-frontend-evidence.cts grows the native branch. * fix(#4658): give the ui plan gate's evidence check a native branch hasStaticFrontendEvidence recognised only JS-ecosystem evidence (a root package.json UI-framework dep, or a .tsx/.jsx/.vue/.svelte file), so computeUiPlanGate could never block for a SwiftUI, Jetpack Compose, Flutter, or .NET MAUI project — the #3312 gate was structurally unreachable for them. Adds a native BFS over the same bounds and skip rules: .xaml is evidence by extension alone (the .tsx analogue), while .swift/.kt/.dart count only when their content carries the ecosystem's UI import marker (import SwiftUI / import UIKit, androidx.compose, package:flutter) — matched on the import, not the extension, so a non-UI Swift package stays silent exactly as the issue's 37-file control case requires. Marker reads are bounded to a 64 KiB prefix; any I/O failure degrades to false per the module contract. The #3718 vocabulary filter, the JS evidence rules, and the weaker-extension exclusion are untouched. * chore(#4658): regenerate the macos conformance tier list The native-evidence additions to tests/check-ui-plan-gate.test.cjs move the file into the macOS conformance tier per the classifier; the committed generated list is a derived artifact and must match the live tests/ tree (the same sync the fragment-single-edit-propagation install test enforces). * fix(#4658): accept both Dart quote styles and extract the shared bounded walk The isolated reviews' remaining findings: the Dart marker carried only the single-quote anchor, missing legal double-quoted imports (a spec-narrowing deviation); the two evidence walks duplicated the subtle MAX_WALK_ENTRIES cap semantics verbatim, so they are extracted into one walkProjectFiles BFS with a visit callback; tests now use the createTempDir helper, shared fixture literals that cannot drift from NATIVE_UI_CONTENT_MARKERS, a double-quoted Flutter import case, and drop a vacuous assertion and a mid-body re-require alias. * docs(#4658): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
5e729445d3 |
fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first proposeConsiderations regression (Danish text + English text_en must classify as its English equivalent, not land in the #1110 unclassified sentinel), proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en validation cases (empty/whitespace/non-string, unconditional under an elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a reference-doc Inputs parity test, and a fast-check property proving any cue-matching prose classifies identically under a cue-free Danish rendering plus text_en. All new assertions are RED until src/ui-consideration-probe.cts and the workflow/reference docs are updated. * fix(#4657): give the ui consideration probe a text_en language channel Element gains an optional text_en; classifyElement's own signature stays untouched (a locked, directly-tested export) and the text_en ?? text selection is pushed to the two classification call sites (proposeConsiderations, proposeElements) instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero kinds. Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the Non-English projects section (mirroring spec-phase Step 5.5) and the ELEMENTS_JSON shape comment documents the field with both zero-applicable guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT predicate (with both derived indexes regenerated), and the nav-override test expectation stay in sync. The ui-phase contract test carries the site-scoped allow-test-rule marker and its cluster is registered in the test-file-count allowlist ratchet. Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657) * docs(#4657): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
3014775a3f |
fix(#4656): expose coverage.unclassified and widen the zero-applicable guards (#4800)
* fix(#4656): expose coverage.unclassified and widen the zero-applicable guards * fix(#4656): regenerate golden coverage fixtures and update the rollup pin Emitted-Drift-Ack-Growth: spec-phase.md — #4656: guard widened to the all-unclassified case, doc claim corrected Emitted-Drift-Ack-Growth: ui-phase.md — #4656: guard widened identically * fix(#4656): sync edge-probe doc blocks and coverage pins with the new field * fix(#4656): key the mandatory confirmation on the widened guard * docs(#4656): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
caecaec62e |
fix(#4648): delegate explore seeds to the plant-seed workflow (#4798)
* fix(#4648): delegate explore seeds to the plant-seed workflow * fix(#4648): delegate explore seeds to the plant-seed workflow Emitted-Drift-Ack-Growth: explore.md — #4648 consumer wiring: the seed output now delegates to /gsd:capture --seed (plant-seed) instead of hand-writing a divergent, reader-invisible shape * fix(#4648): plant-seed extracts an idea-stated trigger into trigger_when Emitted-Drift-Ack-Growth: plant-seed.md — #4648: write-seed sets trigger_when from the idea text when /gsd-explore passes the conversation trigger inline * docs(#4648): backfill changeset PR number * docs(#4648): correct changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
f0a1745e29 |
fix(#4639): exempt the container --env-file value from the secret-read guard (#4789)
* fix(#4639): exempt the container --env-file value from the secret-read guard * test(#4639): adopt the suite assertions and pin the reclassified row * docs(#4639): document and pin the container-env printenv residual * docs(#4639): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
003d982c83 |
fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint Two defects in the covered-input fingerprint (#4155), one issue. 1. `computeCoveredDigest` hashed the whole bytes of every declared path uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` — which every phase rewrites as ordinary bookkeeping, and which the closing phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER the verifier ran — flipped every phase that declared them to `stale` on zero implementation change, and from there `isPhaseComplete` → `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate. Fingerprint v2 leaves any direct child of a planning root out of the hash: `.planning/` itself, plus the phase's own planning root (the parent of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/` layouts are covered without the digest knowing what a workstream is — `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather than a name list so the set cannot drift; a root is accepted only when the phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated exactly as every other covered path (confined, present, a regular file — the fail-closed contract is unchanged); only its bytes are ignored, and a declaration made only of shared documents fails closed like an empty one. A stored digest names its version, and `readVerificationStatus` now recomputes under THAT version (`parseFingerprintVersion`, `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics until it is re-fingerprinted, so the upgrade alone stales nothing; a version this build cannot recompute fails closed. 2. `verification.fingerprint` received a raw positional slice, so `--files a`, `--files "a,b"` and `--files a --files b` all put the literal token into the covered set and failed closed as "a covered file is missing, unreadable, or escapes the project root" — the message that convinced the reporting project the digest was permanently unrecomputable. `parseFingerprintFileArgs` accepts every form (plus `--files=a,b`, freely mixed with bare positionals), treats any other `--flag` and an empty `--files` value as usage errors that say so, and the phase-dir argument must now be an existing directory: omitting it used to take the first covered file as the phase dir and print a plausible digest over the rest at exit 0. Regression tests (tests/verification-status.test.cjs, #4623 block): the cross-phase case from the report, the same-phase `requirements mark-complete` / `phase.complete` cases from the thread, a workstream-scoped root, v1-preserved / unknown-version-stale, the fail-closed cases (missing, directory, escaping symlink, all-shared), every `--files` form against the bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34 fail, the 7 that pass pin behaviour the fix must leave unchanged. Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's covered_files instruction (rewritten in place — the file sits 21 bytes under its LARGE hard cap), gsd-core/templates/verification-report.md. Fixes #4623 Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi * chore(#4623): set changeset fragment pr to 4749 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ad1477d659 |
enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap * enhance(260903-m7p): validate configured entrypoints before success * test(260903-m7p): require pre-success entrypoint validation * enhance(260903-m7p): gate install success on entrypoints * test(260903-m7p): cover configured entrypoints across runtimes * enhance(260903-m7p): cover emitted runtime entrypoints * fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number - finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process before the new configured-entrypoint assertion throws. Without a HOME + config-location-env sandbox that write resolved through the ambient environment and landed in the developer's live ~/.gsd (confirmed absent on origin/next baseline, present only on this branch — full-suite HERMETICITY WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the duration of the test, matching the existing in-process finishInstall/ install() pattern in tests/install.test.cjs (#2665). - .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the fork PR number until the upstream PR number is known. * fix(260903-m7p): repair cross-platform and pre-existing shape fallout - tests/configured-entrypoint-validation.test.cjs: the win32 branch of ensureCodexHooksJsonSessionStart writes a .cmd shim under <codexRoot>/hooks/; create that dir in the test (the real installer only calls this once hooks/gsd-check-update.js already exists) and assert the platform-common entrypoint shape instead of a fixed non-Windows array, since win32 legitimately emits two entries (cmd shim + script). - tests/install.test.cjs: finishInstall's shared settings-json return now carries configuredEntrypoints/rollbackInstallerMigrations for every runtime on that path (trae included, not just Claude/Cursor/Windsurf); update the trae install() exact-shape assertion to match. * fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash configuredEntrypointsForHook's shell branch dropped interpreterCandidates entirely when resolveBashExecutable returned null, unlike the sibling portableHooks runner entry a few lines below (which correctly falls back to the literal 'bash' token). Found via agy adversarial review; verified unreachable through the current call graph (buildHookCommand's own resolveBashRunner==null gate already short-circuits before recordConfiguredHookCommand runs), so this is a defensive consistency fix, not a live-bug patch — kept for the next caller that does not share that gate. * chore(260903-m7p): backfill changeset pr number to the opened upstream PR .changeset/quick-wasps-sing.md carried the fork PR number (16) as a placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now open, so record its real number per CONTRIBUTING.md's changeset pr-field convention. * fix(#4154): track already-registered hooks for entrypoint validation on update applySettingsJsonHooks registers each guard hook only if absent, so a hook already present from a prior install keeps its stale on-disk command. The new entrypoint tracker always records the freshly-computed command for it, which never matches what is actually persisted, so the exact-string filter in finishInstall silently dropped it from validation — the Blocker case this feature exists to catch (an already-installed entrypoint going stale between installs) was exactly the case it never validated. Match on the managed script's basename instead, which the persisted command carries either way, so an already-registered hook stays in the validated set. Regression test forces this path by mutating a freshly-installed hook's persisted command before a second install. * fix(#4154): distinguish an unreadable script from a missing one validateConfiguredEntrypoints folded an EACCES statSync failure into the same 'missing' reason as ENOENT, misreporting a real permission problem as an absent file. Check the error code and report 'unreadable' instead. * docs(#4154): document entrypoint validation's rollback and PATH scope CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite bin/install.js x CONTEXT.md being this repo's strongest co-change pairing. The update-gsd.md how-to overstated what a validation failure undoes: for Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/ config.toml inside install() before the aggregate validation call runs, so there is no rollback path for that write regardless of "where available" phrasing. Also note that interpreter resolution checks the installer's own PATH, not necessarily the PATH a hook fires under later (#2979 launchers). * chore(#4154): point changeset pr field at the fork PR while CI runs there Mirrors the branch's own prior backfill commit: pr: matches whichever PR number changeset-lint is currently validating against (fork PR #16 during the fork-first CI/review loop), flipped back to the upstream PR number right before the final push to open-gsd/gsd-core. * fix(#4249): address adversarial-review findings in entrypoint validation An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR found several real gaps beyond the human reviewer's Blocker, verified against source before fixing: - Codex's install() result bound rollbackInstallerMigrations to the narrow installer-migrations-only rollback instead of restoreCodexSnapshot (#3245), the full pre-install snapshot/restore Codex already owns for exactly this case — a validation failure discovered outside install() reverted nothing of the config.toml/hooks.json that call had already written. - The register-only-if-absent basename match from the prior fix used a bare substring, which an unrelated user command mentioning the same filename could false-positive into GSD's validated set — anchored on the `/hooks/<basename>` path segment instead. - nodeCandidates checked raw process.execPath (always true — we're running in that process) instead of normalizeNodePath's stable version-manager alias, the same one buildNodeRunnerChainToken bakes as its first choice — a false green regardless of whether that alias itself still resolves. - An entry with no interpreterCandidates (Cline's PreToolUse hook, or a Windows-Claude .sh hook invoked without a bash runner) runs via its own shebang; validateConfiguredEntrypoints checked only file-type, never the execute bit. Cline's writer also never reported an entrypoint at all. - Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor hook registered across several events) were validated once per duplicate. Each fix is covered by a new or extended test; the Codex one required inlining runCodexInstall's env sandboxing so the rollback closure — which re-resolves the $HOME-relative skills root live — runs before the sandbox is torn down, matching how installAllRuntimes' real aggregate gate calls it. * docs(#4249): document the round-2 entrypoint-validation fixes Runtime Hooks Surface Module and Installer Module entries now name ConfiguredEntrypoint's not-executable reason, the normalizeNodePath alignment, Cline's tracked hook, and which install() result the finishInstall/installAllRuntimes rollback path actually reverts per runtime (Codex's full snapshot vs. the others' narrow migrations-only rollback). * chore(#4249): point changeset pr field at the upstream PR now that fork CI is green * fix(#4249): address agy adversarial-review findings - validateConfiguredEntrypoints: statSync alone never detects a chmod-000 script (it only needs parent-dir search permission), so an interpreter-invoked entry with an unreadable script passed validation. Add an explicit R_OK check for the interpreterCandidates branch only — the candidate-less/shebang branch already has its own X_OK gate. - docs/how-to/update-gsd.md: the blanket "does not revert" claim was false for Codex, which reverts config.toml/hooks.json via its full pre-install snapshot; qualify it per runtime. - tests/codex-config.test.cjs: the #4249 rollback regression test asserted skills/ and VERSION were reverted but never asserted config.toml/hooks.json were too, despite the test's own stated intent. - CONTEXT.md: qualify which interpreterCandidates entries get normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's Windows .cmd shim as a third candidate-less case that relies on extension dispatch, not a shebang. * fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node', so it needs the execute bit (like any shebang-invoked entry), but its interpreter is looked up on PATH by 'env' at hook-fire time (unlike every other GSD JS hook, which bakes an absolute node path specifically to avoid that dependency). The candidate-less/interpreterCandidates fork treated these as mutually exclusive, so Cline's entry silently skipped interpreter resolution entirely — a completely missing 'node' on PATH would still validate successfully. Add an orthogonal selfExecutable flag so both checks run for entries that need them. (CodeRabbit finding on the fork rehearsal PR.) * fix(#4249): address second-round adversarial review findings (opus + agy) - validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry, not just interpreterCandidates ones — a self-executable shebang script is still opened and read by its kernel-invoked interpreter, so X_OK alone never proved it was readable. - selfExecutable is now the sole, explicit source of truth for the execute-bit check (every producer that needs it sets the flag) instead of being partly inferred from an absent interpreterCandidates, which Cline's hybrid entry also carries. - The execute-bit check now skips explicitly on win32 (matching resolveExecutableBinary's own carve-out) instead of relying on Node's accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows machine and not a test that simulates win32 on a POSIX runner. - bin/install.js: fixed a stale comment claiming no runtime's install()-time writes have a rollback path — Codex's does (restoreCodexSnapshot) — and added the omitted Cline to both that comment and CONTEXT.md's equivalent lists. - CONTEXT.md: fixed the Cline description left stale by the previous commit's selfExecutable addition, and rewrote the validation-mechanism paragraph for clarity (writing-for-agents pass). - docs/how-to/update-gsd.md: split an overloaded 4-clause sentence. - Removed a fault-injection integration test that could not reliably exercise the real installAllRuntimes -> finalize -> rollback wiring without fighting the installer's own pre-registration existence guards; the constituent pieces remain covered individually. * fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner X_OK is a POSIX-only concept, skipped entirely when an entry's platform is win32 (matching production). Two test entries omitted platform, defaulting to process.platform — on an actual windows-latest CI runner that silently skipped the very check they were meant to exercise, turning 'not-executable' into a false pass. Pin platform: 'linux' so these are deterministic regardless of which OS runs the suite. * fix(#4249): classify EPERM the same as EACCES in statSync error handling Windows raises EPERM (not EACCES) for a parent directory that couldn't be traversed into — was falling through to 'missing', misreporting a genuine permission problem as a nonexistent path. * docs(#4249): address final CodeRabbit doc-completeness findings - CONTEXT.md: install()'s documented result shape omitted configuredEntrypoints; the ConfiguredEntrypoint shape omitted selfExecutable. - docs/how-to/update-gsd.md: the failure-mode sentence omitted unreadable and lacks-execute-permission, which the installer also rejects. * fix(#4249): stop double-validating every configured entrypoint on install/update installAllRuntimes' finalize() already runs assertConfiguredEntrypoints once over the aggregate set; finishInstall then re-ran the identical check per runtime in the printSummaries loop right after, so every entrypoint paid its statSync/accessSync/interpreter-resolution cost twice on every install and update. Add entrypointsAlreadyValidated to skip the redundant pass specifically on that path, while leaving the check intact for any caller that invokes finishInstall directly. * chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there * perf(#4249): memoize interpreter candidate resolution across entrypoints resolveExecutableBinary walked PATH once per (entry, candidate) pair; a typical install has a dozen-plus entries sharing the same few candidate lists (process.execPath for JS hooks, bash for shell hooks). Cache by (platform, candidate) so each distinct pair resolves once per validation call instead of once per entry. * chore(#4249): point changeset pr field at the rebased rehearsal fork PR * fix(#4249): drop entrypoint tracking from the now-dead Codex event writer #2586 (landed on next after this branch forked) removed install.js's CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent no longer runs during install or update. The ConfiguredEntrypoint records this branch added inside it were therefore unreachable and untested. Restore the function to its upstream shape; the entrypoints it used to report were never collected by any caller. * refactor(#4249): drop the revalidation bypass flag and the candidate cache Both were this PR's own micro-optimisations over a set of roughly a dozen entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall gate off to save one statSync/accessSync pass; `resolvedCandidateCache` memoised resolveExecutableBinary across entries that are already deduped by (configPath, scriptPath). Neither is measurable, and the flag was the only way to reach finishInstall with validation disabled. finishInstall now always validates what it is given. * chore(#4249): point the changeset pr field back at the upstream PR * refactor(#4249): track settings.json entrypoints without the hooksSurface gate The install-surface writer only tracked configured entrypoints when the runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing asserts that axis agrees with `installSurface`, so a descriptor that broke the coupling would silently pass `configuredEntrypoints: undefined` and drop that runtime out of the validation this PR adds — reintroducing the exact 'reports Done! over a broken entrypoint' failure #4154 exists to close. Remove the dependence rather than test it: everything recorded on this path lands in settings.json by construction, and the registered-command filter already discards entries no persisted hook references. * chore(#4249): put the changeset body in the documented two-part format CONTRIBUTING.md and .changeset/README.md both show `**<bold change>** — <symptom-led explanation>.`; the fragment was a single unbolded sentence. * chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there * fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback #3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-* and gsd-core/VERSION. The install overwrites every other GSD-owned file too — hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest itself — before the entrypoint-validation gate runs, so a validation failure left the new payload sitting on top of the restored old config. Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims, before runInstallerMigrations so the bytes are the true pre-install state, and restore it from both Codex rollback closures ahead of the per-surface restores. Files only the failed install introduced are removed, read from the manifest now on disk. The manifest is already the authoritative record of what GSD owns, so no second hand-written list can drift out of sync, and user-owned files are never snapshotted or removed. Every path is confined through resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into an arbitrary-path write. Non-Codex runtimes are unaffected: the snapshot is gated on the same tomlConfigInstall + non-minimal condition as #3245's. * fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest Two follow-on defects in the previous commit's snapshot: - The capture was gated on `!isMinimalMode`, copied from #3245. A core/ --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the snapshot came back empty while the rollback still ran — and its removal pass would have deleted every file the new manifest lists. Gate on tomlConfigInstall alone, matching where the rollback actually reaches. - An unreadable or unparseable prior manifest was caught alongside ENOENT and treated as a fresh install. That is the same empty-snapshot state, so a failed update over a real install with a corrupt manifest could delete its prior payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means known-empty; any other read error or a parse failure means unknown, and the restore closure returns without touching anything, degrading to #3245's narrower rollback. Deliberately not fatal — a corrupt manifest has to stay repairable by reinstalling over it. Both paths are covered by red-checked regression tests. * fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too commit removed from the manifest snapshot. restoreCodexSnapshot is reachable for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-* skill dir and gsd-* agent file the snapshot does not claim — so with an empty minimal-mode snapshot a rollback deleted the whole skills/agents surface with nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via the ADR-1239 skills-kind home override, so this is also the reason manifest `skills/` keys do not resolve under configDir: that surface belongs to this snapshot, not to the manifest-driven one. Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal mode — doing nothing on an early failure is the non-destructive side. Covered by a red-checked regression test that plants bytes in an alternate-home skill file, reinstalls under the core profile marker, and asserts the rollback restores it. * fix(#4249): never remove on rollback unless a prior manifest proves what predates the install Three defects in the manifest-driven Codex rollback, all in its removal half: - ENOENT marked the snapshot usable, arming the removal pass on a FIRST install. GSD may have overwritten a user's file at a manifest-tracked path there, and no prior manifest records the difference — so rollback deleted it where before it merely left it overwritten. Absent, unreadable and malformed manifests now all leave the prior set UNKNOWN and skip removal entirely. - Membership was tested against the map of files whose pre-install read SUCCEEDED, so a tracked file that existed but was unreadable read as introduced-by-this-install and was removed. Track the prior manifest's paths in their own Set and test against that. - The unreachable "delete the manifest when there was no prior one" branch is gone: usable now implies a parsed prior manifest. Also adds the end-to-end test the aggregate gate was missing — the four Codex rollback tests drove the closure directly, proving the restore but not the wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes Cline's `env node` entry fail validation for real, and asserts Codex's payload comes back. Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper instead of six copies. Both new tests are red-checked. * test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its Windows-EBUSY retry budget. That budget is for directory trees; this removes a single file, which unlinkSync says more precisely and the rule does not flag. * chore(#4249): point the changeset pr field back at the upstream PR * fix(#4249): use an unambiguous dedup key and surface partial-restore failures trek-e's 2026-09-08 adversarial pass flagged two findings in the new entrypoint-validation/rollback code: - assertConfiguredEntrypoints' dedup key already used a raw NUL separator (introduced in ceebb65f2d), but git/Read render NUL as a space, so the key looked like a plain-space join to every reviewer that read the diff. Replace it with JSON.stringify([configPath, scriptPath]) so the separator is visible and unambiguous. - restoreManagedFileSnapshot's per-file restore catch block claimed to 'surface the original error' but only swallowed it, matching (and widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add an actual console.warn using the existing best-effort-warning convention, scoped to just this PR's new function. * fix(#4249): treat a files-less prior manifest as unknown, not known-empty agy's gemini-3.8-flash-high adversarial pass (round 5) found and I reproduced empirically: a structurally-valid manifest missing the files key (e.g. {"version":1}) parses without throwing, so Object.keys(undefined || {}) silently read as 'zero files predate this install' instead of the UNKNOWN state the malformed-manifest guard exists to produce. Rollback's removal pass then deleted every GSD-owned file the failed install's own manifest listed, including ones that predated it — the exact data loss the #4249 CodeRabbit malformed-manifest fix was supposed to prevent, reachable through a JSON.parse success instead of a failure. Route the shapeless case into the same catch-all UNKNOWN path via an explicit shape check. Regression test reproduces the deletion before the fix and confirms the file survives after it. Also extend restoreManagedFileSnapshot's removal-pass rmSync and final manifest-rewrite catches with the same real console.warn trek-e's round-4 review asked for on the per-file restore catch — same rollback function, same operator-facing-signal gap. * docs(#4249): correct which runtimes actually leave a written config on rollback agy's completeness audit (round 5, holistic pass) caught this new paragraph claiming 'for every other runtime, the configuration file(s) already written during that update are left in place' — false for Claude Code and other settings.json-based runtimes, whose write never happens on failure (assertConfiguredEntrypoints runs before finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline actually match that description, since they persist their config file inside install() ahead of the gate. Split the one sentence into the three actual outcomes; matches the PR body's own accurate Before/After wording, which this doc addition had drifted from. * fix(#4249): clean up doc/comment mismatches and dead fields from opus review Opus critical-code-reviewer + ponytail-review pass on the final diff: - assertConfiguredEntrypoints carried finishInstall's old docblock ("Apply statusline config, then print completion message") from before this function was inserted between comment and callee. finishInstall already has its own accurate #4249 comment, so the stale docblock is removed rather than moved. - checked: number on ConfiguredEntrypointValidationResult and error.configuredEntrypointValidation on the thrown error: the first had zero consumers anywhere in the repo, including its own defining file, and is removed. The second matches an existing repo convention (bin/install.js's installerMigrationRollbackFailures, #4249 predates this PR) of attaching structured diagnostic context to a re-thrown Error even before a consumer exists, so it's kept. - finishInstall's own assertConfiguredEntrypoints call is a redundant backstop on the real production path (installAllRuntimes's aggregate call already validates the superset first), but its comment read as though this call alone provided the before-the-write guarantee. Clarified rather than removed — it's the only gate for a caller that invokes finishInstall directly. * chore(#4249): split the manifest-driven rollback engine out into #4544 Issue #4154 asked the installer to consume a validation failure "through the existing rollback mechanism, without a second transaction mechanism". The manifest-driven rollback widening added during review (capture every path the prior gsd-file-manifest.json claims, restore those bytes, remove what only the failed install introduced) is that second mechanism on a plain reading. It is a real fix for a #3245-era gap, but an independent one, so it moves to its own bug report and PR. Removed here: - bin/install.js: the pre-install managed-file capture block and restoreManagedFileSnapshot, plus its call sites in _codexPreConfigRollback and restoreCodexSnapshot (99 lines). - tests/configured-entrypoint-validation.test.cjs: the five tests that exercise the manifest engine. - CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the widened restore. update-gsd.md again documents the #3245 surfaces only. Kept, because it is #4154's own scope: - the entrypoint-validation gate itself; - Codex's install() result binding rollbackInstallerMigrations to restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, gsd-core/VERSION); - the !isMinimalMode gate removal on that snapshot. Binding the closure to the result made it reachable for a core/--minimal install, where its pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot does not claim; an empty minimal-mode snapshot therefore deleted the whole surface with nothing to restore. The surviving aggregate-failure test now asserts on config.toml, a surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which only the manifest engine restored. Refs #4544 * test(#4249): cover configured entrypoints through the packed install path #4154's scope lists install smoke coverage alongside the installer gate — "assert representative configured entrypoints resolve for supported runtime profiles". The gate itself (assertConfiguredEntrypoints / validateConfiguredEntrypoints) is unit-covered by in-process install() calls; nothing proved the property survives npm pack -> npm install -g -> install.js. Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct config surfaces GSD writes launch paths into (settings.json, and hooks.json + config.toml) — run the tarball-installed installer into a throwaway HOME, then re-read that runtime's own written config and return the new ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check becomes a release gate on every matrix host without workflow changes. The scan re-derives paths from the written config instead of reusing the installer's own entrypoint list, and test I shows why that matters: a registration the installer never touched during a run is invisible to the in-process gate, so the install exits 0 and only reading the config back off disk catches the dangling launch path. * ci(#4249): pack a publish-shaped tarball in the install smoke lane `npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs after `npm ci` carries no hook scripts at all — the lane has been smoking a package that differs from the published one in exactly the artifacts the lifecycle smoke is supposed to launch. That went unnoticed because the lane's init runs `--local`, which registers no statusline and therefore registers no hook whose target is missing. A `--global` install on the same tarball exits 1 on #4249's own gate (`gsd-statusline.js (missing)`), which is what the new configured-entrypoint cycle performs, so without this step the cycle would report INIT_FAILED instead of checking anything. Build hooks before packing so the smoked tarball matches prepublishOnly. The CLI now reports 16 configured entrypoints for claude and 1 for codex instead of zero. * fix(#4249): scope Codex's full snapshot restore to entrypoint failures Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage exception un-install a Codex install that had already succeeded and already printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints` call. Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration` scopes finalize-stage rollback to installer *migrations* ("the executor uses the journal to restore modified paths"), and this PR's own operator-facing paragraph in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/ agents/VERSION revert to entrypoint-validation failures specifically ("If a script is missing, unreadable, ... For Codex, this reverts ..."). The wide behaviour is also incoherent as a transaction abort: the same doc says Cursor, Windsurf, Kimi and Cline keep the config they wrote inside install(). Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install bytes — on an update, silently downgrading a working Codex install to the previous version while the user had just been told it was Done. Select the rollback by error kind instead. `assertConfiguredEntrypoints` already tags its error with `configuredEntrypointValidation`, so the full snapshot restore runs for that error (and anything downstream of it, including finishInstall's per-runtime backstop) and the installer-migrations-only closure runs for everything else. The codex result now also exposes that narrow closure as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning the full restore, so the direct-call contract asserted by tests/codex-config.test.cjs is unchanged. Adds a regression test that installs codex+kilo together, injects EACCES on the Kilo permission write by monkeypatching node:fs (restored in a finally — never chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps the bytes the successful install wrote. Verified red against the pre-fix unconditional path. Cline cannot host this test: its plan is writesSharedSettings:false + finishPermissionWriter:null, so its finishInstall performs no write and has no non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally (unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded fs.writeFileSync. * docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing CONTEXT.md still described Codex's rollback as an unconditional bind to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint- validation failures via rollbackInstallerMigrationsOnly and the configuredEntrypointValidation error tag. Caught during the round-6 PR body pass. * fix(#4249): stop rollbackInstallerMigrations meaning its own opposite Codex's install() result bound `rollbackInstallerMigrations` to restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the actual installer-migrations-only closure behind `rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name meant the opposite of what it says, and CONTEXT.md had to concede as much in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure for every runtime, matching both its name and the meaning it already has on next, and the snapshot restore gets its own Codex-only field, `rollbackPreInstallSnapshot`. The selection in rollbackFinalizedInstallerMigrations collapses to one line and no longer needs a fallback chain. Also in this commit, all against the same rollback path: - Correct the rollbackFinalizedInstallerMigrations comment. It read as if the round-6 narrowing prevented any sibling-triggered revert of a Codex install the user has already seen "Done!" for. It does not, and is not meant to: `wide` is true for ANY entrypoint-validation error from ANY runtime, because the aggregate gate is all-or-nothing — an invalid Cline entrypoint reverts Codex's snapshot, which tests/configured-entrypoint-validation.test.cjs's 'an aggregate entrypoint validation failure rolls the Codex install back (#4249)' asserts directly. The discriminator is the error's KIND, not which runtime owns the failing path. Comment and CONTEXT.md now say that. - Name the runtime in the "Configured entrypoint validation failed" error. ConfiguredEntrypointInvalid already carries `runtime`; the message threw it away, leaving an operator of a multi-runtime install unable to tell whose entrypoint broke — which matters precisely because the failure can revert a runtime that was itself fine. - Set `configuredEntrypoints: []` explicitly on the copilot-instructions early return. Every other branch states the key; this one relied on installAllRuntimes' `(result.configuredEntrypoints || [])` defence. `[]` is correct, not a workaround: every Copilot hook is an inline printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no GSD-managed script or interpreter to resolve. No behaviour change beyond the error-message text. * docs(#4249): narrow the smoke scan's config-surface claim to what it checks RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is told to launch is registered in one of settings.json / hooks.json / config.toml, and that nothing else in a config dir is runtime configuration. Both halves are false as stated. Cline registers its hook at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native [[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi), a directory separate from Kimi's own GSD configDir — the same gap installer-migration 007 already documents as structurally unreachable. The scan is in fact correct for what it runs against: entrypointRuntimes defaults to claude + codex, whose launch paths do all live in those three top-level files. Restate the docstring at that scope, name the two known out-of-scope surfaces, and warn that adding either runtime to entrypointRuntimes without teaching scanConfiguredEntrypoints about its surface yields a scan that finds zero entrypoints and proves nothing. The entrypointRuntimes default comment carried the same overgeneralization ("every other runtime reuses one of them") and is corrected with it. Documentation only; no code change. * fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found that install()'s settings-json early return for an unparseable settings.local.json omitted configuredEntrypoints and rollbackInstallerMigrations from its result, unlike every other branch. rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations unconditionally, so this branch silently dropped its own installer-migration rollback on a later finalize-stage failure. Completed the return shape: configuredEntrypoints: [] (matching Copilot's equally-early no-entrypoints-yet return) and rollbackInstallerMigrations (already in closure scope). Red-then-green regression test added. * test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs Same internal adversarial review (round 8): this suite's before() packed the tarball directly, without the ensureHooksDist() guard every sibling install-test suite (install.test.cjs, install-minimal-hooks.test.cjs, mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run in isolation ahead of a suite that builds hooks/dist itself, this suite's pack would ship a tarball with no hook scripts and fail closed on SMOKE.INIT_FAILED instead of testing anything. * fix(#4249): refresh stale test-timings weight for the codex-config split next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the CI shard packer's weight table (tests/test-timings.json) was never updated: codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block this PR extends with its own #4249 install()-pipeline test — had no entry at all, so the packer would silently underestimate it at the table's median weight (roughly a 9x underestimate against its real cost). trek-e's most recent review flagged a Windows shard timeout in-flight on codex-config.test.cjs, plausibly aggravated by this PR's own addition to that file before the rebase moved it. Re-measured both files locally (node --test --test-reporter=tap, max of 3 runs, matching the table's own max-across-streams methodology) and patched just these two entries — not a full regeneration, which would need real multi-lane CI data this session doesn't have access to. * fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists next's platform-conformance-tier classifier (#4591/#4598) landed after this branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's gen-platform-conformance-tier --check and both the Linux and macOS conformance suites. * fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation next's #4733 (landed after this branch's last rebase) replaced the static ISOLATED_HEAVY_FILES set with a threshold derived live from tests/test-timings.json, and pins the current derived result in EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still named codex-config.test.cjs, whose own weight this PR already dropped from 127783ms to 189ms (after splitting its heavy install()-pipeline blocks into codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The live-computed set correctly no longer includes it; the pinned expectation is updated to match. * fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it trek-e's review flagged two Major gaps: the thrown error read identically regardless of which of three real outcomes a runtime hit (nothing persisted / snapshot reverted / config left broken on disk), and no test proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/ Kimi/Cline — only Codex's revert path was ever asserted. assertConfiguredEntrypoints now tags each invalid entry with its actual consequence, mirrored from docs/how-to/update-gsd.md's existing rollback-matrix disclosure. A new test drives the same aggregate failure through Cline (whose own entrypoint is the one that fails) and asserts its hook file is still on disk afterward. * fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR One review pass (gemini-3.8-flash-high via the antigravity review lane) against this PR's full diff against next, findings independently verified against source before fixing: - Copilot's install() return object was the only one of 6 runtime branches missing rollbackInstallerMigrations — reachable now that this PR's own aggregate gate runs rollback across every result on any runtime's entrypoint failure, not just Copilot's own. - buildHookCommand's unresolved-bash early return skipped track() entirely, so a win32 install with no Git Bash silently produced an unregistered .sh hook instead of the 'unresolved-interpreter' validation failure configuredEntrypointsForHook's own comment said it would. - release-tarball-smoke.cjs reported a Cycle 4 install failure under SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing SMOKE.INSTALL_FAILED. - SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's trailing args, which also truncated any configDir containing a space (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan. Anchored the match on the already-known configDir prefix instead of a generic absolute-path guess: removes the ambiguity outright rather than patching the character class, and stays a raw-text scan on purpose (it catches a writer that emits a path without registering it — a JSON.parse of the expected schema would miss exactly that case). One suggested finding (test-timings.json "missing" the new test file) was verified false — that table only holds measured CI timings, populated after a file's first real run — and one Ponytail suggestion (a JSON.stringify dedup key) was rejected as it would reintroduce a real, if narrow, key-collision risk for no benefit. * fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment changeset-lint requires pr: to match the PR it runs on (16 on the fork, not the eventual upstream number) — rehearsal-branch convention already established earlier in this PR's history. lint-docs-guard-registration's quote-pairing heuristic doesn't require the docs/ path itself to be quoted — it flags a file once ANY quote-delimited span containing "docs/" appears anywhere in it, alongside any real fs read call. A comment ending "...update-gsd.md's rollback-matrix paragraph" supplied the closing quote character (the possessive apostrophe) the heuristic paired with an unrelated single-quoted string earlier in the file. Reworded to avoid the unquoted apostrophe next to the path. * chore(#4249): point the changeset pr field back at the upstream PR Fork rehearsal (PR #16) is green; the real target for this changeset is upstream PR #4249. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
49f313d611 |
fix(#4461): make code-review summary extraction shell-safe (#4533)
* fix(#4461): make code-review summary extraction shell-safe Emitted-Drift-Ack-Growth: code-review.md — use a literal heredoc for shell-safe SUMMARY parsing * chore: add changeset for #4533 * fix(#4461): keep heredoc outside command substitution * chore: rerun CI after Windows timeout * test(#4461): execute the summary heredoc adversarially * test(#4461): normalize adversarial paths for Git Bash * test(#4461): keep adversarial fixture valid on Windows * test(#4461): scope heredoc regression claim --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
740ba0d8a3 |
fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file Co-authored-by: sim <sim@local> |
||
|
|
febe6c9885 |
fix(#4685): a directory artifact fails its own entry instead of aborting the check (#4735)
* fix(#4685): a directory artifact fails its own entry instead of aborting the check `must_haves.artifacts` entries are read with `safeReadFile`, which rethrows every errno except ENOENT. A listed path that is a directory therefore threw EISDIR out of the per-artifact loop: `query verify.artifacts` printed Error: EISDIR: illegal operation on a directory, read and reported NOTHING — not the offending entry, and not the plan's other, perfectly checkable artifacts. One directory entry disabled the whole plan's check. Reproduced against a real plan before the fix, and after. A directory is now reported as that entry's own failure, with an issue distinct from `File not found` (the path did resolve; it simply is not the thing an artifact entry can be checked against), and every other artifact in the plan is still checked and reported independently. Anything else the stat or read throws becomes that entry's failure too, carrying its errno, rather than discarding the run — a check that disappears is worse than one that fails, because a failure is visible. Verifying directories properly — matching `contains:`/`min_lines:`/`exports:` across the files inside one — is a feature decision and deliberately not made here, per the issue's stated scope. Authoring-time rejection of a directory path is likewise left alone: the brief raises it as a separate question, and the runtime fix does not depend on it. Also, found in pre-PR review and pre-existing: `safeReadFile(...) || ''` turned a post-stat ENOENT into empty content, so an artifact declaring only `path`/`provides` had no criterion left to fail and passed, having checked nothing. A null read now reports that instead of inheriting a pass. It can only turn a false pass into a failure. Verified: reverting src/verify.cts to the merge-base turns both new rows red with the exact EISDIR message; lint:ci exit 0; full suite 24/24 chunks, 37,280 tests, 0 failures; tests/verify.test.cjs 219/219. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * chore(#4685): backfill changeset PR number to 4735 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * test(#4685): pin the injected-I/O branches, and narrow the guard's comment Review findings from #4735. Major — the two error branches this PR adds were untested, and the PR body claimed the mid-check ENOENT window was "not deterministically reproducible through the CLI seam these tests drive." That was wrong: ADR-3574 records this repo's convention for exactly this — inject filesystem failures by monkeypatching the fs method, never by chmod or mode-bit tricks, which root bypasses and yields a test that passes with zero coverage in root Docker and CI. Both branches are now pinned that way. The injection runs in the CHILD via NODE_OPTIONS=--require, because `output()` writes fd 1 directly (`writeAllSync(1, …)`, io.cjs) rather than through console.log, so an in-process call cannot have its JSON captured. The preload patches the child's own module objects, which the compiled code reads at call time. - a file that disappears between stat and read now fails as that entry rather than passing on empty content - a non-ENOENT errno (EACCES) is reported as that entry's failure, carrying its code, so an operator can tell a permissions problem from an I/O one The injection matches the target by path SUFFIX, not string equality: the first cut compared absolute paths, and a /tmp vs /private/tmp prefix difference silently disarmed it — the test passed while asserting nothing. A disarmed injection test is worse than no test, so the reason is recorded at the call site. Nit — the comment above the try block said the guard "covers what the stat and read below actually throw", which reads as if the min_lines/contains/exports checks inside the same block were deliberately guarded too. They are pure string operations and cannot throw; the comment now says so rather than implying a guarantee it does not make. Verified: reverting src/verify.cts to the merge-base turns all three #4685 rows red — the directory row and both new ones; lint:ci exit 0; full suite 24/24 chunks, 37,389 tests, 0 failures; tests/verify.test.cjs 221/221. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ecc508139a |
fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting (#4755)
* fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting Planners emit <automated> bodies with the chain operator entity-escaped (`&&`), and the executing agent reads the decoded (rendered) form. The grounding probe split the raw text on &&/||/;/newline without decoding first, so `&&` was cut at its semicolons and a cd-form target absorbed the trailing `&` fragment — an existing directory was reported missing_dir (blocker), feeding false blockers into the revision loop. Decode `&` → `&` inside resolveVerifyCommandTarget, after result.command captures the text verbatim and before any segment splitting or target resolution, so the escaped and literal forms of the same command produce identical verdicts. Module-private helper beside splitSegments, mirroring the sibling src/verify.cts decodeEntityAmps (#3611); the two gates are separate modules and neither imports the other. Regression coverage pins escaped/literal verdict parity for: existing dir + manifest (ok), missing dir (missing_dir blocker kept), dir without manifest (no_manifest blocker kept), --prefix form, a literal & inside a quoted dir name, and a full probePhaseVerifyCommands pass whose reported command field stays verbatim. * docs(#4730): use the documented pr:0 placeholder in the changeset fragment The fragment carried pr: 4730 — the ISSUE number, the exact guess-shape DEFECT.CHANGESET-PR-FIELD-DRIFT (#3316, #3325) exists to catch: it parses as a positive integer so local lint passed, but on a real PR run the drift check would fail it against the actual PR number. CONTRIBUTING.md documents pr: 0 as the deliberate unresolved placeholder used during initial commit before the PR number exists (scripts/changeset/new.cjs accepts 0 for exactly this reason); the gate's fail_invalid_fragment on an unbackfilled 0 is the designed backfill enforcement, not a defect. No production or test changes. Backfill pr: with the real PR number once the PR is created. * docs(#4730): backfill changeset pr field with the real PR number pr: 0 → pr: 4755 (the PR carrying this fix), completing the documented placeholder workflow; the changeset gate's content validation can now pass. --------- Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
092d9256b8 |
fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six
Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.
tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.
Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it
The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:
Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.
Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.
Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).
Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.
Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers
The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:
- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
and a `{padded_phase}` binding are the sanctioned shapes.
Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* docs(#4748): add Fixed changeset
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth
The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).
Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* fix(#4748): name padded_phase in execute-phase.md's init parse list
A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form
Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup
The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split
tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling
Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* test(#4748): use t.after() for the composition test's fixture cleanup
CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK
* chore(#4748): set changeset fragment pr to 4768
* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next
Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base
|