Commit Graph

5923 Commits

Author SHA1 Message Date
Tom Boucher
4fe2837ab5 fix(#4774): plan-criteria R4 requires the pipe not to be doubled — a logical-OR fallback is handled, not swallowed (#4877)
* test(#4774): failing-first — R4 must not read a logical-OR fallback as a pipeline stage

* fix(#4774): R4 requires the pipe not to be doubled — a logical-OR fallback is a handled failure, not a swallowed one

Also corrects the rows-3/4 test's makeCriteriaPlan usage (second arg is the
<verify> block, not a second criteria line).

* chore(#4774): backfill changeset PR number (4877)

---------

Co-authored-by: sim <sim@local>
2026-09-19 15:38:16 -04:00
Tom Boucher
a87b83d485 fix(#4764): dep_phases extracts only Phase-prefixed references from Depends-on prose (#4876)
* test(#4764): failing-first — dep_phases must extract only Phase-prefixed references, never dates/shas/ledger ids/self

* fix(#4764): dep_phases anchors phase references to their 'Phase' prose context and never emits the row's own number

* fix(#4764): review fold-ins — hoist the anchored dep-reference grammar to phase-id, cover Oxford lists and hyphen ranges, repair the property test

Adversarial review found: Oxford-comma lists under-extracted ('Phases 1, 2,
and 3' dropped the tail member — a silent real-blocker clear, the dangerous
direction); hyphen ranges ('Phases 1-3') kept only the first endpoint; the
property test called fc.hexaString (absent in fast-check 4.8, threw every
run) and passed the junk arbitrary unspread (vacuous guard) with no
completeness assertion; planning-inspect's extractDependencyTokens carried
the same whole-field scrape (generative-fix divergence). The anchored
grammar now lives beside PHASE_NUMBER_TOKEN_SOURCE in phase-id.cts and both
readers interpolate it.

* chore(#4764): backfill changeset PR number (4876)

---------

Co-authored-by: sim <sim@local>
2026-09-19 13:07:03 -04:00
Tom Boucher
969456c46d fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path (#4875)
* test(#4759): failing-first — preserved-foreign hooks warning must not claim may-not-load for a commonjs package.json

* fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path

* test(#4759): review fold-ins — assert installer output on stdout+stderr (warn is stderr), register cleanup before the spawn, add the type-less foreign case

* chore(#4759): changeset fragment

* chore(#4759): backfill changeset PR number (4875)

---------

Co-authored-by: sim <sim@local>
2026-09-19 10:55:01 -04:00
Tom Boucher
d36514b816 fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() (#4872)
* test(#4758): failing-first — rescue must resolve a relative worktree_path against repoRoot

* fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd()

* test(#4758): review fold-ins — t.after cleanup pattern, post-resolution reader contract comment

* chore(#4758): changeset fragment

* chore(#4758): backfill changeset PR number (4872)

* test(#4758): windows lanes key rescue fakes on resolved path identity, not verbatim strings

win32 path.resolve rewrites driveless-absolute POSIX-style fixture values to the
current drive, so the rescue's (correct) resolved-path handoff stopped matching
verbatim string keys: #3804/#245/#2556/B7/#2852 fakes silently skipped the
rescue and my seam test compared against a POSIX literal. Fakes now key on
path.resolve(repoRoot, …) identity — the same semantics the code and git -C
use — so every rescue test exercises the rescue on every platform.

---------

Co-authored-by: sim <sim@local>
2026-09-19 09:41:51 -04:00
Tom Boucher
63edc777e6 fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading (#4868)
* test(#4588): observed fork-from-HEAD must suppress the stale-origin degrade (failing first)

* fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading

* fix(#4588): a throwing probe git call is an inconclusive observation, not a crash

* test(#4588): name the observed reason in the inconclusive-row assertions

* test(#4588): hermetic state I/O in the inconclusive-row fixtures

* chore(#4588): changeset fragment

* test(#4588): contained cache I/O, emit-payload assertions (review fold-ins)

* fix(#4588): review fold-ins — invalidation fixture, hermetic confirm row, gate comment, cache+scope rows

* chore(#4588): backfill changeset PR number (4868)

---------

Co-authored-by: sim <sim@local>
2026-09-18 23:41:44 -04:00
Tom Boucher
9a41a95212 fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams

* fix(#4717): consult the per-install runtime marker at both identity seams

resolveReportedRuntime (agent_runtime) and loadConfigResolved
(config.runtime) both ignored the per-install .gsd-runtime marker that
resolveRuntime and the model-resolver gate already read. On a
multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing
misreported every Claude Code session as codex, and a shared
defaults.json stamped by the first non-Claude install leaked its runtime
to every other one.

Seam 1: the reported-runtime ladder becomes explicit > install marker >
host detection > claude. Seam 2: loadConfigResolved fills an empty
config.runtime from GSD_RUNTIME then the marker, copy-on-write (the
builtin-defaults branch returns a shared object). Explicit runtimes and
marker-less trees are unchanged.

* fix(#4717): a marker-detected runtime opts into its tier map (decision a)

* fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins

* chore(#4717): backfill changeset PR number (4861)

---------

Co-authored-by: sim <sim@local>
2026-09-18 13:50:01 -04:00
Tom Boucher
58c7bbb16a fix(#4667): rewrite codex @ includes to the codex install root (#4858)
* test(#4667): add failing-first coverage for the codex @-include rewrite

Behavioral end-to-end: a real in-process install(true,'codex') into a temp
CODEX_HOME must leave zero @~/.claude includes in GSD-owned .md artifacts,
rewrite the issue's own example include to @~/.codex/, keep the deliberate
_GSD_RUNTIME_ROOT .claude fallback chains byte-identical, and stay
idempotent across a reinstall (no doubled prefix). All four are RED until
the installer grows the manifest-scoped rewrite pass.

* fix(#4667): rewrite codex @ includes to the codex install root

Codex-installed agents and commands kept @~/.claude/gsd-core/... (and
@/Users/trekkie/.claude/gsd-core/...) include references pointing into the Claude
install: silent wrong-copy reads on dual-runtime machines at divergent
versions, missing files on codex-only ones. Several emitters bypass the
per-runtime converters, so the per-emitter fixes since #570 rotted.

Adds a manifest-scoped rewrite pass in install() beside the leak scanner:
for codex, every manifest-tracked .md/.toml artifact has the @-include
forms rewritten to the codex root. The pass matches the exact include
literal only — the _GSD_RUNTIME_ROOT/$PREFERRED_CONFIG_DIR fallback
chains, prose .claude mentions, and CHANGELOG.md are untouched — and runs
before the scanner, which remains the verification backstop for anything
a future emitter introduces.

* test(#4667): cover the HOME-anchored include form and sync the pass comment

Adds behavioral coverage for the second rewrite literal
(@$HOME/.claude/gsd-core/ -> @$HOME/.codex/gsd-core/) via the
plan-review-convergence command, and corrects the pass's comment: the
agent .tomls are generated after it and prefix themselves, so the .toml
branch of the rewrite is inert by design.

* test(#4667): baseline the offline upgrade against the deployed tree (sanctioned)

* chore(#4667): backfill changeset PR number (4858)

* test(#4667): sandbox the install home in the include-rewrite tests (#3712 guard)

* test(#4667): install via subprocess with an isolated env in the include-rewrite tests

---------

Co-authored-by: sim <sim@local>
2026-09-18 11:36:36 -04:00
Tom Boucher
e1f72cd324 fix(#4741): the plan checkbox tick respects the superseded exclusion (#4851)
* test(#4741): a superseded plan must not be ticked from its summary (failing first)

* fix(#4741): the plan checkbox tick respects the superseded-plan exclusion

* fix(#4741): review fold-ins — changeset typo, dedupe planId derivation

* test(#4741): exercise a halted summary on the active plan in the #2830 pin

* chore(#4741): backfill changeset PR number (4851)

---------

Co-authored-by: sim <sim@local>
2026-09-18 06:33:30 -04:00
Tom Boucher
11b3091df0 fix(#4738): record opencode's staged skills in the install manifest (#4847)
* test(#4738): opencode manifest must record its staged skills (failing first)

* fix(#4738): record opencode's staged skills in the install manifest

* test(#4738): use the centralized temp-dir helper in the manifest tests

* fix(#4738): retire the dead hostBehaviors vocabulary entry, tighten detector asserts, temper changeset

* chore(#4738): backfill changeset PR number (4847)

---------

Co-authored-by: sim <sim@local>
2026-09-18 04:52:32 -04:00
Tom Boucher
c5629bbe74 fix(#4734): degrade worktree isolation when the root has no git repository (#4843)
* test(#4734): non-git root must degrade worktree isolation (failing first)

* fix(#4734): degrade worktree isolation when the root has no git repository

* fix(#4734): review fold-ins — 3972 ladder fixture, parity fixture, docs row, message wording

* chore(#4734): backfill changeset PR number (4843)

---------

Co-authored-by: sim <sim@local>
2026-09-18 03:16:25 -04:00
Tom Boucher
8d0b6868ae fix(#4725): write normalization preserves tight paragraph-list shape (#4842)
* test(#4725): write normalization must not reflow untouched prose (failing first)

* fix(#4725): stop write normalization injecting a blank before a list after prose

* test(#4725): repair ordered-list fixture and list-spacing snapshot

* test(#4725): assert whole-file prose stability, fix heading-list comment

* chore(#4725): backfill changeset PR number (4842)

---------

Co-authored-by: sim <sim@local>
2026-09-18 01:38:14 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
bff99a8bb5 fix(#4731): read hard-wrapped Goal/Requirements fields past the line break (#4826)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; isolated adversarial review round completed (MEDIUM table-bleed finding fixed with RED/GREEN evidence) and sha-pinned bench 46331/0 on the merged head.
2026-09-17 15:59:38 -04:00
Tom Boucher
fb3e228a0d fix(#4724): classify Surefire/Failsafe XML as RED evidence (#4825)
* test(#4724): add failing-first coverage for Surefire XML RED evidence

* fix(#4724): classify Surefire/Failsafe XML as RED evidence

check tdd-red-evidence parsed only node:test TAP, so a JVM project's
genuine Maven red scored INVALID_RED while hand-written synthetic TAP
scored RED_EVIDENCE_OK — the gate was passable only by fabricating its
input (issue #4724's measured repro).

classifyRedEvidence detects Surefire/Failsafe XML (a <testsuite> element)
and parses it by TAG-BOUNDARY scanning: each <testcase> owns its own tag
(self-closing) or the segment up to its </testcase> closer, so the
issue's warned-about spanning trap (a lazy lazy match from a green
self-closing case to the next closing tag) cannot misreport names. A
<failure> or <error> child marks the case failing; the target matches at
class granularity (exact classname, dotted-suffix, or method name). Any
parse anomaly degrades to not-failing — the module stays fail-closed and
PURE (no fs/clock; report freshness remains the workflow's run-start
check per the issue's implementation notes).

TAP classification is byte-identical: all existing fixtures stay green.

* test(#4724): pin the scanner hardening — truncation, TAP-message flip, CDATA phantom

* docs(#4724): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 11:30:23 -04:00
Tom Boucher
7d0c6339d0 fix(#4705): emit Antigravity-native tool names as a YAML sequence (#4822)
* test(#4705): add failing-first coverage for native Antigravity tool sequences

* fix(#4705): emit Antigravity-native tool names as a YAML sequence

convertClaudeAgentToAntigravityAgent and the installer's twin emitted
Gemini CLI tool names as a comma-separated scalar. Antigravity's
documented subagent contract (antigravity.google/docs/subagents) wants a
YAML sequence of native names — view_file, grep_search, run_command,
replace_file_content are the documented examples, and wrong or malformed
grants can hang the subagent per Antigravity's own warning.

Map values move to the native vocabulary where documented (Read ->
view_file, Edit -> replace_file_content, Bash -> run_command, Grep ->
grep_search); undocumented entries keep their best-known grant rather
than being dropped (dropping would silently remove a restriction). The
emitter writes one '- name' item per line; an agent whose every tool was
filtered emits an explicit tools: [] instead of an empty scalar.
Pre-existing pins updated to the native vocabulary.

* test(#4705): update the #4727 map-value pin to the Antigravity-native vocabulary

The #4727-era pin held the map VALUES at the Gemini CLI dialect on the
belief that Antigravity speaks it; the confirmed bug #4705 (with
Antigravity's own documented subagent contract) supersedes that for the
four documented names. Key/shape pinning is preserved; only the values
move.

* docs(#4705): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 08:12:18 -04:00
Tom Boucher
be1b76dddd fix(#4700): queue the headless mempalace mine on the palace lock (#4821)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase

* fix(#4699): skip already-complete phases in the next_phase cascade

Both next-phase scans selected the numerically lowest phase above N
without consulting completion state, so completing a reopened phase
persisted an already-[x] phase as STATE.md current_phase while
roadmap.analyze correctly named the outstanding one (issue repro:
completing 2 with phases 1 and 3 already [x] returned next_phase 03).

The cascade collects the complete phase numbers from the roadmap
checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in
both the disk scan and the roadmap scan; a [x] checkbox row and its
heading sibling both name a phase that is never next. Heading-only and
checkbox-less roadmaps behave exactly as before.

* test(#4699): align the negative-control expectation with the disk spelling

* test(#4699): pin the STATE.md persistence and the all-later-complete tail corner

Review findings: the regression never asserted STATE.md current_phase
(the issue's actual harm), and the all-later-phases-[x] corner
(is_last_phase true, next_phase null) was unpinned. A changeset fragment
is included.

* docs(#4699): backfill changeset PR number

* test(#4700): add failing-first coverage for the queued headless mine

* fix(#4700): queue the headless mempalace mine and surface skipped captures

The capture's mine ran in the foreground with no lock handling: MemPalace
wraps every mine in a per-palace lock, so any concurrent writer (two
phases finishing a stage at once, a git-hook refresh mining the same
palace) made it exit 1 (MineAlreadyRunning) and the onError: skip step
silently dropped the capture — unlost for CONTEXT/PLAN/SUMMARY files that
can be re-filed, unrecoverable for execute:wave:post problem-fix pairs.

The mine now queues via --daemon --background (MemPalace #2029: the daemon
holds a job refused the lock and runs it when the holder exits), and the
report step gains the queued and skipped outcomes per #4700's requirement
that a skipped capture never stay silent. Option 2 (write_routing.cli) is
unreleased at MemPalace 3.9.0; option 3 (retry) re-enters the same lock
race — both declined in the PR body.

* fix(#4700): queue the wave:post problems fragment's headless mine too

The issue names the execute:wave:post problem-fix pair as the
unrecoverable loss (no source file to re-file later); the
capture-problems fragment's headless mine ran foreground like the capture
capability's did. Same fix: --daemon --background, with the lock-deferral
rationale inline.

* docs(#4700): backfill changeset PR number

* fix(#4682): register the stale-reverification part in the capability registry

The new steps/ part is a shipped workflow file; gen-capability-registry
--check requires it in the committed registry.

---------

Co-authored-by: sim <sim@local>
2026-09-17 06:30:36 -04:00
Tom Boucher
d707318e0c fix(#4699): skip already-complete phases in the next_phase cascade (#4820)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase

* fix(#4699): skip already-complete phases in the next_phase cascade

Both next-phase scans selected the numerically lowest phase above N
without consulting completion state, so completing a reopened phase
persisted an already-[x] phase as STATE.md current_phase while
roadmap.analyze correctly named the outstanding one (issue repro:
completing 2 with phases 1 and 3 already [x] returned next_phase 03).

The cascade collects the complete phase numbers from the roadmap
checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in
both the disk scan and the roadmap scan; a [x] checkbox row and its
heading sibling both name a phase that is never next. Heading-only and
checkbox-less roadmaps behave exactly as before.

* test(#4699): align the negative-control expectation with the disk spelling

* test(#4699): pin the STATE.md persistence and the all-later-complete tail corner

Review findings: the regression never asserted STATE.md current_phase
(the issue's actual harm), and the all-later-phases-[x] corner
(is_last_phase true, next_phase null) was unpinned. A changeset fragment
is included.

* docs(#4699): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 04:17:39 -04:00
Tom Boucher
2bfff17ff8 fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing

* fix(#4682): route stale verification to the verifier regeneration path

The stale routing entry sent users to /gsd-verify-work — but verify-work
never rewrites VERIFICATION.md (its only write is the human_needed
canonicalization), so following the advice re-ran UAT, reached the same
stale check, and looped. init's projector and execute-phase's generic
next_command presentation both mirror this entry, so the dead end appeared
on three surfaces.

The stale entry now routes to execute-phase, and execute-phase's
all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping
the spine under its frozen ADR-857 ceiling) mirroring the missing route:
skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at
aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier —
regenerating VERIFICATION.md and its digest, marked phase or not. The
non-stale fall-through. Staleness detection, the digest format (#4623),
every other routing entry, and the #3684 resume arms are untouched.

Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682)
Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682)

* test(#4682): register the stale-reverification part and align projected commands

The new steps/ part must be registered in the inventory manifest and the
per-runtime golden install trees (regen:derived); the projected stale
next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the
runtime surface), the human_needed bare-report probe keeps routing to
verify-work (unchanged semantics), and init-manager's recommended action
follows the new command.

* test(#4682): prefix the remaining stale routing assertions with the runtime surface

Nine stale next_command assertions and the human_needed bare-report probe
still carried the unprefixed or flipped forms from the earlier line-number
edit; all now assert the shipped /gsd-execute-phase <phase> projection,
with the human_needed probe reverted to its unchanged verify-work routing.

* test(#4682): align the last stale projection assertions with the execute-phase route

* docs(#4682): backfill changeset PR number

* test(#4682): refresh the compact-content baseline after the rebase

The rebase onto the #4670 squash brought verify-work.md's bounded
reconciliation text into this branch; the committed compact-content
baseline now reflects the post-rebase split sizes. Local --check is
clean; the previous bench drift (+243) was the baseline, not the diff.

* fix(#4682): carry the response_language directive in the stale-reverification part

The new steps/ part is its own coverage unit for lint-response-language-coverage;
it takes the shared canonical directive line like its sibling execute-phase
parts.

---------

Co-authored-by: sim <sim@local>
2026-09-17 02:23:03 -04:00
Tom Boucher
651511d1e3 fix(#4670): bound the commit-claim window to the plan's own history (#4813)
* test(#4670): add failing-first coverage for the bounded commit-claim window

* fix(#4670): bound the commit-claim window to the plan's own history

The reconciliation measured plan_head_before..HEAD — a window that grows
with every later plan's task and SUMMARY commits plus execute-phase's own
phase-completion commit — so an honest plan flagged commit_claim_mismatch
as soon as anything landed after it (real project: claims 3/5/2 measured
20/10/5).

The executor now also records plan_head_after (HEAD at its measurement
moment, after the last task commit, before the SUMMARY commit), and
verify-work reconciles exactly against plan_head_before..plan_head_after
with a merge-base ancestry check; SUMMARYs without the anchor fall back to
the legacy warning path instead of an unsound BLOCKER. Both #3968 failure
modes (claimed commits never made; task commits lost) still block, driven
by the issue's own fixture scenarios.

Emitted-Drift-Ack-Growth: verify-work.md — reconciliation gains the bounded window and legacy fallback (#4670)
Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* fix(#4670): name the history-rewrite case and keep the executor under its cap

Review findings: the BLOCKER enumeration named only the two #3968 causes,
so an honest plan whose recorded window was rewritten afterwards (rebase,
amend, cherry-pick) got a mislabeled diagnosis — the clause now names that
case with the manual-recount remedy. The executor's growth crossed the
LARGE-tier hard cap (49152), so the plan_head_after documentation is
compressed to the minimal capture + frontmatter write (verify-work.md
carries the semantics), the #2751 PROSE_ALLOWLIST entry is re-pointed at
the shifted line (#4670 moved it from 823 to 825), the compact-content
baseline is regenerated, and the changeset records the two un-established
edges (shared-base waves, subrepo ledgers).

Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* docs(#4670): backfill changeset PR number

* docs(#4670): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 22:06:03 -04:00
Tom Boucher
652796e903 fix(#4665): route --fix past the empty-scope exit (#4810)
* test(#4665): add failing-first contract coverage for the --fix empty-scope recovery

* fix(#4665): route --fix past the empty-scope exit

check_empty_scope exited the entire workflow whenever REVIEW_FILES was
empty — before dispatch-fix — so with #3661's incremental scoping, a phase
whose only post-review changes were planning artifacts could never run
--fix against its standing REVIEW.md findings, and the skip output did not
even mention the flag.

The skip is now a self-contained guarded fence (explicit REVIEW_FILES
emptiness check): it fires only when --fix is absent OR the phase's
REVIEW.md does not exist. Otherwise the workflow proceeds directly to
dispatch-fix, which delegates to code-review-fix.md — the canonical fix
implementation that already documents handling an existing REVIEW.md —
while the fresh-review steps (structural pre-pass, reviewer lanes,
spawn_reviewer, commit_review) are skipped: nothing new to review, nothing
to commit. dispatch-fix.md's route docstring is synced.

Emitted-Drift-Ack-Growth: code-review.md — check_empty_scope gains the --fix recovery branch (#4665)

* docs(#4665): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 19:09:19 -04:00
Tom Boucher
85545a77a5 fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate

verify-work.md's complete_session step flips VERIFICATION.md to passed on
'zero issues' alone, so a session whose every UAT row is blocked (a session
that observed nothing) canonicalizes the report. Pins the deployed contract
the fix must satisfy: the flip runs the unflagged phase uat-passed predicate
inside the human_needed branch, frontmatter.set sits inside a passed==true
guard, a refusal message carries the blocker count and keeps
human_needed, and an indeterminate pre-check fails closed. All four new
assertions are RED until the workflow grows the guard.

* fix(#4663): gate the canonicalization on the uat-passed predicate

complete_session flipped VERIFICATION.md to passed whenever the session
recorded zero issues and the status was human_needed — but blocked rows are
not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked
session (one that observed nothing) rewrote the canonical report to passed.
Every later reader (transition.md's preliminary check, resume paths,
validate-phase, verification.status) then inherited the unearned pass while
the phase-close predicate correctly refused it.

The flip now runs the phase-close predicate in a new --uat-only form before
canonicalizing: UAT rows evaluated (at least one pass, no
pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers
skipped — they must be, because the report still reads human_needed at
pre-check time and that status is itself a blocking verification entry, so
the full predicate could never pass there and the flip would deadlock
(found by isolated review, probed). The --require-verification call stays
the transition gate; the refusal branch reports the blocker count and keeps
human_needed; an indeterminate pre-check fails closed.

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* test(#4663): align the canonicalize pre-check needles with the shipped line

The workflow line carries a 2>/dev/null redirect the needles did not
include, so both pre-check assertions fail against the committed fix
(fixed-string grep verified). Reviewer-found; needle and message aligned.

* fix(#4663): reword the canonicalize prose and refresh its size baseline

The rationale paragraph mentioned the flagged transition-gate call by its
flag, putting a --require-verification literal before the first
phase uat-passed occurrence and breaking the existing ordering pin; the
prose now describes it without the literal. verify-work.md's growth also
drifted the committed compact-content baseline; regenerated via
benchmark-compact-content.cjs --write (derived artifact, report-not-gate
contract).

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* docs(#4663): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 17:26:24 -04:00
Tom Boucher
c72fb34e9a fix(#4658): give the ui plan gate's evidence check a native branch (#4807)
* test(#4658): add failing-first coverage for native frontend evidence

hasStaticFrontendEvidence recognised only JS-ecosystem evidence, so
computeUiPlanGate could never block for a SwiftUI/Compose/Flutter/XAML
project. Adds evidence-level fixtures for the four suggested markers
(import-matched for .swift/.kt/.dart, extension-alone for .xaml), the
reporter's non-UI Swift control case, marker-exactness and SKIP_DIRS and
I/O-degrade negatives, gate-level block assertions through makeProject's
new native frontendEvidence modes, and a pinned-seed fast-check property.
All new assertions are RED until src/ui-frontend-evidence.cts grows the
native branch.

* fix(#4658): give the ui plan gate's evidence check a native branch

hasStaticFrontendEvidence recognised only JS-ecosystem evidence (a root
package.json UI-framework dep, or a .tsx/.jsx/.vue/.svelte file), so
computeUiPlanGate could never block for a SwiftUI, Jetpack Compose, Flutter,
or .NET MAUI project — the #3312 gate was structurally unreachable for them.

Adds a native BFS over the same bounds and skip rules: .xaml is evidence by
extension alone (the .tsx analogue), while .swift/.kt/.dart count only when
their content carries the ecosystem's UI import marker (import SwiftUI /
import UIKit, androidx.compose, package:flutter) — matched on the import, not
the extension, so a non-UI Swift package stays silent exactly as the issue's
37-file control case requires. Marker reads are bounded to a 64 KiB prefix;
any I/O failure degrades to false per the module contract. The #3718
vocabulary filter, the JS evidence rules, and the weaker-extension exclusion
are untouched.

* chore(#4658): regenerate the macos conformance tier list

The native-evidence additions to tests/check-ui-plan-gate.test.cjs move the
file into the macOS conformance tier per the classifier; the committed
generated list is a derived artifact and must match the live tests/ tree
(the same sync the fragment-single-edit-propagation install test enforces).

* fix(#4658): accept both Dart quote styles and extract the shared bounded walk

The isolated reviews' remaining findings: the Dart marker carried only the
single-quote anchor, missing legal double-quoted imports (a spec-narrowing
deviation); the two evidence walks duplicated the subtle MAX_WALK_ENTRIES
cap semantics verbatim, so they are extracted into one walkProjectFiles BFS
with a visit callback; tests now use the createTempDir helper, shared
fixture literals that cannot drift from NATIVE_UI_CONTENT_MARKERS, a
double-quoted Flutter import case, and drop a vacuous assertion and a
mid-body re-require alias.

* docs(#4658): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 15:29:19 -04:00
Tom Boucher
5e729445d3 fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel

Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first
proposeConsiderations regression (Danish text + English text_en must classify
as its English equivalent, not land in the #1110 unclassified sentinel),
proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en
validation cases (empty/whitespace/non-string, unconditional under an
elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a
reference-doc Inputs parity test, and a fast-check property proving any
cue-matching prose classifies identically under a cue-free Danish rendering
plus text_en. All new assertions are RED until src/ui-consideration-probe.cts
and the workflow/reference docs are updated.

* fix(#4657): give the ui consideration probe a text_en language channel

Element gains an optional text_en; classifyElement's own signature stays
untouched (a locked, directly-tested export) and the text_en ?? text
selection is pushed to the two classification call sites (proposeConsiderations,
proposeElements) instead. text_en is validated fail-closed: an empty or
whitespace-only value throws rather than silently winning the ?? fallback and
degrading classification to zero kinds.

Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the
Non-English projects section (mirroring spec-phase Step 5.5) and the
ELEMENTS_JSON shape comment documents the field with both zero-applicable
guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT
predicate (with both derived indexes regenerated), and the nav-override
test expectation stay in sync. The ui-phase contract test carries the
site-scoped allow-test-rule marker and its cluster is registered in the
test-file-count allowlist ratchet.

Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657)

* docs(#4657): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 13:46:40 -04:00
Tom Boucher
3014775a3f fix(#4656): expose coverage.unclassified and widen the zero-applicable guards (#4800)
* fix(#4656): expose coverage.unclassified and widen the zero-applicable guards

* fix(#4656): regenerate golden coverage fixtures and update the rollup pin

Emitted-Drift-Ack-Growth: spec-phase.md — #4656: guard widened to the all-unclassified case, doc claim corrected
Emitted-Drift-Ack-Growth: ui-phase.md — #4656: guard widened identically

* fix(#4656): sync edge-probe doc blocks and coverage pins with the new field

* fix(#4656): key the mandatory confirmation on the widened guard

* docs(#4656): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 10:33:32 -04:00
Tom Boucher
caecaec62e fix(#4648): delegate explore seeds to the plant-seed workflow (#4798)
* fix(#4648): delegate explore seeds to the plant-seed workflow

* fix(#4648): delegate explore seeds to the plant-seed workflow

Emitted-Drift-Ack-Growth: explore.md — #4648 consumer wiring: the seed output now delegates to /gsd:capture --seed (plant-seed) instead of hand-writing a divergent, reader-invisible shape

* fix(#4648): plant-seed extracts an idea-stated trigger into trigger_when

Emitted-Drift-Ack-Growth: plant-seed.md — #4648: write-seed sets trigger_when from the idea text when /gsd-explore passes the conversation trigger inline

* docs(#4648): backfill changeset PR number

* docs(#4648): correct changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 08:24:44 -04:00
Tom Boucher
f0a1745e29 fix(#4639): exempt the container --env-file value from the secret-read guard (#4789)
* fix(#4639): exempt the container --env-file value from the secret-read guard

* test(#4639): adopt the suite assertions and pin the reclassified row

* docs(#4639): document and pin the container-env printenv residual

* docs(#4639): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 06:44:55 -04:00
0xdhx
003d982c83 fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint

Two defects in the covered-input fingerprint (#4155), one issue.

1. `computeCoveredDigest` hashed the whole bytes of every declared path
   uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` —
   which every phase rewrites as ordinary bookkeeping, and which the closing
   phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER
   the verifier ran — flipped every phase that declared them to `stale` on
   zero implementation change, and from there `isPhaseComplete` →
   `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate.
   Fingerprint v2 leaves any direct child of a planning root out of the
   hash: `.planning/` itself, plus the phase's own planning root (the parent
   of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/`
   layouts are covered without the digest knowing what a workstream is —
   `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather
   than a name list so the set cannot drift; a root is accepted only when the
   phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated
   exactly as every other covered path (confined, present, a regular file —
   the fail-closed contract is unchanged); only its bytes are ignored, and a
   declaration made only of shared documents fails closed like an empty one.
   A stored digest names its version, and `readVerificationStatus` now
   recomputes under THAT version (`parseFingerprintVersion`,
   `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics
   until it is re-fingerprinted, so the upgrade alone stales nothing; a
   version this build cannot recompute fails closed.

2. `verification.fingerprint` received a raw positional slice, so
   `--files a`, `--files "a,b"` and `--files a --files b` all put the literal
   token into the covered set and failed closed as "a covered file is
   missing, unreadable, or escapes the project root" — the message that
   convinced the reporting project the digest was permanently
   unrecomputable. `parseFingerprintFileArgs` accepts every form (plus
   `--files=a,b`, freely mixed with bare positionals), treats any other
   `--flag` and an empty `--files` value as usage errors that say so, and
   the phase-dir argument must now be an existing directory: omitting it
   used to take the first covered file as the phase dir and print a
   plausible digest over the rest at exit 0.

Regression tests (tests/verification-status.test.cjs, #4623 block): the
cross-phase case from the report, the same-phase `requirements
mark-complete` / `phase.complete` cases from the thread, a workstream-scoped
root, v1-preserved / unknown-version-stale, the fail-closed cases (missing,
directory, escaping symlink, all-shared), every `--files` form against the
bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and
AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34
fail, the 7 that pass pin behaviour the fix must leave unchanged.

Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's
covered_files instruction (rewritten in place — the file sits 21 bytes under
its LARGE hard cap), gsd-core/templates/verification-report.md.

Fixes #4623

Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi

* chore(#4623): set changeset fragment pr to 4749

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 05:26:44 -04:00
Dennis Alexis Valin Dittrich
ad1477d659 enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap

* enhance(260903-m7p): validate configured entrypoints before success

* test(260903-m7p): require pre-success entrypoint validation

* enhance(260903-m7p): gate install success on entrypoints

* test(260903-m7p): cover configured entrypoints across runtimes

* enhance(260903-m7p): cover emitted runtime entrypoints

* fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number

- finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process
  before the new configured-entrypoint assertion throws. Without a HOME +
  config-location-env sandbox that write resolved through the ambient
  environment and landed in the developer's live ~/.gsd (confirmed absent on
  origin/next baseline, present only on this branch — full-suite HERMETICITY
  WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the
  duration of the test, matching the existing in-process finishInstall/
  install() pattern in tests/install.test.cjs (#2665).
- .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder
  (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the
  fork PR number until the upstream PR number is known.

* fix(260903-m7p): repair cross-platform and pre-existing shape fallout

- tests/configured-entrypoint-validation.test.cjs: the win32 branch of
  ensureCodexHooksJsonSessionStart writes a .cmd shim under
  <codexRoot>/hooks/; create that dir in the test (the real installer only
  calls this once hooks/gsd-check-update.js already exists) and assert the
  platform-common entrypoint shape instead of a fixed non-Windows array,
  since win32 legitimately emits two entries (cmd shim + script).
- tests/install.test.cjs: finishInstall's shared settings-json return now
  carries configuredEntrypoints/rollbackInstallerMigrations for every
  runtime on that path (trae included, not just Claude/Cursor/Windsurf);
  update the trae install() exact-shape assertion to match.

* fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash

configuredEntrypointsForHook's shell branch dropped interpreterCandidates
entirely when resolveBashExecutable returned null, unlike the sibling
portableHooks runner entry a few lines below (which correctly falls back to
the literal 'bash' token). Found via agy adversarial review; verified
unreachable through the current call graph (buildHookCommand's own
resolveBashRunner==null gate already short-circuits before
recordConfiguredHookCommand runs), so this is a defensive consistency fix,
not a live-bug patch — kept for the next caller that does not share that
gate.

* chore(260903-m7p): backfill changeset pr number to the opened upstream PR

.changeset/quick-wasps-sing.md carried the fork PR number (16) as a
placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now
open, so record its real number per CONTRIBUTING.md's changeset pr-field
convention.

* fix(#4154): track already-registered hooks for entrypoint validation on update

applySettingsJsonHooks registers each guard hook only if absent, so a hook
already present from a prior install keeps its stale on-disk command. The
new entrypoint tracker always records the freshly-computed command for it,
which never matches what is actually persisted, so the exact-string filter
in finishInstall silently dropped it from validation — the Blocker case
this feature exists to catch (an already-installed entrypoint going stale
between installs) was exactly the case it never validated.

Match on the managed script's basename instead, which the persisted
command carries either way, so an already-registered hook stays in the
validated set. Regression test forces this path by mutating a
freshly-installed hook's persisted command before a second install.

* fix(#4154): distinguish an unreadable script from a missing one

validateConfiguredEntrypoints folded an EACCES statSync failure into the
same 'missing' reason as ENOENT, misreporting a real permission problem as
an absent file. Check the error code and report 'unreadable' instead.

* docs(#4154): document entrypoint validation's rollback and PATH scope

CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had
no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite
bin/install.js x CONTEXT.md being this repo's strongest co-change pairing.

The update-gsd.md how-to overstated what a validation failure undoes: for
Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/
config.toml inside install() before the aggregate validation call runs, so
there is no rollback path for that write regardless of "where available"
phrasing. Also note that interpreter resolution checks the installer's own
PATH, not necessarily the PATH a hook fires under later (#2979 launchers).

* chore(#4154): point changeset pr field at the fork PR while CI runs there

Mirrors the branch's own prior backfill commit: pr: matches whichever PR
number changeset-lint is currently validating against (fork PR #16 during
the fork-first CI/review loop), flipped back to the upstream PR number
right before the final push to open-gsd/gsd-core.

* fix(#4249): address adversarial-review findings in entrypoint validation

An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR
found several real gaps beyond the human reviewer's Blocker, verified
against source before fixing:

- Codex's install() result bound rollbackInstallerMigrations to the narrow
  installer-migrations-only rollback instead of restoreCodexSnapshot (#3245),
  the full pre-install snapshot/restore Codex already owns for exactly this
  case — a validation failure discovered outside install() reverted nothing
  of the config.toml/hooks.json that call had already written.
- The register-only-if-absent basename match from the prior fix used a bare
  substring, which an unrelated user command mentioning the same filename
  could false-positive into GSD's validated set — anchored on the
  `/hooks/<basename>` path segment instead.
- nodeCandidates checked raw process.execPath (always true — we're running
  in that process) instead of normalizeNodePath's stable version-manager
  alias, the same one buildNodeRunnerChainToken bakes as its first choice —
  a false green regardless of whether that alias itself still resolves.
- An entry with no interpreterCandidates (Cline's PreToolUse hook, or a
  Windows-Claude .sh hook invoked without a bash runner) runs via its own
  shebang; validateConfiguredEntrypoints checked only file-type, never the
  execute bit. Cline's writer also never reported an entrypoint at all.
- Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor
  hook registered across several events) were validated once per duplicate.

Each fix is covered by a new or extended test; the Codex one required
inlining runCodexInstall's env sandboxing so the rollback closure — which
re-resolves the $HOME-relative skills root live — runs before the sandbox
is torn down, matching how installAllRuntimes' real aggregate gate calls it.

* docs(#4249): document the round-2 entrypoint-validation fixes

Runtime Hooks Surface Module and Installer Module entries now name
ConfiguredEntrypoint's not-executable reason, the normalizeNodePath
alignment, Cline's tracked hook, and which install() result the
finishInstall/installAllRuntimes rollback path actually reverts per
runtime (Codex's full snapshot vs. the others' narrow migrations-only
rollback).

* chore(#4249): point changeset pr field at the upstream PR now that fork CI is green

* fix(#4249): address agy adversarial-review findings

- validateConfiguredEntrypoints: statSync alone never detects a
  chmod-000 script (it only needs parent-dir search permission), so an
  interpreter-invoked entry with an unreadable script passed validation.
  Add an explicit R_OK check for the interpreterCandidates branch only —
  the candidate-less/shebang branch already has its own X_OK gate.
- docs/how-to/update-gsd.md: the blanket "does not revert" claim was
  false for Codex, which reverts config.toml/hooks.json via its full
  pre-install snapshot; qualify it per runtime.
- tests/codex-config.test.cjs: the #4249 rollback regression test
  asserted skills/ and VERSION were reverted but never asserted
  config.toml/hooks.json were too, despite the test's own stated intent.
- CONTEXT.md: qualify which interpreterCandidates entries get
  normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's
  Windows .cmd shim as a third candidate-less case that relies on
  extension dispatch, not a shebang.

* fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit

Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node',
so it needs the execute bit (like any shebang-invoked entry), but its
interpreter is looked up on PATH by 'env' at hook-fire time (unlike
every other GSD JS hook, which bakes an absolute node path specifically
to avoid that dependency). The candidate-less/interpreterCandidates
fork treated these as mutually exclusive, so Cline's entry silently
skipped interpreter resolution entirely — a completely missing 'node'
on PATH would still validate successfully.

Add an orthogonal selfExecutable flag so both checks run for entries
that need them. (CodeRabbit finding on the fork rehearsal PR.)

* fix(#4249): address second-round adversarial review findings (opus + agy)

- validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry,
  not just interpreterCandidates ones — a self-executable shebang script
  is still opened and read by its kernel-invoked interpreter, so X_OK
  alone never proved it was readable.
- selfExecutable is now the sole, explicit source of truth for the
  execute-bit check (every producer that needs it sets the flag) instead
  of being partly inferred from an absent interpreterCandidates, which
  Cline's hybrid entry also carries.
- The execute-bit check now skips explicitly on win32 (matching
  resolveExecutableBinary's own carve-out) instead of relying on Node's
  accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows
  machine and not a test that simulates win32 on a POSIX runner.
- bin/install.js: fixed a stale comment claiming no runtime's
  install()-time writes have a rollback path — Codex's does
  (restoreCodexSnapshot) — and added the omitted Cline to both that
  comment and CONTEXT.md's equivalent lists.
- CONTEXT.md: fixed the Cline description left stale by the previous
  commit's selfExecutable addition, and rewrote the validation-mechanism
  paragraph for clarity (writing-for-agents pass).
- docs/how-to/update-gsd.md: split an overloaded 4-clause sentence.
- Removed a fault-injection integration test that could not reliably
  exercise the real installAllRuntimes -> finalize -> rollback wiring
  without fighting the installer's own pre-registration existence
  guards; the constituent pieces remain covered individually.

* fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner

X_OK is a POSIX-only concept, skipped entirely when an entry's platform
is win32 (matching production). Two test entries omitted platform,
defaulting to process.platform — on an actual windows-latest CI runner
that silently skipped the very check they were meant to exercise,
turning 'not-executable' into a false pass. Pin platform: 'linux' so
these are deterministic regardless of which OS runs the suite.

* fix(#4249): classify EPERM the same as EACCES in statSync error handling

Windows raises EPERM (not EACCES) for a parent directory that couldn't
be traversed into — was falling through to 'missing', misreporting a
genuine permission problem as a nonexistent path.

* docs(#4249): address final CodeRabbit doc-completeness findings

- CONTEXT.md: install()'s documented result shape omitted
  configuredEntrypoints; the ConfiguredEntrypoint shape omitted
  selfExecutable.
- docs/how-to/update-gsd.md: the failure-mode sentence omitted
  unreadable and lacks-execute-permission, which the installer also
  rejects.

* fix(#4249): stop double-validating every configured entrypoint on install/update

installAllRuntimes' finalize() already runs assertConfiguredEntrypoints
once over the aggregate set; finishInstall then re-ran the identical
check per runtime in the printSummaries loop right after, so every
entrypoint paid its statSync/accessSync/interpreter-resolution cost
twice on every install and update. Add entrypointsAlreadyValidated to
skip the redundant pass specifically on that path, while leaving the
check intact for any caller that invokes finishInstall directly.

* chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there

* perf(#4249): memoize interpreter candidate resolution across entrypoints

resolveExecutableBinary walked PATH once per (entry, candidate) pair; a
typical install has a dozen-plus entries sharing the same few candidate
lists (process.execPath for JS hooks, bash for shell hooks). Cache by
(platform, candidate) so each distinct pair resolves once per validation
call instead of once per entry.

* chore(#4249): point changeset pr field at the rebased rehearsal fork PR

* fix(#4249): drop entrypoint tracking from the now-dead Codex event writer

#2586 (landed on next after this branch forked) removed install.js's
CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent
no longer runs during install or update. The ConfiguredEntrypoint records
this branch added inside it were therefore unreachable and untested. Restore
the function to its upstream shape; the entrypoints it used to report were
never collected by any caller.

* refactor(#4249): drop the revalidation bypass flag and the candidate cache

Both were this PR's own micro-optimisations over a set of roughly a dozen
entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall
gate off to save one statSync/accessSync pass; `resolvedCandidateCache`
memoised resolveExecutableBinary across entries that are already deduped by
(configPath, scriptPath). Neither is measurable, and the flag was the only
way to reach finishInstall with validation disabled. finishInstall now
always validates what it is given.

* chore(#4249): point the changeset pr field back at the upstream PR

* refactor(#4249): track settings.json entrypoints without the hooksSurface gate

The install-surface writer only tracked configured entrypoints when the
runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing
asserts that axis agrees with `installSurface`, so a descriptor that broke
the coupling would silently pass `configuredEntrypoints: undefined` and drop
that runtime out of the validation this PR adds — reintroducing the exact
'reports Done! over a broken entrypoint' failure #4154 exists to close.

Remove the dependence rather than test it: everything recorded on this path
lands in settings.json by construction, and the registered-command filter
already discards entries no persisted hook references.

* chore(#4249): put the changeset body in the documented two-part format

CONTRIBUTING.md and .changeset/README.md both show
`**<bold change>** — <symptom-led explanation>.`; the fragment was a single
unbolded sentence.

* chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there

* fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback

#3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-*
and gsd-core/VERSION. The install overwrites every other GSD-owned file too —
hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest
itself — before the entrypoint-validation gate runs, so a validation failure
left the new payload sitting on top of the restored old config.

Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims,
before runInstallerMigrations so the bytes are the true pre-install state, and
restore it from both Codex rollback closures ahead of the per-surface restores.
Files only the failed install introduced are removed, read from the manifest
now on disk. The manifest is already the authoritative record of what GSD owns,
so no second hand-written list can drift out of sync, and user-owned files are
never snapshotted or removed. Every path is confined through
resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into
an arbitrary-path write.

Non-Codex runtimes are unaffected: the snapshot is gated on the same
tomlConfigInstall + non-minimal condition as #3245's.

* fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest

Two follow-on defects in the previous commit's snapshot:

- The capture was gated on `!isMinimalMode`, copied from #3245. A core/
  --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the
  manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the
  snapshot came back empty while the rollback still ran — and its removal pass
  would have deleted every file the new manifest lists. Gate on
  tomlConfigInstall alone, matching where the rollback actually reaches.

- An unreadable or unparseable prior manifest was caught alongside ENOENT and
  treated as a fresh install. That is the same empty-snapshot state, so a failed
  update over a real install with a corrupt manifest could delete its prior
  payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means
  known-empty; any other read error or a parse failure means unknown, and the
  restore closure returns without touching anything, degrading to #3245's
  narrower rollback. Deliberately not fatal — a corrupt manifest has to stay
  repairable by reinstalling over it.

Both paths are covered by red-checked regression tests.

* fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too

commit removed from the manifest snapshot. restoreCodexSnapshot is reachable
for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-*
skill dir and gsd-* agent file the snapshot does not claim — so with an empty
minimal-mode snapshot a rollback deleted the whole skills/agents surface with
nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via
the ADR-1239 skills-kind home override, so this is also the reason manifest
`skills/` keys do not resolve under configDir: that surface belongs to this
snapshot, not to the manifest-driven one.

Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal
mode — doing nothing on an early failure is the non-destructive side.

Covered by a red-checked regression test that plants bytes in an alternate-home
skill file, reinstalls under the core profile marker, and asserts the rollback
restores it.

* fix(#4249): never remove on rollback unless a prior manifest proves what predates the install

Three defects in the manifest-driven Codex rollback, all in its removal half:

- ENOENT marked the snapshot usable, arming the removal pass on a FIRST
  install. GSD may have overwritten a user's file at a manifest-tracked path
  there, and no prior manifest records the difference — so rollback deleted it
  where before it merely left it overwritten. Absent, unreadable and malformed
  manifests now all leave the prior set UNKNOWN and skip removal entirely.

- Membership was tested against the map of files whose pre-install read
  SUCCEEDED, so a tracked file that existed but was unreadable read as
  introduced-by-this-install and was removed. Track the prior manifest's paths
  in their own Set and test against that.

- The unreachable "delete the manifest when there was no prior one" branch is
  gone: usable now implies a parsed prior manifest.

Also adds the end-to-end test the aggregate gate was missing — the four Codex
rollback tests drove the closure directly, proving the restore but not the
wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes
Cline's `env node` entry fail validation for real, and asserts Codex's payload
comes back.

Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper
instead of six copies. Both new tests are red-checked.

* test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test

lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its
Windows-EBUSY retry budget. That budget is for directory trees; this removes a
single file, which unlinkSync says more precisely and the rule does not flag.

* chore(#4249): point the changeset pr field back at the upstream PR

* fix(#4249): use an unambiguous dedup key and surface partial-restore failures

trek-e's 2026-09-08 adversarial pass flagged two findings in the new
entrypoint-validation/rollback code:
- assertConfiguredEntrypoints' dedup key already used a raw NUL
  separator (introduced in ceebb65f2d), but git/Read render NUL as a
  space, so the key looked like a plain-space join to every reviewer
  that read the diff. Replace it with JSON.stringify([configPath,
  scriptPath]) so the separator is visible and unambiguous.
- restoreManagedFileSnapshot's per-file restore catch block claimed to
  'surface the original error' but only swallowed it, matching (and
  widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add
  an actual console.warn using the existing best-effort-warning
  convention, scoped to just this PR's new function.

* fix(#4249): treat a files-less prior manifest as unknown, not known-empty

agy's gemini-3.8-flash-high adversarial pass (round 5) found and I
reproduced empirically: a structurally-valid manifest missing the
files key (e.g. {"version":1}) parses without throwing, so
Object.keys(undefined || {}) silently read as 'zero files predate
this install' instead of the UNKNOWN state the malformed-manifest
guard exists to produce. Rollback's removal pass then deleted every
GSD-owned file the failed install's own manifest listed, including
ones that predated it — the exact data loss the #4249 CodeRabbit
malformed-manifest fix was supposed to prevent, reachable through a
JSON.parse success instead of a failure. Route the shapeless case
into the same catch-all UNKNOWN path via an explicit shape check.
Regression test reproduces the deletion before the fix and confirms
the file survives after it.

Also extend restoreManagedFileSnapshot's removal-pass rmSync and
final manifest-rewrite catches with the same real console.warn
trek-e's round-4 review asked for on the per-file restore catch —
same rollback function, same operator-facing-signal gap.

* docs(#4249): correct which runtimes actually leave a written config on rollback

agy's completeness audit (round 5, holistic pass) caught this new
paragraph claiming 'for every other runtime, the configuration file(s)
already written during that update are left in place' — false for
Claude Code and other settings.json-based runtimes, whose write never
happens on failure (assertConfiguredEntrypoints runs before
finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline
actually match that description, since they persist their config file
inside install() ahead of the gate. Split the one sentence into the
three actual outcomes; matches the PR body's own accurate Before/After
wording, which this doc addition had drifted from.

* fix(#4249): clean up doc/comment mismatches and dead fields from opus review

Opus critical-code-reviewer + ponytail-review pass on the final diff:
- assertConfiguredEntrypoints carried finishInstall's old docblock
  ("Apply statusline config, then print completion message") from
  before this function was inserted between comment and callee.
  finishInstall already has its own accurate #4249 comment, so the
  stale docblock is removed rather than moved.
- checked: number on ConfiguredEntrypointValidationResult and
  error.configuredEntrypointValidation on the thrown error: the first
  had zero consumers anywhere in the repo, including its own defining
  file, and is removed. The second matches an existing repo
  convention (bin/install.js's installerMigrationRollbackFailures,
  #4249 predates this PR) of attaching structured diagnostic context
  to a re-thrown Error even before a consumer exists, so it's kept.
- finishInstall's own assertConfiguredEntrypoints call is a redundant
  backstop on the real production path (installAllRuntimes's aggregate
  call already validates the superset first), but its comment read as
  though this call alone provided the before-the-write guarantee.
  Clarified rather than removed — it's the only gate for a caller that
  invokes finishInstall directly.

* chore(#4249): split the manifest-driven rollback engine out into #4544

Issue #4154 asked the installer to consume a validation failure "through
the existing rollback mechanism, without a second transaction mechanism".
The manifest-driven rollback widening added during review (capture every
path the prior gsd-file-manifest.json claims, restore those bytes, remove
what only the failed install introduced) is that second mechanism on a
plain reading. It is a real fix for a #3245-era gap, but an independent
one, so it moves to its own bug report and PR.

Removed here:
- bin/install.js: the pre-install managed-file capture block and
  restoreManagedFileSnapshot, plus its call sites in
  _codexPreConfigRollback and restoreCodexSnapshot (99 lines).
- tests/configured-entrypoint-validation.test.cjs: the five tests that
  exercise the manifest engine.
- CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the
  widened restore. update-gsd.md again documents the #3245 surfaces only.

Kept, because it is #4154's own scope:
- the entrypoint-validation gate itself;
- Codex's install() result binding rollbackInstallerMigrations to
  restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*,
  agents/gsd-*, gsd-core/VERSION);
- the !isMinimalMode gate removal on that snapshot. Binding the closure
  to the result made it reachable for a core/--minimal install, where its
  pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot
  does not claim; an empty minimal-mode snapshot therefore deleted the
  whole surface with nothing to restore.

The surviving aggregate-failure test now asserts on config.toml, a
surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which
only the manifest engine restored.

Refs #4544

* test(#4249): cover configured entrypoints through the packed install path

#4154's scope lists install smoke coverage alongside the installer gate —
"assert representative configured entrypoints resolve for supported runtime
profiles". The gate itself (assertConfiguredEntrypoints /
validateConfiguredEntrypoints) is unit-covered by in-process install() calls;
nothing proved the property survives npm pack -> npm install -g -> install.js.

Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct
config surfaces GSD writes launch paths into (settings.json, and hooks.json +
config.toml) — run the tarball-installed installer into a throwaway HOME, then
re-read that runtime's own written config and return the new
ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a
file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check
becomes a release gate on every matrix host without workflow changes.

The scan re-derives paths from the written config instead of reusing the
installer's own entrypoint list, and test I shows why that matters: a
registration the installer never touched during a run is invisible to the
in-process gate, so the install exits 0 and only reading the config back off
disk catches the dangling launch path.

* ci(#4249): pack a publish-shaped tarball in the install smoke lane

`npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs
build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs
after `npm ci` carries no hook scripts at all — the lane has been smoking a
package that differs from the published one in exactly the artifacts the
lifecycle smoke is supposed to launch.

That went unnoticed because the lane's init runs `--local`, which registers no
statusline and therefore registers no hook whose target is missing. A
`--global` install on the same tarball exits 1 on #4249's own gate
(`gsd-statusline.js (missing)`), which is what the new configured-entrypoint
cycle performs, so without this step the cycle would report INIT_FAILED
instead of checking anything.

Build hooks before packing so the smoked tarball matches prepublishOnly. The
CLI now reports 16 configured entrypoints for claude and 1 for codex instead
of zero.

* fix(#4249): scope Codex's full snapshot restore to entrypoint failures

Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage
exception un-install a Codex install that had already succeeded and already
printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps
the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints`
call.

Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration`
scopes finalize-stage rollback to installer *migrations* ("the executor uses the
journal to restore modified paths"), and this PR's own operator-facing paragraph
in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/
agents/VERSION revert to entrypoint-validation failures specifically ("If a
script is missing, unreadable, ... For Codex, this reverts ..."). The wide
behaviour is also incoherent as a transaction abort: the same doc says Cursor,
Windsurf, Kimi and Cline keep the config they wrote inside install().

Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall
hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install
bytes — on an update, silently downgrading a working Codex install to the
previous version while the user had just been told it was Done.

Select the rollback by error kind instead. `assertConfiguredEntrypoints` already
tags its error with `configuredEntrypointValidation`, so the full snapshot
restore runs for that error (and anything downstream of it, including
finishInstall's per-runtime backstop) and the installer-migrations-only closure
runs for everything else. The codex result now also exposes that narrow closure
as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning
the full restore, so the direct-call contract asserted by
tests/codex-config.test.cjs is unchanged.

Adds a regression test that installs codex+kilo together, injects EACCES on the
Kilo permission write by monkeypatching node:fs (restored in a finally — never
chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps
the bytes the successful install wrote. Verified red against the pre-fix
unconditional path.

Cline cannot host this test: its plan is writesSharedSettings:false +
finishPermissionWriter:null, so its finishInstall performs no write and has no
non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally
(unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded
fs.writeFileSync.

* docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing

CONTEXT.md still described Codex's rollback as an unconditional bind
to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint-
validation failures via rollbackInstallerMigrationsOnly and the
configuredEntrypointValidation error tag. Caught during the round-6
PR body pass.

* fix(#4249): stop rollbackInstallerMigrations meaning its own opposite

Codex's install() result bound `rollbackInstallerMigrations` to
restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the
actual installer-migrations-only closure behind
`rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name
meant the opposite of what it says, and CONTEXT.md had to concede as much
in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure
for every runtime, matching both its name and the meaning it already has
on next, and the snapshot restore gets its own Codex-only field,
`rollbackPreInstallSnapshot`. The selection in
rollbackFinalizedInstallerMigrations collapses to one line and no longer
needs a fallback chain.

Also in this commit, all against the same rollback path:

- Correct the rollbackFinalizedInstallerMigrations comment. It read as if
  the round-6 narrowing prevented any sibling-triggered revert of a Codex
  install the user has already seen "Done!" for. It does not, and is not
  meant to: `wide` is true for ANY entrypoint-validation error from ANY
  runtime, because the aggregate gate is all-or-nothing — an invalid Cline
  entrypoint reverts Codex's snapshot, which
  tests/configured-entrypoint-validation.test.cjs's 'an aggregate
  entrypoint validation failure rolls the Codex install back (#4249)'
  asserts directly. The discriminator is the error's KIND, not which
  runtime owns the failing path. Comment and CONTEXT.md now say that.

- Name the runtime in the "Configured entrypoint validation failed" error.
  ConfiguredEntrypointInvalid already carries `runtime`; the message threw
  it away, leaving an operator of a multi-runtime install unable to tell
  whose entrypoint broke — which matters precisely because the failure can
  revert a runtime that was itself fine.

- Set `configuredEntrypoints: []` explicitly on the copilot-instructions
  early return. Every other branch states the key; this one relied on
  installAllRuntimes' `(result.configuredEntrypoints || [])` defence.
  `[]` is correct, not a workaround: every Copilot hook is an inline
  printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no
  GSD-managed script or interpreter to resolve.

No behaviour change beyond the error-message text.

* docs(#4249): narrow the smoke scan's config-surface claim to what it checks

RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is
told to launch is registered in one of settings.json / hooks.json /
config.toml, and that nothing else in a config dir is runtime
configuration. Both halves are false as stated. Cline registers its hook
at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those
names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native
[[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi),
a directory separate from Kimi's own GSD configDir — the same gap
installer-migration 007 already documents as structurally unreachable.

The scan is in fact correct for what it runs against: entrypointRuntimes
defaults to claude + codex, whose launch paths do all live in those three
top-level files. Restate the docstring at that scope, name the two known
out-of-scope surfaces, and warn that adding either runtime to
entrypointRuntimes without teaching scanConfiguredEntrypoints about its
surface yields a scan that finds zero entrypoints and proves nothing.
The entrypointRuntimes default comment carried the same overgeneralization
("every other runtime reuses one of them") and is corrected with it.

Documentation only; no code change.

* fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json

An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found
that install()'s settings-json early return for an unparseable
settings.local.json omitted configuredEntrypoints and
rollbackInstallerMigrations from its result, unlike every other branch.
rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations
unconditionally, so this branch silently dropped its own installer-migration
rollback on a later finalize-stage failure.

Completed the return shape: configuredEntrypoints: [] (matching Copilot's
equally-early no-entrypoints-yet return) and rollbackInstallerMigrations
(already in closure scope). Red-then-green regression test added.

* test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs

Same internal adversarial review (round 8): this suite's before() packed the
tarball directly, without the ensureHooksDist() guard every sibling
install-test suite (install.test.cjs, install-minimal-hooks.test.cjs,
mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run
in isolation ahead of a suite that builds hooks/dist itself, this suite's
pack would ship a tarball with no hook scripts and fail closed on
SMOKE.INIT_FAILED instead of testing anything.

* fix(#4249): refresh stale test-timings weight for the codex-config split

next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's
heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the
CI shard packer's weight table (tests/test-timings.json) was never updated:
codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the
suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block
this PR extends with its own #4249 install()-pipeline test — had no entry at
all, so the packer would silently underestimate it at the table's median
weight (roughly a 9x underestimate against its real cost).

trek-e's most recent review flagged a Windows shard timeout in-flight on
codex-config.test.cjs, plausibly aggravated by this PR's own addition to that
file before the rebase moved it. Re-measured both files locally (node --test
--test-reporter=tap, max of 3 runs, matching the table's own max-across-streams
methodology) and patched just these two entries — not a full regeneration,
which would need real multi-lane CI data this session doesn't have access to.

* fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists

next's platform-conformance-tier classifier (#4591/#4598) landed after this
branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and
tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's
gen-platform-conformance-tier --check and both the Linux and macOS conformance
suites.

* fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation

next's #4733 (landed after this branch's last rebase) replaced the static
ISOLATED_HEAVY_FILES set with a threshold derived live from
tests/test-timings.json, and pins the current derived result in
EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still
named codex-config.test.cjs, whose own weight this PR already dropped from
127783ms to 189ms (after splitting its heavy install()-pipeline blocks into
codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The
live-computed set correctly no longer includes it; the pinned expectation is
updated to match.

* fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it

trek-e's review flagged two Major gaps: the thrown error read identically
regardless of which of three real outcomes a runtime hit (nothing
persisted / snapshot reverted / config left broken on disk), and no test
proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/
Kimi/Cline — only Codex's revert path was ever asserted.

assertConfiguredEntrypoints now tags each invalid entry with its actual
consequence, mirrored from docs/how-to/update-gsd.md's existing
rollback-matrix disclosure. A new test drives the same aggregate failure
through Cline (whose own entrypoint is the one that fails) and asserts
its hook file is still on disk afterward.

* fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR

One review pass (gemini-3.8-flash-high via the antigravity review lane)
against this PR's full diff against next, findings independently verified
against source before fixing:

- Copilot's install() return object was the only one of 6 runtime branches
  missing rollbackInstallerMigrations — reachable now that this PR's own
  aggregate gate runs rollback across every result on any runtime's
  entrypoint failure, not just Copilot's own.
- buildHookCommand's unresolved-bash early return skipped track() entirely,
  so a win32 install with no Git Bash silently produced an unregistered
  .sh hook instead of the 'unresolved-interpreter' validation failure
  configuredEntrypointsForHook's own comment said it would.
- release-tarball-smoke.cjs reported a Cycle 4 install failure under
  SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing
  SMOKE.INSTALL_FAILED.
- SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's
  trailing args, which also truncated any configDir containing a space
  (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan.
  Anchored the match on the already-known configDir prefix instead of a
  generic absolute-path guess: removes the ambiguity outright rather than
  patching the character class, and stays a raw-text scan on purpose (it
  catches a writer that emits a path without registering it — a
  JSON.parse of the expected schema would miss exactly that case).

One suggested finding (test-timings.json "missing" the new test file) was
verified false — that table only holds measured CI timings, populated
after a file's first real run — and one Ponytail suggestion (a
JSON.stringify dedup key) was rejected as it would reintroduce a real, if
narrow, key-collision risk for no benefit.

* fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment

changeset-lint requires pr: to match the PR it runs on (16 on the fork,
not the eventual upstream number) — rehearsal-branch convention already
established earlier in this PR's history.

lint-docs-guard-registration's quote-pairing heuristic doesn't require the
docs/ path itself to be quoted — it flags a file once ANY quote-delimited
span containing "docs/" appears anywhere in it, alongside any real fs read
call. A comment ending "...update-gsd.md's rollback-matrix paragraph"
supplied the closing quote character (the possessive apostrophe) the
heuristic paired with an unrelated single-quoted string earlier in the
file. Reworded to avoid the unquoted apostrophe next to the path.

* chore(#4249): point the changeset pr field back at the upstream PR

Fork rehearsal (PR #16) is green; the real target for this changeset is
upstream PR #4249.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 04:23:53 -04:00
Michel Moreira
49f313d611 fix(#4461): make code-review summary extraction shell-safe (#4533)
* fix(#4461): make code-review summary extraction shell-safe

Emitted-Drift-Ack-Growth: code-review.md — use a literal heredoc for shell-safe SUMMARY parsing

* chore: add changeset for #4533

* fix(#4461): keep heredoc outside command substitution

* chore: rerun CI after Windows timeout

* test(#4461): execute the summary heredoc adversarially

* test(#4461): normalize adversarial paths for Git Bash

* test(#4461): keep adversarial fixture valid on Windows

* test(#4461): scope heredoc regression claim

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 03:23:11 -04:00
Tom Boucher
740ba0d8a3 fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file

Co-authored-by: sim <sim@local>
2026-09-16 02:38:43 -04:00
Behruz Nassre Esfahani
febe6c9885 fix(#4685): a directory artifact fails its own entry instead of aborting the check (#4735)
* fix(#4685): a directory artifact fails its own entry instead of aborting the check

`must_haves.artifacts` entries are read with `safeReadFile`, which rethrows every
errno except ENOENT. A listed path that is a directory therefore threw EISDIR out
of the per-artifact loop: `query verify.artifacts` printed

    Error: EISDIR: illegal operation on a directory, read

and reported NOTHING — not the offending entry, and not the plan's other,
perfectly checkable artifacts. One directory entry disabled the whole plan's
check. Reproduced against a real plan before the fix, and after.

A directory is now reported as that entry's own failure, with an issue distinct
from `File not found` (the path did resolve; it simply is not the thing an
artifact entry can be checked against), and every other artifact in the plan is
still checked and reported independently. Anything else the stat or read throws
becomes that entry's failure too, carrying its errno, rather than discarding the
run — a check that disappears is worse than one that fails, because a failure is
visible.

Verifying directories properly — matching `contains:`/`min_lines:`/`exports:`
across the files inside one — is a feature decision and deliberately not made
here, per the issue's stated scope. Authoring-time rejection of a directory path
is likewise left alone: the brief raises it as a separate question, and the
runtime fix does not depend on it.

Also, found in pre-PR review and pre-existing: `safeReadFile(...) || ''` turned a
post-stat ENOENT into empty content, so an artifact declaring only
`path`/`provides` had no criterion left to fail and passed, having checked
nothing. A null read now reports that instead of inheriting a pass. It can only
turn a false pass into a failure.

Verified: reverting src/verify.cts to the merge-base turns both new rows red with
the exact EISDIR message; lint:ci exit 0; full suite 24/24 chunks, 37,280 tests,
0 failures; tests/verify.test.cjs 219/219.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* chore(#4685): backfill changeset PR number to 4735

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* test(#4685): pin the injected-I/O branches, and narrow the guard's comment

Review findings from #4735.

Major — the two error branches this PR adds were untested, and the PR body
claimed the mid-check ENOENT window was "not deterministically reproducible
through the CLI seam these tests drive." That was wrong: ADR-3574 records this
repo's convention for exactly this — inject filesystem failures by
monkeypatching the fs method, never by chmod or mode-bit tricks, which root
bypasses and yields a test that passes with zero coverage in root Docker and CI.

Both branches are now pinned that way. The injection runs in the CHILD via
NODE_OPTIONS=--require, because `output()` writes fd 1 directly
(`writeAllSync(1, …)`, io.cjs) rather than through console.log, so an in-process
call cannot have its JSON captured. The preload patches the child's own module
objects, which the compiled code reads at call time.

  - a file that disappears between stat and read now fails as that entry rather
    than passing on empty content
  - a non-ENOENT errno (EACCES) is reported as that entry's failure, carrying its
    code, so an operator can tell a permissions problem from an I/O one

The injection matches the target by path SUFFIX, not string equality: the first
cut compared absolute paths, and a /tmp vs /private/tmp prefix difference
silently disarmed it — the test passed while asserting nothing. A disarmed
injection test is worse than no test, so the reason is recorded at the call site.

Nit — the comment above the try block said the guard "covers what the stat and
read below actually throw", which reads as if the min_lines/contains/exports
checks inside the same block were deliberately guarded too. They are pure string
operations and cannot throw; the comment now says so rather than implying a
guarantee it does not make.

Verified: reverting src/verify.cts to the merge-base turns all three #4685 rows
red — the directory row and both new ones; lint:ci exit 0; full suite 24/24
chunks, 37,389 tests, 0 failures; tests/verify.test.cjs 221/221.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 01:39:52 -04:00
Twisted Fate
ecc508139a fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting (#4755)
* fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting

Planners emit <automated> bodies with the chain operator entity-escaped
(`&amp;&amp;`), and the executing agent reads the decoded (rendered) form.
The grounding probe split the raw text on &&/||/;/newline without decoding
first, so `&amp;&amp;` was cut at its semicolons and a cd-form target
absorbed the trailing `&amp` fragment — an existing directory was reported
missing_dir (blocker), feeding false blockers into the revision loop.

Decode `&amp;` → `&` inside resolveVerifyCommandTarget, after
result.command captures the text verbatim and before any segment splitting
or target resolution, so the escaped and literal forms of the same command
produce identical verdicts. Module-private helper beside splitSegments,
mirroring the sibling src/verify.cts decodeEntityAmps (#3611); the two
gates are separate modules and neither imports the other.

Regression coverage pins escaped/literal verdict parity for: existing dir
+ manifest (ok), missing dir (missing_dir blocker kept), dir without
manifest (no_manifest blocker kept), --prefix form, a literal & inside a
quoted dir name, and a full probePhaseVerifyCommands pass whose reported
command field stays verbatim.

* docs(#4730): use the documented pr:0 placeholder in the changeset fragment

The fragment carried pr: 4730 — the ISSUE number, the exact guess-shape
DEFECT.CHANGESET-PR-FIELD-DRIFT (#3316, #3325) exists to catch: it parses
as a positive integer so local lint passed, but on a real PR run the
drift check would fail it against the actual PR number. CONTRIBUTING.md
documents pr: 0 as the deliberate unresolved placeholder used during
initial commit before the PR number exists (scripts/changeset/new.cjs
accepts 0 for exactly this reason); the gate's fail_invalid_fragment on
an unbackfilled 0 is the designed backfill enforcement, not a defect.

No production or test changes. Backfill pr: with the real PR number
once the PR is created.

* docs(#4730): backfill changeset pr field with the real PR number

pr: 0 → pr: 4755 (the PR carrying this fix), completing the
documented placeholder workflow; the changeset gate's content
validation can now pass.

---------

Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 00:36:10 -04:00
0xdhx
092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00
0xdhx
25d1cb916f fix(#4721): give worktree cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged (#4766)
* fix(#4721): give cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged

`worktree cleanup-wave` ran `git merge --no-ff` under the module-wide
DEFAULT_GIT_TIMEOUT_MS (10 s) that is sized for plumbing calls. The merge is
the one call in the wave that runs user hooks, so a repo whose
pre-merge-commit hook is a test-suite gate lost every code-bearing executor
merge. Three things went wrong at once, each fixed here:

1. Budget. The merge now passes an explicit timeout —
   DEFAULT_MERGE_TIMEOUT_MS (10 min), overridable via deps.mergeTimeoutMs.
   Every other git call in the wave keeps the module default; the shared
   constant is untouched, because every other caller is exactly what its
   10 s comment describes.

2. Reason. A merge that does time out blocks on `merge_timed_out`, and its
   stderr names the budget and says the hook may still be running, instead
   of `merge_failed` carrying whatever the hook had printed before git was
   killed — which made a healthy executor branch look broken.

3. Residue. A merge killed during its hook has already staged the merged
   tree into the primary's index but never wrote MERGE_HEAD, so
   `git merge --abort` finds nothing and repoRootStillMidMerge (#2852)
   reads the primary as clean while the executor's whole diff sits staged
   against the old HEAD; a `git commit` from that state squashes the
   executor's history into one parent. After any failed merge the wave now
   reads `git diff --cached --name-only`; anything staged is the merge's
   own (git refuses to start a merge when the index differs from HEAD), so
   it runs `git reset --merge` — restores exactly those paths, keeps
   unrelated unstaged edits — and re-reads. Restored paths are reported as
   WAVE_CLEANUP_WARNING.MERGE_RESIDUE_RESTORED and the wave continues; a
   still-dirty or unreadable index reports MERGE_RESIDUE_LEFT_STAGED and
   halts the remaining entries, the same repo-level carve-out an
   unfinished merge takes.

Tests: five mock-driven rows (budget wiring incl. the deps override, the
timeout classification with restore, the no-reset control for an ordinary
refused merge, an unrestorable residue halting the wave, an unverifiable
index failing closed) plus a real-git row that runs a sleeping
pre-merge-commit hook under a 1 s budget and asserts HEAD unmoved, index
and worktree clean, the executor branch intact — with the same fixture
merging cleanly under the default budget as its negative control. Two
existing #2852 rows gain a handler for the new post-failure index read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* docs(#4721): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* test(#4721): release the real-git fixtures with t.after, not try/finally

The two real-git rows cleaned up their scratch repo in a `finally` block;
this file's own convention for fixture teardown is the test context's
`t.after(() => cleanup(dir))`, and the house PR ruleset flags `finally` in a
test body. Behaviour-neutral.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): gate the residue restore on the timeout, re-apply a merge autostash, and correct the hook census

Three findings from the pre-file adversarial review of the previous commit,
each driven on real git before changing code:

1. A merge git REFUSED ("your local changes … would be overwritten") also
   leaves no MERGE_HEAD — and that refusal is exactly what a pre-existing
   dirty primary index earns. The residue restore read that index as the
   merge's own and `reset --merge`d the operator's staged work away
   (driven: a staged edit to an unrelated file was discarded and reported
   as "restored"). The restore now runs ONLY when the merge timed out; a
   refusal is an immediate exit, never a timeout, so on that path nothing
   is read or reset.

2. `merge.autoStash=true` lets a merge start on a dirty index by parking
   the work in MERGE_AUTOSTASH, which a killed merge never re-applies.
   `git reset --merge` moves that stash into the stash list; the wave now
   runs `git stash pop --index` afterwards (the outcome `merge --abort`
   gives an autostashed merge), and reports
   WAVE_CLEANUP_WARNING.MERGE_AUTOSTASH_UNRESTORED (path null) when the
   pop fails or the autostash state could not be read — the work stays in
   the stash, the index is clean, the wave continues. Because of this the
   reset runs on a timed-out merge even when the index reads clean.

3. The merge is not the only hook-running git call in the module:
   `worktree add` runs post-checkout and every ref update runs
   reference-transaction. It is the only call that runs the commit-family
   hooks, which is what the budget is for. Comments and docs say so now.

Tests: the "ordinary merge_failed" control becomes the regression row for
finding 1 (strict mock — a `diff --cached` or `reset --merge` on a refused
merge throws), plus a mock row for the autostash pop (dirty and clean
index, pop success and failure), and two real-git rows: a refused merge
over pre-existing staged work leaves it byte-identical, and a killed merge
under merge.autoStash restores the executor residue AND puts the
operator's staged work back. The real-git hook now sleeps 4 s against a
1.5 s budget for margin on slow runners. The two #2852 handlers added
earlier are removed — the residue read no longer fires on their path.
Negative control: 4 of the 10 #4721 rows fail on the previous commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): key the residue restore on a killed merge, and re-read the index after a failed autostash pop

Two more findings from the continuation review, both driven:

1. An externally delivered SIGTERM leaves the same staged/no-MERGE_HEAD
   state as the timeout, and the seam reports it as exitCode null + signal
   with timedOut false — so the timeout-only gate skipped the restore on a
   state it was written for. The gate is now "killed": timedOut, or a null
   exit code with a signal. A refused merge still exits with a code and is
   still never touched. The reason stays merge_failed for a signal kill.

2. A failed `git stash pop --index` keeps the stash entry but can leave
   conflict entries (UU) and partially applied paths, after which the next
   merge fails on "you have unmerged files"; the code returned halt:false
   on the strength of the pre-pop recheck. The index is now re-read after a
   failed pop and a dirty result halts the wave as merge_residue_left_staged
   alongside the merge_autostash_unrestored warning.

Also driven and now documented rather than changed: a kill that lands once
MERGE_HEAD exists (inside commit-msg) is the ordinary #2852 abort path —
`git merge --abort` restores the tree and re-applies an autostash itself,
unstaged, as git does for any aborted autostashed merge.

Tests: the pop-failure mock row now asserts the post-pop re-read and gains
a conflict-leftover variant that halts; a signal-kill mock row; a real-git
row with the sleeping hook moved to commit-msg (timed out, no residue
warnings, MERGE_HEAD cleared, primary clean). 414 pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): key the kill gate on the seam's signal, not on a null exit code

The shell projection seam normalizes a signal death to exitCode 1 and
carries the signal alongside (`_spawnResult`: `result.status ?? 1`), so the
previous `exitCode === null && signal` gate could never fire in production
and the unit row that covered it modelled a shape the seam does not emit
(caught in the round-3 review). The gate is now `timedOut || signal`; a
refused merge exits with a code and no signal. The mock row uses the real
shape, and a mocked spawnSync signal death driven through the compiled seam
reaches `reset --merge` and reports the residue restored.

Also: three comments that still said "at its budget" / "runs user hooks" /
"the index is clean", and the CLI-TOOLS sentence that reserved
`merge_failed` for refusals and conflicts, now name the signal case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* chore(#4721): set changeset fragment pr to 4766

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 22:28:09 -04:00
Tom Boucher
b5fc8061b3 fix(#4624): persist orchestrator-worktree worker lifecycle records (#4778)
* fix(#4624): persist orchestrator-worktree worker lifecycle records

* fix(#4624): address review findings on the worker lifecycle protocol

* fix(#4624): require the summary path and surface torn records on status --path

* fix(#4624): distinguish no-record from torn-record, tolerate older shims in the sweep

* docs(#4624): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-15 20:25:24 -04:00
Tom Boucher
72be41d96d fix(#4601): reap the whole process tree on run-with-timeout's Windows force stage (#4775)
* test(#4601): pin the run-with-timeout tree reap on win32

* docs(#4601): add changeset for the win32 tree reap

* fix(#4601): reap the wrapped command's whole tree on Windows timeouts

* docs(#4601): backfill changeset PR number

* fix(#4601): keep the mediated shim alive without depending on stdin

---------

Co-authored-by: sim <sim@local>
2026-09-15 18:31:05 -04:00
Tom Boucher
0d6bf19bf1 fix(#4600): explicit --converge overrides the convergence feature gate (#4771)
* test(#4600): pin explicit-flag-overrides-gate precedence

* fix(#4600): explicit --converge overrides the convergence feature gate

PLAN_STRATEGY=converge is set only by an explicit --converge/--cross-ai,
so gating it on workflow.plan_review_convergence made an explicit
operator flag lose to a config default and stop the run with a
question-shaped success. The config remains the default for non-flag
invocation; the flag now wins, and the step says so instead of
gate-and-exit.

* test(#4600): pin the precedence contract sentences as written

* fix(#4600): keep the precedence sentence on one line

The pinned contract phrase wrapped across a line break, so the
writer-contract assertion could not match it.

* fix(#4600): document the flag-overrides-gate precedence on user surfaces

commands/gsd/autonomous.md and docs/COMMANDS.md still said --converge
requires workflow.plan_review_convergence=true; both now state the
override. Changeset typed Changed with the docs update alongside.

* docs(#4600): backfill changeset PR number

* fix(#4600): restore the convergence gate mention in user surfaces

* test(#4600): pin the dispatched convergence run against the config gate

* fix(#4600): override the convergence gate on the dispatched run

Emitted-Drift-Ack-Growth: autonomous.md — #4600: the converge dispatch appends --override-gate inside the PLAN_STRATEGY conditional, and the precedence sentences replace the stale fail-fast instruction
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4600: config gate 1.5 honors an explicit --override-gate dispatch (token-anchored) while the standalone veto and config-get default are preserved

---------

Co-authored-by: sim <sim@local>
2026-09-15 15:53:42 -04:00
Tom Boucher
779f67cb11 fix(#4378): mint collision-free SEED-YYMMDD-xxx seed ids instead of a shared count (#4754)
* test(#4378): regression tests for collision-free seed ids

* fix(#4378): mint collision-free SEED-YYMMDD-xxx ids, not a shared count

plant-seed derived the next seed id from 'ls .planning/seeds/SEED-*.md | wc -l'.
.planning/seeds/ is shared but each worktree only sees what has merged, so two
workstreams planting before either merges computed the same id and git merged
both files silently.

The id is now the local date plus a 3-char random base36 suffix -- the shape
.planning/quick/ already uses -- computed from knowledge one worktree has alone,
with a same-day regen guard. deriveSeedIdentity learns the new canonical grammar
alongside legacy SEED-NNN (whose parsing never changes), the --enrich parser and
the filename-prefix fallback keep the full new-format id, and the docs that
state the filename shape move to it.

The prefix fallback previously truncated any non-pure-numeric id at
'SEED-<digits>' -- the same one-id-two-answers ambiguity the issue reports,
reproduced one level down.

* fix(#4378): harden seed id generation per adversarial review

- parse-idea: anchor the --enrich extractor to the flag and capture the
  complete id, uppercase-tolerant; a leftmost 'SEED-[0-9]+' truncated an
  uppercase or malformed suffix to its date and enriched an arbitrary
  same-day seed via head -1. Ambiguous and unmatched targets now fail
  closed instead.
- generate-seed-id: tolerate the expected SIGPIPE under pipefail, abort
  loudly when the suffix cannot be drawn (an empty suffix would collapse
  every seed's id to the bare date), and run the same-day regen guard as
  a find existence test (the 'ls <glob>' shape trips the #3409 drift
  guard and degenerates under a stray nullglob).
- deriveSeedIdentity: document the theoretical legacy/new grammar
  ambiguity (6-digit counter + 3-char base36 slug, no frontmatter).
- changeset: state the residual same-day collision bound instead of
  implying zero.

Emitted-Drift-Ack-Growth: plant-seed.md — the counting step became hardened date+random generation with explicit failure modes; growth is the failure handling, not duplicated logic

* fix(#4378): address standards and spec review findings

- tests: move the allow-test-rule marker to its suppression site (the
  file-header placement was inert per CONTRIBUTING site-scoping); add
  width-boundary coverage (5/7-digit dates, 2/4-char suffixes pin the
  documented branch behavior); add a writer-to-reader parity property
  that parses the mint widths out of the shipped workflow so the two
  grammar owners cannot drift; cover uppercase ids end-to-end in the
  reader.
- plant-seed.md: draw/retry restructured as one loop with a loud
  terminal failure; SEED_SUFX renamed SEED_SUFFIX; regen guard drops
  the redundant head -1; the ambiguity error no longer advises an
  impossible 'complete id' for duplicate legacy ids.
- commands.cts: refresh the cmdListSeeds comment still describing
  SEED-NNN as the only canonical form.
- changeset: drop the audit claim the spec axis showed to be an
  overstatement (audit's id display is filename-derived, pre-existing).
- remove a stray untracked artifact file swept into the tree.

* test(#4378): correct boundary expectations to the module's real branch behavior

The first matrix run on the boundary tests caught my hand-trace of the
regex branches, not a module defect: the slug regex's alternation
backtracks to the legacy branch whenever the canonical branch cannot
complete (so the slug is the remainder after the legacy numeric
prefix), and the 7-digit case fails the canonical branch at its 7th
digit before the dash. Pin the verified values.

* docs(#4378): backfill changeset PR number

* fix(#4378): audit seed identity uses the canonical grammar

Review of this PR found the audit surface publishing a fused filename
stem (SEED-081-region for SEED-081-region.md) where list-seeds reports
the canonical id -- one id, two answers across surfaces, the same
ambiguity class the issue files. scanSeeds now derives identity through
the SAME deriveSeedIdentity the list-seeds gate uses (frontmatter id,
then filename id-prefix, then stem), and audit-open acknowledge
resolves --seed-id by scanning for the derived identity, falling back
to the literal stem so callers scripted against pre-canonical output
keep working. Roll-in per the fix-inline rule: found during this PR's
review, same seed-identity seam.

RED probe: pre-fix audit published seed_id SEED-081-region-becomes /
slug 081-region-becomes for a legacy seeded file; post-fix SEED-081 /
region-becomes, matching list-seeds.

* test(#4378): probe timeout uses the class norm after windows-lane timeout

The windows conformance shard failed its bounded sh -c probes at the
local 5000ms bound (cold sh.exe spawn under shard load) while the
identical code passed this PR's two earlier windows waves. The probe
now uses PROBE_TIMEOUT_MS from the class-norm module instead of a local
override, per the helpers/timeouts.cjs convention.

---------

Co-authored-by: sim <sim@local>
2026-09-15 11:41:29 -04:00
Tom Boucher
cbbde6786a fix(#4546): deferred UAT follow-ups no longer block completion and promote to the backlog (#4769)
* test(#4546): failing-first tests for deferred uat follow-ups

* chore(#4546): regenerate derived lists for the deferred-promotion suite

The new verify-work-deferred-promotion suite changes the tests/ tree the
macOS conformance-tier classifier tracks and is a novel file under the
verify prefix in the test-file-count ratchet; both derived lists are
regenerated/registered per their own guards' instructions.

* fix(#4546): deferred uat follow-ups no longer block, and get promoted

Two halves of one disconnect (#1921's deferral design vs the completion
predicate):

- uat-predicate: the item parser now captures the block's reason: line
  alongside result:. A skipped item whose reason carries the
  verify-work writer's 'Deferred follow-up:' template is a deliberate
  deferral -- non-blocking, flagged deferred in the report. Quote-
  tolerant (the writer wraps the value) and case-insensitive. A
  reasonless skip, a non-deferral reason, pending/blocked/issue/
  failed/missing all still block, exactly as before.
- verify-work complete_session: when the Deferred Follow-Ups section is
  non-empty, offer to promote the items to a ROADMAP.md 999.x backlog
  entry reusing next.md's prior_phase_completeness entry shape, with a
  --files-scoped commit. Offer, not auto-mutation -- matches the
  workflow's interactive convention and next.md's own prompt style.

* chore(#4546): refresh compact-content benchmark baseline

verify-work.md grew (the #4546 deferred-follow-up promotion offer in
complete_session); the registered split's token counts moved with it.
Baseline recomputed with the script's own --write.

Emitted-Drift-Ack-Growth: verify-work.md — complete_session gained the deferred-follow-up promotion offer (detection, [P]/[K] choice, the next.md-shaped 999.x entry template, and the --files-scoped ROADMAP.md commit); the growth is the new contract text, not duplication

* fix(#4546): gate/audit agreement and review fixes for deferred follow-ups

- src/uat.cts categorizeItem: a skipped item carrying the deferred
  follow-up template reason now categorizes as 'deferred' (the category
  already existed for deferred-items.md entries) instead of being
  misfiled into the blocked families by keyword match -- the gate/audit
  agreement #3078-CR expects, restored in the permissive direction the
  #1921 design intends. Checked BEFORE the keyword families so '...
  on the release build next version' is not build_needed.
- verify-work.md promotion step: numbering scans for the smallest free
  999.n (count races + non-contiguous history), one backlog entry per
  deferred follow-up, ROADMAP.md-absent behavior specified, idea text
  newline-flattened, Deferred at placeholder harmonized with next.md.
- DEFERRED_REASON_RE: trust assumption documented (authoring contract,
  not a security boundary; non-matching spellings block fail-closed).
- tests: the property now drives evaluateUatPassed and derives
  expectations from the input spec (never restates the matcher),
  includes the no-result-line branch, and pins its seed; the parity
  test drops try/finally for the approved pattern, uses createTempDir,
  sites its allow-test-rule marker at the suppression site, and asserts
  the literal [P]/[K] choices.

* fix(#4546): close promotion-test docstring, drop fc replay-path misuse, refresh baseline

The final matrix run caught three defects in my own review-fix commit:
the parity test file's JSDoc was left unterminated (the whole file
parsed as one comment -- zero tests registered, hence the file-level
'test failed' the runner reported); fast-check's replay-path parameter
was misused as a label (invalid path at replay); and the workflow-text
ambiguity fixes re-drifted the compact-content benchmark baseline.

* docs(#4546): add Fixed changeset for deferred follow-up coverage

* docs(#4546): backfill changeset PR number

* fix(#4546): use the pattern seam escapeRegex for shape-marker matching

The hand-rolled metacharacter escape in the shape-marker assertion
tripped local/no-adhoc-regex-escape, whose named remedy this adopts.

---------

Co-authored-by: sim <sim@local>
2026-09-15 07:30:38 -04:00
BeeHiggs
0967358b8b enhance(#3638): render bracket phase IDs on progress, stats, manager and statusline surfaces (epic #612 PR-5) (#4111)
* enhance(#3638): render bracket IDs on display surfaces

Gate progress, stats, manager, and statusline projections on the bracket convention; validate phase_id_convention and single-source the convention card.

Forward note: the uat.cts bracket co-change remains deliberately deferred to its owning slice.

* chore(#3638): point the changeset at PR #4111

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3638): close bracket display review gaps

* docs(#3638): register phase display modules

* chore(#3638): re-trigger CI after macOS shard SIGTERM

`full test (macos-latest, 24, shard 3/3)` failed on 20ce98cd1 in
`tests/lint-compiled-artifact-sync.test.cjs` — the spawned
`scripts/lint-compiled-artifact-sync.cjs` was killed at 60024ms
(`exited null (signal SIGTERM)`, stdout and stderr both empty), 24ms past
the test's own `TSC_COMPILE_TIMEOUT_MS`. That is the failure mode the
constant's comment already documents ("under CI shard load that compile
can exceed the budget, dying to a SIGTERM with empty piped stdout").

No content change; this empty commit exists only to re-run the matrix,
since re-running a job needs write access on the upstream repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 04:47:47 -04:00
Michel Moreira
2f0e99f9e0 fix(#4377): opt in to project-relative includes for local installs (#4425)
* enhance(#4377): opt-in project-relative includes for local installs

A local install wrote the includes that point at GSD's own files as absolute
paths — whatever the installer resolved at install time. For one checkout
that is invisible. Across git worktrees it is not: each worktree gets its own
.claude/ copy, but all of them point back at the checkout that ran the
installer, so a worktree runs its own gsd-tools.cjs while reading workflow
prose from a different checkout. Update that one checkout and every other
worktree is running new instructions against an old engine, with nothing to
stage the update with.

--relative-includes (or GSD_RELATIVE_INCLUDES=1) makes a local install emit
`@.claude/gsd-core/...`. Opt-in, and staying opt-in: absolute works for a
single checkout, which is most people, and flipping the default would change
every existing local install to solve a problem those users do not have.

The prefix is the runtime's own localConfigDir descriptor value, never a
literal — the same value resolveScope joins onto the cwd to produce the
install target, and the same one the rewrite engine already uses for its
./.claude/ -> ./<dir>/ substitutions. Copilot and Antigravity have shipped
this shape for local installs since they were added, with hardcoded .github/
and .agents/; this is that behavior, derived rather than written down.

Six seams compute a path prefix and all six had to be threaded, which is why
the opt-in travels through the environment the way --portable-hooks already
does: one variable they all read cannot fall out of sync the way six
signatures can.

The launcher shim deliberately keeps its ABSOLUTE fallbacks. It probes
gsd-tools through ${CLAUDE_CONFIG_DIR:-$HOME/.claude} and one such default
per runtime; those are shell word expansions, not includes, and a relative
value there resolves against the shell's cwd rather than the project.
Trading an include that points at the wrong checkout for a path that points
at nothing is not a fix. All three rewrite paths mask ${VAR:-default} spans
before substituting and restore them after, and the mask only runs when the
prefix is relative, so an absolute install is byte-for-byte unchanged.

Every unexpressible case falls back to absolute: no opt-in, a global install,
a missing dir name, the configHome.kind === 'none' sentinel, an absolute
descriptor value, or one climbing out of the project with '..'.

* chore(#4377): add changeset for project-relative local includes

* fix(#4377): compare against POSIX-normalized roots in the install e2e arms

The emitted prefix is POSIX-normalized by design — it is substituted into
markdown @-references, which use forward slashes universally, so a backslash
would leak into shipped content (#1615). The e2e arms compared against the
raw temp root, which on Windows is `D:\a\...` and appears in no emitted file.

That reddened the control arm on the windows shard, and it was worse than a
red: the negative arm ("nothing references the checkout") was passing
VACUOUSLY there, because a string that cannot occur is trivially absent. Both
now go through the same normalization, so the Windows lane asserts what the
Linux lane does.

* fix(#4377): tolerate a resolved temp root, and make the e2e diff self-diagnosing

Two changes, one confirmed and one to stop guessing.

Confirmed: the emitted content carries the RESOLVED root, not the spelling
mkdtemp handed back. Reproduced on Linux with a symlinked install root —
236 emitted files carry the realpath, zero carry the link path. macOS has
this structurally, since /var is a symlink to /private/var. Comparisons now
go through both spellings, or the negative arms pass vacuously: "nothing
references the checkout" is trivially true when the string being searched
for cannot occur.

Not confirmed: the macOS shard reported ~every workflow file differing in
the "differ ONLY" arm while the five arms around it passed, and the
assertion printed a list of filenames — which says a difference exists
somewhere across 236 files and leaves the reader to guess which bytes. I
cannot reproduce that platform locally, and guessing turns one CI round-trip
into four. The assertion now reports the first divergence as text: the file,
the byte offset, and a bounded window of both sides.

* fix(#4377): strip the longest root spelling first in the install e2e diff

The macOS failure was my test corrupting its own comparison, not a product
defect. /var/folders/…/X is a SUBSTRING of /private/var/folders/…/X, so
stripping the unresolved spelling first matched inside the resolved one and
left the /private prefix glued to what followed:

  @/private/var/…/X/.claude/gsd-core/…  ->  @/private.claude/gsd-core/…

a string present in neither install, which is why all 236 files "differed".
Sorting the spellings longest-first consumes the whole occurrence, and the
short form then has nothing left to match. Proven in isolation on the exact
macOS shapes: short-first yields @/private.claude/…, longest-first yields
@.claude/….

The self-diagnosing assertion added in the previous commit is what found
this — it named the file, the byte offset, and printed both sides, so the
corrupted string was visible rather than inferred from a list of 236
filenames. Keeping it.

* fix(#4377): address review findings

* test(#4377): scan nested shell defaults without regex backtracking

* fix(#4377): close relative include review gaps

* fix(#4377): preserve root-target runtime includes

* fix(#4377): guard project-root relative includes

* test(#4377): normalize Cline fallback roots

* fix(#4377): persist relative include style

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 03:40:04 -04:00
Tom Boucher
7efd9032ee fix(#4544): cover hooks/, scripts/, CHANGELOG and the manifest in codex rollback (#4760)
* test(#4544): failing-first rollback-coverage tests for codex install

* test(#4544): scope the rollback suite to its own requires

The appended describe used bare describe/test/os/cleanup and the folded
block's runCodexInstall — none visible at file scope, so the whole test
file failed to load. Wrap it in its own block with local requires and a
local harness copy, matching the folded-block idiom.

* test(#4544): import beforeEach/afterTest hooks into the suite scope

* fix(#4544): restore manifest-tracked files and hooks/ on codex rollback

restoreCodexSnapshot and _codexPreConfigRollback knew only the five
#3245 targets (config.toml, hooks.json, skills/gsd-*, agents/gsd-*,
VERSION). A Codex install also writes hooks/, gsd-core/CHANGELOG.md,
gsd-core/.gsd-runtime, scripts/changeset|lib + the standalone scripts,
and rewrites gsd-file-manifest.json -- none were captured, so any
rollback left the new payload in place for all of them.

The pre-install capture now records every path the PRIOR install's own
gsd-file-manifest.json lists (bytes when present, absence-marker when
not, so a path deleted between installs is re-deleted rather than
resurrected), the manifest file itself, and the whole hooks/ tree --
wholesale, because the Codex manifest deliberately omits hooks/ and
hooks/ is shared space, so restore returns user files that predated the
install and drops everything the failed install staged. Both rollback
paths share one restore closure; malformed or missing prior state
degrades to today's behavior.

Known residual, documented: files the FAILED install adds under
manifest-tracked dirs survive a rollback that fires before the new
manifest is written (they are named by no prior state). The five
original targets and hooks/ have no residual.

* test(#4544): align malformed-manifest fixtures with pre-install-state semantics

Row 7 seeded VERSION and then asserted its absence -- but a seeded
VERSION is pre-install state the fix must restore, not remove. Row 8
asserted a pre-existing array-shaped manifest must not survive, when
restoring those exact bytes IS the contract. Both were fixture bugs;
the probe-verified installer behavior was correct.

* docs(#4544): add Fixed changeset for codex rollback coverage

* fix(#4544): harden snapshot per adversarial review — minimal mode, symlinks, clean installs

Review (three independent passes) found five defects and one coverage
gap in the first cut; all fixed:

- BLOCKER: the capture gate is off in minimal mode but the restore call
  was not, so a minimal-mode rollback wholesale-deleted the user's
  entire hooks/ directory (empirically confirmed by the reviewer). The
  restore now consults a captured flag: no snapshot means do nothing.
- MAJOR: the hooks/ walk followed file symlinks — a repo-shipped
  .codex/hooks symlink to a FIFO would hang the installer, to
  /dev/zero exhaust memory, or to private data copy that data into the
  snapshot. The walk lstats every entry and captures only true regular
  files; anything else marks the capture incomplete.
- Incomplete captures now downgrade the restore to per-file: put back
  what was captured, remove only the names GSD itself stages (the
  hoisted CODEX_HOOKS_TO_COPY set + CommonJS marker), never wholesale-
  delete a tree the snapshot did not fully see. GSD-owned removal runs
  before the restore so a name in both sets keeps its pre-install
  bytes.
- A pre-existing hooks FILE (not directory) is left alone instead of
  deleted.
- readInstallManifest now rejects a manifest whose files field is a
  JSON array (typeof [] === 'object'), which previously produced
  numeric-key paths.
- Clean FIRST installs: with no prior manifest nothing recorded the
  payload, so a failed clean install rolled back to a half-written
  tree. Capture now enumerates the same source directories the
  installer copies plus the two standalone files (CHANGELOG.md from
  the repo root, generated .gsd-runtime) and records absence — a
  failed clean install now rolls back to actually nothing.

Tests: minimal-mode preservation regression, symlink never-followed
regression, helpers.cjs temp dirs, the injected-failure message is
asserted, residue assertions made unconditional.

* test(#4544): pin symlink-downgrade semantics the final run exposed

The symlink itself marks the capture incomplete, so the restore takes
the per-file downgrade — which preserves uncaptured pre-install state
(the link) rather than wholesale-dropping it. The probe run verified
exactly this; the assertion guessed the wholesale branch. Pin the
verified behavior: referent untouched, link preserved and resolving,
no leak.

* docs(#4544): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-15 02:37:18 -04:00
0xdhx
aeac47b95a fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD (#4472)
* fix(#4465): bound /gsd:undo commit selection to the phase directory and HEAD

`--phase` documented a primary path reading `.planning/.phase-manifest.json`,
but nothing in the repository writes that file, so the documented fallback was
the only reachable path:

    git log --oneline --no-merges --all | grep -E "\(0*${TARGET_PHASE}(-[0-9]+)?\):" | head -50

That selection has no milestone bound and no reachability bound, and it feeds
`git revert --no-commit`. On a project that reuses a phase number it stages
deletion of a previous milestone's files, under a confirmation gate that
displays only `{hash} — {message}` — the one field that carries no milestone
discriminator.

Port the #3995 anchor already live in code-review.md: resolve the phase's own
directory via `find-phase` (which resolves through planningDir, so it is
workstream-correct), take PHASE_START as the first commit adding anything under
it, and select over `PHASE_START^..HEAD`. Drop `--all`. Fail closed when no
anchor resolves rather than widening to a repository-wide search, and report
truncation instead of silently capping at 50.

Also resolve the dead manifest read rather than leaving documented-but-
unreachable behaviour, and hoist dependency_check onto the workstream-resolved
planning root — it read a hardcoded `.planning/ROADMAP.md` and
`.planning/phases/`, which are the wrong tree under an active workstream.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* fix(#4465): close three defects the pre-create adversarial review found

Root-commit off-by-one: `${PHASE_START}..HEAD` EXCLUDES PHASE_START, so when the
phase's first commit is the repository root the selection silently dropped it
and the undo refused legitimate work. The root branch now selects over `HEAD`.

Empty-selection exit status: `grep` exits 1 on no match, and the removed
`| head -50` had been masking that rc. Both selection pipelines now end in
`|| true` so an empty selection reaches the workflow's own Empty check instead
of aborting the block.

Truncation stop was documented for MODE=phase only; MODE=plan could still cap
silently. Both modes now carry it.

Also documents two residuals the review surfaced rather than leaving them
implicit: a revision range is ancestry and not chronology, so a pre-phase side
branch merged in after PHASE_START stays selectable; and the `--diff-filter=A`
anchor does not follow renames, so an archived phase directory under-selects
(reverts too little or refuses, never too much).

Tests: 12 assertions, negative-controlled. A-H and K-L are RED against the
pre-fix workflow; J is RED against the intermediate revision that carried the
off-by-one; I is green in both directions by design.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* chore(#4465): backfill the changeset fragment's pr field

The fragment could not carry `pr:` before the PR existed; both
`scripts/changeset/lint.cjs` and `scripts/lint-docs-required.cjs` require it
and reported `missing_pr` until now. Both are green with it filled in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* chore(#4465): acknowledge the deliberate growth of undo.md

The emitted-attribution gate flags undo.md growing 11881 -> 17435 bytes. The
growth is the fix: a one-line `git log --all | grep` selection is replaced by
an anchored, HEAD-bounded selection for BOTH modes, each with its own
fail-closed branch and truncation stop, plus three residuals documented next to
the code they qualify rather than left implicit. A workflow document is the
executable contract, so the residuals belong in it.

Emitted-Drift-Ack-Growth: undo.md — replaces a one-line unbounded commit-subject grep with an anchored HEAD-bounded selection in both --phase and --plan, each with a fail-closed branch and a truncation stop, plus three residuals documented in-workflow (#4465)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B58zMdMYc3n16mfdqMHMqv

* test(#4465): execute undo.md's selection fences against a git fixture

The shipped regression test matched substrings in undo.md's bash fences. A
substring cannot tell a live invocation from a dead one, and this is the
boundary/reachability defect class this repo's own conventions want driven
with limit-1/limit/limit+1 execution proof. So the fences the runtime runs —
sliced out of undo.md by content anchor, never by position — are now replayed
with `bash -c` inside createTempGitProject fixtures against the real
`gsd-tools.cjs`, the same fence-execution shape #2308 and #2352 use.

Ten executed cases: the fixture's own negative control (the retired `--all`
grep over-selects the archived milestone and a dead branch), `--phase` and
`--plan` selecting only the current milestone's HEAD-reachable instance, the
single-milestone selection unchanged, limit-1 (a matching pre-phase commit
excluded, PHASE_START itself included), the root-commit branch selecting over
`HEAD`, fail-closed on an absent phase in both --phase and --plan (empty PHASE_DIR and UNDO_RANGE), the
active workstream's phase directory winning over the root's, and
dependency_check's `planning inspect --pick generated_from.planning_root --raw`
resolving the workstream root, the project root, and the `.planning` fallback.

Skipped on win32 with the #2352 precedent's reason: the fences are POSIX bash
driven through `bash -c`; the shape half still runs there.

Negative-controlled against the pre-fix undo.md (upstream/next): 11 of 12
shape assertions red, and the executed block fails at fence extraction.

Shape test L now pins the >50 refusal message in both modes, not only its
heading — the cross-AI round review removed the paragraphs under intact
headings and L stayed green; it now fails on that mutation.

* fix(#4465): refuse an archived-milestone PHASE_DIR in both undo modes

`find-phase` searches the live `phases/` directory first, then every
`milestones/v<X.Y>-phases/` directory in ascending version order, and its
ambiguity check is scoped to one directory: the `matches.length > 1` test sits
inside `cmdFindPhase`'s per-`searchDir` loop (`src/phase.cts`). A phase number
that is not live therefore resolves silently to the OLDEST archived milestone
carrying one, with no warning.

Anchoring there is wrong in both directions at once. The oldest commit adding
that path is the archival move, so the phase's real work predates the window
and falls outside it, while the window runs forward from that archival through
every later milestone -- where the subject grep matches THEIR same-numbered
phase. Driven on a two-archived-milestone fixture, `--phase 03` selected v2.0's
`feat(03-01): add search index` and `docs(03-01): v2.0 phase plan` and excluded
v1.0's own `feat(03-01): implement auth endpoint`. On `git revert --no-commit`
that is the cross-milestone contamination this PR exists to close, recurring.

Both modes now blank PHASE_DIR on an archived resolution and refuse with their
own message, so the existing fail-closed rule stays load-bearing even if the
refusal's prose is not honored.

The "archived or renamed phase directory under-selects" residual claimed this
failure could revert "too little or refuse, never too much". That was wrong in
the direction that matters; it is rewritten to cover renames only, and the
archival case is recorded as refused rather than disclosed.

Tests: a two-archived-milestone fixture, a negative control asserting the
unguarded window selects v2.0's two commits and none of v1.0's, refusals in
both --phase and --plan, and an inertness check on a live resolution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): close the three Minor review items on undo.md and its test

Purpose line: it still advertised rolling back "using the phase manifest",
the exact `.planning/.phase-manifest.json` mechanism this PR removes. Test B
already reads the whole file, so scope is not why it missed this -- B greps the
hyphenated `phase-manifest` token, the filename, while the purpose line named
the same dead mechanism in prose. Test N pins the <purpose> block itself, which
is spelling-independent; widening B's pattern to /manifest/i instead would fire
on any future sentence that merely mentions one.

Merge-commit anchors: `git log --diff-filter=A -- "${PHASE_DIR}"` does not walk
merge diffs by default. A directory added on a side branch is still found -- the
side-branch commit that added it is in history -- so the uncovered case is
narrower than "introduced via a merge": it is a directory first appearing in the
merge RESOLUTION. The information is not absent from history, only unrequested:
`git log -m` prints that add once per parent. Recorded as a residual rather than
fixed -- the failure is a refusal, and the evil-merge fixture costs more than a
safe-direction branch is worth.

allow-test-rule marker: suppression is site-scoped, and CONTRIBUTING.md pins the
window at MAX_MARKER_LOOKAHEAD_LINES = 8 with only blanks and comments between.
The file-header marker sat 43 lines above the first `readFileSync` with requires
and a function definition in between, so it was inert for both read sites. Moved
to each site. The markers are belt-and-braces today, and NOT because the rule
ignores `RegExp.test` -- it handles `regex.test(tracked)` explicitly
(no-source-grep.cjs:239, :597-605). Neither read is tracked at all:
`looksLikeSourcePath` (:378-390) admits only .cjs/.cts/.js/.mjs/.mts/.ts and
UNDO_PATH is a .md, and the second site's reader is `readFileNormalized`, which
the rule does not recognise as `readFileSync`. They become load-bearing if
either scope widens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4465): record the archived-milestone refusal in the changeset

The fragment described the PHASE_START bound and the fail-closed rule but not
the archived-resolution refusal added this round, which is a user-visible
behaviour change: `--phase N` on a number that is no longer live now stops
rather than anchoring on an archived directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): disclose the same-number-and-slug residual in undo.md

The round's own cross-model body audit found this stated in the PR description
and nowhere in the deployed workflow -- which is the half that survives merge,
and the half this PR's whole argument says residuals belong in.

The anchor is the CURRENT path and `--diff-filter=A` does not follow renames, so
a later milestone that re-creates the same literal directory (`03-auth` again,
not merely phase `03` again) makes the oldest add at that path the previous
occupant's. The archived-milestone refusal added this round structurally cannot
reach it: `find-phase` returns the LIVE directory, so nothing is under
`milestones/` to refuse.

Driven -- v1.0 and v2.0 both using `.planning/phases/03-auth`, v1.0 archived in
between: PHASE_DIR resolves live, the guard correctly does not fire, the anchor
is `docs(03-01): v1 plan`, and the selection returns all four v1+v2 phase-03
commits. `code-review.md` carries the same residual on the same anchor, where it
is read-only; on `git revert --no-commit` it is not, so the note points at
`/gsd:undo --last N`.

Documented rather than fixed: following renames across a re-created path needs a
phase identity that a directory name does not carry, which is the same wall
residual 1 hits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4465): refuse a LIVE phase directory whose name is also archived

Self-found this round, from the cross-model audit of the previous commit: that
commit disclosed same-number-and-slug reuse as a residual and asserted closing it
needed a phase identity a directory name does not carry. The audit refuted the
second half by producing a fix, and it is cheap and safe-direction, so the case
is now refused rather than documented.

`--diff-filter=A` does not follow renames, so a later milestone that re-creates
the same literal directory (`03-auth` again, not merely phase `03` again) anchors
on the EARLIER occupant's add commit and the window opens there. The archived
refusal added earlier in this round structurally cannot reach it: `find-phase`
returns the LIVE directory, so nothing is under `milestones/` to refuse. Driven
before the guard -- v1.0 and v2.0 both at `.planning/phases/03-auth`, v1.0
archived in between: PHASE_DIR resolves live, the archive guard correctly stays
inert, the anchor is `docs(03-01): v1 plan`, and all four v1+v2 phase-03 commits
are selected. That is a previous milestone's work staged for `git revert
--no-commit`.

The guard needs no identity reconstruction: the same basename present under an
archived `milestones/v*-phases/` means the path has been used before, so the
anchor is untrustworthy and both modes refuse with their own message. It fails
toward refusing a legitimate undo of the reusing milestone, which `--last N`
covers; the alternative is reverting the earlier one's commits.

Tests: a reused-slug fixture, a negative control asserting the unguarded window
reaches back into v1.0 (all four commits), refusals in both modes, and an
inertness check on a distinct slug. All three new checks red against the pre-fix
fence.

Also narrows the changeset, which still claimed selection "no longer" reaches a
previous milestone -- true of the archived route, not of this one until now --
and corrects the concurrent-workstreams residual, which claimed the window
removed the previous-milestone class "entirely" while this case remained open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4465): harden the reused-path refusal against its own false positives

The cross-model audit of the previous commit refuted four of its claims. All
four were right; this closes them.

A REAL BUG: the refusal message interpolated `${PHASE_DIR}` after the fence had
already blanked it, so it would have rendered "Phase 03 resolves to , but that
directory name is also archived at ...". The live path is now preserved in
`PHASE_DIR_LIVE` before blanking, and a test pins that it survives.

THREE FALSE-REFUSAL ROUTES, each of which could block a legitimate undo:

- `[ -e ]` accepted a regular FILE where an archived phase directory would sit.
  The evidence the refusal claims is "an earlier milestone used this path", and
  only a directory is that. Now `[ -d ]`.
- The glob `v*-phases` accepted milestone directory names `cmdFindPhase` itself
  rejects -- its filter is /^v\d+.*-phases$/, so `vnondigit-phases` is not a
  milestone it would ever resolve. Refusing over one is refusing on evidence the
  producer discards. Now `v[0-9]*-phases`, in the collision check and in the
  archived-resolution guard alike.
- `${PHASE_DIR%/phases/*}` silently left the path unchanged when it carried no
  `/phases/` segment, so the scan ran against the wrong root and read as clean.
  The strip must now have fired. Outside today's producer contract either way --
  cmdFindPhase's live output always carries `/phases/` -- so this hardens a claim
  rather than fixing a reachable defect, and is stated as such.

The commit message and the changeset both asserted the guard proves the name is
"archived"; before this commit it proved only that some entry existed. Both are
narrowed to what the fence now actually establishes, and the changeset headline
no longer claims more than the anchor plus the two refusals deliver -- residual 2
(a range is ancestry, not chronology) is untouched by either.

Tests: a file-not-directory twin, a malformed `vnondigit-phases` twin, and the
preserved live path. All three red against the pre-hardening fence, along with
shape test M.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* docs(#4465): disclose what the reused-path scan does not survive

A fifth audit pass refuted three claims in the previous commit. Two are wording,
one is a real fail-open; none is fixed by more shell, so all three are stated.

The wording: `v[0-9]*-phases` was described as MIRRORING cmdFindPhase's
/^v\d+.*-phases$/. It is not equivalent -- the glob's `*` matches a newline where
the regex's `.` does not, so a directory named `v6<newline>-phases` is accepted
here and rejected there. It tracks the filter closely enough to reject the
malformed siblings that motivated it; it does not mirror it, and the comment no
longer says so. The changeset likewise said the twin is "an archived phase
directory", where the check establishes only that a directory of that name exists
under an archived milestone -- narrowed to that.

The fail-open: this collision check is the only fence in undo.md that relies on
pathname expansion (every other one uses `case`, which `set -f` does not affect).
Under a runtime with globbing disabled the scan is skipped silently and a genuine
collision passes; under `shopt -s failglob` a NON-match aborts the fence. Both sit
outside the shell state this workflow assumes throughout, and defending only this
fence while the rest of the file assumes defaults would be inconsistent -- so it
is a documented residual rather than a hardened one.

Also disclosed: the check is conservative at two edges. `[ -d ]` follows symlinks,
and an empty directory of the right name counts, so either can refuse an undo a
stricter ownership test would allow. That direction is the intended one -- refusing
too often costs a `--last N`, refusing too rarely reverts another milestone's work.

No test is added. Asserting behaviour under `set -f` would pin a shell state the
workflow does not otherwise support, and the two conservative edges are the
documented intent rather than defects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4465): register the new regression test in the conformance-tier list

Round 3's only blocker. The platform-conformance-tier classifier, its generated
registry and the test that asserts the registry is fresh all landed together on
`next` in `bcd99696d` (#4591 / #4598) on 2026-09-10 08:36 -0400 -- a day after
this branch's last push, so the gate did not exist when the branch was last
green. This PR adds a unit-suite test file, the classifier's content scan selects
it, and a list regenerated on `next` without it is therefore stale the moment the
two meet.

(The review cites `0b928fe28c` for that landing. That commit is #4592, four hours
later the same day, and `git diff-tree` shows it touched only
`scripts/ci-test-scope.cjs`, `scripts/gen-platform-conformance-tier.cjs` and
those two files' tests -- not the generated registry at all. The date and the
diagnosis hold either way.)

Regenerated with `node scripts/gen-platform-conformance-tier.cjs --write` on the
rebased tree: one line, 267 -> 268 entries. The `--target macos` list is
unaffected -- it writes a different file, `macos-conformance-tier.generated.cjs`,
and its classifier does not select this test (198, already matching) -- so
`lint:generated-sync`'s second conformance-tier link needed nothing.

**Two different baselines, stated so the counts are not read as one.** CI's
failure on the prior head reads `546 !== 547`, and this commit's diff reads
267 -> 268. Both are correct and they are not the same tree: at the merge commit
`54b0b197f` the committed list held 546 entries, so the classifier wanted 547.
`4d65c248e` (#4641, "narrow the tier to 28.5%") then landed on `next` at
2026-09-11 17:00 -0400 -- after that CI run started at 13:26Z -- and collapsed
the list to 266, with `a2331c01f` taking it to 267. Hence 268 here. The two
547-entry lists are not the same file set: upstream's includes
`tests/execute-phase-decimal-arithmetic.test.cjs`.

Order matters and is worth stating: the registry is derived from the `tests/`
tree, and the base range removed 282 entries from it and added 3 (`git diff
--numstat` reports `3 282`; the familiar 279 is the net shrinkage, not the count
of entries changed). Regenerating before the rebase would have produced a
547-entry list against a base carrying 267, and a three-way `git merge-file`
control over that pair does conflict. Rebase first, regenerate second.

This clears both reds on the prior head, not one. `lint-tests` is the one the
review named; `test (ubuntu-latest, 24, shard 3/3)` is the same staleness seen
through `tests/platform-conformance-tier.test.cjs:278`, which asserts the
committed list is fresh. It was the only failure in 1267 tests on that shard, and
it completed at 13:38:33Z -- five minutes after the review was submitted at
13:33:54Z, which is why the review recorded the ubuntu matrix as green.

Control: removing the added line reds `real tests/ tree classification matches
the committed list` (1 of 55, `267 !== 268`); restoring it gives 55/55.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R4EJcuDjaaF1QeeT8AxYbj

* test(#4465): pin what find-phase returns for the workstream archive layout

Review round 4 found the archived-phase handling blind to the second archive
layout, milestones/ws-<name>-<date>/phases/, that `workstream complete` writes.
The suite had no fixture for it and its comments described one archive reader
where there are two.

- Names both readers in the resolver-contract comment: find-phase routes to
  cmdFindPhase, which admits only /^v\d+.*-phases$/ under milestones/, while
  the phase locator (listArchiveVersionDirs) also enumerates ws-*.
- Adds a ws-* fixture built the way `workstream complete` builds it (the whole
  workstream dir moved into milestones/ws-feat-<date>/).
- Negative control: a re-created workstream with the same phase slug, run
  without the collision guard, selects the archived generation's commits.
- Pins the actual find-phase contract: a phase living only in a ws-* archive
  resolves to nothing, and nothing is selected.
- Pins that an ordinary deleted plan file inside a live phase is not treated
  as a previous occupant.

runFences gains an explicit env override so a test can activate a workstream
after the leak scrub. Every test here passes on the pre-fix workflow; the
assertions that need the fix land with it in the next commit.

* fix(#4465): refuse a re-created phase path by its history, not an archive-layout glob

Review round 4: the archived-phase handling encoded the archive layout as a
literal `v[0-9]*-phases` glob, while the phase locator owns two layouts. The
second, milestones/ws-<name>-<date>/phases/, is what `workstream complete`
writes.

Driven on the real fences: a workstream `feat` completed into ws-feat-<date>/
and then re-created with the same 03-auth slug resolves to the live path,
the collision glob finds no v*-phases twin, the anchor opens on the archived
generation's first commit, and --phase 03 selects that generation's commits.

The collision check now asks git whether this exact path went EMPTY somewhere
in HEAD's history and came back: `git log -m --no-renames --diff-filter=D`
over PHASE_DIR, then `git ls-tree -d` at each deleting commit to tell a
vacated directory from an ordinary deleted plan file. That answers for every
way a path can be vacated -- the flat archive, the ws-* archive, and a phase
removed and re-added under the same slug, which no layout glob could see --
without a second reader of the layout to keep in sync. The refusal message
now names the commit that vacated the path.

The archived-resolution refusal names both layouts. find-phase (cmdFindPhase)
searches only the flat archives today, so a ws-*-only phase resolves to
nothing and fails closed on the not-found rule; the ws-* arm keeps the refusal
correct if find-phase is ever taught the locator's second layout, and a
stubbed-resolver test pins it for both modes.

The retired glob's shell-state residual (pathname expansion under set -f /
failglob) is gone with it. Its replacement residual is documented in-workflow:
the history check misses a single commit that both moves the directory away
and re-creates it, and over-refuses when a side branch emptied it and the
merge kept it.

* fix(#4465): run undo's path-scoped git calls from the project root find-phase answers against

Found by this round's pre-push review. find-phase prints PHASE_DIR relative
to the PROJECT ROOT -- gsd-tools resolves the root before it dispatches --
while the workflow's shell stays wherever the user invoked /gsd:undo. A git
pathspec is read relative to git's cwd, so from a subdirectory every
PHASE_DIR-scoped call looked in the wrong place:

- the anchor (`git log --diff-filter=A`) came back empty, so a legitimate
  --phase or --plan undo was refused. Fail-closed, but a refusal of valid
  work, and it dates from round 1;
- the collision check's `git log --diff-filter=D` came back empty, so a
  re-created path was never recognised. The anchor failing too is the only
  reason this did not over-select.

Both modes now take PROJECT_ROOT from `planning inspect --pick
generated_from.cwd` -- the directory find-phase itself resolved against --
and run the three path-scoped calls as `git -C "${PROJECT_ROOT:-.}"`. The
project root, not `git rev-parse --show-toplevel`: a project need not sit at
the top of its repository, and a control test pins that choice. Selection,
revert and rev-parse calls are SHA-only and unchanged.

Tests: every behavioural test ran from the fixture root, where the two
readings coincide. New ones run the fences from sub/dir -- normal selection
in both modes, refusal of a re-created path for both archive layouts in both
modes, and a dropped plan file NOT refused, which is how a mis-rooted
ls-tree would fail (it lists nothing, and nothing reads as "vacated"). Shape
test O pins every PHASE_DIR-scoped git call to the root. The two tests
renamed in the previous commit are relabelled as false-positive controls:
they pass with the collision check deleted, and say so.

Reversion controls, all driven: the pre-change undo.md fails 7 tests; one
mode's ls-tree mis-rooted fails 3 in either mode; PROJECT_ROOT from
--show-toplevel fails 2.

* fix(#4465): refuse a phase planned in another repository, and never widen a foreign anchor

Found by this round's second pre-push review, against the previous commit.
Running the anchor from the project root is right while the project root
and the caller share a repository. In a `sub_repos` project they do not:
.planning/ lives in a parent repository and the code in child ones, and
from a child findProjectRoot returns the parent. The anchor then came from
the PARENT's history -- a commit the child does not hold. `git rev-parse
"${PHASE_START}^"` failed in the child, the root-commit arm read that as
"no parent", set UNDO_RANGE=HEAD, and selection ran over the child's whole
history. Driven: both child commits selected, including one older than the
phase. Before the previous commit that case resolved no anchor and failed
closed; the previous commit made it destructive.

Two layers, both modes:

- a same-repository refusal in the guard fence: the caller's git directory
  and PROJECT_ROOT's, each physically resolved, must be the same one
  (per-worktree, so a linked worktree compares correctly). A different one
  sets PHASE_DIR_FOREIGN, blanks PHASE_DIR, and the workflow stops with a
  --last N pointer;
- in the anchor: a PHASE_START that is not a commit this repository holds is
  blanked before the root-commit arm. That ambiguity -- "no parent" read as
  "root commit" when it can also mean "not a commit here" -- dates from
  round 1; the previous commit made it reachable.

Shape test O now also pins that PROJECT_ROOT is assigned only from planning
inspect: a rooted call is only as good as the root it is handed, and a later
reassignment passed every spelling check (review, claim 8). Shape test P
pins both layers in both modes.

Correction to the previous commit's message: it said an empty anchor was
the only reason a missed collision could not over-select before that
commit. The review drove a counterexample -- a tracked `.planning` path
coinciding with the caller's subdirectory resolves a wrong, non-empty
anchor -- so that sentence overstated it.

Reversion controls, driven: the previous commit's undo.md fails 3 (P, the
sub_repos refusal, defense in depth); the refusal made inert fails 2 (P and
the refusal; the anchor check alone still keeps the range empty); the
anchor check made inert fails 2 (P and defense in depth; the refusal alone
still refuses). A verbatim negative control shows the pre-hardening
root-commit arm widening the parent's anchor to all of HEAD.

* fix(#4465): compare the repository, not the worktree, and anchor only inside HEAD's history

Found by this round's third pre-push review, against the previous commit.
Its repository gate compared per-worktree git directories. gsd-tools maps
a linked worktree with no .planning/ of its own to the MAIN worktree
(resolveMainWorktreeCwd), so from such a worktree PROJECT_ROOT is the main
checkout: one repository, one object database, two per-worktree git dirs
-- and the gate refused a legitimate undo. It now compares the COMMON git
directory, physically resolved, which is the repository's identity; a
sub_repos child is still a different one and is still refused.

Driving that case showed the anchor check was too weak for it. The anchor
comes from the main worktree's branch history, and nothing guaranteed the
linked HEAD contains it: a linked branch that left main before the phase
started would get a window bounded by a commit outside its own history.
The check is now "PHASE_START is an ancestor of HEAD" (`git merge-base
--is-ancestor`), which subsumes the previous "is a commit here" test -- a
missing object is not an ancestor either -- and changes nothing in the
ordinary case, where the anchor is read from HEAD's own log.

Tests: a linked-worktree fixture (the linked checkout has no .planning/,
which is what makes gsd-tools map it to main; the test asserts the mapping
happened) is allowed in both modes and selects the phase; the same shape
branched before the phase resolves no range. Shape test P pins the
common-dir comparison, forbids --absolute-git-dir, and pins the ancestry
check.

Reversion controls, driven: the previous commit's undo.md fails 3 (P and
both linked-worktree tests); per-worktree git dirs restored fails the same
3; the ancestry check replaced by the previous cat-file test fails 2 (P
and the branched-before test, which then selects a linked commit whose
history never held the phase); no anchor check at all fails 3 (P, defense
in depth, the branched-before test).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 01:41:35 -04:00
Michel Moreira
26e650f312 fix(#4462): normalize whitespace workstream environment values (#4509)
* test(#4462): expose whitespace workstream scope

* fix(#4462): normalize the workstream environment scope

* docs(#4462): add changeset for #4509

* fix(#4462): normalize sibling planning scope readers

* test(#4462): cover workstream normalization properties

* fix(#4462): share normalized workstream resolution

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 00:40:15 -04:00
Michel Moreira
e5c4941f2a fix(#4485): isolate hook E2E tests from ambient capabilities (#4529)
* test(#4485): isolate hook E2E capability homes

* test(#4485): isolate the remaining plan hook E2E

* test(#4485): make ambient-home assertions real

* test(#4485): remove stale helper imports

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-14 23:44:18 -04:00
Tom Boucher
e2bfc06558 fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code

AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase
(#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until
this lands.

THE DEFECT, MEASURED

Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking
value, indistinguishable from the same call with a canonical id. Measured on
5d4c98cde7 by executing the built modules:

  getRuntimeLabel('gemini')              -> 'Claude Code'
  getProjectInstructionFile('gemini')    -> 'AGENTS.md'
  getGlobalConfigHomeFragment('gemini')  -> "'.claude'"
  getGlobalConfigDir('gemini')           -> ~/.claude   (byte-identical to 'claude')
  getDirName('gemini')                   -> '.claude'

So asking for a runtime Google sunset on 2026-06-18 wrote into Claude Code's
global config home and labelled the install "Claude Code". Nothing errored and
nothing warned.

AC#1 names four accessors. getDirName is the fifth, found by a reviewer: same
module, same silent-wrong-answer class, and it feeds capability-state's
runtimeConfigDir. Fixing only the four the criterion happened to list would
have left the defect reachable, so it is guarded too.

WHY THE CHECK CANNOT LIVE IN CANONICALIZATION

canonicalizeRuntimeName returns null for 'gemini', 'gemini-cli', 'Gemini',
'GEMINI' AND for ''. After canonicalization a retired id, an unknown id and an
empty string are the same value, so anything keyed off the canonical form
cannot tell them apart — it would have to treat all three alike, which is the
behaviour being fixed. The check therefore runs on the RAW input.

WHAT THIS DELIBERATELY DOES NOT DO

The criterion reads "reject a non-canonical runtime id". Taken literally that
overturns three recorded decisions, so the narrower reading was put to the
maintainer as a blocking question and this implements the answer: RETIRED ids
throw, unknown and future ids keep falling back.

Preserved:

  - The #1529 contract, written into getProjectInstructionFile's own docblock
    as a mapping table ending "unknown / future runtimes -> AGENTS.md (safe
    cross-agent default)". That default exists so a runtime GSD has never heard
    of still gets a working instruction file.
  - ADR-1239 Phase B / #1679, which preserved GLOBAL_CONFIG_HOME_FRAGMENTS
    BYTE-FOR-BYTE when it collapsed a 14-branch chain, with golden install
    parity asserting generated hook output is unchanged across every runtime.
  - The explicit `if (!runtime) return <default>` branch. Empty string is a
    supported input, not a non-canonical id.

The distinction the code encodes: ABSENCE OF KNOWLEDGE IS NOT THE SAME AS
RECORDED RETIREMENT. Unknown means "no information, degrade safely". Retired
means "we know it is gone and we know what replaced it" — and silently
substituting a different product for it is the defect.

ONE INACCURACY IN THE CRITERION, RECORDED RATHER THAN REPEATED

AC#1 says the accessors return "a Claude Code value". True for getRuntimeLabel,
getGlobalConfigHomeFragment, getGlobalConfigDir and getDirName — but
getProjectInstructionFile returns 'AGENTS.md', which is not a Claude value at
all. The defect it points at is real for all of them, so the fix covers all of
them, but the wording is wrong for one.

MATCHING

RETIRED_RUNTIME_DETAILS is a Map keyed by canonical retired id, and
RETIRED_RUNTIME_SPELLINGS maps every spelling to that id. Both are Maps, not
object literals: a literal indexed by a computed key resolves INHERITED
properties, so '__proto__' and 'constructor' were truthy and threw with every
field `undefined`, while isRetiredRuntimeId — which already went through a Set
— correctly answered false for the same input. Two guards disagreeing about one
id is worse than either answer. A Map has no prototype keys, so that hazard is
structural rather than patched. The predicate and the assertion now share one
normaliser and one table and cannot diverge.

Candidates are normalised NFKC + lowercase + strip non-alphanumerics. Folding
the separators makes 'gemini-cli', 'gemini_cli', 'gemini.cli' and 'geminicli'
one key instead of four near-misses found one at a time, and NFKC folds the
full-width 'gemini' a CJK keyboard produces. It stays MEMBERSHIP matching,
never prefix or substring: 'gemini-2.5-pro' folds to 'gemini25pro' and
'gemini-3.1-pro-preview' to 'gemini31propreview', neither a member, so Google's
live model ids — part of Antigravity's real on-disk contract — are untouched.

Homoglyph folding is deliberately not attempted, and a Cyrillic 'і' would slip
through. These values arrive from argv and env, trusted inputs here, and a
mapping broad enough to catch deliberate homoglyphs would start catching
legitimate ids. Stated rather than left for the next reader to discover.

This over-broad-match trap is the recurring shape of the whole epic: an
exclusion or match written wider than its subject. Four occurrences, each
cited: #4716's `gemini-[0-9]` sweep exclusion hid a stale review.models.gemini
row whose value was "gemini-2.5-pro" on the same line; #4753's first
model-display escape was a blanket /^ \d/ that laundered "Gemini 2.5 CLI as a
supported runtime."; its dialect rule then used a +/-24-character window that
let one legitimate reference license a live claim 21 characters away; and its
model rule treated the ABSENCE of a runtime word as a grant, passing five
unqualified live-runtime claims. Earlier drafts of this message and its
artifacts said "five" in one place and "three" in another with nothing cited;
it is four, listed here, and the artifacts now agree.

THE THROW

RetiredRuntimeError carries `code: 'GSD_RETIRED_RUNTIME'` so a caller can
handle this case without string-matching a message that may be reworded, and
the message names the id, the successor and the retiring issue.
assertNotRetiredRuntime runs as the FIRST statement of each accessor, including
before getGlobalConfigDir's explicitDir branch, so an explicit directory cannot
mask a runtime that is gone.

`gsd-tools query project-instruction-file --runtime gemini` answered the new
throw with a raw stack trace — a user-facing regression this change introduced.
Its sibling routeSkillsRoot already emitted a clean single-line error for an
unknown runtime, so that route now maps GSD_RETIRED_RUNTIME through the same
`error()` helper, and a test asserts the contract directly: non-zero exit,
stderr naming Antigravity and #1928, and no stack frame. It was the only
unwrapped call site in that CLI; I checked the rest rather than assuming.

getRuntimeNewProjectCommand is deliberately NOT guarded: its value does not
vary by runtime in a way that makes a retired id a wrong answer, so throwing
would cost callers a crash without correcting anything. Verified by observing
it return the same value across claude, codex, opencode, kimi, antigravity,
copilot and an unknown id.

RECONCILING THE TESTS THAT PINNED THE DEFECT

The full remote matrix went red with 14 failures, and every one was a
pre-existing test asserting the fallback this criterion calls a defect. One had
already been caught locally by review; the matrix found the other thirteen
across four files. They were reconciled by intent, not blanket-inverted:

  - Tests whose SUBJECT is the retired runtime — "gemini falls back on label /
    config-fragment / new-project surfaces", "gemini no longer maps to
    GEMINI.md (defaults to AGENTS.md)", "gemini is no longer a known runtime —
    falls back to AGENTS.md" — had pinned the defect, titles and all. Their
    assertions are INVERTED rather than deleted, so the history of what the
    behaviour used to be stays attached to the test that pinned it.
  - Tests whose SUBJECT is "an unregistered id falls back generically", with
    gemini merely the SAMPLE, still assert a TRUE property that this change
    deliberately preserved. Those keep their assertion and switch the sample to
    a genuinely unknown id, with a retired-id refusal pinned alongside so both
    halves of the distinction sit together.
  - The project-instruction-file parity loop dropped gemini from its
    parametrised runtimes — both sides now refuse, so there is no value to
    agree on — and gained a dedicated refusal-parity test.

A FIFTEENTH was then found by executing the touched suites locally, in process,
one file at a time — `tests/runtime-name-policy.test.cjs:135` asserted
`getProjectInstructionFile('gemini-cli') === 'AGENTS.md'`, and its own comment
read "gemini-cli was an alias for gemini", which is exactly why that spelling
is now a retired one rather than a merely-unrecognised one. Inverted like the
rest.

Two remote runs on this change were avoidable: the first by reconciling the
tests that pinned the old behaviour before shipping, the second by executing
the touched suites locally first. The matrix is the authority; it is not the
discovery mechanism. Local per-file execution is bounded and cheap and is not
the banned `node --test` fan-out.

All five touched suites now pass in process: runtime-name-policy 47/47,
gemini-runtime-removed 32/32, project-instruction-file-parity 12/12,
runtime-homes-legacy-ids-drift-guard 2/2, install 452/452.

COVERAGE

Failing-first, one per accessor as the criterion demands, each proven RED
against 5d4c98cde7 before the fix existed — the table at the top of this
message IS that baseline, and the exports the tests import did not exist yet
either.

Asserting only the throw would pass if every id threw, which would break every
install, so each property is paired with its opposite: every canonical id still
resolves on all five accessors with byte-identical values; '' keeps its
documented branch; a genuinely unknown id keeps 'Claude Code' / 'AGENTS.md' /
'.claude' / ~/.claude. That last one is the load-bearing negative — it is the
decision the maintainer chose to preserve, so a later patch that "tightens" the
guard to reject all non-canonical ids turns it red with the reason attached.

Boundary coverage maps limit-1/limit/limit+1 onto set membership: 'gemin',
'geminix', 'gemini-2.5-pro' and 'gemini-3.1-pro-preview' must NOT throw, the
retired id and its folded spellings must. '__proto__', 'constructor' and
'  CONSTRUCTOR  ' are pinned as must-not-throw, and predicate/assertion
agreement is asserted directly. Several assert.throws calls initially passed a
string as the second argument, which node treats as the MESSAGE rather than a
matcher, so they asserted nothing about the error; they now use a real
predicate checking the code.

The tests live in the owning modules' suites rather than a new issue-named
file: lint-regression-test-names rejects new bug-NNNN/fix-NNNN/issue-NNNN test
files outright and directs the regression to the owning module's suite.
scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own
--write path, since the tracked test-file count moved.

Fixes #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): backfill changeset PR number (#4756)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:22:16 -04:00
Rezolv
85026f6a05 feat(#4668): add StateWriteIntent type surface and opaque-transform guard recognition (ADR-4629 C1) (#4676)
* feat(#4668): add StateWriteIntent type surface and opaque-transform guard recognition (ADR-4629 C1)

Child C1 of epic #4629 — migration-order step (1) of ADR-4629: the guard/type
scaffolding, with NO behavior change and no caller migrated.

1. StateWriteIntent (src/state-transition.cts) extends StateTransaction with the
   ADR-4629 section 8.1 concepts: field/section assertions marked required vs
   best-effort, plus a declared mutation scope (narrow | broad). Frozen like its
   base. createStateWriteIntent builds one from an existing transaction. Nothing
   in production constructs it yet — section 8.1's caller-side rule is Required in
   Phase 2; C2 (the verifying executor) and C3+ (caller migration) consume it.

2. findOpaqueStateTransforms (scripts/lint-state-write-path-drift.cjs) recognizes
   a residual readModifyWriteStateMd(path, (content) => ...) write whose transform
   is an inline anonymous arrow/function — the opaque shape section 8.1 replaces
   with a declared StateWriteIntent. readModifyWriteStateMd goes THROUGH the seam
   (it is not a raw-write bypass, Axis 2's concern), but its opaque body transform
   is neither verified (section 8.2) nor bounded (section 8.3).

   This ships recognition as a CAPABILITY: exported and unit-tested (positive
   control on a seeded fixture) but DELIBERATELY NOT wired into collect()'s
   failing scan. Wiring it now would turn the ~16 residual callers red at once,
   and ADR-3473 section 8.6 retired the ratchet that would otherwise absorb them.
   C2 wires it terminal as the verifying executor lands and callers migrate under
   ADR-3408 section 6 phasing.

No behavior change: the guard is green on the tree (detection not wired), every
state verb's output is unchanged, and the relevant suites (1763 tests) plus
lint:ci pass. Regression tests are failing-first: positive/negative controls for
the guard capability and a shape test for the type, plus a pin that collect() has
no opaque-transform findings (C1 must not enforce; that is C2).

Closes #4668

* chore(#4668): backfill changeset pr field to the real PR number (#4676)

pr: 0 is rejected by parseFragment as invalid_pr (it is not a valid placeholder);
the fragment must carry the real PR number, which fixes both changeset-lint and
docs-lint (fail_invalid_fragment / fail_malformed_fragment).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-14 21:43:49 -04:00
Tom Boucher
5d4c98cde7 chore(#4729): guard the retired-runtime name, and finish the locale residue (#4753)
* chore(#4729): guard the retired-runtime name, and finish the locale residue

Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one
concern: make the tree clean, and keep it clean. The guard is inert until the
tree is clean, and shipping the cleanup without the guard is the
one-bug-at-a-time pattern this epic exists to end.

WHY A GUARD, AND WHY LAST

Nothing in CI answered "does any shipped surface still present a retired
runtime as live?", and the two gates that look like they should cannot.
checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every
declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it
reported 0 violations while all four locale mirrors still documented --gemini
as a live reviewer flag, with usage examples. And
tests/gemini-runtime-removed.test.cjs is scoped by construction - its own
docblock limits it to the installer CLI contract and the runtime-name-policy
exports; it never reads docs/**, gsd-core/workflows/**, commands/** or
agents/**. Every extension to it during this epic was a hand-added assertion
for a surface somebody had already noticed.

A guard written earlier would have red-flagged the very references phases
1b-4b were removing, which is why it lands last.

PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE

Each site was judged against its ENGLISH counterpart, not on its own:

  README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46  English README.md has ZERO
                                                  occurrences -> substituted
                                                  "Antigravity CLI, Kimi CLI"
  how-to/execute-a-phase.md:88  x4 locales        fixed in #4728 -> substitute
  how-to/verify-and-ship.md:89  x4 locales        fixed in #4728 -> substitute
  FEATURES.md cross-AI CLI list                   :1419 no Gemini -> DELETE
  FEATURES.md REQ-MULTI-RT-01                     :1709 -> substitute
  FEATURES.md REQ-SKILLS-03                       :1952 -> rewrite
  FEATURES.md REQ-QUOTA-02                        :3256 deleted upstream -> delete
  VERSIONING.md:133                               stale manifest -> see below

The twelve README occurrences were an adversarial reviewer's BLOCKER, and the
reason they survived my own sweep is structural: root-level *.md was outside
the guard's scan set, so the repo's most-read runtime-advertising surface was
invisible to the guard meant to police it. :46 is a live installer-runtime
claim - it tells the reader the installer will offer a runtime that no longer
exists. Checked for the duplicate-name trap before substituting: neither
Antigravity nor Kimi appears anywhere in those four files.

Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and
verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work,
not a deferral.

Two more show why "substitute Gemini -> Antigravity" is the wrong default: in
the cross-AI list and REQ-QUOTA-02 English DELETES the name, because
Antigravity was already in the list or the classifier had dropped it.
Substituting would have duplicated a name - the identical trap
ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot.

VERSIONING.md:133 is a different and worse defect than translation lag. Under
"Manifest Version Sync" it listed gemini-extension.json as a version-synced
manifest. That file is ABSENT from the repo, and
scripts/sync-manifest-versions.cjs says so in its own comment - "#1928:
gemini-extension.json was removed with the gemini runtime ... it is no longer
a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json,
marketplace.json and vscode/package.json. So the doc named a manifest that
does not exist AND omitted the one that replaced it. Both fixed, verified
against the owning code rather than inferred from the name. The replacement
bullet cites #1942, the issue that actually registered vscode/package.json,
matching the convention of its neighbours.

pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has
no REQ-QUOTA-02 line. Skipped and recorded, never invented.

PART B - THE GUARD

scripts/lint-retired-runtime-name.cjs, modelled on
scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem
shape (forbid a retired token, allowlist frozen content, self-exempt via a
split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1).

Case sensitivity IS the mechanism, not an accident. The naive guard - "the
string gemini must not appear" - is WRONG, not merely noisy: that string is
load-bearing across Antigravity's real on-disk contract. A case-sensitive,
standalone, capitalised name works because every legitimate reference is
spelled differently and therefore cannot match: lowercase config homes
(~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model
ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and
GEMINI.md. Table-driven, so the next retired runtime costs one row.

THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two
rounds of isolated adversarial review reshaped it; both are recorded in
.gsd/bug/chore-4729-gemini-drift-guard/60-review.json.

ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both
Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can
never be complete, so "no runtime word found" silently exempted every phrasing
nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`,
`Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0,
as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and
`Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime`
or `运行时`. The fix was to stop enumerating exceptions and invert the evidence
direction:

  Tier 1(a) - the hook DIALECT Antigravity inherits. Position is
  language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini
  スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo
  Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The
  marker must now form an ADJACENT COMPOUND with the name, not merely sit in a
  +/-24-character window - that window let `| Antigravity | Gemini-style hooks
  | Gemini support is live |` exit 0, one legitimate reference licensing a
  fresh live claim 21 characters later. The runtime-word veto is now
  LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a
  runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files"
  in the same sentence); each is an explicit pin rather than a reason to
  loosen the veto for everyone. Measured: widening it surfaced exactly those
  ten and no others.

  Tier 1(b) - the provider/model axis. A version optionally followed by a
  qualifier, including full-width digits and CJK punctuation, AND positive
  model-axis evidence on the line, AND no runtime word. The positive
  requirement is the part that matters: all eight real model-axis lines in the
  repo name a model explicitly, so requiring it costs nothing on the real tree
  while flagging every laundering attempt. It is also the honest resolution of
  the agent/target tension below - rather than guess at an exhaustive veto
  list, stop treating an empty veto as evidence.

  Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match
  falling INSIDE an occurrence of its own snippet. Line-level containment let
  `Known provider menu update: Gemini CLI is once again a selectable GSD
  runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your
  GSD runtime.` both exit 0, because a short snippet elsewhere on the line
  pre-approved a brand-new claim. Span scoping makes short snippets safe:
  `Google (Gemini)` can only ever excuse the match inside those 15 characters.
  A LOAD-TIME validator now requires every pin to contain a retired name, and
  it immediately caught five of MY OWN pins whose snippets sat BESIDE the name
  rather than covering it - each would have shipped permanently inert and
  permanently reported stale. All pins were then reconciled in one pass.

  A pin is also marked used by PRESENCE on the line now, rather than only on
  the Tier-2 branch. Previously a pinned line that a general rule also matched
  never marked its pin used, producing a provably FALSE "no line matches
  pinned snippet" whose printed remedy told the maintainer to delete a pin
  that was still needed.

  Tier 3 - blanket trust, and a new occurrence inside it IS invisible.
  CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one
  surface - plus six append-only directories. All 21 `.changeset/` hits were
  measured to be fragments DESCRIBING the retirement or a fix to it, 464 of
  them under archived/; a fragment can only describe what already shipped and
  is deleted at release, so pinning them would be friction with no signal. The
  cost is stated in the guard's own header rather than hidden.

THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix
list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and
`references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a
live claim introduced there was invisible at BOTH ends.

The escape hatch must now carry a justification
(`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is
checked first, excuses the whole line, and the failure message advertises it,
so an unexplained one is indistinguishable from a silenced defect.

Plus an anti-vacuity floor counting files actually READ, not files listed - a
candidate count stays healthy-looking even if every read failed.

A FALSE NEGATIVE I INTRODUCED, AND CLOSED

The model-display escape began as a blanket /^ \d/ - "space then a digit" -
which also matched "Install for Gemini 2.5 CLI as a supported runtime.",
laundering a genuine stale-runtime claim through an attached version number.

That was the THIRD appearance of one failure shape in this epic: an exclusion
added to suppress false positives creating a false negative. #4716's sweep
excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby
hid a stale review.models.gemini row whose example value was "gemini-2.5-pro"
ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which
is why the fix this time was to invert the rule's evidence direction rather
than to enumerate more exceptions.

The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI"
matched inside "client" and would have vetoed legitimate model lists - and raw
for CJK terms, where \b is ASCII-word-based and would never fire beside an
ideograph, so anchoring them would silently disable the veto in ja/ko/zh.
"agent" and "target" were deliberately left OUT: both occur throughout
ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so
vetoing on them would red correct content instead of catching runtime claims.
The reasoning is in the guard's comment, not just the omission - and Tier
1(b)'s positive-evidence requirement is what makes that omission safe, since
the rule no longer depends on the veto list being complete.

COVERAGE

tests/lint-retired-runtime-name.test.cjs drives the guard through its
GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring
tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a
guard, and this epic already shipped one that was vacuous for 2 of its 5
files, so properties are paired against BOTH failure modes - too broad
silently absorbs a future defect, too narrow reds on legitimate content. Floor
boundaries are covered at 149/150/151.

The round-2 reviewer's sharpest point was about that claim, and it was right:
the first matrix's pairing was "true of the properties chosen, not of the
predicate's actual surface" - not one of its twenty properties could see the
dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an
over-broad pin colliding with a new line. Every one of those is now a
committed regression using the reviewer's own attack line verbatim, and the
local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0).

That harness earned a finding of its own. Its first run reported PASS=2
FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so
nothing staged, every fixture hit the empty-walk error path, and the two
checks that assert an ABSENCE passed off that error path rather than off real
guard logic. A staging failure is now fatal and every absence-asserting check
first proves the walk ran and the expected violation was flagged. Later, one
case failed because its fixture supplied only one of a pinned file's two
approved lines, so the stale-pin check fired correctly - the expectation was
wrong, not the guard. Telling those two apart is the whole value of running a
matrix rather than reasoning about one.

On the two orthogonal reviews: the isolated adversarial pass executed a great
deal of code, across two rounds, against its own fixture repos. The security
pass did NOT - it self-discloses that it verified by reading only, because
node --test is hard-blocked here. Saying so plainly, because "two orthogonal
reviews" without that caveat overstates what the second one established. It
also raised, and I cleared by measurement, a concern that importing
escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt
clone: six other tracked scripts already require that exact path, three of
them already in lint:ci, and .github/workflows/test.yml:192-193 runs
`npm run build:lib` immediately before it for exactly this reason.

Part A has no new test deliberately - those edits are covered by the guard
itself inside lint:ci, and a separate per-locale assertion would duplicate it
and then drift from it. The one exception is the root README case, which IS
pinned: that residue was invisible to the guard rather than merely unasserted,
so the fix is a scan-set change and needs its own regression test.

No mode-bit read-failure fixture was added on purpose: the benches run as
root, where chmod-based IO injection is vacuous, so such a test would assert
nothing.

The test's fixture helpers write throwaway docs/ paths, which trips
lint-docs-guard-registration's reader-name heuristic. Resolved the way that
lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry
- because the fixtures only WRITE scratch data and never read shipped docs;
the baseline was re-confirmed, not merely extended, each time locale and
adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs
regenerated through its own --write path, since a new test file changes the
count lint:generated-sync reads.

Fixes #4729

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4729): backfill changeset PR number (#4753)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 20:50:26 -04:00
Twisted Fate
39db16a653 fix(#4678): resolve :LINE citations against their underlying files (#4719)
* fix(#4678): check :LINE citations instead of dropping or misreporting them

cmdVerifyReferences had two opposite failures on citations carrying a
trailing line suffix. The backtick extractor anchored the closing
backtick right after the extension, so `src/foo.ts:42` matched neither
regex and was silently dropped -- an all-line-numbered document
reported {valid:true, found:0, missing:[], total:0}, indistinguishable
from one with no citations. The @-extractor treated ':' as a legal path
character, so @src/foo.ts:42 was probed with the suffix glued on and a
resolvable file was reported missing.

Strip the :N / :N-M suffix (stripLineSuffix) before existsSync in both
loops -- for filesystem resolution only; found/missing keep reporting
the original citation text -- and let the backtick regex accept the
optional suffix so those citations are counted at all. URL skip,
template-placeholder skip, dedup, ~/ expansion and the output contract
are unchanged. Whether a line number past EOF counts as missing is an
open design question and stays out of scope.

* chore(#4678): backfill changeset pr field with PR number

The fragment shipped as the documented pr: 0 placeholder; the merge
gate requires pr > 0 before a fragment can land, so backfill 4719 now
that the PR number exists.

---------

Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-14 20:40:46 -04:00
Tom Boucher
77ff068cc3 docs(#4722): record Kiro and Zoo Code runtime requests as out-of-scope (#4750)
* docs(#4722): record Kiro and Zoo Code runtime requests as out-of-scope

Both #4722 and #4746 asked to add a new first-party, in-tree runtime
(Kiro, Zoo Code). The repo's standing policy against expanding the
in-tree runtime set already covers this shape of request
(crush-runtime-in-core.md, omp-runtime-in-core.md); recording these two
so the same asks aren't re-litigated from scratch next time they land.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4722): correct ADR-857 D8 status in the new out-of-scope entries

Reviewer caught that the Re-open criteria section, inherited verbatim
from the crush/omp sibling entries, claimed the ADR-857 D8 external
capability loader "has not been delivered." ADR-1244 (Accepted,
ratified 2026-07-17) explicitly states it delivers that exact gate for
third-party role:"runtime" capabilities. Correct both new entries to
reflect that the mechanism exists today, while keeping the in-tree
non-acceptance verdict itself unchanged (a maintainer bandwidth/scope
call, not a tooling gap).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4722): match the newer out-of-scope house format

Orthogonal review (Standards axis) noted the two newest sibling KB
entries (crush-runtime-in-core.md, codex-native-plugin-skips-preproposal.md)
lead with a dedicated "## Policy (standing, not case-by-case)" section
before "Proposal summary"; these two files folded the same claim into
a bullet instead. Restructure to match, no content change beyond
de-duplicating the policy statement out of the bullet list it was
also stated in.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 19:58:38 -04:00