Commit Graph

1094 Commits

Author SHA1 Message Date
Jakub Zych
792139b5ed chore: sweep Kimi mentions from comments and notes
Some checks failed
Tests / PR mergeability (push) Successful in 19s
Tests / Base branch health (push) Successful in 10s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 1m43s
Tests / plugin-validate (push) Successful in 1m7s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 18s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 19s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 19s
Tests / test (ubuntu-latest, 24) (push) Failing after 17s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 18s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 8s
Tests / Required tests (push) Has been cancelled
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
2026-10-06 20:19:17 +02:00
Jakub Zych
6cfa0c55d2 refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes,
cline, codebuddy and pi end to end: capability descriptors, installer branches
and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters,
hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi
migrations, Kimi payload normalization in the hook guards, dead hostBehaviors
vocabulary, launcher home probes, fixtures, runtime-specific tests and the
prose that presented them as supported.

Installer output for the six kept runtimes is byte-identical to before the
prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and
read-injection-scanner are left in place pending a decision.
2026-10-06 20:02:40 +02:00
Jakub Zych
12ee75a509 chore: point MSD at git.golem15.com/golem15/msd-core
Some checks failed
Tests / PR mergeability (push) Successful in 1m37s
Tests / Base branch health (push) Successful in 11s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 2m16s
Tests / plugin-validate (push) Successful in 1m6s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 25s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24) (push) Failing after 19s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 19s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 9s
Close Draft PRs (sweep) / Sweep open draft PRs (push) Successful in 8s
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
Tests / Required tests (push) Has been cancelled
The fork lives on the golem15 Gitea forge, not GitHub. Package identity now
parses either host and derives in-place raw URLs for Gitea; the identity-drift
lint accepts the new host; README drops GitHub-only badges and the npm
quickstart in favour of the checkout installer.
2026-10-06 10:41:56 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
0xdhx
238bee7b03 fix(#4881): trust worktree.baseRef:"head" in harness mode; a WorktreeCreate hook withholds it and only a measured fork base restores it (#4921)
* fix(#4881): trust worktree.baseRef:"head" in harness mode and keep the #4868 observation as what restores it under a WorktreeCreate hook

The pre-dispatch base check still derived its harness-mode verdict from the
retired #48 premise that the harness never reads worktree.baseRef: with
"head" set and HEAD diverged from origin/HEAD it degraded every wave with
baseref-head-ignored-by-harness. #4868 inserted an observation of a clean
prior harness worktree at HEAD ahead of that comparison, but an
execute-phase run never has one at the moment it checks — the base-check
runs before any dispatch, a degraded wave creates no worktrees, and a wave
that did run in worktrees has them removed and HEAD moved before the next
check — so the common case was unchanged (#4881 repro states 1 and 4).

Re-scope of the closed #4752 onto current next, with #4868 kept:

- branch a trusts "head" in both isolation modes, the way the harness is
  measured to behave (#4588: three settings layers, three OSes), and the
  spawn-time exit-42 guard stays the observation-based backstop;
- a Claude Code WorktreeCreate hook in any settings file the check reads,
  or a file that does not parse, withholds that trust — the hook creates
  the worktree without applying the setting — and the inferred comparison
  runs, degrading with baseref-head-bypassed-by-hook;
- the #4868 observation (b2) now sits behind branch a: it is reached only
  when "head" was not trusted outright, and on a hook host it is what
  restores the trust — a hook that forks from HEAD leaves exactly that
  evidence, one that forks elsewhere never does. Its per-HEAD cache is
  unchanged. It is skipped under an explicit --observed-fork-base, which
  outranks an inference from a prior worktree;
- --observed-fork-base <sha> threads a measured fork base through the
  evaluation (strict full-hex, TypeError otherwise).

Tests: the #4868 rows are unchanged and still reachable (they run with the
setting unset); the #4752 rows re-land, with the one exit-128 row updated
to the degrade #4734 pinned since; a new #4881 block pins the start-of-run
state (baseref-head, git never consulted), the hook + observation
composition in both directions, and that an explicit observation skips
the probe.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* docs(#4881): rewrite the eight surfaces that still state the harness ignores worktree.baseRef:"head"

Every prose surface #4868 left untouched still asserted the retired #48
premise as verified fact, starting with the step file the orchestrator
reads. Each now describes the measured behaviour, the WorktreeCreate-hook
exception, the --observed-fork-base input, and the #4868 observation as
what lifts the hook degrade; docs/CLI-TOOLS.md gains the
fork-from-head-observed reason row #4868 did not document.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* chore(#4881): set changeset fragment pr to 4921

* fix(#4881): withhold the #4868 observation under the hook interlock

The hook interlock this PR added withheld the worktree.baseRef:"head"
trust but still let b2's prior-worktree observation restore it, and that
observation cannot be attributed to the hook. The evidence is a clean
agent worktree sitting at the orchestrator HEAD; nothing on disk records
which creator left it there, so one the plain harness created BEFORE a
WorktreeCreate hook was configured — with HEAD unmoved since — reads as
evidence for the hook. It was the one fail-open branch in a mechanism
documented as fail-closed.

Keying the observation cache by hook configuration does not close it.
observeHarnessForkFromHead has two legs: a HEAD-keyed cache and a live
probe over .claude/worktrees/agent-*. A hook-keyed cache simply misses,
and the miss falls through to the probe, which re-finds the same stale
worktree and re-confirms. The probe takes no hook input at all. So the
observation is not consulted under the interlock rather than re-keyed:
on a hook host the only admissible positive signal is an explicit
--observed-fork-base measurement of the dispatch in hand, and absent one
the inferred comparison runs and a mismatch degrades with
baseref-head-bypassed-by-hook, leaving the spawn-time exit-42 guard as
the backstop.

Scoped to the case branch a. declined to trust: "head" set AND a hook
(or an unparseable layer) in the harness's path. With no "head" setting
b2 is #4868's own arm and is unchanged, hook or not — re-scoping that
trust is a separate question this PR does not open, and a test pins the
boundary.

Cost, stated: a hook host with "head" set, on a branch diverged from
origin/HEAD and passing no observation, now runs sequentially. It still
runs parallel when HEAD matches origin/HEAD. No workflow threads an
observation today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* refactor(#4881): drop the now-unused forkRef message-builder parameter

buildMsgBaserefHeadIgnored stopped reading forkRef when the message was
made mode-neutral, and the parameter was retained with `void forkRef;`
for symmetry with its two sibling builders, which do read it. Symmetry
is not reason enough to keep a dead parameter on a module-private
function with one caller, so drop it (#4921 review).

Behaviour is unchanged; the message text is pinned by an existing
full-string assertion, which is what covers the only real risk here —
transposing the two remaining arguments at the call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* test(#4881): exercise both FULL_SHA_RE alternatives at their boundaries

The invalid-observation list pinned 39 and 41 hex around the 40-hex
SHA-1 arm but left the 64-hex SHA-256 arm's own +/-1 boundary
unexercised, which the repo's boundary-coverage convention asks for
(#4921 review). Adds 63, 65 and a 64-length non-hex string.

The regex already rejected all three -- this is coverage of correct
behaviour, not a fix -- so it carries no negative control against a
pre-fix base. That it is not vacuous was shown instead by widening the
arm to {63,65}, under which the row fails by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* docs(#4881): finish the surface sweep the hook interlock owes

Self-found by this round's own pre-push adversarial review, over six passes.
Eight sites, four classes.

FIVE were prose still asserting the stance the interlock overturned, which
reads as live canon to anyone arriving cold: docs/CLI-TOOLS.md,
docs/CONFIGURATION.md and gsd-core/references/planning-config.md each still
said the hook degrade is lifted "unless/until a clean prior harness worktree is
observed"; observeHarnessForkFromHead's own header still said a qualifying
worktree "can only exist if the harness forked from HEAD"; and a test name
still called the observation "required on a hook host" when it is now
inadmissible there. My own sweep had grepped for "restores"/"lifts" and missed
every one -- the ordinary failure of a grep, which returns what you thought to
search for.

The SIXTH is the same class one step worse: gsd-core/workflows/execute-plan.md
still said flatly that Claude Code's isolation="worktree" "forks from
origin/HEAD, not live local HEAD" -- in a paragraph THIS PR already edits, a
few sentences after the clause it corrected. A tombstone makes only its own
line clean; adjoining text asserting the dead stance is the other half of the
same defect. Now qualified on the setting, with the hook exception named.

The SEVENTH is a proof-strength overstatement that predates this PR, with a
driven counterexample: a worktree created from an older base and since `git
checkout --detach`ed onto HEAD is clean, sits at HEAD, and satisfies the probe
identically, so "can only exist" was false. The worktree's own reflog does
retain that original checkout -- the information is not lost, the probe simply
does not consult it.

The EIGHTH is that the header described only one of the function's two legs. A
cache hit returns the prior conclusion without reading any worktree, so "the
probe reads a worktree's present state" was true of the live probe and false of
the cache. The header now separates them, and names the cache's blindness as a
third reason the observation is inadmissible under a hook.

Gaps seven and eight are inherited from #4868 and accepted there for the
no-hook case. Nothing about the mechanism changes here; only what the header
claims for it.

No behavioural change -- comments, prose, and one test's registered name.

Emitted-Drift-Ack-Growth: execute-plan.md — the Pattern A paragraph gained a qualifying
 clause: it stated flatly that Claude Code forks from origin/HEAD, which is the premise
 this PR retires, a few sentences after the clause the PR had already corrected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 20:15:28 -04:00
Tom Boucher
d7b5b2c2b6 fix(#4906): migrate the ROADMAP.md Plans: line onto the PlanningDoc seam — Phase 2 (#4933)
* feat(#4906): migrate the ROADMAP.md **Plans:** line onto the PlanningDoc seam

Phase 2 of epic #4906. Migrates the two writers of ROADMAP.md's Plans field onto
the seam ADR-4910 locks, and deletes both bespoke regexes per Decision 2 — a
correct copy of a rule the seam now owns is the same divergence risk as an
incorrect one.

src/phase.cts's mutateMilestonePhase carried planCountBodyPattern, one capture
group, replace-to-end-of-line: this is #4852, still live before this change.
src/roadmap.cts's cmdRoadmapUpdatePlanProgress carried the correct three-arm
sibling (the #2853/#3584 correction) that phase.cts never adopted. Both now call
findField/setFieldValue/serialize against a PlanningDoc parsed from the same
milestone/phase-confined substring their existing withPhaseSection /
replaceInCurrentMilestone wrappers already compute — those confinement
wrappers are unchanged, only the field-write mechanism inside them moved.

Verified end-to-end through the real commands against real fixtures, not
against an isolated reimplementation of the classification logic:
cmdRoadmapUpdatePlanProgress and cmdPhaseComplete both preserve a trailing
human annotation across a real count rewrite, and both leave a bracketed
human annotation (the #3584 Finding A discriminator) untouched.

Found and fixed inline, in the already-merged src/planning-document.cts,
rather than deferred: BOLD_FIELD_RE recognized only **Label:** (colon inside
the closing bold). gsd-core/templates/roadmap.md ships every field, Plans
included, as **Label**: (colon outside) -- migrating roadmap.cts onto the
seam as it stood would have silently regressed real generated ROADMAP.md
files back to the bug this migration exists to remove. Widened to recognize
both spellings; deliberately NOT widened to a bare unbolded Label: form,
which would register ordinary prose as a spurious field.

roadmap.cts's writer also recognizes a bare singular/plural count (1 plan /
3 plans, no fraction) as an existing count token to overwrite, not template-
placeholder or freeform prose -- the template's own single-plan-phase shape
and #3584 Finding B's fix. Preserved exactly; this shape is easy to drop by
only porting the more common fraction form.

Two of Phase 2's three originally-cited defects turned out to be already
fixed on next, independent of this epic, and are struck via a dated ADR
amendment rather than silently narrowed: #4862 (stateReplaceField's own
anchoring hardening already preserves sibling fields) and #4499
(spliceFrontmatter's per-key preservation already keeps block sequences
byte-identical). Both reproduced against the built module before being
struck, not assumed. STATE.md's field-write engine is re-scoped out of this
phase entirely -- not because of its get_impact rating alone (measured the
same way, the two sites THIS phase keeps are also CRITICAL, and an earlier
draft of the amendment claimed otherwise without checking; corrected) but
because updateCore is a multi-field transaction with frontmatter sync and
preservation reconciliation that does not map onto PlanningDoc's node model,
where migrating it would mean designing that model, not calling an existing
seam function.

Refs #4852
Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4906): preserve a trailing annotation with no em-dash separator

Test authoring surfaced a real regression against an existing #3584 fixture:
`**Plans**: 0/1 plans executed (11-16 are gap closure from VERIFICATION)` -- a
parenthetical annotation glued on with a bare space, no em-dash -- was left
completely untouched by the migrated code instead of being rewritten with the
count updated and the parenthetical preserved.

Root cause: planning-document.cts's parseBoldFieldLine splits a field's value
from its trailing annotation only on the literal " -- " separator. An
em-dash-separated annotation already lives outside `value` in `trailingSpan`,
untouched by setFieldValue regardless -- that path was never broken. A
parenthetical with no em-dash has nowhere to go but inside `value`, and the
migrated classification required the WHOLE value to match a count-token shape
exactly, so this case fell into "leave untouched."

Fixed in the migrated call sites, not in the seam: prefix-match the count
token against the field's current value, then re-glue whatever textual suffix
follows WITHIN that value onto the new count text before writing. Correct for
both shapes with no special-casing -- the em-dash case's suffix-within-value
is empty by construction (the annotation already lives outside value), the
parenthetical case's suffix is exactly the glued content, preserved verbatim.

Deliberately not fixed by widening planning-document.cts's separator grammar
to also recognize a bare-space-then-parenthesis: that seam is already-merged
and already-tested, and guessing at an open-ended set of annotation shapes at
the seam level is exactly what isTemplatePlaceholder already avoids by
staying caller-side. "What counts as a Plans-field count token" is domain
knowledge about this one field.

phase.cts's writePlansField had no arm-2/arm-3 classification before this
migration -- its original regex unconditionally overwrote whatever value was
present. That unconditional-overwrite behavior is preserved exactly for
values with no recognizable count-token prefix; only the recognized-count
case gained suffix preservation, matching what phase.cts actually did before.

Verified end-to-end via the real cmdRoadmapUpdatePlanProgress and
cmdPhaseComplete commands against real fixtures for both the parenthetical
and em-dash shapes at both sites.

Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4906): cover the migrated Plans-line writers at both sites

15 cases across tests/phase.test.cjs and tests/roadmap.test.cjs, one per row
of the phase test matrix, extending the existing describe blocks and helper
functions those files already use for these two commands rather than
building parallel fixtures.

Covers: trailing-prose preservation on a real count rewrite at both sites
(the #4852 regression, and the #2853/#3584 non-regression); zero-trailing-
content boundary; the bracketed-template-placeholder vs bracketed-human-
annotation discriminator (#3584 Finding A); a missing Plans field not
crashing either command; an unrelated unreadable sibling node in the same
confined section surfacing rather than corrupting the field; confinement
holding across sibling phases and milestones; a round trip through the
command's own read path; CRLF safety; and a parity assertion that both sites
now produce identical Plans-line text for identical inputs, proving one
shared mechanism rather than source-grepping for the deleted regex literals.

Authoring caught a real regression before it could land silently: an
existing #3584 fixture using a parenthetical annotation with no em-dash
separator failed against the first version of the migration. Reported rather
than edited to match the broken behavior -- see the paired fix commit. That
existing test needed no changes once the fix landed; its assertion was
verified independently against the real CLI before this commit.

Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4906): give phase.cts's Plans writer the same arm-2/arm-3 classification as roadmap.cts

Isolated adversarial review (mandatory orthogonal review pass) executed
writePlansField against `[Deferred pending re-scope]` and the fresh-template
placeholder wording and found the first version of this migration kept
phase.cts's OLD unconditional-overwrite behavior for the no-count-prefix
case instead of adopting the same isTemplatePlaceholder / arm-3-untouched
classification roadmap.cts's sibling site already uses. A bracketed human
annotation was being silently rewritten to a new count -- a real
content-destroying regression against this phase's own design-doc behavior
table, not an accepted trade-off, and exactly the kind of divergence between
the two sites Decision 2's parity requirement exists to eliminate.

writePlansField now runs the same template-placeholder check and "no count
prefix and not a placeholder => leave untouched" branch before ever calling
setFieldValue. Added two site-1 tests (rows 4 and 6 of the phase test
matrix) mirroring the existing site-2 coverage for this exact discriminator.

Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4906): restore plain (non-bold) Plans: line support and fix a test helper mismatch

gsd-test (mandatory verification, run before every push) surfaced two real
defects the migration's manual CLI checks had not caught:

1. #1163 regression: a hand-edited/pre-template ROADMAP.md can carry a PLAIN
   (non-bold) `Plans:` line rather than the canonical `**Plans**:`/
   `**Plans:**` bold field. The old deleted regexes tolerated this shape;
   the seam's BOLD_FIELD_RE is deliberately bold-only (widening it would
   register ordinary prose like "Note: see below" as a spurious field
   seam-wide), so the migrated writers silently no-op'd on it instead of
   updating the count -- a real, previously-tested behavior lost.

   Fixed with a caller-side fallback in both src/roadmap.cts (where the
   failing #1163 test lives) and src/phase.cts (added for parity, per
   Decision 2 -- the two sites should not diverge on which legacy shapes
   they tolerate): when findField finds no boldField Plans node, look for a
   plain `Plans:` line directly and apply the same arm-1/2/3 classification
   against it. This is domain knowledge about one field's legacy tolerated
   shape, the same class of thing isTemplatePlaceholder already keeps
   caller-side rather than seam grammar.

2. tests/roadmap.test.cjs's new rows 4/5 (site 2) seeded the colon-outside
   spelling (`**Plans**: ...`) but their plansLineIn() helper only matched
   colon-inside (`**Plans:**`), so both assertions compared against
   `undefined` regardless of whether the write logic was correct -- a
   test-authoring bug, not a source defect. Fixed the helper to recognize
   both BOLD_FIELD_RE spellings, matching what the production code actually
   supports.

Verified via a real gsd-test run before this fix (outcome: failed, 5
failures, all in tests/roadmap.test.cjs) and will be re-verified via a real
gsd-test run on this commit before push, per this repo's non-rationalization
rule: a red gate is fixed, never explained away.

Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4906): backfill the Phase 2 changeset fragment's PR number

pr:0 -> pr:4933, now that gh api POST /pulls has returned the real number.

Refs #4906

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 19:25:39 -04:00
Michel Moreira
1e3e1f7cd8 enhance(#4570): allow disabling planner stall detection (#4585)
* enhance(#4570): allow disabling planner stall detection

* docs(#4570): add changelog fragment

Emitted-Drift-Ack-Growth: plan-phase.md — the explicit opt-out gate covers all five planner and checker spawn classes
Emitted-Drift-Ack-Growth: settings-advanced.md — the toggle prompt and bounded-recovery warning expose the new setting

* chore(#4570): refresh compact-content baseline

* fix(#4570): preserve default-on watchdog fallback

* docs(#4570): qualify the chunked-mode orchestrator rules with the toggle

The two chunked-planning-mode stall-watch imperatives read as
unconditional, with the opt-out stated only in a following bullet. State
the PLANNER_STALL_DETECTION_ENABLED condition inline, matching the three
sites already qualified in plan-phase.md.

* chore(#4570): refresh compact-content baselines against rebased next

* fix(#4570): sync planner stall launcher

Keep the stall-detection helper aligned with the canonical runtime launcher.

* chore(#4570): refresh compact baseline after rebase

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 18:46:23 -04:00
Michel Moreira
76ef60ba25 enhance(#4836): prefer the graphify CLI for planner and researcher graph queries (#4874)
* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries

The planner gets one knowledge-graph query per phase and the researcher two
or three, and that single shot decides which modules the plan treats as
related — and therefore how tasks are ordered into waves. It was spent on
the built-in reader, which seeds by case-insensitive substring match over a
node's label and description and then expands a hardcoded two hops. The
phase "User Authentication" seeds on `author`, `authoring` and
`unauthorized` with the same weight as `authenticate`, and when the
inflated payload exceeds `--budget` the trimmer drops edges by confidence
tier — so the highest-confidence tier can be discarded to fit a payload
that bad seeding inflated in the first place.

The graphify CLI is already a hard dependency of /gsd-graphify build, and
it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context
filters before traversal. Both prompts now prefer it and fall back to the
built-in reader, branching on `command -v graphify` — the same degradation
shape the repo already uses for Context7 to ctx7. Binary presence is a
self-satisfying gate: a graph can only exist if the binary built it, so the
fallback covers edge cases (a CI checkout with a committed graph, a binary
since removed), not the common path. No new config key and no new tool
grant — both agents already have Bash.

The planner additionally runs `graphify affected`. The reference states its
own goal as "which subsystems may be affected by changes in this phase",
which is literally reverse traversal by relation; the built-in reader only
approximates it with undirected two-hop expansion and has no equivalent
verb, so `affected` is skipped on the fallback path.

`graphify status` now reports `graph_path`, the resolved absolute graph
location, on both the present and the missing branch. The CLI takes the
graph location as `--graph`, and the prompts must not re-derive
`.planning/graphs/graph.json` for it: that would point the CLI at a
non-existent local mirror in exactly the umbrella multi-repo setup
`graphify.graph_path` (#1825) exists to serve. For the same reason the
presence gate in both prompts is now the `status` call itself rather than a
bare `ls` of the default location, which was already blind to the override.

Known limit, stated in both prompts rather than implied: the two paths
return different shapes. `graphify query` emits prose and has no `--json`
flag; the built-in emits JSON with per-edge confidence tiers and
budget_met/budget_estimate. `--budget` also counts rendered output on one
and estimated payload bytes on the other (#2738) — same flag name,
different unit. Both are read by a model and nothing machine-parses the
injected block. With graphify absent from PATH the injected context is
byte-identical to before.

Closes #4836

Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point.
Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes.

* docs(#4836): record the CLI-first graph query in the planner and researcher entries

* chore(#4836): add changeset fragment

* enhance(#4836): name the full domain word in the planner's query-term examples

The reference's own example — phase "User Authentication" → term "auth" — is
the exact collision the CLI-first path exists to avoid, and it stays a
collision whenever the fallback path runs, since that path matches the term as
a substring of label and description.

* fix(#4836): surface graph_path on the unparseable-graph status branch

graphifyStatus() returned graph_path on the exists:true and exists:false
outcomes but not on the third, error, outcome (graph.json present but
unparseable). The planner/researcher prompts gate CLI-first dispatch on
exists, not on this outcome, so a corrupt graph file made them fall
through to the CLI-first branch with the literal <graph> placeholder and
no real path to substitute.

* docs(#4836): note graph_path's trust boundary at the --graph interpolation

graph_path is reflected verbatim into a double-quoted --graph argument the
agent executes via Bash. It comes from graphify.graph_path, a config
surface already trusted elsewhere, so this isn't a new trust boundary --
but it is a new injection site (no --graph flag existed on this call
before). One-line caution for anyone hardening this later.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 19:59:45 -04:00
Tom Boucher
585cab41cc fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices (#4920)
* fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices

* docs(#4770): backfill the changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-21 15:30:52 -04:00
Tom Boucher
6a4984cf69 fix(#4763): surface the displaced session record and pass --phase from the executor decision loop (#4919)
* test(#4763): failing-first — replaced-record payload and executor --phase pins

* fix(#4763): surface the displaced session record and pass --phase from the executor decision loop

state record-session keeps its last-writer-wins write (the recorded single-slot
handoff design) but no longer displaces silently: when a non-empty Stopped At or
authored Resume File record is replaced, the payload carries the full prior text
under replacedRecord. Same-value rewrites, the insert path, and the #944
template-default DWIM are not displacements and report nothing.

The executor decision loop now passes --phase "${PHASE}" to state.add-decision,
matching execute-plan.md, so decisions stop inheriting whichever phase the
global pointer names (#4763 case 2). advance-plan is unchanged (#3311 by-design).

Emitted-Drift-Ack-Growth: gsd-executor.md — the decision loop gained its --phase guard and a comment naming why (#4763)

* docs(#4763): add the changeset fragment

* docs(#4763): backfill the changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-21 12:42:11 -04:00
Tom Boucher
55fba5f7ce feat(#4917): add the PlanningDoc parse → mutate → serialize seam — Phase 1 of #4906 (#4918)
* feat(#4917): add the PlanningDoc parse -> mutate -> serialize seam

Phase 1 of epic #4906, implementing ADR-4910 and its 2026-09-21 amendment. Net-new
leaf module; NO call site is migrated, so nothing in the twelve absorbed issues
changes behavior yet.

src/planning-document.cts composes the seams that already exist rather than
reimplementing them: markdown-sectionizer for structure (fences and code spans
come from stripFencedCode / scanInlineCodeSpans, never a second scanner),
markdown-table for tables, frontmatter for frontmatter, write-set for Result<T>.

What is structural rather than conventional:

- A field node carries labelSpan, valueSpan and trailingSpan separately, and the
  only write entry point takes a node id and writes into valueSpan. trailingSpan
  is readable and has no exported writer, so the #2853/#3584/#4852 rule ("the verb
  owns the count token ONLY") stops depending on an author remembering a third
  capture group.
- Mutation is node-addressed. A handle is minted by the parser, so a caller cannot
  name a node the parser did not find. No path strings — a path is a grammar, and a
  grammar needs a parser.
- serialize splices staged spans into the ORIGINAL buffer. A no-edit serialize is
  byte-identical, which is what eliminates the #4499 defect without a targeted fix, and which also
  makes an already-escaped table cell impossible to double-escape on a round trip.
- A node that fails to parse carries its own error and span; siblings stay readable.
- Per the ADR amendment, serialize REFUSES whenever any node carries a parse error,
  even with zero staged edits, naming the offender.

One defect was found while building and fixed in place rather than deferred:
PLANNING_ARTIFACTS first derived from isCanonicalPlanningFile unfiltered, so
config.json, state.json, milestone.lock and skill-manifest.json were accepted and
returned {ok:true, nodes:[]} — "this document records nothing" when the truth was
"I have no grammar for this file". That is the exact empty-vs-error confusion this
epic exists to remove (#4899, #4900), reproduced inside the seam built to prevent
it. The registry now filters to .md while still DERIVING from
isCanonicalPlanningFile, because hand-writing a second list is the divergence this
epic is about, and parsePlanningDoc refuses a non-markdown kind at the document
level per ADR-4910 section 5.

New bin/lib module bookkeeping: .gitignore, eslint.config.mjs ignore (ADR-457 —
lint the .cts, not the emitted .cjs), docs/INVENTORY.md row, regenerated
docs/INVENTORY-MANIFEST.json, and the CONTEXT.md glossary entry.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4917): cover the PlanningDoc seam across 28 input classes

30 cases in 14 describe blocks, one per row of the phase test matrix.

The two load-bearing tests are the fast-check properties (seed 20260921,
numRuns 200):

- a single node mutation leaves every byte outside that node's valueSpan
  identical to the source
- serialize with zero staged edits is the identity function

Both are DOCUMENT-SHAPED per CONTRIBUTING.md fixture provenance (#2371): the
generator assembles arbitrary frontmatter, heading, label and value text with
join(), and never calls serialize or any other function from the module under
test to build a fixture. Seeding the generator from the module's own writer
would make the document shape a constant, and the property could then never
explore a document the writer would not itself emit.

Verified the byte-range assertion is not vacuous with a control run: it passes
against the real writer and FAILS against a simulated #4852 writer (one capture
group, replace-to-end-of-line), which visibly drops the trailing annotation.

Boundary coverage is zero / one / two staged edits. Negative space carries its
own rows — bold emphasis in prose, a field-shaped line inside a fenced block,
the same inside an inline code span, and a horizontal rule mid-body all
correctly mint no node. Row 24 is the Generative-Fix-Divergence parity
assertion: PLANNING_ARTIFACTS must not diverge from isCanonicalPlanningFile.
Row 28 covers the non-markdown canonical file found during the build.

Assertions are structural throughout — a typed Result / NodeRead /
SerializeOutcome shape, or a byte range computed from the node's own Span.
Full-string equality appears only where the contract IS byte equality.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4917): apply review findings — refuse unrepresentable values, compose the layers the seam claimed

Three review passes ran against this branch: /security-review, an isolated
adversarial pass, and a standards+spec pass. Five findings, all fixed here. Every
one is the epic's own failure class reproduced inside the seam built to end it,
which is the thing worth noticing.

1. setFieldValue accepted a value containing a newline. It survived serialization
   and reparsed as a REAL sibling field — one write to Plans forged a second
   Owner into a document that already had one. Content became structure, which
   defeats ADR-4910 Decision 2's "cannot reach past its own token by
   construction": the token boundary is a LINE boundary.

2. setFieldValue accepted a value containing the trailing separator and SILENTLY
   TRUNCATED it. Staged "sneaky - annotation", read back "sneaky", with the
   remainder reclassified as trailing prose. No error, nothing unreadable, both
   resulting nodes parsing perfectly. Worse than (1) because it loses the
   caller's own value rather than adding something visible.

   Both are fixed by ONE general check, deliberately not a blacklist:
   setFieldValue rebuilds the candidate line, re-parses it through the same field
   grammar, and refuses unless the value reads back identical. Blacklisting the
   separator would close this instance and leave the class open for whatever
   separator the grammar grows next. The round-trip check is ADR-4910 Decision 4
   stated executably.

3. The module reimplemented two layers it claims to compose. Checklist detection
   hand-rolled a checkbox regex that markdown-sectionizer's iterateBullets
   already owns. Frontmatter span detection re-derived fence handling because
   frontmatter.cts's frontmatterRegion was module-private — so ADR-4910 section 1's
   stated layering was UNREACHABLE as written, and the first implementation
   routed around it silently instead of surfacing the gap.

   frontmatterRegion is now exported (additive only; ADR-2143 section 2's
   extend-never-mutate lock is inherited) and both layers are consumed.

4. The CONTEXT.md glossary entry asserted "the Frontmatter Module supplies
   frontmatter" while zero frontmatter.cts code was invoked. That was a false
   claim in the repo's vocabulary of record, written by me, and it is now true
   rather than edited away.

5. Adopting iterateBullets narrowed GFM coverage: it classifies only
   dash-prefixed task items as checkboxes, so "* [ ] x" and "+ [x] y" stopped
   becoming checklist nodes. Widening the sectionizer is forbidden by the
   inherited lock, so the task-list MARKER is interpreted in this module while
   bullet STRUCTURE still comes from the sectionizer.

The sharpest finding was not a defect. The fast-check generator constrained
values to [A-Za-z0-9 .,!?], so it could not emit an em-dash, newline, backtick,
pipe or asterisk — precisely where (1) and (2) lived. The property was real,
seeded and non-vacuous, and structurally blind to the module's actual bug class.
The generator now spans the grammar's own metacharacters, and a new property
asserts that every value setFieldValue ACCEPTS round-trips identically. Proven
able to fail: against a scratch copy with the guard stripped it fails after 7
cases on newValue "\n".

Recorded as a measured boundary, not fixed: a bare CR inside a field line leaves
that field unrecognised. Measured — sibling fields still parse, no error node,
and serialize stays byte-identical, so the worst case is an unreadable field and
never a damaged document. Flagged for Phase 4's empty-vs-error census.

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4917): backfill changeset pr number to 4918

Refs #4906

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:17:58 -04:00
Tom Boucher
eea9247c93 enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select

Auto-mode used to auto-select a checkpoint:decision's first <option>
unconditionally, making a decision checkpoint's safety depend on option
presentation order rather than an authored choice. Add an optional
auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode
now escalates to a human exactly like gate="blocking-human" does; present,
it names the option auto-mode selects; naming an id with no matching
<option id> is a hard structural-validation error at plan-parse time
rather than a silent fallback to the first option. gate="blocking-human"
continues to win over everything, unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys

An isolated adversarial review of the auto_select work found that both new
attribute regexes used \b as their left anchor, which is a word boundary,
not a "start of attribute name" boundary. A decoy attribute ending in the
same word (e.g. data-id="...") sitting before the real id="..." on the
same <option> tag matched first, silently corrupting the extracted option
id. Anchor on (?:^|\s) instead so only the real attribute name can match.
Adds a regression test reproducing the exact decoy-attribute shape, plus a
Unicode option-id test from the same review pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane

lint-docs-guard-registration failed: the new test reads docs/reference/
plan-md.md but was not registered, so a future edit to that doc could
silently desync from the test without the guard catching it on the PR
that changed the doc. Registered alongside its direct precedents
(precondition-element.test.cjs, reversibility-tagging.test.cjs), which
read the same file for the same reason.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling

The remote gsd-test run caught what local checks missed: execute-phase.md
carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes
of headroom before this change, and the original checkpoint:decision
wording pushed it to 93772 (over the ceiling). Cascaded into failures in
phase6-capstone-conformance, execute-phase-completion-reconciliation,
claude-orchestration, and the compact-content drift-report test.

Also caught: tests/package-legitimacy-gate.test.cjs anchors a
"decision is conditional, not unconditional" safety check on the literal
phrase "first option" in the decision bullet — which #4095 deliberately
removes, since there is no more unconditional first-option pick. The test
was asserting an assumption this change intentionally makes obsolete;
re-anchored on tokens that still identify the bullet ('decision',
'auto-spawn') without weakening what the test actually verifies (the
bullet must still carry a blocking-human carve-out).

Also fixed a word-order mismatch between my own new test's regex and the
actual doc text it was asserting against (tests/auto-select-attribute.test.cjs).

Regenerated the compact-content benchmark baseline
(tests/fixtures/compact-content-benchmark-baseline.json) to match the new
byte counts.

Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap.
Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4095): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:47:01 -04:00
Tom Boucher
ccfed63355 fix(#4823): the Current Plan reset is scoped to the Current Position section (#4898)
* test(#4823): failing-first — the Current Plan reset must not rewrite prose outside Current Position

* fix(#4823): the Current Plan reset is scoped to the Current Position section — the whole-body 'Plan' fallback matched hard-wrapped prose lines starting with plan:

* chore(#4823): changeset fragment

* chore(#4823): backfill changeset PR number (4898)

---------

Co-authored-by: sim <sim@local>
2026-09-20 11:45:50 -04:00
Tom Boucher
bd79a97df0 fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error, never status missing (#4896)
* test(#4806): failing-first — unparseable VERIFICATION.md frontmatter is a parse error, not status missing / Field not found

* fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error — never 'missing' or 'Field not found'

* test(#4806): census pin 66→67 — cmdFrontmatterGet's unparseable-frontmatter error is a new output({error}) call site

* chore(#4806): backfill changeset PR number (4896)

---------

Co-authored-by: sim <sim@local>
2026-09-20 10:26:53 -04:00
Tom Boucher
e2d879f681 fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse (#4895)
* test(#4802): failing-first — acknowledge must refuse an unparseable-frontmatter target instead of splicing over it

* fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse — never splices the marker-marked object over the file

* chore(#4802): backfill changeset PR number (4895)

---------

Co-authored-by: sim <sim@local>
2026-09-20 08:45:54 -04:00
Tom Boucher
a8394713d1 fix(#4801): init.manager resolves archived phase directories through findPhaseInternal (#4893)
* test(#4801): failing-first — an archived phase directory must resolve and report complete in init.manager

* fix(#4801): init.manager resolves archived phase directories through findPhaseInternal

The private current-milestone-only scan (matchPhaseDirs over listMilestonePhaseDirs)
counted archived phase directories as missing, so an archived phase with a passed
verification reported no_directory/phase_complete:false. findPhaseInternal — already
imported, already the shared primitive for five other init commands — searches the live
directory first and falls back through the workstream-scoped archive (#2855); the private
copy and its single-consumer entries list are retired.

* chore(#4801): backfill changeset PR number (4893)

---------

Co-authored-by: sim <sim@local>
2026-09-20 06:45:38 -04:00
Behruz Nassre Esfahani
2e14b4df17 fix(#4415): treat an absent worktree as removed, not as a branch mismatch (#4612)
* fix(#4415): treat an absent worktree as removed, not as a branch mismatch

Claude Code removes a subagent's worktree the moment the subagent finishes with
a clean tree. A gsd-executor that committed everything — SUMMARY.md included,
under `commit_docs: true` — is exactly that case, so by the time the
orchestrator reaches wave cleanup the directory is routinely gone while the
branch it left behind is intact and mergeable.

`git -C <gone> rev-parse --abbrev-ref HEAD` fails, and nothing distinguished
that filesystem failure from a real branch disagreement: both reached the same
`if`, so the entry blocked `branch_mismatch`, NOTHING merged, and the branch was
left dangling. When the directory instead vanished after the merge landed,
`git worktree remove` failed "is not a working tree" and the entry blocked
`worktree_remove_failed`, leaving the branch undeleted and the operator to run
`git worktree prune` + `git branch -D` + `rm -rf` by hand every wave.

Disambiguated at the point of failure rather than ahead of it. A SUCCESSFUL
in-worktree read still decides identity exactly as before — a present worktree
on the wrong branch blocks, unchanged — and only a FAILED read consults the
filesystem. Two reads can fail, and they are not the same path:

  * The branch read fails with the directory absent. There is no checkout for
    identity to come from, so it falls back to `refs/heads/<branch>` read from
    repoRoot; a missing ref still blocks, so an absent worktree never becomes a
    silent pass. The SUMMARY rescue and the dirty check are then skipped.

  * The branch read succeeded and the later `status` read fails with the
    directory now absent — the harness removed it while the repoRoot-side base,
    deletion and scope checks ran. Identity was already established from the
    checkout and the rescue has already run; only the dirty decision is skipped.
    Without this, a mid-entry removal still blocked `worktree_dirty` with
    nothing merged: the same bug, one window later.

Skipping those reads is not a claim that the worktree was clean. This code
cannot tell who removed the directory, and a forced or manual `rm -rf` of a
DIRTY worktree would already have destroyed an uncommitted SUMMARY before
cleanup ran. The narrow thing that is true either way is that a missing source
cannot be read. The two reads also fail differently: the default SUMMARY finder
catches the unreadable directory and returns no files, while `git -C <gone>
status` errors — and that error is what surfaced as `worktree_dirty`. A rescue
that genuinely FAILS still blocks, since a copy that errored part-way can mean
an uncommitted SUMMARY was really lost.

Teardown prunes the stale .git/worktrees admin entry rather than removing a path
that is not there, re-reading presence instead of reusing the branch-step answer
since the harness can act in between. For an entry accepted as ABSENT it prunes
ONLY and never issues `worktree remove --force`: that entry was merged without
the rescue and dirty checks, so force-removing a checkout recreated at that path
would delete contents that never passed either one — strictly worse than the bug
being fixed. A genuine prune failure still reports `worktree_remove_failed`, and
a blocked teardown still withholds the branch delete. `git worktree prune` is
repository-wide maintenance, not an entry-scoped operation.

The presence probe resolves `worktree_path` against repoRoot, the way git does.
`normalizeCleanupManifestEntry` takes the path from the manifest verbatim, so it
can be relative, and every git call passes it as `-C <path>` with
`cwd: plan.repoRoot`; a bare `fs.existsSync` would have resolved it against the
PROCESS working directory instead. Those differ whenever cleanup runs from
elsewhere, reachable today through gsd-tools' `--cwd` override, and the mismatch
reads both ways: a present checkout reported absent — skipping the dirty check
that would have blocked it — or an absent one reported present.

An earlier cut resolved presence UP FRONT, before the branch read. That broke 52
existing tests: every cleanup-wave test uses a fake path that does not exist on
disk and injects no `existsSync`, so all of them re-routed down the absent
branch. Disambiguating at the point of failure leaves those tests reading as
they did. Three rows still needed their premise stated — each stubs a git
failure against a worktree that is genuinely present — and now inject
`existsSync: () => true`. No assertion in any of the three changed.

Fourteen rows added. Every early row held presence CONSTANT and so could not
reach the windows that matter, since the bug is caused by a directory that
changes state WHILE cleanup runs: removal after the branch read, a present
worktree whose status fails (which must still block), removal between the clean
status read and teardown, a reappeared checkout at teardown, #2852 isolation of
a blocked absent entry from the entries after it, and relative-path resolution.

Verified: ran the issue's own reproduction verbatim against a build of this
branch — `merged_removed`, merge commit present, branch deleted, no prunable
entry in `git worktree list`. The same reproduction against a build at the
merge-base returns blocked/branch_mismatch, no merge, branch present, `wt1 ...
prunable`. Five of the first eight rows go red against the true merge-base file;
the three that stay green are the safety-preservation rows. The rows added after
each review round go red against the commit that round reviewed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X

* chore(#4415): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X

* fix(#4415): build the probe-path expectation with path.resolve, not path.join

The row asserting that the presence probe resolves a relative `worktree_path`
against repoRoot failed on windows-latest while the code under test was correct.
On win32 `path.resolve` prepends the current drive to a drive-less absolute path
(`/repo/main` -> `D:\repo\main`) and `path.join` does not, so a join-built
expectation disagrees with correct behavior:

    expected: '\repo\main\.claude\worktrees\agent-a1'
    actual:   'D:\repo\main\.claude\worktrees\agent-a1'

`path.resolve` is what the fix must use — it is how git resolves `-C <path>`
against `cwd: plan.repoRoot` — so the expectation moves to resolve as well. Two
`notEqual` rows keep that from being circular: the probe must receive neither the
raw relative path nor a process-cwd resolution. Verified by mutation — dropping
the repoRoot anchoring in `worktreeExists` turns the row red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X

* fix(#4415): confirm absence before skipping the rescue and dirty checks

`fs.existsSync` answers false for a genuinely missing path AND for one it
merely cannot traverse — EACCES on a parent directory, an unreachable mount.
Verified: with a parent at mode 000, `existsSync` returns false while
`statSync` throws EACCES.

That distinction carries weight here, because "absent" is what lets an entry
skip the SUMMARY rescue and the dirty check. An unreadable-but-present
worktree read as absent, so cleanup merged over uncommitted work that the
dirty check exists to refuse — and it contradicted this code's own comment
that a present checkout whose git read fails stays blocked. Before this PR a
failed git read blocked unconditionally, so treating unreadable as present is
not a new safety rule; it is the one that was already there.

The default probe becomes `statSync`, which reports WHY it failed. Only
ENOENT is absence; anything else reads as present and blocks. An injected
probe stays authoritative, so tests state presence directly with no hidden
dependency on the real filesystem, and may throw to state that a path is
unreadable.

Two rows added: an unreadable worktree still blocks as branch_mismatch with
no merge and no teardown, and a confirmed-ENOENT probe still takes the absent
path. Verified by mutation — reverting the discrimination to the permissive
`return false` turns the unreadable row RED while the ENOENT row stays green,
which is what distinguishes discrimination from over-blocking. The mutation
was confirmed to reach the compiled artifact the test loads.

Also from this round: the row named for a checkout that "reappeared" never
modeled reappearance (production probes presence once, at identification), so
it is renamed to the unconditional contract it does prove; the comment
crediting the notEqual rows with removing circularity is narrowed to what
they actually establish; and the changeset now says only a confirmed absence
takes the new path.

Found by Codex full-PR review (round 3) before pushing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* fix(#4415): source identity from git's registration, removal from the errno

Maintainer review rejected the premise this fix rested on. It held that once the
worktree directory is gone there is no checkout to read, so identity must fall
back to `refs/heads/<branch>`. Git does not lose the binding — measured, after
`rm -rf`:

    worktree /path/to/wt
    branch refs/heads/feat-x
    prunable gitdir file points to non-existent location

The ref fallback weakened identity from "the checkout registered at this path is
on this branch" to "a branch by this name exists", which let a foreign sibling
branch merge. Identity now comes from `git worktree list --porcelain`, so the
#3677 swap control keeps its teeth on the absent path; the new swap row is what
would have caught this, and dropping the branch conjunct turns only that row red.

Two defects in the first cut of the porcelain rework, both measured rather than
reasoned about:

`prunable` is not a removal test. With a parent directory at mode 000, git prints
`prunable gitdir file points to non-existent location` for a checkout that is
STILL THERE — it cannot traverse the parent, so it reports the gitdir file as
missing. Treating prunable as "removed" would skip the rescue and dirty checks
and merge over uncommitted work in an unreadable worktree, reintroducing the
review's Major finding by another route. Each source now answers only what it can
prove: porcelain for identity, `statSync`'s errno for removal. Only ENOENT is
removal; EACCES/EIO blocks, as it did before this PR.

`git worktree prune` is repository-wide. Measured: two removed worktrees plus ONE
prune leaves neither registration behind. Reading the list per entry therefore let
the first absent entry's teardown erase the identity evidence of every entry after
it, merging one worktree per wave and blocking the rest as branch_mismatch —
worse than the bug being fixed, since a wave of parallel executors is the normal
case. The identity read is now a snapshot, captured lazily on the first entry that
needs it and reused for the wave, which is both pre-prune and off the happy path.

The `existsSync` probe and its dep locals are deleted; the filesystem is consulted
only for the errno. The comment calling repository-wide prune "Harmless" was wrong
under the new identity rule and says so now.

Tests: identity and removal are stated on their own axes rather than through one
present/absent boolean. Added the absent-path #3677 swap row, the two-absent-entry
prune row, a bare `prunable` marker row, and a fail-safe row for an unreadable
worktree list. Three mutations each kill exactly the intended rows, verified
against the compiled artifact the tests load. One fixture that still stated
presence through the removed `existsSync` seam was passing for the wrong reason
and now states both axes.

Verified: lint:ci exit 0; full suite 24/24 chunks, 37,164 tests, 0 failures;
tests/worktree-safety.test.cjs 422/422.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* fix(#4415): re-confirm absence before teardown, and prove the porcelain claim against real git

Maintainer review, Major. Presence was classified once, at identification, and
everything between that point and teardown — the base, deletion and scope gates,
and the merge itself — is a window in which a worktree can reappear. The defence
was "prune only, and a live checkout would make `branch -D` fail visibly", which
holds only while prune's own staleness check is not fooled by the same
filesystem-visibility gap that produced the false absence one call earlier. If it
is, prune clears the admin entry, `branch -D` then SUCCEEDS, and a live,
unreviewed, un-rescued worktree loses its branch.

That asymmetry is the argument for the fix: the bug this PR set out to repair only
ever BLOCKED, while this path could DESTROY state. Absence is now re-confirmed
with `confirmedGone()` immediately before teardown — no new subprocess, just the
statSync already in hand — and a reappeared directory blocks as
`worktree_remove_failed` instead of reaching prune or the branch delete.

The review was also right that the gap was known and unverified: the existing row
said so in its own comment ("it does NOT model the reappearance transition
itself"). It is modelled now, by a stat that answers "gone" at identification and
"present" at teardown. Mutation-verified: removing the re-confirmation turns ONLY
the new row red while the old "prune, never force-remove" row stays green, which
is exactly why that row could not have caught this.

Minor, same review: the #4415 block was entirely mock-based, so the factual claim
the identity mechanism rests on was asserted in comments and measured out of band
but never proved executably. Two real-git rows now prove it — that git keeps the
path -> branch binding after the checkout is deleted and marks the entry prunable,
and that it ALSO reports prunable for an unreadable worktree that is still there,
which is why removal is confirmed by errno rather than by prunable. The second row
skips as root, where mode 000 does not deny traversal.

Minor 2 (rescueSummaryArtifacts resolving worktree_path against process.cwd()
while the new code resolves against plan.repoRoot) is pre-existing and not
reachable through the CLI's same-cwd invocation; left for a follow-up issue rather
than widened into this PR.

Verified: lint:ci exit 0; full suite 27/27 chunks, 37,739 tests, 0 failures, against
the true merge-base; tests/worktree-safety.test.cjs 425/425.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* test(#4415): make the real-git rows platform-correct

The Windows conformance shard caught both rows on their first push, and both
failures were mine, not the code's.

Path separators: git reports porcelain paths with FORWARD slashes on every
platform, while `path.join` yields backslashes on win32, so `includes()` compared
separator styles rather than paths and the registration assertions failed. Both
sides are normalised before comparison now.

Premise setup: the unreadable-worktree row establishes "git cannot traverse the
parent" with mode 000, which win32 does not honour for directory traversal at all
— the row would have asserted `prunable` against a perfectly readable worktree and
failed for a reason unrelated to the behaviour under test. It now skips on win32
for the same reason it already skipped as root, with both reasons stated together.

Verified: lint:ci exit 0; tests/worktree-safety.test.cjs 425/425 locally. The
Windows shard is the real check for the separator fix, since macOS cannot
reproduce it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* fix(#4415): warn when an entry is accepted as absent, giving prunable its consumer

Maintainer review round 3, both Medium findings — they close together, as the
review noted.

The absent path reported `merged_removed`/`ok` indistinguishably from an ordinary
merge. This code cannot tell "the harness cleanly removed a finished executor"
from "an operator or an external process removed this path": git keeps the
path -> branch registration and `statSync` reports ENOENT in both cases. Before
this path existed every anomalous absence blocked loudly, so accepting the routine
case silently took the operator's only signal away from the case that is not
routine. The module already carries an advisory channel for a materially less
risky condition — scope conformance, a few lines below — so withholding one here
was inconsistent with its own pattern.

`WAVE_CLEANUP_WARNING.ACCEPTED_ABSENT_WORKTREE` is now emitted at both acceptance
sites, carrying git's own `prunable` reason. Advisory, never a gate: the entry
still merges.

That also gives `WorktreeEntry.prunable` a consumer. It was parsed, documented as
"worth surfacing to an operator", and then never read — the errno rework made it
unused for the predicate and the parsing stayed behind. Quoting git's reason here
is what it was for.

The bare-marker test was vacuous, as the review said: it asserted
`merged_removed`, which is driven by `confirmedGone` and the branch match, not by
the bare-marker parsing it claimed to cover, so a regression in that parsing would
not have reddened it. It now asserts the parsed value reaches the warning. A bare
`prunable` line normalises to the literal 'prunable' — a truthiness signal, not a
reason — so the warning reports null there rather than quoting a marker back at an
operator as though git had said something.

`WAVE_CLEANUP_WARNING`'s locked code set is updated deliberately, with the reason
recorded in the test: the lock exists so a new advisory code is a decision rather
than something that appears because a branch needed one.

Verified: mutation — suppressing the warning at both sites turns both new rows
red; lint:ci exit 0; full suite 27/27 chunks, 38,245 tests, 0 failures;
tests/worktree-safety.test.cjs 426/426.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-20 04:00:45 -04:00
Tom Boucher
822934c901 fix(#4794): decision-coverage answers an unmeasured shape on could-not-parse; a non-file context path fails closed (#4889)
* test(#4794): failing-first — could-not-parse must answer an unmeasured shape; a directory context path fails closed

* fix(#4794): could-not-parse answers an unmeasured shape (null counts, unreadable ids, no uncovered); a non-file context path fails closed

* chore(#4794): backfill changeset PR number (4889)

* test(#4794): skip the directory-identity probe when the platform cannot discriminate (windows runner volume collapse, measured)

Two consecutive windows conformance runs failed the probe with measured identical
(dev, ino) for two distinct mkdtemp directories (dev=3606225537, ino=9007199255243448
for both) — a runner-volume property, not a regression in the guard. On such a
platform the guard's identity containment degrades to refuse-everything (fail-closed,
documented); the probe asserts capability, so the honest response is an explicit
t.skip carrying the measurement (ADR-2719 §6), not a red lane for every PR.

---------

Co-authored-by: sim <sim@local>
2026-09-20 02:53:58 -04:00
0xdhx
8a5166598c fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next (#4873)
* fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next

Commit 740ba0d8a (#4781) removed every change #4768 had merged for #4748:
the first-non-digit split at execute-phase.md's two arithmetic sites, the
init-emitted `padded_phase` the REVIEW.md lookup binds instead of
`printf "%02d"`, the canonical-grammar extractions in autonomous.md and
plan-review-convergence.md, the `.changeset/zesty-wolves-tumble.md`
fragment, and the three lint-phase-id-drift ratchets with their tests.
The guard and the code it guarded left together, so nothing went red.

This is a cherry-pick of 092d9256b onto current `next`, resolved against
the #4683 threat-id fields on execute-phase.md's Parse-JSON line, with the
changeset `pr:` reset to the placeholder and the compact-content benchmark
baseline regenerated against the current base.

(cherry picked from commit 092d9256b8)

Emitted-Drift-Ack-Growth: autonomous.md — restores #4768's canonical-grammar extraction and its explanatory comment for --from/--to/--only
Emitted-Drift-Ack-Growth: execute-phase.md — restores #4768's first-non-digit split at two arithmetic sites and the padded_phase binding for the REVIEW.md lookup
Emitted-Drift-Ack-Growth: plan-review-convergence.md — restores #4768's canonical-grammar phase extraction and its comment
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AX7LXxc3uAkGki6iaYiAMP

* chore(#4830): set changeset fragment pr to 4873

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-20 00:01:21 -04:00
Tom Boucher
09e1110e68 fix(#4788): code spans in the decision bold lead-in are opaque to the separator grammar (#4883)
* test(#4788): failing-first — code spans in the decision bold lead-in are opaque to the separator grammar

* fix(#4788): the decision lead-in runs are code-span-aware — a backticked span is data, never grammar

* chore(#4788): backfill changeset PR number (4881)

* chore(#4788): correct changeset PR number (4883, was a guessed 4881)

---------

Co-authored-by: sim <sim@local>
2026-09-19 22:46:48 -04:00
Tom Boucher
5906a24ede fix(#4786): plan-row detection accepts the bare planId stem — suffix-less hand-written lists tick in place (#4880)
* test(#4786): failing-first — a suffix-less hand-written plan list is ticked in place, never duplicated

* fix(#4786): plan-row detection accepts the bare planId stem — a suffix-less hand-written list is ticked in place, never duplicated

* chore(#4786): backfill changeset PR number (4880)

---------

Co-authored-by: sim <sim@local>
2026-09-19 21:06:05 -04:00
Tom Boucher
d435723c95 fix(#4782): claude's agents kind skips compact variants — consumed only by the non-claude persona gate (#4878)
* test(#4782): failing-first — claude install must not stage compact agent variants (and must prune stale ones)

* fix(#4782): claude's agents kind skips compact variants — they are consumed only by the non-claude persona-fallback gate

Emitted-Drift-Ack-Hash: agents/gsd-advisor-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-ai-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-assumptions-analyzer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-code-fixer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-code-reviewer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-codebase-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-debug-session-manager.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-doc-classifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-doc-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-doc-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-doc-writer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-dom-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-domain-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-eval-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-eval-planner.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-framework-selector.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-integration-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-intel-updater.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-mempalace-curator.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-nyquist-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-pattern-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-project-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-research-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-roadmapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-security-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-ui-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-ui-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-ui-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies
Emitted-Drift-Ack-Hash: agents/gsd-user-profiler.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies

* chore(#4782): changeset fragment

* chore(#4782): backfill changeset PR number (4878)

---------

Co-authored-by: sim <sim@local>
2026-09-19 17:25:21 -04:00
Tom Boucher
4fe2837ab5 fix(#4774): plan-criteria R4 requires the pipe not to be doubled — a logical-OR fallback is handled, not swallowed (#4877)
* test(#4774): failing-first — R4 must not read a logical-OR fallback as a pipeline stage

* fix(#4774): R4 requires the pipe not to be doubled — a logical-OR fallback is a handled failure, not a swallowed one

Also corrects the rows-3/4 test's makeCriteriaPlan usage (second arg is the
<verify> block, not a second criteria line).

* chore(#4774): backfill changeset PR number (4877)

---------

Co-authored-by: sim <sim@local>
2026-09-19 15:38:16 -04:00
Tom Boucher
a87b83d485 fix(#4764): dep_phases extracts only Phase-prefixed references from Depends-on prose (#4876)
* test(#4764): failing-first — dep_phases must extract only Phase-prefixed references, never dates/shas/ledger ids/self

* fix(#4764): dep_phases anchors phase references to their 'Phase' prose context and never emits the row's own number

* fix(#4764): review fold-ins — hoist the anchored dep-reference grammar to phase-id, cover Oxford lists and hyphen ranges, repair the property test

Adversarial review found: Oxford-comma lists under-extracted ('Phases 1, 2,
and 3' dropped the tail member — a silent real-blocker clear, the dangerous
direction); hyphen ranges ('Phases 1-3') kept only the first endpoint; the
property test called fc.hexaString (absent in fast-check 4.8, threw every
run) and passed the junk arbitrary unspread (vacuous guard) with no
completeness assertion; planning-inspect's extractDependencyTokens carried
the same whole-field scrape (generative-fix divergence). The anchored
grammar now lives beside PHASE_NUMBER_TOKEN_SOURCE in phase-id.cts and both
readers interpolate it.

* chore(#4764): backfill changeset PR number (4876)

---------

Co-authored-by: sim <sim@local>
2026-09-19 13:07:03 -04:00
Tom Boucher
d36514b816 fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() (#4872)
* test(#4758): failing-first — rescue must resolve a relative worktree_path against repoRoot

* fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd()

* test(#4758): review fold-ins — t.after cleanup pattern, post-resolution reader contract comment

* chore(#4758): changeset fragment

* chore(#4758): backfill changeset PR number (4872)

* test(#4758): windows lanes key rescue fakes on resolved path identity, not verbatim strings

win32 path.resolve rewrites driveless-absolute POSIX-style fixture values to the
current drive, so the rescue's (correct) resolved-path handoff stopped matching
verbatim string keys: #3804/#245/#2556/B7/#2852 fakes silently skipped the
rescue and my seam test compared against a POSIX literal. Fakes now key on
path.resolve(repoRoot, …) identity — the same semantics the code and git -C
use — so every rescue test exercises the rescue on every platform.

---------

Co-authored-by: sim <sim@local>
2026-09-19 09:41:51 -04:00
Tom Boucher
63edc777e6 fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading (#4868)
* test(#4588): observed fork-from-HEAD must suppress the stale-origin degrade (failing first)

* fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading

* fix(#4588): a throwing probe git call is an inconclusive observation, not a crash

* test(#4588): name the observed reason in the inconclusive-row assertions

* test(#4588): hermetic state I/O in the inconclusive-row fixtures

* chore(#4588): changeset fragment

* test(#4588): contained cache I/O, emit-payload assertions (review fold-ins)

* fix(#4588): review fold-ins — invalidation fixture, hermetic confirm row, gate comment, cache+scope rows

* chore(#4588): backfill changeset PR number (4868)

---------

Co-authored-by: sim <sim@local>
2026-09-18 23:41:44 -04:00
Tom Boucher
9a41a95212 fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams

* fix(#4717): consult the per-install runtime marker at both identity seams

resolveReportedRuntime (agent_runtime) and loadConfigResolved
(config.runtime) both ignored the per-install .gsd-runtime marker that
resolveRuntime and the model-resolver gate already read. On a
multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing
misreported every Claude Code session as codex, and a shared
defaults.json stamped by the first non-Claude install leaked its runtime
to every other one.

Seam 1: the reported-runtime ladder becomes explicit > install marker >
host detection > claude. Seam 2: loadConfigResolved fills an empty
config.runtime from GSD_RUNTIME then the marker, copy-on-write (the
builtin-defaults branch returns a shared object). Explicit runtimes and
marker-less trees are unchanged.

* fix(#4717): a marker-detected runtime opts into its tier map (decision a)

* fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins

* chore(#4717): backfill changeset PR number (4861)

---------

Co-authored-by: sim <sim@local>
2026-09-18 13:50:01 -04:00
Tom Boucher
e1f72cd324 fix(#4741): the plan checkbox tick respects the superseded exclusion (#4851)
* test(#4741): a superseded plan must not be ticked from its summary (failing first)

* fix(#4741): the plan checkbox tick respects the superseded-plan exclusion

* fix(#4741): review fold-ins — changeset typo, dedupe planId derivation

* test(#4741): exercise a halted summary on the active plan in the #2830 pin

* chore(#4741): backfill changeset PR number (4851)

---------

Co-authored-by: sim <sim@local>
2026-09-18 06:33:30 -04:00
Tom Boucher
c5629bbe74 fix(#4734): degrade worktree isolation when the root has no git repository (#4843)
* test(#4734): non-git root must degrade worktree isolation (failing first)

* fix(#4734): degrade worktree isolation when the root has no git repository

* fix(#4734): review fold-ins — 3972 ladder fixture, parity fixture, docs row, message wording

* chore(#4734): backfill changeset PR number (4843)

---------

Co-authored-by: sim <sim@local>
2026-09-18 03:16:25 -04:00
Tom Boucher
8d0b6868ae fix(#4725): write normalization preserves tight paragraph-list shape (#4842)
* test(#4725): write normalization must not reflow untouched prose (failing first)

* fix(#4725): stop write normalization injecting a blank before a list after prose

* test(#4725): repair ordered-list fixture and list-spacing snapshot

* test(#4725): assert whole-file prose stability, fix heading-list comment

* chore(#4725): backfill changeset PR number (4842)

---------

Co-authored-by: sim <sim@local>
2026-09-18 01:38:14 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
bff99a8bb5 fix(#4731): read hard-wrapped Goal/Requirements fields past the line break (#4826)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; isolated adversarial review round completed (MEDIUM table-bleed finding fixed with RED/GREEN evidence) and sha-pinned bench 46331/0 on the merged head.
2026-09-17 15:59:38 -04:00
Tom Boucher
fb3e228a0d fix(#4724): classify Surefire/Failsafe XML as RED evidence (#4825)
* test(#4724): add failing-first coverage for Surefire XML RED evidence

* fix(#4724): classify Surefire/Failsafe XML as RED evidence

check tdd-red-evidence parsed only node:test TAP, so a JVM project's
genuine Maven red scored INVALID_RED while hand-written synthetic TAP
scored RED_EVIDENCE_OK — the gate was passable only by fabricating its
input (issue #4724's measured repro).

classifyRedEvidence detects Surefire/Failsafe XML (a <testsuite> element)
and parses it by TAG-BOUNDARY scanning: each <testcase> owns its own tag
(self-closing) or the segment up to its </testcase> closer, so the
issue's warned-about spanning trap (a lazy lazy match from a green
self-closing case to the next closing tag) cannot misreport names. A
<failure> or <error> child marks the case failing; the target matches at
class granularity (exact classname, dotted-suffix, or method name). Any
parse anomaly degrades to not-failing — the module stays fail-closed and
PURE (no fs/clock; report freshness remains the workflow's run-start
check per the issue's implementation notes).

TAP classification is byte-identical: all existing fixtures stay green.

* test(#4724): pin the scanner hardening — truncation, TAP-message flip, CDATA phantom

* docs(#4724): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 11:30:23 -04:00
Tom Boucher
7d0c6339d0 fix(#4705): emit Antigravity-native tool names as a YAML sequence (#4822)
* test(#4705): add failing-first coverage for native Antigravity tool sequences

* fix(#4705): emit Antigravity-native tool names as a YAML sequence

convertClaudeAgentToAntigravityAgent and the installer's twin emitted
Gemini CLI tool names as a comma-separated scalar. Antigravity's
documented subagent contract (antigravity.google/docs/subagents) wants a
YAML sequence of native names — view_file, grep_search, run_command,
replace_file_content are the documented examples, and wrong or malformed
grants can hang the subagent per Antigravity's own warning.

Map values move to the native vocabulary where documented (Read ->
view_file, Edit -> replace_file_content, Bash -> run_command, Grep ->
grep_search); undocumented entries keep their best-known grant rather
than being dropped (dropping would silently remove a restriction). The
emitter writes one '- name' item per line; an agent whose every tool was
filtered emits an explicit tools: [] instead of an empty scalar.
Pre-existing pins updated to the native vocabulary.

* test(#4705): update the #4727 map-value pin to the Antigravity-native vocabulary

The #4727-era pin held the map VALUES at the Gemini CLI dialect on the
belief that Antigravity speaks it; the confirmed bug #4705 (with
Antigravity's own documented subagent contract) supersedes that for the
four documented names. Key/shape pinning is preserved; only the values
move.

* docs(#4705): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 08:12:18 -04:00
Tom Boucher
d707318e0c fix(#4699): skip already-complete phases in the next_phase cascade (#4820)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase

* fix(#4699): skip already-complete phases in the next_phase cascade

Both next-phase scans selected the numerically lowest phase above N
without consulting completion state, so completing a reopened phase
persisted an already-[x] phase as STATE.md current_phase while
roadmap.analyze correctly named the outstanding one (issue repro:
completing 2 with phases 1 and 3 already [x] returned next_phase 03).

The cascade collects the complete phase numbers from the roadmap
checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in
both the disk scan and the roadmap scan; a [x] checkbox row and its
heading sibling both name a phase that is never next. Heading-only and
checkbox-less roadmaps behave exactly as before.

* test(#4699): align the negative-control expectation with the disk spelling

* test(#4699): pin the STATE.md persistence and the all-later-complete tail corner

Review findings: the regression never asserted STATE.md current_phase
(the issue's actual harm), and the all-later-phases-[x] corner
(is_last_phase true, next_phase null) was unpinned. A changeset fragment
is included.

* docs(#4699): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-17 04:17:39 -04:00
Tom Boucher
2bfff17ff8 fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing

* fix(#4682): route stale verification to the verifier regeneration path

The stale routing entry sent users to /gsd-verify-work — but verify-work
never rewrites VERIFICATION.md (its only write is the human_needed
canonicalization), so following the advice re-ran UAT, reached the same
stale check, and looped. init's projector and execute-phase's generic
next_command presentation both mirror this entry, so the dead end appeared
on three surfaces.

The stale entry now routes to execute-phase, and execute-phase's
all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping
the spine under its frozen ADR-857 ceiling) mirroring the missing route:
skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at
aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier —
regenerating VERIFICATION.md and its digest, marked phase or not. The
non-stale fall-through. Staleness detection, the digest format (#4623),
every other routing entry, and the #3684 resume arms are untouched.

Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682)
Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682)

* test(#4682): register the stale-reverification part and align projected commands

The new steps/ part must be registered in the inventory manifest and the
per-runtime golden install trees (regen:derived); the projected stale
next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the
runtime surface), the human_needed bare-report probe keeps routing to
verify-work (unchanged semantics), and init-manager's recommended action
follows the new command.

* test(#4682): prefix the remaining stale routing assertions with the runtime surface

Nine stale next_command assertions and the human_needed bare-report probe
still carried the unprefixed or flipped forms from the earlier line-number
edit; all now assert the shipped /gsd-execute-phase <phase> projection,
with the human_needed probe reverted to its unchanged verify-work routing.

* test(#4682): align the last stale projection assertions with the execute-phase route

* docs(#4682): backfill changeset PR number

* test(#4682): refresh the compact-content baseline after the rebase

The rebase onto the #4670 squash brought verify-work.md's bounded
reconciliation text into this branch; the committed compact-content
baseline now reflects the post-rebase split sizes. Local --check is
clean; the previous bench drift (+243) was the baseline, not the diff.

* fix(#4682): carry the response_language directive in the stale-reverification part

The new steps/ part is its own coverage unit for lint-response-language-coverage;
it takes the shared canonical directive line like its sibling execute-phase
parts.

---------

Co-authored-by: sim <sim@local>
2026-09-17 02:23:03 -04:00
Tom Boucher
85545a77a5 fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate

verify-work.md's complete_session step flips VERIFICATION.md to passed on
'zero issues' alone, so a session whose every UAT row is blocked (a session
that observed nothing) canonicalizes the report. Pins the deployed contract
the fix must satisfy: the flip runs the unflagged phase uat-passed predicate
inside the human_needed branch, frontmatter.set sits inside a passed==true
guard, a refusal message carries the blocker count and keeps
human_needed, and an indeterminate pre-check fails closed. All four new
assertions are RED until the workflow grows the guard.

* fix(#4663): gate the canonicalization on the uat-passed predicate

complete_session flipped VERIFICATION.md to passed whenever the session
recorded zero issues and the status was human_needed — but blocked rows are
not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked
session (one that observed nothing) rewrote the canonical report to passed.
Every later reader (transition.md's preliminary check, resume paths,
validate-phase, verification.status) then inherited the unearned pass while
the phase-close predicate correctly refused it.

The flip now runs the phase-close predicate in a new --uat-only form before
canonicalizing: UAT rows evaluated (at least one pass, no
pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers
skipped — they must be, because the report still reads human_needed at
pre-check time and that status is itself a blocking verification entry, so
the full predicate could never pass there and the flip would deadlock
(found by isolated review, probed). The --require-verification call stays
the transition gate; the refusal branch reports the blocker count and keeps
human_needed; an indeterminate pre-check fails closed.

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* test(#4663): align the canonicalize pre-check needles with the shipped line

The workflow line carries a 2>/dev/null redirect the needles did not
include, so both pre-check assertions fail against the committed fix
(fixed-string grep verified). Reviewer-found; needle and message aligned.

* fix(#4663): reword the canonicalize prose and refresh its size baseline

The rationale paragraph mentioned the flagged transition-gate call by its
flag, putting a --require-verification literal before the first
phase uat-passed occurrence and breaking the existing ordering pin; the
prose now describes it without the literal. verify-work.md's growth also
drifted the committed compact-content baseline; regenerated via
benchmark-compact-content.cjs --write (derived artifact, report-not-gate
contract).

Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663)

* docs(#4663): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 17:26:24 -04:00
Tom Boucher
c72fb34e9a fix(#4658): give the ui plan gate's evidence check a native branch (#4807)
* test(#4658): add failing-first coverage for native frontend evidence

hasStaticFrontendEvidence recognised only JS-ecosystem evidence, so
computeUiPlanGate could never block for a SwiftUI/Compose/Flutter/XAML
project. Adds evidence-level fixtures for the four suggested markers
(import-matched for .swift/.kt/.dart, extension-alone for .xaml), the
reporter's non-UI Swift control case, marker-exactness and SKIP_DIRS and
I/O-degrade negatives, gate-level block assertions through makeProject's
new native frontendEvidence modes, and a pinned-seed fast-check property.
All new assertions are RED until src/ui-frontend-evidence.cts grows the
native branch.

* fix(#4658): give the ui plan gate's evidence check a native branch

hasStaticFrontendEvidence recognised only JS-ecosystem evidence (a root
package.json UI-framework dep, or a .tsx/.jsx/.vue/.svelte file), so
computeUiPlanGate could never block for a SwiftUI, Jetpack Compose, Flutter,
or .NET MAUI project — the #3312 gate was structurally unreachable for them.

Adds a native BFS over the same bounds and skip rules: .xaml is evidence by
extension alone (the .tsx analogue), while .swift/.kt/.dart count only when
their content carries the ecosystem's UI import marker (import SwiftUI /
import UIKit, androidx.compose, package:flutter) — matched on the import, not
the extension, so a non-UI Swift package stays silent exactly as the issue's
37-file control case requires. Marker reads are bounded to a 64 KiB prefix;
any I/O failure degrades to false per the module contract. The #3718
vocabulary filter, the JS evidence rules, and the weaker-extension exclusion
are untouched.

* chore(#4658): regenerate the macos conformance tier list

The native-evidence additions to tests/check-ui-plan-gate.test.cjs move the
file into the macOS conformance tier per the classifier; the committed
generated list is a derived artifact and must match the live tests/ tree
(the same sync the fragment-single-edit-propagation install test enforces).

* fix(#4658): accept both Dart quote styles and extract the shared bounded walk

The isolated reviews' remaining findings: the Dart marker carried only the
single-quote anchor, missing legal double-quoted imports (a spec-narrowing
deviation); the two evidence walks duplicated the subtle MAX_WALK_ENTRIES
cap semantics verbatim, so they are extracted into one walkProjectFiles BFS
with a visit callback; tests now use the createTempDir helper, shared
fixture literals that cannot drift from NATIVE_UI_CONTENT_MARKERS, a
double-quoted Flutter import case, and drop a vacuous assertion and a
mid-body re-require alias.

* docs(#4658): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 15:29:19 -04:00
Tom Boucher
5e729445d3 fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel

Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first
proposeConsiderations regression (Danish text + English text_en must classify
as its English equivalent, not land in the #1110 unclassified sentinel),
proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en
validation cases (empty/whitespace/non-string, unconditional under an
elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a
reference-doc Inputs parity test, and a fast-check property proving any
cue-matching prose classifies identically under a cue-free Danish rendering
plus text_en. All new assertions are RED until src/ui-consideration-probe.cts
and the workflow/reference docs are updated.

* fix(#4657): give the ui consideration probe a text_en language channel

Element gains an optional text_en; classifyElement's own signature stays
untouched (a locked, directly-tested export) and the text_en ?? text
selection is pushed to the two classification call sites (proposeConsiderations,
proposeElements) instead. text_en is validated fail-closed: an empty or
whitespace-only value throws rather than silently winning the ?? fallback and
degrading classification to zero kinds.

Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the
Non-English projects section (mirroring spec-phase Step 5.5) and the
ELEMENTS_JSON shape comment documents the field with both zero-applicable
guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT
predicate (with both derived indexes regenerated), and the nav-override
test expectation stay in sync. The ui-phase contract test carries the
site-scoped allow-test-rule marker and its cluster is registered in the
test-file-count allowlist ratchet.

Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657)

* docs(#4657): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 13:46:40 -04:00
Tom Boucher
3014775a3f fix(#4656): expose coverage.unclassified and widen the zero-applicable guards (#4800)
* fix(#4656): expose coverage.unclassified and widen the zero-applicable guards

* fix(#4656): regenerate golden coverage fixtures and update the rollup pin

Emitted-Drift-Ack-Growth: spec-phase.md — #4656: guard widened to the all-unclassified case, doc claim corrected
Emitted-Drift-Ack-Growth: ui-phase.md — #4656: guard widened identically

* fix(#4656): sync edge-probe doc blocks and coverage pins with the new field

* fix(#4656): key the mandatory confirmation on the widened guard

* docs(#4656): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 10:33:32 -04:00
0xdhx
003d982c83 fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint

Two defects in the covered-input fingerprint (#4155), one issue.

1. `computeCoveredDigest` hashed the whole bytes of every declared path
   uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` —
   which every phase rewrites as ordinary bookkeeping, and which the closing
   phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER
   the verifier ran — flipped every phase that declared them to `stale` on
   zero implementation change, and from there `isPhaseComplete` →
   `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate.
   Fingerprint v2 leaves any direct child of a planning root out of the
   hash: `.planning/` itself, plus the phase's own planning root (the parent
   of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/`
   layouts are covered without the digest knowing what a workstream is —
   `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather
   than a name list so the set cannot drift; a root is accepted only when the
   phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated
   exactly as every other covered path (confined, present, a regular file —
   the fail-closed contract is unchanged); only its bytes are ignored, and a
   declaration made only of shared documents fails closed like an empty one.
   A stored digest names its version, and `readVerificationStatus` now
   recomputes under THAT version (`parseFingerprintVersion`,
   `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics
   until it is re-fingerprinted, so the upgrade alone stales nothing; a
   version this build cannot recompute fails closed.

2. `verification.fingerprint` received a raw positional slice, so
   `--files a`, `--files "a,b"` and `--files a --files b` all put the literal
   token into the covered set and failed closed as "a covered file is
   missing, unreadable, or escapes the project root" — the message that
   convinced the reporting project the digest was permanently
   unrecomputable. `parseFingerprintFileArgs` accepts every form (plus
   `--files=a,b`, freely mixed with bare positionals), treats any other
   `--flag` and an empty `--files` value as usage errors that say so, and
   the phase-dir argument must now be an existing directory: omitting it
   used to take the first covered file as the phase dir and print a
   plausible digest over the rest at exit 0.

Regression tests (tests/verification-status.test.cjs, #4623 block): the
cross-phase case from the report, the same-phase `requirements
mark-complete` / `phase.complete` cases from the thread, a workstream-scoped
root, v1-preserved / unknown-version-stale, the fail-closed cases (missing,
directory, escaping symlink, all-shared), every `--files` form against the
bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and
AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34
fail, the 7 that pass pin behaviour the fix must leave unchanged.

Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's
covered_files instruction (rewritten in place — the file sits 21 bytes under
its LARGE hard cap), gsd-core/templates/verification-report.md.

Fixes #4623

Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi

* chore(#4623): set changeset fragment pr to 4749

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 05:26:44 -04:00
Dennis Alexis Valin Dittrich
ad1477d659 enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap

* enhance(260903-m7p): validate configured entrypoints before success

* test(260903-m7p): require pre-success entrypoint validation

* enhance(260903-m7p): gate install success on entrypoints

* test(260903-m7p): cover configured entrypoints across runtimes

* enhance(260903-m7p): cover emitted runtime entrypoints

* fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number

- finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process
  before the new configured-entrypoint assertion throws. Without a HOME +
  config-location-env sandbox that write resolved through the ambient
  environment and landed in the developer's live ~/.gsd (confirmed absent on
  origin/next baseline, present only on this branch — full-suite HERMETICITY
  WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the
  duration of the test, matching the existing in-process finishInstall/
  install() pattern in tests/install.test.cjs (#2665).
- .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder
  (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the
  fork PR number until the upstream PR number is known.

* fix(260903-m7p): repair cross-platform and pre-existing shape fallout

- tests/configured-entrypoint-validation.test.cjs: the win32 branch of
  ensureCodexHooksJsonSessionStart writes a .cmd shim under
  <codexRoot>/hooks/; create that dir in the test (the real installer only
  calls this once hooks/gsd-check-update.js already exists) and assert the
  platform-common entrypoint shape instead of a fixed non-Windows array,
  since win32 legitimately emits two entries (cmd shim + script).
- tests/install.test.cjs: finishInstall's shared settings-json return now
  carries configuredEntrypoints/rollbackInstallerMigrations for every
  runtime on that path (trae included, not just Claude/Cursor/Windsurf);
  update the trae install() exact-shape assertion to match.

* fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash

configuredEntrypointsForHook's shell branch dropped interpreterCandidates
entirely when resolveBashExecutable returned null, unlike the sibling
portableHooks runner entry a few lines below (which correctly falls back to
the literal 'bash' token). Found via agy adversarial review; verified
unreachable through the current call graph (buildHookCommand's own
resolveBashRunner==null gate already short-circuits before
recordConfiguredHookCommand runs), so this is a defensive consistency fix,
not a live-bug patch — kept for the next caller that does not share that
gate.

* chore(260903-m7p): backfill changeset pr number to the opened upstream PR

.changeset/quick-wasps-sing.md carried the fork PR number (16) as a
placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now
open, so record its real number per CONTRIBUTING.md's changeset pr-field
convention.

* fix(#4154): track already-registered hooks for entrypoint validation on update

applySettingsJsonHooks registers each guard hook only if absent, so a hook
already present from a prior install keeps its stale on-disk command. The
new entrypoint tracker always records the freshly-computed command for it,
which never matches what is actually persisted, so the exact-string filter
in finishInstall silently dropped it from validation — the Blocker case
this feature exists to catch (an already-installed entrypoint going stale
between installs) was exactly the case it never validated.

Match on the managed script's basename instead, which the persisted
command carries either way, so an already-registered hook stays in the
validated set. Regression test forces this path by mutating a
freshly-installed hook's persisted command before a second install.

* fix(#4154): distinguish an unreadable script from a missing one

validateConfiguredEntrypoints folded an EACCES statSync failure into the
same 'missing' reason as ENOENT, misreporting a real permission problem as
an absent file. Check the error code and report 'unreadable' instead.

* docs(#4154): document entrypoint validation's rollback and PATH scope

CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had
no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite
bin/install.js x CONTEXT.md being this repo's strongest co-change pairing.

The update-gsd.md how-to overstated what a validation failure undoes: for
Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/
config.toml inside install() before the aggregate validation call runs, so
there is no rollback path for that write regardless of "where available"
phrasing. Also note that interpreter resolution checks the installer's own
PATH, not necessarily the PATH a hook fires under later (#2979 launchers).

* chore(#4154): point changeset pr field at the fork PR while CI runs there

Mirrors the branch's own prior backfill commit: pr: matches whichever PR
number changeset-lint is currently validating against (fork PR #16 during
the fork-first CI/review loop), flipped back to the upstream PR number
right before the final push to open-gsd/gsd-core.

* fix(#4249): address adversarial-review findings in entrypoint validation

An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR
found several real gaps beyond the human reviewer's Blocker, verified
against source before fixing:

- Codex's install() result bound rollbackInstallerMigrations to the narrow
  installer-migrations-only rollback instead of restoreCodexSnapshot (#3245),
  the full pre-install snapshot/restore Codex already owns for exactly this
  case — a validation failure discovered outside install() reverted nothing
  of the config.toml/hooks.json that call had already written.
- The register-only-if-absent basename match from the prior fix used a bare
  substring, which an unrelated user command mentioning the same filename
  could false-positive into GSD's validated set — anchored on the
  `/hooks/<basename>` path segment instead.
- nodeCandidates checked raw process.execPath (always true — we're running
  in that process) instead of normalizeNodePath's stable version-manager
  alias, the same one buildNodeRunnerChainToken bakes as its first choice —
  a false green regardless of whether that alias itself still resolves.
- An entry with no interpreterCandidates (Cline's PreToolUse hook, or a
  Windows-Claude .sh hook invoked without a bash runner) runs via its own
  shebang; validateConfiguredEntrypoints checked only file-type, never the
  execute bit. Cline's writer also never reported an entrypoint at all.
- Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor
  hook registered across several events) were validated once per duplicate.

Each fix is covered by a new or extended test; the Codex one required
inlining runCodexInstall's env sandboxing so the rollback closure — which
re-resolves the $HOME-relative skills root live — runs before the sandbox
is torn down, matching how installAllRuntimes' real aggregate gate calls it.

* docs(#4249): document the round-2 entrypoint-validation fixes

Runtime Hooks Surface Module and Installer Module entries now name
ConfiguredEntrypoint's not-executable reason, the normalizeNodePath
alignment, Cline's tracked hook, and which install() result the
finishInstall/installAllRuntimes rollback path actually reverts per
runtime (Codex's full snapshot vs. the others' narrow migrations-only
rollback).

* chore(#4249): point changeset pr field at the upstream PR now that fork CI is green

* fix(#4249): address agy adversarial-review findings

- validateConfiguredEntrypoints: statSync alone never detects a
  chmod-000 script (it only needs parent-dir search permission), so an
  interpreter-invoked entry with an unreadable script passed validation.
  Add an explicit R_OK check for the interpreterCandidates branch only —
  the candidate-less/shebang branch already has its own X_OK gate.
- docs/how-to/update-gsd.md: the blanket "does not revert" claim was
  false for Codex, which reverts config.toml/hooks.json via its full
  pre-install snapshot; qualify it per runtime.
- tests/codex-config.test.cjs: the #4249 rollback regression test
  asserted skills/ and VERSION were reverted but never asserted
  config.toml/hooks.json were too, despite the test's own stated intent.
- CONTEXT.md: qualify which interpreterCandidates entries get
  normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's
  Windows .cmd shim as a third candidate-less case that relies on
  extension dispatch, not a shebang.

* fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit

Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node',
so it needs the execute bit (like any shebang-invoked entry), but its
interpreter is looked up on PATH by 'env' at hook-fire time (unlike
every other GSD JS hook, which bakes an absolute node path specifically
to avoid that dependency). The candidate-less/interpreterCandidates
fork treated these as mutually exclusive, so Cline's entry silently
skipped interpreter resolution entirely — a completely missing 'node'
on PATH would still validate successfully.

Add an orthogonal selfExecutable flag so both checks run for entries
that need them. (CodeRabbit finding on the fork rehearsal PR.)

* fix(#4249): address second-round adversarial review findings (opus + agy)

- validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry,
  not just interpreterCandidates ones — a self-executable shebang script
  is still opened and read by its kernel-invoked interpreter, so X_OK
  alone never proved it was readable.
- selfExecutable is now the sole, explicit source of truth for the
  execute-bit check (every producer that needs it sets the flag) instead
  of being partly inferred from an absent interpreterCandidates, which
  Cline's hybrid entry also carries.
- The execute-bit check now skips explicitly on win32 (matching
  resolveExecutableBinary's own carve-out) instead of relying on Node's
  accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows
  machine and not a test that simulates win32 on a POSIX runner.
- bin/install.js: fixed a stale comment claiming no runtime's
  install()-time writes have a rollback path — Codex's does
  (restoreCodexSnapshot) — and added the omitted Cline to both that
  comment and CONTEXT.md's equivalent lists.
- CONTEXT.md: fixed the Cline description left stale by the previous
  commit's selfExecutable addition, and rewrote the validation-mechanism
  paragraph for clarity (writing-for-agents pass).
- docs/how-to/update-gsd.md: split an overloaded 4-clause sentence.
- Removed a fault-injection integration test that could not reliably
  exercise the real installAllRuntimes -> finalize -> rollback wiring
  without fighting the installer's own pre-registration existence
  guards; the constituent pieces remain covered individually.

* fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner

X_OK is a POSIX-only concept, skipped entirely when an entry's platform
is win32 (matching production). Two test entries omitted platform,
defaulting to process.platform — on an actual windows-latest CI runner
that silently skipped the very check they were meant to exercise,
turning 'not-executable' into a false pass. Pin platform: 'linux' so
these are deterministic regardless of which OS runs the suite.

* fix(#4249): classify EPERM the same as EACCES in statSync error handling

Windows raises EPERM (not EACCES) for a parent directory that couldn't
be traversed into — was falling through to 'missing', misreporting a
genuine permission problem as a nonexistent path.

* docs(#4249): address final CodeRabbit doc-completeness findings

- CONTEXT.md: install()'s documented result shape omitted
  configuredEntrypoints; the ConfiguredEntrypoint shape omitted
  selfExecutable.
- docs/how-to/update-gsd.md: the failure-mode sentence omitted
  unreadable and lacks-execute-permission, which the installer also
  rejects.

* fix(#4249): stop double-validating every configured entrypoint on install/update

installAllRuntimes' finalize() already runs assertConfiguredEntrypoints
once over the aggregate set; finishInstall then re-ran the identical
check per runtime in the printSummaries loop right after, so every
entrypoint paid its statSync/accessSync/interpreter-resolution cost
twice on every install and update. Add entrypointsAlreadyValidated to
skip the redundant pass specifically on that path, while leaving the
check intact for any caller that invokes finishInstall directly.

* chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there

* perf(#4249): memoize interpreter candidate resolution across entrypoints

resolveExecutableBinary walked PATH once per (entry, candidate) pair; a
typical install has a dozen-plus entries sharing the same few candidate
lists (process.execPath for JS hooks, bash for shell hooks). Cache by
(platform, candidate) so each distinct pair resolves once per validation
call instead of once per entry.

* chore(#4249): point changeset pr field at the rebased rehearsal fork PR

* fix(#4249): drop entrypoint tracking from the now-dead Codex event writer

#2586 (landed on next after this branch forked) removed install.js's
CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent
no longer runs during install or update. The ConfiguredEntrypoint records
this branch added inside it were therefore unreachable and untested. Restore
the function to its upstream shape; the entrypoints it used to report were
never collected by any caller.

* refactor(#4249): drop the revalidation bypass flag and the candidate cache

Both were this PR's own micro-optimisations over a set of roughly a dozen
entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall
gate off to save one statSync/accessSync pass; `resolvedCandidateCache`
memoised resolveExecutableBinary across entries that are already deduped by
(configPath, scriptPath). Neither is measurable, and the flag was the only
way to reach finishInstall with validation disabled. finishInstall now
always validates what it is given.

* chore(#4249): point the changeset pr field back at the upstream PR

* refactor(#4249): track settings.json entrypoints without the hooksSurface gate

The install-surface writer only tracked configured entrypoints when the
runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing
asserts that axis agrees with `installSurface`, so a descriptor that broke
the coupling would silently pass `configuredEntrypoints: undefined` and drop
that runtime out of the validation this PR adds — reintroducing the exact
'reports Done! over a broken entrypoint' failure #4154 exists to close.

Remove the dependence rather than test it: everything recorded on this path
lands in settings.json by construction, and the registered-command filter
already discards entries no persisted hook references.

* chore(#4249): put the changeset body in the documented two-part format

CONTRIBUTING.md and .changeset/README.md both show
`**<bold change>** — <symptom-led explanation>.`; the fragment was a single
unbolded sentence.

* chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there

* fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback

#3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-*
and gsd-core/VERSION. The install overwrites every other GSD-owned file too —
hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest
itself — before the entrypoint-validation gate runs, so a validation failure
left the new payload sitting on top of the restored old config.

Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims,
before runInstallerMigrations so the bytes are the true pre-install state, and
restore it from both Codex rollback closures ahead of the per-surface restores.
Files only the failed install introduced are removed, read from the manifest
now on disk. The manifest is already the authoritative record of what GSD owns,
so no second hand-written list can drift out of sync, and user-owned files are
never snapshotted or removed. Every path is confined through
resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into
an arbitrary-path write.

Non-Codex runtimes are unaffected: the snapshot is gated on the same
tomlConfigInstall + non-minimal condition as #3245's.

* fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest

Two follow-on defects in the previous commit's snapshot:

- The capture was gated on `!isMinimalMode`, copied from #3245. A core/
  --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the
  manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the
  snapshot came back empty while the rollback still ran — and its removal pass
  would have deleted every file the new manifest lists. Gate on
  tomlConfigInstall alone, matching where the rollback actually reaches.

- An unreadable or unparseable prior manifest was caught alongside ENOENT and
  treated as a fresh install. That is the same empty-snapshot state, so a failed
  update over a real install with a corrupt manifest could delete its prior
  payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means
  known-empty; any other read error or a parse failure means unknown, and the
  restore closure returns without touching anything, degrading to #3245's
  narrower rollback. Deliberately not fatal — a corrupt manifest has to stay
  repairable by reinstalling over it.

Both paths are covered by red-checked regression tests.

* fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too

commit removed from the manifest snapshot. restoreCodexSnapshot is reachable
for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-*
skill dir and gsd-* agent file the snapshot does not claim — so with an empty
minimal-mode snapshot a rollback deleted the whole skills/agents surface with
nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via
the ADR-1239 skills-kind home override, so this is also the reason manifest
`skills/` keys do not resolve under configDir: that surface belongs to this
snapshot, not to the manifest-driven one.

Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal
mode — doing nothing on an early failure is the non-destructive side.

Covered by a red-checked regression test that plants bytes in an alternate-home
skill file, reinstalls under the core profile marker, and asserts the rollback
restores it.

* fix(#4249): never remove on rollback unless a prior manifest proves what predates the install

Three defects in the manifest-driven Codex rollback, all in its removal half:

- ENOENT marked the snapshot usable, arming the removal pass on a FIRST
  install. GSD may have overwritten a user's file at a manifest-tracked path
  there, and no prior manifest records the difference — so rollback deleted it
  where before it merely left it overwritten. Absent, unreadable and malformed
  manifests now all leave the prior set UNKNOWN and skip removal entirely.

- Membership was tested against the map of files whose pre-install read
  SUCCEEDED, so a tracked file that existed but was unreadable read as
  introduced-by-this-install and was removed. Track the prior manifest's paths
  in their own Set and test against that.

- The unreachable "delete the manifest when there was no prior one" branch is
  gone: usable now implies a parsed prior manifest.

Also adds the end-to-end test the aggregate gate was missing — the four Codex
rollback tests drove the closure directly, proving the restore but not the
wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes
Cline's `env node` entry fail validation for real, and asserts Codex's payload
comes back.

Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper
instead of six copies. Both new tests are red-checked.

* test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test

lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its
Windows-EBUSY retry budget. That budget is for directory trees; this removes a
single file, which unlinkSync says more precisely and the rule does not flag.

* chore(#4249): point the changeset pr field back at the upstream PR

* fix(#4249): use an unambiguous dedup key and surface partial-restore failures

trek-e's 2026-09-08 adversarial pass flagged two findings in the new
entrypoint-validation/rollback code:
- assertConfiguredEntrypoints' dedup key already used a raw NUL
  separator (introduced in ceebb65f2d), but git/Read render NUL as a
  space, so the key looked like a plain-space join to every reviewer
  that read the diff. Replace it with JSON.stringify([configPath,
  scriptPath]) so the separator is visible and unambiguous.
- restoreManagedFileSnapshot's per-file restore catch block claimed to
  'surface the original error' but only swallowed it, matching (and
  widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add
  an actual console.warn using the existing best-effort-warning
  convention, scoped to just this PR's new function.

* fix(#4249): treat a files-less prior manifest as unknown, not known-empty

agy's gemini-3.8-flash-high adversarial pass (round 5) found and I
reproduced empirically: a structurally-valid manifest missing the
files key (e.g. {"version":1}) parses without throwing, so
Object.keys(undefined || {}) silently read as 'zero files predate
this install' instead of the UNKNOWN state the malformed-manifest
guard exists to produce. Rollback's removal pass then deleted every
GSD-owned file the failed install's own manifest listed, including
ones that predated it — the exact data loss the #4249 CodeRabbit
malformed-manifest fix was supposed to prevent, reachable through a
JSON.parse success instead of a failure. Route the shapeless case
into the same catch-all UNKNOWN path via an explicit shape check.
Regression test reproduces the deletion before the fix and confirms
the file survives after it.

Also extend restoreManagedFileSnapshot's removal-pass rmSync and
final manifest-rewrite catches with the same real console.warn
trek-e's round-4 review asked for on the per-file restore catch —
same rollback function, same operator-facing-signal gap.

* docs(#4249): correct which runtimes actually leave a written config on rollback

agy's completeness audit (round 5, holistic pass) caught this new
paragraph claiming 'for every other runtime, the configuration file(s)
already written during that update are left in place' — false for
Claude Code and other settings.json-based runtimes, whose write never
happens on failure (assertConfiguredEntrypoints runs before
finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline
actually match that description, since they persist their config file
inside install() ahead of the gate. Split the one sentence into the
three actual outcomes; matches the PR body's own accurate Before/After
wording, which this doc addition had drifted from.

* fix(#4249): clean up doc/comment mismatches and dead fields from opus review

Opus critical-code-reviewer + ponytail-review pass on the final diff:
- assertConfiguredEntrypoints carried finishInstall's old docblock
  ("Apply statusline config, then print completion message") from
  before this function was inserted between comment and callee.
  finishInstall already has its own accurate #4249 comment, so the
  stale docblock is removed rather than moved.
- checked: number on ConfiguredEntrypointValidationResult and
  error.configuredEntrypointValidation on the thrown error: the first
  had zero consumers anywhere in the repo, including its own defining
  file, and is removed. The second matches an existing repo
  convention (bin/install.js's installerMigrationRollbackFailures,
  #4249 predates this PR) of attaching structured diagnostic context
  to a re-thrown Error even before a consumer exists, so it's kept.
- finishInstall's own assertConfiguredEntrypoints call is a redundant
  backstop on the real production path (installAllRuntimes's aggregate
  call already validates the superset first), but its comment read as
  though this call alone provided the before-the-write guarantee.
  Clarified rather than removed — it's the only gate for a caller that
  invokes finishInstall directly.

* chore(#4249): split the manifest-driven rollback engine out into #4544

Issue #4154 asked the installer to consume a validation failure "through
the existing rollback mechanism, without a second transaction mechanism".
The manifest-driven rollback widening added during review (capture every
path the prior gsd-file-manifest.json claims, restore those bytes, remove
what only the failed install introduced) is that second mechanism on a
plain reading. It is a real fix for a #3245-era gap, but an independent
one, so it moves to its own bug report and PR.

Removed here:
- bin/install.js: the pre-install managed-file capture block and
  restoreManagedFileSnapshot, plus its call sites in
  _codexPreConfigRollback and restoreCodexSnapshot (99 lines).
- tests/configured-entrypoint-validation.test.cjs: the five tests that
  exercise the manifest engine.
- CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the
  widened restore. update-gsd.md again documents the #3245 surfaces only.

Kept, because it is #4154's own scope:
- the entrypoint-validation gate itself;
- Codex's install() result binding rollbackInstallerMigrations to
  restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*,
  agents/gsd-*, gsd-core/VERSION);
- the !isMinimalMode gate removal on that snapshot. Binding the closure
  to the result made it reachable for a core/--minimal install, where its
  pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot
  does not claim; an empty minimal-mode snapshot therefore deleted the
  whole surface with nothing to restore.

The surviving aggregate-failure test now asserts on config.toml, a
surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which
only the manifest engine restored.

Refs #4544

* test(#4249): cover configured entrypoints through the packed install path

#4154's scope lists install smoke coverage alongside the installer gate —
"assert representative configured entrypoints resolve for supported runtime
profiles". The gate itself (assertConfiguredEntrypoints /
validateConfiguredEntrypoints) is unit-covered by in-process install() calls;
nothing proved the property survives npm pack -> npm install -g -> install.js.

Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct
config surfaces GSD writes launch paths into (settings.json, and hooks.json +
config.toml) — run the tarball-installed installer into a throwaway HOME, then
re-read that runtime's own written config and return the new
ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a
file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check
becomes a release gate on every matrix host without workflow changes.

The scan re-derives paths from the written config instead of reusing the
installer's own entrypoint list, and test I shows why that matters: a
registration the installer never touched during a run is invisible to the
in-process gate, so the install exits 0 and only reading the config back off
disk catches the dangling launch path.

* ci(#4249): pack a publish-shaped tarball in the install smoke lane

`npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs
build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs
after `npm ci` carries no hook scripts at all — the lane has been smoking a
package that differs from the published one in exactly the artifacts the
lifecycle smoke is supposed to launch.

That went unnoticed because the lane's init runs `--local`, which registers no
statusline and therefore registers no hook whose target is missing. A
`--global` install on the same tarball exits 1 on #4249's own gate
(`gsd-statusline.js (missing)`), which is what the new configured-entrypoint
cycle performs, so without this step the cycle would report INIT_FAILED
instead of checking anything.

Build hooks before packing so the smoked tarball matches prepublishOnly. The
CLI now reports 16 configured entrypoints for claude and 1 for codex instead
of zero.

* fix(#4249): scope Codex's full snapshot restore to entrypoint failures

Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage
exception un-install a Codex install that had already succeeded and already
printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps
the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints`
call.

Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration`
scopes finalize-stage rollback to installer *migrations* ("the executor uses the
journal to restore modified paths"), and this PR's own operator-facing paragraph
in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/
agents/VERSION revert to entrypoint-validation failures specifically ("If a
script is missing, unreadable, ... For Codex, this reverts ..."). The wide
behaviour is also incoherent as a transaction abort: the same doc says Cursor,
Windsurf, Kimi and Cline keep the config they wrote inside install().

Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall
hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install
bytes — on an update, silently downgrading a working Codex install to the
previous version while the user had just been told it was Done.

Select the rollback by error kind instead. `assertConfiguredEntrypoints` already
tags its error with `configuredEntrypointValidation`, so the full snapshot
restore runs for that error (and anything downstream of it, including
finishInstall's per-runtime backstop) and the installer-migrations-only closure
runs for everything else. The codex result now also exposes that narrow closure
as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning
the full restore, so the direct-call contract asserted by
tests/codex-config.test.cjs is unchanged.

Adds a regression test that installs codex+kilo together, injects EACCES on the
Kilo permission write by monkeypatching node:fs (restored in a finally — never
chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps
the bytes the successful install wrote. Verified red against the pre-fix
unconditional path.

Cline cannot host this test: its plan is writesSharedSettings:false +
finishPermissionWriter:null, so its finishInstall performs no write and has no
non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally
(unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded
fs.writeFileSync.

* docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing

CONTEXT.md still described Codex's rollback as an unconditional bind
to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint-
validation failures via rollbackInstallerMigrationsOnly and the
configuredEntrypointValidation error tag. Caught during the round-6
PR body pass.

* fix(#4249): stop rollbackInstallerMigrations meaning its own opposite

Codex's install() result bound `rollbackInstallerMigrations` to
restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the
actual installer-migrations-only closure behind
`rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name
meant the opposite of what it says, and CONTEXT.md had to concede as much
in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure
for every runtime, matching both its name and the meaning it already has
on next, and the snapshot restore gets its own Codex-only field,
`rollbackPreInstallSnapshot`. The selection in
rollbackFinalizedInstallerMigrations collapses to one line and no longer
needs a fallback chain.

Also in this commit, all against the same rollback path:

- Correct the rollbackFinalizedInstallerMigrations comment. It read as if
  the round-6 narrowing prevented any sibling-triggered revert of a Codex
  install the user has already seen "Done!" for. It does not, and is not
  meant to: `wide` is true for ANY entrypoint-validation error from ANY
  runtime, because the aggregate gate is all-or-nothing — an invalid Cline
  entrypoint reverts Codex's snapshot, which
  tests/configured-entrypoint-validation.test.cjs's 'an aggregate
  entrypoint validation failure rolls the Codex install back (#4249)'
  asserts directly. The discriminator is the error's KIND, not which
  runtime owns the failing path. Comment and CONTEXT.md now say that.

- Name the runtime in the "Configured entrypoint validation failed" error.
  ConfiguredEntrypointInvalid already carries `runtime`; the message threw
  it away, leaving an operator of a multi-runtime install unable to tell
  whose entrypoint broke — which matters precisely because the failure can
  revert a runtime that was itself fine.

- Set `configuredEntrypoints: []` explicitly on the copilot-instructions
  early return. Every other branch states the key; this one relied on
  installAllRuntimes' `(result.configuredEntrypoints || [])` defence.
  `[]` is correct, not a workaround: every Copilot hook is an inline
  printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no
  GSD-managed script or interpreter to resolve.

No behaviour change beyond the error-message text.

* docs(#4249): narrow the smoke scan's config-surface claim to what it checks

RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is
told to launch is registered in one of settings.json / hooks.json /
config.toml, and that nothing else in a config dir is runtime
configuration. Both halves are false as stated. Cline registers its hook
at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those
names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native
[[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi),
a directory separate from Kimi's own GSD configDir — the same gap
installer-migration 007 already documents as structurally unreachable.

The scan is in fact correct for what it runs against: entrypointRuntimes
defaults to claude + codex, whose launch paths do all live in those three
top-level files. Restate the docstring at that scope, name the two known
out-of-scope surfaces, and warn that adding either runtime to
entrypointRuntimes without teaching scanConfiguredEntrypoints about its
surface yields a scan that finds zero entrypoints and proves nothing.
The entrypointRuntimes default comment carried the same overgeneralization
("every other runtime reuses one of them") and is corrected with it.

Documentation only; no code change.

* fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json

An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found
that install()'s settings-json early return for an unparseable
settings.local.json omitted configuredEntrypoints and
rollbackInstallerMigrations from its result, unlike every other branch.
rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations
unconditionally, so this branch silently dropped its own installer-migration
rollback on a later finalize-stage failure.

Completed the return shape: configuredEntrypoints: [] (matching Copilot's
equally-early no-entrypoints-yet return) and rollbackInstallerMigrations
(already in closure scope). Red-then-green regression test added.

* test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs

Same internal adversarial review (round 8): this suite's before() packed the
tarball directly, without the ensureHooksDist() guard every sibling
install-test suite (install.test.cjs, install-minimal-hooks.test.cjs,
mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run
in isolation ahead of a suite that builds hooks/dist itself, this suite's
pack would ship a tarball with no hook scripts and fail closed on
SMOKE.INIT_FAILED instead of testing anything.

* fix(#4249): refresh stale test-timings weight for the codex-config split

next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's
heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the
CI shard packer's weight table (tests/test-timings.json) was never updated:
codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the
suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block
this PR extends with its own #4249 install()-pipeline test — had no entry at
all, so the packer would silently underestimate it at the table's median
weight (roughly a 9x underestimate against its real cost).

trek-e's most recent review flagged a Windows shard timeout in-flight on
codex-config.test.cjs, plausibly aggravated by this PR's own addition to that
file before the rebase moved it. Re-measured both files locally (node --test
--test-reporter=tap, max of 3 runs, matching the table's own max-across-streams
methodology) and patched just these two entries — not a full regeneration,
which would need real multi-lane CI data this session doesn't have access to.

* fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists

next's platform-conformance-tier classifier (#4591/#4598) landed after this
branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and
tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's
gen-platform-conformance-tier --check and both the Linux and macOS conformance
suites.

* fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation

next's #4733 (landed after this branch's last rebase) replaced the static
ISOLATED_HEAVY_FILES set with a threshold derived live from
tests/test-timings.json, and pins the current derived result in
EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still
named codex-config.test.cjs, whose own weight this PR already dropped from
127783ms to 189ms (after splitting its heavy install()-pipeline blocks into
codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The
live-computed set correctly no longer includes it; the pinned expectation is
updated to match.

* fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it

trek-e's review flagged two Major gaps: the thrown error read identically
regardless of which of three real outcomes a runtime hit (nothing
persisted / snapshot reverted / config left broken on disk), and no test
proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/
Kimi/Cline — only Codex's revert path was ever asserted.

assertConfiguredEntrypoints now tags each invalid entry with its actual
consequence, mirrored from docs/how-to/update-gsd.md's existing
rollback-matrix disclosure. A new test drives the same aggregate failure
through Cline (whose own entrypoint is the one that fails) and asserts
its hook file is still on disk afterward.

* fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR

One review pass (gemini-3.8-flash-high via the antigravity review lane)
against this PR's full diff against next, findings independently verified
against source before fixing:

- Copilot's install() return object was the only one of 6 runtime branches
  missing rollbackInstallerMigrations — reachable now that this PR's own
  aggregate gate runs rollback across every result on any runtime's
  entrypoint failure, not just Copilot's own.
- buildHookCommand's unresolved-bash early return skipped track() entirely,
  so a win32 install with no Git Bash silently produced an unregistered
  .sh hook instead of the 'unresolved-interpreter' validation failure
  configuredEntrypointsForHook's own comment said it would.
- release-tarball-smoke.cjs reported a Cycle 4 install failure under
  SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing
  SMOKE.INSTALL_FAILED.
- SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's
  trailing args, which also truncated any configDir containing a space
  (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan.
  Anchored the match on the already-known configDir prefix instead of a
  generic absolute-path guess: removes the ambiguity outright rather than
  patching the character class, and stays a raw-text scan on purpose (it
  catches a writer that emits a path without registering it — a
  JSON.parse of the expected schema would miss exactly that case).

One suggested finding (test-timings.json "missing" the new test file) was
verified false — that table only holds measured CI timings, populated
after a file's first real run — and one Ponytail suggestion (a
JSON.stringify dedup key) was rejected as it would reintroduce a real, if
narrow, key-collision risk for no benefit.

* fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment

changeset-lint requires pr: to match the PR it runs on (16 on the fork,
not the eventual upstream number) — rehearsal-branch convention already
established earlier in this PR's history.

lint-docs-guard-registration's quote-pairing heuristic doesn't require the
docs/ path itself to be quoted — it flags a file once ANY quote-delimited
span containing "docs/" appears anywhere in it, alongside any real fs read
call. A comment ending "...update-gsd.md's rollback-matrix paragraph"
supplied the closing quote character (the possessive apostrophe) the
heuristic paired with an unrelated single-quoted string earlier in the
file. Reworded to avoid the unquoted apostrophe next to the path.

* chore(#4249): point the changeset pr field back at the upstream PR

Fork rehearsal (PR #16) is green; the real target for this changeset is
upstream PR #4249.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 04:23:53 -04:00
Tom Boucher
740ba0d8a3 fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file

Co-authored-by: sim <sim@local>
2026-09-16 02:38:43 -04:00
Behruz Nassre Esfahani
febe6c9885 fix(#4685): a directory artifact fails its own entry instead of aborting the check (#4735)
* fix(#4685): a directory artifact fails its own entry instead of aborting the check

`must_haves.artifacts` entries are read with `safeReadFile`, which rethrows every
errno except ENOENT. A listed path that is a directory therefore threw EISDIR out
of the per-artifact loop: `query verify.artifacts` printed

    Error: EISDIR: illegal operation on a directory, read

and reported NOTHING — not the offending entry, and not the plan's other,
perfectly checkable artifacts. One directory entry disabled the whole plan's
check. Reproduced against a real plan before the fix, and after.

A directory is now reported as that entry's own failure, with an issue distinct
from `File not found` (the path did resolve; it simply is not the thing an
artifact entry can be checked against), and every other artifact in the plan is
still checked and reported independently. Anything else the stat or read throws
becomes that entry's failure too, carrying its errno, rather than discarding the
run — a check that disappears is worse than one that fails, because a failure is
visible.

Verifying directories properly — matching `contains:`/`min_lines:`/`exports:`
across the files inside one — is a feature decision and deliberately not made
here, per the issue's stated scope. Authoring-time rejection of a directory path
is likewise left alone: the brief raises it as a separate question, and the
runtime fix does not depend on it.

Also, found in pre-PR review and pre-existing: `safeReadFile(...) || ''` turned a
post-stat ENOENT into empty content, so an artifact declaring only
`path`/`provides` had no criterion left to fail and passed, having checked
nothing. A null read now reports that instead of inheriting a pass. It can only
turn a false pass into a failure.

Verified: reverting src/verify.cts to the merge-base turns both new rows red with
the exact EISDIR message; lint:ci exit 0; full suite 24/24 chunks, 37,280 tests,
0 failures; tests/verify.test.cjs 219/219.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* chore(#4685): backfill changeset PR number to 4735

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

* test(#4685): pin the injected-I/O branches, and narrow the guard's comment

Review findings from #4735.

Major — the two error branches this PR adds were untested, and the PR body
claimed the mid-check ENOENT window was "not deterministically reproducible
through the CLI seam these tests drive." That was wrong: ADR-3574 records this
repo's convention for exactly this — inject filesystem failures by
monkeypatching the fs method, never by chmod or mode-bit tricks, which root
bypasses and yields a test that passes with zero coverage in root Docker and CI.

Both branches are now pinned that way. The injection runs in the CHILD via
NODE_OPTIONS=--require, because `output()` writes fd 1 directly
(`writeAllSync(1, …)`, io.cjs) rather than through console.log, so an in-process
call cannot have its JSON captured. The preload patches the child's own module
objects, which the compiled code reads at call time.

  - a file that disappears between stat and read now fails as that entry rather
    than passing on empty content
  - a non-ENOENT errno (EACCES) is reported as that entry's failure, carrying its
    code, so an operator can tell a permissions problem from an I/O one

The injection matches the target by path SUFFIX, not string equality: the first
cut compared absolute paths, and a /tmp vs /private/tmp prefix difference
silently disarmed it — the test passed while asserting nothing. A disarmed
injection test is worse than no test, so the reason is recorded at the call site.

Nit — the comment above the try block said the guard "covers what the stat and
read below actually throw", which reads as if the min_lines/contains/exports
checks inside the same block were deliberately guarded too. They are pure string
operations and cannot throw; the comment now says so rather than implying a
guarantee it does not make.

Verified: reverting src/verify.cts to the merge-base turns all three #4685 rows
red — the directory row and both new ones; lint:ci exit 0; full suite 24/24
chunks, 37,389 tests, 0 failures; tests/verify.test.cjs 221/221.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 01:39:52 -04:00
Twisted Fate
ecc508139a fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting (#4755)
* fix(#4730): decode entity-escaped ampersands before verify-command-paths segment splitting

Planners emit <automated> bodies with the chain operator entity-escaped
(`&amp;&amp;`), and the executing agent reads the decoded (rendered) form.
The grounding probe split the raw text on &&/||/;/newline without decoding
first, so `&amp;&amp;` was cut at its semicolons and a cd-form target
absorbed the trailing `&amp` fragment — an existing directory was reported
missing_dir (blocker), feeding false blockers into the revision loop.

Decode `&amp;` → `&` inside resolveVerifyCommandTarget, after
result.command captures the text verbatim and before any segment splitting
or target resolution, so the escaped and literal forms of the same command
produce identical verdicts. Module-private helper beside splitSegments,
mirroring the sibling src/verify.cts decodeEntityAmps (#3611); the two
gates are separate modules and neither imports the other.

Regression coverage pins escaped/literal verdict parity for: existing dir
+ manifest (ok), missing dir (missing_dir blocker kept), dir without
manifest (no_manifest blocker kept), --prefix form, a literal & inside a
quoted dir name, and a full probePhaseVerifyCommands pass whose reported
command field stays verbatim.

* docs(#4730): use the documented pr:0 placeholder in the changeset fragment

The fragment carried pr: 4730 — the ISSUE number, the exact guess-shape
DEFECT.CHANGESET-PR-FIELD-DRIFT (#3316, #3325) exists to catch: it parses
as a positive integer so local lint passed, but on a real PR run the
drift check would fail it against the actual PR number. CONTRIBUTING.md
documents pr: 0 as the deliberate unresolved placeholder used during
initial commit before the PR number exists (scripts/changeset/new.cjs
accepts 0 for exactly this reason); the gate's fail_invalid_fragment on
an unbackfilled 0 is the designed backfill enforcement, not a defect.

No production or test changes. Backfill pr: with the real PR number
once the PR is created.

* docs(#4730): backfill changeset pr field with the real PR number

pr: 0 → pr: 4755 (the PR carrying this fix), completing the
documented placeholder workflow; the changeset gate's content
validation can now pass.

---------

Co-authored-by: TwistedRiCen <16397953+TwistedRiCen@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 00:36:10 -04:00
0xdhx
092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00
0xdhx
25d1cb916f fix(#4721): give worktree cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged (#4766)
* fix(#4721): give cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged

`worktree cleanup-wave` ran `git merge --no-ff` under the module-wide
DEFAULT_GIT_TIMEOUT_MS (10 s) that is sized for plumbing calls. The merge is
the one call in the wave that runs user hooks, so a repo whose
pre-merge-commit hook is a test-suite gate lost every code-bearing executor
merge. Three things went wrong at once, each fixed here:

1. Budget. The merge now passes an explicit timeout —
   DEFAULT_MERGE_TIMEOUT_MS (10 min), overridable via deps.mergeTimeoutMs.
   Every other git call in the wave keeps the module default; the shared
   constant is untouched, because every other caller is exactly what its
   10 s comment describes.

2. Reason. A merge that does time out blocks on `merge_timed_out`, and its
   stderr names the budget and says the hook may still be running, instead
   of `merge_failed` carrying whatever the hook had printed before git was
   killed — which made a healthy executor branch look broken.

3. Residue. A merge killed during its hook has already staged the merged
   tree into the primary's index but never wrote MERGE_HEAD, so
   `git merge --abort` finds nothing and repoRootStillMidMerge (#2852)
   reads the primary as clean while the executor's whole diff sits staged
   against the old HEAD; a `git commit` from that state squashes the
   executor's history into one parent. After any failed merge the wave now
   reads `git diff --cached --name-only`; anything staged is the merge's
   own (git refuses to start a merge when the index differs from HEAD), so
   it runs `git reset --merge` — restores exactly those paths, keeps
   unrelated unstaged edits — and re-reads. Restored paths are reported as
   WAVE_CLEANUP_WARNING.MERGE_RESIDUE_RESTORED and the wave continues; a
   still-dirty or unreadable index reports MERGE_RESIDUE_LEFT_STAGED and
   halts the remaining entries, the same repo-level carve-out an
   unfinished merge takes.

Tests: five mock-driven rows (budget wiring incl. the deps override, the
timeout classification with restore, the no-reset control for an ordinary
refused merge, an unrestorable residue halting the wave, an unverifiable
index failing closed) plus a real-git row that runs a sleeping
pre-merge-commit hook under a 1 s budget and asserts HEAD unmoved, index
and worktree clean, the executor branch intact — with the same fixture
merging cleanly under the default budget as its negative control. Two
existing #2852 rows gain a handler for the new post-failure index read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* docs(#4721): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* test(#4721): release the real-git fixtures with t.after, not try/finally

The two real-git rows cleaned up their scratch repo in a `finally` block;
this file's own convention for fixture teardown is the test context's
`t.after(() => cleanup(dir))`, and the house PR ruleset flags `finally` in a
test body. Behaviour-neutral.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): gate the residue restore on the timeout, re-apply a merge autostash, and correct the hook census

Three findings from the pre-file adversarial review of the previous commit,
each driven on real git before changing code:

1. A merge git REFUSED ("your local changes … would be overwritten") also
   leaves no MERGE_HEAD — and that refusal is exactly what a pre-existing
   dirty primary index earns. The residue restore read that index as the
   merge's own and `reset --merge`d the operator's staged work away
   (driven: a staged edit to an unrelated file was discarded and reported
   as "restored"). The restore now runs ONLY when the merge timed out; a
   refusal is an immediate exit, never a timeout, so on that path nothing
   is read or reset.

2. `merge.autoStash=true` lets a merge start on a dirty index by parking
   the work in MERGE_AUTOSTASH, which a killed merge never re-applies.
   `git reset --merge` moves that stash into the stash list; the wave now
   runs `git stash pop --index` afterwards (the outcome `merge --abort`
   gives an autostashed merge), and reports
   WAVE_CLEANUP_WARNING.MERGE_AUTOSTASH_UNRESTORED (path null) when the
   pop fails or the autostash state could not be read — the work stays in
   the stash, the index is clean, the wave continues. Because of this the
   reset runs on a timed-out merge even when the index reads clean.

3. The merge is not the only hook-running git call in the module:
   `worktree add` runs post-checkout and every ref update runs
   reference-transaction. It is the only call that runs the commit-family
   hooks, which is what the budget is for. Comments and docs say so now.

Tests: the "ordinary merge_failed" control becomes the regression row for
finding 1 (strict mock — a `diff --cached` or `reset --merge` on a refused
merge throws), plus a mock row for the autostash pop (dirty and clean
index, pop success and failure), and two real-git rows: a refused merge
over pre-existing staged work leaves it byte-identical, and a killed merge
under merge.autoStash restores the executor residue AND puts the
operator's staged work back. The real-git hook now sleeps 4 s against a
1.5 s budget for margin on slow runners. The two #2852 handlers added
earlier are removed — the residue read no longer fires on their path.
Negative control: 4 of the 10 #4721 rows fail on the previous commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): key the residue restore on a killed merge, and re-read the index after a failed autostash pop

Two more findings from the continuation review, both driven:

1. An externally delivered SIGTERM leaves the same staged/no-MERGE_HEAD
   state as the timeout, and the seam reports it as exitCode null + signal
   with timedOut false — so the timeout-only gate skipped the restore on a
   state it was written for. The gate is now "killed": timedOut, or a null
   exit code with a signal. A refused merge still exits with a code and is
   still never touched. The reason stays merge_failed for a signal kill.

2. A failed `git stash pop --index` keeps the stash entry but can leave
   conflict entries (UU) and partially applied paths, after which the next
   merge fails on "you have unmerged files"; the code returned halt:false
   on the strength of the pre-pop recheck. The index is now re-read after a
   failed pop and a dirty result halts the wave as merge_residue_left_staged
   alongside the merge_autostash_unrestored warning.

Also driven and now documented rather than changed: a kill that lands once
MERGE_HEAD exists (inside commit-msg) is the ordinary #2852 abort path —
`git merge --abort` restores the tree and re-applies an autostash itself,
unstaged, as git does for any aborted autostashed merge.

Tests: the pop-failure mock row now asserts the post-pop re-read and gains
a conflict-leftover variant that halts; a signal-kill mock row; a real-git
row with the sleeping hook moved to commit-msg (timed out, no residue
warnings, MERGE_HEAD cleared, primary clean). 414 pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* fix(#4721): key the kill gate on the seam's signal, not on a null exit code

The shell projection seam normalizes a signal death to exitCode 1 and
carries the signal alongside (`_spawnResult`: `result.status ?? 1`), so the
previous `exitCode === null && signal` gate could never fire in production
and the unit row that covered it modelled a shape the seam does not emit
(caught in the round-3 review). The gate is now `timedOut || signal`; a
refused merge exits with a code and no signal. The mock row uses the real
shape, and a mocked spawnSync signal death driven through the compiled seam
reaches `reset --merge` and reports the residue restored.

Also: three comments that still said "at its budget" / "runs user hooks" /
"the index is clean", and the CLI-TOOLS sentence that reserved
`merge_failed` for refusals and conflicts, now name the signal case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9

* chore(#4721): set changeset fragment pr to 4766

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 22:28:09 -04:00
Tom Boucher
b5fc8061b3 fix(#4624): persist orchestrator-worktree worker lifecycle records (#4778)
* fix(#4624): persist orchestrator-worktree worker lifecycle records

* fix(#4624): address review findings on the worker lifecycle protocol

* fix(#4624): require the summary path and surface torn records on status --path

* fix(#4624): distinguish no-record from torn-record, tolerate older shims in the sweep

* docs(#4624): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-15 20:25:24 -04:00
Tom Boucher
779f67cb11 fix(#4378): mint collision-free SEED-YYMMDD-xxx seed ids instead of a shared count (#4754)
* test(#4378): regression tests for collision-free seed ids

* fix(#4378): mint collision-free SEED-YYMMDD-xxx ids, not a shared count

plant-seed derived the next seed id from 'ls .planning/seeds/SEED-*.md | wc -l'.
.planning/seeds/ is shared but each worktree only sees what has merged, so two
workstreams planting before either merges computed the same id and git merged
both files silently.

The id is now the local date plus a 3-char random base36 suffix -- the shape
.planning/quick/ already uses -- computed from knowledge one worktree has alone,
with a same-day regen guard. deriveSeedIdentity learns the new canonical grammar
alongside legacy SEED-NNN (whose parsing never changes), the --enrich parser and
the filename-prefix fallback keep the full new-format id, and the docs that
state the filename shape move to it.

The prefix fallback previously truncated any non-pure-numeric id at
'SEED-<digits>' -- the same one-id-two-answers ambiguity the issue reports,
reproduced one level down.

* fix(#4378): harden seed id generation per adversarial review

- parse-idea: anchor the --enrich extractor to the flag and capture the
  complete id, uppercase-tolerant; a leftmost 'SEED-[0-9]+' truncated an
  uppercase or malformed suffix to its date and enriched an arbitrary
  same-day seed via head -1. Ambiguous and unmatched targets now fail
  closed instead.
- generate-seed-id: tolerate the expected SIGPIPE under pipefail, abort
  loudly when the suffix cannot be drawn (an empty suffix would collapse
  every seed's id to the bare date), and run the same-day regen guard as
  a find existence test (the 'ls <glob>' shape trips the #3409 drift
  guard and degenerates under a stray nullglob).
- deriveSeedIdentity: document the theoretical legacy/new grammar
  ambiguity (6-digit counter + 3-char base36 slug, no frontmatter).
- changeset: state the residual same-day collision bound instead of
  implying zero.

Emitted-Drift-Ack-Growth: plant-seed.md — the counting step became hardened date+random generation with explicit failure modes; growth is the failure handling, not duplicated logic

* fix(#4378): address standards and spec review findings

- tests: move the allow-test-rule marker to its suppression site (the
  file-header placement was inert per CONTRIBUTING site-scoping); add
  width-boundary coverage (5/7-digit dates, 2/4-char suffixes pin the
  documented branch behavior); add a writer-to-reader parity property
  that parses the mint widths out of the shipped workflow so the two
  grammar owners cannot drift; cover uppercase ids end-to-end in the
  reader.
- plant-seed.md: draw/retry restructured as one loop with a loud
  terminal failure; SEED_SUFX renamed SEED_SUFFIX; regen guard drops
  the redundant head -1; the ambiguity error no longer advises an
  impossible 'complete id' for duplicate legacy ids.
- commands.cts: refresh the cmdListSeeds comment still describing
  SEED-NNN as the only canonical form.
- changeset: drop the audit claim the spec axis showed to be an
  overstatement (audit's id display is filename-derived, pre-existing).
- remove a stray untracked artifact file swept into the tree.

* test(#4378): correct boundary expectations to the module's real branch behavior

The first matrix run on the boundary tests caught my hand-trace of the
regex branches, not a module defect: the slug regex's alternation
backtracks to the legacy branch whenever the canonical branch cannot
complete (so the slug is the remainder after the legacy numeric
prefix), and the 7-digit case fails the canonical branch at its 7th
digit before the dash. Pin the verified values.

* docs(#4378): backfill changeset PR number

* fix(#4378): audit seed identity uses the canonical grammar

Review of this PR found the audit surface publishing a fused filename
stem (SEED-081-region for SEED-081-region.md) where list-seeds reports
the canonical id -- one id, two answers across surfaces, the same
ambiguity class the issue files. scanSeeds now derives identity through
the SAME deriveSeedIdentity the list-seeds gate uses (frontmatter id,
then filename id-prefix, then stem), and audit-open acknowledge
resolves --seed-id by scanning for the derived identity, falling back
to the literal stem so callers scripted against pre-canonical output
keep working. Roll-in per the fix-inline rule: found during this PR's
review, same seed-identity seam.

RED probe: pre-fix audit published seed_id SEED-081-region-becomes /
slug 081-region-becomes for a legacy seeded file; post-fix SEED-081 /
region-becomes, matching list-seeds.

* test(#4378): probe timeout uses the class norm after windows-lane timeout

The windows conformance shard failed its bounded sh -c probes at the
local 5000ms bound (cold sh.exe spawn under shard load) while the
identical code passed this PR's two earlier windows waves. The probe
now uses PROBE_TIMEOUT_MS from the class-norm module instead of a local
override, per the helpers/timeouts.cjs convention.

---------

Co-authored-by: sim <sim@local>
2026-09-15 11:41:29 -04:00
Tom Boucher
cbbde6786a fix(#4546): deferred UAT follow-ups no longer block completion and promote to the backlog (#4769)
* test(#4546): failing-first tests for deferred uat follow-ups

* chore(#4546): regenerate derived lists for the deferred-promotion suite

The new verify-work-deferred-promotion suite changes the tests/ tree the
macOS conformance-tier classifier tracks and is a novel file under the
verify prefix in the test-file-count ratchet; both derived lists are
regenerated/registered per their own guards' instructions.

* fix(#4546): deferred uat follow-ups no longer block, and get promoted

Two halves of one disconnect (#1921's deferral design vs the completion
predicate):

- uat-predicate: the item parser now captures the block's reason: line
  alongside result:. A skipped item whose reason carries the
  verify-work writer's 'Deferred follow-up:' template is a deliberate
  deferral -- non-blocking, flagged deferred in the report. Quote-
  tolerant (the writer wraps the value) and case-insensitive. A
  reasonless skip, a non-deferral reason, pending/blocked/issue/
  failed/missing all still block, exactly as before.
- verify-work complete_session: when the Deferred Follow-Ups section is
  non-empty, offer to promote the items to a ROADMAP.md 999.x backlog
  entry reusing next.md's prior_phase_completeness entry shape, with a
  --files-scoped commit. Offer, not auto-mutation -- matches the
  workflow's interactive convention and next.md's own prompt style.

* chore(#4546): refresh compact-content benchmark baseline

verify-work.md grew (the #4546 deferred-follow-up promotion offer in
complete_session); the registered split's token counts moved with it.
Baseline recomputed with the script's own --write.

Emitted-Drift-Ack-Growth: verify-work.md — complete_session gained the deferred-follow-up promotion offer (detection, [P]/[K] choice, the next.md-shaped 999.x entry template, and the --files-scoped ROADMAP.md commit); the growth is the new contract text, not duplication

* fix(#4546): gate/audit agreement and review fixes for deferred follow-ups

- src/uat.cts categorizeItem: a skipped item carrying the deferred
  follow-up template reason now categorizes as 'deferred' (the category
  already existed for deferred-items.md entries) instead of being
  misfiled into the blocked families by keyword match -- the gate/audit
  agreement #3078-CR expects, restored in the permissive direction the
  #1921 design intends. Checked BEFORE the keyword families so '...
  on the release build next version' is not build_needed.
- verify-work.md promotion step: numbering scans for the smallest free
  999.n (count races + non-contiguous history), one backlog entry per
  deferred follow-up, ROADMAP.md-absent behavior specified, idea text
  newline-flattened, Deferred at placeholder harmonized with next.md.
- DEFERRED_REASON_RE: trust assumption documented (authoring contract,
  not a security boundary; non-matching spellings block fail-closed).
- tests: the property now drives evaluateUatPassed and derives
  expectations from the input spec (never restates the matcher),
  includes the no-result-line branch, and pins its seed; the parity
  test drops try/finally for the approved pattern, uses createTempDir,
  sites its allow-test-rule marker at the suppression site, and asserts
  the literal [P]/[K] choices.

* fix(#4546): close promotion-test docstring, drop fc replay-path misuse, refresh baseline

The final matrix run caught three defects in my own review-fix commit:
the parity test file's JSDoc was left unterminated (the whole file
parsed as one comment -- zero tests registered, hence the file-level
'test failed' the runner reported); fast-check's replay-path parameter
was misused as a label (invalid path at replay); and the workflow-text
ambiguity fixes re-drifted the compact-content benchmark baseline.

* docs(#4546): add Fixed changeset for deferred follow-up coverage

* docs(#4546): backfill changeset PR number

* fix(#4546): use the pattern seam escapeRegex for shape-marker matching

The hand-rolled metacharacter escape in the shape-marker assertion
tripped local/no-adhoc-regex-escape, whose named remedy this adopts.

---------

Co-authored-by: sim <sim@local>
2026-09-15 07:30:38 -04:00