Commit Graph

5416 Commits

Author SHA1 Message Date
Tom Boucher
8ed105c8a4 fix(#3684): resume verified-unmarked phases at update_roadmap (#3814)
* test(#3684): failing-first rows for the verified-unmarked resume

* fix(#3684): resume verified-unmarked phases at update_roadmap

* test(#3684): heading-shaped roadmap fixture, plain phase.complete calls

* fix(#3684): fit under the pre-phase-6 margin, fix pins and verify call

* fix(#3684): padding-normalize the marked-complete join, assert STATE idempotency

* test(#3684): anchor fixes, node jq mirror, characterized STATE delta

* chore(#3684): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 10:58:04 -04:00
Tom Boucher
4b84be1da4 fix(#3683): wire gated learnings extraction into completion, align copy path (#3810)
* test(#3683): failing-first rows for learnings source resolution and wiring pins

* fix(#3683): wire gated learnings extraction into completion, align copy path

* test(#3683): register the learnings suite in the docs-guard lane, drop unverified markers

* fix(#3683): close review findings — per-item parsing, readdir guards, docs paths

* fix(#3683): route phase enumeration through the locator seam, fix assertion targets

* fix(#3683): merge execute-phase ack into the 3003 fragment, fix fidelity targets

* fix(#3663): replace the spent execute-phase ack entry with the 3683 re-arm

* chore(#3683): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 09:23:01 -04:00
Tom Boucher
31fcb833ec fix(#3679): gate pr-branch verify on planning-tree deletions (#3803)
* test(#3679): failing-first rows pinning planning preservation and the verify deletion gate

* fix(#3679): gate pr-branch verify on planning-tree deletions

* test(#3679): extract hashes via rev-parse and de-vacuate the pure-code pin

* fix(#3679): close review findings — merged ack, pinned prose gate

* fix(#3679): close two-axis review findings — no-renames gate, structural pin

* chore(#3679): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 03:54:02 -04:00
Tom Boucher
9d65cd5404 fix(#3664): warn when config-dir targets a foreign-agent destination (#3794)
* test(#3664): failing-first rows for the config-dir foreign-agent warning

* fix(#3664): warn when config-dir targets a foreign-agent destination

* test(#3664): fold the foreign-agent warning rows into the install-regressions suite

* fix(#3664): close review findings — kimi-agents kind, gsd.md ownership, e2e gate

* test(#3664): sync boolean call sites and the path-vocab registries

* chore(#3664): backfill changeset pr number

* test(#3663): skip the posix case-pin on win32 where folding is the fix

---------

Co-authored-by: sim <sim@local>
2026-08-24 02:45:28 -04:00
Tom Boucher
314ea20fa4 fix(#3663): fold path casing only on win32 in the w027 active-worktree check (#3793)
* test(#3663): failing-first rows for w027 path-casing normalization

* fix(#3663): fold path casing only on win32 in the w027 active-worktree check

* fix(#3663): close review findings — seam-owned compare key, deterministic case pin

* chore(#3663): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:57:51 -04:00
Tom Boucher
4af59f8dd3 fix(#3662): resolve managed hook node runners at hook-fire time (#3790)
* test(#3662): failing-first suite for runtime-resolving hook runners

* fix(#3662): resolve managed hook node runners at hook-fire time

* fix(#3662): close review findings and document the resolver

* fix(#3662): close adversarial and security review findings

* chore(#3662): backfill changeset pr number

* test(#3662): honor win32 skip return and platform-aware sh runner pin

* test(#3662): pin the bare win32-claude sh-hook shape omitting the bash runner

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:07:06 -04:00
Tom Boucher
cf15682d1c enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules

Stage banners, checkpoints, completion and error panels used fixed-width
runs of box-drawing characters -- a 53-column heavy rule and a 62-column
double-line box. Those runs are ordinary text to a Markdown-rendering
host, so in a narrower pane they wrap and the border comes apart from
the heading it framed.

Shipped content now emits an ATX heading for a titled section and a
blank-line-delimited --- for a break between sections, both of which
adapt to the available width. The same convention is applied to the
three code sites that built these strings at runtime: the UAT
checkpoint renderer, the milestone-close audit report, and the TDD
review checkpoint table.

Removing the box also removes its only reason to exist -- the
east-asian-width padding helpers that kept its right border aligned
(checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE,
CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged.

The convention is specified in gsd-core/references/ui-brand.md and
enforced across all shipped content by tests/responsive-separators.test.cjs.

Refs #3028

* test(#3028): pin the heading form in checkpoint and audit-report assertions

These suites asserted the exact box borders and the 62-column padded
banner interior. With the box gone they assert the ### heading form,
the --- break and the bolded instruction line, and each now carries a
positive assertion that no box character remains -- which is what pins
the fix rather than merely tolerating it.

Language coverage is converted, not dropped: Japanese, Chinese, Korean,
Hindi and Arabic all still assert their rendered banner, and the Arabic
case still asserts the RTL directional isolates the box removal must
not disturb. Adds a case for a banner longer than the old inner width,
which previously produced a ragged border and now has none.

Refs #3028

* chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec

The checkpoint_protocol display spec described the drawn box; it now
describes the heading, the --- break and the bolded action prompt,
which costs 22 bytes (40111 -> 40133, 827 under the cap).

Appended to the existing #3370 fragment rather than filed as a new one:
a growth ack keys on the bare filename and #3370 already declares
execute-plan.md, so a second source naming it would be a hard
duplicate-key error. Same supersede-by-append route #3370 took for the
spent #2652 fragment.

Refs #3028

* docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference

Review found three things.

The rule as first written demanded a blank line above AND below every
---. Only the one above is load-bearing: it is what stops CommonMark
reading the rule as a setext underline for the line above. The one below
is cosmetic, because a thematic break is a leaf block. The rule now says
that, with the reason, instead of asserting a stricter form the content
does not keep.

The zh-CN reference had received the mechanical box-to-heading swap but
none of the prose behind it: it still claimed a 62-character checkpoint
width and still listed --- among forbidden mixed banner styles, so it
contradicted the convention it was translating. It now carries the
separator section, the setext reasoning, the unconditional-vs-per-runtime
rationale and a corrected anti-pattern list, in Chinese.

The user guide asserted that a heading is not a degradation anywhere.
That is an assertion, not a demonstration. It now says what was actually
traded away in a plain terminal, points at the recorded rationale, and
invites the report that would justify the capability flag instead.

Refs #3028

* chore(#3028): backfill changeset PR number

Refs #3028

---------

Co-authored-by: sim <sim@local>
2026-08-23 22:38:12 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Behruz Nassre Esfahani
622f43353c fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode

The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.

Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.

The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.

`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.

Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.

Fixes #3299

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): reconcile the canonical schema table and the stale acceptance test

Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.

1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
   schema reference for the tracer task-type contract, and its Task-types row
   still claimed interactive runs unconditionally present a
   checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
   first round; this one was missed, so the authoritative reference was the
   wrong answer. The row now carries the human_verify_mode-conditional
   behavior and points at the canonical precedence chain.

2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
   tracer ROW EXISTS, never its content, which is why CI could not see the
   drift. It now asserts the row's actual claims and rejects the pre-#3299
   wording. Separately, the #1945 acceptance test named 'interactive run emits
   checkpoint:human-verify after the tracer' kept passing only because its
   substrings still occur in the fallback clause, while its name asserted the
   opposite of shipped behavior. Renamed and narrowed to what #1945 still
   guarantees, plus a new interactiveIsConditional pin so the unconditional
   prose cannot be restored under a passing substring check.

3. plan-md.md's <verify> row now documents that the legacy bare-text form
   (valid, and still shown at :179) does not reach the #3299 auto-continue —
   only a <verify> carrying <automated> does — so the benefit is silently
   unreachable for tracers using that format.

Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions

Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.

MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.

MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:

- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
  'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
  same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
  appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
  presence. Deliberately brittle: CONTEXT.md names that table the canonical
  schema reference, so a wording change must be a conscious edit in both
  places.

Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs

Peer review round 4. Blacklisting did not hold, twice over:

- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
  forms in the auto-continue clause. Round 4 defeated that by appending
  'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
  before expansion' — none of the banned tokens, same restored interruption
  after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
  '<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
  legacy bare form operative. 107/107 passed across tracer, planner and the
  three size-cap suites.

Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.

These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.

Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): strip comments, require uniqueness, pin whole regions

Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):

- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
  wrong copy: every extractor selected the commented decoy. Worked against the
  planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
  checkpoint_protocol before expansion, regardless of the mode-specific rules
  below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
  approval' below the canonical table. The pinned text was untouched, so
  equality held while the shipped meaning inverted.

The shape that holds, applied to every operative surface:
  1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
  2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
     cannot hide behind a correct first one;
  3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
     suffix can override what the pin proves.

Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.

Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.

Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): drop the superseded exact-placeholder planner assertion

Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.

Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.

Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner

Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:

- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
  internally uses the repo's interleaved fence/comment scanner. Instrumenting
  candidate lines as throwaway predicate declarations borrows that scanner with
  no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
  edits silently turned the guards into decoy checks:
    * a forgotten '-->' comments the live rule through to EOF, and the
      balanced-only stripper still saw and accepted the commented rule;
    * a normal fenced documentation example of the rule, plus a whitespace-only
      reformat of the live list item, made the selector choose the example.
  Neither needs intent. A dangling comment is a typo; a fenced example is good
  documentation. Together they reproduce exactly the accidental drift #3299 came
  from — with CI green.

The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.

Verified both ordinary-edit scenarios now fail the suite (each was green before).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): close the operative-selection gaps the maintainer blocked on

trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.

1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
   replaced a matched candidate with an UNINDENTED marker regardless of the
   original line's indentation. A 4-space-indented CommonMark code block is not
   skipped by parsePredicates (it accepts indented declarations by design), so
   stripping the indent PROMOTED an indented decoy to operative — the exact
   inversion of the guard's purpose. The marker now preserves the original
   indent, and a candidate that is itself indented 4+ spaces is never injected.

2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
   pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
   pollute the count. Now filters on a Set of the indexes actually injected on
   this call.

3. RAW FENCE SELECTION (planner). The template test matched the first raw
   ```xml fence after the marker with no fence/comment awareness — the one
   selection in the suite that was not operative-aware — so a commented-out
   decoy template between the marker and the real one would be selected while
   the live template regressed. The opener must now be operative AND the first
   non-blank line after the marker.

4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
   a fenced example containing a ### / <type line truncated the pinned region
   early — a false FAILURE on a legitimate doc edit. End anchors now go through
   the same operative filter as start anchors.

Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.

Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): allow-list operative indentation; pin marker provenance

Review round 9.

BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", "  \t" and "   \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.

MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.

MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.

Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.

KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge

The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.

Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.

Verified: emitted-attribution 175/175 against origin/next@be9329b10.

* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain

Addresses the review on #3390 (B1-B3, M1-M4, minors).

B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.

B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.

B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.

M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.

Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.

Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): correct the size record the 08-22 merge falsified

Review round: one Major, four Minors.

Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.

Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.

Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.

Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.

Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): cite the plan-md example by section, not by a drifting line

Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.

Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.

Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): answer fence-delimiter liveness by insertion, not replacement

Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.

`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.

Measured against the real file rather than argued:

  agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
  262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
  example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
  that question.

It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:

  HEAD as-is                  | anchor 260 | openIdx 262 | ASSERTION PASSES
  +1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***

Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.

The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.

Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.

`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.

Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.

Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.

Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): guard the delimiter class, match the scanner, pin the assignment

Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.

1. The mode-read pin survived the regression it exists to catch.
   `READ` matched the config-get substring only, so rewriting the shipped line
   as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
   nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
   is back with the suite passing. The regex now requires the assignment. A
   lookahead after `end-of-phase` closes the other half: the bare prefix also
   accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
   variable in agents/gsd-executor.md now turns that row red, and did not before.

2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
   It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
   and arbitrary info strings all walk past any list short enough to write down
   — the guard was added precisely because one such regex had already slipped
   through. Now matched against the lines the regex actually selects in the
   document, which cannot go stale and cannot miss a spelling nobody thought of.
   Four such spellings pinned as rows.

3. `isOperativePosition` disagreed with the scanner it delegates to.
   For `<!-- closed --> real content` it stripped the span, found surviving
   content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
   whose trimmed text starts with `<!--`, balanced or not, before it considers
   fences at all. Verified directly against parsePredicates. It now applies the
   scanner's own rule instead of out-reasoning it. Latent for the present caller
   (its anchored ```xml shape cannot match a comment-prefixed line), real in
   general.

Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.

Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.

Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-23 18:43:53 -04:00
Behruz Nassre Esfahani
a44d513566 fix(#3712): confine in-process installs to a sandboxed HOME (#3725)
* fix(#3712): confine in-process installs to a sandboxed HOME

A runtime kind may declare a global `home` override resolved from os.homedir()
rather than from the caller's configDir — codex's skills kind (`home: ".agents"`,
ADR-1239 / #2088) is the only live case. Sandboxing configDir/targetDir does not
contain it, and assertDestWithinConfigHome cannot see the class: that gate
confines a destSubpath to whatever root it is handed, and here the root IS the
escaped home. So an in-process caller that forgot to sandbox HOME wrote to, and
pruned gsd-* entries from, the developer's REAL ~/.agents/skills.

tests/agent-descriptor-parity.install.test.cjs's K1 loop did exactly that: it
iterates every agents-kind runtime (codex included) with a sandboxed targetDir
and an un-sandboxed HOME. Reproduced against a canary home on next @ adb46cdd8 —
71 gsd-* skill dirs deleted, a foreign `cloudflare` skill surviving, suite still
exit 0. It is silent because the runtime's own config home is untouched, so the
manifest keeps reporting a healthy install.

FIVE writers resolve a kind `home` and then destroy under it. Three are reachable
today — installRuntimeArtifacts, uninstallRuntimeArtifacts (install-engine.cts)
and applySurface (surface.cts). Two are descriptor-dependent and guarded against a
future descriptor change rather than a present escape: installOpencodeFamilySkills
(behind the combined-family early return) and installAgentsKindStandalone. Those
two are scoped to the single kind each destroys — passing the whole layout made
codex's unrelated skills override trip a writer that never touches it.

- src/test-home-guard.cts: refuse when a run under a test runner cannot be shown
  to have sandboxed HOME. NODE_TEST_CONTEXT (set by `node --test`) gates it, so
  installs outside a Node test context are untouched; GSD_TEST_MODE is unusable,
  as several candidate files including the offender never set it. Homes are
  compared by FILESYSTEM IDENTITY (st_dev + st_ino), not by pathname:
  path.resolve() resolves neither symlinks nor case, and realpath returns a
  canonical pathname that two routes to one directory can still disagree on (bind
  mounts). Verified on macOS/APFS — HOME=/users/<name> made the strings differ
  while naming the same directory, and the lexical form ALLOWED a write into the
  physical real home. FAILS CLOSED: a pair is "different" only when both identify,
  or one is definitively absent (ENOENT/ENOTDIR) while the other identifies; every
  other errno is "cannot tell" and refuses. Only when neither home identifies is a
  marker consulted, and it carries the sandbox PATH and must equal the home in
  effect — a boolean checked first let an ambient or stale value disarm the guard.
- helpers: promote sandboxHome() out of its two byte-identical private copies,
  which is also what makes them record the sandbox; the three withFakeHome()
  helpers record it too. The marker NAME is duplicated as a bare string rather
  than required from the compiled guard, keeping helpers.cjs's documented
  no-built-lib-at-import-time contract; a test pins the two together.
- agent-descriptor-parity: sandbox HOME across the K1 loop.
- helpers-process-isolation: #3156's canary asserts on <home>/.gsd only, and its
  `--cursor --local` spawn cannot reach `.agents` at all, so an assertion added
  there would pass with all confinement removed. Add a discriminating row — a
  `--codex --global` spawn against a seeded ambient home — which also asserts the
  runtime still declares the override. Its check is a sampled inventory (dir names
  + each SKILL.md), not a tree compare.
- install-write-confinement: predicate rows through the deps seam, covering the
  symlinked HOME, ambient and stale markers, and each sameDirectory branch
  (both-identify, one-absent, neither-identifiable), plus wiring rows that drive
  the REAL entrypoints so deleting a guard call site is red.

Verified: guard fires end-to-end against a real un-sandboxed HOME (exit 1, zero
deletions); the case-variant fail-open reproduced on APFS before the fix and
refuses after; K1 file 29/29 green with skills intact; mutation-tested — each of
the three reachable call sites, lexical-only comparison, and treating an unknown
errno as "absent" each take exactly one row red, with every mutation echoed back;
the process-isolation row negative-controlled by reverting installerEnv to its
pre-#3156 leak (16/0 -> 13/3); a full npm test leaves ~/.agents/skills at 71.

Stated residual: the two descriptor-dependent writers have no wiring test, because
no runtime declares a `home` override on those kinds and neither can be exercised
without inventing a descriptor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3712): add changeset

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3712): let a sandbox nested inside the real home through the guard

All six Windows shards of #3725 failed on legitimately sandboxed
destinations. On Windows os.tmpdir() is %LOCALAPPDATA%\Temp — inside the
user's home — so every sandbox a test creates is a descendant of the real
home, and "does this land inside the real home?" answers yes for the safe
case and the dangerous one alike. POSIX conceals this: /tmp and
/var/folders both sit outside $HOME.

Add the missing conjunct: a destination inside the real home is allowed
only when it also sits beneath a HOME that was sandboxed away from the
passwd home. Both halves are required — dropping the first re-admits a
plain un-sandboxed install, and dropping the second decays into the
"is HOME sandboxed?" check the module rejects, which a layout resolved
before the sandbox walks straight through. Each is mutation-proven by a
row that goes red without it.

Also covers the two fail-closed branches of the new exemption, which
survived mutation to `true` with the suite green, and avoids `<user>` in
a docblock — the prompt-injection scanner reads it as a delimiter tag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): close three writer/rollback gaps found reviewing the whole PR

Cross-AI review of the full PR (not just the round's delta) surfaced
three ways the guard could still be defeated:

- The nested-sandbox exemption trusted the SPELLING of a destination.
  With HOME sandboxed to a directory inside the real home — legitimate on
  Windows — an aliased `.agents` (symlink, junction, subordinate bind
  mount) beneath it redirected an allowed path into the real home. Decide
  containment on the path the write RESOLVES to: walk up to the nearest
  existing ancestor, canonicalize, re-append the tail.

- `migrateLegacyDevPreferencesToSkill` is a SIXTH writer that resolves a
  skills-kind `home` override. It creates rather than prunes, which is
  why it was missed, and `_runLegacyInstallMigrations` runs it before
  `installRuntimeArtifacts`' own assertion. Guarded, scoped to that kind.

- Worst of the three: `bin/install.js` snapshots the resolved skills root
  before installing, and its outer catch rolls back by deleting and
  recreating every snapshotted `gsd-*` directory there. The guard's own
  throw landed in that catch, so refusing an un-sandboxed codex install
  provoked exactly the mutation the guard exists to prevent. Refusals are
  now marked and rethrown without rollback — nothing was written, so
  there is no partial install to undo. Every other error still rolls back.

Also carries the sandbox marker into `installSpawnEnv`, so spawned
installers are not refused on passwd-less CI images, and corrects three
claims that no longer hold: "every writer" (six, and named), the
unconditional "fails CLOSED" (the passwd-less marker branch is a
deliberate weakening, and TOCTOU is out of scope), and the assertion that
Windows os.tmpdir() is always %LOCALAPPDATA%\Temp (Node honors TEMP/TMP).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the guard's two limits instead of overclaiming

Round-2 review found the prose had drifted ahead of the code. Corrected,
with no behavior change:

- The module still said FIVE writers; there are six, and the sixth is
  now named along with why it was missed (it creates rather than prunes)
  and why it carries its own assertion (it runs before the main one).
- The canonicalization docblock listed subordinate bind mounts among the
  aliases it closes. It does not close them: a bind mount is not a link,
  so realpath keeps the mount-point spelling. `sameDirectory` already
  recorded that limit; the new helper now inherits it explicitly rather
  than contradicting it. Closing it needs mount-table introspection.
- "FAILS CLOSED" was unqualified while the passwd-less marker branch is
  a deliberate weakening — with no passwd entry, nothing can contradict a
  marker naming the real home.
- "Refuses BEFORE any write" was too broad: legacy install migrations run
  ahead of the layout-driven ones, which is exactly why the two
  rollbackInstallerMigrations() calls still execute before the rethrow.
  Only the codex skills-root rollback is skipped, and that is the only
  _codexPreConfigRollback() call site — applySurface is never called from
  bin/install.js and uninstall cannot reach it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): sandbox HOME in the opencode-family home-override parity rows

The last two Windows failures, and the same platform asymmetry in a
different disguise. This row drives a skills-kind `home` override on
purpose — precisely what the guard polices — but relied on the override
temp dir happening to sit outside the real home. It does on POSIX
(/tmp, /var/folders); on Windows os.tmpdir() is under %USERPROFILE%, so
the guard correctly refused and only Windows went red.

Declare the sandbox instead of depending on the platform: HOME becomes
the override itself, which is the home the call writes under. This is the
fix the guard's own message prescribes, applied to the test rather than
to the guard.

Both failing Windows shards fail on exactly these two rows and nothing
else; every other shard is green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): make sameDirectory answer NO when it cannot tell

Review Major 1. sameDirectory()'s only caller is the passwd-less marker
branch, which reads a `true` as permission to PROCEED:

    if (marker && sameDirectory(marker, osMod.homedir())) return;

The fallthrough returned `true` whenever neither side identified — two
absent paths, or two stats failing EACCES/EPERM/EIO on a locked-down
host — on the reasoning that "cannot tell" should make the caller refuse.
That reasoning was inverted with respect to this caller: it turned the
passwd-less escape hatch into an unconditional bypass for any marker
value at all, on precisely the hosts the fallback exists to serve. Only
two things now answer yes: one resolved pathname, or two readable
identities that match. Restoring the old fallthrough takes the new row
red.

Also from review:

- Major 2 asked whether st_dev/st_ino discriminate directories on
  Windows, where Node derives them from BY_HANDLE_FILE_INFORMATION. The
  whole guard rests on that primitive, so assert it rather than argue it:
  a row comparing two distinct temp directories, and one directory
  reached by two spellings. It runs on every platform in the matrix, so
  Windows answers the question itself.

- Minor 1: the refusal now names the real home it compared against, not
  just the destination it refused. That is the one fact needed to tell a
  true positive from a false one, and its absence is what made the
  Windows case a CI-log dig rather than a glance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): refuse a HOME that merely spells the real home more widely

Review round: one Blocker, four Minors, a Nit.

N2 (the one with teeth) — a destination's ancestor chain is linear, so
"inside the real home AND inside the effective HOME" admits two
arrangements, not one. The intended `effectiveHome ⊂ realHome` is the
Windows temp shape; `realHome ⊂ effectiveHome` — HOME at /Users, /home,
C:\Users — is not a sandbox at all, it is the real home reached by a
wider spelling, and it was exempting a stale destination pointing
straight at ~/.agents. Third conjunct added; the docblock no longer
claims two conditions suffice. Removing the conjunct reds the new row
and nothing else.

N4 — the migration guard resolved its OWN layout, and without
capabilityRegistry, so a registry-dependent descriptor could make it
vouch for a path the migration does not write: a guard reporting safe
while the unsafe write proceeds. It now guards the destination already
resolved by _resolveDevPreferencesSkillTarget, keyed on
`installRoot !== targetDir` — which is exactly the condition under which
a `home` override was declared, read off that same result.

N1 — CONTEXT.md gains the Test Home Guard Module glossary entry that
contributor-standards.md requires of a new Module. Not CI-enforced, so
green CI was never evidence it was met.

N3 — the docs/INVENTORY.md row was misfiled between install-fs-adapter
and install-model-override-resolver; the table is alphabetical and the
manifest already had it right. Moved, and its text now names six writers
and the third conjunct.

N5 — applySurface's signature docblock was two parameters stale; this PR
added the second of them.

N6 — the duplicated rollbackInstallerMigrations() adjacent to the new
rethrow: two identical consecutive calls, not two phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): guard the sixth writer, and close two false-ALLOW paths

Review round 3 (NEW-1, NEW-2) plus three defects Codex found in the
whole-PR pass, each reproduced before it was fixed.

NEW-1 — migrateLegacyDevPreferencesToSkill called the guard with no
`deps`, so it bound real os/process.env and could not be wiring-tested
the way the other three reachable writers were. It now takes the same
optional `deps: { os?, env? }` tail parameter. The wiring block gains
the missing fourth row, and a fifth pinning the ALLOW half; the test
file's header docblock said "FIVE writers ... the three reachable
today", contradicting the six/four statement this PR already put in
src/test-home-guard.cts, CONTEXT.md, docs/INVENTORY.md and the
changeset. Both directions of the guard's condition now fail a row
when broken — previously neither did.

NEW-2 — derivesFromSandboxedHome's docblock claimed "THREE conditions
are required, and no two of them suffice". False for {2,3}: isInside is
reflexive, so whenever conjunct 1 fires conjunct 2 already returns
false on its own. Reworded as a fast path, which is what it is.

Codex 1 (false ALLOW) — on a host with no readable passwd entry the
marker branch returned as soon as the marker matched the effective
HOME. That attests a caller sandboxed HOME and says nothing about where
an already-resolved destination points, so a layout captured before
sandboxHome() — still naming the real ~/.agents — was waved straight
through: the same stale-layout shape the primary branch refuses by
design. The marker must now identify AND contain every destination.

Codex 2 (false ALLOW) — `installRoot !== targetDir` was the stand-in
for "the skills kind declared a home override". The two are not
equivalent: the inequality is false when the override resolves onto
targetDir itself, which is exactly a configDir of $HOME/.agents. The
guard was skipped and SKILL.md written into the real home under a test
runner. _resolveDevPreferencesSkillTarget now reports hasHomeOverride
off the same resolution instead of inferring it from two paths.

Codex 3 (prose) — the shared refusal message claimed every guarded
writer prunes; the migrate writer only creates. The changeset headline
claimed in-process installer calls can no longer reach the real home,
which is wider than the guard: writeNonClaudeDefaults still writes
~/.gsd/defaults.json through os.homedir(). INVENTORY's and CONTEXT's
fail-closed sentences omitted sameDirectory's pathname-equality
shortcut. All four narrowed to what the code does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): let the sandbox marker follow an overridden HOME

Found by Codex in the whole-PR pass. installSpawnEnv spreads
`overrides` last so an explicit HOME wins — deliberate, and its
docblock tells callers needing per-spawn isolation to pass their own
{ HOME, USERPROFILE }. But the #3712 marker was set before that spread,
so such a caller got HOME=<theirs> and marker=<helper default>. On a
host with no readable passwd entry the guard compares the two and
refuses a legitimately sandboxed spawn — tests/install.test.cjs:7143
and install-shared.cjs's own runInstaller both take that path.

The marker is now derived from the final HOME unless the caller
supplied one explicitly. The contract test asserted HOME after an
override but not the marker, which is why it stayed green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the shipped guard condition, not the deleted one

Review round 4 of #3725. Two artifacts this PR adds still described
`target.installRoot !== targetDir` in the PRESENT tense as the live guard
condition on `migrateLegacyDevPreferencesToSkill`. The shipped condition is
`runtime && target.hasHomeOverride` (src/install-engine.cts:510).

This is not ordinary doc drift. The named condition is the exact false-ALLOW
the previous round closed: a `home` override resolving onto `targetDir` — a
configDir of `$HOME/.agents`, which is where codex's override points — makes
the inequality FALSE while the override is declared, so the guard was skipped.
A maintainer reading CONTEXT.md:290 as authoritative would believe the guard
still skips that case.

  - tests/install-write-confinement.test.cjs — the ALLOW-half row's comment.
    Its "teeth" rationale is unchanged and still correct as written.
  - CONTEXT.md:290 — the Test Home Guard Module glossary entry, a documented
    PR gate. Now states the condition and names the inequality only as what it
    is NOT, with the reason.

The three surviving mentions of the inequality are all past-tense or negated
(src/install-engine.cts:448, :508 and the sibling test comment at :3698) and
are correct as they stand.

Verified: `npm run lint:ci` exit 0; full `npm test` 31327 tests / 31312 pass /
0 fail / 14 skipped, run with TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): canonicalization fails closed, matching identify's errno split

Codex full-PR review of #3725, run against the round-4 head.

`resolveThroughLinks` caught EVERY realpathSync error and fell back to
`path.resolve(dest)` — the lexical spelling. That inverts the function's own
purpose. An aliased `<sandbox>/.agents` that cannot be canonicalized keeps its
sandbox spelling, satisfies the nested-sandbox exemption at :227, and the write
is ALLOWED into the real home — the exact escape this walk exists to close. The
module documents that it fails CLOSED with ONE named exception (the marker
branch); this was a second, unnamed one.

Split by errno, and deliberately by the SAME split `identify` already draws
rather than a second policy in one module — both answer "does this path exist
as named?", so they must not disagree:

  ENOENT / ENOTDIR -> walk up. The ordinary case: a fresh install resolves a
    destination nothing has created yet, so realpath fails on the leaf and on
    every not-yet-created ancestor. Refusing here rejects every install.
  anything else (EACCES, EPERM, ELOOP, EIO) -> refuse. The component exists but
    cannot be resolved, so the guard cannot tell where the write lands.

Three rows in the predicate block, beside the other aliasing rows:
  - a symlink CYCLE in the destination path (ELOOP)   -> REFUSE
  - a destination that does not exist yet (ENOENT)    -> ALLOW
  - a component behind a regular file (ENOTDIR)       -> ALLOW

Teeth checked against the artifact the test loads, not the source: reverting
the condition to the swallow-everything shape in the compiled
test-home-guard.cjs turns row 1 — and only row 1 — red. The ENOTDIR row caught
a stale build during development, which is the point of asserting on the
compiled file.

CONTEXT.md and the resolveThroughLinks docblock both record the new behaviour,
so this does not repeat the prose-vs-code drift the round-4 finding was about.

The changeset's existing scope sentence now bounds "six writers" to the
`installRuntimeArtifacts` call tree and names `cmdGenerateDevPreferences` —
which resolves the same codex `home` override through `getGlobalSkillsBase` and
writes SKILL.md beneath it unguarded. It has no in-process caller today (its
only direct require-and-call is a spawnSync with HOME sandboxed), so it is
latent rather than live, and whether it belongs in this PR is raised with the
maintainer rather than decided here.

Verified: `npm run lint:ci` exit 0; full `npm test` 31330 tests / 31315 pass /
0 fail / 14 skipped, TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 18:27:02 -04:00
Tom Boucher
1178c5f995 test(#3108): make the overlay ENOENT tolerance and hooks/dist readiness check honest (#3772)
* test(#3108): failing-first suite for the overlay vanish-retry and hooks/dist staleness

RED by construction, and deliberately narrower than the issue.

#3108 reports a bare `ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'`
and attributes it to hooks/dist never having been built. That mechanism cannot
produce that error: buildOverlayRepo enumerates with readdirSync and links what
it enumerated, so a directory that never existed yields no names and no link is
ever attempted. The error requires the file to have existed at readdir and
vanished before the link -- which is the atomic-replace race placeVanishableLeaf
was written for in #3285, three days AFTER this issue was filed.

What is still genuinely broken, and what these tests bind to:

placeVanishableLeaf's retry is unguarded. On ENOENT it re-checks existsSync and
calls attempt() once more, bare. A second atomic replace inside that window
throws an unhandled ENOENT of exactly the reported shape. Rows 1/2/3/5/6/7 fence
the surrounding contract -- most already pass, which is the point: they are what
stops the fix from widening into "swallow every error". Row 6 in particular
covers a non-ENOENT on the RETRY, the exact path the new code will live on.

ensureHooksDist's staleness predicate is extension-blind. It rebuilds only when
hooks/dist is absent or holds zero .js files, while build-hooks.js also ships
.sh -- including gsd-session-state.sh, the very file in the report. A dist with
.js present and every .sh missing reads as populated and the rebuild is skipped.
Rows 10 and 11 are mirrored on purpose: asserting only the .sh direction would
permit swapping one extension heuristic for another, so both directions force
the predicate to be about the expected set (HOOKS_TO_COPY, which build-hooks.js
already exports) rather than about counting an extension.

Rows 12 and 13 are cost guards. ensureHooksDist runs per suite and its rebuild
is a real subprocess, so a predicate that over-triggers turns a correctness fix
into a throughput regression nobody attributes to it; and hooks/dist legitimately
carries files the expected list does not name, so an exact-set match would
rebuild forever.

The warning-text row is t.skip()'d rather than faked: the only ways to assert it
were a source-grep (banned by local/no-source-grep) or a full overlay build, and
a test that cannot be written honestly is better skipped visibly than written
vacuously.

Filename note: this started as fix-3108-*.test.cjs and tripped
lint-regression-test-names, which bans new fix/bug/issue-NNNN files, then as
install-overlay-helpers.test.cjs and tripped lint-test-file-count, whose `install`
bucket is already at its limit. Module-named under the overlay bucket satisfies
both. The allowlists were left untouched -- both are empty, so nothing here is
grandfathered and adding an entry would have been the wrong instinct.

* fix(#3108): guard the overlay retry and make the hooks/dist check see .sh

Two holes, both reachable from the failure #3108 reports, neither of them the
cause it names.

placeVanishableLeaf's retry was bare. On ENOENT it re-checked existsSync and
called attempt() once more with no catch, so a second atomic replace landing
inside that window threw an unhandled ENOENT -- exactly the reported
`ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'`. Two vanishes inside
the window means the same thing one does: the path is going away and is not part
of the snapshot. It now returns false and skips the leaf, reaching the conclusion
the single-vanish case already reached.

Still ONE retry. No loop, no backoff, no sleep -- the existing comment argues
that a timing-based wait here would be the flake rather than the fix, and that
reasoning did not change. Non-ENOENT still propagates from either attempt, which
is the invariant a careless widening would eat; the suite pins it on the retry
path specifically, because a fix that guarded only the first attempt would look
right and be wrong.

ensureHooksDist could not see the file class that caused the report. It rebuilt
only when hooks/dist was absent or held zero .js files, while build-hooks.js also
ships .sh -- including gsd-session-state.sh itself. A dist with .js present and
every .sh missing read as populated and the rebuild was skipped. The predicate is
now membership against build-hooks.js's own exported HOOKS_TO_COPY, so it asks
"is everything expected present" instead of counting an extension, and it cannot
be blind to a file class again.

It is extracted as isHooksDistStale(dir) with ensureHooksDist calling it, so
there is one predicate rather than two that can drift. Extra unexpected entries
are explicitly not stale -- hooks/dist legitimately accumulates subdirectory and
hooks/lib output, and an exact-set match would rebuild forever. One readdirSync
into a Set, no per-entry existsSync, no stat: it runs per suite and its rebuild
is a real subprocess, so an over-triggering predicate would turn this into a
throughput regression nobody would attribute to it.

The skipped-leaf warning now names `npm run build:hooks`. It already named the
cause; a reader still had to know what produces that directory.

Deliberately NOT done: nothing here makes an absent hooks/dist fail. Absence is a
legitimate package shape that bin/install.js:11191 treats as "nothing to verify",
and six of the seven install suites never read the directory at all.

* fix(#3108): count hooks/dist subdirectories, and never throw out of the predicate

Two gaps found reviewing the predicate I had just written.

It ignored HOOKS_SUBDIRS_TO_COPY. That is ["lib"], and hooks/dist/lib carries
gsd-graphify-rebuild.sh, so a dist with all 27 top-level files but no lib/ read
as populated. That is precisely the blindness the .js-count heuristic had, one
level down: a whole file class invisible to the check. Fixing the extension case
and leaving the subdirectory case would have been half a fix, and the half left
behind is the one nobody would look at again.

Subdir names are bare (no slashes), so they slot into the same top-level readdir
Set — no second readdirSync, no stat. Whether lib is really a directory is not
checked; that would cost a stat per entry and buys nothing, since the build owns
that.

It could also throw. existsSync passing does not make readdirSync safe: the path
may be a regular file (ENOTDIR), unreadable (EACCES), or retired in the race
between the two calls. This helper runs in every install suite's before(), so an
unhandled throw there fails a suite on a condition it cannot act on. Unreadable
is indistinguishable from unusable for this question, and rebuilding is
idempotent, so it now reports stale instead.

Both are the same shape as the original bug and the same shape as each other: a
guard that answers "is this ready" must not have a blind spot or a hard edge,
because every caller treats a false negative as "carry on".

Two existing tests asserted a "complete" dist without lib and had to be corrected
to keep meaning what their names claim, rather than being left passing against a
definition of complete that no longer holds.

* test(#3108): close the review findings, including a half-closed subdir check

An isolated correctness reviewer found no blockers and four real gaps.

The wiring was untested. Every Group-2 test exercised the pure predicate; none
called ensureHooksDist. So restoring the old inline .js-count check INSIDE
ensureHooksDist -- keeping isHooksDistStale exported and correct -- left the
whole suite green, and that wiring is the actual #3108 defect. Two tests now
drive ensureHooksDist itself through the process seam: build invoked exactly once
when stale, never when fresh. The second is the one a permissive revert fails.

Reaching that seam meant requiring process-seam as a module object rather than
destructuring runNode, so a test can replace it in place. That is a testability
affordance in a test helper, not a production change, and it is commented as such
so it does not read as an accident later.

The subdir check was only half closed, and the half left open was the important
one. It required `lib` to be PRESENT in the top-level readdir, never looked
inside -- so an EMPTY dist/lib, missing gsd-graphify-rebuild.sh, still read as
populated. That is precisely the missing-file-class case the subdir check was
added to catch, which made the fix a gesture at the problem rather than a fix.
Each subdir entry must now be a readable, NON-EMPTY directory. A stray regular
file named `lib` throws ENOTDIR into the same try/catch and reads stale too.

Cost stayed honest: one extra readdirSync total (there is exactly one subdir
entry), no stat, no per-expected-file syscall. Probed against the real
hooks/dist -- still reports fresh, so no suite gains a rebuild.

Two nits, both real: the error-code sweep re-tested EACCES already covered
standalone, and the property ignored presentAtFinalAttempt whenever vanishCount
was not 1, making roughly half the 200 runs duplicates. The flag now varies
meaningfully across the whole range and the assertions depend on it.

One reviewer finding was already stale: the subdir and ENOTDIR work was
uncommitted when the reviewer snapshotted the tree, and had landed in a67aefb9c
before the report arrived. Verified rather than assumed.

* fix(#3108): stop the vanish tolerance from swallowing a dest-side ENOENT

A defect this PR introduced, caught by an isolated security reviewer.

linkSync(src, dest) throws ENOENT for the DESTINATION path too, not only for a
vanished source. The widened retry caught that, saw the source still present,
retried, got the same dest-side ENOENT, and returned false -- recording the leaf
as "vanished mid-walk" and printing a warning that tells the reader to run
`npm run build:hooks`. A remedy with nothing to do with the actual cause, an
overlay quietly short a file, and the install under test proceeding against an
incomplete tree.

It also falsified the function's own documented invariant, which says in as many
words: "Returns false only when the path left the source tree entirely." Widening
the tolerance without re-reading the sentence above it is how that happens.

The retry now re-checks existsSync(srcPath) before tolerating: source still
present means the ENOENT was about something else and it propagates untouched.
Chose the existsSync re-check over comparing retryErr.path to srcPath -- err.path
normalization is not guaranteed across platforms, and a path-equality test is a
subtler thing to get wrong later.

The FIRST catch was probed and is already correct: for a dest-side ENOENT the
source is present, so it falls through to the retry rather than returning false.
Left unchanged rather than "fixed" symmetrically.

Two regression pins, deliberately opposed: a dest-side ENOENT with the source
present must THROW, and a genuinely absent source must still return false. The
second exists because the obvious over-correction -- always rethrow on the retry
-- passes the first and silently undoes what this PR set out to fix.

Also closed the skipped placeholder. It claimed no non-flaky seam existed for
asserting the warning text; the reviewer pointed out an injectable `warn` param
is trivial, and they were right. buildOverlayRepo now takes opts.warn defaulting
to console.warn (byte-identical for every existing caller) and the skip is
replaced by real tests: fires with the remedy named on a skipped leaf, silent on
a clean walk. "No seam exists" was a design choice presented as a constraint.

Recorded the sequential-only constraint at the two sites that monkeypatch fs
process-wide: adding { concurrency: true } to this file would cross-contaminate
every other suite in the process. Better written down than rediscovered.

Known limit, disclosed rather than fixed here: a legitimately dropped leaf can
still pass vacuously downstream -- agent-fragments-emission asserts a negative
over filesContaining, and mcp-catalog-parity has only an anti-vacuity floor of
one. That is a pre-existing property of those suites and the tolerance #3285
already chose; this change narrows which drops are possible rather than adding
the completeness assertion those suites lack.

* fix(#3108): discriminate ENOENT by the dest parent, not by re-checking the source

The previous commit's dest-side guard was wrong, and the remote run said so:

  "a leaf that vanishes again during the retry is skipped, not a bare ENOENT"
  Got unwanted exception. Actual message: "ENOENT: no such file or directory"

That test was right and the guard was wrong. It rethrew when existsSync(srcPath)
was still true, on the theory that a present source means the ENOENT was about
the destination. But in the genuine race the source is being atomically REPLACED,
so it is legitimately present again at the re-check while the ENOENT was entirely
source-side. The gate therefore threw on precisely the race #3285 exists to
tolerate -- trading one misclassification for a worse one, since the old bug was
a bare crash and the new one broke the working tolerance.

The security reviewer's alternative discriminator does not work either, and a
probe settles it. Node populates BOTH `path` and `dest` on a link ENOENT, and
`err.path` is the SOURCE in both directions:

  linkSync(existingSrc, missingDir/a.txt) -> ENOENT path=<source> dest=<dest>
  linkSync(missingSrc,  validDest)        -> ENOENT path=<source> dest=<dest>

So the error object cannot tell you which side failed.

What CAN: the dest parent. buildOverlayRepo builds its own dest tree --
place() mkdirSync's recursively into a private mkdtempSync root no other process
touches -- so a missing dest parent is always a bug (Windows MAX_PATH, a
concurrent cleanup, a bad dest), never the replace race. A present dest parent
means the ENOENT was about the source, which is the case we tolerate.

placeVanishableLeaf therefore takes an optional destPath and uses the dest
parent as the sole discriminator when it has one; with no destPath it behaves
exactly as before. linkOrCopyFile and the copy-mode call site both pass it,
because those are the two places that actually know the destination.

The doc comment now records BOTH failed discriminators and why each fails --
existsSync because the source is legitimately replaced mid-race, err.path
because it names the source either way. Those are the two things a future reader
reaches for first, and both look correct until they are not.

The dest-side regression pin was rewritten to drive the real mechanism: a real
temp source and a dest whose parent does not exist, through linkOrCopyFile.
Previously it forced a throw through a present source, which is what encoded the
wrong theory into a test and made it look verified.

---------

Co-authored-by: sim <sim@local>
2026-08-22 23:04:44 -04:00
Tom Boucher
004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00
Tom Boucher
3fd03bec4c fix(#3760): refuse a legacy-key migration into a non-object config section (#3767)
* test(#3760): failing-first regression for non-object config section

Locks the contract from the issue's Expected section before any fix exists:
a legacy-key section holding a string, number, boolean or array must be
preserved verbatim, reported, and never persisted in an expanded form.

Covers both blocks the issue names (branching_strategy -> git.*, sub_repos ->
planning.*), the migrateOnDisk multiRepo branch that shares the shape, and the
loader write paths that are what actually reach the user's config.json.
Includes the negative-space cases that must keep hoisting ({} , null, absent
section, canonical-nested-wins) and two fast-check properties.

Refs #3760

* fix(#3760): refuse a legacy-key migration into a non-object config section

normalizeLegacyKeys hoisted a legacy top-level key into its canonical nested
section by spreading `result[section] ?? {}`. `??` guards only null and
undefined, so a section holding a string was enumerated by index —
`{...'main'}` is `{0:'m',1:'a',2:'i',3:'n'}` — while a number or boolean
spread to `{}` and the value vanished. Because a fired block always pushed a
Normalization, and every caller treats a non-empty normalizations array as
'config is dirty', that shape was written back to .planning/config.json and
the original value became unrecoverable.

Both blocks the issue names are fixed via one shared hoistLegacyKey helper,
plus the two further sites that share the shape and are reachable from the
same input: migrateOnDisk's multiRepo branch, and the loader's two
`if (!planning) planning = {}` guards, where a non-empty string is truthy and
the following assignment threw a strict-mode TypeError that the enclosing
catch swallowed — discarding the user's entire config.

A present non-object section now blocks its own migration. The section, the
legacy key, and the file are left byte-identical; no Normalization is pushed,
so nothing marks the config dirty; the refusal is reported in-band as
`skipped[]` and out-of-band through the ADR-1411 warnUnusableInput seam
(new frozen reason config_section_not_object). null and undefined keep their
long-standing 'absent' meaning and still create the section.

This is the nested-section analog of the ADR-227 shape check _readConfigFile
already performs on the top-level document: valid JSON is not a config object.
isConfigSection is exported and shared by both modules rather than copied.

Fixes #3760

* fix(#3760): keep the multiRepo marker when planning cannot receive it

Follow-up from the isolated adversarial review, and the same defect class as
the two blocks the issue names — in the block it did not name.

normalizeLegacyKeys block 3 deleted `multiRepo` and pushed a Normalization
before anything consulted the planning section, deferring 'can this section
receive sub_repos?' to the caller that runs filesystem detection. By then the
marker was already gone and the config was already dirty, so with
{"multiRepo":true,"planning":"docs"} the loader wrote the file back with
multiRepo removed, the sub_repos injection silently no-opped against the
string, and no diagnostic was emitted at all. migrateOnDisk warned for the
same input; the ~30-caller loadConfig path did not.

Section validity is knowable from the parsed config alone — detection is only
needed for the VALUE, not for whether the destination can hold it. The refusal
moves into block 3: the marker is kept, no Normalization is pushed, and a
skipped entry is recorded, so all three callers inherit the preservation and
the diagnostic together. The caller-side guards drop to pure narrowing.

Also from review: skipped[] now reports sectionType ('string' | 'number' |
'boolean' | 'array') instead of sectionValue. migrateOnDisk's report is printed
verbatim by `migrate-config`, and this module already masks config values on
the set/unset output path; the type is the whole diagnostic and the value is
still in the file.

And `migrate-config --raw` no longer answers a refused migration with 'No
legacy keys found — config is already canonical.' Legacy keys WERE found and
declined, and the decline is the one thing only the user can fix by hand.

Refs #3760

* fix(#3760): keep configuration.cjs dependency-free; emit from its callers

The remote matrix caught a regression my own change introduced: adding
`require('./unusable-input.cjs')` to configuration.cts broke the #3571
install-layout contract. `configuration.cjs` must load from a layout holding
only itself plus bin/shared/*.manifest.json — the installer does not co-locate
arbitrary siblings — so the new require failed at load time:

  Cannot find module './unusable-input.cjs'
  Require stack:
  - /tmp/gsd-3571-.../.codex/gsd-core/bin/lib/configuration.cjs

pinned by 'co-located bin/shared manifests let configuration.cjs load without
sdk/shared' in tests/install.test.cjs (3 failures).

The contract is deliberate and the test is right, so the module goes back to
zero sibling requires and the out-of-band diagnostic moves to the callers that
already carry a dependency budget and hold the resolved path: cmdMigrateConfig
(config.cts) and loadConfigResolved (config-loader.cts). normalizeLegacyKeys
keeps reporting refusals in-band via skipped[], which is what lets it be pure
and dependency-free at the same time.

The emission-count and dedup assertions move to tests/config-loader.test.cjs,
where the diagnostic now originates. A new assertion pins the inverse for the
module itself — migrateOnDisk must emit ZERO diagnostics while still reporting
skipped[] — so regrowing a sibling require fails a unit test instead of only
the install suite. CONTEXT.md records why the emitter is the caller.

Refs #3760

* chore(#3760): backfill changeset pr number to 3767

---------

Co-authored-by: sim <sim@local>
2026-08-22 18:57:19 -04:00
Tom Boucher
b6977d9d11 fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps (#3766)
* test(#3762): failing-first roster gate for docs/INVENTORY.md rows

Anchors the human half of the inventory-drift rule: every entry in
docs/INVENTORY-MANIFEST.json must have a hand-written row in
docs/INVENTORY.md. Expected RED on this commit -- next carries 32
unrostered surfaces, which is the defect the gate exists to catch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps

docs/INVENTORY.md calls itself the authoritative roster of every shipped
GSD surface, and CLAUDE.md's inventory-drift rule requires both a roster
row and a manifest regen. Only the manifest half was anchored, so a PR
could ship a surface, regenerate the manifest, omit the row, and stay
green -- as PR #3758 did with gsd-core/references/planner-coupling.md.

Adds the roster half to tests/inventory-manifest-sync.test.cjs, backed by
a pure matcher in tests/helpers/inventory-roster.cjs, and backfills the 32
surfaces already missing rows on next.

Fixes #3762

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3762): harden roster matcher against fenced blocks and trim exports

Skips fenced code regions when splitting level-2 sections so a documented
'## ' example inside a fence cannot truncate a family section (false red)
or contribute a phantom row (false pass); makes the heading pattern linear
rather than a backtracking lazy match; narrows the module surface to the
three names the gate consumes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3762): close two false-RED gaps found in orthogonal review

Indented headings: CommonMark permits an ATX heading to carry 1-3 leading
spaces, and a ^##-anchored pattern read such a document as having no family
sections at all -- reporting all six missing, a structural red for zero real
drift. Reproduced, then fixed and pinned.

Link-wrapped cells: a row written as [`x.md`](../x.md) was not recognized,
though docs/INVENTORY.md already uses that form elsewhere. Unwrapping now
peels a whole-cell link and a whole-cell code span, and only layers that
wrap the cell entirely -- a file mentioned mid-prose is still not a row, so
the false red is not traded for a false pass.

Adds a second fast-check property over the commands source-link rule, the
one family whose matching rule differs from the other five.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3762): boundary trio for the fence-marker length limit

RULESET.TESTS.boundary-coverage — the matcher's only numeric limit is the
fence marker's {3,}. Exercises 2 (inline markup, swallows nothing), 3, and
4 characters, for both backtick and tilde delimiters.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: remove stray pwned_cmdsub injection-test canary from the repo root

A zero-byte file committed by 0e6fa2e2c (#3124) while remediating the
#3118 command-substitution injection — the marker a test wrote into cwd to
prove a substitution had NOT executed, left behind when the run ended.
Nothing in the tree references it (verified by Grep across the repo),
package.json's files array excludes root-level files so it never shipped,
and lint-removed-but-needed confirms no surviving reference.

Found while auditing the repo root for #3762; fixed inline rather than
deferred, per CLAUDE.md's no-deferrals rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: backfill changeset pr number to 3766

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 18:24:25 -04:00
Tom Boucher
9ade7926ca test(#3761): replace vacuous wave/files_modified gate with an anchored one (#3764)
The test 'planner has validation step or quality gate for wave/files_modified
consistency' asserted a six-token disjunction over whole-file substrings, five
of which occur ZERO times in agents/gsd-planner.md (validate_waves,
wave_validation, same wave — the last killing the entire quality-gate arm,
since that file has no quality_gate block). The sixth matched one incidental
pseudocode comment, so deleting the wave-ordering gate outright left the suite
green.

Per RULESET.TESTS.delete-bad-tests (CONTEXT.md:595) the vacuous test is DELETED
and replaced, not patched. The replacement anchors on the normative **Rule:**
sentence inside <step name="assign_waves"> — scoped to that step rather than to
the whole document, and requiring all four semantic clauses (wave scope,
files_modified subject, prohibition, overlap predicate) to hold within a single
sentence, so co-occurring words scattered across a paragraph do not pass.

agents/gsd-planner.md is unchanged: the anchor already exists there and survives
PR #3758's edits to the same region.

Teeth proven on the remote runner: a control commit paired these tests with a
deliberately gutted assign_waves step that retained the incidental phrase. The
new anchored test failed while the sibling test using the old predicate's
surviving arm stayed green — same input, opposite verdicts.

Fixes #3761

Co-authored-by: sim <sim@local>
2026-08-22 17:52:23 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
738f42f4fd feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs (#3755)
* test(#2398): failing-first suite for the CYCLE_SUMMARY consensus gate

Binds the gate before it exists, so the suite is RED against next.

The load-bearing rows are the two the closed PR #2417 did not have. The B2
regression row asserts a judgment-class lone HIGH counts WITHOUT corroboration
when its raiser is unmarked — if anyone re-couples that class to corroboration,
more reviewers again produce a weaker gate than one, which is what closed #2417.
The parity row asserts every marker literal the gate names is one
review-lane-runner actually emits, so the gate cannot key on a signal nothing
produces; a mutation row and a seeded fast-check property prove that guard runs
its failure branch rather than only reading a correct tree.

Also pinned: gate position before Counting rules, the untouched CYCLE_SUMMARY
line shape the orchestrator greps, fence balance, the single-reviewer no-op,
classification by what a claim asserts rather than by citation presence, the
all-marked fail-open, current_actionable staying out of scope, and the
leading-marker requirement that stops a review which merely quotes a marker
from suppressing its own findings.

* feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs

With review.reviewer_instances running several reviewer identities off one
adapter, any single instance's fabricated HIGH could force a full replan cycle
on its own. Across ~9 real cycles on two projects each of four instances
fabricated at least once, and each was also the most accurate reviewer in some
other cycle, so dropping to fewer reviewers trades away real signal.

The gate engages only when 2+ reviewers actually ran, and weighs a lone HIGH by
what the claim asserts rather than by whether anyone agreed with it. An
existence claim -- a symbol, file, flag, commit or ID exists, is absent, or says
something specific -- counts only if source-grounding confirms it or another
reviewer raised the same concern. A judgment claim -- a design or correctness
property -- counts unless that reviewer's own section opens with an
evidence-quality discount marker the review lane already stamps
([reviewed-without-source-citations] #3194, [reviewed-without-repo-access]
#2176, or a diff-only lane).

That split is what resolves B2, the finding that closed PR #2417. B2 showed the
approved wording made more reviewers produce a WEAKER gate than one: condition
(a) pointed at the source-grounding pass, which verifies every symbol THE PLAN
cites and never takes reviewer claims as input, so a genuine architectural HIGH
that one reviewer caught and another missed was neither groundable nor
corroborated and stopped gating. Judgment-class findings are therefore exempt
from corroboration entirely -- reviewers catch materially different classes of
issue, and demanding two of them independently raise the same architectural
concern suppresses exactly what a multi-reviewer setup exists to surface.

Guards on the gate itself: an all-marked cycle disengages it, so a cycle in
which nothing was verified can never be counted as converged; the marker must
OPEN a reviewer's section, so a review that merely quotes a marker does not
suppress its own findings; a suppressed HIGH stays listed and tagged rather
than dropped; current_actionable is untouched; and a single-reviewer run is
unchanged.

No new command, config key, or dependency -- the gate reads signals that
already exist. The CYCLE_SUMMARY line shape the orchestrator greps is
unchanged; only the integer it computes moves, and only for 2+ reviewers.

Known limit, inherited rather than introduced: SOURCE_CITATION_RE checks
citation presence, not resolution, which src/review-lane-runner.cts records as
a deliberate #3194 scope boundary. A fabricated but plausible file:line still
gates.

Scope revised and re-approved on the issue before any code was written.

* test(#2398): make marker parity behavioral, and stop overclaiming the gate

Review found the parity tests were vacuous: they asserted a marker STRING
appeared in review-lane-runner.cjs's source text, never requiring the module or
calling the stampers, so they would pass even if stampUngroundedReview were
broken or never invoked. They now invoke the real exported functions and assert
what those functions PRODUCE — that an uncited review gains a leading marker
blockquote, that a review carrying a file:line does not, that a self-reported
blind review is stamped, and that stamping is idempotent. Removing the source
read also removes an incidental no-source-grep evasion via a parameterized path.

Review also found the changeset headline false for the class it matters most
in. The discount markers detect 'cited nothing' and 'had no repo access'; they
cannot detect 'drew a wrong conclusion from a real citation', so a judgment-class
finding invented by an evidence-bearing reviewer still counts alone. That is the
deliberate side of the tradeoff jags-faith named when closing #2417 — the
alternative is requiring corroboration for design findings, which is B2 — but
the changeset claimed lone hallucinations no longer force a cycle, full stop.
Corrected there, and stated plainly in docs/COMMANDS.md and the design record.

Also dropped the reviewer-instances.md entry from the emitted-drift ack: the
growth ratchet's currentSizes() scans only gsd-core/workflows/ and agents/
(tests/helpers/emitted-runtime.cjs:916-929), so references/ is outside it and
that entry acknowledged a delta the gate cannot see.

* chore(#2398): backfill changeset pr number to 3755

---------

Co-authored-by: sim <sim@local>
2026-08-22 10:53:14 -04:00
Behruz Nassre Esfahani
444069d601 fix(#3613): copy component dirs into the plugin-validate fixture (#3627)
* fix(#3613): copy component dirs into the plugin-validate fixture

C2 builds its synthetic plugin root by symlinking commands/, hooks/ and
skills/ into a temp dir. `claude plugin validate` (>=2.1.233) reads
component directories without following symlinks and warns on each one,
and `--strict` promotes a warning to a non-zero exit — so the assertion
failed on how the fixture was built, not on the manifest under test.
Copy the directories with fs.cpSync instead. The validated tree still
holds only plugin.json and the three component directories, so nothing
else from the repo root is placed where the validator can read it, and
the #2665 CLI-config isolation is untouched.

The construction helper is self-cleaning: its callers' try/finally only
begins once it returns, so a throw partway through would strand a
half-built root on disk. The pre-refactor code ran these same steps inside
C2's own try, and that teardown guarantee is preserved rather than
narrowed. Its cleanup is best-effort so a teardown error cannot replace
the real construction error.

Add C3, an unconditional tripwire asserting the fixture exposes real,
non-symlinked component directories at any depth. C2 never runs in CI —
no job under .github/workflows/ installs the `claude` binary — so a revert
to symlinks would pass every CI lane and surface only as a red suite on
contributor machines. C3 is not a substitute for C2's end-to-end check and
cannot be shown red against this base, since the helper it calls arrives
with it; C2 is the failing-first artifact.

C3 enforces a symlink-free fixture throughout, deliberately stricter than
the CLI's own boundary. Measured on 2.1.234, `plugin validate --strict`
exits 1 for a symlinked component dir and for a symlink one level inside
skills/, and 0 for two levels in or for symlinks under commands/ or
hooks/. Encoding that external, undocumented line would be more fragile
than a superset costing one walk over ~176 files. The walk uses an
explicit stack rather than readdirSync's `recursive: true`, which follows
directory symlinks — a link pointing outward would otherwise traverse an
unrelated tree, or a cycle, before the assertion ran.

Correct the Section C docstring, which claimed C2 provides defence-in-depth
coverage it cannot provide in CI. Whether to provision the CLI in a CI job
is a maintainer call and is left open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): skip hooks/dist and .dist-staging when building the fixture

The recursive copy added for #3613 races the hook builders. build-hooks.js
writes atomically through a per-PID hooks/.dist-staging-<pid> and removes
it when done, and nine test files invoke that script from their before()
hooks, so a recursive walk can enumerate a staging directory and then
lstat it after the owning process deleted it — the ENOENT that #3656 just
fixed in the cold-tree fixture.

Copy entry-by-entry and skip by NAME BEFORE anything stats it, reusing
shouldCopyHookEntry from tests/helpers/cold-runtime-lib-fixture.cjs rather
than re-deriving the rule. A filter applied after the stat would not close
it. The predicate is already pinned including its over-match cases, which
a loose startsWith('dist') would get wrong: dist-staging-no-dot and
distant.js must both be kept.

Excluding hooks/dist is independently right for this fixture — a real
marketplace install contains neither dist nor a transient staging dir,
the same reasoning that keeps the repo-root CLAUDE.md out of the
validated tree. C2 still validates clean without it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): validate agents/ too, and scope the hooks predicate

Three review Minors.

agents/ ships in package.json `files` and IS auto-validated by the CLI —
verified: a frontmatter-less agents/*.md exits 1. Including it was
pointless while the fixture symlinked, because the CLI read nothing
through a symlink; now that the tree is real it is the last shipped
component tree C2 could not see. Proven to buy coverage rather than
bytes: planting a frontmatter-less agent now reds C2, which it could
not do before this change.

shouldCopyHookEntry is documented as a hooks/ entry filter, so it now
runs only for hooks/. Applying it to the other trees was harmless today
but silently encoded a hooks-shaped exclusion into them — a future
commands/dist would have vanished from the validated tree with no signal.

assert.deepEqual -> deepStrictEqual on the nested-symlink assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): scope the fixture to the three directories the issue names

Review findings, all five.

Major 1 — `agents/` dropped from COMPONENT_DIRS. The observation behind
adding it was right (the CLI does auto-validate it; a frontmatter-less
agents/*.md exits 1), but it is a NEW gate the issue does not ask for, on
the largest of the trees, and since C2 never runs in CI it would be red
only on contributor machines with `claude` installed — the same
worst-of-both-states #3613 exists to remove. Worth having as its own
issue, where "should CI provision the CLI" gets answered for it too.

Major 2 — C3 no longer passes on a fixture that validates nothing.
buildValidationPluginRoot() mkdir's every component dir before the entry
loop, so a copy that stops happening leaves real, EMPTY directories and
all three structural assertions still hold (an empty tree contains no
symlinks). Adds a non-empty assertion plus a known entry per tree, so a
partial copy is caught as well. Stubbing the copy now reds C3 with
"commands/ is EMPTY"; dropping just commands/gsd reds it too. Neither
did before.

Minor 1 — the measured table was wrong, and the mistake was measuring an
inert file. Re-measured on CLI 2.1.239, reproducing the review's 2.1.237
result: a symlinked skills/<name>/SKILL.md at depth 2 exits 1, while a
stray symlinked *.md at the same depth exits 0. The boundary is not depth
at all, it is whether the symlink is a file the CLI reads as a component.
Table replaced with that.

Minor 2 — the hooks name filter is now local instead of importing
cold-runtime-lib-fixture.cjs's, whose docstring scopes it to the cold-tree
fixture and which had exactly one caller. Two consumers across fixtures
with different requirements and nothing asserting they stay compatible is
how a later cold-tree change silently alters what this fixture validates.

Nit 1 — cost figures re-measured: ~1.06 MB over 176 files, not 2.1 MB
over 243.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): correct the hooks row and the count the agents/ removal left behind

Three comment-prose items from review, no code change.

N1 — the corrected table gained a new wrong row, erring unsafe. "symlink
inside commands/ or hooks/ -> 0" is false for hooks/hooks.json, which is
the single most likely thing anyone would symlink there. Re-measured on
CLI 2.1.239, matching the review's 2.1.237 figures on every cell:

    symlink inside commands/ (a dir, or a component *.md)  0
    stray symlinked dir or *.js inside hooks/              0
    symlinked hooks/hooks.json                             1

hooks/hooks.json is now its own row, and is named in the C3 assertion
message alongside SKILL.md — that message is what the next engineer
actually reads.

N2 — "the four component directories" was residue from the revision that
also copied agents/; it survived the very edit that removed it, three
lines above a sentence saying three. Now three.

N3 — the promised agents/ follow-up is filed as #3751 and the comment
cites it, rather than promising an issue that did not exist. It frames
the real decision (should CI provision the claude CLI) rather than just
asking for the directory back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-22 08:06:52 -04:00
Tom Boucher
0062f6d033 fix(#2845): stop the dimension parity guard counting a back-reference (#3752)
next went red on dacae9273 against documentation that was correct.

parseDeclaredCounts read any '<numeral> ... dimensions' collocation as a claim
about the gsd-ui-checker dimension TOTAL. The how-to sentence 'Dimension 7 is a
rule gsd-ui-checker follows, the same as the other six dimensions' refers to the
other members of a seven-member set; the guard counted it as that document
declaring six, and reported a count-mismatch on prose that was accurate. A drift
guard that fires on correct prose is a false positive, which is how guards end
up switched off.

A negative lookbehind now excludes a numeral introduced by 'other' or
'remaining'. It is declared once as NOT_A_BACK_REFERENCE and shared by both
English scans rather than written at each site — the same two-surface divergence
class this suite exists to catch. Both scans apply it case-insensitively; the
first cut of this fix had the digit scan case-sensitive and the word scan not,
so a sentence-initial 'Other 6 dimensions' still slipped through. The exclusion
is word-anchored, so 'another six dimensions' — a real claim about a second set
— still counts.

parseDeclaredCounts takes an excludeBackReferences opt-out so a test can prove
against the REAL shipped how-to that the exclusion is load-bearing: with it off
the file reports [6,7], with it on [7]. That replaces a raw substring match on
prose, which the suite's own header forbids, with a typed before/after.

The docs prose is deliberately unchanged. It is the only instance of the pattern
in the tree, so keeping it means the real-tree assertion exercises the path this
fix exists for instead of asserting only on fixtures.

Regression tests cover both polarities: eight back-reference shapes including
all four sentence-initial cases, five real count claims that must still count,
the 'another' word-boundary case, and the shipped how-to itself.

Known limits recorded in the code: the exclusion is English-only, because
translated docs here are corrected to match English rather than authored, so
there is no instance to model the grammar on; and it cannot distinguish a
back-reference from a genuine total opening with the same word ('Other 6
dimensions were added'), where the false negative is the safer side of the trade.

Why it reached next at all: the docs PR (#3746) was green. A doc-only diff
inert-skips the test matrix in the PR lane, so the guard that reads docs never
ran against the docs change that broke it — it fired on push to next, after
merge.

Co-authored-by: sim <sim@local>
2026-08-22 06:39:48 -04:00
Tom Boucher
dacae92730 docs(#2845): record the inventory-provenance limits where readers meet them (#3746)
The limits shipped with #2845 were disclosed only in the PR body, which is
read once at merge and then buried. They are properties of what the feature
does, so they belong in the documentation.

Three surfaces, each at the point a reader forms an expectation:
docs/how-to/design-a-ui-phase.md gains a 'What this check is and is not'
subsection under the provenance how-to; docs/explanation/security-model.md
gains a residual-risk pair matching the section's existing shape; and
docs/AGENTS.md notes them where gsd-ui-checker's behavior is described.

The substance: a provenance line makes an inventory's origin falsifiable
rather than verified, since nothing re-runs the command or compares the
count; the rule is agent-applied like the other six dimensions, not a schema
check; and 'the checker never runs the recorded command' is an instruction
rather than a capability boundary, because the checker holds a Bash grant it
genuinely needs for the agent-skills bootstrap and tool grants here are not
command-scoped.

Co-authored-by: sim <sim@local>
2026-08-21 13:22:50 -04:00
Tom Boucher
4918c62d76 feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance

Binds two shared formats before either exists, so the suite is RED against
next: the gsd-ui-checker dimension roster (asserted independently on twelve
surfaces, eight English and four translated) and the provenance-line grammar
the UI-SPEC template emits and Dimension 7 consumes.

Every parity assertion is paired with a synthetic mutation case, so the guard's
failure branch executes rather than only reading a correct tree: limit-1 (a
surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a
label that drifts on one surface only, a non-contiguous roster, a duplicated
number, and a surface that stops declaring a count at all. A seeded fast-check
property renders the roster under formatting noise (CRLF, padding, interleaved
sections) and asserts the parse round-trips and is strictly sensitive to a
dropped heading.

Assertions are on parsed typed records, never raw substrings.

* docs: normalize design-a-ui-phase how-to to American English

House style for docs/ is American English (CLAUDE.md). This file carried
colour/initialisation/initialise/artefact throughout. Spelling only — no
content change; kept separate from the #2845 feature commit so the
release-notes classifier and the hotfix cherry-pick filter see it for what
it is.

* feat(#2845): require provenance for UI-SPEC component inventories

A UI-SPEC's component inventory was treated downstream as a closed allowlist
while the document recorded nothing about whether the list had been enumerated
from the installed design system or recalled from memory. A recalled inventory
is indistinguishable from an enumerated one, so an executor complying with the
spec builds against a fraction of what the package offers, and every gate stays
green because they assert semantics rather than composition.

The UI-SPEC template gains a Component Inventory slot carrying one of two
provenance lines: the command that enumerated the list, the count it returned,
the resolved package@version and the date; or a Could not enumerate record with
a real reason. gsd-ui-researcher gains an enumeration ladder and must record
the line rather than write the list from recall.

gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count
with no command, an empty could-not-enumerate reason, or a line still carrying
the template's unfilled placeholders BLOCKs; a partial line, a line placed below
its table, or an honest negative record FLAGs; a complete line passes, and so
does a spec carrying no inventory at all, which keeps every UI-SPEC predating
the dimension validating unchanged. Whatever the verdict, an unsourced inventory
is reported as a non-exhaustive list of known-good components rather than a
closed allowlist, so the executor is never blocked from a component the spec
merely failed to mention. The checker never runs the recorded command.

The dimension count moved on all thirteen surfaces that assert it, across five
languages. Also corrects the claim in the English, Korean and Portuguese how-tos
that this checker applies a scored six-pillar rubric — that rubric belongs to
/gsd-ui-review's retroactive audit.

* chore(#2845): backfill changeset pr number to 3745

---------

Co-authored-by: sim <sim@local>
2026-08-21 11:59:56 -04:00
Tom Boucher
2b42b28687 fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows

* fix(#3659): make baseref-head suppress mode-aware and thread isolation mode

* fix(#3659): review fixes - stale advice purge, message pins, mode alias

* fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin

* test(#3659): rewrite set-baseref pin, fix writeSync row stub

* chore(#3659): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 04:58:33 -04:00
Tom Boucher
1f76861202 fix(#3657): tolerate commonmark fence widths in ledger readers (#3733)
* test(#3657): fence-width tolerance regression rows

* test(#3657): fix pure-row fixtures to use appendWindow result shape

* fix(#3657): tolerate commonmark fence widths in ledger readers

* fix(#3657): restore throw-block indentation in parseJsonBlock

* chore(#3657): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:45:06 -04:00
Tom Boucher
72819a4616 fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi (#3731)
* test(#3031): failing-first coverage for opt-in ~/.kimi legacy reclaim

Drives the user-reachable installer surface against a sandbox HOME seeded
with the pre-#2755 wreckage: a GSD [[hooks]] block, hooks bundle and
CommonJS marker orphaned in ~/.kimi by a --kimi-code install.

Covers the reclaim itself plus the four guards the diagnosis identified as
negative space: opt-in only (no flag, no deletion), user-authored TOML and
hook files preserved, a --kimi install never reclaiming its own root, and
the KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory. Adds a fast-check property that stripping the block never
destroys user content.

Red until --reclaim-kimi-legacy exists.

Refs #3031

* fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi

A --kimi-code install older than 1.10.0 wrote its GSD [[hooks]] block, hook
bundle and CommonJS marker into Kimi CLI's ~/.kimi. #2755 fixed the
destination but could not reclaim what the old bug already wrote: the stale
block is byte-identical to a legitimate Kimi CLI one — both runtimes render
the same bytes for the same root, since the command paths derive from the
hooks root, not the runtime — so no inspection can tell litter from a working
install.

Cleanup is therefore opt-in. `--reclaim-kimi-legacy` on a --kimi-code install
removes GSD's own artifacts from the legacy root; without it nothing is
touched, so a dual-product machine keeps Kimi CLI's hooks and #2755's
acceptance criterion holds.

Extracts the uninstall path's removal sequence into reclaimKimiHooksRoot() and
drives both callers through it, so the reclaim removes precisely what a real
uninstall removes rather than a hand-copied second implementation. Guards the
wrong-runtime case (a --kimi install would delete its own hooks) and the
KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory.

Also corrects two pre-#2755 leftovers in the same surface that told users to
run `--kimi --config-dir ~/.kimi-code` — the form that produces this very
defect, since --config-dir moves only the skills root — and adds the missing
--kimi-code entry to the installer's own help.

Regression coverage folded into tests/kimi-upgrades.test.cjs beside the #2755
cases, per the regression-test-naming lint.

Fixes #3031

* fix(#3031): never reclaim ~/.kimi when this run also installs kimi

Found by the isolated adversarial review pass and independently while tracing
--all ordering, then reproduced.

selectRuntimesFromArgs orders 'kimi' before 'kimi-code' in both --all and an
explicit --kimi --kimi-code, and installAllRuntimes installs in that order. So
--all --reclaim-kimi-legacy installed a fresh, legitimate Kimi CLI hooks block
into ~/.kimi and then deleted it moments later from the kimi-code leg — exiting
0 and reporting success while leaving the user with no Kimi CLI hooks at all.
The collision guard could not catch it: kimi-code's own root is ~/.kimi-code, a
genuinely different directory.

The flag asserts "I only use Kimi Code"; installing kimi in the same invocation
falsifies that, so the reclaim is skipped with a notice.

Also hardens the collision guard itself. It compared path.resolve strings,
which returns false for two spellings of ONE directory — measured, not assumed:
a symlinked alias and a case variant on a case-insensitive filesystem both
compared unequal, so the guard would not have fired and the install would have
deleted its own freshly-written hooks. isSameDirectory now compares directories
via resolve, then dev+ino identity, then realpath.

Regression tests for all three cases; the two alias tests probe the real
filesystem and t.skip() where the alias cannot exist.

Refs #3031

* docs(#3031): reattach reclaimKimiHooksRoot's JSDoc to its own function

Inserting isSameDirectory anchored on the function name, which placed the
helper between reclaimKimiHooksRoot's doc block and the function it documents.
isSameDirectory ended up with two stacked doc blocks above it and
reclaimKimiHooksRoot with none.

Refs #3031

* fix(#3031): warn when --reclaim-kimi-legacy cannot apply

The flag only acts inside the kimi-code GLOBAL install branch. Passed with any
other runtime, or with --local, it was consumed in silence: exit 0, no cleanup,
no message. For a cleanup the user explicitly asked for, silence is
indistinguishable from "it ran and found nothing".

The scope warning is raised at argument-resolution time rather than inside
install(). kimi-code declares hostBehaviors.localInstallDeferred, so install()
returns early at the deferral check long before the kimi-hooks-toml branch — a
guard placed there is unreachable, which is both dead code and a linted drift
shape in this repo. Verified reachable by spawning the real installer.

Neither case is a hard error: the flag stays composable with --all, where it is
legitimately inert for the other seventeen runtimes.

Refs #3031

* docs(#3031): document every case where --reclaim-kimi-legacy skips

Refs #3031

* fix(#3031): resolve local config dirs from RUNTIME_META alone in the install harness

The remote runner surfaced this: the #3031 warning test drives a local
kimi-code install and died with "The path argument must be of type string.
Received undefined".

runMinimalInstall carried a SECOND, hand-maintained local-dir map beside
RUNTIME_META, and it had drifted — four runtimes present in RUNTIME_META
(hermes, kimi, kimi-code, zcode) were missing from it, so scope:'local' for any
of them resolved path.join(root, undefined) and threw a bare TypeError naming
neither the runtime nor the map at fault. #3023 had already hit exactly this
for pi and fixed it by adding one more entry, which left the divergence itself
in place for the next runtime to rediscover.

Local scope now reads RUNTIME_META.localDir, the same table the global branch
already reads, with the same loud named error the global branch raises. Parity
verified for all 14 previously-supported runtimes: every one resolves to a
byte-identical configDir. cline keeps its ternary — its local artifacts land at
the project root itself, which is a real exception, not a directory name.

Guarded in golden-parity-single-source.test.cjs beside the buildParityManifest
anti-divergence test, and both arms of that guard were proven able to fail.

Refs #3031

* chore(#3031): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:05:58 -04:00
Tom Boucher
65d2839b11 fix(#3651): prescribe only writes config-set accepts in integrations flow (#3732)
* test(#3651): regression rows for workflow config-write prescriptions

* fix(#3651): prescribe only writes config-set accepts in integrations flow

* fix(#3651): review fixes - single-source lane list, configSchema-derived test set

* fix(#3651): canonical cite, review wording fixes, one-element array pin

* chore(#3651): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 02:20:53 -04:00
Tom Boucher
94bc492f57 fix(#3645): tracked-source rule for planner/pattern-mapper path resolution (#3728)
* test(#3645): failing-first agent tracked-source contract rows

* fix(#3645): tracked-source rule for planner and pattern-mapper paths

* Revert "fix(#3645): tracked-source rule for planner and pattern-mapper paths"

This reverts commit 61f05e947bbdaf3b4897240c3819d349215744fb.

* fix(#3645): tracked-source rule at the spawn seam and mapper gate

* fix(#3645): review fixes - bounded block, ack merge assertion, git wording

* fix(#3645): fit the tracked-source block under the 1168 ceiling

* chore(#3645): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 23:24:42 -04:00
Tom Boucher
95f7c14413 fix(#3642): stop the single-section total_phases leak into an absent milestone (#3727)
* test(#3642): failing-first single-section leak rows

* fix(#3642): gate the unbounded total on any-milestone-section, not >=2

* test(#3642): rewrite the 3185 wrapper row to the withhold contract

* docs(#3642): glossary amendment for the >=1 sibling; changeset

* chore(#3642): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 22:10:52 -04:00
Tom Boucher
072b97d276 fix(#3641): make v005/v004 see bracket-convention phase entries (#3723)
* test(#3641): failing-first bracket-window validate rows

* fix(#3641): thread phase convention into hasphaseentries for v004/v005

* test(#3641): review rows - digit-anchor, decoy, probe parity, t.after

* fix(#3641): digit-anchor bracket entry token; thread probe scope axis

* fix(#3641): align frontmatter bound with probe; changeset

* chore(#3641): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 17:48:50 -04:00
Tom Boucher
03a3f779ce fix(#3640): isolate the drift cli e2e fixture via a --root override (#3722)
* test(#3640): failing-first --root isolation rows for the drift cli e2e

* fix(#3640): add --root override to the drift cli scan root

* fix(#3640): harden --root validation; typed exit-2 usage rows

---------

Co-authored-by: sim <sim@local>
2026-08-20 16:15:22 -04:00
Tom Boucher
9a69a86f42 enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter

Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe
for /gsd-pr-branch across six layers: pure classification and forbidden-path
predicates, real-git fixtures that run the cherry-pick filter loop end to end,
config-key registration through the real CLI and both manifests, the executed
worktree-materialization claim the issue's triage asked to establish, fast-check
properties over arbitrary path sets, and a drift guard over the shipped workflow.

Two live defects in today's shipped recipe are pinned as regressions, both
reproduced empirically first: `git rm -r --cached` stages a deletion of any
.planning/ path the target branch already tracks, so the generated PR removes the
base branch's planning files; and the same command leaves the cherry-picked file
untracked on disk, so a second commit touching that path aborts the pick with
"untracked working tree files would be overwritten" and every remaining commit is
silently dropped.

The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md
rather than restating them, so the workflow stays the single source of truth and the
suite cannot drift from what ships.

Refs #2971

* feat(#2971): strict planning filter mode for /gsd-pr-branch

Adds planning.pr_strict — a boolean, default false, that selects what
/gsd-pr-branch means by "filtered". Default mode is unchanged: structural
planning state survives into the PR branch and the nine transient
subdirectories do not. Strict mode drops every .planning/ path, structural
files included, and carries a commit over only when it touches at least one
file outside .planning/.

Strict mode is what makes planning.commit_docs: true safe for a project that
versions its planning tree locally but publishes none of it. The alternative
posture, commit_docs: false, silently costs parallel executor isolation — a
worktree is checked out from a commit, so an untracked or ignored .planning/
is simply absent inside it and the executor has no PLAN.md to read. That claim
is now established by an executed fixture rather than inherited.

The two path lists are declared once and both projections derived from them,
so create_pr_branch and verify can no longer disagree about what the filter
promised. verify previously counted every .planning/ path against a documented
success criterion of zero while create_pr_branch was specified to preserve five
structural files, so a correct run reported itself as failed on every phase
that touched STATE.md — which is every phase. It now asserts against the active
mode, and names the .planning/ paths default mode deliberately keeps rather
than trading a wrong signal for silence.

Two verified defects in the same recipe are fixed alongside, because strict
mode would have amplified both. `git rm -r --cached` staged a deletion for any
.planning/ path the target branch already tracked, so the generated PR removed
the base branch's planning files — under strict mode that would have been the
entire tree. The same command left the picked file untracked on disk, so a
second commit touching that path aborted the cherry-pick with "untracked
working tree files would be overwritten" and every remaining commit was
silently dropped. Both were reproduced against real git before being fixed.
The filter now forces excluded paths back to what the PR branch's HEAD carries,
in the index and the working tree; a conflict outside the filter halts instead
of being improvised past; a commit left empty by filtering is skipped rather
than failing. A clean-working-tree precondition makes the worktree half safe.

Closes #2971

* fix(#2971): unwind the checkout on a conflict halt, and test the real recipe

Two review findings, both fixed in place.

The isolated adversarial pass found that the conflict-outside-the-filter branch
exited while leaving the user checked out on the half-built PR branch with
cherry-pick state still live — this loop runs in the user's own working
directory, so stranding them there is a real cost even though it is not a
vulnerability. The branch now aborts the pick, returns to the original branch,
removes the partial PR branch, and says so before exiting.

The standards pass found the L2 fixtures executed a hand-written mirror of the
cherry-pick filter recipe rather than the recipe itself, so a reordering in the
workflow would not have been caught — and the order is load-bearing, since
restoring a path from HEAD before removing it inverts the filter. The helper now
extracts the canonical loop from the shipped workflow and the fixtures execute
that verbatim, which also gives the conflict-halt unwind above real coverage.
The drift guard additionally pins the two commands' relative order and asserts
the workflow carries exactly one canonical loop.

Also records the publication gate in the CONTEXT.md glossary next to the commit
gate it is distinct from.

Refs #2971

* fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form

The remote matrix caught two defects in the previous commit.

The halt path claimed to restore the original branch but did not. `git
cherry-pick --abort` does not apply to a single `--no-commit` pick with no
sequencer file, and the fallback left the unmerged index in place, which makes
`git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so
the user was told they had been restored while still sitting on the half-built
PR branch. The unwind now drops sequencer state, hard-resets the disposable PR
branch to clear the unmerged index, and only claims a restore when the checkout
actually succeeded; when it does not, it says where the user is and gives them
the two commands to finish it by hand. Verified against real git: exit 1, the
conflict named, HEAD back on the original branch, the partial branch gone, a
clean tree and no CHERRY_PICK_HEAD.

Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form,
which names a command no runtime registers. The canonical authoring token for
workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the
documentation added in this branch is unaffected. The comment in src/config.cts
moves to the colon form too, since it propagates into the generated lib.

Refs #2971

* docs(#2971): backfill PR number into the changeset fragments (#3720)

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:40 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8df5cb36c2 enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)
* test(#2951): pin the absent-evidence provenance contract (failing first)

17 tests / 22 anchors on the deployed agent text. Measured against the parent
commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity
guards on the package-name and in-repo-value rules -- green before and after is
their intended signature.

Refs #2951

* enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata

A claim of the form "X does not support Y" drawn from MISSING metadata -- no
python_requires, no engines field, no per-version classifier, no changelog entry,
no matching support-matrix row -- no longer earns [VERIFIED] however
authoritative the source consulted. An absence is silence about every value, so
the same evidence would "prove" both the version being ruled out and the version
being standardized on. The only route from an absence to [VERIFIED] is a positive
falsification attempt with its failing output pasted; everything short of that is
[ASSUMED], which the file already routes to "needs user confirmation before
becoming a locked decision".

Third member of the family beside the package-name and in-repo-value provenance
rules, mirroring PR #2768's shape. A present declared constraint and an
affirmatively documented incompatibility are untouched.

Closes #2951

* fix(#2951): close the allow-list ambiguity and the mutation gap review found

Findings from the isolated adversarial pass and the two-axis review, all fixed:

MAJOR (x2, one root cause) -- the absence clause and the present-constraint
carve-out gave opposite verdicts on the same evidence for the commonest real
case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A
researcher could read the enumerated list as a "declared" positive constraint
and re-earn [VERIFIED], which is also the evasion vector. The rule now states
the decision procedure -- does the declaration bound EVERY value or only the
ones it names -- and closes the positive-reframing restatement explicitly. New
contract test pins all four clauses.

MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]'
asserted two independent substrings and never that the route lands on
[VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the
rule and survived all 17 tests. Now pinned as one joined sentence.

MINOR -- the attributable-failure test regex-matched illustrative examples
("a missing certificate, a wrong host"), so a copy-edit would break it for no
reason; relaxed to the substantive clause. The no-paraphrase guard counted only
the heading, missing the drift mode in its own name; it now also pins the core
proposition to one occurrence, and the test name matches what it checks. An
off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed.

MINOR -- docs/AGENTS.md listed four of the five governed absence forms while
the agent prose, docs/COMMANDS.md and the changeset listed five; three copies
disagreeing on list membership is the drift this repo treats as a defect.

SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its
docs/README.md index line. Both reviewers flagged them as a seventh and eighth
surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill
change" row requires only docs/AGENTS.md. The actionable four-case guidance is
retained in docs/COMMANDS.md, which is inside the approved scope.

Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550
bytes headroom under the LARGE cap of 49152.

Refs #2951

* docs(#2951): restore the how-to the phase gate requires

Reverses the removal in 6404b43d3. Both /code-review axes had flagged
docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md
index line as a seventh and eighth surface beyond the six the requester capped,
and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names
only docs/AGENTS.md, so they were dropped.

gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded
enablement sequence has more than one step and the how-to quadrant is empty.
The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED]
claim, probe or cite or accept it unlocked, then answer discuss-phase's
checkpoint -- and the last step lands on a different capability's surface, so a
reference table cannot carry it. Compressing the sequence to one step to unlock
howToSkipReason would be gaming the gate, which is the same Goodhart failure
this whole change exists to close.

A machine-enforced repo gate outranks two reviewers' scope preference and my own
reading, so the page is restored and the PR body discloses the two extra
surfaces instead of hiding them. Reverting is a one-file change if a maintainer
prefers the tighter scope.

Refs #2951

* chore(#2951): backfill the changeset PR number

pr: 0 -> 3718 now that the real PR exists. The placeholder fails
scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks
lint-docs-required from consuming the fragment.

Refs #2951

---------

Co-authored-by: sim <sim@local>
2026-08-20 14:35:14 -04:00
Tom Boucher
8da2dd3ad2 feat(#2790): add read-only planning.inspect schema-v1 snapshot query (#3708)
* feat(#2790): add read-only planning.inspect schema-v1 snapshot query

Adds a read-only query emitting a schema-versioned JSON projection of .planning/
so downstream harness UIs can consume planning state without parsing GSD's
Markdown a second time.

Composed strictly from the ADR-3180 section 7 owners plus parsePlanDocument,
parseRequirements and parseUatItems; markdown structure is read through the
Markdown Sectionizer and Markdown Table Model seams. It declares its own flat
external schema rather than serializing PlanningSnapshot, which is the
diagnostic-rule subject and still growing.

Extracts plan-document parsing out of cmdPhasePlanIndex into a shared leaf
module so phase.plan-index and planning.inspect cannot drift, including the
plan-id derivation both surfaces report.

Also fixes parseRequirements dropping the separator delimiter used by the
shipped requirements template, surfaced while wiring the requirement rows.

* fix(#2790): close spec gaps and a raw-text test assertion found in review

Review findings from the standards, spec and security passes:

- phases[] rows carry goal and dependencies, the two per-phase elements the
  issue Summary names that had no corresponding field. Goal is bounded to the
  section's leading prose so the Depends-on line, the Plans checklist and the
  wave annotations are not duplicated into it.
- requirement rows carry their own diagnostic codes, so a consumer no longer
  has to string-parse the global diagnostics subject to correlate.
- roadmap_acceptance.checkbox is looked up through the phase-id key owners.
  It was compared raw against the on-disk directory name, so it read null for
  every real-world slugged phase directory and the evidence channel was inert.
- the hostile-input test asserts the structured payload instead of matching the
  raw stdout string. The absence proof over raw stdout is kept deliberately.

* fix(#2790): register planning in the runtime usage list and repair fixtures

Remote runner reported 9 failures on 9b3f9aa. Two root causes, both fixed:

- gsd-tools.cjs registered the planning family in HOST_COMMAND_ROUTERS but
  never added it to TOP_LEVEL_USAGE's Commands list. Those are two surfaces a
  parity test guards, and the top-of-file block comment is not the runtime
  help string. A real wiring gap that every local gate and three review passes
  missed.

- the new suite's fixtures could not produce a resolvable phase set. STATE.md
  frontmatter omitted the milestone field, which ADR-3180 7.2 rule 1 makes the
  primary milestone selector, so the phase set scoped unscoped and every
  percentage was correctly withheld. Separately declarePhase returned a path
  without creating the directory, so a phase declared but never written to left
  phases empty. Both reproduced against the built module before fixing.

No assertion was weakened. The withholding path is still exercised and still
returns null when the roadmap is absent.

* chore(#2790): backfill changeset pr number

* test(#2790): cover every enumerated matrix row and contain a symlink escape

Reverses a silent deferral. An earlier revision left 23 of the 78 enumerated
matrix rows unimplemented and 7 more as one-off manual checks, with a paragraph
in the artifact and the PR body describing the gap. CLAUDE.md is explicit that
such a note is not a fix and is not surfacing. The rows are implemented instead
and the manual-evidence bucket is gone: 49 test cases become 88, covering all 78.

Writing the symlink row proved a real leak: a *-PLAN.md symlinked outside
.planning/ had its content emitted into the payload, confirmed via a direct call
and the spawned CLI. readDocument now resolves target and planning root with
realpathSync and rejects an escape, returning the ordinary unreadable-document
shape. Tested both ways, because a containment check that over-rejects is its own
defect: an escaping symlink leaks nothing and degrades that plan alone, while a
legitimately relocated .planning/ symlink stays fully readable.

The three new modules are registered in the mutation COVERED registry, which had
been reporting has_work false and skipping the Stryker gate entirely. Provisional
non-binding floors so the shards run and report; raised to the measured value
before merge, since the registry forbids calibrating from a local run.

* fix(#2790): satisfy the mutation ratchet contract and scope the 1MB test

Remote runner reported 16 failures on 8c451ed. Two causes.

The COVERED registry has a paired contract the earlier commit violated: every
module needs a matching RATCHET_BASELINE entry, and minScore must be between 50
and 100 with minScore === baseline. The provisional floor of 1 was illegal on
both counts. All three modules now sit at 50 — the registry's own enforced
minimum — with matching baselines. The score cannot be measured locally: the
shard runs node --test, which this repo hard-blocks, so CI is the only source.
Floors are raised to the measured value once this PR's shards report; a shard
below 50 means the tests need strengthening, since the floor cannot go lower.

The 1MB test was measuring the test harness rather than the product. The command
handles the oversized payload correctly by spilling to a tmpfile and resolving it
back, but the resolved stdout then exceeds runGsdTools' maxBuffer and the helper
reports ENOBUFS. It now uses --pick so stdout stays one byte while the full 1MB
document is still read and parsed end to end.

* fix(#2790): wire containment across every document read this command drives

An isolated security review of the containment control found the boundary logic
sound but not comprehensively wired: two content reads reached the filesystem
without it.

An escaped phase DIRECTORY could enumerate external filenames into the file
fields and diagnostic subjects. Both enumeration sites now containment-check the
directory before reading. Worth recording that the leak was already prevented one
layer earlier than the review claimed: Dirent#isDirectory() reports false for a
directory symlink, so such a directory never becomes a phase row at all. The
guard is defense-in-depth for a direct caller and for platforms where a reparse
point reports as a directory.

A *-VERIFICATION.md symlinked outside the root leaked one frontmatter value
verbatim, because readVerificationStatus does its own read and copies an
unrecognized status into the payload's next_action. Closed from the consumer
side through that function's existing fs injection seam, so src/verification.cts
keeps its signature and its other callers are untouched.

The reviewer additionally rated a forged status: passed as an integrity bypass.
It is not: anyone able to plant the symlink can plant a real VERIFICATION.md
saying the same thing. The incremental risk is confidentiality, which is what
these fixes close.

src/plan-scan.cts is deliberately unchanged: isPlanSuperseded reads
symlink-followed content but yields only a derived boolean, no document text.

* test(#2790): give the mutation shards an in-process surface

Two Stryker shards were CANCELLED at the 15-minute cap, not failed on score.
CI log: 640 mutants instrumented, and the dry run reported 'Ran 1 tests in 20
seconds' because the shards pointed at the integration suite, where nearly every
case spawns a gsd-tools subprocess and Stryker's command runner treats the whole
test-runner invocation as a single test. 640 x 20s cannot finish in 15 minutes;
at the kill it was 27/640 with an ETA over an hour.

Every other COVERED module points at a property or unit file, and the workflow's
own paths filter lists exactly those two patterns. In-process is the intended
mutation surface; the shards were pointed at the wrong shape of test.

Adds tests/planning-inspect.unit.test.cjs — 39 cases in 10 describes that spawn
nothing and call the built modules directly. plan-document and the router need no
filesystem at all, one being a pure content-to-object parser and the other taking
an injected mock. The three shards now point here. The 91-case integration suite
is untouched and still runs in the normal test job.

* chore(#2790): ratchet mutation floors to the measured CI scores

CI run 32392791843 measured all three shards, which is the only source the
registry accepts — local runs count timeouts as kills and inflate badly.

  planning-command-router  95.65 -> floor 94
  plan-document            76.58 -> floor 75
  planning-inspect         57.03 -> floor 56

Applied the registry's own rule, floor(score) - 1, and updated RATCHET_BASELINE
to match, since the ratchet test enforces equality.

planning-inspect sits well below the file's target of 80 and is the obvious
ratchet candidate as its tests improve. planning-command-router already exceeds
the target. The placeholder comment about floors pending measurement is removed
rather than left standing as a false statement.

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:42:43 -04:00
Tom Boucher
77fa08f1e8 fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input

Locks the Step 5.5 contract that a response_language project must feed the
edge probe an English translation of each requirement's text, and binds that
advice to measured engine behavior: the same requirement classifies to zero
shapes in Portuguese and to collection/adjacency/empty/ordering in English.

Also pins the honest limit — the issue's own repro sentence classifies to []
in English too, so translation is necessary but not sufficient and the
authored shapes override is the documented fallback.

Red before the doc change; the assertions are all false today.

Refs #2773

* fix(#2773): feed the spec-phase edge probe English-translated requirement text

The shape cues in src/edge-probe.cts are English word-boundary regexes, so a
project running with response_language set wrote its SPEC requirements into the
Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero
shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy
contributed nothing and --auto left it all unresolved — the probe was a silent
no-op for exactly the spec type it exists to harden.

Step 5.5 now states that the $REQS_JSON payload is engine input rather than
user-facing output, so the response_language rule does not govern it: each
requirement's text carries a faithful English translation, the SPEC keeps its
original language, and requirement ids are never translated or renumbered. The
instruction sits before the heredoc on purpose — the downstream APPLICABLE=0
warning fires only when every requirement is unclassified, so a partly-classified
non-English spec would otherwise slip through with no signal at all.

Measured against the compiled engine: the same requirement returns [] in
Portuguese and collection -> adjacency/empty/ordering in English. Also measured:
the issue's own repro sentence returns [] in English too, so translation is
necessary but not sufficient — the instruction therefore points at the authored
shapes override for prose carrying no cue in any language rather than promising
that translation restores classification.

Doc scope only, per the triage disposition on the issue. The compiled engine is
untouched; the lang-hint / per-language cue-set fix is a separate follow-up.

Closes #2773

* fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path

Surfaced by the isolated security review of this branch. Between the mktemp and
the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about
their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before
exiting, while the empty/placeholder guard directly above it exited without one.
A spec run that tripped the placeholder check therefore stranded a temp file
holding the SPEC's requirement text in TMPDIR, once per failed run.

The added contract test walks the region between the mktemp and the
unconditional cleanup and asserts no exit path leaves the file behind, so the
two guards can no longer drift apart. Proven to bind: run against the pre-fix
file the walker reports the leaking exit; against the fixed file it reports none.

Refs #2773

* docs(#2773): record the edge probe's English-cue input constraint in the predicate store

The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and
docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real
gaps rather than incidental coupling.

CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for
classifyShape but did not record that SHAPE_CUES are English word-boundary
patterns — so the predicate store implied text was language-agnostic, which is
what a future agent reads before touching this seam.

docs/CONFIGURATION.md's response_language row is what a non-English project
reads when it turns the setting on; it now names the one deliberate exception
and links to the FEATURES.md explanation, so the interaction is discoverable
from the config key rather than only from the workflow.

CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack
fragment is updated for the final byte range and now also records the
placeholder-guard cleanup fix folded into the same block.

Refs #2773

* fix(#2773): append the growth rationale to the existing spec-phase.md ack entry

The remote runner caught this: emitted-attribution.test.cjs pins the
0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration
regression test asserts the exact '31987 -> 31997' delta text survives), so
removing it to avoid a duplicate-key collision with a new fragment broke that
test instead of satisfying the ratchet.

The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102
were each appended to the same reason string by later PRs, which is how a shared
growth key coexists with the rule that two ack sources may never name the same
path. This appends the #2773 rationale the same way and drops the separate
fragment, whose spec-phase.md key was the collision.

Verified locally by reproducing both affected tests against the real fragment
before re-dispatching: the pinned delta survives, grown[0].acked is true,
staleAcks is empty, and all 35 entries still read as spent.

Refs #2773

* docs(#2773): add a how-to for probing edges in a non-English project

The phase gate's enablementSequence check caught a wrong call of mine. I had
recorded that no how-to was owed because the user takes zero extra steps — the
workflow translates the probe input itself. Written out, though, the sequence
from off to value is two steps and step 1 depends on response_language, a
setting owned by a different capability than the edge probe, which is exactly
the condition the how-to test names.

There is also real task content a reference table cannot carry: the three-way
split between a few unclassified rows (the classifier's recall gap), every row
unclassified (the probe could not read the spec at all), and the silent
partly-classified case where the APPLICABLE=0 warning never fires. That last
one is what a user would otherwise misread as a clean bill of health.

Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard
siblings and indexed from docs/README.md next to its closest relative.

Refs #2773

* chore(#2773): backfill the changeset PR number

pr:0 placeholder replaced with the real PR number now that #3713 exists.

Refs #2773

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:36:00 -04:00
Tom Boucher
adb46cdd85 feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker

Binds the contract before any hook change exists: a `state ~N commits back`
segment gated on the state_head stamp landed by #2622, firing at the same
advisory threshold /gsd-health's W024 uses rather than at > 0.

Covers all five acceptance criteria — threshold parity (19/20/21 boundaries),
both renderers including formatGsdStateCompact, an exact spawn-count assertion,
repo-pinning and sub_repos degradation, and behavioral parity against
readStateHeadFreshness rather than a source-grep of the two fence copies.

52 example-based tests plus 5 seeded fast-check properties. Red now by design.

* feat(#2734): surface STATE.md commit-age on the statusline

Adds an opt-in `state ~N commits back` marker to the GSD-state segment,
consuming the `state_head` stamp and freshness contract landed by #2622.
A solo developer returning to a project reads "Phase 4, executing" in
STATE.md and acts on it, without noticing the codebase moved 40 commits
since that line was written. /gsd-health reports it as W024, but only if
you think to run it; the statusline is the surface you see without asking.

Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses,
not at > 0: with commit_docs:true the commit carrying a STATE.md sync
advances HEAD by one, so > 0 would alarm permanently on a fresh project.

Costs exactly one bounded git subprocess per render and none when
disabled. `rev-list --left-right --count` answers ancestry and distance
together, and repo pinning is a filesystem check mirroring
projectOwnsItsRepo rather than a --show-toplevel compare, which is
unreliable on macOS /private/var and Windows 8.3 paths.

Every unresolvable input degrades to the tri-state unknown -- the marker
is absent, never a "fresh" claim the project cannot substantiate: a
malformed stamp, a root that does not own its .git, a sub_repos
workspace, history rewound past the stamp, or git being unavailable.

Also collapses statusline config resolution onto one resolveStatuslineOptions()
seam. runStatusline() and renderStatusline() duplicated it byte-for-byte;
one copy is what keeps a newly-added key from reaching only one of them.

* test(#2734): route the e2e spawn through the process seam and fix fixture leaks

Review findings from the two orthogonal passes:

- `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched
  its stdout to test a pure function. It now calls resolveStatuslineOptions()
  directly — no subprocess, no text matching.
- `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate
  lives in runStatusline, which reads stdin), so it now spawns through
  tests/helpers/process-seam.cjs and proves the negative with a filesystem fact:
  the git shim appends to a marker file on every invocation, and the assertion is
  that the marker never appears. Stronger than asserting text is missing, and it
  drops the last stdout substring match in the block.
- Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead
  of a trailing cleanup(dir), which leaked the temp repo on assertion failure.
  derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it
  binds each directory at scheduling time rather than cleaning only the last.

Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong
expectation rather than finding a code defect: `percent` drives the progress bar
too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends
after it, which is what the test exists to prove.

CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well
as the new key; both are now enumerated.

* docs(#2734): backfill changeset PR number (#3700)

---------

Co-authored-by: sim <sim@local>
2026-08-20 00:35:01 -04:00
Tom Boucher
2fca0e17e4 enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides

Binds the not-yet-built code-review-depth module: segment-aware path-prefix
matching of a changed-file set against ordered {paths,depth} rules, resolution
order flag > strongest matching rule > global > standard, typed validation
errors, and the large-scope downgrade boundary. Also proves behaviorally that
workflow.code_review_depth_overrides is not yet a registered config key.

Refs #2554

* feat(#2554): resolve code review depth from path-scoped override rules

Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth}
rules matched against a review's changed-file set by segment-aware path-prefix
comparison. Resolution order is --depth= flag, then the strongest matching rule,
then workflow.code_review_depth, then standard; a matching rule replaces the
global rather than being max'd with it, so quick and standard rules stay
meaningful. Glob metacharacters are a hard configuration error rather than sugar
for a prefix, and malformed rules halt the review instead of degrading to
standard. The resolver is pure and reports its own provenance, so the workflow
can print the resolved depth and the rule that matched. The pre-existing
>50-file deep-to-standard downgrade moves into the module and now names the rule
it overrode.

The key is registered centrally rather than as a capability config slice: the
federated slice channel admits only boolean/string/number/enum, so an array
slice would be dropped as malformed.

Closes #2554

* test(#2554): correct depth-provenance assertions and pin out-of-repo paths

Two corrections to the failing-first suite. The source assertion for a
non-matching rule with no global configured expected 'config'; with no global
set the depth comes from the default, and a companion assertion tolerated
either value, so both passed against an implementation that derived provenance
from whether any rules existed rather than from where the depth came from.

The out-of-repo absolute-path case used a home-directory path that matched
neither implementation, so it never exercised the defect it named. It now pins
the discriminating cases: an absolute path outside the repo root must not match
a repo-relative rule, and one under the root must.

* docs(#2554): document path-scoped code review depth overrides

Reference rows for workflow.code_review_depth_overrides in the configuration,
features and commands references plus the locale copies that carry those tables,
and in the planning-config reference. Explanation of why escalation is
whole-review rather than per-file and why v1 is prefix-only. New how-to for
scoping review depth by path, carrying the configuration-error reason table and
the distinction between nothing to report and could not look. CONTEXT.md
glossary entry and the INVENTORY row for the new CLI module.

ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and
pt-BR FEATURES.md carry no code-review config table, so those files are
deliberately untouched.

* fix(#2554): make the depth-misconfiguration halt executable and reject control chars

Three review findings, all in this change.

The misconfiguration halt was prose rather than shell: the error-printing fence
was followed by an unconditional extraction fence, so an ok:false result threw
and left the depth empty instead of stopping the review. Prose is not a guard —
the two fences are now one block with a real conditional, and anything that is
not the literal string true fails closed.

An interior control character in a rule path survived validation and reached the
provenance string and the summary box; rule paths now reject control characters
via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is
unchanged. That in turn makes the field record safe to delimit, so the seven
node invocations that each re-parsed the same result to read one field collapse
to one.

Also corrects the glossary entry's illustrative paths, which the glossary-ref
check read as real repository references.

* fix(#2554): use the fast-check v4 string API and acknowledge workflow growth

Two failures from the remote matrix on d3111f45, both this branch's.

The property block built its segment arbitrary with fc.stringOf, removed in
fast-check v4. Because the arbitrary is constructed in the describe body, the
throw took out all four property tests rather than one — they had never
executed. Rewritten to fc.string({unit, ...}), the form this repo already uses
in emitted-attribution.test.cjs. Every other fast-check helper in the file was
audited against the installed module.

The emitted-attribution growth arm needed an acknowledgment for code-review.md,
which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is
spent — its ripple was absorbed when #3503 merged, and the base file is exactly
the 34435-byte baseline this growth is measured against — so it cannot clear
anything, while the ack lint hard-fails on a duplicate key across two sources.
Removed it in favor of the new fragment, which is exactly how #3503 itself
replaced the spent 3191 fragment.

* docs(#2554): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-19 22:45:35 -04:00
Tom Boucher
71e00d426e fix(#3639): dir-aware sentinel recognition for the disk-side guards (#3698)
* test(#3639): pin bracket sentinel recognition in disk-side guards

* fix(#3639): dir-aware sentinel recognition for the disk-side guards

* chore(#3639): add changeset

* fix(#3639): disclose the digit-continuation residual, join phases-clear, load-bearing over-suppression guard

* test(#3639): match the token form W007 reports

* chore(#3639): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 21:55:18 -04:00
Tom Boucher
66228a89cf fix(#3637): carry the full executor contract in the orchestrator-worktree spawn (#3694)
* test(#3637): pin the executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): carry the full executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): role-definition embed, embed-performance gates, drop stale ack

* chore(#3637): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 20:50:12 -04:00
Tom Boucher
7fc1561806 fix(#3611): decode entity-escaped ampersands and split shell segments quote-aware (#3693)
* test(#3611): pin entity-escaped ampersand chains in the negative-grep gate

* fix(#3611): decode entity-escaped ampersands before the negative-grep gate scans

* chore(#3611): add changeset

* test(#3611): pin entity chains in the 968 detector and quote-aware splits

* fix(#3611): quote-aware segment split + entity decode in both plan gates

* chore(#3611): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 19:53:45 -04:00
Tom Boucher
bad1f045b1 fix(#3610): hoist surviving top-level codex config keys to file scope on merge (#3690)
* test(#3610): pin top-level key hoisting above the codex managed block

* fix(#3610): hoist surviving top-level keys above the codex managed block

* chore(#3610): add changeset

* fix(#3610): hoist to file scope (before the first table header) with reviewer-driven coverage

* chore(#3610): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 18:38:13 -04:00
Tom Boucher
8526bd46f8 enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes

The guard matched only `resolve-execution ... --effort\s`, but the CLI also
accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written
with the equals form was a live invocation-override caller the guard passed
silently, along with `--effort` at end-of-input.

Lift the matcher to a shared predicate and assert it directly against every
shape the CLI accepts, plus the decoys it must not fire on (--effortless, a
bare --effort with no resolve-execution, item 6's --attempt caller, a call and
flag split across lines).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR

ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator
invocation override needed a caller in shipped orchestration. Per the
maintainer's ruling, take unblock path (b) for item 1 only -- record that the
override is an operator-facing CLI surface, deliberately not driven by shipped
orchestration, and ratify.

The ADR's own path (b) wording is not adopted verbatim: it says the scope is
limited to static install-time propagation, which is false on both counts --
item 6 has a live caller (#2296) and #2481 delivered a live invocation-time
argv channel. Only one precedence step is narrowed.

No consumer was invented to clear the gate: nobody has asked for a per-run
effort override, and #2475's actual complaint is already closed by the
cascade-to-argv path. The amendment states explicitly that --effort remains
supported and is not deprecated, so the scoping is not read as dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run`

There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run
and gsd-mcp-server. Both sites presented a command that cannot run as written.

One is in this branch's own new ADR-443 amendment; the other is a pre-existing
error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than
deferred. No translated copy carries either line, so no i18n drift is created.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2475): record the guard/CLI divergence risk on the item-1 matcher

The predicate independently models gsd-tools.cjs's argument parser rather than
sharing a constant with it, so a third --effort spelling would leave the guard
reporting green while ADR-443's ratifying invariant silently stopped holding.
Name that risk where the next editor will meet it.

Also restores the bounded-prose rationale that was attached to the eslint
directive removed in 39793079c -- the directive went unused once the regex moved
to a const, but the reasoning it carried is still worth having.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 17:56:04 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
Tom Boucher
79781e68eb enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands

Adds a deterministic resolvability probe over each PLAN.md <automated> verify
command and surfaces the nearest prior phase's proven commands to the planner
at every context window.

- src/verify-command-grounding.cts: recognizer (not a shell interpreter) that
  grounds a leading cd <literal> chain and npm --prefix <literal>, and reports
  unresolvable rather than guessing. Never executes command text.
- gsd-tools check verify-command-paths <N>: per-phase probe, wired into
  plan-phase.md before the plan-check pass.
- init.plan-phase gains prior_verify_commands, ungated by context_window.
- gsd-plan-checker: new Verify Command Path Resolvability dimension that
  reports the failing target and never prescribes a replacement.

Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs
(readdir order is not stable across platforms, so a module whose name extends
another's with a hyphen bucketed differently on Linux than on macOS).

Closes #2401

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets

Independent review found three defects in the recognizer:

- npm --prefix DIR run SCRIPT never reached the script-existence check,
  because the pattern required npm and run to be adjacent. That is the
  form the docs tell planners to prefer, so script_missing never fired
  for it. The prefix flag and its value are now stripped before matching.
- --prefix captured with \S+, so a quoted path containing a space was
  truncated to a stray opening quote and reported as a missing directory
  - a false blocker, worse than the bug this feature fixes. The capture
  is now quote-aware.
- A chained cd whose later segment was absolute concatenated instead of
  resetting, producing a nonsense path and another false blocker. The
  fold now resets on an absolute segment.

Also replaces the bespoke phase-directory regex with the canonical
phase-id helpers. Real phase directories are NN-slug, not phase-N-slug,
so the prior-command harvest matched nothing outside its own fixtures
and the planner-inheritance half of this feature was dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(#2401): source task blocks from the canonical sectionizer

The module carried its own copy of the <task>-block grammar - a fourth
hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps
its copy only because it needs the type= attribute the canonical helper
discards; this module never reads that attribute, so it can share the
owner outright instead of adding a test around a copy.

extractAutomatedCommands now takes task bodies from extractTaggedBlocks
and the out-of-task remainder from stripTaggedBlocks. A task-grammar
parity test pins the attributed task-name set against the canonical
helper across six awkward task shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): extract agent-file overflow to references and repair the property arbitrary

The remote matrix run came back red with 19 failures, four root causes:

- agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the
  49152 agent cap. Their bodies move to gsd-core/references/, leaving
  @-reference stubs, per the documented overflow pattern.
- The new checker dimension invoked gsd_run before the canonical
  preamble that defines it. The call is deleted outright: plan-phase.md
  already runs the probe and hands the result in as {VERIFY_PATHS}, so
  the dimension consumes that rather than re-running anything.
- fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with
  fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range.
- Three runtime-loaded files grew; acknowledged in the existing ack
  fragments that already own those bare filenames, since two ack sources
  may never name the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2401): regenerate golden install-tree fixtures for the new references

Adding two files under gsd-core/references/ changes what the installer
emits into every runtime's tree, so all 19 golden install-parity
fixtures went stale. Regenerated with npm run gen:install-tree; the
delta is exactly the two new reference paths per runtime, no removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2401): backfill changeset pr number to 3678

* fix(#2401): treat ~ as a home expansion only at the start of a path

Windows CI caught this on both shards; the Linux-only remote matrix
cannot see it. The dynamic-path refusal rejected ~ anywhere, and a
GitHub Windows runner's tmpdir is an 8.3 short name -
C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows
path came back unresolvable/dynamic_path.

This was a production bug, not a test artifact: any Windows user whose
project path carries an 8.3 short name, or any literal ~, silently lost
the probe entirely - every command degrading to unresolvable with no
explanation.

~ is a home expansion only at the start of a path; elsewhere it is an
ordinary literal. The check is now split: $, backtick, *, ? and newline
stay refused anywhere (substitution and globs, and the glob characters
are illegal in Windows path components regardless), while ~ is refused
only leading, tolerating one leading quote since the check runs before
quote stripping.

The prior tests only caught this on Windows because only Windows puts a
~ in tmpdir. Four new tests pin it on every platform via a fixture
directory literally named RUNNER~1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 15:21:15 -04:00
Tom Boucher
4e60dba717 fix(#3604): make glossary ref visibility independent of backtick parity (#3680)
* test(#3604): pin parity-dependent ref visibility in the glossary gate

* fix(#3604): make glossary ref visibility independent of backtick parity

* chore(#3604): regenerate CONTEXT-INDEX for corrected predicates

* chore(#3604): regenerate examples CONTEXT-INDEX for corrected predicates

* fix(#3604): complete retired-family exemptions and pin the guard rails

* chore(#3604): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 13:51:40 -04:00
Tom Boucher
7cf6a079fa fix(#3602): bind model resolution for every workflow subagent spawn (#3670)
* test(#3602): guard every spawned gsd-* subagent has a model resolution

* fix(#3602): bind model resolution for every workflow subagent spawn

* test(#3602): merge drift-ack entries into their owning fragments

* fix(#3602): address review findings - docs-update verifier binding, ack merge, guard residuals

* chore(#3602): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:35:45 -04:00
Tom Boucher
dae134b960 test(#2650): bound the stall-watch fan-out with the class norm that reddened windows shard 1/3 (#3671)
* test(#2650): failing-first coverage for the mis-sized stall-watch bound

Lands the regression matrix before the fix. Two assertions fail deterministically on this commit because the helper still calls raw spawnSync with a hard-coded 10000ms bound:

1. The process seam is never reached, so a mock.method spy on it records nothing and the class-norm bound cannot be observed.

2. A raw spawnSync result carries no outcome/timedOut field at all — which is precisely why the windows shard 1/3 failure printed 'null !== 0' instead of naming the timeout or the bound it exceeded.

The bound is sized for the wrong class of call. extractStallHelpersBash slices the entire bash fence, whose runtime-launcher preamble resolves gsd-tools.cjs and really runs two config-get lines — two Node spawns, measured 236ms against 2ms for the fallback the existing comment claims fires (118x). tests/helpers/timeouts.cjs already owns this class as HOOK_FANOUT_TIMEOUT_MS and its comment records the identical failure on PR #3285.

Refs #2650

* test(#2650): bound the stall-watch fan-out with the class norm and report it typed

Drives the failing-first coverage from 15aaef4a7 green (5 failures -> 0). runBashScript now routes through tests/helpers/process-seam.cjs instead of a hand-rolled spawnSync, which CONTRIBUTING.md already requires of anything that shells out.

The bound moves from a hard-coded 10000ms literal to HOOK_FANOUT_TIMEOUT_MS. That constant already exists for exactly this shape — a bash invocation that fans out to nested subprocesses — and its own comment records the identical failure on PR #3285: a bound sized for the wrong class, not a slow machine. This script is that shape: the extracted fence opens with the runtime-launcher preamble, which resolves gsd-tools.cjs and really runs two config-get lines.

The result is now typed. spawnSync reports a kill as status:null, so an exceeded bound reached the call sites as 'null !== 0' — naming neither the timeout nor the bound it exceeded. OUTCOME.TIMED_OUT names itself. status is still aliased from exitCode so the five existing assertions read unchanged.

Corrects the comment that caused the mis-sizing: it claimed gsd_run is undefined so the '|| echo' fallback fires. It is not — the preamble defines it, and the two Node spawns are real (236ms vs 2ms, 118x). They are deliberately left in place; removing them would change what the extracted script executes.

Also drops two now-dead '{ timeout: 10000 }' call-site options. The seam reads only its own documented keys, so those would have been silently ignored while still reading as a 10s bound.

Refs #2650

* test(#2650): make the boundary test a real value-domain boundary

Standards review flagged that the previous 'bound boundary' test repeated the TIMED_OUT arm the test above it already covers, and was not a limit-1/limit/limit+1 in any meaningful sense — an exact-millisecond timing edge would have been a race, not a boundary.

Replaced with a boundary on the VALUE DOMAIN of the bound itself: 0 and -1 must be rejected with TypeError, and 1 (the smallest positive value) must be accepted. Zero is the load-bearing case — spawnSync reads it as 'no timeout at all', which is exactly the unbounded-spawn hazard local/no-unbounded-spawn exists to prevent, so the seam rejects it rather than honouring it.

Also asserts the rejection path does not leak the script temp dir, since the throw escapes through runBashScript's finally. Soak: zero=TypeError, negative=TypeError, one=no-throw, leaked dirs=0.

Refs #2650

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:04:24 -04:00
Tom Boucher
8d1f770dfe test(#3395): pin the clock and scope the stale-prose scan that reddened windows shard 2/3 (#3669)
* test(#3395): failing-first coverage for the silently-ignored clock pin

Lands the regression matrix BEFORE the fix so the failure is proven rather than asserted. Three assertions fail deterministically on this commit:

1. PINNED_ENV does not actually pin. `_pinnedNowMs()` (src/clock.cts) returns null unless GSD_TEST_MODE is set, so GSD_NOW_MS alone is discarded and last_updated is stamped from the live wall clock. An instant ending ...:35.149Z contains the substring 35.1, which is what reddens the windows-latest shard 2/3 lane roughly 1 run in 600.

2. The colliding-instant regression cannot reach its instant, for the same reason.

3. The #3052 same-date test never lands on 2020-09-10, so it has been exercising the different-date path and passing for the wrong reason.

Also adds currentPositionBlock() plus boundary (ms 099/100/199/200, second 34/35/36, LF and CRLF) and two-arm fast-check coverage for the scoped read the fix will switch to.

Refs #3395

* test(#3395): pin the clock and scope the stale-prose scan to the body

Drives the failing-first coverage from a5a919ffb green. Two changes, both needed:

1. PINNED_ENV now sets GSD_TEST_MODE alongside GSD_NOW_MS. _pinnedNowMs() (src/clock.cts:44) returns null without it, so the pin was silently discarded and last_updated carried a live wall-clock instant. src/clock.cts is deliberately NOT changed: requiring both keys is what stops an ambient GSD_NOW_MS from freezing a production clock, so the caller was the side that was wrong.

2. The stale-prose assertion now reads currentPositionBlock(stateContent) instead of the whole document. Frontmatter is not phase prose, and an instant ending ...:35.149Z contains the substring 35.1 — which is exactly how a document with no stale prose in it produced 'the stale 35.1 phase prose must be refreshed away'.

Confirmed hypothesis: the two defects compose. The inert pin supplies a live timestamp; the whole-document scan turns it into a failure. Either alone is latent, which is why this sat unnoticed for five days and then reddened a lane the release never touched.

Also corrects two things the failing-first run exposed. The property test used fc.date() without noInvalidDate, so ~1 sample in 300 was an Invalid Date whose toISOString() threw (counterexample: new Date(NaN)); re-soaked at 5000 runs. And a precondition assertion added to the #3052 block was measured to pass with or without the pin, so it was removed rather than shipped as vacuous truth — last_activity there is body-derived, not clock-derived.

Refs #3395

* test(#3395): apply review findings — pin #3052, one fixture builder, CRLF coverage

Spec-axis review caught a real slip: the #3052 block carried a comment saying its pin was being added as hygiene, but the RED-state revert had removed GSD_TEST_MODE and the fix commit never restored it. A comment describing an action that was not taken is worse than either doing it or leaving it alone — the pin is now actually there.

Standards-axis review flagged the same frontmatter+heading fixture shape being rebuilt in three tests. Extracted one stateDoc({iso, lines, eol}) builder; eol is a parameter rather than a constant because the helper's CRLF behavior is a claim under test.

Self-review finding: CRLF was only exercised on a single-heading document, and the following-heading case only under LF — so the exact claim the helper's comment rests on (`\n## ` matches inside `\r\n## ` because the CR precedes the newline) was never actually run. The control test now loops both line endings WITH a following heading.

Also drops a comment that restated the PINNED_INSTANT rationale verbatim.

Refs #3395

---------

Co-authored-by: sim <sim@local>
2026-08-19 11:25:24 -04:00
Tom Boucher
cd22667b27 Merge pull request #3668 from open-gsd/chore/backmerge-main-to-next-b0ccf790
chore: back-merge main → next (b0ccf790)
2026-08-19 09:52:35 -04:00
github-actions[bot]
02fd1cfc8e chore: back-merge main into next (b0ccf790) 2026-08-19 13:52:12 +00:00
Tom Boucher
552d146086 Merge pull request #3667 from open-gsd/chore/sync-next-version-1.11.0
chore: sync next package version to 1.11.0
2026-08-19 09:51:59 -04:00