Commit Graph

930 Commits

Author SHA1 Message Date
Tom Boucher
7a02f98574 fix(#3493): confine key_links from:/to: to the project directory (#3506)
`cmdVerifyKeyLinks` resolved `from:` and `to:` with `path.join(cwd, <value>)` where the value comes verbatim from plan YAML. `path.join` normalizes `../` rather than rejecting it, so a plan travelling with a repository could name any file the process can read, and the command reports whether the link's `pattern` matched it — an arbitrary-file-read oracle reachable from `verify-phase`.

Both reads now go through `validatePath` in src/security.cts, the existing realpath-based confinement seam already used at 11 call sites. Not a missing capability — a bypassed one.

Two defects in that seam, found by adversarial review and fixed here because they affect all 11 callers:

1. A dangling in-project symlink escaped confinement. A link to an EXISTING outside path was refused (realpath lands outside) while a link to a MISSING outside path took the parent-resolution fallback and was accepted — an existence oracle for arbitrary absolute paths. lstat succeeds on a dangling link and throws ENOENT on a truly absent path; an unresolvable link is now refused. A symlink resolving inside the project is still accepted.

2. A canonicalized base was compared against an uncanonicalized path when a file and its parent were both missing, wrongly refusing legitimate in-project paths on any non-canonical cwd (every macOS temp dir). This was a live regression in this PR: the wave-pending classification (#1202) depends on the not-yet-created case. Resolution now walks up to the nearest existing ancestor.

Two adjacent aborts fixed: the from: read sat outside the per-link try, so a non-ENOENT errno killed the whole command; and an empty from: read the cwd directory, throwing EISDIR. Both now fail per-link.

The issue was filed as a fourth ADR-0174 consolidation loss. It is not one — validatePath/requireSafePath never went away, only the SDK's name for the concept did. This is an instance of epic #3473's F2 family. The ADR-0174 loss count is three.

Closes #3493

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:29:58 -04:00
Tom Boucher
e2f4c16d9e refactor(#3469): one composition for the STATE.md write seam (#3501)
* docs(#3469): amend ADR-3408 section 8.3 — the pipeline has sanctioned exceptions

Section 8.3 read 'Every STATE.md write applies the pipeline.' That is false by
design for two commands, and acting on it would have inverted a shipped
feature.

Preservation makes curated frontmatter win over a re-derived body value.
state sync exists to do the opposite — #905's 'body annotation beats existing
frontmatter when both are present'; it re-derives frontmatter FROM the body.
REGENERATE_STATE is a factory reset that rebuilds STATE.md from scratch.
Applying the pipeline to either would re-lock exactly what the command was
invoked to replace.

This issue's own scope line, inherited from the epic, said to route the direct
writeStateMd callers through the pipeline. For cmdStateSync that would have
shipped silently, with every gate green, because no test asserts that sync
LETS the body win. Caught by reading the helper's docstring and then verifying
the claim against the code — a stale comment had already misdirected this epic
once.

Both commands are now named in a closed exception list and are permanent
ratchet entries.

Consequence recorded rather than left to bite Phase 4: the 'drive the ratchet
to 0 and delete the file' target in this ADR and in #3471 is wrong. Two
entries are permanent, so the correct end state is 2, and the honest report is
'0 removable bypasses, 2 sanctioned'. A guard reaching 0 here would only do so
by having stopped looking at two real writers.

* refactor(#3469): one composition for the write seam, not one per caller

Implements ADR-3408 section 8.3 as amended.

syncAndPreserveStateMd is now the single composition of syncStateFrontmatter
and applyPostSyncPreservation. readModifyWriteStateMd and cmdPhaseComplete
both CALL it instead of each assembling the two steps themselves.
cmdPhaseComplete keeps its own writePlanningFileSet envelope — the
composition returns content, it does not take over the write, so STATE.md
still commits atomically with ROADMAP and REQUIREMENTS.

Assembling the stages at a call site is a re-derivation even when every step
calls an owner. Upstream's fix(#3374) routed cmdPhaseComplete through
applyPostSyncPreservation but left it calling syncStateFrontmatter directly
first, so the composition was duplicated and free to diverge with both guards
green. That is ADR-3180 Amendment 2's finding repeating on the write side.

cmdMilestoneComplete gains preservation. It wrote through writeStateMd, so it
got sync and no preservation — the identical shape #3374 reported for
phase.complete, and flagged upstream as a follow-up in the helper's own
docstring. This is that follow-up.

Divergence is now visible: preservation_warnings names each field restored
over a disagreeing derived value. Deliberately NOT named warnings —
cmdPhaseComplete already exposes warnings as a prose string array, and two
sibling commands carrying that name with different element types is
Generative Fix Divergence, the class this epic exists to remove.

patchCore stops running stateReplaceField over the whole document. One
observable consequence, intended per design row 9: a frontmatter-shaped patch
key with no body counterpart now reports failed instead of silently
succeeding, because the old whole-document match was literally hitting the
YAML line case-insensitively.

The guard closes Phase 1's DECLARED KNOWN GAP as promised rather than
re-deferring it: section 8.3(b) detection is tractable now the composition
exists. Scoped by two factors to avoid Phase 1's measured 29-to-1 false
positive rate — a variable field-name argument AND a content argument whose
nearest preceding assignment is not stripFrontmatter. Verified 0 findings and
0 false positives across all 33 call sites, plus 5 synthetic shapes. It also
detects the re-assembly shape above.

Ratchet: 4 entries to 2, both sanctioned-permanent. cmdStateSync's owner
changes from #3471 to sanctioned-permanent per Amendment 2 — routing it
through preservation would invert the #905 contract.

Also fixed inline rather than deferred: cmdMilestoneComplete's STATE.md read
now happens inside withStateLock. It previously read outside any lock before
writeStateMd took its own, leaving a TOCTOU window under concurrent writers.

* test(#3469): characterization coverage for the single write seam

Matrix sections A-E. Criterion 6 was amended by maintainer decision — all five
instances closed by point fixes while Phase 1 was in flight — so these are
characterization tests at the consumer's output per ADR-3180 Decision 4(b)/(c),
paired with the drift guard's count, never either alone.

Section C is the one that earns its keep. cmdStateSync is a sanctioned
permanent exception: state sync exists to re-derive frontmatter FROM the body,
so preservation there re-locks exactly what the command was invoked to
replace. C1 pins that the body wins; C4 pins that this phase left the command
byte-identical. Nothing else in the suite would notice if a future change made
sync start preserving, and the natural reading of 'one write seam' is to make
precisely that change.

Section E pins the guard's false-positive scoping. E4 (updateCore's
strip-then-replace) and E5 (sectionBody-scoped calls) must NOT be reported —
the naive detector measured 29 false positives to 1 true positive in Phase 1.
E7 is the inverse: a sanctioned-permanent entry disappearing must FAIL,
because a guard reaching zero here would only do so by having stopped looking
at two real writers.

Also corrects a stale test that asserted patchCore's old whole-document
behavior, which this phase deliberately changes.

One honest limitation, flagged rather than papered over: A1's 'byte-identical
to pre-refactor' cannot be diffed against real pre-refactor bytes from inside
the suite. It is implemented as the seeded fast-check property that
cmdPhaseComplete's composed output equals readModifyWriteStateMd's for the
same inputs — the strongest available proxy, not the literal claim.

* docs(#3469): refresh the seam glossary entry and add the changeset

Two spec-review gaps, both real.

CONTEXT.md's STATE.md Transition Module entry named three direct writeStateMd
callers including cmdMilestoneComplete. This phase routed that one through the
composition, so the line was false the moment the refactor landed.

Worth recording plainly: I wrote that sentence in Phase 0, correcting an
older stale pointer in it, and my own Phase 2 change invalidated it again
within the same epic. That is the exact drift this epic exists to remove,
demonstrated on the epic's own documentation — and it is why the entry now
ends by saying the whole-repo drift guard, not this line, is the authoritative
count.

The entry now records the composition (syncAndPreserveStateMd) and states that
exactly two direct callers remain, both SANCTIONED PERMANENT rather than debt.

Changeset: type Changed, because milestone complete's observable output moves.
Tier-2 per ADR-3180 Decision 3 — a stale body line no longer wins over fresher
frontmatter, and the command gains preservation_warnings. Docs requirement is
met by the ADR amendment already in this diff.

* test(#3469): register property-test temp-dir cleanup at creation time

Standards review, minor but real: the new fast-check property cleaned up its
temp dirs in a loop AFTER fc.assert returned. A genuine property failure
throws, so that line never ran and every dir from the failing run — including
all of fast-check's shrinking iterations — leaked.

The failure path is exactly when a littered machine hurts most, and a failing
property test is the case the test exists for.

Cleanup is now registered with t.after() at dir-creation time, so teardown
happens however the test exits. Not try/finally — CONTRIBUTING.md:356 bans it
inside test bodies, which is why the after-the-assertion shape existed in the
first place.

Swept the rest of the branch's test diff for the same shape; phase.test.cjs
already uses registered teardown and nothing else matched.

* fix(#3469): patchCore routes frontmatter writes instead of dropping them

Checkpoint returned 10 failures of 33880. One implementation defect, three
test defects, one stale test — all fixed, and the implementation defect is the
one that matters.

patchCore stripped frontmatter and then reconstructed it VERBATIM, applying no
patches to it. An arbitrary custom frontmatter key with no body counterpart and
no FIELD_CLASSIFICATION row — risk_level in the upstream fix(#3351) test —
therefore always reported failed and silently never wrote. It worked before,
via the old whole-document match on the raw YAML line.

That is a regression against this phase's own design row 9, which requires
frontmatter changes to ROUTE THROUGH the seam — still work, policy-governed —
not to stop working. Removing a capability is not routing it. An upstream test
caught it, which is the argument for running the checkpoint before believing
the refactor.

patchCore now partitions by frontmatter shape, decided structurally from the
parsed frontmatter's own keys rather than a naming heuristic:
  - classified keys still report failed — policy owns them and a raw patch may
    not bypass it;
  - unclassified keys apply to the frontmatter object and report updated —
    Phase 1's behavior-table row 19, a field with no row is not this contract's
    business;
  - body-shaped keys are unchanged.

The property 'failure' was my own test breaking the repo's Clock Seams rule.
The two paths agree byte-for-byte; the only difference was last_updated,
stamped from the wall clock on two invocations milliseconds apart, so it could
never pass. Time is now frozen with mock.timers across both — not by excluding
last_updated from the comparison, which would have silently stopped comparing
a field the composition writes.

B4's fixture could not discriminate: normalizeStateStatus maps any text
containing 'complete' to 'completed', and milestone complete's own new body
value derives to exactly that — which was also the fixture's stale value. The
stale value is now 'executing' so the assertion can tell 'body correctly won'
from 'stale survived'.

B5's fixture tripped a pre-existing unstarted-phase guard before reaching any
write-seam code; it now has the matching phase directory.

D9 asserted the old exempt set. readModifyWriteStateMd now calls one symbol
rather than assembling two, so it needs no exemption; syncAndPreserveStateMd
is the sole legitimate composition site.

* fix(#3469): patchCore resolves body-first, so the body wins a name collision

Re-verification returned 2 failures of 33880, both D4 — the hostile row for a
key that exists as BOTH a frontmatter key and a body field.

The partition checked frontmatter first, so 'status' — classified in
FIELD_CLASSIFICATION and also present as a body 'Status:' line — routed to the
frontmatter branch, was rejected as classified, and reported failed.

Wrong order. Patching 'status' means the body field, and upstream fix(#3351)
says so in its own comment: 'the legitimate working case for state.patch is
display-cased BODY fields — Status, Current Plan, Phase.' The body is
authoritative in this model; frontmatter is the projection. D4 asserted
exactly that and was right.

Resolution order is now body, then frontmatter:
  1. resolves to a body field -> apply to body, updated
  2. else an own key of the frontmatter:
       classified   -> failed  (policy owns it)
       unclassified -> apply to frontmatter, updated
  3. else -> failed

Verified by probe against the compiled lib for all four cases rather than
asserted: risk_level (frontmatter-only, unclassified) still lands;
current_phase still fails; display-cased Status unchanged; D4's lower-cased
status now lands via the body with the frontmatter untouched.

The current_phase case was the one that could have regressed silently, so its
fixture was read rather than assumed — D1's body carries 'Phase: 3 (alpha)'
and no 'Current Phase:' line, so body-first cannot reach it.

* chore(#3469): backfill pr number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:04:09 -04:00
Tom Boucher
71180983a0 fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag

* test(#3423): flip tag assertions, extend consistency guard to spawner surfaces

* fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test

* chore(#3423): acknowledge tag-rename emitted ripples and workflow growth

* chore(#3423): broaden emitted-ripple acknowledgment to all embedders

* chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys)

* chore(#3423): drop stale ripple acks, ack execute-phase growth

* chore(#3423): restore pristine 3004 fragment, keep only consumed appends

* chore(#3423): backfill changeset pr number

* chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297)

* chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple

* fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:03:48 -04:00
Tom Boucher
895d9df96d fix(#3477): run untrusted key_links patterns on a linear-time engine (#3496)
`cmdVerifyKeyLinks` compiled `must_haves.key_links[].pattern` from plan frontmatter with `new RegExp()` and tested it against whole file contents, so a nested-quantifier pattern such as `(a+)+$` hung `verify-phase` indefinitely (CWE-1333). JavaScript has no regex-execution timeout.

Untrusted patterns now run on RE2 (re2js), whose match time is linear in input length — the class is closed by the engine, not by a heuristic screen. The screen lost in the ADR-0174 consolidation was deliberately NOT restored: it never worked, since `(a|a)*$`, `((a+))+$`, `(a+){2,}$` and `(a{1,3})+$` all evade it. A refused pattern's matcher returns false for every input, so it cannot report a match no matter what the caller does.

The engine is vendored at gsd-core/bin/lib/vendor/re2js.cjs because gsd-core/bin/** is copied into installed trees with no node_modules; runtime dependencies are unchanged. New ESLint rule local/no-external-require-in-bin enforces that invariant, which had been documented in a comment since the #3024/#2071 bug class and enforced nowhere.

Backreferences and look-around are unsupported by RE2 by construction — disclosed in a Changed changeset.

Closes #3477

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 14:34:36 -04:00
Tom Boucher
1218d76d62 refactor(#3468): dispatch state preservation on the declared policy, not the field (#3495)
* test(#3468): add write-path drift guard, ratcheted at its measured baseline

Guard-first, per ADR-3180 Amendment 3's standing rule that a phase builds
and runs its guard BEFORE its scope is fixed, and states its copy count as
'N found by the guard', never 'N per the epic'.

Measured, not assumed:

  Axis 1 (policy dispatch, ADR-3408 section 8.1) — 7 violations, RED by
  design. 5 field-name-keyed getFieldClassification('literal') branches
  plus 2 declared FieldPreservation members with no executor at all
  (derive, clear). This is the fail-first evidence for the refactor.

  Axis 2 (write seam, section 8.3) — 4 bypasses, ratcheted. Epic #3408
  scoped this at two writers; the whole-repo scan found four, and one the
  epic named (patchCore) is not among them because it bypasses via
  stateReplaceField rather than the seam calls. Fourth consecutive time an
  epic's copy count proved a lower bound.

Two detectors were written and removed again before this commit, both
recorded in the file header rather than silently dropped:

  - A prompt-layer detector that reported 5 backticked prose mentions as
    drift. That is ADR-3180 Amendment 3's recorded false-positive class,
    and CONTRIBUTING.md already settles it: a backticked command reference
    is a mention. Now gated on inline-code spans.

  - A stateReplaceField co-occurrence detector for section 8.3(b). Measured
    at 29 false positives to 1 true positive — it matched the function's own
    definition and ~20 calls on frontmatter-free body slices. Banking 29
    non-defects to catch one is the 'ratchet as a parking lot' gaming route
    Decision 5 names, so it is a DECLARED KNOWN GAP owned by Phase 2
    (#3469), which both fixes it and makes its detection tractable.

* test(#3468): failing-first coverage for policy dispatch and the loud failure

Matrix sections A, B and C from 50-test-matrix.md.

Expected RED against this tree, confirmed by static trace rather than
assumed:

  B1, B2, B3 — an unwired declared preserve-when-unchanged row must throw
  with code STATE_PRESERVATION_UNWIRED_ROW and a structured .field. Today
  src/state-transition.cts:314 silently continues.

  A4 — a whitespace-only snapshot is restored today, because the guard is
  .length > 0. Required behavior is skip.

Everything else is characterization, locking in behavior the refactor must
preserve. C1 is table-driven over every FIELD_CLASSIFICATION key; C2 pins
current_phase_name's exact outputs as literals, because its row is being
reclassified preserve-always to preserve-when-unchanged as a
behavior-preserving change and nothing else would catch a drift. C3 is a
seeded fast-check property (seed 3468, 200 runs, replay data on failure).

A22 is deliberately NOT a behavioral test. Whether 'derive' has an explicit
executor is not observable through applyStatePreservation's public API — it
is a structural property, and the drift guard's unimplemented_policy axis is
what enforces it. That split is ADR-3408 Decision 5's own pairing: the lint
is the structural metric, the test is the outcome metric, and neither is
reported alone.

* refactor(#3468): dispatch preservation on the declared policy, not the field

Implements ADR-3408 sections 8.1, 8.2 and 8.6.

applyStatePreservation is now one loop over FIELD_CLASSIFICATION dispatching
on the row's preservation value, with four small executors — one per
FieldPreservation member. No branch is selected by field name. Zero
literal-argument getFieldClassification calls remain.

Behavior-preserving for 16 of 20 input classes. The four that change:

  - An unwired declared preserve-when-unchanged row now THROWS
    (code STATE_PRESERVATION_UNWIRED_ROW, structured .field) instead of
    silently continuing. This fires only on an internal invariant violation
    with both ends in our own source; a drifted, malformed or unparseable
    user STATE.md must never reach it, which is section 8.2's bright line
    and what test B8 proves through the real CLI.
  - derive gained an explicit no-op executor. That is what makes the throw
    decidable: 'policy says do nothing' is now distinguishable from 'nobody
    wired this'.
  - current_phase_name's row is corrected from preserve-always to
    preserve-when-unchanged. The row was wrong, not the code — it has always
    been delta-gated on the body Phase line, so preserve-always had two
    divergent implementations. Behavior is unchanged and test C2 pins it.
  - A whitespace-only snapshot is no longer restored; the check is trimmed.

clear is deleted from the FieldPreservation union — no row used it and no
executor existed. Speculative Generality: a policy invented for a need that
never arrived. Verified zero dependents.

The caller folds six dedicated pre/post parameters into one bodyDeltas map
keyed by field, so all seven preserve-when-unchanged rows travel one channel
instead of two. Two shapes for one kind of data is why the executor needed
per-field branches at all.

Also fixed, found while reviewing the refactor rather than deferred:

  - applyPreserveIfPlaceholder opened with a field-name literal test, which
    section 8.1 forbids outright. The executor is idempotent, so the test
    bought nothing. The drift guard could not see it, so Axis 1 is widened
    to catch field-variable comparisons against literals — the guard
    reported zero while a violation sat in the file it polices, which is
    Goodhart's gaming-by-indirection.
  - loadBaseline conflated an unreadable baseline with an absent one. A
    guard whose own diagnostic collapses two states into one identical
    result reproduces the exact failure shape this epic exists to remove.

* docs(#3468): record Phase 1 validation as ADR-3408 Amendment 1

Amendment 1 records what Phase 1 found, per ADR-3408 section 8's rule that a
behavior it does not state is not decided:

- preserve-always had TWO divergent implementations; current_phase_name's
  row was wrong and is reclassified, behavior unchanged.
- section 8.6 resolved: clear is deleted, zero dependents.
- the closed guard vocabulary is real and has exactly one true member,
  because stopped_at's scoping turned out to be caller-side extraction.
- copy count found by the guard: 4 write-seam bypasses where the epic
  scoped 2, and patchCore — one of the two it named — is not among them.
- two detectors built and removed again, with their measured false-positive
  rates, so nobody re-attempts them.
- a DECLARED KNOWN GAP for section 8.3(b), owned by Phase 2.
- Decision 5's anti-gaming list earned itself twice in one phase.

Also adds the changeset fragment.

* test(#3468): fix review findings — try/finally, stale clear allowlist, ratchet owners

Standards axis, both hard violations:

  - tests/state-write-path-drift-guard.test.cjs wrapped stdout/argv/exitCode
    restoration in try/finally inside the test body. CONTRIBUTING.md:356
    forbids it outright, and the correct t.after() pattern was already in
    use two lines up in the same test.

  - tests/state-transition.test.cjs still listed 'clear' as an allowed
    FieldPreservation value in the row-enumeration test AND the
    getFieldClassification property test, after this PR deleted it. A stale
    allowlist weakens the property's negative space — it would accept a
    resurrected clear row as valid.

Contract tension, resolved rather than left:

  ADR-3408 section 8.3 requires each ratchet entry carry the issue owning
  its removal. All four shipped with owner: null. The guard was right not to
  INVENT one, but the owners are known from the phase plan, so recording
  them is not inventing: phase.cts -> #3469, state.cts and milestone.cts ->
  #3471, health-diagnostic.cts -> sanctioned-permanent.

  Rather than a JSDoc caveat, --baseline now MERGES prior owner values on
  the (file, source) key, so a mechanical regeneration can no longer
  silently discard curated provenance. Verified by regenerating twice.

* fix(#3468): sanitize attacker-controlled fields on every guard output path

Isolated security review, MEDIUM, confidence 8/10.

findSeamBypasses and findPromptSeamUses built findings with an UNSANITIZED
`file`, while the co-located `source` on the same object was correctly
wrapped in sanitizeForReport. On a fork PR a filename is exactly as
attacker-controlled as a source fragment — a repo can legally track a
filename carrying C1 control bytes or bidi overrides.

The raw value reached two paths: --json stdout, and the COMMITTED baseline
JSON via buildBaselineEntries. JSON.stringify neutralizes C0 controls but
does NOT escape C1 (0x7f-0x9f) nor the bidi/zero-width range
sanitizeForReport exists to strip — which is the precise threat the guard's
own header names. Only the human formatter was safe.

Sanitization now happens at CONSTRUCTION, so every consumer inherits it
rather than each output path having to remember. The same defect was present
on `field` and `policy` and is fixed alongside. Double-sanitization in the
formatter is left in place, verified idempotent: escaped output is ASCII and
cannot re-match the control/bidi classes.

Also: the guard was not referenced anywhere in package.json, so nothing ran
it. A drift guard nobody runs is not a guard, and ADR-3408 Decision 5 assumes
it runs. Wired into lint:ci beside its sibling drift guards; it was already
green on this tree, so the chain stays green.

* chore(#3468): re-curate ratchet after an upstream rewording of a tracked bypass

The rebase onto origin/next turned the guard red on its first real day, which
is the ratchet working rather than a defect.

c90ae479f fix(#3350) reworded cmdPhaseComplete's syncStateFrontmatter call
onto one line and changed its third argument. Because entries are keyed on
(file, trimmed source text) rather than a line number, that single upstream
edit registered as BOTH a stale acknowledgment and an unrecorded site — the
two-sided signal the design intends, forcing a human to look rather than
letting a tracked bypass drift out of view.

The owner-preserving merge behaved exactly as designed: three owners survived
because their keys were unchanged, and phase.cts's dropped to null because its
source text is genuinely a different key. Re-curated to #3469, the phase that
owns its removal.

Note for Phase 2: c90ae479f is #3350's fix landing independently on next —
one of the two instances Phase 2 was scoped to drive fail-first. Surfaced to
the epic rather than absorbed silently.

* test(#3468): derive B1's fixture from the table so it cannot go stale

Checkpoint 2 came back with 2 failures of 33803, both B1:

  actual   'current_phase_name'
  expected 'current_plan'

The implementation was right and the test was stale. B1 hand-built a
bodyDeltas literal intending current_plan to be the ONLY unwired row, but it
also omitted status, stopped_at and current_phase_name — all three of which
became preserve-when-unchanged rows in THIS PR. Table order puts
current_phase_name first, so the throw correctly named it.

B1 now builds from neutralBodyDeltas() and deletes exactly one key, which is
what its own comment always claimed it did. A future table change can no
longer silently make it assert the wrong field.

Audited every other bodyDeltas literal in the file: four exist, all correct —
two enumerate all seven rows explicitly, two pass {} where the emptiness is
the point of the test. Roughly thirty other sites already derive from the
helper.

Also renames the local unchchangedChanged to lastActivityDescChangedDeltas.
A typo'd identifier that happens to work is still a Mysterious Name; noted
during research and fixed now that this change touches the file.

* chore(#3468): re-curate ratchet and fold the seam channel into the shared helper

The rebase onto be9329b10 fix(#3374) was a true semantic conflict, resolved
rather than handed back, because the resolution was determinable:

That PR extracted the post-sync preservation pass into a shared
applyPostSyncPreservation helper — which is ADR-3408 section 8.3, i.e. a
piece of Phase 2's own deliverable, landing upstream. Its structure is kept
wholesale; this branch's contribution is applied INSIDE it.

That combination had to be checked rather than assumed. Upstream's helper
wires only FOUR bodyDeltas keys and still passes status / stopped_at /
current_phase_name through six dedicated parameters. This branch reclassifies
current_phase_name to preserve-when-unchanged, deletes those six parameters
from StatePreservationInput, and makes an unwired declared row THROW. Taking
upstream's file as-is would therefore have thrown on EVERY STATE.md write.

The helper now wires all seven rows through the single channel. Verified
7-to-7 against FIELD_CLASSIFICATION, with a clean tsc — which is the real
proof the dedicated parameters are gone, since they no longer exist on the
input type.

The ratchet also caught the same phase.cts call being reworded a second time,
reporting it as both a stale acknowledgment and an unrecorded site. Re-curated
to #3469. Recording the tradeoff plainly: keying on (file, source text) means
an upstream reword of a tracked line needs re-curation, where keying on line
numbers would churn on every unrelated edit. ADR-3180 Decision 4(e) chose
source text deliberately, and the owner-preserving merge added earlier covers
the common case where the text is unchanged.

* chore(#3468): backfill pr number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-08-14 13:25:30 -04:00
Tom Boucher
dbc8b4077a docs(#3467): adr-3408 state.md write-path behavior contract (#3474)
* docs(#3467): adr-3408 state.md write-path behavior contract

* docs(#3467): correct write-seam caller list and adr heading depth

Review findings from the Standards and Spec axes, fixed in place:

- ADR behavior contract demoted from H2 to `### 8` with `#### 8.x`
  subsections, matching ADR-3180's `### 7` / `#### 7.1` precedent the
  front matter claims to follow.
- CONTEXT.md placed the STATE.md factory-reset primitive at
  verify.cts:1925. It moved to health-diagnostic.cts:337 when
  cmdValidateHealth migrated onto the rule table (#3309); verify.cts
  now has no writeStateMd call. The design intent was correct — only
  the address was stale.
- A repo-wide scan found three direct writeStateMd callers, not two:
  cmdStateSync, cmdMilestoneComplete, and the REGENERATE_STATE remedy.
  The last is documented as a sanctioned permanent exception — it is
  a factory reset, so preservation would restore the values it was
  invoked to discard.
- phase.cts comment citation corrected to :2953-2957.
- Amendment 4 attribution corrected: the recorded owner-file exemption
  failure is roadmap-parser.cts; the state.cts transfer is this ADR's
  own extrapolation.

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:03:48 -04:00
Tom Boucher
29c27e948b fix(#3312): require static frontend evidence before ui-plan-gate blocks (#3451)
* fix(#3312): require static frontend evidence before ui-plan-gate blocks

* fix(#3312): fill changeset pr reference

* fix(#3312): align regression assertions with section heading and legacy artifact shape

---------

Co-authored-by: sim <sim@local>
2026-08-14 09:27:48 -04:00
Tom Boucher
6b34557ba3 fix(#3311): milestone lock makes parallel-phase state conflicts visible (#3455)
* fix(#3311): milestone lock makes parallel-phase state conflicts visible

Two sessions running different phases in one working tree silently
clobbered STATE.md's single un-scoped ## Current Position slot: byte-level
serialization already existed (STATE.md.lock since #464), but nothing ever
surfaced that two sessions claimed two different phases, and
state.advance-plan — which takes no phase argument — kept advancing whatever
plan the just-clobbered position named.

Adds the maintainer-chosen milestone lock (issue #3311 comment): an advisory
.planning/milestone.lock claim keyed by phase + session id (session identity
via getWorkstreamSessionKey: env-first, then controlling TTY). begin-phase
claims (inside the STATE.md lock), advance-plan detects a claim/position
mismatch and heartbeats a matching claim, phase.complete warns via warnings[]
and releases the claim when the claimed phase completes. Conflicts warn
(stderr + typed milestone_conflict JSON field) instead of blocking, per the
decision's blocking/warning latitude; TTL 4h with heartbeat liveness expires
abandoned claims. milestone.lock is registered in the canonical artifact
registry so validate.health W019 recognizes it.

* chore(#3311): point changeset fragment at pr 3455

---------

Co-authored-by: sim <sim@local>
2026-08-14 03:04:09 -04:00
Tom Boucher
7976b1ca0d feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing

Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing.

- src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint)
- agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names
- gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json)
- execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling)
- execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree)
- config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator)
- docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests)

* chore(#1689): backfill changeset PR number (#3417)

* chore(#1689): regenerate install-tree fixtures for new workflow fragment

* chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing)

* test(#1689): SPAWN contract allows parameterized subagent_type placeholder

agent-frontmatter's spawn-type checks scanned subagent_type="..." as a
concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}"
(a runtime placeholder resolved via resolve-agent, default gsd-executor).
Skip {TOKEN} placeholders in both the known-type and <available_agent_types>
checks; execute-phase still lists the built-in roster incl. gsd-executor.

* fix(#1689): CI conformance for the routing fragment

- per-plan-executor-routing.md: add the canonical runtime-launcher preamble to
  its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments.
- agent-install-check.cts: drop a literal ~/.claude/agents path from the
  resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs
  (cline install leak guard).

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:10:52 -04:00
Tom Boucher
470389f3a2 chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169

Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes
hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"):
tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and
indentWidth (bullet-nesting depth).

git-cmd.js migrates onto tokenizeShellLike with zero behavior change
(parity-asserted against every existing #3129 fixture in
tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases
1-3 (env-prefix skip, executable check, global-option consume) extracted
into skipToSubcommand, shared with the new extractBranchArgument (git
checkout -b / git branch <name>) — a new capability exercising the seam
on the domain the ADR names, not a migration of existing duplicated logic
(none existed).

Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish
a cross-reference bullet nested under an open decision from a fresh
malformed declaration attempt. An earlier bold-run-content-classification
design was tried and disproven against the repo's own existing FIX-B
fixtures (D-02, "no colon no dash") before being adopted — both have
identical shape under any content-only rule. Nesting depth (via
indentWidth) is the actual distinguishing signal: a bullet indented
deeper than the currently-open decision's own bullet is elaboration,
folded into its text like a continuation line, never tested against the
parse-miss guard. A bullet at the same-or-shallower indent is unchanged.

Scope-narrowing disclosed, not silent: of the ADR's four named bugs
(#3197, #3169, #2570, #2528), three no longer need this phase's work.
were independently fixed and closed since the ADR was authored — #2570's
fix is already a correctly-bounded regex per the ADR's own decidability
test (no scanner needed); #2528's fix is a deliberate, twice-reviewed
non-scanner design (its own code comment records a scanner-based attempt
that regressed a symmetric case and was reverted) that this phase does
not disturb. Only #3169 required new work.

get_impact: isGitSubcommand CRITICAL/196 affected symbols,
parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence).

Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md,
docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary.

Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md
Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3414): add required fast-check property tests per code review

TESTING-STANDARDS.md:169 requires at least one fast-check property test
for any module that implements parsing — src/token-scanner.cts had none,
an orthogonal Standards-axis review finding. Adds two seeded property
tests (mirroring Phase 1/2's fast-check-setup.cjs convention):
indentWidth counts exactly a generated leading-space run; tokenizeShellLike
round-trips a generated array of whitespace/quote-free words joined with
single spaces.

The design doc's own "no property test needed" rationale was wrong — it
argued no algebraic law applied, but the standard is unconditional for
parsing modules regardless of whether one "feels" applicable. Corrected
in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md.

Also fixes two Spec-axis wording drifts the same review found between
the design doc and the shipped code (doc-only, no behavior change):
extractBranchArgument's documented signature dropped an unused
subVariants parameter that was never implemented, and the #3169
fail-first fixture description corrected from "15-decision plan via
cmdDecisionCoverageVerify" to the actual compact 3-decision analog via
the real blocking gate, check.decision-coverage-plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): add changeset for #3169 fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): backfill changeset pr number to 3424

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-13 23:08:34 -04:00
Tom Boucher
d30c99bc92 chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference

* test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md

* chore(#1892): reword retired-workflow mentions for removed-but-needed lint

* test(#1892): correct stale surface labels in retargeted suites

* docs(#1892): add verifier-phase-gates row to locale inventories

* chore(#3421): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-13 21:22:03 -04:00
Tom Boucher
0624c5da6f chore(#3212): src/text-lines.cts is the sole owner of line-terminator handling — Phase 2 (#3420)
* test(#3413): failing-first suite for the line-terminator seam

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts
does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND
at its require line, which is the intended RED.

The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first:
parseMustHavesBlock currently returns [] for every must_haves block on a
CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript
and two /m-anchored \s* patterns can absorb it, inflating a captured indent
by one character and tripping the "not nested under must_haves" guard.
Verified locally against the current (unfixed) compiled module: both the
direct repro and the silent-exit "blank line before must_haves:" variant
return [] today. A parity property test (crlf vs lf must deep-equal for
every block name) matches a pattern this maintainer has required repeatedly
for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held).

The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's
future fix-hint text (pointing at splitLines()) and its self-reference
non-violation (the seam's own correct \r?\n split must never flag itself).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* chore(#3413): src/text-lines.cts owns line-terminator handling

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/
detectEol/joinLines and migrates frontmatter.cts onto it.

parseMustHavesBlock (#3360, confirmed-bug) returned [] for every
must_haves block on a CRLF plan file. Root cause: \r is its own
LineTerminator in ECMAScript, so under /m two \s*-anchored indentation
lookups could match at the position INSIDE a \r\n pair and absorb the
terminator, inflating the captured indent by one character and tripping
the "not nested under must_haves" guard. Two silent exits, one with a
diagnostic and one without (a blank line before must_haves: hits the
silent path). Fixed by converting both lookups from a whole-string /m
match to split-then-scan — splitLines first, then a per-line, non-/m
match — the same structural pattern parseYamlRegion (30 lines away in
the same file) already used safely. Nothing downstream of the two
lookups changed; blockLines is now sliced from the already-split array
instead of re-splitting a substring, but its contents are unchanged for
LF input, and the per-line dash/kv parsing loop is untouched.

A parity property test (CRLF and LF plans parse to identical must_haves
for every block name) matches a pattern this maintainer has required
repeatedly for prior CRLF fixes in this file's neighborhood (Cortex:
7 recorded decisions, verify_intent -> held).

frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion,
isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter)
are rerouted onto splitLines — a literal 1:1 substitution, zero behavior
change, since splitLines IS that same regex plus a type guard.

The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host-
contract, gen-capability-registry, gen-context-index) are deleted and
rerouted onto normalizeEol, which strips a bare unpaired \r exactly like
the deleted copies did (not just \r\n pairs) -- verified against each
script's own --check mode against its real generated output.

local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its
fix-hint message now naming splitLines() instead of the raw regex --
the prohibition finally has a primitive to point at. Detection logic
unchanged in this phase (deliberate scope limit, see design doc Known
limits: the rule doesn't yet recognize safeReadFile/platformReadSync as
a content source, and has no detector for the \s-adjacent-to-anchor
shape that is #3360's actual mechanism -- the CLASS is converged by the
direct fix + regression test regardless).

joinLines/detectEol are NOT wired into frontmatter.cts's own write path
(cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that
platformWriteSync already, unconditionally converts CRLF->LF on every
.md write today as a pre-existing policy owned by a different module,
and ADR-3212's backward-compatibility clause rules out a file-format
change in any phase. Stated explicitly in Known limits rather than left
for a reader to discover.

Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block),
docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md
glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found

Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the
previous commit) immediately surfaced 13 real, pre-existing violations
across 10 files -- undetected until now because the rule never scanned
src/. This is the exact defect class ADR-3212 exists to close, playing
out again one phase after Phase 1 hit the same shape ("the new lint
rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule,
fixed inline rather than deferred or suppressed; there is no
established suppression convention for this rule in src/ and inventing
one now would undermine the point of widening it.

audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts,
phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or
regex character classes widened to \r?\n / [^\r\n], each following the
same pattern already established migrating frontmatter.cts.

phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` --
already semantically CRLF-safe, but the rule's lexical scanner doesn't
recognize \r? guarding a group (only \r? immediately before a literal
\n). Verified the two forms are equivalent across all four EOL/EOF
cases before restructuring, not assumed.

roadmap-upgrade.cts needed two coupled sites, not the one flagged line:
computeMigrationPlan and applyMigration must agree on line
representation for the lines[edit.lineIndex] === edit.from equality
check to hold, and the write-back needed joinLines + detectEol -- a
plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF
wholesale on every migration. This is the first real production
consumer of joinLines/detectEol in this epic (frontmatter.cts's own
write path doesn't use them -- see the previous commit's Known limits).

Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites
the rule doesn't track (.search() and new RegExp(dynamicString) aren't
in its tracked call/construction set). Investigated each empirically --
hand-tracing this exact bug class already produced one wrong conclusion
earlier in this phase (a detectEol design-doc arithmetic error), so
these were verified with real CRLF fixtures rather than reasoned about
on paper:

  - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split
    ('\n')[0]` leaked a trailing \r into a user-visible todo summary on
    CRLF input -- .trim() only strips the string's outer edges, not a
    \r sitting mid-string before the first bare \n. Now splitLines(...)
    [0].
  - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed.
    [^\n]* in targetBulletPattern swallowed a line's trailing \r on
    CRLF input, shifting the computed insert position to land INSIDE
    the \r\n pair; combined with a hardcoded '\n' bullet separator, a
    CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an
    insert. Fixed with two coupled changes (either alone still
    corrupts, verified both ways): [^\r\n]* in the pattern, and the new
    bullet's leading terminator now comes from detectEol(rawContent).
  - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan):
    investigated, genuinely safe, left untouched. The .search(/\n#{2,4}
    .../) boundary-finder and the [^\n]*-based heading match were
    empirically verified on a 3-phase CRLF fixture -- the only stray \r
    ends up at the tail of an intermediate phaseSection string that is
    only ever used for .test()-based idempotency checks, never for an
    exact-match comparison or written back to disk. No corruption on
    round-trip.

Every fix re-verified: npm run build:lib clean, npx eslint
'src/**/*.cts' --no-cache reports 0 problems (was 13), and each
fixed function's existing LF-input tests were spot-checked unchanged.

* fix(#3413): apply orthogonal review findings

Two isolated review engines (correctness + security) ran against the
full diff and found three majors, one real security issue, and several
disclosure-worthy minors. All fixed or explicitly disclosed with
evidence; nothing deferred.

MAJOR — detectEol's tie-break contradicted its own documented contract.
Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc,
CONTEXT.md, the function's own comment) says ties resolve to '\r\n'.
The existing test masked this by reusing the same tie fixture the
buggy code happened to satisfy, rather than a genuine LF-majority
case. Root cause: an Edit attempted earlier in this phase to fix this
exact arithmetic error was blocked by the tier guard, and a later
dispatch was incorrectly told it had already landed. Fixed: condition
is now crlfCount >= bareLfCount; the test fixture corrected to a
genuine 2:1 majority, with a new explicit tie-case test.

MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via
detectEol(rawContent), justified by a comment claiming a hardcoded
'\n' corrupts a CRLF ROADMAP.md. False: this write goes through
platformWriteSync, whose normalizeContent/_normalizeMd unconditionally
converts CRLF->LF for any .md target — the templating was inert dead
code, erased before the file is ever written. Reverted to hardcoded
'\n', comment corrected to state the true reasoning. The separate
[^\n]* -> [^\r\n]* widening one function up (a real splice-position
fix, independent of final EOL) was kept.

MAJOR — roadmap-upgrade.cts's stated rationale for switching onto
splitLines/joinLines was wrong (both functions always agreed on line
representation, before and after — the claimed equality-check risk
never existed), and the change it justified introduced a real
regression: forcing every line onto one dominant terminator silently
rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This
write path uses raw fs.writeFileSync, not platformWriteSync, so unlike
the phase.cts case above the regression is genuinely live.

Fixing this took two attempts. The first attempt (revert to
split('\n')/join('\n') plus a suppression comment) was correctly
blocked by an agent that discovered local/no-crlf-fragile-split is a
PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs —
a hard, out-of-band, ADR-1703-governed guardrail banning any
eslint-disable of this rule anywhere in src/**/*.cts. That agent also
detected and correctly disregarded an injected instruction that
appeared in tool output during a git operation, per this session's
untrusted-content policy. The actual fix: computeMigrationPlan
reverted to roadmapContent.split('\n') (confirmed lint-clean — the
rule's data-flow tracking only follows a variable's initializer, and
this one is declared empty then reassigned in a try block).
applyMigration's write-back now splices edits against the ORIGINAL
content string via indexOf('\n', pos) boundary-walking instead of a
full split/rejoin, so every untouched character — including every
line's own terminator — is copied byte-for-byte. A capture-group split
(/(\r\n|\n)/, preserving terminators inline) was tried first and
empirically confirmed to still trip the rule before this approach was
chosen instead.

MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used
the STRING form of String#replace, so $&, $`, $', $1-$9 inside
must_haves.truths content (author-controlled) were interpreted as
replacement directives, splicing unrelated ROADMAP.md text into the
result. Fixed with the function-replacement form, which is never
pattern-interpreted. Verified before/after with the reviewer's exact
repro.

Also disclosed rather than silently left: test matrix row 31 (four
planned CRLF-materialized regression tests) was never implemented as
separate files — corrected to record the actual verification (a
manual --check run plus incidental existing coverage via each script's
normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF
behavior was claimed byte-for-byte unchanged but the old
yaml.indexOf(blockMatch[0]) substring search could match an unrelated
earlier occurrence of the header text (e.g. inside a quoted value) —
the split-then-scan fix incidentally also closes this, a strict
improvement now recorded in the design doc rather than left implicit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error

Checkpoint 2 came back red with 5 failures on the reviewed sha, both
gaps genuinely undetectable by any local gate.

eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs'
ignores-list entry (ADR-457: generated .cjs artifacts are excluded from
direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits
two lines above it and was the exact precedent read while researching
the six-gate ripple for this module -- missed anyway. Caught by
tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which
only runs on the remote suite.

tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both
`messageId` and `message` on the same RuleTester error assertion --
ESLint's RuleTester rejects that combination outright. This existed
since the test was first authored and was never caught locally: `npx
eslint` only lints the file's syntax, it does not execute RuleTester,
and local `node --test` is hard-blocked in this repo -- the assertion
had never actually RUN before this checkpoint. It was even present in
checkpoint 1's failure list, listed there as one of the "expected RED"
tests; I matched it against my expected-failures list by test NAME
only and never inspected the actual failure detail closely enough to
notice it was failing for the wrong reason (a RuleTester config error,
not the intended message-text mismatch). Fixed by keeping `message`
(the exact-text assertion the test exists to make) and dropping
`messageId`. Verified the crlfFragileSplit message string in
eslint-rules/no-crlf-fragile-split.cjs matches this assertion
character-for-character, and swept every other invalid case in the
file for the same double-specification bug (none found -- all
pre-existing cases use messageId alone).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix

The sole user-visible effect of this phase. No breaking-change label
or Changed fragment needed — ADR-3212's Backward Compatibility section
names the Node floor (Phase 1, already shipped) as the epic's only
breaking change; Phase 2 has none.

* chore(#3413): backfill changeset pr number to 3420

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 20:27:48 -04:00
Weibao Chen
430e66dac4 docs(registries): add gsd-reasonix EoS entry (#3403)
Adds one `type: "eos"` entry for a Reasonix host integration and regenerates
docs/registries/eos-registry.md.

Reasonix is a Go, single-static-binary, DeepSeek-native terminal coding agent.
The integration renders the GSD workflow set as native Reasonix slash-command
launchers; GSD workflows are not reimplemented.

Every axis is sourced from Reasonix's own docs per the never-infer rule.
`effortSurface` is omitted: Reasonix resolves effort from config and the
/effort command with no argv route, so neither allowed value fits and the
field is optional.

Supersedes #3346, which proposed an in-tree capabilities/reasonix entry and
was closed as filed against the wrong surface.

Co-authored-by: onionviolet <apexofficial21@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-13 18:22:45 -04:00
Tom Boucher
dc3c81e93d chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam

Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts
and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both
suites fail with MODULE_NOT_FOUND, which is the intended RED.

Locks the measured behavior rather than the assumed behavior:
RegExp.escape hex-escapes the leading character of nearly every string
("abc" -> "\x61bc"), so the suite asserts match-equivalence against an
inlined historical oracle (the implementation being deleted) rather
than byte-equivalence of pattern text — 200 seeded fast-check runs plus
a fixed corpus, 0 mismatches. Also locks the latent character-class
range bug this phase fixes as a side effect: a hyphen-bearing value
interpolated into [...] currently forms a real range and matches an
unintended character; post-migration it must not.

* chore(#3412): src/pattern.cts owns runtime-value regex construction

Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam
delegating to the built-in RegExp.escape, deletes every hand-rolled
copy, and raises the Node floor to the Active LTS line.

The census was low, three times over. ADR-3212 counted 10 copies; a
graph query found 12; the new lint rule — once live — found 27 more.
The difference is that the census counted named helper FUNCTIONS while
the rule counts the escape SHAPE, so inline .replace(<class>, '\$&')
copies were never in scope. ADR §1's actual requirement is that no
module outside the seam escapes a value for regex use, so all of them
are, and CLAUDE.md's no-defer rule makes them this change's work.
Fourth consecutive epic here whose copy count was low — the argument
for ADR-3180 Amendment 3's "state N found by the guard" rule.

Also corrected mid-implementation: the survey reported phase-id.cts's
escapeRegex had 0 external importers. It had 8 production importers,
making its removal a public-surface change to an ADR-2121-owned module
and requiring an update to that ADR's locked-surface test. Blast
radius revised Medium-High -> High.

RegExp.escape is match-equivalent but NOT text-equivalent: it
hex-escapes the leading char of nearly every string ("abc" ->
"\x61bc"). Equivalence is proven by a seeded fast-check property test
against the deleted implementation as oracle. It also fixes a latent
bug: a hyphen-bearing value interpolated into a character class
previously formed a real range and matched an unintended character.

Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines,
.nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate
`required-tests` context is unchanged and no job was added or removed,
so branch protection cannot be orphaned by the dropped lanes.

Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with
structural provenance for reviewed pattern-fragment constants rather
than a name heuristic) plus a whole-tree companion guard covering the
directories ESLint's globs miss.

* fix(#3412): close the _SOURCE guard evasion, correct two false claims

Three findings from the orthogonal review pass, all fixed.

1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier-
   name matching with no binding check, so `new RegExp(userInput_SOURCE)`
   — a function parameter — sailed past the guard. That is the same
   rename-evasion class issue #3410 documents, reopened by the very
   fallback meant to complement the structural check. Now bound to the
   identifier's actual binding kind: import, require-derived const, or
   module-scope const; parameters, `let`/`var`, and unresolvable
   bindings fail closed. Four RuleTester cases cover the evasion and
   prove the legitimate cross-module case still passes.

2. src/pattern.cts's own header carried the stale pre-correction counts
   (12 copies / 17 call sites) while CONTEXT.md and the design doc
   carried the corrected ones (~39 / ~44) — a self-contradiction inside
   the PR whose entire purpose is deleting divergent copies. Rewritten,
   preserving the durable lesson: a named-function census cannot see
   inline copies; only a shape-matching guard can.

3. The claim that all deleted copies threw TypeError on non-string was
   false. phase-id.cts's copy — the one with 8 external importers — did
   String(value).replace(...) and never threw. The seam's locked
   signature does not coerce, so this is a real, now-disclosed behavior
   change rather than the pure preservation the tests asserted. Audited
   all 32 invocations across the 8 importers and 6 in-file callers:
   every one is safe by construction (upstream truthy guard or a
   string-producing derivation), verified by runtime probe against the
   compiled modules rather than by TS compilation, which cannot see a
   runtime undefined. Corrected the false claim in both the test comment
   and the design doc, and added it to Known limits.

* docs(#3412): add Changed changeset for the Node 24 floor

The only user-visible break in this phase. The escape-behavior change
is internal and match-equivalent, so it carries no user-facing note.

* fix(#3412): resolve the seam's require graph in script fixtures and packaging

Checkpoint 2 came back red with 90 failures on the node24 lane. Three
distinct defects, all introduced by routing scripts/ through the new
pattern seam, none reproducible by any local gate:

1. ~82 failures — tests/adr-index-gate.test.cjs and
   tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an
   mkdtemp fixture and spawn it there (necessary: those scripts resolve
   their scan root from __dirname/.., so running the real script would
   scan the real repo). Each harness hand-listed the dependencies to
   copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to
   gen-adr-index.cjs made both lists silently incomplete ->
   MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON'
   failures from the same crash.

   Fixed as a class, not an instance: new tests/helpers/copy-script-
   fixture.cjs walks a script's transitive static relative-require graph
   and copies it, so dependencies are derived and never re-declared. It
   throws (naming the unbuilt artifact) instead of letting the child die
   with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming
   scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host-
   contract, sync-runtime-launcher.

2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so
   the new scripts/lint-no-adhoc-regex-escape.cjs would be
   MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from
   the tarball, matching the existing precedent for gen-emitted-
   baseline.cjs, which is excluded for the identical reason, and locked
   with a test modeled on that one. Confirmed against a real npm pack:
   890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs
   present (so the other four scripts' requires are legitimate).

3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped
   source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to
   the retired hand-rolled escaper but NOT text-equivalent: it hex-
   escapes the leading character and all hyphens ('0*\x329',
   '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match
   decisions across all three real interpolation prefixes, zero
   divergence. Those tests now compile each source into the same heading
   regex src/roadmap.cts's searchPhaseInContent builds and assert what
   matches and what does not, including the 'i'-flag canonicalization
   the hex escape has to preserve. Re-pinning the new literals would
   have rebuilt the same brittleness one layer down. Adds a test for the
   property the escape exists for: a dot in '1.2' must not act as a
   wildcard.

Also shares one definition of 'a require' between the packaging guard
and the fixture copier, so the two cannot disagree about what they scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): refuse to copy a fixture dependency outside the fixture root

copyScriptWithDeps resolved each relative require and joined the
repo-relative result onto fixtureRoot. A require resolving OUTSIDE the
repo yields a '../'-prefixed relative path, so path.join climbed out of
the fixture and wrote into the surrounding temp dir (verified:
repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd).

No script in the tree does this today, so this closes an available
escape rather than an active one. Refuses via the existing unresolved-
require path so the failure names the offending specifier. Covered by a
negative proof that the guard fires and that nothing lands outside the
fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract

Applies all findings from the second orthogonal review round, re-run
because real code changed after round 1.

HIGH (security) — extractRequires stripped BLOCK comments before LINE
comments, so a '//' comment containing '/*' opened a phantom block
comment, and a '//' inside a string literal truncated the line. Both
hid real requires: 'const u="http://x"; require("./real.cjs")'
returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were
invisible. Replaced with a real AST parse via espree.

This is ADR-3212's own Decision 4 — tokenizer-first for stateful
grammars — applied to the case it describes; comment/string/regex
nesting is exactly such a grammar, which is why the regex version was
wrong. The function was moved byte-identical out of the #2858 packaging
guard, so the bug PRE-DATES this branch and has been a live blind spot
there: a shipped script could have required an unshipped path
undetected. Fixing it makes that guard strictly stronger than on next.

espree is promoted from a transitive eslint dependency to an explicit
devDependency rather than relying on hoisting. The script parse attempt
sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a
function, making a top-level return legal — scripts/check-coverage-gate
.cjs relies on it, and without the flag the guard throws on a file it
is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js
under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a
real npm pack, so the exact extractor does not newly fail the guard.

MEDIUM (security) — the repo-containment check guarded dependencies but
not the entry path. One escapesContainment predicate now guards both.

LOW (security) — containment was lexical while fs follows symlinks, and
a directory symlink could mint a fresh dedupe key per level. realpath
now resolves both repoRoot and each dependency before the decision, and
the realpath-derived path is the dedupe key. Destination layout still
uses the original repo-relative path, so copied trees are unchanged.

MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests
lost the foreign-prefix contract: every assertion was satisfied by an
impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599
bug class the exact-source prevents. The literal assertions it replaced
were catching this. Now asserts the compiled regex REJECTS a different
prefix with the same number.

MAJOR (standards) — the test hand-duplicated production's heading regex
with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed
the parallel surface instead of policing it: src/roadmap.cts exports
buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports
it. Byte-identical .source and .flags verified for both escaped forms.

MINOR — '..foo' no longer false-flagged as an escape; the inverted
spurious-vs-missing doc claim corrected; the dead allow-test-rule
header removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3412): backfill changeset pr number to 3416

* fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision

Two CI failures on PR #3416, both in code this branch added.

CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a
bracket run be consumed EITHER by the character-class branch OR one
character at a time by the trailing catch-all, so a failing match
explored both parses of every pair. Measured on the real regex:
n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script
scans repo source, so a file with a long bracket run after '.replace(/'
would hang CI outright — a guard against undisciplined pattern
construction was itself the worst pattern in the diff.

Fixed the way ADR-3212 already prescribes: the catch-all branch now
excludes '[' and ']' so a bracket can only be consumed by the class
branch (this is what makes it linear), and every quantifier is bounded
(the locked bounded-quantifiers decision) as a second line of defense.
Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the
constant: a regex literal with a BARE unescaped ']' outside a class is
no longer matched by this backstop. No census shape has that form, and
the AST rule remains the primary detector.

Verified the guard did not go blind doing it: a real census-shape
violation is still reported, and an allow-adhoc-regex-escape
suppression comment is still honored.

Regression test drives the exported findViolations on a
2000-repetition adversarial input and asserts the RESULT. It makes no
wall-clock assertion — elapsed-time tests are forbidden — so a
regression surfaces as a harness timeout, which is the correct signal.

Prompt injection scan — 'must not act as a regex wildcard' in a test
comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if|
my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a
whole test file over one phrase would blunt the scanner permanently,
and the comment has nothing to do with injection.

Neither failure was reachable from the remote runner — CodeQL and the
injection scan are not in that matrix, so the sha it passed was green
and still wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 16:19:57 -04:00
Tom Boucher
6dbc124018 enhance(#3180): the sibling validators share one envelope and one owner — Phase 12 (#3407) 2026-08-13 11:31:16 -04:00
sim
580a0c95a4 docs(#3309): update FEATURES.md's Health Validation requirements
REQ-HEALTH-05 said --repair auto-fixes recoverable issues without
qualification — now inaccurate since DESTRUCTIVE-risk remedies are
reported but never auto-applied. Adds REQ-HEALTH-06 for --backfill,
previously unmentioned in this requirements register.
2026-08-13 05:18:33 -04:00
sim
c67992de87 docs(#3309): document --backfill and the --repair DESTRUCTIVE-refusal change
docs/COMMANDS.md's /gsd-health section never documented --backfill at
all, and predates this phase's breaking changes: --repair no longer
auto-applies resetConfig/regenerateState (both destructive), and W021/
W017 split into W026/W027 for their previously-conflated second
subjects. Required by lint:docs, which needs a docs/ touch alongside
any Changed-type changeset fragment.
2026-08-13 05:14:01 -04:00
sim
42729b21fa fix: scan bin/lib subdirectories in the inventory-manifest generator
gen-inventory-manifest.cjs's cli_modules family did a flat readdirSync
of gsd-core/bin/lib/, invisible to anything shipped in a subdirectory.
Found while registering this phase's 8 health-diagnostic-rules/*.cjs
files in docs/INVENTORY.md (Standards-axis review) — the automated
manifest cross-check couldn't see them even though the manual
INVENTORY.md rows were correct.

Adds collectOneLevelSubdirs (mirrors the existing collectNested's
defensive statOrNull style) and merges flat + one-level-subdirectory
results into cli_modules's single sorted array, using the same
<subdir>/<file>.cjs key format INVENTORY.md's rows already use.

Regenerating the manifest surfaced that three OTHER existing
subdirectories (installer-migrations/, host-integration-adapters/,
observability/ — pre-existing, unrelated to this phase) were equally
invisible and had zero docs/INVENTORY.md rows at all. Added all 15
missing rows rather than leave a gap the fix itself just exposed.

Also fixes 3 pre-existing lint-legacy-dir-name violations in the
installer-migrations rows (legitimate references to the historical
get-shit-done -> gsd-core rename these migrations clean up — marked
with the guard's own gsd-allow-legacy-name exemption) and a stale
health-diagnostic.cjs row that still said "RULES ships empty."
2026-08-13 02:56:12 -04:00
sim
e0021a2fed refactor(#3309): extract health-diagnostic-types leaf module, wire RULES
Wiring all 8 rule-group files into health-diagnostic.cts's RULES array
created a genuine CJS circular dependency: each group file required
health-diagnostic.cjs back for the shared enums, and health-diagnostic.cjs
now required the group files forward, so the enums were undefined
mid-load (destructuring health-diagnostic.cjs's still-unassigned
exports).

Fixes it by splitting the enums/types (SEVERITY, REMEDY_ACTION,
REMEDY_RISK, Remedy, Diagnostic, Rule) into a dependency-free leaf
module, health-diagnostic-types.cts, that both sides import instead of
each other. health-diagnostic.cts re-exports the enums for existing
consumers. RULES is now the real concatenation of all 8 groups (31
codes — E001 intentionally stays a pre-check outside the table).
2026-08-13 01:46:20 -04:00
sim
cc1ec5b0fd chore(#3309): register health-diagnostic-rules/*.cjs as generated artifacts
Mirrors the existing health-diagnostic.cjs / planning-snapshot.cjs
pattern: gitignore the compiled output and exclude it from eslint so
the generated JS isn't linted as hand-written source. Also adds the
CONTEXT.md glossary entry and INVENTORY.md rows for the new
src/health-diagnostic-rules/ directory.
2026-08-13 01:37:13 -04:00
sim
ef10bba707 refactor(#3309): add health-diagnostic.cts skeleton (types + evaluator)
Phase 11 of epic #3180 (ADR-3180 §8.2/§8.3/§8.5). New src/health-diagnostic.cts:
SEVERITY/REMEDY_ACTION (7 members: 6 real repair actions + ADVISE)/REMEDY_RISK
(NONE/DESTRUCTIVE) frozen enums, Diagnostic/Remedy/Rule types, an empty RULES
table (rules land in the next commits), evaluateRules (with a duplicate-code
defense-in-depth check ahead of the lint guard), and applyRepairs (the
DESTRUCTIVE-risk-refusal dispatcher — §8.3 rule 3 — with stub handlers; real
repair bodies port in the migration step).

Six-gate .cts ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md +
manifest, CONTEXT.md glossary entry.
2026-08-13 00:51:12 -04:00
sim
97876a77a4 docs(#3308): ADR-3180 §8.1 Enforced, Amendment 9 — Phase 10 validation
Updates docs/adr/3180-planning-semantic-model-single-owner.md to
reflect Phase 10 shipping: §8.1 status Required -> Enforced (Phase 10,
#3308), guard-roster row contract only -> enforced, phase-index table
issue/status backfilled, and a new Amendment 9 recording the guard's
real baseline (15 distinct raw-read sites, 21 total acknowledged
occurrences in cmdValidateHealth) against the issue's own vaguer
estimate, per Amendment 4a's standing "N found by the guard, never
per the epic" rule. Also records the intended reading of an absent
STATE.md as UNREADABLE-without-diagnostic, symmetric with every other
§7 owner's absence-vs-corruption distinction.
2026-08-12 21:52:50 -04:00
sim
2538fd6344 refactor(#3308): add planning-snapshot.cts parsed projection per ADR-3180 §8.1
Phase 10 of epic #3180. src/planning-snapshot.cts is a new parsed
projection of .planning/, composed exclusively from the already-
consolidated §7 owners (getMilestoneInfo, listMilestonePhaseDirs,
isPhaseComplete, scanPhasePlans, stateFieldValue, planningPaths) plus
the frozen SCOPE enum. No new semantic derivation is introduced beyond
worstScope, a pure combinator folding several independently-scoped
owner answers into one composite signal.

Adds STATE_UNREADABLE to src/unusable-input.cts's UNUSABLE_REASON
(seventh #1879 site) for STATE.md exists-but-unreadable, distinct
from absent.

Adds scripts/lint-planning-snapshot-bypass-drift.cjs, a ratcheted
drift guard (ADR-3180 Decision 4(e)) scoped to DIAGNOSTIC_RULE_FUNCTIONS
(currently cmdValidateHealth in src/verify.cts only) preventing new
raw .planning/ reads from bypassing the snapshot, while acknowledging
cmdValidateHealth's existing 15 raw-read sites as debt owned by
Phase 11 (#3309).

Six-gate .cts ripple: .gitignore, eslint.config.mjs,
docs/INVENTORY.md + manifest regen, CONTEXT.md glossary entry.

Breaking changes: none. This phase adds the subject only; Phase 11
migrates cmdValidateHealth onto it.
2026-08-12 21:43:39 -04:00
Behruz Nassre Esfahani
f0abdb1b89 fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic

On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for
worktree isolation and persisted workflow.use_worktrees: true — the exact
value the execution workflows fail closed on (#1521 guards). Branch the
question on the same stamped config-get runtime read the guards use:
Claude keeps the unchanged question; non-Claude offers only
"No (Recommended)" / "Leave unchanged", never persists true, and warns
when the config carries an inherited explicit true. /gsd:health gains
W020, surfacing such a config with the guards' own predicate before
execution-time failure. Docs state the runtime-conditional default.

Fixes #2486

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2486): add changeset for PR #2531

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision)

The workflow-level check collided with the live W020 (git-worktree-list
health) emitted by cmdValidateHealth in src/verify.cts — invisible from
health.md's error_codes table, which stops at W019 and under-represents
the real namespace (W010-W017, W020-W023 all live). W024 verified free.
Adds a regression test pinning the chosen code against src/verify.cts so
a future assignment cannot silently collide, a table note naming the
namespace owner, and the changeset body reworded to house style.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): pre-select the recommended repair in the broken-inheritance case

Review round 2: at settings.md:142 the pre-selection rule left "Leave
unchanged" as the default when the config carried an explicit
non-false use_worktrees — the exact broken state the adjacent notice
warns about, so accepting the default kept a config that fails closed
at execution time. "Leave unchanged" is now the default only when the
key is absent (nothing to repair); explicit false AND explicit
non-false both pre-select "No (Recommended)", aligning the default,
the label, and the notice.

Pinned by two source-contract assertions in the #2486 regression
block. Goldens (settings.md hash x19) + size baseline regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name

Review round 2: #2584 Phase 3 replaced the runtime-name test with a
declared `dispatch.isolation` capability, invalidating this PR's
premise. cursor declares harness-worktree and codex/opencode/kimi/
kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate
blocked a supported configuration on five runtimes and false-warned in
health.

- settings.md + health.md read `query dispatch-isolation` and branch on
  `ISOLATION = none`; the runtime-name read is gone from both, and the
  capability read needs no per-runtime stamping (it fail-closes unknown/
  undocumented internally)
- all "Claude Code-only primitive" prose rewritten, including the two
  gates the shell-syntax check missed (config-key list, JSON schema
  comment)
- W024 reconciled across health.md + CONFIGURATION.md + planning-config.md
  (docs still said W020, which collides with a verify.cts code)
- health.md error-codes table fixed: the namespace note no longer sits
  between rows orphaning I001
- the asymmetry note for the two workflows #2584 has not migrated
  (quick.md, diagnose-issues.md) is enforced by a set-equality test with
  a self-check table, so it cannot go stale in either direction

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2486): restore the Executor isolation section clobbered by #2661

`46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md
to a pre-#2584 state: it restored the old "Non-Claude note" wording on the
workflow.use_worktrees row and deleted the whole "Executor isolation per
runtime" section. The change is unrelated to that PR's phase-estimation
feature and looks like a stale-copy edit.

This PR's use_worktrees row links to #executor-isolation-per-runtime, so the
deletion leaves a dangling anchor. Restored byte-for-byte from a40ee8a5 (the
text #2584 Phase 3 originally shipped). No link/anchor checker exists in
scripts/ or tests/, so CI would not have caught the dead link.

The reverted row wording is outside this PR's scope and is still wrong on
next; reported upstream as #2668.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2486): acknowledge the emitted size growth for settings.md and health.md

The #2723 differential attribution check (epic #2719 Phase 3) landed on next
after this branch opened and gates emitted-file growth behind a committed
acknowledgment. Both grown files are attributable to source this PR changes;
the ack names them and says why, per ADR-2719 §3.

Verified load-bearing: removing tests/emitted-drift-ack.json reproduces the
same failure; restoring it passes 59/59.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): reword the W024 remediation so it names no bare gsd-tools call

The W024 warning told the user to run `gsd-tools query config-set`, a
command-position bare invocation that fails with "command not found" on a
shim-only install (#2751) — the diagnostic sent the user into a second,
more confusing error than the one it reported. The remediation now names
only `/gsd:settings` and the config key itself, both of which work on
every install layout, and the #2751 command-position gate goes green.

Fixes #2486

* fix(#2486): drop the Known-asymmetry note and its guard test — #2728 migrates both workflows

Review Major 2: once #2728 lands, the note describes a gap that no longer
exists and instructs maintainers not to do the thing that was just done —
with the set-equality guard test pinning the stale prose green. This PR now
depends on #2728 (declared in the PR body), so the note and its guard go.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2486): make the #2728 merge order enforceable instead of advisory

settings.md recommends and persists `workflow.use_worktrees: true` for every
runtime whose declared dispatch.isolation is not `none` — cursor
(harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree).
quick.md and diagnose-issues.md consume that value at dispatch time and, while
they still gate on the runtime NAME, FATAL for all five. Merged first, this PR
reintroduces #2486's own shape for the exact runtimes it exists to help — on
the path it now labels "Recommended".

The dependency was stated only as prose in the PR body. A merge-order note is
not a gate: an automated batch merge never reads it. This adds the check that
makes the ordering structural — red while either sibling is still name-gated,
green the moment #2728 lands.

The predicate matches a RUNTIME-vs-"claude" comparison, not the legitimate
runtime-identity read, and accepts either the inline canonical
dispatch-isolation read or a reference to dispatch-isolation-gate.md, which is
the shape #2728 gives quick.md. Verified both directions against real sources:
red against the current workflows, green against #2728's.

This also closes W024's coverage window. W024 fires only when ISOLATION is
`none`, so it is structurally blind to these five runtimes — /gsd:health would
report healthy right up until quick.md FATALs. W024 cannot see the hazard, so
the hazard is prevented by making the unsafe ordering unmergeable rather than
by warning after the fact.

Separately, the W-code namespace-collision test now declares its source read
instead of passing lint silently: verify.cts emits its codes as inline string
literals across ~25 addIssue() calls and exports no enumerable registry, so
there is nothing to require() and assert against. The gap, and what it costs,
are stated in the test.

* docs(#2486): use the hyphen command form for /gsd-health in CONFIGURATION.md

docs/ is never passed through the install-time slash-form converters, so the
colon form names a command no runtime registers. Clears the sole
lint-docs-command-form violation attributable to this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): resolve isolation via new side-effect-free inspect-dispatch-isolation query; scope the settings change to isolation-none runtimes

Review round 4:

- B1: /gsd:health and /gsd:settings no longer call the recording
  dispatch-isolation query — on current next it persists the resolved
  decision to the executor-isolation sentinel as an unconditional #3045
  side effect, letting a read-only diagnostic hard-block executor
  dispatch for the sentinel's lifetime across sessions. Both surfaces now
  use inspect-dispatch-isolation, a new read-only verb sharing the exact
  resolution implementation (extracted resolveDispatchIsolationDecision)
  with zero writes. Behavioral tests pin: no sentinel write, per-runtime
  parity with the recording verb, recording knobs ignored, --json shape.

- B2: the #2728 merge-order interlock test is deleted — a repo test
  cannot sequence merges; it only made this PR unmergeable on its own
  schedule. The settings behavior change is scoped entirely to the
  ISOLATION=none branch, which needs nothing from #2728; the != none
  path is base behavior unchanged.

- M1: the W024-vs-verify.cts namespace test (an admitted source-grep) is
  deleted per RULESET.TESTS.delete-bad-tests, without a standing
  exemption; the namespace claim lives as guidance in health.md.

- M2: remaining allow-test-rule exemptions re-derived into documented
  categories (source-text-is-the-product, integration-test-input) with
  issue refs per ADR-456.

- Minor: settings.md's current-value read drops the stampable
  --default/fallback so key-absence stays distinguishable from an
  explicit false on non-Claude emits (the pre-selection rule depends on
  the tri-state); docs now present inspect-dispatch-isolation as the
  inspection command and name dispatch-isolation as the recording
  resolver.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): put the install-marker rung in the canonical runtime resolver

Round-7 review (independent cross-AI pass over the whole PR).

BLOCKER — the gate did not fire in the default case. `resolveRuntime`
stopped at GSD_RUNTIME > config.runtime > 'claude', and
`config-new-project` writes NO `runtime` key — so on a real non-Claude
install every consumer believed it was on Claude: isolation reported
`harness-worktree`, /gsd:settings still offered "Yes (Recommended)" and
W024 stayed silent. That is #2486's own symptom, surviving the fix meant
to remove it. Verified end-to-end on a real `--qwen` install with a
runtime-neutral config and GSD_RUNTIME unset: pre-fix `harness-worktree`,
fixed `none`.

The rung lives in `resolveRuntime` (src/runtime-slash.cts), the ONE
canonical resolver, not in the isolation call site. A per-consumer fix
forks precedence: `inspect-dispatch-isolation` would answer `cursor` while
`dispatch-should-flatten` and `resolve-dispatch-type` still answered
`claude` for the same install. All three now agree. `readInstallRuntimeMarker`,
its cache and its test seams MOVED from model-resolver.cts to
runtime-slash.cts, with model-resolver re-exporting the seams — one marker
read and one cache, not two that drift.

Deliberately NOT `resolveActiveRuntime`/`loadConfig`: an intermediate
revision of this fix routed through `loadConfig`, which normalizes and
rewrites legacy keys back to disk. That gave `inspect-dispatch-isolation`
— the verb whose entire purpose is being side-effect-free — a write side
effect, which is the defect the verb exists to avoid. `resolveRuntime`
reads .planning/config.json directly. Re-verified: the inspect query
against a real install creates no `.gsd/`.

Tests, both of which were too weak in the first attempt and are now
fail-first proven:
- The W024 behavioral stub returned success-with-empty-output for an absent
  key. Real `config-get` EXITS NON-ZERO, and the `|| echo "true"` fallback
  only triggers on failure — so reintroducing the fallback would have passed.
  The stub now returns 1 for the absent case.
- The marker regression asserted on the exported helper, so reverting
  gsd-tools.cjs to a marker-blind resolver still passed. It now drives a real
  `--qwen` install through the shipped `inspect-dispatch-isolation` query.
  (GSD_TEST_MODE must be cleared for that child, or install.js no-ops while
  still exiting 0 — a green test over an install that wrote nothing.)
  Spawned through `installSpawnEnv()` so ambient GSD_HOME cannot leak in.

Also: the three PR-added `try/finally` test bodies converted to `t.after()`
per CONTRIBUTING; CONTEXT.md's `worktree create` UNCONSUMED claim corrected
(Phase 3 calls it in executor-isolation-dispatch.md); docs/CONFIGURATION.md
now states the current non-Claude `use_worktrees` default rather than
describing capability-scoped stamping that lands with #2652; resolver
precedence comments updated to name the marker rung.

Refs #2486
Refs #2668

* fix(#2486): split the marker rung out; make the use_worktrees doc row order-independent

Round-8 review. trek-e's Blocker was procedural — commit 10fba7b8 moved the
per-install .gsd-runtime marker rung into the canonical resolver, a large
blast-radius change that arrived undisclosed and unreviewed. They offered
two remedies; taking the second: SPLIT IT OUT.

src/runtime-slash.cts and src/model-resolver.cts are reverted to their next
state and the marker regression test is removed. What remains is what this
PR was filed for: the settings.md / health.md / gsd-tools.cjs isolation
query, W024, and the #2668 docs restoration.

COST, stated plainly rather than buried: without that rung this PR's gate
resolves 'claude' on a non-Claude install whose project config carries no
`runtime` key — which is every config config-new-project writes. On those
installs W024 stays quiet and /gsd-settings still offers Worktrees. The gate
is correct whenever the runtime IS resolvable (GSD_RUNTIME set, or an
explicit config.runtime). That is why `Fixes #2486` is already downgraded to
`Refs` — #2486 must not close until the residual lands. A KNOWN LIMITATION
comment at the resolver call site names #2395 so this does not read as an
oversight.

The rung itself belongs to #2395, which reports this exact defect and was
closed by #2446 — a PR that touched only bin/install.js and fixtures and
never runtime-slash.cts, so it persisted the identity into
~/.gsd/defaults.json, a tier resolveRuntime does not read. Evidence and a
reopen request are posted there. That same tier is #2566's B1.

DOCS — the reviewer flagged that this row and #2728's are order-dependent:
whichever merges second falsifies the other. Removed the dependency instead
of picking an order. The "Current default … until #2652" paragraph is now a
plain troubleshooting note, true before and after #2728 lands.

CONFLICT — one hunk in CONTEXT.md: next added the #2596 scope-conformance
interface to the same Worktree-Safety paragraph where this PR corrected the
`worktree create` UNCONSUMED claim. Resolved keeping both.

Verified: lint:ci green. Full suite clean apart from the pre-existing
#1160 installed-runtime capability surface. (emitted-attribution also failed
until the fork's stale next was fast-forwarded — it defaults to origin/next,
which was 131 commits behind; 175/175 against the current base.)

Refs #2486
Refs #2668

* fix(#2486): address round-9 review — distinguish "cannot resolve" from "declares none", reject recording-only args on the read verb

Codex review of the whole PR, five findings, all verified against source first.

Major 3 (the one real defect). Both surfaces read isolation as
`ISOLATION=$(… || echo "none")`, so a resolver failure became indistinguishable
from a genuine `dispatch.isolation: none` declaration — and W024's text then
asserts the latter, telling a user their runtime declares no executor-isolation
primitive when GSD simply could not find out. health.md and settings.md now
capture the raw value and track ISOLATION_RESOLVED, the same shape
references/dispatch-isolation-gate.md already uses (#2652 review). W024 gained a
second message for the unresolved case; settings still fails closed there — it
must never persist a `true` it cannot justify — but reports what happened rather
than a verdict it never reached.

Major 4. `dispatch-isolation` applies --force-isolation AFTER the shared
resolver returns; inspection accepted the flag and silently ignored it, so the
same argv yielded 'none' from one verb and the declared capability from the
other. inspect-dispatch-isolation now rejects --force-isolation/--phase/--plan
as usage errors. Verified by mutation: disabling the guard reds exactly the
three new rejection tests.

Major 1. settings.md claimed this flow and the execution guards "always reach
the same verdict". False while quick.md still gates on the runtime name: an
orchestrator-worktree host can be offered a `true` that /gsd:quick rejects as
fatal, W024 silent because isolation is not none. Scoped the claim to the
capability gate, named #2728 as the conversion, and added the same caveat to the
CONFIGURATION.md use_worktrees row. Not a behavior change — the `!= none` branch
is untouched base behavior.

Major 2. "side-effect-free" overstated it: every gsd-tools invocation runs the
shared bootstrap, and getActiveWorkstream unlinks a stale workstream pointer.
That is pre-existing, verb-independent and cannot block a dispatch; writing the
sentinel can. Renamed the claim to "sentinel-free" everywhere and documented
precisely what is and is not asserted.

Minor 5. The JSON parity test asserted against a handwritten key list and never
invoked the recording verb, so the two could diverge and stay green. It now runs
`dispatch-isolation --json` in a separate project dir and deep-compares, with a
control asserting that verb DID write a sentinel. Added the orchestrator exec
branch (--cwd-target) the registry parity test never covered.

runtime-converters.test.cjs pinned the old `|| echo "none"` line as canonical —
that literal was the Major 3 defect. Repinned as two invariants (the raw read is
present; the collapsing fallback is absent) plus an ISOLATION_RESOLVED
requirement, so reformatting does not fail the test but a semantic regression
does.

settings.md sits at 40777 bytes against the 40960 DEFAULT hard cap. The
round-9 prose was compressed to fit rather than raising the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d (#1160 _resolveManifest, and
the #3053 quick_id tests, which compare a local-time expectation against a
TZ=UTC child and so only pass on a UTC host).

* fix(#2486): second review pass — drop a dangling reference, correct the whole unresolved branch, pin the branch behaviorally

Codex re-review of the full PR after the first round-9 pass. It confirmed Majors
1, 2 and 4 and Minor 5 fixed, and found seven more. All verified against source.

Minor 5 was mine and the worst of them: settings.md and health.md pointed at
`gsd-core/references/dispatch-isolation-gate.md`, which exists only on the #2728
branch — not in this PR and not on the merge-base. Merging this alone would have
shipped three dangling canonical-source references. Removed; the surrounding
text now stands on its own.

Major 2. The unresolved branch corrected only the two option descriptions. The
pre-selection rationale still claimed an absent key already resolves to false on
this runtime, and the explicit-true notice still stated the runtime declares no
primitive — both capability verdicts that were never reached. The substitution
rule now covers every place in that branch that asserts one.

Major 3. The repinned invariants could not catch the mutation they existed to
catch: flipping the shipped block's ISOLATION_RESOLVED=true to false left all
three green, since they only assert the raw read is present, the token appears,
and the collapsing fallback is gone. Added a behavioral test that drives the
shipped W024 block twice — resolver answering vs resolver exiting non-zero — and
asserts the two emit different text, that the unresolved branch says "could not
resolve", and that it does NOT claim the runtime has no primitive. Verified: the
true->false mutation now reds it.

Major 1. The verb fail-closes an unknown or undeclared runtime to `none` and
exits 0, so ISOLATION_RESOLVED=true means "the query answered", not "the runtime
published a declaration" — the shell cannot see an internal fallback. Exposing
provenance is an API change and out of scope here, so W024's resolved-case text
is instead written to be true of every path that reaches it ("no usable
executor-isolation primitive — declared or fail-closed from an unknown value"),
and both workflows state the limit of the signal explicitly.

Minor 4. The rejection tests regex-matched human prose, so swapping
ERROR_REASON.USAGE for UNKNOWN would have kept them green while breaking machine
consumers. They now pass --json-errors and assert reason === 'usage' on the
parsed envelope.

Minor 6. CONTEXT.md still described the verb as side-effect-free. Now
sentinel-free, consistent with the router comment and the other two surfaces.

Minor 7. 183 bytes of headroom under the 40960 hard cap was called
unacceptable, and it was — a routine 184-byte edit would have failed CI.
Compressed this PR's own settings.md prose (the #2486 block was carrying ~7.1KB
of rationale) to 40499 bytes, 458 free. Two phrases other tests pin verbatim
were restored after the first compression pass reworded them.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): third review pass — placeholder table replaces the fragile substitution rule

Codex round 3 blocked the push on two Majors. Both verified and real.

Major 2 was the substantive one and my error. The round-2 fix told the model to
find and replace the string "this runtime declares no executor-isolation
primitive (dispatch.isolation: none)" — text that does not appear anywhere in
the branch. The first option description has no parenthetical, the second
asserts "absent, it already resolves to false on this runtime" (not covered at
all), and the pre-selection rationale and notice carried three more unsupported
assertions. A model following the instruction literally would have found no
match and changed nothing.

Replaced with three named placeholders — {FINDING}, {CONSEQUENCE}, {ABSENCE} —
and a two-column table giving each one's resolved and unresolved wording. Every
assertion in the branch now flows from the table, so none of them can outrun
what was actually established, and there is no string-matching to get wrong. It
is also shorter than the prose it replaced.

Major 1: the resolver catches thrown errors and returns none successfully, a
path the round-2 wording ("declared, or unknown/undeclared") did not cover.
Both surfaces now say "declared as none, or fail-closed because the capability
could not be determined", which is true of the thrown path too, and both state
that an internal resolution error is among the things the verb fail-closes.
Still not provenance — the verb cannot distinguish these for the caller — but no
longer a claim the code contradicts.

Minor 3: the behavioral test asserted only that the two branches differ and that
the resolved one omits "could not resolve", so arbitrary resolved text stayed
green. Now pins what it must positively say: the capability finding, the
fail-closed consequence, the offending key, and (both branches) the repair
command.

Minor 4: two stale artifacts the round-2 sweep missed — a test comment still
citing the #2728-only references/dispatch-isolation-gate.md, and the emitted
drift ack still describing inspection as side-effect-free. Also aligned the two
remaining router comments to sentinel-free.

settings.md 40420 bytes, 540 free under the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): fourth review pass — stop recommending a repair whose effect depends on the emit

Codex round 4, one Major and three Minors. The Major is a real hole and worth
the round.

W024 advised "remove the key from .planning/config.json so the runtime default
(false) applies". That default is not false everywhere: execute-phase.md:102,
quick.md:155 and diagnose-issues.md:62 all read the key as
`|| echo "true"`, and only `_stampNonClaudeRuntimeDefaults` rewrites them to
`--default false` on a non-Claude emit. So on an un-stamped emit the advised
repair leaves use_worktrees resolving to TRUE against isolation none — exactly
the state W024 exists to flag. Reachable in the round-3 gap: a caught internal
resolver error returns none with rc=0, so ISOLATION_RESOLVED is true and the
confident branch fires on a host whose emit was never stamped.

Fixed by removing the dependence rather than the symptom: both surfaces now say
to set workflow.use_worktrees false explicitly, and say why deleting the key is
not equivalent. The {ABSENCE} placeholder no longer claims absence resolves to
false unconditionally — it is scoped to an emit that stamped that default.

Minor: the notice and W024 hardcoded .planning/config.json, which is the wrong
file in an active workstream (settings itself resolves $GSD_CONFIG_PATH
correctly). Both now say "the project config", naming the workstream case.

Minor: the new repair pin required the literal /gsd:settings, a form documented
as no longer routable. Relaxed to accept the canonical slash and $-prefixed
forms so a correct rewording cannot fail the test.

Nit: two test descriptions still said side-effect-free.

Not fixed, deliberately: the verb still cannot tell a caught resolver error from
a declared none, so ISOLATION_RESOLVED remains a signal about the CALL, not the
declaration. Both workflows now state that limit outright. Exposing provenance
is a change to the query contract that belongs with the runtime-resolution work
in #2395, not here.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d. The only change after that
suite run was removing a redundant regex escape flagged by eslint; that test was
re-run individually and eslint is clean.

* fix(#2486): fifth review pass — stop recommending "leave it absent" as the safe default

Codex round 5, one Major: the round-4 fix corrected the {ABSENCE} wording but
left the pre-selection rule built on the assumption it had just falsified. An
absent key still pre-selected "Leave unchanged" on the rationale that there was
nothing to repair. Absence resolves to false only where the emit stamped that
default; on an un-stamped emit it reads as true, so accepting the recommended
default could preserve the exact isolation-none-plus-true state this branch
exists to prevent.

The isolation-none branch now pre-selects "No (Recommended)" in every case,
including an absent key. Writing an explicit false is correct under either emit;
"Leave unchanged" stays available for the deliberate shared-config case but is
never the default. The test that pinned the old rule pinned a falsified
assumption, so it now pins the new one and asserts the old sentence is gone.

Also tightened the round-4 repair regex, which had been relaxed far enough to
accept `$gsd:settings` and `/gsd:settings-bogus`. It now matches only the
canonical `/gsd-settings`, `/gsd:settings` and `$gsd-settings` forms.

settings.md 40685 bytes, 275 free.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): make the inspection exemption block-scoped, and drop the last pre-#2728 claim

Codex review of the conflict resolution: one Major, two Minors, all real.

Major. The inspection-surface exemption I added to #2728's "every dispatch-site
degrade block re-records" guard was FILE-wide. Codex probed it by adding an
unrecorded executor block to health.md: the exemption assertions still passed,
the whole file was skipped, and the offender went unreported. That is the same
hole the hand-listed revision of this guard had — a promise of "every dispatch
site" with a silent carve-out.

Now scoped per block. A block earns the exemption only by resolving through
`inspect-dispatch-isolation` AND carrying no dispatch primitive (`Agent(`,
`harnessFlag`/`HARNESS_FLAG`, `isolation="worktree"`, or the recording verb).
The file-level assertions remain as a precondition on top. Verified with Codex's
own probe: the appended dispatch block is now reported at health.md:285, while
the legitimate inspection block still passes.

Minor. settings.md still carried the caveat saying `quick.md` gates on the
runtime name and "#2728 converts that surface". #2728 has landed. Same obsolete
premise already removed from docs/CONFIGURATION.md in the merge commit; this was
the copy I missed. Removing it also returns 413 bytes, so settings.md now has
688 free under the 40960 cap rather than 275.

Minor. Four scratch-directory prefixes still read `w024` after the renumber.
Non-functional, but the whole point of the rename was that W024 now means
someone else's warning.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next.

* fix(#2486): name the diagnostics' state INSPECTED_ISOLATION and delete the exemption

Codex probed the per-block exemption and it was still escapable: DISPATCH_PRIMITIVE
matched only `Agent(`, two harness-flag spellings, `isolation="worktree"` and the
recording verb, so `Task(`, `spawn_agent(`, a `codex exec` process dispatch, and —
unfixably by any regex — the repo's normal split shape (isolation resolved in one
bash fence, dispatch in a later one) all still qualified as exempt.

Took Codex's suggested fix, which is better than the one it replaces. The
diagnostics now name their state `INSPECTED_ISOLATION` / `INSPECTED_RESOLVED` /
`_INSPECTED_RAW` instead of reusing the dispatch-site names. #2728's guard scans
for a literal `ISOLATION=none`, so health.md and settings.md fall outside it BY
CONSTRUCTION and the exemption is deleted outright — no carve-out to escape, and
a diagnostic that ever writes a real `ISOLATION=none` is caught like any other
site. The name is also just more accurate: an inspection result is not a dispatch
decision.

Pinned so the reasoning cannot be lost: the #2486 suite now asserts neither
diagnostic assigns the bare `ISOLATION` name, with the rationale in the comment.
Verified by mutation — renaming back makes #2728's guard flag health.md:107 AND
fails the new naming pin.

Net effect on the guard's coverage is positive: before this PR it scanned two
fewer files by exemption; now it scans everything.

settings.md 40352 bytes, 608 free.

Validated: lint:ci clean; full suite green except the two failures that reproduce
identically on pristine next.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 20:45:08 -04:00
JusticeWay
77374cfc32 fix(#2528): resolve digit-leading phase directories by bare number (#2559)
* fix(#2528): resolve digit-slug phase dirs by bare number — tokenizer rewind, shared bare-integer fallback, resolution-path parity gate

extractPhaseToken welded 2-digit slug words onto the phase token (phase 10
named "24/7 Autonomy" -> dir 10-24-7 -> token 10-24), making digit-prefixed
phase names unresolvable by bare number across every phase verb.

- phase-id: continuation segments must be the PURE 2-digit zero-padded form
  the write side emits; a 1-digit terminator rewinds the absorbed run
  (10-24-7 -> 10) while >=2-digit terminators keep the locked #2232
  round-trip (14-06-2026-photos -> 14-06).
- phase-id: new matchPhaseDirs owner — primary exact-token match plus a
  bare-integer leading-digit-run fallback for shapes the tokenizer cannot
  rewind (05-80-20-cleanup); collisions stay #2237-loud.
- locator/find-phase/phase-plan-index all delegate selection to the owner;
  plan-index gains the previously missing multi-match guard.
- tests: #2528 unit + fast-check metamorphic blocks; new 9-scenario
  resolution-path parity gate across all three paths.

Fixes #2528

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2528): add changeset for PR #2559

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2528): align validation token grammar

* fix: address phase token review

* fix: restore phase grammar parity for numeric slugs

* docs: document digit-leading phase resolution

* docs: clarify ambiguous phase resolution behavior

* docs: register canonical phase directory selectors

* fix: align prefixed deep phase token parsing

* fix(#2528): route the fourth resolution site through matchPhaseDirs

Review BLOCKER. smart-entry.cts::detectVerifyFailed resolved the current
phase's directory with its own `.find(phaseTokenMatches)` and never
reached the shared selection, so the bare-integer-fallback family the
issue names — `05-80-20-cleanup`, `30-12-factor-refactor` — resolved
nowhere. The miss is silent by construction: an unresolved phase reports
"not failed", which is byte-identical to a healthy one, so a failed
verification simply never surfaced in /gsd or /gsd:progress.

`entries` is already sorted and matchPhaseDirs filters without
reordering, so matches[0] reproduces the previous selection exactly
wherever the old code resolved at all.

Wiring it into phase-resolution-parity.test.cjs as a fourth path then
exposed a second, older defect in the same function: phaseTokenFromDirName
shape-probed the UNSTRIPPED token, so a project-code-prefixed directory
(`MEM-05-…`, tokenizing to `MEM-05-80-20`) failed the leading-digit test
and was dropped before any resolution ran — every phase in a
project-coded plan was invisible to this check. The probe now runs on the
stripped token; the returned value is unchanged, so the comparePhaseNum
sort is untouched.

Path 4 has no JSON surface to compare, so the gate observes selection
indirectly: plant the failing artifact in exactly one directory and a
passing one everywhere else, then read the boolean. Reverting either fix
turns 5 of the 10 corpus scenarios red.

* refactor(#2528): collapse the duplicated extractCanonicalPlanId

Review MAJOR. The function existed as two independent, byte-identical
copies — src/core-utils.cts and src/phase.cts — and this PR had to patch
BOTH with the same single-digit-slug rewind rule. That is the generative
fix divergence CLAUDE.md names, and only the core-utils copy was under
test, so a future one-sided patch would have silently split plan-id
canonicalization between the plan listing and everything else.

Removed rather than parity-tested: core-utils was already the leaf owner
and already exported it, and phase.cts already imported that module, so
there is no second surface left for a parity test to police.

* test(#2528): pin matchPhaseDirs at the digit-width boundaries

Review MAJOR. The bare-integer fallback's correctness rests entirely on
capturing each directory's whole leading digit run before the zero-strip
compare; a regex that stopped short would turn every query into a prefix
match, and "1" would claim 10, 100, and 12 alike. The existing coverage
was example-based and never touched that boundary.

Adds the explicit 9/10 and 1/10/100 cases — including the forms where
only the wider directories exist, so an exact-width neighbour cannot
satisfy the assertion — plus a fast-check property over arbitrary
distinct leading runs. The property is stated as an invariant on the
result (every returned directory's leading run IS the query) rather than
an expected list, so it covers primary and fallback matches alike and
cannot be satisfied by reimplementing the selection in the test.

Both fail when the fallback regex is degraded to a prefix match.

* fix(#2528): route the remaining eight consumers through matchPhaseDirs

phaseTokenMatches had eight consumers left that each rebuilt the directory
selection around it by hand: phases-list, next-decimal, phase-remove, the
W021 milestone-consistency check, schema-drift, the init-manager overview,
milestone-complete's disk check, and roadmap analyze. Every one of them
reproduced the reported symptom in full after the tokenizer was fixed.

None of them derives a displayed phase number from the matched directory,
so none needs phaseNumberForMatch; the change at each site is the
selection and nothing else. matchPhaseDirs filters without reordering, so
matches[0] reproduces the prior .find() choice wherever the old code
resolved at all.

phaseTokenMatches now has no call sites outside phase-id.cts. It stays
exported as the primitive matchPhaseDirs is built from and as a pinned
canonical surface, but no consumer reaches past the owner to it.

* test(#2528): extend the parity gate to the migrated consumers

Each of the eight is observed through the surface a user sees, not
through the matcher, with a no-directory control so the assertions cannot
be satisfied by a consumer that resolves unconditionally. init-manager
and roadmap-analyze are additionally asserted to agree with each other.

* refactor(#2528): own the case-flexible phase grammar and the leading-digit-run fragment

validate.cts derived its case-flexible regex sources by running
`replaceAll('A-Z', 'A-Za-z')` over two constants exported by phase-id.cts.
That passes lint-phase-id-drift.cjs — there is no literal copy of the
grammar — but it depends on the owner rendering that exact substring. The
day phase-id.cts expresses the same class any other way the replaceAll
silently no-ops and validate.cts narrows to uppercase-only. The failure
mode is a NON-match, so nothing throws and no uppercase-only fixture
notices. Both variants are now derived once, beside the sources they
widen, and imported.

Also names the leading digit run the bare-integer fallback selects on.
It was spelled `/^(\d+)(?:-|$)/` where the fallback filters and `/^\d+/`
where phaseNumberForMatch reads the number back off the winner; selecting
on one run and displaying another would resolve a directory and then label
it with a number that never matched it.

* fix(#2528): refuse to remove a phase when two directories claim its number

cmdPhaseRemove was the only migrated site taking matches[0] with no
multi-match guard. Every sibling resolution path returns ambiguous_matches
and refuses to choose; this one is the DESTRUCTIVE path, so choosing
silently is strictly worse than anywhere else. With 05-80-20-a and
05-90-till-late on disk, `phase remove 5 --force` deleted one of them and
renumbered every phase after it — where the base resolved nothing, deleted
nothing, and the corpus in tests/phase-resolution-parity.test.cjs already
declared that exact input ambiguous.

The refusal is emitted before any file is touched and carries both
candidates. CONSUMER_SCENARIOS could not express the case — every row is
binary, resolving to one directory or to none — so the gate gains a
dedicated ambiguous test. It asserts on the filesystem, not only on the
reported directory_deleted: a null printed after an rmSync would satisfy
every other check.

* fix(#2528): pair digit-leading phase directories with their roadmap phase in validate health

W006/W007 are the ninth site of this bug class and the one a
`phaseTokenMatches` grep could never surface: they resolve roadmap↔disk by
intersecting TOKEN SETS, which is a dir→token labelling rather than the
query→dir selection matchPhaseDirs owns. On the canonical fixture the
label is wrong in both directions at once, so `validate health` reported
"Phase 5 in ROADMAP.md but no directory on disk" AND "Phase 05-80-20
exists on disk but not in ROADMAP.md" for the same directory.

collectDiskPhases now keeps the directory names behind each token, so
W006 can ask the canonical matcher whether a roadmap phase resolves to a
real directory, and W007 — which iterates directories and therefore has no
query to resolve — gets the inverse mapping it never had: a directory is
claimed when some roadmap phase resolves to it.

Both checks are additive: the token intersection still decides every shape
it already decided, and the resolution can only REMOVE a warning. The
regression test carries controls in the other direction — a roadmap phase
with no directory must still raise W006, an unclaimed directory must still
raise W007 — so it cannot be satisfied by a check that stopped reporting.

* docs(#2528): state and pin the directory-side scope of the bare-integer fallback

The matchPhaseDirs docblock claimed deep-decomposition lookups were
untouched. That is true of the QUERY side only — no non-bare query enters
the fallback — but the DIRECTORY side is what changed classification: a
bare `5` now reaches a lone `05-01-auth` and resolves it (phase_number
"05", phase_name "01-auth") where the base found nothing.

The widening is irreducible from directory names alone. `05-01-auth`
(sub-phase 5.1) and `30-12-factor-refactor` (phase 30 named "12-Factor
Refactor") are the same `NN-NN-<slug>` shape, and the discriminator that
would separate them — "is the second segment a valid decimal sub-phase" —
accepts `5.1` and `30.12` equally. Any rule strong enough to exclude the
first excludes the second, which is the defect #2528 exists to fix. So the
tie is broken in favour of resolving, the docblock now says so, and the
consequence is bounded where it matters: two such directories are two
matches, and every caller (including phase remove) refuses to choose.

Pins both directions, since nothing observed the directory side before.

* fix(#2528): count surviving phases by identity in phase remove's STATE resync

#2640 landed on `next` while this branch was open. Its STATE.md phase-count
resync re-derives "which directory was removed" from the query with
`phaseTokenMatches`, which is the tenth site of this issue's defect: the
bare-integer fallback resolves `05-80-20-cleanup` for query `5`, but the
token predicate does not, so the just-deleted directory is counted as still
present and the written `Total Phases` is one too high.

`targetDir` already IS the directory that was removed, and the block is gated
on it being non-null, so identity answers the question exactly — which is also
what the comment above the filter already claimed it did. This keeps
`phaseTokenMatches` out of `phase.cts` rather than re-importing it to satisfy
one call site: the module's public surface should not grow for a question that
does not need re-derivation.

Pinned in the parity gate with a control on a directory the tokenizer reads
correctly, so the assertion is about the digit-leading shape and not about the
counting rule changing for everything.

* fix(#2528): let the resolution layer own the digit-leading slug family alone

The tokenizer rewind this fix carried — pop the last absorbed continuation when
the segment that stopped the scan is a bare single digit — reads
"10-24-7-autonomy" (phase 10 named "24/7 Autonomy") correctly and silently
re-reads "10-24-7-zip" (sub-phase 10.24 named "7-Zip Integration") from "10-24"
to "10". The two names are string-identical in shape, so no local signal
separates them; the rule traded the reported ambiguity for the symmetric one a
level down, on a 15-caller chokepoint whose output also feeds query-less
derivations (STATE.md phase counts, W007, the #2562 key surface). A well-formed
sub-phase directory became unresolvable by its own id — the very symptom #2528
was filed about.

It also bought nothing. The bare-integer fallback in matchPhaseDirs already
resolves "10-24-7-autonomy" for query "10" whatever the token is: no primary
match, bare query, leading digit run "10". The reported case was covered twice,
by two rules, and the two disagreed about the case nobody reported.

So the rewind is removed rather than narrowed, in the tokenizer and in the five
surfaces kept in lockstep with it (BRACKET_PHASE_TOKEN_SOURCE,
PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem's pair grammar and its collision
branch, roadmap-parser's numericRe, extractCanonicalPlanId), together with the
SINGLE_DIGIT_RUN_SEGMENT_SOURCE owner constant they shared. Disambiguation now
lives only where a QUERY exists to disambiguate against, which is the same
bounded mechanism the "05-80-20-cleanup" shape already used.

Measured, not argued: over 29800 generated directory names, extractPhaseToken is
byte-identical to `next` on every input except the lowercase-continuation class
("01-20a", "05-80-20-25abc") — a rule about the segment itself, not a guess about
its neighbour.

Both readings now stay reachable by their own ids:
  matchPhaseDirs(['10-24-7-autonomy'], '10')    -> the dir   (fallback)
  matchPhaseDirs(['10-24-7-zip'],      '10')    -> the dir   (fallback)
  matchPhaseDirs(['10-24-7-zip'],      '10-24') -> the dir   (primary)

* test(#2528): pin the one-continuation boundary the rewind had no coverage for

The regressing shape was invisible to the suite by construction, not by luck:
the deep-rewind property built its cases from `continuationArb` with
`minLength: 2`, so it never exercised the single-continuation case — exactly one
genuine sub-phase level before a digit-leading slug — and every hand-written
fixture used the ambiguous shape only where "phase-plus-slug" was the intended
reading.

`continuationArb` is now `minLength: 1` and the property states the invariant
instead of the old rule: for any prefix, phase, 1-5 continuations and any
one-digit terminator, the token equals the FULL continuation run on both the
imperative and the regex surface, and `matchPhaseDirs([dir], token)` returns that
dir. That third assertion is the one that catches the class on its own — the old
behaviour made a well-formed directory unresolvable by its own id, which is a
property, not a fixture.

Around it: "10-24-7-zip" and "10-24-3d-printer" now sit beside
"10-24-7-autonomy" everywhere the family is pinned, so the two readings can never
diverge again; the 24/7 metamorphic property asserts the RESOLUTION result rather
than the token (the token is precisely the part no surface may decide); the
end-to-end parity corpus gains "a sub-phase with a digit-leading slug resolves by
its full id" across all four resolution paths; and the milestone-scoping residual
is pinned in three directions rather than left to prose.

Mutation: re-inserting the rewind and rebuilding turns 9 tests red, the
`minLength: 1` property first, and nothing else. Build success checked separately
(build:lib reports 0 `error TS`), so the mutation reached the artifact under test.

* test(#2528): pin the #2946 guard against digit-leading phase directories

The #2946 fix makes the milestone-complete unstarted-phase guard run
unconditionally, so whether it fires now rides entirely on the
directory-resolution owner this PR replaces. Two cases, both with STATE.md
carrying no `milestone:` field so the #2946 path is the one exercised:

  - ROADMAP Phase 5, disk `05-80-20-cleanup` → guard must stay silent.
    RED on next (fail-closed: the guard blocks a legitimate one-way-door
    operation because phaseTokenMatches resolves neither 05 nor 80 for
    that directory).
  - ROADMAP Phase 80, same directory → guard must still fire. Green on
    both sides; it pins the fail-open direction against a future widening
    of the matcher.

* fix(#3175): stop the injection scanner reading RegExp.exec as code execution

The apostrophe fix in 27aa40f6 replaced ["\x27] with a real ["'] class.
The old class never contained an apostrophe at all (POSIX bracket
expressions do not honour backslash escapes, so it was the set ", \, x,
2, 7), so only exec(" matched. Single-quoted method calls now match for
the first time, and RegExp.prototype.exec takes a subject string, not
code: any PR touching a file that tests a regex goes red. On next, six
files match the scanner's own pattern across 16 method calls.

A plain [^[:alnum:]] boundary cannot separate the two forms because . is
not alnum, so exec gets [^[:alnum:].] and the command-execution vector
moves to a dedicated member-call pattern. Bare exec('rm -rf /'),
cp.exec(...) and child_process.exec(...) all still fire.

Mutation: reverting the boundary reds 1 test and only it; removing the
member-call pattern reds the 2 non-weakening tests and only them.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(#2528): keep exec( detection receiver-blind, allowlist the two grammar suites

The left boundary [^[:alnum:].] added in 58a7b560 excluded a preceding dot,
which dropped every member-position .exec('…') from the scanner. The follow-up
receiver pattern only restored three literal spellings (child_process,
childProcess, cp), so require('child_process').exec('…') — the most common Node
spelling of the vector this pattern exists to catch — became invisible, along
with any opaque receiver (conn.exec, shelljs.exec).

Revert the pattern to its receiver-blind form and handle the RegExp.prototype
.exec false positive where the script already handles this class: per-file
ALLOWLIST entries for the two phase-token grammar suites. Mutation-checked —
removing the two entries reds exactly those two files and nothing else.

The four assertions written around the old patterns are replaced by a
table-driven set covering all six spellings, including the three the narrowed
pattern silently lost.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(#2528): pin the three undeclared grammar edges, correct the ambiguity claim

Review round 10 asked for declaration, not behavior change, on four items. All
four have zero production consumers or preserve their caller's prior rule, so
each is pinned as a test or corrected in prose rather than reverted.

- BRACKET_PHASE_TOKEN_SOURCE: the (?=-|$) terminator is what keeps the bracket
  read path in step with the other surfaces, and it costs the display shapes
  (`05.03: Title`, `12A: X`, `05.03]` no longer tokenize). Pinned so widening
  the terminator class is a deliberate act rather than a lookahead deletion.
- canonicalPlanStem: uppercase plan suffixes still strip, lowercase and dotted
  sub-plans now fall through. Dead export; pinned as a decision on record.
- getMilestonePhaseFilter: `12A-01-foo` now yields `12A-01`, matching what
  `12-01-foo` has always yielded. The letter suffix was the only reason a
  sub-phase directory folded into its parent phase's milestone window; the two
  shapes now agree. Not named in the review — found auditing the same commit.
- matchPhaseDirs docblock claimed every caller refuses on multi-match. Four do;
  five take matches[0]. Replaced the claim with the actual two-tier policy and
  the honest caveat that the bare fallback makes multi-match newly reachable
  for queries that previously found nothing.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-12 20:45:04 -04:00
sim
c6e49a5729 chore(#3331): fix stale CONTEXT.md predicates left by the no-elapsed-assertion promotion
Standards-axis code review caught 3 stale predicates (CONTEXT.md:470,
536, 541) still describing local/no-elapsed-assertion as warn and
citing the superseded epic #1885 (subsumed into #3053 and closed
stale). Regenerated the derived CONTEXT-INDEX.json snapshots.
2026-08-12 16:53:05 -04:00
sim
23e26a9de8 chore(#3339): extend BUG_FILE_RE to fix-/issue- prefixes, close H3 (#3315)
Widens the identity ratchet in scripts/lint-regression-test-names.cjs
from banning only new bug-NNNN-*.test.cjs files to also banning
fix-NNNN-*.test.cjs and issue-NNNN-*.test.cjs, per epic #3053's own
recorded decision: extend only after the 74-file fix-*/issue-* backlog
fully drains, never before.

That precondition is now real, not assumed: this branch is rebased
onto origin/next with Waves 1-6 (#3341, #3342, #3373, #3376, #3378,
#3383) all merged, and a repo-wide check confirms zero tests/fix-*.test.cjs
or tests/issue-<N>-*.test.cjs files remain -- only the two known false
matches (issue-dedupe.test.cjs, issue-version-gate.test.cjs, feature-named
suites with no digits after the prefix) survive the glob, and the widened
regex correctly excludes both.

Allowlist snapshotted at zero (scripts/lint-regression-test-names.allowlist.json
was already [] and stays [] -- confirmed via --update reporting "already
in sync"). Updated docs/TESTING-SUITES.md's Regression tests section to
name fix-/issue- alongside bug- as banned new-file prefixes.

This is the last commit of H3 (#3315)'s 7-wave decomposition.
2026-08-12 08:28:17 -04:00
sim
1bc7f7e6b0 test(#3339): fold the state/phase/dispatch & model-profile issue-* cluster — Wave 7
Folds 9 legacy issue-*.test.cjs regression files (140 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
LAST of 4 issue-* waves — closes out the 74-file fix-*/issue-* backlog
(pending BUG_FILE_RE extension, held for a follow-up commit until Wave 6
is confirmed merged, per the epic's own zero-backlog precondition).

- issue-2828-flat-roadmap-total-phases.test.cjs (1) + issue-3204-state-
  writer-phase-count.test.cjs (21): both target state-document.cjs
  buildStateFrontmatter via different CLI entrypoints — merged jointly
  into state-document.test.cjs, 0 dropped.
- issue-2945-phase-complete-checkbox-rollback.test.cjs (4) + issue-2949-
  phase-complete-stage3-sentinel.test.cjs (4): both target phase.cts
  cmdPhaseComplete; issue explicitly warned of overlap — verified
  disjoint fixtures/assertions, 0 dropped, merged into phase.test.cjs.
- issue-2927-reviewer-lane-overlay-invocation.test.cjs (10) merged into
  review-lane-descriptor.test.cjs.
- issue-2939-dispatch-flatten-maxdepth.test.cjs (9) merged into
  host-integration.test.cjs, 2 dropped as verified exact duplicates.
- issue-2977-frontmatter-bom.test.cjs (5) merged into frontmatter.test.cjs.
- issue-2045-third-party-skills-surface.test.cjs (6) merged into
  capability-loader.test.cjs.
- issue-2517-runtime-aware-profiles.test.cjs (80, the largest single
  fold in the epic) merged into model-resolver.test.cjs, 1 dropped as a
  verified true duplicate (checked against src/model-resolver.cts logic,
  not just title similarity).

Fixed a genuine eslint irregular-whitespace finding: a literal BOM
character embedded in a doc comment (pre-existing content from the
original #2977 source, illustrating what a BOM looks like) — replaced
with a readable U+FEFF notation.

3 stale doc references found and fixed (docs/adr/2313, 3180, 443).

Zero net test-coverage loss. No production code changed.
2026-08-12 08:25:14 -04:00
Tom Boucher
4a0b5e26b4 test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6 (#3383)
* test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6

Folds 10 legacy issue-*.test.cjs regression files (75 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
Third of 4 issue-* waves.

- issue-2701-nul-corrupted-validators.test.cjs (9) + issue-429-comment-
  text-gate.test.cjs (31, incl. fast-check property tests): both target
  verify.cjs/validate.cjs via different calling styles (CLI vs direct-
  require) — merged jointly into verify.test.cjs, 0 dropped.
- issue-2762-plan-reviews-chunked.test.cjs (3) merged into
  plan-phase-drift-guard.test.cjs.
- issue-2771-advisor-subagent-type.test.cjs (1) + issue-2772-discuss-
  phase-text-inconsistencies.test.cjs (6) merged jointly into
  discuss-phase-power.test.cjs.
- issue-498-update-backup-runtime-dir.test.cjs (3, rename basis) +
  issue-815-update-next-channel.test.cjs (7, merged in): both concern
  update.md workflow-text contracts, now update-workflow.test.cjs.
- issue-498-update-context.test.cjs (13): pure rename to
  update-context.test.cjs, sole comprehensive suite for its module.
- issue-2765-brace-expansion-lockfile.test.cjs (1, rename basis) +
  issue-3238-js-yaml-lockfile.test.cjs (1, merged in): two distinct CVE
  regression pins against package-lock.json, now lockfile-cve-audit.test.cjs,
  shared ROOT/npmLs helpers deduped instead of double-declared.

Ratchet upkeep to keep this wave's own gates green: pruned 3 stale
allow-test-rule allowlist entries, cited 2 previously-uncited comments
that surfaced in the folded content (#3338), tightened the exemption-file
ceiling 305 -> 297. Fixed one stale filename reference in production code
(src/init.cts) plus one in docs/reference/workflow-fragments.md.

Zero net test-coverage loss. No production code BEHAVIOR changed.

* test(#3338): fix orthogonal-review findings — Wave 6 fold

Standards-axis review found a real structural defect in two files, both
the same root cause and both fixed here:

- tests/update-workflow.test.cjs: the folded:issue-815-update-next-channel
  wrapper's closing brace was placed after the file's pre-existing tail
  instead of before it, making the already-established folded:bug-2470
  and folded:bug-3130 wrappers CHILDREN of issue-815's block in the test
  hierarchy instead of independent siblings — confirmed via an actual
  node --test run showing the mislabeled TAP nesting. Moved the closing
  brace to the correct position; all three fold wrappers are now
  top-level siblings again (verified via node --test, TAP hierarchy
  correct, 10/10 tests, 5 suites, identical count before and after).
- tests/plan-phase-drift-guard.test.cjs: same mistake in the other
  direction — folded:issue-2762-plan-reviews-chunked was spliced inside
  the pre-existing folded:bug-2492-context-coverage-gate wrapper instead
  of after it. Fixed the same way (227/227 tests, 38 suites, identical
  count before and after).
- Reverted unnecessary 815-suffixed local renames (assert815/fs815/etc.)
  introduced by the fold — the wrapper is genuinely block-scoped once
  correctly closed, so no collision existed (same class as Wave 4's
  __foldSetNested finding).

No test() count changed in either file. No production code touched.

---------

Co-authored-by: sim <sim@local>
2026-08-12 00:51:19 -04:00
Tom Boucher
aee83c8e9e test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 (#3378)
* test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5

Folds 6 legacy issue-*.test.cjs regression files (120 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
Second of 4 issue-* waves.

- issue-766-plugin-manifest.test.cjs (50 tests): pure rename (git mv) into
  plugin-manifest.test.cjs, sole comprehensive suite for its module.
- issue-844-manifest-version-sync.test.cjs (14 tests, rename basis) +
  issue-1855-marketplace-manifest.test.cjs (17 tests, merged in): both
  target scripts/sync-manifest-versions.cjs via non-overlapping describe
  blocks (generic engine vs. marketplace.json-specific), now
  manifest-version-sync.test.cjs, 31 tests, 0 dropped.
- issue-498-identity-drift-lint.test.cjs (8 tests, rename basis) +
  issue-498-package-identity.test.cjs (17 tests, merged in): both target
  package-identity.cjs's surface from different angles (lint-side drift
  detection vs. derive/slugify), now package-identity.test.cjs, 25 tests,
  0 dropped.
- issue-607-cache-lineage.test.cjs (14 tests) merged into the existing
  gsd-statusline.test.cjs, matching its already-established fold-wrapper
  convention from prior waves.

Learned from Wave 4: fold agents checked src/** (not just docs/ and
gsd-core/references/) for stale filename references — found and fixed 5
across 3 ADR docs (766, 2121, 457), zero in src/ this time.

Zero net test-coverage loss. No production code changed.

* test(#3337): fix orthogonal-review findings — Wave 5 fold

Standards-axis review found real issues in the just-folded
manifest-version-sync.test.cjs, both fixed here:

- The folded:issue-1855-marketplace-manifest block omitted its own
  test/describe/assert/fs/path/os requires, silently closing over the
  #844 basis section's outer-scope bindings instead of declaring its own
  dependencies — inconsistent with every other fold block in this wave.
  Restored the local requires the original issue-1855 file declared.
- Merging two independently-lettered legacy files (each A-F) left
  duplicate top-level describe() labels (two "A:", two "B:", two "C:").
  Renamed the folded-in #1855 labels (A2/B4/C2) to make every top-level
  label in the file unique.

No test() count changed (31). No production code touched.

---------

Co-authored-by: sim <sim@local>
2026-08-11 23:55:11 -04:00
Tom Boucher
ad07f76a31 test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 (#3376)
* test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4

Folds 10 legacy issue-*.test.cjs regression files (79 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
First of 4 issue-* waves (following the 3 fix-* waves, all merged).

- 1 file with no prior target coverage: renamed (git mv) into
  legacy-cleanup.test.cjs (sole comprehensive suite for that module).
- 9 files merged into 6 pre-existing suites: golden-parity-single-source,
  runtime-artifact-layout-surface, codex-config (4 sources merged jointly
  in one pass per the issue's own instruction, to catch overlap between the
  4 sources themselves, not just against the pre-existing target — zero
  overlap found, all 20 blocks additive), runtime-config-adapter-registry
  (1 of 10 source blocks dropped as a proven subset of existing coverage),
  cline-install, install.test.cjs.

Incidental fixes required to keep this wave's own ratchets green:
- Fixed a stale ADR doc reference (docs/adr/1235) to a folded-away filename.
- scripts/lint-allow-test-rule-refs: pruned 4 stale allowlist entries for
  renamed/merged-away files, cited 2 previously-uncited allow-test-rule
  comments that surfaced as "new" only because their file path changed,
  added 1 fresh allowlist entry for a pre-existing uncited comment that
  predates this PR, and tightened the exemption-file ceiling 309 -> 305
  to match the real post-fold high-water mark.

Zero net test-coverage loss. No production code changed.

* test(#3336): fix orthogonal-review findings — Wave 4 fold

Standards-axis review + Memtrace graph pass found real issues in the
just-folded suites, all fixed here:

- Standardized the fold-wrapper convention (block-scoped __foldDescribe)
  across golden-parity-single-source.test.cjs, runtime-artifact-layout-
  surface.test.cjs, runtime-config-adapter-registry.test.cjs, and
  cline-install.test.cjs to match the pattern already used by
  codex-config.test.cjs and install.test.cjs in this same wave (and by
  earlier folds elsewhere in the epic) — repeats the exact inconsistency
  Wave 3 (#3335) already fixed once in this epic.
- Fixed a stale allowlist entry's alphabetical position (cosmetic, not
  tool-gated, caught by review anyway).
- Fixed two stale test-filename references in PRODUCTION code comments
  (src/capability-writer.cts, src/runtime-config-adapter-registry.cts)
  caught by lint-removed-but-needed — a class of stale reference this
  wave's fold agents didn't check for, since they were scoped to docs/
  and gsd-core/references/ only, not src/. First fix attempt wrongly
  edited the gitignored gsd-core/bin/lib/*.cjs BUILD OUTPUT instead of
  the tracked .cts source; caught and corrected before commit.
- Fixed one remaining stale doc reference in docs/adr/1235 (a prior
  partial fix in this same wave missed it).

No test() count changed in any file. No production code BEHAVIOR
changed — comment-only fixes in src/.

---------

Co-authored-by: sim <sim@local>
2026-08-11 23:35:44 -04:00
0xdhx
0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00
Behruz Nassre Esfahani
2076d450d7 fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name

quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"`
worktree gate, so every non-Claude runtime failed closed regardless of the
capability it negotiated — including Codex, which declares
orchestrator-worktree. Route both through the negotiated dispatch.isolation
seam via a new shared reference, and migrate the two execute-phase reference
fragments that carried the same runtime-name gate.

- new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION
  resolution, harness-flag resolution, single-agent degrade rule
- quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag}
  placeholder rather than a hardcoded isolation="worktree"
- execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate
  [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ]
- every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one
  dispatched an isolated agent with no base guard and no manifest
- parity guard in host-integration.test.cjs scans workflows AND references and
  matches six reintroduction shapes
- migrate four tests that pinned the pre-#2584 runtime-name contract

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages

The degrade warnings cited /gsd-execute-phase, the retired hyphen form that
slash-command-namespace.test.cjs rejects in Claude-facing source. Same length,
so the quick.md size budget is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2652): add changeset for PR #2728

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): normalize dispatch-site paths to forward slashes for Windows

path.relative() returns backslash-separated paths on Windows, so the
#2652 dispatch-site parity test compared "gsd-core\workflows\quick.md"
against the hardcoded forward-slash literal "gsd-core/workflows/quick.md"
and failed on every windows-latest CI lane. Normalize with
.replace(/\\/g, '/'), matching the existing convention used elsewhere in
this suite (e.g. tests/branch-no-track-guard.test.cjs:37).

* test(#2652): restore the size-growth acknowledgment

The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the
ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md
(+2086) and quick.md (+230) still need an ack naming them and saying why.

Verified: 65/66 without it (both files named), 66/66 with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate

Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal —
gated only on `workflow.use_worktrees`, with no capability negotiation at
all. It is the same defect #2652 fixes at the other four sites, just a
different shape: the file contains no RUNTIME variable, so the new detector
correctly does not flag it.

Concrete break: a Codex user who follows this PR's own newly-documented
pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via
/gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an
unconverted path — either an Agent() call erroring on an unrecognized
parameter or silent unisolated execution, depending on host tolerance.

Pattern A is a single-agent dispatch site through the host's own subagent
tool, so it takes the same treatment as quick.md and diagnose-issues.md:
resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to
sequential on orchestrator-worktree hosts, and substitute the host's declared
{harnessFlag} instead of Claude Code's literal.

while the area was open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT

Two bookkeeping gaps flagged in review:

INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md.
INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a
live directory scan against the committed manifest, so CI passed regardless —
but gen-inventory-manifest.cjs's own stderr guidance says to add the matching
INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed
with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch)
rather than alphabetically, matching how that table is grouped.

CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation
as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of
#2584". That was already stale before this PR (execute-phase graduated in Phase 3)
and more so now with three single-agent dispatch sites consuming it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): detect reversed-operand runtime gates; add a permutation property

All five reintroduction regexes assumed $RUNTIME on the LEFT of the
comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written
backwards — evaded every one of them. Verified against the old patterns
before fixing: all four reversed shapes (single bracket, double bracket,
test builtin, JS template) scored EVADED.

Each comparison shape is now generated in both operand orders from a single
template, so a shape cannot be added in one order and forgotten in the other.
The mutation table gains the four reversed cases.

Also adds the fast-check property review suggested in place of the hand-rolled
cases: it generates the cross product of the axes an author actually varies —
bracket form, operator, operand order, quoting, spacing, runtime id — so a
permutation the hand-written patterns miss surfaces here rather than in
production. The 11 explicit cases stay as named regression anchors.

execute-plan.md joins the scan's required-identities list now that it is a
converted dispatch site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): acknowledge the execute-plan.md size growth

The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the
size axis needs its own acknowledgment, independent of attribution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase

The #2751 command-position gate pins its prose exemptions by line number.
This branch inserts the dispatch-isolation resolution above the
`validated downstream by gsd-tools uat classify-coverage` sentence, moving
it from execute-plan.md:387 to :397 — which fired the gate twice for one
displacement (an un-allowlisted mention at 397, a stale entry at 387).
The prose itself is unchanged from next; only the pin moves.

Fixes #2652

* fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md

The rebase onto next merged #2649's pre-dispatch base-check textually, but
its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so
the degrade never reached the dispatch decision. Gate the block on
ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same
pairing quick.md already uses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal

Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and
its skip clause (l.839) all conditioned on the literal isolation="worktree" —
Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a
newly-unblocked isolated Cursor run created a worktree whose committed work
was never merged back and never cleaned up, silently. All three now key on
ISOLATION = "harness-worktree" at dispatch.

The existing parity detector cannot catch this class (its ISOLATION_TOKEN
treats the literal as a legitimate marker), so this adds a dedicated
literal-condition detector with a discrimination proof against both pre-fix
sentences, a benign-mention control, and a positive pin on all three
re-keyed conditions. Verified fail-first against the pre-fix quick.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes

`_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's
`workflow.use_worktrees` read to `--default false`. That default resolved
before `gsd_run query dispatch-isolation` was ever consulted, so the five
runtimes that declare worktree support — cursor (harness-worktree) and
codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none
regardless of what they negotiated. The gate this PR migrates dispatch onto
was therefore still deciding isolation by runtime name, one layer down.

The stamp's #1521 premise was that worktree isolation *was* Claude Code's
isolation="worktree" spawn parameter, which no other host honored. #2584
replaced that premise with the negotiated capability. The stamp is now scoped
to runtimes whose negotiated isolation really is `none`, where the default it
writes is the outcome the resolver reaches anyway.

`_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution
against the same registry — closed vocabulary, a harness-worktree host must
declare its flag, an orchestrator-worktree host must carry a descriptor that
resolves — and fails closed to `none` on anything else, so an undeclared or
unknown runtime keeps today's behavior.

Two #1515 tests pinned the superseded premise for codex and are re-pointed at
the new contract rather than deleted: the safety property they protect is now
held by the isolation gate's fail-closed resolution, not by a name-scoped
install-time default. Verified fail-first — all five assertions red against
the pre-fix source, green after.

* test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof

Scoping the use_worktrees stamp changes emitted output, and two gates caught it.

`gsd-core/workflows/execute-phase.md` now differs at emit time for the five
hosts that declare worktree support (cursor harness-worktree; codex, opencode,
kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only
the stamp is gone. Acknowledged in this PR's fragment.

`tests/install.test.cjs`'s real-install assertion pinned the superseded premise
end-to-end, asserting codex receives `--default false`. Re-pointed rather than
deleted, matching the two unit tests: it now proves codex keeps the unstamped
`true` read. A second arm installs windsurf — which declares isolation `none` —
and asserts the false stamp is still applied there, so the change cannot
silently degrade into "never stamp" without a test noticing.

The ack entry collides with `2658-trae-instruction-file-path.json`, which is
fully spent (merged via #2925, so all 25 of its entries are present at base and
gate nothing) and is pruned for the same reason and by the same rule as the
spent `2649-*` fragment this PR already removed. #2566 prunes the same file for
the same collision on `new-project.md`; a delete/delete merges cleanly either
way, and the base-side cleanup would make both unnecessary.

* fix(#2652): re-record the sentinel when a dispatch site degrades isolation

Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided
in shell, where routeDispatchIsolation cannot see it. That resolver persists
whatever it resolved to the run-scoped sentinel as an unconditional side effect
(#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel
asserting harness-worktree while the dispatch correctly omits the harness flag.
The shipped PreToolUse guard reads the sentinel at the instant of the Agent()
call and denies that mismatch with exit 2 — the work does not run unisolated,
it does not run at all. Latent on this branch and lands on rebase, since
8f75e275 (#3045) is not yet in the merge-base.

Four sites now push the final shell-computed value through the same single
write path with --force-isolation, matching the idiom #3045 established in
executor-isolation-dispatch.md:

  - quick.md, after the #1941 base-check degrade
  - diagnose-issues.md, after the config-gate degrades and after #2649's
  - execute-plan.md Pattern A, before spawning
  - references/dispatch-isolation-gate.md, both degrade paths, plus a new
    "Re-record after every degrade" section — the canonical file taught the
    defect, so fixing only the call sites would leave the source of truth wrong

Tests assert the RECORDED value, not $ISOLATION. Asserting the local variable
is what let this class through: $ISOLATION was already `none` at every site and
the defect was entirely in what reached the sentinel. Each workflow's own
degrade block is executed under a gsd_run stub that captures the write, with a
fail-first proof that the pre-fix shape records nothing (while $ISOLATION reads
`none` in both), plus a coverage guard so a new degrade site cannot skip it.

Also corrects the drift-ack rationale (review Minor 5): @-references are
eagerly inlined, so extracting the gate does not reduce loaded context. The
reason to extract is single-sourcing across five dispatch sites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): satisfy the new CRLF-portability and cleanup lint rules in host-integration.test.cjs

next's local/no-crlf-fragile-split and no-raw-rmsync-in-tests rules now
cover the fenced-block regexes, log-line split, and temp-dir removal this
suite added: bash-fence matchers and line counting accept \r\n, and the
raw fs.rmSync becomes helpers.cleanup (Windows-EBUSY retry budget).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): restore next's 2658 ack fragment, minus the one colliding key

trek-e (PR #2728, 2026-08-07): the branch deleted
tests/emitted-drift-acks/2658-trae-instruction-file-path.json wholesale
while next had modified it. That was correct against the 08-03 base, where
the fragment was fully spent; it is wrong against next @ 1d208e5a, which
still carries 23 live entries.

next's copy is restored byte-identical except for the single key that
genuinely collides with this PR's own fragment,
gsd-core/workflows/execute-phase.md. Both acks name that path for
different deltas -- 2658's is the trae CLAUDE.md replacement-target
rewrite, ours is the emit-time _stampNonClaudeRuntimeDefaults ripple from
review round 3. Per the ack-lifecycle law (#2789), an entry already at the
base is spent and inert, so this PR's entry is the live one and 2658's is
dropped.

This follows the guidance given on #2566 in the 08-06 round: "Regenerate
rather than delete -- the collision is one entry."

Verified: lint-emitted-drift-ack ok (0 problems, 357 keys, no cross-source
duplicates); emitted-attribution 170/170 with GSD_EMITTED_BASE=upstream/next;
host-integration 222/222; runtime-converters 130/130.

* fix(#2652): bound the degrade-harness spawn and close the round-6 majors

B1 (CI red, ours): tests/host-integration.test.cjs spawned bash with no
timeout, violating local/no-unbounded-spawn. `next` deleted the allowlist
outright (#3148), so the merge-commit run flags it even though this branch
still carries the file's grandfathered entry. Bounded at 15s, with a named
failure on timeout/signal rather than an opaque `exited null`.

M1: add a parity test between `_negotiatedDispatchIsolation` (install time)
and `routeDispatchIsolation` (dispatch time). Both read the same capability
registry and the same `resolveOrchestratorExec`, but duplicate the DECISION
on top of them across two surfaces with no call edge between them, so neither
symbol appears in the other's impact graph and nothing static can catch them
drifting apart. The resolver leg drives the real gsd-tools CLI per registered
runtime, both ways it is really called: with `--cwd-target` (the executor
spawn, which resolves the orchestrator descriptor — the same question install
time asks) and without it (the `Resolve ISOLATION` call every dispatch site
makes first, which does not). The second leg is what catches an orchestrator
host whose descriptor stops resolving: the install would stamp
`use_worktrees=false` while the workflow gate still reported
`orchestrator-worktree`.

M2: add install-level Cursor coverage. A real `--cursor` install, then the
gate blocks that install emitted, run against the gsd-tools that install
emitted, with the runtime declared through `.planning/config.json` — the tier
`resolveRuntime` actually reads — and any ambient GSD_RUNTIME blanked, so the
install has to reach the right resolver on its own. It then performs the
documented `{harnessFlag}` substitution against the `Agent()` call that
install emitted and asserts on the rendered dispatch: exactly one emitted
Agent() call carries the slot, it is the gsd-executor / gsd-debugger dispatch
rather than some other call in the same file, and rendering it yields
`--worktree` with no residual placeholder and no `isolation="worktree"`.
This is artifact-level — it proves the emitted wiring, not a live Cursor
host invocation. Asserting the shell variable alone would have stayed green
if the placeholder were deleted from the emitted dispatch, or drifted onto
the reviewer call beside it.

M3: diagnose-issues.md inlined a reordered copy of the reference this PR
introduces as the single source of truth. It now reads the reference the
same way quick.md and execute-plan.md do; the drift-ack entry is corrected
to describe what the file actually contains, and to name the four files that
reference the gate rather than claiming five.

M4: CONTEXT.md still called `resolveOrchestratorExec` UNCONSUMED in the same
paragraph this PR edits. It has been consumed since #2584 Phase 3 — routed
through `query dispatch-isolation --json` and process-spawned by
executor-isolation-dispatch.md — and #2652 adds a second consumer.

Every new assertion verified fail-first against a real mutation: cursor's
negotiated isolation (breaks the target-bound parity leg), the no-target
orchestrator branch in routeDispatchIsolation (breaks the gate parity leg),
cursor's harnessIsolationFlag (breaks the resolved value), deleting
`{harnessFlag}` from quick.md's emitted Agent() call (breaks the slot), and
moving it onto the code-reviewer dispatch (breaks the wrong-call guard).

Refs #2652

* fix(#2652): serialize unisolated diagnosis, scope the execute-plan gate to dispatching patterns

Round-7 review findings (independent cross-AI pass over the whole PR against
the current base).

BLOCKER — diagnose-issues.md announced sequential mode and then fanned out.
The `orchestrator-worktree` degrade sets ISOLATION=none and prints "debug
agents run sequentially on the main working tree", but the spawn step still
said "All agents spawn in single message (parallel execution)". On Codex,
OpenCode and Kimi that dispatched N unisolated debuggers concurrently against
the primary checkout — the exact outcome the degrade exists to prevent, and
reachable only because this PR removed the FATAL that used to stop those
hosts earlier. Fan-out is now keyed on ISOLATION: parallel only when each
agent has its own worktree, one at a time otherwise.

BLOCKER — execute-plan.md Pattern B could not dispatch at all on Claude or
Cursor. The gate recorded `harness-worktree` to the #3045 sentinel, but only
Pattern A carries `{harnessFlag}`; Pattern B's segment executors carry none,
and `hooks/gsd-agent-isolation-guard.js` blocks precisely that mismatch with
exit 2. Segments are unisolated BY DESIGN — each continues on the working
tree the previous one left behind, so per-agent worktrees would break the
sequence — so Pattern B now records `none` before its first dispatch and
dispatches without the flag.

MAJOR — the same gate ran before routing was chosen, so an isolation-`none`
host with `use_worktrees=true` hit the fail-closed FATAL even when routing
would have selected Pattern C, which is fully inline and dispatches nothing.
Resolution now happens after the pattern is known, and Pattern C skips it.

MAJOR — tests/host-integration.test.cjs fed `fs.readFileSync` output straight
to bash. The `\r?\n` fence regex guards only the delimiter, leaving embedded
CR on every line of the captured body — DEFECT.WINDOWS-CRLF-TEST-PORTABILITY,
which helpers.cjs documents by name. Now reads through `readFileNormalized`.

MAJOR — the "every dispatch-site degrade block re-records" test hand-listed
three files, so its name was a claim its scan could not support. The scan is
now derived from the workflow/reference tree (SCAN_ROOTS/collectMarkdown
hoisted to module scope so there is one definition, not two). Verified
fail-first against execute-plan.md — a file the previous scan never opened.
The two wave fragments are exempt because they delegate the re-record to
per-plan-worktree-gate.md via USE_WORKTREES_FOR_PLAN; that delegation is now
ASSERTED, so deleting the delegate fails this test instead of widening a hole
silently.

MINOR — the changeset claimed the FATAL was gone for "non-Claude runtimes"
full stop. Narrowed: isolation-`none` hosts still fail closed when worktrees
are explicitly enabled, which is the contract rather than the defect.

Two further findings were investigated and rejected, with evidence:
- Raw `spawnSync` vs `tests/helpers/process-seam.cjs`: the seam exposes
  runNode/runGit/runHook and cannot express the `bash -c` harness these tests
  need; `installAndRead` in this same file is byte-identical to the base and
  still uses raw spawnSync with an explicit timeout, which is the form the
  lint sanctions. Migrating only the new call sites would split the file's
  convention for no safety gain.
- `pending-migration-to-typed-ir` on the runtime-converters parity test: the
  annotation and the rendered-text loop both exist at the merge-base under
  #3090. This PR extends an already-tracked test rather than adding a new one
  under a category CONTRIBUTING closes to new tests.

Refs #2652

* test(#2652): re-point the execute-plan prose allowlist at its shifted line

`PROSE_ALLOWLIST` in tests/no-bare-gsd-tools-command-position.test.cjs keys
entries by LINE NUMBER. The previous commit added the post-routing isolation
block to execute-plan.md, which pushed the `validated downstream by
gsd-tools uat classify-coverage` prose mention from line 397 to 414. That
broke the guard in both directions at once: the entry at 397 went stale, and
the real mention at 414 became an unallowlisted offender.

Caught by CI (7 red jobs, all shard 3/3 plus ubuntu-22) rather than locally,
because I verified only the suites I believed the change touched. Any edit to
a workflow .md shifts line numbers, and this repo carries line-keyed
allowlists — so a workflow edit needs the full suite, not a subset.

Refs #2652

* test(#2652): route the new subprocesses through the process seam

Retracting a rejection I made on the record. In the round-6 response I
argued these call sites could keep a hand-rolled `spawnSync` because the
seam exposes only runNode/runGit/runHook and cannot express `bash -c`, and
because `installAndRead` in the same file uses that shape. The first half
was true and irrelevant, the second half is not a licence: CONTRIBUTING is
unambiguous — "Anything that shells out goes through
tests/helpers/process-seam.cjs — never a hand-rolled spawnSync/execFileSync
in your suite", and "Never use try/finally inside test bodies."

`runHook` already documents `interpreter: 'bash'` for running a shell
script, so writing the harness to a file complies without extending the
seam. I had the rule and the seam's own documentation in front of me and
reasoned around both.

Converted:
- host-integration.test.cjs degrade harness: spawnSync('bash', ['-c', …])
  → runHook(scriptFile, [], { interpreter: 'bash' }).
- install.test.cjs cursor gate: `which bash` probe → process.platform;
  the installer spawn → runNode(…, { env: installSpawnEnv({HOME,
  USERPROFILE}) }), which also blanks ambient GSD_HOME/runtime-location
  vars that could otherwise make capability discovery host-dependent;
  the emitted-gate spawn → runHook(gateScript, [], { interpreter: 'bash' }).
- Three try/finally test bodies → t.after().

Class-norm timeouts: tests/helpers/timeouts.cjs arrived with this branch's
latest base merge, so the literals written earlier (15000/120000/60000) now
duplicate PROBE_TIMEOUT_MS and INSTALL_TIMEOUT_MS. Imported instead — that
module exists because INSTALL_TIMEOUT_MS had already drifted 60s→120s once
after a real bench ETIMEDOUT.

Deliberately NOT converted: `installAndRead`'s spawnSync, which is
byte-identical to the merge-base and predates this PR — converting shared
scaffolding is an unrelated change.

Verified equivalent, not assumed: argv/cwd/env/encoding/timeout and every
assertion are preserved; t.after() still cleans up on the assertion-failure
path the try/finally covered; and the cursor test still resolves Cursor
under a hostile ambient GSD_RUNTIME=claude.

Refs #2652

* fix(#2652): replace the falsified use_worktrees doc row; distinguish an unresolvable gate from a declared none

Round-8 review findings.

BLOCKER — docs/CONFIGURATION.md's `Non-Claude note` asserted three things
this PR overturns: that worktree isolation "no other runtime honors"
(Cursor declares harness-worktree with `--worktree`, and this PR's own
install test asserts that flag reaching the emitted Agent() slot), that
non-Claude installs default the key to `false`, and that forcing `true`
always fails closed. Replaced with the capability-based description, and
the `#1515, #1521` citation dropped — those are the two issues whose
premise this PR removes.

The reviewer flagged that the fix is merge-order dependent, because #2531
rewrites the same row and its replacement text is written in anticipation
of this PR landing. Rather than pick an order, BOTH sides are now
order-independent: #2531's "Current default … until #2652" paragraph
becomes a plain troubleshooting note, and this row states the capability
rule without asserting a stamp state. Whichever merges first, the row is
correct; the second merge is a textual conflict at worst.

MINOR — the gate reported a capability verdict the tool never returned.
`ISOLATION=$(… || echo "none")` made a shim-resolution failure, a non-zero
exit and an empty stdout indistinguishable from a declared `none`, so a
transient query failure aborted /gsd:quick on Claude Code with "runtime
'claude' declares no executor-isolation primitive" — false. The gate now
tracks ISOLATION_RESOLVED separately: both paths still fail closed, but only
a real verdict claims the host declares nothing; the unresolved branch says
it could not resolve and points at the shim. Fixed in the canonical
reference so every dispatch site inherits it.

MINOR — quick.md:527 cited #2649 for its own degrade; that is #1941, and
#2649 is the diagnose-issues/execute-plan gate. Corrected, and the
distinction stated so the next reader does not chase it.

MINOR — quick.md:413 (manifest init) and :429 (worktree_branch_check embed)
still branched on USE_WORKTREES while dispatch, manifest-append, merge-back
and the skip clause had all moved to ISOLATION. Safe only by coincidence —
both now key on ISOLATION.

MINOR — the diff removes a second drift-ack entry (the execute-phase.md key
from 2658-trae-instruction-file-path.json), forced by the same duplicate-key
lint rule as the 2649 removal. Disclosed in the PR comment; the earlier
disclosure covered only one of the two.

Verified: lint:ci green; 300/300 across host-integration,
fix-1941-quick-worktree-stale-base, execute-phase-wave and workflow-guard.

Refs #2652

* test(#2652): anchor the emitted-gate finder on the heading, not the assignment

`b89c3fbf` added a finder that located the gate's `Resolve ISOLATION` block by
the literal `ISOLATION=$(gsd_run query dispatch-isolation --raw`. `f3bccf21`
then split that assignment into `_ISOLATION_RAW`/`ISOLATION_RESOLVED` so a shim
failure stops masquerading as a declared `none` — and the finder stopped
matching. The test did not report the drift it exists to catch; it reported
"emitted dispatch-isolation-gate.md has no Resolve ISOLATION bash block" and
went red, and stayed red because the earlier full-suite run was read from a
truncated log.

Anchored on the heading instead. The workflows tell a dispatch site to run the
`Resolve ISOLATION` and `Resolve the harness flag` blocks BY NAME, so the
heading is the contract and the body is free to change under it.

Verified: 413/413 in tests/install.test.cjs. The test still bites — mutating the
gate's `ISOLATION="$_ISOLATION_RAW"` to `ISOLATION=none` turns it red (the
emitted gate then resolves cursor to none and exits 1 instead of printing
harness-worktree), and reverting restores green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): wire the canonical resolver at the one dispatch site that still inlined the old shape

Codex review of the whole PR on the new base found one Major, and it was real.

executor-isolation-dispatch.md declares references/dispatch-isolation-gate.md
canonical at line 10, then kept the OLDER resolver inline: `|| echo "none"`, no
ISOLATION_RESOLVED. So the one site that resolves isolation for the wave path
collapsed a shim failure into a declared `none` and aborted with "runtime
'$RUNTIME' declares no executor-isolation primitive" — false for a Claude or
Cursor user whose resolver merely failed to answer. Still fail-closed, so not an
unsafe-dispatch hole, but the correction this PR is about was unwired at the
site that matters most.

Replaced with the gate's exact shape: capture the raw value, track
ISOLATION_RESOLVED, and emit the "could not resolve" FATAL when no verdict was
learned.

Added a regression test in the #2652 dispatch-site parity suite: every file that
ASSIGNS from `gsd_run query dispatch-isolation --raw` must carry
ISOLATION_RESOLVED, must not use the collapsing form, and must have a distinct
unresolved message. Nothing covered this before — install.test.cjs checks the
emitted REFERENCE, not each site's own inline copy, which is exactly how the two
drifted apart.

The test's first draft also flagged quick.md, diagnose-issues.md and
execute-plan.md. That was a false positive worth recording: those three
@-reference the gate and only make `--force-isolation` re-record calls, which
carry no verdict. The predicate now matches an assignment from the resolver, not
any mention of it, so it flags sites that can actually be wrong.

Verified by mutation: restoring the collapsing line reds the new test.

Validated: lint:ci clean; full suite shows the same 7 known failures as the
pre-change baseline — #1160 _resolveManifest and the #3053 quick_id
host-timezone tests (both reproduce on pristine next @ 33fca50d), plus
helpers-cleanup "outside os.tmpdir()", which fails only in a worktree.

* test(#2652): close two vacuous-pass holes in the new inline-resolver guard

Codex cleared the push and flagged the guard test itself. Both holes were real.

SCAN_ROOTS already yields references/dispatch-isolation-gate.md, and the test
appended it a second time, so the candidate list was [executor, gate, gate] and
`length >= 2` was satisfiable by the gate alone. If the executor site had
dropped out of the predicate — the exact regression the test exists to catch —
it would still have passed. Paths are deduped and the assertion now pins the two
expected inliner identities instead of a count.

The collapse detector keyed on `ISOLATION=$(…)`, so `_ISOLATION_RAW=$(… || echo
"none")` restored the identical defect while satisfying every other assertion
(ISOLATION_RESOLVED still appears in the file). Codex mutation-probed exactly
that and it passed. The pattern now matches any assignment target and any
`|| … echo` tail. Verified: that mutation now reds the test.

Test-only change; the workflow bash is byte-identical to the commit the full
suite ran green against. lint:ci clean, host-integration 223/223.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:10:28 -04:00
Rezolv
e87fb409ee enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint

Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age
freshness proxy (state_commits_behind / state_commit_stale) through
state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as
health W024. The proxy is advisory: classify() deliberately does NOT consume
it (ADR-1787 locks the classification/routing boundary — a signal, not a route).

Composes with #3099 and #1882 (both merged to next after this branch): the
commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE
diagnostic reads `last_activity` — two different fields, not "two staleness
signals on one field." A new regression test asserts a STATE.md carrying both
an unparseable last_activity AND a valid state_head resolves each independently
(diagnostic fires once; freshness reads state_head, commits_behind 0).

Rebased onto next (flattened): resolved the add/add conflicts in
src/smart-entry.cts (kept both the #2573 freshness import/derivation and the
#3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both
describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B).
Tests: smart-entry 62, state/state-transition/health/verify 639, all pass.

* chore(#2573): allowlist health-validation test in the prompt-injection scan

The scanner's `exec('` code-execution pattern matches the benign
`re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests
(pre-existing: 16 such calls on next, this PR adds none). The file entered the
diff-mode scan's changed-file set only because #2573's W024 state_head
assertions touch it. Allowlist it alongside the other test files that carry
pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION
class). Scanner self-test 38/0; diff scan 14 files, 0 findings.
2026-08-11 17:10:23 -04:00
Tom Boucher
9341d8b8d3 test(#3334): fold the workflow-dispatch & review-lane fix-* cluster — Wave 2 (#3342)
Folds 15 tests/fix-*.test.cjs regression files (191 test() blocks) into
their module's main suite, per the wave decomposition of #3315 (H3 of
epic #3053). 187 blocks land in 8 existing suites (4 exact-duplicate
cases dropped, documented inline); 4 blocks move via git mv into 2 new
suite files with no prior coverage to merge into. Zero production
behavior change.

Also tightens two H1 (#3313) ratchets that the fold's own file-count
reduction moved past their grace window, per the ratchets' documented
dual failure mode (a stale/too-loose baseline fails exactly like a
novel violation):
- lint-test-file-count.allowlist.json: removes the stale "audit" entry
  (folding fix-2766 into tests/uat.test.cjs drops that module back to
  its 2-file cap).
- lint-allow-test-rule-refs.ceiling.json: lowers maxFiles 314 -> 309,
  the real post-fold high-water mark (gsd-test's own repo-baseline
  test caught this — CI, not a human, found it).

Two orthogonal review passes (Standards+Spec code-review, isolated
security-review) found and this commit fixes two issues before push:
a genuinely-distinct #2287 test case (file-absent vs. file-present-
resolved) that a prior fold pass had wrongly dropped as a duplicate —
restored verbatim into tests/uat.test.cjs; and a missing same-line
allow-test-rule citation on the #2196 block in
tests/debug-session-management.test.cjs, added for consistency with
its sibling #2257 block.

lint-removed-but-needed also caught two stale doc references to the
now-folded-away fix-2285-claude-orchestration-wiring.test.cjs filename
(docs/adr/1143-claude-orchestration-capability.md,
gsd-core/references/execute-phase-response-language.md) — updated both
to point at tests/claude-orchestration.test.cjs, its new home.

Co-authored-by: sim <sim@local>
2026-08-10 20:51:40 -04:00
Tom Boucher
7a7bf19fc1 enhance(#2872): record scope and runtime in the install manifest (#3323)
* enhance(#2872): record scope and runtime in the install manifest

gsd-file-manifest.json gains manifestVersion, runtime and scope, and a new
read-only Installed Surface Resolver Module reads both install scopes for a
runtime in one call -- the first code path in the repo that does.

Phase 3 of epic #2866 (ADR-2866). Blocks Phase 4 (#2873), which resolves
#2218: the resolver's shadowedBy field is that defect expressed as a value
for the first time. It ships computed-and-unread here.

Installed-ness is decided by manifest PRESENCE, never by the new fields, so
a manifest written by an older GSD stays fully functional and no user needs
to reinstall. Recorded runtime/scope are corroboration: a disagreement with
the probed config dir is reported as declaredScopeMatchesProbe: false, never
silently corrected.

readInstallManifest is widened additively -- version/timestamp/mode/files
keep their exact names, types and meanings for all four existing callers.
manifestVersion is a new field rather than a reinterpretation of version,
which holds the package version and is read by the golden-parity fixtures.

Stems are derived from the installed manifest's own file keys, the inverse
of Phase 2's filename composition, guarded by a fast-check round-trip
property plus a kebab-case charset check so a crafted manifest key cannot
put a traversal segment, control character or ANSI escape into a trigger
that Phase 4 renders back to the user.

Also fixes two defects found while working:
- bin/install.js hardcoded manifestVersion: 2 while the reader owned
  MANIFEST_SCHEMA_VERSION = 2. Now single-sourced, with a parity test.
- docs/installer-migrations.md documented an install-state schema of five
  snake_case fields that have never been written; InstallState has only ever
  been { schemaVersion, appliedMigrations }. Corrected with a dated note.

Verification runs on the remote runner.

* fix(#2872): fold review findings from three independent engines

Standards axis:
- convert the manifest-schema suite from a hybrid setup(t) closure to
  beforeEach/afterEach (CONTRIBUTING.md:319-354 Pattern 1). The hybrid was
  neither approved pattern and a new test forgetting the call got no warning.
- SCOPE_ORDER was declared twice with no parity test -- this repo's recorded
  generative-fix-divergence class. Give the ordering one owner: install-scope
  exports it frozen, the layout module and the resolver both import it, and a
  test locks it against scopeRank so the constant and the ranks cannot drift.
- drop the defaultReadManifest passthrough (Middle Man).

Spec axis:
- add the VOLATILE_FILES exclusion test and source comment the acceptance
  table promised and did not deliver. gsd-file-manifest.json stays excluded:
  the new fields are deterministic, but timestamp -- the original reason --
  is unchanged.

Security axis:
- bound the reported manifest runtime at 64 chars, matching the
  truncatePostureValue convention already used in this subsystem. It reached
  declaredRuntime unbounded while the adjacent stems were gated by SAFE_STEM;
  an inconsistent posture on the same attacker-influenceable document. The
  charset stays ungated on purpose -- declaredRuntimeMatchesProbe needs to see
  the real value -- so Phase 4 must sanitize before rendering, recorded in the
  design's Known limits.

Both new parity tests were verified to FAIL when the two sides are made to
disagree, then pass again on revert. Verification runs on the remote runner.

* chore(#2872): backfill changeset pr number to 3323

* fix(#2872): give git fixture construction its own timeout class

PR #3323's full test (windows-latest, 22, shard 2/3) failed with

  gitOrThrow: 'git init' failed -- outcome=timed_out exitCode=null
  gitOrThrow: 'git commit --allow-empty' failed -- outcome=timed_out

from drift-detection.test.cjs's beforeEach, a file this branch never touched.
Every other lane passed the same commit, including windows-latest node 24 on
all three shards, and next is green.

Root cause is a bound sized for the wrong class. DEFAULT_GIT_TIMEOUT_MS is
15000 and its own comment scopes it to plumbing READS -- rev-parse, branch,
log -- against an existing repo. createFixture uses it for six sequential
repo-CONSTRUCTION spawns: init, three config writes, add -A, commit. init and
commit each write dozens of files, and on Windows every spawn is
Defender-scanned. Sibling tests in the failing block took 15.6-22.0s against
a 15000ms bound.

This repo already diagnosed this exact shape once: timeouts.cjs's
HOOK_FANOUT_TIMEOUT_MS records PR #3285 failing in the SAME job with the SAME
outcome=timed_out exitCode=null signature at the SAME bound while every other
lane passed, and concludes 'a bound sized for the wrong class, not a slow
machine'. It was fixed by splitting out a heavier class-norm at 60000. Same
remedy here: GIT_FIXTURE_TIMEOUT_MS = 60000, 4x the bound that failed and half
INSTALL_TIMEOUT_MS.

DEFAULT_GIT_TIMEOUT_MS deliberately stays at 15000 -- a blanket raise would
stop a genuinely hung plumbing read from surfacing fast.

Verified the value reaches the spawn rather than being an ignored option:
spawnSync was monkeypatched before requiring the fixture module, and all six
git construction calls were captured carrying timeout: 60000.

This branch's two new test files shift shard composition, which is how a
pre-existing fragility landed in the heaviest shard on the slowest lane.
Fixed here rather than deferred, per the no-defer rule.

Verification runs on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-10 15:50:55 -04:00
Tom Boucher
80734a9694 chore(#3053): clock-seam ADR-456 amendment + backfill — H2 (#3332)
* test(#3314): backfill deterministic clock-seam coverage (failing-first)

Replaces loose regex/range assertions with exact-value and boundary
tests for the CLI-subprocess and in-process clock-touching call sites
identified by H2's audit (epic #3053): cmdCurrentTimestamp,
_wsParseRetryAfter (commands.cts), cmdInitManager's is_active gate and
cmdInitQuick's quick_id generation (init.cts), and reapStaleTempFiles
(io.cts). The CLI-subprocess-pinned tests are expected RED until a
follow-up commit routes those call sites through realClock so
GSD_TEST_MODE+GSD_NOW_MS can reach them.

* refactor(#3314): route CLI-subprocess clock reads through realClock

cmdCurrentTimestamp, cmdInitManager's is_active gate, and cmdInitQuick's
quick_id generation read Date directly, which the GSD_TEST_MODE+GSD_NOW_MS
subprocess pin cannot reach (it only fires inside realClock.now()).
Behavior-preserving: realClock.now() falls through to Date.now() whenever
GSD_TEST_MODE is unset, which is every real invocation.

* docs(#3314): amend ADR-456 with reachability-based clock-control rule

ADR-456 §(a) documented one mechanism (injected {clock=Date} +
t.mock.timers). Adds the two this repo already relies on: t.mock.timers
for in-process direct-Date reads, and the GSD_TEST_MODE+GSD_NOW_MS
subprocess pin (routed through realClock) for CLI-spawned code. Updates
TESTING-STANDARDS.md's matching passages, which already referenced this
issue by number as the no-elapsed-assertion promotion precondition.

* fix(#3314): address orthogonal review findings

Spec-axis findings: file and link the no-elapsed-assertion promotion
follow-up (#3331) instead of leaving TESTING-STANDARDS.md pointing at
a dead #1885, and ship the module-by-module audit table in the ADR
itself rather than only in a gitignored phase artifact.

Standards-axis finding: pin the "hour-old file = not active" test via
GSD_TEST_MODE+GSD_NOW_MS for consistency with the reachability rule
this PR's own ADR amendment now documents.

---------

Co-authored-by: sim <sim@local>
2026-08-10 15:24:39 -04:00
Tom Boucher
8b4545f3c0 feat(#3218): the prompt layer asks the CLI for plan counts (#3327)
* feat(#3218): the prompt layer asks the CLI for plan counts

Seven sites across four workflows counted plans with ls and wc -l instead of
asking the CLI. A shell glob is not scanPhasePlans, so every fix that landed on
the owner missed all seven: they counted superseded plans as live, reported zero
for the nested plans layout, and missed loosely-named files. The 1762 figure of
30 plans and 24 summaries came from here.

phase find is extended rather than a verb added - 3218 is an enhancement whose
own checklist says it adds no new command, and CONTRIBUTING makes a new verb a
feature needing approved-feature. It gains plan_count and summary_count for the
live set and plan_count_all for the physical one, additively; the existing arrays
are untouched.

Both sets are exposed because the sites need different ones. Amendment 1 names
two cases; three of these sites ask a third - did the planner write files to disk
- and take the physical set, because a superseded plan is still a file the
planner wrote.

The progress.md dead route is fixed and was worse than the issue said. It read
.plans and .summaries arrays that roadmap.analyze has never emitted, so the
fallback always fired, both counts were always zero, and Route 0's
resume-incomplete-phase check had never fired at all.

The ratchet baseline is empty. Its own stale-entry check makes that
self-enforcing.

Verified on the remote runner.

* test(#3218): acknowledge the workflow growth and update the stale guard

The emitted-attribution gate named its own remedy, so it was followed rather
than pre-guessed: four workflow files grew between 200 and 770 bytes because each
replaced a shell glob with a find-phase call plus its jq extraction. plan-phase
grew most - two sites, and it takes the physical count for its did-the-planner-
write-files question. progress also carries the Route 0 dead-path fix.

plan-phase-drift-guard asserted the literal old ls shape. Updated rather than
deleted: what it protects is that a filesystem fallback exists and is reachable,
and that is intact. It is not a regression - gsd_run is already load-bearing
throughout plan-phase.md long before step 9, so the 9a and 11a fallback never
existed to survive gsd_run being unavailable; it guards against the planner
subagent's return hanging.

Three ack sources collided with the new fragment, which the gate treats as a hard
error rather than last-wins. Only the three colliding keys were removed, not the
421 spent entries, and two fragments left entryless were deleted per the
convention that an empty fragment signals nothing.

Verified on the remote runner.

* docs(#3218): document the live and physical plan counts

docs/CLI-TOOLS.md gains a find-phase counts section covering plan_count and
summary_count for the live set against plan_count_all for the physical one, plus
the null-not-zero not-found behavior. The live-versus-physical distinction is
spelled out because a caller picking the wrong one gets a plausible number, which
is the trap Amendment 1 records.

Changeset leads with what a user sees: progress and execute-plan stop counting
superseded plans as outstanding, a nested plans layout stops reporting zero, and
Route 0 resume routing starts working after never having worked.

No how-to. Nothing is enabled and nothing is sequenced - the user runs the same
command and the number is simply correct. The one new distinction is field
semantics, which is what a reference entry is for.

* chore(#3218): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3327 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-10 12:56:38 -04:00
Tom Boucher
bcf7b04864 chore(#2896): convert CONTEXT.md prose defect registry into enforced gates (#3325)
* chore(#2896): convert CONTEXT.md prose defect registry into enforced gates

Squashes the prior 4-commit sequence and fixes defects found while
resuming this branch: 5 orphaned/corrupted DEFECT fragment lines left
by an earlier botched edit, 17 "Source of truth: Memtrace `find_symbol`"
placeholders that had destroyed real file-path citations, and 3
DEFECT.GENERATIVE-* entries merged into one RULESET.GENERATIVE-FIX
predicate (policy, not an unenforced defect) to satisfy the zero
DEFECT.<NAME>.<field>= acceptance criterion.

Six mechanizable defects get real gates: DEFECT.UNBOUNDED-SUBPROCESS
(eslint-rules/require-subprocess-timeout.cjs), DEFECT.CANARY-VERSION-LEAK
(scripts/lint-canary-version-leak.cjs + version-gate.yml),
DEFECT.CHANGESET-PR-FIELD-DRIFT (findPrFieldDrift in changeset/lint.cjs),
DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, DEFECT.REMOVED-BUT-NEEDED, and
DEFECT.DEFAULT-FLIP-DOCUMENTATION (new lint scripts, wired into lint:ci).
Already-enforced and unenforceable prose entries are deleted; the gate
is the record.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2896): route the new lint tests' subprocess calls through the bounded process-seam helper

The 4 new test files for this PR's lint checks called cp.spawnSync/
execFileSync directly with no timeout, tripping this repo's own
existing local/no-unbounded-spawn ESLint rule. Route every one through
runNode/gitOrThrow (tests/helpers/process-seam.cjs,
tests/helpers/git-fixture.cjs) instead, matching the pattern already
used elsewhere in the suite (e.g. tests/changeset-lint.test.cjs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: register claude-orchestration.cjs and regenerate stale generated indexes

Pre-existing drift on next, unrelated to this PR's own change, surfaced
by running lint:ci as part of verifying #2896: two cli_modules
(claude-orchestration.cjs, write-set.cjs) landed without a manifest
regen, and CONTEXT.md's own edits in this PR staled its two generated
indexes. Adds the missing docs/INVENTORY.md row for
claude-orchestration.cjs (write-set.cjs already had one — only its
manifest entry was stale) and regenerates
docs/INVENTORY-MANIFEST.json, docs/CONTEXT-INDEX.json, and
examples/dynamic-context-management/CONTEXT-INDEX.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2896): default-flip-documentation lint's local fallback base was main, not next

Found in review: every other base-ref fallback in this repo (see
scripts/changeset/lint.cjs's DEFAULT_BASE, #2988) defaults to `next`,
the integration branch every PR actually targets — `main` is the
release branch. This script's local fallback (used only when
GITHUB_BASE_REF is unset, i.e. never in CI, but potentially on a local
or direct invocation) diffed against the wrong ref. No test exercised
the unset-env-var path, so it shipped unnoticed; every e2e test sets
GITHUB_BASE_REF explicitly and is unaffected by this fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2896): stale eslint comment, overclaiming CONTEXT.md wording, and an incompletely-regenerated manifest

Found by the isolated Standards code-review pass:
- eslint.config.mjs's require-subprocess-timeout comment said "'warn'
  for now... flip to 'error' once migrated" while the rule already
  shipped as 'error' with all 8 sites migrated in the same commit —
  described a state that never existed.
- The CONTEXT.md pointer block claimed the rule's bounded call sites
  "never throw", but roadmap-upgrade.cts's pre-mutation clean-tree
  check correctly still throws on failure (it gates a destructive
  real-run migration; degrading to "assume clean" would risk clobbering
  uncommitted work) — softened the claim to describe both shapes
  accurately instead of overclaiming one.
- docs/INVENTORY-MANIFEST.json's claude-orchestration.cjs/write-set.cjs
  entries from the prior "fix: register claude-orchestration.cjs..."
  commit didn't actually land — re-running the generator now includes
  them; lint:generated-sync is green.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#2896): backfill changeset pr field with the real PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#2896): normalize buildCorpus file paths to POSIX in lint-removed-but-needed

Windows CI caught it: path.relative(root, abs) returns backslash-
separated paths on Windows, but findSurvivingReferences's package-lock
special case does file.startsWith('.github/workflows') — a forward-
slash literal. On Windows the check silently never matched, so
tests/removed-but-needed-lint.test.cjs's real-defect-shape fixture got
exit 0 instead of the expected exit 1. Normalize at the production
source (RULESET.CONTENT-PATH-NORMALIZATION) rather than the test side.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-10 12:55:52 -04:00
Tom Boucher
aceea3ce4a refactor(#3217): withhold a percentage when its scope is not complete (#3318)
* wip(#3217): rule-4 scope withholding — parked, two open findings

Implemented but NOT shippable. An isolated review found buildStateFrontmatter
still hardcodes SCOPE.COMPLETE, so state json reports percent 0 where roadmap
analyze, stats and query progress all correctly report null on the same disk
state - rule 4 reintroduced at a site this phase claims to close. Also: roadmap
analyze emits scope complete beside progress_percent null with nothing
explaining it.

Parked to build Phase 4 (#3186) first, which is unblocked. Findings recorded in
.gsd/phase/refactor-3217-completion-ratio-scoping/60-review.json.

* fix(#3217): withhold the sync percentage on a non-complete scope

The parked blocker is fixed - buildStateFrontmatter no longer hardcodes
SCOPE.COMPLETE, and the prose Progress fallback is gated too, which was a second
leak found while tracing the first. roadmap analyze exposes progress_scope so a
consumer can tell WHY a percentage is absent from the JSON alone.

Then a residual gap was reproduced rather than assumed. cmdStateSync carried the
same hardcode behind a written reason claiming it did not reproduce. It did: on a
TRUNCATED window and on UNSCOPED row 4, state sync wrote Progress 0 percent to 100
percent while state json, roadmap analyze, stats and query progress all withheld -
and it persisted a self-contradictory file, body claiming 100 percent while its own
frontmatter correctly omitted percent.

The excuse was also wrong. syncRoadmapRaw is already parsed in that function and is
exactly what produces a real scope, so there was a scope to pass. Threaded through
listMilestonePhaseDirs; a non-complete scope now skips the write with a reason in
changes. milestoneBounded stays as the orthogonal 1761 guard for row 5.

Second time this epic a does-not-reproduce claim was too generous. Recorded in
ADR Amendment 8 as a correction rather than a quiet rewrite.

Verified on the remote runner.

* test(#3217): give the withholding fixtures a resolvable scope

40 matrix failures, all fixture drift - no code regression. My own hypothesis
that this was over-withholding was wrong and is recorded as such: the worry case,
a plain ROADMAP with Phase entries and no version heading, resolves to complete
exactly as ADR 7.1 says it should.

The real causes were two fixture shapes. Most had no ROADMAP.md at all, which is
unreadable via a pre-existing graceful path, and asserted a numeric percent. The
five vscode, pi-extension, mcp-server and shell-projection failures were that
shape - bare temp dirs using progress json as a reachability proxy while
asserting typeof percent is number, which under rule 4 is now null.

The rest had a version token in a title or heading with no STATE.md milestone
pointer to resolve it, which is classification row 4, versioned but unresolved,
so withholding is correct per the contract.

Verified on the remote runner.

* test(#3217): make the LM-tools reachability tests dispatch against their fixture

The gsd_progress reachability test was never testing its fixture. invoke()
resolves cwd from vscode.workspace.workspaceFolders by design (the real
LanguageModelToolInvocationOptions has no cwd field, per the 2103 fix in
extension.js), the mock had no workspace at all, and the test passed a cwd option
nothing reads - so it dispatched against the repo working directory. Writing a
ROADMAP into the temp dir had no effect. Rule 4 only made it visible.

Fixed by mocking workspaceFolders. The two siblings in the same file carried the
identical dead cwd and were dispatching against the repo too; they were not
failing only because their assertions did not touch scope-dependent output. Both
now use their own fixture with assertions unchanged - the no-planning fallback
paths already satisfy them honestly.

Re-scanned the other five reachability files: no further instances. They thread
cwd into parameters that genuinely read it, not through an options shape that
ignores it.

Verified on the remote runner.

* chore(#3217): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3318 exists.

* ci(#3217): give the coverage merge enough heap for the merged shards

The coverage gate OOMed at exit 134. c8 report merges three shard artifacts,
roughly 358MB of V8 dumps in coverage/tmp, and died holding their per-file
position maps at the ~4GB default heap. Verified as this branch's delta rather
than pre-existing: the same job succeeded on next at 14:18, after phases 4 and 5
merged.

Both coverage-gate steps get the bump because both re-slice the same merged data.
8192 doubles what failed and leaves headroom on a 16GB ubuntu runner, matching
the idiom the shard step already uses at 6144.

This is a memory bound, not a change to what is measured. No threshold was
touched. The test file was checked for gratuitous subprocess spawning and is
already reasonable at 43 spawns, each a distinct fixture-by-surface pairing.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-10 11:59:51 -04:00
Tom Boucher
e201cde73c refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4

The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a
ROADMAP checkbox is a human annotation with no machine authority. Section 7.4
still carried the OPEN QUESTION and was marked blocked, so the contract said one
thing and the tracker another.

Recorded per section 7's own rule - a behavior not stated there is not decided,
and amending a rule is an ADR amendment rather than a code change with a comment.
The decision comment names Phase 4's PR as the carrier of this edit and makes it
an acceptance criterion that the text be in the tree before implementation
begins, so this lands first, alone, ahead of any code.

Also clears the stale blocked-on-2957 row in the guard roster.

* refactor(#3186): one shared phase-completion predicate, disk-strict

isPhaseComplete in verification.cts becomes the single owner. It calls
readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a
zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init
gated the read on a plan count and synthesized a not_required sentinel, so
phase.complete succeeded while init.manager reported incomplete for the same
phase.

The guard, built and run before scope was fixed per Amendment 3, found 9
re-derivations where the ADR named 3. Four were unnamed, including one in the
prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under
disk-strict is the divergence itself.

Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no
machine authority. The overrides in roadmap analyze and init manager are deleted
rather than generalized; the user's checkbox stays in ROADMAP.md, only its
authority goes.

scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded
- they answer 'are all plans summarized', which is a different question, and
folding them would either over-report completion or invert the dependency
direction between Phase 1's owner and this one.

Verified on the remote runner.

* fix(#3186): close seven review findings and record the missing-verdict rule

The isolated review reproduced a write-path regression I introduced: migrating
cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase
with a fresh passing verification plus a newly-added unsummarized plan reported
complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The
owner stays right per 7.4 - plan count is not a completion precondition - so the
gate is restored at the write site as an explicit composition, mirroring the
separate 2648 unexecuted-plan gate cmdPhaseComplete already carries.

The spec axis was right that my 0.x-split reasoning was too permissive. The 2957
decision names buildStateFrontmatter as one of the three that must converge, and
buildWorkstreamInventory combined a summaries-met local with verification data to
decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in
a third surface. Both now route through the owner. The raw scanPhasePlans helper
stays: it answers are-plans-summarized, which genuinely is a different question.

Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so
an absent VERIFICATION.md means not complete everywhere. That retires 2645's
verifier-disabled tolerance and inverts its Goodhart incentive - deleting the
evidence now lowers completion instead of raising it.

Guard hardened: block-form count gates and algebraic restatements are caught, and
the header now discloses its remaining limits instead of overclaiming.

Verified on the remote runner.

* fix(#3186): route state sync through the owner and catch bare completed reads

The matrix found 52 failures. 51 were fixtures asserting the old semantics: a
phase with plans and summaries but no VERIFICATION.md used to count complete and
correctly no longer does. Each fixture now carries a passing verification where
that is what the test was actually about, rather than having its assertion
weakened.

The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync
destructured scanPhasePlans().completed directly - a bare field read, not a
comparison - and used it as a completion verdict, so state sync and state json
disagreed on completed_phases for identical disk state. Routed through the owner.

Guard gains shape (d): any read of .completed off a scanPhasePlans() result
outside plan-scan.cts, in chained, destructured and indirect forms, function
scoped with no line window. It cannot tell a summaries-met read from a completion
read - that is data flow - so it flags every one and requires a written-reason
exemption, which is the same discipline shapes a-c already use. The blind spot is
disclosed in the header rather than overclaimed.

The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md
checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks.

Verified on the remote runner.

* test(#3186): give the nested-plans sync fixture a passing verification

Last 3 matrix failures were one failure echoing up two describe levels. Phase
01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict
completed stayed 0 and no Progress change was emitted - correct new behavior, not
a regression.

Added the passing verification rather than dropping the Progress expectation, so
the test still covers what #3257 is about: that a nested plans/ layout is counted
and not undercounted. Probe against the built lib confirms
Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3.

* chore(#3186): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3306 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-10 10:18:29 -04:00
Tom Boucher
96a82bbffb enhance(#3245): report the detected host runtime in init (#3307)
* test(#3245): failing-first coverage for host runtime detection in init

Locks the behavior epic #2313 Phase 5 must produce before any of it exists: init reports the detected host, explicit GSD_RUNTIME and config runtime still outrank detection, non-Codex sessions are untouched, and nothing is ever written to shared defaults (#2297).

* enhance(#3245): report the detected host runtime in init

init reported agent_runtime: claude inside a Codex session, and resolved agents_dir to the Claude agents root with agents_installed: true — a spuriously healthy triple. Runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json.

Adds a detection rung beneath both explicit sources, in a new pure module. Codex is identified from its own documented session environment (CODEX_SANDBOX / CODEX_SANDBOX_NETWORK_DISABLED), else an explicitly exported CODEX_HOME whose config.toml exists. The default ~/.codex is never probed: that file exists on every machine that has run Codex, so probing it would misreport other runtimes' sessions.

resolveRuntime keeps its exact contract and all 71 dependents, including formatGsdSlash command-style emission; only withProjectRoot consumes the new rung. Nothing is written on any path (#2297). Explicit config still wins (#2517).

* fix(#3245): make the parity guard real and single-source the marker

Four independent review passes found the generative-fix-divergence guard was vacuous: it asserted agreement at the one input where inferPreferredRuntime and detectHostRuntime do not differ, so it could not fail. It now pins the actual divergence point (CODEX_HOME set, config.toml absent) and records that the asymmetry is deliberate.

The config.toml marker is now single-sourced from update-context.cts and imported, rather than carried independently by two surfaces. tests/helpers.cjs now scrubs CODEX_SANDBOX and CODEX_SANDBOX_NETWORK_DISABLED: GSD reads them, so an ambient Codex session would otherwise make the non-codex control test fail non-deterministically.

Also: detection is throw-safe end to end rather than only around the fs probe; the Windows-join test is replaced with one that can actually fail (trailing-separator, catches hand-rolled concatenation); the #2297 no-write proof now wraps resolveReportedRuntime, the function that ships, across all three ladder outcomes.

* chore(#3245): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-10 09:48:26 -04:00
Tom Boucher
d28ab7c8f7 enhance(#3243): sync installed codex .toml model/effort to the passive posture (#3296)
* feat(#3243): sync installed codex .toml model/effort to the passive posture

Implements ADR-2313 D7, and owns the Codex .toml typed IR that Phase 1's
review assigned to this phase.

The IR exists for a structural reason, not tidiness: this phase has to
PARSE these files, and a parser kept bug-compatible with a separate
renderer is the generative-fix-divergence shape this epic already dealt
with once for the model predicate. So Phase 2's parsing MOVES here
rather than being copied — agent-install-check now imports it, and its
test file passing unchanged is the proof the extraction altered nothing.

The load-bearing property is byte-identical round-trip: render(parse(x))
=== x. Without it a sync silently reformats a user's file — line
endings, key order, BOM, trailing newline — turning a two-line repair
into a whole-file diff in their dotfile repo. The IR keeps original
lines and removes targeted ones rather than reconstructing from parsed
fields, which is what makes that property hold.

It also reconciles a real contradiction between Phase 2 and ADR-2313. An
unterminated developer_instructions block: the reader excludes the rest
of the file, deliberately failing toward a false positive, because
misreading prose as a pin only wastes a user's time. The writer must
refuse, because proceeding on a malformed document rewrites it. A false
positive is the safe direction for a reader and the dangerous one for a
writer. So the parse reports the fact and the two consumers branch on
it — one parse, one truth, two policies, instead of two parsers that
agree today.

The sync leaves a legal real-Codex pin and its coupled effort untouched,
reported skipped rather than synced; strips a stale Anthropic or tier
model and an orphaned effort; keeps dry-run as the default; refuses any
file whose parse fails; and skips symlinks exactly as the Claude path
already did. The Claude path itself is byte-identical.

PARSE_REASON.NO_HEADER from the ADR's illustrative snippet is
deliberately not implemented — a missing header is legal, not an error,
so it would be a dead enum member that the enum-lock test then pins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3243): preserve per-line endings and make the codex write atomic

Two findings from an isolated review, both in the write path.

BLOCKER: mixed line endings broke the byte-identical round-trip. `eol`
was a single whole-file flag and split(/\r?\n/) discarded each line's
own terminator, so render re-joined with ONE style and normalized every
line — even with zero strips performed. A file with one CRLF line and
the rest LF came back fully converted. That falsified the A14 guarantee,
violated the design's "must not silently rewrite every line", and made
the CONTEXT.md glossary claim wrong. It was untested because A12 and B15
only cover PURE CRLF; no mixed-ending fixture existed anywhere.

Fixed by keeping each line's terminator alongside its content, so render
is a plain concatenation and a strip removes only the target line and
its own terminator. `eol` survives as informational metadata that render
never reads. Seven fixtures added for the paths nothing exercised:
mixed endings unmodified and with a strip, a lone \r, a file ending on
the block's closing ''' with no newline, multiple trailing newlines, a
BOM-only file, and an empty file.

MINOR, but it contradicted this phase's own contract: the write was
in-place open-truncate, so a failure between truncate and completion
leaves a truncated .toml — exactly what ADR-2313 says must never happen.
The Codex path now writes a sibling temp file and renames over the
target, which is atomic on one filesystem, with cleanup on failure. It
uses the repo's existing retryRenameSync rather than a hand-rolled
rename, and deliberately NOT platformWriteSync, whose normalizeContent
would mangle the very CRLF and trailing-newline bytes the round-trip
property exists to preserve.

The Claude path keeps its in-place write untouched. It has the same
shape, but changing it is not this phase's business and its tests must
stay byte-identical.

B20 previously mocked writeFileSync to throw BEFORE touching anything,
so it proved nothing about a mid-write failure — its passing comment was
true only because of how the mock was built. It now performs a real
truncated write wherever writeFileSync is called, catching both the
naive direct-to-target path and the new temp path, and asserts the
target is byte-identical afterwards with no stray temp file left.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3243): preserve the trailing-newline state when stripping a last line

Caught by B17, one of this phase's own tests — the suite working, not a
test problem.

Content is reconstructed as the concatenation of lines[k] +
terminators[k], so a file with no trailing newline has '' as its last
terminator. removeLine spliced out both arrays at the same index, which
is right for a middle line but wrong for the last one: it dropped the
empty terminator and left the PREVIOUS line's newline in place. A file
ending `...\nmodel = "sonnet"` with no trailing newline came back
as `...\n`, gaining a newline the user never wrote.

The new last line now inherits the removed line's terminator, so a
removal leaves the file exactly as if that line had never been written.
Removing the only line yields an empty file rather than a stray
terminator.

Both stripModel and stripReasoningEffort funnel through the one
removeLine, confirmed rather than assumed, so a single fix covers both —
including the row-B7 shape where a stale model and its orphaned effort
are removed in sequence and the second removal targets the last line.

Two of the four new cases are honestly not red-first and say so in
their comments: removing a last line that HAS a trailing newline only
exposes the bug under mixed EOL, since uniform files coincidentally
have equal terminators on both sides; and removing the only line
already degenerated correctly through Array.slice. They are kept as
guards for the new branch rather than dressed up as catches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3243): document the codex repair path and close the loop

How-to: the Codex-400 entry added in Phase 2 told users to re-run the
installer, because that was the only repair available then. It now leads
with `effort sync` and keeps the reinstall as the alternative, with the
reason to prefer one — a reinstall regenerates the agent files
wholesale, so anyone who hand-edited theirs loses those edits. Detect,
preview, apply is now one continuous path in one place.

Reference: docs/COMMANDS.md had no `effort sync` entry at all — the same
gap `validate agents` had in Phase 2, found the same way. The entry
documents BOTH runtimes, because the command genuinely forks on runtime
and describing only the new half would misdescribe it.

The write flag is `--apply`. The design doc and test matrix both said
`--no-dry-run` throughout, which does not exist — verified against the
actual arg parser in gsd-tools.cjs before writing. Documenting a flag
that does not exist is worse than documenting nothing, because it fails
at the moment someone needs it.

Both surfaces state that only the targeted lines are removed and every
other byte is preserved. That is a user-visible guarantee rather than an
implementation note: it is the difference between a two-line diff and a
reformatted file in someone's dotfile repo, it is what the IR's
round-trip property exists to deliver, and writing it down makes it a
contract a future change has to break knowingly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3243): inherit the trailing-newline state, not the line ending style

My previous rule was subtly wrong and this phase's own test caught it.

"The new last line inherits the removed line's terminator" copies the
removed line's STYLE as well as its presence. A26 uses mixed endings on
purpose — line one terminated \r\n, the model line terminated \n — so
inheriting silently rewrote line one's ending to \n. That is precisely
the defect class the mixed-EOL blocker fix existed to eliminate,
reintroduced one layer down by the fix for it.

The correct rule inherits the EMPTINESS only. If the removed line had no
terminator, the new last line loses its own, preserving "this file has
no trailing newline". Otherwise the new last line keeps its own
terminator: it is already a newline, and already the right style for
that line.

A26's assertion moved too, and that deserves saying plainly rather than
burying: it previously encoded my wrong rule. Changing a test to match
the implementation is usually the mistake, so it was checked from first
principles instead — a file whose first line ends \r\n and whose last
line ends \n, with that last line removed entirely, must be the first
line with its own \r\n intact. The new expectation is what the user's
file should actually look like; the old one was wrong.

A29 adds the interaction nothing covered: the compounding case (strip a
stale model, then its orphaned effort, the second removal landing on the
last line) with non-uniform endings either side. The two fixes meet
there and nothing exercised the meeting point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3243): drop the phantom trailing line from the IR representation

Root cause, not another patch on the removal rule. Three consecutive
fixes there each surfaced the next issue, which was the signal that the
data model was wrong.

splitPreservingTerminators left a phantom empty final entry for any file
ending in a newline: "a\nb\n" became lines ['a','b','']. So for the
common case the real last content line was NOT the last array element,
removeLine's isLastLine check never matched it, and every rule I gave
was reasoning about the wrong element.

What hid it: render was already a plain concatenation, so a phantom
empty line with an empty terminator contributes nothing to the output.
A14's byte-identical round-trip could never have caught it — the defect
is byte-neutral until a removal shifts the index arithmetic under it.
That is worth recording, because "the round-trip test is green" was
exactly the reassurance that kept the search pointed elsewhere.

The representation is now 1:1 — terminators[i] follows lines[i] and may
be '' — with no phantom, verified across empty, no-trailing-newline,
trailing-newline, blank-line and mixed-CRLF inputs. render stays a plain
concat and needs no special cases. With the phantom gone the removal
rule is correct as stated and finally applies to the genuinely last
element.

Consumers checked rather than assumed: the block-range detector and
header scanner are agnostic to array shape, and Phase 2's reader uses
its own independent split, so tests/agent-install-check.test.cjs is
untouched and still passes unchanged.

One test expectation was wrong and is corrected rather than quietly
adjusted: A18 asserted a 7-element terminators array whose trailing ''
was the phantom itself. It now asserts the six real terminators, which
is what the invariant lines.length === terminators.length requires.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3243): backfill changeset pr number (#3296)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 01:31:15 -04:00
Tom Boucher
c28134ab39 fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host

Adds local/no-duplicate-fold-marker, an AST rule that reports the second and
every subsequent `folded:<name>` marker in a host file, plus RuleTester cases
and a tree-wide regression assertion.

Failing-first on purpose: the rule is registered at error and the 25 duplicated
regions are still present, so eslint and the new tree-wide test are RED. The
deletions land in the next commit.

The marker key is the whitespace-delimited token after `folded:` — not the
issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on
tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and
feat-443-effort-fast-mode are two distinct folded suites.

Refs #3271

* fix(#3271): delete 25 duplicated folded suites from three install hosts

Three consolidated install suites each carried a verbatim second copy of a
contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and
green — each duplicated block registered and ran twice on every lane.

  tests/install.test.cjs                   5981-9937  (3957 lines, 18 blocks)
  tests/install-minimal-hooks.test.cjs     2734-4015  (1282 lines,  5 blocks)
  tests/install-write-confinement.test.cjs 1754-2321  ( 568 lines,  2 blocks)

Introduced by 6d072435d (#1975 re-applying #1970's hunks on a tree that already
had them, 2026-07-03) — one stale-base re-application, three files, one commit.
Verified by marker-count bisect: 1 at 4f779eda4 and 0cc7a1a42, 2 from 6d072435d
onward.

The later copy is deleted in each case, so every file returns to what its
authoring batch produced and blame on the surviving lines stays accurate.
local/no-duplicate-fold-marker, red on the previous commit, is now green.

tests/model-resolver.test.cjs is untouched: the issue lists it, but its two
blocks are folded from two different files and are not identical. It is a false
positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`.

Fixes #3271

* test(#3271): property-test marker identity and pin the alias non-goal

Three review findings, all fixed inline:

1. foldMarkerOf is a parser and carried no fast-check property test. Raised
   independently by the /code-review standards axis and the isolated adversarial
   pass; the file already establishes the fc.property-driving-ruleTester idiom
   for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]:
   the same marker twice always reports exactly once against firstLine 1, and
   two distinct markers never collide. The alphabet includes `.` on purpose —
   an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and
   fails arm 2, which is exactly the model-resolver false positive.

2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local
   rules use 'Best Practices', 'Portability' or 'Reliability'. Now
   'Best Practices'.

3. A call through a further alias (const d = __foldDescribe) was unreported and
   undocumented — accidental rather than deliberate. It is now the fourth entry
   in the rule's documented non-goals, with the reason, and pinned by a valid
   RuleTester case so it cannot drift silently.

Refs #3271

* test(#3271): name the step and elapsed time when a baseline build fails

buildBaselineAtRef runs four bounded steps and, when one exceeded its bound,
threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long
anything took. Diagnosing one real failure took four separate experiments to
recover information the throw already had.

Each step is now timed, and any throw carries the breakdown: which step failed,
its elapsed time, the timings of every step that completed before it, all three
bounds, and the tail of the child's captured stdout/stderr.

The failure message is deliberately the carrier. On the remote runner the
captured output field comes back empty in failures.json while error and stack
survive verbatim, so the message is the only channel that reaches a reader of a
remote verdict.

Refs #3271

* fix(#3271): size the baseline generator bound for the machine it runs on

Instrumentation from a real remote-runner failure gave the breakdown:

  git-worktree-add=15.1s  npm-run-build-lib=19.8s  gen-emitted-baseline=FAILED@300.1s

Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is
not hung — it needs more than 300s there.

Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a
clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real
suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly
where a container degrades worst (3.9x slower than host, against 1.1x for file
IO) — which is why a CPU-only load test did not reproduce it and why four
earlier hypotheses (container slowness, network, shallow clone, CPU contention)
all measured clean.

The 300s bound was sized on an idle machine for a step that never runs on one.
Under the remote runner the on-disk baseline cache is structurally absent — CI
restores it via actions/cache keyed on github.event.pull_request.base.sha, a key
that exists only inside GitHub Actions — so this slow path runs on every remote
verification. The result: this gate has passed 0 times in 754 runs, failing 80
times and never once executing successfully.

Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the
largest meaningful bound; the other two bounds are untouched. This makes the
gate RUN, which is the point: the alternative considered and rejected was
degrading the timeout to a skip, and that was measured to turn the suite green
with the gate silently not running at all.

The real remedy is making the cache reachable from the remote runner so the
in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a
gsd-test-runner change, not one this repo can make.

Refs #3271

* fix(#3271): tolerate an overlay source that vanishes mid-walk

Observed on the remote runner, three runs across three different branches:

  ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js'
    -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js'

buildOverlayRepo enumerates names with readdirSync and then acts on each one, so
statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is
regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so
any concurrently running test that rebuilds hooks retires a just-listed name
mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and
EPERM; ENOENT went straight through.

On ENOENT the source is now re-examined ONCE rather than slept on. An atomic
rename is a single syscall, so by the time the failure surfaces the successor is
either already in place (the retry succeeds) or the path has genuinely left the
tree, in which case there is nothing to mirror and the leaf is skipped. No sleep
and no spin: a timing-based wait here would be the very flake being fixed. Every
other errno still propagates untouched, so a real permission or IO fault stays a
hard failure.

Five tests hold the boundary: gone-for-good skips without retrying, mid-replace
retries exactly once and places the file, EACCES still throws, a real linkSync
ENOENT is injected by monkeypatching fs and restoring it in a finally (never a
mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT.

Refs #3271

* fix(#3271): order the timeout ladder inward-out and lock it

Two review blockers, both real.

The generator bound had been raised to 600000ms — exactly the whole-chunk timeout
in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the
race: the chunk is killed first and the failure arrives as an opaque "no failed
step" kill, so the per-step diagnostic added a commit earlier was built and then
made unreachable in the same change.

Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the
inner bound it was meant to permit, so it could still die at the exact 300s
ceiling this was supposed to lift — via node:test's timeout rather than
spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the
same opaque-kill hazard from the other direction.

The three bounds only produce a useful failure if they fire inward-out, so they
now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation
(91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of
chunk headroom for every other file sharing it. Four tests lock the ordering,
including a drift guard on the exported values — without it, editing a call
site's literal timeout would leave the ordering assertions passing while the real
ladder inverted.

Also from review:

- err.gsdBaselineStep and err.gsdBaselineTimings were written and never read
  anywhere in the tree; only the rewritten message is consumed. Removed rather
  than kept as speculative surface.
- buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so
  a vanished leaf left the overlay with no accounting at all. It now collects the
  skipped paths and warns once. Not thrown: a source that left the tree really is
  not part of the snapshot, and throwing would reintroduce the crash the
  tolerance removes — but silence would let a dropped leaf resurface later as an
  unrelated missing-file assertion.
- The instrumentation commit shipped no test. One now drives a real failure and
  asserts the message names the step, its elapsed time, and the bounds.

Refs #3271

* chore(#3271): backfill the changeset PR number

* fix(#3271): bound a hook fan-out as its own class, not as a bare probe

CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every
other lane green, including windows-latest node 24 across all three shards:

  not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex
    error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr=
    duration_ms: 15040.2168

A bound, not a hang: the test supplies stdin via input:, so the hook is not
blocked reading its ref list, and the duration lands exactly on the 15000ms
bound.

The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a
single short CLI query or node -e probe against a temp fixture". This is not
that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK
git that is itself a bash script, so one runHook is roughly four Git Bash spawns.
On Windows each is Defender-scanned and the first hook test in a file pays cold
start on top. That module's own docstring warns against precisely this: a call
site that differs from its class must not be forced onto a shared value that does
not describe it.

HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed
and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much
lighter than a full installer run and far heavier than reading back a version
string. Two tests lock the ordering against both neighbours, including one
asserting real margin over the censored 15040ms observation, since a bound that
merely matched what was measured would be the same defect again.

Scoped deliberately: the other ~360 runHook sites keep their current bounds. This
adds the norm and applies it where a real failure demonstrated the need, rather
than sweeping a value across sites with no evidence for any of them.

Refs #3271

---------

Co-authored-by: sim <sim@local>
2026-08-09 23:41:44 -04:00
Tom Boucher
95d0da9060 docs(#3287): add ADR-3180 decision 8, the diagnostic contract (#3292)
Co-authored-by: sim <sim@local>
2026-08-09 23:32:11 -04:00
Rezolv
2dbee3ebdd enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229.

Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone.

Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined.

To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged.

Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change).

Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes.
2026-08-09 23:05:41 -04:00
Tom Boucher
693f12ad56 refactor(#3187): give state field extraction one canonical owner (#3283)
* refactor(#3187): give state field extraction one canonical owner

stateFieldValue in state-document.cts becomes the single owner of the #1760
frontmatter-then-body fallback chain. The new whole-repo guard found 14
independent re-derivations where the epic scoped 5, all now routed through it:
cmdStateSnapshot (11), cmdStatePrune (2) and smart-entry fmScalar (1).

state validate was a gate that could not fail. Every warning it could emit sat
behind a phase resolved without the frontmatter tier, so a STATE.md whose phase
lives only in frontmatter skipped the drift scan entirely and returned
valid:true. It also read unstripped content, letting a frontmatter status: key
shadow the body field (#1255 class). Both fixed; output gains a scope field so
could-not-look stops being output-identical to looked-and-clean.

Verified on the remote runner.

* docs(#3187): document the state validate scope field and its reason codes

Adds docs/how-to/interpret-state-validate-results.md so a reader can tell
nothing-to-report from could-not-look, updates the COMMANDS.md and USER-GUIDE.md
entries, corrects the CONTEXT.md glossary overstatement about Current Position
sole ownership, and drops the changeset fragment.

* fix(#3187): close three drift-guard evasion shapes and test the refuse path

The isolated adversarial review found the ladder detector was evadable by
ordinary reformatting, not just deliberately: a member or computed operand
(fm.key / fm[key]) missed the bare-identifier backreference, a swapped tier
order missed a hardcoded number-then-boolean sequence, and a ladder wrapped
across lines missed single-line detection. All three now caught, each with its
own test plus a proven boundary control.

The frontmatter-parse refuse path on the destructive complete-phase route was
unreachable and therefore untested. It is now driven by an injected parse
failure and asserts STATE.md is byte-identical after the refusal, rather than
shipping untested defensive code on a path that rewrites user state.

Verified on the remote runner.

* fix(#3187): widen the drift guard to the prompt layer and disclose tier-2 changes

The code-review spec axis found the guard's scan surface was src/ only, which is
Decision 4(d)'s forbidden allowlist one directory wide - and it had a live miss:
gsd-core/workflows/smart-entry.md tells an agent to read status from frontmatter
or the body, a prose expression of this same chain. The surface now covers the
prompt layer. That one site carries a permanent written exemption rather than a
ratchet: it is the gsd-tools-is-down fallback, so it cannot call the owner by
construction, and a ratchet would imply removable debt that does not exist.

Two tier-2 output changes shipped undisclosed and are now named in the changeset
and docs: complete-phase's idempotency guard consulting frontmatter, and the
workstream inventory resolving frontmatter-only fields. docs/COMMANDS.md gains a
state complete-phase entry, which it never had.

Also records Amendment 5 on ADR-3180, extracts the duplicated frontmatter-parse
block the epic's own thesis forbids, and re-points two assertions from free-form
warning prose onto the structured drift object.

Verified on the remote runner.

* chore(#3187): backfill changeset PR number

pr:0 placeholder replaced with the real PR number now that #3283 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:49:39 -04:00
Tom Boucher
cf6de5e1c0 feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution

23 tests over the 50-test-matrix rows. RED by construction:
resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet,
and the validator silently ignores triggerPrecedence today.

Written in the per-runtime describe idiom the other four
runtime-artifact-layout suites use, not a table.

The rows that carry the weight: windsurf must NOT report a shadow it
does not have, since its global scope emits only agents and agents are
not trigger-bearing; agents and kimi-agents must be absent from the
output for every runtime; and reordering a runtime's triggerPrecedence
must flip the winner, which is the only assertion that proves the axis
is read rather than decorative.

Stems are injected, never scanned, so the surface is assertable with no
filesystem.

* feat(#2871): resolve triggers and host precedence, not just placement

resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger
a runtime emits, with the scope and kind that produced it, whether the
host registers it directly or only through a router, and which artifact
shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers
need placement only and the issue requires them unchanged.

AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The
host-integration matrix models command and dispatch as separate interface
points: an agent is invoked through the Agent tool's subagent_type, not
by typing a slash trigger, and _copyStaged never applies the kind prefix
to an agents entry. So agents and kimi-agents are excluded from the
surface entirely, and this commit amends ADR-2866 with a dated
correction. #2218's conclusion is unchanged -- the collision is strictly
commands-vs-skills, and claude's local /gsd-* trigger surface is still
fully shadowed -- but the ADR implied the local agents surface was lost
too, and it is not.

That correction is what makes windsurf come out right. Its global scope
emits only agents, so it has no global trigger and its local commands
are unshadowed. Model agents as trigger-bearing and windsurf falsely
reports a full shadow.

The triggerPrecedence axis lands on all 19 descriptors as an ordered
kind list, one value with one owner, rather than a numeric rank spread
across N kind entries with nothing keeping them consistent. Validation
uses a required-with-default shape that has no precedent in this
validator -- every existing axis is hard-required -- so a third-party
capability.json omitting the field still validates, which is what
ADR-894's additive-only contract promises.

Winner resolution reads Phase 1's scope rank first, then the kind
ordering. A test reorders the axis and asserts the winner flips, since
an axis that is added, validated and never consulted would pass every
other assertion.

shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this
issue's out-of-scope note.

Verified via the remote runner.

* fix(#2871): single-source namespacedByDir and close two test gaps

Four findings from the isolated adversarial review.

The namespacedByDir rule had reached three copies -- install-engine,
surface, and the new trigger resolver -- one of which carried a
hand-written keep-in-sync comment and no assertion. That is this repo's
generative-fix-divergence class. Extracted to one exported predicate all
three now call. Verified by diverging one copy deliberately: the existing
#816 parity test failed, and passes again on revert.

The omission test was vacuous. Row 16 asserted that a descriptor without
triggerPrecedence still validates, but built its fixture from claude's
shipped descriptor -- which this PR had just added the axis to. It now
clones and deletes the key, following the shippedDescriptorWithout
pattern, and asserts both that validation passes and that the resolver
still picks the right winner from the default. The second half is what
makes it prove anything.

resolveTriggerSurface silently dropped an unrecognized scope while every
sibling in this epic throws. Two phases of one epic should not disagree
about whether an invalid scope is an error, so it now rejects through the
same shared validator; an empty scope list still returns empty rather
than throwing.

The ADR amendment had been spliced into the middle of the References
list, orphaning its last bullet. Moved to the top, after the header
block, which is where ADR-3660 and ADR-1016 both put dated amendments.
No lint checks markdown structure, so this was green while malformed.

* fix(#2871): single-source the command filename composition too

The earlier fix shared the namespacedByDir boolean but left the
filename composition around it written twice -- once in _copyStaged as
what actually gets written, once in resolveTriggerSurface as what gets
predicted. The predictor could go stale silently.

One exported helper now composes it for both. The entry.name asymmetry
that looked like it would block extraction does not: entry.name is
filtered to end in .md and stem is entry.name minus those three
characters, so the two branches are the same string by construction.

Divergence proven to fail: injecting a marker into the helper broke the
trigger-surface suite; reverting restored 25/25. The four sibling layout
suites hold at 227 unchanged.

* docs(#2871): correct the ADR timing notes that this phase makes stale

The Amended by back-links on ADR-3660 and ADR-1016 were written in
Phase 0, when the widenings they describe had not shipped. Each carried
a forward-looking clause -- "the module changes at Phase 2, not before,
until then this module resolves placement only" -- which becomes false
the moment this PR merges. ADR-2866's own Amends header and its
reciprocal-notes section carried the same tense.

All four now describe what shipped. This is a tense and status
correction on Accepted ADRs, not a change to any decision.

Worth stating because it is the failure mode this epic keeps meeting:
gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI
would have caught either the missing back-link in Phase 0 or these stale
clauses now. They stay correct only because someone checks.

* chore(#2871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:25:42 -04:00
Tom Boucher
4a1ed2531f enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check

Specifies ADR-2313 D6 before the implementation exists, so the tests
bind to the contract rather than to whatever the code happens to do.

RED is established by construction, not by a remote run:
checkCodexModelPosture and POSTURE_REASON are absent from the compiled
lib today, so every row fails on the missing export. A remote checkpoint
here would prove only that the function is missing, which is already
known — so the run is deliberately deferred to the combined green
checkpoint rather than spent proving a tautology.

That makes the NEGATIVE PROOFS the rows that carry real signal. Every
positive row passes even for a naive implementation that greps
/model\s*=/ over the whole file. Six rows fail it: light-tier
service_tier/model_verbosity decoupling (#774), hand-added keys, a
commented pin, the model_verbosity key-prefix collision, the runtime
no-op ordering, and the headline case — a literal `model = "sonnet"`
inside the developer_instructions ''' block, which the emitter fills
with agent prompts that discuss models constantly.

Row 14's fixture was verified to discriminate before being written: a
whole-file scan matches it and a header-slice scan does not. Without
that check the test would pass trivially and prove nothing, which is the
vacuous-test failure this epic has already hit repeatedly.

Adversarial TOML fixtures are hand-authored against the real Codex shape
rather than generated by generateCodexAgentToml, per #2371 — a fixture
from the writer can only confirm what the writer already believed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3242): validate codex .toml model posture, not just presence

Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export,
not a branch inside checkAgentsInstalled — that function carries 33
upstream dependents, cyclomatic 25, and sits in two traced process
flows, so it is deliberately left untouched.

It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf.
That is what Phase 1's constant move bought: agent-install-check is
documented as pure read/verify and imports only leaves, so reaching the
rule through model-resolver would have dragged config-loader into it.

Reads liberally, judges strictly, and never guesses. Tolerates comments,
key order, whitespace, CRLF, and a BOM; anchors on full key names so
model_verbosity does not satisfy a `model` probe; treats extra
hand-added keys as none of its business, since the check is a predicate
on the two fields the posture owns rather than a whitelist over the
document. An unreadable file becomes a named violation and the loop
keeps going.

The scan covers only the header slice — the lines before the
developer_instructions ''' marker. The emitter writes agent prompts into
that block and GSD's prompts discuss models constantly, so a whole-file
scan reports violations for prose. This is the highest-risk defect in
the phase and the reason its fixture was verified to discriminate before
being written.

The non-codex short-circuit runs before any filesystem call, so a stray
.toml under another runtime is never inspected.

Wired through cmdValidateAgents as an additive codex_posture key, so a
violating install is visible from a command a user actually runs rather
than only from a library nothing calls.

Also fixes a test defect found while implementing: .gitattributes forces
`* text=auto eol=lf` repo-wide, so the committed CRLF fixture was
normalized to LF in the index — `git ls-files --eol` reported `i/lf
w/crlf`, the working copy being stale pre-normalization bytes. The CRLF
row was asserting against a file that could not survive a fresh clone.
CRLF is now derived at runtime, which puts it under the test's control
rather than git's, instead of adding a .gitattributes exception that
fights a deliberate repo-wide policy and that anyone could re-normalize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3242): document the codex posture check where a user will look

Three quadrants, filed by where the reader actually arrives.

How-to (recover-and-troubleshoot.md, under Install and update problems)
is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400
about an unsupported model" — and opens with the verbatim error string.
Someone hitting this does not know the words "posture" or "ADR-2313";
they have a 400 in their terminal and will search for that.

Reference (COMMANDS.md) had no `validate agents` entry at all, though
sibling gsd-tools subcommands are documented. Adding user-visible output
to an undocumented command and then linking to it from the new how-to
would have left a dangling reference. The entry carries the
violation-reason table, since the frozen POSTURE_REASON enum is the
machine-readable contract a reader needs rather than the prose.

Both surfaces state that presence and posture are separate verdicts — a
missing agent lands in `missing`, never as a posture violation. That is
a deliberate design decision and would otherwise be invisible to someone
watching one command emit both.

Explanation stays in ADR-2313, which already covers D6 and the
liberal-parse/strict-judge boundary. Pointing at it beats duplicating it
into COMMANDS.md and creating two copies to drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): close two false negatives in the posture scan

Both found by an isolated reviewer and reproduced before fixing. Both
made the check report clean when it was not — the worst direction for
this function, since the how-to tells users an empty violations list
means the install is posture-clean.

A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML,
but the key pattern required a bare identifier, so the pin was silently
invisible. Bare, "double" and 'single' quoted forms now normalize to the
same key name.

The block marker was found by unanchored whole-content search and used
to truncate the header. A `description` value merely containing the
literal text `developer_instructions = '''` truncated the scan before a
real pin, and a user who hand-reordered `model` to sit after the block —
still legal TOML — was never scanned at all.

Fixed by changing the strategy rather than the regex: find the block's
range, anchored at line start, and scan every line OUTSIDE it. That
covers both failures and is strictly more correct than truncation, while
still never reading prompt prose. An unterminated block excludes the
rest of the file, which fails toward a false positive — the safe
direction, since misreading prose as a pin wastes a user's time while
the alternative hides a real one.

Also corrects two overclaims of mine. The how-to named "v1.11", a
version that does not exist — package.json is 1.10.0 and unreleased — so
it now describes the boundary by behavior and links the ADR. And the
test matrix asserted that a naive whole-file scan "fails exactly rows
12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16
produce the correct result against that baseline too. They guard real
but *different* mistakes, and the matrix now says which one each catches
instead of attributing them all to the header-slice defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): skip symlinked agent files instead of following them

Security review finding. The scan listed entries with readdirSync and
read them with readFileSync, which follows symlinks — so a symlink in
the agents directory pointing anywhere would have its contents read, and
any line matching the model pattern echoed into cmdValidateAgents'
output through the `value` field. A read-and-echo primitive on an
arbitrary path.

It needs write access to the agents directory, so it crosses no new
trust boundary today. Fixed anyway, for two reasons.

This repo already does it correctly next door: cmdEffortSync filters
with lstatSync().isFile() and the comment "Skip symlinks — only write
regular files to avoid clobbering symlink targets." Being inconsistent
with a sibling in the same subsystem IS the defect.

And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these
files. Establishing symlink-following as the house pattern for Codex
.toml handling here would hand Phase 3 a worse starting point while it
writes rather than reads.

Skipped silently rather than reported, matching the sibling: a symlinked
agent file is a structural install choice, which checkAgentsInstalled
owns, not a model-content posture defect. An lstat that itself throws
excludes the file rather than crashing the scan.

That does narrow the guarantee slightly, so the how-to now says an empty
list means every REGULAR .toml is clean, and tells anyone symlinking
their configs to check the targets by hand. Claiming a clean bill of
health over files the check declined to open would be the same kind of
false confidence the two false negatives above produced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3242): backfill changeset pr number (#3290)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 21:50:49 -04:00
Tom Boucher
c2f24265f2 feat(#2870): resolve install scope as a value (#3278)
* test(#2870): failing-first suite for the Install Scope Module

19 tests over the 50-test-matrix rows 1-19. RED by construction: the
module under test does not exist yet, so the suite fails at require with
MODULE_NOT_FOUND until src/install-scope.cts lands.

Every row asserts a returned value with injected env/home/existsSync --
no filesystem, per the issue's acceptance criterion that tests assert the
resolved value directly.

Row 7 asserts the RELATION rank(global) > rank(local) rather than a
literal, so Phase 2 (#2871) can re-base the numbers without a fixture
edit. Row 4 iterates the real runtime registry rather than a hardcoded
list, excluding vscode, which declares configHome.kind none and is never
CLI-installed.

* feat(#2870): add the Install Scope Module

Scope becomes one resolved value instead of a bare string re-derived at
every layer. resolveScope({id, runtime, ...}) returns
{id, configHome, settingsFile, consentRequired, hostPrecedenceRank}.

It COMPOSES resolveConfigHomeFromDescriptor rather than extending it.
That function has 60 dependents across 13 files and 2 process flows -- a
CRITICAL blast radius -- so adding a scope parameter to it, which the
issue's framing invites, would ripple through all of them. Composing
costs nothing and leaves every existing caller byte-identical.

The module owns the InstallScope type name, which previously lived
privately in runtime-artifact-install-plan.cts; that module now imports
it. A fifth spelling of the same concept would have defeated the phase.

settingsFile is null for the 18 runtimes that declare no
settingsFileByScope -- absence is a value, not an error, and inventing a
Claude-shaped default would leak that host's shape onto every other one.

hostPrecedenceRank ships unread: Phase 2 (#2871) is its first consumer.
It is carried as data only, per this issue's out-of-scope note that
precedence semantics belong to that phase.

Vocabulary: the install axis standardizes on local. ConsentRecord.scope
keeps project deliberately -- that literal is serialized into consent
records in the user's home, and renaming it would silently deactivate
every project-scoped capability on the machine. CONTEXT.md records the
boundary mapping instead.

Every environmental input is injectable (env, home, existsSync, cwd), so
the resolved value is assertable with no filesystem at all.

Registration ripple: .gitignore, eslint.config.mjs, CONTEXT.md glossary,
docs/INVENTORY.md, and the inventory manifest (regenerated after
build:lib, never before).

Verified via the remote runner.

* refactor(#2870): route scope re-derivations through the module

bin/install.js resolves scope once per function instead of inline at
each of its 12 sites, and the settingsFileByScope consumer reads it
through resolveScope().

Seven downstream boolean re-derivations now call the module's
isGlobalScope() instead of comparing the literal independently:
runtime-artifact-install-plan, both runtime-artifact-layout kind
builders, dispatchKindEntry, surface, and two install-engine sites. The
fifth through seventh were not named in the issue -- they are the same
re-derivation class, and leaving them would have made the acceptance
criterion false.

_computePathPrefix keeps its isGlobal boolean API, so the projection is
centralized rather than eliminated. resolveScope and isGlobalScope share
one validator, so the two surfaces cannot drift.

TWO SITES DELIBERATELY NOT ROUTED: runtime-artifact-conversion's
rewriteStagedSkillBodies and rewriteStagedCommandBodies. ADR-1508 fixes
the direction as installer/layout -> conversion, never upward, and
install-scope composes runtime-homes, so importing it into the
conversion module would invert that direction. Left as-is on purpose.

Behavior-preserving throughout. Each step was proven by capturing full
layout and plan output -- including every kind's home field and the
hashed contents of emitted files -- before and after, across both scopes
for claude, codex, opencode, hermes, kimi and kilo. Byte-identical.

surface.cts keeps a scope ?? 'global' default before the call because
Layout.scope is optional there; isGlobalScope throws where the old
inline compare returned false, and that difference would have been a
placement regression.

Verified via the remote runner.

* fix(#2870): cover the no-config-home throw and document the strictness

Two findings from the isolated adversarial review.

The vscode case was implemented but untested. resolveScope throws for a
runtime whose descriptor declares configHome.kind 'none', which is the
design's own behavior-table row 13, but the registry sweep excluded
vscode rather than asserting the throw -- so the behavior shipped with
no test. The exclusion is now legitimate because the case has its own
test naming the runtime in the assertion.

isGlobalScope throws where the inline compare it replaced returned
false. No reachable caller can deliver an out-of-union value today, but
the types are not enforced at runtime, so a future caller passing an
optional Layout.scope would crash rather than silently misroute. That is
the better failure -- misrouting writes artifacts to the wrong place --
but it was undocumented, so the reason is now on the function.

Adds the changeset the acceptance criteria require.

* refactor(#2870): route the last two sites; correct the ADR-1508 claim

The previous commit declined to route runtime-artifact-conversion's
rewriteStagedSkillBodies and rewriteStagedCommandBodies, claiming
ADR-1508's dependency direction forbade the import. That reasoning was
wrong, and this commit corrects it.

Two independent reviewers checked the actual import graph:
runtime-artifact-conversion already imports capability-registry,
command-roster, runtime-name-policy and shell-command-projection -- it
depends on leaf-tier siblings today. install-scope imports only
runtime-homes plus node builtins, and runtime-homes imports only node
builtins, so there is no cycle at any depth. ADR-1508 governs the
installer/layout to conversion boundary, not a leaf-to-leaf sibling
import of the same shape conversion already makes.

With those two routed, every isGlobal re-derivation in the tree now goes
through one owner and acceptance criterion 1 is fully met rather than
partially. Nine sites, not the four the issue enumerated.

Also from the review:

Tests were falling through to the real process.cwd() at five local-scope
call sites, which contradicts the acceptance criterion that the resolved
value be assertable with no filesystem. Every one now injects a cwd. One
of the five was a site the review had not spotted.

bin/install.js carried two near-identical copies of the guarded
resolveScope block, one in install() and one in uninstall() -- duplicated
scope logic in the phase whose purpose is removing it. Extracted to one
helper, and the new sites use the file's existing ternary idiom rather
than the if/else that replaced it.

Equivalence re-proven across both scopes for claude, codex, opencode,
kilo and hermes, now including the staged skill and command body
rewrites hashed per file, since those decide the literal spec-root path
baked into every emitted artifact. Byte-identical.

Verified via the remote runner.

* fix(#2870): assert configHome portably instead of with a native separator

The windows-latest node24 shard failed on two install-scope assertions.
The module was right and the tests were wrong: they built their expected
value with path.join, which emits \fake\home\.claude on Windows, while
resolveScope normalizes separators unconditionally to /fake/home/.claude.

That unconditional normalization is deliberate -- backslash paths arrive
on Linux too, so normalizing via path.sep is the documented defect this
repo guards against. Weakening it to make the assertion pass would have
inverted the fix.

Every path.join-built expectation in the suite now goes through
toPosixPath from tests/helpers.cjs, which is the pattern the
no-path-literal-in-assert rule's own valid-case list sanctions. It splits
on the running platform's path.sep and rejoins with forward slashes, so
it reverses whatever path.join produced on that same platform and the
expectation is invariant everywhere.

Two more call sites had the same latent problem and passed on Linux and
macOS by luck; they are fixed too.

This is the class of defect the remote runner structurally cannot catch
-- its matrix is Linux-only, so a green pass there is not evidence of
portability, and CI's Windows lane is the only place it surfaces.

Verified via the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-09 20:16:23 -04:00