Commit Graph

2815 Commits

Author SHA1 Message Date
Tom Boucher
3fd03bec4c fix(#3760): refuse a legacy-key migration into a non-object config section (#3767)
* test(#3760): failing-first regression for non-object config section

Locks the contract from the issue's Expected section before any fix exists:
a legacy-key section holding a string, number, boolean or array must be
preserved verbatim, reported, and never persisted in an expanded form.

Covers both blocks the issue names (branching_strategy -> git.*, sub_repos ->
planning.*), the migrateOnDisk multiRepo branch that shares the shape, and the
loader write paths that are what actually reach the user's config.json.
Includes the negative-space cases that must keep hoisting ({} , null, absent
section, canonical-nested-wins) and two fast-check properties.

Refs #3760

* fix(#3760): refuse a legacy-key migration into a non-object config section

normalizeLegacyKeys hoisted a legacy top-level key into its canonical nested
section by spreading `result[section] ?? {}`. `??` guards only null and
undefined, so a section holding a string was enumerated by index —
`{...'main'}` is `{0:'m',1:'a',2:'i',3:'n'}` — while a number or boolean
spread to `{}` and the value vanished. Because a fired block always pushed a
Normalization, and every caller treats a non-empty normalizations array as
'config is dirty', that shape was written back to .planning/config.json and
the original value became unrecoverable.

Both blocks the issue names are fixed via one shared hoistLegacyKey helper,
plus the two further sites that share the shape and are reachable from the
same input: migrateOnDisk's multiRepo branch, and the loader's two
`if (!planning) planning = {}` guards, where a non-empty string is truthy and
the following assignment threw a strict-mode TypeError that the enclosing
catch swallowed — discarding the user's entire config.

A present non-object section now blocks its own migration. The section, the
legacy key, and the file are left byte-identical; no Normalization is pushed,
so nothing marks the config dirty; the refusal is reported in-band as
`skipped[]` and out-of-band through the ADR-1411 warnUnusableInput seam
(new frozen reason config_section_not_object). null and undefined keep their
long-standing 'absent' meaning and still create the section.

This is the nested-section analog of the ADR-227 shape check _readConfigFile
already performs on the top-level document: valid JSON is not a config object.
isConfigSection is exported and shared by both modules rather than copied.

Fixes #3760

* fix(#3760): keep the multiRepo marker when planning cannot receive it

Follow-up from the isolated adversarial review, and the same defect class as
the two blocks the issue names — in the block it did not name.

normalizeLegacyKeys block 3 deleted `multiRepo` and pushed a Normalization
before anything consulted the planning section, deferring 'can this section
receive sub_repos?' to the caller that runs filesystem detection. By then the
marker was already gone and the config was already dirty, so with
{"multiRepo":true,"planning":"docs"} the loader wrote the file back with
multiRepo removed, the sub_repos injection silently no-opped against the
string, and no diagnostic was emitted at all. migrateOnDisk warned for the
same input; the ~30-caller loadConfig path did not.

Section validity is knowable from the parsed config alone — detection is only
needed for the VALUE, not for whether the destination can hold it. The refusal
moves into block 3: the marker is kept, no Normalization is pushed, and a
skipped entry is recorded, so all three callers inherit the preservation and
the diagnostic together. The caller-side guards drop to pure narrowing.

Also from review: skipped[] now reports sectionType ('string' | 'number' |
'boolean' | 'array') instead of sectionValue. migrateOnDisk's report is printed
verbatim by `migrate-config`, and this module already masks config values on
the set/unset output path; the type is the whole diagnostic and the value is
still in the file.

And `migrate-config --raw` no longer answers a refused migration with 'No
legacy keys found — config is already canonical.' Legacy keys WERE found and
declined, and the decline is the one thing only the user can fix by hand.

Refs #3760

* fix(#3760): keep configuration.cjs dependency-free; emit from its callers

The remote matrix caught a regression my own change introduced: adding
`require('./unusable-input.cjs')` to configuration.cts broke the #3571
install-layout contract. `configuration.cjs` must load from a layout holding
only itself plus bin/shared/*.manifest.json — the installer does not co-locate
arbitrary siblings — so the new require failed at load time:

  Cannot find module './unusable-input.cjs'
  Require stack:
  - /tmp/gsd-3571-.../.codex/gsd-core/bin/lib/configuration.cjs

pinned by 'co-located bin/shared manifests let configuration.cjs load without
sdk/shared' in tests/install.test.cjs (3 failures).

The contract is deliberate and the test is right, so the module goes back to
zero sibling requires and the out-of-band diagnostic moves to the callers that
already carry a dependency budget and hold the resolved path: cmdMigrateConfig
(config.cts) and loadConfigResolved (config-loader.cts). normalizeLegacyKeys
keeps reporting refusals in-band via skipped[], which is what lets it be pure
and dependency-free at the same time.

The emission-count and dedup assertions move to tests/config-loader.test.cjs,
where the diagnostic now originates. A new assertion pins the inverse for the
module itself — migrateOnDisk must emit ZERO diagnostics while still reporting
skipped[] — so regrowing a sibling require fails a unit test instead of only
the install suite. CONTEXT.md records why the emitter is the caller.

Refs #3760

* chore(#3760): backfill changeset pr number to 3767

---------

Co-authored-by: sim <sim@local>
2026-08-22 18:57:19 -04:00
Tom Boucher
b6977d9d11 fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps (#3766)
* test(#3762): failing-first roster gate for docs/INVENTORY.md rows

Anchors the human half of the inventory-drift rule: every entry in
docs/INVENTORY-MANIFEST.json must have a hand-written row in
docs/INVENTORY.md. Expected RED on this commit -- next carries 32
unrostered surfaces, which is the defect the gate exists to catch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps

docs/INVENTORY.md calls itself the authoritative roster of every shipped
GSD surface, and CLAUDE.md's inventory-drift rule requires both a roster
row and a manifest regen. Only the manifest half was anchored, so a PR
could ship a surface, regenerate the manifest, omit the row, and stay
green -- as PR #3758 did with gsd-core/references/planner-coupling.md.

Adds the roster half to tests/inventory-manifest-sync.test.cjs, backed by
a pure matcher in tests/helpers/inventory-roster.cjs, and backfills the 32
surfaces already missing rows on next.

Fixes #3762

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3762): harden roster matcher against fenced blocks and trim exports

Skips fenced code regions when splitting level-2 sections so a documented
'## ' example inside a fence cannot truncate a family section (false red)
or contribute a phantom row (false pass); makes the heading pattern linear
rather than a backtracking lazy match; narrows the module surface to the
three names the gate consumes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3762): close two false-RED gaps found in orthogonal review

Indented headings: CommonMark permits an ATX heading to carry 1-3 leading
spaces, and a ^##-anchored pattern read such a document as having no family
sections at all -- reporting all six missing, a structural red for zero real
drift. Reproduced, then fixed and pinned.

Link-wrapped cells: a row written as [`x.md`](../x.md) was not recognized,
though docs/INVENTORY.md already uses that form elsewhere. Unwrapping now
peels a whole-cell link and a whole-cell code span, and only layers that
wrap the cell entirely -- a file mentioned mid-prose is still not a row, so
the false red is not traded for a false pass.

Adds a second fast-check property over the commands source-link rule, the
one family whose matching rule differs from the other five.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3762): boundary trio for the fence-marker length limit

RULESET.TESTS.boundary-coverage — the matcher's only numeric limit is the
fence marker's {3,}. Exercises 2 (inline markup, swallows nothing), 3, and
4 characters, for both backtick and tilde delimiters.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: remove stray pwned_cmdsub injection-test canary from the repo root

A zero-byte file committed by 0e6fa2e2c (#3124) while remediating the
#3118 command-substitution injection — the marker a test wrote into cwd to
prove a substitution had NOT executed, left behind when the run ended.
Nothing in the tree references it (verified by Grep across the repo),
package.json's files array excludes root-level files so it never shipped,
and lint-removed-but-needed confirms no surviving reference.

Found while auditing the repo root for #3762; fixed inline rather than
deferred, per CLAUDE.md's no-deferrals rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: backfill changeset pr number to 3766

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 18:24:25 -04:00
Tom Boucher
9ade7926ca test(#3761): replace vacuous wave/files_modified gate with an anchored one (#3764)
The test 'planner has validation step or quality gate for wave/files_modified
consistency' asserted a six-token disjunction over whole-file substrings, five
of which occur ZERO times in agents/gsd-planner.md (validate_waves,
wave_validation, same wave — the last killing the entire quality-gate arm,
since that file has no quality_gate block). The sixth matched one incidental
pseudocode comment, so deleting the wave-ordering gate outright left the suite
green.

Per RULESET.TESTS.delete-bad-tests (CONTEXT.md:595) the vacuous test is DELETED
and replaced, not patched. The replacement anchors on the normative **Rule:**
sentence inside <step name="assign_waves"> — scoped to that step rather than to
the whole document, and requiring all four semantic clauses (wave scope,
files_modified subject, prohibition, overlap predicate) to hold within a single
sentence, so co-occurring words scattered across a paragraph do not pass.

agents/gsd-planner.md is unchanged: the anchor already exists there and survives
PR #3758's edits to the same region.

Teeth proven on the remote runner: a control commit paired these tests with a
deliberately gutted assign_waves step that retained the incidental phrase. The
new anchored test failed while the sibling test using the old predicate's
surviving arm stayed green — same input, opposite verdicts.

Fixes #3761

Co-authored-by: sim <sim@local>
2026-08-22 17:52:23 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
738f42f4fd feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs (#3755)
* test(#2398): failing-first suite for the CYCLE_SUMMARY consensus gate

Binds the gate before it exists, so the suite is RED against next.

The load-bearing rows are the two the closed PR #2417 did not have. The B2
regression row asserts a judgment-class lone HIGH counts WITHOUT corroboration
when its raiser is unmarked — if anyone re-couples that class to corroboration,
more reviewers again produce a weaker gate than one, which is what closed #2417.
The parity row asserts every marker literal the gate names is one
review-lane-runner actually emits, so the gate cannot key on a signal nothing
produces; a mutation row and a seeded fast-check property prove that guard runs
its failure branch rather than only reading a correct tree.

Also pinned: gate position before Counting rules, the untouched CYCLE_SUMMARY
line shape the orchestrator greps, fence balance, the single-reviewer no-op,
classification by what a claim asserts rather than by citation presence, the
all-marked fail-open, current_actionable staying out of scope, and the
leading-marker requirement that stops a review which merely quotes a marker
from suppressing its own findings.

* feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs

With review.reviewer_instances running several reviewer identities off one
adapter, any single instance's fabricated HIGH could force a full replan cycle
on its own. Across ~9 real cycles on two projects each of four instances
fabricated at least once, and each was also the most accurate reviewer in some
other cycle, so dropping to fewer reviewers trades away real signal.

The gate engages only when 2+ reviewers actually ran, and weighs a lone HIGH by
what the claim asserts rather than by whether anyone agreed with it. An
existence claim -- a symbol, file, flag, commit or ID exists, is absent, or says
something specific -- counts only if source-grounding confirms it or another
reviewer raised the same concern. A judgment claim -- a design or correctness
property -- counts unless that reviewer's own section opens with an
evidence-quality discount marker the review lane already stamps
([reviewed-without-source-citations] #3194, [reviewed-without-repo-access]
#2176, or a diff-only lane).

That split is what resolves B2, the finding that closed PR #2417. B2 showed the
approved wording made more reviewers produce a WEAKER gate than one: condition
(a) pointed at the source-grounding pass, which verifies every symbol THE PLAN
cites and never takes reviewer claims as input, so a genuine architectural HIGH
that one reviewer caught and another missed was neither groundable nor
corroborated and stopped gating. Judgment-class findings are therefore exempt
from corroboration entirely -- reviewers catch materially different classes of
issue, and demanding two of them independently raise the same architectural
concern suppresses exactly what a multi-reviewer setup exists to surface.

Guards on the gate itself: an all-marked cycle disengages it, so a cycle in
which nothing was verified can never be counted as converged; the marker must
OPEN a reviewer's section, so a review that merely quotes a marker does not
suppress its own findings; a suppressed HIGH stays listed and tagged rather
than dropped; current_actionable is untouched; and a single-reviewer run is
unchanged.

No new command, config key, or dependency -- the gate reads signals that
already exist. The CYCLE_SUMMARY line shape the orchestrator greps is
unchanged; only the integer it computes moves, and only for 2+ reviewers.

Known limit, inherited rather than introduced: SOURCE_CITATION_RE checks
citation presence, not resolution, which src/review-lane-runner.cts records as
a deliberate #3194 scope boundary. A fabricated but plausible file:line still
gates.

Scope revised and re-approved on the issue before any code was written.

* test(#2398): make marker parity behavioral, and stop overclaiming the gate

Review found the parity tests were vacuous: they asserted a marker STRING
appeared in review-lane-runner.cjs's source text, never requiring the module or
calling the stampers, so they would pass even if stampUngroundedReview were
broken or never invoked. They now invoke the real exported functions and assert
what those functions PRODUCE — that an uncited review gains a leading marker
blockquote, that a review carrying a file:line does not, that a self-reported
blind review is stamped, and that stamping is idempotent. Removing the source
read also removes an incidental no-source-grep evasion via a parameterized path.

Review also found the changeset headline false for the class it matters most
in. The discount markers detect 'cited nothing' and 'had no repo access'; they
cannot detect 'drew a wrong conclusion from a real citation', so a judgment-class
finding invented by an evidence-bearing reviewer still counts alone. That is the
deliberate side of the tradeoff jags-faith named when closing #2417 — the
alternative is requiring corroboration for design findings, which is B2 — but
the changeset claimed lone hallucinations no longer force a cycle, full stop.
Corrected there, and stated plainly in docs/COMMANDS.md and the design record.

Also dropped the reviewer-instances.md entry from the emitted-drift ack: the
growth ratchet's currentSizes() scans only gsd-core/workflows/ and agents/
(tests/helpers/emitted-runtime.cjs:916-929), so references/ is outside it and
that entry acknowledged a delta the gate cannot see.

* chore(#2398): backfill changeset pr number to 3755

---------

Co-authored-by: sim <sim@local>
2026-08-22 10:53:14 -04:00
Behruz Nassre Esfahani
444069d601 fix(#3613): copy component dirs into the plugin-validate fixture (#3627)
* fix(#3613): copy component dirs into the plugin-validate fixture

C2 builds its synthetic plugin root by symlinking commands/, hooks/ and
skills/ into a temp dir. `claude plugin validate` (>=2.1.233) reads
component directories without following symlinks and warns on each one,
and `--strict` promotes a warning to a non-zero exit — so the assertion
failed on how the fixture was built, not on the manifest under test.
Copy the directories with fs.cpSync instead. The validated tree still
holds only plugin.json and the three component directories, so nothing
else from the repo root is placed where the validator can read it, and
the #2665 CLI-config isolation is untouched.

The construction helper is self-cleaning: its callers' try/finally only
begins once it returns, so a throw partway through would strand a
half-built root on disk. The pre-refactor code ran these same steps inside
C2's own try, and that teardown guarantee is preserved rather than
narrowed. Its cleanup is best-effort so a teardown error cannot replace
the real construction error.

Add C3, an unconditional tripwire asserting the fixture exposes real,
non-symlinked component directories at any depth. C2 never runs in CI —
no job under .github/workflows/ installs the `claude` binary — so a revert
to symlinks would pass every CI lane and surface only as a red suite on
contributor machines. C3 is not a substitute for C2's end-to-end check and
cannot be shown red against this base, since the helper it calls arrives
with it; C2 is the failing-first artifact.

C3 enforces a symlink-free fixture throughout, deliberately stricter than
the CLI's own boundary. Measured on 2.1.234, `plugin validate --strict`
exits 1 for a symlinked component dir and for a symlink one level inside
skills/, and 0 for two levels in or for symlinks under commands/ or
hooks/. Encoding that external, undocumented line would be more fragile
than a superset costing one walk over ~176 files. The walk uses an
explicit stack rather than readdirSync's `recursive: true`, which follows
directory symlinks — a link pointing outward would otherwise traverse an
unrelated tree, or a cycle, before the assertion ran.

Correct the Section C docstring, which claimed C2 provides defence-in-depth
coverage it cannot provide in CI. Whether to provision the CLI in a CI job
is a maintainer call and is left open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): skip hooks/dist and .dist-staging when building the fixture

The recursive copy added for #3613 races the hook builders. build-hooks.js
writes atomically through a per-PID hooks/.dist-staging-<pid> and removes
it when done, and nine test files invoke that script from their before()
hooks, so a recursive walk can enumerate a staging directory and then
lstat it after the owning process deleted it — the ENOENT that #3656 just
fixed in the cold-tree fixture.

Copy entry-by-entry and skip by NAME BEFORE anything stats it, reusing
shouldCopyHookEntry from tests/helpers/cold-runtime-lib-fixture.cjs rather
than re-deriving the rule. A filter applied after the stat would not close
it. The predicate is already pinned including its over-match cases, which
a loose startsWith('dist') would get wrong: dist-staging-no-dot and
distant.js must both be kept.

Excluding hooks/dist is independently right for this fixture — a real
marketplace install contains neither dist nor a transient staging dir,
the same reasoning that keeps the repo-root CLAUDE.md out of the
validated tree. C2 still validates clean without it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): validate agents/ too, and scope the hooks predicate

Three review Minors.

agents/ ships in package.json `files` and IS auto-validated by the CLI —
verified: a frontmatter-less agents/*.md exits 1. Including it was
pointless while the fixture symlinked, because the CLI read nothing
through a symlink; now that the tree is real it is the last shipped
component tree C2 could not see. Proven to buy coverage rather than
bytes: planting a frontmatter-less agent now reds C2, which it could
not do before this change.

shouldCopyHookEntry is documented as a hooks/ entry filter, so it now
runs only for hooks/. Applying it to the other trees was harmless today
but silently encoded a hooks-shaped exclusion into them — a future
commands/dist would have vanished from the validated tree with no signal.

assert.deepEqual -> deepStrictEqual on the nested-symlink assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): scope the fixture to the three directories the issue names

Review findings, all five.

Major 1 — `agents/` dropped from COMPONENT_DIRS. The observation behind
adding it was right (the CLI does auto-validate it; a frontmatter-less
agents/*.md exits 1), but it is a NEW gate the issue does not ask for, on
the largest of the trees, and since C2 never runs in CI it would be red
only on contributor machines with `claude` installed — the same
worst-of-both-states #3613 exists to remove. Worth having as its own
issue, where "should CI provision the CLI" gets answered for it too.

Major 2 — C3 no longer passes on a fixture that validates nothing.
buildValidationPluginRoot() mkdir's every component dir before the entry
loop, so a copy that stops happening leaves real, EMPTY directories and
all three structural assertions still hold (an empty tree contains no
symlinks). Adds a non-empty assertion plus a known entry per tree, so a
partial copy is caught as well. Stubbing the copy now reds C3 with
"commands/ is EMPTY"; dropping just commands/gsd reds it too. Neither
did before.

Minor 1 — the measured table was wrong, and the mistake was measuring an
inert file. Re-measured on CLI 2.1.239, reproducing the review's 2.1.237
result: a symlinked skills/<name>/SKILL.md at depth 2 exits 1, while a
stray symlinked *.md at the same depth exits 0. The boundary is not depth
at all, it is whether the symlink is a file the CLI reads as a component.
Table replaced with that.

Minor 2 — the hooks name filter is now local instead of importing
cold-runtime-lib-fixture.cjs's, whose docstring scopes it to the cold-tree
fixture and which had exactly one caller. Two consumers across fixtures
with different requirements and nothing asserting they stay compatible is
how a later cold-tree change silently alters what this fixture validates.

Nit 1 — cost figures re-measured: ~1.06 MB over 176 files, not 2.1 MB
over 243.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): correct the hooks row and the count the agents/ removal left behind

Three comment-prose items from review, no code change.

N1 — the corrected table gained a new wrong row, erring unsafe. "symlink
inside commands/ or hooks/ -> 0" is false for hooks/hooks.json, which is
the single most likely thing anyone would symlink there. Re-measured on
CLI 2.1.239, matching the review's 2.1.237 figures on every cell:

    symlink inside commands/ (a dir, or a component *.md)  0
    stray symlinked dir or *.js inside hooks/              0
    symlinked hooks/hooks.json                             1

hooks/hooks.json is now its own row, and is named in the C3 assertion
message alongside SKILL.md — that message is what the next engineer
actually reads.

N2 — "the four component directories" was residue from the revision that
also copied agents/; it survived the very edit that removed it, three
lines above a sentence saying three. Now three.

N3 — the promised agents/ follow-up is filed as #3751 and the comment
cites it, rather than promising an issue that did not exist. It frames
the real decision (should CI provision the claude CLI) rather than just
asking for the directory back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-22 08:06:52 -04:00
Tom Boucher
0062f6d033 fix(#2845): stop the dimension parity guard counting a back-reference (#3752)
next went red on dacae9273 against documentation that was correct.

parseDeclaredCounts read any '<numeral> ... dimensions' collocation as a claim
about the gsd-ui-checker dimension TOTAL. The how-to sentence 'Dimension 7 is a
rule gsd-ui-checker follows, the same as the other six dimensions' refers to the
other members of a seven-member set; the guard counted it as that document
declaring six, and reported a count-mismatch on prose that was accurate. A drift
guard that fires on correct prose is a false positive, which is how guards end
up switched off.

A negative lookbehind now excludes a numeral introduced by 'other' or
'remaining'. It is declared once as NOT_A_BACK_REFERENCE and shared by both
English scans rather than written at each site — the same two-surface divergence
class this suite exists to catch. Both scans apply it case-insensitively; the
first cut of this fix had the digit scan case-sensitive and the word scan not,
so a sentence-initial 'Other 6 dimensions' still slipped through. The exclusion
is word-anchored, so 'another six dimensions' — a real claim about a second set
— still counts.

parseDeclaredCounts takes an excludeBackReferences opt-out so a test can prove
against the REAL shipped how-to that the exclusion is load-bearing: with it off
the file reports [6,7], with it on [7]. That replaces a raw substring match on
prose, which the suite's own header forbids, with a typed before/after.

The docs prose is deliberately unchanged. It is the only instance of the pattern
in the tree, so keeping it means the real-tree assertion exercises the path this
fix exists for instead of asserting only on fixtures.

Regression tests cover both polarities: eight back-reference shapes including
all four sentence-initial cases, five real count claims that must still count,
the 'another' word-boundary case, and the shipped how-to itself.

Known limits recorded in the code: the exclusion is English-only, because
translated docs here are corrected to match English rather than authored, so
there is no instance to model the grammar on; and it cannot distinguish a
back-reference from a genuine total opening with the same word ('Other 6
dimensions were added'), where the false negative is the safer side of the trade.

Why it reached next at all: the docs PR (#3746) was green. A doc-only diff
inert-skips the test matrix in the PR lane, so the guard that reads docs never
ran against the docs change that broke it — it fired on push to next, after
merge.

Co-authored-by: sim <sim@local>
2026-08-22 06:39:48 -04:00
Tom Boucher
4918c62d76 feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance

Binds two shared formats before either exists, so the suite is RED against
next: the gsd-ui-checker dimension roster (asserted independently on twelve
surfaces, eight English and four translated) and the provenance-line grammar
the UI-SPEC template emits and Dimension 7 consumes.

Every parity assertion is paired with a synthetic mutation case, so the guard's
failure branch executes rather than only reading a correct tree: limit-1 (a
surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a
label that drifts on one surface only, a non-contiguous roster, a duplicated
number, and a surface that stops declaring a count at all. A seeded fast-check
property renders the roster under formatting noise (CRLF, padding, interleaved
sections) and asserts the parse round-trips and is strictly sensitive to a
dropped heading.

Assertions are on parsed typed records, never raw substrings.

* docs: normalize design-a-ui-phase how-to to American English

House style for docs/ is American English (CLAUDE.md). This file carried
colour/initialisation/initialise/artefact throughout. Spelling only — no
content change; kept separate from the #2845 feature commit so the
release-notes classifier and the hotfix cherry-pick filter see it for what
it is.

* feat(#2845): require provenance for UI-SPEC component inventories

A UI-SPEC's component inventory was treated downstream as a closed allowlist
while the document recorded nothing about whether the list had been enumerated
from the installed design system or recalled from memory. A recalled inventory
is indistinguishable from an enumerated one, so an executor complying with the
spec builds against a fraction of what the package offers, and every gate stays
green because they assert semantics rather than composition.

The UI-SPEC template gains a Component Inventory slot carrying one of two
provenance lines: the command that enumerated the list, the count it returned,
the resolved package@version and the date; or a Could not enumerate record with
a real reason. gsd-ui-researcher gains an enumeration ladder and must record
the line rather than write the list from recall.

gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count
with no command, an empty could-not-enumerate reason, or a line still carrying
the template's unfilled placeholders BLOCKs; a partial line, a line placed below
its table, or an honest negative record FLAGs; a complete line passes, and so
does a spec carrying no inventory at all, which keeps every UI-SPEC predating
the dimension validating unchanged. Whatever the verdict, an unsourced inventory
is reported as a non-exhaustive list of known-good components rather than a
closed allowlist, so the executor is never blocked from a component the spec
merely failed to mention. The checker never runs the recorded command.

The dimension count moved on all thirteen surfaces that assert it, across five
languages. Also corrects the claim in the English, Korean and Portuguese how-tos
that this checker applies a scored six-pillar rubric — that rubric belongs to
/gsd-ui-review's retroactive audit.

* chore(#2845): backfill changeset pr number to 3745

---------

Co-authored-by: sim <sim@local>
2026-08-21 11:59:56 -04:00
Tom Boucher
2b42b28687 fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows

* fix(#3659): make baseref-head suppress mode-aware and thread isolation mode

* fix(#3659): review fixes - stale advice purge, message pins, mode alias

* fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin

* test(#3659): rewrite set-baseref pin, fix writeSync row stub

* chore(#3659): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 04:58:33 -04:00
Tom Boucher
1f76861202 fix(#3657): tolerate commonmark fence widths in ledger readers (#3733)
* test(#3657): fence-width tolerance regression rows

* test(#3657): fix pure-row fixtures to use appendWindow result shape

* fix(#3657): tolerate commonmark fence widths in ledger readers

* fix(#3657): restore throw-block indentation in parseJsonBlock

* chore(#3657): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:45:06 -04:00
Tom Boucher
72819a4616 fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi (#3731)
* test(#3031): failing-first coverage for opt-in ~/.kimi legacy reclaim

Drives the user-reachable installer surface against a sandbox HOME seeded
with the pre-#2755 wreckage: a GSD [[hooks]] block, hooks bundle and
CommonJS marker orphaned in ~/.kimi by a --kimi-code install.

Covers the reclaim itself plus the four guards the diagnosis identified as
negative space: opt-in only (no flag, no deletion), user-authored TOML and
hook files preserved, a --kimi install never reclaiming its own root, and
the KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory. Adds a fast-check property that stripping the block never
destroys user content.

Red until --reclaim-kimi-legacy exists.

Refs #3031

* fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi

A --kimi-code install older than 1.10.0 wrote its GSD [[hooks]] block, hook
bundle and CommonJS marker into Kimi CLI's ~/.kimi. #2755 fixed the
destination but could not reclaim what the old bug already wrote: the stale
block is byte-identical to a legitimate Kimi CLI one — both runtimes render
the same bytes for the same root, since the command paths derive from the
hooks root, not the runtime — so no inspection can tell litter from a working
install.

Cleanup is therefore opt-in. `--reclaim-kimi-legacy` on a --kimi-code install
removes GSD's own artifacts from the legacy root; without it nothing is
touched, so a dual-product machine keeps Kimi CLI's hooks and #2755's
acceptance criterion holds.

Extracts the uninstall path's removal sequence into reclaimKimiHooksRoot() and
drives both callers through it, so the reclaim removes precisely what a real
uninstall removes rather than a hand-copied second implementation. Guards the
wrong-runtime case (a --kimi install would delete its own hooks) and the
KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory.

Also corrects two pre-#2755 leftovers in the same surface that told users to
run `--kimi --config-dir ~/.kimi-code` — the form that produces this very
defect, since --config-dir moves only the skills root — and adds the missing
--kimi-code entry to the installer's own help.

Regression coverage folded into tests/kimi-upgrades.test.cjs beside the #2755
cases, per the regression-test-naming lint.

Fixes #3031

* fix(#3031): never reclaim ~/.kimi when this run also installs kimi

Found by the isolated adversarial review pass and independently while tracing
--all ordering, then reproduced.

selectRuntimesFromArgs orders 'kimi' before 'kimi-code' in both --all and an
explicit --kimi --kimi-code, and installAllRuntimes installs in that order. So
--all --reclaim-kimi-legacy installed a fresh, legitimate Kimi CLI hooks block
into ~/.kimi and then deleted it moments later from the kimi-code leg — exiting
0 and reporting success while leaving the user with no Kimi CLI hooks at all.
The collision guard could not catch it: kimi-code's own root is ~/.kimi-code, a
genuinely different directory.

The flag asserts "I only use Kimi Code"; installing kimi in the same invocation
falsifies that, so the reclaim is skipped with a notice.

Also hardens the collision guard itself. It compared path.resolve strings,
which returns false for two spellings of ONE directory — measured, not assumed:
a symlinked alias and a case variant on a case-insensitive filesystem both
compared unequal, so the guard would not have fired and the install would have
deleted its own freshly-written hooks. isSameDirectory now compares directories
via resolve, then dev+ino identity, then realpath.

Regression tests for all three cases; the two alias tests probe the real
filesystem and t.skip() where the alias cannot exist.

Refs #3031

* docs(#3031): reattach reclaimKimiHooksRoot's JSDoc to its own function

Inserting isSameDirectory anchored on the function name, which placed the
helper between reclaimKimiHooksRoot's doc block and the function it documents.
isSameDirectory ended up with two stacked doc blocks above it and
reclaimKimiHooksRoot with none.

Refs #3031

* fix(#3031): warn when --reclaim-kimi-legacy cannot apply

The flag only acts inside the kimi-code GLOBAL install branch. Passed with any
other runtime, or with --local, it was consumed in silence: exit 0, no cleanup,
no message. For a cleanup the user explicitly asked for, silence is
indistinguishable from "it ran and found nothing".

The scope warning is raised at argument-resolution time rather than inside
install(). kimi-code declares hostBehaviors.localInstallDeferred, so install()
returns early at the deferral check long before the kimi-hooks-toml branch — a
guard placed there is unreachable, which is both dead code and a linted drift
shape in this repo. Verified reachable by spawning the real installer.

Neither case is a hard error: the flag stays composable with --all, where it is
legitimately inert for the other seventeen runtimes.

Refs #3031

* docs(#3031): document every case where --reclaim-kimi-legacy skips

Refs #3031

* fix(#3031): resolve local config dirs from RUNTIME_META alone in the install harness

The remote runner surfaced this: the #3031 warning test drives a local
kimi-code install and died with "The path argument must be of type string.
Received undefined".

runMinimalInstall carried a SECOND, hand-maintained local-dir map beside
RUNTIME_META, and it had drifted — four runtimes present in RUNTIME_META
(hermes, kimi, kimi-code, zcode) were missing from it, so scope:'local' for any
of them resolved path.join(root, undefined) and threw a bare TypeError naming
neither the runtime nor the map at fault. #3023 had already hit exactly this
for pi and fixed it by adding one more entry, which left the divergence itself
in place for the next runtime to rediscover.

Local scope now reads RUNTIME_META.localDir, the same table the global branch
already reads, with the same loud named error the global branch raises. Parity
verified for all 14 previously-supported runtimes: every one resolves to a
byte-identical configDir. cline keeps its ternary — its local artifacts land at
the project root itself, which is a real exception, not a directory name.

Guarded in golden-parity-single-source.test.cjs beside the buildParityManifest
anti-divergence test, and both arms of that guard were proven able to fail.

Refs #3031

* chore(#3031): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:05:58 -04:00
Tom Boucher
65d2839b11 fix(#3651): prescribe only writes config-set accepts in integrations flow (#3732)
* test(#3651): regression rows for workflow config-write prescriptions

* fix(#3651): prescribe only writes config-set accepts in integrations flow

* fix(#3651): review fixes - single-source lane list, configSchema-derived test set

* fix(#3651): canonical cite, review wording fixes, one-element array pin

* chore(#3651): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 02:20:53 -04:00
Tom Boucher
94bc492f57 fix(#3645): tracked-source rule for planner/pattern-mapper path resolution (#3728)
* test(#3645): failing-first agent tracked-source contract rows

* fix(#3645): tracked-source rule for planner and pattern-mapper paths

* Revert "fix(#3645): tracked-source rule for planner and pattern-mapper paths"

This reverts commit 61f05e947bbdaf3b4897240c3819d349215744fb.

* fix(#3645): tracked-source rule at the spawn seam and mapper gate

* fix(#3645): review fixes - bounded block, ack merge assertion, git wording

* fix(#3645): fit the tracked-source block under the 1168 ceiling

* chore(#3645): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 23:24:42 -04:00
Tom Boucher
95f7c14413 fix(#3642): stop the single-section total_phases leak into an absent milestone (#3727)
* test(#3642): failing-first single-section leak rows

* fix(#3642): gate the unbounded total on any-milestone-section, not >=2

* test(#3642): rewrite the 3185 wrapper row to the withhold contract

* docs(#3642): glossary amendment for the >=1 sibling; changeset

* chore(#3642): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 22:10:52 -04:00
Tom Boucher
072b97d276 fix(#3641): make v005/v004 see bracket-convention phase entries (#3723)
* test(#3641): failing-first bracket-window validate rows

* fix(#3641): thread phase convention into hasphaseentries for v004/v005

* test(#3641): review rows - digit-anchor, decoy, probe parity, t.after

* fix(#3641): digit-anchor bracket entry token; thread probe scope axis

* fix(#3641): align frontmatter bound with probe; changeset

* chore(#3641): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 17:48:50 -04:00
Tom Boucher
03a3f779ce fix(#3640): isolate the drift cli e2e fixture via a --root override (#3722)
* test(#3640): failing-first --root isolation rows for the drift cli e2e

* fix(#3640): add --root override to the drift cli scan root

* fix(#3640): harden --root validation; typed exit-2 usage rows

---------

Co-authored-by: sim <sim@local>
2026-08-20 16:15:22 -04:00
Tom Boucher
9a69a86f42 enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter

Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe
for /gsd-pr-branch across six layers: pure classification and forbidden-path
predicates, real-git fixtures that run the cherry-pick filter loop end to end,
config-key registration through the real CLI and both manifests, the executed
worktree-materialization claim the issue's triage asked to establish, fast-check
properties over arbitrary path sets, and a drift guard over the shipped workflow.

Two live defects in today's shipped recipe are pinned as regressions, both
reproduced empirically first: `git rm -r --cached` stages a deletion of any
.planning/ path the target branch already tracks, so the generated PR removes the
base branch's planning files; and the same command leaves the cherry-picked file
untracked on disk, so a second commit touching that path aborts the pick with
"untracked working tree files would be overwritten" and every remaining commit is
silently dropped.

The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md
rather than restating them, so the workflow stays the single source of truth and the
suite cannot drift from what ships.

Refs #2971

* feat(#2971): strict planning filter mode for /gsd-pr-branch

Adds planning.pr_strict — a boolean, default false, that selects what
/gsd-pr-branch means by "filtered". Default mode is unchanged: structural
planning state survives into the PR branch and the nine transient
subdirectories do not. Strict mode drops every .planning/ path, structural
files included, and carries a commit over only when it touches at least one
file outside .planning/.

Strict mode is what makes planning.commit_docs: true safe for a project that
versions its planning tree locally but publishes none of it. The alternative
posture, commit_docs: false, silently costs parallel executor isolation — a
worktree is checked out from a commit, so an untracked or ignored .planning/
is simply absent inside it and the executor has no PLAN.md to read. That claim
is now established by an executed fixture rather than inherited.

The two path lists are declared once and both projections derived from them,
so create_pr_branch and verify can no longer disagree about what the filter
promised. verify previously counted every .planning/ path against a documented
success criterion of zero while create_pr_branch was specified to preserve five
structural files, so a correct run reported itself as failed on every phase
that touched STATE.md — which is every phase. It now asserts against the active
mode, and names the .planning/ paths default mode deliberately keeps rather
than trading a wrong signal for silence.

Two verified defects in the same recipe are fixed alongside, because strict
mode would have amplified both. `git rm -r --cached` staged a deletion for any
.planning/ path the target branch already tracked, so the generated PR removed
the base branch's planning files — under strict mode that would have been the
entire tree. The same command left the picked file untracked on disk, so a
second commit touching that path aborted the cherry-pick with "untracked
working tree files would be overwritten" and every remaining commit was
silently dropped. Both were reproduced against real git before being fixed.
The filter now forces excluded paths back to what the PR branch's HEAD carries,
in the index and the working tree; a conflict outside the filter halts instead
of being improvised past; a commit left empty by filtering is skipped rather
than failing. A clean-working-tree precondition makes the worktree half safe.

Closes #2971

* fix(#2971): unwind the checkout on a conflict halt, and test the real recipe

Two review findings, both fixed in place.

The isolated adversarial pass found that the conflict-outside-the-filter branch
exited while leaving the user checked out on the half-built PR branch with
cherry-pick state still live — this loop runs in the user's own working
directory, so stranding them there is a real cost even though it is not a
vulnerability. The branch now aborts the pick, returns to the original branch,
removes the partial PR branch, and says so before exiting.

The standards pass found the L2 fixtures executed a hand-written mirror of the
cherry-pick filter recipe rather than the recipe itself, so a reordering in the
workflow would not have been caught — and the order is load-bearing, since
restoring a path from HEAD before removing it inverts the filter. The helper now
extracts the canonical loop from the shipped workflow and the fixtures execute
that verbatim, which also gives the conflict-halt unwind above real coverage.
The drift guard additionally pins the two commands' relative order and asserts
the workflow carries exactly one canonical loop.

Also records the publication gate in the CONTEXT.md glossary next to the commit
gate it is distinct from.

Refs #2971

* fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form

The remote matrix caught two defects in the previous commit.

The halt path claimed to restore the original branch but did not. `git
cherry-pick --abort` does not apply to a single `--no-commit` pick with no
sequencer file, and the fallback left the unmerged index in place, which makes
`git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so
the user was told they had been restored while still sitting on the half-built
PR branch. The unwind now drops sequencer state, hard-resets the disposable PR
branch to clear the unmerged index, and only claims a restore when the checkout
actually succeeded; when it does not, it says where the user is and gives them
the two commands to finish it by hand. Verified against real git: exit 1, the
conflict named, HEAD back on the original branch, the partial branch gone, a
clean tree and no CHERRY_PICK_HEAD.

Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form,
which names a command no runtime registers. The canonical authoring token for
workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the
documentation added in this branch is unaffected. The comment in src/config.cts
moves to the colon form too, since it propagates into the generated lib.

Refs #2971

* docs(#2971): backfill PR number into the changeset fragments (#3720)

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:40 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8df5cb36c2 enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)
* test(#2951): pin the absent-evidence provenance contract (failing first)

17 tests / 22 anchors on the deployed agent text. Measured against the parent
commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity
guards on the package-name and in-repo-value rules -- green before and after is
their intended signature.

Refs #2951

* enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata

A claim of the form "X does not support Y" drawn from MISSING metadata -- no
python_requires, no engines field, no per-version classifier, no changelog entry,
no matching support-matrix row -- no longer earns [VERIFIED] however
authoritative the source consulted. An absence is silence about every value, so
the same evidence would "prove" both the version being ruled out and the version
being standardized on. The only route from an absence to [VERIFIED] is a positive
falsification attempt with its failing output pasted; everything short of that is
[ASSUMED], which the file already routes to "needs user confirmation before
becoming a locked decision".

Third member of the family beside the package-name and in-repo-value provenance
rules, mirroring PR #2768's shape. A present declared constraint and an
affirmatively documented incompatibility are untouched.

Closes #2951

* fix(#2951): close the allow-list ambiguity and the mutation gap review found

Findings from the isolated adversarial pass and the two-axis review, all fixed:

MAJOR (x2, one root cause) -- the absence clause and the present-constraint
carve-out gave opposite verdicts on the same evidence for the commonest real
case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A
researcher could read the enumerated list as a "declared" positive constraint
and re-earn [VERIFIED], which is also the evasion vector. The rule now states
the decision procedure -- does the declaration bound EVERY value or only the
ones it names -- and closes the positive-reframing restatement explicitly. New
contract test pins all four clauses.

MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]'
asserted two independent substrings and never that the route lands on
[VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the
rule and survived all 17 tests. Now pinned as one joined sentence.

MINOR -- the attributable-failure test regex-matched illustrative examples
("a missing certificate, a wrong host"), so a copy-edit would break it for no
reason; relaxed to the substantive clause. The no-paraphrase guard counted only
the heading, missing the drift mode in its own name; it now also pins the core
proposition to one occurrence, and the test name matches what it checks. An
off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed.

MINOR -- docs/AGENTS.md listed four of the five governed absence forms while
the agent prose, docs/COMMANDS.md and the changeset listed five; three copies
disagreeing on list membership is the drift this repo treats as a defect.

SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its
docs/README.md index line. Both reviewers flagged them as a seventh and eighth
surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill
change" row requires only docs/AGENTS.md. The actionable four-case guidance is
retained in docs/COMMANDS.md, which is inside the approved scope.

Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550
bytes headroom under the LARGE cap of 49152.

Refs #2951

* docs(#2951): restore the how-to the phase gate requires

Reverses the removal in 6404b43d3. Both /code-review axes had flagged
docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md
index line as a seventh and eighth surface beyond the six the requester capped,
and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names
only docs/AGENTS.md, so they were dropped.

gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded
enablement sequence has more than one step and the how-to quadrant is empty.
The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED]
claim, probe or cite or accept it unlocked, then answer discuss-phase's
checkpoint -- and the last step lands on a different capability's surface, so a
reference table cannot carry it. Compressing the sequence to one step to unlock
howToSkipReason would be gaming the gate, which is the same Goodhart failure
this whole change exists to close.

A machine-enforced repo gate outranks two reviewers' scope preference and my own
reading, so the page is restored and the PR body discloses the two extra
surfaces instead of hiding them. Reverting is a one-file change if a maintainer
prefers the tighter scope.

Refs #2951

* chore(#2951): backfill the changeset PR number

pr: 0 -> 3718 now that the real PR exists. The placeholder fails
scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks
lint-docs-required from consuming the fragment.

Refs #2951

---------

Co-authored-by: sim <sim@local>
2026-08-20 14:35:14 -04:00
Tom Boucher
8da2dd3ad2 feat(#2790): add read-only planning.inspect schema-v1 snapshot query (#3708)
* feat(#2790): add read-only planning.inspect schema-v1 snapshot query

Adds a read-only query emitting a schema-versioned JSON projection of .planning/
so downstream harness UIs can consume planning state without parsing GSD's
Markdown a second time.

Composed strictly from the ADR-3180 section 7 owners plus parsePlanDocument,
parseRequirements and parseUatItems; markdown structure is read through the
Markdown Sectionizer and Markdown Table Model seams. It declares its own flat
external schema rather than serializing PlanningSnapshot, which is the
diagnostic-rule subject and still growing.

Extracts plan-document parsing out of cmdPhasePlanIndex into a shared leaf
module so phase.plan-index and planning.inspect cannot drift, including the
plan-id derivation both surfaces report.

Also fixes parseRequirements dropping the separator delimiter used by the
shipped requirements template, surfaced while wiring the requirement rows.

* fix(#2790): close spec gaps and a raw-text test assertion found in review

Review findings from the standards, spec and security passes:

- phases[] rows carry goal and dependencies, the two per-phase elements the
  issue Summary names that had no corresponding field. Goal is bounded to the
  section's leading prose so the Depends-on line, the Plans checklist and the
  wave annotations are not duplicated into it.
- requirement rows carry their own diagnostic codes, so a consumer no longer
  has to string-parse the global diagnostics subject to correlate.
- roadmap_acceptance.checkbox is looked up through the phase-id key owners.
  It was compared raw against the on-disk directory name, so it read null for
  every real-world slugged phase directory and the evidence channel was inert.
- the hostile-input test asserts the structured payload instead of matching the
  raw stdout string. The absence proof over raw stdout is kept deliberately.

* fix(#2790): register planning in the runtime usage list and repair fixtures

Remote runner reported 9 failures on 9b3f9aa. Two root causes, both fixed:

- gsd-tools.cjs registered the planning family in HOST_COMMAND_ROUTERS but
  never added it to TOP_LEVEL_USAGE's Commands list. Those are two surfaces a
  parity test guards, and the top-of-file block comment is not the runtime
  help string. A real wiring gap that every local gate and three review passes
  missed.

- the new suite's fixtures could not produce a resolvable phase set. STATE.md
  frontmatter omitted the milestone field, which ADR-3180 7.2 rule 1 makes the
  primary milestone selector, so the phase set scoped unscoped and every
  percentage was correctly withheld. Separately declarePhase returned a path
  without creating the directory, so a phase declared but never written to left
  phases empty. Both reproduced against the built module before fixing.

No assertion was weakened. The withholding path is still exercised and still
returns null when the roadmap is absent.

* chore(#2790): backfill changeset pr number

* test(#2790): cover every enumerated matrix row and contain a symlink escape

Reverses a silent deferral. An earlier revision left 23 of the 78 enumerated
matrix rows unimplemented and 7 more as one-off manual checks, with a paragraph
in the artifact and the PR body describing the gap. CLAUDE.md is explicit that
such a note is not a fix and is not surfacing. The rows are implemented instead
and the manual-evidence bucket is gone: 49 test cases become 88, covering all 78.

Writing the symlink row proved a real leak: a *-PLAN.md symlinked outside
.planning/ had its content emitted into the payload, confirmed via a direct call
and the spawned CLI. readDocument now resolves target and planning root with
realpathSync and rejects an escape, returning the ordinary unreadable-document
shape. Tested both ways, because a containment check that over-rejects is its own
defect: an escaping symlink leaks nothing and degrades that plan alone, while a
legitimately relocated .planning/ symlink stays fully readable.

The three new modules are registered in the mutation COVERED registry, which had
been reporting has_work false and skipping the Stryker gate entirely. Provisional
non-binding floors so the shards run and report; raised to the measured value
before merge, since the registry forbids calibrating from a local run.

* fix(#2790): satisfy the mutation ratchet contract and scope the 1MB test

Remote runner reported 16 failures on 8c451ed. Two causes.

The COVERED registry has a paired contract the earlier commit violated: every
module needs a matching RATCHET_BASELINE entry, and minScore must be between 50
and 100 with minScore === baseline. The provisional floor of 1 was illegal on
both counts. All three modules now sit at 50 — the registry's own enforced
minimum — with matching baselines. The score cannot be measured locally: the
shard runs node --test, which this repo hard-blocks, so CI is the only source.
Floors are raised to the measured value once this PR's shards report; a shard
below 50 means the tests need strengthening, since the floor cannot go lower.

The 1MB test was measuring the test harness rather than the product. The command
handles the oversized payload correctly by spilling to a tmpfile and resolving it
back, but the resolved stdout then exceeds runGsdTools' maxBuffer and the helper
reports ENOBUFS. It now uses --pick so stdout stays one byte while the full 1MB
document is still read and parsed end to end.

* fix(#2790): wire containment across every document read this command drives

An isolated security review of the containment control found the boundary logic
sound but not comprehensively wired: two content reads reached the filesystem
without it.

An escaped phase DIRECTORY could enumerate external filenames into the file
fields and diagnostic subjects. Both enumeration sites now containment-check the
directory before reading. Worth recording that the leak was already prevented one
layer earlier than the review claimed: Dirent#isDirectory() reports false for a
directory symlink, so such a directory never becomes a phase row at all. The
guard is defense-in-depth for a direct caller and for platforms where a reparse
point reports as a directory.

A *-VERIFICATION.md symlinked outside the root leaked one frontmatter value
verbatim, because readVerificationStatus does its own read and copies an
unrecognized status into the payload's next_action. Closed from the consumer
side through that function's existing fs injection seam, so src/verification.cts
keeps its signature and its other callers are untouched.

The reviewer additionally rated a forged status: passed as an integrity bypass.
It is not: anyone able to plant the symlink can plant a real VERIFICATION.md
saying the same thing. The incremental risk is confidentiality, which is what
these fixes close.

src/plan-scan.cts is deliberately unchanged: isPlanSuperseded reads
symlink-followed content but yields only a derived boolean, no document text.

* test(#2790): give the mutation shards an in-process surface

Two Stryker shards were CANCELLED at the 15-minute cap, not failed on score.
CI log: 640 mutants instrumented, and the dry run reported 'Ran 1 tests in 20
seconds' because the shards pointed at the integration suite, where nearly every
case spawns a gsd-tools subprocess and Stryker's command runner treats the whole
test-runner invocation as a single test. 640 x 20s cannot finish in 15 minutes;
at the kill it was 27/640 with an ETA over an hour.

Every other COVERED module points at a property or unit file, and the workflow's
own paths filter lists exactly those two patterns. In-process is the intended
mutation surface; the shards were pointed at the wrong shape of test.

Adds tests/planning-inspect.unit.test.cjs — 39 cases in 10 describes that spawn
nothing and call the built modules directly. plan-document and the router need no
filesystem at all, one being a pure content-to-object parser and the other taking
an injected mock. The three shards now point here. The 91-case integration suite
is untouched and still runs in the normal test job.

* chore(#2790): ratchet mutation floors to the measured CI scores

CI run 32392791843 measured all three shards, which is the only source the
registry accepts — local runs count timeouts as kills and inflate badly.

  planning-command-router  95.65 -> floor 94
  plan-document            76.58 -> floor 75
  planning-inspect         57.03 -> floor 56

Applied the registry's own rule, floor(score) - 1, and updated RATCHET_BASELINE
to match, since the ratchet test enforces equality.

planning-inspect sits well below the file's target of 80 and is the obvious
ratchet candidate as its tests improve. planning-command-router already exceeds
the target. The placeholder comment about floors pending measurement is removed
rather than left standing as a false statement.

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:42:43 -04:00
Tom Boucher
77fa08f1e8 fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input

Locks the Step 5.5 contract that a response_language project must feed the
edge probe an English translation of each requirement's text, and binds that
advice to measured engine behavior: the same requirement classifies to zero
shapes in Portuguese and to collection/adjacency/empty/ordering in English.

Also pins the honest limit — the issue's own repro sentence classifies to []
in English too, so translation is necessary but not sufficient and the
authored shapes override is the documented fallback.

Red before the doc change; the assertions are all false today.

Refs #2773

* fix(#2773): feed the spec-phase edge probe English-translated requirement text

The shape cues in src/edge-probe.cts are English word-boundary regexes, so a
project running with response_language set wrote its SPEC requirements into the
Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero
shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy
contributed nothing and --auto left it all unresolved — the probe was a silent
no-op for exactly the spec type it exists to harden.

Step 5.5 now states that the $REQS_JSON payload is engine input rather than
user-facing output, so the response_language rule does not govern it: each
requirement's text carries a faithful English translation, the SPEC keeps its
original language, and requirement ids are never translated or renumbered. The
instruction sits before the heredoc on purpose — the downstream APPLICABLE=0
warning fires only when every requirement is unclassified, so a partly-classified
non-English spec would otherwise slip through with no signal at all.

Measured against the compiled engine: the same requirement returns [] in
Portuguese and collection -> adjacency/empty/ordering in English. Also measured:
the issue's own repro sentence returns [] in English too, so translation is
necessary but not sufficient — the instruction therefore points at the authored
shapes override for prose carrying no cue in any language rather than promising
that translation restores classification.

Doc scope only, per the triage disposition on the issue. The compiled engine is
untouched; the lang-hint / per-language cue-set fix is a separate follow-up.

Closes #2773

* fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path

Surfaced by the isolated security review of this branch. Between the mktemp and
the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about
their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before
exiting, while the empty/placeholder guard directly above it exited without one.
A spec run that tripped the placeholder check therefore stranded a temp file
holding the SPEC's requirement text in TMPDIR, once per failed run.

The added contract test walks the region between the mktemp and the
unconditional cleanup and asserts no exit path leaves the file behind, so the
two guards can no longer drift apart. Proven to bind: run against the pre-fix
file the walker reports the leaking exit; against the fixed file it reports none.

Refs #2773

* docs(#2773): record the edge probe's English-cue input constraint in the predicate store

The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and
docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real
gaps rather than incidental coupling.

CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for
classifyShape but did not record that SHAPE_CUES are English word-boundary
patterns — so the predicate store implied text was language-agnostic, which is
what a future agent reads before touching this seam.

docs/CONFIGURATION.md's response_language row is what a non-English project
reads when it turns the setting on; it now names the one deliberate exception
and links to the FEATURES.md explanation, so the interaction is discoverable
from the config key rather than only from the workflow.

CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack
fragment is updated for the final byte range and now also records the
placeholder-guard cleanup fix folded into the same block.

Refs #2773

* fix(#2773): append the growth rationale to the existing spec-phase.md ack entry

The remote runner caught this: emitted-attribution.test.cjs pins the
0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration
regression test asserts the exact '31987 -> 31997' delta text survives), so
removing it to avoid a duplicate-key collision with a new fragment broke that
test instead of satisfying the ratchet.

The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102
were each appended to the same reason string by later PRs, which is how a shared
growth key coexists with the rule that two ack sources may never name the same
path. This appends the #2773 rationale the same way and drops the separate
fragment, whose spec-phase.md key was the collision.

Verified locally by reproducing both affected tests against the real fragment
before re-dispatching: the pinned delta survives, grown[0].acked is true,
staleAcks is empty, and all 35 entries still read as spent.

Refs #2773

* docs(#2773): add a how-to for probing edges in a non-English project

The phase gate's enablementSequence check caught a wrong call of mine. I had
recorded that no how-to was owed because the user takes zero extra steps — the
workflow translates the probe input itself. Written out, though, the sequence
from off to value is two steps and step 1 depends on response_language, a
setting owned by a different capability than the edge probe, which is exactly
the condition the how-to test names.

There is also real task content a reference table cannot carry: the three-way
split between a few unclassified rows (the classifier's recall gap), every row
unclassified (the probe could not read the spec at all), and the silent
partly-classified case where the APPLICABLE=0 warning never fires. That last
one is what a user would otherwise misread as a clean bill of health.

Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard
siblings and indexed from docs/README.md next to its closest relative.

Refs #2773

* chore(#2773): backfill the changeset PR number

pr:0 placeholder replaced with the real PR number now that #3713 exists.

Refs #2773

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:36:00 -04:00
Tom Boucher
adb46cdd85 feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker

Binds the contract before any hook change exists: a `state ~N commits back`
segment gated on the state_head stamp landed by #2622, firing at the same
advisory threshold /gsd-health's W024 uses rather than at > 0.

Covers all five acceptance criteria — threshold parity (19/20/21 boundaries),
both renderers including formatGsdStateCompact, an exact spawn-count assertion,
repo-pinning and sub_repos degradation, and behavioral parity against
readStateHeadFreshness rather than a source-grep of the two fence copies.

52 example-based tests plus 5 seeded fast-check properties. Red now by design.

* feat(#2734): surface STATE.md commit-age on the statusline

Adds an opt-in `state ~N commits back` marker to the GSD-state segment,
consuming the `state_head` stamp and freshness contract landed by #2622.
A solo developer returning to a project reads "Phase 4, executing" in
STATE.md and acts on it, without noticing the codebase moved 40 commits
since that line was written. /gsd-health reports it as W024, but only if
you think to run it; the statusline is the surface you see without asking.

Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses,
not at > 0: with commit_docs:true the commit carrying a STATE.md sync
advances HEAD by one, so > 0 would alarm permanently on a fresh project.

Costs exactly one bounded git subprocess per render and none when
disabled. `rev-list --left-right --count` answers ancestry and distance
together, and repo pinning is a filesystem check mirroring
projectOwnsItsRepo rather than a --show-toplevel compare, which is
unreliable on macOS /private/var and Windows 8.3 paths.

Every unresolvable input degrades to the tri-state unknown -- the marker
is absent, never a "fresh" claim the project cannot substantiate: a
malformed stamp, a root that does not own its .git, a sub_repos
workspace, history rewound past the stamp, or git being unavailable.

Also collapses statusline config resolution onto one resolveStatuslineOptions()
seam. runStatusline() and renderStatusline() duplicated it byte-for-byte;
one copy is what keeps a newly-added key from reaching only one of them.

* test(#2734): route the e2e spawn through the process seam and fix fixture leaks

Review findings from the two orthogonal passes:

- `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched
  its stdout to test a pure function. It now calls resolveStatuslineOptions()
  directly — no subprocess, no text matching.
- `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate
  lives in runStatusline, which reads stdin), so it now spawns through
  tests/helpers/process-seam.cjs and proves the negative with a filesystem fact:
  the git shim appends to a marker file on every invocation, and the assertion is
  that the marker never appears. Stronger than asserting text is missing, and it
  drops the last stdout substring match in the block.
- Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead
  of a trailing cleanup(dir), which leaked the temp repo on assertion failure.
  derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it
  binds each directory at scheduling time rather than cleaning only the last.

Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong
expectation rather than finding a code defect: `percent` drives the progress bar
too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends
after it, which is what the test exists to prove.

CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well
as the new key; both are now enumerated.

* docs(#2734): backfill changeset PR number (#3700)

---------

Co-authored-by: sim <sim@local>
2026-08-20 00:35:01 -04:00
Tom Boucher
2fca0e17e4 enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides

Binds the not-yet-built code-review-depth module: segment-aware path-prefix
matching of a changed-file set against ordered {paths,depth} rules, resolution
order flag > strongest matching rule > global > standard, typed validation
errors, and the large-scope downgrade boundary. Also proves behaviorally that
workflow.code_review_depth_overrides is not yet a registered config key.

Refs #2554

* feat(#2554): resolve code review depth from path-scoped override rules

Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth}
rules matched against a review's changed-file set by segment-aware path-prefix
comparison. Resolution order is --depth= flag, then the strongest matching rule,
then workflow.code_review_depth, then standard; a matching rule replaces the
global rather than being max'd with it, so quick and standard rules stay
meaningful. Glob metacharacters are a hard configuration error rather than sugar
for a prefix, and malformed rules halt the review instead of degrading to
standard. The resolver is pure and reports its own provenance, so the workflow
can print the resolved depth and the rule that matched. The pre-existing
>50-file deep-to-standard downgrade moves into the module and now names the rule
it overrode.

The key is registered centrally rather than as a capability config slice: the
federated slice channel admits only boolean/string/number/enum, so an array
slice would be dropped as malformed.

Closes #2554

* test(#2554): correct depth-provenance assertions and pin out-of-repo paths

Two corrections to the failing-first suite. The source assertion for a
non-matching rule with no global configured expected 'config'; with no global
set the depth comes from the default, and a companion assertion tolerated
either value, so both passed against an implementation that derived provenance
from whether any rules existed rather than from where the depth came from.

The out-of-repo absolute-path case used a home-directory path that matched
neither implementation, so it never exercised the defect it named. It now pins
the discriminating cases: an absolute path outside the repo root must not match
a repo-relative rule, and one under the root must.

* docs(#2554): document path-scoped code review depth overrides

Reference rows for workflow.code_review_depth_overrides in the configuration,
features and commands references plus the locale copies that carry those tables,
and in the planning-config reference. Explanation of why escalation is
whole-review rather than per-file and why v1 is prefix-only. New how-to for
scoping review depth by path, carrying the configuration-error reason table and
the distinction between nothing to report and could not look. CONTEXT.md
glossary entry and the INVENTORY row for the new CLI module.

ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and
pt-BR FEATURES.md carry no code-review config table, so those files are
deliberately untouched.

* fix(#2554): make the depth-misconfiguration halt executable and reject control chars

Three review findings, all in this change.

The misconfiguration halt was prose rather than shell: the error-printing fence
was followed by an unconditional extraction fence, so an ok:false result threw
and left the depth empty instead of stopping the review. Prose is not a guard —
the two fences are now one block with a real conditional, and anything that is
not the literal string true fails closed.

An interior control character in a rule path survived validation and reached the
provenance string and the summary box; rule paths now reject control characters
via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is
unchanged. That in turn makes the field record safe to delimit, so the seven
node invocations that each re-parsed the same result to read one field collapse
to one.

Also corrects the glossary entry's illustrative paths, which the glossary-ref
check read as real repository references.

* fix(#2554): use the fast-check v4 string API and acknowledge workflow growth

Two failures from the remote matrix on d3111f45, both this branch's.

The property block built its segment arbitrary with fc.stringOf, removed in
fast-check v4. Because the arbitrary is constructed in the describe body, the
throw took out all four property tests rather than one — they had never
executed. Rewritten to fc.string({unit, ...}), the form this repo already uses
in emitted-attribution.test.cjs. Every other fast-check helper in the file was
audited against the installed module.

The emitted-attribution growth arm needed an acknowledgment for code-review.md,
which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is
spent — its ripple was absorbed when #3503 merged, and the base file is exactly
the 34435-byte baseline this growth is measured against — so it cannot clear
anything, while the ack lint hard-fails on a duplicate key across two sources.
Removed it in favor of the new fragment, which is exactly how #3503 itself
replaced the spent 3191 fragment.

* docs(#2554): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-19 22:45:35 -04:00
Tom Boucher
71e00d426e fix(#3639): dir-aware sentinel recognition for the disk-side guards (#3698)
* test(#3639): pin bracket sentinel recognition in disk-side guards

* fix(#3639): dir-aware sentinel recognition for the disk-side guards

* chore(#3639): add changeset

* fix(#3639): disclose the digit-continuation residual, join phases-clear, load-bearing over-suppression guard

* test(#3639): match the token form W007 reports

* chore(#3639): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 21:55:18 -04:00
Tom Boucher
66228a89cf fix(#3637): carry the full executor contract in the orchestrator-worktree spawn (#3694)
* test(#3637): pin the executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): carry the full executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): role-definition embed, embed-performance gates, drop stale ack

* chore(#3637): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 20:50:12 -04:00
Tom Boucher
7fc1561806 fix(#3611): decode entity-escaped ampersands and split shell segments quote-aware (#3693)
* test(#3611): pin entity-escaped ampersand chains in the negative-grep gate

* fix(#3611): decode entity-escaped ampersands before the negative-grep gate scans

* chore(#3611): add changeset

* test(#3611): pin entity chains in the 968 detector and quote-aware splits

* fix(#3611): quote-aware segment split + entity decode in both plan gates

* chore(#3611): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 19:53:45 -04:00
Tom Boucher
bad1f045b1 fix(#3610): hoist surviving top-level codex config keys to file scope on merge (#3690)
* test(#3610): pin top-level key hoisting above the codex managed block

* fix(#3610): hoist surviving top-level keys above the codex managed block

* chore(#3610): add changeset

* fix(#3610): hoist to file scope (before the first table header) with reviewer-driven coverage

* chore(#3610): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 18:38:13 -04:00
Tom Boucher
8526bd46f8 enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes

The guard matched only `resolve-execution ... --effort\s`, but the CLI also
accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written
with the equals form was a live invocation-override caller the guard passed
silently, along with `--effort` at end-of-input.

Lift the matcher to a shared predicate and assert it directly against every
shape the CLI accepts, plus the decoys it must not fire on (--effortless, a
bare --effort with no resolve-execution, item 6's --attempt caller, a call and
flag split across lines).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR

ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator
invocation override needed a caller in shipped orchestration. Per the
maintainer's ruling, take unblock path (b) for item 1 only -- record that the
override is an operator-facing CLI surface, deliberately not driven by shipped
orchestration, and ratify.

The ADR's own path (b) wording is not adopted verbatim: it says the scope is
limited to static install-time propagation, which is false on both counts --
item 6 has a live caller (#2296) and #2481 delivered a live invocation-time
argv channel. Only one precedence step is narrowed.

No consumer was invented to clear the gate: nobody has asked for a per-run
effort override, and #2475's actual complaint is already closed by the
cascade-to-argv path. The amendment states explicitly that --effort remains
supported and is not deprecated, so the scoping is not read as dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run`

There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run
and gsd-mcp-server. Both sites presented a command that cannot run as written.

One is in this branch's own new ADR-443 amendment; the other is a pre-existing
error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than
deferred. No translated copy carries either line, so no i18n drift is created.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2475): record the guard/CLI divergence risk on the item-1 matcher

The predicate independently models gsd-tools.cjs's argument parser rather than
sharing a constant with it, so a third --effort spelling would leave the guard
reporting green while ADR-443's ratifying invariant silently stopped holding.
Name that risk where the next editor will meet it.

Also restores the bounded-prose rationale that was attached to the eslint
directive removed in 39793079c -- the directive went unused once the regex moved
to a const, but the reasoning it carried is still worth having.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 17:56:04 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
Tom Boucher
79781e68eb enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands

Adds a deterministic resolvability probe over each PLAN.md <automated> verify
command and surfaces the nearest prior phase's proven commands to the planner
at every context window.

- src/verify-command-grounding.cts: recognizer (not a shell interpreter) that
  grounds a leading cd <literal> chain and npm --prefix <literal>, and reports
  unresolvable rather than guessing. Never executes command text.
- gsd-tools check verify-command-paths <N>: per-phase probe, wired into
  plan-phase.md before the plan-check pass.
- init.plan-phase gains prior_verify_commands, ungated by context_window.
- gsd-plan-checker: new Verify Command Path Resolvability dimension that
  reports the failing target and never prescribes a replacement.

Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs
(readdir order is not stable across platforms, so a module whose name extends
another's with a hyphen bucketed differently on Linux than on macOS).

Closes #2401

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets

Independent review found three defects in the recognizer:

- npm --prefix DIR run SCRIPT never reached the script-existence check,
  because the pattern required npm and run to be adjacent. That is the
  form the docs tell planners to prefer, so script_missing never fired
  for it. The prefix flag and its value are now stripped before matching.
- --prefix captured with \S+, so a quoted path containing a space was
  truncated to a stray opening quote and reported as a missing directory
  - a false blocker, worse than the bug this feature fixes. The capture
  is now quote-aware.
- A chained cd whose later segment was absolute concatenated instead of
  resetting, producing a nonsense path and another false blocker. The
  fold now resets on an absolute segment.

Also replaces the bespoke phase-directory regex with the canonical
phase-id helpers. Real phase directories are NN-slug, not phase-N-slug,
so the prior-command harvest matched nothing outside its own fixtures
and the planner-inheritance half of this feature was dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(#2401): source task blocks from the canonical sectionizer

The module carried its own copy of the <task>-block grammar - a fourth
hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps
its copy only because it needs the type= attribute the canonical helper
discards; this module never reads that attribute, so it can share the
owner outright instead of adding a test around a copy.

extractAutomatedCommands now takes task bodies from extractTaggedBlocks
and the out-of-task remainder from stripTaggedBlocks. A task-grammar
parity test pins the attributed task-name set against the canonical
helper across six awkward task shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): extract agent-file overflow to references and repair the property arbitrary

The remote matrix run came back red with 19 failures, four root causes:

- agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the
  49152 agent cap. Their bodies move to gsd-core/references/, leaving
  @-reference stubs, per the documented overflow pattern.
- The new checker dimension invoked gsd_run before the canonical
  preamble that defines it. The call is deleted outright: plan-phase.md
  already runs the probe and hands the result in as {VERIFY_PATHS}, so
  the dimension consumes that rather than re-running anything.
- fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with
  fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range.
- Three runtime-loaded files grew; acknowledged in the existing ack
  fragments that already own those bare filenames, since two ack sources
  may never name the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2401): regenerate golden install-tree fixtures for the new references

Adding two files under gsd-core/references/ changes what the installer
emits into every runtime's tree, so all 19 golden install-parity
fixtures went stale. Regenerated with npm run gen:install-tree; the
delta is exactly the two new reference paths per runtime, no removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2401): backfill changeset pr number to 3678

* fix(#2401): treat ~ as a home expansion only at the start of a path

Windows CI caught this on both shards; the Linux-only remote matrix
cannot see it. The dynamic-path refusal rejected ~ anywhere, and a
GitHub Windows runner's tmpdir is an 8.3 short name -
C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows
path came back unresolvable/dynamic_path.

This was a production bug, not a test artifact: any Windows user whose
project path carries an 8.3 short name, or any literal ~, silently lost
the probe entirely - every command degrading to unresolvable with no
explanation.

~ is a home expansion only at the start of a path; elsewhere it is an
ordinary literal. The check is now split: $, backtick, *, ? and newline
stay refused anywhere (substitution and globs, and the glob characters
are illegal in Windows path components regardless), while ~ is refused
only leading, tolerating one leading quote since the check runs before
quote stripping.

The prior tests only caught this on Windows because only Windows puts a
~ in tmpdir. Four new tests pin it on every platform via a fixture
directory literally named RUNNER~1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 15:21:15 -04:00
Tom Boucher
4e60dba717 fix(#3604): make glossary ref visibility independent of backtick parity (#3680)
* test(#3604): pin parity-dependent ref visibility in the glossary gate

* fix(#3604): make glossary ref visibility independent of backtick parity

* chore(#3604): regenerate CONTEXT-INDEX for corrected predicates

* chore(#3604): regenerate examples CONTEXT-INDEX for corrected predicates

* fix(#3604): complete retired-family exemptions and pin the guard rails

* chore(#3604): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 13:51:40 -04:00
Tom Boucher
7cf6a079fa fix(#3602): bind model resolution for every workflow subagent spawn (#3670)
* test(#3602): guard every spawned gsd-* subagent has a model resolution

* fix(#3602): bind model resolution for every workflow subagent spawn

* test(#3602): merge drift-ack entries into their owning fragments

* fix(#3602): address review findings - docs-update verifier binding, ack merge, guard residuals

* chore(#3602): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:35:45 -04:00
Tom Boucher
dae134b960 test(#2650): bound the stall-watch fan-out with the class norm that reddened windows shard 1/3 (#3671)
* test(#2650): failing-first coverage for the mis-sized stall-watch bound

Lands the regression matrix before the fix. Two assertions fail deterministically on this commit because the helper still calls raw spawnSync with a hard-coded 10000ms bound:

1. The process seam is never reached, so a mock.method spy on it records nothing and the class-norm bound cannot be observed.

2. A raw spawnSync result carries no outcome/timedOut field at all — which is precisely why the windows shard 1/3 failure printed 'null !== 0' instead of naming the timeout or the bound it exceeded.

The bound is sized for the wrong class of call. extractStallHelpersBash slices the entire bash fence, whose runtime-launcher preamble resolves gsd-tools.cjs and really runs two config-get lines — two Node spawns, measured 236ms against 2ms for the fallback the existing comment claims fires (118x). tests/helpers/timeouts.cjs already owns this class as HOOK_FANOUT_TIMEOUT_MS and its comment records the identical failure on PR #3285.

Refs #2650

* test(#2650): bound the stall-watch fan-out with the class norm and report it typed

Drives the failing-first coverage from 15aaef4a7 green (5 failures -> 0). runBashScript now routes through tests/helpers/process-seam.cjs instead of a hand-rolled spawnSync, which CONTRIBUTING.md already requires of anything that shells out.

The bound moves from a hard-coded 10000ms literal to HOOK_FANOUT_TIMEOUT_MS. That constant already exists for exactly this shape — a bash invocation that fans out to nested subprocesses — and its own comment records the identical failure on PR #3285: a bound sized for the wrong class, not a slow machine. This script is that shape: the extracted fence opens with the runtime-launcher preamble, which resolves gsd-tools.cjs and really runs two config-get lines.

The result is now typed. spawnSync reports a kill as status:null, so an exceeded bound reached the call sites as 'null !== 0' — naming neither the timeout nor the bound it exceeded. OUTCOME.TIMED_OUT names itself. status is still aliased from exitCode so the five existing assertions read unchanged.

Corrects the comment that caused the mis-sizing: it claimed gsd_run is undefined so the '|| echo' fallback fires. It is not — the preamble defines it, and the two Node spawns are real (236ms vs 2ms, 118x). They are deliberately left in place; removing them would change what the extracted script executes.

Also drops two now-dead '{ timeout: 10000 }' call-site options. The seam reads only its own documented keys, so those would have been silently ignored while still reading as a 10s bound.

Refs #2650

* test(#2650): make the boundary test a real value-domain boundary

Standards review flagged that the previous 'bound boundary' test repeated the TIMED_OUT arm the test above it already covers, and was not a limit-1/limit/limit+1 in any meaningful sense — an exact-millisecond timing edge would have been a race, not a boundary.

Replaced with a boundary on the VALUE DOMAIN of the bound itself: 0 and -1 must be rejected with TypeError, and 1 (the smallest positive value) must be accepted. Zero is the load-bearing case — spawnSync reads it as 'no timeout at all', which is exactly the unbounded-spawn hazard local/no-unbounded-spawn exists to prevent, so the seam rejects it rather than honouring it.

Also asserts the rejection path does not leak the script temp dir, since the throw escapes through runBashScript's finally. Soak: zero=TypeError, negative=TypeError, one=no-throw, leaked dirs=0.

Refs #2650

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:04:24 -04:00
Tom Boucher
8d1f770dfe test(#3395): pin the clock and scope the stale-prose scan that reddened windows shard 2/3 (#3669)
* test(#3395): failing-first coverage for the silently-ignored clock pin

Lands the regression matrix BEFORE the fix so the failure is proven rather than asserted. Three assertions fail deterministically on this commit:

1. PINNED_ENV does not actually pin. `_pinnedNowMs()` (src/clock.cts) returns null unless GSD_TEST_MODE is set, so GSD_NOW_MS alone is discarded and last_updated is stamped from the live wall clock. An instant ending ...:35.149Z contains the substring 35.1, which is what reddens the windows-latest shard 2/3 lane roughly 1 run in 600.

2. The colliding-instant regression cannot reach its instant, for the same reason.

3. The #3052 same-date test never lands on 2020-09-10, so it has been exercising the different-date path and passing for the wrong reason.

Also adds currentPositionBlock() plus boundary (ms 099/100/199/200, second 34/35/36, LF and CRLF) and two-arm fast-check coverage for the scoped read the fix will switch to.

Refs #3395

* test(#3395): pin the clock and scope the stale-prose scan to the body

Drives the failing-first coverage from a5a919ffb green. Two changes, both needed:

1. PINNED_ENV now sets GSD_TEST_MODE alongside GSD_NOW_MS. _pinnedNowMs() (src/clock.cts:44) returns null without it, so the pin was silently discarded and last_updated carried a live wall-clock instant. src/clock.cts is deliberately NOT changed: requiring both keys is what stops an ambient GSD_NOW_MS from freezing a production clock, so the caller was the side that was wrong.

2. The stale-prose assertion now reads currentPositionBlock(stateContent) instead of the whole document. Frontmatter is not phase prose, and an instant ending ...:35.149Z contains the substring 35.1 — which is exactly how a document with no stale prose in it produced 'the stale 35.1 phase prose must be refreshed away'.

Confirmed hypothesis: the two defects compose. The inert pin supplies a live timestamp; the whole-document scan turns it into a failure. Either alone is latent, which is why this sat unnoticed for five days and then reddened a lane the release never touched.

Also corrects two things the failing-first run exposed. The property test used fc.date() without noInvalidDate, so ~1 sample in 300 was an Invalid Date whose toISOString() threw (counterexample: new Date(NaN)); re-soaked at 5000 runs. And a precondition assertion added to the #3052 block was measured to pass with or without the pin, so it was removed rather than shipped as vacuous truth — last_activity there is body-derived, not clock-derived.

Refs #3395

* test(#3395): apply review findings — pin #3052, one fixture builder, CRLF coverage

Spec-axis review caught a real slip: the #3052 block carried a comment saying its pin was being added as hygiene, but the RED-state revert had removed GSD_TEST_MODE and the fix commit never restored it. A comment describing an action that was not taken is worse than either doing it or leaving it alone — the pin is now actually there.

Standards-axis review flagged the same frontmatter+heading fixture shape being rebuilt in three tests. Extracted one stateDoc({iso, lines, eol}) builder; eol is a parameter rather than a constant because the helper's CRLF behavior is a claim under test.

Self-review finding: CRLF was only exercised on a single-heading document, and the following-heading case only under LF — so the exact claim the helper's comment rests on (`\n## ` matches inside `\r\n## ` because the CR precedes the newline) was never actually run. The control test now loops both line endings WITH a following heading.

Also drops a comment that restated the PINNED_INSTANT rationale verbatim.

Refs #3395

---------

Co-authored-by: sim <sim@local>
2026-08-19 11:25:24 -04:00
Tom Boucher
249c586a40 fix(3582): stop the cold-tree guard from racing the builders it guards against (#3665)
The guard added in #3656 asserts that this test leaves the real repo hooks/ directory
alone. It compared a RAW listing before and after — and just failed on the runner:

    the real repo hooks/ directory listing must be unchanged by this test
    -   '.dist-staging-20858',
        'dist',
        ...

Nothing was wrong with the test's own behaviour. A CONCURRENT scripts/build-hooks.js —
nine test files invoke it from before() hooks — created hooks/.dist-staging-20858 inside
the comparison window. The assertion assumed the shared hooks/ directory is stable for the
duration of a test, which is precisely the assumption this line of work exists to
disprove. The race-detector raced.

The intent is right and is kept: this test must not add or remove anything in the repo.
Only the comparison changes — both snapshots are now filtered through shouldCopyHookEntry,
the same rule the fixture itself uses, so transient build scratch that is not this test's
doing and is excluded from the fixture anyway no longer registers as a difference.

Proven by execution: with a .dist-staging dir injected mid-window the filtered listings
compare equal, while the unfiltered listings provably differ by exactly that entry — so
the old comparison would have failed and the new one is immune rather than merely quieter.
The injected directory is removed afterwards and hooks/ is confirmed byte-identical.

Checked for the same shape elsewhere: this is the only raw listing comparison of the live
hooks/ directory in the file or the repo. The second title in the failure output is the
describe() wrapper around this same test, not a sibling.

Refs #3582

Co-authored-by: sim <sim@local>
2026-08-19 09:00:39 -04:00
Tom Boucher
1bf73d957b enhance(#2295): record the resolved model per reviewer in REVIEWS.md frontmatter (#3649)
* test(#2295): failing-first coverage for per-lane resolved-model recording

* feat(#2295): record the resolved model per reviewer lane

* docs(#2295): document the recorded reviewer model and its provenance

* fix(#2295): refuse control characters in a recorded model value

* test(#2295): correct watermark assertions for the widened mark shape

* fix(#2295): anchor the role-manipulation injection pattern at a word boundary

* feat(#2295): record the applied reasoning effort in the model value

* chore(#2295): backfill changeset pr number

* chore(#2295): restore em-dash in changeset body

---------

Co-authored-by: sim <sim@local>
2026-08-19 08:59:37 -04:00
Tom Boucher
4e70b245e8 fix(#3582): stop the cold-tree fixture racing concurrent hook builds (#3656)
* fix(3582): stop the cold-tree fixture racing concurrent hook builds

Two tests in tests/gsd-check-update-worker-platform-gate.test.cjs failed a verification run
with `ENOENT: no such file or directory, lstat '/work/hooks/.dist-staging-20836'`. This is
a race I introduced in #3582, not a flake, and it passed when #3582 merged because it only
fires when the timing lines up.

buildColdInstallTree() copied the LIVE repo hooks/ directory with a filter that excluded
only the basename 'dist'. scripts/build-hooks.js writes atomically through a per-PID
staging dir, hooks/.dist-staging-<pid>, and removes it when finished — and the archived
build-hooks-atomic-write changeset records that NINE test files invoke build-hooks.js from
their before() hooks. So several test processes create and delete staging directories
inside hooks/ while other tests are reading it. cpSync enumerated one, and the owning
process removed it before cpSync got to it.

The helper's own header already states the rule it needed: hooks/dist is excluded because
it "is not present in a raw marketplace checkout either". hooks/.dist-staging-* is
gitignored (.gitignore:21) and equally absent from a raw checkout — it was simply missed.

Fixed by enumerating hooks/ explicitly and skipping 'dist' and any '.dist-staging' prefix
BY NAME, before anything stats or copies the entry, then copying each surviving entry
individually. A name-first skip means a vanishing staging dir is never touched at all.

Worth recording because it corrects the assumption this fix was written under: cpSync's
filter IS invoked before the entry is lstat'd, and returning false leaves it untouched
(verified by deleting inside the callback and returning false — no throw). So merely adding
'.dist-staging' to the old filter would also have closed the race. The explicit enumeration
was kept anyway so correctness does not depend on that Node implementation detail.

Proven by execution both ways: with a staging dir planted in hooks/, the OLD
cpSync-with-filter form copied it straight through into the fixture, while the new form
succeeds and produces no .dist-staging entry with the real hook set intact.

Regression test added beside the existing cold-tree tests: it plants a real
hooks/.dist-staging-test-<random>, asserts the fixture builds clean without it, and removes
only the directory it created.

Repo swept for the same exposure: this helper is the only place doing a bulk enumeration of
the whole live hooks/ tree. The other hooks/-touching tests reference specific named files
or hooks/dist/ and are not exposed. scripts/build-hooks.js is deliberately untouched — its
per-PID staging is what makes its own writes atomic and is correct.

Refs #3582

* fix(3582): make the race regression test hermetic instead of mutating the live tree

The regression test added in the previous commit failed the runner with "failed running
after hook", and it was wrong in two ways — the second one worse than the first.

cleanup() (tests/helpers.cjs:452-487) deliberately THROWS for any path outside the known
temp roots. The test planted hooks/.dist-staging-test-<random> inside the repo and then
asked cleanup() to remove it, so the after-hook threw. That guard is correct and is left
alone.

The real problem is that the test mutated the LIVE hooks/ directory while other test files
concurrently read it — the exact shared-state hazard this change exists to remove. A
regression test for a race must not introduce one.

buildColdInstallTree now takes an optional opts.repoRoot (defaulting to the real REPO_ROOT
and used for both copies it performs), so the test builds a fake repo root under the temp
dir, plants representative hooks plus dist/ and .dist-staging-99999/ THERE, and asserts the
fixture excludes both. All six pre-existing callers pass no arguments and are unaffected.
The test also asserts the real hooks/ listing is identical before and after, so a future
edit that reintroduces live-tree mutation fails loudly.

The name rule is now pinned directly rather than only through the copy. shouldCopyHookEntry
is exported and asserted, including the two cases a sloppier implementation would get
wrong: 'dist-staging-no-dot' and 'distant.js' must both be KEPT. Anything matching on a
loose 'dist' substring or startsWith passes every other case and fails those two.

Also corrected the issue number on the tests introduced here: they were labelled #3631,
which is the unrelated capability-consent bytecode work. This is #3582.

Verified by execution: the predicate rule holds on all nine cases; a fake-root fixture
yields exactly the representative hooks with dist and .dist-staging excluded; the no-arg
default still copies the real tree (29 entries); and the real hooks/ listing is byte-identical
before and after.

Refs #3582

* chore(3582): re-trigger CI after an orphaned Validate Branch Name run

The Validate Branch Name run for this branch (32211622051) sat queued from 03:17 and was
never picked up — updatedAt never advanced past createdAt while the same workflow completed
normally for other branches. `gh run rerun` refused it ("already running") and
`gh run cancel` returned HTTP 500, so the run is orphaned on the GitHub side.

Closing and reopening the PR re-fired the other pull_request workflows but not that one,
whose triggers evidently do not include reopened. An empty commit is the remaining way to
get a fresh run.

No file changes: the tree is identical to dc71534b6, whose remote-runner pass carries
forward unchanged.

Recording this rather than admin-merging past the pending check. Everything else was green
(24 pass, 0 fail), but admin merge is sanctioned only for the missing-secondary-reviewer
case, never to skip a gate that has not actually run.

Refs #3582

---------

Co-authored-by: sim <sim@local>
2026-08-19 01:33:40 -04:00
Tom Boucher
9e4f0e99ad fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash

bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.

Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.

The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.

Refs #3631

* fix(3631): exclude derived bytecode caches from the consent digest

RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.

collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.

Three properties were preserved deliberately, each pinned by a test:

  - The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
    first would have turned the exclusion into a way to smuggle a symlink past the check;
    a symlink named x.pyc still throws.
  - Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
    caps guard the WALK; the digest answers a different question, and exclusion must not
    become an unbounded-bytes hole.
  - An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
    directory marker was the sharper half of this bug: an empty __pycache__ flipped the
    hash before any .pyc existed, so a *.pyc-only filter would have left it live.

The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.

Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.

Fixes #3631

* fix(3631): narrow the digest exclusion after two isolated security reviews

The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.

HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.

FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.

Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.

Narrowed to what is actually defensible:
  - a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
    the walk still recurses and hashes every non-excluded child.
  - .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
  - a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
  - declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
    basename are now rejected in both validator copies — a file named .pyc can contain
    perfectly valid JavaScript, so the exclusion must not be reachable from a declared
    surface.

Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.

Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.

Refs #3631

* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind

Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.

HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.

The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.

The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.

Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.

Changeset rewritten: it still described the rejected wholesale-exclusion semantics.

Refs #3631

* chore(3631): backfill changeset PR number (#3650)

---------

Co-authored-by: sim <sim@local>
2026-08-18 22:35:54 -04:00
Tom Boucher
02a36d3db9 fix(3618): update the Windows fallow assertions to the behavior #3618 chose (#3654)
full test (windows-latest, 24, shard 1/3) has been RED on next since ac1b6d679 and blocks
every PR. Three assertions fail, Windows-only, and all three are stale TESTS rather than
resolver defects.

Two of them create a BARE extensionless 'fallow' and assert it resolves. #3618 deliberately
stopped resolving that on Windows, and said so in its own commit message: the extensionless
file is npm's POSIX sh shim, which CreateProcess cannot run (#3275). The code did what it
meant to; the tests were never updated to match.

Both now assert BOTH platform contracts rather than skipping a lane — the precedent #3618
set when it fixed its own F4 ('a t.skip on one lane would have been green and would have
left the win32 carve-out unpinned'). POSIX keeps the original bare+chmod fixture verbatim.
win32 creates the artifact npm actually writes there, fallow.cmd, and asserts resolution
finds it; each also pins that a bare extensionless fallow ALONE still resolves to null on
win32, so the carve-out is pinned rather than merely stepped around. The precedence test
keeps testing precedence: on win32 both node_modules/.bin and the PATH dir get a
fallow.cmd, and .bin must still win.

The third failure was case, not logic: the test wrote fallow.cmd and asserted exact string
equality, but candidates are built by appending PATHEXT entries, and PATHEXT is uppercase.
Confirmed at the source — DEFAULT_PATHEXT = '.EXE;.CMD;.BAT;.COM'
(src/shell-command-projection.cts:651), read verbatim and never lowercased, with the
candidate returned as-is. So the resolver returned fallow.CMD, which is the same file on
case-insensitive NTFS. Fixed the assertion to compare case-insensitively on win32; the
resolver is untouched, because its return value has to stay usable verbatim.

Not verifiable here: this repo's remote runner matrix is Linux-only and the host is macOS,
so the win32 branches are reasoned from the resolver source and proven only by GitHub CI.
The POSIX lanes were verified by execution (bare+chmod in .bin resolves, .bin beats PATH,
non-executable in PATH yields null).

No changeset: changeset-lint's USER_FACING_PREFIXES does not include tests/, so this is
OK_NO_USER_FACING_CHANGES rather than a missing fragment.

Refs #3618

Co-authored-by: sim <sim@local>
2026-08-18 22:08:52 -04:00
Tom Boucher
2972da4c9d enhance(#3619): ratchet the platform seam with local/no-private-binary-resolution (epic #3411 Phase 3) (#3636)
* chore(#3619): ratchet the platform seam with local/no-private-binary-resolution

Epic #3411 Phase 3, the ratchet. Scope revised with maintainer approval and
recorded on the issue: the epic's literal ask was a rule rejecting a bare-name
spawn outside the seam. Surveyed at ac1b6d679, ~30 such sites exist and none is
a defect — git, gh and npm ship native .exe that CreateProcess resolves unaided,
and the rest are POSIX-only tools. ADR-1703 rules 2 and 3 forbid grandfathering
and escape hatches, so a literal rule would be unsuppressable and would force
rewriting 30 correct calls.

The epic's actual thesis was four private RESOLVERS, not four bare spawns. So
the rule flags re-implementing resolution: reading PATHEXT in any casing from
any object, and a hardcoded list carrying two or more of .exe/.cmd/.bat/.com —
precisely the shapes fallow-runner's candidateNames and gsd-tools' PATHEXT
string had before Phases 1 and 2 deleted them.

Three boundaries were arrived at rather than assumed:

  two-or-more   a single .endsWith('.cmd') is a classification, not a candidate
                set; runtime-hooks-surface derives .cmd shim paths that way
  boundary-aware  a naive substring test flags .execute and .compacting, caught
                on src/host-integration.cts before it could become a false
                positive nobody could suppress
  suffix-anchored  the seam exemption matches src/shell-command-projection.cts
                exactly; a substring match would also exempt the dispatch test
                file. Case I9 pins it.

PATH scans are deliberately NOT flagged — membership checks (bin/install.js)
are indistinguishable from resolution scans, and an unsound rule in a
zero-escape-hatch architecture is worse than no rule.

To make the ratchet strict with no carve-out, resolveExecutableBinary gained
pathOverride: search THIS PATH, read everything else including PATHEXT from the
ambient environment. resolveFallowBinary now supplies its own search path
without hand-threading PATHEXT, which would itself have been a private read.
The three alternatives were all worse: exempting the file is grandfathering,
exempting the AST shape is a carve-out every future caller must replicate, and
dropping the pass-through would silently ignore a user's real PATHEXT — buying
a lint rule with a correctness regression.

eslint-rules/** is outside the rule's globs rather than exempted, because
portability-vocab.cjs owns the extension set. scripts/**/*.cjs got its own block
so that exclusion does not leave a hole in the ratchet.

Started green with nothing suppressed. Proven able to fail: a fixture with both
signals reports two errors.

Refs #3411

* fix(#3619): close the PATHEXT destructuring evasion and correct two overclaims

Adversarial review found a trivial evasion of the rule's primary signal: the
visitor only handled MemberExpression, so

  const { PATHEXT } = process.env
  const { PATHEXT: exts } = process.env
  const { Pathext } = opts.env

were all unflagged. That is a common idiom, not an exotic bypass. An ObjectPattern
visitor now catches it in every form — renamed, any casing, any receiver, string
keys — while leaving a computed key alone, since it is not statically decidable.
I10-I13 pin the invalid forms and V9/V10 pin PATH and the computed key.

Two overclaims corrected, both mine:

Standards review proved the docs were factually wrong. Both the ADR amendment and
the CONTEXT.md entry asserted that tests/shell-command-projection-dispatch.test.cjs
is still linted by this rule. It is not — the rule's surface is src, gsd-core/bin,
scripts and hooks, and tests/** is deliberately outside it because test setup
legitimately assigns process.env.PATHEXT (fallow-runner's P3 does exactly that).
The suffix-vs-substring distinction is therefore proven by RuleTester case I9
feeding a synthetic filename, NOT by real coverage of that file. Both documents now
say so.

The rule's own docstring claimed the seam exemption matches the seam path
'exactly'. It is a suffix match, so a nested foo/src/shell-command-projection.cts
would also be exempt. Suffix matching is kept — it is how sibling rules resolve
paths and the nested case does not exist — but the docstring now states the
boundary rather than overstating the precision.

The evasion fix was verified by executing eslint against both destructuring forms
in scripts/, not by inspection. Probe: 31/31.

Refs #3411

* chore(#3619): backfill changeset pr number 3636

---------

Co-authored-by: sim <sim@local>
2026-08-18 16:40:17 -04:00
Tom Boucher
0f417aa6d0 fix(#3584): the verb owns the count token and nothing else (#3635)
* test(3584): failing-first coverage for Plans-line trailing text

roadmap update-plan-progress preserves trailing text only when the line begins with a
canonical count token. Every other phrasing — including the TBD value the shipped
template itself suggests — is replaced to end-of-line, and a sentence wrapping onto a
second line has its first line deleted, leaving the continuation standing alone so the
roadmap asserts something nobody wrote. Exit 0, updated:true, and the diff reads as a
routine count bump.

These tests fail on that, and pin the arms that must keep working: the template
placeholder is still replaced, the #2853 token-plus-annotation path is unchanged, CRLF is
neither stranded nor duplicated, and a run that leaves the line alone still updates the
Progress table and checkboxes rather than becoming a no-op.

* fix(3584): the verb owns the count token and nothing else

RED proven at 105bdf7c: 7 failures — the preserving cases (freeform prose, wrapped
continuation, TBD, CRLF) failed while the template-placeholder and #2853 token arms
passed on base.

The trailing-text guard fired only when the line began with a canonical count token:
 dropped the rest of the line whenever the
regex's count group did not match. #2853 fixed end-of-line truncation on that one path
only. The in-code comment justified the rest as 'the fresh-template bracketed placeholder
or other freeform guidance, not user prose' — a heuristic that misreads ordinary human
phrasing and destroys even TBD, the value the shipped template itself suggests at
templates/roadmap.md:37.

The sharper failure was the wrapped sentence: only the first line is inside the match, so
the verb deleted line one and left line two standing alone, leaving the roadmap asserting
something nobody wrote — at exit 0, updated:true, in a diff that reads as a routine count
bump.

Inverted the default into three arms. A real count token is rewritten with its annotation
preserved (unchanged, #2853). A bracketed placeholder is detected POSITIVELY and replaced.
Everything else — freeform prose, TBD, a wrapped sentence's first line, an empty value —
returns the match untouched. That last arm resolves the wrapped case by construction: an
untouched first line cannot orphan its continuation.

Positive detection is the load-bearing part. Implemented as 'not a count token, therefore
disposable', rows 1-3 come straight back; the detector instead asks whether the value IS a
bracketed placeholder.

Leaving the line alone does not make the verb a no-op: the phase checkbox, the
Progress-table cells and the plan-checklist row still update in the same run, and that is
asserted. CRLF is unaffected in every arm — the pattern's [^\r\n]* never consumes the
\r, so it sits outside the match regardless of which arm runs.

Fixes #3584

* fix(3584): detect the template placeholder by its text, not by its brackets

Two defects in the arm-2 detector shipped in fc23e49e, both found in review.

Finding A: isBracketedPlaceholder asked only whether the trimmed value was wrapped
in [...]. Brackets are ordinary prose punctuation in a roadmap, so any hand-written
bracketed note — '[Deferred pending re-scope]', '[blocked on #1234]' — was classified
as the fresh-template placeholder and destroyed. That is the very defect #3584 is
about, reintroduced one arm over. The detector now matches the placeholder's TEXT
(/^\[\s*Number of plans\b[\s\S]*\]$/i), so it recognizes the shipped template
value and its short form and nothing else.

Finding B: the count group matched '\\d+\\s+plans' only. The plural is not the
template's own output shape — templates/roadmap.md:62 ships '1 plan' — so a
single-plan phase fell through every arm and its line froze permanently, never
updating again. Widened to 'plans?'. This one was introduced by the arm-3 default:
before it, the singular fell through to the old replace-everything path and at
least stayed current.

Cases 11-14 cover both: a bracketed human note preserved, the short placeholder
still replaced, '1 plan' rewritten, and '1 plan (annotation)' rewritten with the
annotation intact. Verified against the live binary, not just re-read.

Also converted all 15 cases in this block from try/finally to t.after(), per
CONTRIBUTING.md:356-370 which bans try/finally in test bodies. The existing #2853
block above is untouched — it is not in this change's scope and its conversion is
not this fix's concern.

Refs #3584

* chore(3584): add changeset fragment

Fixed-type fragment for the roadmap Plans-line trailing-text fix. pr:0 placeholder, backfilled once the PR number exists.

Refs #3584

* chore(3584): backfill changeset PR number (#3635)

---------

Co-authored-by: sim <sim@local>
2026-08-18 16:39:17 -04:00
Tom Boucher
ac1b6d679f enhance(#3618): fold fallow-runner onto the canonical binary resolver (epic #3411 Phase 2) (#3633)
* chore(#3618): fold fallow-runner onto the canonical binary resolver

Epic #3411 Phase 2. src/fallow-runner.cts was the fourth divergent
implementation of Windows binary resolution the epic enumerated —
candidateNames, isExecutableFile, findInPath, findInNodeModules, 40 lines.
All four are deleted; resolveFallowBinary is one seam call.

Two OPT-IN options were added to resolveExecutableBinary to make the fold
behavior-preserving, both defaulting off so Phase 1's callers are byte-identical:

  prependPaths      dirs searched before env.PATH, in order, through the
                    identical per-directory candidate logic. This expresses
                    node_modules/.bin-first precedence without env surgery —
                    the rejected alternative re-introduced the
                    spread-loses-the-proxy hazard the Windows lane caught in
                    Phase 1, at every future call site instead of once.
  requireExecutable POSIX-only accessSync(X_OK); a no-op on win32 where mode
                    bits do not mean execute. Opt-in rather than default
                    because unconditional X_OK breaks #3445's suite, which
                    stages candidates with plain writeFileSync and never sets
                    an exec bit — the repo bans chmod in tests — so every one
                    would resolve to null on POSIX.

Deliberate behavior change on Windows: fallow's prior candidate list ended in a
BARE fallow. The seam never tries a bare name there, so an extensionless file
beside fallow.cmd is no longer resolved. That is the fix, not a regression — the
extensionless file is npm's POSIX sh shim, which CreateProcess cannot run
(#3275). Rows 7 and 8 of the design record it.

Defect found while working, fixed inline: the resolution order was documented
BACKWARDS as PATH-then-.bin in structural-pre-pass.md, docs/INVENTORY.md and
four INVENTORY translations. The code has always been .bin first, and .bin first
is correct — a project-local tool should beat a global one. The archived
changeset is left alone as a historical record.

fallow-runner had no test file at all. tests/fallow-runner.test.cjs is new
(F1-F15) and the seam options are pinned by S1-S12 folded into the existing
dispatch suite. RED proven by execution: with both source files stashed and
build:lib re-run, 7 of 27 probe cases failed.

Refs #3411

* chore(#3618): backfill changeset pr number 3633

* fix(#3618): assert both platform contracts in F4 instead of a POSIX-only premise

Windows CI on #3633 failed F4. The test monkeypatched accessSync to throw and
asserted resolveFallowBinary returned null — but that premise, that the X_OK
check is consulted at all, is POSIX-only by design. requireExecutable is a
deliberate no-op on win32 because Windows mode bits do not mean execute, so the
staged fixture correctly resolved there.

40-design.md's negative-space section already states this carve-out verbatim.
The test contradicted the design it was written from: fixtures were made
platform-adaptive in the previous commit, and this assertion was left
platform-blind.

F4 now asserts BOTH contracts — null on POSIX, resolves on win32 — rather than
skipping either. A t.skip on one lane would have been green and would have left
the win32 carve-out unpinned by fallow's own entry point.

Audited every other row for the same class. F1-F3, F5, F6, F11-F15 hold on both
platforms; F7-F10 and S1-S12 inject platform explicitly and are unaffected. F4
was the only row with a single-platform premise.

The local probe runs on one platform and structurally cannot catch this, which
is why it was green — that limitation is now stated at the top of the probe so a
green probe is not mistaken for platform coverage. The win32 branch was proven
by injecting platform:'win32' with accessSync throwing and asserting it still
resolves.

Refs #3411

---------

Co-authored-by: sim <sim@local>
2026-08-18 15:32:57 -04:00
Tom Boucher
46f14c621e fix(#3583): one percent per write — route update-progress through the shared computation (#3634)
* test(3583): failing-first coverage for one percent per write

state update-progress computes plan throughput (summaries/plans) for stdout and the
body Progress bar, while the same write re-derives frontmatter progress.percent as
min(planFraction, phaseFraction). Neither consults the other, so mid-phase the file
contradicts itself and state json disagrees with the verb that just wrote it.

These tests fail on that: equality across stdout, body bar, frontmatter and state json
on fixtures where the two fractions differ, plus a derivation-parity test that fails if
completedPhases is ever derived by summary parity instead of verification-passed status.

Also updates three pre-existing tests that pinned stdout to the plan-throughput value
(50->0, 50->0, 100->0). Those fixtures have summarized-but-unverified phases, so the old
expectations encoded the bug; changing them IS the fix, as the issue states explicitly.

* fix(3583): one percent per write — route the verb through the shared computation

RED proven at 7dbbb2d2: 9 failures — the new cross-surface equality tests, the withhold
test, and the pre-existing tests whose expectations encoded the bug.

state update-progress computed plan throughput (summaries/plans) for stdout and the body
Progress bar, while the SAME write re-derived frontmatter progress.percent as
min(planFraction, phaseFraction) through a separate path. Neither consulted the other, so
on any project where plan throughput ran ahead of phase completion — the normal mid-phase
state — the file contradicted itself and state json disagreed with the verb that had just
written it. Exit 0, no signal.

This is not a dispute about which metric is right. The min cap is deliberate (#3242 Bug B)
and is untouched; the fix aligns the printed and body values WITH it. Verified by diff:
computeProgressPercent's definition and cmdStateSync are both unmodified.

The verb now takes its percent from buildStateFrontmatter — the single owner of the
isPhaseComplete-based completedPhases count and the ROADMAP-union totalPhases logic that
the frontmatter sync later uses inside the same read-modify-write. Both calls hit the same
disk-scan cache against the same on-disk state, so they cannot disagree. Reusing that
owner, rather than re-deriving completedPhases locally, is the point: a second
almost-identical derivation is the very defect class being fixed, and a parity test now
fails if anyone swaps it for summary parity.

The first cut fell back to plan throughput when the shared computation withheld. That
reintroduced the defect in a rarer case — stdout would print a number the frontmatter
deliberately did not contain — so it is gone. The verb now withholds in the same shape as
its existing #3217 and #3233 guards. That path is reachable, not theoretical: a bare vX.Y
token in ROADMAP prose with no versioned heading leaves the milestone unbounded while both
existing guards see a COMPLETE scope. Covered by a test that also asserts state json omits
the percent, proving it is the same withhold rather than a divergent local computation.

Three pre-existing tests pinned stdout to plan throughput (50->0, 50->0, 100->0); their
fixtures have summarized-but-unverified phases, so those expectations encoded the bug.
Updating them is the fix, as the issue states.

Fixes #3583

* fix(3583): source the reported counts from the same milestone window as the percent

The adversarial pass found the first cut left the SAME defect one field over.

cmdStateUpdateProgress still reported completed/total from the top-of-function scan,
which calls listMilestonePhaseDirs with NO versionOverride — the auto-derived current
milestone — while percent now came from buildStateFrontmatter, whose scan scopes by
versionOverride: storedMilestone. getMilestonePhaseFilter shows those can select
different milestone windows, and #3017's own comment warns about exactly that mis-bind.
So a single JSON object could report a percent inconsistent with its own counts: the
self-contradiction this issue was filed to close, relocated rather than removed.

Counts now come from the same buildStateFrontmatter result as the percent. Proven on a
real divergent-milestone fixture where a preamble phase leaks into the auto-derived scan
but is excluded from the stored-milestone-scoped one: with the fix stashed the verb emits
{percent:0, completed:1, total:2}; with it applied, {percent:0, completed:1, total:1}.
The guard scan remains, gating only the #3217/#3233 withholds.

Also corrected a comment that overstated caching. Only the phase/plan disk scan is shared
between the two buildStateFrontmatter calls; getMilestoneInfo re-reads and re-parses
ROADMAP.md and readGitHeadSha spawns a bounded git rev-parse, and both now run twice per
invocation. Threading a precomputed frontmatter through the write seam to avoid it was
rejected: that seam is the shared ADR-3408 §8.3 composition with three other callers and
heavily-documented invariants, and this is not the change to renegotiate it. The comment
now says what is and is not cached instead of implying the second call is free.

Standards: six new assertions matched raw STATE.md body text the code under test had just
produced — the pattern CONTRIBUTING bans by name. They now extract the body Progress field
with the repo's own field extractor and assert the parsed percent, so the check survives
rewording of the rendered bar. The acceptance criterion still verifies the bar; only what
it asserts on moved.

Also trimmed ~50 lines of narration around a ~15-line change into a named helper, and
fixed a stale test comment that still claimed 100% next to assertions expecting 0%.

* chore(3583): add changeset fragment

* chore(3583): backfill changeset PR number (#3634)

---------

Co-authored-by: sim <sim@local>
2026-08-18 15:24:07 -04:00
Tom Boucher
bf87dd4156 enhance(#3617): one canonical Windows binary resolver in the platform seam (epic #3411 Phase 1) (#3621)
* feat(#3411): one canonical Windows binary resolver in the platform seam

CONTEXT.md declares src/shell-command-projection.cts the single OS-facing seam,
but Windows binary resolution had grown four divergent implementations outside
it. #3445 folded two of them together — inside gsd-core/bin/gsd-tools.cjs, not
the seam — so the declaration stayed untrue and execTool still had no handling
at all.

Lift the resolver into the seam as resolveExecutableBinary, and export the half
that actually executes as projectSpawnInvocation: CreateProcess cannot run a
.cmd/.bat, so the cmd.exe mediation is inseparable from the lookup and splitting
them is how the copies accumulated. cmd.exe is invoked with an explicit argv
array, never shell:true — CVE-2024-27980's vector and Node 26's DEP0190.

execTool now resolves on win32. POSIX is a strict no-op by construction, which
matters: execTool rates CRITICAL blast radius (167 symbols, 53 files).
gsd-tools.cjs deletes its private scan and its private mediation and delegates.

Two semantics grown beyond #3445's resolver, both additive: a name already
carrying a PATHEXT-listed extension is tried as-is before the append loop, and a
suffix outside PATHEXT is not treated as an extension.

Refs #3411

* fix(#3411): keep mediating a declared .cmd that PATH resolution misses

Standards review caught a narrowing against the code this replaces. gsd-tools.cjs
computed `target = resolveSpawnBinary(binary) || binary` and keyed the shim test
on `target`, so a declared .cmd mediated whether or not PATH resolution found it.
That is load-bearing: resolveExecutableBinary scans PATH only, while `cmd.exe /c`
also finds a batch file in the current directory.

Mediation now keys on the target — resolved path, else declared name. The ENOENT
contract still holds for BARE unresolved names, which is the case it was written
for. P9/P10 pin both halves.

Spec review found E1/E2/E3/E5 promised by 50-test-matrix.md but never written;
added. E3 is the integration proof that the CVE-relevant mediation fires through
execTool, not only through projectSpawnInvocation in isolation.

Also adds the CONTEXT.md glossary entry for the seam's new resolution ownership
(a PR gate) and the changeset fragment.

Refs #3411

* fix(#3617): pass mediated cmd.exe arguments verbatim so metacharacters cannot inject

The isolated security pass found the mediation shape carried an argument-injection
surface. libuv's quote_cmd_arg force-quotes an argv element only when it contains
a space, tab, or quote — never for a cmd metacharacter — and cmd.exe re-parses
everything after /c. So an arg of a&calc arrived unquoted and cmd ran calc.
Node's own CVE-2024-27980 escaping cannot help: it fires only when the spawned
FILE is the .bat/.cmd, and here the file is cmd.exe.

Caret-escaping is not a fix. It is correct only when libuv does not quote, and
libuv quotes whenever the arg also contains a space — no per-arg transform is
right in both cases. So build the command line and pass it through verbatim, the
shape Rust's std adopted for the sibling CVE-2024-24576: one outer quote pair
that cmd /c strips, every token inside force-quoted, embedded quotes doubled.

An argument containing CR or LF is refused rather than mediated — a newline
cannot be represented in a Windows command line, so mediating would silently
truncate. Failing visibly is correct.

Known limit, documented at the seam: %VAR% still expands inside a /c string and
has no escape outside a batch file. That is information disclosure, not arbitrary
execution, and is the same limit Rust's std documents.

This was byte-for-byte the shape #3445 shipped, so the fix closes it for the
reviewer-lane spawn path too, not only for execTool's newly reachable route.

Refs #3411

* docs(#3617): document the subprocess-execution security posture

Adds Layer 4 to the security model: why GSD never uses shell:true for binary
invocation (CVE-2024-27980, Node 26 DEP0190), why resolution is explicit and
never tries the bare name on Windows (the npm extensionless-shim trap behind
#3275), and why .cmd/.bat mediation builds a verbatim force-quoted command line
rather than relying on default escaping — Node's own CVE protection cannot fire
once the started program is cmd.exe.

The residual %VAR% expansion limit is stated plainly under Trade-offs rather
than left implicit: it is information disclosure, not arbitrary execution, and
callers passing untrusted text to a Windows .cmd should not assume the value
arrives byte-identical.

Docs-only; no code change.

Refs #3411

* chore(#3617): backfill changeset pr number 3621

* fix(#3617): read PATH, PATHEXT and ComSpec case-insensitively

The Windows CI lane on #3621 failed E5, and the root cause was a defect in the
implementation, not the assertion.

Windows names the variable Path, not PATH. process.env is a case-insensitive
proxy, so process.env.PATH works — but execTool builds
{ ...process.env, ...opts.env } whenever a caller supplies opts.env, and
spreading discards the proxy while keeping the OS's actual casing. The exact-case
env['PATH'] lookup then returned undefined, the PATH scan saw zero segments,
resolution returned null, and the change degraded to precisely the spawn ENOENT
it exists to fix. ComSpec and PATHEXT had the same exposure.

#3445's tests never caught it because they pass uppercase keys explicitly, and
neither did the Linux remote runner — this is a defect only the Windows lane
could see.

_envGet resolves a variable by exact match first (so a canonical caller pays no
scan) and falls back to a case-insensitive sweep.

R23 and P16 pin it and were proven RED by execution: with the fix stashed and
build:lib re-run, R23 returned null and P16 returned the cmd.exe default.

R24 was rewritten because the first version was vacuous — it staged foo.CMD, so
the default PATHEXT already contained .CMD and it passed against the broken code
for the wrong reason. It now stages foo.XYZ, an extension absent from the
default, and carries a negative control asserting that dropping the Pathext key
yields null. Re-proven RED the same way.

E5's assertion was corrected alongside the fix: 'PATH' in options.env expressed
the wrong contract. It now checks case-insensitively for the key.

Refs #3411

* fix(#3617): execTool spawns the declared name unless mediation is required

The Windows full-test lane on #3621 failed tests/graphify.test.cjs — the python3
identity check asserted 'python3' and got the absolute resolved path
C:\hostedtoolcache\windows\Python\3.12.10\x64\python3.EXE instead.

Those tests are correct and the change was wrong. They pin a long-standing
contract — execTool spawns the program name it was given — by spying on
spawnSync's first argument, and routing every win32 call through the projected
invocation broke it.

Resolving a .exe buys nothing. libuv's CreateProcess path already performs
PATH + PATHEXT search, which is why spawning a bare 'node' has always worked on
Windows. The only case the OS genuinely cannot spawn is a .cmd/.bat. So execTool
now adopts the projection only when mediation actually happened —
windowsVerbatimArguments is exactly that flag — and otherwise passes the declared
program and args through untouched.

40-design.md already rejected gratuitous change for this reason: symmetry is not
worth a behavior change to 53 files that fixes nothing. That reasoning was
applied to POSIX and missed the win32 non-batch case. Rows 5 and 20 now record
it, and the CONTEXT.md glossary states the caller-choice rule.

deps.spawn deliberately still adopts the resolved path: its hasBinary probe
answers from the same resolver, so probe and spawn must agree on the exact file
(#3445). The asymmetry is now documented at both call sites rather than latent.

E7 pins the restored contract and was verified by executing execTool against a
monkeypatched spawnSync: python3 in, python3 spawned.

Refs #3411

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:12:44 -04:00
Tom Boucher
bf2332e67c fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint

On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately
absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw
tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before
requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in
its fail-closed catch and is misreported as an unreadable dispatch-isolation
configuration, blocking every executor dispatch.

These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor
guard and update worker, plus the seam's actionable build error surfacing instead of
the generic misreport.

Also adds the drift lint the acceptance criteria require, with a fixture proving it
CAN fail — a guard never shown to fail is worthless. It is red here by design: it
flags today's unfixed hooks, which is exactly the defect.

* fix(3582): route every hook's compiled-module require through the self-heal seam

RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard,
cursor guard and update worker, the fail-closed-with-actionable-message assertion, and
the lint's own real-tree check.

The compiled runtime library is produced by build:lib and gitignored (ADR-457,
build-at-publish). The npm package builds before publishing; a plugin-marketplace or
git-clone install materializes the raw tree and never does, so on that channel those
modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly
this, and the CLI entrypoint already calls it — no hook did. The isolation guard's
Cannot-find-module therefore landed in its fail-closed catch and was reported as
'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was
misdiagnosed as an unreadable project config and every executor dispatch was blocked.

All SEVEN affected files now call the seam before their first compiled require. The issue
named four; a scan found six; implementing it surfaced a seventh — the shared isolation
sentinel helper, used by BOTH guards, which requires two compiled modules itself and
would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so
fixed here rather than left as a known-broken remainder.

Failure posture is deliberately split by hook kind:
- Gates (agent isolation guard, cursor subagent start) surface the seam's actionable
  build error distinctly instead of swallowing it into the generic text, and stay
  fail-closed — a genuinely unreadable project config still DENIES exactly as before.
- Cosmetic and detached hooks (statusline, update worker, update check, update banner)
  DEGRADE rather than crash: the statusline draws on every render and the worker is a
  detached process, so a build failure there must not take down the prompt.

The npm path is untouched: the seam's already-built fast path returns immediately, so
prebuilt installs pay nothing and behave bit-for-bit as before.

Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than
remembered — without it the next hook to add a compiled require reintroduces the class
silently. It is proven able to fail: a fixture hook requiring a compiled module without
the seam is flagged, and one that uses the seam is not. Verified directly — on the
unfixed tree it named all seven offenders; with the fix it passes.

While writing the lint's comment stripper, a naive whole-text block-comment regex ate its
own fixture, because this repo's comments legitimately spell the compiled-lib glob whose
star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression
test pinning that case.

* fix(3582): test the three untested seam call sites and assert typed reason codes

Two independent reviews converged on the same major gap: the fix wired the seam into
seven files but only four had cold-tree tests. The adversarial pass put it plainly —
deleting the shared isolation-sentinel helper's seam call would not have failed any test
in the diff. That file was my own addition beyond the issue's four, so it shipped
untested; that is now closed.

- Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT
  directly under cwd, and every existing cold-tree fixture puts it there, so the early
  return always fired first. Now covered, and proven load-bearing by mutation: with the
  call removed the spy records zero seam invocations and the test fails.
- update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED
  VERDICT — the fallback cache filename, and silent suppression when the package name
  degrades to null — rather than merely 'did not throw'. The banner hook previously had
  no test file at all.

Standards violation fixed: two tests asserted on free-form prose via assert.match against
a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly
that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch
it. Both isolation guards now emit a machine-readable reason_code from a frozen enum,
following the repo's existing REASON convention, and the tests assert that instead. The
human-readable message is unchanged for operators; only the assertion target moved.

The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT
extracted, and the reason is recorded at each site: both viable shapes — a
path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file
literal co-occurrence check, so extracting would require the lint to special-case its own
helper. Triplication is the lesser evil while the lint stays a co-occurrence scan.

The lint's header now states what it does and does not catch (literal quoted requires
only; hooks/ scan root), so a future reader does not over-trust a guard that a
concatenated path or a require inside a non-hooks helper would evade.

* chore(3582): regenerate the committed install-tree fixtures

Adding a new shipped hook helper changed the install tree, and those fixtures are
committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>'
tests failed on 541a1913. Regenerated rather than hand-edited.

The delta across all 15 runtime fixtures is exactly two lines — the new helper under both
its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in
no unrelated drift.

This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from
lint:ci, which passed both before and after.

* chore(3582): backfill changeset PR number (#3629)

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:11:23 -04:00
Tom Boucher
9de4d67118 fix(#3579): a pointer-less session inherits the repo active-workstream marker (#3616)
* test(3579): failing-first coverage for repo-marker inheritance

A session that carries an identity but has never run 'workstream use' reads an
absent session pointer, resolves null, and composes the flat .planning tree even
when .planning/active-workstream names a live workstream. These tests fail on that
and pin the invariants the fix must not break: a session with its own pointer is
never repointed, and a session that merely lacked a pointer must never clear the
shared marker on another session's behalf.

* fix(3579): a pointer-less session inherits the repo active-workstream marker

RED proven at 157cae26: the three inheritance tests failed while every isolation and
negative control passed on base — the gap, and nothing else.

pickActiveWorkstreamAdapter returned exactly ONE adapter: the session-scoped one
whenever a session key existed, so the shared .planning/active-workstream marker was
never consulted. getWorkstreamSessionKey resolves a key from ~13 env vars or the
controlling TTY, so on any normal interactive terminal a key almost always exists —
which is why a session that had never run 'workstream use' read an absent pointer,
resolved null, and composed the FLAT planning tree even though the repo marker named a
live workstream. Reads misreported; writes corrupted the superseded flat STATE. Silent,
because the stale tree is well-formed.

This was a genuine design fork, not an oversight: references/workstream-flag.md
documented step 4 as a fallback 'when no session key exists', and the session isolation
that buys is deliberate (#2850). The issue's Agent Brief left the choice open and said
the reference doc should match whatever semantics ship. The maintainer ruled in chat for
inheritance.

Resolution now walks an ORDERED chain — session adapter first, shared second — and only
a null from the session adapter falls through to the marker. Strictly additive: it can
only turn a null into a name, never change a name that already resolves.

The dangerous part is clear() ownership. resolveFromChain treats chain[0] as owned: only
it is ever cleared, and only under selfHeal (getActiveWorkstream, never peek). An
INHERITED marker is read-only — a stale value there resolves null and the file is left
alone. Without that, one pointer-less session's read would delete the repo marker for
every other session, which is a worse bug than the one being fixed. Covered by a test
that asserts the marker still exists on disk after such a read.

peekActiveWorkstream inherits but still mutates nothing (#2850 — the statusline draws on
every render).

references/workstream-flag.md's Resolution Priority is rewritten to match, keeping the
session-isolation rationale and noting that inheritance does not weaken it: a session
that owns a pointer is never repointed.

Fixes #3579

* fix(3579): correct the guard diagnostics and lock the clear-semantics

Three review passes; every finding fixed inline.

MISSING ACCEPTANCE CRITERION (spec pass). The brief requires refusal diagnostics that
distinguish 'marker present but the session lookup missed it' from 'no workstream set at
all', and the two workstream-mode fail-safe guards were byte-for-byte untouched — still
emitting a generic 'no active workstream is set' even when a marker exists and merely
names a missing directory. Both guards (cmdPhaseComplete, cmdInitProgress) now branch on
a new read-only diagnoseUnresolvedActiveWorkstream, which reuses the SAME
resolvesToExistingWorkstream predicate resolveFromChain uses, so the diagnosis and the
resolution cannot disagree. Two typed reasons added to ERROR_REASON; both arms still
refuse — the fail-closed behavior is unchanged, only the message is now true.

REAL TEST FAILURE, not a flake. The remote run failed 'clearing one session does not
clear another session pointer'. That describe uses before() rather than beforeEach, so
one tmpDir is shared and an earlier test writes active-workstream=beta into it; under
inheritance the just-cleared session picks that marker up and resolves beta instead of
null. The failure is a CORRECT consequence of Option A surfaced through an
order-dependent fixture. The test now establishes its own marker state explicitly — its
real intent (clearing A must not disturb B's pointer) is preserved and not weakened — and
a new test pins the semantic deliberately: clearing a session pointer returns that
session to INHERITING the marker, it does not force flat mode. Documented in
references/workstream-flag.md, including how to actually get flat behavior.

Also from review: partial activeWorkstreamAdapters injection no longer silently
synthesizes a REAL filesystem adapter for the missing half (a latent test-isolation
trap); the duplicated validate-then-existsSync logic is factored into one predicate; and
the two try/finally test bodies are converted to t.after per CONTRIBUTING.

New coverage: whitespace/empty shared marker; a session whose OWN pointer is stale while
the marker names a different valid workstream (must self-heal to null, never inherit —
the isolation guarantee at its sharpest); and both new diagnostic arms asserted on
structured --json-errors output rather than prose.

* fix(3579): read resolvability with the non-mutating peek, not the self-healing resolver

Three of our own new tests failed on 7f5e706a. All three had ONE root cause, and none
was fixed by relaxing an assertion.

gsd-tools.cjs's bootstrap called the MUTATING getActiveWorkstream unconditionally on
every invocation, purely to populate routing env. On an unresolvable pointer that
self-healed — cleared it — BEFORE the dispatched command ran its own resolution. A second
read in the same process then observed already-cleared state:

- Isolation violation: a session whose own pointer was stale had it cleared by the
  bootstrap, so cmdWorkstreamGet's own resolution found a pointer-LESS session and
  inherited the shared marker ('beta' instead of null). Exactly the guarantee #2850 exists
  to protect, defeated across two calls rather than within one.
- Guard diagnostics: the guards' own truthiness check also used the mutating resolver, so
  it cleared the invalid marker and the immediately-following read-only diagnosis found
  nothing and reported none_active instead of marker_unresolved.

So a single invocation's answer depended on how many times it resolved. The bootstrap
self-heal is PRE-EXISTING and was harmless while pointer-less meant flat — inheritance is
what made it answer-changing, so this fix belongs here.

Every call site that only CHECKS resolvability — the bootstrap, both fail-safe guards'
truthiness check, and two informational init report fields — now uses the non-mutating
peekActiveWorkstream. Self-heal is unchanged in active-workstream-store and still fires
exactly once, at whichever site actually consumes the workstream.

Verified by driving the real CLI against temp fixtures, since the suite cannot run
locally: stale-own-pointer resolves null with the marker intact; both guard arms report
marker_unresolved with missing_workstream_dir / invalid_name and the marker survives;
no-marker still reports none_active; identity-less self-heal still deletes an invalid
marker byte-identically to pre-#3579; and a session with a valid own pointer still wins.

* chore(3579): backfill changeset PR number (#3616)

* test(3579): kill the surviving mutants in the new resolution code

CI's Stryker gate failed: active-workstream-store scored 79.45% against a break
threshold of 80 — 259 killed, 67 survived, at 'Ran 1.00 tests per mutant on average'.
The survivors cluster in the code this PR added (pickActiveWorkstreamAdapterChain,
resolvesToExistingWorkstream, resolveFromChain, diagnoseUnresolvedActiveWorkstream):
the CLI-level tests exercise those paths but do not DISCRIMINATE their branches, which
is precisely what a surviving mutant means.

Raised by strengthening assertions, never by touching the threshold. 21 unit tests added
to the existing unit suite, each written to fail under a specific named mutant, using the
module's injected adapter seams and createMemoryPointerAdapter so they stay hermetic
under Stryker's per-mutant reruns:

- chain shape with and without a session key, asserting length AND element identity
  (kills the if(false), the ': []' array mutant, and the block removal)
- partial adapter injection, asserting the missing half is an inert memory adapter that
  never touches the filesystem (kills the three '??' -> '&&' mutants)
- both arms of '!name || !validateWorkstreamName(name)' as SEPARATE tests — an absent
  name and a non-empty invalid one — which is what kills the '||' -> '&&' mutant
- self-heal discrimination: getActiveWorkstream must clear an unresolvable owned pointer
  and peekActiveWorkstream must not, asserted on adapter state after each
  (kills if(selfHeal) -> if(true))
- fallback arm both ways: a fallback that resolves and one that does not
- diagnoseUnresolvedActiveWorkstream asserted as a full object per case, with the reason
  strings compared exactly (kills present:true -> false and both StringLiteral mutants)

One mutant is deliberately left: 'if (chain.length === 0)' -> 'if (false)'. The branch is
structurally unreachable — the only chain source always returns a 1- or 2-element array
literal — and resolveFromChain is not exported. Killing it would mean exporting an
internal or deleting a defensive guard; neither is worth doing for a mutant, and the
score clears 80 without it. Recorded here rather than left unexplained.

Every new assertion was evaluated against the built module with real fixtures before
committing, since the suite cannot run locally.

---------

Co-authored-by: sim <sim@local>
2026-08-18 11:37:44 -04:00
Tom Boucher
682eaae3f0 enh(#2876): retire the dead and pass-through exports from bin/install.js (#3615)
* enh(#2876): retire the dead and pass-through exports from bin/install.js

The installer exported 197 names and had zero production consumers - every
non-test require of it repo-wide sits inside a comment. Its interface was
shaped by test access, not by callers.

Removes 9 dead exports and 61 pass-throughs, repointing their tests onto the
extracted modules' own interfaces. 197 down to 127.

Every count in the issue was wrong: 197 exports not 188, 9 dead not 12, 61
pass-throughs not 49, 44 test files not 42 - and the audit itself then missed
7 more consumer files. restoreUserArtifacts was on the dead list but ceased to
exist in phase 6, and two _GSD_EFFORT_MANIFEST_* names listed as dead are now
genuinely asserted, so acting on that list would have deleted live exports.

7 of the 9 dead names collide with an independent declaration that install.js
delegates TO. Each removal was justified by which declaration a reference
resolves to, never by whether the name appears somewhere.

Coverage parity was the gate rather than test greenness: per-file counts were
captured before any edit and diffed after. 44 of 45 files are byte-identical;
the single delta is one added assertion, not a loss.

The sweep for scattered require sites found two forms static grep misses -
require(VARIABLE) and multi-line require() - plus tests asserting that
install.js re-exports the SAME object, which now assert retirement instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2876): close review findings — restore the duplicate-body guard, sweep orphaned code

Both review engines found real defects in the first cut.

The DEFECT.GENERATIVE-FIX single-owner guard from #1511 had been repointed
from a reference-identity check to install.X === undefined. Those are not
equivalent: the guard exists to catch a duplicate function body reintroduced
into install.js, and the replacement passes cleanly if that duplicate is used
internally and never exported. It now walks bin/install.js's real top-level
bindings, so it catches a duplicate under either shape, exported or not -
strictly stronger than the check it replaced. Proved by injecting a duplicate
and watching it go red.

That weakening survived the coverage-parity gate because the assertion count
never moved. The gate compares counts, so an assertion that changes meaning
rather than number is invisible to it.

Removing the exports had orphaned their wrapper bodies: 14 dead wrappers, 9
consts and 9 destructure entries, several pre-existing and found by the same
sweep. Dead code left in the file this phase exists to shrink.

Three more comments claimed re-exports this phase removed, and tests were
reading Cursor and Windsurf hook constants from install.js's local copy while
calling functions from the hooks surface - equal today, with nothing holding
them equal. The local consts now reference the owning module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2876): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 09:23:44 -04:00
Tom Boucher
bcefffc132 fix(#3578): derive milestone status from phase counters, not phase-completion prose (#3614)
* test(3578): failing-first coverage for milestone status on partial completion

Completing phase 2 of a 4-phase milestone sets frontmatter status: completed while
the same call correctly writes completed_phases: 2 / total_phases: 4. These tests
fail on that conflation and pin the boundary either side of it (3-of-4 must not
complete, 4-of-4 must), plus milestone_name byte-identity and the 1-of-1 case that
legitimately does complete.

* fix(3578): derive milestone status from phase counters, not phase-completion prose

RED proven at 253843b4 (tests-only): the 2-of-4 and 3-of-4 cases failed while the
4-of-4, milestone_name and 1-of-1 controls passed — the conflation, and nothing else.

state complete-phase writes body prose `Phase N complete`. normalizeStateStatus
matches 'complete' as a case-insensitive SUBSTRING, so phase-level prose collapsed
into milestone-level frontmatter status: completed — even while the same call
correctly derived completed_phases: 2 / total_phases: 4 / percent: 50.

Check ORDER is why the sibling surface stays correct: completePhaseCore writes
'Ready to plan' for non-final phases, hitting the 'planning' arm before 'complete'.
The two phase-completion surfaces disagreed and this was the conflated one — a
violation of ADR-2207, which gives milestone termination solely to
milestoneCompleteCore.

buildStateFrontmatter now honors a 'completed' normalization from phase-completion
prose only when the counters it already derived agree. Scoped deliberately:

- anchored to bare `Phase <token> complete`, so 'All phases complete' and
  '<version> milestone complete' are untouched (both out of scope). Verified by
  executing the guard's own regex from source against both forms.
- gated on counter trustworthiness (COMPLETE disk scope, finite counts, positive
  denominator) so an unknown scope withholds rather than guessing 'not complete',
  which would be the mirror-image bug
- normalizeStateStatus itself is NOT modified — it feeds every state.* write and
  the read path

A 1-of-1 milestone still yields 'completed' by the rule, not by exemption, so the
#1255 pinning test stays green on its merits.

Fixes #3578

* fix(3578): gate the guard on milestone boundedness and close the review gaps

Review findings from two orthogonal passes, all fixed inline.

GUARD (correctness, from the standards pass): the guard omitted `milestoneUnbounded`,
which is the established trust authority for these very counters in this same function
— it nulls progressPercent at :2286 and gates the prose fallback at :2294. An unbounded
milestone yields a conflated/understated total, so `completedPhases < totalPhases` could
be an artifact of a bad denominator and demote a genuinely-complete milestone. Now gated.

TESTS:
- Prose/guard parity assertion. The guard regex-matches prose emitted from a DIFFERENT
  file; if that prose drifts the guard silently stops firing and the bug returns
  undetected. Per the repo's generative-fix-divergence rule, a test now asserts the
  emitted body Status still matches the guard's pattern — asserting the emitted value
  against the pattern rather than duplicating the string.
- limit+1: completedPhases > totalPhases must NOT fire; inconsistent counters fall
  through rather than guessing.
- Untrustworthy counters (no phases dir → totalPhases null) must NOT fire.
- AC4: MCP invoke-command dispatch parity via handleMessage, the criterion both
  reviewers independently flagged as asserted-but-untested.
- Hand-rolled STATE.md writes routed through the existing writeState fixture helper.

The adversarial pass independently verified, by reading rather than trusting the diff's
own comments, that: paused/stopped short-circuit before 'completed' so a paused milestone
can never be clobbered; only cmdStateCompletePhase emits the targeted prose, so no
sibling caller over-fires; the counters come from a fresh disk scan independent of this
write, so there is no pre/post off-by-one; and the #1255 pinning fixture creates no
phases dir, leaving completedPhases null and the guard inert — so that test is provably
unaffected rather than assumed to be.

* chore(3578): add changeset fragment

* chore(3578): backfill changeset PR number (#3614)

---------

Co-authored-by: sim <sim@local>
2026-08-18 08:33:46 -04:00
Tom Boucher
b42cb4fb29 fix(#3597): count scenario expectation failures in the QA gate, and fix the workstream scope split it exposed (#3607)
* fix(#3597): count scenario expectation failures in the QA ratchet gate

buildReport counts totals.violations as oracle violations PLUS scenario
expectFailures, but collectFindings read only step.violations. A scenario
whose declared expect failed therefore produced ok:false and violations:1
in the report while the ratchet printed "0 violations" and exited 0.

multi-workstream has failed that way on every CI run since 2026-08-10,
when #3217 (PR #3318) made computeProgressPercent withhold a percentage
whose scope is not COMPLETE. The walk detected the change the day it
landed; nothing was listening.

- collectFindings returns a third bucket, expectationFailures, carrying no
  fingerprint so it can never be baselined or acked away
- both modes of main() print and gate on it; the summary line reports it
- guard runMain(main) behind require.main === module, so the QA suite can
  require the script to test collectFindings without running a real walk
  (that import side effect is why the gate logic had no test)
- multi-workstream now asserts the true contract: phase_scope unreadable
  and percent null, per ADR-3180 7.6 rule 4
- the perturbation test asserts scenario ok, closing the test-side half

Closes #3597

* fix(#3597): resolve the milestone window against the active workstream

listMilestonePhaseDirs defaulted its ws option to null. planningDir
treats undefined as "resolve the ambient workstream" and null as
"force the project root", so that default suppressed the ambient
resolution every other planning-path read uses.

All 18 call sites derive phasesDir ambiently via planningPaths(cwd),
so the counts came from the workstream while the milestone window came
from the root .planning/ROADMAP.md — the exact numerator/denominator
scope split ADR-3180 7.6 rule 3 forbids. workstream create migrates
that root roadmap away, so the read threw and scope stayed UNREADABLE,
and rule 4 then correctly withheld the percentage.

Proof: with a workstream tree byte-unchanged, copying its own ROADMAP
to the project root flipped --ws alpha progress from
phase_scope:unreadable/percent:null to complete/100.

This is the defect the loop QA walk was pointing at all along; the
scenario expectation is restored to percent:100 rather than bent to
match the bug.

- pass ws through as undefined so ambient resolution applies
- multi-workstream asserts phase_scope complete + percent 100
- regression test in completion-ratio-scope-withholding covers a
  workstream-only project with no root ROADMAP
- replace the vacuous require.main test: runMain defers through a
  promise, so the in-process timing check passed against the unguarded
  file too; a child-process spawn now observes the guard for real
- tie the oracle-violation test to expectationFailures, and cover the
  absent-key, multi-scenario and zero-step report shapes in parity
- flatten scenario-authored strings before rendering them into the
  step summary and CI logs (forged markdown / ANSI injection)
- widen the scenario contract assertions past perturbation-* so
  multi-workstream is actually covered test-side

Closes #3597

* fix(#3597): flatten scenario-authored strings on the CI-log output path

The step-summary path already routed findings through flattenUntrusted;
the check-mode NEW-smell and STALE-entry console.error blocks, and the
repro line in both printers, still interpolated raw.

detail carries a scenario-authored expect[].path verbatim, and
reason/scenario/id come from contributor-authored baseline and ack
fragments validated only as non-empty strings. A crafted path could
print a forged summary line into the CI log directly above the real
one, plus ANSI repaint and unbounded length.

Exit codes are unaffected — this is log spoofing, not gate bypass.

* fix(#3597): refuse to archive on an unreadable milestone window; close review gaps

Resolving the milestone window against the active workstream can leave
the window UNREADABLE when that workstream has no ROADMAP of its own.
getMilestonePhaseFilter throws, the window degrades to a pass-all
fallback, and milestone complete would then move every phase dir --
breaking the guarantee stated at the archive site that no out-of-window
directory is touched.

milestone complete now refuses to archive when the window is UNREADABLE
and reports the refusal; --dry-run previews the same refusal from the
same shared derivation.

The guard is scoped to UNREADABLE, not to every non-COMPLETE scope. A
broader condition regressed ordinary root projects: the QA walk caught
milestone-rollover leaving 01-parser on disk, which then tripped the
#1447 abort in phases clear. UNSCOPED and TRUNCATED are pre-existing
classifications and keep their existing behavior.

Review fixes:
- the workstream regression test asserted complete/100 but its fixture
  wrote no workstream STATE.md, so it resolved unscoped/null and the
  test failed; it now asserts a milestone and genuinely fails-first
- the parity test hand-supplied totals.violations, hardcoding the very
  formula under test; at least one case now goes through the real
  buildReport
- drop a vacuous qa-report.json assertion (jsonOut defaults to null, so
  no report is written by either shape)
- buildRepro emitted a repo-relative binary path after cd-ing into a
  temp project, so every repro died with MODULE_NOT_FOUND; it now
  resolves an absolute path
- flattenUntrusted truncated the repro to 300 chars, handing reviewers a
  command that looks complete and is not; length capping is now opt-out
  for repro while newline/control/backtick stripping still applies

* chore(#3597): backfill changeset pr number (#3607)

---------

Co-authored-by: sim <sim@local>
2026-08-18 07:26:28 -04:00
Tom Boucher
fe64704ace enhance(#3588): add an opt-in commit_docs pre-commit hook (#3609)
* feat(#3588): add an opt-in commit_docs pre-commit hook

Final phase of epic #2292, scope narrowed to opt-in by maintainer decision:
default-on installation and the bin/install.js wiring it would have required
are explicitly out of scope.

Enabling is an explicit verb call. The hook is written to the repo's real hooks
dir resolved via git rev-parse --git-path hooks, so a linked worktree or
submodule whose .git is a FILE works rather than getting a literal .git/hooks
path. It refuses rather than overwrite a foreign pre-commit, refuses to delete
one it did not write, and refuses outright when core.hooksPath is already set --
a written-but-ignored hook is worse than a refusal. Ownership is detected by
marker presence, not byte-equality, so a user who appends a line does not make
it unrecognizable.

Deliberately NOT included: teaching cmdCheckCommit the per-phase commit_docs
tier. #3587 was still unmerged when this landed, and implementing precedence
against helpers that did not yet exist would have meant a second copy of the
resolution chain -- the divergence class this epic has spent three phases
fighting. That follows as its own change now that #3587 is on next.

The ordering constraint is recorded in the design doc: this must not merge
before #3587, or the hook would block a commit cmdCommit itself allows.

* fix(#3588): teach the commit_docs guard the per-phase tier and -z paths

Part 1, deferred until #3587 merged.

cmdCheckCommit read only project-level commit_docs, so once #3587 landed, a
phase with phase_commit_docs true under project false was ALLOWED by
query commit and BLOCKED by this guard -- and the hook shipped in this same
branch shells out to it. It now derives the staged phase via the single-owner
detectPhaseNumberFromFiles and resolves through #3587's own
resolveCommitDocsPolicy rather than a second precedence copy.

Also fixes a proven false negative in the harm direction. git diff --cached
--name-only C-style-quotes any path with non-ASCII or special characters, so a
staged .planning/cafe.md was emitted as a quoted string, failed
startsWith('.planning/'), and slipped past the guard entirely under
commit_docs:false. Reading with -z and splitting on NUL removes the quoting at
the source. The f.startsWith('.planning\\') branch was dead code under that
read -- git emits /-separated paths on every platform -- and is removed rather
than left implying coverage it never provided.

The earlier C7 test pinned the buggy behavior as intended; it now asserts the
file is detected and the commit refused.

Self-caught: the commit-docs-guard verb was wired into the routers by this
branch's earlier pass but missing from the top-level help listing.

* test(#3588): replace try/finally with t.after, add negative-routing cases

Standards review findings.

CONTRIBUTING bans try/finally inside a test body outright -- it masks failures
-- and B8 used one for worktree cleanup. Now t.after(), assertions unchanged.

The new commit-docs-guard command family had zero negative-routing coverage,
which CONTRIBUTING requires for any change to command dispatch. B11-B15 cover
no subcommand, unknown, empty string, whitespace-only and a flag-shaped value,
each asserting non-zero exit, a structured error, no stack trace, and -- the
one that matters for a command that writes into a user's repo -- that NO hook
is written in any of them.

Those tests were verified to fail when routeCommitDocsGuard's else-branch is
neutered, so they exercise the routing guard rather than any convenient error
path.

Also made two error() calls' control flow explicit with a return; they were
safe only because error() is typed never two files away.

* chore(#3588): backfill changeset pr number to 3609

* test(#3588): skip Windows-unrepresentable fixtures on win32

CI's Windows shards caught two of my own tests: fixtures whose filenames
contain a quote and a backslash. Both are illegal on Windows -- backslash is
the path separator, quote is invalid on NTFS -- so fixture creation failed
before any assertion ran.

Test-portability defect, not a production one. Those inputs cannot exist on
that platform, so the guard has nothing to detect there.

Both now check process.platform FIRST, before any fs or git call, and use
t.skip() rather than a bare return -- a bare return registers as a PASS and
would hide the gap it is meant to record. Each carries a comment saying the
input is unrepresentable rather than unverified, so nobody later re-enables it.

No padding added: the cafe.md case already exercises git's C-quoting path on
every platform, since non-ASCII names are legal on NTFS.

This is exactly the coverage the Linux-only remote matrix cannot provide, which
the PR body already stated -- CI's Windows shards are what caught it.

---------

Co-authored-by: sim <sim@local>
2026-08-18 00:25:32 -04:00