Commit Graph

5400 Commits

Author SHA1 Message Date
Tom Boucher
738f42f4fd feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs (#3755)
* test(#2398): failing-first suite for the CYCLE_SUMMARY consensus gate

Binds the gate before it exists, so the suite is RED against next.

The load-bearing rows are the two the closed PR #2417 did not have. The B2
regression row asserts a judgment-class lone HIGH counts WITHOUT corroboration
when its raiser is unmarked — if anyone re-couples that class to corroboration,
more reviewers again produce a weaker gate than one, which is what closed #2417.
The parity row asserts every marker literal the gate names is one
review-lane-runner actually emits, so the gate cannot key on a signal nothing
produces; a mutation row and a seeded fast-check property prove that guard runs
its failure branch rather than only reading a correct tree.

Also pinned: gate position before Counting rules, the untouched CYCLE_SUMMARY
line shape the orchestrator greps, fence balance, the single-reviewer no-op,
classification by what a claim asserts rather than by citation presence, the
all-marked fail-open, current_actionable staying out of scope, and the
leading-marker requirement that stops a review which merely quotes a marker
from suppressing its own findings.

* feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs

With review.reviewer_instances running several reviewer identities off one
adapter, any single instance's fabricated HIGH could force a full replan cycle
on its own. Across ~9 real cycles on two projects each of four instances
fabricated at least once, and each was also the most accurate reviewer in some
other cycle, so dropping to fewer reviewers trades away real signal.

The gate engages only when 2+ reviewers actually ran, and weighs a lone HIGH by
what the claim asserts rather than by whether anyone agreed with it. An
existence claim -- a symbol, file, flag, commit or ID exists, is absent, or says
something specific -- counts only if source-grounding confirms it or another
reviewer raised the same concern. A judgment claim -- a design or correctness
property -- counts unless that reviewer's own section opens with an
evidence-quality discount marker the review lane already stamps
([reviewed-without-source-citations] #3194, [reviewed-without-repo-access]
#2176, or a diff-only lane).

That split is what resolves B2, the finding that closed PR #2417. B2 showed the
approved wording made more reviewers produce a WEAKER gate than one: condition
(a) pointed at the source-grounding pass, which verifies every symbol THE PLAN
cites and never takes reviewer claims as input, so a genuine architectural HIGH
that one reviewer caught and another missed was neither groundable nor
corroborated and stopped gating. Judgment-class findings are therefore exempt
from corroboration entirely -- reviewers catch materially different classes of
issue, and demanding two of them independently raise the same architectural
concern suppresses exactly what a multi-reviewer setup exists to surface.

Guards on the gate itself: an all-marked cycle disengages it, so a cycle in
which nothing was verified can never be counted as converged; the marker must
OPEN a reviewer's section, so a review that merely quotes a marker does not
suppress its own findings; a suppressed HIGH stays listed and tagged rather
than dropped; current_actionable is untouched; and a single-reviewer run is
unchanged.

No new command, config key, or dependency -- the gate reads signals that
already exist. The CYCLE_SUMMARY line shape the orchestrator greps is
unchanged; only the integer it computes moves, and only for 2+ reviewers.

Known limit, inherited rather than introduced: SOURCE_CITATION_RE checks
citation presence, not resolution, which src/review-lane-runner.cts records as
a deliberate #3194 scope boundary. A fabricated but plausible file:line still
gates.

Scope revised and re-approved on the issue before any code was written.

* test(#2398): make marker parity behavioral, and stop overclaiming the gate

Review found the parity tests were vacuous: they asserted a marker STRING
appeared in review-lane-runner.cjs's source text, never requiring the module or
calling the stampers, so they would pass even if stampUngroundedReview were
broken or never invoked. They now invoke the real exported functions and assert
what those functions PRODUCE — that an uncited review gains a leading marker
blockquote, that a review carrying a file:line does not, that a self-reported
blind review is stamped, and that stamping is idempotent. Removing the source
read also removes an incidental no-source-grep evasion via a parameterized path.

Review also found the changeset headline false for the class it matters most
in. The discount markers detect 'cited nothing' and 'had no repo access'; they
cannot detect 'drew a wrong conclusion from a real citation', so a judgment-class
finding invented by an evidence-bearing reviewer still counts alone. That is the
deliberate side of the tradeoff jags-faith named when closing #2417 — the
alternative is requiring corroboration for design findings, which is B2 — but
the changeset claimed lone hallucinations no longer force a cycle, full stop.
Corrected there, and stated plainly in docs/COMMANDS.md and the design record.

Also dropped the reviewer-instances.md entry from the emitted-drift ack: the
growth ratchet's currentSizes() scans only gsd-core/workflows/ and agents/
(tests/helpers/emitted-runtime.cjs:916-929), so references/ is outside it and
that entry acknowledged a delta the gate cannot see.

* chore(#2398): backfill changeset pr number to 3755

---------

Co-authored-by: sim <sim@local>
2026-08-22 10:53:14 -04:00
Behruz Nassre Esfahani
444069d601 fix(#3613): copy component dirs into the plugin-validate fixture (#3627)
* fix(#3613): copy component dirs into the plugin-validate fixture

C2 builds its synthetic plugin root by symlinking commands/, hooks/ and
skills/ into a temp dir. `claude plugin validate` (>=2.1.233) reads
component directories without following symlinks and warns on each one,
and `--strict` promotes a warning to a non-zero exit — so the assertion
failed on how the fixture was built, not on the manifest under test.
Copy the directories with fs.cpSync instead. The validated tree still
holds only plugin.json and the three component directories, so nothing
else from the repo root is placed where the validator can read it, and
the #2665 CLI-config isolation is untouched.

The construction helper is self-cleaning: its callers' try/finally only
begins once it returns, so a throw partway through would strand a
half-built root on disk. The pre-refactor code ran these same steps inside
C2's own try, and that teardown guarantee is preserved rather than
narrowed. Its cleanup is best-effort so a teardown error cannot replace
the real construction error.

Add C3, an unconditional tripwire asserting the fixture exposes real,
non-symlinked component directories at any depth. C2 never runs in CI —
no job under .github/workflows/ installs the `claude` binary — so a revert
to symlinks would pass every CI lane and surface only as a red suite on
contributor machines. C3 is not a substitute for C2's end-to-end check and
cannot be shown red against this base, since the helper it calls arrives
with it; C2 is the failing-first artifact.

C3 enforces a symlink-free fixture throughout, deliberately stricter than
the CLI's own boundary. Measured on 2.1.234, `plugin validate --strict`
exits 1 for a symlinked component dir and for a symlink one level inside
skills/, and 0 for two levels in or for symlinks under commands/ or
hooks/. Encoding that external, undocumented line would be more fragile
than a superset costing one walk over ~176 files. The walk uses an
explicit stack rather than readdirSync's `recursive: true`, which follows
directory symlinks — a link pointing outward would otherwise traverse an
unrelated tree, or a cycle, before the assertion ran.

Correct the Section C docstring, which claimed C2 provides defence-in-depth
coverage it cannot provide in CI. Whether to provision the CLI in a CI job
is a maintainer call and is left open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): skip hooks/dist and .dist-staging when building the fixture

The recursive copy added for #3613 races the hook builders. build-hooks.js
writes atomically through a per-PID hooks/.dist-staging-<pid> and removes
it when done, and nine test files invoke that script from their before()
hooks, so a recursive walk can enumerate a staging directory and then
lstat it after the owning process deleted it — the ENOENT that #3656 just
fixed in the cold-tree fixture.

Copy entry-by-entry and skip by NAME BEFORE anything stats it, reusing
shouldCopyHookEntry from tests/helpers/cold-runtime-lib-fixture.cjs rather
than re-deriving the rule. A filter applied after the stat would not close
it. The predicate is already pinned including its over-match cases, which
a loose startsWith('dist') would get wrong: dist-staging-no-dot and
distant.js must both be kept.

Excluding hooks/dist is independently right for this fixture — a real
marketplace install contains neither dist nor a transient staging dir,
the same reasoning that keeps the repo-root CLAUDE.md out of the
validated tree. C2 still validates clean without it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): validate agents/ too, and scope the hooks predicate

Three review Minors.

agents/ ships in package.json `files` and IS auto-validated by the CLI —
verified: a frontmatter-less agents/*.md exits 1. Including it was
pointless while the fixture symlinked, because the CLI read nothing
through a symlink; now that the tree is real it is the last shipped
component tree C2 could not see. Proven to buy coverage rather than
bytes: planting a frontmatter-less agent now reds C2, which it could
not do before this change.

shouldCopyHookEntry is documented as a hooks/ entry filter, so it now
runs only for hooks/. Applying it to the other trees was harmless today
but silently encoded a hooks-shaped exclusion into them — a future
commands/dist would have vanished from the validated tree with no signal.

assert.deepEqual -> deepStrictEqual on the nested-symlink assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): scope the fixture to the three directories the issue names

Review findings, all five.

Major 1 — `agents/` dropped from COMPONENT_DIRS. The observation behind
adding it was right (the CLI does auto-validate it; a frontmatter-less
agents/*.md exits 1), but it is a NEW gate the issue does not ask for, on
the largest of the trees, and since C2 never runs in CI it would be red
only on contributor machines with `claude` installed — the same
worst-of-both-states #3613 exists to remove. Worth having as its own
issue, where "should CI provision the CLI" gets answered for it too.

Major 2 — C3 no longer passes on a fixture that validates nothing.
buildValidationPluginRoot() mkdir's every component dir before the entry
loop, so a copy that stops happening leaves real, EMPTY directories and
all three structural assertions still hold (an empty tree contains no
symlinks). Adds a non-empty assertion plus a known entry per tree, so a
partial copy is caught as well. Stubbing the copy now reds C3 with
"commands/ is EMPTY"; dropping just commands/gsd reds it too. Neither
did before.

Minor 1 — the measured table was wrong, and the mistake was measuring an
inert file. Re-measured on CLI 2.1.239, reproducing the review's 2.1.237
result: a symlinked skills/<name>/SKILL.md at depth 2 exits 1, while a
stray symlinked *.md at the same depth exits 0. The boundary is not depth
at all, it is whether the symlink is a file the CLI reads as a component.
Table replaced with that.

Minor 2 — the hooks name filter is now local instead of importing
cold-runtime-lib-fixture.cjs's, whose docstring scopes it to the cold-tree
fixture and which had exactly one caller. Two consumers across fixtures
with different requirements and nothing asserting they stay compatible is
how a later cold-tree change silently alters what this fixture validates.

Nit 1 — cost figures re-measured: ~1.06 MB over 176 files, not 2.1 MB
over 243.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3613): correct the hooks row and the count the agents/ removal left behind

Three comment-prose items from review, no code change.

N1 — the corrected table gained a new wrong row, erring unsafe. "symlink
inside commands/ or hooks/ -> 0" is false for hooks/hooks.json, which is
the single most likely thing anyone would symlink there. Re-measured on
CLI 2.1.239, matching the review's 2.1.237 figures on every cell:

    symlink inside commands/ (a dir, or a component *.md)  0
    stray symlinked dir or *.js inside hooks/              0
    symlinked hooks/hooks.json                             1

hooks/hooks.json is now its own row, and is named in the C3 assertion
message alongside SKILL.md — that message is what the next engineer
actually reads.

N2 — "the four component directories" was residue from the revision that
also copied agents/; it survived the very edit that removed it, three
lines above a sentence saying three. Now three.

N3 — the promised agents/ follow-up is filed as #3751 and the comment
cites it, rather than promising an issue that did not exist. It frames
the real decision (should CI provision the claude CLI) rather than just
asking for the directory back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-22 08:06:52 -04:00
Tom Boucher
0062f6d033 fix(#2845): stop the dimension parity guard counting a back-reference (#3752)
next went red on dacae9273 against documentation that was correct.

parseDeclaredCounts read any '<numeral> ... dimensions' collocation as a claim
about the gsd-ui-checker dimension TOTAL. The how-to sentence 'Dimension 7 is a
rule gsd-ui-checker follows, the same as the other six dimensions' refers to the
other members of a seven-member set; the guard counted it as that document
declaring six, and reported a count-mismatch on prose that was accurate. A drift
guard that fires on correct prose is a false positive, which is how guards end
up switched off.

A negative lookbehind now excludes a numeral introduced by 'other' or
'remaining'. It is declared once as NOT_A_BACK_REFERENCE and shared by both
English scans rather than written at each site — the same two-surface divergence
class this suite exists to catch. Both scans apply it case-insensitively; the
first cut of this fix had the digit scan case-sensitive and the word scan not,
so a sentence-initial 'Other 6 dimensions' still slipped through. The exclusion
is word-anchored, so 'another six dimensions' — a real claim about a second set
— still counts.

parseDeclaredCounts takes an excludeBackReferences opt-out so a test can prove
against the REAL shipped how-to that the exclusion is load-bearing: with it off
the file reports [6,7], with it on [7]. That replaces a raw substring match on
prose, which the suite's own header forbids, with a typed before/after.

The docs prose is deliberately unchanged. It is the only instance of the pattern
in the tree, so keeping it means the real-tree assertion exercises the path this
fix exists for instead of asserting only on fixtures.

Regression tests cover both polarities: eight back-reference shapes including
all four sentence-initial cases, five real count claims that must still count,
the 'another' word-boundary case, and the shipped how-to itself.

Known limits recorded in the code: the exclusion is English-only, because
translated docs here are corrected to match English rather than authored, so
there is no instance to model the grammar on; and it cannot distinguish a
back-reference from a genuine total opening with the same word ('Other 6
dimensions were added'), where the false negative is the safer side of the trade.

Why it reached next at all: the docs PR (#3746) was green. A doc-only diff
inert-skips the test matrix in the PR lane, so the guard that reads docs never
ran against the docs change that broke it — it fired on push to next, after
merge.

Co-authored-by: sim <sim@local>
2026-08-22 06:39:48 -04:00
Tom Boucher
dacae92730 docs(#2845): record the inventory-provenance limits where readers meet them (#3746)
The limits shipped with #2845 were disclosed only in the PR body, which is
read once at merge and then buried. They are properties of what the feature
does, so they belong in the documentation.

Three surfaces, each at the point a reader forms an expectation:
docs/how-to/design-a-ui-phase.md gains a 'What this check is and is not'
subsection under the provenance how-to; docs/explanation/security-model.md
gains a residual-risk pair matching the section's existing shape; and
docs/AGENTS.md notes them where gsd-ui-checker's behavior is described.

The substance: a provenance line makes an inventory's origin falsifiable
rather than verified, since nothing re-runs the command or compares the
count; the rule is agent-applied like the other six dimensions, not a schema
check; and 'the checker never runs the recorded command' is an instruction
rather than a capability boundary, because the checker holds a Bash grant it
genuinely needs for the agent-skills bootstrap and tool grants here are not
command-scoped.

Co-authored-by: sim <sim@local>
2026-08-21 13:22:50 -04:00
Tom Boucher
4918c62d76 feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance

Binds two shared formats before either exists, so the suite is RED against
next: the gsd-ui-checker dimension roster (asserted independently on twelve
surfaces, eight English and four translated) and the provenance-line grammar
the UI-SPEC template emits and Dimension 7 consumes.

Every parity assertion is paired with a synthetic mutation case, so the guard's
failure branch executes rather than only reading a correct tree: limit-1 (a
surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a
label that drifts on one surface only, a non-contiguous roster, a duplicated
number, and a surface that stops declaring a count at all. A seeded fast-check
property renders the roster under formatting noise (CRLF, padding, interleaved
sections) and asserts the parse round-trips and is strictly sensitive to a
dropped heading.

Assertions are on parsed typed records, never raw substrings.

* docs: normalize design-a-ui-phase how-to to American English

House style for docs/ is American English (CLAUDE.md). This file carried
colour/initialisation/initialise/artefact throughout. Spelling only — no
content change; kept separate from the #2845 feature commit so the
release-notes classifier and the hotfix cherry-pick filter see it for what
it is.

* feat(#2845): require provenance for UI-SPEC component inventories

A UI-SPEC's component inventory was treated downstream as a closed allowlist
while the document recorded nothing about whether the list had been enumerated
from the installed design system or recalled from memory. A recalled inventory
is indistinguishable from an enumerated one, so an executor complying with the
spec builds against a fraction of what the package offers, and every gate stays
green because they assert semantics rather than composition.

The UI-SPEC template gains a Component Inventory slot carrying one of two
provenance lines: the command that enumerated the list, the count it returned,
the resolved package@version and the date; or a Could not enumerate record with
a real reason. gsd-ui-researcher gains an enumeration ladder and must record
the line rather than write the list from recall.

gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count
with no command, an empty could-not-enumerate reason, or a line still carrying
the template's unfilled placeholders BLOCKs; a partial line, a line placed below
its table, or an honest negative record FLAGs; a complete line passes, and so
does a spec carrying no inventory at all, which keeps every UI-SPEC predating
the dimension validating unchanged. Whatever the verdict, an unsourced inventory
is reported as a non-exhaustive list of known-good components rather than a
closed allowlist, so the executor is never blocked from a component the spec
merely failed to mention. The checker never runs the recorded command.

The dimension count moved on all thirteen surfaces that assert it, across five
languages. Also corrects the claim in the English, Korean and Portuguese how-tos
that this checker applies a scored six-pillar rubric — that rubric belongs to
/gsd-ui-review's retroactive audit.

* chore(#2845): backfill changeset pr number to 3745

---------

Co-authored-by: sim <sim@local>
2026-08-21 11:59:56 -04:00
Tom Boucher
2b42b28687 fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows

* fix(#3659): make baseref-head suppress mode-aware and thread isolation mode

* fix(#3659): review fixes - stale advice purge, message pins, mode alias

* fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin

* test(#3659): rewrite set-baseref pin, fix writeSync row stub

* chore(#3659): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 04:58:33 -04:00
Tom Boucher
1f76861202 fix(#3657): tolerate commonmark fence widths in ledger readers (#3733)
* test(#3657): fence-width tolerance regression rows

* test(#3657): fix pure-row fixtures to use appendWindow result shape

* fix(#3657): tolerate commonmark fence widths in ledger readers

* fix(#3657): restore throw-block indentation in parseJsonBlock

* chore(#3657): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:45:06 -04:00
Tom Boucher
72819a4616 fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi (#3731)
* test(#3031): failing-first coverage for opt-in ~/.kimi legacy reclaim

Drives the user-reachable installer surface against a sandbox HOME seeded
with the pre-#2755 wreckage: a GSD [[hooks]] block, hooks bundle and
CommonJS marker orphaned in ~/.kimi by a --kimi-code install.

Covers the reclaim itself plus the four guards the diagnosis identified as
negative space: opt-in only (no flag, no deletion), user-authored TOML and
hook files preserved, a --kimi install never reclaiming its own root, and
the KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory. Adds a fast-check property that stripping the block never
destroys user content.

Red until --reclaim-kimi-legacy exists.

Refs #3031

* fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi

A --kimi-code install older than 1.10.0 wrote its GSD [[hooks]] block, hook
bundle and CommonJS marker into Kimi CLI's ~/.kimi. #2755 fixed the
destination but could not reclaim what the old bug already wrote: the stale
block is byte-identical to a legitimate Kimi CLI one — both runtimes render
the same bytes for the same root, since the command paths derive from the
hooks root, not the runtime — so no inspection can tell litter from a working
install.

Cleanup is therefore opt-in. `--reclaim-kimi-legacy` on a --kimi-code install
removes GSD's own artifacts from the legacy root; without it nothing is
touched, so a dual-product machine keeps Kimi CLI's hooks and #2755's
acceptance criterion holds.

Extracts the uninstall path's removal sequence into reclaimKimiHooksRoot() and
drives both callers through it, so the reclaim removes precisely what a real
uninstall removes rather than a hand-copied second implementation. Guards the
wrong-runtime case (a --kimi install would delete its own hooks) and the
KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one
directory.

Also corrects two pre-#2755 leftovers in the same surface that told users to
run `--kimi --config-dir ~/.kimi-code` — the form that produces this very
defect, since --config-dir moves only the skills root — and adds the missing
--kimi-code entry to the installer's own help.

Regression coverage folded into tests/kimi-upgrades.test.cjs beside the #2755
cases, per the regression-test-naming lint.

Fixes #3031

* fix(#3031): never reclaim ~/.kimi when this run also installs kimi

Found by the isolated adversarial review pass and independently while tracing
--all ordering, then reproduced.

selectRuntimesFromArgs orders 'kimi' before 'kimi-code' in both --all and an
explicit --kimi --kimi-code, and installAllRuntimes installs in that order. So
--all --reclaim-kimi-legacy installed a fresh, legitimate Kimi CLI hooks block
into ~/.kimi and then deleted it moments later from the kimi-code leg — exiting
0 and reporting success while leaving the user with no Kimi CLI hooks at all.
The collision guard could not catch it: kimi-code's own root is ~/.kimi-code, a
genuinely different directory.

The flag asserts "I only use Kimi Code"; installing kimi in the same invocation
falsifies that, so the reclaim is skipped with a notice.

Also hardens the collision guard itself. It compared path.resolve strings,
which returns false for two spellings of ONE directory — measured, not assumed:
a symlinked alias and a case variant on a case-insensitive filesystem both
compared unequal, so the guard would not have fired and the install would have
deleted its own freshly-written hooks. isSameDirectory now compares directories
via resolve, then dev+ino identity, then realpath.

Regression tests for all three cases; the two alias tests probe the real
filesystem and t.skip() where the alias cannot exist.

Refs #3031

* docs(#3031): reattach reclaimKimiHooksRoot's JSDoc to its own function

Inserting isSameDirectory anchored on the function name, which placed the
helper between reclaimKimiHooksRoot's doc block and the function it documents.
isSameDirectory ended up with two stacked doc blocks above it and
reclaimKimiHooksRoot with none.

Refs #3031

* fix(#3031): warn when --reclaim-kimi-legacy cannot apply

The flag only acts inside the kimi-code GLOBAL install branch. Passed with any
other runtime, or with --local, it was consumed in silence: exit 0, no cleanup,
no message. For a cleanup the user explicitly asked for, silence is
indistinguishable from "it ran and found nothing".

The scope warning is raised at argument-resolution time rather than inside
install(). kimi-code declares hostBehaviors.localInstallDeferred, so install()
returns early at the deferral check long before the kimi-hooks-toml branch — a
guard placed there is unreachable, which is both dead code and a linted drift
shape in this repo. Verified reachable by spawning the real installer.

Neither case is a hard error: the flag stays composable with --all, where it is
legitimately inert for the other seventeen runtimes.

Refs #3031

* docs(#3031): document every case where --reclaim-kimi-legacy skips

Refs #3031

* fix(#3031): resolve local config dirs from RUNTIME_META alone in the install harness

The remote runner surfaced this: the #3031 warning test drives a local
kimi-code install and died with "The path argument must be of type string.
Received undefined".

runMinimalInstall carried a SECOND, hand-maintained local-dir map beside
RUNTIME_META, and it had drifted — four runtimes present in RUNTIME_META
(hermes, kimi, kimi-code, zcode) were missing from it, so scope:'local' for any
of them resolved path.join(root, undefined) and threw a bare TypeError naming
neither the runtime nor the map at fault. #3023 had already hit exactly this
for pi and fixed it by adding one more entry, which left the divergence itself
in place for the next runtime to rediscover.

Local scope now reads RUNTIME_META.localDir, the same table the global branch
already reads, with the same loud named error the global branch raises. Parity
verified for all 14 previously-supported runtimes: every one resolves to a
byte-identical configDir. cline keeps its ternary — its local artifacts land at
the project root itself, which is a real exception, not a directory name.

Guarded in golden-parity-single-source.test.cjs beside the buildParityManifest
anti-divergence test, and both arms of that guard were proven able to fail.

Refs #3031

* chore(#3031): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 03:05:58 -04:00
Tom Boucher
65d2839b11 fix(#3651): prescribe only writes config-set accepts in integrations flow (#3732)
* test(#3651): regression rows for workflow config-write prescriptions

* fix(#3651): prescribe only writes config-set accepts in integrations flow

* fix(#3651): review fixes - single-source lane list, configSchema-derived test set

* fix(#3651): canonical cite, review wording fixes, one-element array pin

* chore(#3651): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 02:20:53 -04:00
Tom Boucher
94bc492f57 fix(#3645): tracked-source rule for planner/pattern-mapper path resolution (#3728)
* test(#3645): failing-first agent tracked-source contract rows

* fix(#3645): tracked-source rule for planner and pattern-mapper paths

* Revert "fix(#3645): tracked-source rule for planner and pattern-mapper paths"

This reverts commit 61f05e947bbdaf3b4897240c3819d349215744fb.

* fix(#3645): tracked-source rule at the spawn seam and mapper gate

* fix(#3645): review fixes - bounded block, ack merge assertion, git wording

* fix(#3645): fit the tracked-source block under the 1168 ceiling

* chore(#3645): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 23:24:42 -04:00
Tom Boucher
95f7c14413 fix(#3642): stop the single-section total_phases leak into an absent milestone (#3727)
* test(#3642): failing-first single-section leak rows

* fix(#3642): gate the unbounded total on any-milestone-section, not >=2

* test(#3642): rewrite the 3185 wrapper row to the withhold contract

* docs(#3642): glossary amendment for the >=1 sibling; changeset

* chore(#3642): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 22:10:52 -04:00
Tom Boucher
072b97d276 fix(#3641): make v005/v004 see bracket-convention phase entries (#3723)
* test(#3641): failing-first bracket-window validate rows

* fix(#3641): thread phase convention into hasphaseentries for v004/v005

* test(#3641): review rows - digit-anchor, decoy, probe parity, t.after

* fix(#3641): digit-anchor bracket entry token; thread probe scope axis

* fix(#3641): align frontmatter bound with probe; changeset

* chore(#3641): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 17:48:50 -04:00
Tom Boucher
03a3f779ce fix(#3640): isolate the drift cli e2e fixture via a --root override (#3722)
* test(#3640): failing-first --root isolation rows for the drift cli e2e

* fix(#3640): add --root override to the drift cli scan root

* fix(#3640): harden --root validation; typed exit-2 usage rows

---------

Co-authored-by: sim <sim@local>
2026-08-20 16:15:22 -04:00
Tom Boucher
9a69a86f42 enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter

Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe
for /gsd-pr-branch across six layers: pure classification and forbidden-path
predicates, real-git fixtures that run the cherry-pick filter loop end to end,
config-key registration through the real CLI and both manifests, the executed
worktree-materialization claim the issue's triage asked to establish, fast-check
properties over arbitrary path sets, and a drift guard over the shipped workflow.

Two live defects in today's shipped recipe are pinned as regressions, both
reproduced empirically first: `git rm -r --cached` stages a deletion of any
.planning/ path the target branch already tracks, so the generated PR removes the
base branch's planning files; and the same command leaves the cherry-picked file
untracked on disk, so a second commit touching that path aborts the pick with
"untracked working tree files would be overwritten" and every remaining commit is
silently dropped.

The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md
rather than restating them, so the workflow stays the single source of truth and the
suite cannot drift from what ships.

Refs #2971

* feat(#2971): strict planning filter mode for /gsd-pr-branch

Adds planning.pr_strict — a boolean, default false, that selects what
/gsd-pr-branch means by "filtered". Default mode is unchanged: structural
planning state survives into the PR branch and the nine transient
subdirectories do not. Strict mode drops every .planning/ path, structural
files included, and carries a commit over only when it touches at least one
file outside .planning/.

Strict mode is what makes planning.commit_docs: true safe for a project that
versions its planning tree locally but publishes none of it. The alternative
posture, commit_docs: false, silently costs parallel executor isolation — a
worktree is checked out from a commit, so an untracked or ignored .planning/
is simply absent inside it and the executor has no PLAN.md to read. That claim
is now established by an executed fixture rather than inherited.

The two path lists are declared once and both projections derived from them,
so create_pr_branch and verify can no longer disagree about what the filter
promised. verify previously counted every .planning/ path against a documented
success criterion of zero while create_pr_branch was specified to preserve five
structural files, so a correct run reported itself as failed on every phase
that touched STATE.md — which is every phase. It now asserts against the active
mode, and names the .planning/ paths default mode deliberately keeps rather
than trading a wrong signal for silence.

Two verified defects in the same recipe are fixed alongside, because strict
mode would have amplified both. `git rm -r --cached` staged a deletion for any
.planning/ path the target branch already tracked, so the generated PR removed
the base branch's planning files — under strict mode that would have been the
entire tree. The same command left the picked file untracked on disk, so a
second commit touching that path aborted the cherry-pick with "untracked
working tree files would be overwritten" and every remaining commit was
silently dropped. Both were reproduced against real git before being fixed.
The filter now forces excluded paths back to what the PR branch's HEAD carries,
in the index and the working tree; a conflict outside the filter halts instead
of being improvised past; a commit left empty by filtering is skipped rather
than failing. A clean-working-tree precondition makes the worktree half safe.

Closes #2971

* fix(#2971): unwind the checkout on a conflict halt, and test the real recipe

Two review findings, both fixed in place.

The isolated adversarial pass found that the conflict-outside-the-filter branch
exited while leaving the user checked out on the half-built PR branch with
cherry-pick state still live — this loop runs in the user's own working
directory, so stranding them there is a real cost even though it is not a
vulnerability. The branch now aborts the pick, returns to the original branch,
removes the partial PR branch, and says so before exiting.

The standards pass found the L2 fixtures executed a hand-written mirror of the
cherry-pick filter recipe rather than the recipe itself, so a reordering in the
workflow would not have been caught — and the order is load-bearing, since
restoring a path from HEAD before removing it inverts the filter. The helper now
extracts the canonical loop from the shipped workflow and the fixtures execute
that verbatim, which also gives the conflict-halt unwind above real coverage.
The drift guard additionally pins the two commands' relative order and asserts
the workflow carries exactly one canonical loop.

Also records the publication gate in the CONTEXT.md glossary next to the commit
gate it is distinct from.

Refs #2971

* fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form

The remote matrix caught two defects in the previous commit.

The halt path claimed to restore the original branch but did not. `git
cherry-pick --abort` does not apply to a single `--no-commit` pick with no
sequencer file, and the fallback left the unmerged index in place, which makes
`git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so
the user was told they had been restored while still sitting on the half-built
PR branch. The unwind now drops sequencer state, hard-resets the disposable PR
branch to clear the unmerged index, and only claims a restore when the checkout
actually succeeded; when it does not, it says where the user is and gives them
the two commands to finish it by hand. Verified against real git: exit 1, the
conflict named, HEAD back on the original branch, the partial branch gone, a
clean tree and no CHERRY_PICK_HEAD.

Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form,
which names a command no runtime registers. The canonical authoring token for
workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the
documentation added in this branch is unaffected. The comment in src/config.cts
moves to the colon form too, since it propagates into the generated lib.

Refs #2971

* docs(#2971): backfill PR number into the changeset fragments (#3720)

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:40 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8df5cb36c2 enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)
* test(#2951): pin the absent-evidence provenance contract (failing first)

17 tests / 22 anchors on the deployed agent text. Measured against the parent
commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity
guards on the package-name and in-repo-value rules -- green before and after is
their intended signature.

Refs #2951

* enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata

A claim of the form "X does not support Y" drawn from MISSING metadata -- no
python_requires, no engines field, no per-version classifier, no changelog entry,
no matching support-matrix row -- no longer earns [VERIFIED] however
authoritative the source consulted. An absence is silence about every value, so
the same evidence would "prove" both the version being ruled out and the version
being standardized on. The only route from an absence to [VERIFIED] is a positive
falsification attempt with its failing output pasted; everything short of that is
[ASSUMED], which the file already routes to "needs user confirmation before
becoming a locked decision".

Third member of the family beside the package-name and in-repo-value provenance
rules, mirroring PR #2768's shape. A present declared constraint and an
affirmatively documented incompatibility are untouched.

Closes #2951

* fix(#2951): close the allow-list ambiguity and the mutation gap review found

Findings from the isolated adversarial pass and the two-axis review, all fixed:

MAJOR (x2, one root cause) -- the absence clause and the present-constraint
carve-out gave opposite verdicts on the same evidence for the commonest real
case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A
researcher could read the enumerated list as a "declared" positive constraint
and re-earn [VERIFIED], which is also the evasion vector. The rule now states
the decision procedure -- does the declaration bound EVERY value or only the
ones it names -- and closes the positive-reframing restatement explicitly. New
contract test pins all four clauses.

MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]'
asserted two independent substrings and never that the route lands on
[VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the
rule and survived all 17 tests. Now pinned as one joined sentence.

MINOR -- the attributable-failure test regex-matched illustrative examples
("a missing certificate, a wrong host"), so a copy-edit would break it for no
reason; relaxed to the substantive clause. The no-paraphrase guard counted only
the heading, missing the drift mode in its own name; it now also pins the core
proposition to one occurrence, and the test name matches what it checks. An
off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed.

MINOR -- docs/AGENTS.md listed four of the five governed absence forms while
the agent prose, docs/COMMANDS.md and the changeset listed five; three copies
disagreeing on list membership is the drift this repo treats as a defect.

SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its
docs/README.md index line. Both reviewers flagged them as a seventh and eighth
surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill
change" row requires only docs/AGENTS.md. The actionable four-case guidance is
retained in docs/COMMANDS.md, which is inside the approved scope.

Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550
bytes headroom under the LARGE cap of 49152.

Refs #2951

* docs(#2951): restore the how-to the phase gate requires

Reverses the removal in 6404b43d3. Both /code-review axes had flagged
docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md
index line as a seventh and eighth surface beyond the six the requester capped,
and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names
only docs/AGENTS.md, so they were dropped.

gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded
enablement sequence has more than one step and the how-to quadrant is empty.
The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED]
claim, probe or cite or accept it unlocked, then answer discuss-phase's
checkpoint -- and the last step lands on a different capability's surface, so a
reference table cannot carry it. Compressing the sequence to one step to unlock
howToSkipReason would be gaming the gate, which is the same Goodhart failure
this whole change exists to close.

A machine-enforced repo gate outranks two reviewers' scope preference and my own
reading, so the page is restored and the PR body discloses the two extra
surfaces instead of hiding them. Reverting is a one-file change if a maintainer
prefers the tighter scope.

Refs #2951

* chore(#2951): backfill the changeset PR number

pr: 0 -> 3718 now that the real PR exists. The placeholder fails
scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks
lint-docs-required from consuming the fragment.

Refs #2951

---------

Co-authored-by: sim <sim@local>
2026-08-20 14:35:14 -04:00
Tom Boucher
8da2dd3ad2 feat(#2790): add read-only planning.inspect schema-v1 snapshot query (#3708)
* feat(#2790): add read-only planning.inspect schema-v1 snapshot query

Adds a read-only query emitting a schema-versioned JSON projection of .planning/
so downstream harness UIs can consume planning state without parsing GSD's
Markdown a second time.

Composed strictly from the ADR-3180 section 7 owners plus parsePlanDocument,
parseRequirements and parseUatItems; markdown structure is read through the
Markdown Sectionizer and Markdown Table Model seams. It declares its own flat
external schema rather than serializing PlanningSnapshot, which is the
diagnostic-rule subject and still growing.

Extracts plan-document parsing out of cmdPhasePlanIndex into a shared leaf
module so phase.plan-index and planning.inspect cannot drift, including the
plan-id derivation both surfaces report.

Also fixes parseRequirements dropping the separator delimiter used by the
shipped requirements template, surfaced while wiring the requirement rows.

* fix(#2790): close spec gaps and a raw-text test assertion found in review

Review findings from the standards, spec and security passes:

- phases[] rows carry goal and dependencies, the two per-phase elements the
  issue Summary names that had no corresponding field. Goal is bounded to the
  section's leading prose so the Depends-on line, the Plans checklist and the
  wave annotations are not duplicated into it.
- requirement rows carry their own diagnostic codes, so a consumer no longer
  has to string-parse the global diagnostics subject to correlate.
- roadmap_acceptance.checkbox is looked up through the phase-id key owners.
  It was compared raw against the on-disk directory name, so it read null for
  every real-world slugged phase directory and the evidence channel was inert.
- the hostile-input test asserts the structured payload instead of matching the
  raw stdout string. The absence proof over raw stdout is kept deliberately.

* fix(#2790): register planning in the runtime usage list and repair fixtures

Remote runner reported 9 failures on 9b3f9aa. Two root causes, both fixed:

- gsd-tools.cjs registered the planning family in HOST_COMMAND_ROUTERS but
  never added it to TOP_LEVEL_USAGE's Commands list. Those are two surfaces a
  parity test guards, and the top-of-file block comment is not the runtime
  help string. A real wiring gap that every local gate and three review passes
  missed.

- the new suite's fixtures could not produce a resolvable phase set. STATE.md
  frontmatter omitted the milestone field, which ADR-3180 7.2 rule 1 makes the
  primary milestone selector, so the phase set scoped unscoped and every
  percentage was correctly withheld. Separately declarePhase returned a path
  without creating the directory, so a phase declared but never written to left
  phases empty. Both reproduced against the built module before fixing.

No assertion was weakened. The withholding path is still exercised and still
returns null when the roadmap is absent.

* chore(#2790): backfill changeset pr number

* test(#2790): cover every enumerated matrix row and contain a symlink escape

Reverses a silent deferral. An earlier revision left 23 of the 78 enumerated
matrix rows unimplemented and 7 more as one-off manual checks, with a paragraph
in the artifact and the PR body describing the gap. CLAUDE.md is explicit that
such a note is not a fix and is not surfacing. The rows are implemented instead
and the manual-evidence bucket is gone: 49 test cases become 88, covering all 78.

Writing the symlink row proved a real leak: a *-PLAN.md symlinked outside
.planning/ had its content emitted into the payload, confirmed via a direct call
and the spawned CLI. readDocument now resolves target and planning root with
realpathSync and rejects an escape, returning the ordinary unreadable-document
shape. Tested both ways, because a containment check that over-rejects is its own
defect: an escaping symlink leaks nothing and degrades that plan alone, while a
legitimately relocated .planning/ symlink stays fully readable.

The three new modules are registered in the mutation COVERED registry, which had
been reporting has_work false and skipping the Stryker gate entirely. Provisional
non-binding floors so the shards run and report; raised to the measured value
before merge, since the registry forbids calibrating from a local run.

* fix(#2790): satisfy the mutation ratchet contract and scope the 1MB test

Remote runner reported 16 failures on 8c451ed. Two causes.

The COVERED registry has a paired contract the earlier commit violated: every
module needs a matching RATCHET_BASELINE entry, and minScore must be between 50
and 100 with minScore === baseline. The provisional floor of 1 was illegal on
both counts. All three modules now sit at 50 — the registry's own enforced
minimum — with matching baselines. The score cannot be measured locally: the
shard runs node --test, which this repo hard-blocks, so CI is the only source.
Floors are raised to the measured value once this PR's shards report; a shard
below 50 means the tests need strengthening, since the floor cannot go lower.

The 1MB test was measuring the test harness rather than the product. The command
handles the oversized payload correctly by spilling to a tmpfile and resolving it
back, but the resolved stdout then exceeds runGsdTools' maxBuffer and the helper
reports ENOBUFS. It now uses --pick so stdout stays one byte while the full 1MB
document is still read and parsed end to end.

* fix(#2790): wire containment across every document read this command drives

An isolated security review of the containment control found the boundary logic
sound but not comprehensively wired: two content reads reached the filesystem
without it.

An escaped phase DIRECTORY could enumerate external filenames into the file
fields and diagnostic subjects. Both enumeration sites now containment-check the
directory before reading. Worth recording that the leak was already prevented one
layer earlier than the review claimed: Dirent#isDirectory() reports false for a
directory symlink, so such a directory never becomes a phase row at all. The
guard is defense-in-depth for a direct caller and for platforms where a reparse
point reports as a directory.

A *-VERIFICATION.md symlinked outside the root leaked one frontmatter value
verbatim, because readVerificationStatus does its own read and copies an
unrecognized status into the payload's next_action. Closed from the consumer
side through that function's existing fs injection seam, so src/verification.cts
keeps its signature and its other callers are untouched.

The reviewer additionally rated a forged status: passed as an integrity bypass.
It is not: anyone able to plant the symlink can plant a real VERIFICATION.md
saying the same thing. The incremental risk is confidentiality, which is what
these fixes close.

src/plan-scan.cts is deliberately unchanged: isPlanSuperseded reads
symlink-followed content but yields only a derived boolean, no document text.

* test(#2790): give the mutation shards an in-process surface

Two Stryker shards were CANCELLED at the 15-minute cap, not failed on score.
CI log: 640 mutants instrumented, and the dry run reported 'Ran 1 tests in 20
seconds' because the shards pointed at the integration suite, where nearly every
case spawns a gsd-tools subprocess and Stryker's command runner treats the whole
test-runner invocation as a single test. 640 x 20s cannot finish in 15 minutes;
at the kill it was 27/640 with an ETA over an hour.

Every other COVERED module points at a property or unit file, and the workflow's
own paths filter lists exactly those two patterns. In-process is the intended
mutation surface; the shards were pointed at the wrong shape of test.

Adds tests/planning-inspect.unit.test.cjs — 39 cases in 10 describes that spawn
nothing and call the built modules directly. plan-document and the router need no
filesystem at all, one being a pure content-to-object parser and the other taking
an injected mock. The three shards now point here. The 91-case integration suite
is untouched and still runs in the normal test job.

* chore(#2790): ratchet mutation floors to the measured CI scores

CI run 32392791843 measured all three shards, which is the only source the
registry accepts — local runs count timeouts as kills and inflate badly.

  planning-command-router  95.65 -> floor 94
  plan-document            76.58 -> floor 75
  planning-inspect         57.03 -> floor 56

Applied the registry's own rule, floor(score) - 1, and updated RATCHET_BASELINE
to match, since the ratchet test enforces equality.

planning-inspect sits well below the file's target of 80 and is the obvious
ratchet candidate as its tests improve. planning-command-router already exceeds
the target. The placeholder comment about floors pending measurement is removed
rather than left standing as a false statement.

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:42:43 -04:00
Tom Boucher
77fa08f1e8 fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input

Locks the Step 5.5 contract that a response_language project must feed the
edge probe an English translation of each requirement's text, and binds that
advice to measured engine behavior: the same requirement classifies to zero
shapes in Portuguese and to collection/adjacency/empty/ordering in English.

Also pins the honest limit — the issue's own repro sentence classifies to []
in English too, so translation is necessary but not sufficient and the
authored shapes override is the documented fallback.

Red before the doc change; the assertions are all false today.

Refs #2773

* fix(#2773): feed the spec-phase edge probe English-translated requirement text

The shape cues in src/edge-probe.cts are English word-boundary regexes, so a
project running with response_language set wrote its SPEC requirements into the
Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero
shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy
contributed nothing and --auto left it all unresolved — the probe was a silent
no-op for exactly the spec type it exists to harden.

Step 5.5 now states that the $REQS_JSON payload is engine input rather than
user-facing output, so the response_language rule does not govern it: each
requirement's text carries a faithful English translation, the SPEC keeps its
original language, and requirement ids are never translated or renumbered. The
instruction sits before the heredoc on purpose — the downstream APPLICABLE=0
warning fires only when every requirement is unclassified, so a partly-classified
non-English spec would otherwise slip through with no signal at all.

Measured against the compiled engine: the same requirement returns [] in
Portuguese and collection -> adjacency/empty/ordering in English. Also measured:
the issue's own repro sentence returns [] in English too, so translation is
necessary but not sufficient — the instruction therefore points at the authored
shapes override for prose carrying no cue in any language rather than promising
that translation restores classification.

Doc scope only, per the triage disposition on the issue. The compiled engine is
untouched; the lang-hint / per-language cue-set fix is a separate follow-up.

Closes #2773

* fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path

Surfaced by the isolated security review of this branch. Between the mktemp and
the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about
their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before
exiting, while the empty/placeholder guard directly above it exited without one.
A spec run that tripped the placeholder check therefore stranded a temp file
holding the SPEC's requirement text in TMPDIR, once per failed run.

The added contract test walks the region between the mktemp and the
unconditional cleanup and asserts no exit path leaves the file behind, so the
two guards can no longer drift apart. Proven to bind: run against the pre-fix
file the walker reports the leaking exit; against the fixed file it reports none.

Refs #2773

* docs(#2773): record the edge probe's English-cue input constraint in the predicate store

The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and
docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real
gaps rather than incidental coupling.

CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for
classifyShape but did not record that SHAPE_CUES are English word-boundary
patterns — so the predicate store implied text was language-agnostic, which is
what a future agent reads before touching this seam.

docs/CONFIGURATION.md's response_language row is what a non-English project
reads when it turns the setting on; it now names the one deliberate exception
and links to the FEATURES.md explanation, so the interaction is discoverable
from the config key rather than only from the workflow.

CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack
fragment is updated for the final byte range and now also records the
placeholder-guard cleanup fix folded into the same block.

Refs #2773

* fix(#2773): append the growth rationale to the existing spec-phase.md ack entry

The remote runner caught this: emitted-attribution.test.cjs pins the
0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration
regression test asserts the exact '31987 -> 31997' delta text survives), so
removing it to avoid a duplicate-key collision with a new fragment broke that
test instead of satisfying the ratchet.

The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102
were each appended to the same reason string by later PRs, which is how a shared
growth key coexists with the rule that two ack sources may never name the same
path. This appends the #2773 rationale the same way and drops the separate
fragment, whose spec-phase.md key was the collision.

Verified locally by reproducing both affected tests against the real fragment
before re-dispatching: the pinned delta survives, grown[0].acked is true,
staleAcks is empty, and all 35 entries still read as spent.

Refs #2773

* docs(#2773): add a how-to for probing edges in a non-English project

The phase gate's enablementSequence check caught a wrong call of mine. I had
recorded that no how-to was owed because the user takes zero extra steps — the
workflow translates the probe input itself. Written out, though, the sequence
from off to value is two steps and step 1 depends on response_language, a
setting owned by a different capability than the edge probe, which is exactly
the condition the how-to test names.

There is also real task content a reference table cannot carry: the three-way
split between a few unclassified rows (the classifier's recall gap), every row
unclassified (the probe could not read the spec at all), and the silent
partly-classified case where the APPLICABLE=0 warning never fires. That last
one is what a user would otherwise misread as a clean bill of health.

Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard
siblings and indexed from docs/README.md next to its closest relative.

Refs #2773

* chore(#2773): backfill the changeset PR number

pr:0 placeholder replaced with the real PR number now that #3713 exists.

Refs #2773

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:36:00 -04:00
Tom Boucher
adb46cdd85 feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker

Binds the contract before any hook change exists: a `state ~N commits back`
segment gated on the state_head stamp landed by #2622, firing at the same
advisory threshold /gsd-health's W024 uses rather than at > 0.

Covers all five acceptance criteria — threshold parity (19/20/21 boundaries),
both renderers including formatGsdStateCompact, an exact spawn-count assertion,
repo-pinning and sub_repos degradation, and behavioral parity against
readStateHeadFreshness rather than a source-grep of the two fence copies.

52 example-based tests plus 5 seeded fast-check properties. Red now by design.

* feat(#2734): surface STATE.md commit-age on the statusline

Adds an opt-in `state ~N commits back` marker to the GSD-state segment,
consuming the `state_head` stamp and freshness contract landed by #2622.
A solo developer returning to a project reads "Phase 4, executing" in
STATE.md and acts on it, without noticing the codebase moved 40 commits
since that line was written. /gsd-health reports it as W024, but only if
you think to run it; the statusline is the surface you see without asking.

Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses,
not at > 0: with commit_docs:true the commit carrying a STATE.md sync
advances HEAD by one, so > 0 would alarm permanently on a fresh project.

Costs exactly one bounded git subprocess per render and none when
disabled. `rev-list --left-right --count` answers ancestry and distance
together, and repo pinning is a filesystem check mirroring
projectOwnsItsRepo rather than a --show-toplevel compare, which is
unreliable on macOS /private/var and Windows 8.3 paths.

Every unresolvable input degrades to the tri-state unknown -- the marker
is absent, never a "fresh" claim the project cannot substantiate: a
malformed stamp, a root that does not own its .git, a sub_repos
workspace, history rewound past the stamp, or git being unavailable.

Also collapses statusline config resolution onto one resolveStatuslineOptions()
seam. runStatusline() and renderStatusline() duplicated it byte-for-byte;
one copy is what keeps a newly-added key from reaching only one of them.

* test(#2734): route the e2e spawn through the process seam and fix fixture leaks

Review findings from the two orthogonal passes:

- `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched
  its stdout to test a pure function. It now calls resolveStatuslineOptions()
  directly — no subprocess, no text matching.
- `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate
  lives in runStatusline, which reads stdin), so it now spawns through
  tests/helpers/process-seam.cjs and proves the negative with a filesystem fact:
  the git shim appends to a marker file on every invocation, and the assertion is
  that the marker never appears. Stronger than asserting text is missing, and it
  drops the last stdout substring match in the block.
- Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead
  of a trailing cleanup(dir), which leaked the temp repo on assertion failure.
  derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it
  binds each directory at scheduling time rather than cleaning only the last.

Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong
expectation rather than finding a code defect: `percent` drives the progress bar
too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends
after it, which is what the test exists to prove.

CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well
as the new key; both are now enumerated.

* docs(#2734): backfill changeset PR number (#3700)

---------

Co-authored-by: sim <sim@local>
2026-08-20 00:35:01 -04:00
Tom Boucher
2fca0e17e4 enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides

Binds the not-yet-built code-review-depth module: segment-aware path-prefix
matching of a changed-file set against ordered {paths,depth} rules, resolution
order flag > strongest matching rule > global > standard, typed validation
errors, and the large-scope downgrade boundary. Also proves behaviorally that
workflow.code_review_depth_overrides is not yet a registered config key.

Refs #2554

* feat(#2554): resolve code review depth from path-scoped override rules

Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth}
rules matched against a review's changed-file set by segment-aware path-prefix
comparison. Resolution order is --depth= flag, then the strongest matching rule,
then workflow.code_review_depth, then standard; a matching rule replaces the
global rather than being max'd with it, so quick and standard rules stay
meaningful. Glob metacharacters are a hard configuration error rather than sugar
for a prefix, and malformed rules halt the review instead of degrading to
standard. The resolver is pure and reports its own provenance, so the workflow
can print the resolved depth and the rule that matched. The pre-existing
>50-file deep-to-standard downgrade moves into the module and now names the rule
it overrode.

The key is registered centrally rather than as a capability config slice: the
federated slice channel admits only boolean/string/number/enum, so an array
slice would be dropped as malformed.

Closes #2554

* test(#2554): correct depth-provenance assertions and pin out-of-repo paths

Two corrections to the failing-first suite. The source assertion for a
non-matching rule with no global configured expected 'config'; with no global
set the depth comes from the default, and a companion assertion tolerated
either value, so both passed against an implementation that derived provenance
from whether any rules existed rather than from where the depth came from.

The out-of-repo absolute-path case used a home-directory path that matched
neither implementation, so it never exercised the defect it named. It now pins
the discriminating cases: an absolute path outside the repo root must not match
a repo-relative rule, and one under the root must.

* docs(#2554): document path-scoped code review depth overrides

Reference rows for workflow.code_review_depth_overrides in the configuration,
features and commands references plus the locale copies that carry those tables,
and in the planning-config reference. Explanation of why escalation is
whole-review rather than per-file and why v1 is prefix-only. New how-to for
scoping review depth by path, carrying the configuration-error reason table and
the distinction between nothing to report and could not look. CONTEXT.md
glossary entry and the INVENTORY row for the new CLI module.

ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and
pt-BR FEATURES.md carry no code-review config table, so those files are
deliberately untouched.

* fix(#2554): make the depth-misconfiguration halt executable and reject control chars

Three review findings, all in this change.

The misconfiguration halt was prose rather than shell: the error-printing fence
was followed by an unconditional extraction fence, so an ok:false result threw
and left the depth empty instead of stopping the review. Prose is not a guard —
the two fences are now one block with a real conditional, and anything that is
not the literal string true fails closed.

An interior control character in a rule path survived validation and reached the
provenance string and the summary box; rule paths now reject control characters
via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is
unchanged. That in turn makes the field record safe to delimit, so the seven
node invocations that each re-parsed the same result to read one field collapse
to one.

Also corrects the glossary entry's illustrative paths, which the glossary-ref
check read as real repository references.

* fix(#2554): use the fast-check v4 string API and acknowledge workflow growth

Two failures from the remote matrix on d3111f45, both this branch's.

The property block built its segment arbitrary with fc.stringOf, removed in
fast-check v4. Because the arbitrary is constructed in the describe body, the
throw took out all four property tests rather than one — they had never
executed. Rewritten to fc.string({unit, ...}), the form this repo already uses
in emitted-attribution.test.cjs. Every other fast-check helper in the file was
audited against the installed module.

The emitted-attribution growth arm needed an acknowledgment for code-review.md,
which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is
spent — its ripple was absorbed when #3503 merged, and the base file is exactly
the 34435-byte baseline this growth is measured against — so it cannot clear
anything, while the ack lint hard-fails on a duplicate key across two sources.
Removed it in favor of the new fragment, which is exactly how #3503 itself
replaced the spent 3191 fragment.

* docs(#2554): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-19 22:45:35 -04:00
Tom Boucher
71e00d426e fix(#3639): dir-aware sentinel recognition for the disk-side guards (#3698)
* test(#3639): pin bracket sentinel recognition in disk-side guards

* fix(#3639): dir-aware sentinel recognition for the disk-side guards

* chore(#3639): add changeset

* fix(#3639): disclose the digit-continuation residual, join phases-clear, load-bearing over-suppression guard

* test(#3639): match the token form W007 reports

* chore(#3639): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 21:55:18 -04:00
Tom Boucher
66228a89cf fix(#3637): carry the full executor contract in the orchestrator-worktree spawn (#3694)
* test(#3637): pin the executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): carry the full executor contract in the orchestrator-worktree spawn prompt

* fix(#3637): role-definition embed, embed-performance gates, drop stale ack

* chore(#3637): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 20:50:12 -04:00
Tom Boucher
7fc1561806 fix(#3611): decode entity-escaped ampersands and split shell segments quote-aware (#3693)
* test(#3611): pin entity-escaped ampersand chains in the negative-grep gate

* fix(#3611): decode entity-escaped ampersands before the negative-grep gate scans

* chore(#3611): add changeset

* test(#3611): pin entity chains in the 968 detector and quote-aware splits

* fix(#3611): quote-aware segment split + entity decode in both plan gates

* chore(#3611): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 19:53:45 -04:00
Tom Boucher
bad1f045b1 fix(#3610): hoist surviving top-level codex config keys to file scope on merge (#3690)
* test(#3610): pin top-level key hoisting above the codex managed block

* fix(#3610): hoist surviving top-level keys above the codex managed block

* chore(#3610): add changeset

* fix(#3610): hoist to file scope (before the first table header) with reviewer-driven coverage

* chore(#3610): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 18:38:13 -04:00
Tom Boucher
8526bd46f8 enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes

The guard matched only `resolve-execution ... --effort\s`, but the CLI also
accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written
with the equals form was a live invocation-override caller the guard passed
silently, along with `--effort` at end-of-input.

Lift the matcher to a shared predicate and assert it directly against every
shape the CLI accepts, plus the decoys it must not fire on (--effortless, a
bare --effort with no resolve-execution, item 6's --attempt caller, a call and
flag split across lines).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR

ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator
invocation override needed a caller in shipped orchestration. Per the
maintainer's ruling, take unblock path (b) for item 1 only -- record that the
override is an operator-facing CLI surface, deliberately not driven by shipped
orchestration, and ratify.

The ADR's own path (b) wording is not adopted verbatim: it says the scope is
limited to static install-time propagation, which is false on both counts --
item 6 has a live caller (#2296) and #2481 delivered a live invocation-time
argv channel. Only one precedence step is narrowed.

No consumer was invented to clear the gate: nobody has asked for a per-run
effort override, and #2475's actual complaint is already closed by the
cascade-to-argv path. The amendment states explicitly that --effort remains
supported and is not deprecated, so the scoping is not read as dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run`

There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run
and gsd-mcp-server. Both sites presented a command that cannot run as written.

One is in this branch's own new ADR-443 amendment; the other is a pre-existing
error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than
deferred. No translated copy carries either line, so no i18n drift is created.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2475): record the guard/CLI divergence risk on the item-1 matcher

The predicate independently models gsd-tools.cjs's argument parser rather than
sharing a constant with it, so a third --effort spelling would leave the guard
reporting green while ADR-443's ratifying invariant silently stopped holding.
Name that risk where the next editor will meet it.

Also restores the bounded-prose rationale that was attached to the eslint
directive removed in 39793079c -- the directive went unused once the regex moved
to a const, but the reasoning it carried is still worth having.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 17:56:04 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
Tom Boucher
79781e68eb enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands

Adds a deterministic resolvability probe over each PLAN.md <automated> verify
command and surfaces the nearest prior phase's proven commands to the planner
at every context window.

- src/verify-command-grounding.cts: recognizer (not a shell interpreter) that
  grounds a leading cd <literal> chain and npm --prefix <literal>, and reports
  unresolvable rather than guessing. Never executes command text.
- gsd-tools check verify-command-paths <N>: per-phase probe, wired into
  plan-phase.md before the plan-check pass.
- init.plan-phase gains prior_verify_commands, ungated by context_window.
- gsd-plan-checker: new Verify Command Path Resolvability dimension that
  reports the failing target and never prescribes a replacement.

Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs
(readdir order is not stable across platforms, so a module whose name extends
another's with a hyphen bucketed differently on Linux than on macOS).

Closes #2401

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets

Independent review found three defects in the recognizer:

- npm --prefix DIR run SCRIPT never reached the script-existence check,
  because the pattern required npm and run to be adjacent. That is the
  form the docs tell planners to prefer, so script_missing never fired
  for it. The prefix flag and its value are now stripped before matching.
- --prefix captured with \S+, so a quoted path containing a space was
  truncated to a stray opening quote and reported as a missing directory
  - a false blocker, worse than the bug this feature fixes. The capture
  is now quote-aware.
- A chained cd whose later segment was absolute concatenated instead of
  resetting, producing a nonsense path and another false blocker. The
  fold now resets on an absolute segment.

Also replaces the bespoke phase-directory regex with the canonical
phase-id helpers. Real phase directories are NN-slug, not phase-N-slug,
so the prior-command harvest matched nothing outside its own fixtures
and the planner-inheritance half of this feature was dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(#2401): source task blocks from the canonical sectionizer

The module carried its own copy of the <task>-block grammar - a fourth
hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps
its copy only because it needs the type= attribute the canonical helper
discards; this module never reads that attribute, so it can share the
owner outright instead of adding a test around a copy.

extractAutomatedCommands now takes task bodies from extractTaggedBlocks
and the out-of-task remainder from stripTaggedBlocks. A task-grammar
parity test pins the attributed task-name set against the canonical
helper across six awkward task shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): extract agent-file overflow to references and repair the property arbitrary

The remote matrix run came back red with 19 failures, four root causes:

- agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the
  49152 agent cap. Their bodies move to gsd-core/references/, leaving
  @-reference stubs, per the documented overflow pattern.
- The new checker dimension invoked gsd_run before the canonical
  preamble that defines it. The call is deleted outright: plan-phase.md
  already runs the probe and hands the result in as {VERIFY_PATHS}, so
  the dimension consumes that rather than re-running anything.
- fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with
  fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range.
- Three runtime-loaded files grew; acknowledged in the existing ack
  fragments that already own those bare filenames, since two ack sources
  may never name the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2401): regenerate golden install-tree fixtures for the new references

Adding two files under gsd-core/references/ changes what the installer
emits into every runtime's tree, so all 19 golden install-parity
fixtures went stale. Regenerated with npm run gen:install-tree; the
delta is exactly the two new reference paths per runtime, no removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2401): backfill changeset pr number to 3678

* fix(#2401): treat ~ as a home expansion only at the start of a path

Windows CI caught this on both shards; the Linux-only remote matrix
cannot see it. The dynamic-path refusal rejected ~ anywhere, and a
GitHub Windows runner's tmpdir is an 8.3 short name -
C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows
path came back unresolvable/dynamic_path.

This was a production bug, not a test artifact: any Windows user whose
project path carries an 8.3 short name, or any literal ~, silently lost
the probe entirely - every command degrading to unresolvable with no
explanation.

~ is a home expansion only at the start of a path; elsewhere it is an
ordinary literal. The check is now split: $, backtick, *, ? and newline
stay refused anywhere (substitution and globs, and the glob characters
are illegal in Windows path components regardless), while ~ is refused
only leading, tolerating one leading quote since the check runs before
quote stripping.

The prior tests only caught this on Windows because only Windows puts a
~ in tmpdir. Four new tests pin it on every platform via a fixture
directory literally named RUNNER~1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 15:21:15 -04:00
Tom Boucher
4e60dba717 fix(#3604): make glossary ref visibility independent of backtick parity (#3680)
* test(#3604): pin parity-dependent ref visibility in the glossary gate

* fix(#3604): make glossary ref visibility independent of backtick parity

* chore(#3604): regenerate CONTEXT-INDEX for corrected predicates

* chore(#3604): regenerate examples CONTEXT-INDEX for corrected predicates

* fix(#3604): complete retired-family exemptions and pin the guard rails

* chore(#3604): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 13:51:40 -04:00
Tom Boucher
7cf6a079fa fix(#3602): bind model resolution for every workflow subagent spawn (#3670)
* test(#3602): guard every spawned gsd-* subagent has a model resolution

* fix(#3602): bind model resolution for every workflow subagent spawn

* test(#3602): merge drift-ack entries into their owning fragments

* fix(#3602): address review findings - docs-update verifier binding, ack merge, guard residuals

* chore(#3602): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:35:45 -04:00
Tom Boucher
dae134b960 test(#2650): bound the stall-watch fan-out with the class norm that reddened windows shard 1/3 (#3671)
* test(#2650): failing-first coverage for the mis-sized stall-watch bound

Lands the regression matrix before the fix. Two assertions fail deterministically on this commit because the helper still calls raw spawnSync with a hard-coded 10000ms bound:

1. The process seam is never reached, so a mock.method spy on it records nothing and the class-norm bound cannot be observed.

2. A raw spawnSync result carries no outcome/timedOut field at all — which is precisely why the windows shard 1/3 failure printed 'null !== 0' instead of naming the timeout or the bound it exceeded.

The bound is sized for the wrong class of call. extractStallHelpersBash slices the entire bash fence, whose runtime-launcher preamble resolves gsd-tools.cjs and really runs two config-get lines — two Node spawns, measured 236ms against 2ms for the fallback the existing comment claims fires (118x). tests/helpers/timeouts.cjs already owns this class as HOOK_FANOUT_TIMEOUT_MS and its comment records the identical failure on PR #3285.

Refs #2650

* test(#2650): bound the stall-watch fan-out with the class norm and report it typed

Drives the failing-first coverage from 15aaef4a7 green (5 failures -> 0). runBashScript now routes through tests/helpers/process-seam.cjs instead of a hand-rolled spawnSync, which CONTRIBUTING.md already requires of anything that shells out.

The bound moves from a hard-coded 10000ms literal to HOOK_FANOUT_TIMEOUT_MS. That constant already exists for exactly this shape — a bash invocation that fans out to nested subprocesses — and its own comment records the identical failure on PR #3285: a bound sized for the wrong class, not a slow machine. This script is that shape: the extracted fence opens with the runtime-launcher preamble, which resolves gsd-tools.cjs and really runs two config-get lines.

The result is now typed. spawnSync reports a kill as status:null, so an exceeded bound reached the call sites as 'null !== 0' — naming neither the timeout nor the bound it exceeded. OUTCOME.TIMED_OUT names itself. status is still aliased from exitCode so the five existing assertions read unchanged.

Corrects the comment that caused the mis-sizing: it claimed gsd_run is undefined so the '|| echo' fallback fires. It is not — the preamble defines it, and the two Node spawns are real (236ms vs 2ms, 118x). They are deliberately left in place; removing them would change what the extracted script executes.

Also drops two now-dead '{ timeout: 10000 }' call-site options. The seam reads only its own documented keys, so those would have been silently ignored while still reading as a 10s bound.

Refs #2650

* test(#2650): make the boundary test a real value-domain boundary

Standards review flagged that the previous 'bound boundary' test repeated the TIMED_OUT arm the test above it already covers, and was not a limit-1/limit/limit+1 in any meaningful sense — an exact-millisecond timing edge would have been a race, not a boundary.

Replaced with a boundary on the VALUE DOMAIN of the bound itself: 0 and -1 must be rejected with TypeError, and 1 (the smallest positive value) must be accepted. Zero is the load-bearing case — spawnSync reads it as 'no timeout at all', which is exactly the unbounded-spawn hazard local/no-unbounded-spawn exists to prevent, so the seam rejects it rather than honouring it.

Also asserts the rejection path does not leak the script temp dir, since the throw escapes through runBashScript's finally. Soak: zero=TypeError, negative=TypeError, one=no-throw, leaked dirs=0.

Refs #2650

---------

Co-authored-by: sim <sim@local>
2026-08-19 12:04:24 -04:00
Tom Boucher
8d1f770dfe test(#3395): pin the clock and scope the stale-prose scan that reddened windows shard 2/3 (#3669)
* test(#3395): failing-first coverage for the silently-ignored clock pin

Lands the regression matrix BEFORE the fix so the failure is proven rather than asserted. Three assertions fail deterministically on this commit:

1. PINNED_ENV does not actually pin. `_pinnedNowMs()` (src/clock.cts) returns null unless GSD_TEST_MODE is set, so GSD_NOW_MS alone is discarded and last_updated is stamped from the live wall clock. An instant ending ...:35.149Z contains the substring 35.1, which is what reddens the windows-latest shard 2/3 lane roughly 1 run in 600.

2. The colliding-instant regression cannot reach its instant, for the same reason.

3. The #3052 same-date test never lands on 2020-09-10, so it has been exercising the different-date path and passing for the wrong reason.

Also adds currentPositionBlock() plus boundary (ms 099/100/199/200, second 34/35/36, LF and CRLF) and two-arm fast-check coverage for the scoped read the fix will switch to.

Refs #3395

* test(#3395): pin the clock and scope the stale-prose scan to the body

Drives the failing-first coverage from a5a919ffb green. Two changes, both needed:

1. PINNED_ENV now sets GSD_TEST_MODE alongside GSD_NOW_MS. _pinnedNowMs() (src/clock.cts:44) returns null without it, so the pin was silently discarded and last_updated carried a live wall-clock instant. src/clock.cts is deliberately NOT changed: requiring both keys is what stops an ambient GSD_NOW_MS from freezing a production clock, so the caller was the side that was wrong.

2. The stale-prose assertion now reads currentPositionBlock(stateContent) instead of the whole document. Frontmatter is not phase prose, and an instant ending ...:35.149Z contains the substring 35.1 — which is exactly how a document with no stale prose in it produced 'the stale 35.1 phase prose must be refreshed away'.

Confirmed hypothesis: the two defects compose. The inert pin supplies a live timestamp; the whole-document scan turns it into a failure. Either alone is latent, which is why this sat unnoticed for five days and then reddened a lane the release never touched.

Also corrects two things the failing-first run exposed. The property test used fc.date() without noInvalidDate, so ~1 sample in 300 was an Invalid Date whose toISOString() threw (counterexample: new Date(NaN)); re-soaked at 5000 runs. And a precondition assertion added to the #3052 block was measured to pass with or without the pin, so it was removed rather than shipped as vacuous truth — last_activity there is body-derived, not clock-derived.

Refs #3395

* test(#3395): apply review findings — pin #3052, one fixture builder, CRLF coverage

Spec-axis review caught a real slip: the #3052 block carried a comment saying its pin was being added as hygiene, but the RED-state revert had removed GSD_TEST_MODE and the fix commit never restored it. A comment describing an action that was not taken is worse than either doing it or leaving it alone — the pin is now actually there.

Standards-axis review flagged the same frontmatter+heading fixture shape being rebuilt in three tests. Extracted one stateDoc({iso, lines, eol}) builder; eol is a parameter rather than a constant because the helper's CRLF behavior is a claim under test.

Self-review finding: CRLF was only exercised on a single-heading document, and the following-heading case only under LF — so the exact claim the helper's comment rests on (`\n## ` matches inside `\r\n## ` because the CR precedes the newline) was never actually run. The control test now loops both line endings WITH a following heading.

Also drops a comment that restated the PINNED_INSTANT rationale verbatim.

Refs #3395

---------

Co-authored-by: sim <sim@local>
2026-08-19 11:25:24 -04:00
Tom Boucher
cd22667b27 Merge pull request #3668 from open-gsd/chore/backmerge-main-to-next-b0ccf790
chore: back-merge main → next (b0ccf790)
2026-08-19 09:52:35 -04:00
github-actions[bot]
02fd1cfc8e chore: back-merge main into next (b0ccf790) 2026-08-19 13:52:12 +00:00
Tom Boucher
552d146086 Merge pull request #3667 from open-gsd/chore/sync-next-version-1.11.0
chore: sync next package version to 1.11.0
2026-08-19 09:51:59 -04:00
github-actions[bot]
97ca69a53e chore: sync next package version to 1.11.0 2026-08-19 13:51:49 +00:00
Tom Boucher
b0ccf790f8 Merge pull request #3666 from open-gsd/release/1.11.0
chore: merge release v1.11.0 to main
2026-08-19 09:51:46 -04:00
github-actions[bot]
182f60b4c1 chore: promote CHANGELOG for v1.11.0 2026-08-19 13:51:02 +00:00
github-actions[bot]
61e30c465b chore: bump version to 1.11.0 for release 2026-08-19 13:06:16 +00:00
Tom Boucher
249c586a40 fix(3582): stop the cold-tree guard from racing the builders it guards against (#3665)
The guard added in #3656 asserts that this test leaves the real repo hooks/ directory
alone. It compared a RAW listing before and after — and just failed on the runner:

    the real repo hooks/ directory listing must be unchanged by this test
    -   '.dist-staging-20858',
        'dist',
        ...

Nothing was wrong with the test's own behaviour. A CONCURRENT scripts/build-hooks.js —
nine test files invoke it from before() hooks — created hooks/.dist-staging-20858 inside
the comparison window. The assertion assumed the shared hooks/ directory is stable for the
duration of a test, which is precisely the assumption this line of work exists to
disprove. The race-detector raced.

The intent is right and is kept: this test must not add or remove anything in the repo.
Only the comparison changes — both snapshots are now filtered through shouldCopyHookEntry,
the same rule the fixture itself uses, so transient build scratch that is not this test's
doing and is excluded from the fixture anyway no longer registers as a difference.

Proven by execution: with a .dist-staging dir injected mid-window the filtered listings
compare equal, while the unfiltered listings provably differ by exactly that entry — so
the old comparison would have failed and the new one is immune rather than merely quieter.
The injected directory is removed afterwards and hooks/ is confirmed byte-identical.

Checked for the same shape elsewhere: this is the only raw listing comparison of the live
hooks/ directory in the file or the repo. The second title in the failure output is the
describe() wrapper around this same test, not a sibling.

Refs #3582

Co-authored-by: sim <sim@local>
2026-08-19 09:00:39 -04:00
Tom Boucher
1bf73d957b enhance(#2295): record the resolved model per reviewer in REVIEWS.md frontmatter (#3649)
* test(#2295): failing-first coverage for per-lane resolved-model recording

* feat(#2295): record the resolved model per reviewer lane

* docs(#2295): document the recorded reviewer model and its provenance

* fix(#2295): refuse control characters in a recorded model value

* test(#2295): correct watermark assertions for the widened mark shape

* fix(#2295): anchor the role-manipulation injection pattern at a word boundary

* feat(#2295): record the applied reasoning effort in the model value

* chore(#2295): backfill changeset pr number

* chore(#2295): restore em-dash in changeset body

---------

Co-authored-by: sim <sim@local>
2026-08-19 08:59:37 -04:00
Tom Boucher
1adf6d2245 fix(#3620): point the docs at files that actually exist (#3658)
* fix(3620): point the docs at files that actually exist

docproof found 34 stale references; the reporter hand-read all 34 and reported the 8 that
are real, explaining why the other 26 are deliberate (files the documents themselves label
legacy or "superseded by", and one pre-Diataxis link label whose target still resolves).
Those 26 are left alone — re-touching them would contradict the issue's own analysis.

Every claim was re-verified against git ls-files at HEAD before editing.

docs/INVENTORY.md said its roster is anchored by six drift-control tests. Five are gone
(commands-doc-parity, agents-doc-parity, cli-modules-doc-parity, hooks-doc-parity in
5d8a8c4d; command-count-sync in fbf30792), so the sentence now names the one that exists.
Whether one test is sufficient coverage is a maintainer question the issue explicitly
declined to answer, so no new drift tests are proposed here.

The four translations were a revision further behind, each naming a seventh test deleted in
ae8bb707 that the English file had already dropped. All four now match.

Renamed targets corrected in CONTEXT.md, VERSIONING.md, docs/CONFIGURATION.md and the
update workflow. The new test names carry no issue-NNN- prefix, which is what
lint-regression-test-names requires, so they are the correct targets.

docs/TESTING-SUITES.md is the one that could cost somebody time: it INSTRUCTED contributors
to add an acknowledgment to the legacy drift-ack file, which CONTRIBUTING.md says to never
use. Rewritten from the real workflow — per-PR fragments under the drift-acks directory,
and a spent base-side ack is re-armed by rewording that fragment's reason in place, never
by adding a duplicate, since two sources naming one path is a hard error.

docs/skills/discovery-contract.md's heading named a query module deleted in 11918dcc. The
section was REMOVED rather than retargeted: its documented behavior (skip the deprecated
root) is not what the surviving code does — skill-manifest includes that root marked
deprecated:true — so retargeting would have documented something false.

Found and fixed inline, same class: VERSIONING.md described an SDK bundling step the
release workflow does not have (zero such mentions in that file); CONFIGURATION.md and four
translations named a dead model-catalog triple collapsed by ADR-457.

Dead config removed: the changeset lint's user-facing prefix list still carried two retired
sdk entries. git ls-files -- 'sdk/*' returns nothing. No test pins that array.

Left deliberately: the comment explaining the retired catalog path, the install regression
test that reconstructs the old broken layout to prove it fails, and the generated
test-timings cache. Each is a legitimate mention of a dead path, not drift.

Note lint-removed-but-needed cannot catch this class: it diffs baseRef...HEAD, so it only
sees files deleted in the change under review. These were orphaned by PRs that predate the
lint. A repo-wide existence audit would need a suppression mechanism for the 26 deliberate
mentions above; that is a feature, not part of this fix.

Fixes #3620

* chore(3620): backfill changeset PR number (#3658)

---------

Co-authored-by: sim <sim@local>
2026-08-19 01:54:01 -04:00
Tom Boucher
4e70b245e8 fix(#3582): stop the cold-tree fixture racing concurrent hook builds (#3656)
* fix(3582): stop the cold-tree fixture racing concurrent hook builds

Two tests in tests/gsd-check-update-worker-platform-gate.test.cjs failed a verification run
with `ENOENT: no such file or directory, lstat '/work/hooks/.dist-staging-20836'`. This is
a race I introduced in #3582, not a flake, and it passed when #3582 merged because it only
fires when the timing lines up.

buildColdInstallTree() copied the LIVE repo hooks/ directory with a filter that excluded
only the basename 'dist'. scripts/build-hooks.js writes atomically through a per-PID
staging dir, hooks/.dist-staging-<pid>, and removes it when finished — and the archived
build-hooks-atomic-write changeset records that NINE test files invoke build-hooks.js from
their before() hooks. So several test processes create and delete staging directories
inside hooks/ while other tests are reading it. cpSync enumerated one, and the owning
process removed it before cpSync got to it.

The helper's own header already states the rule it needed: hooks/dist is excluded because
it "is not present in a raw marketplace checkout either". hooks/.dist-staging-* is
gitignored (.gitignore:21) and equally absent from a raw checkout — it was simply missed.

Fixed by enumerating hooks/ explicitly and skipping 'dist' and any '.dist-staging' prefix
BY NAME, before anything stats or copies the entry, then copying each surviving entry
individually. A name-first skip means a vanishing staging dir is never touched at all.

Worth recording because it corrects the assumption this fix was written under: cpSync's
filter IS invoked before the entry is lstat'd, and returning false leaves it untouched
(verified by deleting inside the callback and returning false — no throw). So merely adding
'.dist-staging' to the old filter would also have closed the race. The explicit enumeration
was kept anyway so correctness does not depend on that Node implementation detail.

Proven by execution both ways: with a staging dir planted in hooks/, the OLD
cpSync-with-filter form copied it straight through into the fixture, while the new form
succeeds and produces no .dist-staging entry with the real hook set intact.

Regression test added beside the existing cold-tree tests: it plants a real
hooks/.dist-staging-test-<random>, asserts the fixture builds clean without it, and removes
only the directory it created.

Repo swept for the same exposure: this helper is the only place doing a bulk enumeration of
the whole live hooks/ tree. The other hooks/-touching tests reference specific named files
or hooks/dist/ and are not exposed. scripts/build-hooks.js is deliberately untouched — its
per-PID staging is what makes its own writes atomic and is correct.

Refs #3582

* fix(3582): make the race regression test hermetic instead of mutating the live tree

The regression test added in the previous commit failed the runner with "failed running
after hook", and it was wrong in two ways — the second one worse than the first.

cleanup() (tests/helpers.cjs:452-487) deliberately THROWS for any path outside the known
temp roots. The test planted hooks/.dist-staging-test-<random> inside the repo and then
asked cleanup() to remove it, so the after-hook threw. That guard is correct and is left
alone.

The real problem is that the test mutated the LIVE hooks/ directory while other test files
concurrently read it — the exact shared-state hazard this change exists to remove. A
regression test for a race must not introduce one.

buildColdInstallTree now takes an optional opts.repoRoot (defaulting to the real REPO_ROOT
and used for both copies it performs), so the test builds a fake repo root under the temp
dir, plants representative hooks plus dist/ and .dist-staging-99999/ THERE, and asserts the
fixture excludes both. All six pre-existing callers pass no arguments and are unaffected.
The test also asserts the real hooks/ listing is identical before and after, so a future
edit that reintroduces live-tree mutation fails loudly.

The name rule is now pinned directly rather than only through the copy. shouldCopyHookEntry
is exported and asserted, including the two cases a sloppier implementation would get
wrong: 'dist-staging-no-dot' and 'distant.js' must both be KEPT. Anything matching on a
loose 'dist' substring or startsWith passes every other case and fails those two.

Also corrected the issue number on the tests introduced here: they were labelled #3631,
which is the unrelated capability-consent bytecode work. This is #3582.

Verified by execution: the predicate rule holds on all nine cases; a fake-root fixture
yields exactly the representative hooks with dist and .dist-staging excluded; the no-arg
default still copies the real tree (29 entries); and the real hooks/ listing is byte-identical
before and after.

Refs #3582

* chore(3582): re-trigger CI after an orphaned Validate Branch Name run

The Validate Branch Name run for this branch (32211622051) sat queued from 03:17 and was
never picked up — updatedAt never advanced past createdAt while the same workflow completed
normally for other branches. `gh run rerun` refused it ("already running") and
`gh run cancel` returned HTTP 500, so the run is orphaned on the GitHub side.

Closing and reopening the PR re-fired the other pull_request workflows but not that one,
whose triggers evidently do not include reopened. An empty commit is the remaining way to
get a fresh run.

No file changes: the tree is identical to dc71534b6, whose remote-runner pass carries
forward unchanged.

Recording this rather than admin-merging past the pending check. Everything else was green
(24 pass, 0 fail), but admin merge is sanctioned only for the missing-secondary-reviewer
case, never to skip a gate that has not actually run.

Refs #3582

---------

Co-authored-by: sim <sim@local>
2026-08-19 01:33:40 -04:00
Tom Boucher
9e4f0e99ad fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash

bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.

Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.

The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.

Refs #3631

* fix(3631): exclude derived bytecode caches from the consent digest

RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.

collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.

Three properties were preserved deliberately, each pinned by a test:

  - The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
    first would have turned the exclusion into a way to smuggle a symlink past the check;
    a symlink named x.pyc still throws.
  - Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
    caps guard the WALK; the digest answers a different question, and exclusion must not
    become an unbounded-bytes hole.
  - An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
    directory marker was the sharper half of this bug: an empty __pycache__ flipped the
    hash before any .pyc existed, so a *.pyc-only filter would have left it live.

The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.

Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.

Fixes #3631

* fix(3631): narrow the digest exclusion after two isolated security reviews

The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.

HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.

FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.

Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.

Narrowed to what is actually defensible:
  - a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
    the walk still recurses and hashes every non-excluded child.
  - .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
  - a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
  - declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
    basename are now rejected in both validator copies — a file named .pyc can contain
    perfectly valid JavaScript, so the exclusion must not be reachable from a declared
    surface.

Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.

Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.

Refs #3631

* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind

Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.

HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.

The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.

The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.

Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.

Changeset rewritten: it still described the rejected wholesale-exclusion semantics.

Refs #3631

* chore(3631): backfill changeset PR number (#3650)

---------

Co-authored-by: sim <sim@local>
2026-08-18 22:35:54 -04:00
Tom Boucher
02a36d3db9 fix(3618): update the Windows fallow assertions to the behavior #3618 chose (#3654)
full test (windows-latest, 24, shard 1/3) has been RED on next since ac1b6d679 and blocks
every PR. Three assertions fail, Windows-only, and all three are stale TESTS rather than
resolver defects.

Two of them create a BARE extensionless 'fallow' and assert it resolves. #3618 deliberately
stopped resolving that on Windows, and said so in its own commit message: the extensionless
file is npm's POSIX sh shim, which CreateProcess cannot run (#3275). The code did what it
meant to; the tests were never updated to match.

Both now assert BOTH platform contracts rather than skipping a lane — the precedent #3618
set when it fixed its own F4 ('a t.skip on one lane would have been green and would have
left the win32 carve-out unpinned'). POSIX keeps the original bare+chmod fixture verbatim.
win32 creates the artifact npm actually writes there, fallow.cmd, and asserts resolution
finds it; each also pins that a bare extensionless fallow ALONE still resolves to null on
win32, so the carve-out is pinned rather than merely stepped around. The precedence test
keeps testing precedence: on win32 both node_modules/.bin and the PATH dir get a
fallow.cmd, and .bin must still win.

The third failure was case, not logic: the test wrote fallow.cmd and asserted exact string
equality, but candidates are built by appending PATHEXT entries, and PATHEXT is uppercase.
Confirmed at the source — DEFAULT_PATHEXT = '.EXE;.CMD;.BAT;.COM'
(src/shell-command-projection.cts:651), read verbatim and never lowercased, with the
candidate returned as-is. So the resolver returned fallow.CMD, which is the same file on
case-insensitive NTFS. Fixed the assertion to compare case-insensitively on win32; the
resolver is untouched, because its return value has to stay usable verbatim.

Not verifiable here: this repo's remote runner matrix is Linux-only and the host is macOS,
so the win32 branches are reasoned from the resolver source and proven only by GitHub CI.
The POSIX lanes were verified by execution (bare+chmod in .bin resolves, .bin beats PATH,
non-executable in PATH yields null).

No changeset: changeset-lint's USER_FACING_PREFIXES does not include tests/, so this is
OK_NO_USER_FACING_CHANGES rather than a missing fragment.

Refs #3618

Co-authored-by: sim <sim@local>
2026-08-18 22:08:52 -04:00
Tom Boucher
2972da4c9d enhance(#3619): ratchet the platform seam with local/no-private-binary-resolution (epic #3411 Phase 3) (#3636)
* chore(#3619): ratchet the platform seam with local/no-private-binary-resolution

Epic #3411 Phase 3, the ratchet. Scope revised with maintainer approval and
recorded on the issue: the epic's literal ask was a rule rejecting a bare-name
spawn outside the seam. Surveyed at ac1b6d679, ~30 such sites exist and none is
a defect — git, gh and npm ship native .exe that CreateProcess resolves unaided,
and the rest are POSIX-only tools. ADR-1703 rules 2 and 3 forbid grandfathering
and escape hatches, so a literal rule would be unsuppressable and would force
rewriting 30 correct calls.

The epic's actual thesis was four private RESOLVERS, not four bare spawns. So
the rule flags re-implementing resolution: reading PATHEXT in any casing from
any object, and a hardcoded list carrying two or more of .exe/.cmd/.bat/.com —
precisely the shapes fallow-runner's candidateNames and gsd-tools' PATHEXT
string had before Phases 1 and 2 deleted them.

Three boundaries were arrived at rather than assumed:

  two-or-more   a single .endsWith('.cmd') is a classification, not a candidate
                set; runtime-hooks-surface derives .cmd shim paths that way
  boundary-aware  a naive substring test flags .execute and .compacting, caught
                on src/host-integration.cts before it could become a false
                positive nobody could suppress
  suffix-anchored  the seam exemption matches src/shell-command-projection.cts
                exactly; a substring match would also exempt the dispatch test
                file. Case I9 pins it.

PATH scans are deliberately NOT flagged — membership checks (bin/install.js)
are indistinguishable from resolution scans, and an unsound rule in a
zero-escape-hatch architecture is worse than no rule.

To make the ratchet strict with no carve-out, resolveExecutableBinary gained
pathOverride: search THIS PATH, read everything else including PATHEXT from the
ambient environment. resolveFallowBinary now supplies its own search path
without hand-threading PATHEXT, which would itself have been a private read.
The three alternatives were all worse: exempting the file is grandfathering,
exempting the AST shape is a carve-out every future caller must replicate, and
dropping the pass-through would silently ignore a user's real PATHEXT — buying
a lint rule with a correctness regression.

eslint-rules/** is outside the rule's globs rather than exempted, because
portability-vocab.cjs owns the extension set. scripts/**/*.cjs got its own block
so that exclusion does not leave a hole in the ratchet.

Started green with nothing suppressed. Proven able to fail: a fixture with both
signals reports two errors.

Refs #3411

* fix(#3619): close the PATHEXT destructuring evasion and correct two overclaims

Adversarial review found a trivial evasion of the rule's primary signal: the
visitor only handled MemberExpression, so

  const { PATHEXT } = process.env
  const { PATHEXT: exts } = process.env
  const { Pathext } = opts.env

were all unflagged. That is a common idiom, not an exotic bypass. An ObjectPattern
visitor now catches it in every form — renamed, any casing, any receiver, string
keys — while leaving a computed key alone, since it is not statically decidable.
I10-I13 pin the invalid forms and V9/V10 pin PATH and the computed key.

Two overclaims corrected, both mine:

Standards review proved the docs were factually wrong. Both the ADR amendment and
the CONTEXT.md entry asserted that tests/shell-command-projection-dispatch.test.cjs
is still linted by this rule. It is not — the rule's surface is src, gsd-core/bin,
scripts and hooks, and tests/** is deliberately outside it because test setup
legitimately assigns process.env.PATHEXT (fallow-runner's P3 does exactly that).
The suffix-vs-substring distinction is therefore proven by RuleTester case I9
feeding a synthetic filename, NOT by real coverage of that file. Both documents now
say so.

The rule's own docstring claimed the seam exemption matches the seam path
'exactly'. It is a suffix match, so a nested foo/src/shell-command-projection.cts
would also be exempt. Suffix matching is kept — it is how sibling rules resolve
paths and the nested case does not exist — but the docstring now states the
boundary rather than overstating the precision.

The evasion fix was verified by executing eslint against both destructuring forms
in scripts/, not by inspection. Probe: 31/31.

Refs #3411

* chore(#3619): backfill changeset pr number 3636

---------

Co-authored-by: sim <sim@local>
2026-08-18 16:40:17 -04:00
Tom Boucher
0f417aa6d0 fix(#3584): the verb owns the count token and nothing else (#3635)
* test(3584): failing-first coverage for Plans-line trailing text

roadmap update-plan-progress preserves trailing text only when the line begins with a
canonical count token. Every other phrasing — including the TBD value the shipped
template itself suggests — is replaced to end-of-line, and a sentence wrapping onto a
second line has its first line deleted, leaving the continuation standing alone so the
roadmap asserts something nobody wrote. Exit 0, updated:true, and the diff reads as a
routine count bump.

These tests fail on that, and pin the arms that must keep working: the template
placeholder is still replaced, the #2853 token-plus-annotation path is unchanged, CRLF is
neither stranded nor duplicated, and a run that leaves the line alone still updates the
Progress table and checkboxes rather than becoming a no-op.

* fix(3584): the verb owns the count token and nothing else

RED proven at 105bdf7c: 7 failures — the preserving cases (freeform prose, wrapped
continuation, TBD, CRLF) failed while the template-placeholder and #2853 token arms
passed on base.

The trailing-text guard fired only when the line began with a canonical count token:
 dropped the rest of the line whenever the
regex's count group did not match. #2853 fixed end-of-line truncation on that one path
only. The in-code comment justified the rest as 'the fresh-template bracketed placeholder
or other freeform guidance, not user prose' — a heuristic that misreads ordinary human
phrasing and destroys even TBD, the value the shipped template itself suggests at
templates/roadmap.md:37.

The sharper failure was the wrapped sentence: only the first line is inside the match, so
the verb deleted line one and left line two standing alone, leaving the roadmap asserting
something nobody wrote — at exit 0, updated:true, in a diff that reads as a routine count
bump.

Inverted the default into three arms. A real count token is rewritten with its annotation
preserved (unchanged, #2853). A bracketed placeholder is detected POSITIVELY and replaced.
Everything else — freeform prose, TBD, a wrapped sentence's first line, an empty value —
returns the match untouched. That last arm resolves the wrapped case by construction: an
untouched first line cannot orphan its continuation.

Positive detection is the load-bearing part. Implemented as 'not a count token, therefore
disposable', rows 1-3 come straight back; the detector instead asks whether the value IS a
bracketed placeholder.

Leaving the line alone does not make the verb a no-op: the phase checkbox, the
Progress-table cells and the plan-checklist row still update in the same run, and that is
asserted. CRLF is unaffected in every arm — the pattern's [^\r\n]* never consumes the
\r, so it sits outside the match regardless of which arm runs.

Fixes #3584

* fix(3584): detect the template placeholder by its text, not by its brackets

Two defects in the arm-2 detector shipped in fc23e49e, both found in review.

Finding A: isBracketedPlaceholder asked only whether the trimmed value was wrapped
in [...]. Brackets are ordinary prose punctuation in a roadmap, so any hand-written
bracketed note — '[Deferred pending re-scope]', '[blocked on #1234]' — was classified
as the fresh-template placeholder and destroyed. That is the very defect #3584 is
about, reintroduced one arm over. The detector now matches the placeholder's TEXT
(/^\[\s*Number of plans\b[\s\S]*\]$/i), so it recognizes the shipped template
value and its short form and nothing else.

Finding B: the count group matched '\\d+\\s+plans' only. The plural is not the
template's own output shape — templates/roadmap.md:62 ships '1 plan' — so a
single-plan phase fell through every arm and its line froze permanently, never
updating again. Widened to 'plans?'. This one was introduced by the arm-3 default:
before it, the singular fell through to the old replace-everything path and at
least stayed current.

Cases 11-14 cover both: a bracketed human note preserved, the short placeholder
still replaced, '1 plan' rewritten, and '1 plan (annotation)' rewritten with the
annotation intact. Verified against the live binary, not just re-read.

Also converted all 15 cases in this block from try/finally to t.after(), per
CONTRIBUTING.md:356-370 which bans try/finally in test bodies. The existing #2853
block above is untouched — it is not in this change's scope and its conversion is
not this fix's concern.

Refs #3584

* chore(3584): add changeset fragment

Fixed-type fragment for the roadmap Plans-line trailing-text fix. pr:0 placeholder, backfilled once the PR number exists.

Refs #3584

* chore(3584): backfill changeset PR number (#3635)

---------

Co-authored-by: sim <sim@local>
2026-08-18 16:39:17 -04:00
Tom Boucher
ac1b6d679f enhance(#3618): fold fallow-runner onto the canonical binary resolver (epic #3411 Phase 2) (#3633)
* chore(#3618): fold fallow-runner onto the canonical binary resolver

Epic #3411 Phase 2. src/fallow-runner.cts was the fourth divergent
implementation of Windows binary resolution the epic enumerated —
candidateNames, isExecutableFile, findInPath, findInNodeModules, 40 lines.
All four are deleted; resolveFallowBinary is one seam call.

Two OPT-IN options were added to resolveExecutableBinary to make the fold
behavior-preserving, both defaulting off so Phase 1's callers are byte-identical:

  prependPaths      dirs searched before env.PATH, in order, through the
                    identical per-directory candidate logic. This expresses
                    node_modules/.bin-first precedence without env surgery —
                    the rejected alternative re-introduced the
                    spread-loses-the-proxy hazard the Windows lane caught in
                    Phase 1, at every future call site instead of once.
  requireExecutable POSIX-only accessSync(X_OK); a no-op on win32 where mode
                    bits do not mean execute. Opt-in rather than default
                    because unconditional X_OK breaks #3445's suite, which
                    stages candidates with plain writeFileSync and never sets
                    an exec bit — the repo bans chmod in tests — so every one
                    would resolve to null on POSIX.

Deliberate behavior change on Windows: fallow's prior candidate list ended in a
BARE fallow. The seam never tries a bare name there, so an extensionless file
beside fallow.cmd is no longer resolved. That is the fix, not a regression — the
extensionless file is npm's POSIX sh shim, which CreateProcess cannot run
(#3275). Rows 7 and 8 of the design record it.

Defect found while working, fixed inline: the resolution order was documented
BACKWARDS as PATH-then-.bin in structural-pre-pass.md, docs/INVENTORY.md and
four INVENTORY translations. The code has always been .bin first, and .bin first
is correct — a project-local tool should beat a global one. The archived
changeset is left alone as a historical record.

fallow-runner had no test file at all. tests/fallow-runner.test.cjs is new
(F1-F15) and the seam options are pinned by S1-S12 folded into the existing
dispatch suite. RED proven by execution: with both source files stashed and
build:lib re-run, 7 of 27 probe cases failed.

Refs #3411

* chore(#3618): backfill changeset pr number 3633

* fix(#3618): assert both platform contracts in F4 instead of a POSIX-only premise

Windows CI on #3633 failed F4. The test monkeypatched accessSync to throw and
asserted resolveFallowBinary returned null — but that premise, that the X_OK
check is consulted at all, is POSIX-only by design. requireExecutable is a
deliberate no-op on win32 because Windows mode bits do not mean execute, so the
staged fixture correctly resolved there.

40-design.md's negative-space section already states this carve-out verbatim.
The test contradicted the design it was written from: fixtures were made
platform-adaptive in the previous commit, and this assertion was left
platform-blind.

F4 now asserts BOTH contracts — null on POSIX, resolves on win32 — rather than
skipping either. A t.skip on one lane would have been green and would have left
the win32 carve-out unpinned by fallow's own entry point.

Audited every other row for the same class. F1-F3, F5, F6, F11-F15 hold on both
platforms; F7-F10 and S1-S12 inject platform explicitly and are unaffected. F4
was the only row with a single-platform premise.

The local probe runs on one platform and structurally cannot catch this, which
is why it was green — that limitation is now stated at the top of the probe so a
green probe is not mistaken for platform coverage. The win32 branch was proven
by injecting platform:'win32' with accessSync throwing and asserting it still
resolves.

Refs #3411

---------

Co-authored-by: sim <sim@local>
2026-08-18 15:32:57 -04:00
Tom Boucher
46f14c621e fix(#3583): one percent per write — route update-progress through the shared computation (#3634)
* test(3583): failing-first coverage for one percent per write

state update-progress computes plan throughput (summaries/plans) for stdout and the
body Progress bar, while the same write re-derives frontmatter progress.percent as
min(planFraction, phaseFraction). Neither consults the other, so mid-phase the file
contradicts itself and state json disagrees with the verb that just wrote it.

These tests fail on that: equality across stdout, body bar, frontmatter and state json
on fixtures where the two fractions differ, plus a derivation-parity test that fails if
completedPhases is ever derived by summary parity instead of verification-passed status.

Also updates three pre-existing tests that pinned stdout to the plan-throughput value
(50->0, 50->0, 100->0). Those fixtures have summarized-but-unverified phases, so the old
expectations encoded the bug; changing them IS the fix, as the issue states explicitly.

* fix(3583): one percent per write — route the verb through the shared computation

RED proven at 7dbbb2d2: 9 failures — the new cross-surface equality tests, the withhold
test, and the pre-existing tests whose expectations encoded the bug.

state update-progress computed plan throughput (summaries/plans) for stdout and the body
Progress bar, while the SAME write re-derived frontmatter progress.percent as
min(planFraction, phaseFraction) through a separate path. Neither consulted the other, so
on any project where plan throughput ran ahead of phase completion — the normal mid-phase
state — the file contradicted itself and state json disagreed with the verb that had just
written it. Exit 0, no signal.

This is not a dispute about which metric is right. The min cap is deliberate (#3242 Bug B)
and is untouched; the fix aligns the printed and body values WITH it. Verified by diff:
computeProgressPercent's definition and cmdStateSync are both unmodified.

The verb now takes its percent from buildStateFrontmatter — the single owner of the
isPhaseComplete-based completedPhases count and the ROADMAP-union totalPhases logic that
the frontmatter sync later uses inside the same read-modify-write. Both calls hit the same
disk-scan cache against the same on-disk state, so they cannot disagree. Reusing that
owner, rather than re-deriving completedPhases locally, is the point: a second
almost-identical derivation is the very defect class being fixed, and a parity test now
fails if anyone swaps it for summary parity.

The first cut fell back to plan throughput when the shared computation withheld. That
reintroduced the defect in a rarer case — stdout would print a number the frontmatter
deliberately did not contain — so it is gone. The verb now withholds in the same shape as
its existing #3217 and #3233 guards. That path is reachable, not theoretical: a bare vX.Y
token in ROADMAP prose with no versioned heading leaves the milestone unbounded while both
existing guards see a COMPLETE scope. Covered by a test that also asserts state json omits
the percent, proving it is the same withhold rather than a divergent local computation.

Three pre-existing tests pinned stdout to plan throughput (50->0, 50->0, 100->0); their
fixtures have summarized-but-unverified phases, so those expectations encoded the bug.
Updating them is the fix, as the issue states.

Fixes #3583

* fix(3583): source the reported counts from the same milestone window as the percent

The adversarial pass found the first cut left the SAME defect one field over.

cmdStateUpdateProgress still reported completed/total from the top-of-function scan,
which calls listMilestonePhaseDirs with NO versionOverride — the auto-derived current
milestone — while percent now came from buildStateFrontmatter, whose scan scopes by
versionOverride: storedMilestone. getMilestonePhaseFilter shows those can select
different milestone windows, and #3017's own comment warns about exactly that mis-bind.
So a single JSON object could report a percent inconsistent with its own counts: the
self-contradiction this issue was filed to close, relocated rather than removed.

Counts now come from the same buildStateFrontmatter result as the percent. Proven on a
real divergent-milestone fixture where a preamble phase leaks into the auto-derived scan
but is excluded from the stored-milestone-scoped one: with the fix stashed the verb emits
{percent:0, completed:1, total:2}; with it applied, {percent:0, completed:1, total:1}.
The guard scan remains, gating only the #3217/#3233 withholds.

Also corrected a comment that overstated caching. Only the phase/plan disk scan is shared
between the two buildStateFrontmatter calls; getMilestoneInfo re-reads and re-parses
ROADMAP.md and readGitHeadSha spawns a bounded git rev-parse, and both now run twice per
invocation. Threading a precomputed frontmatter through the write seam to avoid it was
rejected: that seam is the shared ADR-3408 §8.3 composition with three other callers and
heavily-documented invariants, and this is not the change to renegotiate it. The comment
now says what is and is not cached instead of implying the second call is free.

Standards: six new assertions matched raw STATE.md body text the code under test had just
produced — the pattern CONTRIBUTING bans by name. They now extract the body Progress field
with the repo's own field extractor and assert the parsed percent, so the check survives
rewording of the rendered bar. The acceptance criterion still verifies the bar; only what
it asserts on moved.

Also trimmed ~50 lines of narration around a ~15-line change into a named helper, and
fixed a stale test comment that still claimed 100% next to assertions expecting 0%.

* chore(3583): add changeset fragment

* chore(3583): backfill changeset PR number (#3634)

---------

Co-authored-by: sim <sim@local>
2026-08-18 15:24:07 -04:00
Tom Boucher
924f649f87 docs(#3625): record the spawn-library evaluation as ADR-3625 (#3632)
Spike outcome for #3625: evaluate cross-spawn / nano-spawn / execa against
the hand-rolled Windows binary resolution and cmd.exe mediation that epic
#3411 Phase 1 (PR #3621) is landing in the platform seam.

Verdict: stay hand-rolled, with revisit-if conditions recorded so the call
is not re-litigated in a future PR.

Evidence, per the issue's "Done when" list:

- Sync/async verdict per candidate. nano-spawn is async-only — settled by
  its own README, which lists "synchronous execution" among the features
  execa has and it does not. cross-spawn exposes `.sync`. execa exposes
  execaSync, which its own docs discourage.
- CVE-2024-27980 escaping verdict per candidate. None uses shell:true.
  cross-spawn independently arrives at the SAME mechanism the seam uses:
  cmd.exe /d /s /c with a pre-escaped line and windowsVerbatimArguments.
  That validates the seam's approach rather than superseding it.
- Maturity axes scored with measured data (registry metadata 2026-08-18,
  transitive footprint measured by install).

Two premises in the issue did not survive measurement, both recorded:

- The CRITICAL 167-symbol/53-file blast radius is a `direction:both`
  measurement, inflated by downstream callees. A call-shape change ripples
  to CALLERS: upstream at depth 15 is 14 symbols / 7 files, MEDIUM, and it
  terminates at depth 4. The decision does not rest on the CRITICAL figure.
- The cited vendoring precedent path does not exist; the real one is
  gsd-core/bin/lib/vendor/re2js.cjs.

Decisive against the only structurally-eligible candidate (cross-spawn):
it resolves process.cwd() first on Windows even when an explicit PATH is
supplied, calls process.chdir() during resolution, and keys escape depth on
a node_modules/.bin/*.cmd path regex of the same shape #3411 was filed to
delete. Adoption would also break execTool's observable not-found contract
across 53 dependent files.

Doc-only: docs/adr/** plus a root-level CONTEXT.md pointer. The CONTEXT.md
addition is a new line rather than an edit to the seam's glossary paragraph,
which PR #3621 rewrites wholesale — same-line edits would conflict on merge.
ADR index regenerated via scripts/gen-adr-index.cjs --write.

lint:ci exit 0; lint-docs-required and changeset/lint both run with
GITHUB_BASE_REF=next and report ok_no_user_facing_changes.

Closes #3625

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 14:24:45 -04:00
Tom Boucher
bf87dd4156 enhance(#3617): one canonical Windows binary resolver in the platform seam (epic #3411 Phase 1) (#3621)
* feat(#3411): one canonical Windows binary resolver in the platform seam

CONTEXT.md declares src/shell-command-projection.cts the single OS-facing seam,
but Windows binary resolution had grown four divergent implementations outside
it. #3445 folded two of them together — inside gsd-core/bin/gsd-tools.cjs, not
the seam — so the declaration stayed untrue and execTool still had no handling
at all.

Lift the resolver into the seam as resolveExecutableBinary, and export the half
that actually executes as projectSpawnInvocation: CreateProcess cannot run a
.cmd/.bat, so the cmd.exe mediation is inseparable from the lookup and splitting
them is how the copies accumulated. cmd.exe is invoked with an explicit argv
array, never shell:true — CVE-2024-27980's vector and Node 26's DEP0190.

execTool now resolves on win32. POSIX is a strict no-op by construction, which
matters: execTool rates CRITICAL blast radius (167 symbols, 53 files).
gsd-tools.cjs deletes its private scan and its private mediation and delegates.

Two semantics grown beyond #3445's resolver, both additive: a name already
carrying a PATHEXT-listed extension is tried as-is before the append loop, and a
suffix outside PATHEXT is not treated as an extension.

Refs #3411

* fix(#3411): keep mediating a declared .cmd that PATH resolution misses

Standards review caught a narrowing against the code this replaces. gsd-tools.cjs
computed `target = resolveSpawnBinary(binary) || binary` and keyed the shim test
on `target`, so a declared .cmd mediated whether or not PATH resolution found it.
That is load-bearing: resolveExecutableBinary scans PATH only, while `cmd.exe /c`
also finds a batch file in the current directory.

Mediation now keys on the target — resolved path, else declared name. The ENOENT
contract still holds for BARE unresolved names, which is the case it was written
for. P9/P10 pin both halves.

Spec review found E1/E2/E3/E5 promised by 50-test-matrix.md but never written;
added. E3 is the integration proof that the CVE-relevant mediation fires through
execTool, not only through projectSpawnInvocation in isolation.

Also adds the CONTEXT.md glossary entry for the seam's new resolution ownership
(a PR gate) and the changeset fragment.

Refs #3411

* fix(#3617): pass mediated cmd.exe arguments verbatim so metacharacters cannot inject

The isolated security pass found the mediation shape carried an argument-injection
surface. libuv's quote_cmd_arg force-quotes an argv element only when it contains
a space, tab, or quote — never for a cmd metacharacter — and cmd.exe re-parses
everything after /c. So an arg of a&calc arrived unquoted and cmd ran calc.
Node's own CVE-2024-27980 escaping cannot help: it fires only when the spawned
FILE is the .bat/.cmd, and here the file is cmd.exe.

Caret-escaping is not a fix. It is correct only when libuv does not quote, and
libuv quotes whenever the arg also contains a space — no per-arg transform is
right in both cases. So build the command line and pass it through verbatim, the
shape Rust's std adopted for the sibling CVE-2024-24576: one outer quote pair
that cmd /c strips, every token inside force-quoted, embedded quotes doubled.

An argument containing CR or LF is refused rather than mediated — a newline
cannot be represented in a Windows command line, so mediating would silently
truncate. Failing visibly is correct.

Known limit, documented at the seam: %VAR% still expands inside a /c string and
has no escape outside a batch file. That is information disclosure, not arbitrary
execution, and is the same limit Rust's std documents.

This was byte-for-byte the shape #3445 shipped, so the fix closes it for the
reviewer-lane spawn path too, not only for execTool's newly reachable route.

Refs #3411

* docs(#3617): document the subprocess-execution security posture

Adds Layer 4 to the security model: why GSD never uses shell:true for binary
invocation (CVE-2024-27980, Node 26 DEP0190), why resolution is explicit and
never tries the bare name on Windows (the npm extensionless-shim trap behind
#3275), and why .cmd/.bat mediation builds a verbatim force-quoted command line
rather than relying on default escaping — Node's own CVE protection cannot fire
once the started program is cmd.exe.

The residual %VAR% expansion limit is stated plainly under Trade-offs rather
than left implicit: it is information disclosure, not arbitrary execution, and
callers passing untrusted text to a Windows .cmd should not assume the value
arrives byte-identical.

Docs-only; no code change.

Refs #3411

* chore(#3617): backfill changeset pr number 3621

* fix(#3617): read PATH, PATHEXT and ComSpec case-insensitively

The Windows CI lane on #3621 failed E5, and the root cause was a defect in the
implementation, not the assertion.

Windows names the variable Path, not PATH. process.env is a case-insensitive
proxy, so process.env.PATH works — but execTool builds
{ ...process.env, ...opts.env } whenever a caller supplies opts.env, and
spreading discards the proxy while keeping the OS's actual casing. The exact-case
env['PATH'] lookup then returned undefined, the PATH scan saw zero segments,
resolution returned null, and the change degraded to precisely the spawn ENOENT
it exists to fix. ComSpec and PATHEXT had the same exposure.

#3445's tests never caught it because they pass uppercase keys explicitly, and
neither did the Linux remote runner — this is a defect only the Windows lane
could see.

_envGet resolves a variable by exact match first (so a canonical caller pays no
scan) and falls back to a case-insensitive sweep.

R23 and P16 pin it and were proven RED by execution: with the fix stashed and
build:lib re-run, R23 returned null and P16 returned the cmd.exe default.

R24 was rewritten because the first version was vacuous — it staged foo.CMD, so
the default PATHEXT already contained .CMD and it passed against the broken code
for the wrong reason. It now stages foo.XYZ, an extension absent from the
default, and carries a negative control asserting that dropping the Pathext key
yields null. Re-proven RED the same way.

E5's assertion was corrected alongside the fix: 'PATH' in options.env expressed
the wrong contract. It now checks case-insensitively for the key.

Refs #3411

* fix(#3617): execTool spawns the declared name unless mediation is required

The Windows full-test lane on #3621 failed tests/graphify.test.cjs — the python3
identity check asserted 'python3' and got the absolute resolved path
C:\hostedtoolcache\windows\Python\3.12.10\x64\python3.EXE instead.

Those tests are correct and the change was wrong. They pin a long-standing
contract — execTool spawns the program name it was given — by spying on
spawnSync's first argument, and routing every win32 call through the projected
invocation broke it.

Resolving a .exe buys nothing. libuv's CreateProcess path already performs
PATH + PATHEXT search, which is why spawning a bare 'node' has always worked on
Windows. The only case the OS genuinely cannot spawn is a .cmd/.bat. So execTool
now adopts the projection only when mediation actually happened —
windowsVerbatimArguments is exactly that flag — and otherwise passes the declared
program and args through untouched.

40-design.md already rejected gratuitous change for this reason: symmetry is not
worth a behavior change to 53 files that fixes nothing. That reasoning was
applied to POSIX and missed the win32 non-batch case. Rows 5 and 20 now record
it, and the CONTEXT.md glossary states the caller-choice rule.

deps.spawn deliberately still adopts the resolved path: its hasBinary probe
answers from the same resolver, so probe and spawn must agree on the exact file
(#3445). The asymmetry is now documented at both call sites rather than latent.

E7 pins the restored contract and was verified by executing execTool against a
monkeypatched spawnSync: python3 in, python3 spawned.

Refs #3411

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:12:44 -04:00