Commit Graph

5831 Commits

Author SHA1 Message Date
sim
cd58aaabf4 refactor(#4653): drain the containment duplicates and record the two rulings
Phase 3 of epic #4636, stage 3c. ADR-4650 decision 6: a wrapper may decide HOW
to degrade, never WHETHER a path is contained. Four implementations are drained
on that rule; two are retained, with the reasons recorded rather than assumed.

DRAINED — the containment decision now comes from the canonical predicate:

  scripts/check-glossary-refs.cjs   local isWithinRoot deleted outright.
  src/installer-migrations.cts      ensureInsideConfig keeps its throw and its
                                    lexical fullPath; only the decision moves.
  src/planning-inspect.cts          isPathContained keeps must-exist as its own
                                    condition; only the decision moves.

Two of those are wrappers rather than deletions, and each is a wrapper for a
reason that would have been a silent behavior change if collapsed naively:

- `isPathContained` returns FALSE for a path that does not exist, because
  fs.realpathSync throws ENOENT and its catch swallows it. The canonical
  predicate does the opposite: for a missing target it walks up to the nearest
  existing ancestor and ACCEPTS a not-yet-created path under the root. Its
  callers at planning-inspect.cts:747 and :839 guard a phaseDir immediately
  before readdirSync, so under a naive swap a missing phaseDir would stop
  reporting scope UNREADABLE and start throwing ENOENT out of readdirSync.
  Existence is therefore kept as an explicit local requirement.

- `ensureInsideConfig` returns a LEXICAL fullPath that both callers consume for
  existsSync and for journal entries. The canonical predicate realpath-resolves,
  so if configDir is itself a symlink the two differ. The decision is canonical;
  the returned value stays lexical. Its message is likewise preserved verbatim,
  which is why this uses tryWithinRoot plus an explicit throw rather than
  assertWithinRoot.

`isWithinRoot` in planning-inspect is left in place and documented: it is a pure
comparison over paths the CALLER has already resolved, which readDocument does
inline specifically to keep a third degradation shape (exists-but-unreadable vs
absent) that neither isPathContained nor the canonical predicate expresses. It
is the comparison step of one implementation, not a second implementation.

RETAINED, DELIBERATELY — gsd-core/bin/gsd-tools.cjs. My own design document said
"collapse" and that was wrong. The file carries an explicit comment forbidding
it, and the comment is correct: its three checks reject symlinks OUTRIGHT, which
is strictly stricter than the canonical predicate, not a reimplementation of it.
The canonical predicate accepts a link whose target lands inside the root — for
a restore that is still wrong, because writing through the link overwrites
whatever it points at instead of materializing a regular file. Collapsing would
have reintroduced that hole. The comment is updated to name the current exported
predicate, to record that this was reviewed under this phase and deliberately
not collapsed, and to note that isInsideDir treats target === root as NOT
contained — the one implementation in the repo that does.

THE configHome RULING — retained lexical, and a false safety claim corrected.
isPathConfined stays lexical because two of its callers must validate a
destSubpath BEFORE the mkdirSync that creates it (install-engine.cts:1608,
install-profiles.cts:880), where realpath cannot resolve and a realpath-based
predicate would reject every legitimate install.

Its docstring's justification, however, did not survive being checked. It cited
capability-source.cts:491,577,675 as the upstream symlink rejection that made
the lexical form safe. Read directly: :491 is a blank line before assertSafeId's
JSDoc and :577 is an entry-count budget check. Neither is a symlink check. The
real guards are :585-586 and :671-674. Worse than stale line numbers, the claim
that this "keeps every caller of this function's callers symlink-safe" is false:
that rejection lives in capability-source's staging path and covers only the
capability-loader route to assertDescriptorConfined. Three other callers do not
reach it, and only retired-artifact-cleanup.cts:69 carries its own defense
(its lstatSync check at :77). The docstring now states what is actually true and
cites the lines that actually exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 13:29:54 -04:00
sim
7e7a239a65 refactor(#4653): replace the allowAbsolute flag with a named acceptance policy
Phase 3 of epic #4636, stage 3b. Satisfies #4653's criterion that
`opts.allowAbsolute` become "a named acceptance policy on the predicate, not a
per-call-site boolean".

The flag was actively misleading at the call site. `{ allowAbsolute: true }`
reads as "containment is relaxed here". It never was: an absolute path that
resolves outside the root is rejected exactly as a traversal is. The flag only
ever controlled whether an absolute candidate was CONSIDERED. On a security
predicate that is the wrong thing for a reviewer to have to infer, and 31 call
sites were asking them to infer it.

    PathAcceptance.RelativeOnly         relative candidates only
    PathAcceptance.AbsoluteInsideRoot   absolute accepted, containment unchanged

The three exported wrappers take the policy and translate it inward.
validatePath keeps its internal `{ allowAbsolute }` opts and its body untouched —
the engine is not re-derived here either, only the exported surface is renamed.

MEASURED, NOT ESTIMATED. 31 call sites across 10 files, counted by walking the
AST with the repo's own @typescript-eslint/parser rather than grepping: a text
match would have folded in the options-type declaration, default parameter
values and comments. All 31 pass the literal `true`; none passes `false` or a
dynamic value, so the migration is uniform and `RelativeOnly` is purely the
existing default made nameable. audit.cts alone holds 18 of them.

This migration is compiler-verified in a way the containment-value migration in
the previous commit was not: the parameter type changed from an object to a
string union, so any missed site is a build error rather than a silent
behavioral difference. That is why a 31-site mechanical edit is acceptable in
the phase whose stated risk is the width of mechanical change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 12:31:01 -04:00
sim
26384ca988 refactor(#4653): make validatePath module-internal
Phase 3 of epic #4636, stage 3a. ADR-4650 decision 2: the engine stops being a
public shape. The only exported containment surface is now assertWithinRoot /
tryWithinRoot / requireSafePath, none of which can hand a caller a usable path
when the answer is unsafe.

The src/security.cts diff is one keyword. The engine body is byte-identical —
the dangling-symlink existence-oracle closure, the ancestor canonicalization and
the separator-aware boundary test are untouched, which is the whole constraint
this phase operates under.

WHAT THE TRANSLATION COST, AND THE RULE THAT KEPT IT AT ZERO. Roughly thirty test
call sites consumed validatePath directly, including the two BLOCKER regressions
that are this refactor's safety net. Translating them all to
`tryWithinRoot(...) === null` would have looked correct and silently destroyed
one of them: BLOCKER-1 asserts the rejection reason contains "unresolvable
symbolic link", which is what distinguishes a DANGLING symlink from an ordinary
escape. tryWithinRoot returns a bare null and cannot tell those apart, so that
assertion would have degenerated into "it failed somehow" — and the
existence-oracle closure could regress with the test still green.

So the rule applied throughout is: an assertion on the rejection REASON goes
through assertWithinRoot, whose throw carries the engine's message verbatim; only
assertions on the boolean go through tryWithinRoot. Under that rule no coverage
is lost. BLOCKER-1 still pins "unresolvable symbolic link" and BLOCKER-2 still
pins the exact canonicalized resolved value.

Three success-path tests came out BETTER than they went in. They previously
carried `expected safe:true, got error: ${result.error}` as an assertion message;
routing them through assertWithinRoot means an engine regression now surfaces the
real reason in the failure itself rather than as a hand-built string.

The two describe blocks named after validatePath are renamed — a block named for
a symbol the module no longer exports is a false signpost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 12:22:22 -04:00
sim
4ebcec7750 refactor(#4653): narrow the containment export and migrate all 15 validatePath sites
Phase 3 of epic #4636, stages 1-2. Implements ADR-4650 decisions 1 and 2: one
containment predicate, and an exported shape that cannot hand a caller a usable
path when the answer is unsafe.

THE EXPORT IS NOW A PAIR, BOTH RETURNING A BRANDED TYPE:

  assertWithinRoot(candidate, root, label?, opts?) -> ContainedPath   (throws)
  tryWithinRoot(candidate, root, opts?)            -> ContainedPath | null

ADR-4650 names only the throwing form. That does not survive contact with the
call sites: findPhaseArtifact probes a direct path, then a .planning/ path, then
each readdir entry, and throwing on the first miss breaks it outright. Six of the
fifteen sites need a non-throwing check. Recorded here rather than papered over.

WHY BRANDED. validatePath returns { safe, resolved, error } and populates
resolved with the ESCAPING path on the traversal branch — so a caller who skips
the boolean gets an attacker-controlled value precisely in the dangerous case.
tryWithinRoot returns exactly null there; assertWithinRoot throws. A plain string
is not assignable to ContainedPath, so a migrated site that validates one path
and then passes a different one is now a type error rather than a silent bug.
That is the defect this epic exists to close, and I introduced it twice in
Phase 2.

THE ENGINE IS UNTOUCHED. validatePath's body is not re-derived — the diff shows
zero edits to the dangling-symlink existence-oracle closure, the ancestor
canonicalization (macOS /var vs /private/var), or the separator-aware boundary
test. Each was acquired as a bug fix and a re-derivation would silently lose one.
requireSafePath now delegates to assertWithinRoot, so there is one implementation
beneath both names; its return type is branded, which is why its 13 call sites
compile unchanged.

A TYPESCRIPT LIMITATION, FIXED AT THE ROOT RATHER THAN WORKED AROUND. TS applies
never-return control-flow narrowing only when the callee is a function
declaration or a const with an EXPLICIT type annotation. Both routers do
`const { error } = io` — destructured, unannotated — so `error(...)` did not
narrow ContainedPath | null and four sites wanted a dead
`throw new Error('unreachable')` after it. Annotating the const
(`const error: typeof io.error = io.error`) makes TS narrow properly and the dead
throws are gone.

That annotation has a large, deliberate consequence: with narrowing working,
`@typescript-eslint/no-unnecessary-type-assertion` fires at 37 sites in
commands.cts where `as string` / `!` existed ONLY to paper over the missing
narrowing. They are removed. The rule is type-aware and fires only where the
assertion changes nothing, and both forms erase at compile time, so the emitted
behavior is unchanged — but the module loses 37 unchecked casts over
string | undefined, which is the same class of "trust me" the containment work
is removing. Widening the diff here buys that.

TWO SITES LOSE DIAGNOSTIC TEXT, deliberately. tryWithinRoot has no error channel,
so init.cts's agent-skills warning and cmdPrSubrepo's rejection now name the
condition rather than echoing validatePath's message. verify.cts had already
stopped echoing it on purpose — the message embeds absolute host paths — so this
makes the three agree instead of two-of-three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 12:16:24 -04:00
sim
889c7efba0 test(#4653): failing-first coverage for the narrowed containment export
Phase 3 of epic #4636. Tests only; no implementation. These MUST fail.

A refactor changes what a good test looks like: the behavior under test must be
IDENTICAL before and after, so most of this phase's safety comes from
invariance rather than new assertions. That safety net already exists and is
untouched here — tests/security.test.cjs already pins the two engine behaviors
a re-derivation would silently lose:

  :177  a DANGLING symlink to a non-existent OUTSIDE target stays safe:false
        (the existence-oracle closure)
  :213  a not-yet-created file in a not-yet-created subdir under a
        non-canonical base stays safe:true (ancestor canonicalization)

plus traversal, absolute in/out, null bytes, empty, non-string, and
requireSafePath's throw. Those 0 deletions are the point: if any of them had to
change, the engine would have changed, and the engine is not supposed to.

What is new is the export surface Phase 3 introduces:

  assertWithinRoot(candidate, root, label?, opts?) -> ContainedPath  (throws)
  tryWithinRoot(candidate, root, opts?)            -> ContainedPath | null

Two shapes rather than one, because several call sites need a NON-throwing
check — findPhaseArtifact probes a direct path, then a .planning/ path, then
each readdir entry, and throwing on the first miss would break it outright.
ADR-4650 names only the throwing form; this is the gap between the ADR and the
call sites, recorded rather than papered over.

Rows that exist because they are the ones nobody enumerates:

- tryWithinRoot must return EXACTLY null on escape, and its return must not
  contain the escaping path's basename. The shape being replaced populates its
  "resolved" field with the escaping path precisely on the traversal branch, so
  a caller who ignores the boolean gets a usable attacker-controlled value.
  That is the defect the narrowing exists to remove, so it is asserted
  directly.
- A seeded parity property: tryWithinRoot returns non-null if and only if
  assertWithinRoot does not throw, and the values agree. Two exported shapes
  over one engine is a divergence pair by construction.
- The rejection text still contains the phrase "escapes allowed directory".
  Another suite surfaces it through a user-facing "reason" field, and a
  refactor is exactly where wording drifts unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 11:54:49 -04:00
Tom Boucher
a0270a7945 Merge pull request #4666 from open-gsd/fix/4652-containment-at-boundaries 2026-09-12 11:44:24 -04:00
Rezolv
7bcfbe4542 docs(#4629): add ADR-4629 — STATE.md write intent beyond frontmatter (#4645) 2026-09-12 11:43:16 -04:00
sim
e1f169a378 chore(#4652): backfill changeset PR number to 4666
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 10:41:11 -04:00
sim
f64f8a0e7b test(#4652): correct the absolute-filename test to match the basename guard
The test asserted that an absolute filename is folded under the pending dir
and fails as "not found" rather than being rejected. That held for exactly
one commit. The basename guard rejects any name containing a separator before
any join happens, so an absolute path never reaches containment or the
filesystem at all.

Now asserts the USAGE rejection the CLI actually emits, verified by running it.
All four outside-file protections are kept unchanged — the file still exists,
its content is byte-identical, it never lands in completed/, and cleanup runs
in finally. Those are the assertions that carry the security value; only the
claim about HOW the rejection happens was stale.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 10:24:48 -04:00
sim
c7956c162b fix(#4652): a todo name must be a basename — containment alone cannot say that
The verification checkpoint caught a real design gap, not a flaky test.

Confining sourcePath/targetPath within todosRoot correctly rejects
`../../escaped`, which leaves the root. It does NOT reject these, because they
all land inside it:

  ../sibling.md   -> <todosRoot>/sibling.md     escapes pending/, not the root
  a/../../b.md    -> <todosRoot>/b.md           same
  sub/name.md     -> <pendingDir>/sub/name.md   inside pending/, but nested

Three committed tests asserted these must be rejected and were right: the
design says "a todo name is a basename, not a path", and #4327 requires the
resolved path stay inside "the todos root (pending and completed subdirs)".
Containment against a root is structurally incapable of expressing "basename" —
it answers "is this inside?", and all three are. The wrong tool was reaching
for the wrong question.

A basename guard now runs BEFORE any path is joined: reject on a `/` or `\`
separator, on a path.basename / path.win32.basename mismatch, on `.` / `..`,
and on a NUL byte. Both separators are checked explicitly because on POSIX a
literal backslash is an ordinary filename character to path.basename but not to
path.win32.basename or to the user's intent — this repo has a documented bug
class for exactly that asymmetry. Same predicate shape as findPhaseArtifact in
check-command-router.cts, so the two agree.

Containment is kept as defense-in-depth rather than replaced. The basename
guard is the specific rule; containment is the backstop.

Message wording matters here and is deliberate: `sub/name.md` does NOT escape
its allowed directory, so reusing the escape message would have stated
something false. It now says the name must be a plain filename, not a path.

docs/CLI-TOOLS.md corrected again, in the opposite direction from last time.
The previous revision said an absolute filename is "folded under the root" and
404s — true then, false now: the basename guard rejects it before any join
happens. Two corrections to one paragraph in one phase is the cost of
documenting behavior while it is still moving; the paragraph now matches the
shipped code.

Verified through the real CLI, not by calling the built function directly:
all eight rejection cases produce the new USAGE message; `ok.md` still
completes and moves to completed/; `missing.md` still gives "Todo not found".

Also corrected the now-stale comment above the isFile() check — it described
`.`/`..` reaching that line, which the basename guard now prevents. The check
itself stays: a bare basename can still name a directory, FIFO or socket in
pending/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 10:07:40 -04:00
sim
9c20b7b40a fix(#4652): use the validated path, collapse the duplication, correct two false claims
Seven findings from the two-axis review, all fixed in place.

THE ONE THAT MATTERS: cmdTodoComplete validated sourcePath and targetPath and
then ran every fs call against the RAW strings — existsSync, statSync,
readFileSync, platformWriteSync, unlinkSync, and the dry-run path payload —
never sourceCheck.resolved / targetCheck.resolved. That is the exact
"validate one path, use another" shape ADR-4650 names as the defect this epic
exists to prevent, and it is the same bug this phase had just fixed in
check-command-router. Committed inside the fix for it. All I/O now uses the
resolved paths; user-facing messages still echo the raw filename, never a
resolved absolute path.

A VACUOUS TEST, and the false doc claim it was propping up. The test
"[RED #4327] an absolute path outside the project is rejected" would have
passed with ZERO containment logic: path.join(pendingDir, '/abs/outside/x')
yields <pendingDir>/abs/outside/x — Node does not let a later absolute segment
escape — so the name is FOLDED under the root, passes containment, and simply
404s. The test only ever observed "Todo not found". It now asserts what is
actually true and actually valuable: an absolute name is neutralized, and the
real outside file is not read, not moved, and still present afterward.
docs/CLI-TOOLS.md claimed such a path "is rejected as a usage error", which
was false; it now describes the fold-under-root behavior. Traversal and
embedded separators ARE rejected, and those claims stand.

DUPLICATION THIS EPIC EXISTS TO REMOVE. resolvePath already did
isAbsolute-or-join + validatePath + reject; cmdGapAnalysisPlanPost and
cmdCheckPredicate each re-inlined the identical triplet in the same file. Both
now call resolvePath. Cost, stated rather than hidden: its generic message
replaces the two sites' distinct "phase-dir escapes…" wording. The message
still names the offending input, and one predicate with one message is the
point.

SYMLINK COVERAGE was required by #4652's "Done when" and was missing. Added
for both the todos root and --phase-dir, skipping cleanly on EPERM so the
Windows lanes do not fail where unprivileged symlink creation is disallowed.

Both fast-check properties were UNSEEDED. Seeded now.

The changeset named "check decision-coverage-plan" as a boundary; that is a
caller of the shared resolvePath, which the body never mentioned. Corrected.

DISCLOSED, not hidden: ctx.phaseDir is now always the resolved ABSOLUTE path,
so ${PHASE_DIR} interpolation and the "not found in <targetDir>" message show
an absolute value where a relative --phase-dir previously produced a relative
one. That is an observable output change. A test pins it and
docs/reference/gate-predicates.md states it.

Also regenerated scripts/lib/platform-conformance-tier.generated.cjs and its
macos twin — the new tests changed check-predicate.test.cjs's tier
classification. Caught by npm run lint:ci locally rather than by a bench run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 09:43:50 -04:00
sim
374300da17 fix(#4652): confine every boundary that joins argv to a managed root
Phase 2 of epic #4636, absorbing #4327 and #4354. Implements ADR-4650
decision 3: containment is a boundary concern — the predicate runs where
external input enters, not at whichever interior call site remembered.

Four boundaries now validate against their managed root and reject with a
USAGE-shaped error before touching the filesystem:

  todo complete <name>                 -> todosDir(cwd)
  check predicate --phase-dir <dir>    -> projectDir
  check decision-coverage-plan <dir>   -> projectDir   (via resolvePath)
  check gap-analysis.plan-post <dir>   -> projectDir

#4327 understated its own severity. It reports that a traversal name
"resolves outside the todos root", which reads as an information leak.
Measured, it was destructive: the command exited 0, MOVED the outside file
into completed/, and unlinked the original. cmdTodoComplete ends in
fs.unlinkSync(sourcePath), so an unconfined name consumed across the
boundary rather than merely reading across it. Validation now precedes every
fs call — existsSync, readFileSync, ensureDir, writeSync, unlinkSync — and
both halves of the move are confined, so neither source nor destination can
land outside the root. --dry-run is rejected on the same terms; a preview
must not leak a resolved outside path either.

#4354 reproduces exactly: a BLOCKING gate returned block:false sourced
entirely from a SECURITY.md in a caller-chosen directory outside the project.

THE HARDER HALF, found by the isolated adversarial review of the first
attempt: validating a path and then using a DIFFERENT one closes nothing.
The first fix validated `--phase-dir` joined against `--cwd`, then passed the
RAW unjoined value into the predicate context. gate-predicate-evaluator uses
it as-is and findPhaseArtifact resolves a relative path against the REAL
process cwd — so validation and the read used two different roots whenever
process.cwd() differed from --cwd. Reproduced: running from a directory
holding a plan with `secret_field: LEAKED_VALUE`, a predicate declared
against an empty --cwd project exited 0 and returned "actual":"LEAKED_VALUE".

The rule now applied at all three router sites: **use the validated resolved
path, never the raw input.** Independently re-verified after the fix — the
lookup resolves in the --cwd project and no value leaks.

gate-predicate-evaluator.cts is untouched and still imports no fs. Confining
in the router is what keeps that pure-leaf contract intact AND covers
${PHASE_DIR} interpolation into command-exit-zero, which an evaluator-local
fix would have missed entirely.

Also fixed, same review: `todo complete .` and `..` passed containment
(they resolve to the pending dir, which IS inside the root) and then threw an
uncaught EISDIR with an absolute-path stack trace. Now a clean USAGE
rejection naming the real reason — "todo name is not a file" — rather than
borrowing the escape message, which would have stated something false.

Ripples discharged BEFORE the verification checkpoint rather than after, per
the Phase 1 retrospective: docs/reference/gate-predicates.md and
docs/CLI-TOOLS.md document the new constraints, CONTEXT.md records why
containment lives at the router rather than the evaluator, the changeset is
written, and the install-tree goldens were regenerated to confirm unchanged
(no new shipped file) rather than assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 09:43:50 -04:00
sim
3925839f2a test(#4652): failing-first coverage for the four unconfined boundaries
Phase 2 of epic #4636, absorbing #4327 and #4354. Tests only; no fix. These
MUST fail.

Four CLI boundaries join externally-supplied input to a managed root with no
containment validation. Each was driven through the real CLI and confirmed
unconfined before the assertions were written:

  todo complete <name>                     src/commands.cts cmdTodoComplete
  check predicate --phase-dir <dir>        check-command-router cmdCheckPredicate
  check decision-coverage-plan <dir>       check-command-router resolvePath
  check gap-analysis.plan-post <dir>       check-command-router

Boundary 1 is worse than the issue describes. #4327 reports that a traversal
name "resolves outside the todos root", which reads as an information leak.
Measured, it is destructive: `todo complete ../../../../b1out/leak.md` exited
0, MOVED the outside file into completed/, and unlinked the original. The file
was gone. cmdTodoComplete ends in fs.unlinkSync(sourcePath), so an unconfined
name does not merely read across the boundary, it consumes across it.

Boundary 2 reproduces #4354 exactly: a BLOCKING gate returned
{"block":false,"details":{"match":true}} sourced entirely from a SECURITY.md
in a directory the caller chose, outside the project.

Boundaries 3 and 4 are not named in the epic. Both accepted an outside phase
dir and exited 0.

Rows that exist because they are the ones nobody enumerates:

- ORDERING. A real file is created outside the todos root, then the traversal
  name targeting it is asserted rejected AND the outside file asserted still
  present and unmoved. #4327 notes the existence check and the move target
  BOTH follow the unvalidated join, so a rejection that lands after the read
  has already leaked — and, per the finding above, after the unlink has
  already destroyed.
- `a/../../b.md` — looks balanced, resolves outside.
- --dry-run must reject too; a preview must not leak a resolved outside path.
- ${PHASE_DIR} interpolation into a command-exit-zero predicate is the SECOND
  predicate kind, which a fix inside gate-predicate-evaluator.cts would miss.
- An absolute path INSIDE the project must still be accepted at every
  boundary — absolute is not a synonym for escaping.

Cross-boundary rows loop over one shared list of escaping inputs and assert
all four reject with the same shape, so four sites adopting one predicate
cannot drift into four rejection contracts.

Property tests cover BOTH directions — outside is always rejected, inside is
always accepted. A property asserting only rejection is satisfied by a
predicate that rejects everything, which is the degenerate-implementation trap
found in Phase 1's review. Both are seeded.

Regressions fold into the owning module suites rather than a new
tests/fix-NNNN-*.test.cjs, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 09:43:50 -04:00
Tom Boucher
c99d7bb2be test(#4522): migrate core CLI/domain state batch to named timeout constants (#4662)
Batch 11 of the ad hoc timeout literal migration (epic #4445). Replaces
every bare numeric timeout/timeoutMs object-literal property in
tests/state-document.test.cjs, tests/phase.test.cjs, tests/commands.test.cjs,
tests/pattern.test.cjs, tests/adr-612-bracket-coherence.test.cjs,
tests/adr-612-bracket-read-tolerance.test.cjs, tests/milestone-lock.test.cjs,
tests/init.test.cjs, tests/state-todos-render.test.cjs,
tests/quick-batch.test.cjs, tests/graphify.test.cjs, and
tests/effort-surface-axis.test.cjs with a named constant, per
eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 12 files from the
rule's allowlist.

Ground truth via eslint found 25 sites, not the issue's stated 24 (phase.test.cjs
has 5, not 4) -- disclosed in the PR body.

Reuses PROBE_TIMEOUT_MS, GIT_TIMEOUT_MS, and LOOP_HOOK_POINT_CLI_TIMEOUT_MS
across 8 files. Adds two new shared constants to tests/helpers/timeouts.cjs
(each independently arrived at by 2 files in this batch, crossing the
promotion bar): PATHOLOGICAL_INPUT_TEST_TIMEOUT_MS (node:test's own per-test
timeout option, not a subprocess bound) and GSD_TOOLS_CLI_MODERATE_TIMEOUT_MS
(a single gsd-tools.cjs CLI subcommand spawn, distinct tier from
PROBE_TIMEOUT_MS/LOOP_HOOK_POINT_CLI_TIMEOUT_MS). Adds 3 file-local constants
for values used by only 1 file in this batch. No src/bin file touched, no
numeric value changed anywhere.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 09:12:31 -04:00
Tom Boucher
241646a43a fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification

Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail.

The guard classifies a name by comparing everything after `.env.` as one
token against a set whose members are FINAL EXTENSIONS. So `.env.local.example`
yields suffix `local.example`, which is not a member, and a committed
secret-free template is refused. That is a category error, not strictness.

Two arms are covered because the same classification is hand-rolled twice in
one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret`
(`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a
guard that allows `cat .env.local.example` while refusing
`Grep --glob '.env.local.example'` — the same file, the same hook, opposite
answers. A cross-arm parity loop over one shared list asserts the two cannot
drift.

Rows that exist because they are the ones nobody enumerates:

- `.env.example.local` must stay BLOCKED. Final extension is `local`; this is
  dotenv's documented local-override convention and a real secret. Any fix
  shaped as "contains example" admits it.
- `.env.local.` must stay BLOCKED — empty final extension is not a member.
- `.env.` must stay ALLOWED. Note #4580's proposed patch adds
  `if (suffix === '') return true;`, which flips it to blocked; that breaks the
  existing `allows` assertion in this suite and broadens the protected set,
  which epic #4636's non-goals forbid. Not applied.
- `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select
  `.env.local`, and a partial literal cannot be classified.
- `*.example` and `*` must stay ALLOWED — regression protection on the arm
  that already works.

Local behavioral repro of the current guard, confirming the tests fail for the
right reason rather than by construction:

  .env.local.example  rc=2 (blocked)   <- the defect
  .env.example        rc=0 (allowed)
  .env.local          rc=2 (blocked)
  .env.example.local  rc=2 (blocked)
  .env.               rc=0 (allowed)
  glob .env.local.example  rc=2        <- the second arm

Regressions are folded into the owning module's suite rather than a new
tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): classify by final extension so .env.<name>.example is readable

Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5.

The guard compared everything after `.env.` as ONE token against a set whose
members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`,
which is not a member, so a committed, secret-free template was refused — the
guard blocked the one file that exists so nobody has to open the real `.env`.

That is a category error, not strictness. The fix is not "add local.example to
the set"; it is to compare the right token. hooks/lib/filename-classification.js
now owns that distinction and is the only place it is expressed.

Both arms are fixed, because the same classification was hand-rolled twice in
this one file:

  - isSecretBasename (Read/Bash) now tests finalExtension(suffix).
  - globAltSelectsSecret (Grep --glob) split its first branch. With no
    wildcard the alternative IS a whole filename, so it is classified exactly
    via isSecretBasename. With a wildcard present the literal is only a
    PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and
    cannot be classified, so the original conservative rule stays.

Fixing only the first would have shipped a self-contradicting guard: `cat
.env.local.example` allowed while `Grep --glob '.env.local.example'` refused —
same file, same hook, opposite answers. A cross-arm parity loop over one shared
list now asserts the two cannot drift.

Two deliberate departures from #4580's suggested patch, both verified:

  - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.`
    from allowed to blocked, breaking an existing assertion in this suite and
    broadening the protected set, which epic #4636's non-goals forbid.
  - `fullSuffix` was drafted alongside finalExtension and removed before
    commit: zero production consumers, and none planned (Phases 2-4 are
    containment, duplicate draining and the path-join ratchet, none of which
    classify filenames). A zero-caller export is dead code. The distinction is
    pinned instead by a test asserting finalExtension('local.example') is
    'example' and explicitly NOT 'local.example'.

The protected set is unchanged. `.env.example.local` stays BLOCKED — its final
extension is `local`, dotenv's local-override convention and a real secret;
any fix shaped as "contains example" admits it.

Scoped out by measurement, not assumption: src/validate.cts:395 and
src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase
identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding
them in would repeat this same category error in the opposite direction.

Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed,
26 failures / 45330, all 26 in the two new test files, zero pre-existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): document the widened template exemption and cover the Bash arm

Two findings from the isolated adversarial review, both fixed in place.

1. The header's "Stated cost" passage named only the four literal template
   names, but since this change the exemption keys on the FINAL EXTENSION, so
   the trusted set is `.env.<anything>.{example,sample,template,dist}` — an
   unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example`
   is allowed. That is the deliberate and necessary cost of fixing #4580, but
   it was materially larger than what the header disclosed, and a silent
   expansion of a security guard's trusted set is not acceptable. The passage
   now states the family, the concrete bypass, and that it applies across
   Read, Grep and Bash alike.

2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but
   not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely
   distinct. The Bash arm was covered only by two one-off tests outside the
   shared table, so the table could not have caught a drift there. The loop now
   drives all three arms from the same TEMPLATES/SECRETS arrays.

No classification logic changed. The Read-arm behavioral table is byte-identical
before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for
.env.local / .env.example.local / .env / .secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): treat trailing dots and spaces as aliases of the protected file

Closes a Windows path-alias bypass surfaced by the isolated adversarial review
of this phase. Maintainer-approved as in scope.

Win32 strips trailing dots and spaces from every path component, so `.env.`,
`.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve
to the real `.env` / `.secrets` on Windows. The guard allowed every one of them
— a bypass of a file it already protects, reachable from Read, Grep and Bash
alike. `isSecretBasename` now normalizes the basename before classifying.

The whole class is fixed, not the reported name. `.env.` alone would have left
`.secrets.` and the trailing-space forms open, which is the same
one-cause-explains-every-failure trap this epic exists to close.

Two consequences, both measured rather than assumed:

  - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted
    `.env.example` template, so this is correct; it was previously blocked only
    because the trailing dot broke final-extension parsing.
  - A Bash token that is exactly `.env` plus trailing whitespace flips
    allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive
    class: the bare `.env` token was ALREADY blocked as an operand in the same
    position before this change, so the alias now simply behaves like the thing
    it aliases.

The header's "No whitespace trimming" guarantee is preserved and now stated
precisely: leading and interior whitespace is still never trimmed, so prose
like a commit message mentioning `.env` in a sentence stays prose and stays
allowed. Only TRAILING dots and spaces are stripped. Two tests pin that.

This lands at the same behavior #4580's proposed `if (suffix === '') return
true;` would have produced for `.env.`, which this phase earlier rejected. The
rejection was correct on its stated grounds — that line broadens the protected
set, which epic #4636's non-goals forbid. The Windows framing is different:
normalizing an alias of an already-protected file is not a broadening, and the
fix is reached by normalization rather than by special-casing an empty suffix,
so it generalizes to `.secrets.` and the space forms.

Cannot be reproduced on this host — the remote matrix is Linux-only and Windows
coverage arrives from CI — so this ships on the Win32 path-normalization
contract plus the CI lane, and that limitation is stated rather than implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): one owner for path segmentation, closing a Read/Grep divergence

Four findings from the two-axis review, all fixed in place.

The real one: the guard had TWO path-segmentation rules. `lastSegment` (used
by Read and Bash via `namesSecret`) splits on both `/` and `\`, while
`classifyGrepGlob` hand-rolled its own on `/` only. Measured:

  Read  of `config\.env`      rc=2  BLOCKED
  Grep  --glob 'config\.env'  rc=0  ALLOWED

Same logical file, opposite answers — precisely the divergence this epic
exists to remove, sitting inside the file this phase was already fixing.
`lastSegment` now lives in hooks/lib/filename-classification.js and both arms
call it. All five path-bearing cases (both separators) now agree.

Note on how this was nearly missed: the first measurement of it reported
"both allow", which looked like the reviewer was wrong. That reading was a
measurement artifact — `config\.env` inside a printf'd JSON payload is an
invalid escape, so the hook fails open at rc=0 and the test was observing
JSON breakage rather than the predicate. Re-measured with correct escaping,
the divergence is real. The tests added here use properly escaped literals
and were verified by running, not by reasoning about the escaping.

Also fixed:

  - Both fast-check properties were satisfied by a degenerate
    always-return-'' implementation: "never contains a dot / is a suffix" and
    "never ends with dot-or-space / is a prefix" are both trivially true of
    the empty string. They now additionally pin content preservation — the
    removed tail must match /^[. ]*$/, and a name with nothing to strip must
    come back unchanged.
  - The cross-arm parity loop used only bare basenames, so it could not have
    caught the divergence above. It now covers path-bearing names with both
    separators.
  - That loop's description overclaimed: Read and Bash BOTH route through
    `namesSecret`, so they are not independent paths; only the Grep glob arm
    is genuinely separate. The description now says so rather than implying
    three-way independence.
  - `normalizeWindowsBasename` runs on every platform, not only Windows. Its
    doc now states that explicitly: the guard must answer identically
    everywhere, and a name is judged by what Win32 would resolve it to.

No classification logic changed; the 12-name regression sweep is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs

Three things, all consequences of the fix rather than new behavior.

1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED
   file — package.json `files` includes `hooks` — so every per-runtime install
   tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in
   tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via
   scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11
   failures one-for-one.

   This ripple was identified at design time and then not acted on. Fleet's
   impact preview named golden-install-tree.test.cjs before any code was
   written, and 40-design.md records it under "Ripples identified". Writing a
   risk down is not the same as discharging it, and a full matrix run was spent
   discovering something already known.

2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's
   protected set: it named `.env.example` / `.sample` / `.template` / `.dist`
   as the four exempt names. The exemption keys on the FINAL EXTENSION, so the
   exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`.
   The page now states that family, the widened residual, that order matters
   and only the last segment counts (`.env.example.local` is a secret), and
   that trailing dots and spaces are stripped because Windows resolves them to
   the protected file. A wrong user-facing model of what a security guard
   protects is worth correcting even though Fixed/Security changesets are
   exempt from the required-docs rule.

   docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as
   `.env.example` exempt" — non-exhaustive, still true, deliberately left
   alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same
   hedged phrasing; hand-translating a security description unreviewed is not
   something to do silently.

3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass
   closed is `Security`. Folding the second into the first would under-report
   it in the release notes. Both carry `pr: 0` for backfill once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): backfill changeset PR number to 4659

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:42:43 -04:00
Tom Boucher
8fa2c3dbcf docs(#4650): record the path-containment and filename-classification design lock (#4655)
Phase 0 of epic #4636. ADR-4650 fixes the decisions Phases 1-4 inherit, so that
four phases do not each invent them independently.

The load-bearing decision is that the engine and the exported shape are
separable. A resolver-based, symlink-safe containment predicate already exists
as validatePath, and building the epic's literal assertWithinRoot() from scratch
would create a sixth implementation of the very thing this epic consolidates --
while risking silent loss of behavior validatePath acquired as bug fixes (a
closed dangling-symlink existence oracle, ancestor canonicalization for
non-canonical roots, a separator-aware boundary test).

But the epic's other clause is correct and lands on the current export:
validatePath returns a boolean a caller can forget to check, and populates
`resolved` with the escaping path precisely on the traversal branch. While that
form stays exported, the Phase-4 ratchet could only assert that a helper was
called -- validatePath(x, root).resolved would pass the rule.

So: preserve the engine, narrow the export. assertWithinRoot becomes the only
export and yields a branded ContainedPath.

Also recorded, each found by measurement rather than from the epic text:

- The rejection message text is a real contract. tests/quick-batch.test.cjs
  asserts a user-facing `reason` field matches /escapes allowed directory/, so
  the string reaches CLI consumers and Phase 3 must preserve it verbatim.
- Two further unconfined boundaries the epic does not enumerate: resolvePath
  and gap-analysis.plan-post, both in check-command-router.cts.
- --phase-dir also interpolates into ${PHASE_DIR} for command-exit-zero, so
  confining at the boundary covers both predicate kinds; the evaluator stays
  fs-free.
- opts.allowAbsolute is a per-call-site liberality knob, which is an acceptance
  policy living exactly where this ADR says it must not.
- planning-inspect's isWithinRoot is deliberately pure-string with no I/O; its
  contract differs, so Phase 3 decides rather than assumes.

The acceptance policy is stated once: conservative about the resource, exact
about the classification. #4580's guard was not too strict, it was wrong -- a
category error comparing a whole suffix against a set of final extensions.

ADR opens as Proposed; ratified at Phase 4 closeout per docs/adr/README.md.
No changeset: the diff touches docs/adr/ only, which is outside
USER_FACING_PREFIXES in scripts/changeset/lint.cjs.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 23:45:04 -04:00
Tom Boucher
a2331c01f1 fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites

Manually traced against the current tree: the validating regex at
code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's
extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites

Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical
behavior for 1- and 2-segment ids, character class unchanged): code-review.md,
code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating
sites, plus their comment/error-message text), execute-plan.md's plan-filename
extraction, and plan-phase.md's --research-phase flag capture.

Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor,
which was matching an unrelated earlier `--research-phase` occurrence (line
77's generic-value capture) instead of the targeted site (line 131).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/

Adds findSingleSegmentPhaseRegexDrift, banning the bounded
`[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any
phase-carrying line across gsd-core/workflows/**/*.md,
gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned
the same way as the existing shell-arith rule. Wired into scanAll; confirmed
zero violations against the real tree post-#4568 fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4568): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

The emitted-attribution gate also flags 4 files growing: code-review-fix.md
(+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9
bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each
site's validation regex widened from a bounded single-optional-dotted-segment
shape to the unbounded form, and the accompanying comment/error-message text
grew by a few characters to mention the new 3-segment example.

Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4568): backfill changeset pr number to 4646

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:04:28 -04:00
Tom Boucher
4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00
Tom Boucher
db4d8a9bae fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic

$((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is
decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is
valid shell-arithmetic syntax at all, and the failed expansion aborts the
rest of the snippet in a non-interactive shell. safe_resume_gate runs
unconditionally before trusting STATE.md or dispatching any executor, so
execute-phase failed at its own gate before the first executor on any
decimal phase, regardless of workflow.tdd_mode. Regression from #4194.

Fixes all 4 sites: safe_resume_gate and the TDD gate in
workflows/execute-phase.md, the completion-signal spot-check fallback in
workflows/execute-phase/steps/completion-reconciliation.md, and the
executor gate validation example in references/tdd.md. Each now zero-strips
only the leading integer segment into a *_INT variable (via %%.* / #
parameter expansion — always valid shell syntax regardless of what follows)
and keeps the remainder as an escaped-dot string for the anchored commit-
scope regex, exactly as issue #4619 verified in both bash and zsh. A plain
integer phase (12, 01) computes byte-identically to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug

Behavioral coverage via real bash execution: the old $((10#01.1)) form
throws (characterizes the bug, matching the issue's own reproduction); the
new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain
integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches
feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):,
feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring
issue #4619's own verified table exactly. Updates
safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions
(one per site) to the new fixed text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic

With #4619's fix in place, the guard's original "ban $((10#... outright,
match any occurrence" was too blunt: it flagged a comment merely mentioning
the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an
already-%%.*-stripped integer, and the always-safe plan-id arithmetic
(plan ids are plain integers, never decimal). Refines the detector to skip
full-line comments and to only flag a captured variable/placeholder name
that contains "phase" and does NOT end in _INT/_int — the naming convention
the #4619 fix establishes at all four sites for "already reduced to a safe
integer." A plan-id variable was never phase-number arithmetic in the first
place and is excluded on the same basis.

This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new
exemptions") and D7 ("a decimal and N-segment phase id survive an
end-to-end execute-phase selection without error") for real — the guard now
reports zero violations across all five .cts/.md rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): cover the plain-padded-integer near-miss matrix too

Review found the anchored-ERE near-miss coverage only exercised the
decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also
verifies the plain padded-integer case (01 -> PHASE_N=1) against its own
near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the
missing assertion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test

The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote
only 2 backslash characters in JS source, which single-quoted-string parsing
collapses to 1 real backslash at runtime -- but the workflow/reference files
actually contain 2 raw backslash bytes at that position (needed so bash's
${var//pattern/replacement} produces the correct single-backslash output).
Write 4 backslash characters in the JS source at all 4 occurrences so the
runtime string matches the files' real bytes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): refresh the committed compact-content benchmark baseline

The new PHASE_INT/PHASE_FRAC arithmetic lines added to
gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio
baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): note the safe_resume_gate arithmetic growth in the test header

The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846
bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED
block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a
decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading
integer segment via base-10 arithmetic instead of forcing the whole value
through $((10#...)) and hitting a hard shell syntax error on the first dot.

A blank line previously separated the Emitted-Drift-Ack-Growth trailer from
the Co-Authored-By trailer below it, which splits git's trailer-block
detection: only the last contiguous non-blank run of Key: Value lines at the
end of a commit message is recognized as trailers, so the growth ack was
silently read as ordinary body text and the differential-attribution gate
failed with the growth unacknowledged. Joining the two trailers into one
contiguous block fixes it.

Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4208): replace chmod-based restore-failure injection with a root-proof git shim

`tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated
an unwritable index via a `post-index-change` hook running `chmod a-w` on
the git dir. That relies on the OS enforcing the *owner's own* permission
bits against itself, which uid 0 (a routine identity inside this repo's
Docker-based gsd-test benches) does not: every DAC check short-circuits true
for root, so the write the chmod meant to block silently succeeds, the
restore comes back clean, and the disclosure/rollback behavior under test
never actually gets exercised.

This is CLAUDE.md's own named anti-pattern for I/O-failure injection
("Cross-platform test IO-failure injection" — chmod tricks fail under root
Docker/CI). It is confirmed as the actual root cause here, not a production
defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure
logic (added by #4253, merged just before this run) was hand-traced and
manually reproduced end to end on an unprivileged workstation against a
freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces
exactly the `staging_failed` + "could not be restored" / "could NOT be
restored during rollback" results both tests assert. The other
`post-index-change`-based tests in this file (a `sleep` to force a timeout;
a real `update-index` to flip a restored entry's mode) are unaffected
because neither depends on a permission check — consistent with only the
two chmod-based tests failing on the real remote run.

Replaces the chmod fixture with a fake `git` placed ahead of the real one on
PATH that fails only `update-index --add --cacheinfo` — the one call the
restore makes — unconditionally, regardless of privilege level. Every other
git invocation execs straight through to the real binary, so the rest of
each scenario (`rm --cached`, the restore's own `ls-files` verification,
etc.) is exercised exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): backfill changeset pr number to 4644

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI

Passing the script as a `-c "<script>"` argv element made it subject to
Windows' CreateProcess command-line argument encoding, which silently
dropped the escaped-dot backslashes before bash ever saw them (observed on
PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the
same script via stdin instead removes argv entirely from the transport, so
there is nothing for Windows to re-encode. POSIX behavior is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 15:47:14 -04:00
Tom Boucher
5e0a7b1b56 fix(#4433,#4569,#4126): consolidate the phase-identity seam at name-validity, allocation, and branch-slug (#4640)
* fix(#4433): apply the name-validity guard symmetrically to every milestone-name capture

extractMilestoneHeadingName already refused a punctuation-only captured name
(#4134), but its two sibling capture sites in getMilestoneInfo — the
STATE.md-anchored 🚧-bullet match and the no-STATE.md in-progress 🚧-bullet
fallback — skipped straight to a bare truthiness check, so a malformed bullet
whose only content past the version was punctuation passed through as a real
milestone name.

Extracts the existing inline /[\p{L}\p{N}]/u check into a single shared
hasNameableContent predicate and applies it at all three capture sites, so
the guard is one owner rather than a copy that happened to land at only one
of them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4433): pin the name-validity guard at all three milestone-name capture sites

Failing-first coverage for the hasNameableContent extraction: a
punctuation-only 🚧-bullet name must not surface as a real milestone name,
either on the STATE.md-anchored path or the no-STATE.md in-progress
fallback, while a real name (including a digits-only one) still resolves
COMPLETE exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): consolidate decimal-phase-number allocation into one function

cmdPhaseInsert allocated its next decimal sub-phase number by scanning only
on-disk phases/ directories and ### Phase N.M: headings, never the roadmap
summary checklist — so a decimal that existed only as a checklist bullet
(no heading yet, no on-disk directory yet) was invisible, and phase insert
could silently reallocate an already-used number. It also always nested one
level deeper under afterPhase, with no way to request a sibling.

cmdPhaseNextDecimal had its own separate, near-identical two-source scan
(missing the checklist source too) — the exact "duplicate implementations
kept in sync instead of deleted" pattern this issue exists to close.

Extracts scanExistingDecimalPhaseNumbers (directories + headings + checklist
bullets, in one place) and migrates both cmdPhaseInsert and
cmdPhaseNextDecimal onto it — deleting cmdPhaseNextDecimal's own copy rather
than patching it in parallel. Adds an allocation: 'nested' | 'sibling'
argument to cmdPhaseInsert (default 'nested', matching every existing
caller's behavior); a top-level phase with no existing decimal segment falls
back to nested since there is no sibling level to join. No CLI flag wires
'sibling' yet — that is a separate, disclosed follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): pin decimal-allocation coverage across phase insert and next-decimal

Failing-first coverage for scanExistingDecimalPhaseNumbers: a checklist-only
decimal must not be reallocated by phase insert; a decimal present in
heading, checklist, and on-disk directory simultaneously must count once;
an unrelated phase family's checklist bullet must not cross-pollute; and
phase next-decimal (migrated onto the same shared helper) must see a
checklist-only decimal too, closing the same gap in a second command.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the phase-id drift guard for name-validity and shell arithmetic

The epic's ratchet requirement: lint-phase-id-drift.cjs must cover the two
new predicates this PR introduces, and must also scan shell inside
gsd-core/workflows/**/*.md and gsd-core/references/**/*.md for
integer-coercing phase-number arithmetic ($((10#...)) and friends), which
neither the canonical TypeScript module nor a source-only lint can reach.

Adds findNameValidityDrift (bans re-deriving /[\p{L}\p{N}]/u outside
hasNameableContent's owner file) and findShellPhaseArithDrift +
scanMarkdownShellArith (bans $((10#...)) in workflow/reference markdown,
sanctioned via <!-- phase-id-owner: --> on the preceding line). scanRepo
keeps its existing, narrower contract (src/**/*.cts only) so the
already-passing "the live repo is clean" test is untouched; a new scanAll
merges both for the CLI's full report.

Running the guard directly against this tree correctly reports the 7
pre-existing #4619 shell sites (workflows/execute-phase.md x4,
workflows/execute-phase/steps/completion-reconciliation.md x2,
references/tdd.md x1) as violations — demonstrating the ratchet works, not
fixing them. #4619 is a live regression tracked and fixed separately; this
PR does not touch those markdown files. A characterization test pins the
current count of 7 so a future change to that number is investigated rather
than silently absorbed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4569): wire --sibling through phase insert's CLI so the argument is reachable

cmdPhaseInsert's allocation parameter had no CLI path to 'sibling' — shipped,
untested, unreachable code (code-review finding: a guaranteed surviving
mutant). Adds --sibling to phase insert's argument parsing, threads it
through, and documents the flag in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4569): exercise --sibling end-to-end through the real CLI

Confirms --sibling joins afterPhase's parent decimal level rather than
nesting, and falls back to nested when afterPhase has no existing decimal
segment (no sibling level to join).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4634): demonstrate the two new drift detectors end-to-end via a planted violation

The epic asks for the guard to be "demonstrated by watching it go red" on a
reintroduced copy. The two new detectors (name-validity, shell-arith) had
only unit-level fixture tests; mirrors the existing bracket-rule's
planted-violation-in-a-temp-tree test for both, proving they're actually
wired into scanRepo/scanMarkdownShellArith end-to-end, not just correct in
isolation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): consolidate the drift guard's own owner-sanction-check logic

Standards review flagged the "walk to nearest preceding non-blank line,
check for a phase-id-owner comment" logic as duplicated across all four
detector functions in a PR whose whole point is eliminating exactly that
pattern. Extracts isSanctionedByPrecedingComment, shared by all four;
behavior-preserving (verified: identical output before/after, same 7 known
#4619 violations, zero token/bracket/name-validity).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): add Fixed changeset for the name-validity guard and allocation consolidation

pr:0 placeholder — backfilled once the real PR number exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4126): consolidate branch-name slug substitution into one shared renderer

cmdCommit (commands.cts) and cmdInitExecutePhase (init.cts) each
independently implemented branch-name template substitution, and both
substituted the literal string 'phase' when phase_slug was empty or
undeliverable — producing a non-identifying branch name (gsd/phase-08-phase)
that contradicted the honestly-reported phase_slug: null in the same
payload. Same structural defect as the other three gaps in this epic: two
consumers reimplementing one concept independently instead of sharing an
owner.

Adds renderPhaseBranchName (src/phase-id.cts) as the sole owner: a real slug
substitutes normally; an empty/undeliverable one drops the {slug} token plus
one adjacent separator (collapsing/trimming the result) rather than
substituting a placeholder word, for the shipped default template and any
user-configured shape alike. Both call sites now delegate to it; the old
inline duplicates are deleted, not kept in sync. {project} substitution
stays a separate step in init.cts, unchanged, since it is a config-level
field with its own fallback contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4126): pin renderPhaseBranchName and both migrated call sites

Property-based coverage for the shared renderer's degrade-path invariant
(output, when non-null, never contains {slug} and never starts/ends with a
separator), plus example coverage for real-slug substitution, empty/null/
non-string slug, token position at either edge, a doubled-separator
template, and the only-{slug} -> null case. One regression test each in
commands.test.cjs and init.test.cjs confirms a phase with no derivable slug
no longer produces a branch name ending in the literal '-phase'.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: route scanExistingDecimalPhaseNumbers through the canonical enumeration owner

Caught by an actual gsd-test run, not a hypothesis: the new decimal-scan
helper (fix(#4569)) enumerated phases/ directories via a raw
fs.readdirSync, which the pre-existing phase-enumeration drift guard
(#3185/#3882) correctly flags as an unsanctioned re-derivation outside its
canonical owner (listAllPhaseDirs / isSentinelPhaseId). Ironic given this
epic's own thesis, and exactly why the guard exists: consolidating one seam
can reintroduce drift in an adjacent one if the new code doesn't route
through what's already there. Migrates the enumeration to listAllPhaseDirs;
identical decimal-detection output for every existing case.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend the drift guard for branch-slug fallback; fix a real regex bug

Adds the fourth detector the epic's ratchet section names ("both
branch-name sites"): bans a `.replace('{slug}', ... || 'phase')` call
outright, sanctioned via renderPhaseBranchName or a dedicated comment.
Wired into scanRepo (no per-file exemption — this is a banned anti-pattern
everywhere, not a grammar with one legitimate owner). Now that #4126's fix
(prior commit) has landed, scanRepo reports zero violations across all four
.cts-scanning rules, restoring the simple "the live repo is clean" assertion
instead of a pinned-known-count characterization.

Also fixes a real bug an actual gsd-test run caught: findNameValidityDrift's
regex didn't tolerate the doubled-backslash template-string form its own
test claimed to cover (0 !== 1) — widened to \{1,2} matching
TOKEN_DRIFT_RE's existing tolerance for the same two forms.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4126): document the {slug} degrade behavior; update changeset for the full seam

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: detectPhaseNumberFromFiles wrongly rejected bare, slug-less phase directories

Caught by an actual gsd-test run on the #4126 regression test, not a
hypothesis: a bare phase directory with no slug remainder (e.g.
.planning/phases/01/) has extractPhaseToken correctly return "01" — which is
simply identical to the directory name in that case, not its no-match
fallback. A stale `token !== phaseDir` check treated that equality as "no
numeric token found" and rejected it regardless, leaving phaseNum null and
silently skipping cmdCommit's phase-branching block entirely (the commit
proceeded on whatever branch was already checked out instead of the
phase branch).

phaseTokenShape.test(normalized) already excludes every genuine non-phase
case on its own: extractPhaseToken's real no-match fallback only fires for a
dirName that doesn't start with a digit or short letter+digit prefix, and
normalizePhaseName's leading-\d+ requirement rejects those regardless. The
equality check was redundant for real rejections and actively wrong for
bare-numeric directories.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: backfill changeset PR number to 4640

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 12:42:04 -04:00
0xdhx
4cc2a466b5 fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec

`cmdCommit`'s `--files` list can stage an addition but never a deletion:
the #2014 guard skips a missing explicit entry because the filesystem
cannot tell "moved away" from "not written yet". A caller that moves a
file therefore had two forms, both wrong — a directory entry records the
move but also commits every unrelated file in that directory (a
concurrent session's in-flight todo, in the unattended execute-phase
sweep), and a file entry leaves the old path's deletion dangling with the
todo tracked at both paths.

`--files-removed <paths>` is the caller-declared delete intent. Each entry
names a file, or a directory whose tracked-but-absent files are the
removals; those paths are staged with `git rm --cached` and join the
commit pathspec. `--files` keeps its skip-if-missing contract untouched.
A file entry still present on disk fails the commit closed with the
existing staging-failure rollback; a never-tracked path is a no-op.
`--files-removed` alone is a declared scope, not the unscoped .planning/
sweep.

The dispatcher previously folded every non-flag token after `--files`
into that list, so a second list flag could not exist; each list now
runs from its flag to the next `--` token.

The execute-phase todo sweep names the moved todos on both sides from
CLOSED[@], and cleanup's archive commit moves .planning/phases/ and
.planning/quick/ under --files-removed.

Fixes #4208

Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries

* chore(#4208): set changeset fragment pr to 4253

* fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard

Three CI failures, all consequences of this PR's own change.

1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the
   ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The
   three-line rationale comment plus the four-line array-building block
   added 318 bytes to a file that had only 141 of headroom on next.

   Move the rationale to docs/CLI-TOOLS.md -- which this PR already
   extends with the --files-removed contract, and which is where the
   ADR-857 gate wants call-site detail to live rather than in the host
   workflow -- and fold the array build onto one line. 93,577 -> 93,372.

2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415
   guarantee to its old MECHANISM: it regex-matched the literal
   .planning/todos/{completed,pending}/ directory pathspecs in the
   commit --files list. This PR deliberately replaced those with named
   files (a directory entry also committed an unrelated todo a
   concurrent session dropped in mid-close), so the guard failed on a
   change it should have accepted.

   Re-point it at the new mechanism without weakening it: assert the
   ADDED array reaches --files, the REMOVED array reaches
   --files-removed, STATE.md is still committed, and -- newly -- that
   the two arrays are built from $COMPLETED_DIR and $PENDING_DIR
   respectively. Verified by negative control: deleting
   --files-removed "${REMOVED[@]}" from the workflow still fails the
   test, so the #2415 regression remains caught.

Note for the merge queue: #4233 also grows execute-phase.md (+114). The
two are additive -- different regions, no textual conflict -- so with
both landed the file reaches ~93,486, over the 93,400 margin though
under the 93,600 hard ceiling. Whichever merges second will need to
reclaim ~86 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv

* fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin

Rebasing onto next surfaced the byte-gate collision flagged earlier on
this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354),
so this PR's +113 landed at 93,467 against the <= 93,400 margin in
tests/claude-orchestration.test.cjs.

Compact the close_phase_todos step this PR already edits -- drop the
PHASE_NUM indirection, fold the normaliser and the match guard, print
the closed list with one printf, shorten the step's prose -- without
touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the
plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no
longer spends any of next's 46 bytes of headroom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback

Review of #4253 found three Majors with one root cause: the removal
side judged presence by fs.lstatSync alone, where the addition side
already reads `git ls-files -v` state. Absence from the worktree is not
removal:

- a submodule gitlink (mode 160000) whose directory was deleted by hand
  lists like a file and was `rm --cached` with no .gitmodules cleanup;
- a skip-worktree path is never materialised by a cone-mode sparse
  checkout, so a directory entry over a sparse-excluded tree dropped
  that whole tree from the index;
- an assume-unchanged path's worktree state is not something git
  itself consults;
- an intent-to-add entry (`git add -N`) renders as a plain cached entry
  on the empty blob, yet nothing tracked exists to remove and no
  rollback can restore the flag.

The index listing now carries each entry's `ls-files -v -s` tag, mode
and stage. Only a plain cached (H), stage-0, non-gitlink entry is a
removal candidate; every other state is left alone under a directory
entry (exactly like a present file) and fails closed when named
directly, with the state in the error. "Named directly" is decided on
RESOLVED paths, not strings -- realpath of the longest existing prefix
with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a
symlinked spelling of the tree (macOS `/var` ->
`/private/var`, where `process.cwd()` is the real path and the caller's
absolute path is not -- CI on this round's first push) all resolve to the
same entry, where a string compare against git's cwd-relative output
silently took the directory polarity (pre-push review, driven; the
symlink case is driven with an aliased fixture directory). The enumeration's domain is what
`ls-files -v -s` can emit for an index entry, stated at the classifier.

The third Major -- on an unborn HEAD a successful `rm --cached` was
never rolled back when a later entry failed -- is fixed differently
from the review's suggestion. Pushing the path into stagedPaths would
put it on the commit pathspec, which a root commit refuses ("pathspec
did not match", driven), and `git reset -- <path>` cannot restore an
entry with no HEAD anyway. Instead every index entry this call removes
is recorded (mode, blob) before the `rm` and put back with
`update-index --cacheinfo` on rollback. That also restores a
caller-pre-staged blob at a removed path exactly, where a reset would
have silently replaced it with HEAD's version. The rollback is
best-effort, as the addition-side reset already was, and the docs say
so.

Eight tests: gitlink under a directory entry, named directly, and named
by absolute path; skip-worktree both forms; intent-to-add both forms;
assume-unchanged named; unborn-HEAD partial failure restores the
removal; pre-staged blob survives the rollback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step

Review nit on #4253: inserting the --files-removed rationale between the
original bash block and its closing fence left an empty ```bash``` pair
before </step>. Harmless at runtime, a formatting artifact of this PR's
own diff; removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): a boolean flag inside a commit path list no longer ends the list

Review minor on #4253: collectList stopped at the next `--` token, so a
positional wedged between a boolean flag and the next list flag
(`--files a --amend b --files-removed c`) was claimed by neither list
and silently dropped -- a regression in shape against the old
slice-to-end parse, which filtered `--` tokens and kept `b`. No current
call site interleaves that way, but the gap was real.

A list now runs to the next LIST flag (`--files` / `--files-removed`)
and skips boolean flags on the way, and a REPEATED list flag merges
its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse
did -- a first cut stopped at the repeat and dropped `b`, the same
silent-drop shape one level over (pre-post comment audit). The only
change #4208 makes to parsing is that a second list flag can exist.
Tests: STATE.md wedged between --no-verify and --files-removed lands
in the commit; both runs of a repeated --files reach it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* test(#4208): drive the reappearance window with a post-index-change hook

Review nit on #4253: the defensive re-check for a file recreated between
the absence test and `git rm --cached` -- the concurrent-session race
this PR's own changeset names -- had no test. git fires
post-index-change the moment `rm --cached` writes the index, so a hook
that copies the file back exactly then exercises the window
deterministically. The call reports staging_failed / "reappeared on
disk", commits nothing, and the rollback restores the removed entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): restore a staged removal when the call records nothing

A `git rm --cached` that succeeds mutates the index whether or not a commit
follows. Only the staging-failure rollback put those entries back, so a call
that reached `nothing_to_commit` reported no state change while the removal sat
staged -- riding along on the caller's next commit.

The review named the unborn-HEAD, removal-only shape. Keying on `headExists`
would have fixed half of it: the guard also fires with a real HEAD when the
removed path is index-only (added, never committed), because `diff HEAD` reads
clean with the path absent on both sides. Both shapes now restore, at both
`nothing_to_commit` exits. The failure exits are deliberately left alone --
they report a failure rather than no-change, and the addition side leaves its
own staged paths there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot

`cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had
inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)`
now owns the index-state classification, path canonicalisation and entry
recording, returning the pathspec entries and the recorded removals its caller
merges.

Pure motion: no branch, message or probe changed. Only the two accumulators
became local names, and `restoreRemovedEntries` stays with the caller because
the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the
extracted helper is 277.

(Figures corrected after publication: an earlier version of this message said
854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both
were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit
is 581, so this is above it, not below.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): property-test the two-list commit parser

RULESET.TESTS.property-based-testing asks a parser for at least one property
test asserting a domain invariant; `collectList` had only hand-picked examples,
one per shape a review round had already broken.

Hoisted it to module scope as `collectListFlagValues` and exported it in the
file's existing exported-for-tests convention -- a parser reachable only by
spawning the CLI can be tested one example at a time and no faster.

Three properties over generated argv: every positional lands in exactly the run
open at it whatever the flag order or count; no positional after the first list
flag is dropped or double-claimed; and with `--files-removed` absent the parse
equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run
ending at any `--` token, and a repeated list flag that does not merge -- each
of which the properties catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin cleanup.md's archive commit to --files-removed

execute-phase.md's rewrite is pinned by the #2415 guard in this file;
cleanup.md's equivalent was not, so reverting its routing would have been
caught by nothing -- the mechanism's unit tests never read this file and pass
either way.

Asserts the two archived directories are under --files-removed and NOT under
--files (where a directory entry sweeps in a concurrent session's in-flight
writes), and that the destinations and STATE.md stay on the additive half.
Controlled by restoring the pre-#4208 sweep, which fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin that a symlink to a directory is one tracked path

Review of #4253 read the `lstatSync(...).isDirectory()` test as a
symlink-following defect. Driving it says the opposite: git tracks the link as
a single blob (mode 120000) and does not traverse it, so the tracked paths
"under" it live at the real directory and were never named by the caller.
Following the link would stage those -- the directory sweep #4208 exists to
remove -- while the named entry still sat present on disk.

Pinned rather than changed, with the premise driven in the test body. Swapping
`lstatSync` for `statSync` -- the prescription as written -- fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit

The base range added `tests/benchmark-compact-content.test.cjs` and a committed
token baseline over the compacted workflows. This PR edits
`gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on
that entry and on the aggregate.

Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is
those two entries and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): report a removal the call could not put back

Round review of this round found the restore itself unchecked: the helper
ignored `update-index`'s exit code, so a FAILED restore still reported
`nothing_to_commit` -- the same false "no state changed" the restore exists to
prevent, surviving one level down on the restore-failure path.

It now returns a boolean. The two no-change exits report `staging_failed`
naming the paths left staged; the staging-failure rollback still ignores it,
deliberately, because it is already reporting a failure and an unwritable index
is usually the failure being reported.

Driven with a post-index-change hook that makes the git dir unwritable the
moment `rm --cached` lands, so the restore cannot take its lock. Reverting both
guards fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): disclose a removal the rollback could not restore

Round review refuted the reasoning behind leaving the rollback path's restore
unchecked. The claim was that this exit is already reporting a failure, so the
restore's result adds nothing. The counterexample is the ordinary case: the
reported failure is usually a DIFFERENT cause -- a contradictory declaration, a
reappeared path -- so a caller reading `failures` sees only that cause and
learns nothing about the removal still sitting in its index.

The rollback now appends a disclosure entry per un-restored removal, naming the
path. The reason and `file` still report the failure that caused the rollback;
the disclosure is additive.

Also moves the restore-failure test's chmod into a `finally`: `t.after` runs
AFTER the parent `afterEach`, so a throw before it left the fixture undeletable.

Both driven; reverting the disclosure fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): decide index state by observation, never by an exit code

The restore added two commits earlier keyed both its record decision and its
success verdict on git's exit code. An exit code answers "did the command
succeed", never "did the index change" -- execGit collapses a spawn timeout to
a non-zero exit, and a killed git can already have written the index. Round
review drove four failures from that one assumption, in both directions:

  - a failed `rm` still contributed an entry, so the rollback disclosed a
    removal that was never staged (stale index.lock);
  - a timed-out `rm` whose write DID land contributed none, so a real mutation
    was neither restored nor disclosed;
  - a timed-out `update-index` whose write landed reported failure, publishing
    a "could NOT be restored" disclosure that was false;
  - and the read-back that replaced it omitted `-z`, so core.quotePath rendered
    `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not
    restored -- the same quoting defect this PR already fixed for `preStaged`.

Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the
path: gone means this call owns the removal and records it; still there means
nothing was staged; a probe that cannot answer becomes its own failure entry
rather than an assumption. The restore verifies the same way, comparing the
WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a
path-only test accepts an entry that came back as something else.

The verdict is three-valued -- `restored` / `not-restored` / `unverified` --
and the unverified wording says the restore could not be VERIFIED rather than
that it failed. The rm's own failure is pushed ahead of any probe diagnostic so
a timed-out removal keeps `timed_out: true` and its own message as the reported
cause.

Five regression cases, each negative-controlled against the shape it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): treat a declared removal path as a path, not a pathspec

An index path handed back to git is parsed as a PATHSPEC, and the removal side
handed several back. Three driven harms, all of them the sweep-in this flag
exists to remove, arriving through the operand rather than through a directory
entry:

  - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it
    removed `peer.md` and `stays.md` too, only the declared entry was recorded,
    so the rollback restored one of three and the other two rode out as staged
    deletions the result disclosed nowhere;
  - the same name reached `git commit -- <paths>`, which globbed and committed
    an undeclared `M peer.md` alongside the declared removal;
  - and the intent-to-add probe (`diff --cached` over the path) matched a
    STAGED PEER instead of itself, so an `add -N` entry was misclassified as
    ordinary content, removed, and restored by `--cacheinfo` -- which cannot
    restore the intent flag. It came back as a real staged addition.

Every operand on this path is now `:(literal)`: the `rm`, both index probes,
the intent-to-add probe, the restore read-back, the entry-level `ls-files` /
`ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream
`ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep
whatever pathspec behaviour they have today; that is not this change's to
alter. `:(literal)` still resolves a directory to its descendants (driven), so
the directory form is unchanged.

Closes what an earlier cut of this commit declared as a residual: a filename
beginning with `:` is now removable end to end, because the commit pathspec no
longer reinterprets it.

Also fixes a MINOR from the same review: cleanup.md's contract test checked the
destinations' position relative to `--files-removed` but never that `--files`
was present at all, so deleting the flag still passed.

Un-literalising the seven sites fails three of the new tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): scope the rollback to the caller's own name space

Round review drove a rollback that destroyed the caller's own staged work. Two
causes, one of them pre-existing:

  - `git diff --cached` prints REPO-relative paths whatever the cwd, while
    `stagedPaths` holds the caller's cwd-relative names. In a project nested
    inside its repo (`<repo>/sub/.planning/...`) the two name spaces never
    intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`,
    and the reset unstaged a caller-staged deletion and modification that this
    call had never touched. `--relative` makes the two sets comparable, and is a
    no-op when the project IS the repo root. This governs the `--files` side too
    and predates this flag.
  - the rollback's `reset` was the last place a removal-derived name reached git
    as a bare pathspec; it takes `asPathspec` like every other site.

Driven on a nested fixture; dropping `--relative` fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): gate six fixtures that Windows cannot construct

CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two
primitives the new fixtures rely on do not exist on Windows, both driven on a
real Windows host rather than inferred:

  - a filename containing `*` or `:` cannot be created at all (`IOException` /
    `FileNotFoundException`), which is four of the pathspec fixtures;
  - `chmod` cannot make a directory unwritable — a write into a ReadOnly
    directory succeeds — so the two restore-failure fixtures cannot drive the
    failure they exist to drive.

Each is skipped on win32 with its measured reason, in the repo's existing
`{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The
behaviours they pin are platform-independent; only the fixtures are not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): build git's index-syntax path with forward slashes

The remaining Windows red was mine, not the platform's: `git rev-parse :<path>`
takes a forward-slash path, and `path.join` yields backslashes there, so git
rejected it as an ambiguous argument. The hook in the same test already used
the slash form.

Not gated — the behaviour it pins is portable; only the argument was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline against the rebased base

`next` moved the `new-project` split and the aggregate under this PR's
execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs
--write` so the only leaves differing from the base's copy are the
execute-phase split and the aggregate it feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

* chore(#4208): regenerate the macOS conformance tier for this PR's fixtures

`next` gained the macOS-specific conformance tier (#4593) after this branch
was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target
macos`) now selects `tests/commit-files-deletion.test.cjs` on the
`chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the
chmod-driven failed-restore cases and the symlink-to-directory case).
Regenerated with `--target macos --write`; the platform tier was already in
sync. The file was modified, not added, which is why the added-files check
did not surface it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 12:26:55 -04:00
Tom Boucher
9f6f0d27dd docs(#4625): record the Codex native supervisor adapter as out-of-scope (#4637)
Codex's native subagent status is a fixed two-value enum
(CollabAgentToolCallStatus::{InProgress, Completed}) with no custom status
field, so the proposal's executing/verifying/completed triad cannot be
rendered in that view by anyone. Records the decision, and preserves the
reachable path -- codex exec --json ThreadEvents persisted by #4624, then
read at execute:wave:pre/post by an EoS capability.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 10:55:51 -04:00
Michel Moreira
1f84f45ed5 fix(#4259): fold shell continuations before the T6 docs-parity site scan (#4423)
* fix(#4259): fold shell continuations before the T6 docs-parity site scan

The scan is $-anchored with [^\n]* on both sides of --grep=, so `git log` and
`--grep=` had to share a physical line. A backslash-continued derivation —
the natural way to write a git log carrying a long ERE — produced zero hits
and T6 passed on it.

Both generations of the assertion were defeated. The current anti-revert ban
let a wrapped site through outright; at v1.12.0, where T6 instead asserted
pattern conformance, a wrapped site was silently exempted from the very
checks written to catch the macOS \b-no-op class, so it could have carried
exactly the malformed pattern T6 exists to reject. A real candidate
implementation for #3926 wrapped its derivation, passed T6, and was caught
only by later manual review.

Fold the continuations before matching rather than widening the regex: the
assertion's message and its PHASE_SCOPE_NUM filter both assume one site is
one string, and a [\s\S]*? would run the scan across unrelated statements.
Correcting the input repairs everything built on the scan at once.

The fold uses [ \t]* after the newline rather than \s* — the shell's own
rule, and it cannot swallow a blank line and glue two unrelated statements.
The scan is hoisted to findGrepSites so the controls can drive it directly:
the continued form is caught, the same-line form still is, a benign --grep
stays clean, a wrapped unrelated assignment is not glued, a continuation does
not cross a blank line, and the live workflow files still report nothing — so
this lands without editing any workflow to appease it.

* fix(#4259): honor shell continuation boundaries

* test(#4259): use canonical fenced-block scanner

* test(#4259): centralize shell continuation scanning

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 10:40:37 -04:00
Tom Boucher
aad96e0b5f test(#4521): migrate capability subsystem batch to named timeout constants (#4627)
Batch 10 of the ad hoc timeout literal migration (epic #4445). Replaces
every bare numeric timeout/timeoutMs object-literal property in
tests/adr857-core-without-capabilities.test.cjs, tests/capability-cli.test.cjs,
tests/capability-probe-fallback.test.cjs, tests/capability-state.test.cjs,
tests/capability-trust.test.cjs,
tests/capability-validator-task-content-resolver.test.cjs, and
tests/capability-writer.test.cjs with a named constant, per
eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 7 files from the
rule's allowlist.

Reuses the existing PROBE_TIMEOUT_MS constant at 8 sites across 3 files.
Adds 7 new file-local constants (no promotion to the shared helper needed
this batch -- every new class is confined to exactly one file, below the
two-file promotion bar): GSD_TOOLS_CLI_TIMEOUT_MS,
FRAGMENT_PROBE_SNIPPET_TIMEOUT_MS, INSTALLED_RUNTIME_CLI_TIMEOUT_MS,
FIXTURE_MCP_SERVER_TIMEOUT_VALUE, TASK_RESOLVER_FIXTURE_TIMEOUT_MS,
TASK_RESOLVER_TIMEOUT_CEILING_MS, and TASK_RESOLVER_TIMEOUT_CEILING_PLUS_ONE_MS
(the last two forming a boundary-coverage limit/limit+1 pair). No src/bin
file touched, no numeric value changed anywhere.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 09:55:55 -04:00
0xdhx
7249bdddac fix(#4294): reserve a full progress bar for 100% and give the render half one owner (#4473)
* fix(#4294): reserve a full progress bar for 100% and give the render half one owner

Six call sites each carried `Math.round((percent / 100) * width)` inline, and
every copy rounded to a full bar before the percent reached 100: from 95 up at
width 10, from 98 up at width 20. A project at 19/20 plans drew the same bar as
a shipped one beside a number that said otherwise, and an out-of-range percent
threw `RangeError` from the unguarded `'░'.repeat`.

ADR-3180 Decision 7 gave the completion-RATIO derivation one owner
(`clampPercentFromFraction`). This gives the RENDER half the same:
`progressBarFilledCells` / `renderProgressBar` in phase-lifecycle.cts, with the
`progress` table and bar renderers, the stats renderer, the gsd2 import writer,
and #4231's `formatProgressMachineSegment` (which now serves both STATE.md
writers) all drawing through it.

Contract: below 100 the fill is held one cell short of the width, so only the
saturating percents move (95-99 at width 10, 98-99 at width 20) and every other
value in 0-100 renders as before — pinned by an exhaustive comparison against
the legacy formula at both widths. Null / non-finite renders an empty bar;
out-of-range is clamped, never thrown.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hbFn24VWaxJBw8DUWjAmU

* chore(#4294): set changeset fragment pr to 4473

* docs(#4294): correct the pre-fix inline call-site count to five

The kernel's doc comment said SIX call sites carried their own
`Math.round((percent / 100) * width)`. The base tree has five: three in
`commands.cts` plus one each in `gsd2-import.cts` and
`formatProgressMachineSegment`, the latter two using `/ 10` with the width
already substituted (`pct` and `clamped` respectively).

The six is #4294's count of consumers -- it counts `cmdStateUpdateProgress`
and `syncCore` separately, but #4231 had already routed both through
`formatProgressMachineSegment` (as it does `applyPostSyncPreservation`), so by
this branch's base they share one copy. The comment now states the tree's count
and records where the six comes from, so neither number reads as an error
later.

Comment-only; no behaviour change, and no change to compiled output.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 00:42:55 -04:00
Tom Boucher
523be34133 fix(#4282): register PATTERNS.md as a canonical .planning/ artifact (#4618)
* test(#4282): prove PATTERNS.md is unrecognized by the artifact registry

Regression test only, no fix yet: CANONICAL_EXACT in src/artifacts.cts was
never updated when workflows/graduation.md started writing .planning/
PATTERNS.md, same omission class as the already-fixed #3224 (WINDOWS.md).

Expected RED on this commit (src/artifacts.cts is unchanged).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4282): register PATTERNS.md as a canonical .planning/ artifact

CANONICAL_EXACT in src/artifacts.cts was never updated when
workflows/graduation.md started writing .planning/PATTERNS.md for the
`patterns` graduation-target category -- same omission class as the
already-fixed #3224 (WINDOWS.md). validate.health's W019 falsely flagged it
as unrecognized on every repo that has run the graduation scan.

Also backfilled 5 other pre-existing stale rows in
gsd-core/templates/README.md's artifact table (WINDOWS.md, STATE-ARCHIVE.md,
milestone.lock, state.json, skill-manifest.json) that were already in the
source registry but missing from the docs table -- found while fixing this
exact drift class, cheap to close alongside it.

RED proven on f3dd791fb8cda18196803e7144ce20e506d6490b (test-only commit,
gsd-test outcome:failed, exactly the new PATTERNS.md test failing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): fix stale function name in skill-manifest.json comment

Review finding: both the source comment and the new docs row said
"routeSkillManifest" -- no such symbol exists (verified via Memtrace); the
actual function is cmdSkillManifest (src/init.cts). Copied verbatim from a
pre-existing comment, not introduced by this PR, but cheap to fix alongside.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): add changeset fragment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4282): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: isolate lint-vendored-deps-manifest.test.cjs's fixRow tests from the real vendor file

Genuine, pre-existing defect found and fixed per this repo's no-defer policy
(discovered while investigating a real CI failure during this PR's own
merge attempt, user-directed investigation -- not deferred to a separate
issue since it was actively blocking work and root-caused with concrete
evidence, not speculation).

Root cause: fixRow(row) (scripts/lint-vendored-deps.cjs) unconditionally
does fs.copyFileSync(upstreamCjs, vendoredCjs) as its first line. All three
tests in the #4573 describe block called fixRow(row) with the REAL js-yaml
row, so all three wrote to the real, shared gsd-core/bin/lib/vendor/
js-yaml.cjs -- a file other test files' require() calls can read at any
moment, since node --test runs files concurrently in this repo.
fs.copyFileSync's write is not atomic against a concurrent reader on every
filesystem; a concurrent require() elsewhere caught the file mid-overwrite
and read a truncated file, crashing an entirely unrelated test
(m9-statelock-write-error-orphan.test.cjs) with a SyntaxError.

Confirmed via two real CI log fetches, not assumed: the exact same shard
grouping (same 308 files) ran clean ~90 minutes earlier during PR #4615's
own final merge CI, with the identical #3660 reap-fix code already present
-- ruling out a deterministic connection to that change and confirming a
genuine, non-deterministic timing race in this pre-existing test design.

Fix: all three tests now redirect row.vendoredCjs to a private os.tmpdir()
path via a cloned row object before calling fixRow, so the real vendored
file is never touched. upstreamCjs stays pointed at the real node_modules
copy (read-only, safe to share).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: also isolate fixRow's package.json pin-rewrite from the real file

Review finding (major) on the previous race-condition fix: fixRow's pin
-rewrite path still hardcoded path.join(ROOT, 'package.json'), so the third
#4573 test still wrote the real, shared package.json -- read at module
top-level by dozens of other test files, the same concurrent-file race
class already fixed for the vendored .cjs copy.

Adds an optional pkgRoot parameter (defaults to the real ROOT) threaded
through readPinState/checkRow/fixRow -- fully backward-compatible, every
existing call site (the CLI --fix path, any other caller) is unaffected
since the default is unchanged. The pin-rewrite test now builds an isolated
temp root (its own package.json + node_modules/js-yaml/package.json) and
passes it explicitly, so the real package.json is never touched either.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: use helpers.cleanup instead of raw fs.rmSync in test cleanup

CI caught it: local/no-raw-rmsync-in-tests flagged the three t.after temp-dir
cleanup calls added for the fixRow isolation fix. helpers.cleanup() carries
the Windows-EBUSY retry budget (maxRetries/retryDelay) that raw fs.rmSync
lacks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 22:46:18 -04:00
Dennis Alexis Valin Dittrich
1316e03b84 fix(#4424): assert the launcher snippet's env surface is covered by the scrub lists (#4503)
* chore(#4424): assert launcher snippet env surface is covered by scrub lists

SNIPPET_SCRUB is hand-maintained for vars TEST_ENV_BASE's registry-derived
list can't carry. Nothing asserted the union actually covers every
${VAR:-default} arm in _runtime-launcher.snippet.sh, so a new runtime-home
arm with no scrub entry could drift silently — the #4205 shape, one door
over. Adds (A2): extracts every ${[A-Z_]+:-} capture from the snippet and
checks membership in TEST_ENV_BASE, SNIPPET_SCRUB, or the two vars the
snippet/fixtures set themselves (RUNTIME_DIR, GSD_TOOLS).

* fix(#4424): allow digits in the (A2) fallback-var regex

CodeRabbit review on fork PR #37: [A-Z_]+ silently drops any \${VAR:-...}
capture whose name contains a digit (e.g. CLAUDE2_CONFIG_DIR) instead of
flagging it uncovered, defeating the guard's own purpose. Matches bash
identifier syntax instead: leading letter/underscore, then alnum/underscore.

* fix(#4424): rename SELF_ASSIGNED to reflect RUNTIME_DIR's real provenance

Gemini adversarial review (agy) on fork PR #37: RUNTIME_DIR is an external
input the snippet reads via \${RUNTIME_DIR:-...}, never assigns — every
fixture sets it in-script before sourcing the snippet. Only GSD_TOOLS is
truly snippet-self-assigned. SELF_ASSIGNED conflated the two; renamed to
CALLER_OR_SELF_ASSIGNED. No behavior change.

Reviewed and rejected: moving RUNTIME_DIR into SNIPPET_SCRUB (blanking it
is indistinguishable from unset to the resolver's own \${RUNTIME_DIR:-...}
fallback, re-opening the #4205 ambient-leak this suite guards against —
see the existing comment at line ~1801); widening the regex to mixed-case,
colon-less \${VAR-default}, or \${VAR:=default} forms (none exist in the
snippet, and the issue's own spec scopes this to \${[A-Z_]+:- captures);
stripping bash comments before matching (the snippet is one physical line
with zero '#' characters, so no comment can exist in it).

* fix(#4424): guard CALLER_OR_SELF_ASSIGNED against silent future additions

trek-e review on PR #4503: a future ${VAR:-default} arm could be dropped
into this set without confirming it is genuinely caller-supplied/
self-assigned rather than a real coverage gap. Adds a comment requiring
justification for any addition, pointing to SNIPPET_SCRUB as the default
when in doubt. No behavior change.

* fix(#4424): guard (A2) against a vacuous pass on empty extraction

agy adversarial review (gemini-3.8-flash-high, /gsd-review lane) on PR
#4503: if the snippet becomes unreadable/truncated/renamed, matchAll
yields zero matches, uncovered stays [], and assert.deepStrictEqual
passes vacuously — same "guards the guard" gap the sibling (E)-adjacent
tests already close with assert.ok(files.length > 0, ...). Asserts
extracted.length >= 15 before filtering.

Reviewer's second finding (regex misses colon-less ${VAR-default}) is not
applied: no such form exists in the snippet today, and 688cc1c already
recorded this exact widening as scope creep the issue's own spec (${[A-Z_]+:-)
does not ask for.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-10 22:12:22 -04:00
Michel Moreira
eadcba5f53 fix(#4481): anchor bold STATE field reads to line start (#4510)
* test(#4481): reproduce mid-sentence state field reads

* fix(#4481): anchor bold STATE field reads to line start

* docs(#4481): add changeset for #4510

* fix(#4481): align bold field readers with anchored writers
2026-09-10 20:29:10 -04:00
Tom Boucher
43c48ce92e Merge pull request #4621 from open-gsd/test/4520-batch9-generators-doc-gates
test(#4520): migrate generators/doc-gates/attribution batch to named timeout constants
2026-09-10 20:29:05 -04:00
sim
2eef8ada4f test(#4520): migrate generators/doc-gates/attribution batch to named timeout constants
Batch 9 of 17 in the ad hoc timeout literal migration (epic #4445).
Replaces every bare numeric timeout/timeoutMs object-literal property in
14 files with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs.
Removes the 14 files from the rule's allowlist.

The issue's own guess ("all run a scripts/*.cjs generator or lint script
once, BUILD_TIMEOUT_MS class") needed two corrections found by reading
every site directly. First, BUILD_TIMEOUT_MS's own doc comment scopes it
specifically to scripts/build-hooks.js, which none of this batch's
generator/lint-script sites run — a new shared constant,
GENERATOR_SCRIPT_TIMEOUT_MS, covers the class instead. Second,
no-pending-3212-markers.test.cjs's single site spawns `git ls-files`
directly, not a scripts/*.cjs script at all — routed to a second new
shared constant, REAL_REPO_GIT_TIMEOUT_MS, promoted once
emitted-attribution.test.cjs's own git-plumbing sites were found sharing
the same class and value.

Isolated Standards-axis review caught a further misclassification: one
of REAL_REPO_GIT_TIMEOUT_MS's three emitted-attribution.test.cjs sites
actually builds a fresh throwaway temp repo (createTempDir + git init),
contradicting that constant's own real-repo-tree-only scope. Fixed with
a new file-local FRESH_FIXTURE_GIT_TIMEOUT_MS holding the exact
pre-existing value under an honest name, rather than reusing the shared
GIT_FIXTURE_TIMEOUT_MS (which would have doubled the bound).

emitted-attribution.test.cjs also gets two more file-local constants:
HEAVY_REAL_TREE_TEST_TIMEOUT_MS (node:test's own per-test timeout
option, not a spawn bound) and BUILD_HOOKS_UNDER_LOAD_TIMEOUT_MS (the
same build-hooks.js script as the shared norm, at 4x its bound inside
the suite's heaviest test). emitted-ack-trailer.test.cjs gets
IMPOSSIBLY_SHORT_GIT_TIMEOUT_MS — the one value in this migration that
is deliberately tiny (20ms), used to force a timeout in a negative test,
not generous headroom.

No src/bin file touched, no numeric value changed anywhere.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 19:22:04 -04:00
Tom Boucher
a1a4182bda Merge pull request #4616 from open-gsd/test/4519-batch8-security-scanners 2026-09-10 19:01:23 -04:00
Michel Moreira
411199d08b fix(#4480): require a name column in roadmap phase tables (#4511)
* test(#4480): reproduce unnamed roadmap table phases

* fix(#4480): require named roadmap phase tables

* docs(#4480): add changeset for #4511

* test(#4480): cover phase name columns generatively

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-10 18:32:55 -04:00
Tom Boucher
5e2055ab90 fix(#3660): reap a bounded check's orphaned worker after its own timeout kill (#4615)
* test(#3660): prove a bounded node-test check orphans its worker on timeout

Regression test only, no fix yet: `execFileSync`'s timeout kills the direct
`node --test` runner but never the per-file worker it forks by default since
Node 22 (`--test-isolation=process`). The worker is reparented to PID 1 and
can busy-loop forever while the bounded-check verdict still reports a clean
fail-closed timeout.

Adds three tests driven through the real, uninjected defaultRunCheck path:
a hanging subject's worker must not survive the call, a control proving the
liveness probe can actually distinguish alive-vs-dead, and a non-hanging
failure proving the reap-gating logic added by the next commit doesn't
change the ordinary-failure return shape.

Expected RED on this commit (src/prohibition-enforcement.cts is unchanged).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): reap a bounded check's descendant worker after its own timeout kill

execFileSync's timeout only signals the direct child (the node --test runner);
since Node 22, node --test forks a per-file WORKER by default
(--test-isolation=process), so a hung subject's worker survives the bound,
gets reparented to PID 1, and busy-loops forever while the verdict still
reports a clean fail-closed timeout.

Adds execFileSyncReaping (wraps execFileSync, detached:true on POSIX) and
reapDescendants(pid): POSIX process.kill(-pid, 'SIGKILL') against the
process group, Windows an absolute-path taskkill /PID <pid> /T /F (never a
bare PATH-resolved name -- PR #3681 review minor-9). The reap fires ONLY
when this call's own timeout killed the child (the thrown error carries a
signal) -- an ordinary non-zero-exit failure has signal:null and is left
alone, which is the fix for PR #3681's Blocker-3 (that attempt reaped on
every throw, risking a PGID-reuse collateral kill on a ordinary red run).

All four execFileSync(process.execPath, ...) call sites now route through
execFileSyncReaping: runNodeTestWithSubject, defaultRunCheck's node-test and
lint-rule arms, defaultProveFailFirst's lint-rule arm (its node-test arm
reuses runNodeTestWithSubject).

RED proven on 0bd741fbc2b2ad0792fdf2361de14b68b2e3aea3 (test-only commit,
gsd-test outcome:failed, exactly the new orphan-detection test failing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): address code-review nits on the reap doc comments

- Clarify execFileSyncReaping's gate covers a maxBuffer-triggered kill too,
  not just a timeout -- both set .signal, both are "this call's own bound".
- Note reapDescendants' POSIX catch swallows any errno, not only ESRCH.

No behavior change (tsc --noEmit clean, no-op for the compiler).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* debug(#3660): fix hardcoded Windows path + add temp diagnostics for CI reap failure

Real defect #1 (fixed for good): taskkillPath() had a 'C:\Windows' literal
fallback, tripping tests/hardcoded-paths.test.cjs's repo-wide scanner. Now
returns null when neither SystemRoot nor windir is set, and the caller skips
the Windows reap rather than guessing a path.

Real defect #2 (under investigation): the prior GREEN gsd-test run showed the
#3660 orphan-detection test STILL failing on linux-node24 even with the fix
applied -- the worker survived. Isolated diagnostic scripts against the exact
same execFileSync({detached:true})+process.kill(-pid) mechanism, including
one using a REAL node --test worker, both confirm the mechanism works
correctly on macOS (group-kill reaches the worker). This commit adds
TEMPORARY stderr instrumentation (GSD-DEBUG-3660 tags) around the reap
attempt to get direct evidence from the actual Linux CI environment before
guessing further. Will be removed once the root cause is confirmed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): gate the reap on err.code === 'ETIMEDOUT', not err.signal

Root cause of the prior GREEN run's failure, confirmed with real evidence
from linux-node24 CI: execFileSync's thrown error on a genuine timeout-kill
does NOT reliably set `.signal` -- on that environment it came back
`signal: null, code: 'ETIMEDOUT', status: 7`, so the reap gate never fired.
A separate macOS/Node run of the identical scenario showed `signal: 'SIGTERM'`
for the same case -- neither field alone is safe across platforms/versions,
but `code === 'ETIMEDOUT'` was present and correct in both. Verified via
temporary stderr instrumentation on a real gsd-test run (now removed) before
landing this, rather than guessing from the macOS-only result.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* debug(#3660): round-2 instrumentation -- ETIMEDOUT gate fix alone didn't work

The err.code === 'ETIMEDOUT' gate fix (previous commit) did not resolve the
failure -- same test still red on real Linux CI. Adding probes around the
actual process.kill(-pid, 'SIGKILL') call itself to see whether it throws,
and whether the group is observably alive/dead before and after, since the
gate may now be firing correctly but the kill may not be reaching the
worker's process group on this environment. Temporary, will be removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): the fix was already correct -- the TEST's liveness probe was not

Root cause of the two prior red rounds, confirmed via process-group probes on
real Linux CI: process.kill(-pid, 'SIGKILL') succeeds (no throw) every time
the ETIMEDOUT gate fires -- the worker genuinely IS killed. But
process.kill(pid, 0) cannot tell a truly-running process from an
already-killed ZOMBIE stuck unreaped: this bench's container has no init
process collecting arbitrary orphans, so a killed worker (reparented to PID 1
on death) sits as a zombie forever, still answering kill(pid,0) with "exists"
even though it is fully dead and burning zero CPU -- which is the actual harm
#3660 is about.

Test now reads /proc/<pid>/stat's process-state field on Linux and treats 'Z'
(zombie) as dead, falling back to the plain kill(pid,0) probe elsewhere (no
/proc on macOS/Windows).

Also strips the round-2 GSD-DEBUG-3660b instrumentation now that its evidence
has been used and the real root cause is fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): merge duplicate doc comment, fix stale err.signal reference

Leftover artifacts from the multi-round debugging: taskkillPath had two
stacked doc comments (an edit only replaced the function body, not the
original comment above it); a test comment still said "err.signal" after
the gate was changed to err.code === 'ETIMEDOUT'. Comment-only, no behavior
change (tsc --noEmit no-op).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): close a vacuous-test gap and harden isAlive's error handling

Code review finding (major): none of the three #3660 regression tests ever
asserted isAlive(pid) === true for a genuinely running process -- the real
code path is fully synchronous, so there's no natural window to observe
"alive" before "dead" inside those tests. A probe that always returned false
would have passed all three vacuously. Added a standalone test proving
isAlive(process.pid) reports true, using this test's own unambiguously-alive
process, running before the three existing tests.

Also hardened isAlive's /proc read-failure handling (minor finding): only
ENOENT (process genuinely gone) now means "dead"; any other read error
(EACCES, EIO, ...) reports "alive" (inconclusive) rather than risking a
false "dead" that would silently mask a real regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): add changeset fragment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3660): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3660): bound the taskkill spawnSync with a timeout

CI caught it: local/require-subprocess-timeout (DEFECT.UNBOUNDED-SUBPROCESS)
flagged the new spawnSync(taskkill, ...) call in reapDescendants' Windows
branch for having no timeout. 5s bound -- a local OS command, not a network
call; reapDescendants already treats any failure (including a hypothetical
hang) identically via its existing try/catch, so the bound costs nothing and
just prevents a stuck taskkill from blocking the caller forever.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 17:55:05 -04:00
sim
150acd78c1 test(#4519): migrate security-scanner batch to named timeout constants
Batch 8 of 17 in the ad hoc timeout literal migration (epic #4445).
Replaces every bare numeric timeout/timeoutMs object-literal property in
tests/secret-scan-lint.security.test.cjs, tests/security-scan.security.test.cjs,
tests/security-prompt-injection.security.test.cjs, tests/prompt-injection-scan.security.test.cjs,
tests/read-injection-scanner.security.test.cjs, tests/read-injection-scanner.property.test.cjs,
and tests/security.test.cjs with a named constant, per
eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 7 files from the
rule's allowlist.

The issue guessed this batch "most likely needs its own named
SCAN_TIMEOUT_MS." Reading every one of the 19 call sites directly found a
more specific picture: 8 sites across 3 files scan exactly one small temp
fixture file and match the existing QUICK_SPAWN_TIMEOUT_MS class exactly
(reused, no new constant). Two new shared constants cover genuinely
distinct classes that happen to coincide in value:
SCAN_USAGE_ERROR_TIMEOUT_MS (a bash scan script given missing arguments)
and MALFORMED_INPUT_HOOK_TIMEOUT_MS (a Node hook fed malformed JSON) --
kept as separate names per this migration's standing rule that numeric
coincidence is never identity. Three file-local constants cover a real
multi-file directory scan, a property-fuzzing safety net, and a
path-traversal hook test, each with its own pre-existing rationale
preserved.

No bound is lowered or raised anywhere in this batch, honoring the
issue's explicit caution that security-scan timing margins deserve
extra scrutiny. No src/bin file touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 17:09:23 -04:00
Tom Boucher
b25fe4dbaf Merge pull request #4614 from open-gsd/test/4518-batch7-loop-hook-point-e2e 2026-09-10 16:59:25 -04:00
sim
6d6e3eea73 test(#4518): migrate loop/hook-point e2e batch to named timeout constants
Batch 7 of 17 in the ad hoc timeout literal migration (epic #4445).
Replaces every bare numeric timeout/timeoutMs object-literal property in
tests/loop-render-hooks.test.cjs, tests/loop-walk.qa.test.cjs,
tests/loop-hooks-empty-points-e2e.test.cjs, tests/loop-hooks-ship-pre-e2e.test.cjs,
tests/loop-hooks-verify-post-e2e.test.cjs, tests/check-gap-analysis-plan-post-e2e.test.cjs,
tests/check-tdd-review-checkpoint-e2e.test.cjs, tests/execute-wave-post-gate-pipeline-e2e.test.cjs,
tests/plan-pre-hook-e2e.test.cjs, and tests/qa/tdd-walk.cjs with a named
constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 10
files from the rule's allowlist.

Promotes a new shared class norm to tests/helpers/timeouts.cjs,
LOOP_HOOK_POINT_CLI_TIMEOUT_MS: 7 files independently arrived at the same
value for a single gsd-tools.cjs CLI subcommand invocation with no
confirmed subprocess fan-out. Reuses the existing PROBE_TIMEOUT_MS for
loop-render-hooks.test.cjs's 11 sites (same class, exact value match).
Adds 4 file-local constants for values that share a class with the new
norm or an existing one but diverge in pre-existing value, or that
numerically coincide with an unrelated existing constant without
matching its actual operation.

Isolated Standards-axis review caught that the new shared constant's doc
comment exhaustively enumerated 3 verb families while a 7th genuine site
(an `init new-project` invocation) also correctly belonged to the class;
fixed by rewording the comment to state the class definition (call
shape) first and list all 4 representative verbs, explicitly
illustrative rather than exhaustive. No value changed, no site's
classification changed.

No src/bin file touched, no numeric value changed anywhere.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 15:50:52 -04:00
Tom Boucher
138e70d734 test(#4517): migrate hook/guard invocation batch to named timeout constants (#4608)
Batch 6 of 17 in the ad hoc timeout literal migration (epic #4445).
Replaces every bare numeric timeout/timeoutMs object-literal property in
tests/read-guard.test.cjs, tests/feat-2483-review-claude-mds-guard.test.cjs,
tests/gsd-write-guard.test.cjs, tests/gsd-secret-read-guard.test.cjs, and
tests/hooks-crash-policy.test.cjs with a named constant, per
eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 5 files from the
rule's allowlist.

Reuses existing tests/helpers/timeouts.cjs class norms where the shape
matches (QUICK_SPAWN_TIMEOUT_MS x2, PROBE_TIMEOUT_MS x1). Adds two
file-local constants for classes not shared across files:
READ_GUARD_HOOK_TIMEOUT_MS (read-guard.test.cjs's 4 sites, a tighter
no-fan-out bound than QUICK_SPAWN_TIMEOUT_MS with no bench data to widen
it) and FIXTURE_PROBE_CAPABILITY_TIMEOUT_MS (fixture data, not a real
spawn timeout).

The fifth site (feat-2483-review-claude-mds-guard.test.cjs's review-lane
invoke spawn) was initially classified onto HOOK_FANOUT_TIMEOUT_MS;
isolated Spec-axis review caught that this call is one nested spawn, not
the multi-spawn git-hook fan-out shape that constant's own doc comment
defines. Corrected to a new file-local REVIEW_LANE_INVOKE_TIMEOUT_MS
holding the exact pre-existing 60000ms value under an honest name,
rather than narrowing to PROBE_TIMEOUT_MS with no bench justification.

No src/bin file touched, no numeric value changed anywhere.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 15:36:43 -04:00
Michel Moreira
93ba63aeff fix(#4482): strip Copilot notes from OpenCode artifacts (#4532)
* fix(#4482): strip Copilot notes from OpenCode artifacts

* chore: add changeset for #4532

* fix(#4482): filter runtime notes across emitted surfaces

* fix(#4482): make runtime note filtering runtime-neutral

* test(#4482): keep ack fixture outside registered provenance

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-10 15:29:08 -04:00
Tom Boucher
1e47560e34 feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the
same 546-file, Windows-oriented conformance-tier list as windows-latest --
built from signals like windows-shell-token/windows-env-var that have
nothing to do with macOS. Issue #4593 asked for macOS coverage sized to
its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin-
specific behavior) instead.

Issue #4593 was filed before Phase 5 (#4603) existed and referenced
updating test-full's macOS legs -- that job is gone. Corrected the issue's
body before any code was touched: the "shrink from full replay" half of
the original ask was already done by Phase 5; what remained was narrowing
the still-Windows-oriented tier macOS was inheriting.

Two design assumptions were measured and rejected before accepting a
design (documented in docs/adr/4593-macos-conformance-tier-architecture.md):
- Reusing the general tier's signals minus its 3 Windows-specific
  categories barely narrows anything (546 -> 424, 78% retained) -- most
  files match multiple signals and only need one to survive exclusion.
- A standalone CRLF/autocrlf signal, despite the issue naming
  "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit
  143/930 files. Root cause: CRLF is primarily a Windows checkout concern
  in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST-
  PORTABILITY), so the signal was really re-selecting Windows-relevant
  files already covered by the general tier, not narrowing macOS
  specifically.

Built 5 new, genuinely macOS-specific signals instead: darwin-literal
(darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch,
case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim
from the general tier (genuinely Unix-relevant, not Windows-motivated).
Measured against the real tree: 196 of 930 eligible unit-suite files
(21%), versus the general tier's 546 (59%) -- a real, evidence-backed
narrowing.

scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/
classifyMacosTree/renderMacosGeneratedFile and a --target windows
(default, unchanged)/--target macos CLI flag, so the same generator
produces two independent, gated outputs rather than needing a second
script. New committed output: scripts/lib/macos-conformance-tier.
generated.cjs. .github/workflows/test.yml's test-conformance job: only
the macos-latest leg's file-list source changes; windows-latest is
byte-for-byte untouched. New shipped-file ripples handled proactively
(19 install-tree fixtures regenerated, bin/install.js registered).

An isolated code-review pass found one real defect: the ADR's per-
category count table had drifted by 1 (zsh-dispatch, case-sensitivity)
because the new test file's own fixture strings joined the tree it
classifies after the table was authored -- fixed, with the union total
(196, what CI actually gates on) confirmed unaffected. An isolated
security-review pass found no qualifying findings.

The ADR also records an explicit requirement for any future widening
proposal: check whether the motivating regression is already covered by
Phase 1's no-rendered-text-length-assert lint rule (#4590) before
re-proposing full macOS/Linux parity, since that is exactly what #4421's
root cause was (a rendered-text-length assertion, not a real behavioral
divergence).

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 14:37:15 -04:00
Tom Boucher
557a0b2876 test(#4516): migrate per-runtime install/upgrade adapters to named timeout constants (#4582)
Batch 5 of the ad hoc timeout literal migration (epic #4445). Replaces
every bare numeric timeout/timeoutMs object-literal property across
tests/augment-upgrades.test.cjs, tests/shared-hooks-dir-resolution.test.cjs,
tests/antigravity-upgrades.test.cjs, tests/cursor-hooks.test.cjs,
tests/cursor-hook-workspace-roots.test.cjs, tests/gemini-runtime-removed.test.cjs,
tests/kilo-upgrades.test.cjs, tests/kimi-upgrades.test.cjs,
tests/kimi-variant-disambiguation.test.cjs, tests/opencode-plugin-adapter.test.cjs,
tests/windsurf-hooks-bridge.test.cjs, tests/effort-sync-installed-runtime.test.cjs,
and tests/hooks-commonjs-marker.test.cjs with a named constant, per
eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 13 files from the
rule's allowlist.

Adds one new shared constant to tests/helpers/timeouts.cjs for spawning a
single already-staged hook script directly (a heavier class than the
existing quick-spawn norm), shared across three files in this batch. One
new file-local constant covers a property-test driver process distinct
from any existing class. Every other site reuses an existing shared norm.
No src/bin file touched, no numeric timeout value changed anywhere.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 14:11:26 -04:00
Tom Boucher
0b928fe28c feat(#4592): replace the blanket test-file full_matrix rule with reachability (#4602)
scripts/ci-test-scope.cjs's classify() previously set full_matrix=true for
ANY changed tests/**/*.test.cjs file, unconditionally (restored by #4421
after #962's narrowing let a real macOS-only regression, PR #4384, land
undetected). This replaces that blanket rule with a reachability check
against real data instead of a path prefix:

- A changed test file forces full_matrix only when it is present in Phase
  2's committed CONFORMANCE_TIER_FILES list (scripts/lib/platform-
  conformance-tier.generated.cjs) -- direct membership, not a graph walk.
- A changed src/ file forces full_matrix when its own content carries a
  genuine platform-conditional signal, reusing gen-platform-conformance-
  tier.cjs's classifyContent with a narrowed, source-code-safe signal
  subset (excludes two categories -- hardcoded-path-vs-path-call and
  symlink-keyword -- empirically found to flag 100/235 src/ files when
  applied verbatim, versus 28/235 with the narrow subset, all verified to
  carry genuine platform branches). New export: NOISY_FOR_SOURCE_REACHABILITY.
- A change to the classification mechanism's own definition files
  (gen-platform-conformance-tier.cjs, the generated tier list, or
  suite-detection.cjs) always forces full_matrix -- the mechanism being
  changed cannot presume its own new output is safe.
- Any computation error (a require/read failure, a malformed module) fails
  safe to full_matrix=true, per the issue's explicit requirement.

The existing RULES array entries with their own fullMatrix:true (workflow
automation, installer/package layout, hooks, environment/dependency gates,
test harness) are deliberately left untouched -- they are curated,
narrowly-scoped triggers for "this diff changes the CI/installer/hooks
mechanism itself," a different and still-valid reason than "product code
might reach a platform branch." Disclosed in .gsd/phase/.../40-design.md
as a scope decision, since the issue's "Done when" wording read broader
than its "Proposed work" bullets.

Two design assumptions were caught and corrected before any code was
written (rubber-duck pass, documented in 40-design.md): (1) reusing Phase
2's classifyContent verbatim against src/ was far too noisy; (2) a single
hardcoded seam file (src/shell-command-projection.cts only, per CLAUDE.md's
"single platform seam" framing) would have silently missed genuine,
independent platform branches in src/runtime-hooks-surface.cts,
src/capability-lock.cts, src/capability-ledger.cts, and src/surface.cts --
reintroducing the #4421 failure shape inside src/ instead of tests/.

An isolated code-review pass found and fixed one real defect (a dead,
untested branch that would have survived Stryker mutation testing) and one
design-doc completeness gap (2 of 8 "narrow" signal categories were left
implicitly rather than explicitly audited). An isolated security-review
pass found no qualifying findings.

tests/ci-test-scope.test.cjs gains the full #4592 boundary-case matrix
(.gsd/phase/.../50-test-matrix.md), including a named #4421 regression case
proving tests/state-todos-render.test.cjs still forces full_matrix, now for
the documented reason instead of the removed blanket rule. Two pre-existing
tests were corrected: one used a nonexistent fixture path (src/semver.cts
-> src/semver-compare.cts, a real file); one (A3) asserted the exact old
blanket-rule behavior this issue removes, updated to the new, verified-
correct expectation.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:21 -04:00
Tom Boucher
181c4c8659 chore(#4603): retire the test-full CI job (#4604)
* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:04 -04:00
Tom Boucher
bcd99696d3 chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) 2026-09-10 08:36:13 -04:00
Tom Boucher
2cefa5a5ac enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes (#4587)
* enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes

ADR-4139's final phase. workflow.compact_content already defaulted to false
(Phase 1's buildNewProjectConfig hardcoded default), but nothing surfaced it:
/gsd-new-project never asked, and /gsd-settings/config had no toggle path for
an already-initialized project — config-set/config-get were the only route.

new-project.md gains a fourth question in the existing Round 2 AskUserQuestion
array (grouped with the other general-workflow-behavior toggles, not the
per-agent capability questions above it) and threads compact_content into the
config-new-project CLI JSON literal. settings.md mirrors the exact pattern
every other non-capability workflow.* key already follows: read_current bullet,
question block, update_config write, the safe-merge non-capability-keys list,
save_as_defaults, and the confirm summary table — seven edits, zero new
src/*.cts code, since Phase 1's merge logic is a generic passthrough. Its
success_criteria question-count ("24 settings") is bumped to 25 to match the
now-25-entry main AskUserQuestion batch.

settings-advanced.md deliberately does NOT get a duplicate question: no other
boolean toggle in this repo is asked in both settings.md and
settings-advanced.md, and there's no reason to start with this one.

docs/CONFIGURATION.md, docs/USER-GUIDE.md, and a new docs/features/4139-compact-
content.md fragment (regenerated into docs/FEATURES.md) document the toggle.

ADR-4139 itself: Status flips Proposed -> Accepted, the acceptance-criteria
section becomes a guard ledger — a 13-row table covering all 12 of #4139's
original checkboxes plus the shipped-content guard criterion, each with real
evidence (the merged PR that satisfied it, fetched via `gh issue view
--json closedByPullRequestsReferences` rather than asserted from phase
numbers) — and both "Open questions for the implementation phases" are
resolved rather than left dangling: discuss-phase was never converted to
spine+detail shape (verified: no detail/ subdir exists) — a genuine gap, not a
reasoned decline; the disjointness check is confirmed line-based by reading
compact-content-split.cjs's normalizeNonTrivialLines directly.

Orthogonal review (isolated Standards/Spec code-review + security-review
sub-agents) found and this fixes two real defects: the changeset fragment's
body didn't match CONTRIBUTING.md's single em-dash-sentence format (was
multi-sentence prose naming implementation file paths); and settings.md's own
success_criteria still said "24 settings" after the new question pushed the
main batch to 25. Also fixed, found by the Spec pass while confirming
commands/gsd/settings.md correctly needed no sync edit: that file and its
skills/gsd-settings/SKILL.md twin both still described "Interactive 5-question
prompt (model, research, plan_check, verifier, branching)", stale since long
before this phase (the batch has had far more than 5 questions for a while) —
replaced with a description that names the current set without hardcoding a
count that will drift again.

gsd-test (real run, sha 1da78fe2) caught a third real regression the local
sweep missed: new-project.md is a registered spine+detail split for Phase 4's
token-reduction benchmark (scripts/benchmark-compact-content.cjs), and the new
question's +167 tokens drifted the committed baseline
(tests/fixtures/compact-content-benchmark-baseline.json). The benchmark itself
is designed never to fail CI on drift, but the test asserting the COMMITTED
baseline is currently non-drifted correctly caught it. Regenerated via
`node scripts/benchmark-compact-content.cjs --write`; re-verified --check now
reports "up to date" and the test file passes 27/27.

Closes #4408.
Closes #4139.

Emitted-Drift-Ack-Growth: new-project.md — new 4th Round-2 AskUserQuestion entry (Compact Content, #4139) plus the config-new-project CLI JSON field and explanatory sentence; a new opt-in toggle needs new prose.
Emitted-Drift-Ack-Growth: settings.md — new workflow.compact_content read_current bullet, question block, update_config write, safe-merge key, save_as_defaults field, and confirm summary row (the same seven-edit pattern every other non-capability workflow.* toggle already follows), plus the 24->25 success_criteria count fix found in review.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4408): backfill changeset PR number

pr:0 -> pr:4587 now that gh pr create has returned the real number.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 00:04:51 -04:00
Tom Boucher
9770258558 chore(#4590): add no-rendered-text-length-assert ESLint rule (#4595)
* test(#4590): add no-rendered-text-length-assert ESLint rule

Enforces ADR-456's typed-surface mandate for one specific bug shape: a test
assertion whose pass/fail depends on the length/substring content of a
template literal that interpolates an OS-derived path (os.tmpdir(),
os.homedir(), path.join/resolve/..., or a PATH_RETURNING_FNS resolver).
Because macOS's default tmpdir prefix is longer than Linux's, such an
assertion can pass on one runner and fail on another -- the defect class
behind #4421's incident (git show 4e75b836e9), already fixed there by
pinning to a typed field per ADR-456 Sec(c) before this rule existed to
catch a recurrence.

Two repo-wide sweeps against the real tests/ tree narrowed the rule to a
sound scope: an initial design that traced call arguments (to approximate
the historical incident's cross-file render-function shape) produced false
positives on ordinary fs.readFileSync(path.join(...)) + assert.match
patterns; a second design that matched any bare direct path-returning call
produced 45 false positives on path suffix/prefix/non-emptiness checks. The
shipped rule matches only a path-returning expression interpolated into a
template literal, directly or via one identifier hop -- disclosed in the
rule's own "Known boundaries" as not covering the literal cross-file
incident shape, which would require tracing into a callee's body.

Phase 1 of epic #4589 (CI test-matrix Linux-primary migration) -- Phase 2's
safety argument depends on this class of OS-dependent test assertion being
enforced going forward, not merely fixed once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4590): address code-review findings on no-rendered-text-length-assert

Reletter the "Known boundaries" doc-comment list (a)-(e), fixing a gap left
by an earlier edit pass and every stale cross-reference to it. Collapse
isDirectPathTaint/isTaintedInterpolation's duplicated TemplateLiteral-walk
into one recursive relationship (isTaintedInterpolation now delegates a
nested-template-literal case back to isDirectPathTaint instead of
re-implementing the .some() traversal) -- behavior unchanged, confirmed by
re-running the repo-wide sweep (still zero false positives).

Found by the Standards-axis /code-review pass on this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 21:55:51 -04:00
Tom Boucher
b2d50ffd83 fix(#4489): make capability-registry.test.cjs's extractShellBlocks CRLF-safe (#4584)
* fix(#4489): make capability-registry.test.cjs's extractShellBlocks CRLF-safe

Second, independent copy of the #4409 CRLF-fragile line-splitting bug,
explicitly flagged as out of scope there ("other duplicated helper in file
sibling test files not part of the shadowing chain, tracked separately if
divergent"). Same fix: content.split('\n') -> content.split(/\r?\n/),
matching src/text-lines.cts's splitLines() and the already-fixed sibling
copy in tests/runtime-launcher-parity.test.cjs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4488): regenerate INVENTORY-MANIFEST.json for tdd-red-evidence.cjs's ADR-457 untracking

Discovered while validating #4489's push: #4488's merge (untracking
gsd-core/bin/lib/tdd-red-evidence.cjs per ADR-457) left docs/INVENTORY-
MANIFEST.json stale, since that file was previously listed as a tracked
shipped artifact. Removed via node scripts/gen-inventory-manifest.cjs
--write. docs/INVENTORY.md already described this file as gitignored
(no update needed there -- it already documented the intended state).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* revert: undo incorrect INVENTORY-MANIFEST.json edit from 57f1457a27

The prior commit removed tdd-red-evidence.cjs's manifest entry based on a
false premise: a stale tsconfig.build.tsbuildinfo (gitignored, untouched
by git checkout/rebase) told tsc the file's compilation was already
current even though git's own checkout had deleted the actual output file
during this branch's rebase onto #4488's merge (a tracked-in-old-tree,
untracked-in-new-tree transition deletes the working-tree file regardless
of the new .gitignore entry). tsc's incremental cache doesn't verify its
recorded output still exists on disk, so it silently skipped re-emitting
it. Confirmed real root cause: deleting tsconfig.build.tsbuildinfo and
rebuilding fresh correctly re-emits gsd-core/bin/lib/tdd-red-evidence.cjs
(it is gitignored now, not deleted -- src/tdd-red-evidence.cts is
unaffected by ADR-457's tracked-vs-gitignored distinction and always
compiles). The manifest's own purpose (per its docstring) is 'every
shipped surface derived entirely from the filesystem' -- this file still
ships via the normal build, so it belongs in the manifest regardless of
git-tracking status. Net result matches next's own INVENTORY-MANIFEST.json
byte-for-byte; this correction should not have been needed at all had the
build cache been fresh when the prior commit was made.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 21:53:37 -04:00
Tom Boucher
7fe440a838 fix(#4488): report state update as successful when the value is already correct (#4581)
* fix(#4488): report state update as successful when the value is already correct

`cmdStateUpdate` unconditionally overwrote `updateCore`'s own `updated:true`
signal with `reconcileReportedFields`'s disk-diff result. That diff reports
`[]` -- by design -- whenever `readModifyWriteStateMd`'s #948 no-op guard
fires because the transform's output was byte-identical to the input, which
happens precisely when the requested value already equals what's on disk.
The field genuinely was found and matched; there was simply nothing left to
change. Collapsing that into the same `false`/"not found" response as a
genuine miss produced an actively wrong diagnostic message and a silent
same-day no-op in gsd-ship + gsd-extract-learnings, which both write
`Last Activity` to today's date.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4488): backfill changeset pr number to 4581

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4488): untrack tdd-red-evidence.cjs, completing its ADR-457 gitignore migration

Bundled discovery from this PR's own CI run: tests/lint-compiled-artifact-
sync.test.cjs's full tsc compile (which runs whenever ANY compiled artifact
remains tracked) SIGTERM'd under shard contention. gsd-core/bin/lib/tdd-red-
evidence.cjs (introduced by #3770/PR #4279) was the sole remaining tracked
artifact -- a tenth, later, separate instance of the #2657/#2653
migration-gap defect class this test file's closed nine-item list doesn't
cover. Untracked it and added the .gitignore entry, same fix shape as the
original nine. This eliminates the slow tsc-compile path entirely (verified:
0.1s vs ~7s locally) rather than papering over a timeout. Added a generic
regression test asserting the tracked set is fully empty.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 18:32:54 -04:00
Tom Boucher
615b74ff45 fix(#4460): correct two stale changesets left by an admin-merge race (#4572)
* fix: address orthogonal-review findings on the new work in this PR

Isolated code-review + security-review of everything added to this PR
since its original review (hono override, check-env.cjs rewrite/revert,
new lib file, its test, installer enumeration). Security review: clean,
no findings. Code review found:

- BLOCKER: .changeset/silly-hens-relax.md described a hono override
  this PR no longer actually makes -- PR #4560 landed the identical fix
  on next first, and this branch's own hono commit became a genuine
  no-op the moment it was rebased onto that updated next (git diff
  origin/next -- package.json package-lock.json is empty). Deleted the
  orphaned changeset; next already carries #4560's equivalent one
  (.changeset/zesty-seals-click.md).
- HIGH: .changeset/tame-hens-jump.md's body still described the
  execNpm-routing approach that was tried and reverted -- stale text
  from before that revert, would have shipped a release note for code
  that isn't actually in the diff. Rewritten to describe what actually
  shipped (self-contained spawnSync, 15s timeout, accurate ENOENT vs.
  timeout vs. non-zero-exit diagnosis).
- LOW: no comment explaining why the spawnSync call has no try/catch
  (safe -- its documented contract routes failures through the returned
  result, never a throw -- but worth stating given this file's whole
  purpose is graceful degradation). Added one.
- nit: exitCode 0 + empty stdout fell through to "npm binary not found
  on PATH", misdescribing a real npm binary that simply printed
  nothing. Gave it its own message; updated the corresponding test.

Manually re-verified describeNpmVersionCheckFailure's branches and the
real check:env success path before re-running gsd-test, since this
repo blocks local node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: rest of the orthogonal-review fixes (previous commit only caught the deletion)

Tooling mistake in the previous commit: a git add with the already-staged
deleted changeset mixed into the same pathspec list errored out and
silently skipped staging the other four files, so only the changeset
deletion actually committed. This commit carries the rest of that same
change: tame-hens-jump.md's rewritten body, check-env.cjs's no-try/catch
comment, npm-version-check-diagnosis.cjs's exitCode-0-empty-stdout fix,
and the corresponding test update.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4460): fix changeset pr field to point at this PR, not the original

.changeset/tame-hens-jump.md's pr field still said 4552 (the PR its
original text was authored under), but this PR (#4572) is what's
actually landing the corrected body -- changeset-lint's own
DEFECT.CHANGESET-PR-FIELD-DRIFT check caught it: "pr: 4552, expected
pr: 4572".

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 16:26:25 -04:00
Tom Boucher
385ed619f1 fix(#4487): stamp broken-windows ledger entries with the resolved milestone (#4583)
* enhance(#4487): stamp windows-ledger entries with the resolved milestone

Broken-windows ledger entries (`.planning/WINDOWS.md`) carry `phase` as
a bare number. Phase numbers are unique only within one active phases/
directory -- `milestone complete` archives phases and frees their
numbers for reuse, so two milestones routinely produce entries sharing
the same phase value with nothing distinguishing them. Since
`/gsd-ship` blocks while any entry is open, an already-archived
milestone's open entries could silently block shipping the CURRENT
milestone, with no supported way to attribute which entry belonged to
which milestone short of manually cross-referencing MILESTONES.md
timestamps against decision IDs that happened to appear in description
prose.

Added an optional `milestone: string | null` field to WindowEntry,
stamped by `windows append` (cmdWindowsAppend, which already does file
I/O) from the workstream's resolved milestone version. Reused the
existing `readCurrentMilestoneVersion` (workstream-inventory.cts --
STATE.md `milestone:` frontmatter first, ROADMAP.md in-progress marker
as fallback) rather than writing a parallel implementation: exported it
via that module's existing `export = {...}` CJS-interop convention
(matching the `import ... = require(...)` pattern already used in
workstream.cts/init.cts). appendWindow itself stays pure -- it accepts
milestone as an optional input field and passes it through; only the
CLI-facing cmdWindowsAppend resolves it from disk.

Backward compatible by construction: validateEntryShape does NOT add
`milestone` to its required fields, so an existing ledger entry with no
milestone key at all parses without error and reads back as null --
exactly "recorded before this change," no migration needed. The
rendered markdown table is deliberately left unchanged (the issue's own
words: "the JSON is the source of truth"); adding a table column would
be a separate, larger change than adding an optional JSON field.

Two smaller gaps the issue itself flags as separable ("happy to split
them out") are explicitly NOT addressed here: no verb to amend an
entry's description, and the table/JSON drift-repair advice that can
destroy table-only edits on a parse failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4487): preserve absent-vs-null milestone through parse/render roundtrip

validateEntryShape stamped an explicit `milestone: null` onto every
entry lacking the key, so a pre-#4487 ledger entry gained permanent
JSON churn ("milestone": null) the first time ANY entry in the ledger
was touched -- breaking the pure parse/render roundtrip-identity
property test (render(parse(render(ledger))) must equal render(ledger))
and, in real usage, contaminating unrelated entries' diffs on every
append/waive/fixed of an old ledger.

Fixed by distinguishing "key genuinely absent" (undefined -- JSON.
stringify drops it, matching pre-#4487 behavior exactly) from "recorded
but unresolvable" (explicit null, the real signal appendWindow stamps
on brand-new entries). WindowEntry.milestone is now optional
(`milestone?: string | null`) so returning undefined type-checks.

Updated tests/broken-windows.test.cjs's roundtrip property generator to
exercise all three states (absent/null/string) -- its prior silence on
this field is exactly what let the regression through. Also corrected
the earlier backward-compatibility test's assertion: a pre-#4487 entry
reads as milestone: undefined, not null, and re-rendering it must not
introduce a milestone key at all.

Also ran npm run regen:derived: docs/features/broken-windows-ledger.md
(edited in an earlier commit) had never been propagated to its
generated docs/FEATURES.md projection, which is what was independently
failing tests/features-index-gate.test.cjs and, as a side effect of
staleness, tripping tests/fragment-single-edit-propagation.install.
test.cjs's second-source-surface check.

Manually verified via the compiled lib (500 fast-check iterations plus
direct legacy/new-entry roundtrip checks) before wiring the test file,
since this repo blocks local node --test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4487): materialize milestone via conditional spread, not undefined assignment

An object literal property set to `milestone: undefined` is still an
OWN property -- `'milestone' in entry` reads true regardless of the
assigned value, only JSON.stringify treats undefined specially. My
prior commit's own new backward-compat test asserted `'milestone' in
entry === false` for a pre-#4487 entry and failed on exactly this.
Switched to conditionally spreading the key in only when the source
object actually had it, so a genuinely absent milestone is not
materialized at all -- matching both the `in` check and JSON
serialization. Re-verified via the compiled lib (500 fast-check
roundtrip iterations, plus the specific in/undefined/JSON assertions
the failing test makes) before re-running gsd-test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4487): backfill changeset pr number to 4583

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 15:37:50 -04:00
Tom Boucher
8deb40722a fix(#4478): anchor collectAnalyzePhases's phase-heading regex to line start (#4578)
* fix(#4478): anchor collectAnalyzePhases's phase-heading regex to line start

phasePattern (src/roadmap.cts, backing `gsd-tools roadmap analyze`) had no
line anchor -- #{2,4} could match a `### Phase N:`-shaped mention ANYWHERE
the global regex scan reached: mid-sentence prose, inside a blockquote,
inside an inline code span (backtick-quoted on the same line, not a fenced
code block tokenizeHeadings would exclude). Any such line minted a phantom
phase entry, inflating phase_count and able to collide on a phase NUMBER
with a real heading nearby.

Two correctly-anchored reference implementations already exist for the
same heading grammar in this codebase: tokenizeHeadings
(src/markdown-sectionizer.cts:453) and findRoadmapPhaseInContent
(src/roadmap-parser.cts:1385), which anchors against the tokenizer's own
output. collectAnalyzePhases was the one path scanning raw content
directly instead. Anchored to line start with the same 0-3 leading-space
tolerance tokenizeHeadings uses (rather than routing through the
tokenizer, which the issue offers as the more thorough fix but which
would require re-deriving this function's bracket/number/name capture
groups and section-boundary lookup from tokenized output instead of a
single combined regex scan -- a materially larger refactor than a bug fix
warrants; the issue itself offers anchoring as the sufficient fallback).

Added coverage to the existing tests/roadmap.test.cjs "roadmap analyze
command" describe block (not a new file -- the roadmap module already
had 4 test files and lint-test-file-count.cjs's own remedy is to
consolidate, not add a 5th) against the issue's own 5-row prose-lookalike
table, its duplicate-number consequence, and a boundary case (a
legitimately-indented real heading must still count).

Independent code review on this same diff found one more consequence:
the "next heading" section-boundary lookup (nextHeader, a few lines
below phasePattern) lacked the SAME {0,3} leading-space tolerance --
a legitimately-indented NEXT phase heading was invisible to it, letting
the prior phase's own goal/mode/depends_on extraction bleed across the
section boundary into the next phase's body. Confirmed via a targeted
repro (**Depends on:** -- a field only the second phase has, so the
bleed is directly observable) and fixed with the same tolerance, plus
its own regression test.

CI-adjacent findings caught by gsd-test on a stale sha, fixed inline:
(1) my own explanatory comment block was inserted BETWEEN a pre-existing
`phase-id-owner:` sanction comment and the regex it sanctions, pushing
it out of the "line directly above" position lint-phase-id-drift.cjs
requires -- reordered so the sanction stays immediately above the
regex; (2) the standalone test file this fix originally added tripped
lint-test-file-count.cjs's per-module cap -- consolidated into the
existing describe block as described above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4478): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 14:50:02 -04:00