Commit Graph

45 Commits

Author SHA1 Message Date
Tom Boucher
7a02f98574 fix(#3493): confine key_links from:/to: to the project directory (#3506)
`cmdVerifyKeyLinks` resolved `from:` and `to:` with `path.join(cwd, <value>)` where the value comes verbatim from plan YAML. `path.join` normalizes `../` rather than rejecting it, so a plan travelling with a repository could name any file the process can read, and the command reports whether the link's `pattern` matched it — an arbitrary-file-read oracle reachable from `verify-phase`.

Both reads now go through `validatePath` in src/security.cts, the existing realpath-based confinement seam already used at 11 call sites. Not a missing capability — a bypassed one.

Two defects in that seam, found by adversarial review and fixed here because they affect all 11 callers:

1. A dangling in-project symlink escaped confinement. A link to an EXISTING outside path was refused (realpath lands outside) while a link to a MISSING outside path took the parent-resolution fallback and was accepted — an existence oracle for arbitrary absolute paths. lstat succeeds on a dangling link and throws ENOENT on a truly absent path; an unresolvable link is now refused. A symlink resolving inside the project is still accepted.

2. A canonicalized base was compared against an uncanonicalized path when a file and its parent were both missing, wrongly refusing legitimate in-project paths on any non-canonical cwd (every macOS temp dir). This was a live regression in this PR: the wave-pending classification (#1202) depends on the not-yet-created case. Resolution now walks up to the nearest existing ancestor.

Two adjacent aborts fixed: the from: read sat outside the per-link try, so a non-ENOENT errno killed the whole command; and an empty from: read the cwd directory, throwing EISDIR. Both now fail per-link.

The issue was filed as a fourth ADR-0174 consolidation loss. It is not one — validatePath/requireSafePath never went away, only the SDK's name for the concept did. This is an instance of epic #3473's F2 family. The ADR-0174 loss count is three.

Closes #3493

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:29:58 -04:00
Tom Boucher
895d9df96d fix(#3477): run untrusted key_links patterns on a linear-time engine (#3496)
`cmdVerifyKeyLinks` compiled `must_haves.key_links[].pattern` from plan frontmatter with `new RegExp()` and tested it against whole file contents, so a nested-quantifier pattern such as `(a+)+$` hung `verify-phase` indefinitely (CWE-1333). JavaScript has no regex-execution timeout.

Untrusted patterns now run on RE2 (re2js), whose match time is linear in input length — the class is closed by the engine, not by a heuristic screen. The screen lost in the ADR-0174 consolidation was deliberately NOT restored: it never worked, since `(a|a)*$`, `((a+))+$`, `(a+){2,}$` and `(a{1,3})+$` all evade it. A refused pattern's matcher returns false for every input, so it cannot report a match no matter what the caller does.

The engine is vendored at gsd-core/bin/lib/vendor/re2js.cjs because gsd-core/bin/** is copied into installed trees with no node_modules; runtime dependencies are unchanged. New ESLint rule local/no-external-require-in-bin enforces that invariant, which had been documented in a comment since the #3024/#2071 bug class and enforced nowhere.

Backreferences and look-around are unsupported by RE2 by construction — disclosed in a Changed changeset.

Closes #3477

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 14:34:36 -04:00
Tom Boucher
ff2d08d453 fix(#3193): tolerate attributes on plan-task child tags (#3433)
* fix(#3193): tolerate attributes on plan-task child tags

* chore(#3193): add changeset

* chore(#3193): set changeset pr to 3433

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:14:07 -04:00
Tom Boucher
6dbc124018 enhance(#3180): the sibling validators share one envelope and one owner — Phase 12 (#3407) 2026-08-13 11:31:16 -04:00
sim
d1760e3c31 refactor(#3309): migrate cmdValidateHealth onto the rule table
Replaces cmdValidateHealth's hand-rolled addIssue/switch accumulation
(961 lines) with buildPlanningSnapshot -> evaluateRules -> map to the
legacy {code, message, fix, repairable} shape, bucketed by severity.
Two pre-checks (home-dir E010/I010, .planning/-root-missing E001) stay
outside the rule table entirely, per ADR-3180 §8.2 rule 4 ("no
precedence system") — building "some rules suppress others" into the
table would itself be the forbidden precedence system.

W024 (STATE.md commit-age freshness) also stays outside the table:
its committed rule is a documented permanent no-op (readStateHeadFreshness's
git-log shell-out is ambient I/O a Rule.check may never perform, and no
PlanningSnapshot field carries a commits-behind count). Migrating onto
the rule table as designed would have silently regressed 7 passing
tests in tests/health-validation.test.cjs — found while wiring this
function, kept as a real check in the wrapper instead (same I/O
license applyRepairs already relies on), fixed inline per this repo's
no-defer policy rather than accepted as a silent loss.

Ports the real repair-handler bodies (createConfig/resetConfig,
regenerateState, addNyquistKey/addAiIntegrationPhaseKey,
backfillMilestones) into health-diagnostic.cts's applyRepairs,
replacing the skeleton's stub. DESTRUCTIVE-risk remedies
(resetConfig/regenerateState) are refused by --repair — a disclosed
breaking change; repairable now means "an automatic repair will
actually run," not merely "a remedy exists to describe," so E004/E005
now report repairable:false. --backfill alone now actually triggers
backfillMilestones, fixing a latent bug where its gate was unreachable
without --repair also being set (verify.cts:2504, confirmed dead code
pre-migration).

Test updates distinguish the two explicitly-authorized behavior
changes (DESTRUCTIVE refusal, backfill-alone fix, W021->W026 split)
from preservation — every changed assertion is commented with why, and
new regression tests were added for both changes plus W021/W026
mutual independence. Drift-guard bookkeeping (bypass-baseline shrunk
to the one disclosed W024 exception, milestone-window and
phase-enumeration exemptions, test-file-count allowlist) updated for
the relocated/new functions this migration introduces.
2026-08-13 02:28:49 -04:00
JusticeWay
77374cfc32 fix(#2528): resolve digit-leading phase directories by bare number (#2559)
* fix(#2528): resolve digit-slug phase dirs by bare number — tokenizer rewind, shared bare-integer fallback, resolution-path parity gate

extractPhaseToken welded 2-digit slug words onto the phase token (phase 10
named "24/7 Autonomy" -> dir 10-24-7 -> token 10-24), making digit-prefixed
phase names unresolvable by bare number across every phase verb.

- phase-id: continuation segments must be the PURE 2-digit zero-padded form
  the write side emits; a 1-digit terminator rewinds the absorbed run
  (10-24-7 -> 10) while >=2-digit terminators keep the locked #2232
  round-trip (14-06-2026-photos -> 14-06).
- phase-id: new matchPhaseDirs owner — primary exact-token match plus a
  bare-integer leading-digit-run fallback for shapes the tokenizer cannot
  rewind (05-80-20-cleanup); collisions stay #2237-loud.
- locator/find-phase/phase-plan-index all delegate selection to the owner;
  plan-index gains the previously missing multi-match guard.
- tests: #2528 unit + fast-check metamorphic blocks; new 9-scenario
  resolution-path parity gate across all three paths.

Fixes #2528

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2528): add changeset for PR #2559

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2528): align validation token grammar

* fix: address phase token review

* fix: restore phase grammar parity for numeric slugs

* docs: document digit-leading phase resolution

* docs: clarify ambiguous phase resolution behavior

* docs: register canonical phase directory selectors

* fix: align prefixed deep phase token parsing

* fix(#2528): route the fourth resolution site through matchPhaseDirs

Review BLOCKER. smart-entry.cts::detectVerifyFailed resolved the current
phase's directory with its own `.find(phaseTokenMatches)` and never
reached the shared selection, so the bare-integer-fallback family the
issue names — `05-80-20-cleanup`, `30-12-factor-refactor` — resolved
nowhere. The miss is silent by construction: an unresolved phase reports
"not failed", which is byte-identical to a healthy one, so a failed
verification simply never surfaced in /gsd or /gsd:progress.

`entries` is already sorted and matchPhaseDirs filters without
reordering, so matches[0] reproduces the previous selection exactly
wherever the old code resolved at all.

Wiring it into phase-resolution-parity.test.cjs as a fourth path then
exposed a second, older defect in the same function: phaseTokenFromDirName
shape-probed the UNSTRIPPED token, so a project-code-prefixed directory
(`MEM-05-…`, tokenizing to `MEM-05-80-20`) failed the leading-digit test
and was dropped before any resolution ran — every phase in a
project-coded plan was invisible to this check. The probe now runs on the
stripped token; the returned value is unchanged, so the comparePhaseNum
sort is untouched.

Path 4 has no JSON surface to compare, so the gate observes selection
indirectly: plant the failing artifact in exactly one directory and a
passing one everywhere else, then read the boolean. Reverting either fix
turns 5 of the 10 corpus scenarios red.

* refactor(#2528): collapse the duplicated extractCanonicalPlanId

Review MAJOR. The function existed as two independent, byte-identical
copies — src/core-utils.cts and src/phase.cts — and this PR had to patch
BOTH with the same single-digit-slug rewind rule. That is the generative
fix divergence CLAUDE.md names, and only the core-utils copy was under
test, so a future one-sided patch would have silently split plan-id
canonicalization between the plan listing and everything else.

Removed rather than parity-tested: core-utils was already the leaf owner
and already exported it, and phase.cts already imported that module, so
there is no second surface left for a parity test to police.

* test(#2528): pin matchPhaseDirs at the digit-width boundaries

Review MAJOR. The bare-integer fallback's correctness rests entirely on
capturing each directory's whole leading digit run before the zero-strip
compare; a regex that stopped short would turn every query into a prefix
match, and "1" would claim 10, 100, and 12 alike. The existing coverage
was example-based and never touched that boundary.

Adds the explicit 9/10 and 1/10/100 cases — including the forms where
only the wider directories exist, so an exact-width neighbour cannot
satisfy the assertion — plus a fast-check property over arbitrary
distinct leading runs. The property is stated as an invariant on the
result (every returned directory's leading run IS the query) rather than
an expected list, so it covers primary and fallback matches alike and
cannot be satisfied by reimplementing the selection in the test.

Both fail when the fallback regex is degraded to a prefix match.

* fix(#2528): route the remaining eight consumers through matchPhaseDirs

phaseTokenMatches had eight consumers left that each rebuilt the directory
selection around it by hand: phases-list, next-decimal, phase-remove, the
W021 milestone-consistency check, schema-drift, the init-manager overview,
milestone-complete's disk check, and roadmap analyze. Every one of them
reproduced the reported symptom in full after the tokenizer was fixed.

None of them derives a displayed phase number from the matched directory,
so none needs phaseNumberForMatch; the change at each site is the
selection and nothing else. matchPhaseDirs filters without reordering, so
matches[0] reproduces the prior .find() choice wherever the old code
resolved at all.

phaseTokenMatches now has no call sites outside phase-id.cts. It stays
exported as the primitive matchPhaseDirs is built from and as a pinned
canonical surface, but no consumer reaches past the owner to it.

* test(#2528): extend the parity gate to the migrated consumers

Each of the eight is observed through the surface a user sees, not
through the matcher, with a no-directory control so the assertions cannot
be satisfied by a consumer that resolves unconditionally. init-manager
and roadmap-analyze are additionally asserted to agree with each other.

* refactor(#2528): own the case-flexible phase grammar and the leading-digit-run fragment

validate.cts derived its case-flexible regex sources by running
`replaceAll('A-Z', 'A-Za-z')` over two constants exported by phase-id.cts.
That passes lint-phase-id-drift.cjs — there is no literal copy of the
grammar — but it depends on the owner rendering that exact substring. The
day phase-id.cts expresses the same class any other way the replaceAll
silently no-ops and validate.cts narrows to uppercase-only. The failure
mode is a NON-match, so nothing throws and no uppercase-only fixture
notices. Both variants are now derived once, beside the sources they
widen, and imported.

Also names the leading digit run the bare-integer fallback selects on.
It was spelled `/^(\d+)(?:-|$)/` where the fallback filters and `/^\d+/`
where phaseNumberForMatch reads the number back off the winner; selecting
on one run and displaying another would resolve a directory and then label
it with a number that never matched it.

* fix(#2528): refuse to remove a phase when two directories claim its number

cmdPhaseRemove was the only migrated site taking matches[0] with no
multi-match guard. Every sibling resolution path returns ambiguous_matches
and refuses to choose; this one is the DESTRUCTIVE path, so choosing
silently is strictly worse than anywhere else. With 05-80-20-a and
05-90-till-late on disk, `phase remove 5 --force` deleted one of them and
renumbered every phase after it — where the base resolved nothing, deleted
nothing, and the corpus in tests/phase-resolution-parity.test.cjs already
declared that exact input ambiguous.

The refusal is emitted before any file is touched and carries both
candidates. CONSUMER_SCENARIOS could not express the case — every row is
binary, resolving to one directory or to none — so the gate gains a
dedicated ambiguous test. It asserts on the filesystem, not only on the
reported directory_deleted: a null printed after an rmSync would satisfy
every other check.

* fix(#2528): pair digit-leading phase directories with their roadmap phase in validate health

W006/W007 are the ninth site of this bug class and the one a
`phaseTokenMatches` grep could never surface: they resolve roadmap↔disk by
intersecting TOKEN SETS, which is a dir→token labelling rather than the
query→dir selection matchPhaseDirs owns. On the canonical fixture the
label is wrong in both directions at once, so `validate health` reported
"Phase 5 in ROADMAP.md but no directory on disk" AND "Phase 05-80-20
exists on disk but not in ROADMAP.md" for the same directory.

collectDiskPhases now keeps the directory names behind each token, so
W006 can ask the canonical matcher whether a roadmap phase resolves to a
real directory, and W007 — which iterates directories and therefore has no
query to resolve — gets the inverse mapping it never had: a directory is
claimed when some roadmap phase resolves to it.

Both checks are additive: the token intersection still decides every shape
it already decided, and the resolution can only REMOVE a warning. The
regression test carries controls in the other direction — a roadmap phase
with no directory must still raise W006, an unclaimed directory must still
raise W007 — so it cannot be satisfied by a check that stopped reporting.

* docs(#2528): state and pin the directory-side scope of the bare-integer fallback

The matchPhaseDirs docblock claimed deep-decomposition lookups were
untouched. That is true of the QUERY side only — no non-bare query enters
the fallback — but the DIRECTORY side is what changed classification: a
bare `5` now reaches a lone `05-01-auth` and resolves it (phase_number
"05", phase_name "01-auth") where the base found nothing.

The widening is irreducible from directory names alone. `05-01-auth`
(sub-phase 5.1) and `30-12-factor-refactor` (phase 30 named "12-Factor
Refactor") are the same `NN-NN-<slug>` shape, and the discriminator that
would separate them — "is the second segment a valid decimal sub-phase" —
accepts `5.1` and `30.12` equally. Any rule strong enough to exclude the
first excludes the second, which is the defect #2528 exists to fix. So the
tie is broken in favour of resolving, the docblock now says so, and the
consequence is bounded where it matters: two such directories are two
matches, and every caller (including phase remove) refuses to choose.

Pins both directions, since nothing observed the directory side before.

* fix(#2528): count surviving phases by identity in phase remove's STATE resync

#2640 landed on `next` while this branch was open. Its STATE.md phase-count
resync re-derives "which directory was removed" from the query with
`phaseTokenMatches`, which is the tenth site of this issue's defect: the
bare-integer fallback resolves `05-80-20-cleanup` for query `5`, but the
token predicate does not, so the just-deleted directory is counted as still
present and the written `Total Phases` is one too high.

`targetDir` already IS the directory that was removed, and the block is gated
on it being non-null, so identity answers the question exactly — which is also
what the comment above the filter already claimed it did. This keeps
`phaseTokenMatches` out of `phase.cts` rather than re-importing it to satisfy
one call site: the module's public surface should not grow for a question that
does not need re-derivation.

Pinned in the parity gate with a control on a directory the tokenizer reads
correctly, so the assertion is about the digit-leading shape and not about the
counting rule changing for everything.

* fix(#2528): let the resolution layer own the digit-leading slug family alone

The tokenizer rewind this fix carried — pop the last absorbed continuation when
the segment that stopped the scan is a bare single digit — reads
"10-24-7-autonomy" (phase 10 named "24/7 Autonomy") correctly and silently
re-reads "10-24-7-zip" (sub-phase 10.24 named "7-Zip Integration") from "10-24"
to "10". The two names are string-identical in shape, so no local signal
separates them; the rule traded the reported ambiguity for the symmetric one a
level down, on a 15-caller chokepoint whose output also feeds query-less
derivations (STATE.md phase counts, W007, the #2562 key surface). A well-formed
sub-phase directory became unresolvable by its own id — the very symptom #2528
was filed about.

It also bought nothing. The bare-integer fallback in matchPhaseDirs already
resolves "10-24-7-autonomy" for query "10" whatever the token is: no primary
match, bare query, leading digit run "10". The reported case was covered twice,
by two rules, and the two disagreed about the case nobody reported.

So the rewind is removed rather than narrowed, in the tokenizer and in the five
surfaces kept in lockstep with it (BRACKET_PHASE_TOKEN_SOURCE,
PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem's pair grammar and its collision
branch, roadmap-parser's numericRe, extractCanonicalPlanId), together with the
SINGLE_DIGIT_RUN_SEGMENT_SOURCE owner constant they shared. Disambiguation now
lives only where a QUERY exists to disambiguate against, which is the same
bounded mechanism the "05-80-20-cleanup" shape already used.

Measured, not argued: over 29800 generated directory names, extractPhaseToken is
byte-identical to `next` on every input except the lowercase-continuation class
("01-20a", "05-80-20-25abc") — a rule about the segment itself, not a guess about
its neighbour.

Both readings now stay reachable by their own ids:
  matchPhaseDirs(['10-24-7-autonomy'], '10')    -> the dir   (fallback)
  matchPhaseDirs(['10-24-7-zip'],      '10')    -> the dir   (fallback)
  matchPhaseDirs(['10-24-7-zip'],      '10-24') -> the dir   (primary)

* test(#2528): pin the one-continuation boundary the rewind had no coverage for

The regressing shape was invisible to the suite by construction, not by luck:
the deep-rewind property built its cases from `continuationArb` with
`minLength: 2`, so it never exercised the single-continuation case — exactly one
genuine sub-phase level before a digit-leading slug — and every hand-written
fixture used the ambiguous shape only where "phase-plus-slug" was the intended
reading.

`continuationArb` is now `minLength: 1` and the property states the invariant
instead of the old rule: for any prefix, phase, 1-5 continuations and any
one-digit terminator, the token equals the FULL continuation run on both the
imperative and the regex surface, and `matchPhaseDirs([dir], token)` returns that
dir. That third assertion is the one that catches the class on its own — the old
behaviour made a well-formed directory unresolvable by its own id, which is a
property, not a fixture.

Around it: "10-24-7-zip" and "10-24-3d-printer" now sit beside
"10-24-7-autonomy" everywhere the family is pinned, so the two readings can never
diverge again; the 24/7 metamorphic property asserts the RESOLUTION result rather
than the token (the token is precisely the part no surface may decide); the
end-to-end parity corpus gains "a sub-phase with a digit-leading slug resolves by
its full id" across all four resolution paths; and the milestone-scoping residual
is pinned in three directions rather than left to prose.

Mutation: re-inserting the rewind and rebuilding turns 9 tests red, the
`minLength: 1` property first, and nothing else. Build success checked separately
(build:lib reports 0 `error TS`), so the mutation reached the artifact under test.

* test(#2528): pin the #2946 guard against digit-leading phase directories

The #2946 fix makes the milestone-complete unstarted-phase guard run
unconditionally, so whether it fires now rides entirely on the
directory-resolution owner this PR replaces. Two cases, both with STATE.md
carrying no `milestone:` field so the #2946 path is the one exercised:

  - ROADMAP Phase 5, disk `05-80-20-cleanup` → guard must stay silent.
    RED on next (fail-closed: the guard blocks a legitimate one-way-door
    operation because phaseTokenMatches resolves neither 05 nor 80 for
    that directory).
  - ROADMAP Phase 80, same directory → guard must still fire. Green on
    both sides; it pins the fail-open direction against a future widening
    of the matcher.

* fix(#3175): stop the injection scanner reading RegExp.exec as code execution

The apostrophe fix in 27aa40f6 replaced ["\x27] with a real ["'] class.
The old class never contained an apostrophe at all (POSIX bracket
expressions do not honour backslash escapes, so it was the set ", \, x,
2, 7), so only exec(" matched. Single-quoted method calls now match for
the first time, and RegExp.prototype.exec takes a subject string, not
code: any PR touching a file that tests a regex goes red. On next, six
files match the scanner's own pattern across 16 method calls.

A plain [^[:alnum:]] boundary cannot separate the two forms because . is
not alnum, so exec gets [^[:alnum:].] and the command-execution vector
moves to a dedicated member-call pattern. Bare exec('rm -rf /'),
cp.exec(...) and child_process.exec(...) all still fire.

Mutation: reverting the boundary reds 1 test and only it; removing the
member-call pattern reds the 2 non-weakening tests and only them.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(#2528): keep exec( detection receiver-blind, allowlist the two grammar suites

The left boundary [^[:alnum:].] added in 58a7b560 excluded a preceding dot,
which dropped every member-position .exec('…') from the scanner. The follow-up
receiver pattern only restored three literal spellings (child_process,
childProcess, cp), so require('child_process').exec('…') — the most common Node
spelling of the vector this pattern exists to catch — became invisible, along
with any opaque receiver (conn.exec, shelljs.exec).

Revert the pattern to its receiver-blind form and handle the RegExp.prototype
.exec false positive where the script already handles this class: per-file
ALLOWLIST entries for the two phase-token grammar suites. Mutation-checked —
removing the two entries reds exactly those two files and nothing else.

The four assertions written around the old patterns are replaced by a
table-driven set covering all six spellings, including the three the narrowed
pattern silently lost.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(#2528): pin the three undeclared grammar edges, correct the ambiguity claim

Review round 10 asked for declaration, not behavior change, on four items. All
four have zero production consumers or preserve their caller's prior rule, so
each is pinned as a test or corrected in prose rather than reverted.

- BRACKET_PHASE_TOKEN_SOURCE: the (?=-|$) terminator is what keeps the bracket
  read path in step with the other surfaces, and it costs the display shapes
  (`05.03: Title`, `12A: X`, `05.03]` no longer tokenize). Pinned so widening
  the terminator class is a deliberate act rather than a lookahead deletion.
- canonicalPlanStem: uppercase plan suffixes still strip, lowercase and dotted
  sub-plans now fall through. Dead export; pinned as a decision on record.
- getMilestonePhaseFilter: `12A-01-foo` now yields `12A-01`, matching what
  `12-01-foo` has always yielded. The letter suffix was the only reason a
  sub-phase directory folded into its parent phase's milestone window; the two
  shapes now agree. Not named in the review — found auditing the same commit.
- matchPhaseDirs docblock claimed every caller refuses on multi-match. Four do;
  five take matches[0]. Replaced the claim with the actual two-tier policy and
  the honest caveat that the bare fallback makes multi-match newly reachable
  for queries that previously found nothing.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-12 20:45:04 -04:00
Tom Boucher
ae7dc52972 fix(#3225): guard W006/W007 + consistency loops with isSentinelPhaseId (#3371)
* test(#3225): sentinel phase dirs no longer trigger W007 / consistency warnings / gaps

The W006/W007 (validate health) and the parallel consistency disk↔roadmap and
gap-numbering loops never got the isSentinelPhaseId guard that phase.cts has
(#2786/#2949), so every sentinel phase dir (999.x/0.x — never-on-roadmap by
convention) produced a spurious W007 and a spurious 'Gap in phase numbering:
N → 999'. Add failing-first regressions for both surfaces + the gap check, each
with a non-sentinel orphan negative-space guard.

RED — fails on next; fix follows.

* fix(#3225): guard W006/W007 + consistency + gap loops with isSentinelPhaseId

cmdValidateHealth's W006/W007 loops, cmdValidateConsistency's parallel disk↔
roadmap loops, AND its gap-numbering check never got the isSentinelPhaseId guard
that phase.cts has at 10+ sites (#2786/#2949). So any repo using the sentinel-id
convention (999.x backlog/interim, 0.x drafts) got a permanent spurious W007 and
a spurious 'Gap in phase numbering: N → 999', with advice to add-to-roadmap
(violates the convention) or delete (destroys archived work).

Add isSentinelPhaseId to the phaseIdMod destructure and skip sentinel ids in:
W006 + W007 (cmdValidateHealth); the two plain-warning disk↔roadmap loops and
the gap-numbering integerPhases filter (cmdValidateConsistency — same bug family,
folded in inline per no-silent-defer). Additive only: non-sentinel orphans and
real numbering gaps still warn. The gap-numbering guard was surfaced by the
isolated review (a 999-interim dir would otherwise create a false 'N → 999' gap).
Same family as #3167 (since fixed).

* chore(#3225): add changeset fragment

* chore(#3225): backfill changeset PR number (#3371)

---------

Co-authored-by: sim <sim@local>
2026-08-11 21:49:39 -04:00
Rezolv
e87fb409ee enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint

Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age
freshness proxy (state_commits_behind / state_commit_stale) through
state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as
health W024. The proxy is advisory: classify() deliberately does NOT consume
it (ADR-1787 locks the classification/routing boundary — a signal, not a route).

Composes with #3099 and #1882 (both merged to next after this branch): the
commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE
diagnostic reads `last_activity` — two different fields, not "two staleness
signals on one field." A new regression test asserts a STATE.md carrying both
an unparseable last_activity AND a valid state_head resolves each independently
(diagnostic fires once; freshness reads state_head, commits_behind 0).

Rebased onto next (flattened): resolved the add/add conflicts in
src/smart-entry.cts (kept both the #2573 freshness import/derivation and the
#3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both
describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B).
Tests: smart-entry 62, state/state-transition/health/verify 639, all pass.

* chore(#2573): allowlist health-validation test in the prompt-injection scan

The scanner's `exec('` code-execution pattern matches the benign
`re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests
(pre-existing: 16 such calls on next, this PR adds none). The file entered the
diff-mode scan's changed-file set only because #2573's W024 state_head
assertions touch it. Allowlist it alongside the other test files that carry
pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION
class). Scanner self-test 38/0; diff scan 14 files, 0 findings.
2026-08-11 17:10:23 -04:00
Tom Boucher
4a1ed2531f enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check

Specifies ADR-2313 D6 before the implementation exists, so the tests
bind to the contract rather than to whatever the code happens to do.

RED is established by construction, not by a remote run:
checkCodexModelPosture and POSTURE_REASON are absent from the compiled
lib today, so every row fails on the missing export. A remote checkpoint
here would prove only that the function is missing, which is already
known — so the run is deliberately deferred to the combined green
checkpoint rather than spent proving a tautology.

That makes the NEGATIVE PROOFS the rows that carry real signal. Every
positive row passes even for a naive implementation that greps
/model\s*=/ over the whole file. Six rows fail it: light-tier
service_tier/model_verbosity decoupling (#774), hand-added keys, a
commented pin, the model_verbosity key-prefix collision, the runtime
no-op ordering, and the headline case — a literal `model = "sonnet"`
inside the developer_instructions ''' block, which the emitter fills
with agent prompts that discuss models constantly.

Row 14's fixture was verified to discriminate before being written: a
whole-file scan matches it and a header-slice scan does not. Without
that check the test would pass trivially and prove nothing, which is the
vacuous-test failure this epic has already hit repeatedly.

Adversarial TOML fixtures are hand-authored against the real Codex shape
rather than generated by generateCodexAgentToml, per #2371 — a fixture
from the writer can only confirm what the writer already believed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3242): validate codex .toml model posture, not just presence

Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export,
not a branch inside checkAgentsInstalled — that function carries 33
upstream dependents, cyclomatic 25, and sits in two traced process
flows, so it is deliberately left untouched.

It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf.
That is what Phase 1's constant move bought: agent-install-check is
documented as pure read/verify and imports only leaves, so reaching the
rule through model-resolver would have dragged config-loader into it.

Reads liberally, judges strictly, and never guesses. Tolerates comments,
key order, whitespace, CRLF, and a BOM; anchors on full key names so
model_verbosity does not satisfy a `model` probe; treats extra
hand-added keys as none of its business, since the check is a predicate
on the two fields the posture owns rather than a whitelist over the
document. An unreadable file becomes a named violation and the loop
keeps going.

The scan covers only the header slice — the lines before the
developer_instructions ''' marker. The emitter writes agent prompts into
that block and GSD's prompts discuss models constantly, so a whole-file
scan reports violations for prose. This is the highest-risk defect in
the phase and the reason its fixture was verified to discriminate before
being written.

The non-codex short-circuit runs before any filesystem call, so a stray
.toml under another runtime is never inspected.

Wired through cmdValidateAgents as an additive codex_posture key, so a
violating install is visible from a command a user actually runs rather
than only from a library nothing calls.

Also fixes a test defect found while implementing: .gitattributes forces
`* text=auto eol=lf` repo-wide, so the committed CRLF fixture was
normalized to LF in the index — `git ls-files --eol` reported `i/lf
w/crlf`, the working copy being stale pre-normalization bytes. The CRLF
row was asserting against a file that could not survive a fresh clone.
CRLF is now derived at runtime, which puts it under the test's control
rather than git's, instead of adding a .gitattributes exception that
fights a deliberate repo-wide policy and that anyone could re-normalize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3242): document the codex posture check where a user will look

Three quadrants, filed by where the reader actually arrives.

How-to (recover-and-troubleshoot.md, under Install and update problems)
is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400
about an unsupported model" — and opens with the verbatim error string.
Someone hitting this does not know the words "posture" or "ADR-2313";
they have a 400 in their terminal and will search for that.

Reference (COMMANDS.md) had no `validate agents` entry at all, though
sibling gsd-tools subcommands are documented. Adding user-visible output
to an undocumented command and then linking to it from the new how-to
would have left a dangling reference. The entry carries the
violation-reason table, since the frozen POSTURE_REASON enum is the
machine-readable contract a reader needs rather than the prose.

Both surfaces state that presence and posture are separate verdicts — a
missing agent lands in `missing`, never as a posture violation. That is
a deliberate design decision and would otherwise be invisible to someone
watching one command emit both.

Explanation stays in ADR-2313, which already covers D6 and the
liberal-parse/strict-judge boundary. Pointing at it beats duplicating it
into COMMANDS.md and creating two copies to drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): close two false negatives in the posture scan

Both found by an isolated reviewer and reproduced before fixing. Both
made the check report clean when it was not — the worst direction for
this function, since the how-to tells users an empty violations list
means the install is posture-clean.

A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML,
but the key pattern required a bare identifier, so the pin was silently
invisible. Bare, "double" and 'single' quoted forms now normalize to the
same key name.

The block marker was found by unanchored whole-content search and used
to truncate the header. A `description` value merely containing the
literal text `developer_instructions = '''` truncated the scan before a
real pin, and a user who hand-reordered `model` to sit after the block —
still legal TOML — was never scanned at all.

Fixed by changing the strategy rather than the regex: find the block's
range, anchored at line start, and scan every line OUTSIDE it. That
covers both failures and is strictly more correct than truncation, while
still never reading prompt prose. An unterminated block excludes the
rest of the file, which fails toward a false positive — the safe
direction, since misreading prose as a pin wastes a user's time while
the alternative hides a real one.

Also corrects two overclaims of mine. The how-to named "v1.11", a
version that does not exist — package.json is 1.10.0 and unreleased — so
it now describes the boundary by behavior and links the ADR. And the
test matrix asserted that a naive whole-file scan "fails exactly rows
12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16
produce the correct result against that baseline too. They guard real
but *different* mistakes, and the matrix now says which one each catches
instead of attributing them all to the header-slice defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): skip symlinked agent files instead of following them

Security review finding. The scan listed entries with readdirSync and
read them with readFileSync, which follows symlinks — so a symlink in
the agents directory pointing anywhere would have its contents read, and
any line matching the model pattern echoed into cmdValidateAgents'
output through the `value` field. A read-and-echo primitive on an
arbitrary path.

It needs write access to the agents directory, so it crosses no new
trust boundary today. Fixed anyway, for two reasons.

This repo already does it correctly next door: cmdEffortSync filters
with lstatSync().isFile() and the comment "Skip symlinks — only write
regular files to avoid clobbering symlink targets." Being inconsistent
with a sibling in the same subsystem IS the defect.

And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these
files. Establishing symlink-following as the house pattern for Codex
.toml handling here would hand Phase 3 a worse starting point while it
writes rather than reads.

Skipped silently rather than reported, matching the sibling: a symlinked
agent file is a structural install choice, which checkAgentsInstalled
owns, not a model-content posture defect. An lstat that itself throws
excludes the file rather than crashing the scan.

That does narrow the guarantee slightly, so the how-to now says an empty
list means every REGULAR .toml is clean, and tells anyone symlinking
their configs to check the targets by hand. Claiming a clean bill of
health over files the check declined to open would be the same kind of
false confidence the two false negatives above produced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3242): backfill changeset pr number (#3290)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 21:50:49 -04:00
Tom Boucher
86bebcefa2 refactor(#3216): bind milestone identity to the canonical locator (#3226)
* refactor(#3216): widen milestone-window guard to literal-## matchers

The guard keyed only on the `#{N,M}` quantifier plus a literal version or
phase-lookahead token. getMilestoneInfo hand-rolls its milestone-heading match
with a literal `^##`/`## ` and an interpolated ${escapedVer}, so it satisfied
neither token and the guard reported a clean zero on a file carrying live
re-derivations (#3171, #3197) — a zero it did not earn.

Widen token (a) to a literal 2-6 `#` run, admitted ONLY inside a heading-MATCHER
literal (a regex literal, or a string/template handed to new RegExp) so a
heading-BUILDING template is not mistaken for a re-derivation. Widen token (b)
with the grouped `v(\d+(?:\.\d+)+)` shape and an interpolated version
placeholder.

Ships BEFORE the consolidation per ADR-3180 s7.2: a guard widened afterwards
measures an already-cleaned surface. It is expected to be RED until the
consolidation lands.

* test(#3216): failing-first milestone-identity single-owner suite

63 tests across two files, from the matrix in .gsd/phase/. Section H of
milestone-window-single-owner.test.cjs covers the 21 input classes of the
design's behavior table plus its negative space; milestone-window-drift-guard
covers the widened tokens and proves the exemption is function-scoped, not
file-scoped.

Copy count is 3 found by the guard, not 1 per the epic (ADR-3180 Amendment 3's
standing rule, holding for the fourth consecutive phase): both getMilestoneInfo
sites plus cmdRoadmapAnalyze's milestone enumeration at roadmap.cts:454, which
carries the same #3171 truncation and #3197 phase-heading confusion.

Expected RED until the consolidation lands.

* refactor(#3216): bind milestone identity to the canonical locator

getMilestoneInfo hand-rolled two milestone-heading regexes inside the owner's
own file. Both were wrong, differently: the STATE-version site's ^## anchor is
level-blind so [^\n]* absorbs a third #, and the fallback site had no anchor at
all, so '## ' matched from the second # of '###'. Against
'### Phase 7: Close v3.3 gaps' the fallback returned {v3.3, gaps} (#3197). Both
captured names with [^\n(], truncating at a parenthetical (#3171).

Bind both to the canonical grammar. locateMilestoneHeadings becomes a
version-filtered view over one shared source, and a new version-agnostic
listMilestoneHeadings enumerates milestone headings for callers that need all
of them. getMilestoneInfo returns ScopedResult<MilestoneInfo|null>; the
{v1.0,'milestone'} default, which was output-identical to a real v1.0 project,
is deleted. The #2245 never-throws invariant is preserved.

Copy count: 3 found by the guard, not 1 per the epic. The third was
cmdRoadmapAnalyze's own milestone enumeration (roadmap.cts:454), carrying both
defects in the implementation the epic blessed.

buildStateFrontmatter and archivePhaseDirectories branch on scope: the first
writes null rather than a fabricated identity, the second falls through to its
dated-label fallback. A fabricated v3.3 passes ARCHIVE_VERSION_LABEL_RE, so it
would otherwise misfile phase history.

Also fixes an unsafe cast in init.cts that masked these type errors across five
call sites, which would have shipped undefined milestone fields under green tsc.

* fix(#3216): restore the #1761 unbounded guard and bullet precedence

Review and the first full-matrix run surfaced five real defects in the
consolidation, all fixed here rather than by relaxing the tests that caught
them:

- buildStateFrontmatter gated its isMilestoneBoundedInRoadmap check on the
  scope-gated milestone value, which is null on any non-COMPLETE scope, so the
  #1761 unbounded guard was silently skipped and state json reported a percent
  it must omit. It now gates on the STATE-asserted version, independent of
  identity scope.
- The rewrite lost #2135's precedence: the name-bearing progress-marker bullet
  is consulted before the heading again.
- A single-segment version (v3, no dot) did not resolve; the name-extraction
  fallback now accepts it.
- A version carrying regex metacharacters, or a $& / $1 replacement pattern,
  is matched literally.
- listMilestoneHeadings' heading field trimmed, so a CRLF roadmap no longer
  leaks a trailing carriage return into roadmap analyze's output.

Also emits milestone_version / milestone_name / current_milestone as explicit
null rather than omitting the key, so the prompt layer cannot render a bare
placeholder, and corrects an init.cts comment plus a cast left inconsistent.

* test(#3216): update milestone-identity expectations to the scoped contract

getMilestoneInfo returns ScopedResult<MilestoneInfo|null> and the
{v1.0,'milestone'} default is deleted, so the suites asserting the old shape
assert removed behavior. Updated rather than weakened: every touched call site
now asserts the scope explicitly against the frozen SCOPE enum.

roadmap-parser.test.cjs: 20 expectations moved to {value,scope}. The #1881
unreadable-vs-absent diagnostic assertions are untouched and still prove their
original point — only the return shape moved. One pre-existing assert.ok(info)
is now a specific UNSCOPED assertion, so that case is stronger than before.

new-milestone-clear-phases.test.cjs: the test asserting phases clear archives
under the v1.0 default now asserts the dated archived-<YYYYMMDD> fallback,
which is the deliberate consequence of deleting that default.

Two of this branch's own tests were also corrected after they drove the
implementation the wrong way: the parity test compared raw heading text and so
pushed a stray ## prefix into roadmap analyze's public output, and the hostile
metacharacter row demanded a pathological version resolve, which pushed a
widening of the ADR-locked \b boundary. Both now assert what the contract
actually requires.

* docs(#3216): document milestone identity and correct the CONTEXT.md entry

ADR-3180 s7.2 moves to Enforced and gains two rules that were unstated: the
name derives from the heading's own version token and drops a trailing status
marker, and a free-form legacy ROADMAP with no version anywhere is UNSCOPED
with no identity rather than a defaulted v1.0 (decided by the maintainer before
implementation, per s7's own rule that an unstated behavior is not decided).
Amendment 4 records Phase 6's validation, including that the copy count was a
lower bound for the fourth consecutive phase.

CONTEXT.md's Roadmap Parser entry described locateMilestoneHeadings as
boundary-matched with (?![\w.-]) — the alternative Amendment 2 tried and
REVERTED. The code uses \b and says so, and the ADR agrees; the revert updated
code and ADR and missed CONTEXT.md, which is the epic's own fixed-on-one-copy
failure class in the docs layer, on a file that is itself a PR gate.

* fix(#3216): persist the real version on a truncated identity

buildStateFrontmatter wrote null for BOTH milestone and milestone_name on any
non-COMPLETE scope, discarding a real version. ADR-3180 s7.2 rule 6: a version
known with no resolvable name is TRUNCATED carrying {version, name: null} —
'the version is a real answer, the name is a non-answer, and collapsing the two
is the failure this contract exists to prevent.'

The two fields are now gated by what is actually known: the version whenever one
exists (COMPLETE or TRUNCATED), the name only on COMPLETE. Never fabricated.

Caught by this phase's own Decision 4(c) consumer-output test, which is the
argument for asserting at the consumer rather than the owner — the owner was
correct throughout; only the consumer collapsed its answer.

* refactor(#3216): extract helpers and make cmdCommit's scope gate explicit

From the two-axis code review:

- init.cts repeated the identical getMilestoneInfo cast at five sites with
  copy-pasted comments — duplication inside a PR whose thesis is that duplicates
  get deleted. Extracted milestoneRecord(cwd); the one site-specific comment is
  kept, the four generic copies removed.
- getMilestoneInfo hand-built its { value, scope } literal at ten return points;
  a local scoped() constructor now does it once. Every per-branch rationale
  comment is preserved and no returned value or scope changed.
- cmdCommit gated the milestone branch name on plain truthiness, which is also
  true for TRUNCATED, so an unresolved identity drove branch creation
  incidentally rather than deliberately. It now gates on the SCOPE enum,
  accepting COMPLETE or TRUNCATED because both carry a real version, and the
  comment records why that differs from archivePhaseDirectories — which demands
  COMPLETE because it uses the value as a filesystem path component.

* test(#3216): cover the bare-version-in-prose truncated path

The spec review found the bareVersionMatch path — no STATE version, no
milestone heading, a version token only in prose — returning TRUNCATED with no
test exercising that exact shape, violating Decision 4's boundary-coverage
requirement.

* docs(#3216): record the missed Tier-2 surfaces and rule 5's corollary

Decision 3 requires an explicit call-out for EVERY Tier-2 change, and Amendment
4's first draft named eight surfaces while the change touched thirteen. Adds
cmdCommit's branch-name construction and the four init JSON bundles, an
incomplete list being the same defect in miniature that this epic removes.

s7.2 rule 5 gains a corollary separating two cases the original wording ran
together: no version token ANYWHERE is UNSCOPED, while a bare version token in
prose or a non-milestone heading is weak but real evidence and yields TRUNCATED
under rule 6.

* chore(#3216): set changeset fragment pr to 3226

---------

Co-authored-by: sim <sim@local>
2026-08-08 19:06:13 -04:00
Tom Boucher
343835facc refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans

scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one
independent re-derivations across seven modules now route through it, and
scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather
than an allowlist (ADR-3180 Decision 4a).

The epic scoped this at three copies. A whole-repo guard found twenty-six sites
across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3
narrows to window plus sentinel enumeration.

Two sites are exempt with a documented reason rather than a bare allowlist:
audit.cts scans one quick task's own directory for a single completion record,
and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time
import. Neither is a phase directory.

scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one
owner answers both questions: verify.cts's numbering-gap check wants every plan
on disk, its pairing check wants the live set. Both fields are additive.

Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave
scheduling, was scheduling status:superseded plans into waves and reporting zero
plans for the post-#3139 nested layout.

filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned
them and only their own tests still called them.

New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator,
with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and
the CONTEXT.md glossary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice

The contract held; the phase boundary did not. The whole-repo drift guard found
26 re-derivations across 9 files against the epic's estimate of 3, and
cmdProgressRender re-derives both enumeration and plan counting on adjacent
lines, so DW4 was unsatisfiable within Phase 1's original file scope.

Records the amended scope, scanPhasePlans's new allPlanFiles field,
findOrphanSummaries, the two documented exemptions, the re-derived Tier-2
table, and the describeNonCanonicalPlans trap for later phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): complete the canonical pairing rule and gate the naming diagnostic

The remote runner went red with 13 deterministic failures on both lanes,
and they were right: replacing verify.cts's canonicalPlanStem pairing with
summaryCandidates dropped a case the bespoke rule covered. A plan carrying a
descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its
canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no
such candidate, so the plan read unsummarized.

The fix is to complete the one rule rather than restore a second:
summaryCandidates gains a canonical-id candidate, narrowed to fire only when an
id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans
and findOrphanSummaries all inherit it. The two-plans-one-summary collision
behaviour of the original rule is preserved deliberately and documented in
place.

Second defect, independently root-caused while verifying: routing the #2893
naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i
fallback, which is correct for counting and wrong for a naming check — a
non-canonically-named file was accepted as a valid plan and the diagnostic
went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect
with a strict isCanonicalPlanFile predicate before reporting names.

Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180:
a question about file naming wants the physical, strictly-matched set; only a
question about outstanding work wants the live set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): register planning-scope.cjs in the eslint migration list

tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor
ignored per its ADR-457 migration state. The new planning-scope module closed
five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md
and the CONTEXT.md glossary - but not eslint, because that one is enforced by a
test rather than by lint:ci, so the local pipeline stayed green while it was
missing.

Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is
linted, matching every other migrated module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): replace the plan-count drift detector with a literal tokenizer

CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the
backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an
escaped \.md". Five review rounds found it had two defects, not one:

  - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.`
    pair be consumed either as one escape or as two class characters, which is
    exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`.
    Excluding `\` from the class killed that but left a cubic path — 23ms at
    N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with
    no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI
    runs on fork pull requests, so a crafted src/*.cts could stall the job.
  - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g.
    `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the
    literal at that `/`, so the scan never reached `\.md` and the guard missed
    it entirely. (Classes holding an ESCAPED `\/` were already matched; the
    tests cover those separately as parity, not as regressions.)

Both defects have one root cause: regex-literal grammar — `\x` escapes, and
`/` inside `[...]` not terminating — is not expressible in a backtracking
regex. So the detector is now a tokenizer, not a regex.

readRegexLiteralAt reads the literal at a given `/` in a single left-to-right
pass with no backtracking, treating escapes as two-character units and
suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch
restarts it at every `/` on the line, preserving the old "find anywhere"
behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the
trailing-flag scan — which keeps the whole-line cost linear.

Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential
shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now
caught. Differential against the old regex over 28,474 lines (those matching
FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/
gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the
old regex hangs): 6 differences, all the tokenizer returning the fuller or
newly-correct literal, 0 old-only misses. The `\.md` token stays
case-insensitive, matching the `/i` the old regex carried.

Also closes three holes in the same new file:

  - walk() tested entry.isFile(), false for a symlink, so a symlinked
    src/*.cts was silently unscanned — an evasion of a guard whose stated
    principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist.
    It now resolves symlinks, but confined: file links must resolve inside the
    repo root, directory links inside the scanned dir itself. Every sibling
    drift guard in scripts/ uses the Dirent classification and never follows
    links, so following them unconfined would have made this the only linter
    able to read outside the tree — on fork PRs an arbitrary out-of-repo read
    whose matched fragments reach a public CI log. The narrower directory rule
    additionally stops `src/up -> ..` from sweeping the whole repo, and the
    skip list is now checked against resolved paths so `src/g -> ../.git`
    cannot reach .git/** or node_modules/**. Real paths are de-duplicated and
    files reported canonically, so a symlink alias cannot shift which
    FUNCTION_SCOPED_EXEMPTIONS key applies.
  - Both the reported fragment and the reported FILE PATH are attacker-
    controlled source text written straight to a CI log, and git permits
    control bytes in a filename. Both are now escaped — C0/C1/DEL plus the
    bidi and zero-width controls — so a crafted literal or filename cannot
    recolour the log, overwrite a line with CR, or fabricate a line that looks
    like this guard's own success output.

Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process
probe over both pathological shapes (catastrophic backtracking is synchronous
and would freeze the suite rather than fail one test), the bare-`/` class
shapes verified to fail against the parent-commit blob, root-confinement tests
covering the outside-file, outside-directory, cycle, broken-link and duplicate
cases, direct isInsideRoot coverage including the sibling-prefix case that a
bare startsWith would let through, sanitizeForReport coverage, and
limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the
exported constant. The earlier structural assertion was dropped — it checked
for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats
while staying exponential.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): backfill changeset PR number

Restores b77931869, which a force-push during the ReDoS remediation dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 01:20:55 -04:00
Tom Boucher
53ea8e0664 fix(#3057): make a guard's failure distinguishable from its benign result — Wave 1 (#3088)
* fix(#3057): refuse the write when the duplicate scan cannot complete

writeManifest documents itself as a fail-closed duplicate guard: if any
existing manifest shares plan_id with a different, non-terminal job_id it must
refuse, because dispatching again would duplicate the external job.

It could not honour that. The scan reads every sibling manifest looking for the
duplicate, and an unreadable or unparseable sibling was `continue`d past. If
the corrupt file was the one holding the live duplicate, the scan found nothing
and a duplicate external job dispatched.

The asymmetry is what gives it away: a malformed TARGET refused with
malformed_existing because clobbering is unacceptable, while a malformed
SIBLING was skipped — yet siblings are the only thing the duplicate check
reads.

Adds a scan_incomplete verdict that refuses and names the offending file, so an
operator can quarantine or repair it. Fail-closed alone would let one stale
corrupt manifest wedge every dispatch for that planning dir permanently; naming
the file is what makes refusing survivable. malformed_existing is untouched, so
the target/sibling distinction stays visible. The docstring is updated — it
previously stated a rule the function did not keep.

memFs() gains an optional failReads map so these branches are reachable at all;
they had zero coverage because the fake could not express a per-file read
fault. The signature is additive and every existing caller is unchanged.

The regression is proved by a pair, not a single test. A control writes a
readable sibling holding a genuine non-terminal duplicate and asserts
duplicate_plan_id, establishing the scenario is real; the regression then makes
that same path unreadable and asserts scan_incomplete. A first draft of this
test used a corrupt-JSON fixture containing no plan_id at all while its comment
claimed otherwise — it duplicated the unparseable-sibling case and proved
nothing, which is the defect class this phase exists to remove.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): make a guard's failure distinguishable from its benign result

Wave 1 of the negative-space backfill: the branches where a guard that could
not verify something reported the same value it reports when everything is
fine. That indistinguishability is the defect; every fix here makes the two
states tellable apart, and every test proves it with a pair — one for the
failure, one for the benign case. A single test cannot establish that two
states are distinguishable, which is the whole property being fixed.

state.cts phaseInventoryProvider returned null for both a real disk-scan
failure and a genuinely empty phases dir, so `state rebuild` could report
success while phase-table reconciliation never ran. It now returns a
discriminated result and the CLI surfaces phase_inventory_scan_failed plus a
reason. The reason field turned out never to have been wired into the emitted
JSON at all — it existed only as an internal variable — so a test could only
assert on the operator-facing note. It is a real field now.

state.cts treated an unreadable lock body the same as an empty one, applying
the 1-second stealable floor. A lock we cannot read is not a lock we know is
stale; an unreadable body is now held to the deadman ceiling like a live
holder.

verification.cts findStaleVerificationSummary returned null on any fs, scan or
clock failure — meaning "not stale". It now returns a discriminated
StaleCheckResult and the caller records that the check was indeterminate.

git-base-branch resolveBaseBranch returned 'main' both when no candidate branch
existed and when every git tier timed out. A diagnostics variant now reports
whether the answer was verified, and the CLI writes an unverified-fallback note
to stderr. The stdout contract five workflows parse is untouched.

worktree-safety snapshotWorktreeInventory left exists:true when statSync threw,
so a guard that could not check reported the worktree present; exists is now
tri-state and a stat failure surfaces as an 'unverified' finding.
planWorktreePrune reported 'no_worktrees' for a parse failure, which is not the
same as an empty list — and it drives a prune. It now reports 'parse_failed'.

Fixing the inventory change exposed a second fail-open in verify.cts: the
validate-health consumer silently dropped findings whose kind it did not
recognise, so the new kind would have vanished. That is closed too — worth
noting that the survey enumerated producers of degraded verdicts, not consumers
that discard them.

worktree-base-ref and state-transition gain the distinguishing signal without
changing what they do: headAbsenceVerified, and a phase-inventory scan meta.
Whether those guards should ACT differently is a product question this change
does not answer, and both are flagged rather than quietly settled.

rescueSummaryArtifacts is left alone: rescuing on an uncertain cat-file is
deliberate per #2556. It now has tests proving it, and a recorded negative
finding — git cat-file -e returns 128 for both "absent from HEAD" and a fatal
error, so "uncertain" and "certain-and-fine" are not separable at the git
level.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): assert typed values, not rendered text

Ten assertions in the rebuild CLI suite matched substrings of produced output —
STATE.md body fields, a markdown table row, an audit-log heading, and JSON keys
read as text. CONTRIBUTING prohibits that: if the code under test produces
text, the test asserts on its structured surface instead.

No production surface had to be built. Every one already existed and was
already compiled into bin/lib: stateExtractField for body fields,
parseMarkdownTable for the phase table, collectSection for the audit-log
section, and result.data.log — already a typed RebuildLogEntry[]. The tests
were matching rendered text sitting next to the structured data.

One of those assertions was passing for the wrong reason. `stdout.includes
('rebuilt')` matched the JSON KEY name, not a value: the dry-run path emits
`mutated` and the real path emits `rebuilt`, so it would have passed whether
the value was true or false. It now asserts the value.

external-job's refusal already had to name the offending file — that naming is
why the fail-closed variant is survivable rather than a permanent wedge — but
the tests proved it by substring of a prose message. The failure result now
carries offendingPath as its own field and the tests assert it by value. The
human message is unchanged; operators read it.

Array membership is left alone. `phaseIds.includes('99')` and
`result.updated.includes('Completed Phases')` are membership checks on real
arrays, not text matching, and converting them would weaken nothing and clarify
nothing.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): execute acquireStateLock instead of grepping its source

The non-EEXIST lock test asserted on the TEXT of the built .cjs and never
called acquireStateLock. It carried an allow-test-rule: architectural-invariant
exemption to permit that. A source grep proves a literal is present in a file,
not that the behaviour works — it is weaker than a liveness test, which at
least runs the code, and it was the only coverage the fatal-errno path had.

Replaced with tests that inject the errno through fs and assert what actually
happens: a fatal EACCES propagates out of acquireStateLock with zero backoff
sleeps, while EAGAIN/EINTR/EINVAL/EIO/ENOENT/ESTALE/EPERM/EBUSY retry once and
succeed. The exemption is removed and its allowlist entry with it.

One old assertion is deliberately not carried over: it checked the retryable
errnos were expressed as a Set rather than an inline literal. That is a shape
check with no runtime signature; the behavioural tests fail if the code reverts
to the old inline check, which is the regression it was really guarding.

The #3057 lock-body tests move into that same file rather than a new one, which
is what lint-test-file-count asks for and puts every acquireStateLock test in
one place.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): surface an indeterminate staleness check to its callers

An isolated review caught an inconsistency inside this wave. Two of the three
"add the distinguishing signal" fixes wire through to something a user sees:
git base-branch writes an unverified-fallback diagnostic to stderr, and an
unverifiable worktree surfaces as a W020 finding. The third set
staleCheckIndeterminate on readVerificationStatus's result and nothing read it.

A signal nobody consumes leaves the fail-open exactly as silent as before: the
staleness check could fail and the operator saw precisely what they would see
if the answer were genuinely "not stale". That is the defect this issue exists
to remove, so it is not defensible as scaffolding when its two siblings in the
same change already wire through.

All five callers now surface it, each through the channel it already had rather
than a mechanism imposed uniformly: phase complete adds it to its existing
warnings array and, on the blocked path, as an additive note on the error text;
init and roadmap carry it as a field on output they already emit; the UAT
report carries it without ever gating passed/blockers; workstream inventory
takes an injectable writeDiagnostic mirroring the git base-branch idiom,
because its return shape had nowhere to hang a per-phase field without
rippling the builder's types.

The routing decision is unchanged everywhere. What changes is only that a
caller and an operator can now tell a failed check from a completed one.

That diagnostic carries structured meta rather than being asserted by regex —
the default still writes only the human message to stderr, but tests assert
phaseDir and reason by value. Two earlier assertions in this branch were
converted the same way; this was the last raw-text assertion left.

Also records a scope correction: the completePhaseCore guards now compare
stateReplaceField's result to the body instead of testing truthiness, so a
field whose substitution produced identical text no longer reports as updated.
That is a real behaviour fix, not the signal-only change this file was
described as carrying, and its tests cover both the changed and unchanged
cases.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): bound two heavy subprocesses for a loaded bench, not an idle one

The remote matrix surfaced three failures unrelated to this branch's changes.
All were bad tests, and a re-run would have hidden every one of them.

The reviewer-flags parse block bounded bash -> node -> a full gsd-tools cold
start at 5 seconds. On a bench running thirty thousand tests in parallel that
is not a hang, it is a busy machine. Raised to 30s, matching the convention
sibling suites already use for script invocations, with a comment saying what
the budget covers so nobody tightens it back. Two further copies of the same
5-second spawn in the same file had the identical defect and are raised too —
they were not in the failure report, but they will be next time.

The fragment-propagation test bounded npm run regen:derived — a full build plus
eight generators, the heaviest subprocess in the suite — at five minutes, and
node22 was killed near the end. The captured output proves it: every generator
had written its files and gen:install-tree had emitted all fifteen runtimes
before the kill. Raised to fifteen minutes.

That failure read as `null !== 0`, which says nothing. status null means killed,
not a non-zero exit, and the two want different responses: one is a timeout to
size correctly, the other is a real build break. The assertion now distinguishes
them and names the signal.

Neither test's assertions were weakened and no retry was added. A retry here
would suppress exactly the signal the timeout exists to produce.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture fd 1 through the mock tracker, not a raw reassignment

The phase suite reported zero test results on both lanes while running for five
and a half minutes and exiting 1. No assertion text, no stderr, four events for
the whole file: enqueue, start, dequeue, complete. That shape is not a failing
assertion — it is the runner being unable to read the child at all, because it
parses its event stream from the child's stdout.

The cause was the capture helper reassigning fs.writeSync directly. Proven
rather than assumed: a standalone probe patched fs.writeSync and called
process.stdout.write, and the interception fired only when fd 1 resolved to a
FILE, not when it was a pipe. The remote runner captures the event stream to a
file, so a helper that was invisible against a pipe swallowed the reporter's own
output on the bench. That is also why the two sibling suites wired the same way
in this change pass cleanly — they use the mock tracker, the seam io.test.cjs
established for this exact function.

The helper now uses t.mock.method with an explicit restore after each call, so
teardown belongs to node:test rather than a second hand-rolled implementation,
and the interception cannot outlive the one synchronous call it wraps even if
that call throws. Ten call sites thread the test context through; three test
callbacks gained the parameter they lacked.

The three B3 tests are untouched — same assertions, same fault injection. Only
how the context reaches the helper changed.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture phase-complete output from a subprocess, not fd 1

Two attempts to make in-process fd-1 interception safe both failed on the
bench. The suite reported zero test results on either lane while exiting 1 —
four events for the whole file — because the runner parses its event stream
from the child's stdout, and process.stdout.write routes through fs.writeSync
whenever fd 1 resolves to a file, which is how the runner captures. Patching
that seam anywhere in a file can therefore destroy the file's own reporting,
and tightening the window only moved the runtime from 326s to 125s without
recovering a single event.

So the interception is gone rather than tuned. The helper now spawns gsd-tools
as a real subprocess and reads stdout the way the OS already gives it to us,
which is what the rest of the suite does. It asserts the command succeeded
before parsing, so a genuine failure can no longer present as a JSON parse
error.

The two fault-injecting tests could not survive that move as written: a
subprocess cannot see a mock installed in the parent. Instead of reinstating
the interception they now produce the fault on disk — the summary artifact is
created as a dangling symlink, so the staleness check's real statSync throws
inside the child. That is a more honest fixture than a mock in any case, since
it is a condition a user's tree can actually be in. Skipped on Windows, matching
the existing symlink precedent in the write-guard suite.

Three further call sites turned out to depend on parent-process writeFileSync
mocks the subprocess could not see. Those call the CJS function directly, which
is what they always wanted — they never needed stdout at all.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): one name for one signal, one encoding for one distinction

Standards review found four things this branch introduced, all of them
inconsistencies with itself rather than with the repo.

One upstream bit reached its consumers under three names —
verification_stale_check_indeterminate in two modules, the same value with
"stale" dropped in a third, and stderr only in the fourth. Standardised on the
long name wherever it is a field. The workstream inventory keeps its stderr
channel, since its return shape has nowhere to hang a per-phase field without
rippling the builder's types, but it now says the same word for the same thing.

worktree-safety encoded one three-way distinction two ways in a single file: a
named union for a finding's kind, and boolean|null for an inventory entry's
existence. The second is now a named union too.

Two assertions matched human prose because the blocked and non-blocked
completion paths carried no typed field for the signal. Both now assert typed
values. The first round of this fix added the field but left the regex beside
it, which is the banned pattern sitting next to its own replacement; the second
removed it and added an assertion on the reason enum so nothing was lost.

The remaining two were reasoned away before being fixed, and both reasons were
bad. "No typed surface exists" is the condition CONTRIBUTING says to fix by
adding one — it took three lines. "The file already does this dozens of times"
is not licence to add instance number thirty-one; a convention that violates a
documented rule is debt, not precedent.

Vocabulary differing across DIFFERENT modules is left alone: CONTEXT.md rejects
a single shared result envelope, so per-module shapes are precedented, and a
baseline smell does not outrank a documented standard.

A census of every line this branch adds to a test file now finds no regex or
substring assertion on produced prose: 87 strictEqual, 25 ok (all non-empty or
shape guards), 12 equal, 3 throws (all typed err.code predicates), 3
deepStrictEqual, 2 notStrictEqual.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3057): backfill changeset pr number to 3088

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 16:00:52 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
Tom Boucher
42f4f184c0 fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context

verify-summary's Pattern 1 matched any backticked path-like token with no
context check, so a prose mention of a future deliverable (`shared/types.ts`
in a 'next phase will add…' sentence) was checked for existence and its absence
failed the verdict on a healthy phase. #2685 added shape filtering but no
context check.

- src/verify.cts: both extraction patterns now require a claim label on the line
  (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no
  longer matches; genuine labeled claims still do.
- gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION
  so relative claim paths resolve against the project root, not the raw cwd
  (subdirectory invocation no longer manufactures missing files).

Regression tests: prose mention not treated as a claim; prose-only SUMMARY
passes; absent claimed file still fails.

* chore(#2844): backfill changeset PR 2910

---------

Co-authored-by: Test <test@example.com>
2026-07-31 03:16:15 -04:00
Tom Boucher
6229f0e55c fix(#2701): reject NUL-corrupted plan/state artifacts at the validator entry points (#2829)
* test(#2701): failing-first regression for NUL-corrupted plan/state validators

* fix(#2701): reject NUL-corrupted plan/state artifacts at the validator entry points

* fix(#2701): seed STATE.md in test (writeState); add NUL-path guards to validate/verify for parity (review)

* docs(changeset): #2701 validators reject NUL-corrupted artifacts

* docs(changeset): backfill #2701 PR number to 2829
2026-07-29 12:29:45 -04:00
Tom Boucher
d626dbc6e3 fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans (#2802)
* test(#1883): failing-first regression for findContextMdIn / listMilestoneArchiveDirs swallowing EACCES

Adds failing-first regression tests proving an unreadable dir is currently
swallowed as empty/null instead of surfacing the permission error. Covers
EACCES, EIO, the unchanged ENOENT empty path, the array fast-path, and both
CONTEXT.md forms. listMilestoneArchiveDirs is exercised in-process via a new
_listMilestoneArchiveDirs test seam (the validate command runs in a subprocess,
so an fs monkeypatch in the test process cannot reach it).

* fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans

findContextMdIn (src/planning-workspace.cts) and listMilestoneArchiveDirs
(src/verify.cts) catch-alled every readdirSync error into the empty marker
(null / []), conflating a genuine ENOENT ('nothing there') with an EACCES/EIO
failure ('can't read this'). An unreadable phase dir was silently reported as
'no CONTEXT.md' (discuss/plan gates wrongly skipped context) and an unreadable
milestones/ dir as 'no archives' (active-milestone resolution / archived-phase
filtering misbehaved).

Narrow each catch to ENOENT only — keep the long-standing null/[] contract for
genuine absence (Hyrum: empty path unchanged) and re-throw every other error so
it propagates to the caller's existing try/catch. All six findContextMdIn
callers either pass a pre-read string[] (no readdir) or sit inside a try block
that already handles readdir failures; the two listMilestoneArchiveDirs callers
live in the validate command path where errors reach the command error handler.

Exposes a _listMilestoneArchiveDirs test seam so the permission-error path can
be unit-tested in-process (the validate command runs in a subprocess, so an fs
monkeypatch in the test process cannot reach the private helper).

* fix(tests): delete stale emitted-drift ack for gsd-phase-researcher.md

Pre-existing base-branch defect, not part of #1883: commit 6932fb16d
(enhance(#1699): require read-and-cite provenance, #2768) landed the
gsd-phase-researcher.md growth onto next WITHOUT removing its now-obsolete
emitted-drift-ack.json entry. An ack explains a ripple in a PR diff; once the
change merges to next the ripple becomes part of the baseline, so the ack is
permanently stale for every subsequent PR — the emitted-attribution gate fails
on every branch off next with '1 stale acknowledgment(s) — the ripple they
explained is gone'.

The gate's own error message prescribes the fix: delete the stale entry, and
since gsd-phase-researcher.md was the only entry, delete the file (an empty ack
file signals nothing; its presence is the alarm). This unblocks the red base for
all in-flight PRs, not just this one. Folded inline per the bug-fixer directive
(one-line test-prescribed fix, does not bury the #1883 fix).

* docs(changeset): #1883 permission-error-not-empty

* test(#1883): use t.mock.method + t.after per CONTRIBUTING test conventions

Address code-review finding: replace inline try/finally in test bodies and
manual fs.readdirSync reassignment with t.mock.method (auto-restored) and
t.after for tmp-dir cleanup, matching CONTRIBUTING.md test patterns and the
existing t.mock.method idiom in tests/phase.test.cjs.

* docs(changeset): backfill #1883 PR number to 2802
2026-07-28 21:11:48 -04:00
Rezolv
7e8f6a6d7d enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (#2685)
* enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (W025)

The artifact<->git check has existed since the beginning but was only ever
pointed at .planning/research/SUMMARY.md (new-project.md:1145,
new-milestone.md:425). Phase summaries -- the ones that actually claim
"I created these files" -- were never checked.

- extract verifySummaryCore from cmdVerifySummary: same checks, lifted out of
  the output() wrapper so callers consume {passed, checks, errors} directly
  instead of shelling out and re-parsing JSON; cmdVerifySummary is now a thin
  adapter over it
- validate.health gains advisory W025 per phase SUMMARY with missing files

Advisory only: appends to warnings[], never touches status escalation beyond
the channel's own warning semantics, the repair set, or readVerificationStatus.

Resolves both open questions from triage: (a) commits_exist is deliberately NOT
surfaced -- its hash pattern matches any hex-shaped token in prose, too loose to
show a user; (b) a phase carries N per-plan summaries plus a legacy bare
SUMMARY.md, so all of them are checked via the repo-wide filter.

* chore(#2572): add changeset fragment

* feat(#2572): move the SUMMARY artifact check to phase completion

Responds to the #2685 review. Three substantive changes.

Seam (Blocker 2). The check now runs in cmdPhaseComplete, the seam the
issue body cited (src/phase.cts:~1745), not validate.health. That channel
does exist: cmdPhaseComplete declares warnings[], populates it from the
UAT/VERIFICATION pre-scan, and emits it. The cycle objection raised
against the earlier deviation holds for state.cts only -- verify.cts has
no transitive import path to phase.cts, so phase.cts -> verify.cjs adds
no cycle (verified over every src/*.cts). Moving it also retires the
retroactive firing across all historical phases: this fires once, at
completion, for the phase being completed.

Extraction (Blocker 1). Pattern 2 now excludes [ and ] from its path
class. The SUMMARY templates prescribe a YAML flow sequence
(key-files.created: [a.ts, b.ts]) and the label matches case-insensitively,
so the class previously captured the literal [ and produced a candidate
that can never exist on disk -- firing on healthy projects built from the
templates GSD itself ships. Stripping frontmatter was the other offered
remedy; measured across all three shipped templates it is a no-op on top
of the exclusion, so it is not carried. Consequence named in-code: the
key-files block still is not read, which needs a real frontmatter parse.

Also narrowed to the noise classes confirmed in review -- globs, bare
hostnames, and paths resolving outside the project are skipped rather
than reported, and the containment guard the old comment claimed now
actually exists.

Budget (Majors 1 and 3). verifySummaryCore takes a checkCommits option;
phase completion passes false, so the discarded git cat-file probes are
not spawned at all. It also passes Infinity, so every referenced file is
reported instead of the first two -- a summary listing twelve files of
which nine are missing now says nine, not zero. The verb keeps its
historical 2-file default.

Tests (Major 2). The vacuous fixtures are gone with the health block.
The replacements use /-bearing paths that genuinely extract, and each
fix was mutation-checked: un-anchoring pattern 2, dropping the glob,
hostname or containment filter, forcing commit checking on, and
re-capping at 2 each fail at least one test.

* docs(#2572): describe the phase-completion SUMMARY artifact check

The W025 text under /gsd-health is withdrawn with the health seam; the
check is documented where it now runs, under `phase complete` in
docs/CLI-TOOLS.md.

Both the docs and the changeset previously overclaimed: they said a
referenced file not on disk is warned about, while at most two candidates
per SUMMARY were ever examined. The cap is gone at this seam, so the
claim now holds -- and the text states the limits that remain, rather
than leaving them to be discovered: the key-files frontmatter block is
not read, commit hashes are not resolved, and globs, URLs, bare
hostnames and out-of-project paths are skipped rather than reported.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-07-28 18:31:33 -04:00
Tom Boucher
9a76ca6783 fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter

extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.

Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.

The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.

Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (3eb1cede2), making it
the only non-text file under src. file(1) reported it as data and text tools
silently skipped it, defeating the audit rule that says to search the authored
source; tsc passed because the bytes sat inside a comment, so no gate caught it.
It is live on next.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): pin unterminated-frontmatter detection and its negative space

Covers the discriminator on both sides. The positive rows are the issue's own
repro (LF and CRLF) plus the key-count boundary 0/1/2 around the ">= 1 parsed
key" threshold. The negative rows are the documents that reach the same branch
and must stay silent -- above all a Markdown thematic break at byte 0, which is
how this class of check has previously shipped a false positive on valid
Markdown.

Deduplication is tested on both halves of the composite key: a repeat of the
same (path, cause) is suppressed, a genuine second failure in a different file
is not, and a Windows and POSIX spelling of one path resolve to a single key.
The reset seam is asserted to actually clear -- #2674 is the precedent where a
reset that silently failed to clear made every later dedup assertion a vacuous
pass, and the cases only passed because each happened to pick an unused key, so
every case here uses a path unique to itself.

Assertions are on typed surfaces throughout -- the frozen reason enum and the
dedup-set size -- never on diagnostic prose. The one CLI-level case asserts a
differential between two runs (whether stderr is empty) rather than matching a
message, and is the wired user-reachable surface for this fix. Stream failure is
injected by overriding process.stderr.write and restoring it, never chmod 0o000,
which root bypasses.

Two properties guard the ~50 call sites of the changed function: the new
optional path argument is inert with respect to the parsed value, and LF/CRLF
spellings of a document still parse identically.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): raise the truncation threshold and repair the dedup key

Isolated adversarial review found the one-key discriminator false-positives on
ordinary Markdown: a thematic break above a single labelled line -- `Note:`,
`Author:`, `TODO:`, `See:` -- parses as exactly one key and was reported as
corruption, which is the precise failure the design claimed to prevent and the
changeset promised was fixed. The threshold is now two keys. A file truncated
after exactly one key becomes a false negative; that is the same
precision-over-recall direction already taken at zero keys, and every GSD
artefact this guards carries two or more frontmatter keys.

Three dedup-key defects, each of which could silently swallow a real diagnostic:

- Backslash normalization is removed. A backslash is a legal filename character
  on Linux and macOS, so folding it to a forward slash made two genuinely
  different files share one key. Two spellings of one Windows path may now
  report twice; two distinct files can never silence each other. Lost signal is
  the worse failure.
- The key namespaces are tagged so a file literally named like the unnamed
  digest fallback can no longer collide with a path-less caller whose content
  hashes to that digest -- computable for any predictable content, no brute
  force needed.
- The source identity is computed once rather than hashed twice per emission.

Corrects the previous commit. The two NUL bytes in src/config-loader.cts were
NOT in a JSDoc comment as that message claimed; they were deliberate separators
in the live dedup key, and stripping them degraded it to bare concatenation.
They are restored as escape sequences -- byte-identical runtime string, and the
file is text again so grep can see it. The diagnostic script that misled me
indexed a character-offset string with a byte offset.

Also threads sourcePath through the STATE.md and PLAN.md readers so the two
artefacts epic #1879 is actually about name their file rather than reporting
under a content digest.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): correct fixtures and assertions left behind by the review fixes

The previous commit changed two behaviours deliberately and the suite still
encoded the old ones, so gsd-test came back red with six failures across both
lanes -- all of them mine.

Fixtures carrying a single frontmatter key no longer clear the two-key
truncation threshold, so the CLI differential and the two path-less dedup cases
were asserting a diagnostic that is now correctly withheld. They now carry two
keys, which is what a real interrupted write of a GSD artefact looks like.

The Windows/POSIX case asserted that two spellings of one path collapse to a
single key -- the exact folding that was removed because it also collapsed
genuinely distinct POSIX files whose names contain a backslash. Inverted to
assert they now report separately, with the reasoning recorded inline so the
trade is not silently reversed later: mild duplicate noise on one Windows path
is acceptable, a swallowed diagnostic is not.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): name the file at every read site, and report each file once

The diagnostic reached only the four frontmatter CLI verbs, so ~47 of 53 call
sites reported a truncated file under an anonymous content digest instead of
naming it. Since naming the file is the whole point -- it is what an operator
can act on -- that was a gap in the deliverable, not a scoping choice. 43 of 53
sites now pass the resolved path.

Closing it surfaced a defect the original design missed. A single truncated
STATE.md is parsed twice in a normal run: once by the read wrapper, which holds
the path, and again by a pure core downstream, which is handed only the string
and cannot know it. Those two parses keyed separately, so one file produced two
diagnostics -- and wiring more sites made the collision more likely, not less.
Every emission now registers both identities the input could be known by and
checks both before writing, so whichever caller arrives first speaks and the
other is suppressed. Distinct files with distinct content still report
separately, which is the property ADR-1411 actually requires; two files whose
truncated content is byte-identical collapse to one report, which stays the
documented limit.

Ten call sites deliberately keep no path. Two are frontmatter's own round-trip
checks during set and merge, where passing a path would report on every write.
The other eight are the state-transition pure cores, which ADR-1769 defines as
(content, intent, deps) -> newContent with injected I/O; threading a path
through them would contradict that recorded decision, so it is surfaced rather
than taken unilaterally. With the widened key they no longer double-report, and
in the normal flow the named parse runs first, so the file is still named.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): inject the STATE.md path into the transition cores

The six state-transition cores parsed STATE.md frontmatter without knowing
which file it came from, so a truncated STATE.md reached the operator as an
anonymous content digest on exactly the artefact epic #1879 is named for.

ADR-1769 section 3 shapes these as (content, intent, deps) -> newContent with
injected deps, and deps is the seam for precisely this: something the core
cannot derive without doing I/O. It already carries roadmapProvider and a
phase-inventory provider on that basis, each documented as injected rather than
imported so the core stays pure and testable without disk access. A resolved
path is data, not I/O, so an optional sourcePath member extends the established
pattern rather than contradicting it, and every existing stub keeps compiling
because the member is optional.

updateCore and reconcileCurrentPosition take no deps and are left alone. With
the widened dedup key they cannot double-report, and in the normal flow the read
wrapper has already named the file by the time they run.

Also regenerates gsd-core/bin/lib/state-transition.cjs. That artifact is tracked
rather than gitignored, unlike most of its siblings, so leaving it stale would
have shipped a runtime without this change to anyone reading the repo without
building. tsc had skipped the re-emit because its incremental build info still
recorded an emit that had since been reverted, so the stale output survived a
clean build; clearing tsconfig.build.tsbuildinfo forced it. The
compiled-artifact-sync gate is what surfaced the drift and now reports all nine
tracked artifacts matching their source.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): stop the widened dedup key from hiding a second file

The previous commit widened the dedup guard so one file parsed twice -- once by
a read wrapper holding the path, once by a pure core holding only the string --
reported once instead of twice. It did that by checking BOTH keys before
emitting, which silently traded one defect for a worse one: two DIFFERENT files
whose truncated content happened to be byte-identical now collided on the shared
content digest, and the second file's diagnostic was swallowed. That is the
over-coarse keying ADR-1411 explicitly forbids, reintroduced while fixing
something else.

The guard now checks only the key matching what the caller actually knows -- a
named read checks its path key, a path-less read checks its digest key -- while
still recording every key the input could later be identified by. The redundant
path-less re-parse of an already-named file stays silent, and two distinct files
always both report.

Verified across all six orderings: same file named-then-anonymous reports once;
two different files with identical content report twice; two different files
with different content report twice; the same path twice reports once; two
path-less parses of identical content report once; two path-less parses of
different content report twice.

The suite caught this -- twenty failures, all in the unusable-input tests that
reuse one truncated fixture across different paths. The local probe written
alongside the broken change did not, because it compared two files with
different content and could therefore only confirm the expected behaviour.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): count diagnostics emitted, not identities interned

The suite measured the size of the dedup set as a stand-in for "how many
diagnostics were emitted". That held only while one emission recorded exactly
one key. Once an emission began recording every identity the input could later
be matched by -- a path key and a content key for the same file -- the set grew
by two per write and twenty assertions read 2 where they expected 1.

The production behaviour was correct throughout; the proxy was not. Set size
counts identities, which is an implementation detail of the guard. The
behavioural claim these tests exist to make is how many diagnostics an operator
actually saw, so the module now exposes that directly as an emission counter and
the suite asserts on it. The set-size accessor stays for assertions genuinely
about key shape.

The local probe written alongside the change did not catch this because it
counted process.stderr.write calls -- the right thing -- while the suite counted
set growth. Verification now asserts both and requires them to agree, so a
future divergence between the counter and real writes fails immediately rather
than being discovered a bench run later.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): retire two assertions that outlived the behaviour they described

Both tests encoded assumptions the dedup fix invalidated, and both were caught
by the suite rather than by the probe written alongside the change.

The forged-path case asserted that a file named like the anonymous digest
fallback must not suppress a later path-less report. That premise is gone: an
emission now records every identity its input could be matched by, so ANY named
report of some content silences the anonymous re-parse of that same content --
which is the same-file guard working as intended, and has nothing to do with the
crafted name. The property still worth defending is that a crafted filename can
never silence a real file reported under its own path, so that is what the test
now asserts, with the deliberate suppression documented beside it.

The reset-seam case ended by reading the size of the dedup set and expecting 1.
Set size counts interned identities, not diagnostics written, and one emission
now interns two. It asserts the emission counter for the event and keeps a
weaker set-size check for the interning.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): close the review findings on the discriminator, dry-run and counter

Three orthogonal review passes ran against the final diff. Their findings:

A labelled preamble under a leading rule was still misreported. Raising the key
threshold to two only moved the boundary, because two colon-labelled lines are
as common in ordinary prose as one -- a document opening with a rule over an
Author and a Reviewed-by line, then prose, was called corrupt. Key count alone
cannot separate the two. What does is what follows: a write interrupted part way
through a frontmatter block ends mid-block, so every line of the region is still
frontmatter-shaped, whereas a document merely opening with a rule goes on to
prose. Both conditions are now required, and each closes a false-positive class
the other leaves open. Nested list values and indented continuations stay
frontmatter-shaped, so legitimate truncations are unaffected.

`state rebuild --dry-run` reported a truncated STATE.md anonymously. The write
path is named only because readModifyWriteStateMd parses with the path first;
the dry-run branch reads the file directly and never did. Dry-run is the
read-only mode an operator reaches for first when they suspect corruption, so it
is the one that most needed to name the file. reconcileCurrentPosition takes the
path as an optional argument now and rebuildCore passes it down. That function
was previously left alone on the grounds that a read wrapper always names the
file first -- this is the flow that disproves it.

The emission counter counted write attempts rather than writes, so on a broken
stderr it claimed a diagnostic had reached the operator when nothing had. It is
incremented only after a write that completed, and the broken-stderr test now
asserts the count as well as the return value.

Two documentation defects. The module described a guarantee it does not keep:
one file yields one diagnostic only when the named read comes first. The reverse
ordering emits twice, and that is deliberate -- a path-less caller cannot
identify its file, so suppressing the later named report would also suppress a
genuine second failure in a different file whenever two files share identical
truncated bytes, which ADR-1411 ranks the worse failure. The comment now states
the asymmetric guarantee and a test pins it. Separately, the CONTEXT.md glossary
entry still described backslash normalization that a later commit removed, and
asserted the opposite of what the tests pin; no lint checks prose against code,
so nothing caught it.

Also converts three body-level try/finally blocks to t.after(), per
CONTRIBUTING.md's rule that try/finally belongs only in helpers with no test
context -- the file's own emissionsDuring helper already did this correctly.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#1882): tell the operator what the truncated-frontmatter warning means

A user who has just seen the new warning is acting, not studying, so this lands
in the How-To quadrant beside the other "if you see X" branches in
debug-a-failed-execution, not in reference or explanation. It gives them what
the warning means for this run, three steps to restore the file, and the fact
that the warning changes no return value or exit code.

It also states the case that matters more than the warning itself: silence does
not prove the file is intact. GSD says nothing when the partial block carries
fewer than two fields or reads as prose, because a Markdown document opening
with a horizontal rule is indistinguishable from one of those. A reader chasing
missing metadata needs to know not to treat quiet as clean. Why that threshold
exists is explanation and deliberately stays out of a how-to.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1882): backfill changeset pr number to 2712

* test(#1882): constrain each branch of the frontmatter-shape check

CI's mutation gate came in at 61.56 against a threshold of 62, and the surviving
mutants were concentrated in isFrontmatterShaped -- the function added last, in
response to review, and the only one never given tests of its own. It was
exercised solely through extractFrontmatter, which covers the composite decision
but leaves each branch of the predicate unconstrained: drop the blank-line
filter, or any one of the three shape alternatives, and every existing assertion
still passed.

Four cases now pin the halves independently. A blank line inside an interrupted
block must not disqualify it, which constrains the filter and its comparison. An
unindented list item and an indented folded-scalar continuation each exercise one
shape alternative that no other case reaches on its own -- the folded line is
neither a key nor a list item, so it is the only input that distinguishes the
indented branch. And two keys followed by prose must stay silent, which is the
negative half: it fails if the predicate is ever mutated to accept everything,
and it is the case that proves key count alone was never sufficient.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): register the unusable-input suite with the frontmatter mutation shard

The mutation gate reported an identical 61.56 across two runs whose only
difference was four added tests. That is the tell: the tests were never
executed. The frontmatter shard runs a fixed file list in stryker.config.mjs and
scripts/mutation-matrix.cjs, and tests/unusable-input.test.cjs was in neither, so
the entire suite covering the new unterminated-fence branch was invisible to the
gate while passing perfectly well in the normal run.

So the score was not measuring weak tests, it was measuring absent ones: #1882
added mutants to frontmatter.cjs and no test in the shard covered them. Both
lists gain the file; the config already notes they must stay in sync.

This is a registration ripple a new test file carries when it covers a
mutation-tracked module, alongside the .gitignore, eslint, inventory, glossary
and size-baseline ripples a new module carries. Nothing warned about it, which
is why two runs were spent before the identical score gave it away.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:50:12 -04:00
Tom Boucher
c5e0371775 feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging

Red phase for issue #1951 (reversibility tagging: classify decisions by
undo cost, gate one-way doors behind a checkpoint:decision).

Tests assert, per the issue's acceptance criteria:
- discuss-phase CONTEXT.md template records a **Reversibility:** field with
  a rationale on captured decisions, and states it is optional
- gsd-planner @-references planner-reversibility.md and stays under the
  49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs)
- a one-way rating inserts a checkpoint:decision before the dependent task;
  reversible inserts none; costly is flagged but never blocks
- the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard)
  and inserting a checkpoint implies autonomous: false
- docs/reference/plan-md.md documents <reversibility> as optional with all
  three ratings
- --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected
  into the planner prompt, and is advertised in the command argument-hint
  and help full mode (argument-hint parity)
- the override suppresses the gate but still persists the rating
- cmdVerifyPlanStructure accepts every rating and the absent case
  (additive-validator guarantee, behavioral via runGsdTools)
- parity: thinking-models-planning.md #4 adopts the canonical three-level
  taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone
- no content loss from the planner extraction made to fit under the cap

Prose-contract assertions are Red until the implementation lands. The
behavioral validator assertions pass immediately — regression guards
proving the validator already accepts unknown optional tags.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1951): reversibility tagging — gate one-way-door decisions

Classify planning decisions by what undoing them would cost, and give a
one-way door a human beat before the agent walks through it (issue #1951,
The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way
door framing).

Acceptance criteria met:
- discuss-phase records an optional reversibility rating with a rationale
  on <decisions> entries in the phase CONTEXT.md template. Unrated
  decisions are treated as reversible, so existing phases are unaffected.
- a one-way rating makes gsd-planner insert a checkpoint:decision before
  the task that implements the decision, reusing the existing checkpoint
  mechanism -- no new checkpoint machinery.
- reversible ratings trigger no checkpoint; costly ratings are flagged in
  the plan but never block.
- the rating persists on the task as the optional <reversibility rating=>
  element. cmdVerifyPlanStructure accepts every rating and the absent
  case; the structural validator does not reject unknown optional tags.
- --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses
  checkpoint insertion for intentionally-unattended runs while still
  recording ratings -- the override changes what stops the run, not what
  the plan remembers.

Single taxonomy, not two: references/thinking-models-planning.md #4
already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is
loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the
canonical three-level vocabulary and now points at planner-reversibility.md
as the taxonomy owner, with a parity test that fails if the surfaces
diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE).

agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the
checkpoint DO/DON'T guidance was relocated verbatim into
planner-antipatterns.md -- already @-referenced from the same section for
the same topic, so the planner still loads it and nothing was dropped. A
test guards the relocation against content loss.

Files: gsd-core/references/planner-reversibility.md (NEW, canonical
taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase
workflow/command/help (flag wiring + parity), plan-md.md schema,
discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest,
size baselines, install goldens, plugin skills regen, changeset.

Closes #1951

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1951): address orthogonal review findings

Two isolated reviewers (correctness + security), neither of which authored
the change. Every finding fixed:

Security — the rationale is untrusted input (ADR-1577). It originates in
conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each
hop an LLM reading the previous hop's output, with no validation on the
path. planner-reversibility.md and the discuss-phase template now state
it is data and never instructions, and name the </reversibility>
early-termination hazard explicitly -- a rationale that closes its own
element injects sibling structure the executor reads as real tasks.
Four tests guard it.

Correctness 1 — nothing machine-enforced the feature's own promise: a task
rated one-way with no preceding checkpoint:decision validated as fully
clean, so a planner error silently reopened the gap this feature exists to
close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A
warning, not an error: <reversibility> stays additive and the plan stays
valid. Four tests cover ungated (warns), gated (silent), still-valid, and
reversible/costly never flagged.

Correctness 2 — pass-always test. The --no-reversibility-gates parse test
substring-matched the whole workflow file, and plan-phase.md prose mentions
both tokens in one sentence, so it passed with the bash conditional
deleted: it was testing the documentation, not the parser. Now scoped to
the fenced bash blocks and matched as one physical line, with a negative
control confirming prose alone cannot satisfy it.

Correctness 3 — costly had no itemized emission rule, only one-way did, so
two agents could diverge on whether to tag costly at all.

Correctness 4 — template convention break: the example ratings were bare
while every sibling field uses [...] to signal substitution, inviting an
LLM to copy one-way/costly forward as boilerplate. Now bracketed.

Correctness 5 — latent false-green: .includes('reversible') also matches
inside irreversible/irreversibility, which appear in anti-pattern
prose, so a surface that dropped the real taxonomy entry would still pass.
Now word-boundary matched.

ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md
1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on
next). The ceiling may only rise for privileged host machinery, and
reversibility gating is optional-feature logic, so the wiring was slimmed
to its minimum and the explanatory prose moved to the reference files the
planner already loads. plan-phase.md is now 94400 bytes -- 119 under the
ceiling and 70 bytes SMALLER than on next, so the host loop shrank while
gaining the feature, which is what phase 6 ratchets toward. The tracer
contract (tests/tracer-bullet.test.cjs) is unchanged.

Lint — fixed an unnecessary non-null assertion in verify.cts and a
CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST-
PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture must carry the common task elements

The gated-one-way fixture built a checkpoint:decision task from the
abbreviated skeleton in gsd-planner.md, which shows only the
checkpoint-specific elements (<decision>/<context>/<resume-signal>).
cmdVerifyPlanStructure requires <name> and <action> on EVERY task
regardless of type, so the fixture failed validation for reasons that had
nothing to do with reversibility:

  errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"]

Caught by gsd-test on 14d14a39 (2 failures, both this fixture).

The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task
carries <name>/<files>/<action>/<verify> like any other. Fixture corrected
to match. Verified behaviorally against the real gsd-tools CLI across all
four cases: gated one-way (valid, silent), ungated one-way (valid, warns),
costly (valid, silent), absent (valid, silent).

Not a product defect: the validator's every-task contract is intentional
and pre-existing, and docs/reference/plan-md.md scopes its required-element
list to type=auto/tracer only because those are the elements a planner must
author, not because checkpoints are exempt from <name>.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1951): backfill changeset pr number to 2471

* fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision

Both CI failures were real defects in code this PR added, not false
positives.

CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 —
the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`,
which escapes the hyphen but not backslash, so the escape was incomplete.
It was also unnecessary: `-` carries no special meaning outside a character
class. Replaced with a complete metacharacter escape (backslash included).
Word-boundary behavior verified unchanged across all three ratings — notably
that "irreversible" prose still does not satisfy a "reversible" match, which
is the false-green this helper exists to prevent.

Prompt injection scan — the checkpoint fixture used the human-verification
child element inside <verify>. That tag name is a fake-instruction-boundary
pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed
files, so copying the shape from tests/verify.test.cjs (unflagged only
because it is not in this diff) tripped the gate. Switched to the documented
plain-prose <verify> form.

The first attempt at that fix failed the same gate a second time: the
comment explaining the collision quoted the offending tag literally. The
comment now names it in prose instead — the scanner does not care whether a
match is code or commentary, which is the whole point of the
DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md.

Verified locally before push: scan reports 0 findings across 57 changed
files, eslint clean, and both fixtures still validate as designed (gated
one-way silent, ungated one-way warns, neither errors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): record measured cost and halve gsd-tools spawns

The Windows shard 1/3 job timeout was traced to the sharding layer, not to
this PR's assertions — see #2472. Two contributing factors were this file's
own, and are fixed here.

1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs,
   so scripts/run-tests.cjs weighted it at the table's median fallback
   (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x
   under-weight. Recorded the measured value from the green gsd-test run
   (max across the node22/node24 lanes, per gen-test-timings.cjs's
   convention). Only this one entry: a full regen churns 634 entries of
   run-to-run drift, and the table is explicitly advisory and un-gated, so
   a 637-line diff does not belong in a feature PR.

2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost.
   Spawns cut from 9 to 6 with no coverage lost:
   - the ungated-one-way warning and its stays-valid assertion now share
     one plan instead of building the same plan twice;
   - the reversible/costly never-flagged-as-ungated test was strictly
     subsumed by the additive suite, which already runs those two ratings
     ungated and asserts no /reversibilit/ warning at all — and the gate
     warning's text contains both "reversibility" and "one-way", so the
     broader assertion catches it. It only re-spawned gsd-tools twice to
     prove the same thing.

Both are symptom fixes. The shard imbalance itself (19/11/10 minutes
against a 20-minute cap, from a cost-blind round-robin partition that also
reshuffles downstream files whenever one is inserted) is tracked in #2472.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture adopts the #2444 type-branched contract

Surfaced by rebasing onto next, which gained #2444 (branch plan-structure
validation on task type=checkpoint:*) while this PR was in review.

cmdVerifyPlanStructure no longer applies one required-element set to every
task. A checkpoint:decision now requires <name> + <resume-signal> +
<decision> + <options>, and is exempt from the <action>/<verify>/<done>/
<files> set that auto and tracer tasks carry. The gated-one-way fixture
predated that split and failed on the new requirement:

  errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"]

Fixture rewritten to mirror the checkpoint:decision contract exactly — real
<options> with two <option> children — rather than padding it with fields
checkpoints no longer need. That also drops the plain-prose <verify> the
earlier revision carried purely to dodge the prompt-injection scan; a
checkpoint task has no <verify> requirement at all, so the workaround is
moot.

Verified against the real gsd-tools CLI across all four cases: gated one-way
(valid, silent), ungated one-way (valid, warns), costly (valid, silent),
absent (valid, silent).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 10:44:55 -04:00
Tom Boucher
46ba9ed464 fix(#2444): branch plan-structure validation on task type=checkpoint:* (#2473)
* test(#2444): failing-first regression for checkpoint:* plan-structure validation

Add acceptance-criteria tests covering the three canonical checkpoint task
types (human-verify, decision, human-action) plus an unknown-subtype
forward-compat case. Each canonical type must pass verify plan-structure
when it carries its type-specific required fields (per
gsd-core/references/checkpoints.md), and must be flagged when those
fields are missing. Non-checkpoint tasks keep the existing
<action>/<verify>/<done>/<files> requirements unchanged (AC3 regression
guards).

The existing 'errors when checkpoint task but autonomous is true' fixture
is updated to use the canonical checkpoint:human-verify triple
(<what-built>/<how-to-verify>/<resume-signal>) so it does not collide
with the new per-type validator; the assertion (autonomous is not false)
is unchanged.

* fix(#2444): branch plan-structure validation on task type=checkpoint:*

cmdVerifyPlanStructure unconditionally required <action>/<verify>/<done>/
<files> on every task, so every checkpoint:* task — which uses the
checkpoint convention's type-specific fields instead — was reported as a
structural error. Checkpoint-heavy phases produced walls of false findings.

The fix introduces two pure helpers in verify.cts:

  - extractPlanTaskInfos(content): single ReDoS-safe pass over
    <task ...>...</task> blocks that captures BOTH the opening-tag
    attribute string (so the type= selector is not lost, as it is with
    extractTaggedBlocks) and the body, returning a typed PlanTaskInfo.

  - validatePlanTaskStructure(task): branches on the task's type.
    checkpoint:human-verify requires <what-built>/<how-to-verify>/
    <resume-signal> (the canonical triple).
    checkpoint:decision requires <decision>/<options>/<resume-signal>.
    checkpoint:human-action requires <action>/<instructions>/
    <verification>/<resume-signal>.
    Unknown checkpoint:* subtypes require only the universal
    <resume-signal> (forward-compat). All other types keep the historical
    <action>/<verify>/<done>/<files> requirements unchanged.

Canonical reference: gsd-core/references/checkpoints.md. Per-type field
sets validated against the documented templates in
agents/gsd-planner.md and gsd-core/templates/phase-prompt.md.

* fix(#2444): re-resolve body-parser to 2.3.0 in lockfile (GHSA-v422-hmwv-36x6)

GHSA-v422-hmwv-36x6 (body-parser DoS via invalid limit value, low severity,
published 2026-07-20T23:23:26Z) made tests/npm-integrity-gate.test.cjs
(#3588: root workspace production tree has no advisories) fail any subsequent
npm audit --omit=dev. The advisory affects body-parser >=2.0.0 <2.3.0 pulled
transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk
-> express -> body-parser@2.2.2.

express@5.2.1 already declares body-parser as ^2.2.1, so 2.3.0 is a valid
re-resolution within express's own compatibility range — no override needed.
Regenerated the lockfile via 'npm audit fix --omit=dev' which re-resolves
transitive deps within their declared ranges; package.json is unchanged.

Verified: npm audit --omit=dev reports 0/0/0/0/0 advisories; body-parser
now reads as 2.3.0 in 'npm ls body-parser --omit=dev'.

* test(#2444): close review gap-closure tests + harden type-attr charset

Orthogonal review (code-review + security-review subagents) returned APPROVE
on Standards and Spec. Per the playbook's zero-tolerance policy, address
every Low finding:

Spec gap-closures:
- AC3 verbatim: add explicit <done> and <files> regression tests for
  non-checkpoint tasks (pre-existing tests only covered <action> and
  <verify>).
- AC2: add checkpoint:decision missing <decision>, checkpoint:human-action
  missing <action>, checkpoint:human-action missing <verification> cases
  (the implementation enforces all of these; only one missing-field case
  per type was previously tested).
- Remove the duplicate 'returns error for nonexistent file' test that
  leaked into the new describe block from the insertion edit.

Security hardening (Low-sev, defense-in-depth):
- Tighten the task type= attribute extractor in src/verify.cts from
  [^"'>\s]+ to [\w:-]+ so a hostile type= attribute cannot carry
  markup fragments (e.g. type=evil<fragment) into the verifier's typed
  JSON output. All legitimate type values (auto, tracer, manual,
  checkpoint:human-verify, checkpoint:decision, checkpoint:human-action,
  checkpoint:tdd-review) match the tighter charset.
- Add adversarial regression test asserting type=evil<fragment surfaces
  as 'evil' (capture stops at '<'), with no markup chars (< > ( ) &)
  in the surfaced type field.

* docs(changeset): add Fixed fragments for #2444 PR

Two fragments:
- sturdy-jays-tumble.md: the verify plan-structure checkpoint fix
- witty-badgers-hum.md: the body-parser 2.3.0 re-resolution

PR number backfilled to 0 placeholder per CLAUDE.md 'PR Number Handling';
will backfill to the real PR number immediately after gh pr create returns.

* docs(changeset): backfill PR number to 2473

Per CLAUDE.md 'PR Number Handling': backfill the placeholder pr:0 with the
real PR number returned by gh pr create.
2026-07-21 08:12:52 -04:00
Tom Boucher
be5113abfe fix(#2408): fold colliding phase statuses + add W023 collision warning (#2461)
* fix(#2408): fold colliding phase statuses + add W023 collision warning

Two coupled bugs from #2408:

1. cmdStats last-write-wins status (src/commands.cts:1610-1618): when two
   on-disk phase directories normalize to the same phase key (e.g.
   `05-real/` and `05-real-stray/`), the directory-scan merge overwrote
   `status` with whatever the *current* directory in scan order computed,
   discarding `existing?.status` entirely. fs.readdirSync order is not
   stable across platforms, so /gsd-stats could silently report `Not
   Started` for a phase that is actually `Complete`. Plan/summary counts
   were already additively merged; only `status` was wrong.

   Fix: introduced a `foldPhaseStatus(a, b)` helper that returns whichever
   status is further along the precedence ladder
   `Complete > Needs Review > Executed > In Progress > Planned > Not Started`
   (with `Pending` and unrecognized statuses ranked after). The merge site
   now calls `existing ? foldPhaseStatus(existing.status, status) : status`.
   The fold is commutative, so the result is identical regardless of read
   order.

2. cmdValidateHealth had no collision-detection pass (src/verify.cts):
   codes W001-W022 cover every condition except normalized-key collisions,
   so an operator got zero signal that anything was wrong.

   Fix: added W023 — groups phaseDirEntries by their normalized phase key
   (via the existing normalizePhaseName + extractPhaseToken helpers from
   phase-id.cjs — same normalization cmdStats uses) and emits a warning
   for any group with ≥2 dirs. The warning names the normalized key, both
   directory names (sorted by comparePhaseNum for stable output), and each
   directory's independently-computed status (via determinePhaseStatus
   imported from commands.cjs). Wording is deliberately neutral — never
   guesses which directory is the real one. The optional --repair path
   from the issue is intentionally NOT implemented in this PR (the issue
   marked it lower priority and acceptance criterion 4 is vacuously
   satisfied by omission).

Triage correction applied: the issue proposed W022, but that code is
already in use for config.json model-tier validation (src/verify.cts
:1372-1391). The next free code is W023, used here.

Tests:
- tests/commands.test.cjs: integration test that 05-real/ (Complete) +
  05-real-stray/ (empty/Not Started) collide and stats reports the merged
  phase as Complete regardless of read order; plus a direct unit test of
  foldPhaseStatus asserting commutativity + correct precedence for every
  status pair + correct handling of unrecognized statuses.
- tests/health-validation.test.cjs: integration test that W023 fires on
  the collision naming both dirs + their statuses (and uses neutral
  wording), plus a negative test that no W023 fires when only one dir
  exists per key.

References: #2408; reporter's three-layer triage + acceptance criteria;
triage correction that W022 is already in use (model-tier validation).

* chore(#2408): backfill pr:2461 in .changeset/graceful-koalas-forage.md
2026-07-20 15:49:57 -04:00
Tom Boucher
b2961c3f69 fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation

Encodes the three acceptance criteria from #2070 plus the boundary cases the
resolver silently ignores today (non-string values, empty string, mistyped
phase-type key), and pins VALID_TIERS to a catalog-derived set.

Red phase: these fail against current src/ by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022)

W004 sourced its profile list from a hand-maintained literal that predated the
adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now
reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json.

models.<phase_type> was validated nowhere: the resolver's tier gate silently
drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable
no-op. A new W022 flags unknown phase-type keys and invalid tier values
(including non-string values, which the same gate also drops).

VALID_TIERS moves from a function-local literal in model-resolver.cts to a
catalog-derived export, so health and the resolver cannot disagree by
construction rather than by parity test. Object.values(adaptiveTierMap) is
['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal,
so resolution behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): changeset for validate health adaptive profile + W022

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate

Review of the initial fix surfaced three real defects, folded in per the
no-defer rule:

1. verify.cts: the W022 guard skipped a top-level `models` that is present but
   not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those
   identically, so they were the same undiagnosable no-op #2070 targets — just
   one level up. They now warn; absent/null/{} stay silent.

2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the
   tier vocabulary this change had just de-hardcoded elsewhere. It now derives
   from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides
   resolve to a concrete tier). Byte-equivalent to the old literal.

3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457
   the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib,
   so the `gsd-core/` prefix is dead coverage for library code and a src/-only
   PR could merge with no release note — including this one. Adding `src/`
   closes the gate; tests/ stays non-user-facing.

Also corrects a false docstring in the VALID_TIERS test: value-equality cannot
detect a re-hardcoded literal, so the test no longer claims it does.

Two review findings were rejected with evidence rather than actioned:
- W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ...
  kept, message-disambiguated"), not a defect.
- Global-defaults validation would be a false-positive generator: config-loader
  reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health
  early-returns E001 without .planning/, so those values provably never affect
  resolution in any context health can run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2070): regenerate install goldens for the changeset-lint change

scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content
hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to
USER_FACING_PREFIXES changed that hash and tripped every golden parity check.
Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): backfill PR number 2336 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:04 -04:00
Tom Boucher
e0f969af6a refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:

- toPosixPath(p)   — this machine's native path → POSIX (running-OS relative;
                     for local filesystem paths).
- toNativePath(p)  — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
                     to a POSIX/bash TARGET (which may differ from the running
                     OS) and for parsing mixed-separator input.

core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.

- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
  runtime-artifact-install-plan, drift, init, worktree-safety,
  installer-migrations, installer-migration-authoring, install-engine, surface,
  verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
  corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.

Closes #2246

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 15:36:17 -04:00
Tom Boucher
58b20c31f8 fix(#2136): migrate ALL operator-facing date sites to localToday (anti-pattern elimination)
The UTC-slice anti-pattern (deriving a calendar day a human reads by slicing a
UTC instant) remained in several operator-facing sites beyond the original
seamed last_activity set. Eliminate it everywhere a human reads the value as a
calendar day — do not leave known bad code in place:

- commands.cts cmdTodoComplete + cmdScaffold: completion/scaffold dates.
- gsd2-import.cts: migrated STATE.md 'Last activity' / 'Last session'.
- template.cts: plan frontmatter 'completed:' date.
- verify.cts: health --repair session-log date + '(Backfilled: <date>)' header.
- workstream.cts: workstream-create 'Last Activity' / 'created' + archive dirname.
- init.cts: JSON-bundle 'date' (-> localToday) + 'timestamp' (-> nowIso) at all
  three sites; drop the now-dead 'const now = new Date()' in cmdInitTodos /
  cmdInitMapCodebase (cmdInitQuick keeps it for the local branch-id derivation).
- state.cts: prune-archive '## Pruned <date>' header.
- state-transition.cts: the 7 seamed last_activity writes (prior commit).

No realClock.today() / .clock.today() / raw new Date().toISOString().split('T')
operator-facing sites remain in src/. Rebuilds the tracked
bin/lib/state-transition.cjs artifact to match.
2026-07-12 15:34:59 -04:00
Tom Boucher
8e4ebb49e4 fix(#2128): address shared-seam review regressions (#557, action-scan, comment-strip)
Final convergence review found the shared <tag> seam introduced 3 behavior
regressions; fixed all + locked with tests:

- #557 REGRESSION: stripTaggedBlocks's attribute-tolerance stripped `<details open>`
  (the ACTIVE-milestone marker) that the old `<details>`-only regex preserved. The
  seam now takes `allowAttributes` (default false) — details/decisions strip is
  attr-INTOLERANT (preserves `<details open>`); only `<task type="…">` opts in.
  Regression test added to roadmap-parser + markdown-sectionizer suites.
- verify.cts actionZones (negative-grep-echo security scan): reverted to a bounded
  to-first-close scan `<action>([\s\S]{0,20000}?)</action>` so a grep-echo trick
  can't hide behind an unterminated inner <action> (the seam's stop-at-next-open
  would drop it). ReDoS-safe via the cap.
- check-command-router HTML-comment strip: `(?:-->|$)` fallback wiped to EOF
  (fail-closed spurious gate block) — replaced with stop-at-next-open so an
  unclosed `<!--` leaves downstream tags intact.
- Updated the extractTaggedBlocks nested-tag tests to the new (stop-at-next-open)
  behavior: `<x><x>inner</x></x>` -> ['inner'].

All vectors still linear; every fix verified in-process.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 11:12:45 -04:00
Tom Boucher
01b691fae8 fix(#2128): route <tag> scans through a hardened shared seam (root-cause ReDoS fix)
A convergence audit showed the `<tag>[\s\S]*?</tag>` lazy-scan ReDoS was pervasive
(a dozen+ bespoke copies across roadmap-parser/check-command-router/verify), each
a distinct quadratic vector on a large document with unclosed tags. Rather than
whack-a-mole, single-source them (maintainer-directed):

- markdown-sectionizer: extractTaggedBlocks now shares one ReDoS-safe
  `taggedBlockPattern` (stop-at-next-open, bounded optional attributes) and gains
  a `stripTaggedBlocks` companion for block removal.
- roadmap-parser: 3 `<details>` strips -> stripTaggedBlocks (behavior-identical —
  no <details> here carries attributes).
- verify: actionZones + both <task> loops + their nested <name>/<files>/gate/req
  extractions -> extractTaggedBlocks (behavior byte-equivalent, verified).
- check-command-router: the objective|tasks?|action alternation hardened in place
  (distinct multi-tag shape); HTML-comment strip gains a `$` fallback.

Every vector now linear (<3ms on 1.5MB adversarial); real content unchanged
(end-to-end verify/roadmap resolution + task extraction confirmed). The only
remaining `<!--…-->` scan (uat.cts:201) is anchored + non-global — one scan, safe.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 10:42:18 -04:00
Tom Boucher
2f6662d195 fix(#2128): bound the remaining lazy-scan ReDoS vectors (files_modified, Plans-count, <tag>)
A ReDoS-completeness audit surfaced a distinct class beyond the tag/bracket
clause: unbounded `[\s\S]*?` / `[^\]]*` lazy-scans searching for a literal
terminator that may never appear, driven quadratic by REPEATED structures in a
large PLAN.md/ROADMAP.md. Folded all 7 in at maintainer direction:

- files_modified `[^\]]*` -> `[^\]]{0,8000}` (commands.cts, verify.cts): 39.7s -> 0.9s.
- Plans-count `[\s\S]*?` -> section-local `(?:(?!\n#{1,4}\s)[\s\S])*?` — stops at the
  next heading (semantically correct: Plans: belongs to the phase's own section)
  (roadmap.cts x3, phase.cts): 36s -> 4ms.
- <tag> extraction `([\s\S]*?)` -> stop at the next same-tag opening
  `((?:(?!<tag>)[\s\S])*?)` (verify.cts x3, markdown-sectionizer.cts): ~6s -> 2ms.

Every vector is now linear (comprehensively re-measured); real content matches
(end-to-end `roadmap get-phase` still resolves Plans-counted phases). Pre-existing;
byte-behavior preserved for realistic inputs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 10:11:43 -04:00
Tom Boucher
b321bc04f4 fix(#2128): bound the sibling bracket-prefix clause — complete the ReDoS fix
Review caught that the prior commit bounded only the paren tag clause and left
the SIBLING bracket-prefix `(?:\[[^\]]+\]\s*)?` (same host regexes, before Phase)
UNBOUNDED — the identical quadratic reachable via a `[...]` run (measured ~16s at
1.7MB). Bound `[^\]]+`/`[^\]]*` -> {1,200}/{0,200} across all 19 phase/milestone
heading prefixes. Comprehensive re-measurement now shows EVERY vector linear
(bracket/paren/id/name/milestone all ~2-44ms at 2.45MB; bracket scaling
2k->2ms, 4k->5ms, 8k->10ms). Also: update the #1729 literal-mirror parity test
off its stale unbounded constant, and add limit-1 (199) boundary coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 09:44:49 -04:00
Tom Boucher
c1cd43a39f fix(#2128): bound the phase-tag clause to {0,200} — kill quadratic ReDoS
The canonical OPTIONAL_PHASE_TAG_SOURCE tag clause `(?:\s*\([^)\n]*\))?` (and its
inlined literal mirrors across 11 modules) had an UNBOUNDED body, making the
optional-group + /g header scan quadratic on adversarial ROADMAP.md/STATE.md — a
long run of `(` after a header ran ~18.8s at 1.7MB. Bound the body to {0,200} in
the constant AND every mirror in lockstep (the #1729 "both forms change together"
contract), so the scan is linear: the same 1.7MB input now resolves in ~9ms
(measured), while real tags (a handful of chars) still match and a 201-char tag
is rejected. Added a #2128 boundary regression to the #1729 suite.

Pre-existing (byte-identical before/after the Phase 4 migrations); folded in at
maintainer direction rather than deferred.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 09:31:17 -04:00
Tom Boucher
dfad3a7510 refactor(#2128): single-source 23 phase-token re-derivations; sanction 14 context-specific sites
Route 23 literal re-derivations of the canonical phase-number token through
phase-id.cjs `PHASE_NUMBER_TOKEN_SOURCE` (via new RegExp). Each conversion was
proven BYTE-IDENTICAL (old.source === new.source && old.flags === new.flags), so
the runtime regexes are unchanged — zero behavior change by construction.

The remaining 14 phase-token sites are genuine but context-specific and stay
literal with a `// phase-id-owner: <reason>` sanction: dir-name parses whose
dash-continuation semantics differ from extractPhaseToken, and the [A-Za-z]
case-variant / [.-] dot-or-dash separator forms that are not source-byte-equal
to the canonical token.

Scanner (`npm run check:phase-id-drift`) is now green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 08:49:36 -04:00
Behruz Nassre Esfahani
12b35eeeaf fix(#1729): resolve phase headers with a pre-colon parenthetical tag (#1765)
* fix(#1729): resolve phase headers with a pre-colon parenthetical tag

A phase header may carry a parenthetical tag between the number and the
colon, e.g. `### Phase 26 (Cluster B): Title`. Every phase-header regex
built `Phase\s+<num>` immediately against the colon delimiter, so the
tagged phase was invisible: the resolver returned found:false and, just
as bad, the capture-all enumeration/parse paths (roadmap analyze,
milestone listing + milestone-scope filter, verify, init/import, state
total_phases, validate, the command router, preamble stripping, and the
phase-remove renumbering rewrite) silently dropped, miscounted, or
failed to renumber it — wrong phase_count, progress_percent, next_phase,
or corrupt numbering after a removal.

The fix tolerates the tag at the header seam. Parameterized resolver
sites compose the exported OPTIONAL_PHASE_TAG_SOURCE fragment; literal
enumeration sites inline its character-for-character mirror
`(?:\s*\([^)\n]*\))?`, placed immediately before the colon so it cannot
alter an existing match (optional, single-line, one paren pair, no
capture-group shift). In the renumber-on-removal rewrite the tag is
folded into the re-emitted suffix capture so it survives verbatim. Both
forms are documented to change together and a drift-guard test asserts
their behavioral equivalence over a header corpus.

Deliberately excluded: roadmap-upgrade.cts (legacy one-time migration),
where tolerating the tag would silently drop it on header rewrite — that
needs its own data-preserving treatment. Known boundaries left for
follow-up: checklist/bullet-style phase entries (`- [ ] Phase N (tag):`)
and a malformed space-before-colon variant, both pre-existing.

Validated empirically against the issue's reproduction: `roadmap
get-phase 26` resolves and `roadmap analyze` lists Phase 26 with the tag
excluded from the name (phase_count 2, next 26); an all-tagged versioned
roadmap now scopes correctly instead of falling back to a pass-all
filter. Regression coverage in tests/phase.test.cjs asserts resolver
parity (pre- vs post-colon), padding tolerance (#3537), decimal
sub-phases, no cross-phase false match, the shared seam, enumeration
coherence, renumber-preserves-tag, and seam/mirror drift. Full unit
suite green (7291 pass, 0 fail); eslint + regression-name +
resolution-provenance + changeset lints pass. Reviewed by Codex
(no critical/high; the two enumeration misses it surfaced are folded in).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1729): add changeset for pre-colon phase-tag fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 11:55:44 -04:00
Behruz Nassre Esfahani
3870fafe74 fix(#1571): resolve schema-drift phase by token, not substring (#1640)
* fix(#1571): resolve schema-drift phase by token, not substring

verify schema-drift <phase> resolved the phase directory with a naive
entry.name.includes(phaseArg) test, so a non-existent phase could
silently match a different phase whose directory name merely contained
the requested token (e.g. "1" matched "11-expansion"), running the drift
gate against the wrong phase. Use the canonical phaseTokenMatches +
normalizePhaseName, matching find-phase, verify phase-completeness, and
this file's own unstarted-phase check.

Regression coverage folded into tests/schema-drift.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1571): add changeset for schema-drift token-match fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 17:03:01 -04:00
Tom Boucher
6425a3cb72 fix(#1493): read workflow.drift_action/drift_threshold from nested config shape in verify.cts
loadConfig() returns a flattened object with no nested `workflow` key, so
config?.workflow was always undefined, making drift_action permanently 'warn'
and drift_threshold permanently 3 regardless of .planning/config.json. Fixes
by reading the raw config.json directly (matching the pattern in
check-command-router.cts:readWorkflowConfig). Adds two behavioral regression
tests that fail under the old code and pass under the fix.

Closes #1493

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 16:37:54 -04:00
Tom Boucher
bb88a78faa fix(#1472,#1454): workstream-aware health paths; exclude active worktree from W017 (#1483)
* fix(#1472,#1454): validate health workstream-aware paths; exclude active worktree from W017

#1472: cmdValidateHealth now uses planningRoot(cwd) for shared-root files
(PROJECT.md, config.json, MILESTONES.md) and planningDir(cwd) for
workstream-scoped files (ROADMAP.md, STATE.md, phases/). Previously a
single planningDir() call was used for all paths, causing false
E002/E003/E004/W003 when GSD_WORKSTREAM is set.

#1454: W017 no longer fires for a stale worktree whose path equals or is
an ancestor of process.cwd(), preventing advice to remove the active
session's own worktree.

Regression tests added for both bugs; all 40 existing health tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:09 -04:00
Tom Boucher
2d264da661 enhance(#968): region-scoped negative-grep idiom + cross-task conflict warning (#1320)
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.

Closes #968

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 02:20:34 -04:00
Tom Boucher
a5f213e73e refactor(#1281): T2 — migrate 12 single-leaf callers off the core spine (batch 1) (#1282)
Per the T1 design rubber-duck, batch by FILE so each tranche drops
convergence-lint allowlist entries. Migrate 12 files' core imports to the
leaf modules directly (behaviour-identical — leaves are the objects core
re-exports by reference):
- io (output/error/ERROR_REASON): agent-command-router, capability-state,
  capability-writer, frontmatter, gsd2-import, learnings, loop-resolver,
  task-command-router
- roadmap-command-router -> config-loader; workstream-inventory -> core-utils
- milestone, verify -> their full leaf sets (both were multi-leaf, not
  single-leaf as first scoped; migrated completely)

All 12 files now import zero core symbols and are removed from the
allowlist (30 -> 18). core.cts re-exports untouched (still serve the
remaining 18 files); teardown is T-final. Stale core.cjs docstrings in the
migrated files corrected to reference io.cjs. No behaviour change.

Closes #1281

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 16:18:54 -04:00
Tom Boucher
54420bae9e refactor(#1277): T1 — decouple agent-install-check + git-base-branch from the core spine (#1280)
First leaf-migration tranche of epic #1267 (after T0 #1268). Migrate the
via-core callers of the two leaves T0 created to import from the leaf
modules directly, and stop core re-exporting their symbols:
- checkAgentsInstalled: docs.cts, verify.cts, init.cts -> agent-install-check.cjs
- gitWorktreeInfoInternal: init.cts -> git-base-branch.cjs
- getAgentsDir had no external via-core caller (internal to the leaf)

core no longer re-exports getAgentsDir / checkAgentsInstalled /
gitWorktreeInfoInternal; the now-unused agent-install-check + git-base-branch
requires are dropped from core; the shim-identity assertions for these are
deleted (behaviour tests retained). Convergence lint stays green (0 new).
No behaviour change.

Closes #1277

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:46:41 -04:00
Tom Boucher
2e8f4f6de1 fix(#1202): make verify key-links wave-aware for planned future files (#1219)
* fix(#1202): make verify key-links wave-aware for planned future files

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1219)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:28:14 -04:00
Tom Boucher
b10e56818b feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink

The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).

Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)

First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:

- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.

Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate profile-pipeline to a command-family Capability

ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).

Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1167): wire execute:wave:post + implement ui.safety-gate check

Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.

Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates

Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.

Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.

Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)

tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.

BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.

Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate schema-gate to a plan:pre contribution Capability

The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)

Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN

Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.

Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract

Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.

Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.

Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)

The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.

Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.

Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)

3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).

Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)

Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.

Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests

The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):

Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
  the inline schema_drift_gate step was removed — non-Claude runtimes would
  stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
  cite the execute:wave:post contract (loop body shrinks below the frozen
  pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
  baked verbatim into the committed capability-registry.cjs and leaked the
  install path on 11 non-Claude runtimes (registry .cjs is copied, not
  path-converted). Made the fragment path-free; regenerated the registry. The
  phase-6 conformance gate now guards this (no ~/.claude install path in any
  capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
  step 6 (schema-gate is a plan:pre capability, §5.7 is gone).

Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).

profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.

Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables

Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).

Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.

ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.

federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): remove dead drifted converter dups + address adversarial review

Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.

Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
  disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
  `node -e` form (node is guaranteed; matches the file's other node-e usages) so
  a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
  the orphan assertion). Now asserts the removed capability's key is genuinely
  not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
  Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
  `~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
  stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
  structured-output check.

Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): make runtime-homes-descriptor-drive titles environment-independent

The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.

Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)

The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).

Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.

Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:07:55 -04:00
Tom Boucher
2023f47c64 fix(#1058): cross-reference install manifest in validate agents to catch .md/.toml pair drift (#1079)
* fix(#1058): cross-reference install manifest in validate agents to catch pair drift

`validate agents` considered an agent installed if ANY supported file format was
present on disk. The Codex installer generates a per-agent PAIR (agents/gsd-*.md
AND agents/gsd-*.toml) and records both in gsd-file-manifest.json, so a partial
generated install — one side of the pair missing — was reported as healthy
(agents_found: true, missing: []), masking an incomplete Codex agent install.

checkAgentsInstalled now cross-references the install manifest beside the agents
dir (path.dirname(agentsDir)/gsd-file-manifest.json): for each expected agent, if
the manifest tracks files for it and any tracked file is absent on disk, the
agent is reported in a new `incomplete` list and agents_found becomes false. The
check no-ops when no manifest is present (preserves bundled/claude behavior) and
is scoped to expected agents so retired/stale manifest entries cannot false-flag.

Regression cases added to tests/agent-install-validation.test.cjs cover the
drift case, the complete-pair (no false positive), and the no-manifest no-op.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1058): add changeset for validate-agents manifest pair-drift fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 20:50:08 -04:00
Tom Boucher
8813ee5f95 feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.

`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
  <action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope

Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).

Closes #429

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 15:46:49 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
808df9110c fix(#892): parse checklist-style roadmap phases in validate/verify (#908)
buildRoadmapPhaseVariants() only matched heading-style phases (## Phase N:),
silently skipping the supported checklist format (- [x] **Phase N: name**).
This caused W007 false-positives for every on-disk phase dir when the project
uses a checklist ROADMAP. Fix adds a second regex pass (mirroring the existing
buildNotStartedPhaseVariants() approach). Also refactors the duplicate
inline heading-only regex in cmdValidateConsistency() to delegate to
buildRoadmapPhaseVariants() (DRY). Regression test in
tests/bug-892-validate-checklist-roadmap-phases.test.cjs covers both paths.

Closes #892

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:52:19 -04:00
Tom Boucher
b238baddbf fix(#663): resolve open CodeQL/Dependabot security alerts (ReDoS, prototype pollution, workflow perms, qs DoS) (#665)
* fix(#663): resolve open CodeQL/Dependabot security alerts

- ReDoS: collapse ambiguous nested quantifiers in phase-heading regexes
  (verify/validate/commands) and the plan-filename lookahead (phase) to
  provably-equivalent non-backtracking forms
- prototype pollution: guard __proto__/constructor/prototype in setConfigValue
- remove dead no-op .replace(/-/g,'-') in phase.cts
- escape all regex metachars in bug-2839 test
- add contents:read permissions to security-scan + install-smoke workflows
- pin qs >= 6.15.2 via overrides (DoS GHSA)
- broaden prompt-injection allowlist to translated security-model docs

Closes #663

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#663): regression tests for prototype-pollution guard and roadmap-phase ReDoS

Behavioral test that config-set rejects __proto__/constructor/prototype keys
without polluting Object.prototype, plus a ReDoS guard (timing-bound) and
behavior-preservation assertions for the collapsed phase-heading regexes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#663): make ReDoS regression assert structured result, not elapsed time

Replace elapsed-time assertions (which tripped local/no-elapsed-assertion
ESLint rule and were unsound for synchronous ReDoS) with structured-result
assertions on adversarial inputs: assert that malformed phase headings/
unchecked-item lines without a terminating colon/space yield an empty Set,
which is both the correct behavior and an exercise of the fixed linear regex
on the catastrophic-backtracking input shape.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#663): add Security changeset fragment for #665

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#663): fold prototype-pollution regression into config.test.cjs

The standalone bug-663-config-prototype-pollution.test.cjs was a 9th
config-module test file, tripping lint-test-file-count (the allowlist is
ratcheted and must not grow). Consolidated into config.test.cjs instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 08:54:15 -04:00
Tom Boucher
df04aae5e4 enhancement(#537): migrate all hand-written bin/lib/*.cjs to TypeScript source of truth (ADR-457) (#602)
* enhancement(#537): migrate code-review-flags to TS source of truth

Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a
TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a
gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish).
Second module after the semver-compare pilot (#541).

Behaviour is preserved byte-for-behaviour (characterization test added in
tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time
type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union.
The require() path is unchanged, so code-review.md and the bug-3727 test keep
working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth

ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps):
001-legacy-orphan-files, context-utilization, redaction, artifacts,
command-arg-projection, clock, ui-safety-gate, review-reviewer-selection,
clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a
gitignored .cjs at the same require() path; behaviour preserved byte-for-
behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): add @types/node, drop hand-rolled node-globals shim

ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient
shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks
migrating the ~49 remaining bin/lib modules that use node:fs/path/os/
child_process. Build + full suite (3030 pass) + lint all green; no .cts type
changes were needed (real Node types matched the shim).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2)

ADR-457 build-at-publish. Clean leaves: installer-migration-report,
prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now
strict-typed and removed from that exclude list): secrets, phase-lifecycle,
workstream-name-policy, decisions, validate, schema-detect. Plus
runtime-name-policy. Strict type fixes narrow unknown->concrete domain types
(no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate runtime-slash to TS (cross-import proof)

ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports
./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict
(no declaration files; NodeNext .cjs->.cts mapping), emitting a correct
require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules,
which must be migrated in dependency order (leaves-up). Suite green, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3)

ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder,
plan-scan, fallow-runner, project-root, installer-migration-authoring,
update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict
typing fixed real issues (narrowing unknown, qualified fs/path calls, removed
unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped
from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4)

ADR-457 build-at-publish: configuration, state-document, shell-command-
projection (42 dependents), security, command-aliases. shell-command-
projection keeps a namespace child_process import for mock-intercept
testability. loadConfig/migrateOnDisk emit synchronously (every caller uses
them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) —
full suite (3030 pass) confirms behaviour preserved. configuration/
state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes
the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5)

ADR-457 build-at-publish: config-schema, model-profiles,
002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser.
First batch importing already-migrated siblings (configuration, model-catalog,
shell-command-projection, redaction, security) via ./sibling.cjs specifiers.
Strict type narrowing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6)

ADR-457 build-at-publish: graphify, install-profiles, intel,
installer-migrations, worktree-safety. installer-migrations preserves its
dynamic require() loader for numbered migration modules (scoped lint
suppressions). Strict typing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors. Wave 2 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate Wave-3 modules to TS (batch 7)

ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout,
command-routing-hub, drift. Uses `import x = require()` for export= siblings;
drift's lazy require of runtime-slash hoisted to a top-level import (verified
non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate small Wave-4 modules to TS (batch 8)

ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router,
surface, roadmap-upgrade. Typed the hub router handler results as the HubResult
discriminated union; surface drops 4 genuinely-unused imports. Behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9)

ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts,
preserving all 63 exports via export=. All sibling deps already migrated
(shell-command-projection, model-profiles, model-catalog, worktree-safety,
planning-workspace, project-root, configuration, config-schema). Strict types,
no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour
preserved (independently verified: core's shard 3030 pass / 0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): make ESLint-coverage + test-sprawl checks migration-aware

#551 test hardcoded 12 now-migrated modules as "hand-written, must be linted";
that invariant is obsoleted by the ADR-457 migration. Rewrite it to a
filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be
eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted
(covers package-identity, which has no TS source). Also eslint-ignore
config-types.cjs (has a src counterpart) and drop the redundant
tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests),
which tripped the lint-test-file-count ratchet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10)

ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state
command routers + workstream-inventory. Router handler results typed against
core's exported shapes; behaviour preserved (caught+fixed a --verify boolean
flag regression mid-migration). Full suite green across all shards (only the 4
local gpg-env changeset-notes failures remain; CI passes them).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11)

ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter,
learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green
across all shards (only the 4 local gpg-env failures remain). Also broadens
atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form
while still asserting platformWriteSync is called (safety guard intact).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate config + profile-output to TS (batch 12)

ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All
exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited
caller tolerates it). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures). Wave 5 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13)

ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour
preserved (dead toPosixPath import dropped from audit; inline requires hoisted).
Suite green across all shards (only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate commands + state hubs to TS (batch 14)

ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents).
All exports preserved; inner requires kept non-hoisted where load-order matters
(install.js, per-call security); acquireStateLock cast inlined to preserve the
err.code source token a structural test inspects. Behaviour preserved; suite
green across all shards (only the 4 local gpg-env failures). Wave 6 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate milestone to TS (batch 15a, hand-authored)

ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly
(subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to
optional, matching its real always-optional call contract (unblocks remaining
2-arg output callers). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules)

ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify
(1615), init (2113). Adds src/package-identity.d.cts so verify can import the
permanently value-baked package-identity.cjs under strict TS.

Fixes two regressions the migration introduced in verify: restore
cmdValidateHealth's `return result` (callers/tests read result.warnings — it is
NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the
tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour
intact). Full suite green across all shards (only the 4 local gpg-env failures);
lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json

All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs
stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is
deleted per ADR-457's final step. Also gitignore the tsc-generated
config-types.cjs (was still committed) for consistency with every other
emitted artifact. package-identity.cjs stays value-baked (declared via
src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts

ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The
prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g
<dir>` and git installs run the `prepare` lifecycle — which was missing — so the
unpacked install shipped without the compiled .cjs and failed at startup with
"Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job).
Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does
NOT run for registry consumers (they get the pre-built tarball), only for
source/local/pack installs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts

ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are
always present on disk:
- check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the
  `prepare` build (tsc) — but it runs before deps are installed, so tsc is
  absent and it misreported the lockfile as out of sync. Add --ignore-scripts
  (a lockfile check must not build).
- the lint-tests job installs with --ignore-scripts (no prepare build), but
  lint:skill-deps require()s the built install-profiles.cjs. Add an explicit
  `npm run build:lib` step after install.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step)

prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`,
which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT —
breaking it with "Invalid format". build:lib (tsc) is silent on success and is
all the unpacked/source install needs (the smoke-unpacked assertions exercise
gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches
prepack. build:hooks still runs on prepublishOnly for real publishes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): wire Stryker mutation gate to build-at-publish layout

The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were
generated artifacts and (b) included modules with no coverage in the command's
test set. Rework: mutation.yml now derives changed COVERED modules from
src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those
built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc
was ~3x over the 30-min CI budget).

NOTE: with the gate now correctly measuring the covered modules, their actual
mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap
(adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving
migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver —
a maintainer decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): raise mutation coverage of covered modules above the 50 gate

Adds focused example-based unit tests that kill surviving mutants in the two
lowest-scoring covered modules:
- tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9%
- tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4%
Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered
modules now scores 82.25% (>= break 50); every covered module is >= 68%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix

The serial Stryker run timed out at 30 min once the migration's added tests
made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the
gate completes well under budget — folded into this PR (was tracked as #609)
because it's a prerequisite for this PR's mutation gate to pass.

- scripts/mutation-matrix.cjs: single source of truth (covered-module -> test
  files) computing changed covered modules from git diff -> {has_work, matrix}.
- mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one
  parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that
  module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name
  "Stryker mutation score (changed files only)" so branch protection is
  unchanged. Per-shard jobs report as "Stryker (<module>)".
- stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back
  to the full command locally).

Closes #609.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note

Per-module sharding revealed that active-workstream-store (46.5%) and
frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from
the bloated 300-test command; on their own tests they were below the gate. Add
focused unit tests:
- tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9%
- tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4%
Both wired into scripts/mutation-matrix.cjs (per-module test map) and
stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50
with only their own tests (config-schema/context-utilization/prompt-budget/
adr-parser already did). Also removes the leftover blacksmith TODO comment —
GitHub-hosted runners only; speed comes from parallel per-module shards.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests

prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the
way the per-module CI shard runs it) — an earlier ~98% reading was inflated by
accidentally running the full multi-module command. Add 96 targeted tests to
tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation
arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans):
scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear
the gate on their own tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:45:01 -04:00