Commit Graph

5454 Commits

Author SHA1 Message Date
Tom Boucher
4e8927b0b9 fix(#3707): degrade the fold for every UAT gap class, and stop line endings hiding rows from the audit and the acceptance gate (#3903)
* test(#3707): failing-first coverage for reverting the fence-shortfall fold shield

Pins the post-revert contract: a phase whose only gap is a fence shortfall must
degrade the fold and withhold the milestone percentages, like every other gap class.

Five of the eight rows are CONTROLS that pass before the change, and they carry more
weight than the failing row. The failure mode of this revert is degrading TOO MUCH:
a revert that sets foldScope outside the headingsSeen > 0 branch would withhold every
percentage in the project, and only the no-gap control catches that. Another control
catches a revert that collapses the two scopes into one and loses the distinction
between what a phase reports and what the fold folds -- uat.scope must stay TRUNCATED
for every gap either way, which it already is.

The row that pinned the shielded behavior is rewritten rather than deleted. Deleting
a test because the behavior it asserts is being reversed leaves the reversal
unguarded.

* fix(#3707): degrade the fold for every UAT gap class, reverting the fence-shortfall shield

Maintainer decision. The two orthogonal engines split on this during #3707 and
neither filed it as blocking, so it shipped in the shape the engine that raised the
objection endorsed after verifying seven fixtures. The call has now gone the other
way, restoring the fail-safe direction chosen twice already on this issue.

The shield exempted one gap class from the fold's teeth. It could not do that
safely: shortfallBlocks is a single tally incremented at exactly one site and spans
BOTH a harmless fenced documentation sample AND a genuinely fence-straddled
result: blocked row. Exempting it therefore could not exempt only the harmless case
-- it also published a milestone percentage over a real, unread outstanding row.
SCOPE.TRUNCATED means the scan could not SEE part of the evidence, which is exactly
that case.

scope and foldScope now agree: every gap class degrades both. The accepted
over-report documented in uat.cts is unchanged and still documented there; what
changed is only that it no longer buys an exemption from the fold.

The comment block above it argued FOR the shield and is rewritten, because a
comment defending behavior the code no longer has is worse than no comment.
shortfallBlocks leaves this function's destructure but is untouched upstream, where
audit-uat still consumes it.

* fix(#3707): correct the caller comment, add the changeset, and name what the order tests guard

Review found a SECOND comment still documenting the removed shield -- the caller's,
beside the worstScope fold, stating that foldScope differs from scope for exactly
one case which must not raise phase_scope_degraded or withhold the milestone's
percentages. That is now the opposite of what the code does. I rewrote the
buildUatRows comment in the previous commit and asserted in its message that a
comment defending behavior the code no longer has is worse than no comment, then
left exactly that one standing a few hundred lines away.

The change had no changeset. It is user-visible: a milestone's percentage goes from
published to withheld whenever any phase has a fence-shortfall-only gap. PR gates
hard-fail a user-facing code diff without one.

The two scopes are now identical at every return site. They are NOT collapsed --
that would change the return shape and the caller on what is meant to be a
one-condition revert, and the seam is worth keeping if the distinction is ever
wanted again -- but the declaration now says plainly that they agree by decision
rather than by accident, so a reader does not have to re-derive it.

The two order-independence tests were renamed. foldScope is monotonic with no reset
path, so file order is structurally irrelevant and those rows could never have
failed for the ordering reason their names promised. They do guard something real --
a multi-file phase degrading when any one file has a shortfall-only gap -- so they
now say that instead.

* test(#3707): failing-first coverage for the lone-CR UAT false-clean

The parser splits on newline only, and the heading tokenizer agrees with it, so a
lone carriage return is not a line boundary anywhere in it. CommonMark treats a lone
CR as a line ending, so such a row renders to a human reader while being invisible
to BOTH sides of the parser's symmetry invariant: no item, no shortfall, no
headingsSeen. A phase hiding a result: blocked row this way reports 100 percent with
zero diagnostics.

Found by the security review of the fold-shield revert. It is the one false-clean
class that revert does not reach, and it is the same bug class this issue exists to
fix -- an unreadable row reported as clean.

Nine rows. The LF control is what proves this is a separator defect rather than a
content defect: identical bodies, one separator apart, and only one of them hides
the row. CRLF and CR-inside-a-fence controls guard the coming normalization against
double-counting or tearing content that legitimately contains a carriage return.
Two further manifestations turned up while writing them: a leading CR breaks
column-0 anchoring of the first heading, and an all-CR document flags a shortfall it
cannot attribute to any row.

* fix(#3707): treat a lone carriage return as a line ending in the UAT parser

A lone CR was not a line boundary anywhere in the parser -- it split on newline
only, and the heading tokenizer agreed with it. CommonMark treats a lone CR as a
line ending, so such a row rendered to a human reader while being invisible to BOTH
sides of the parser's symmetry invariant: no item, no shortfall, no headingsSeen. A
phase hiding a result: blocked row that way reported 100 percent with zero
diagnostics.

Line endings are now normalized once at document ingress -- CRLF and lone CR both to
newline -- at the two independent entry points, rather than teaching each split site
about CR. Every downstream scan, offset and span therefore reads one convention.
That single-frame property is deliberate: this issue already cost a HIGH when two
scans read the same document through different frames.

MY OWN END-TO-END TEST WAS WRONG and is replaced rather than weakened. It asserted
that a lone-CR document must withhold its percentage, which reasons from the
pre-fix symptom: after the fix the row is not hidden, it is surfaced, and this
module deliberately keeps visible outstanding UAT work separate from completion
percentages -- only unreadable evidence degrades scope. The success of the fix is
what made the assertion false. The implementing agent refused to satisfy both it and
the architecture and asked instead of bending either; it was right.

What replaces it is a stronger contract: a lone-CR document and its LF twin, built
from one source, must produce identical audit output -- scope, percent, every
unresolved row by identity, and the diagnostic set. That is what 'a line-ending
convention must not change what the audit reports' actually means, and it carries a
non-vacuity check so it cannot pass with both sides empty.

shortfallBlocks keeps being returned, now documented as currently unconsumed. An
earlier reviewer told me audit-uat still consumed it and I passed that on as an
instruction; it was wrong, and it was caught by checking rather than by me.

* fix(#3707): normalize at the document read boundary, not at two call sites

The lone-CR fix was half-applied and both review engines caught it independently.
cmdAuditUat has four document ingresses, not the two I normalized: VERIFICATION.md
and deferred-items.md still handed raw text to newline-only splitters, and the
frontmatter extract in the UAT loop read raw content while its parser read
normalized -- one audit entry mixing the two frames the fix exists to unify.
Measured: a phase written twice from one source gave total_files 2 / total_items 4
under LF and results [] / total_items 0 under lone CR, with zero diagnostics.

Normalizing two call sites and declaring it done is exactly why two were missed, so
this moves it to the read boundary: every document now enters through a helper that
normalizes, in audit-uat, in planning-inspect's readDocument, and in the shared
verification-status read. Future parsers downstream get normalized text by
construction rather than because someone remembered.

That last seam also fixes an under-reporting case of the same root: a lone-CR
VERIFICATION.md saying status: passed was read as missing, telling the user a verify
step that had completed never ran.

The parity test's load-bearing assertion is now marked as such. Four of its five
equality checks still pass with the bug present -- only the unresolved-row identity
differs -- so trimming that one as redundant would make the row vacuous.

Second changeset added: the CR fix is user-visible independently of the fold revert,
and one fragment covering both would have described neither.

* test(#3707): failing-first coverage for the U+2028 and duplicate-result false-cleans

Two more of the same class, both found by the security review of this branch and
both reproduced before writing a line.

normalizeLineEndings folds only carriage returns, but a JS /m anchor also treats
U+2028 and U+2029 as line terminators while split on newline does not. That is the
identical asymmetry the carriage-return bug exploited, one separator over, and worse
in one respect: these are not CommonMark line endings, so a reader still sees the
column-0 result: blocked that the tool discards. Measured: a scalar-internal
result: pass placed after U+2028 wins over the real blocked line and the row
disappears with no gap raised.

Separately, and independent of any separator, a block with two column-0 result:
lines resolves to the first with no ambiguity signalled. Prepending result: pass to
a block therefore deletes an outstanding row silently; reversing the order surfaces
it. Order deciding meaning is the defect, so the pair of rows pins the contract as
ambiguity-is-a-gap rather than last-one-wins, leaving the fix room to implement the
gap sensibly.

Four controls: an ordinary marker in the same position (proving separator not
content), legitimate U+2028 inside prose that must not be torn, a single result line,
and a result line inside a fence that must not count as a second occurrence.

* fix(#3707): scan result lines by split, not by a multiline anchor

Two more false-cleans from the security review, both closed by the same change.

A JS /m anchor treats U+2028 and U+2029 as line terminators while split on newline
does not. A scalar-internal result: pass placed after one of those separators
therefore matched as a line start and beat the real column-0 result: blocked, and
the row vanished at 100 percent with no gap. Worse than the carriage-return case in
one respect: these are not CommonMark line endings, so a reader still saw the
blocked row the tool discarded.

Separately, the non-global match returned the leftmost hit, so a block with two
column-0 result: lines silently resolved to the first. Prepending result: pass
deleted an outstanding row; reversing the order surfaced it. Order deciding meaning
was the defect.

Both close by scanning lines produced by split rather than by anchoring a regex
inside the whole document: each line is tested on its own, and a count other than
exactly one is reported as a parse gap instead of resolved to either candidate.

I asked for U+2028 to be folded in normalizeLineEndings and that was wrong. Folding
is length-preserving, so it would have made the U+2028 fixture byte-identical to the
genuine two-result-line fixture -- while one requires a confident item and the other
requires an ambiguity gap. No implementation can satisfy both once the distinguishing
character is erased. The agent proved that and deviated rather than forcing it, which
is why normalizeLineEndings still folds only carriage returns, now with a comment
saying why.

* fix(#3707): bound the ambiguity scan at the next heading-shaped line

The split-based result scan regressed four pre-existing #3078/#3707 guards, each
off by exactly one gap.

My diagnosis was wrong. I read the off-by-one as double counting -- zero-result
blocks taking both the new path and the pre-existing one -- and said to change the
ambiguity condition from not-equal-one to greater-than-one. The agent checked and
refused: the zero path was never duplicated. The real cause is double ATTRIBUTION.
A block is sliced to the next TOKENIZED heading, so when the next row is untokenized
-- hidden by a straddling fence, or indented and already counted by the shortfall
scan -- that row's own result: line is absorbed into the previous block. The scan
then saw two result lines across what are really two rows and raised a second,
redundant gap on top of the one already counted elsewhere.

Had the greater-than-one change gone in, the counts would have matched while the
double attribution stayed. That is the compensating-adjustment failure I had asked
it to refuse, and it did.

The scan is now bounded at the first following heading-shaped line, either indent
class, so a genuine same-block ambiguity is untouched while spillover from a row
counted elsewhere is excluded.

* fix(#3707): keep the U+2028 immunity, revert the ambiguity detection

The ambiguity half of this change regressed the suite twice and is coming out.

Attempt one double-attributed: a block is sliced to the next TOKENIZED heading, so
when the real next row is untokenized its result: line was absorbed into the
previous block and raised a second gap on a row already counted elsewhere. Four
guards broke.

Attempt two bounded the scan at the next heading-shaped line and broke thirty. An
indented ### N. inside a block scalar is legitimate scalar CONTENT, not a heading,
and truncating there defeats every #3078 guard that exists to stop scalar bodies
being read as rows. Telling a genuinely hidden indented row apart from indented
scalar text is a classification countUnattributedIndentedRows already owns; a raw
regex does not have that information.

What survives is the half that is sound and was never implicated in either
regression: the result scan tests each line produced by split rather than anchoring
a regex with the multiline flag over the whole block. split never treats U+2028 or
U+2029 as a delimiter, so those separators can no longer manufacture a line start
and steal a row. Everything else returns to first-match-wins, byte-identical to
origin/next.

The two tests pinning ambiguity-as-a-gap are removed with it, since the contract is
no longer implemented here. The defect they described is real, pre-existing and
independent of any separator -- result: pass before result: blocked silently deletes
an outstanding row -- and it needs its own change with a scalar-aware counter rather
than being wedged into a branch already carrying three fixes.

* fix(#3707): correct the shared-seam rationale and restore U+2028 trailing text

The revert left a stale rationale in core-utils, justifying the decision not to fold
U+2028 by claiming uat.cts must tell a fake line start apart from a real second
column-0 result: declaration that gets flagged as ambiguous. Nothing flags ambiguity
any more; that behavior was reverted and the same file says so a few lines away. The
decision is still right, the stated reason was false.

This is the third stale comment this branch has shipped and had to fix, and the worst
placed of them: core-utils is a shared leaf that every future document consumer will
read for guidance. Rewritten to the true reason -- the scan tests each split line
individually rather than anchoring over the block, so an exotic separator cannot
manufacture a line start and folding is unnecessary.

Also a real behavior delta I had not noticed. Dropping the multiline flag left the
pattern's trailing .*$ in place, and dot never matches U+2028, so a genuine column-0
result: blocked whose TRAILING text contained one stopped parsing entirely -- a
visible parse gap rather than a false clean, so fail-safe, but a regression against
origin/next that nothing pinned. The trailing portion now matches any character and
a test pins it by identity against its plain-LF twin.

Plus the JSDoc orphaned when normalizeLineEndings moved to core-utils, and the
changeset, which described neither the separator fix nor planning-inspect surfacing
lone-CR rows.

* fix(#3707): harden the acceptance gate, which had both halves of the same bug

uat-predicate is a SECOND, independent UAT parser, and it is the one that decides
phase uat-passed. It read raw bytes and anchored a multiline regex over unsplit
text -- exactly the two defects this branch closed one module away in uat.cts.

The consequence is worse than the audit surface it mirrors. Measured on identical
bytes: a U+2028 scalar injection made the gate return passed true while planning
inspect reported the same row as blocked and outstanding. The hardened surface and
the gate disagreed, and the gate was the permissive one -- so a phase could be
accepted over a row the audit could see and the gate could not.

Both raw reads now go through the shared normalize seam and both scans test lines
produced by split rather than anchoring over the document. First-match-wins,
matching uat.cts; no ambiguity counting is reintroduced. Tests assert the AGREEMENT
between the two surfaces rather than each separately, because divergence is the
defect.

Also finishes the same root cause one module over: phase complete's advisory
pre-scan read raw bytes, so a lone-CR VERIFICATION.md lost its human_needed or
gaps_found warning -- the fix verification.cts already got on this branch.

And narrows the core-utils rationale I reworded last commit, which claimed consumers
already avoid multiline anchors. uat.cts still has five over unsplit text. That is
the fourth comment on this branch to assert something the code does not do, so it
now states only what is true of core-utils itself.

* fix(#3707): give structure and attribution different line frames, normalize the close audit

Two more from review, and the first was a regression I introduced one commit
earlier.

Converting the gate's heading scan to split-then-match removed a detection
origin/next had: a ### N. heading delimited by U+2028 was found by the old multiline
scan and was not found after. So hardening the result scan quietly weakened the
heading scan, and the gate stopped blocking on rows origin/next blocked -- the
permissive direction, on the surface that decides acceptance.

The insight I had missed is that the two scans need DIFFERENT frames. Heading
detection is structure: there is no distinction to preserve, so it splits on newline
or either exotic separator and finds a heading however it is delimited. The result
scan is attribution: the newline-only frame is exactly what stops a scalar-internal
result: from being read as a column-0 line, so it stays. One frame applied uniformly
was the error.

Second, a THIRD unnormalized parser family: the milestone-close audit read every
artifact raw. A lone-CR VERIFICATION.md degraded to status unknown and was skipped,
and deferred entries vanished outright -- measured as three items requiring
decisions under LF and one under CR, on identical bytes. All nine scanner reads now
normalize; six of them had the identical defect beyond the three review named. The
acknowledge path stays deliberately raw, since it splices by byte offset, and now
says so.

Also pins the cross-newline result: divergence, and replaces three raw U+2028
literals in test source with escapes. A raw separator in a fixture is one formatter
away from becoming an ordinary-character control that still passes -- vacuous in the
only test pinning the separator fix.

* fix(#3707): share one frame between the acknowledge writer and the audit reader

Normalizing the audit scanners left the writer and the reader on different frames.
cmdAuditAcknowledge derives its stored snapshot values from raw content -- correct
for the SPLICE, which rewrites by byte offset -- but scanUatGaps and
scanContextQuestions now recompute those same values from normalized content. For a
lone-CR artifact the two can never match, so an acknowledgement never suppresses its
item and it resurfaces on every audit: acknowledge became a silent no-op.

Fail-safe in direction, since the item stays visible rather than being wrongly
suppressed, but it is the writer and reader disagreeing about what a line is -- the
exact class this branch exists to eliminate, and the fourth instance of it here.

The derive functions now read a normalized copy while the splice keeps raw bytes and
raw offsets, so both sides share one frame and the byte-offset rewrite is untouched.
Round trip pinned for lone-CR and LF, with an existing LF marker asserted still
recognised so the change cannot silently invalidate acknowledgements already in
users' files.

Also tightens an assertion that pinned this branch's own heading fix with a proxy:
notStrictEqual against 'passed' also passes on 'pass', which IS a passing token, so
it could not have caught a regression attributing a passing result to the recovered
heading. It now pins the exact token.

* chore(#3707): backfill changeset pr numbers

Both fragments still carried the pr: 0 placeholder, which failed changeset-lint and
docs-lint on PR 3903. The review had flagged the backfill as pending and I opened
the PR without doing it.

---------

Co-authored-by: sim <sim@local>
2026-08-26 20:17:35 -04:00
Tom Boucher
6b7df61938 enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises

ADR-3473 §8.1 carries a blocking open question with a forcing function: it must
be answered before any implementation PR for the rule opens. Answered here as (a),
a string-coercing adapter, with the measurement that settles it.

The sequencing note bet that §8.8's schema would make (b) tractable. Measured
against merged reality it does not: only 33 of extractFrontmatter's 78 non-test
call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists
survive real types, leaving ~31 lines across 3 call sites as the actual prize.

Also corrects three claims verified false while answering it. §8.1's justifying
sentence names #3349 and #3360 as defects a real parser would fix; both are
already fixed on next, confirmed by executing the compiled parser rather than
reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an
expected casualty of this rule, but it guards shell grep idioms in workflow bash
fences and never touches our parser. The same roster calls lint-vendored-deps.cjs
reusable as-is; it is hardcoded to re2js throughout.

The last two were caught by applying the rule this amendment records -- a factual
claim in this ADR is a hypothesis until the implementing phase executes it -- on
its first use.

Refs #3881

* docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable

An adversarial pass on the Phase 4 design established by execution that
extractFrontmatter is not a YAML parser but a line-oriented scanner whose output
is a function of raw source text. Four spellings of the same value collapse to
one js-yaml tree but produce four distinct legacy strings, one of them mangled.
No adapter over a tree can choose among outputs the tree does not distinguish,
so fork (a) -- keep a string-coercing adapter so the existing contract holds --
cannot be built. For any document with a non-scalar value, (a) collapses into
(b); about 26 percent of frontmatter-carrying documents have one.

Also records three design defects and one new attack surface, all confirmed by
execution: catching a parse failure and returning {} would delete the frontmatter
block on the next write at eight call sites that conflate empty with unparseable;
an empty value yields null where legacy yields {}, and reconstructFrontmatter
omits null-valued keys, so the shipped state template's empty progress key would
vanish; the #1882 truncation probe is parseYamlRegion itself rather than a
pre-parse heuristic, so it cannot both stay unchanged and survive that deletion;
and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB.

The rule is not deferred. The measurement is the deliverable and the re-scoping
is recorded as an open question with a forcing function, per section 8's own rule.

Refs #3881

* test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix

Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed.

Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states.

B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today.

B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today.

B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today.

Refs #3881

* chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest

Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496).

gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own.

src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare.

scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored.

docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds.

Refs #3881

* feat(#3881): parse .planning frontmatter with the vendored js-yaml

ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled
line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and
parseQuotedScalar are deleted (not patched); parsing now goes through the
vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA,
json: true }. Everything js-yaml does not do is layered on top, in one
place, carrying the seven design-doc consequences:

1. Empty value: a null js-yaml value is coerced to {} (matching legacy's
   own empty-value contract) so reconstructFrontmatter — which omits
   null-valued keys — still round-trips a bare `key:` line instead of
   deleting it. Verified live: progress: with no value survives
   parse -> reconstruct -> re-parse.

2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE
   Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS
   channel, is carried on the {} returned for malformed/refused YAML. Invisible
   to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that
   never inspect it are unaffected; wiring the 8 hasFrontmatter sites to
   consult it is a separate change, not done here.

3. Non-scalar object-list items (the four spellings of `- test: a b` that
   js-yaml collapses into one tree shape) are rendered as a canonical
   `key: value[, key2: value2]` string per item, keeping the existing
   array-of-strings value SHAPE. A full corpus differential over all 1702
   tracked markdown files found 11 residual divergences from the legacy
   parser (enumerated in the PR/report), most of them the parser now being
   MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key
   defect, a dropped Unicode key).

4. The #1882 truncation probe still runs the one real parser, but derives
   its key count from js-yaml's own thrown error and mark.line when the
   whole region doesn't parse cleanly (the dominant real truncation shape:
   fence opened, well-formed keys, no closing fence). Verified against both
   the clean-parse and the exception-fallback path.

5. The #3257 comment channel now attributes each pending column-0 comment
   against js-yaml's own parsed top-level key list (matched by literal key
   text, in document order) instead of the legacy ASCII-only key regex, so
   a comment above a Unicode key attaches correctly.

6. Anchors, aliases and merge keys are refused outright (a raw-text
   pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences
   today: zero. A 7-line billion-laughs fixture is verified refused rather
   than expanded.

7. A literal U+0000 is swapped for a private-use sentinel before the parse
   and restored in every resulting string afterward, since js-yaml rejects
   NUL unconditionally under every schema.

escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump()
(forced double-quoted style), with control-char hex escapes lowercased to
keep serialized output byte-stable (#1779 emitted lowercase); it keeps its
exported name and signature for its two other call sites (commands.cts,
runtime-artifact-conversion.cts), which need no change.

frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments,
regenerateFrontmatterKey's guard, noOpObjectListSetError and
parseMustHavesBlock are all unchanged — retiring them is fork (b) and is
not this phase.

Refs #3881

* fix(#3881): quote template placeholders and preserve unparseable frontmatter

SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as
bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax
under the vendored js-yaml parser, not the literal placeholder text
intended. Quote them so they parse as strings.

Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the
8 call sites in state.cts/state-transition.cts that compute
hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and
reassemble the document without a frontmatter block when false. That
check conflated 'no frontmatter' with 'unparseable frontmatter' (both
parse to {}), so a document with a merge-conflict marker or refused
alias in its frontmatter had that block silently dropped on write.
Each site now preserves the exact raw bytes stripFrontmatter removed
when the marker is set, leaving the genuinely-empty case unchanged.

Refs #3881

* test(#3881): consequence and boundary coverage for the js-yaml migration

Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix.

Refs #3881

* docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to

Refs #3881

* docs(#3881): correct the frontmatter glossary entry

Two errors in the entry as first written: it named parseYamlRegion as part of
the read path when that function is deleted, and it recorded the eight
hasFrontmatter call sites as unwired follow-on work when they were wired in
e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the
frontmatter block independently, so the marker binds at the transform layer.

Refs #3881

* docs(#3881): record the semantic-migration decision and the counted guard ledger

The maintainer chose the full semantic migration over splitting the rule into
its own epic or patching the scanner, so section 8.1 is answered as "the fork
was ill-posed and the migration is semantic" rather than as (a) or (b).

Also replaces the pre-implementation guess that this phase would shrink the
guard surface with the counted result: excluding vendored third-party lines the
hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines
despite four functions being deleted, because the compatibility layer over
js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is
therefore not delivered as written; what improved is the kind of code
maintained, not the amount. Decision 6 requires recording that rather than
netting it away.

Refs #3881

* chore(#3881): changeset for the vendored YAML parser migration

Refs #3881

* test(#3881): golden parity, round-trip property and packaging coverage

Refs #3881

* fix(#3881): refuse anchors structurally and fold in review findings

ADR-3473 §8.1 review findings, addressed inline:

Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched
only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping
({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor
mechanics while never matching that line shape, so the exact expansion the
guard exists to stop went straight through unrefused (a 303-byte quoted-key
bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener`
callback, which reports `state.anchor` for every event belonging to an
anchored node in every spelling, and throws from inside the callback to abort
before any expansion (~1-2ms vs full expand-then-discard). A merge key with
an alias is still refused (merge always requires a previously anchored node,
so the alias itself trips the listener); a bare merge key with NO alias is no
longer separately refused, documented as intentional: FAILSAFE_SCHEMA never
resolves `!!merge`, so it carries no expansion risk. Table-driven tests added
for all four bypass spellings + merge key, plus a quoted-key-spelled
billion-laughs fixture registered in the adversarial matrix and README.

Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/
aliases were "simply UNREACHABLE from typed code" through the twin. Corrected
to state the truth: anchor/alias resolution is document-level `load`
mechanics reachable through exactly the declared surface, and refusal is
enforced at RUNTIME (Finding 1's listener), not by the type surface.

Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was
non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree
back to NUL, including one the document author legitimately wrote, silently
corrupting it. Now refuses outright whenever the raw region already contains
U+E000 (consistent with the existing anchor/merge-key refusal path), making
the substitution provably injective. Tests added for a real NUL alone
(preserved), a pre-existing U+E000 alone (refused, not corrupted), and both
together (refused, not merged into one byte).

Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead
for a hand-authored row (only read inside the upstream-verbatim branch) —
exactly how Finding 2's stale docblock drifted unnoticed. Added
checkHandAuthoredTwin: every value-level export the twin DECLARES must be an
actual own property of the vendored runtime module at require-time. Tests
added, including a sensor that a declared-but-nonexistent export IS caught.

Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/
unparseableFm/reassemble preamble, copy-pasted at 7 sites in
state-transition.cts plus a sixth hand-inlined copy in state.cts's
cmdStateCompletePhase, is now one exported helper
(beginFrontmatterReassembly) every site routes through, including the
hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore)
keep a literal `body = stripFrontmatter(content)` assignment alongside the
helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward
scan (which does not chase aliases) still sees the strip; stripFrontmatter is
pure/idempotent so the extra call changes nothing observable.

Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a
separate change" claim (the 8 call sites are wired on this branch) and the
changeset's backlink from (#3473) to (#3881).

Finding 7: fixed the lint:ci failures blocking the gate — an
@typescript-eslint/only-throw-error violation from throwing a bare Symbol as
the anchor-detected signal (now a real Error subclass), unused-var warnings
left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by
two migration-specific test files (allowlisted with justification), and the
lint-state-write-path-drift false positive from Finding 5's helper (fixed
above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already
carried an explicit timeout; no change was needed there.

Golden fixture: added a golden entry for the new
anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line
scanner would also produce, since it independently dropped every quoted
top-level key). No other corpus document diverges: real .planning/ documents
carry zero anchors/aliases/merge keys/U+E000 today.

Refs #3881

* fix(#3881): fold in second-round review findings

Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration
ASCII-only key regex for the Unicode fixture; updated to require the 相
key's value now that js-yaml has no such restriction. Audited the rest of
the file for other pre-migration pins (block scalars, quoted keys,
flattened values, empty values, duplicate keys, unclosed blocks, null
bytes) by execution against real fixtures; found none regressed.

Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to
parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts
so no function still answers to the deleted hand-rolled scanner's name
(ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three
external call sites (src/commands.cts, src/runtime-artifact-conversion.cts)
updated in the same change — a mechanical rename, not an ADR-amendment
matter.

Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found
by execution. A top-level key named constructor/__proto__/toString/
valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read
resolving an inherited Object.prototype member); a key literally named
__proto__ was silently DROPPED entirely (bracket assignment on an
ordinary {} invoked the inherited __proto__ setter instead of creating a
data property). Fixed by building every parsed Frontmatter object with
Object.create(null), and replacing an `in` check with hasOwnProperty.call
in propagateCommentChannel. Added round-trip tests for all five hostile
keys, each with its own leading comment.

Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed
full byte-stability across the migration. Verified by execution: BEL/NUL/
NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw-
literal forms. Proved round-trip equivalence (each escape re-parses to the
exact source codepoint) and corrected the docstring. Found and fixed a
related real defect while verifying: a lone UTF-16 surrogate was emitted
BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely
unparseable YAML that silently collapsed to {} on re-read — extended
scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped
path.

Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real
truncation shapes (unquoted colon, open flow collection, mis-indented
sibling key, refused anchor). Root cause: the mark-based prefix recovery
excluded the very line whose key needed counting, and a mark-less refusal
never entered the recovery branch at all. Fixed by taking the max of two
lower bounds: the longest parser-verified line-prefix, and a raw-text
count of key-shaped lines (reusing the same key-shape pattern this file
already uses for isFrontmatterShaped). Extended test-matrix row A5
table-driven over all 4 regressed shapes.

Finding 6: the design doc's claim that no test owned the #3594 adversarial
fixture corpus was false — consolidation epic #1969 had already folded it
into tests/frontmatter.test.cjs. An earlier commit on this branch
re-created a standalone duplicate under that false premise; folded its
genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures,
block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the
duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the
ADR's §8.1 note, including the roadmap-sibling claim (no such file exists).

Finding 7: the golden serializer sorted object keys, making it structurally
blind to the key-order-parity invariant ADR-3473 §8.1 actually claims.
Made it order-preserving and regenerated the golden fixture from a
standalone compile of the legacy (pre-#3881) parser at ddde001af; the
current parser matches it with zero undocumented divergences, confirming
key-order parity genuinely holds. Extended row A2 table-driven across 6 of
the remaining 7 transitionCore kinds (all pass) plus documented, by
execution, a newly-discovered 8th-site regression: state.cts's
cmdStateCompletePhase calls the same preservation helper but its result is
clobbered by a later unconditional resync — filed as a distinct finding
rather than fixed here (touches syncAndPreserveStateMd, outside this
change's verified scope).

Refs #3881

* fix(#3881): preserve unparseable frontmatter through the CLI write path

Characterization (executed, before/after shown): case (b), not (a). The
frontmatter FENCE survives — `state complete-phase` on a conflict-marked
STATE.md returns success and a well-formed, freshly-derived frontmatter
block, not a document with no frontmatter at all. But the block's actual
content (the merge-conflict markers, and with them any signal to a human
that the document was in conflict) is silently discarded and replaced.

Root cause was two clobber sites, not one:

1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved
   `transformedContent` from readModifyWriteStateMd, finds {} + the
   FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh
   frontmatter block from the body anyway.
2. Even after (1) is fixed, applyPostSyncPreservation's own
   postFm/applyStatePreservation/authoritativeFm-reassertion machinery
   re-extracts frontmatter from syncedContent, restores curated fields
   from the pre-write snapshot, and reconstructs a NEW block again —
   confirmed live via `state begin-phase`, which still lost the markers
   after fixing (1) alone.

Both are now guarded by the same predicate (isUnparseableFrontmatter,
checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not
parse and the caller is not on ADR-3408 §8.3's closed "body wins" list,
both functions return their input content unchanged rather than
re-deriving over it. The closed list (cmdStateSync #905,
/gsd-health --repair's REGENERATE_STATE, both routed only through
writeStateMd, which never reaches applyPostSyncPreservation and passes
sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is
untouched — neither widened nor narrowed; verified by execution that
`state sync` still overwrites the conflict-marked block exactly as before.

Other verbs sharing the same readModifyWriteStateMd path were checked and
were equally affected before this fix: state update, query state.patch,
and state begin-phase all lost the conflict markers (RED, shown by
execution), and all three now preserve them (GREEN). Covered table-driven
in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe
block, which drives the real CLI verbs via runGsdTools — not just the pure
transitionCore layer the earlier A2 rows exercised — plus a control
asserting state sync's body-wins contract is unchanged.

Refs #3881

* fix(#3881): restore the parse surface's prototype and fix remote-runner failures

Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched.

Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write.

Refs #3881

* chore(#3881): backfill changeset PR number

Refs #3881

* test(#3881): make golden parity resistant to unrelated tree churn

A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus.

Refs #3881

* test(#3881): make the parser golden hermetic instead of tree-keyed

This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the
exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md,
docs/*.md) the prior golden pinned by tracked path. Any PR editing one of
those files' frontmatter for reasons unrelated to the parser (an
argument-hint addition, an allowed-tools tweak) turned the suite red, and the
reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to
catch a real regression. Excluding .changeset/** was not enough; the design
itself was wrong: a regression fixture must not be keyed to mutable repo
paths, and a single 376-entry JSON every such PR touches is also a
guaranteed merge-conflict surface.

Rebuilt the fixture to carry its own documents: each of 51 entries stores a
stable id, literal documentText (shrunk from a real ddde001af-era corpus
document), and an expectedParse captured independently from the
pre-migration legacy parser (git show ddde001af:src/frontmatter.cts,
compiled standalone against its byte-identical sibling modules). The test
reads no tracked path, shells out to no git command, and enumerates no tree
-- a PR editing commands/gsd/help.md cannot affect it. Every entry's
reconstruction was verified at capture time to reproduce both the current
and legacy parser's output on the original document; 0 of 51 candidates
were dropped by that check (1, the deliberately-unterminated
unclosed-block.md adversarial fixture, has no closing fence to truncate at
and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now
diverges:true entries) and the D2 order-preserving structural serializer
that keeps the comparison from passing vacuously; dropped the
tree-enumeration helpers, the coverage floor, the post-capture-skip logic,
and the vanished-file check -- all artifacts of the path-keyed design.

Refs #3881

* fix(#3881): resolve vendored-deps paths independently of cwd shape

Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on
windows-latest CI: the test passed absolute scratch-file paths into
compareFiles()/checkRow(), whose helpers joined every input onto ROOT
via path.join(ROOT, rel), producing garbage when the input was already
absolute. It surfaced on windows-latest specifically because GitHub's
Windows runners checkout the repo on a different drive than TEMP, so
path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged
(no relative traversal is representable across drives) rather than the
relative form the test assumed. The remote gsd-test runner this repo
gates pushes on is Linux-only and could never have caught this;
GitHub CI's windows-latest job is the only signal that does, and it did.

Fixed the helper itself (scripts/lint-vendored-deps.cjs's new
resolvePath()) to treat an already-absolute input as absolute-in,
absolute-out instead of silently mis-joining it, and updated the test
to pass the scratch file's absolute path directly rather than relying
on a relative conversion that is not always representable. Kept every
mutation-sensor assertion intact and added coverage proving
resolvePath is a no-op for relative inputs and correctly passes
absolute ones through unchanged.

Refs #3881

* fix(#3881): warn when state sync regenerates over unparseable frontmatter

state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly
overwrites an unparseable frontmatter block per its 'body wins'
contract — that overwrite behavior is unchanged here. The defect was
the silence: synced:true/exit 0 gave no signal that the existing
block (including git merge-conflict markers) could not be parsed and
was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be
reported as authoritative when the derivation dropped input it could
not resolve') and §8.4 ('failure is a value').

Adds a gsd: warning — ... (#3881) line on stderr, matching the
existing #3573 precedent, and surfaces the same disclosure in the
JSON result's existing changes[] array so a machine consumer sees it
too. Exit code and synced:true are left unchanged — sync did what its
contract says.

REGENERATE_STATE (/gsd-health --repair's sibling on the same
sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally
refused by applyRepairs's dispatcher before runRepairAction ever runs
(src/health-diagnostic.cts), so it is not a live path today and is not
in scope for this fix.

Refs #3881

* fix(#3881): exit non-zero when a state command returns an error

Refs #3881

* chore(#3881): changeset for the state exit-code fix

Refs #3881

* fix(#3881): honor the documented --project-dir flag

Refs #3881

* revert(#3881): restore exit-0 result envelopes for state errors

Reverts 9638f2936 and its changeset. The change was wrong and the revert is
the correction.

This repo distinguishes two error mechanisms deliberately. error() in
src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure
path. output({error: ...}) writes a JSON result envelope to stdout and returns
normally with exit 0. The reverted commit converted 23 result-envelope sites
into hard failures, which is a different contract, not a bug fix.

tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope
contract directly -- a failing command exits 0 with a JSON error envelope and
must not publish state.json -- and the remote matrix run caught it along with
four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that
the original commit rewrote were encoding that real contract, not the bug it
claimed; they are restored.

Whether an error envelope on stdout with exit 0 is the right CLI design is a
genuine question, and it is section 8.4's rule ('failure is a value') with its
own phase. It is not something to flip inside this PR.

Refs #3881

* chore(#3881): backfill changeset PR number for the project-dir fix

Refs #3881

* test(#3881): keep the frontmatter mutation shard inside its time budget

The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard
cap. Root cause is NOT row-level spawn overhead (contrast the #2790/
core-utils precedent): the three shard test files' own logic runs in
~413ms total (356+30+27ms) with all 392 assertions passing. Instead,
src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to
the vendored YAML parser, proportionally growing the mutant count Stryker
generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner
bills the full 'node --test <3 files>' invocation once per mutant, and
node:test's default per-file process isolation forks a child process for
each of the three files on every one of those invocations — pure fork
overhead multiplied by a much larger mutant population.

Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares
isolation: 'none', and .github/workflows/mutation.yml passes
--test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e.
unchanged behavior — for the other 8 shards, which were not individually
audited for cross-file state leakage under shared-process execution).
Measured locally via node:test's run() API on the exact 3-file set:
isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same
392 passing assertions. The true CI-shard number can only be confirmed
on the GitHub Actions run (Stryker cannot run locally, and 'node --test'
is hard-blocked in this environment).

Refs #3881

* test(#3881): register the vendored-parser tests in the frontmatter mutation shard

stryker.config.mjs's own rule ("Keep this list in sync with the tests
arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew
src/frontmatter.cts from ~825 to 1496 lines but its new tests
(tests/feat-3881-yaml-parser-consequences.test.cjs,
tests/frontmatter-golden-parity.test.cjs,
tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in
tests/frontmatter.test.cjs) were never added to the frontmatter shard's
tests array, so Stryker's mutants in the new vendored-js-yaml adapter had
nothing constraining them. PR #3888 measured 55.8% against the 65 floor
(748 killed / 593 survived / 17 timeout) and the shard was separately
cancelled at 15m04s against the 15-minute per-shard cap.

Registers all four files (each earns its slot on evidence of a unique
constraining assertion, documented inline), gives the shard a
measured/projected 180-minute budget via a new per-module
timeoutMinutes field threaded through mutation.yml's job-level
timeout-minutes the same way isolation is threaded, and removes the
prior isolation:'none' override (re-measured at this file-set size, its
savings are within run-to-run noise, not worth the unaudited
cross-file-state-leakage risk).

Refs #3881

* feat(#3881): derive the mutation test list and ratchet the score floor

Refs #3881

* test(#3881): ratchet five stale mutation floors and close the frontmatter gap

Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1):
config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78,
context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both
scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs
RATCHET_BASELINE in the same diff per the ratchet's own contract.

Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in
tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented
near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order
semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's
leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter),
repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via
extractFrontmatter), and the null-byte sentinel round-trip surviving at region
offset 1. Did not lower minScore.

Refs #3881

* test(#3881): decouple the ratchet test from real module floors

The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour.

Refs #3881

* refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency

Refs #3881

* fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs

A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives.

Closes #2571
Refs #3881

---------

Co-authored-by: sim <sim@local>
2026-08-26 19:29:32 -04:00
Tom Boucher
a3d5841117 fix(#3714): deliver an explicitly pinned model to the Codex worktree executor, and drop an unusable one (#3891)
* test(#3714): failing-first coverage for the dropped Codex worktree model override

Pins the argv contract for the orchestrator-worktree process dispatch: an explicit
model_overrides pin must reach the child as --model, while an unpinned, empty,
inherit, or profile-only configuration must emit no flag at all.

Five of the eight matrix rows are CONTROLS that pass before the fix. They carry the
weight here because this is an over-emission bug waiting to happen: resolve-model
returns 'sonnet' for the unpinned, empty and profile-only cases, so a fix that
threads its return value into argv would satisfy the positive row and emit
--model sonnet to Codex on every unpinned install -- the documented 400 that
ADR-2313 exists to prevent. The controls are what separate the correct fix from
the obvious one.

* fix(#3714): deliver an explicitly pinned model to the Codex worktree executor

resolveOrchestratorExec had no model input at all -- the descriptor carried only
command/args/cwdFlag/promptFlag -- so a resolved override had no way to reach the
spawned process even in principle. The baked gsd-executor.toml could not compensate
because this path spawns a process rather than dispatching a named agent.

Three parts, and the third is the load-bearing one.

The descriptor gains modelFlag (codex: --model), keeping the per-host knowledge as
descriptor data exactly as cwdFlag and promptFlag already are, so the scheduler
grows no per-host branch. No other runtime declares it.

The seam appends [modelFlag, model] and stays MECHANICAL: it does not know the
inherit sentinel, does not know which models Codex rejects, and reads no config.
Its sibling codex-agent-toml states that rule outright -- callers decide what to
strip. Argv order is baseArgs, model, cwd, prompt so the prompt remains the final
positional token. Omitting the model is byte-identical to before. A model starting
with '-' now fails closed as unsafe_leading_dash_model, the same hazard the prompt
and cwd guards already reject and which was silently accepted before.

The policy lives at the caller and passes ONLY an explicit, non-sentinel per-agent
pin. Passing null as the runtime resolver is what keeps profile and tier derived
models out of argv, which is what Codex's session-only model posture requires: the
model resolver returns 'sonnet' for the unpinned, empty and profile-only cases, and
emitting that revives the documented 400 from #2310/#2311 that ADR-2313 removed. So
the gate is the presence of an explicit pin, never that a value came back.

* fix(#3714): enforce the real-Codex value policy the sibling surface already applies

Review found one root cause behind a BLOCKER, two MAJORs, two MINORs and an
argv-injection finding: the dispatch path gated on the PRESENCE of an explicit pin
but never applied the VALUE policy that generateCodexAgentToml already applies to
the same config key. Textbook generative divergence -- and the parity row I wrote
tested resolver parity, not this policy, so it could never have caught it.

BLOCKER: a global ~/.gsd/defaults.json model_overrides.gsd-executor of 'sonnet',
'opus' or 'claude-sonnet-4-5' reached codex exec --model verbatim. That is the
documented 400 from #2310/#2311, arriving through the explicit-pin door rather
than the tier door. The issue asks for an explicit REAL-CODEX pin; real-Codex was
unenforced. Now dropped with a warning via the isAnthropicFlavoredModel predicate
#3241 single-sourced for exactly this reason.

Also: values are trimmed, so a whitespace-only pin is blank rather than
--model "   "; 'inherit' is matched case- and whitespace-insensitively, so
'Inherit' and ' inherit ' no longer reach the wire; and a value outside a model-id
charset is dropped with a warning. That last one closes the injection surface --
.planning/config.json travels with a clone, and values like
'gpt-5 -c approval_policy=never' or a command-substitution value previously reached argv verbatim,
where the spawner is an agent writing bash.

Every rejection DROPS AND WARNS rather than failing closed. An unusable exec is not
degraded, it is fatal: the dispatch step halts the wave after the worktree already
exists, so a config typo would have aborted execute-phase. The stale comment
claiming it degrades to sequential is corrected.

Separately, the seam's own empty-model handling contradicted the committed contract
and failed four tests on the remote runner. An absent, null or empty model is not an
error -- it means use the host default, the same degradation cwdFlag:null already
expresses. Unlike a prompt, where empty is a hang rather than a degraded run, so
that one stays fail-closed. Non-string values still fail closed.

Tests: the CLI rows were not HOME-hermetic and read the developer's real
~/.gsd/defaults.json, which is precisely the file the BLOCKER is about; HOME and
USERPROFILE are now sandboxed per call. Adds the global-pin regression, the
injection shapes, the case-variant inherit rows, and a real cross-surface
divergence guard.

* fix(#3714): close the flag-shaped pin, the case-variant alias, and the warning sink

Round two of review found three more defects, two of which both engines reached
independently, and all three were mine.

The charset allowlist put the dash INSIDE the character class, so a value made only
of allowed characters passed the pin policy silently and then tripped the seam's
leading-dash guard, producing exec:null. That is the wave-fatal path the whole
drop-and-warn design exists to avoid, reachable from a committed config file: -c,
--config, -p and --dangerously-skip-permissions all reproduced it. It also regressed
hosts with no model flag at all, where a dash pin turned a previously working
kimi-code dispatch into exec:null. The first character is now anchored, so a
flag-shaped value is dropped and warned like every other rejection, and the comment
that claimed this path was unreachable is corrected.

isAnthropicFlavoredModel folded case on its substring arm but not on its alias-set
arm, so SONNET, Sonnet, OPUS and HAIKU all reached Codex argv while lowercase sonnet
was correctly dropped -- the same 400 the drop exists to prevent. The predicate is
the one #3241 single-sourced so these surfaces cannot diverge, so folding case there
fixes the install-side .toml surface too.

The warning wrote the rejected value RAW to stderr. Every value that fails the
charset test contains by definition the characters the charset excludes, so it was a
guaranteed-reachable raw-to-terminal sink: an OSC sequence in a committed config
reached the operator's terminal byte for byte, and truncation could sever an escape
before its reset. The value is now sanitized before truncation.

Also from review: the allowlist rejected Vertex version pins like text-bison@002, a
false positive on a real model id; the policy ran host-neutrally so hosts with no
model flag printed a misleading drop warning on every dispatch; the invalid_model
branch had no test at all; and the changeset disclosed only that a pin is delivered,
not that an unusable one is now dropped with a warning.

* fix(#3714): single-source the model-id charset, bound the pin, keep the flag diagnosis

Round three found no blocking findings on either engine. These are the three
correctness items left in code I added.

The charset existed TWICE -- once to accept a pin, once to render a rejected one in
the warning -- and the two copies had already drifted inside a single commit: '@'
was added to the accept class and not the render class, so a Vertex-shaped value
rejected for some other reason rendered as text-bison?002. Both are now derived
from one definition, with a parity test asserting every character the matcher
accepts survives the sanitizer unchanged, so they cannot drift again.

A pin reached argv unbounded. CLAUDE.md documents the hazard: execFileSync aborts
on Windows above 32,767 characters of argv. A model id has no reason to be long, so
a pin over 200 characters is dropped and warned rather than truncated -- a truncated
model id is a different model id. Boundary rows at 199, 200 and 201.

The first character is now required to be alphanumeric, so '@evil' and '/c' no
longer reach argv. Security rates both inert on codex today, so this is hardening
rather than a live defect; it is here because it is one character of regex and
resolveOrchestratorExec documents itself as a general descriptor-to-argv seam that
other hosts may adopt.

Tightening the anchor made the leading-dash branch unreachable and, with it,
regressed the diagnosis: '-c' began reporting 'unsafe characters' instead of
'looks like a flag/option'. The dash check now runs before the charset test, which
both restores the actionable message for the most likely user typo and keeps the
branch live. A test pins the distinction between the three rejection messages so
the branch cannot silently die again.

* test(#3714): make the charset parity guard actually guard, and remove a false-green trap

Review proved by mutation that my parity test could not do what its own comment
claimed. It bound the expected character set to a local that was assigned and
discarded, and the shared definition was not exported, so widening that definition
left the test green. A comment overstating what a test guards is worse than no
comment, because the next reader trusts it. The shared body is now exported and
asserted equal, and I watched the assertion fail against a widened definition
before keeping it.

The sanitizer regex was module-scope, carried the g flag and was exported.
Production only uses it with replace, which resets lastIndex, so there was no live
bug -- but test() on a g-flagged regex alternates between calls, so any future test
reaching for it would false-green. The g-flagged copy is now internal to the one
call site that needs it and the exported companion carries no flags, which removes
the footgun rather than documenting it.

Also notes at the shared definition that it is interpolated into both a positive and
a negated character class, so only plain characters and ranges are safe to add.

No runtime behavior changes: the accept set, the anchor, the length cap, the
sanitizer output, the rejection ordering and all three message texts are
byte-identical, re-verified through the real CLI.

* chore(#3714): backfill changeset pr number

* chore(#3714): re-trigger CI after the GitHub Actions outage

The workflow runs for this branch were created during the Actions major outage on
2026-08-26 and never got scheduled. They are wedged: GitHub reports them queued,
refuses to cancel them, and refuses to rerun them because it believes the workflow
is already running. The Tests run is also pinned to a superseded sha, so no test run
exists for the current head at all.

Actions is operational again and the repo-wide queue has drained, so a fresh push is
what creates schedulable runs. This commit is empty on purpose: nothing about the
change is being altered, and the verified content is byte-identical to 7860c6ccc.

---------

Co-authored-by: sim <sim@local>
2026-08-26 17:01:09 -04:00
Tom Boucher
c1f755d75d docs(#3889): ADR-3889 — the process-exit contract, nothing fails with success (#3890)
* docs(#3889): ADR-3889 — the process-exit contract

Records the decision for epic #3889: one Outcome vocabulary, one pure
projection to an integer, two terminators sharing both.

Doc-only. Regenerates docs/adr/README.md via gen-adr-index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3889): rewrite ADR-3889 as an exit-code registry, on measured evidence

The first draft's census was grep-derived and roughly 2x inflated: it counted
comment prose and also matched process.exitCode, the correct pattern. An AST
census (@typescript-eslint/parser, CallExpression on process.exit) gives 128
deduped call sites, 71% of them in hooks/, against 88 process.exitCode
assignments already in use. scripts/ is effectively migrated (39:1).

Corrects the framing accordingly: the seam (src/cli-exit.cts) already exists
and is already adopted; what is missing is an allocator and a hooks adapter.

Replaces the fixed six-value enum with a registry: 0 and 1 free, 2 reserved to
the Claude Code hook protocol, 3-13 forbidden (Node reserves them), 64+
allocated with one meaning and one owner per code. Domain codes permitted.

Adds three findings obtained by reading and executing the code:
- ui-safety-gate/api-coverage report empty stdin as an authoritative negative
  verdict (exit 1) — reproduced with a control
- four modules hand-roll the same three-outcome convention, documented only in
  a comment at src/ui-safety-gate.cts:137
- two capability fragments fabricate {"detected":false} on probe failure, while
  the honest {"skipped":true,"reason":"sdk-failed"} form already exists in-tree

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 16:53:38 -04:00
Tom Boucher
389bc86e02 enhance(#3883): route every slug call site through its canonical owner (#3896)
* docs(#3883): correct section 8.3 against the tree

I wrote this section in Phase 0 stating rules I had not executed against the
tree. Measured at 832dcbb75, three statements are wrong.

The slug count is 11 across 5 files, not 13; two of the thirteen were an
unrelated tokenizer regex. The divergence is real and reproduced.

resolveRuntime reads no install marker at all and has no cache; PR #3382,
cited as prior art for that rung, is closed unmerged.

The Codex sandbox is still a hand-maintained subset map with a silent
read-only fallback, and validate agents checks file presence only -- both
halves of that claim are false.

The shortFormToId rule is accurate. The guard roster names no casualty for
this rule.

Section 8.3 is therefore a work list, not a conformance check.

Refs #3883

* test(#3883): failing-first rows for slug re-implementation divergence

Refs #3883

* feat(#3883): route every slug call site through its canonical owner

Migrated commands.cts:209, init.cts:176/1935/1957/3109, phase-id.cts:229/380, phase-locator.cts:269, workstream-name-policy.cts:75 to generateSlugInternal (core-utils.cts:107).

Declared different: gsd2-import.cts:97 (no truncation), active-workstream-store.cts:97 (fixed ASCII env-key domain, already matches).

Fixed a latent circular-require bug: core-utils.cts top-level destructured comparePhaseNum/scopeToPhase from phase-id.cjs; switched both sides of the new circular require to lazy function-body requires.

Refs #3883

* fix(#3883): restore per-site truncation contracts broken by the slug consolidation

Refs #3883

* fix(#3883): close review findings — changeset, miscounts, loose rows, cyclic destructure

Refs #3883

* docs(#3883): correct the same slug miscount in the Context table

Section 8.3's rule text was corrected to 11 copies across 7 files; the Context
table above it still carried the original 13 across 5. Same wrong claim, second
location, both mine.

Refs #3883

* chore(#3883): backfill changeset PR number

Refs #3883

* test(#3883): keep the core-utils mutation shard inside its time budget

PR #3896's Stryker (core-utils) shard was cancelled at the 15-minute
shard cap. Stryker's command-runner bills one `node --test <file>`
invocation as a single unit costing whatever the file's slowest run
costs, re-run once per mutant (documented in
tests/state-contract.test.cjs's header, #2790 precedent). A3/A5 drove
every CLI-reachable slug site through runGsdTools (a real child-process
spawn per call, ~85-170ms each across ~70 calls), accounting for ~7.1s
of the file's ~7.9s wall time.

commands.cts:cmdGenerateSlug and init.cts:cmdInitExecutePhase/
cmdInitPhaseOp/cmdInitProgress are plain functions reachable in-process
from the built gsd-core/bin/lib/*.cjs, so this calls them directly
instead of spawning gsd-tools, capturing their fd-1 JSON output with the
bug #1008 fs.writeSync-mock pattern already used in tests/io.test.cjs
and tests/init.test.cjs. A hermetic-env helper reproduces the isolation
runGsdTools's { HOME: tmpDir } + testEnvBase() gave the child process.

File wall time drops from ~7.97s to ~1.1s (246/246 passing, same
coverage), well inside the 15-minute cap even at hundreds of mutants.

Refs #3883

---------

Co-authored-by: sim <sim@local>
2026-08-26 16:03:11 -04:00
Tom Boucher
a638ca4332 enhance(#3882): stop sentinel phases skewing estimation calibration (#3893)
* test(#3882): failing-first rows for sentinel phases skewing calibration

Adds A1a/A1b/A2/A3 to tests/estimate-calibrate.test.cjs, the module's
existing test file, rather than a new bug-NNNN file. collectCalibrationSamples
(src/estimate-cli.cts:206) does a raw readdirSync over .planning/phases and
never applies isSentinelPhaseId, so a sentinel phase (milestone 0 or 999)
carrying a PLAN estimate / SUMMARY actuals pair contributes a phantom
calibration sample.

computeCalibration is median-based, so a single 50x outlier among three
samples leaves the factor unmoved — asserting "the factor is unchanged"
against one sentinel would pass on the broken code for the wrong reason.
Each row instead asserts the WHOLE computed CalibrationResult object
(factor, applied, confidence, sampleCount, clamped) for a sentinel-free
project against its sentinel-injected twin:

- A1a: one sentinel flips applied false->true and confidence low->med on
  phantom evidence (calibration switches on with zero real signal).
- A1b: two sentinels corrupt the factor itself (1 -> 3, clamped false->true).
- A2: the sentinel's own sample is verified absent from the returned list.
- A3: the two genuine phases still contribute their own unchanged samples
  (regression pin — stops A1/A2 passing by filtering everything).

Verified RED on today's code (node tests/estimate-calibrate.test.cjs):
A1a/A1b/A2 fail with the exact differing objects; A3 and all pre-existing
rows in the file remain green (no collateral).

Refs #3882

* feat(#3882): route phase enumeration through its owner and name the sentinel axis

Task 1: collectCalibrationSamples (src/estimate-cli.cts) hand-rolled a raw readdirSync over .planning/phases, treating every directory (including sentinel phases, milestone 0/999) as a completed phase and feeding phantom PLAN/SUMMARY samples into the estimation calibration factor. Routed through the existing owner, listMilestonePhaseDirs(phasesRoot) with no cwd -- already 'all milestones, sentinels excluded', exactly the combination this caller needs; no new API was required for this half. It now also surfaces the scope discriminator: an unreadable phases directory throws PhasesUnreadableError instead of silently returning zero samples, and cmdEstimateCalibrate reports it via a new ERROR_REASON.ESTIMATE_PHASES_UNREADABLE instead of persisting a phantom empty calibration document.

Task 2: added listAllPhaseDirs(phasesDir, { includeSentinels }) to src/phase-locator.cts -- the one genuinely missing axis: 'physical set, sentinels INCLUDED'. includeSentinels has no default and is required, so a call site cannot obtain sentinel-inclusion by omission (compile-time refusal, not just documentation). Mirrors listMilestonePhaseDirs's absent/unreadable scope handling.

Task 3: migrated the two exemptions whose written reason maps cleanly onto 'physical set, sentinels included' -- cmdRoadmapAnalyze's _phaseDirNames (src/roadmap.cts) and cmdInitMilestoneOp's diskPhaseDirs (src/init.cts), both heading->directory lookup indexes. Left the rest: archivePhaseDirectories's own body has no readdirSync to migrate (its callers already resolve dirs before calling it, and both current callers deliberately EXCLUDE sentinels -- migrating it would be an unauthorized behavior change, not an API swap); cmdValidateHealth's exemption is vestigial (its actual physical-set sweep already lives in planning-snapshot.cts's buildAllPhaseDirNamesField, a pre-existing near-duplicate of the new axis, flagged as a finding, not restructured); cmdPhasesClear/cmdMilestoneComplete/cmdVerifySchemaDrift/detectHasPriorPhases/detectUiPhaseActive want a different combination (sentinels excluded, or a single-phase lookup) and are unaffected.

Task 4: detector 2 (sentinel literal) is untouched and retained. Removed exemption entries only for the two migrated call sites; every other function-scoped exemption is preserved. Guard exits 0.

Refs #3882

* refactor(#3882): delegate the snapshot phase-dir scan to its owner

buildAllPhaseDirNamesField duplicated listAllPhaseDirs's own
readdirSync + directory-filter + absent/unreadable handling — the
'one implementation per rule' defect ADR-3473 SS8.3 names, introduced
by this branch's own #3882 work. Delegate to listAllPhaseDirs and
re-apply the field's existing lexicographic sort on top, since W007's
observable order must not change.

Refs #3882

* docs(#3882): document the sentinel axis and the enumeration consolidation

Records listAllPhaseDirs in the Phase Locator glossary entry, and the fact
that the owner already answers the all-milestones sentinel-free question when
called without a cwd -- the call collectCalibrationSamples was missing.

Also notes that buildAllPhaseDirNamesField now delegates rather than carrying a
second readdir, and that exactly one readdirSync over the phases directory
remains across the two modules.

Refs #3882

* test(#3882): close review findings — real order proof, unreadable coverage, collision fixtures

Refs #3882

* chore(#3882): backfill changeset PR number

Refs #3882

---------

Co-authored-by: sim <sim@local>
2026-08-26 15:03:12 -04:00
Tom Boucher
832dcbb751 fix(#3707): surface UAT rows audit-uat silently dropped, and never report a clean result for a file it could not read (#3887)
* test(#3707): failing-first coverage for the three parseUatItems false negatives

Nine tests that must be red and three controls that must already be green.

The controls are the point of the split. `result: pass` staying unsurfaced is
what stops the fix inverting the filter so eagerly that every passing test
becomes an outstanding item, and the classic single-line shape is the no-churn
control for rewriting the adjacency regex. Both were confirmed green against
the current build before being written down; a control that is red today would
be a second bug, not a control.

Each failing fixture was run through the built parser first and returns []
for its stated cause — the issue row matched then filtered, the block-scalar
and wrapped rows never matched at all, the all-unparseable file vanishing whole.
That evidence is in 50-test-matrix.md rather than asserted.

Tests target ../gsd-core/bin/lib/uat.cjs, the built live module, and drive the
real CLI through runGsdTools. #3706 lost a full RED/GREEN cycle to tests that
imported a different copy of the function under test, so the import target was
verified before anything was written.

* fix(#3707): stop parseUatItems dropping outstanding UAT rows

Three independent false negatives, all in the audit path, plus one the issue
did not mention.

The matcher no longer requires `expected:` and `result:` to be adjacent single
lines. It slices each `### N.` block to the next heading and reads the first
`result:` line within it, taking `expected:` from parseExpectedFromTestBlock —
the seam that already parsed both the block-scalar and inline forms correctly
and was sitting unused two hundred lines away. Two parsers in one module read
the same field with different grammars; now there is one.

The result filter is inverted from an inclusion list of three to an exclusion
of a minimal PASS set. This was the issue's one open design question, which the
reporter explicitly declined to answer for the maintainer; it was asked and
decided deliberately. The fail-safe direction is what parseGapsItems documents
seventy lines below for this same false-negative class (#2286): a token nobody
recognised surfaces rather than vanishing. The trade is a visible, correctable
false positive if a project invents a novel pass-word, against today's silent
and invisible drop.

`issue` also needed a category. It is template-sanctioned with its own `issues:`
counter, but categorizeItem fell through to `unknown` — surfacing it in the
wrong bucket would have been a half-fix.

Finally, a file parsing to zero items no longer vanishes with its frontmatter
`status:`. One with a non-terminal status is reported with `parse_gap: true`,
so the reader gets a cue to look; a `complete` one stays omitted as before.
That is what made the first two defects dangerous rather than merely lossy —
the audit omitted the phase instead of under-counting it.

* fix(#3707): close the review blockers, including a regression I introduced

The remote suite was RED on the previous commit and both reviews found real
defects. Everything below was verified by execution, not by reading.

I introduced a regression against origin/next. The rewritten result matcher was
END-anchored where the old one was not, so `result: pending (blocked on
staging)`, `result: [skipped] # no device` and `result: blocked - waiting` all
returned a row before this branch and returned nothing on it — me reproducing
the exact defect class this issue exists to kill, in the fix for it. The anchor
is gone and each shape has a regression test; trailing text now falls back to
`reason` when the block has none.

`parse_gap` was inferred from the wrong signal. It fired for ANY zero-item file
whose status was not `complete`, which asserted something false about a
perfectly-parsed all-pass file, swept in archived phases left at `testing`, and
is what turned the #2286 Gaps tests red — a control this change was supposed to
keep green. It now derives from headings SEEN BUT UNYIELDED, reported by a new
parseUatItemsWithStats, so an all-pass file and a Gaps-only file are not parse
gaps and a file whose blocks carry no `result:` line is.

The fix was also invisible end to end, which both reviewers caught
independently. parse_gap entries carry no items, and both audit-uat.md and
progress.md gate on `total_items === 0` — so the headline symptom, the phase
vanishing, still reproduced for a user and only the raw JSON had changed. There
is now a `parse_gap_files` counter and both workflows gate and report on it.

Also: categorizeItem compared case-sensitively while the new PASS check
lowercased, so `result: PENDING` surfaced as `unknown`; blocks are bounded at
the next heading of any level, so a trailing `## Gaps` entry no longer bleeds
its `reason` onto the preceding test; dead unreachable fallbacks removed; and
the all-pass control was strengthened, since asserting only `total_items === 0`
let it stay green through the bogus parse_gap entry.

* chore(#3707): acknowledge the workflow growth the fix required

The emitted-attribution guard went red because audit-uat.md and progress.md
grew, and it is right to ask: runtime-loaded workflow prose is the product,
so growth there is a real change to what an executing agent reads.

The growth is not incidental to this fix, it IS the fix reaching a user. Both
reviewers found independently that emitting `parse_gap` in the JSON changed
nothing observable, because both workflows gated their output on
`total_items === 0` and parse-gap entries carry no items — so a phase whose
rows could not be parsed still printed "All Clear" and still vanished from the
progress report. The widened gates and the branches that name the unparsed
files with their phase and path are what close that.

Acks exactly the two paths the guard reported, keyed on the bare filename. The
three spent acknowledgments it also listed are inert by its own description —
the base already absorbs them — so they are left alone rather than swept up
here, where they would just add unrelated churn to this diff.

* fix(#3707): close the mixed-file blocker and the second false-clean surface

The suite was GREEN and the isolated review still found a blocker, which is
the useful part: none of this was covered by a test.

A MIXED file dropped its unparseable rows silently. `parse_gap` sat behind an
`else if` on `items.length > 0`, so one parseable row was enough to discard
`headingsSeen` entirely — a file with one pending row and two unreadable blocks
reported one item and no gap. That is the exact class this issue exists to kill,
reappearing inside its own fix for the third time. The flag is now set
independently of item count and the entry carries `unparsed_blocks`, so the
count is quantified rather than merely flagged.

A `result:` inside a fenced code block was being read as real, so a PASSING
test could be reported as outstanding from a value in a code sample — another
regression against origin/next, whose adjacency regex ignored it. Field scans
now run against a fence-stripped copy while `expected:` still reads the raw
block, since a block scalar may legitimately contain fenced-looking text.

The workflow report was still unreachable whenever anything else was
outstanding: the unparsed table lived in the all-clear branch, so a project with
one pending row in phase 01 and an unreadable phase 02 rendered phase 02
nowhere. It now fires on `parse_gap_files > 0` from the `present` step.

planning-inspect was the second surface making a false-clean claim — for
exactly the files audit-uat now flags it emitted `scope: 'complete'` with an
empty unresolved list, positively asserting completeness over a file it could
not read. It consumes the stats now and reports SCOPE.TRUNCATED with a
`uat_unreadable` diagnostic, reusing the vocabulary already used two lines above
for an unreadable file rather than inventing a token.

Also: headings with no name no longer vanish whole; the trailing-text-to-reason
synthesis I had added is removed, since it was never required by the blocker and
silently changed categorization for `result: [skipped] # no device`; the emitted
`result` token is normalized to lower case so it agrees with `category` in a
published contract; and an O(n^2) indexOf is gone from the heading loop.

* fix(#3707): stop rows stealing each other's fields, on all three surfaces

The suite was green when the security review found these. Two are
blocker-severity and one of them is a direct hit on my own verification.

A `### N.` line indented two spaces inside an `expected: |` scalar is a valid
ATX heading, so it became a phantom row that STOLE the real row's result token
while the real row vanished. I had probed this shape and declared it fixed — my
probe asserted the item COUNT and the result token, both of which the phantom
satisfied, so it passed for exactly the reason it should have failed. Block
scalar bodies are now masked to blank lines (line count preserved, so offsets
still line up) before headings are tokenized, and the tests assert row IDENTITY
— number and name — not presence.

Feeding parseExpectedFromTestBlock the raw slice let one row publish another
row's `expected:` from inside a fence the stripped view had correctly excluded.
Blocks handed to it are now clipped at the first fence opener. This was not
cosmetic on the render-checkpoint path: a checkpoint banner a HUMAN reads and
answers was rendering a different row's expected text.

A balanced fence pair straddling a test block made that heading invisible, so
an outstanding row disappeared with no item, no gap and no count — the exact
false-clean this issue exists to close, and a regression against origin/next.
Suppressed `### N.` lines now count toward headingsSeen so the file is flagged.

An unterminated fence swallowed the rest of the document including `## Gaps`,
producing a whole-file false clean. Such a file is now treated as a parse gap,
following what uat-predicate already does.

Found and fixed inline while there: parseExpectedFromTestBlock's scalar opener
required a bare newline, so a CRLF `expected: |` fell through to the inline arm
and published `expected: "|"`, silently discarding the entire value. The same
fall-through hit `|-` and `|+`.

parseFirstPendingTest had the identical exposure on the render-checkpoint path
and now shares the same masking and clipping. Five legitimate fixtures — inline
expected, a real block scalar, CRLF, bracketed pending, and a first-pending
that is not the first test — are byte-identical before and after.

Also from the code review: the admit condition disagreed with the terminal
status guard, so a `status: complete` file with an unparseable block was
emitted as an empty entry that rendered nowhere but inflated total_files; a
control test was vacuous because its fixture filename did not match its phase
dir, so #3511 scoping meant the file was never opened — and that vacuity is why
the admit regression shipped green; an unterminated fence discarded the flag
that would have caught it; `### 1.2.3` parsed as test 1; and planning-inspect
did not share the terminal-status rule.

* test(#3707): assert what the render-checkpoint fix actually does

The suite went red on three of my own tests and the source was right — the
assertions were wrong, in a way worth naming.

One forbade the rendered checkpoint from containing `### 3. Fake Row`. But in
that fixture the string IS row 1's legitimate `expected:` block-scalar value; a
heading-shaped line inside a scalar is inert text and rendering it is correct.
The test was forbidding correct output. It now asserts row IDENTITY — the
checkpoint is for test 1 named Alpha and never test 3 named Fake Row — which is
the property that actually distinguishes the fix from the bug.

The other expected success where the correct outcome is a clean error: row 1 in
that fixture has no `expected:` of its own and only ever appeared to have one by
stealing row 2's from inside a fence. Depending on the bug to produce a pass is
how a test ends up pinning the defect. The fixture now gives row 1 its own
value and asserts the checkpoint carries it and never the fence-hidden text, and
the error path gets its own test asserting it fails cleanly without leaking.

All three were checked against the real rendered output before the assertion was
written, and each was reasoned through for whether it can fail: the identity
test breaks if a phantom row is parsed, the clipping test breaks if the raw
block is read again, and the error test would pass-not-fail under the old
stealing behavior.

* fix(#3707): correct the scalar masking frame and cover every YAML block opener

Two reviews independently found the same blocker, and it is the sharpest defect
on this issue: maskBlockScalarBodies computed line offsets in UTF-16 units but
spliced them into Array.from(content), a CODE POINT array. One emoji anywhere
earlier in the file shifted every later mask write, so the mask blanked the
wrong characters and spilled past line ends. Measured: at two astral characters
a result token truncated `pending` to `pendi` and recategorized to unknown; at
six the real row vanished; at twelve the FOLLOWING row's `result: blocked`
disappeared and the file reported clean. That is the false-clean class this
issue exists to close, reintroduced by the mitigation written to prevent it, and
defeating both new detectors at once. The mask is rebuilt line by line now,
which is frame-agnostic and length-preserving by construction.

The opener grammar was also incomplete. YAML block scalar headers take an
optional indentation indicator and an optional chomping indicator in either
order, so `|2`, `|2-`, `|-2`, `>2`, `>2+` are all valid — and none were matched.
An unmasked `expected: |2` body meant a `### N.` line inside the value became a
real heading: reproduced, row 1 disappeared and a fabricated row 2 named
"Phantom" took its identity.

Fixing that exposed a third instance of the same family, found by my own probe
rather than by review: the value extractor understood only the `|` openers, so
every `>` folded scalar published the LITERAL OPENER as its value — `expected`
came back as ">" or ">2+" and the whole scalar was discarded. The extractor now
shares the opener grammar and implements real folding, joining paragraph lines
with a space and turning a blank line into a newline, rather than pretending `>`
means `|`.

Also from the reviews: the shortfall counter scanned the masked copy but not a
fence-stripped one, so a `### N.`-shaped line inside a properly closed
documentation fence — the ordinary way to document the row format inside a UAT
file — counted as a suppressed row and flagged the file against nothing; and
clipping at the first fence discarded a legitimate `expected:` that appeared
after a closed fence, which is silent field loss.

Every opener now verified for both row identity and exact extracted value, in
LF and CRLF, alongside the emoji fixtures at 1/2/6/12.

* refactor(#3707): replace the scalar masking with a column-0 heading rule

The fix had grown to five helpers whose only job was undoing one
over-permissive rule: tokenizeHeadings treats a heading indented up to three
spaces as real, so a `### N.` inside an `expected: |` body was parsed as a row
and stole the real row's identity. Every blocker in the last three review
rounds came out of that machinery rather than the reported bug — worst of all
a UTF-16-versus-code-point frame mismatch that corrupted any document
containing an emoji.

A UAT test heading is at column 0. The shipped template puts all of them there,
no `*UAT*.md` in the repo has an indented one, and the only indented `### N.`
lines in the tree are the adversarial fixtures that must not parse. Requiring
column 0 makes a scalar-interior heading a non-heading by construction, so
maskBlockScalarBodies, indentWidthOf and BLOCK_SCALAR_OPENER_RE are gone along
with the mask-invariant test that existed only to guard them. The frame bug is
now structurally unreachable: no code-point array or offset splicing remains.

The premise was incomplete and the reviewer caught it rather than forcing it
through. Masking had been doing double duty — it also hid indented FENCE
delimiters from the tokenizer, so removing it let a two-space fence inside a
scalar body swallow a later column-0 row. The alternative on offer was to
rewrite that test to assert the row is merely counted, which is a behavior
regression dressed as a passing suite. Instead there is a small line-based pass
that blanks only indented fence delimiters — same "column 0 is structure" rule
extended consistently, no YAML knowledge, and line-based by construction so the
frame bug cannot come back. It was proven load-bearing by a negative control:
reverting the wiring reproduces the regression exactly.

Kept, because they fix defects column-0 does not touch: the fence clipping that
stops one row reading another's `expected:` from inside a fence, and the folded
scalar handling that stopped `expected: >` publishing the literal ">".

Also corrects a comment left pointing at a symbol this commit deletes.

* fix(#3707): blank neutralized fence bodies, and give both parse paths one grammar

Reviewing the simplification found two more, and the first is row theft again —
the sixth time this class has surfaced on this issue, and the second time
inside a mitigation written to stop it.

Neutralizing an indented fence blanked only its DELIMITER lines. If the block's
body held a column-0 `### N.`, un-hiding the delimiters made that line a real
heading, which then took the preceding row's fields: an `### 1. Alpha` document
came back as a single row 9 named Phantom, with Alpha gone. Neutralized blocks
are now blanked open-to-close, body included, which is the honest reading of
the intent — an indented fence inside a scalar is content, so nothing in it
should be able to produce structure. The raw block is still what the expected
extractor reads, so a legitimate `expected: |` carrying a fenced code sample
keeps its full text.

The two parse paths also disagreed about what a test row IS.
parseFirstPendingTest filtered on `^\d+\.\s+` while parseUatItemsWithStats used
`^\d+\.(?!\d)`, so `### 3.Foo` was a row when audited and not a row when
resumed. Both now share one predicate and one extractor. The extractor mattered
as much as the filter: the checkpoint path's name-mandatory pattern would have
skipped exactly the shapes the widened filter admits, so fixing the filter alone
would have moved the divergence down a line rather than closing it.

Also from the security pass: parseUatItems had become an export with no callers
and no direct test once both consumers moved to the stats form. It stays, since
deleting an exported symbol from a shipped module is a contract change and not
this issue's business, but it is now documented as the items-only wrapper and
has a test. And the PASS check lowercased a value that extraction had already
lowercased; normalization now happens once.

* fix(#3707): revert the whole-block fence blanking, and pin the rule instead

My previous commit over-corrected and the suite caught it. Blanking a
neutralized fence block open-to-close destroys content legitimately living
between the delimiters, and on an UNTERMINATED opener it blanks to EOF and
deletes every later row. No framing makes that correct, and it is what turned
two earlier tests red.

The "blocker" that prompted it was my own misreading. This change adopted the
rule that column 0 is structure and indentation is content. Under that rule an
indented fence delimiter is not a fence, so a column-0 `### 9.` sitting between
two indented delimiters genuinely IS a heading, and a `result:` after it
genuinely belongs to it. That document is malformed and the parser reading it
that way is consistent, not stealing. Nothing is silently lost either: the row
whose result was taken surfaces as the parse gap.

So the blanking is back to delimiters only, and rather than leaving the
question open, the behavior is now pinned by a test asserting the rows by
identity, with the rule stated at the site — so the next person does not
oscillate the way I just did.

Kept from the reverted commit: the shared row-heading grammar and extractor
across both parse paths, the parseUatItems wrapper documentation and its test,
and the single point of lowercasing.

* fix(#3707): scope the shortfall scan, and stop neutralized content becoming structure

The security review found a HIGH that is the earlier frame-mismatch bug wearing
different clothes. The shortfall scan compared a SECTION-scoped raw line count —
the `## Tests` body — against a DOCUMENT-wide token count. So a single legal
`### N.` row anywhere outside `## Tests` decremented the shortfall and switched
the fence-straddle detector off: an identical `## Tests` section went from
`headingsSeen 1, parse_gap true` to `headingsSeen 0, no gap` purely because a
`## Prior` section existed. A document whose rows render as ordinary blocked rows
in any CommonMark renderer audited as totally clean. Both sides of the
comparison now come from the same surface, by filtering tokens to the scan
span rather than re-tokenizing, so there is no second offset basis to keep in
step.

Neutralizing a fence could also promote its former CONTENT into structure: a
column-0 delimiter run inside an indented pair became an opener once the
enclosing delimiters were blanked, hiding every later heading to EOF. Column-0
delimiter-shaped lines inside a neutralized block are now blanked too — and only
those, so a column-0 heading between neutralized delimiters is still a heading
(the pinned behaviour) and the field lines of a row living between two scalars
still survive. The reviewer corrected my repro while fixing it: an even number of
inner runs re-pairs and hides nothing, so the live shape needs an odd one, and
both are now tests.

Two more from that pass. A legal scalar header carrying a trailing comment
(`expected: | # sample`) failed the end-anchored grammar, publishing the literal
header and raising a false gap. And the indented-row counter walked backwards
per row: 3.6 seconds at sixteen thousand rows, now 15ms, via one forward pass —
though the reviewer also established my example was not the quadratic shape,
which needs an uninterrupted scalar body.

Carried in from the previous round: the indented-row counter keys on any block
scalar rather than only `expected:`, so a template-sanctioned `reported: |`
holding user prose with a heading-shaped line no longer raises a false gap; and
`reason:`/`blocked_by:` read block scalars through the same shared extractor
instead of publishing the literal `"|"`, which also means a multi-line reason
can finally reach categorizeItem — a `reason:` mentioning a server now
categorizes as server_blocked, which was impossible while the value was thrown
away.

* fix(#3707): make the shortfall scan whole-document on both sides

Second HIGH in this area, and the diagnosis is the useful part: I closed the
first one by making the two sides agree, but I did it by NARROWING the token
side to the `## Tests` span while the parse side stayed whole-document. Rows
outside that section are still parsed and surfaced when visible, so when a fence
straddled one it fell through both sides of the comparison — no item, no gap,
file never entered the results at all. A `## Regression Tests` section, or a
second `## Tests` (collectSection takes the first), audited as totally clean
while origin/next surfaced those rows.

Both sides are whole-document now. Symmetry is the property that matters here;
every attempt to be clever about which scope to compare has produced one of
these, twice at HIGH severity.

That reinstates a known over-report, deliberately: a `### N.`-shaped line inside
a closed fence in a `## Notes` section — the ordinary way to document the row
format — counts as a suppressed row and raises a gap on a file with nothing
missing. Noisy, but visible and fail-safe, against two silent false-cleans on
the other side of the trade. This issue exists to eliminate false cleans, so the
trade goes that way, and the reasoning is written at the site so it does not get
optimized back.

Three existing tests encoded the retired scoping and are replaced rather than
worked around: two now assert the accepted over-report, and one asserting a
4-space row is "not counted" was already contradicted by widening the counter to
any indentation — refusing to PARSE a 4-space heading is right, refusing to
COUNT it reopened the hole the counter exists to close.

Also in this commit, from the same review round: the inner-delimiter sweep tested
a column-0-anchored pattern, so an INDENTED delimiter inside a neutralized block
was still promoted to structure and lost a row; it is indent-tolerant now.

A refinement was identified and deliberately not taken — keying the
documentation-sample exemption on the fence info string rather than on section
scope. It is content-based and symmetric, so it would not reintroduce the
asymmetry, but it belongs in its own change rather than riding this one.

* docs(#3707): correct two claims in the over-report justification

Both from review, both comment-only, and both matter because they would
mislead the next person into "fixing" something correct.

The over-report note called the triggering shape "the ordinary way to document
the row format". It is narrower than that: the scan requires literal digits, so
the conventional placeholder `### N. Name` does not trigger it at all — only a
sample written with real numbers does, and no phase UAT file in-tree has one,
only the shipped template, which selectPhaseUatFiles never scans. A maintainer
who tested the documented placeholder form would find no over-report and could
reasonably conclude the pin was stale. That is now stated, and it also makes the
trade look better than I claimed: the real-world frequency is lower.

The attribution guard is described as structural rather than positional. It is
positional in one respect: the walk stops at the nearest column-0 line, so a
block scalar nested inside a `## Gaps` bullet is transparent to it and a
heading-shaped line in that value gets counted. Same accepted over-report,
reached by a path the comment did not mention — recorded so it is not later
mistaken for a new defect.

* fix(#3707): a complete status no longer switches off the parse-gap detector

The security review named this as the last silent-clean path in the change, and
its phrase is the right one: a self-declared kill switch over the very detector
this issue built. A file whose frontmatter said `status: complete` was omitted
unconditionally, so one containing a fence-straddled `result: blocked` computed
headingsSeen = 1 — the detector fired — and then emitted no entry at all. The
audit reported nothing.

The predicate is now status-independent: a file is surfaced when blocks were
seen but yielded nothing, whatever it claims about itself. A terminal status is
an assertion by the author, and an assertion is exactly what must not be allowed
to suppress the signal that would contradict it. What does not change is the
thing the status is actually for — a complete file with nothing to parse, and a
complete file whose rows all parse and all pass, both stay silent, verified
through the real CLI.

I replaced a control test of mine, and it is worth saying why that is not a
weakening: its name was already false. "A zero-item file with a complete status
is still omitted" used a fixture with a `### 1.` block carrying no `result:`
line, so headingsSeen was 1 — it was never a zero-item file, it was the kill
switch itself, pinned. The intent it claimed is now covered by two stricter
tests, one for a file with no blocks and one for a file where every row parses
and passes, each asserting both that no entry exists and that no items are
counted, where the old test asserted only the former.

Everything else is byte-identical: 61 regression cases and the non-complete
equivalents of all four shapes produce exactly the same output as before, with
the delta confined to the two cells this change is meant to move.

* fix(#3707): close the moved kill switch, and keep the archive out of the live gate

Both reviewers independently found that closing the kill switch on one surface
left it standing on the other. cmdAuditUat dropped the terminal-status guard,
but buildUatRows in planning-inspect kept it — and its comment justified that by
claiming to mirror a guard cmdAuditUat no longer had. One byte-identical file
with `status: complete` and a fence-straddled `result: blocked` reported
parse_gap through audit-uat while planning-inspect published
`uat.scope: "complete"` with no diagnostic at all. That is the repo's own
generative-fix-divergence class, and no test pinned that arm, which is why it
survived. The clause and the false comment are gone and the arm now has tests.

Removing it exposed a MAJOR the security pass had not reached: archived phase
dirs are deliberately not milestone-filtered and archived UAT files are
`complete` by definition, so status-independence newly admitted the entire
project archive. One live pending row plus four signed-off milestones produced
parse_gap_files 4 — and since progress.md gates Verification Debt on that
counter, a mature project would have warned on every run, forever, about closed
history no user action can clear. Warning fatigue that buries the next real gap
is the feature defeating itself.

So the counter is split rather than suppressed: `parse_gap_files` counts live
phases only and remains the gate, `archived_parse_gap_files` carries the rest,
and every archived entry stays in `results` with its parse_gap and its
milestone. Nothing became silent; the live signal stayed actionable. Both
workflows report the archived bucket as closed history rather than as something
to act on.

The scope cascade is also decoupled, on the security reviewer's advice that it
is load-bearing here rather than a follow-up: `uat.scope` still reports
TRUNCATED honestly so no completeness is claimed over an unread row, while the
accepted fence-shortfall over-report no longer flips the aggregate fold that
withholds a phase's percentage. A genuinely unreadable file degrades as before.

Also corrects a frequency claim of mine: "no phase UAT file in-tree triggers
this" was true over a sample of zero, since the only UAT file in the tree is the
shipped template. The comment now says the shape is uncommon, which is what I
can actually support.

* fix(#3707): state that the live/archived split does not extend to outstanding_debt

Review MINOR: the split's rationale read as though it governed every counter, but
`summary.total_items` was never split — so a single archived `result: pending` row
re-trips the same Verification Debt warning the split exists to stop. The asymmetry
is deliberate: an archived parse gap is a row nobody can read, so the warning can
never be cleared, whereas an archived pending row is legible work someone can still
pay down by retesting. Debt that can be settled stays counted. The prose now says so
at the point of the claim, instead of leaving the next reader to file it as a miss.

Also rewrites the changeset, which described only the secondary fixes and omitted all
three defects the issue actually reports: `result: issue` dropped, any wrapped or
block-scalar `expected:` never matched at all, and the phase vanishing outright.

* fix(#3707): count every parse gap, dropping the live/archived split

The split had two regressions, both reproduced through the real CLI, and its
premise was false.

uat.cts carries #2766's rationale ~190 lines above the code I added: 'Outstanding
UAT items do not stop mattering when a milestone closes: a deferred human-UAT
scenario or a skipped live-stack test is exactly what gets archived still-open.'
So 'archived UAT files are complete by definition' was never true, and the split
rested on it.

Regression 1: archived-ness was inferred from path shape alone. A phase in the
CURRENT milestone, status in_progress, filed under .planning/milestones/v1.1-phases/
was classified archived and demoted out of the gate — live work reported as closed
history that needs no action.

Regression 2: the split was one-sided. total_items has no archived split, so an
archived outstanding row that PARSES gates Verification Debt while the identical
row that fails to parse was informational. The parse failure was what buried the
debt — the exact bug class this issue exists to fix, re-created one surface over.

parse_gap_files counts every parse_gap entry again, archived or not, so it agrees
with total_items on what archived means. The pre-existing archived_milestone field
and archived-phase scanning are untouched. Regression tests added for both cases.

* fix(#3707): correct the changeset clause left behind by the split revert

The changeset was rewritten before the split was removed, so its final clause
still claimed archived parse gaps are counted separately because signed-off
history is not work anyone can act on. There is one counter now, and that premise
is the one uat.cts refutes and the revert was made over. This text lands in
CHANGELOG.md verbatim, so it would have shipped a description of behavior the
code does not have.

* chore(#3707): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-26 09:16:22 -04:00
Tom Boucher
ddde001af6 enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move

Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md
reference carries a Status lifecycle section that is missing from all four
translations — the section documenting the status enum whose clobbering is
#3853. The test derives the heading set rather than hard-coding the missing
one, and names the locale and the heading when it fails.

Two tripwires that must pass today and after. The field-drift guard still
catches a re-derived fallback ladder: §8.8 instructs deleting that script, and
that instruction rests on a wrong premise about what it guards, so the test
stops a future reader from deleting it on the ADR's word. And last_activity's
label resolution is pinned to what ships today, because it is declared in one
of the two tables this phase consolidates and not the other — the
consolidation must not silently pick a side.

The locale test buckets under docs rather than state, which is what it tests;
that bucket is allowlisted with justification rather than folded into an
unrelated docs suite. It reads only markdown, so it carries no allow-test-rule
marker — a marker there would suppress nothing and would grow the unverified
pool against its ceiling.

Refs #3873

* feat(#3873): one schema owns the STATE.md key set, three tables become projections

ADR-3473 §8.8. The key set was declared in four places that had to agree by
hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE,
FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One
frozen null-prototype schema now declares each key's type, enum, cardinality,
source, preservation, body source, body label, accepted parse shapes and
whether it is emitted unconditionally; the three tables are derived from it at
module load.

The projections are byte-identical to the literals they replace, key order
included, and the parity tests compare against verbatim copies of today's
tables rather than re-deriving both sides from the schema — a parity test fed
from one source proves nothing, which is how a consolidation ships a changed
policy under a green test.

last_activity was the live disagreement: present in one table, absent from the
other. The schema declares what ships today rather than the tidier answer, and
a test pins it.

The schema is a leaf module and owns the four field-policy types, re-exported
from state-transition so existing importers are untouched — the same split
health-diagnostic-types made to break a CJS require cycle.

Refs #3873

* feat(#3873): generate the schema-derived regions, parity-check the prose tables

ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in
the shipped template and all five reference docs, follows gen-features.cjs's
fail-closed contract, and is wired into regen:derived and lint:generated-sync.

The Status lifecycle section was missing from all four translations — the
section documenting the status enum behind #3853 — and is now generated into
every locale. Field cardinality is a new generated table: pure schema data,
no prose, so nothing to lose.

The Field-reference and Status-values tables are parity-CHECKED rather than
generated. Their Purpose, When-populated and Matched-text columns are
genuinely hand-translated per locale, and §8.8 itself says prose stays
hand-translated; generating them from an English registry would overwrite four
locales' translations on every write. The row set is checked against the schema
instead, so a key added to one and not the other fails, which is what field
drift actually means. Building that check found last_activity_desc
undocumented in all five tables.

Three keys the docs describe are absent from the schema — active_phase,
next_action, next_phases. They are grandfathered by name, not by wildcard, so a
fourth fails: a declared gap with a forcing function rather than a silent one.

Refs #3873

* fix(#3873): declare what the parsers do, and close the shape-parity gap

Two declarations in the new schema described intended behavior rather than
actual — the defect class this epic exists to end, committed inside the epic.
Both were caught by executing the parsers instead of reading their docstrings.

current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid
shape errors; the path that looks like support is parseInt truncating '2 of 5'
to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT
fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands
this row must widen, and the shape test will go red until it does — the schema
and the parser cannot drift apart quietly, which is what §8.8's checked-not-
generated rule is for.

STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold.
normalizeStateStatus passes unrecognized prose through unchanged, so it is not
closed at runtime. The seven members are the canonical values it maps onto; the
docstring now says that and the test asserts the real lenient contract.

Closes the acceptance item that a test asserts the parsers accept exactly the
declared shapes: the check is table-driven over every row carrying
acceptedShapes, guarded against passing vacuously on an empty set, and fails
loudly if a future row has no registered driver. Adds the unwired-label throw
and the fast-check property that every projection agrees with its schema row.

Refs #3873

* fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail

The remote matrix caught 12 failures with one cause. Making the template's
frontmatter a generated region wrapped it in its own yaml fence ahead of the
markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which
both match the single markdown block — found the heading first, not the
frontmatter. That breaks the contract every new project's STATE.md is created
from: bug #21 and epic #1969 B8 pin that the File Template block starts with
frontmatter and carries gsd_state_version.

The markers now sit inside the single markdown fence, so the fence opens before
the frontmatter and the region still ends ahead of the heading. Same layout as
before this phase, with markers embedded rather than a second fence.

Row 27 existed to catch exactly this and did not, because it was writer-seeded:
it asserted against the generator's own output shape, so it passed on the broken
template. It now parses the fence the way production does and was verified to
fail against the broken shape before being trusted against the fixed one. A test
that would not have caught the bug it exists to prevent is worse than no test.

The emitted-attribution failure was separate and the fragment was the wrong
remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy
identity rule, so a diff touching it needs no acknowledgment. Fragment deleted
rather than left explaining nothing.

Refs #3873

* docs(#3873): how to change the STATE.md schema

The phase gate was right and my docs artifact was wrong. I listed
lint:generated-sync as the second enablement step, which is a verification
command dressed as one, and then claimed a one-step sequence owed no how-to.

The real sequence is build:lib then regen:derived, and the ordering is a trap:
the generator reads the COMPILED schema, so regenerating before building
regenerates against the previous schema and commits artifacts that look
plausible while disagreeing with the code just written. A reference table
cannot carry an ordering dependency; that is what the how-to test is for.

The page covers adding, changing and removing a key, every reason code the
check emits and what to do about each, what is generated versus hand-translated
and why the two prose-bearing tables are parity-checked instead of generated,
adding a language, and the three grandfathered keys. Indexed from docs/README.md.

Refs #3873

* chore(#3873): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-26 01:57:47 -04:00
Tom Boucher
3b18eff388 enhance(#3872): what a command reports it wrote — the transaction diff (#3878)
* test(#3872): failing-first regressions for what a command reports it wrote

Pins ADR-3473 §8.7 at the consumer's output. state planned-phase advances
current_phase on disk and never reports it, and reports progress.total_plans
which reconcileReportedFields silently drops because it cannot resolve a
dotted key against nested frontmatter. Both directions of #3818's own
before/after diff, reproduced against the real CLI.

Also pins the two properties the change must not break: a fully-failed patch
still reports an empty updated array, which is what state.cts:607's success
boolean depends on; and two content-identical writes differ in last_updated
alone. That second one measured state_head NOT to be ambient — it is
recomputed every write but only changes when git HEAD moved — so the
provenance exclusion is a one-element set, with a companion test pinning that
state_head does change when HEAD moves.

Refs #3872

* feat(#3872): derive what a command reports from the transaction diff

ADR-3473 §8.7. reconcileReportedFields compared the transform's own output
against persisted bytes and then filtered what preservation had restored by
its FIELD_CLASSIFICATION policy. Both are replaced by one comparison of
persisted against the pre-write state the transaction already holds, surfaced
to the command through the same caller-allocates out-param idiom divergedFields
established.

Both of the old directions fall out of that single comparison: a field the
transform reported but the pipeline discarded is persisted-equals-snapshot and
drops out, and a field nobody reported but the write moved is different and
appears. The classification filter is deleted, not relocated — no policy test
remains anywhere in the reporting path.

Reporting is at dotted-leaf granularity, enumerated from the progress.* rows
FIELD_CLASSIFICATION already declares rather than by walking user data to
arbitrary depth. That closes a live defect: plannedPhaseCore already pushed
progress.total_plans and reconcileReportedFields silently dropped it, because
a flat hasOwnProperty cannot resolve a dotted key against nested frontmatter.
Current Position was lost the same way and is fixed in the same place.

The exclusion is one field, last_updated, and it is by provenance rather than
by classification: it is the only field measured to change on every write
regardless of content. state_head was measured NOT to qualify — it is
recomputed every write but only changes when git HEAD moved. Without that
exclusion state.patch's success boolean, which is updated.length > 0, would be
permanently true and a fully-failed patch would report success.

Refs #3872

* fix(#3872): cover the matrix, and close a prototype-chain read the coverage found

Review found 20 of 29 test-matrix rows uncovered. Covering them found two real
defects rather than merely documenting the intended behavior.

bodyLabelFor read FRONTMATTER_KEY_TO_BODY_LABEL with a bare bracket index on a
plain object literal, so a field named __proto__, constructor or toString
resolved to the inherited prototype member and leaked a non-string value into
the updated array. Fixed with an own-property check, mirroring the discipline
resolveFrontmatterPath already had. The security-relevant matrix row proved it
before the fix.

applyPostSyncPreservation still carried its own inline copy of the value
comparison alongside the new stateFieldValuesDiffer, which is two live copies
of one rule introduced by the epic that exists to remove them. Routed through
the single owner.

Adds the fast-check property that a field appears iff its persisted value
differs from the snapshot, the string-versus-number representation boundary,
dotted paths into missing parents and into scalars, deleted and added keys,
and the preserve-if-placeholder pair that proves no classification test
survives in the reporting path.

Refs #3872

* docs(#3872): document the transaction diff on the write path

The updated array's contract belongs where the write path is described. States
the iff rule, leaf granularity, the single provenance exclusion and why
state_head is deliberately not one, and closes with the consequence a reader
actually needs: these arrays are longer than they used to be, because they used
to under-report.

Refs #3872

* fix(#3872): a derived leaf materializing is not a change the caller made

The remote matrix caught 17 failures with two causes. The substantive one is
that progress is source: disk, and the disk cannot change during a STATE.md
write — the write only touches STATE.md. So a progress block appearing where
the snapshot had none is the scanner populating a document that had never been
synced. The bytes moved; nothing the caller did moved them.

That is the same shape as last_updated one level up, so the provenance rule is
generalized rather than special-cased: a field appears iff its persisted value
changed for a reason attributable to this write's action, and two cases are not
attributable — a field stamped unconditionally on every save, and a declared
derived leaf materializing from a source that did not change. Crucially this
does not consult the preservation policy, so the filter §8.7 deleted stays
deleted; it uses the declared leaf set to know which keys are derived.

This had a second production consumer the earlier review concluded did not
exist: cmdStatePlannedPhase gates publishStateContract on updated.length, and
its own inline comment predicts exactly this failure. A no-op call was
publishing state.json.

advancePlanNoOpDoesNotPublish genuinely encoded pre-§8.7 behavior and moves. E2
and E6 had carved out total_plans as reportable-on-materialization, an error
introduced earlier on this branch rather than a pre-existing pin, and are
corrected with it.

Refs #3872

* chore(#3872): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 23:51:16 -04:00
Tom Boucher
8f674281fd chore(#3875): sweep the spent ack fragments and automate the sweep (#3877)
* chore(#3875): sweep the spent ack fragments and automate the sweep

next has been red on every push since a84f75630 (#3823) — 24 consecutive pushes
over two days — on two fully-spent emitted-drift-ack fragments nobody swept.

#3823 introduced guard-no-ack-on-next together with a 45-fragment sweep, but
computed that sweep as a static set of deletions fixed at its branch point.
#3809's fragment merged to next while #3823 was in flight, so the guard reds on
its own merge commit. The condition is evaluated dynamically at merge time and
remediated statically at branch time; on a moving branch the second can never
reliably satisfy the first.

- delete tests/emitted-drift-acks/3809-* and 3866-* (3034-* and 3172-* stay --
  the #3842 open-PR hold correctly defers them)
- runGuardNext returns `sweepable`, the set the guard actually reasoned about,
  plus `legacyPresent` for the legacy document, which is a fixed path rather
  than a fragment basename and would otherwise be invisible to any sweeper
- new --sweep-plan mode turns the guard into a work list: plan on stdout, prose
  on stderr, exit 0 so a non-empty plan does not fail the step that asked for it
- main() is injectable in BOTH lanes; a half-injected seam lets a test that
  passes cwd silently read the real repository instead of its fixture
- ack-fragment-sweep.yml derives its deletion list from that plan on a timer and
  opens a reviewable PR, so the sweep can no longer go stale between branch
  and merge

Hardening found in review, each verified against a live reproduction:

- git rm reads its arguments as PATHSPECS with wildmatch semantics, so a
  fragment named a bare-star .json name -- legal, and admitted by
  listFragmentFiles since it filters only on the suffix -- expanded to every
  fragment in the directory, including ones the #3842 hold withheld. Confirmed
  in a scratch repo: one such file deleted all three. Closed with a literal
  allowlist and a :(literal) pathspec, two independent layers.
- an apostrophe inside a heredoc nested in a command substitution is an
  unterminated quote and a hard syntax error at runtime, not just under bash -n.
- an empty plan no longer reports success unconditionally: the guard is re-run
  without the hold to tell "next is clean" from "everything is held", the
  commonest holder being the sweep PR from the previous run, which touches
  exactly the fragments it proposed to delete.
- a branch pushed by a run that died before it could open the PR wedged every
  later run on a non-fast-forward push; re-pointed under a lease instead.
- a guard crash in plan mode no longer reads as "nothing to sweep".

Refs #3875

* chore(#3875): regenerate CONTEXT-INDEX.json for the glossary entry

lint:generated-sync failed on CI: gen-context-index.cjs derives
docs/CONTEXT-INDEX.json from CONTEXT.md, and the RULESET.EMITTED_ATTRIBUTION
entry added in the previous commit left it stale.

Refs #3875

* chore(#3875): regenerate the example CONTEXT-INDEX for the glossary entry

CONTEXT.md feeds TWO committed indexes, not one: docs/CONTEXT-INDEX.json via
scripts/gen-context-index.cjs, and the examples/dynamic-context-management copy
that lint-example-parser-parity holds to a fresh parse. The previous commit
regenerated only the first, so the parity check stayed red.

Refs #3875

---------

Co-authored-by: sim <sim@local>
2026-08-25 23:17:29 -04:00
Tom Boucher
1863f5569c enhance(#3871): the state transaction — mandatory snapshot, open()/rebuild() (#3874)
* test(#3871): failing-first regressions for the dropped curated progress block

Pins ADR-3473 §8.6 / #3756 at the consumer's output: state record-session and
state add-decision on an archived-milestone project drop the curated progress
frontmatter entirely, exit 0, and report nothing. Reproduced against the real
CLI before writing the tests, not inferred from the issue text.

Also adds the unit-level probe that applyStatePreservation's preserve-always
row is inert on a resyncing write, and an over-preservation guard that an
empty project is never inflated.

Refs #3871

* feat(#3871): make the STATE.md pre-write snapshot mandatory via open()/rebuild()

ADR-3473 §8.6. StatePreservationInput's nullable preFm and the always-present
preFmSnapshot were the same extractFrontmatter call, one of them nulled on
resync — a policy flag baked into a snapshot. Both collapse into a single
StateTransaction whose snapshot cannot be absent: openStateTransaction()
applies preservation, rebuildStateTransaction() does not, and both carry the
snapshot because the reporting phase needs it either way. An absent snapshot
is now a construction failure; an empty one stays legal, because that is what
a document with no parseable frontmatter honestly has.

writeStateMd requires a rebuild transaction, which types ADR-3408 §8.3's
closed exception list at both call sites (state sync, health --repair) instead
of matching them as strings in a ratcheted baseline.

Fixes the dropped curated progress block: an all-zero or absent derived total
set is an unmeasured scan, not a measurement, so the curated block stands.
Also fixes two defects surfaced while building — preserve-always reported a
mutation even when it restored an identical value, and it re-entered the
curated object by reference, which would alias the snapshot the next phase
diffs against.

Refs #3871

* fix(#3871): close the three remaining subsumed defects and restore the arm the type does not replace

Review of the first two commits found four things.

The guard shrink deleted the seam-bypass axis whole, but only its
writeStateMd( arm became redundant. Its other arm catches a call site
re-assembling syncStateFrontmatter + applyPostSyncPreservation instead of the
owned composition, which the transaction type does not make unrepresentable
and which #3469 found live. Restored as findCompositionBypasses, terminal
rather than ratcheted.

Three of the four issues this phase claims were untouched. All three are the
epic's own shape and are fixed at the seam: current_phase_name is reasserted
from the curated value when the caller names none, and cmdStateJson stops
carrying a hand-maintained list parallel to FIELD_CLASSIFICATION and projects
it instead.

The construction failure that is the point of this phase had no test. Every
enumerated matrix row now has one, including the measured-versus-unmeasured
coercion boundary and a seeded property that no curated key is ever dropped.

ADR-3473 §8.6 said the guard 'keeps only its raw-write check'. Verified
against next: there was no raw-write check, and four other checks it does not
name. Amended in place with the evidence. ARCHITECTURE.md separately
advertised a preservation policy the code had deleted.

Refs #3871

* fix(#3871): do not let the unmeasured-scan rule block an explicitly-requested resync

The remote matrix caught over-preservation, the failure this phase's own
negative space says must not happen. state update Progress re-derives the
block from the body the caller just rewrote; on a project with no phase dirs
the derivation yields zero totals, the unmeasured rule read that as 'the scan
measured nothing', and the stale curated percent was restored over the resync
the user asked for.

preserve-always already said what the missing condition was: never overwrite
unless the caller explicitly names this field. explicitProgressField carries
it and is derived from shouldResyncStateProgress, not set by hand at a call
site, so it cannot drift from what the caller asked for.

Two defects found in the same mechanism and fixed with it. readModifyWriteStateMd
enumerates its option keys, so a new option was silently dropped rather than
rejected. And the raw-write axis captured its first argument up to the first
comma, which lands inside a nested path.join, so a write to a STATE.md literal
was invisible to it — the prove-it-can-fail test caught that one immediately.

No test assertion was weakened; all three frontmatter rows encode #3242, #1969
B3 and #1972 and stand unchanged.

Refs #3871

* docs(#3871): record why the raw-write check is kept, not why it was named

The amendment justified findRawStateWrites as 'written because §8.6 requires
it to exist', which is cargo-culting the contract and would have been the
wrong reason to keep anything. The real reason is that writeStateMd acquires
the STATE.md lockfile and a raw fs.writeFileSync acquires nothing, so this is
a lock bypass and lost-update is the #500/#905/#1230 family — and after this
phase it is the one reachable path into the file that nothing else covers.

Also records why ADR-3408 §8.6's deletion of the 'clear' policy is not the
precedent it looks like: 'clear' was dead vocabulary in a closed enum, this is
coverage of a reachable path.

Refs #3871

* chore(#3871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 20:38:12 -04:00
Tom Boucher
382bf7c423 fix(#3706): deliver the resolved reasoning effort to OpenCode subagents (#3867)
* test(#3706): failing-first coverage for OpenCode variant emission and frontmatter escaping

* fix(#3706): emit the resolved reasoning effort as OpenCode's variant key

`query resolve-execution` resolved an effort level for every agent, but the
OpenCode bake wrote only `model:` — the effort never reached the generated
agent, so subagents ran at whatever the runtime defaulted the model to. This
is the effort-side twin of the model-side defect fixed in #3705.

The key is written only when an `effort` block is actually configured.
`resolveInstallTimeEffort` always returns a level (the catalog default is
`high`), so gating on its return value would stamp `variant: high` into every
existing OpenCode install — and OpenCode resolves a variant name against a
`variants` map in the user's `opencode.jsonc`, so a value nobody declared is
not a safe default. Gating on `readGsdEffectiveEffortConfig` keeps installs
that never asked for effort routing byte-identical.

Kilo does not receive the key: `EFFORT_ARGV` declares surfaces for claude,
opencode and codex and has no kilo entry. This is deliberately asymmetric with
the model side, where #2794 J8 requires the two runtimes to resolve alike.

Both frontmatter sinks now route through `frontmatterScalar`, which quotes and
escapes any value that is not a plain scalar. The raw interpolation predates
this change, but it was already shown by execution during the #3705 security
review to let a config value containing a newline inject additional top-level
keys (`tools:`, `permission:`) into a generated agent file. This change adds a
second write to that sink, so it is closed here rather than doubled.

* fix(#3706): quote frontmatter values YAML would not read back verbatim

Self-review of the predicate added in the previous commit. Treating
/^[A-Za-z0-9._:/@+-]+$/ as 'safe to emit bare' answers the wrong question:
a value can match it and still not round-trip.

  - A leading '@' is a YAML *reserved* indicator and may not open a plain
    scalar at all, so a scoped ID like '@org/model' emitted bare is a parse
    error, not an ambiguity — the whole agent file becomes unreadable.
  - 'no' / 'y' / 'off' / 'null' resolve to booleans and null, so a variant
    with one of those names would match no entry in the user's variants map.
  - '12:30' resolves to 750 under YAML 1.1 sexagesimal, and ':' is legal
    mid-identifier here, so the form is reachable rather than contrived.

Real model IDs pass every clause and stay bare, so already-generated files
remain byte-identical.

* fix(#3706): route variant through the declared effort seam and cover the live path

Addresses six findings from the isolated review, all confirmed by execution.

The tests were the serious one: they required `../bin/install.js` while the fix
landed in src/, which compiles to gsd-core/bin/lib/. They exercised a different
copy of the converter than the one the bake actually uses, so the whole suite
was green-by-construction against unchanged code and the remote run failed all
13. Every case now runs against BOTH copies from one table, which doubles as the
parity assertion the generative-fix note in runtime-artifact-conversion.cts asks
for, and bin/install.js carries the mirrored change.

Emission no longer hand-rolls the value. It goes through `renderEffortArgv`,
the declared OpenCode effort seam (EFFORT_ARGV.opencode: its own supported set
and clamp). That is what rejects a level that is not a wire value — above all
`inherit`, which per #3533 (10d) means "omit the key and follow the host
default" and was previously written literally, naming a variant that cannot
resolve. Reachable two ways, both now pinned: an agent_overrides entry and a
routing_tier_defaults entry. A bare effort.default does NOT reach a tiered
agent (the #3531 tier ladder answers first), so a test written against
`default` alone asserts nothing — that is pinned too.

The plain-scalar decision moved into frontmatter.cts beside
`scalarNeedsDoubleQuoting` rather than sitting next to it as a second, weaker
predicate. `agentScalarNeedsDoubleQuoting` is a documented superset: it adds a
trailing `:` (read as a nested mapping key, which fails the whole frontmatter),
boolean/null words, and numeric-looking values including YAML 1.1 sexagesimal.

Docs now state the cascade plainly: the gate is on effort being configured at
all, not on the individual agent being named, so every generated OpenCode agent
gets a variant line once any effort block exists.

* test(#3706): assert the two frontmatterScalar copies cannot diverge

A hand-picked adversarial corpus plus a fast-check property over
YAML-significant strings, both run against bin/install.js and the live
src copy. Verified the property can actually fail: mutating one copy's
quoting rule is killed well inside the run budget.

* fix(#3706): close the review findings — predicate, seam, and dead mirror

Third review round; every item below was confirmed by execution.

The scalar predicate was wrong in two families, both found by a round-trip
property test rather than by reading. Basing it on scalarNeedsDoubleQuoting
dropped the "first character must be alphanumeric" clause, so `~`, `.inf`,
`.nan`, `+1`, `-0` and `.5` went out bare and came back as null/floats/ints;
and that base predicate only inspects the FIRST character, so an embedded `: `
(a nested mapping, i.e. a parse error) or ` #` (a comment, i.e. silent
truncation) also passed. Dates round out the set: `2026-08-25` opens
alphanumeric, survives every other clause, and YAML resolves it to a Date.
The property now asserts the contract directly over generated values instead
of trusting an enumerated character list.

The bin/install.js mirror is gone. Its premise was false — install.js already
requires bin/lib at :65 — and it was unreachable besides: install.js's
convertClaudeToOpencodeFrontmatter has no `isAgent: true` call site, because
its agents path resolves converters from the compiled module. It was a third
copy of the YAML rules serving a test rather than a caller, so the file is
back to origin/next and the tests target the live copy only.

Effort clamping moved to `clampEffortForHost`, which renderEffortArgv now
delegates to. The layout was calling renderEffortArgv with a hardcoded 'argv'
to borrow its clamp, which read as if the frontmatter key were gated on the
invocation-time axis. It is not: claude declares effortSurface "argv" and
independently bakes an effort: key. One capability table, one clamp, two
channels that no longer pretend to be each other.

Also corrects an earlier claim of mine: adding EFFORT_RENDERING.opencode would
NOT have made `effort sync` write the wrong key, because it guards on the
runtime name before it ever renders. The seam choice stands on other grounds.
`effort sync` still skips OpenCode, but its stated reason claimed OpenCode
"does not use effort: frontmatter", which this change makes false — so the
message now says what is actually true.

* docs(#3706): restate the changeset around the round-trip contract

* fix(#3706): restore the changeset fragment belonging to #3809

An earlier commit in this branch picked the first file in .changeset/ by
glob order instead of the fragment created for this issue, and overwrote
agile-geese-squeak.md (PR 3815 / #3809) with this change's body. Restored
verbatim from origin/next; this change's text now lives in its own
patient-cranes-parade.md, where it was created.

* feat(#3706): maintain the OpenCode variant key from effort sync

Install bakes the resolved effort into OpenCode agent frontmatter as
`variant:`, so `effort sync` has to maintain it or a config change only takes
effect on reinstall — and its skip message claimed OpenCode does not use
frontmatter effort at all, which this issue made false.

cmdEffortSyncOpencode mirrors the codex branch: resolve per agent, clamp
through the declared OpenCode capability, then write, strip, or skip. A null
target means the key must not exist, which covers both "no effort configured"
and "resolved to inherit or to an unsupported level" — the same states under
which install writes nothing, so sync and install agree by construction.

The frontmatter line-editors are key-parameterised rather than copied:
setEffortFrontmatter / removeEffortFrontmatter are now thin wrappers over the
same internals the variant path uses, and a test pins that the claude `effort:`
behavior did not move. The child-process test harness fixes both HOME and
USERPROFILE, so the hermetic-config assertions cannot pass vacuously on Windows.

* fix(#3706): scope the frontmatter line editors to the matched block

Found by the security review of the sync path, reported as correctness rather
than vulnerability, and reproduced against pre-fix code before being fixed.

Both editors matched the frontmatter with a regex that can match a block after
a preamble, then derived the EOL and the opening-fence length from the START OF
THE FILE. On a CRLF document with a preamble those disagree, the offsets shift
by one byte, and the reassembled document comes back with a mangled fence
(`---\rname: x`). Both now take the EOL from the matched block.

`setFrontmatterKeyLine` additionally did a whole-file `/m` replace when the key
already existed, gated only on the key being present in the frontmatter body —
so a preamble line starting with the same key was rewritten instead of the
frontmatter one. It now replaces inside the frontmatter span only, which is the
hazard `removeFrontmatterKeyLine` already documented and guarded against.

Neither is reachable from an install-written `gsd-*.md` (those begin at byte 0
with `---`), and both predate this change — but the editors are in this diff
because #3706 key-parameterised them, so they are fixed here rather than left
for the next caller to trip over. Three regression tests, each confirmed to
fail against the pre-fix build.

* fix(#3706): treat a present-but-empty key as present, and pin the real seam

Fourth review round.

The MAJOR one: both sync branches read the current value with `(.+?)`, which
needs at least one character, so a key present with an EMPTY value read as
"key absent". When the target was also null the code concluded "already
correct" and skipped — leaving the key in the file, where it reads back as
YAML `null`: exactly the unresolvable-variant state this change exists to
prevent. Whitespace decided whether it fired, since `variant:   ` matched and
`variant:` did not. Presence and value are now separate questions at both the
opencode and the claude branch.

The OpenCode writer now follows the codex branch rather than the claude one:
tmp file plus retryRenameSync with orphan cleanup, and a write failure skips
that agent and is reported instead of aborting the sweep. Same granularity,
same transient-Windows-lock exposure, so the hardened sibling was the right
precedent.

Also: the generic line-editors escape their interpolated key, the JSDoc
stranded by the clampEffortForHost extraction is back on renderEffortArgv, and
a cast that declared a nullable function as non-nullable is corrected.

Tests close the gaps the review listed — empty value (both spellings), CRLF
round-trip through write and strip, the symlink guard, a body line starting
`variant:`, a file with no frontmatter, and the YAML classes that actually
broke the predicate. The new layout-seam test drives the real stage() path and
was verified to FAIL when `variant` is removed from the converter call; a seam
test that survives cutting the seam is worse than none.

* fix(#3706): clear the round-five review findings

No blockers or majors this round; the repo's review gate is zero-tolerance, so
the minors are cleared too.

A duplicated key was only half-stripped: the strip regex had no `g` flag, so a
frontmatter carrying the key twice lost one occurrence, reported success, and
left the "a null target means the key must not exist" invariant false on disk —
converging only on a second run. Such a document is already invalid YAML, so
this is robustness rather than a live corruption path, but a successful sync
has to leave the invariant true.

A run in which every write failed still summarised as `ok`, so a caller could
not tell "nothing to do" from "everything failed". The OpenCode branch now
reports `failed` when any write failed. The write-failure path was also the
newest code in the change with no coverage at all; it now has a test that
injects the failure by monkeypatching the write, per CLAUDE.md §4, rather than
by chmod — mode bits do not bite under root in CI.

`CodexEffortSyncWriteFailure` is renamed `EffortSyncWriteFailure` now that two
branches share it. Removed a guard on the claude concrete path that was
provably unreachable — no member of EFFORT_SET renders null there, so it read
as protection that did not exist. The claude inherit path's presence check is
load-bearing and untouched.

Three stale statements corrected: the OpenCode result shape matches codex's,
not claude's, now that it emits write_failures; the `thread()` test helper now
calls `clampEffortForHost` so it genuinely mirrors the layout instead of
merely claiming to; and a test helper restored `USERPROFILE` by assignment,
writing the literal string "undefined" into the environment on POSIX — it
deletes now.

* fix(#3706): converge the set path, degrade on unreadable files, preserve mode

Rounds five and six of review. No blockers or majors; the review gate is
zero-tolerance, so the minors are cleared too.

`setFrontmatterKeyLine` was the mirror of a defect already fixed in its
sibling: `remove` was made global, `set` was not, so on a frontmatter carrying
the key twice it rewrote the first and left a stale second. Last-wins YAML
readers honour the stale value while the sync's own first-occurrence read
reports "in sync" — permanently non-converging. It now collapses to exactly one
occurrence, in the position of the first, so ordinary single-occurrence
documents stay byte-identical (verified across seven shapes before and after).

An unreadable agent file used to throw and abort the entire sweep, while a
failed WRITE in the same loop degraded into a report. The OpenCode branch now
reports read failures alongside write failures; the claude branch degrades to a
skip without a new result field, because its shape is long-standing and widely
consumed and one bad file aborting the sweep is the actual defect.

The tmp+rename publish dropped the original file's mode — a plain writeFileSync
preserves it, a rename does not — so a 0600 agent came back 0644. Both the
OpenCode and the codex branch now carry the original's permission bits across
the publish, masked with 0o7777: the raw stat mode includes the file-type bits,
and POSIX leaves those unspecified for chmod. Linux is the only OS the remote
matrix runs, so relying on Darwin's tolerance would have been untestable here.

Also documents the `from` contract on EffortSyncChange (null means the key was
absent, '' means present with an empty value — a distinction earlier rounds
introduced and then collapsed in the output), adds OpenCode to the docs
paragraph enumerating where the key is omitted under inherit, and records in a
comment that the 'failed' summary reaches only raw mode and does not change the
exit code, which is a CLI-contract change affecting all three branches and is
deliberately not made here.

* fix(#3706): guard the codex read, close the tmp permission window, rename the failure type

Round seven, plus one thing I found myself.

`cmdEffortSyncCodex` still had an unguarded `fs.readFileSync` — a read fault on
one agent exited 1 and aborted the whole sweep. The claude and opencode
branches were both guarded earlier this round and codex was missed, with the
unguarded read sitting ten lines above the chmod block the previous commit did
edit. It now reports read failures the way the OpenCode branch does, and a read
failure flips its summary to `failed` — which write failures did not do there
either, so both are corrected for consistency.

The tmp file was created at the default mode and only tightened afterwards, so
a 0600 agent's contents sat in a 0644 file for the length of the publish. I
measured the window rather than assuming it, then closed it by passing the
mode at creation. The chmod after the write is deliberately RETAINED and
commented: the `mode` option only applies when the file is actually created, so
a leftover tmp from an earlier crashed run would be truncated and reused at its
old mode, and the chmod is what corrects that.

`EffortSyncWriteFailure` is renamed `EffortSyncFileFailure` — it was typing a
`read_failures` array, the same naming-lie the `Codex…` prefix had last round.

Also pins the codex mode preservation with a test. It only writes on a path
that genuinely rewrites the file, so the fixture is an Anthropic-flavoured
model pin the sync strips, and the test asserts the content changed before
checking the mode — otherwise it would pass on a sync that did nothing.

* fix(#3706): guard the claude writes and share one escaping rule

The security sign-off caught a comment of mine that was factually wrong: the
new claude read guard said the failure is folded in "like the write path in
this same loop does", and there was no write guard in that loop. Rather than
correct the sentence, both claude write sites are now guarded the way the read
is — a failed file is skipped, the sweep continues, and the raw summary token
flips to `failed`. The JSON shape stays frozen deliberately, because it is
long-standing and widely consumed; the token is the channel that can carry the
signal without a compatibility risk, which is the reviewer's own suggestion.

That makes all three branches consistent: reads and writes guarded everywhere,
per-file failures degrade instead of aborting, and every branch reports
`failed` rather than `ok` when something did not sync.

`setFrontmatterKeyLine` interpolated its value raw while the install-side
writer quoted through the shared helpers — two writers of the same frontmatter
key disagreeing on escaping, the divergence class this repo requires closed.
They now share one rule. Verified no churn: all six effort levels are plain
scalars and emit byte-identically, with claude's documented minimal-to-low
clamp the only difference in the table, exactly as before.

* fix(#3706): publish claude agent writes atomically too

Both reviewers found this independently, and it is data loss rather than a
reporting gap. The claude branch wrote in place, so `fs.writeFileSync`'s
O_TRUNC meant a post-open fault left the agent file truncated or half-written:
an injected ENOSPC produced an empty file, and under `ulimit -f` a 60000-byte
agent came back as 512 bytes of wrong content. The guard added earlier this
round then counted that destroyed file as `skipped`, which in JSON mode is
indistinguishable from "already in sync" — so a caller would have read the
sweep as clean while an agent on disk was corrupt.

It now publishes the way the codex and opencode branches already do: write to
a tmp file created at the original's masked mode, chmod, then retryRenameSync,
with the tmp unlinked and the agent skipped on any failure. The corrupting case
is gone rather than merely reported, which matters because this branch
deliberately takes no new result key.

I had claimed all three branches were consistent after the previous commit.
That was true for degradation and reporting and not for atomicity; the reviewer
caught the overclaim. It is true now.

Also sorts the claude file list, which the other two branches already did —
readdir order is platform-dependent, so leaving it unsorted made the reported
`changes` ordering differ across machines for identical inputs.

* chore(#3706): backfill the changeset PR number

pr:0 placeholder replaced with the real PR now that gh api returned it.

* test(#3706): kill the frontmatter mutants this change introduced

CI's Stryker frontmatter shard scored 60.58 against a break floor of 62.
The cause is documented in the lane's own config, from #1882: this PR added a
multi-clause predicate to frontmatter.cts and exported the escaper, but the
tests constraining them live in tests/runtime-converters.test.cjs, which that
shard does not run — so every mutant in the new code was uncovered there even
though the behaviour is tested elsewhere.

The fix is assertions that kill real mutants, per the repo's own instruction,
not a lowered floor and not a Stryker disable: scripts/mutation-matrix.cjs is
untouched. Each clause of agentScalarNeedsDoubleQuoting now has a true case AND
a near-miss that must answer the opposite way, so flipping the clause fails a
specific named test — alnum-first against `a-b`, trailing `:` against `foo:bar`,
embedded `: ` against `a:b`, embedded ` #` against `a#b`, the word list against
`yes1`/`nullish`, the numeric forms against `1a`/`0xzz`, the timestamp against
`2026-08-25x`, plus the case-insensitive spellings that pin the `i` flag.
escapeDoubleQuoted is pinned on exact output, including a case constructed so
that escaping in the wrong ORDER yields a different string.

Two of my expectations were wrong and are asserted as the code actually
behaves: `12:99` is NOT quoted, because the sexagesimal alternative never
range-checks minutes and so does not match — which is right, since YAML would
not read it as sexagesimal either; and `20260825` is quoted by the numeric
clause rather than the timestamp one, being a bare integer.

* chore(#3706): ratchet the frontmatter mutation floor to 65

The lane measured 66.67 on PR 3867 after the mutant-killing unit tests landed —
above its pre-change 63.35 baseline, not merely recovered. Step 3 of this
file's own HOW TO UPDATE procedure says to set minScore = floor(measured) - 1
in the same diff, so 62 becomes 65 and the improvement is locked in rather than
left free to slide back.

The ledger of measured scores now records the new measurement, why the shard
broke in the first place (logic added to frontmatter.cts whose only tests lived
in a file this lane does not run — the same trap the #1882 note describes), and
one discrepancy: step 3 also says to update "the matching RATCHET_BASELINE
entry", but no such declaration exists in this file. The name appears only in
that comment, so minScore and the ledger are all there is to update.

* fix(#3706): update RATCHET_BASELINE alongside the raised floor

The ratchet test caught the previous commit: it raised COVERED['frontmatter']
.minScore to 65 without updating the baseline that mirrors it, which is exactly
the mismatch that guard exists to make visible in review.

I had claimed RATCHET_BASELINE did not exist. It does — in
tests/mutation-matrix-ratchet.test.cjs, not in scripts/mutation-matrix.cjs,
which is the only file I searched before concluding it was a stale reference.
The ledger comment is corrected to say where it lives and to record that the
guard caught the error rather than leaving my wrong claim on the record.

* docs(#3706): put the mutation ledger entries back under their own dates

The 2026-08-25 measurement was spliced into the middle of the 2026-06-14 list,
so adr-parser, config-schema, active-workstream-store and core-utils ended up
sitting under the wrong heading and misattributing their measurement dates.
That ledger is what a future change reads to calibrate a floor, so a wrong date
there is not cosmetic. Each measurement is now under the date it was taken.

Also drops the first-person account of my own mistake from the entry — the
factual half (where RATCHET_BASELINE lives, and that it is updated in the same
diff) is what a reader needs; the confession is not.

---------

Co-authored-by: sim <sim@local>
2026-08-25 19:54:30 -04:00
Tom Boucher
86fa2917d7 enh(#3866): dispatch step and contribution hooks at verify:pre (#3869)
* test(#3866): pin that verify:pre must dispatch every hook kind

verify-work.md's verify_pre_hooks step dispatches only `kind == "gate"`, so
getWiredKinds reports verify:pre -> {gate} and gen-capability-registry rejects
any capability declaring a step or contribution there. The verify lane is
therefore closed to capabilities that want to contribute to what UAT covers
rather than refuse to let it start.

Failing-first: the step, contribution, and exact-kind-set rows are RED; the
pre-existing gate row is a green regression pin so the new arms cannot orphan
the arm verify:pre already had.

Refs #3866

* feat(#3866): dispatch step and contribution hooks at verify:pre

verify_pre_hooks dispatched `kind == "gate"` only, so getWiredKinds reported
verify:pre -> {gate} and gen-capability-registry's validateHooksWired rejected
any capability declaring a step or contribution there. A capability could
refuse to let UAT start; it could not contribute to what UAT covers.

Add contribution and step arms mirroring execute:wave:post, deferring to
references/loop-hook-dispatch.md and carrying its ref.command in-context
validation guard ahead of any shell-use prose. A verify:pre step is advisory:
it never blocks the start of UAT and an erroring step is routed by its own
onError. The gate arm and its check guard are untouched.

Give extract_tests an additive consumption seam for the artefacts those steps
declare via the existing steps[].produces field -- no new registry field, no
new ordering, no invented filename. Manifest-supplied artefact names are
validated in-context against an allowlist and resolved only inside PHASE_DIR.
With no producing step the derivation is unchanged, pinned by test rather than
asserted in prose.

Review findings folded in: the artefact-name allowlist (isolated adversarial
pass), the artefact-shape contract and the seam-inertness tests (spec axis),
and the reference/how-to split so one constraint has one source of truth
(standards axis).

Closes #3866

* chore(#3866): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 18:07:29 -04:00
Tom Boucher
7bbbe495d7 docs(#3473): ADR-3473 — enforcement by construction, one owner per invariant (Phase 0) (#3870)
Phase 0 of epic #3473 ships this ADR alone. Every rule in §8 carries a
status of Enforced or Required — Phase N.

§8.1–§8.5 state the epic's B1–B5 criteria as contract rules with no phase
assigned yet. §8.6–§8.8 assign Phases 1–3 to the STATE.md family that
ADR-3180 and ADR-3408 left between them: the write seam's unvalidated
input, its classification-excluded reporting, and the schema transcribed
by hand into nine artifacts.

Derivation authority is explicitly NOT in scope here — it lands as
ADR-3180 Amendment 8, generalizing §7.5's already-locked sentence.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 17:50:34 -04:00
Tom Boucher
e40e9670f8 fix(#3705): consult model_policy in the install-time bake so agent frontmatter matches dispatch (#3863)
* test(#3705): failing-first coverage for model_policy in the install-time bake

* fix(#3705): consult model_policy in the install-time bake so frontmatter matches dispatch

* fix(#3705): inject the effective runtime into the policy so runtime_tiers is reached

* test(#3705): use assert.doesNotMatch, the assertion that exists

* chore(#3705): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 12:11:15 -04:00
sim
f7df920681 chore(#3841): backfill changeset PR number
Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 11:36:59 -04:00
sim
67335c498c fix(#3841): express the anchor semantically so the pretty payload still verifies
The remote matrix caught a regression I introduced in the previous commit. The
anchor check was implemented as a byte-prefix match against IDENTITY_RAW_PREFIX,
which describes the `--raw` wire format -- but `cmdRuntimeIdentity` without that
flag pretty-prints at indent 2, and the classifier is handed BOTH serializations.
Only the shell is restricted to `--raw`. The pre-existing test that runs the real
verb with no flag went red: `+ 'unparseable' - 'ok'`.

No local gate caught it. build:lib, eslint and lint:ci were green throughout,
because none of them execute tests.

The anchor now reproduces its two properties semantically instead of byte-wise,
and both hold for either serialization: the payload begins at the first byte of
stdout, and `packageName` serializes first. IDENTITY_RAW_PREFIX stays exported
with its own tests -- it is the wire contract for the shell, not a general
classifier predicate, and conflating those was the error.

Adds the two rows the matrix was missing: the default pretty serialization
verifies, and a pretty payload with `packageName` not first does not. The design
and matrix now record that they enumerated only the inputs the SHELL produces
and assumed the classifier's input set was the same.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 11:36:59 -04:00
sim
5e997de5f0 fix(#3841): honor the payload anchor in the classifier, dedup the fixture
Three review findings, all fixed.

The isolated security review found a SECOND divergence the design missed. The
classifier parses structurally, so it accepted `packageName` at any key
position; the shell's `case` is anchored at the start of stdout. For
`{"note":"x","packageName":"<us>",...}` the classifier said ok and the shell
said unverified -- a fail-open disagreement, and none of the original ten parity
rows caught it because every one put `packageName` first. Certifying agreement
that does not hold would have been worse than shipping no parity suite. The
classifier now honors the anchor on its ok arm, which is what the module already
claimed to do: IDENTITY_RAW_PREFIX is documented as "ANCHORED, never a substring
search". A foreign packageName stays identity_mismatch wherever it appears, so
that arm is untouched. Parity rows P11-P13 added.

The spec review found the test matrix marked the empty-string `packageName` row
as already covered. It was not -- the `.length > 0` guard is a distinct path
from "no packageName key at all", which is tested. Added, and the matrix
corrected to say it was wrong.

The standards review flagged ~45 lines of fixture helpers duplicated verbatim
between the two preamble describes. Extracted to one `makeIdentityFixture()`.

Changeset rewritten: it described only the exit-code half of the diff.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 11:36:59 -04:00
sim
3b8e4f3e5a fix(#3841): parse the identity payload before consulting the probe exit code
`classifyIdentityProbe` short-circuited on a non-zero exit BEFORE it looked at
stdout, so a tool that had already proved its identity and merely exited
non-zero was classified `no_identity_verb`. The launcher preamble that this
module speaks for reads stdout only -- its command substitution discards the
status -- so the two surfaces disagreed on exactly that input: shell `ok`,
classifier `no_identity_verb`. Measured against the real snippet, not inferred.

That matters because the classifier is the engine for the announced hard-fail
phase and has no production caller yet. An install verified by today's warn
phase would have been refused the moment hard-fail landed, in the phase where
that stops the run rather than printing a line. Nothing recorded or tested the
difference.

The classifier moves rather than the shell: the shell is the shipped path with
observable dependents, the classifier has none. The predecessor defence is
untouched -- a usage screen yields no usable payload, so it still falls through
to the exit-code branch.

Adds the cross-surface parity test the gauntlet requires for two surfaces
implementing one decision: ten probe behaviors driven through both the real
snippet and the classifier, asserting they agree with each other.

Surfaced by re-running the feature-implementation directive's design and QA
steps against the code merged in #3848, which shipped without them.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 11:36:59 -04:00
Tom Boucher
8bfb5c47c9 fix(#3842): recover ack paths from pull requests past the per-PR file cap (#3857) 2026-08-25 11:36:27 -04:00
Tom Boucher
27318ecec3 fix(#3704): rewrite fnm versioned node paths to the stable alias on macOS and Linux (#3856)
* test(#3704): failing-first coverage for fnm versioned-path normalization on POSIX

* fix(#3704): rewrite fnm versioned node paths to the stable alias on POSIX

* fix(#3704): share one root normalizer across both fnm branches so no baked path carries a doubled separator

* chore(#3704): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 09:44:48 -04:00
Tom Boucher
aa6c332b5a fix(#3701): resolve next_phase from the roadmap, selecting the numerically lowest successor (#3852)
* test(#3701): failing-first coverage for roadmap-order next_phase resolution

* fix(#3701): resolve next_phase from roadmap order, keeping the disk scan for spelling and fallback

* fix(#3701): select the numerically lowest successor in both scans, not the first row encountered

* chore(#3701): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 08:58:10 -04:00
Tom Boucher
308c17505c fix(#3840): reject a malformed feature order instead of coercing it (#3851)
* fix(#3840): reject a malformed feature `order` instead of coercing it

`scripts/gen-features.cjs` was the one field validated by coercion rather than
by shape. `Number('')` is 0, and `0x10`, `0b11`, `0o17`, `1e3`, `1.` and `.5`
all coerce to finite numbers, so a fragment declaring a bare `order:` sorted to
position 0 -- ahead of every real feature, in both the body and the generated
table of contents -- with zero violations, a clean `--check` and `--write`
exiting 0. That is a fail-open in a gate whose entire contract is a typed
violation rather than a silent guess.

`order` is now shape-checked against an optionally-signed decimal literal
before coercion, mirroring how ID_RE guards `id`. The finite check stays: the
regex alone would admit a literal long enough to overflow to Infinity.

Surfaced by re-running the feature-implementation directive's design and QA
steps against the code merged in #3845, which shipped without them.

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3840): backfill changeset PR number

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 08:36:33 -04:00
Tom Boucher
fb2d122d7f feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb

only this package publishes. The path-based branches — a project-local install,
a runtime config directory — had no such guarantee; they trusted their
configured location. This closes them.

Mechanism: once resolution finishes, and before any verb runs, the preamble
probes the tool it picked with `runtime-identity --raw` and matches the answer
with a shell `case` pattern ANCHORED to the start of the compact payload
(`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts
the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which
any colliding package could publish. The outcome is exported as the two-valued
`GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE
rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling:
`unverified` prints one line naming BOTH causes and continues, because
`no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core`
older than the verb, and at rollout the old-version case is the common one.

The blocker was byte budget, not design. The preamble is inlined into 112
shipped files and several sat within single-digit bytes of frozen ceilings
(`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a
first attempt broke five of them. What made room was collapsing the resolver's
twenty near-identical `elif [ -f … ]` arms into one candidate-list helper
(`_gsd_at`), which buys far more than the assertion costs. The preamble is now
2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every
capped file moved away from its ceiling rather than toward it. No cap raised, no
size-budget exception added, no override token emitted.

Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source
fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all
preserved byte-for-byte in substring terms; the snippet still begins with
`_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor
on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal.

Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both
described an `[ -x ]` guard as the load-bearing re-source defense. That guard
was tried and REMOVED in #3831 — it rejected the bare function name, fell
through every branch, and hit `exit 1`, which kills a sourced caller's shell.
`unset -f gsd_run` is the actual mechanism.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3841): pair the anchor's brace by requiring a closed identity payload

The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md
has unbalanced braces: net depth 2" — plus a knock-on report from its parent
`bug #1516` describe, which is the same failure counted once at the child and
once at the block.

Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and
increments on `{`, decrements on `}`, with no awareness of shell quoting. It
scans `new-project.md` PLUS every `new-project/steps/*.md`, and both
`new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy
— hence net 2 from a snippet that was off by exactly one. The unpaired brace was
the `{` inside the single-quoted `case` pattern of the identity anchor, which is
correct shell and invisible to a text scanner.

Fix in the snippet, not the guard. The pattern now anchors at BOTH ends:
`'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that
does real work rather than a cosmetic pair — a truncated payload whose prefix
matches now fails too, where before it verified. Safe for any future additive
field: a JSON object's own closing brace is always the last character, whatever
type the last value has, which is pinned by two negative-space tests (a nested
object and an array-valued last key must both still verify). Cost: +3 bytes,
against the 1,873 the resolver fold already gave back.

The alternative considered and rejected was dropping the literal `{` for a `?`
glob. It balances too, but weakens the anchor from "must be an opening brace" to
"must be any one character", and the anchor is the entire point.

Two guards added so this cannot recur silently:
- runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next
  edit to that pattern fails on the file it broke instead of surfacing three
  files downstream in a test whose name mentions neither the launcher nor this
  issue. It also asserts depth never goes negative, since a `}` preceding its
  `{` nets to zero while being unbalanced at every prefix.
- runtime-identity gains behavioral truncated-payload and trailing-garbage
  fixtures, so the added `}` is proven load-bearing rather than merely present.

Verified: snippet 51/51 braces; new-project combined net depth 0; the seven
other preamble-bearing files with nonzero depth are unchanged from merged next
(their own prose, not the preamble, and not in any guard's scan set); all 112
inlined copies and the resolver reference re-synced byte-equal; sync:launcher
idempotent on the second run.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3841): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 01:05:53 -04:00
Tom Boucher
8d8e9ef5eb fix(#3842): stage the emitted-drift-ack sweep so it does not conflict in-flight PRs (#3847)
* fix(#3842): stage the emitted-drift-ack sweep around open PRs

The guard-no-ack-on-next sweep (#3078) deleted every all-spent fragment
under tests/emitted-drift-acks/ unconditionally. When an open PR still
modified the same fragment file, that delete became a modify/delete
conflict on the PR's next merge attempt -- the exact shared-file
conflict fragments were adopted (#2914) to eliminate, reintroduced by
the sweep itself. The first real sweep hit three open, outside-
contributor PRs simultaneously (#3330, #3774, #3648), each with the
swept fragment as its only conflicting path.

assertNoAllSpentFragments now accepts an optional openPrTouchedPaths
set (or the sentinel 'unknown') and partitions all-spent fragments
into "safe to sweep" and "held" -- a held fragment is reported
informationally, never as a failure, and is swept once the touching
PR merges or closes. fetchOpenPrTouchedAckPaths computes the touched
set with a single `gh pr list --json number,files` call. The guard-next
codepath is factored out of main() into the exported, dependency-
injectable runGuardNext() so this wiring is unit-testable without a
real, network-dependent `gh` binary.

The new behavior is strictly opt-in via a --defer-to-open-prs flag,
wired only from the guard-no-ack-on-next job in test.yml (which also
gains pull-requests: read and a GH_TOKEN env for the `gh` call). Every
pre-#3842 caller -- including every existing test -- is unaffected
when the flag is omitted.

CONTRIBUTING.md's "Why fragments, not one file (#2914)" section now
documents the staged-sweep policy so it no longer reads as though
fragments are unconditionally conflict-free once spent.

Refs #3842
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3842): unify the failed-open-PR-check message to one greppable phrase

The remote matrix (commit 4ec4527fb) failed two tests on the fail-closed
path: assertNoAllSpentFragments's 'unknown' sentinel branch described the
failure as "...could not be determined this run...", while runGuardNext's
catch around fetchOpenPrTouchedAckPaths described the same condition as
"open-PR check failed (<err>)". Both messages were genuinely present and
informative (not missing or empty), but they used different wording for
the same fail-closed condition, so there is no single string a human
scanning CI output can search for to find out why nothing got swept.

Unify both sites on "open-PR check unavailable" -- runGuardNext's line
now reads "open-PR check unavailable — <err.message>", and
assertNoAllSpentFragments's holdAll message now leads with "deferred
(open-PR check unavailable): ...". This is a real fix to the fail-safe
path's diagnosability, not a relaxed test assertion: the two now-failing
tests already expected this exact phrase, and the fix makes the code
say what the tests (correctly) expected instead of loosening them to
match arbitrary prior wording.

Refs #3842
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3842): backfill changeset PR number and retype to Fixed

pr:0 backfilled to 3847. Retyped Changed -> Fixed: the change repairs broken sweep behaviour rather than adding any, and its contributor-facing documentation lives in CONTRIBUTING.md at the repo root, which lint-docs-required does not count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:08:14 -04:00
Tom Boucher
de95c03f72 fix(#3699): report why a derived frontmatter key was not written, and repair a missing body source (#3846)
* test(#3699): failing-first coverage for derived-key reporting and the case-D fallback

* fix(#3699): report why a derived frontmatter key was not written, and repair a missing body source

* fix(#3699): scope session-field writes to ## Session so an archived line cannot absorb the update

* fix(#3699): resolve the session writer from body labels only, so a frontmatter key never writes the body

* chore(#3699): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-24 23:20:42 -04:00
Tom Boucher
36375513b9 feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments

docs/FEATURES.md was hand-maintained, and every feature PR wrote into two
shared mutable cells: the '### N.' heading whose integer was hand-allocated at
authoring time, and the hand-maintained table of contents. Concurrent PRs all
picked the same next integer, and two PRs adding differently numbered features
still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across
successive rebases, each collision also costing a full matrix verification run
because the sha-keyed pass marker dies with the rebase.

Mechanism: one fragment per feature at docs/features/<slug>.md carrying
id/title/group (and an optional order) in frontmatter, consolidated by
scripts/gen-features.cjs --write|--check into a marker-delimited region of
docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings
and their order are derived too - a group sorts by its lowest-ordered member -
so there is no shared registry to edit either; optional per-group prose lives in
docs/features/_groups/<slug>.md. A contributor adds exactly one new file.
Wired into regen:derived and lint:generated-sync alongside the eight existing
generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting.

Migration froze all 168 existing numbers verbatim: identical section set,
identical order, identical bodies. Two defects found in the tree are fixed
inline rather than carried forward - the '## Related' block had been spliced
into the middle of the document, orphaning §142's Reference line, and four
inbound anchors were already broken on next (FEATURES.md#runtime-identity in
two files, and #143-spec-phase-edge-completeness-probe off by one). Since the
repo has no link checker, --check now validates every inbound
FEATURES.md#anchor by resolved target, so that class cannot ship silently
again; locale FEATURES.md files resolve elsewhere and stay out of scope.

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3840): carry upstream §69 delta into its fragment and harden the generator

Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus
origin/next. Root cause was a stale base, not extraction loss: those lines
landed in 394bf384b (#3844) AFTER this branch forked at 63abcface, and
'git diff 63abcface origin/next -- docs/FEATURES.md' is exactly that hunk.
Merging origin/next auto-applied the hunk into the GENERATED region, which
--check immediately reported as stale; the delta is now carried in
docs/features/statemd-consistency-gates.md and regenerated from there.

--write is now fail-closed. It previously rendered the region even with
violations outstanding, warning only on stderr and exiting 0, so a
'--write && git commit' chain could commit a FEATURES.md carrying two
colliding sections. It now refuses and exits 1; --force is the explicit
override and says so in the report. The test that pinned the old behavior now
pins the refusal, plus the --force override and its scoping.

Marker forgery is rejected at two layers. A fragment body containing
'<!-- FEATURES:START' or '<!-- FEATURES:END' is a typed
body_forges_region_marker violation (fragments and group notes alike), and
spliceIntoFeatures anchors the end boundary with lastIndexOf instead of
indexOf, so a marker that reaches the document by any other route can only
make the generated region grow, never shrink. Matching is on marker PREFIXES,
so a decorated variant comment cannot slip past.

Symlinked corpus entries are refused with a typed dirent_not_regular_file
rather than read. A fork PR could otherwise commit docs/features/evil.md as a
symlink to any readable path and have the generator inline those bytes into
the committed docs/FEATURES.md on the next regen.

Equivalence re-verified with a method that cannot cancel out. The first
check extracted both operands with the same body-normalising helper, so
anything that helper dropped was dropped on both sides. The replacement runs
two independent passes: a global content-line multiset diff with no
per-section logic at all (0 gained, 19 lost, all 19 the stale hand-written
mini-TOC links this change deliberately deletes), and a per-section
byte-exact body diff carrying a coverage assertion that fails loudly per file
when the extractor accounts for fewer lines than the file contains. That
assertion caught two blind spots in the checker itself. 168/168 sections
present, order identical, one intended body difference (§142 regains the
Reference line orphaned by the misplaced '## Related' block).

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3840): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 22:49:00 -04:00
Tom Boucher
394bf384be fix(#3696): report the last_activity invariant and make the verdict gateable with --strict (#3844)
* test(#3696): failing-first coverage for the last_activity invariant and --strict exit status

* fix(#3696): report the last_activity invariant and make the verdict gateable with --strict

* fix(#3696): agree with the real reader on last_activity, and stop reporting structure as truncation

* chore(#3696): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-24 21:33:48 -04:00
Tom Boucher
6905726d9c chore(#3833): gate every PR compute lane behind a mergeability preflight (#3843)
* test(#3833): failing-first suite for the PR mergeability preflight

* chore(#3833): gate every PR compute lane behind a mergeability preflight

* fix(#3833): assert the preflight gate on parsed yaml and guard status-function if

* fix(#3833): run stub-backed cli tests in-process to avoid a spawnsync deadlock

* fix(#3833): fail the preflight open when its script is absent at the base sha

---------

Co-authored-by: sim <sim@local>
2026-08-24 21:13:55 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
aaf47c5fc2 fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget

No prompt cap can reach any CLI reviewer lane, by any configuration. Two
independent defects compound: all nine `transport: spawn` lanes declare
`promptBudgetKey: null`, so `budgetFor` returns on its first line; and the
documented global `review.max_prompt_tokens` is advertised in the schema
manifest but declared nowhere, so the resolver never materializes it and
`budgetFor`'s fallback is dead code.

Adds to tests/reviewer-config-federation.test.cjs, which already owns the
per-reviewer budget config-set/config-get idiom:

- a CLI lane inherits the global cap (RED: reports null)
- an http lane with the -1 sentinel inherits the global cap (RED: reports null)
- the resolved review surface carries max_prompt_tokens at all (RED: absent)
- per-lane overrides the global on a CLI lane
- the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as
  unset, 1 is the smallest real budget — the regression budgetFor's own comment
  warns about
- anti-tightening pins that must stay green: an empty config leaves every lane
  null, the three existing budgeted lanes are unchanged, and config-set still
  rejects a per-reviewer key naming something that is not a declared lane
- a fast-check property over the resolution contract itself, with -1, 0 and
  non-finite inputs generated explicitly rather than left to chance

Every row was reproduced by hand against the real CLI before being written, so
the RED/GREEN split is observed rather than predicted.

Refs #3691

* fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve

No prompt cap could reach any CLI reviewer lane, by any configuration. Two
independent defects compounded.

The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor,
gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so
`budgetFor` returned on its first line and `review-lane plan` reported
`promptBudget: null` no matter what was configured. Each now declares
`review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset
sentinel the three local-server lanes already use.

Separately, the central `review.max_prompt_tokens` was listed in the schema
manifest's validKeys and documented as a supported setting, but declared
nowhere — the resolved surface is built from capability declarations plus the
defaults manifest, and neither carried it. `configGet` returned undefined and
`budgetFor`'s documented fallback was dead code. It is now declared with a
`null` default, exactly as docs/CONFIGURATION.md already specified, so the
default behavior is unchanged: nothing configured means nothing trims.

Two things the diagnosis had not predicted, found and fixed while implementing:

- `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded
  registration site that `mergeReviewerLanes` prefers over the capability
  registry on a slug collision. Editing only the capability files left every
  CLI lane still null. Both sites now agree.
- The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked
  the capability edits; regenerated with `npm run gen:capability-registry`
  rather than hand-edited.

docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one —
today ollama, lm_studio and llama_cpp". That is false as of this change and is
corrected rather than left to rot.

The trim-versus-refuse question the issue raises is deliberately not taken up
here: the refusal path already exists for the case that matters — a reviewer
whose minimum set exceeds its budget is skipped rather than sent a misleading
prompt — and trimming above that floor is the documented, shipped design of the
feature. Changing it would alter behavior for the three lanes that already
work, which is not what the issue asks for.

Fixes #3691

* fix(#3691): document the new global and narrow an invariant this change obsoleted

The full suite surfaced two consequences of giving every CLI lane a budget key.

`review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in
the planning-config reference, which config-field-docs guards. Documented,
including the sentinel semantics a reader needs: a per-lane value overrides the
global, `-1` means unset and inherits it, and `0` means "do not trim that lane"
and is not unset.

The #2797 federation guard asserted that "a lane with no model flag and no host
owns no config keys". That held only because budget keys existed solely on the
three local-server lanes, all of which have hosts. A lane can now legitimately
own a config key for a third reason, so qwen tripped it.

The assertion is narrowed rather than weakened: such a lane must still own no
model key and no host key, and may own at most its own
`review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is
strictly more specific in the dimensions that still matter. Proven to still
bite: hypothetically giving qwen a `review.models.qwen` key fails it with
`model/host: review.models.qwen`. The name and comment cite #3691 for why the
premise changed, so a reader sees a deliberate narrowing, not erosion.

Checked the sibling assertions in that describe block; the other three do not
rest on the obsolete premise and are untouched.

Refs #3691

* fix(#3685): port the write-flag content-change contract to its three sibling sites

#3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which
reported `fs.existsSync(path)` rather than whether the transaction wrote
anything. Three sibling sites carried the identical defect and are ported here.

- `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded.
  `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and
  the flag reports it. #2640/#2974 already fixed `state_updated` at this same
  call site and left this one behind, so the correct shape was adjacent.
- `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` —
  byte-identical to #3685's bug in a different command.
- `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never
  consulting the MILESTONES.md write.

`gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display
and never branches on it, so the flip from always-true to content-based changes
no workflow behavior. Verified by reading the step, not assumed.

One trap found while implementing: the obvious in-memory
`finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s
shipped shape verbatim — gives a FALSE POSITIVE for milestone completion.
`platformWriteSync` normalizes Markdown at write time, and the milestone-closure
transform regenerates `## Current Position` fresh on every call, so its
pre-normalize output always differs from the already-normalized file on disk
even when the persisted bytes are identical. The comparison is therefore made
against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are
left untouched — their repeat-no-op tests pass, so they are not exposed to this
artifact.

`milestones_updated` has no reachable no-op: the MILESTONES.md write
unconditionally appends an entry every call. Only the true direction is pinned,
documented inline rather than faked with a passing test.

Refs #3685

* fix(#3685): compare write-flag content through the writer's own normalizer

An independent reviewer disproved a claim made while porting #3685's contract
to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to
the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`.

`platformWriteSync` normalizes on write — CRLF stripped, blank-line runs
collapsed, a blank line inserted after a heading, a single trailing newline
enforced. Every flag that compares the PRE-normalization in-memory string
against the on-disk pre-image can therefore report a change when the persisted
bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading
the file after the write; the other sites compared raw strings.

All of them now go through one exported seam,
`contentChangedAfterNormalize(filePath, before, after)`, which normalizes both
sides exactly as the writer does. That removes the extra disk read the milestone
workaround needed, and makes the sites agree by construction rather than by
four independent implementations of one rule — the divergence the repo names as
an anti-pattern.

Reachability, stated precisely rather than uniformly: the seam is load-bearing
at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and
`stateUpdated`, where section-rewrite logic genuinely regenerates content into a
different-but-normalization-equivalent shape. At
`updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch
never reassigns `content`, so the raw comparison was already correct there. The
first analysis claimed the reverse; this is the corrected finding.

Also fixes an unsound test premise the remote suite caught. The byte-identity
precondition in `roadmap_updated is false when ROADMAP.md comes out
byte-identical` asserted against a hand-authored, un-normalized fixture — so the
very first write reformatted it and the file could not come back identical. The
fixture is now written already-normalized, so the assertion compares a
normalized pre-image against a normalized post-image and still fails if the flag
regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS
too, and the earlier local check simply never exercised it.

The sibling true-direction and milestone tests were checked for the same premise
and do not share it — they assert `notEqual`, or compare two post-write states
produced through the same normalizing seam.

Refs #3685

* chore(changeset): backfill PR number for #3691 fragment

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:39:51 -04:00
Tom Boucher
c933184b97 enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe

Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel
exemption, degraded-read contract, CLI arm and the plan-authoring contract text.
RED by construction: the module exports it requires do not exist yet.
Executed on the remote runner.

* feat(#3172): require a stated failing direction for every automated acceptance command

Every runnable <automated> command now carries a <fails_when> sibling naming
what output constitutes failure. A command with no expressible failure mode is
not an acceptance test: it reads as rigour and is not falsifiable.

- verify-command-grounding gains a failing-direction probe sharing the existing
  <automated> grammar, MISSING sentinel and walk guard rather than copying them
- gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches
  it and hands the JSON to gsd-plan-checker check 8f
- Dimension 8 detail extracted to references to stay under the agent size cap

Verified on the remote runner.

* fix(#3172): close four review findings in the failing-direction probe

- MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so
  a real command was exempted from the new blocking gate. Tightened the SHARED
  constant rather than adding a second copy.
- Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k).
  Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The
  pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too.
- probePhaseFailingDirections reported status 'ok' when one plan was unreadable,
  conflating 'could not look' with 'nothing to report'.
- Extracted the phase-resolution block both check arms had copied verbatim.

Also corrects a docs/AGENTS.md dimension list stale since #2401.
Verified on the remote runner.

* fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping

The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen
under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars
of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled
where such a rule goes: the planner spawn contract in plan-phase.md, beside
<tracked_source_paths>. The agent file is reverted to origin/next verbatim.

- plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the
  contract there and row 30b guards the freeze in both directions
- plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the
  precedent that two ack sources may never name the same path
- install-tree fixtures regenerated for the three new reference files

Verified on the remote runner.

* chore(#3172): backfill PR number into the changeset fragment

pr:0 -> pr:3825 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:05:11 -04:00
Tom Boucher
7a41248c4f fix(#3685): report phase-complete write flags from the transaction, not the filesystem (#3826)
* test(#3685): failing-first regression coverage for phase-complete write flags

`phase complete` reports `roadmap_updated`/`state_updated` from
`fs.existsSync(path)`, so both read `true` whenever the file merely exists —
including when the transaction wrote nothing. Add the regression tests that
prove it, plus the negative-space and true-direction pins, before the fix.

New in tests/phase.test.cjs:
- roadmap_updated is false when the transaction rewrites nothing (FAILS today)
- state_updated is false when the transaction rewrites nothing, and stays
  false on a third consecutive run (FAILS today)
- both flags are true when the transaction genuinely rewrites (pins the true
  direction so the fix cannot be tightened into always-false)
- each flag stays false when its file is absent

The STATE.md cases pin the clock via GSD_TEST_MODE + GSD_NOW_MS
(src/clock.cts:43-70) because syncStateFrontmatter stamps a
millisecond-resolution `last_updated:` on every write, which would otherwise
make the no-op unobservable.

Also strengthens four pre-existing `=== true` assertions on these fields that
passed vacuously: each now pairs the flag assertion with a content-changed
assertion against a pre-call snapshot, so the `true` is earned.

Refs #3685

* fix(#3685): report phase-complete write flags from the transaction, not the filesystem

`phase complete` computed `roadmap_updated` and `state_updated` as
`fs.existsSync(path)`, so both read `true` for any project that had the file at
all — including a run that rewrote nothing. The flags are the only signal a
caller has that the rollup landed, so a no-op was indistinguishable from a
successful write and a stale ROADMAP went unnoticed until something downstream
read wrong numbers.

Both flags now reflect whether that file's content actually changed in the
transaction, computed at the existing `writes.push({filePath, before, after})`
sites — the same contract `requirements_updated` has honored since #2316-3, and
the same correction #2640/#2974 already applied to `phase remove`.

Nothing about what gets written changes; only what gets reported.

Fixes #3685

* chore(changeset): backfill PR number for #3685 fragment

---------

Co-authored-by: sim <sim@local>
2026-08-24 18:05:35 -04:00
Tom Boucher
596540f864 feat(#3227): publish machine-readable state contract at step boundaries (#3824)
* feat(#3227): publish machine-readable state contract at step boundaries

Adds src/state-contract.cts, a best-effort publisher that writes
.planning/state.json (contract 1.0.0) at 11 step-boundary commands, so
external tools read a versioned contract instead of parsing STATE.md and
ROADMAP.md heuristically.

Composes existing owners rather than re-deriving: phase rows come from a
new locateProgressTable extracted from deriveProgressFromRoadmap (so the
snapshot can never disagree with GSD's own progress counters), milestone
identity from getMilestoneInfo, and next from classifyProject. Owners are
required lazily to avoid the state -> state-contract -> smart-entry ->
state require cycle.

Also fixes a pre-existing defect in scripts/lint-test-file-count.cjs
(maintainer-approved as a second concern): testEffectivePrefix never
stripped the suite qualifier, so 65 dotted test files counted against no
module and 9 mis-bucketed into a shorter one. Allowlist re-baselined for
the 74 files the gate can now see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): backfill PR number into the changeset fragment

pr:0 -> pr:3824 now that the PR exists. Doc-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3227): shape hostile-name fixtures away from the scan corpus

The two hostile-input fixtures used a literal phrase from
scripts/prompt-injection-scan.sh's corpus, so CI's Security Scan redded on
this file. These tests assert that an arbitrary phase name round-trips into
state.json as inert data -- the property holds for any string, so the
injection flavor is illustrative, not load-bearing.

Reshaped to a hyphenated fake instruction tag, which stays hostile-looking
while matching none of the scanner's patterns. Allowlisting the file was
rejected: that mechanism is for suites whose subject IS injection defense,
and it would blind the scanner to this whole file permanently.
See DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): ratchet the state-contract mutation floor to its measured score

The module was registered at minScore 50, the ratchet's minimum permitted
floor for a newly-registered module whose score had not been measured. This
PR's own Stryker shard measured 66.25% (run 32769289750, job 97565813640),
so the floor moves to floor(measured) - 1 = 65, per the rule the registry
documents.

66.25 is below TARGET_MUTATION_SCORE (80), so this stays a ratchet
candidate: raise as the tests improve, never lower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 17:56:02 -04:00
Tom Boucher
fb9823e1e1 fix(#3689): refuse a ledger write when the rendered table disagrees with its JSON (#3828)
* test(#3689): failing-first coverage for the ledger table/JSON agreement guard

`.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is
its source of truth, but nothing checks the two still agree before a write
overwrites the table. `windows append` / `waive` / `fixed` therefore discard a
drifted cell silently, and erase a table-only row entirely, both at exit 0.

Adds to tests/broken-windows.test.cjs:
- five refusal cases that fail today, covering all three write commands, a
  drifted cell, a table-only row, and drift on a non-first row; each asserts the
  typed reason via GSD_JSON_ERRORS and that the file is byte-identical after
  the refusal, so a guard that refuses only after writing cannot pass
- six anti-tightening pins that must stay green: an agreeing ledger, the
  first-write ENOENT path, #2893 trailing-prose preservation, #3657 3-backtick
  fence tolerance, escaped pipes and backslashes in a description, and the
  zero-entry placeholder table
- a fast-check property pinning the round trip the guard depends on —
  extractTableRegion(renderLedger(l)) === renderTable(l.entries) — because a
  false refusal on a clean ledger would be worse than the bug

Fixtures are built by running the real CLI and then perturbing only the table,
so frontmatter and JSON stay consistent and the pre-existing counts cross-check
still passes; a hand-written ledger would pass these for the wrong reason.

Refs #3689

* fix(#3689): refuse a ledger write when the rendered table disagrees with its JSON

`.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is
its source of truth, and `writeLedgerAtomic` regenerated that table on every
`windows append` / `waive` / `fixed` without ever checking the two still agreed.
A hand-edited cell was silently reverted; a row that existed only in the table
vanished entirely. Both at exit 0, with nothing on stdout to say so.

The write seam now compares the on-disk table against
`renderTable(<entries parsed from the on-disk JSON>)` before regenerating
anything, and refuses with a typed `windows_ledger_table_drift` error naming the
drifted row ids and the remedy. Because the check sits at the single write seam,
all three commands inherit it, and the file is left byte-identical on refusal.

Deliberately not enforced in `parseLedger`: hardening the read would break
`windows status` and the ship gate on exactly the ledgers an operator needs to
inspect to diagnose the drift.

Two hazards handled explicitly, both discovered in review of the first draft:

- The pre-image read now distinguishes ENOENT from every other errno, per the
  #1950-H2 fail-closed-on-unreadable invariant `readLedgerOrNull` already
  honors. A bare catch would have let an unreadable pre-image skip the guard
  and write anyway.
- Both the entries baseline and the table extraction pass the pre-image's own
  frontmatter `total_count` to `locateJsonBlock`. Without that hint the
  no-expectation fallback binds to the LATEST fenced JSON array in the file,
  which is the operator's prose block whenever that prose contains one — the
  exact case #2893 exists for — refusing every write on a ledger that never
  drifted. A regression test covers it.

Also extends the CONTEXT.md Broken Windows Ledger glossary entry: the table is
a third projection of the same source, cross-checked at the write seam, and the
frozen REASON enum gains WINDOWS_LEDGER_TABLE_DRIFT.

Fixes #3689

* fix(#3689): bind prose preservation to the pre-image's own ledger block

Found while reviewing the table drift guard: the #2893 trailing-prose
preservation in `writeLedgerAtomic` passed `ledger.total_count` — the
POST-mutation count — as the disambiguation hint for a lookup over the
PRE-image. On an append the pre-image holds N entries while the hint says N+1,
so the hint can never match and `locateJsonBlock` falls through to its
last-array-shaped-span fallback.

When the operator's trailing prose itself contains a fenced JSON array — the
ordinary case #2893 was written to protect — that prose block wins the
fallback. The preserved region is then computed from the prose fence rather
than the ledger fence, and everything between them, including the operator's
own text above the array, is silently dropped on the next write.

Reproduced against the real CLI: a prose block reading "Operator notes above
the array, IMPORTANT DO NOT LOSE THIS TEXT." plus a fenced 3-element array came
back empty after one `windows append`.

Both the prose lookup and the drift guard now share one pre-image-derived
`preImageExpectedTotal`, taken from the pre-image's own frontmatter, so they
bind to the same and correct block. The existing trailing-prose regression test
is strengthened to assert the prose survives byte-for-byte rather than merely
that the command exited 0 — asserting only the exit code is why this was
invisible.

Refs #3689

* fix(#3689): anchor table extraction on the header row, not a line-prefix scan

Independent review found the drift guard could brick a ledger nobody had
hand-edited. `validateDescription` accepts a description containing a raw
newline, and `renderTable`'s cell escaping covers backslash and pipe but not
newlines — so such a description renders a row that physically spans two file
lines, the second of which does not begin with `|`.

`extractTableRegion` bounded the table by walking backward over the contiguous
run of `|`-prefixed lines, so it stopped at that split. In the common case
where the row's tail is the last line before the fence it returned null, and
every subsequent append/waive/fixed was refused with "table region could not be
located" — permanently, with no CLI recovery path, on a ledger that never
drifted. A false refusal is worse than the bug this guard exists to fix.

The region is now anchored on the header row `renderTable` always emits,
running from its last line-start occurrence to the end of the pre-fence text.
The boundary is the fence rather than a line prefix, so a multi-line row is
captured whole, re-renders byte-identically, and compares equal. The header
literal is hoisted to one constant both `renderTable` branches and the
extractor share, so the two surfaces cannot drift apart.

Deliberately unchanged: `cell()` and `validateDescription`. The cosmetic
corruption a newline causes in the rendered table is pre-existing, and either
escaping it or rejecting the input would change what existing ledgers render to
or what input is accepted.

Also closes a coverage gap the standards review raised: the non-ENOENT
pre-image read branch — the one that stops an unreadable file from bypassing
the guard — now has a behavioral test that injects EACCES by monkeypatching
`fs.readFileSync` for that one path and restoring it in a `finally`, never by
`chmod 0o000` (root ignores mode bits, so that would pass with zero coverage).
The #3689 property generator no longer strips newlines out of descriptions,
which is why this was invisible to it.

Refs #3689

* chore(changeset): backfill PR number for #3689 fragment

* chore(changeset): backfill PR number for #3689 fragment

* fix(#3689): terminate the header scan when the match sits at index 0

`extractTableRegion`'s backward search for the table header could loop
forever. On a rejected match at index 0 it set `searchFrom = idx - 1`, i.e.
`-1`; `String.prototype.lastIndexOf` clamps its position argument into
`[0, length]`, so the next iteration searched from 0, found the same match,
rejected it identically, and set `-1` again. The loop made no progress.

Reachable only through the exported `extractTableRegion` — `writeLedgerAtomic`
reaches it after `parseFrontmatterStrict` has already succeeded, so the
candidate region begins with the `---` frontmatter fence and a match at index 0
is impossible. Latent rather than live, but an exported `for(;;)` that can fail
to advance is not something to ship.

Confirmed by running the pre-fix compiled function on
`TABLE_HEADER_LINE + 'X\n' + <a valid json fence>` as a backgrounded child: it
was still alive after five seconds having printed nothing, and had to be killed.
Post-fix the same input returns `null` promptly — correct, since the sole
header occurrence fails the end-of-line test and no valid header exists.

A regression here would stall the suite rather than fail it, so the new test
also asserts the returned value rather than relying on termination alone. No
wall-clock assertion is involved.

Refs #3689

* test(#3034): publish the lane trace before the done-file that releases dependents

`preservesSelectionOrderParallelDespiteCompletionOrder` forces a reverse
completion order with a dependency chain rather than sleeps: each stub lane
waits on `done-<dep>` before finishing. It then ended with

    touch "$RUN_DIR/done-$slug"
    echo "end:$slug" >> "$TRACE"

Those are two unsynchronized operations in separate shell processes. A
dependent's `wait_for_file` unblocks the instant the upstream's `touch` lands,
but the upstream's own `echo` has not necessarily run — so if the upstream is
descheduled between the two, the dependent can run its whole body and append
its `end:` line first. The done-file was published before the state it signals.

Observed on the remote runner as `[end:claude, end:codex, end:gemini]` where
selection order demands `[end:claude, end:gemini, end:codex]`. The failure was
in the fixture's own self-check, before it reached the assertion #3034 exists to
make.

Not a flake and not a wall-clock margin: this branch passed the full suite twice
at 14f494644 and 90c5d7a03, and the only delta in the failing run was one added
test in tests/broken-windows.test.cjs — an unrelated module. Adding load
elsewhere in the suite was enough to invert it, which is what a real race does.

Swapping the pair establishes a genuine happens-before: anything a dependent can
observe is written before the file that releases it. A comment records why, so
the order is not tidied back.

The production path is unaffected and was independently confirmed correct —
`invoke_reviewers` joins every lane with `wait`, then aggregates by iterating
DISPATCH_SLUGS in selection order, reading per-slug result files. It consumes no
completion-order signal at all.

Refs #3034

---------

Co-authored-by: sim <sim@local>
2026-08-24 17:43:25 -04:00
Tom Boucher
a2387a0545 feat(#3034): add opt-in parallel reviewer lanes (#3822)
* test(#3034): failing-first coverage for opt-in parallel reviewer lanes

Executes the real invoke_reviewers dispatch block from review.md against a
stubbed gsd_run seam rather than pattern-matching the workflow text, so the
two properties that actually carry risk are observable: that every lane is
joined before aggregation, and that concurrent lanes cannot tear a line in
gsd-review-lane-results.jsonl.

Concurrency is proven by a barrier fixture, not by elapsed time -- each stub
lane blocks until all lanes have checked in, which can only complete if they
overlap.

Red against the current sequential dispatch, by design.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3034): add opt-in parallel reviewer lanes

Reviewer lanes within one review pass inspect the same immutable plan
snapshot and have no dependency on one another, but were dispatched strictly
one at a time, so a multi-reviewer pass cost roughly the sum of its lanes.
The serialization is a deliberate protection against provider rate limits,
so it stays the default; review.parallel_lanes opts a project out of it.

The loop body is hoisted into run_review_lane so the sequential and
concurrent paths share one body -- two hand-synced dispatch bodies is the
divergence class ADR-2782 spent a phase deleting. Each lane writes a
slug-scoped result file, concatenated in selection order after the join:
concurrent O_APPEND is atomic only below PIPE_BUF, and write_reviews parses
that JSONL to render the models:/model_sources: frontmatter, so a torn line
is a broken REVIEWS.md rather than a cosmetic log defect. Aggregating in
selection order also keeps the artifact byte-identical between the two paths.

The guard is strict equality on "true" and falls back to sequential when
config-get fails -- the opposite polarity from the commit_docs guard,
because failing open here fires the very requests the default prevents.

Also corrects docs/COMMANDS.md and its four locale mirrors, which described
--all as running every configured reviewer in parallel when dispatch was in
fact sequential.

Closes #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3034): de-duplicate dispatch slugs and scope lane locals

Review finding (Standards axis): a slug repeated in SELECTED_REVIEWERS would
put two concurrent background jobs on the same > -truncated per-lane result
file. The shared-append form this replaced could not corrupt itself that way,
so de-duplicating is what keeps the concurrent path no worse than the
sequential one.

Selection de-dupes today -- the roster is a Set and review.default_reviewers
normalizes lowercase-unique -- but reachability analysis is not a contract,
which is the same reason the roster derivation itself is guarded.

Splitting once into DISPATCH_SLUGS also removes the duplicated tr-split the
same review flagged: the dispatch and aggregation loops now share one list,
which is what guarantees they walk the same slugs in the same order. A plain
string accumulator rather than an array, because zsh and bash disagree on
array indexing and this block runs under both.

Also scopes run_review_lane's locals. Not a live fix -- each dispatched call
already forks its own subshell -- but it makes the isolation a property of the
function rather than of the dispatch mechanism happening to fork.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3034): acknowledge review.md growth, drop spent 2295 ack

The differential attribution gate reported review.md growing 4173 bytes
(30712 -> 34885) with no live acknowledgment. Adds the per-PR fragment it
asks for, naming only the one path it reported.

Deleting tests/emitted-drift-acks/2295-resolved-model.json is required, not
opportunistic. That fragment declared review.md and nothing else, and its
ripple is already absorbed into the base, so it is spent -- it can no longer
clear anything, which is why the gate still reported review.md as
unacknowledged. It could not simply be left alone either: two ack sources may
never name the same path, so it blocked this PR's fragment outright.
CONTRIBUTING is explicit that a fragment whose last entry is removed gets
deleted with it, because an empty fragment signals nothing while its presence
reads as a live alarm.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3034): backfill changeset PR number

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:25:22 -04:00
Tom Boucher
a84f756303 fix(#3078): sweep all-spent ack fragments on next, name the collision remedy (#3823)
* fix(#3078): sweep all-spent ack fragments on next, name the collision remedy

`guard-no-ack-on-next` only ever watched the legacy tests/emitted-drift-ack.json.
#2914 exempted the fragment directory on the premise that a persisting fragment
"cannot conflict with any other PR". Fragments do not share a FILE, but they do
share a PATH KEY SPACE, and a path claimed by two sources is a hard failure in
the same script -- so a fully-spent fragment on next owns keys it can no longer
gate, and the next PR to grow one of those paths can declare it neither there
(spent) nor in its own fragment (duplicate). Measured at the sweep: 45 fragments
owning 403 paths, up from 13/272 at triage 19 days earlier.

- `assertNoAllSpentFragments` fails a fragment only when EVERY surviving entry is
  spent against the copy at HEAD^, so a partially spent fragment -- and the
  re-arm-by-appending route #2639/#2993 ship on -- keeps working.
- `ackProse` duplicates the gate's zero-width/whitespace stripping across the
  scripts-ship/tests-do-not line, bounded by a prose-parity test.
- The guard job's checkout takes fetch-depth: 2; at depth 1 HEAD^ is absent and
  every fragment reads as brand-new, i.e. the guard passes vacuously.
- The duplicate-ack error now names both resolutions, since the guard is
  post-merge by design and cannot stop the colliding PR.
- All 45 spent fragments deleted, 0000-legacy-migration.json included, and the
  three tests that pinned its permanence corrected.

Verification is the remote runner (gsd-test), not a local suite.

Closes #3078

* fix(#3078): make the prose-parity test two-sided, cover the git seam, base on the pre-push tip

Three review findings, all fixed:

- The parity test was a tautology: it checked ACK_INVISIBLE against a
  hardcoded list matching its own definition, never against the gate. The
  gate's INVISIBLE and its reason normalizer (hoisted out of diffEmitted as
  normalizeAckReason) are now exported for that sole purpose, and the test
  sweeps 0x00-0xFFFF against both surfaces. Mutation-checked: adding a
  codepoint to one side and not the other now fails.
- resolveBaseRef, readFragmentAtRef and assertUsableBaseRef had zero direct
  coverage -- the tests reimplemented the git reads in a local helper, so the
  ls-tree-vs-show discrimination, the root-commit fallback and the
  option-injection guard were never executed. All are exported and tested
  against real temp repositories now, plus an end-to-end --base-ref subprocess.
- HEAD^ is not 'the state of next before this push'. The default branch allows
  REBASE merges, so one push can carry N commits, and a 2-commit rebase-merge
  whose first commit adds a fragment would be told to git rm it on the very
  push that introduced it. CI now passes github.event.before via --base-ref and
  fetches it explicitly; HEAD^ remains only the local fallback.

Also adds the safe.directory guard every other git call in this repo carries
(#2767), and stops naming the deleted migration fragment by filename in
CONTEXT.md, which tripped lint-removed-but-needed.

Refs #3078

* fix(#3078): keep the fragment directory alive after the sweep empties it

Sweeping every fragment leaves the directory untracked, and check-glossary-refs
then fails: CONTEXT.md references tests/emitted-drift-acks, which no longer
exists. The empty directory IS the intended steady state, so it has to survive
its own remedy.

Adds tests/emitted-drift-acks/README.md documenting the create/use/delete
lifecycle where a contributor actually meets it, matching the existing
tests/qa/smell-acks/README.md precedent. Every reader filters on .json, so the
README is invisible to the gate.

Also sweeps #3809's ack fragment, which the rebase onto origin/next brought in
and the new guard immediately reported as all-spent -- its own remedy applied.

Refs #3078

* fix(#3078): guard the added tests' git calls, drop a second fragment-existence pin

Both defects surfaced by the remote runner (linux-node24, 4/37445 failed).

- The new --base-ref E2E test ran `git rev-parse HEAD` against the checkout
  without the #2767 safe.directory guard. The runner mounts the repo at a path
  owned by another uid, so git refused every operation there with 'detected
  dubious ownership'. Every git call the new tests make now names its own
  specific directory as safe, via one local helper, mirroring safeDirArgs in
  helpers/emitted-runtime.cjs.
- tests/agent-tracked-source-rule.test.cjs pinned the existence and contents of
  the 3645 and 3409 ack fragments. That is a merged PR's paperwork, not live
  behavior: once the growth is in next's baseline the acks are spent and this
  PR's guard sweeps them. The third assertion pinned the hand-appended
  workaround for the exact collision #3078 removes. Deleted; #3645's real
  protection is the two behavioral tests above it, untouched.

Also restores #3809's ack fragment, which merged one commit before this branch.
Deleting an ack in the same window as its introducing PR races any consumer
whose baseline predates it -- the runner's container proved it, resolving
origin/next to 8ed105c8a where the file is still 13847. The backlog sweep is
this PR's scope; that fragment is left for the guard's own first run.

Adds the rule to the fragment README so the class stops recurring.

Refs #3078

* test(#3078): derive the E2E guard expectation from the fragment inventory

The --base-ref E2E test asserted exit 0 while passing the checkout's own HEAD
as the base ref. HEAD-as-base makes every present fragment byte-identical to
itself, so all of them are trivially all-spent and the guard correctly exits 1.
The test only ever passed because the directory happened to be empty when it
was written; restoring #3809's fragment made it fail. The script was right and
the test was wrong.

The degenerate base ref is kept deliberately -- it is what makes 'spent'
trivially true and therefore deterministic -- but the expectation is now
derived from listFragmentFiles() at runtime: zero fragments means exit 0 and
the no-survivors line, N fragments means exit 1 with every name and its git rm.
Proven state-independent by running the suite with the fragment present, with
the directory emptied, and with it restored.

The option-shaped --base-ref rejection is split into its own test, unchanged.

Refs #3078

* chore(#3078): backfill PR number into the changeset fragment (pr:0 -> pr:3823)

---------

Co-authored-by: sim <sim@local>
2026-08-24 15:24:30 -04:00
Tom Boucher
8442d984b9 fix(#3809): route runtime-loaded markdown through the gsd_run launcher (#3815)
* test(#3809): generalize dead-ref guard into a rule table (failing first)

The #2020 guard hardcoded `sdk/(src|dist|handlers)/` — the three dead paths
that had caused that storm. That proved those three paths were gone and said
nothing about the class, so #3809 reproduced the identical Windows find.exe
storm under a different token and the guard could not see it.

Replaces the single regex with a rule table over the same runtime-loaded
markdown surface, adds `commands/` to the scan set (previously uncovered),
and adds rule B: the runtime shim filename must never appear in command
position, because it is not a PATH command and an agent that meets it falls
back to locating the file.

Rule B's matcher is deliberately lenient — the launcher's own resolver
assignment, `node <path>/<shim>` calls, bare paths, and prose that names the
file all stay unflagged, each pinned by a negative-space row.

This commit is expected to FAIL: 50 offenders across 23 files remain in the
tree. The remediation lands next.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): route every workflow call through the gsd_run launcher

50 places across 23 runtime-loaded workflow, agent, reference, and command
files instructed the agent to run the runtime shim by filename. That filename
is not on PATH under any name -- package.json ships gsd-core, gsd-tools,
gsd_run and gsd-mcp-server -- so the call exited 127, the file-shaped token
sent the agent looking for the file, and on Git Bash for Windows the resulting
`find /` walked the entire drive (7268 CPU-seconds in the report) until
somebody killed it by hand.

CONTEXT.md -> Runtime Launcher Module already makes gsd_run the single entry
point: "Canonical space-safe shell preamble (`gsd_run`) used by every workflow
bash block to invoke the GSD runtime CLI." These sites predate that rule --
they trace to 0e6907050 (docs(#195): migrate workflow markdown off gsd-sdk
query), which swapped one non-PATH token for another.

Two further instances of the same class surfaced during remediation and are
fixed here rather than left for later:

  - references/model-profiles.md prescribed `node <shim> effort sync` with no
    path at all; node resolves a bare filename against cwd, so it fails the
    same way.
  - references/universal-anti-patterns.md rule 25 instructed every agent to
    "use <shim>" when shelling out. That rule did not contain the defect, it
    prescribed it repo-wide.

Five "(or legacy <shim>)" parentheticals left dangling by the substitution are
removed; after the rewrite they offered the non-resolving form as an
alternative.

The guard from the previous commit now passes. Its node-prefix exemption was
tightened to require a path separator, which is what exposed model-profiles.

Fixes #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): key the guard on the CLI's whole verb roster, not observed usage

Review found the first cut of rule B repeating the very mistake it exists to
prevent. Its verb set held query, commit and effort -- the verbs that happened
to appear in the tree -- so it could not see `<shim> phase add`,
`<shim> state load`, `<shim> verify ...` or twenty-odd other real single-word
subcommands. A guard that only recognises yesterday's offenders is not a guard.

The set is now the CLI's full advertised roster, unioned from the usage banner
and HOST_COMMAND_ROUTERS (which carries verification, planning, uat, stats,
todo and windows, all absent from the banner).

Widening it immediately caught a live offender the first pass had missed:
references/planning-config.md prescribed `node <shim> worktree set-baseref`
with no path. Fixed here.

Also drops the "a hyphen or a dot means subcommand" heuristic, which was
unsound for prose -- it flagged `built-in` and `v1.2`. Detection now keys
entirely on the roster, testing the first dot-segment so that phase.add and
state.patch still match while prose does not. Both false positives are pinned
as negative-space rows.

Guard verified against the pre-fix tree at origin/next: 52 offenders across 25
files, and 0 after this branch's remediation.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): derive the verb roster from the router; repair launcher parity

Standards review caught the guard repeating the defect it exists to prevent.
Its verb list was a hand-copied literal -- and worse, transcribed from an
INSTALLED older binary, so it was missing 22 verbs this tree actually ships
(websearch, windows, state-snapshot, context-predicates and the dispatch-*
family among them). gsd-tools.cjs already carries three hand-maintained
rosters whose drift is a named defect pinned by the parity test in
tests/commands.test.cjs; a hand-copied fourth was that same defect wearing a
guard's clothes.

The roster is now derived from HOST_COMMAND_ROUTERS + TOP_LEVEL_USAGE, lazily
and memoised, with `query` supplemented explicitly -- it dispatches through
the routing hub ahead of the host-router table, so it appears in neither
export, yet 45 of the 50 offenders used it. A parity test pins the derivation.

Two regressions this branch introduced, both caught by the remote runner:

  - runtime-launcher-parity: rewriting a comment in gsd-research-synthesizer.md
    put a `gsd_run` token at line 65 while the canonical preamble sits at 158,
    breaking "exactly ONE preamble, before the first gsd_run call". The comment
    is descriptive and needs no command token at all; it now names none.
  - The #2751 guard's PROSE_ALLOWLIST entry for that same line went stale once
    the line stopped carrying a bare mention. Pruned, exactly as that guard's
    own stale-entry test instructs.

Also corrects git-planning-commit.md, where the first pass rewrote only the
trailing "legacy" clause and left the sentence reading backwards.

Note the #2751 guard and this one are complementary, not duplicates: its regex
requires whitespace immediately after `gsd-tools`, so it cannot match the
`.cjs` form, and this one only matches the `.cjs` form.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2751): extend the bare-command guard to references/ and commands/

The #2751 guard has only ever scanned agents/ and gsd-core/workflows/. Two
runtime-loaded directories were never in its scan set, and 47 bare
`gsd-tools <verb>` calls had accumulated there unseen -- the same defect that
guard exists to catch, in the rooms it never entered.

  - gsd-core/references/: 37 calls, all rewritten to gsd_run. references are
    fragments inlined into a parent that defines the launcher, which is why 21
    of the 22 files already using gsd_run carry no local preamble.
  - commands/gsd/: 10 operative calls rewritten. The remaining 10 are
    descriptive prose ("resolved inside the workflow via ...") and are
    allowlisted with reasons, bringing PROSE_ALLOWLIST to 15.

commands/ also came under launcher propagation. sync-runtime-launcher.cjs
walked only WORKFLOWS_DIR and AGENTS_DIR, so every preamble under commands/
was a hand-pasted copy nothing propagated and no test checked -- graphify.md
had accumulated five. It now walks COMMANDS_DIR too, which collapses those
five to the canonical one-per-file, and runtime-launcher-parity gains a
(B-commands) arm mirroring (B-agents) exactly so the placement stays honest.

The parity arm keys on shell blocks, so commands/gsd/workstreams.md and
config.md -- which name gsd_run only in inline backtick prose -- are exempt,
as they should be. gsd_run is itself a shipped npm bin, so those inline
instructions resolve from PATH exactly as the gsd-tools form they replace did.

skills/ is deliberately NOT added to either guard's scan set: it is generated
from commands/ and pinned by lint:generated-sync, so guarding the source
guards both, and scanning the mirror would double-report every future
offender. Regenerated here.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3809): acknowledge the one emitted file this change grows

The emitted-attribution gate failed on the previous sha: gsd-research-synthesizer.md
grew 3 bytes (13847 -> 13850) with no acknowledgment. The substitution SHRANK the
other 19 emitted files, which is why the growth arm was not expected to fire at all.

The 3 bytes are unavoidable. Line 65 is a descriptive comment inside a fenced block;
naming any command there puts a gsd_run token ahead of the file's canonical preamble
at line 158, which runtime-launcher-parity's (B-agents) arm correctly rejects. So the
comment names no command and says where the config is actually loaded instead, which
reads longer than the token it replaced.

Acks only the path the gate reported, per the fragment rules.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#2751): drop the commands/ half — three contracts pin it in place

The remote runner refuted the commands/ extension outright. Reverting it and
keeping the references/ conversion, which passed.

What broke, all of it caused by bringing commands/ under launcher propagation:

  - graphify.md's five per-block preambles are LOAD-BEARING, not accumulated
    drift. tests/graphify-visualization.test.cjs extracts individual Step-3
    shell chains and executes them standalone, so each fenced block needs its
    own definition of gsd_run. Collapsing them to the canonical one-per-file
    produced `bash: gsd_run: command not found`, exit 127, across four tests.
    The "define once per file" contract holds for workflows and agents because
    nothing extracts their blocks in isolation; commands/ is not like that.
  - explore.md broke "the preamble that DEFINES gsd_run must appear before the
    first USE of gsd_run anywhere in the file".
  - tests/gsd-tools-path-refs.test.cjs (#1766) ASSERTS that
    commands/gsd/workstreams.md contains the literal string
    `gsd-tools query workstream.list`. Rewriting it to gsd_run contradicts a
    test that pins the opposite, so the two guards disagree about that file by
    construction.

So commands/ is not a scan-set widening. It needs those contracts reconciled
first, and that is its own change. SCAN_DIRS keeps gsd-core/references/ and
drops commands/, the ten commands/ allowlist entries go with it (back to 5),
and the reasoning is recorded in the guard itself so the next person does not
rediscover it by burning a matrix run.

commands/gsd/import.md keeps its #3809 fix — that one is the .cjs form this
PR exists to remove, and it is untouched by any of the above.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#3809): restore explore.md's Step 1 preamble placement

Running the launcher sync script processed workflows/ and agents/ too, not
just the commands/ directory the run was aimed at, and it MOVED
gsd-core/workflows/explore.md's preamble from Step 1 down to Step 3.

The script inserts into the first bash block that USES gsd_run. explore.md's
Step 1 block only DEFINES it, and that placement is deliberate -- the file
says so on the line above: "Placed in Step 1 rather than Step 3 so declining
the research offer cannot leave Step 5's commit call unbootstrapped."
tests/explore-command.test.cjs pins it.

explore.md carried no #3809 offender, so reverting it costs this fix nothing.
This was collateral from invoking the sync script at all, not from the
COMMANDS_DIR change, which is why the earlier commands/ revert did not catch it.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3809): backfill PR number into changeset fragments

pr:0 -> pr:3815 for both fragments now that the PR exists.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): drop the hand-rolled regex escaper CodeQL flagged

CodeQL raised js/incomplete-sanitization (HIGH) on the guard's pattern build:
`SHIM.replace(/\./g, '\\.')` escapes the dot and nothing else, so it does not
escape backslashes. It blocked PR #3815.

The repo already bans this shape -- local/no-adhoc-regex-escape exists exactly
to stop hand-rolled escapers, with the canonical one in src/pattern.cts. Rather
than reach for that helper, the pattern now carries no escaping logic at all:
SHIM is a compile-time constant whose only metacharacter is the dot, so the
regex source is spelled out literally. The generated source string is
byte-identical to what the replace() produced, verified before and after --
0 offenders on this tree, 52 against origin/next, unchanged.

A drift pin asserts SHIM_PATTERN still matches SHIM exactly, and that the dot
is escaped rather than acting as a wildcard, so the two cannot separate.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:47:37 -04:00
Tom Boucher
8ed105c8a4 fix(#3684): resume verified-unmarked phases at update_roadmap (#3814)
* test(#3684): failing-first rows for the verified-unmarked resume

* fix(#3684): resume verified-unmarked phases at update_roadmap

* test(#3684): heading-shaped roadmap fixture, plain phase.complete calls

* fix(#3684): fit under the pre-phase-6 margin, fix pins and verify call

* fix(#3684): padding-normalize the marked-complete join, assert STATE idempotency

* test(#3684): anchor fixes, node jq mirror, characterized STATE delta

* chore(#3684): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 10:58:04 -04:00
Tom Boucher
4b84be1da4 fix(#3683): wire gated learnings extraction into completion, align copy path (#3810)
* test(#3683): failing-first rows for learnings source resolution and wiring pins

* fix(#3683): wire gated learnings extraction into completion, align copy path

* test(#3683): register the learnings suite in the docs-guard lane, drop unverified markers

* fix(#3683): close review findings — per-item parsing, readdir guards, docs paths

* fix(#3683): route phase enumeration through the locator seam, fix assertion targets

* fix(#3683): merge execute-phase ack into the 3003 fragment, fix fidelity targets

* fix(#3663): replace the spent execute-phase ack entry with the 3683 re-arm

* chore(#3683): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 09:23:01 -04:00
Tom Boucher
31fcb833ec fix(#3679): gate pr-branch verify on planning-tree deletions (#3803)
* test(#3679): failing-first rows pinning planning preservation and the verify deletion gate

* fix(#3679): gate pr-branch verify on planning-tree deletions

* test(#3679): extract hashes via rev-parse and de-vacuate the pure-code pin

* fix(#3679): close review findings — merged ack, pinned prose gate

* fix(#3679): close two-axis review findings — no-renames gate, structural pin

* chore(#3679): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 03:54:02 -04:00
Tom Boucher
9d65cd5404 fix(#3664): warn when config-dir targets a foreign-agent destination (#3794)
* test(#3664): failing-first rows for the config-dir foreign-agent warning

* fix(#3664): warn when config-dir targets a foreign-agent destination

* test(#3664): fold the foreign-agent warning rows into the install-regressions suite

* fix(#3664): close review findings — kimi-agents kind, gsd.md ownership, e2e gate

* test(#3664): sync boolean call sites and the path-vocab registries

* chore(#3664): backfill changeset pr number

* test(#3663): skip the posix case-pin on win32 where folding is the fix

---------

Co-authored-by: sim <sim@local>
2026-08-24 02:45:28 -04:00
Tom Boucher
314ea20fa4 fix(#3663): fold path casing only on win32 in the w027 active-worktree check (#3793)
* test(#3663): failing-first rows for w027 path-casing normalization

* fix(#3663): fold path casing only on win32 in the w027 active-worktree check

* fix(#3663): close review findings — seam-owned compare key, deterministic case pin

* chore(#3663): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:57:51 -04:00
Tom Boucher
4af59f8dd3 fix(#3662): resolve managed hook node runners at hook-fire time (#3790)
* test(#3662): failing-first suite for runtime-resolving hook runners

* fix(#3662): resolve managed hook node runners at hook-fire time

* fix(#3662): close review findings and document the resolver

* fix(#3662): close adversarial and security review findings

* chore(#3662): backfill changeset pr number

* test(#3662): honor win32 skip return and platform-aware sh runner pin

* test(#3662): pin the bare win32-claude sh-hook shape omitting the bash runner

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:07:06 -04:00
Tom Boucher
cf15682d1c enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules

Stage banners, checkpoints, completion and error panels used fixed-width
runs of box-drawing characters -- a 53-column heavy rule and a 62-column
double-line box. Those runs are ordinary text to a Markdown-rendering
host, so in a narrower pane they wrap and the border comes apart from
the heading it framed.

Shipped content now emits an ATX heading for a titled section and a
blank-line-delimited --- for a break between sections, both of which
adapt to the available width. The same convention is applied to the
three code sites that built these strings at runtime: the UAT
checkpoint renderer, the milestone-close audit report, and the TDD
review checkpoint table.

Removing the box also removes its only reason to exist -- the
east-asian-width padding helpers that kept its right border aligned
(checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE,
CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged.

The convention is specified in gsd-core/references/ui-brand.md and
enforced across all shipped content by tests/responsive-separators.test.cjs.

Refs #3028

* test(#3028): pin the heading form in checkpoint and audit-report assertions

These suites asserted the exact box borders and the 62-column padded
banner interior. With the box gone they assert the ### heading form,
the --- break and the bolded instruction line, and each now carries a
positive assertion that no box character remains -- which is what pins
the fix rather than merely tolerating it.

Language coverage is converted, not dropped: Japanese, Chinese, Korean,
Hindi and Arabic all still assert their rendered banner, and the Arabic
case still asserts the RTL directional isolates the box removal must
not disturb. Adds a case for a banner longer than the old inner width,
which previously produced a ragged border and now has none.

Refs #3028

* chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec

The checkpoint_protocol display spec described the drawn box; it now
describes the heading, the --- break and the bolded action prompt,
which costs 22 bytes (40111 -> 40133, 827 under the cap).

Appended to the existing #3370 fragment rather than filed as a new one:
a growth ack keys on the bare filename and #3370 already declares
execute-plan.md, so a second source naming it would be a hard
duplicate-key error. Same supersede-by-append route #3370 took for the
spent #2652 fragment.

Refs #3028

* docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference

Review found three things.

The rule as first written demanded a blank line above AND below every
---. Only the one above is load-bearing: it is what stops CommonMark
reading the rule as a setext underline for the line above. The one below
is cosmetic, because a thematic break is a leaf block. The rule now says
that, with the reason, instead of asserting a stricter form the content
does not keep.

The zh-CN reference had received the mechanical box-to-heading swap but
none of the prose behind it: it still claimed a 62-character checkpoint
width and still listed --- among forbidden mixed banner styles, so it
contradicted the convention it was translating. It now carries the
separator section, the setext reasoning, the unconditional-vs-per-runtime
rationale and a corrected anti-pattern list, in Chinese.

The user guide asserted that a heading is not a degradation anywhere.
That is an assertion, not a demonstration. It now says what was actually
traded away in a plain terminal, points at the recorded rationale, and
invites the report that would justify the capability flag instead.

Refs #3028

* chore(#3028): backfill changeset PR number

Refs #3028

---------

Co-authored-by: sim <sim@local>
2026-08-23 22:38:12 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Behruz Nassre Esfahani
622f43353c fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode

The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.

Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.

The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.

`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.

Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.

Fixes #3299

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): reconcile the canonical schema table and the stale acceptance test

Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.

1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
   schema reference for the tracer task-type contract, and its Task-types row
   still claimed interactive runs unconditionally present a
   checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
   first round; this one was missed, so the authoritative reference was the
   wrong answer. The row now carries the human_verify_mode-conditional
   behavior and points at the canonical precedence chain.

2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
   tracer ROW EXISTS, never its content, which is why CI could not see the
   drift. It now asserts the row's actual claims and rejects the pre-#3299
   wording. Separately, the #1945 acceptance test named 'interactive run emits
   checkpoint:human-verify after the tracer' kept passing only because its
   substrings still occur in the fallback clause, while its name asserted the
   opposite of shipped behavior. Renamed and narrowed to what #1945 still
   guarantees, plus a new interactiveIsConditional pin so the unconditional
   prose cannot be restored under a passing substring check.

3. plan-md.md's <verify> row now documents that the legacy bare-text form
   (valid, and still shown at :179) does not reach the #3299 auto-continue —
   only a <verify> carrying <automated> does — so the benefit is silently
   unreachable for tracers using that format.

Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions

Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.

MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.

MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:

- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
  'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
  same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
  appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
  presence. Deliberately brittle: CONTEXT.md names that table the canonical
  schema reference, so a wording change must be a conscious edit in both
  places.

Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs

Peer review round 4. Blacklisting did not hold, twice over:

- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
  forms in the auto-continue clause. Round 4 defeated that by appending
  'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
  before expansion' — none of the banned tokens, same restored interruption
  after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
  '<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
  legacy bare form operative. 107/107 passed across tracer, planner and the
  three size-cap suites.

Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.

These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.

Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): strip comments, require uniqueness, pin whole regions

Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):

- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
  wrong copy: every extractor selected the commented decoy. Worked against the
  planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
  checkpoint_protocol before expansion, regardless of the mode-specific rules
  below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
  approval' below the canonical table. The pinned text was untouched, so
  equality held while the shipped meaning inverted.

The shape that holds, applied to every operative surface:
  1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
  2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
     cannot hide behind a correct first one;
  3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
     suffix can override what the pin proves.

Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.

Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.

Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): drop the superseded exact-placeholder planner assertion

Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.

Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.

Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner

Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:

- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
  internally uses the repo's interleaved fence/comment scanner. Instrumenting
  candidate lines as throwaway predicate declarations borrows that scanner with
  no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
  edits silently turned the guards into decoy checks:
    * a forgotten '-->' comments the live rule through to EOF, and the
      balanced-only stripper still saw and accepted the commented rule;
    * a normal fenced documentation example of the rule, plus a whitespace-only
      reformat of the live list item, made the selector choose the example.
  Neither needs intent. A dangling comment is a typo; a fenced example is good
  documentation. Together they reproduce exactly the accidental drift #3299 came
  from — with CI green.

The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.

Verified both ordinary-edit scenarios now fail the suite (each was green before).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): close the operative-selection gaps the maintainer blocked on

trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.

1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
   replaced a matched candidate with an UNINDENTED marker regardless of the
   original line's indentation. A 4-space-indented CommonMark code block is not
   skipped by parsePredicates (it accepts indented declarations by design), so
   stripping the indent PROMOTED an indented decoy to operative — the exact
   inversion of the guard's purpose. The marker now preserves the original
   indent, and a candidate that is itself indented 4+ spaces is never injected.

2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
   pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
   pollute the count. Now filters on a Set of the indexes actually injected on
   this call.

3. RAW FENCE SELECTION (planner). The template test matched the first raw
   ```xml fence after the marker with no fence/comment awareness — the one
   selection in the suite that was not operative-aware — so a commented-out
   decoy template between the marker and the real one would be selected while
   the live template regressed. The opener must now be operative AND the first
   non-blank line after the marker.

4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
   a fenced example containing a ### / <type line truncated the pinned region
   early — a false FAILURE on a legitimate doc edit. End anchors now go through
   the same operative filter as start anchors.

Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.

Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): allow-list operative indentation; pin marker provenance

Review round 9.

BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", "  \t" and "   \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.

MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.

MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.

Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.

KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge

The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.

Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.

Verified: emitted-attribution 175/175 against origin/next@be9329b10.

* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain

Addresses the review on #3390 (B1-B3, M1-M4, minors).

B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.

B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.

B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.

M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.

Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.

Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): correct the size record the 08-22 merge falsified

Review round: one Major, four Minors.

Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.

Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.

Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.

Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.

Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): cite the plan-md example by section, not by a drifting line

Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.

Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.

Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): answer fence-delimiter liveness by insertion, not replacement

Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.

`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.

Measured against the real file rather than argued:

  agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
  262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
  example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
  that question.

It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:

  HEAD as-is                  | anchor 260 | openIdx 262 | ASSERTION PASSES
  +1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***

Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.

The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.

Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.

`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.

Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.

Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.

Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): guard the delimiter class, match the scanner, pin the assignment

Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.

1. The mode-read pin survived the regression it exists to catch.
   `READ` matched the config-get substring only, so rewriting the shipped line
   as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
   nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
   is back with the suite passing. The regex now requires the assignment. A
   lookahead after `end-of-phase` closes the other half: the bare prefix also
   accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
   variable in agents/gsd-executor.md now turns that row red, and did not before.

2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
   It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
   and arbitrary info strings all walk past any list short enough to write down
   — the guard was added precisely because one such regex had already slipped
   through. Now matched against the lines the regex actually selects in the
   document, which cannot go stale and cannot miss a spelling nobody thought of.
   Four such spellings pinned as rows.

3. `isOperativePosition` disagreed with the scanner it delegates to.
   For `<!-- closed --> real content` it stripped the span, found surviving
   content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
   whose trimmed text starts with `<!--`, balanced or not, before it considers
   fences at all. Verified directly against parsePredicates. It now applies the
   scanner's own rule instead of out-reasoning it. Latent for the present caller
   (its anchored ```xml shape cannot match a comment-prefixed line), real in
   general.

Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.

Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.

Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-23 18:43:53 -04:00
Behruz Nassre Esfahani
a44d513566 fix(#3712): confine in-process installs to a sandboxed HOME (#3725)
* fix(#3712): confine in-process installs to a sandboxed HOME

A runtime kind may declare a global `home` override resolved from os.homedir()
rather than from the caller's configDir — codex's skills kind (`home: ".agents"`,
ADR-1239 / #2088) is the only live case. Sandboxing configDir/targetDir does not
contain it, and assertDestWithinConfigHome cannot see the class: that gate
confines a destSubpath to whatever root it is handed, and here the root IS the
escaped home. So an in-process caller that forgot to sandbox HOME wrote to, and
pruned gsd-* entries from, the developer's REAL ~/.agents/skills.

tests/agent-descriptor-parity.install.test.cjs's K1 loop did exactly that: it
iterates every agents-kind runtime (codex included) with a sandboxed targetDir
and an un-sandboxed HOME. Reproduced against a canary home on next @ adb46cdd8 —
71 gsd-* skill dirs deleted, a foreign `cloudflare` skill surviving, suite still
exit 0. It is silent because the runtime's own config home is untouched, so the
manifest keeps reporting a healthy install.

FIVE writers resolve a kind `home` and then destroy under it. Three are reachable
today — installRuntimeArtifacts, uninstallRuntimeArtifacts (install-engine.cts)
and applySurface (surface.cts). Two are descriptor-dependent and guarded against a
future descriptor change rather than a present escape: installOpencodeFamilySkills
(behind the combined-family early return) and installAgentsKindStandalone. Those
two are scoped to the single kind each destroys — passing the whole layout made
codex's unrelated skills override trip a writer that never touches it.

- src/test-home-guard.cts: refuse when a run under a test runner cannot be shown
  to have sandboxed HOME. NODE_TEST_CONTEXT (set by `node --test`) gates it, so
  installs outside a Node test context are untouched; GSD_TEST_MODE is unusable,
  as several candidate files including the offender never set it. Homes are
  compared by FILESYSTEM IDENTITY (st_dev + st_ino), not by pathname:
  path.resolve() resolves neither symlinks nor case, and realpath returns a
  canonical pathname that two routes to one directory can still disagree on (bind
  mounts). Verified on macOS/APFS — HOME=/users/<name> made the strings differ
  while naming the same directory, and the lexical form ALLOWED a write into the
  physical real home. FAILS CLOSED: a pair is "different" only when both identify,
  or one is definitively absent (ENOENT/ENOTDIR) while the other identifies; every
  other errno is "cannot tell" and refuses. Only when neither home identifies is a
  marker consulted, and it carries the sandbox PATH and must equal the home in
  effect — a boolean checked first let an ambient or stale value disarm the guard.
- helpers: promote sandboxHome() out of its two byte-identical private copies,
  which is also what makes them record the sandbox; the three withFakeHome()
  helpers record it too. The marker NAME is duplicated as a bare string rather
  than required from the compiled guard, keeping helpers.cjs's documented
  no-built-lib-at-import-time contract; a test pins the two together.
- agent-descriptor-parity: sandbox HOME across the K1 loop.
- helpers-process-isolation: #3156's canary asserts on <home>/.gsd only, and its
  `--cursor --local` spawn cannot reach `.agents` at all, so an assertion added
  there would pass with all confinement removed. Add a discriminating row — a
  `--codex --global` spawn against a seeded ambient home — which also asserts the
  runtime still declares the override. Its check is a sampled inventory (dir names
  + each SKILL.md), not a tree compare.
- install-write-confinement: predicate rows through the deps seam, covering the
  symlinked HOME, ambient and stale markers, and each sameDirectory branch
  (both-identify, one-absent, neither-identifiable), plus wiring rows that drive
  the REAL entrypoints so deleting a guard call site is red.

Verified: guard fires end-to-end against a real un-sandboxed HOME (exit 1, zero
deletions); the case-variant fail-open reproduced on APFS before the fix and
refuses after; K1 file 29/29 green with skills intact; mutation-tested — each of
the three reachable call sites, lexical-only comparison, and treating an unknown
errno as "absent" each take exactly one row red, with every mutation echoed back;
the process-isolation row negative-controlled by reverting installerEnv to its
pre-#3156 leak (16/0 -> 13/3); a full npm test leaves ~/.agents/skills at 71.

Stated residual: the two descriptor-dependent writers have no wiring test, because
no runtime declares a `home` override on those kinds and neither can be exercised
without inventing a descriptor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3712): add changeset

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3712): let a sandbox nested inside the real home through the guard

All six Windows shards of #3725 failed on legitimately sandboxed
destinations. On Windows os.tmpdir() is %LOCALAPPDATA%\Temp — inside the
user's home — so every sandbox a test creates is a descendant of the real
home, and "does this land inside the real home?" answers yes for the safe
case and the dangerous one alike. POSIX conceals this: /tmp and
/var/folders both sit outside $HOME.

Add the missing conjunct: a destination inside the real home is allowed
only when it also sits beneath a HOME that was sandboxed away from the
passwd home. Both halves are required — dropping the first re-admits a
plain un-sandboxed install, and dropping the second decays into the
"is HOME sandboxed?" check the module rejects, which a layout resolved
before the sandbox walks straight through. Each is mutation-proven by a
row that goes red without it.

Also covers the two fail-closed branches of the new exemption, which
survived mutation to `true` with the suite green, and avoids `<user>` in
a docblock — the prompt-injection scanner reads it as a delimiter tag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): close three writer/rollback gaps found reviewing the whole PR

Cross-AI review of the full PR (not just the round's delta) surfaced
three ways the guard could still be defeated:

- The nested-sandbox exemption trusted the SPELLING of a destination.
  With HOME sandboxed to a directory inside the real home — legitimate on
  Windows — an aliased `.agents` (symlink, junction, subordinate bind
  mount) beneath it redirected an allowed path into the real home. Decide
  containment on the path the write RESOLVES to: walk up to the nearest
  existing ancestor, canonicalize, re-append the tail.

- `migrateLegacyDevPreferencesToSkill` is a SIXTH writer that resolves a
  skills-kind `home` override. It creates rather than prunes, which is
  why it was missed, and `_runLegacyInstallMigrations` runs it before
  `installRuntimeArtifacts`' own assertion. Guarded, scoped to that kind.

- Worst of the three: `bin/install.js` snapshots the resolved skills root
  before installing, and its outer catch rolls back by deleting and
  recreating every snapshotted `gsd-*` directory there. The guard's own
  throw landed in that catch, so refusing an un-sandboxed codex install
  provoked exactly the mutation the guard exists to prevent. Refusals are
  now marked and rethrown without rollback — nothing was written, so
  there is no partial install to undo. Every other error still rolls back.

Also carries the sandbox marker into `installSpawnEnv`, so spawned
installers are not refused on passwd-less CI images, and corrects three
claims that no longer hold: "every writer" (six, and named), the
unconditional "fails CLOSED" (the passwd-less marker branch is a
deliberate weakening, and TOCTOU is out of scope), and the assertion that
Windows os.tmpdir() is always %LOCALAPPDATA%\Temp (Node honors TEMP/TMP).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the guard's two limits instead of overclaiming

Round-2 review found the prose had drifted ahead of the code. Corrected,
with no behavior change:

- The module still said FIVE writers; there are six, and the sixth is
  now named along with why it was missed (it creates rather than prunes)
  and why it carries its own assertion (it runs before the main one).
- The canonicalization docblock listed subordinate bind mounts among the
  aliases it closes. It does not close them: a bind mount is not a link,
  so realpath keeps the mount-point spelling. `sameDirectory` already
  recorded that limit; the new helper now inherits it explicitly rather
  than contradicting it. Closing it needs mount-table introspection.
- "FAILS CLOSED" was unqualified while the passwd-less marker branch is
  a deliberate weakening — with no passwd entry, nothing can contradict a
  marker naming the real home.
- "Refuses BEFORE any write" was too broad: legacy install migrations run
  ahead of the layout-driven ones, which is exactly why the two
  rollbackInstallerMigrations() calls still execute before the rethrow.
  Only the codex skills-root rollback is skipped, and that is the only
  _codexPreConfigRollback() call site — applySurface is never called from
  bin/install.js and uninstall cannot reach it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): sandbox HOME in the opencode-family home-override parity rows

The last two Windows failures, and the same platform asymmetry in a
different disguise. This row drives a skills-kind `home` override on
purpose — precisely what the guard polices — but relied on the override
temp dir happening to sit outside the real home. It does on POSIX
(/tmp, /var/folders); on Windows os.tmpdir() is under %USERPROFILE%, so
the guard correctly refused and only Windows went red.

Declare the sandbox instead of depending on the platform: HOME becomes
the override itself, which is the home the call writes under. This is the
fix the guard's own message prescribes, applied to the test rather than
to the guard.

Both failing Windows shards fail on exactly these two rows and nothing
else; every other shard is green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): make sameDirectory answer NO when it cannot tell

Review Major 1. sameDirectory()'s only caller is the passwd-less marker
branch, which reads a `true` as permission to PROCEED:

    if (marker && sameDirectory(marker, osMod.homedir())) return;

The fallthrough returned `true` whenever neither side identified — two
absent paths, or two stats failing EACCES/EPERM/EIO on a locked-down
host — on the reasoning that "cannot tell" should make the caller refuse.
That reasoning was inverted with respect to this caller: it turned the
passwd-less escape hatch into an unconditional bypass for any marker
value at all, on precisely the hosts the fallback exists to serve. Only
two things now answer yes: one resolved pathname, or two readable
identities that match. Restoring the old fallthrough takes the new row
red.

Also from review:

- Major 2 asked whether st_dev/st_ino discriminate directories on
  Windows, where Node derives them from BY_HANDLE_FILE_INFORMATION. The
  whole guard rests on that primitive, so assert it rather than argue it:
  a row comparing two distinct temp directories, and one directory
  reached by two spellings. It runs on every platform in the matrix, so
  Windows answers the question itself.

- Minor 1: the refusal now names the real home it compared against, not
  just the destination it refused. That is the one fact needed to tell a
  true positive from a false one, and its absence is what made the
  Windows case a CI-log dig rather than a glance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): refuse a HOME that merely spells the real home more widely

Review round: one Blocker, four Minors, a Nit.

N2 (the one with teeth) — a destination's ancestor chain is linear, so
"inside the real home AND inside the effective HOME" admits two
arrangements, not one. The intended `effectiveHome ⊂ realHome` is the
Windows temp shape; `realHome ⊂ effectiveHome` — HOME at /Users, /home,
C:\Users — is not a sandbox at all, it is the real home reached by a
wider spelling, and it was exempting a stale destination pointing
straight at ~/.agents. Third conjunct added; the docblock no longer
claims two conditions suffice. Removing the conjunct reds the new row
and nothing else.

N4 — the migration guard resolved its OWN layout, and without
capabilityRegistry, so a registry-dependent descriptor could make it
vouch for a path the migration does not write: a guard reporting safe
while the unsafe write proceeds. It now guards the destination already
resolved by _resolveDevPreferencesSkillTarget, keyed on
`installRoot !== targetDir` — which is exactly the condition under which
a `home` override was declared, read off that same result.

N1 — CONTEXT.md gains the Test Home Guard Module glossary entry that
contributor-standards.md requires of a new Module. Not CI-enforced, so
green CI was never evidence it was met.

N3 — the docs/INVENTORY.md row was misfiled between install-fs-adapter
and install-model-override-resolver; the table is alphabetical and the
manifest already had it right. Moved, and its text now names six writers
and the third conjunct.

N5 — applySurface's signature docblock was two parameters stale; this PR
added the second of them.

N6 — the duplicated rollbackInstallerMigrations() adjacent to the new
rethrow: two identical consecutive calls, not two phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): guard the sixth writer, and close two false-ALLOW paths

Review round 3 (NEW-1, NEW-2) plus three defects Codex found in the
whole-PR pass, each reproduced before it was fixed.

NEW-1 — migrateLegacyDevPreferencesToSkill called the guard with no
`deps`, so it bound real os/process.env and could not be wiring-tested
the way the other three reachable writers were. It now takes the same
optional `deps: { os?, env? }` tail parameter. The wiring block gains
the missing fourth row, and a fifth pinning the ALLOW half; the test
file's header docblock said "FIVE writers ... the three reachable
today", contradicting the six/four statement this PR already put in
src/test-home-guard.cts, CONTEXT.md, docs/INVENTORY.md and the
changeset. Both directions of the guard's condition now fail a row
when broken — previously neither did.

NEW-2 — derivesFromSandboxedHome's docblock claimed "THREE conditions
are required, and no two of them suffice". False for {2,3}: isInside is
reflexive, so whenever conjunct 1 fires conjunct 2 already returns
false on its own. Reworded as a fast path, which is what it is.

Codex 1 (false ALLOW) — on a host with no readable passwd entry the
marker branch returned as soon as the marker matched the effective
HOME. That attests a caller sandboxed HOME and says nothing about where
an already-resolved destination points, so a layout captured before
sandboxHome() — still naming the real ~/.agents — was waved straight
through: the same stale-layout shape the primary branch refuses by
design. The marker must now identify AND contain every destination.

Codex 2 (false ALLOW) — `installRoot !== targetDir` was the stand-in
for "the skills kind declared a home override". The two are not
equivalent: the inequality is false when the override resolves onto
targetDir itself, which is exactly a configDir of $HOME/.agents. The
guard was skipped and SKILL.md written into the real home under a test
runner. _resolveDevPreferencesSkillTarget now reports hasHomeOverride
off the same resolution instead of inferring it from two paths.

Codex 3 (prose) — the shared refusal message claimed every guarded
writer prunes; the migrate writer only creates. The changeset headline
claimed in-process installer calls can no longer reach the real home,
which is wider than the guard: writeNonClaudeDefaults still writes
~/.gsd/defaults.json through os.homedir(). INVENTORY's and CONTEXT's
fail-closed sentences omitted sameDirectory's pathname-equality
shortcut. All four narrowed to what the code does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): let the sandbox marker follow an overridden HOME

Found by Codex in the whole-PR pass. installSpawnEnv spreads
`overrides` last so an explicit HOME wins — deliberate, and its
docblock tells callers needing per-spawn isolation to pass their own
{ HOME, USERPROFILE }. But the #3712 marker was set before that spread,
so such a caller got HOME=<theirs> and marker=<helper default>. On a
host with no readable passwd entry the guard compares the two and
refuses a legitimately sandboxed spawn — tests/install.test.cjs:7143
and install-shared.cjs's own runInstaller both take that path.

The marker is now derived from the final HOME unless the caller
supplied one explicitly. The contract test asserted HOME after an
override but not the marker, which is why it stayed green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#3712): name the shipped guard condition, not the deleted one

Review round 4 of #3725. Two artifacts this PR adds still described
`target.installRoot !== targetDir` in the PRESENT tense as the live guard
condition on `migrateLegacyDevPreferencesToSkill`. The shipped condition is
`runtime && target.hasHomeOverride` (src/install-engine.cts:510).

This is not ordinary doc drift. The named condition is the exact false-ALLOW
the previous round closed: a `home` override resolving onto `targetDir` — a
configDir of `$HOME/.agents`, which is where codex's override points — makes
the inequality FALSE while the override is declared, so the guard was skipped.
A maintainer reading CONTEXT.md:290 as authoritative would believe the guard
still skips that case.

  - tests/install-write-confinement.test.cjs — the ALLOW-half row's comment.
    Its "teeth" rationale is unchanged and still correct as written.
  - CONTEXT.md:290 — the Test Home Guard Module glossary entry, a documented
    PR gate. Now states the condition and names the inequality only as what it
    is NOT, with the reason.

The three surviving mentions of the inequality are all past-tense or negated
(src/install-engine.cts:448, :508 and the sibling test comment at :3698) and
are correct as they stand.

Verified: `npm run lint:ci` exit 0; full `npm test` 31327 tests / 31312 pass /
0 fail / 14 skipped, run with TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3712): canonicalization fails closed, matching identify's errno split

Codex full-PR review of #3725, run against the round-4 head.

`resolveThroughLinks` caught EVERY realpathSync error and fell back to
`path.resolve(dest)` — the lexical spelling. That inverts the function's own
purpose. An aliased `<sandbox>/.agents` that cannot be canonicalized keeps its
sandbox spelling, satisfies the nested-sandbox exemption at :227, and the write
is ALLOWED into the real home — the exact escape this walk exists to close. The
module documents that it fails CLOSED with ONE named exception (the marker
branch); this was a second, unnamed one.

Split by errno, and deliberately by the SAME split `identify` already draws
rather than a second policy in one module — both answer "does this path exist
as named?", so they must not disagree:

  ENOENT / ENOTDIR -> walk up. The ordinary case: a fresh install resolves a
    destination nothing has created yet, so realpath fails on the leaf and on
    every not-yet-created ancestor. Refusing here rejects every install.
  anything else (EACCES, EPERM, ELOOP, EIO) -> refuse. The component exists but
    cannot be resolved, so the guard cannot tell where the write lands.

Three rows in the predicate block, beside the other aliasing rows:
  - a symlink CYCLE in the destination path (ELOOP)   -> REFUSE
  - a destination that does not exist yet (ENOENT)    -> ALLOW
  - a component behind a regular file (ENOTDIR)       -> ALLOW

Teeth checked against the artifact the test loads, not the source: reverting
the condition to the swallow-everything shape in the compiled
test-home-guard.cjs turns row 1 — and only row 1 — red. The ENOTDIR row caught
a stale build during development, which is the point of asserting on the
compiled file.

CONTEXT.md and the resolveThroughLinks docblock both record the new behaviour,
so this does not repeat the prose-vs-code drift the round-4 finding was about.

The changeset's existing scope sentence now bounds "six writers" to the
`installRuntimeArtifacts` call tree and names `cmdGenerateDevPreferences` —
which resolves the same codex `home` override through `getGlobalSkillsBase` and
writes SKILL.md beneath it unguarded. It has no in-process caller today (its
only direct require-and-call is a spawnSync with HOME sandboxed), so it is
latent rather than live, and whether it belongs in this PR is raised with the
maintainer rather than decided here.

Verified: `npm run lint:ci` exit 0; full `npm test` 31330 tests / 31315 pass /
0 fail / 14 skipped, TMPDIR unset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 18:27:02 -04:00
Tom Boucher
1178c5f995 test(#3108): make the overlay ENOENT tolerance and hooks/dist readiness check honest (#3772)
* test(#3108): failing-first suite for the overlay vanish-retry and hooks/dist staleness

RED by construction, and deliberately narrower than the issue.

#3108 reports a bare `ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'`
and attributes it to hooks/dist never having been built. That mechanism cannot
produce that error: buildOverlayRepo enumerates with readdirSync and links what
it enumerated, so a directory that never existed yields no names and no link is
ever attempted. The error requires the file to have existed at readdir and
vanished before the link -- which is the atomic-replace race placeVanishableLeaf
was written for in #3285, three days AFTER this issue was filed.

What is still genuinely broken, and what these tests bind to:

placeVanishableLeaf's retry is unguarded. On ENOENT it re-checks existsSync and
calls attempt() once more, bare. A second atomic replace inside that window
throws an unhandled ENOENT of exactly the reported shape. Rows 1/2/3/5/6/7 fence
the surrounding contract -- most already pass, which is the point: they are what
stops the fix from widening into "swallow every error". Row 6 in particular
covers a non-ENOENT on the RETRY, the exact path the new code will live on.

ensureHooksDist's staleness predicate is extension-blind. It rebuilds only when
hooks/dist is absent or holds zero .js files, while build-hooks.js also ships
.sh -- including gsd-session-state.sh, the very file in the report. A dist with
.js present and every .sh missing reads as populated and the rebuild is skipped.
Rows 10 and 11 are mirrored on purpose: asserting only the .sh direction would
permit swapping one extension heuristic for another, so both directions force
the predicate to be about the expected set (HOOKS_TO_COPY, which build-hooks.js
already exports) rather than about counting an extension.

Rows 12 and 13 are cost guards. ensureHooksDist runs per suite and its rebuild
is a real subprocess, so a predicate that over-triggers turns a correctness fix
into a throughput regression nobody attributes to it; and hooks/dist legitimately
carries files the expected list does not name, so an exact-set match would
rebuild forever.

The warning-text row is t.skip()'d rather than faked: the only ways to assert it
were a source-grep (banned by local/no-source-grep) or a full overlay build, and
a test that cannot be written honestly is better skipped visibly than written
vacuously.

Filename note: this started as fix-3108-*.test.cjs and tripped
lint-regression-test-names, which bans new fix/bug/issue-NNNN files, then as
install-overlay-helpers.test.cjs and tripped lint-test-file-count, whose `install`
bucket is already at its limit. Module-named under the overlay bucket satisfies
both. The allowlists were left untouched -- both are empty, so nothing here is
grandfathered and adding an entry would have been the wrong instinct.

* fix(#3108): guard the overlay retry and make the hooks/dist check see .sh

Two holes, both reachable from the failure #3108 reports, neither of them the
cause it names.

placeVanishableLeaf's retry was bare. On ENOENT it re-checked existsSync and
called attempt() once more with no catch, so a second atomic replace landing
inside that window threw an unhandled ENOENT -- exactly the reported
`ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'`. Two vanishes inside
the window means the same thing one does: the path is going away and is not part
of the snapshot. It now returns false and skips the leaf, reaching the conclusion
the single-vanish case already reached.

Still ONE retry. No loop, no backoff, no sleep -- the existing comment argues
that a timing-based wait here would be the flake rather than the fix, and that
reasoning did not change. Non-ENOENT still propagates from either attempt, which
is the invariant a careless widening would eat; the suite pins it on the retry
path specifically, because a fix that guarded only the first attempt would look
right and be wrong.

ensureHooksDist could not see the file class that caused the report. It rebuilt
only when hooks/dist was absent or held zero .js files, while build-hooks.js also
ships .sh -- including gsd-session-state.sh itself. A dist with .js present and
every .sh missing read as populated and the rebuild was skipped. The predicate is
now membership against build-hooks.js's own exported HOOKS_TO_COPY, so it asks
"is everything expected present" instead of counting an extension, and it cannot
be blind to a file class again.

It is extracted as isHooksDistStale(dir) with ensureHooksDist calling it, so
there is one predicate rather than two that can drift. Extra unexpected entries
are explicitly not stale -- hooks/dist legitimately accumulates subdirectory and
hooks/lib output, and an exact-set match would rebuild forever. One readdirSync
into a Set, no per-entry existsSync, no stat: it runs per suite and its rebuild
is a real subprocess, so an over-triggering predicate would turn this into a
throughput regression nobody would attribute to it.

The skipped-leaf warning now names `npm run build:hooks`. It already named the
cause; a reader still had to know what produces that directory.

Deliberately NOT done: nothing here makes an absent hooks/dist fail. Absence is a
legitimate package shape that bin/install.js:11191 treats as "nothing to verify",
and six of the seven install suites never read the directory at all.

* fix(#3108): count hooks/dist subdirectories, and never throw out of the predicate

Two gaps found reviewing the predicate I had just written.

It ignored HOOKS_SUBDIRS_TO_COPY. That is ["lib"], and hooks/dist/lib carries
gsd-graphify-rebuild.sh, so a dist with all 27 top-level files but no lib/ read
as populated. That is precisely the blindness the .js-count heuristic had, one
level down: a whole file class invisible to the check. Fixing the extension case
and leaving the subdirectory case would have been half a fix, and the half left
behind is the one nobody would look at again.

Subdir names are bare (no slashes), so they slot into the same top-level readdir
Set — no second readdirSync, no stat. Whether lib is really a directory is not
checked; that would cost a stat per entry and buys nothing, since the build owns
that.

It could also throw. existsSync passing does not make readdirSync safe: the path
may be a regular file (ENOTDIR), unreadable (EACCES), or retired in the race
between the two calls. This helper runs in every install suite's before(), so an
unhandled throw there fails a suite on a condition it cannot act on. Unreadable
is indistinguishable from unusable for this question, and rebuilding is
idempotent, so it now reports stale instead.

Both are the same shape as the original bug and the same shape as each other: a
guard that answers "is this ready" must not have a blind spot or a hard edge,
because every caller treats a false negative as "carry on".

Two existing tests asserted a "complete" dist without lib and had to be corrected
to keep meaning what their names claim, rather than being left passing against a
definition of complete that no longer holds.

* test(#3108): close the review findings, including a half-closed subdir check

An isolated correctness reviewer found no blockers and four real gaps.

The wiring was untested. Every Group-2 test exercised the pure predicate; none
called ensureHooksDist. So restoring the old inline .js-count check INSIDE
ensureHooksDist -- keeping isHooksDistStale exported and correct -- left the
whole suite green, and that wiring is the actual #3108 defect. Two tests now
drive ensureHooksDist itself through the process seam: build invoked exactly once
when stale, never when fresh. The second is the one a permissive revert fails.

Reaching that seam meant requiring process-seam as a module object rather than
destructuring runNode, so a test can replace it in place. That is a testability
affordance in a test helper, not a production change, and it is commented as such
so it does not read as an accident later.

The subdir check was only half closed, and the half left open was the important
one. It required `lib` to be PRESENT in the top-level readdir, never looked
inside -- so an EMPTY dist/lib, missing gsd-graphify-rebuild.sh, still read as
populated. That is precisely the missing-file-class case the subdir check was
added to catch, which made the fix a gesture at the problem rather than a fix.
Each subdir entry must now be a readable, NON-EMPTY directory. A stray regular
file named `lib` throws ENOTDIR into the same try/catch and reads stale too.

Cost stayed honest: one extra readdirSync total (there is exactly one subdir
entry), no stat, no per-expected-file syscall. Probed against the real
hooks/dist -- still reports fresh, so no suite gains a rebuild.

Two nits, both real: the error-code sweep re-tested EACCES already covered
standalone, and the property ignored presentAtFinalAttempt whenever vanishCount
was not 1, making roughly half the 200 runs duplicates. The flag now varies
meaningfully across the whole range and the assertions depend on it.

One reviewer finding was already stale: the subdir and ENOTDIR work was
uncommitted when the reviewer snapshotted the tree, and had landed in a67aefb9c
before the report arrived. Verified rather than assumed.

* fix(#3108): stop the vanish tolerance from swallowing a dest-side ENOENT

A defect this PR introduced, caught by an isolated security reviewer.

linkSync(src, dest) throws ENOENT for the DESTINATION path too, not only for a
vanished source. The widened retry caught that, saw the source still present,
retried, got the same dest-side ENOENT, and returned false -- recording the leaf
as "vanished mid-walk" and printing a warning that tells the reader to run
`npm run build:hooks`. A remedy with nothing to do with the actual cause, an
overlay quietly short a file, and the install under test proceeding against an
incomplete tree.

It also falsified the function's own documented invariant, which says in as many
words: "Returns false only when the path left the source tree entirely." Widening
the tolerance without re-reading the sentence above it is how that happens.

The retry now re-checks existsSync(srcPath) before tolerating: source still
present means the ENOENT was about something else and it propagates untouched.
Chose the existsSync re-check over comparing retryErr.path to srcPath -- err.path
normalization is not guaranteed across platforms, and a path-equality test is a
subtler thing to get wrong later.

The FIRST catch was probed and is already correct: for a dest-side ENOENT the
source is present, so it falls through to the retry rather than returning false.
Left unchanged rather than "fixed" symmetrically.

Two regression pins, deliberately opposed: a dest-side ENOENT with the source
present must THROW, and a genuinely absent source must still return false. The
second exists because the obvious over-correction -- always rethrow on the retry
-- passes the first and silently undoes what this PR set out to fix.

Also closed the skipped placeholder. It claimed no non-flaky seam existed for
asserting the warning text; the reviewer pointed out an injectable `warn` param
is trivial, and they were right. buildOverlayRepo now takes opts.warn defaulting
to console.warn (byte-identical for every existing caller) and the skip is
replaced by real tests: fires with the remedy named on a skipped leaf, silent on
a clean walk. "No seam exists" was a design choice presented as a constraint.

Recorded the sequential-only constraint at the two sites that monkeypatch fs
process-wide: adding { concurrency: true } to this file would cross-contaminate
every other suite in the process. Better written down than rediscovered.

Known limit, disclosed rather than fixed here: a legitimately dropped leaf can
still pass vacuously downstream -- agent-fragments-emission asserts a negative
over filesContaining, and mcp-catalog-parity has only an anti-vacuity floor of
one. That is a pre-existing property of those suites and the tolerance #3285
already chose; this change narrows which drops are possible rather than adding
the completeness assertion those suites lack.

* fix(#3108): discriminate ENOENT by the dest parent, not by re-checking the source

The previous commit's dest-side guard was wrong, and the remote run said so:

  "a leaf that vanishes again during the retry is skipped, not a bare ENOENT"
  Got unwanted exception. Actual message: "ENOENT: no such file or directory"

That test was right and the guard was wrong. It rethrew when existsSync(srcPath)
was still true, on the theory that a present source means the ENOENT was about
the destination. But in the genuine race the source is being atomically REPLACED,
so it is legitimately present again at the re-check while the ENOENT was entirely
source-side. The gate therefore threw on precisely the race #3285 exists to
tolerate -- trading one misclassification for a worse one, since the old bug was
a bare crash and the new one broke the working tolerance.

The security reviewer's alternative discriminator does not work either, and a
probe settles it. Node populates BOTH `path` and `dest` on a link ENOENT, and
`err.path` is the SOURCE in both directions:

  linkSync(existingSrc, missingDir/a.txt) -> ENOENT path=<source> dest=<dest>
  linkSync(missingSrc,  validDest)        -> ENOENT path=<source> dest=<dest>

So the error object cannot tell you which side failed.

What CAN: the dest parent. buildOverlayRepo builds its own dest tree --
place() mkdirSync's recursively into a private mkdtempSync root no other process
touches -- so a missing dest parent is always a bug (Windows MAX_PATH, a
concurrent cleanup, a bad dest), never the replace race. A present dest parent
means the ENOENT was about the source, which is the case we tolerate.

placeVanishableLeaf therefore takes an optional destPath and uses the dest
parent as the sole discriminator when it has one; with no destPath it behaves
exactly as before. linkOrCopyFile and the copy-mode call site both pass it,
because those are the two places that actually know the destination.

The doc comment now records BOTH failed discriminators and why each fails --
existsSync because the source is legitimately replaced mid-race, err.path
because it names the source either way. Those are the two things a future reader
reaches for first, and both look correct until they are not.

The dest-side regression pin was rewritten to drive the real mechanism: a real
temp source and a dest whose parent does not exist, through linkOrCopyFile.
Previously it forced a throw through a present source, which is what encoded the
wrong theory into a test and made it look verified.

---------

Co-authored-by: sim <sim@local>
2026-08-22 23:04:44 -04:00
Tom Boucher
004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00