b90eef28e832a08e4c60e30ee96648a2abee22c3
280 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b90eef28e8 |
enhance(#3829): report code review severity counts and record a per-finding disposition (#3861)
* enhance(#3829): report code review severity counts and record a per-finding disposition `code_review_gate` extracted `status:` from REVIEW.md's frontmatter and discarded the `critical`/`warning`/`info`/`total` values sitting in the same range, so its output was byte-identical for a review with one `info` finding and a review with a Critical. Nothing anywhere recorded what happened to a finding: no file under `gsd-core/workflows/` branches on `issues_found`, and `gsd-verifier.md` has zero references to REVIEW.md. A phase therefore reached `phase.complete` with Criticals standing and no trace they had been seen. Both halves were approved on the issue; the gate stays advisory. A — severity surfacing, in `execute-phase.md`. The gate states the breakdown it already parsed, accepting `blocker:` as the documented tier-equivalent of `critical:`. The breakdown is shown only when all four counts are numeric (`REVIEW_COUNTS_OK`); otherwise the countless message stands, because gating on the total alone still emits `6 findings — critical` for a review carrying a total and nothing else. Frontmatter is extracted by an `awk` that emits only when it saw the CLOSING delimiter, after stripping CR. A `sed` range re-opens on a body `---` and runs to EOF: first-match protects a key the frontmatter always carries, but not an optional one, so a review with no `findings:` block and a body `total:` line would have reported the body's number. An unterminated block would leak the whole body the same way. Every read is guarded and `|| true`-terminated. This step is advisory, and under `set -e`/`pipefail` a non-matching `grep` exits 1 — an assignment whose command substitution fails would take the step down with it. A REVIEW.md that is missing, a directory, or unreadable now leaves the counts empty and execution continues. B — per-finding disposition, in a new lazily-read step file, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, referenced from the gate in the established plain read-and-execute form. One row per finding ID, defaulting to `open`, and: - `fixed`/`skipped` are reconciled from REVIEW-FIX.md, whose section headings are matched WHOLE — a prefix match let `## Fixed Issues Verification` classify every finding beneath it as fixed — and only when the fix report names the SAME finding. Finding ids are reused across re-reviews, so matching on the id alone let a stale fix report declare a brand-new CR-01 already fixed. - headings inside fenced blocks are ignored; a quoted example is not a finding. - an id listed under both sections resolves by first occurrence, not row order. - a recorded disposition is preserved together with the reason in its Source cell, escaped pipes included, and a hand-mangled row missing its trailing pipe still keeps its decision. - a decided finding the current review no longer reports is CARRIED and marked; `--auto` rewrites REVIEW.md each iteration, so this is routine, and dropping the row would erase the record that it was seen. An untriaged `open` row for a vanished finding is not carried. A review reporting nothing still reconciles an existing ledger rather than freezing it. - a run that changes no disposition rewrites nothing, so a re-executed phase does not produce a docs commit whose only delta is a timestamp. The record is a sibling artifact, not a section inside REVIEW.md: `--auto`'s re-review loop rewrites REVIEW.md every iteration, so a ledger kept inside it would not survive the next pass, and REVIEW.md has a single writer that this step is not. B lives in an extracted step file because `execute-phase.md` was 91,493 bytes against a 98,304 hard cap the size-budget test calls a red line, and because `scanWiredKinds` caps a call site's dispatch-coverage region at 6000 characters — an inline version pushed the `kind == "gate"` paragraph out of that window, which silently drops `gate` from the covered set and fails `gen-capability-registry --check` while pointing at the capability rather than at the prose that displaced it. Extraction is what that size test's own message prescribes, and it leaves the file at 93,854 bytes. The tests execute the shipped script rather than modelling it. Three adversarial review rounds each refuted "the mirror is faithful", and mutation testing agreed: with a hand-written model, deleting the carried-row logic from the shipped file turned nothing red. The suite now extracts the embedded script — undoing exactly the four shell double-quote escapes — and runs it, so all ten mutations of its behaviour are caught. * chore(#3829): set changeset fragment pr to 3861 * fix(#3829): keep execute-phase.md under both size ceilings and propagate the launcher probe The first push failed `full test (macos-latest, 24, shard 3/3)`. Two things it caught that the CI-selected scope for this diff does not run, and that I therefore did not run either: 1. `execute-phase.md` is governed by TWO ceilings, not one. The XL hard cap in `tests/workflow-size-budget.test.cjs` (98304) was satisfied at 95179, but the frozen ADR-857 pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs` (93600) was not. The whole budget from base is 2107 bytes, which the inline reporting half alone did not fit. That half now lives in the extracted step file alongside the disposition half, and the parent carries only the paragraph that reads and executes it — 91529 bytes, 36 over base. 2. The step file calls `gsd_run`, so it owes the hermes runtime-home probe that `tests/runtime-launcher-parity.test.cjs` (E) requires of every workflow file that does. Propagated with `node scripts/sync-runtime-launcher.cjs`, the remedy that test names. Verified with the FULL unit suite this time rather than the scoped selection — 14 shards, 0 failures — plus `npm run lint:ci`, and a re-run of the ten mutations of the shipped disposition script, all still caught. * fix(#3829): stop the disposition step instructing the agent to execute itself Blocker 1 and Minor 7 of the round-1 review are one defect. The step file carried a copy of execute-phase.md's pointer paragraph, so it named its own path as something to "read and execute" — unbounded self-recursion at runtime — and that copy is also the duplicated paragraph, sitting immediately above the full instruction it duplicates. Removing the copy resolves both. execute-phase.md remains the only surface that points here, which is what it always intended. Two structural tests guard it. Both are red against the pre-fix file: no behavioural test could see either defect, because they execute the node script through the process seam and so never read the prose that tells the agent what to load. * fix(#3829): re-derive the ledger paths in the block that uses them Blocker 2. The disposition block reads REVIEW_FILE, DISPOSITION_FILE and PADDED, all derived in the step's FIRST shell block. Each fenced block is dispatched as its own shell, so all three are empty by the time the second block runs: the ledger write lands on a bare `-REVIEW-DISPOSITION.md` path and the review read finds nothing. The step then reports success having produced no artifact — the feature's central acceptance criterion, silently unmet, with no error to notice. The tell was already in the file: the gsd_run shim preamble is re-emitted in the second block for exactly this reason. These three paths belong beside it, and now are. The guard test asserts the general property rather than the instance — every block derives what it reads, inheriting only the step's declared inputs (PHASE_DIR, PHASE_NUMBER) — so a third block added later cannot reintroduce it. Red against the pre-fix file. * test(#3829): assert the counts mirror against the shipped shell, and execute its guards Major 4, with Minor 6 and part of Minor 9. The disposition builder stopped being a mirror three rounds ago, and the reason given then was that a hand model of a shell-embedded script drifts while the tests stay green. parseGateCounts kept its mirror anyway. That argument does not stop applying at the boundary between the step's two shell blocks, so the mirror now loses its authority: it is asserted against the shipped awk and greps, run under `set -euo pipefail` in a real shell, across every fixture it is exercised on. Negative-controlled in both directions. Dropping `blocker:` from the mirror alone fails the parity test; replacing the shipped awk with the leaky `sed -n '/^---$/,/^---$/p'` range fails it on the unterminated-frontmatter fixture. Divergence in either half is now red, which is what the finding asks for. Skipped on win32, where there is no bash to compare against. Minor 6: the zero-count edge is covered — `0` is numeric, so a zero-finding review reports `0 findings — 0 critical, …` rather than falling back to the countless form. A guard written against truthiness would have failed here silently, and now cannot. Minor 9, partially: running the block makes its advisory guards behavioural, so the four `src.includes()` assertions that stood in for them are retired — a missing and an unreadable REVIEW.md are now proven not to abort under `set -e`, rather than asserted to contain a string. The remaining docs-parity assertions are kept deliberately; see the PR discussion. * test(#3829): add the render/re-parse fixed-point property for the ledger Major 3. RULESET.TESTS.property-based-testing asks for at least one fc property on a parsing/transformation contract, and the ledger is one with a fixed point stated in its own prose: re-running the gate preserves every disposition except `open`, and rewrites nothing when nothing changed. Two properties, both driving the SHIPPED script rather than a model of it: idempotency — a second run reports `unchanged` and leaves the file byte-identical. Without it, the timestamp alone dirties the tree on every phase re-run. round-trip — a hand-recorded decision AND the reason beside it survive render -> re-parse -> render, escaped pipes included. The Source cell is where a human writes why something was deferred, so losing it loses the only thing that instruction asks for. Negative-controlled per property: disabling the unchanged-check fails the first and only the first; discarding the carried source cell fails the second and only the second. numRuns is 40 rather than the shared 200 because each case spawns the shipped script twice through the process seam. The seed stays pinned, so a failure still reproduces; the deviation is stated in the file header rather than made silently. * fix(#3829): state a stale fix-report match instead of dropping it silently Minor 5, plus the finding-id census this round owes. Exact-title coupling stays — ids are reused across re-reviews, so a stale REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed. What changes is the silence. A row that stays `open` because the report named a different finding under the same id is indistinguishable, to any reader, from a row that stays open because no report mentioned it. The gate now names the ids it could not reconcile, on both report paths, and stays advisory throughout. The census (RV4, self-found — the review did not ask for this). The script enumerates finding-id prefixes in three places: the heading matcher, the ledger re-parser, and the severity map's keys. The DOMAIN those enumerate is owned elsewhere — gsd-code-reviewer.md's body template and its Label-equivalence paragraph — so it can acquire a member without this script changing. Reached: CR, BL, WR, IN — 4 of 4, all present. Not reached: none today. What follows if that changes is the payload: an unlisted prefix is not mis-tiered, it is INVISIBLE — the finding never enters the order list and gets no row at all, so the artifact silently under-reports the review it is meant to record. Adding a prefix to two of the three copies fails the same way, and additionally drops carried rows on the next run. Two guards rather than a rewrite: hoisting the alternation into one constant means rebuilding three regexes inside a double-quoted shell string, which is the exact class of edit that produced both of this round's blockers. The guards make the drift loud instead, and are negative-controlled against each of the two ways it can happen. * docs(#3829): keep the feature reference descriptive, not instructional Minor 8 — a Diataxis mode mix. "Set `deferred` by hand and put the reason in the Source cell" is a how-to instruction sitting in a reference doc. The information belongs there (a reader needs to know the field exists and what preserves it); the imperative does not. Rewritten to describe the field instead: `deferred` is the one disposition the gate never writes, and the reason recorded beside it survives re-runs. The same pass records Minor 5's new behaviour, since the reference described the title coupling but not what happens when it misses. The imperative form is kept where it belongs — inside the ledger the gate renders, which is where a reader meets the field and the only place an instruction has an audience. docs/FEATURES.md regenerated from it; `gen-features.cjs --check` is green. * fix(#3829): close six defects found by reviewing this round's own fixes None of these came from the maintainer's review. They came from adversarially reviewing the five commits above before pushing them, and two are worse than anything the round was opened to fix. 1. A foreign fence marker swapped an example for a finding. The heading scanner toggled fenced/not-fenced on ANY fence marker, so a ~~~ line inside a ``` example closed the fence and the example's real close reopened one. Driven: a review quoting ~~~ inside a fenced example produced a ledger recording CR-77, the illustration, and omitting CR-01, the actual finding. A confidently-written artifact wrong in both directions at once. The open marker's character and length are now remembered, and a fence closes only on the same character at least as long, per CommonMark. 2. The disposition block had no status gate at all. The prose above it says it runs only when the review reports issues — but block 1 computes REVIEW_STATUS, emits nothing, and its shell is discarded, so no later block could act on that condition even in principle. A prose gate on a value nothing downstream can see is not a gate, and a clean re-review would rewrite a ledger it was never meant to touch. Re-derived in block 2's own shell. 3. A numeric breakdown could still be internally false. `total: 0` beside `critical: 1` is four valid numbers rendering `0 findings — 1 critical, …`. Numeric was necessary and not sufficient; an inconsistent breakdown is now withheld for the same reason a partial one is. 4. The carried-marker strip ate hand-written prose. It removed a trailing `(not in the current review)` unboundedly and unconditionally, so a deferral reason that merely ENDED in that phrase lost it — the one field a human writes into this artifact. Now bounded to one occurrence, and only on rows the marker can legitimately be on. The no-growth property it exists for is re-pinned. 5. parseGateCounts diverged from the shipped pipeline in two ways no fixture reached. The shipped reads are `cut -d: -f2 | tr -d ' '`: `tr` removes INTERNAL spaces (`1 0` -> `10`) where `.trim()` keeps them, and `cut` takes only the second colon-field where a tail capture keeps the rest. The mirror models the pipeline now, and both counterexamples are fixtures — a parity assertion that agrees only on well-formed input asserts very little. 6. The prefix census guards were both partly vacuous. The drift guard read the two regex alternations and not the severity map, so a set could agree in both regexes while mis-tiering in the map. The domain guard scanned only `### XX-01:` headings — and BL appears in no heading at all, only in the Label-equivalence prose, so the guard passed purely because BL happened to be hard-coded and would have missed the next prose-defined prefix exactly as it missed BL. Both widened; the domain the guard now sees is BL, CR, IN, WR. Each fix fails a named test on reversion and none fires on the ordinary path. The property generator now deliberately produces the reserved suffix from (4), which a generator drawn only from innocuous characters could never reach. Also corrected: the previous commit's account of the empty-path failure. The script did not write a bare `-REVIEW-DISPOSITION.md`; it threw on reading the empty review path and the trailing `|| echo` swallowed it as a non-blocking skip. Same silent outcome, different mechanism, and the comment said the wrong one. * fix(#3829): the tests now run what bash runs — and six fixes to the fixes A second adversarial pass over the previous commit. It found a regression that commit introduced, and the reason it slipped through is the finding worth keeping. THE FIDELITY GAP. Every test here extracts the embedded script as TEXT and runs it. Bash does not: it expands the double-quoted `node -e "..."` argument first, so a backtick inside it is COMMAND SUBSTITUTION. The previous commit put one in a code comment. Bash duly ran it, failed with `+: command not found`, and handed Node a script two bytes shorter than the one 122 green tests were exercising. No behavioural test could see this, because none of them ever asked bash what it would actually pass. One now does, and it is the general guard: it catches an unescaped backtick, an unescaped $, and any other expansion the extractor cannot model. Then, in the shipped step: - A padded count silently disabled the sum check. `$((08 + …))` fails on base inference; it does not abort — the expansion sits in an `if` condition, where set -e does not fire — so the check simply never ran and an inconsistent breakdown passed with a stray diagnostic as its only trace. `10#` on every operand. - The status guard made the script's own reconciliation unreachable. A clean review with an EXISTING ledger must still be reconciled — decided rows carried, stale `open` rows dropped — or the ledger freezes showing findings as open that the review no longer reports. The guard now skips only when there is nothing to reconcile. - The carried marker is no longer stripped at parse time at all. Bounding the strip still ate a carried row's human-written reason. No-growth is a property of the RENDER, so it is enforced there: a marker already present is not appended again. Nothing is stripped, nothing doubles. - Fence openers are bounded to three leading spaces, per CommonMark. - parseGateCounts matched `[ \t]` where the shipped grep uses `[[:space:]]`, which covers form feed and vertical tab. Third counterexample of the same class, and a fixture. - The census drift guard checked only one direction, so a tier for a prefix the regexes never admit stayed green as dead code that reads as coverage. TWO OF MY OWN TESTS WERE VACUOUS, and the controls are what said so. The leading-zero test asserted exit 0 and a consistent verdict — both true before the fix. The clean-review test drove the node script directly, which never executes the shell guard at all: it passed unchanged with the guard made unconditional. Both are rewritten to test the layer the defect lives on, and both now fail when their fix is reverted. Every fix in this commit fails a named test on reversion, each mutation verified to have applied before its verdict was read. * fix(#3829): the carried marker can no longer outlive the carry A third adversarial pass. Its most important finding is a defect the SECOND pass talked me into, which is worth recording as plainly as the fix. THE MARKER BECAME A LIE. Pass 2 objected that bounding the carried-marker strip still altered a human-written reason, and proposed storing the cell verbatim instead. That objection was a preference, not a defect — its own driven output showed exactly one marker, which is correct — and adopting it created a real one: once the generated marker is stored it can never leave, so a carried finding that REAPPEARS in a later review still renders "not in the current review". The ledger then contradicts its own contents. Driven both runs. The strip is back, bounded to one occurrence and unconditional. The residual ambiguity is irreducible — a reason ending in exactly that phrase is indistinguishable from the marker — and it costs nothing real: on a carried row the render puts the phrase straight back, and on a current row the phrase was self-contradictory to begin with. The unbounded quantifier is what had to go, not the strip. The property now states that contract rather than asserting a verbatim survival the code deliberately does not provide. Also: - An ABSENT REVIEW.md abandoned the ledger it was meant to reconcile. The guard proceeds when a ledger exists, then the script read the review unconditionally, threw, and the trailing fallback swallowed it — the freeze the reconciliation path exists to prevent, reached through the door the guard opened. - Counts are length-bounded as well as digit-only. Bash integers wrap at 2^64, so a 20-digit count arrived at the sum as 0 and an inconsistent breakdown passed. - A closing fence must carry only whitespace after its marker; a line with an info string is an opener's shape and ended the fence early. - parseGateCounts matched [ \t\n\v\f\r] where the shipped grep uses [[:space:]], which under this UTF-8 locale matches EM SPACE. `\s` is the faithful model. Fourth counterexample of that class, and a fixture. - The agent-domain scan required [A-Z]{2,}, so a one-letter prefix like `C-01` — explicit and parseable, not prose — was invisible to it. AND THE FIDELITY GUARD PAID FOR ITSELF INSIDE ONE SESSION: writing this round's first draft I put backticks around a token in a code comment again, in the very commit whose subject is that mistake. The probe failed, named it, and no test of behaviour could have. Two of my own tests also had to be rewritten: one asserted things true before its fix, and one drove the node script directly where the defect lived in the shell. 383 pass across the touched files and the two size ceilings; ten lint gates green; every fix fails a named test on reversion, each mutation verified to have applied before its verdict was read. * docs(#3829): the Source reason is preserved, but not verbatim — say so Found by claim-auditing the response comment before posting it, which is the one place this would have been caught: the doc and the code were written in different commits and only a reader holding both notices they disagree. The feature reference said the hand-written reason is "preserved verbatim across re-runs". It is not, and deliberately so — a reason ending in the literal phrase "(not in the current review)" loses that trailing phrase, because it is indistinguishable from the carried marker the gate appends. The exception is stated rather than dropped, with the reason it is the better trade: storing the marker instead means it never leaves, and a carried finding that later reappears goes on claiming it is absent from the very review that reports it. A ledger wrong about its own contents beats losing a duplicated phrase, but only if the doc admits which one it chose. FEATURES.md regenerated; gen-features --check and lint:docs green. * fix(#3829): the gate now emits the counts it computes (B1a/B1b) Block 1 computed REVIEW_STATUS and the four counts and printed none of them, then the prose below asked the agent to display four of them. The shell exits at the closing fence and the agent sees only stdout, so those values were unobtainable: REQ-REVIEW-08 was unreachable in every shipped path and the fence was decorative. The rule was already stated one block down -- "a prose-only gate on a value no later block can see is not a gate" -- and applied only to block 2. It now governs the block that is this step's primary deliverable. Both arms emit, and the status gate is mechanical rather than prose: a clean/skipped/absent review prints nothing, an inconsistent or partial breakdown prints the countless form, and the full breakdown prints otherwise. Driven against the review's own case (critical: 1, warning: 9, info: 8, total: 18) with no appended emitter: Code review: 18 findings - 1 critical, 9 warning, 8 info. Consider running: /gsd:code-review 1 --fix * test(#3829): the counts harness stops manufacturing the output it asserts on (B2) runShippedGateCounts extracted the shipped fence and then APPENDED its own printf of the six internal variables before running it. Every counts assertion was green against a script that existed only inside the test process: the shipped fence emitted nothing, the tested fence emitted six lines because the test added them. That is why B1a shipped past a suite that looks like it covers exactly that surface -- the green was structurally incapable of turning red for it. The emitter now lives in the fence, so the harness reads the fence's own stdout and synthesizes nothing. Parity with the mirror moved up a level with it: renderGateMessage() renders both arms from the mirror's parsed counts and the assertion compares the WHOLE emitted message, so a drift in any parsed value changes the string or the arm it selects. Asserting on the observable is strictly stronger than asserting on five intermediates, and it can express what the old probe could not -- an absent review now reports NOTHING, which is a different fact from reporting a countless review. A fifth src.includes() assertion converted with it (round 1 retired four). It pinned the PROSE stating the countless condition, so it went red when the emitter moved into the fence while the behaviour it named was untouched -- the pin arguing for its own conversion. Negative control: reverting the shipped echo now turns 16 tests red. Before this commit the same reversion turned zero red, which is the finding. * fix(#3829): the disposition column is an enum, not any lowercase token (B3) ADR-227 requires a trust boundary to validate semantic SHAPE and to coerce a failure to the contract's safe default. The ledger is a trust boundary by construction -- the rendered instruction tells a human to hand-edit it -- and the prior-row parser captured column 3 as ([a-z]+), checked against nothing. One transposed character was enough. `| CR-01 | critical | opne | - |` is not the literal 'open', so it beat the default, was excluded from the `open:` headline count, and was carried forward forever. The ledger then reported the phase fully triaged off a typo. The asymmetry is what made this a correctness bug rather than a style point: a typo OUTSIDE [a-z] ('Deferred') already failed to match, lost the decision and reset the row to open -- safe. A typo INSIDE [a-z] was unsafe. The parser failed open in the one direction that matters. A row that fails the enum now yields no prior entry and the row falls back to 'open', by the same path the capital-D case already took. The property test could not have caught this: DECIDED is drawn from the vocabulary, so no property built on it can present an out-of-vocabulary token. Added JUNK, the arbitrary for the complement, deliberately lowercase so it stays inside the old capture's own character set -- the unsafe half is the token that LOOKS like a decision and is not. The new property also asserts the headline count agrees with the row it renders, which is the half the defect actually reported wrongly. Negative control: the new property fails against the ([a-z]+) capture and passes against the enum. * fix(#3829): a finding the heading parser cannot match is surfaced, not dropped (B4) Two independent parsers produce two numbers one paragraph apart -- the counts from REVIEW.md's frontmatter, the rows from `### <ID>:` heading matches against a closed CR|BL|WR|IN alternation -- and nothing reconciled them. A finding the alternation could not reach contributed no row, no note and no diagnostic, and the ledger then declared `open: 3 of 3` over a set strictly smaller than the console line had reported one paragraph earlier. Two findings recorded nowhere, and neither artifact said so. The PR's own argument for the closed alternation -- that an unlisted prefix produces no row rather than a MIS-CLASSIFIED one -- is the wrong trade under this repo's fail-safe rule. A dropped finding is demoted below every finding that parsed, and an unparseable finding is precisely the one a human most needs to see. Block 2 now derives the frontmatter total (anchored inside the findings: mapping, digit-and-length-bounded like block 1's) and hands it to the script, which reconciles it against the CURRENT review's matched findings -- order.length, never rows.length, which also counts carried rows and would either understate the shortfall or invent one. Surfaced exactly as the stale fix-report case already is: a non-blocking `unparsed: N` key plus the console line, both naming the two numbers so the claim is checkable. Code review disposition recorded: 3 of 3 finding(s) open (2 finding(s) recorded NOWHERE: the review reports 5, but only 3 matched the expected heading shape `### <CR|BL|WR|IN>-NN: <title>`) The key is emitted only when there IS a shortfall, so an ordinary ledger gains no noise key and the unchanged-run check is unaffected. Four tests, including three negative controls the round owed itself: a clean review gains no key, an absent/non-numeric total reconciles nothing rather than fabricating a shortfall, and a total SMALLER than the row count cannot render `unparsed: -1`. Reversion control: dropping the key turns the first red. * fix(#3829): pass --raw to the commit_docs config-get (#3763) Not from the review -- from a gate the base range added after it. #3763 lands `tests/config-get-raw-guard.test.cjs`, and this branch was its sole offender: a config-get command substitution without --raw feeds JSON.stringify output into a bash string comparison, where it silently never matches for string values. The consumer here is exactly that: if [ "$COMMIT_DOCS" = "true" ] Every other shipped call site in the tree already passes --raw (spike.md, fast.md, new-milestone.md, sketch-wrap-up.md, ...), so this is sibling convention, not a new posture. Worth recording because the two readings are both correct and they disagree: round 2's review cleared this exact line under ADR-3409 as "the safe member of that family", since `query config-get <key>` with no --pick exits 1 on absence and the fallback arm is reachable. That is still true -- --raw does not change it. The base then moved and added a gate that reads the same line for a different property. * fix(#3829): scope the count reads to the findings: mapping, not just the frontmatter (m1) `^[[:space:]]*total:` matches any indented key anywhere in the block, so a top-level key later named `total:`, `info:` or `critical:` was picked up ahead of the nested one. The block's own extensive comment is about scoping the FRONTMATTER, and the scoping stopped one level short of the mapping the values actually belong to. `status:` was never exposed -- it is anchored to column 0 because it IS top-level. The reads now run over the `findings:` block alone, selected by awk and cut at the next column-0 key. Block 2's REVIEW_TOTAL derivation (added with B4) already used that filter; this brings block 1 to it, so the two agree by construction rather than by coincidence. The mirror models the same scoping, and two fixtures drive it: a top-level `total: 999` ahead of a nested `total: 1`, and top-level `critical:`/`info:` ahead of theirs. Reversion control: unanchoring the shipped reads turns them red. * fix(#3829): severity comes from the section heading, not just the id prefix (M3) gsd-code-reviewer.md emits findings under '## Critical Issues' / '## Warnings' / '## Info', and that heading is the reviewer's own statement of a finding's severity. The walker already visits every line -- the fix-report path tracks '## ' sections -- so the signal was in hand and discarded in favour of the id prefix alone. A reviewer who mis-numbers a Critical as WR-04 while filing it under '## Critical Issues' produced a row reading 'warning'. The ledger's Severity column is the whole basis for triaging it, and it then disagreed both with the review it summarizes and with the frontmatter count line block 1 prints from findings.critical. Section first, prefix as fallback: a finding under no recognized section -- a review that does not use the documented headings, and every row carried from an earlier review -- keeps the prefix mapping, BL- included. Sections are matched WHOLE, exactly as the fix-report sections are, so '## Critical Issues Verification' does not re-tier what sits under it, and a heading inside a fenced example does not govern. Five tests: both mis-numbering directions, the prefix fallback across all four prefixes, the lookalike heading, and the fenced-example case. Reversion control: prefix-only turns the first two red. Sixth src.includes() assertion converted with it -- it pinned the exact source LINE of the enumeration loop, so it went red when that loop was reformatted while the property it names was strictly widened. It now asserts the property: every finding id, in order, once each. * fix(#3829): an untriaged row is carried too, not silently deleted (M1) The carry-forward kept a prior row only when its disposition was not 'open', so an untriaged row for a finding the current review no longer reports was dropped entirely. Combined with the reconciliation gap that left EVERY row open, a re-review deleted the whole ledger. The re-review loop rewrites REVIEW.md on every iteration, so REVIEW.md does not retain it either: run 1 records CR-01 open, the re-review renumbers it to CR-02, run 2's ledger contains neither. That is #3829's complaint verbatim -- "no trace of what happened to them" -- reproduced by the artifact built to prevent it. The old justification, "nothing was decided about it", is exactly the state #3829 says must leave a trace. Every prior row is now carried, and the carried marker is what keeps it honest: the row does not claim the finding is live, it records that it was seen and never triaged. Two costs, stated rather than discovered: a renumbered finding shows twice until the old row is triaged, and a carried untriaged row persists until decided. Both are bounded by the phase's own findings, both are legible from the marker, and both beat a silent delete. Five tests updated -- they encoded the dropped-untriaged behaviour as the contract -- plus one new test for the renumbering case M1 names. Reversion control: restoring the guard turns six red. Two self-inflicted defects caught while writing this, both by probes round 1 built: - Four unescaped backticks in a comment inside the double-quoted node -e argument, which bash ran as command substitution. The extractor-parity probe fired ("--auto: command not found"). Third time that trap has been sprung in this PR, third time the probe caught it. - The reworded ledger footer contained the literal carried-marker phrase, and the marker-accumulation assertion counts it across the whole file, so a doc line read as a second marker. The assertion was right. * test(#3829): cover the count-length threshold at limit-1, limit and limit+1 (M2) The guard is `?????????*` -- nine or more characters -- so the limit is 8 digits accepted, 9 rejected. The only cases were 'x', single digits and a 20-digit value, none of which pins the boundary. RULESET.TESTS boundary-coverage is a hard rule here and it was unmet. All three points asserted, with the sum kept consistent at each so the LENGTH rule is what decides the verdict rather than the sum check incidentally agreeing. Reversion control is the off-by-one M2 names: dropping one `?` moves the limit to 7 digits, which no test could previously notice, and now turns this one red. * feat(#3829): wire the disposition ledger into the fix path (B1c/B1d) REQ-REVIEW-09 was unreachable in every shipped path. execute-phase.md's code_review_gate invokes review with neither --fix nor --auto, so <NN>-REVIEW-FIX.md cannot exist when the gate runs and every row it writes is `open` by construction. The operator then runs /gsd:code-review N --fix by hand -- the very suggestion the step prints -- which writes REVIEW-FIX.md and never touched the ledger. A phase with 23 findings, all fixed, ended at `open: 23 / total: 23`: the artifact that exists to distinguish a triaged finding from a forgotten one asserted that 23 triaged findings were forgotten. Worse than recording nothing, because it looks authoritative and is inverted. Taking remedy (i), not (ii). Narrowing the docs to say the ledger reflects the previous phase execution is a legitimate choice, but it ships a feature whose central artifact is inert and then documents the inertness. ONE ADAPTATION, because the prescribed site does not exist. The review says to wire code-review.md's --fix/--auto path. code-review.md is not the writer (gsd-code-fixer writes the report, code-review-fix.md commits it), and more decisively it has no point that is AFTER the report exists: it delegates through code-review/steps/dispatch-fix.md, which calls Workflow(code-review-fix.md) and then exits the workflow. There is nothing downstream of that call to wire to. The site is code-review-fix.md, immediately after commit_fix_report. That is where the report is on disk and committed, it is the canonical implementation for all fix logic by dispatch-fix.md's own statement, and it additionally covers a direct invocation of that workflow -- which a wiring in code-review.md would have missed. The same step, not a second copy: it consumes PHASE_DIR and PHASE_NUMBER, both already parsed from the init JSON, and it is idempotent, so a phase that reaches the gate and then a fix run ends with one ledger reflecting both rather than two competing ones. Driven end to end: the gate writes `open: 2 of 2`, the fix path reconciles to fixed/skipped and `open: 0`. Two tests -- one pins the wiring and its ordering relative to commit_fix_report and present_results, one drives the two call sites in sequence. Reversion control: removing the step turns the first red; the second covers the reconciliation the wiring makes reachable rather than the wiring itself. * fix(#3829): a reflowed fix-report title is the same title (m2) The stale-fix-report guard compared titles with trim() equality. The strict instinct is right -- ids are reused across re-reviews, so a stale REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed -- but gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that the title is copied byte-for-byte from REVIEW.md. A fixer that reflows a long title produced a spurious mismatch note, left a genuinely-fixed row 'open', and told the reader the report named a different finding. That false-positive mode was acknowledged nowhere. Whitespace is normalized, and only whitespace: a wrapped title is the same title, and it is the one divergence that carries no information. Case changes and truncation stay strict on purpose -- they are the shapes a genuinely DIFFERENT finding takes, and widening to them would trade a visible false positive for the silent false negative the strict match exists to prevent. The residual is now stated in the step rather than left to be rediscovered. The note's wording changed with it. It asserted the report "names a different finding"; both causes reach that branch and the step cannot tell them apart, so it now reports the observation -- "titles its finding differently from the review ... a stale report, or a re-titled one" -- rather than a conclusion it has not earned. Three tests: the reflow case reconciles cleanly, the re-cased case still reports, and the stale case still reports with the new wording. Reversion control: restoring the strict comparison turns the reflow test red. Seventh src.includes() converted -- it pinned the comparison EXPRESSION, so it went red when the comparison gained normalization while the property it names was unchanged. * docs(#3829): describe the flow that ships, not the one implied (m3) Both reference pages said "/gsd-code-review <N> --fix records fixed and skipped, which the gate reconciles from REVIEW-FIX.md" -- true in the abstract, materially misleading in practice, because no shipped path performed that reconciliation. With B1c/B1d wired it is now real, and the pages say WHERE it happens rather than leaving a reader to assume the in-phase gate does it: the gate runs before any fix report exists and writes all-open, and --fix is what records what happened. The round's other behaviour changes land here too, since a reference page that lags the artifact is worse than none: - the disposition column is a closed vocabulary, and a value outside it falls back to open rather than being treated as a decision - severity comes from the section heading when the review uses one, and from the ID prefix otherwise - an unparsed shortfall is stated rather than dropped - titles are compared ignoring whitespace, so a reflowed title still reconciles, and a mismatch is reported as an observation rather than as a claim that the report is stale - EVERY row is carried now, triaged or not, with the cost of the renumbered-finding double-entry stated rather than left to be found docs/FEATURES.md regenerated from the fragment; lint:generated-sync and lint:docs both exit 0. * chore(#3829): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 landed on next in #3954: the acknowledgment is a git commit trailer now, and tests/emitted-drift-acks/ no longer exists. Worth noting for anyone reading the rebase: this did NOT surface as the modify/delete conflict the migration guidance predicts. This branch ADDED its fragment rather than modifying an existing one, and the base deleted only the files that were already there, so the replay was clean and the fragment survived silently into a directory that no longer exists. Quieter than a conflict, and worse -- the gate is what catches it, not git. Two Growth keys rather than the fragment's one: round 2 wired the ledger into code-review-fix.md, so that file grew too. Both key on the bare filename, per the Growth namespace. Emitted-Drift-Ack-Growth: code-review-fix.md — #3829 review round 2, blocker 1c/1d: REQ-REVIEW-09 was unreachable in every shipped path because the in-phase gate runs before any REVIEW-FIX.md exists, so every ledger row it wrote was open and nothing ever reconciled them. This file gains one step, record_disposition, that reads and executes the same lazily-read step after commit_fix_report. It is the only point in the fix flow that is after the report is on disk: code-review.md delegates here through steps/dispatch-fix.md and exits, so it has no such point at all. Growth is one step of prose, no logic is duplicated, and the step is idempotent so the two call sites converge on one ledger. * chore(#3829): the changeset describes the round's behaviour, not round 1's It renders into CHANGELOG, so it carries the same misleading implication minor 3 was about: "the gate ... reconciling fixed/skipped from REVIEW-FIX.md" reads as though the in-phase gate does it, when the gate runs before any fix report exists. Says where it happens, and picks up the round's other user-visible changes -- carried untriaged rows, section-based severity, the disposition vocabulary, and the unparsed shortfall. * fix(#3829): a dotted phase number no longer aborts the step Found by this round's own adversarial review, in its MISSED section: no finding asked about it, and it is the most serious thing in the round after the two blockers. Both callers explicitly accept a dotted phase -- code-review.md:60 and code-review-fix.md:36 both validate ^[0-9]+(\.[0-9]+)?$ and name "03.1" in their own error text -- and both fences reconstructed the path with `printf "%02d" "${PHASE_NUMBER}"`, which cannot format one. Driven with PHASE_NUMBER=3.1: bash prints `invalid number` and exits 1, and under `set -euo pipefail` that aborts the step on its FIRST line. An advisory gate that promises never to block took the phase's entire review report down with it, and the newly wired fix-path call site inherited the same defect. Pad the integer part and carry the sub-number verbatim, so 3.1 -> 03.1 and 3 -> 03, with both arms falling back to the raw value rather than aborting. Driven: 3.1 now reads 03.1-REVIEW.md and writes 03.1-REVIEW-DISPOSITION.md; the integer path is unchanged. Two other findings from the same review, both about claims rather than code: MINOR 2's TEST WAS MIS-NAMED, and the reviewer was right to refute the claim. It called itself the "reflowed" case while substituting triple spaces, which is not a reflow. Driven: a genuinely WRAPPED heading is still not reconciled, because a `###` heading is one line by definition and the continuation is a separate paragraph. Not widened -- absorbing whatever follows a heading into the title would swallow arbitrary prose and make the stale-report check meaningless, and the kept failure mode is the safe one (a visible mismatch note, never a wrong "fixed"). The test is renamed to what it covers and the bound is now pinned by its own test. THE SHELL-SHARING GUARD DID NOT GUARD. Negative-controlling it -- rather than reading it -- showed that deleting block 2's real REVIEW_FILE derivation left it GREEN, on the exact defect it was written for. Block 2 prefixes its `node -e` with `REVIEW_FILE="${REVIEW_FILE}" ...` to put the values in the child's environment, and the detector counted that self-referential pass-through as a derivation. Pass-throughs are now excluded, and the control fires. Pre-existing, not introduced here: the original column-0 anchor matched that same line. Also worth recording: my first attempt at that control silently patched nothing and reported clean. Same lesson this PR already learned once. * fix(#3829): validate the phase number before formatting it, and make the shell guard executable Three findings from the round review's continuation pass, all confirmed by driving them. 1. MY OWN DOTTED-PHASE FIX WAS WRONG on the fallback path. `printf "%02d" abc` writes `00` to stdout BEFORE it fails, so `$(printf ... || printf %s ...)` CONCATENATES the two: `abc` became `00abc`, empty became `00`, and a legitimate `08.1` became `0008.1` because bash reads the leading zero as octal. An unset PHASE_NUMBER also aborted under `set -u` -- in the step that promises never to abort. Validate, then format: never format and fall back on failure. Driven across every edge the review named -- 3.1 -> 03.1, 3 -> 03, 08.1 -> 08.1, 09 -> 09, 1.2.3 -> 01.2.3, and abc / empty / -1 / unset carried verbatim with exit 0. 2. THE SHELL-SHARING GUARD STILL DID NOT GUARD. Excluding pass-throughs was not enough: a structural predicate recognises assignment TOKENS, never assignments that derive a usable value, so `REVIEW_FILE=`, `REVIEW_FILE=$REVIEW_FILE` and a commented-out assignment all evaded it. No regex closes that class. The authority moves to execution -- the third time this PR has learned that lesson. The real second fence now runs in a fresh shell with nothing but the step's two declared inputs and must write the ledger at the correct derived path. All four mutations are caught: empty assignment, self-reference, commented-out, and deletion. The textual check stays as a cheap fast-fail and is labelled as one. 3. THE TITLE-BOUND CORRECTION HAD NOT REACHED THE DOCS. The step comment and both docs pages still said a reflowed title reconciles, contradicting the bound pinned one commit earlier. Superseded prose left standing reads as current to anyone arriving cold, so all three surfaces are rewritten rather than annotated, and FEATURES.md regenerated. Also hoisted `HAS_BASH` to the file's other top-level constants. `const` is in the temporal dead zone until its declaration runs, and a `{ skip: !HAS_BASH }` option object is evaluated eagerly, so a bash-gated test added above the old mid-file declaration threw a ReferenceError that aborted its whole describe and CANCELLED its siblings -- while the summary line still read `fail 0`. It caught three separate additions in this round before I stopped moving tests and moved the constant. * fix(#3829): refuse an out-of-shape phase number instead of carrying it into a path Self-found while writing the prompt for the next review pass, which is the honest provenance: I asked the reviewer whether a path traversal was reachable through PHASE_NUMBER, then checked before dispatching. It was, and I had introduced it. The previous commit's fallback carried an unusable phase number VERBATIM, and PHASE_NUMBER is interpolated into a file path: PHASE_NUMBER='../../etc/passwd' -> REVIEW_FILE=/tmp/phase/../../etc/passwd-REVIEW.md The `printf "%02d"` it replaced had at least mangled that to `00`. A fix that makes a path more reachable than the bug it replaced is a regression, whatever it does for the case it was written for. Both callers already validate ^[0-9]+(\.[0-9]+)?$ (code-review.md:60, code-review-fix.md:36), so this is defense in depth rather than a live exploit -- but the step has two call sites now and should not take either caller's word for its own inputs. It validates the WHOLE value and, on failure, builds no path at all: PADDED is empty and each fence refuses by name rather than coercing. Block 1 declines to report counts read from a path made out of the bad value; block 2 declines to write, which also keeps it clear of the bare-name ledger defect round 1 closed. Driven across the shape boundary: 3.1 / 3 / 08.1 / 09 accepted; abc, empty, unset, 1.2.3, -1, 3., .1, +1, "3 1" and ../../etc/passwd all refused with exit 0 and a named diagnostic. Reversion control: restoring carry-verbatim turns the traversal test red. * fix(#3829): bound the phase number's length, and make the shell guard prove derivation Third adversarial pass. Two of its three refutations were already closed by the previous commit (the ../escape and 1/../../escape traversals, and the unset-input abort); these two were not. 1. A 54-DIGIT PHASE NUMBER WRAPPED SILENTLY. The validator accepted any all-digit value, so `$((10#$_int))` overflowed 64 bits and PADDED became `-7908320945662590977`. Length-bounded now at 8 digits, exactly as the counts already are and for the identical reason -- and the counts guard sitting twenty lines away is why this one is embarrassing rather than subtle. Driven at the boundary: 8 digits accepted, 9 rejected. The bare `${PHASE_NUMBER}` in the suggestion line is hardened to `${PHASE_NUMBER:-}` while here. The empty-PADDED guard makes it unreachable today, but it is one refactor away from an unbound-variable abort under `set -u`, in the step that promises not to abort. 2. THE EXECUTED SHELL GUARD PROVED THE FENCE WORKS, NOT THAT IT DERIVES. A single-phase probe is satisfied by a hardcode, and the review demonstrated exactly that: replacing the derivation with `case ... in 1) PADDED=01 ;; 7) PADDED=07 ;; *) PADDED=07 ;; esac` breaks every real phase and passed the entire suite. It now runs two distinct phases, 7 and 3.1 -- a hardcode cannot satisfy both, and the dotted one additionally pins the integer-part split. The claim "given only the declared inputs" was also overstated: the test spreads `...process.env` (it needs PATH and HOME). The DERIVED names are now explicitly deleted from that environment, so the claim is true rather than merely intended. Also rewrote a comment that had become false: it pinned a describe to the end of the file because of the HAS_BASH temporal-dead-zone constraint, which the hoist removed. Superseded prose left standing reads as current to anyone arriving cold. The changeset's "the gate stays advisory and never blocks" is now verified rather than asserted: both fences exit 0 under an unset PHASE_NUMBER and a traversal-shaped one. * fix(#3829): validate both inputs, refuse before building a path, and never write through a symlink Fourth adversarial pass. Four findings, all confirmed by driving them. 1. PHASE_DIR WAS NOT VALIDATED AT ALL. Unset, both fences died with `PHASE_DIR: unbound variable` under `set -u` -- the same class as PHASE_NUMBER, which I had just spent two commits fixing while its sibling input sat one line away. The step declares two inputs; it now validates two. 2. THE LENGTH BOUND WAS ON THE WRONG THING. The nine-character glob applied to the WHOLE value rather than the integer part, so it falsely rejected `12345678.1` (a legal 8-digit phase) while accepting `1.123456`. Each component is bounded on its own now; the sub-number is bounded too, since it is likewise interpolated into a filename. 3. REJECTED VALUES STILL HAD PATHS BUILT FROM THEM. The refusal guard sat AFTER the assignments, so an unusable input still assembled `${PHASE_DIR}/-REVIEW.md` and stat'ed it before refusing. The guard is now the first thing after validation, and both fences construct paths from validated locals rather than from the raw environment. 4. THE LEDGER WRITE FOLLOWED SYMLINKS. From the review's MISSED section, and the sharpest thing in it: `fs.writeFileSync` follows a symlink, so a pre-existing symlink at the ledger path replaced the contents of whatever it pointed at -- outside the phase directory, with the link left intact so nothing looked wrong. Driven, and the target's contents were gone. This PR introduces the artifact, so it owns the check: an existing ledger that is not a regular file is not a ledger, and the advisory gate says so and steps over. The executed shell guard now draws its phases AT RUN TIME. Fixed fixtures cannot establish derivation -- the review defeated the one-phase version with a hardcode, then defeated the two-phase version by adding one more arm to the same case. Any finite sample loses that race. A phase picked per run cannot be enumerated in advance; the drawn values print in every assertion message so a failure stays reproducible. Control: the review's three-value hardcode now fails on three consecutive runs. Eighth src.includes() converted -- it pinned the literal `${PHASE_DIR}` interpolation and went red when construction moved to a validated local, while "writes a REVIEW-DISPOSITION sibling" was untouched. It now asserts that property, and that REVIEW.md is not written. The changeset's "stays advisory and never blocks" is verified rather than asserted: 8 of 8 hostile-input cases across both fences exit 0 -- both inputs unset, PHASE_DIR unset, a traversal-shaped phase, and a missing phase directory. * fix(#3829): check the ledger path before reading it, and pin the write-safety behaviour Fifth adversarial pass, and the last one this round. Three fixes, three disclosed residuals. FIXED 1. A FIFO AT THE LEDGER PATH BLOCKED FOREVER. readFileSync on a FIFO never returns, so the step documented as "advisory, never blocks" blocked indefinitely -- the literal counterexample to its own headline claim. The non-regular-file check ran after that read. 2. THE UNCHANGED-RUN FAST PATH BYPASSED THE CHECK. A symlink whose target already matched the rendered ledger read through the link, reported `unchanged`, and never reached the refusal. Both fixed by the same move: the check is now the FIRST thing the script does, before any read or write of that path. Ordering was the defect, not the predicate. 3. THE COMMIT TEST FOLLOWED THE LINK the script had just refused. `[ -f ]` resolves symlinks, so the guard and its consumer disagreed about the same path and the helper could still be handed one. `[ ! -L ]` added. And the behaviour shipped with NO regression control -- I hand-drove it last commit and did not pin it, which the review caught by grepping for the words. Five tests now: symlink, symlink-with-matching-target, FIFO, directory, and an ordinary ledger as the negative control so the refusal is not a blanket one. mkfifo goes through the process seam like every other spawn here. DISCLOSED, NOT FIXED -- these are stated in the step rather than carried silently: - TOCTOU between the lstat and the write. Node exposes no portable O_NOFOLLOW write, and an attacker who can write into the phase directory mid-run already has what the check would protect. It narrows a real accident; it is not a security boundary and the docs claim none. - A hard link passes isFile() by construction. - The REVIEW.md and REVIEW-FIX.md reads still resolve symlinks. They are reads of files the operator owns, in their own phase directory. Also narrowed a comment that overclaimed. The randomized guard's domain is FINITE -- 88 integer and 792 dotted values -- so a mutation enumerating all 880 passes forever, and Math.random() is unseeded, so "reproducible" means only that the drawn values are printed on failure. Raising the bar is what it buys; proving derivation is not, and nothing short of reading the fence is. The previous comment claimed otherwise and was refuted. * test(#3829): make the write-safety controls portable to the Windows lane CI caught what neither the local suite nor five adversarial review passes could: every one of those ran on Linux. The FIFO test gated on `mkfifo`'s exit code. On the Windows lane mkfifo EXISTS and exits 0 while producing something that is not a FIFO, so the guard passed, the test ran against an ordinary path, the ledger wrote normally, and the assertion failed for a reason unrelated to the behaviour under test. It now gates on `lstatSync().isFIFO()` -- what was actually created, not what the command claimed. Control: with the shipped guard disabled the test still goes red on Linux, where the FIFO is real. The two symlink tests are skipped on win32, following this repo's existing convention for symlink-planting tests (tests/settings-jsonc.test.cjs:389 skips the same class; tests/unreachable-guard-drift.test.cjs:726 records the reason -- symlink creation requires elevated privileges on Windows CI). The privilege happened to be available on the lane this round, which is exactly why the convention is not "try it and see". * fix(#3829): a bare `|` in a deferral reason is prose, not a parse failure Review round 3, the one blocker. The Source cell is the one field this ledger asks a human to hand-edit, and "waiting on team A | team B to align" is an ordinary thing to type there. The prior-row capture admitted a pipe only when escaped, so a bare one failed the WHOLE line: prior.get() was undefined, the row fell through to `open` with an empty Source, and the console line read "1 of 1 finding(s) open" — a Critical a human explicitly deferred, with a documented reason, rendered indistinguishable from one never triaged, and the reason gone. The exact ambiguity #3829 exists to remove, reachable by one missing backslash. The Source cell is the LAST column, so it is now captured through to the end of the line, less an optional trailing pipe; a bare `|` inside it is prose. The render escapes a bare pipe on the next write so the table stays a table, and the escaped form re-parses to itself, so the second run reports `unchanged` — the fixed point holds. The ledger's own instruction line says so instead of asking the human to escape. Why the property never caught it: SOURCE_CELL only ever appended a PRE-ESCAPED pipe, so the arbitrary built to stress this cell could not reach the one input that broke it. It now also emits a bare pipe, and the round-trip expectation is the escaped form of what the human wrote. A fixed regression case drives the reviewer's exact input through two runs and asserts the decision, the reason, the headline count and convergence. Negative-controlled: both new tests fail against the previous capture. The src.includes() pin on the old capture text is retired for the behavioural case — it was pinning the defect. * fix(#3829): escape every bare pipe in one write, whatever precedes it Round 3, found by the adversarial pass over the round's own fix rather than by the review. The first escape used /(^|[^\\])\|/g, which CONSUMES the character before the pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third run to converge, breaking the advertised second-run fixed point), and an escaped backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare in the rendered table. The property generator emits at most one bare pipe, which is the one case the old form got right, so no property reached either. Scan as pairs instead: an escaped pair (backslash + anything) is kept verbatim and only a pipe outside one is escaped. One write, then a fixed point. Regression case drives `A||B and C\\|D` through two runs; it fails against the previous escape. * fix(#3829): the script leaves by return, so an explicit exit cannot drop its verdict line Round 3, from the adversarial pass over the round's own fix. The embedded node script printed its verdict and then called process.exit(0) -- on the 'unchanged' branch only; the 'recorded' branch fell off the end. Node's "A note on process I/O" documents process.stdout writes to pipes and sockets as asynchronous on POSIX, and process.exit() as forcing exit before pending asynchronous stdout writes complete -- so on a POSIX lane the caller can see exit 0 with no verdict line. This is a hardening against that documented hazard, not a reproduced defect: the reviewer's empty-second-run stdout, which first pointed here, turned out to be its own sandbox -- a bare console.log child printed nothing there either -- and that attribution is withdrawn. The script now runs inside main() and leaves by return on all four early-exit paths, so the event loop drains stdout before the process ends. Same exit status either way, and the || echo fallback is unaffected. A structural test pins the absence of the call (comment-stripped; dotted, bracketed and whitespace-split spellings). The empty-review docs-parity pin that asserted the literal process.exit(0) line is retired -- the round-2 describe drives that property behaviourally. Also widens the property generator: SOURCE_CELL now reaches adjacent pipes and a backslash of either parity before a pipe, the two shapes the first render escape got wrong while passing every input the generator could then produce -- checked against an independent parity-walk oracle rather than a copy of the render's own scan. * fix(#3829): record what an --auto iteration fixed, instead of reporting it open Round 5's major. `record_disposition` runs once, after the whole capped-at-3 `--auto` loop converges — but this workflow keeps ONE final version of REVIEW.md and REVIEW-FIX.md rather than per-iteration copies, and deletes the .iterN.md backups on convergence. A finding fixed in iteration 1 was therefore absent from the final review (it was fixed, so the re-review stopped reporting it) AND from the final fix report (overwritten by the last iteration), so the row fell back to the gate's `open` and rendered `open ... (not in the current review)` — the same bytes a finding that vanished for an unrelated reason produces. That is the one distinction #3829 exists to make, undone by the artifact built to make it. The precise site was the two-arm `applied` construction: for an id the current review does not report, `sameTitle(undefined, h.title)` is false and `title.has(id)` is false too, so the entry entered NEITHER `applied` NOR `staleFix`. It was dropped in silence. Four changes, one defect: - A third arm. When the review does not report an id at all there is no title to disagree with, so this is not the stale-report case — it is what a finding looks like once it has been acted on. Record it. The id-reuse hazard stays closed by the arm below it: when the review DOES report the id, a title mismatch still goes to `staleFix` and is never applied, so a renumbered finding cannot inherit an earlier iteration's `fixed`. - Rows for decided ids the review no longer reports, carried and marked. A decision the ledger cannot render is a decision lost — the same silent drop the carry-forward loop already refuses for prior rows, one source over. - The .iterN.md fix-report backups are read alongside the final report, newest first, so the most recent statement about an id wins — the precedence a duplicate id already gets within one report. - The shell guard proceeds on a fix report, not only on an existing ledger. A direct `/gsd-code-review N --auto` writes no gate ledger, and a converged loop leaves `status: clean`, so a fully successful multi-iteration run recorded nothing at all. And the backups now go in `cleanup_iteration_backups`, after the ledger has read them. #3190's rule is untouched — spent scratch on convergence, retained on degradation — only the timing moved; deleting them inside the loop erased every early fix before anything read it. `CONVERGED` does not survive the loop's shell and is re-derived from the final review's status, which is exactly how the loop sets it; anything but a proven-clean review retains. Seven new regression tests plus an ordering test, all eight reversion-controlled against pre-fix code — every one fires. One is the negative control that matters: a reused id whose title differs must stay `open`, never inherit `fixed`. Residual, stated: an id appearing only in an iteration fix report takes its severity from the id prefix rather than a section heading, because `sectionSev` is built from the current review. That is the documented fallback for carried rows, not a new gap. * test(#3829): pin the two PADDED derivations against a silent desync Round 5's minor 1. Each fenced block runs in a fresh shell and must derive what it reads, so the PADDED derivation — the traversal fence between an attacker-influenceable phase number and a file path, plus the per-component length bound — is duplicated verbatim. Both copies were independently tested and nothing asserted they stay in step, which is the shared-parallel-surface shape CLAUDE.md requires a parity test for, on security-relevant validation logic rather than incidental repetition. Compared line by line rather than through a normalizing rewrite: a normalizer has to be told what may differ, and whatever it is told to tolerate stops being asserted. Exactly one line may differ — each block refuses by its own name — and the test names both forms. It also asserts the slice is substantial, since a parity test over an empty slice passes vacuously. Control: dropping one `?` from block 2's length bound, which moves that copy's limit to 7 digits while block 1 keeps 8, turns it red. That is the exact silent divergence the finding describes. One correction to the finding's own statement, since it is worth recording: the cited lines are :324 and ~:480, which are node-script lines; the derivations are at :48-83 and :211-246. And they are 35-of-36 identical rather than byte-identical — the refusal message differs, deliberately. * docs(#3829): state the PHASE_DIR trust boundary instead of carrying it Round 5's minor 2 asked that the assumption behind PHASE_DIR's validation be confirmed rather than silently carried forward at the two new call sites. It is confirmed, and the comment that stood here was wrong about it: "PHASE_DIR is the step's other declared input and gets the same treatment" describes something the code does not do. Both inputs have the SAME provenance — each caller binds them from `gsd_run query init.phase-op` (code-review-fix.md:7,17; execute-phase.md the same) — so neither is raw user input and neither is more trusted. The asymmetry is not about trust. It is that only one of them has a shape: PHASE_NUMBER carries a documented contract, `^[0-9]+(\.[0-9]+)?$`, asserted by both callers, so a value outside it is provably wrong and is refused. PHASE_DIR's contract is "a filesystem path", which admits `..`, absolute and relative forms and symlinked parents alike; no predicate separates a legitimate planning directory from an illegitimate one, so a shape check would reject working setups while proving nothing. So the emptiness check is adopted as what it actually is — the guard against `PHASE_DIR: unbound variable` aborting a step that promises never to block — and the shape check is declined, with the reason written where the next reader meets it rather than left to be re-derived. The residual is restated in place rather than left in a PR comment: PHASE_DIR may itself be a symlink and the ledger is then written through it, outside the phase directory, deterministically. Left alone deliberately — the write goes where the caller pointed. Not a security boundary, and nothing here claims one. * docs(#3829): record why HAS_BASH is a platform assumption, not a probe Round 5's minor 3 is DECLINED, and the reason is the repo's own contract rather than a judgement call — written at the constant so the next reader does not "fix" it and re-enable what the rule exists to prevent. The gap is real and confirmed: 22 tests carry `{ skip: !HAS_BASH }`, so block 1's bash severity-reporting path has no Windows-lane coverage. But `local/no-unguarded-nonportable-exec` (eslint-rules/no-unguarded-nonportable-exec.cjs, DEFECT.WINDOWS-TEST-PORTABILITY) REQUIRES this guard around `sh -c` / `bash -c` in tests, and its own remedy text names `if (process.platform !== 'win32')` as the sanctioned form, because these constructs fail under Windows Git Bash. So the constant is the repo's answer to this question, not an oversight in this PR. Swapping it for a runtime `bash` probe would light 22 tests up on a lane the rule has already determined they cannot pass — trading a legible, rule-encoded skip for a red matrix. Reversing that is the rule's decision; a change here belongs with a change there. * docs(#3829): describe how --auto's iterations reach the disposition ledger The reconciliation section described the `--fix` path accurately and said nothing about `--auto`, which is where round 5's major lived. It now states that the loop overwrites its fix report each pass, that the re-review drops a finding once it is fixed, that the gate therefore reads the per-iteration backups newest-first, and that the backups are removed after the ledger has read them rather than before. It also states the converged-with-no-ledger case: a fix report on disk is reason enough to record. FEATURES.md regenerated (176 features / 21 groups). Changeset extended to name the shipped behaviour rather than only the `--fix` half. * fix(#3829): clear lint-workflow-shellcheck, a gate the base range added Not from the review. The rebase onto `next` brought in `lint-workflow-shellcheck` (#4109), whose baseline was generated before this PR's new step file existed — so that file's findings are new by construction and `lint:ci` exited 1 on the rebased head before this round touched anything. The last green CI run predates the gate. Caught locally rather than by a red push. Three fixes and one baseline entry, split by whether the finding is real: - STRUCTURAL (not ShellCheck, not baselineable): the guard's `for _f in "…${PADDED}-REVIEW-FIX.iter"*.md` is the bare `for x in $VAR` shape that word-splits differently under bash and zsh. Wrapped in `$(printf '%s' "$PADDED")`, the linter's own prescribed remedy. - SC2097/SC2098, and this one was a genuine latent bug rather than a lint nit: `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"` sat in the same env-prefix list that sets `PADDED`, so its `${PADDED}` expanded the OUTER variable, not the one two entries earlier. Both happen to hold the same value here, which is exactly why it would have kept being wrong quietly. Built before the command now. - SC2317 ×3 is baselined, not fixed. It fires on `return 0 2>/dev/null || exit 0` — the deliberate idiom that lets a fence refuse whether it is sourced or executed — and the verdict is a false positive: the `exit 0` is reached precisely in the executed case. Rewriting a dual-mode refusal to satisfy a wrong unreachability claim trades a real behaviour for a clean report. Baseline 207 -> 210. `lint:ci` exits 0. 173 tests pass across the two touched files. * fix(#3829): a reused finding id no longer inherits the old finding's decision Found by this round's own adversarial review, which drove it rather than reasoned about it — and it refuted the arm I had named as my strongest suspicion, so it is recorded as a correction, not a discovery. Finding ids are reused across re-reviews: the --auto loop renumbers. `row()` inherited a prior decision on an id MATCH ALONE, with nothing checking it was the same finding. Driven: a prior `CR-01 fixed` row against a review reporting a brand-new CR-01 rendered the NEW finding `fixed`. A false decision in the artifact whose entire purpose is telling triaged from forgotten — the same failure mode round 4's blocker was, reached by the other door. I had argued this was closed by the stale-report arm. It is not: that arm guards the FIX-REPORT path only. The PRIOR-LEDGER path had no title check at all. - The ledger now records each finding's title, in the FRONTMATTER rather than a fifth table column: the Source cell is the field a human hand-edits and the one that must escape pipes, and a second free-text column doubles that surface for no reader benefit. - A prior decision is inherited only when the recorded title still matches. An ABSENT prior title inherits, deliberately — a ledger written before titles were recorded carries none, and refusing there would reset every decision in it, which is the loss this guard exists to prevent, caused by the guard. - A decision whose id has been reused is PRESERVED under a `superseded:` key rather than dropped. The review's driven refutation was precisely that the mismatch was surfaced while the decision was lost. It cannot keep a row — the id is taken, and two rows under one id is an ambiguity, not a record — so it is carried in the frontmatter, re-emitted every run, deduped by id+title, and named on the console. - And an iteration-derived decision now cites the report it actually came from. The Source cell hard-coded the unsuffixed `<NN>-REVIEW-FIX.md`, so a decision read out of an iteration backup cited a file that may not exist. A citation the reader cannot follow is worse than none. Also the review's finding. Five new tests. Four fail against the pre-fix step; the fifth — that a ledger with no recorded title still inherits — is a BACK-COMPAT guard and passes both ways by construction. It is not a reversion control and is not counted as one. * fix(#3829): follow the cleanup move through, and stop miscalling a converged run Three loose ends the earlier cleanup relocation left, two of them found by the round's own review and one by the suite. **#3190's own test still pinned the old placement.** T6 asserted the `.iterN.md` removal lives inside `auto_iteration_loop` — exactly what moving it broke. Its SEMANTICS are unchanged and still asserted: removed on convergence, retained on degradation, creation intact. What it now pins additionally is the ordering that forced the move — the ledger reads the backups BEFORE they are removed — and that the loop no longer removes what it just wrote. Rewritten rather than deleted: the assertion was superseded, the guarantee was not. **`CONVERGED` had become a decoy.** With the removal gone from the loop, the flag was set in two places and read in none. Deleted, and the prose that still said "the loop sets it" rewritten to what is true: the loop breaks on exactly one condition, a clean re-review, which leaves REVIEW.md at `status: clean` — and that is what `cleanup_iteration_backups` re-derives from. **A converged final iteration reported the opposite of what happened.** The post-loop message keyed on the iteration COUNTER alone, so a run that converged ON iteration 3 exited with `ITERATION == MAX_ITERATIONS` and printed "Reached maximum iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which every finding was fixed. Convergence is re-derived from the review the loop left behind — the same signal the cleanup step reads, so the two cannot disagree. * docs(#3829): retract two claims this round made and could not support Both were caught by the round's own adversarial review, both were driven, and both would have reached the maintainer. Recording the retraction where the claim was made, rather than only in a PR comment. **The env-prefix "latent bug" does not exist.** An earlier commit in this round claimed that `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"`, sitting in the same `node -e` env-prefix list that sets `PADDED`, expanded the OUTER variable rather than the one two entries earlier — reading ShellCheck's SC2097/SC2098 as a defect report. Driven in bash and in dash: assignments in one prefix list take effect left to right, and the later entry DOES see the earlier one. The warning is a false positive here. The split is kept, but for readability only; the comment no longer describes it as a fix. **The HAS_BASH decline rested on a rule that does not govern these call sites.** It cited `local/no-unguarded-nonportable-exec` as REQUIRING the `process.platform !== 'win32'` guard. Checked, and wrong on both halves: the rule fires only on a file that also chmods an exec bit with an octal literal, and this file has none — so it never runs here — while `eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus `os.platform()`, not one. A constraint that exists is not a constraint that applies, and I did not check which. The decline stands on narrower and honest grounds: whether these fences PASS on the Windows lane is UNVERIFIED. What evidence there is points at divergence rather than absence — the rule's subject line is that `bash -c` constructs "fail on Windows Git Bash", and this PR already measured `mkfifo` existing on that runner, exiting 0, and creating no FIFO. So a probe would not be a clean win; it would light 22 tests on a lane whose shell semantics are known to differ and unknown in detail. That is a measurement to make deliberately, not a change to make in passing. The gap is real and is now stated as a gap. * fix(#3829): close four defects the review drove out of the first title fix The round's own adversarial review re-ran against the reworked tree and refuted two more claims. Every item below is its finding, verified before acting. **An iteration-only decision recorded no title, so the reuse guard leaked.** `applied` stored `{d, src}` and the row took its title from the current review — which does not report the finding at all. The row shipped with no title, and the next review reusing that id hit the title-ABSENT back-compat exception and inherited the old `fixed`. The exact defect the title machinery exists to close, surviving through the hole opened for legacy ledgers. `applied` now carries the title it was decided under. **A changed decision was dropped in favour of the obsolete one.** The dedupe was a has()-guard, so re-superseding a finding whose decision had since changed left the older record standing. It now replaces. **Re-spaced titles double-recorded.** The dedupe keyed on the raw title while `sameTitle()` collapses whitespace; the key now agrees with the comparison. **And the frontmatter was not valid YAML.** `title: Parser: loses data` is rejected outright by a real reader, and the `superseded:` line format was not YAML at all. Values are emitted as JSON scalars — YAML 1.2 is a JSON superset — and superseded records are properly nested. Round-tripped through js-yaml in the tests. One more, self-inflicted while fixing the above: the parse registered each carried superseded record TWICE, once at `- id:` under an empty-title key and again at `title:`. Records doubled on every run. They are collected during the walk and registered once, complete. **T6 was vacuous.** The review flipped `= "clean"` to `!=` in the cleanup and the rewritten T6 still passed — it greps for `FINAL_STATUS`, `rm` and "retained" occurring somewhere, never wiring them to a branch. T6b now EXECUTES the fence in both directions against real files. It fails on that exact mutation. **And a converged final iteration printed two success messages** — the loop's break already reported it. This branch now stays silent and exists only to withhold the degradation warning. Three CI gates the base range brought in, all tripped by this round's own text: - `/gsd-code-review` in a comment — runtime workflow artifacts take the colon form. Now `/gsd:code-review`. - The preamble-ordering parity test: my PHASE_DIR comment wrote the literal `gsd_run` before the shim preamble. Reworded. - Prompt-stuffing: the file passed 50K. I trimmed 5.8K of my own commentary first; even removing every added comment leaves the added CODE over the line, and the file entered this round at 44,523 — 89% of the budget. Added to SIZE_ONLY_WORKFLOWS with the same reasoning the two existing entries carry, and the same acknowledgement: splitting is the real fix. * test(#3829): extract the cleanup fence without an ad-hoc markdown regex T6b's helper used `/```bash\n([\s\S]*?)\n```/`, which trips two of the repo's own rules: `local/no-adhoc-markdown-parsing` (use the sectionizer, not a hand-rolled fence regex) and `local/no-crlf-fragile-split` (a bare `\n` against readFileSync content is wrong under Windows autocrlf). Line-scanned now, CRLF-normalized first — the same shape `bashFences()` in tests/code-review-pipeline-regression.test.cjs already uses, which solved this first. `npm run lint` is clean and T6b still fails on the inverted-branch mutation it exists to catch. * fix(#3829): withdraw the superseded-decision store; keep the identity guard Three adversarial passes over this round each found real defects, and passes 2 and 3 were entirely inside the `superseded:` block added in pass 1 — a second identity scheme, keyed on (id, title), living beside the row store keyed on id. Pass 3 refuted it on three separate counts: a legacy title that merely looked like JSON lost its quotes and fabricated a record; a finding that was deferred, superseded, then returned and fixed left an active row and an obsolete superseded record standing together, reporting `unchanged` forever; and my own test for the replacement path never passed the earlier ledger in, so it guarded nothing. The construct had no terminal state. It is withdrawn. **What survives is the safety property.** The ledger records each finding's title, and a recorded decision is carried forward only while the id still names the same finding. That is what stops a renumbered `CR-01` inheriting an earlier `CR-01`'s `fixed` — a false decision in the artifact whose purpose is telling triaged from forgotten, and the same class as round 4's blocker. **What is given up, and it is disclosed rather than hidden.** On a detected reuse the earlier decision loses its row. The drop is reported on the console naming the id and what had been decided, the previous ledger is committed so the row remains in git, and docs/features/code-review-pipeline.md states the limitation. Two defects from pass 3 are fixed rather than deleted, because they are in the guard and not the store: - **Known-empty and NOT-KNOWN were conflated.** `### CR-01:` yields an empty title; that is a title. While it emitted no `title:` key it read back as a pre-format ledger and inherited across a reused id — the same leak, three passes running. Emitted whenever the title is known, empty included; a carried row no source knows stays absent, which is the legacy-compatible read. Underneath it was a falsy fallback: `(act && act.t) || priorTitle.get(id)` discards `''`. Now a typeof check. - **JSON.parse ran on legacy values.** A pre-format ledger whose bare title was written `"quoted"` was parsed and lost its quotes, so the decision stopped matching. The frontmatter now declares `titles: json` and the parse is gated on it; a ledger without the marker keeps its scalars. One defect from pass 3 is NOT mine and is not fixed here: a converged run prints a success message from the loop break AND another from `present_results`. Both predate this round. My earlier claim that "the duplicate is gone" was true only of the pair I introduced; the pre-existing pair stands, and widening this round into `present_results` is not warranted. 188 tests pass. The three new tests fire against the pre-simplification step. `lint:ci` exits 0. The step file is 55,590 chars, down from a 62,220 peak. * docs(#3829): stop the ledger promising a preservation it no longer makes Fourth review pass. No machinery defects this time — both findings are claims in text this step SHIPS, which is the class this whole stack exists to prevent. **The rendered ledger still said "Re-running the gate preserves every row and every disposition."** That was true until the same round gave the step an intentional drop for a reused finding id, and then it was false in the artifact's own user-facing footer. It now states what the step does, including the one exception, where a reader actually meets it. **And the console asserted "the previous ledger is in git."** Committing the ledger is gated on `commit_docs`, and a failed commit is swallowed — so under `commit_docs=false` the overwritten decision may exist nowhere. The note reports the drop and stops there; asserting a recovery path that may not be there is the same overclaim in a smaller font. Two residuals from the same pass are DECLINED and documented rather than fixed, because both would need the second identity scheme just withdrawn: - A pre-titles ledger carries no titles, so its decisions inherit on the id alone. Refusing there resets every decision in every existing ledger, which is the loss the guard exists to prevent. - Two genuinely distinct findings sharing both an id and a title are indistinguishable to an (id, title) key. The pass also refuted the `titles: json` marker on a ledger written by `b86ea6065^`, which emitted JSON titles before the marker existed. Declined: that revision is an intermediate commit on this unpushed branch and has never been released. The PR's published head writes no titles at all, so a real ledger is either pre-titles (unmarked, bare — handled) or written by the shipped version (marked). The unmarked-JSON state cannot reach a user. Test pinned, and it fails against the pre-correction step. * docs(#3829): fix four wrong citations and one false size justification All four came out of a claim-audit of this round's own response comment — an audit of the text, not the code, which is where the remaining errors were. - **The caller citation was wrong.** The in-code note said both inputs bind from `gsd_run query init.phase-op`. `execute-phase.md:85` uses `init.execute-phase`; only `code-review-fix.md:22` uses `init.phase-op`. The substantive point is unchanged — both are orchestrator-derived, neither is raw user input — but the citation was not checked. - **A leftover "the prior row is in git."** Removed from the console note last commit, left standing in the comment two lines above it. - **The docs still carried the promise the ledger had just dropped.** The rendered footer was corrected; the same sentence in `docs/features/code-review-pipeline.md` was not. - **The SIZE_ONLY_WORKFLOWS justification was false.** It claimed the added CODE alone exceeded the threshold. Removing every round-added comment leaves 47,148 chars against a 50,000 limit, so the file CAN fit — the claim was wrong, and an exemption defended on a wrong premise is worse than no exemption. So the entry is re-justified on what is actually true, and earned first: another **10,188 chars** of this round's own commentary are cut (62,220 → 52,032, from a 44,523 baseline that was already 89% of the budget). Fitting under is possible only by stripping essentially all remaining explanation from logic three review passes found defects in. That is the wrong trade in a file whose house style is heavy in-fence documentation, and the entry says so rather than implying the file had no choice. One measurement corrected while checking: the Windows-lane skip count is **37**, not the 22 the review cited nor the 26 I first counted. Twenty-two and 26 count `{ skip: !HAS_BASH }` CALL SITES; a skip on a `describe` cancels its subtests. Forced the constant false and counted what actually skips. 236 tests pass. `lint:ci` exits 0. * fix(#3829): the drop report is conditional, and two published claims were not A fifth adversarial pass, run against the two commits that went out AFTER the fourth pass and were never reviewed, refuted three claims this round published. 1. The drop is NOT reported unconditionally. `row()` reports only a RECORDED decision (`was.d !== 'open'`); a prior row still at `open` is replaced in silence. The behaviour is right — `open` records no decision to lose — but the shipped ledger legend and BOTH feature docs asserted the report happens every time. Text corrected in all three places, which is the same defect class this round already corrected once for the preservation promise. 2. The test guarding that console wording was VACUOUS: it ran with no prior ledger, so no reuse occurred and its `is in git` assertion could not have failed however the console was worded. Driven through a real drop now, with the drop asserted as a precondition. A new test covers the `open` arm and fails on the pre-fix legend. 3. The HAS_BASH gap is now MEASURED rather than assumed, on native Windows with Git Bash 5.2.37 / MINGW64 first on PATH, node v25.2.1: HAS_BASH left alone: 179 tests, 127 pass, 0 fail, 52 skipped HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped So 37 of the skips are this guard's, confirming the count the round published — and unskipping is NOT a clean win: 24 fail, clustered on `bash -c` quoting and spawn failures, exactly the divergence the eslint rule's subject line names. The guard stays; it now documents a measured gap. The stale "the count is 22" comment is gone. 4. The size-exemption justification was wrong a second time. The overshoot is ~2.3K normalized chars, not "essentially all remaining explanation": the round's committed peak was 59,246 chars (not 62,220, which was never committed), and it entered at 44,466 chars, not 44,523 — both earlier figures mixed bytes into a character measurement. Rewritten to the numbers the scanner actually produces. Also: the shipped comment said both callers validate the phase shape without naming that they validate PADDED_PHASE, not the raw PHASE_NUMBER this step is handed. * fix(#3829): renumber this PR's two REQs, which #3661 took while the branch sat The rebase onto current `next` surfaced a REQ-number collision, not a text conflict. #3661 landed `REQ-REVIEW-08` (`workflow.code_review_point`) on `docs/features/code-review-pipeline.md` while this branch also claimed 08 and 09 for severity surfacing and the per-finding disposition. Two different requirements under one identifier is the kind of thing that reads as correct in both diffs and is wrong in the merged tree. Base numbering wins, because it shipped: `REQ-REVIEW-08` stays #3661's. This PR's two become **REQ-REVIEW-09** (severity surfacing) and **REQ-REVIEW-10** (per-finding disposition). Swept the whole tree rather than the conflict hunk — two references sat in files git merged cleanly and never flagged: - `gsd-core/workflows/code-review-fix.md:450`, the prose stating why `record_disposition` is the step's only reachable call site. - `tests/code-review-pipeline-regression.test.cjs:1782`, the comment on the test that pins that call site. `docs/FEATURES.md` is regenerated from the fragment rather than hand-edited; `node scripts/gen-features.cjs --check` is green (178 features, 21 groups) and `lint:generated-sync` exits 0. Two things stated rather than quietly carried. The `Emitted-Drift-Ack-Growth` trailer on the round-2 commit still reads `REQ-REVIEW-09` for what is now REQ-REVIEW-10 — it is a historical acknowledgment of that commit's growth, and its purpose is unaffected, so it is left rather than rewritten across 52 replayed commits. And `docs/INVENTORY-MANIFEST.json` appeared stale immediately after the replay, reporting two missing `cli_modules/` entries; that was the lane's pre-rebase build output, not manifest drift. Rebuilding in the replayed lane and re-checking shows it in sync and unmodified. Regenerating before the build would have committed the deletion of two base-added entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): reach the title round-trip with a generator that can break it Round 6's only finding. The round-5 title tracking introduced a fresh parser (the `titles: json` / ` - id:` / ` title:` frontmatter walk) and a fresh bijective contract (`JSON.stringify(oneLine(t))` out, `/^ title: (.*)$/` plus `JSON.parse` back in), and `tests/code-review-disposition.property.test.cjs` was untouched since round 4 with no reference to `title` at all. Every heading the generator built was `'### <id>: finding number <i>'` — never a colon, a quote, a backslash, or the empty string. You were right that this is the round-3 shape again, and I would rather demonstrate that than assert it. Two mutations to the shipped step, each a plausible edit rather than a contrived one: A. render `titles: raw` instead of `titles: json`, so the re-parser never JSON.parses and stores the quoted scalar as the title; B. `yv = (t) => oneLine(t)` — the bare scalar, no JSON at all. mutation A — new generator: FAIL old generator: pass (3/3) mutation B — new generator: FAIL old generator: pass (3/3) Both ship past the pre-round suite. The gap was reachable, not theoretical. What changed: - `TITLE`, a new arbitrary drawn from the class the render's own comments say the escaping is for — `:` (why `yv()` exists), `"` and `\` (what stringify/parse must round-trip), the empty string (the known-empty vs not-known distinction the render draws explicitly) — plus scalars that MIMIC the ledger's own frontmatter grammar (`findings:`, `titles: json`, a nested ` title: ` line, ` - id: CR-99`), unicode, surrounding whitespace, and one title long enough to outrun a scanner assuming short scalars. - `FINDINGS` now carries a title per id, so all four properties run the cycle over the title contract instead of over a constant. `IDS` keeps the old id-only shape it is built from. - A fourth property asserting the round trip in the two places it is observable: the stored scalar must `JSON.parse` back to the trimmed heading title, and a hand-recorded decision must survive the next run. The second half is the one that matters, and its construction is the point. The decision is made by EDITING THE RENDERED LEDGER IN PLACE, never by writing a bare row the way the existing properties do. A bare row carries no frontmatter, so `priorTitle` is empty, `sameFinding()` returns true through its `!priorTitle.has(id)` back-compat arm, and the title contract is never consulted — the property would pass over a completely broken round-trip. Both mutations above go green against the bare-row form. That collapse is why the property is written this way, and the comment says so in place. So the assertion is the consequence, not the JSON: a lossy round-trip does not corrupt a title, it makes `sameFinding()` false and resets a human's `deferred` to `open` with the reason gone — this PR's own founding failure mode, reached through the field the round-5 work added. BOUND, stated rather than quietly omitted: the generator emits no CR or LF. A `###` heading is one line by definition, so a newline is not an input the heading parser can be handed; `oneLine()` guards the value's other producers, not this one. Two things found while writing it, both corrected here rather than left: - `runOnce` now returns stdout. The reuse report is a CONSOLE note, not a ledger key, so my first draft's `assert.doesNotMatch(ledger, /^reused:/m)` was vacuously true forever — a test that cannot fail. - `expectedTitle` is a TRIM, not a `\s+` collapse. Collapsing is `sameTitle`'s COMPARISON rule; `oneLine()` is the STORAGE rule and preserves internal whitespace. The collapse form fails on an internal tab against entirely correct code, which is how a test gets weakened instead of believed the first time it goes red. The file header claimed "two properties" while three were running; it now states four, one line each. 239 tests pass across the four pipeline files, 0 skipped. `lint:ci` exits 0 (`lint-workflow-shellcheck`: 203 baseline findings, 0 new). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): the prefix census guard now says four sites, because round 5 added one Self-found, from re-deriving the round-1 finding-id census this round rather than carrying the round-1 verdict forward. The census guard's comment says the prefix set is "written out three times — the heading matcher, the ledger re-parser, and (by its keys) the severity map". That was true when it was written. Round 5's title tracking added a fourth copy: the frontmatter `- id: ((?:CR|BL|WR|IN)-\d+)` matcher that rebuilds `priorTitle`. The guard itself did not fall behind, and the reason is worth keeping visible: `idAlternations()` scans the extracted script by PATTERN rather than walking a fixed list of sites, so the new alternation was absorbed with no edit. Verified by running the extractor at this head — three alternations found, one distinct set, severity map keys `CR,BL,WR` with `IN` on the documented `info` default, 0 domain members not reached. Only the prose fell behind. Corrected, with the pattern-scan rationale stated in place so the next reader does not helpfully convert it into the hand-listed enumeration it deliberately is not — which would be exactly the defect this guard exists to catch, in the guard. Census discharge for this round: re-derived at the rebased head over the extracted shipped script, 3 enumeration sites reached, 0 not reached; the domain (the prefixes `gsd-code-reviewer.md` can emit, walked across both its heading template and its prose Label-equivalence paragraph) is unchanged since round 1 at 4 of 4. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): catch a duplicate REQ id in a fragment, since nothing did Not from your review — this is the test the round owed itself, and I would rather say why than let it look like scope creep. The renumber commit earlier in this round has no reversion control without it. I reverted that fix to check, and the first attempt LOOKED controlled: reverting only the fragment turned `gen-features --check` red. That is the generated-sync gate noticing the projection went stale, not anything noticing the collision. Reverting CONSISTENTLY — fragment plus a regenerated `docs/FEATURES.md` — is silent: gen-features --check rc=0 lint:ci rc=0 pipeline suite rc=0 with two `REQ-REVIEW-08` entries standing in one requirement list. Nothing in the repo reads REQ ids at all, so there was no second place for it to be caught. The failure this guards is a MERGE, not an edit, which is why review does not see it: two PRs open at once each append "the next" REQ number to the same list, and whichever lands second is rebased onto a list that already used it. git merges them as different lines of one file and reports nothing. Neither PR's diff shows a collision — each is correct against the tree it was written on. That is exactly how #3661 and this PR both ended up claiming REQ-REVIEW-08. Scope, stated because it is the part that could be wrong: the check is WITHIN a fragment, never across the corpus. Two different features legitimately both carry `REQ-REVIEW-01..07` — the cross-AI review feature and the code-review pipeline — so corpus-wide uniqueness would be false on the committed tree and would have to be weakened the day it first ran. A requirement list belongs to its feature; that is the scope of the identifier. It lives in `describe('the committed docs/features/ corpus')` because it is an invariant over the committed corpus, which is that block's stated job, and it pins no count — the file's own header rules out counts as shared mutable cells that every feature PR would have to edit. Control: green on the committed tree (no fragment carries a duplicate today); red on the restored collision, naming the file and the id. 85 tests pass in this file. Happy to drop this if you would rather the round stayed inside the review's four corners — but then the renumber ships uncontrolled, and I would rather put that choice in front of you than make it quietly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): finish the census comment correction, which stopped one line short Found by this round's own pre-push adversarial review, which refuted the claim the previous commit made about itself. `7f019d985` said the census comment correction was complete. It corrected one site and left two, both in the helper block twelve lines above the test it belongs to: - `severityMapKeys`' header still read "The THIRD copy: the severity map's keys". With three alternations the map is the FOURTH copy, and has been since round 5. - `idAlternations`' header said "adding a prefix to only two of them is silent", written when there were two alternations and never updated to three. This is the defect the original correction was ABOUT, committed inside the correction: a fragment of prose carries no supersession marker, so a reader landing on line 2810 gets the dead count stated as current fact, and the fixed comment eighty lines down does not reach them. Fixing one surface and leaving its neighbour is not a partial fix, it is the same fix not done. The region is now consistent end to end, and both headers say the thing that actually matters — the scan is by PATTERN, not a fixed list of sites, which is why round 5's new matcher needed no edit here and why converting it to an enumeration would reintroduce exactly the drift it guards. WHILE HERE, a disclosure that was narrower than the truth. `7a6680e8f` said the `Emitted-Drift-Ack-Growth` trailer still names REQ-REVIEW-09 for what is now REQ-REVIEW-10, and left it deliberately rather than rewrite 52 replayed commits. That is right, but it is not the whole set: the message BODIES of `c94106568` ("wire the disposition ledger into the fix path") and `06282f668` ("migrate the emitted-drift ack") both state "REQ-REVIEW-09 was unreachable in every shipped path", meaning the disposition requirement, which is now REQ-REVIEW-10. Same decision, stated at its real size: three historical references, not one. They are commit history rather than living documentation — git is the record of what was believed when — and rewriting the branch to correct a number in a message would cost every review round its correspondence to the commits it reviewed. The TREE carries no stale reference; `docs/`, the workflows and the tests all read REQ-REVIEW-09 for severity surfacing and REQ-REVIEW-10 for the disposition. Regression file: 181 tests pass, 0 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): the third stale count, and a disclosure that over-counted itself Both found by re-running this round's pre-push review after the last fix. It refuted the commit that claimed the region was consistent — for the second time in a row — and it was right again. **The third site.** `:2917` said "And the third copy, which is not an alternation" and `:2919` said "Without this, both regexes can gain a prefix". Written when there were two alternations; there are three, so the map is the fourth copy and it is three regexes that can drift. Worth saying how it survived two passes, because the mechanism is the point and it is the same one this PR keeps re-learning. Both earlier passes VERIFIED with a grep built from the strings I had just fixed — `THIRD copy`, case-sensitive, plus a handful of phrasings I expected. `the third copy` in lowercase matched none of them, and `both regexes` was not a phrasing I thought to look for. A grep returns what you already thought of; that is not a verification of prose, it is a re-statement of your own assumption. The region is now checked by reading it end to end, and all four count statements agree: three alternations (heading matcher, ledger row re-parser, frontmatter `- id:` matcher), with the severity map as the fourth copy. **And the disclosure over-counted.** The previous commit widened the historical REQ-REVIEW-09 references from one to three. Three is wrong. There are TWO underlying statements: - `c94106568`'s message body, and - the `Emitted-Drift-Ack-Growth` trailer on `06282f668`. I counted `06282f668` twice — once as "the trailer" and once as "a body" — when its only mention IS that trailer (`git show -s --format=%B 06282f668 | grep -c REQ-REVIEW-09` outside the trailer line: 0). Over-counting is the safe direction and it is still a wrong number in a message, which is the thing this round has been correcting all along. The decision is unchanged: both are commit history rather than living documentation, and rewriting the branch to fix a number in a message would cost every review round its correspondence to the commits it reviewed. The TREE carries no stale reference — 08 is #3661's `workflow.code_review_point`, 09 is severity surfacing, 10 is the per-finding disposition. Comment-only in one test file; no assertion, regex or extracted-script expectation moved. Regression file: 181 tests pass, 0 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * fix(#3829): join the disposition-step dispatch so REQ-LANG-04 inheritance is provable `lint-response-language-coverage` (#2529, which landed on `next` after this PR was approved) reported `execute-phase/steps/code-review-disposition.md` as having no response-language coverage. The step does inherit it: `execute-phase.md` imports `references/execute-phase-response-language.md` and dispatches the step with `Read and execute`. The dispatch stub wrapped, leaving the verb at the end of one line and the path at the start of the next, and `namesFragmentAsEntryPoint` matches within a single line — so a genuine inheritance was unprovable to the linter. Rejoining the verb and the path restores it: `namesFragmentAsEntryPoint` goes false -> true and the lint reports `OK (165 workflows covered)`. Only line breaks move — the word stream is identical to the previous revision, and the file is unchanged at 93,390 bytes, so no growth acknowledgment is owed. This takes the third coverage form the lint documents — inheritance — rather than the inline directive the CI message names first. Where inheritance is provable the lint's own comments say a second copy "buys no coverage and adds a sentence that can drift", and the step file already sits over the prompt-stuffing threshold. Swept all 76 fragments in the catalog: this is the only one whose parent's previous line ends with a dispatch verb. The 17 others that are mentioned without a provable entry point are table-routed or bare prose references carrying no dispatch verb at all, and correctly hold the pinned inline directive instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JgX6QQmygeZnQqbc3o8RNC * chore(#3829): regenerate derived artifacts after rebase onto next The rebase onto current `next` conflicted on the 19 install-tree goldens and `docs/FEATURES.md`. Those are generated, so the conflicts were resolved arbitrarily and the generators re-run (`npm run regen:derived`) rather than hand-merged — a clean textual merge of a generated file attests the merge, never the content. Reconciled per artifact against the base's own committed copy rather than against the pre-regen tree, because the pre-regen tree is the arbitrary resolution: - all 19 `tests/fixtures/install-tree/*.json` now differ from `upstream/next` by exactly one key, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`; - `docs/FEATURES.md` differs by exactly REQ-REVIEW-09/10 and this PR's own reference section; - `docs/INVENTORY-MANIFEST.json` differs by exactly the same one step file, and needed no regeneration to get there. Nothing the base added was dropped by the arbitrary resolution: the restored entries (the `gsd-core/agents/` and `gsd-core/commands/gsd/` families, the compact templates, the `detail/elaboration.md` files, `gsd-secret-read-guard.js`) are all base-owned and came back through the generator, which is what the resolve-arbitrarily-then-regenerate discipline is for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV * fix(#3829): repair the rebase's conflict resolution in the regression suite The rebase onto current `next` hit one add/add conflict in this file: #4209's external-reviewer-evidence describe and this PR's #3829 block were added at the same insertion point. Resolving it by keeping both sides was correct in substance and wrong in mechanics — the conflict boundary cuts through two open blocks that the SHARED trailing ` });\n});` closes, so each side carries +2 unbalanced braces on its own and concatenating them left the file with 683 `{` against 680 `}`. `node --check` fails outright, so the whole file deregistered rather than failing a test — 188 tests silently stopped existing. Rebuilt the region as a real three-way merge (ancestor |
||
|
|
fac0e9de86 |
docs(#4906): ADR-4910 amendment — a write refuses on a document with any unreadable node (#4913)
* docs(#4906): ADR-4910 amendment — a write refuses on a document with any unreadable node §5 scoped the parse error to the node. That is correct for reads and was never examined for writes. §5 was reasoned entirely from #4899, a read bug. Node-scoping is right there: a ragged Progress table must not make phase list and init.progress fail, because those are the commands a user needs to see what to repair. But the rule was stated unqualified, and it licensed something never considered — phase.complete mutating a ROADMAP.md whose Progress table it could not read. A partial-view write persists a wrong answer rather than merely returning one. Amendment: reads stay node-scoped; a write refuses when ANY node in the document carries a parse error, whether or not the mutation targets it. A reader answers a bounded question from a bounded region; a writer asserts that the document it emits is the document it read, and cannot make that assertion about a region it could not parse. §3's byte-stability does not rescue it — splicing untouched bytes faithfully is not the same as knowing they were consistent with the change. Rejected: refusing only when the mutation's own target node is unreadable. It is the appealing middle and it does not hold — phase.complete writes **Plans:** while deriving that value from plan/summary counts in a different region. The regions a write depends on are not statically the regions it touches. Downstream: Phase 1 ships both scopes plus a hasUnreadableNodes predicate the serializer consults; Phase 2 gains a new acceptance criterion (a refused write leaves the file byte-identical on disk); Phase 4's census gains a second axis and must publish both counts. Appended as a dated section per docs/contributor-standards.md pattern 1 — the original Decision body is unchanged. Refs #4906 Refs #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4906): correct the amendment's motivating example and state the limit it exposes An isolated review pass falsified the first draft's central example. The correction is recorded in the amendment rather than quietly patched, because it bounds what the rule can promise. - The draft claimed phase.complete derives the value it writes into **Plans:** from a different REGION of the same document. It does not. planCount and summaryCount come from findPhaseInternal at src/phase.cts:3429-3434 — a filesystem scan of the phase directory. No PlanningDoc node holds them. That widens the dependency rather than narrowing it, and it makes an explicit limit necessary: a document-scoped write refusal protects the document's own consistency and says nothing about the correctness of a value sourced from outside the document. Recorded as a stated limit and named out of scope for this epic. - The rejected alternative (refuse only when the mutation's own target node is unreadable) is re-grounded on two arguments that survive: it would almost never fire, since a verb locates the node in order to write it; and it guards something smaller than the operation performs, because serialization re-emits the whole file. - §5's body names `phase list` as a consumer an unreadable Progress table would block. It is not one — cmdPhasesList (src/phase.cts:207) enumerates phases/ and never opens ROADMAP.md. init.progress and roadmap.analyze are the real instances; the argument never needed three. §5's body left unmodified per the append-only rule, correction carried in the amendment, fold in at ratification. - Clarified that the new WriteOutcome refusal sits on a different axis from §5's reserved document-level Result<PlanningDoc> parse failure, so no reader can conclude that reservation was reopened. A document can be valid to open and still refuse to be written. Refs #4906 Refs #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
01dbda9c49 |
docs(#4910): ADR-4910 — the PlanningDoc parse → mutate → serialize seam — Phase 0 of #4906 (#4911)
* docs(#4910): ADR-4910 — the PlanningDoc parse → mutate → serialize seam — Phase 0 of #4906 Design lock for epic #4906. Docs-only; no production code lands here. Eight decisions: one PlanningDoc seam composing markdown-sectionizer, markdown-table and frontmatter as layers; node-replacement writes so a field write cannot reach past its own value; byte-stable serialization for untouched regions; escape-or-refuse shared between each artifact's writer and its reader, with an explicit accepted-superset-of-emittable split; a typed parse error scoped to the node rather than the document; a type-narrowed write boundary paired with a lint; a positive control per accepted grammar; and one implementation per shared pattern. Two corrections to the epic's stated mechanism, both load-bearing for later phases: - The epic asks for a ratchet where a reintroduced content.replace() "does not typecheck". It cannot: fs.writeFileSync(p, s.replace(...)) typechecks fine because fs has never heard of PlanningDoc. Enforcement is type-narrowing plus a lint that owns the bypass, and the ADR says so rather than shipping a guarantee one require() defeats. - The epic names scripts/lint-planning-artifact-writer-drift.cjs as the drain point. That script is a registry-completeness guard that states "No ratchet / no baseline" by design and never inspects how a write is performed. It is correct on its own axis and left untouched; the right home is local/no-adhoc-markdown-parsing. ADR-2143's Phase 4 already shipped the table-regex and replace-mutation detectors, so the ADR scopes the remaining gap precisely rather than asking for that work twice: adhocReplaceMutation keys on a table-or-section regex, and #4852's pattern is a bold-label field regex — which is why the defect sits in src/phase.cts, a file the rule does lint, with lint:ci green. Per docs/adr/README.md lifecycle rule 3, ADR-1372 and ADR-2143 are NOT given the reciprocal `Subsumed by` back-link here: a Proposed ADR's Subsumes claim is prospective, so its targets are not marked until ratification. Both back-links land in the Phase 6 ratification PR. ADR index regenerated. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4910): apply review findings — link every ADR cross-reference and state the node-scoping interpretation Three review passes ran against the ADR: an isolated adversarial fact-check of every citation, and both axes of /code-review as separate sub-agents. Standards axis (hard violation of docs/adr/README.md lifecycle rule 2): 14 bare ADR-1372 / ADR-2143 / ADR-1411 references in body prose. gen-adr-index.cjs does not catch this — its bare-id check runs only over relation-field values, never body prose — so a green lint:generated-sync did not clear it. Every bare cross-reference is now a file link; only the self-reference ADR-4910 remains bare, which is not a cross-reference. Spec axis: §5 scopes the parse error to the node, while #4906's criterion reads "an unparseable shape surfaces could-not-parse with the offending span" with no document-or-node qualifier. Both readings are faithful to that sentence and they produce materially different Phase 1 and Phase 4 work. §5 now records the distributive reading explicitly, names the document-scoped alternative, states what it would cost (one bad table failing phase list and init.progress alongside roadmap.analyze), and names what changes if the epic meant the other one. A silent narrowing became a stated one. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4910): make each phase's acceptance a structural property, not a list of fixed issues Phases 2-5 read as "fixes #4852, #4862, #4499" — a point-fix list with a seam attached. That is the failure the epic names in its own words: "an implementation that does that has not closed this epic, even with every symptom gone and CI green." New section "What makes a phase done" locks three criteria every phase carries: - Census -> zero. A phase enumerates every instance of its anti-pattern in the tree, publishes the count in its PR, and closes when it is zero. Not "the reported ones". - Deletion, not coexistence. Bespoke implementations are removed, not kept in sync beside the seam — including src/roadmap.cts:1196's three-arm planCountPattern, which is correct today and still goes, because a correct copy of a rule the seam owns is the two-copies-that-agree case. - Unrepresentable by construction. Each phase ships one property that makes its class impossible rather than currently absent: a property over generated documents, a type that does not admit the wrong shape, or a drift guard. The absorbed issues are demoted to fail-first regression evidence. A phase may not close on those tests alone. Each phase restated accordingly, with its own census / deletion / unrepresentable / evidence breakdown. The six community point-fix PRs and their issues were closed unmerged (#4762/#4736, #4897/#4837, #4848/#4661, #4610/#4605, #4609/#4606, #4530/#4499). The ADR now records that as executed rather than pending, which is what lets census-to-zero be an acceptance criterion at all — a landed point fix would make the tree look healthier than it is. Phase PRs reference those six with Refs, since CONTRIBUTING forbids a closing keyword against an already-closed issue. Refs #4906 Closes #4910 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
06845717fe |
feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition Failing-first coverage for the Loop Host Contract role partition. At this commit crossCheckRoleFamilies does not exist, so the rows throw "crossCheckRoleFamilies is not a function" -- the RED proof they bind to behavior rather than restating it. ADR-894 section 3 assigns roles per step but parenthesises the assignment as "(illustrative roles)", and nothing enforced it. The only thing standing in the way was a single deepEqual in this same file, which is editable prose. Rows cover: each step's own family accepted; a strict subset accepted; a foreign role rejected at every step; an unknown role rejected; an unknown step failing CLOSED; capitalization not silently matched; every offending role reported rather than only the first; and purity, because buildContract puts the same array into the generated contract. Two rows exist because an earlier cut of this suite was vacuous. The purity fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot fail an in-place sort(), and the mutant was being killed by three unrelated rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the exact same role-name domain, both directions: they are parallel constants over one domain, so divergence is the generative-fix class CLAUDE.md names. Every negative row asserts the offending ROLE NAME and the STEP NAME appear in the message. A count-only assertion survives a mutant that reports the wrong role, which the 80% Stryker gate would surface only after a full CI round-trip. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#4740): reject a cross-family agent-role declaration Orchestration and execution are distinct functions of the loop and must not drift into one another. That partition was real but unenforced: ADR-894 section 3 calls its own role assignment "illustrative", and the generator accepted anything. Adding orchestrator to execute-phase.md's agent-roles line compiled, --check passed once regenerated, and capability-validator.cjs then began accepting into:"orchestrator" at every execute point. ROLE_FAMILY maps every role to one of orchestration, planning or execution. EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family. crossCheckRoleFamilies rejects a cross-family role, a role outside the vocabulary, and an unknown step. It reports every offender, not the first. It fails CLOSED on an unknown step, deliberately diverging from assertPointsCoverage's "unknown step -- caught elsewhere". For points that is true: the canonical-set and duplicate checks catch it. For roles there is no second net, so failing open would leave an unknown step as the one input that bypasses the gate. crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles to agent FILES and the orchestrator is the host, owning none -- admissibility and agent-file presence are separate concerns with separate checks. Additive to section 3's existing rule that contribution.into must be a member of the step's agentRoles, which is unchanged. That governs what a CAPABILITY may target; this governs what a WORKFLOW may declare. No capability is affected, and all five workflows already declare single-family sets, so the gate is green on the commit that introduces it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4740): make the ADR-894 role assignment normative Section 3 parenthesises its per-step role assignment as "(illustrative roles)". That word was accurate about the list's PURPOSE -- it illustrated the shape of a generated contract entry -- and wrong about its STATUS, because the assignment was load-bearing from the moment the generator consumed it. Read literally it makes the partition an example rather than a rule. Appended as a dated in-place section per docs/contributor-standards.md, which records that an accepted ADR is never rewritten and names this the default pattern. Section 3's original body is untouched. The amendment states the three disjoint families, the one family each step admits, that a step may declare a strict subset but never outside it, and why this is a clarification rather than a new decision: the contract is generated from the workflow markers "so it cannot drift into a lie", and all five workflows have always declared single-family sets. What was absent was any statement that it is required, and any check that it holds. It also pins the distinction that is easy to re-merge: contribution.into being a member of agentRoles governs what a CAPABILITY may target and is unchanged; the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary entry for the Loop Host Contract records the same, beside the agent-reference drift guard it already documented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): backfill changeset pr number Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both go from invalid_pr(0) to ok. Without that env both report success without evaluating the branch at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): stop injecting the orchestrator procedure into executors claude-orchestration declared a contribution at execute:wave:pre with into:"executor". loop-hook-dispatch.md defines a contribution as "inject fragment.inline verbatim into the context for the role named in into", so its 267 lines were injected into EXECUTOR prompts whenever the capability was enabled. Those lines are orchestration end to end -- construct a wave manifest, resolve the dispatch backend, invoke the Workflow tool to spawn executors, bridge per-agent results into the merge chain. An executor can act on none of it. Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT carries no orchestrator entry by design: the orchestrator IS the host, and the host's procedure lives in execute-phase.md. A step's agentRoles enumerates agents a capability may inject context INTO, so adding orchestrator there would model the host as an injectable agent -- the same category error pointed the other way, and it would need an exception carved into the partition the same issue just made normative. So the defect is the mechanism, not the label. A contribution injects into an agent's context; "replace step 3's inline dispatch loop" is a change to what the HOST does. The contribution channel was serving as a host-behaviour directive because it was the only channel available at an execute point. The entry is removed. plan:post into:"planner" is correct and untouched. The procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the capability -- it is the only copy in the repo -- and is no longer injected anywhere. Consequence, not softened: the Workflow backend now has no loop wiring. Detection, emission and config remain and the design is intact, but nothing dispatches it. Under the separation ADR-1143 itself asserts it never had a legitimate channel; ADR-1143's own audit already records the end-to-end path has never been exercised. Wiring it properly needs a host-level mechanism that does not exist today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): invert the stale execute:wave:pre registry assertions Removing the contribution left four surfaces asserting or describing the old state. Caught by an isolated review before a verification run was spent, which is the point of reviewing first: the first of these was a guaranteed CI red. execute-wave-post-gate-pipeline-e2e asserted against the REAL generated registry that byLoopPoint['execute:wave:pre'] held exactly one contribution with capId claude-orchestration. It now holds zero. Inverted to assert exactly 0 -- not a vague >= 0 -- and the #2285 comment above it now explains the current state rather than the one it was written for. CONTEXT.md's Claude Orchestration entry claimed two contributions at wired points. It is now one, and the entry's execute:wave:post label was already wrong before this change: the manifest said execute:wave:pre. Rewritten to one plan:post contribution, why the execute-point one was removed, and where the procedure now lives. One assertion in claude-orchestration.test.cjs could not fail. It tested for the prose "(into the executor)" while the doc says "(`into: executor`)", so no plausible wording matched it and the paired plan:post assertion was carrying the row. Replaced with a check on the structural claim, and proved RED by restoring the two-contribution wording before reverting. The moved procedure keeps section headings that speak as a live contribution -- "When this contribution is active", "Why execute:wave:pre". Preserving the body verbatim was deliberate, so the headings stay and an editor's note under the header explains why they read that way. A sweep of all 17 files referencing byLoopPoint found no further siblings: the remaining hits are a synthetic capability fixture and an empty-points test that already expected no active hooks, both correct before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b647e28313 |
chore(#4727): name the tool-conversion helpers for the runtime that uses them (#4732)
* chore(#4727): name the tool-conversion helpers for the runtime that uses them
GSD has had no Gemini runtime since #1928 removed it (Google sunset Gemini CLI
on 2026-06-18, shipped 1.8.0), yet two helpers were still named for it:
claudeToGeminiTools -> claudeToAntigravityTools
convertGeminiToolName -> convertAntigravityToolName
The sole consumer is convertClaudeAgentToAntigravityAgent, whose own comment
read "Map tools to Gemini equivalents (reuse existing convertGeminiToolName)".
Nothing named Gemini consumes them, because nothing named Gemini exists. The
new names follow the convention the file already sets with its neighbouring
Copilot pair, claudeToCopilotTools / convertCopilotToolName.
Zero behavior change. Every mapped VALUE is byte-identical, deliberately:
read_file, write_file, replace, run_shell_command, glob,
search_file_content, google_web_search, web_fetch, write_todos
Those are Gemini's built-in tool dialect and Antigravity genuinely speaks it.
This rename covers only the identifiers, which are the one part of the surface
that was GSD's choice rather than Google's contract.
Renamed in BOTH copies. CLAUDE.md labels bin/install.js "(generated)", but no
script emits it -- build:lib is tsc -p tsconfig.build.json and writes only
gsd-core/bin/lib/**. These converters are the #1099/#1173/#1182 situation: they
were extracted into src/runtime-artifact-conversion.cts while bin/install.js
kept its own working inline copies, so each symbol existed twice in two
independently hand-maintained files. Renaming one would have left two names for
one concept. Verified first that no capability descriptor resolves either by
name -- antigravity's descriptor names only convertClaudeCommandToAntigravitySkill
and convertClaudeAgentToAntigravityAgent, neither of which moved.
Comments keep their reasoning and their issue refs (#3362 AskUserQuestion,
#1394 Skill/SlashCommand); only the subject is corrected, from "Gemini CLI" to
Antigravity speaking the Gemini dialect. Those describe the dialect's behavior,
which is still Antigravity's behavior, so deleting them would destroy the record
of two real bugs.
docs/research/gemini-to-antigravity-migration.md is left unedited and carries a
dated addendum instead: it is pinned to
|
||
|
|
c0b2a05d2f |
fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser The isolation guards decided whether a run-scoped sentinel applied to a dispatch by regex-scraping model-authored prose. The scrape returned values in a different namespace from the ones the sentinel records, so the comparison could never succeed: sentinel { phase: "03", plan: "03-02-hardening" } <- $PHASE_NUMBER / $plan_id prose "Execute plan 02 of phase 03-auth." scraped { phase: "03-auth.", plan: "02" } <- greedy (\S+), both wrong #4594 reports only the phase half. Measured against a real phase-plan-index run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose carries a bare in-phase plan number — so the plan field mismatches too, and the Claude path is dead rather than latent. A fresh sentinel was therefore discarded on every executor dispatch and every legitimate ISOLATION=none degrade was denied, leaving the work unrun. hooks/lib/dispatch-identity.js is now the single owner of both halves. The two prompt-body producers emit a canonical marker carrying the same shell values the sentinel records, so producer and consumer agree by construction. The prose frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and deliberately reporting no plan — an absent identifier means "cannot compare" and is safe; a wrong one is a false mismatch and is not. The prose sentence itself is byte-identical: the executor agent reads it too, so the marker is purely additive (Hyrum's Law). An inapplicable sentinel is now named in the guards' deny reason instead of being dropped silently — the silence is why this survived three producers and two consumers unnoticed. Interpolated values come from a sentinel file and from prompt text, so both are length-bounded and stripped of control characters. ADR-4630 locks the seam and maps the epic's three phases. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4594): resolve eight review findings across the dispatch-identity seam Three orthogonal review engines ran on 43418af144 — the code-review skill's Standards and Spec axes, and an isolated adversarial security pass — plus a self-review of the committed diff. Every finding is fixed here; none deferred. F1 (major, reproduced). A keyless or unknown-key-only marker — the literal "[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and returned source:'marker' with both fields null, suppressing the prose fallback entirely. Any prompt text containing that literal silently disabled identity narrowing, so a fresh sentinel applied to a dispatch it was never scoped to, defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker that yields neither recognized key is no longer a marker: the scan continues to later markers, then later texts, then prose. Forward-compatible tolerance of unknown keys is unchanged. F2/F3 (major). The first cut duplicated sanitizeForReason, describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across both guards — the exact defect class this epic exists to delete, and with no cold-load justification, since both hooks already require hooks/lib/. They now live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in isolation-sentinel.js beside the comparison it mirrors, returning the nested {sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke four-field bag that renamed the pairs already flowing through the seam. F4 (hard violation). The visibility test asserted on the deny reason's prose. CONTRIBUTING.md prohibits raw text matching on hook output, which is why every deny carries a stable reason_code. The discard is now a structured sentinel_discarded field on each hook's stdout JSON, and the test asserts that; the sentence stays for the operator but is no longer the contract. F5 (hard violation). The 64-character truncation limit had no boundary coverage. 63/64/65 are now exercised against the single consolidated helper. F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or the bidi overrides, so a crafted value could still reflow or reverse the message. Both classes are stripped, with a test each. F7 (major). The producer/template parity test was vacuous — it rendered a marker and re-parsed its own output, and would have passed with both templates deleted. It now reads the two workflow templates, extracts each marker line, substitutes the measured values and asserts the owner's parser returns them. Proven red by deleting one template's marker line before being proven green. F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the orchestrator-worktree path because that prompt is built in shell. It is not: executor-isolation-dispatch.md:131 says plainly that those are template placeholders, not shell variables, so {plan_id} is model-substituted there too. A false guarantee in a design lock is worse than a stated limit. Both documents now say the marker is model-substituted on both paths and that the prose fallback is the real floor everywhere. The "3 workflow templates" count was also wrong — 3 prose sites across 2 files, 2 of which carry the marker. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth Refs #4630. The dispatch-identity marker and its substitution note grew gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts two real-tree guards that lint:ci does not run: - tests/benchmark-compact-content.test.cjs asserts the committed baseline is "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on 23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline regenerated with scripts/benchmark-compact-content.cjs --write. - tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer for any emitted file that grows, keyed on the bare filename. Added below. The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line inside the Agent() prompt's <objective>, and the note telling the orchestrator to substitute {plan_id} with the plan's id verbatim. Both are load-bearing -- the marker is what lets a guard hook match a dispatch to the sentinel the per-plan gate wrote, and without the note the orchestrator has no instruction telling it the value must not be paraphrased. Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): set changeset fragment pr to 4693 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7bcfbe4542 | docs(#4629): add ADR-4629 — STATE.md write intent beyond frontmatter (#4645) | ||
|
|
8fa2c3dbcf |
docs(#4650): record the path-containment and filename-classification design lock (#4655)
Phase 0 of epic #4636. ADR-4650 fixes the decisions Phases 1-4 inherit, so that four phases do not each invent them independently. The load-bearing decision is that the engine and the exported shape are separable. A resolver-based, symlink-safe containment predicate already exists as validatePath, and building the epic's literal assertWithinRoot() from scratch would create a sixth implementation of the very thing this epic consolidates -- while risking silent loss of behavior validatePath acquired as bug fixes (a closed dangling-symlink existence oracle, ancestor canonicalization for non-canonical roots, a separator-aware boundary test). But the epic's other clause is correct and lands on the current export: validatePath returns a boolean a caller can forget to check, and populates `resolved` with the escaping path precisely on the traversal branch. While that form stays exported, the Phase-4 ratchet could only assert that a helper was called -- validatePath(x, root).resolved would pass the rule. So: preserve the engine, narrow the export. assertWithinRoot becomes the only export and yields a branded ContainedPath. Also recorded, each found by measurement rather than from the epic text: - The rejection message text is a real contract. tests/quick-batch.test.cjs asserts a user-facing `reason` field matches /escapes allowed directory/, so the string reaches CLI consumers and Phase 3 must preserve it verbatim. - Two further unconfined boundaries the epic does not enumerate: resolvePath and gap-analysis.plan-post, both in check-command-router.cts. - --phase-dir also interpolates into ${PHASE_DIR} for command-exit-zero, so confining at the boundary covers both predicate kinds; the evaluator stays fs-free. - opts.allowAbsolute is a per-call-site liberality knob, which is an acceptance policy living exactly where this ADR says it must not. - planning-inspect's isWithinRoot is deliberately pure-string with no I/O; its contract differs, so Phase 3 decides rather than assumes. The acceptance policy is stated once: conservative about the resource, exact about the classification. #4580's guard was not too strict, it was wrong -- a category error comparing a whole suffix against a set of final extensions. ADR opens as Proposed; ratified at Phase 4 closeout per docs/adr/README.md. No changeset: the diff touches docs/adr/ only, which is outside USER_FACING_PREFIXES in scripts/changeset/lint.cjs. Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d65c248e5 |
fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector Tests only, committed ahead of the implementation so the RED run is real. - tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative cases proving seam calls and path-call-plus-slash-literal are not platform signals; positive pins that genuine platform content, seam-bypassing spawns, chmod and symlink still classify in; macOS signal set and generated list unchanged. - tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest rows and test-conformance still has 3 windows + 1 macOS. - tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint test does, and resolveSelection rejects the retired windows scope. Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): delete the second Windows selector and narrow the conformance tier Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped conformance tier on real Windows/macOS — was not met. Measured on PR #4640 (run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7 changed test files running on a real Windows runner twice. Two selectors, only one in the epic's scope. The test job's three scope:windows shards predate the epic (#494, sharded #3057) and gate on product_changed, not full_matrix, so they fire on every product PR whatever Phase 3's classifier decides. They are deleted; test-conformance becomes the sole Windows selector, as it already was for macOS. Non-Linux jobs 7 -> 4. Gating the lane instead was rejected as provably redundant: for a test file reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file), and that same predicate sets full_matrix, which turns test-conformance on. Every file a gated lane would run is already covered in the same run. The lane's one non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix instead of feeding a parallel lane. Two detectors matched the repo's own test idiom rather than any platform signal and carried 226 of the tier's sole-signal membership against 41 for the other eight: process-seam-subprocess (335 files, 118 unique) matches the tests/helpers.cjs entry points nearly every CLI test uses, and going through the seam is the opposite of a platform signal since shell-command-projection takes platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs only a path call anywhere plus a slash literal anywhere, and that class is already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed. Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured. Adds the size gate Phase 2 never had, as a ratio against a live denominator so it cannot stop binding as the suite grows. 292 files leave real-OS Windows execution. The drop-out set was audited: 14 have a platform-suggestive filename and all 14 are static source-text analyses or seam-mediated CLI tests. raw-child-process was investigated as a suspected false negative and left unchanged -- relaxing it adds 13 files, all false positives. macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated macos-conformance-tier.generated.cjs is byte-identical at 196 files. Fixes #4641 Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): register the new ADR path in the docs-guard exempt baseline tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md in a comment justifying the retired windows scope; lint-docs-guard-registration tracks that reference set, so the baseline needs the new path. Verified the exemption still holds: the path is prose, not a filesystem read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): make the escalation tier-backed and drop every hardcoded count Three follow-ups from measuring the first pass rather than trusting it. The windows-hint escalation now requires tier membership as well as the filename hint. Setting full_matrix runs test-conformance, which runs only the tier; escalating on a test that is NOT in the tier costs four jobs and still never runs that test on Windows. Measured over the 16 RULES entries the narrowed predicate fires on exactly the same rules today, so this is correct-by-construction rather than a behavior change. The broader variant -- escalate on any tier member a rule pulls in, ignoring the hint -- was measured at 14/16 rules and rejected as over-broad. Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as soon as the rebase pulled in one new test file from #4253, which is the whole argument against them: the ceilings are ratios against a live denominator, the committed lists are pinned by comparison against a fresh classification of the live tree, and the three named probe files now assert on their SIGNAL rather than on membership in a literal list -- asserting by filename is the exact error this PR fixes in the classifier. Regenerates both lists against the rebased tree. Same-tree figures are now 547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none added; macOS is unchanged at 197 with a zero-line diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke An isolated adversarial review found a real false negative. Removing the blanket process-seam-subprocess detector also removed the only coverage for tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook spawns options.interpreter via real spawnSync, so runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary executing a shell script extracted from workflow markdown. The seam argument holds for src/shell-command-projection.cts, which takes platform as an injected parameter; it does NOT hold for the test helpers, which spawn real binaries. Conflating the two is what made the blanket detector look purely noisy -- it was 99% noise wrapping a real signal. Adds a narrow shell-interpreter-spawn category keyed on a real interpreter option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33% ceiling. All 9 confirmed by reading the matching source line, zero comment or fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured and rejected -- each adds 9 files but misses the counterexample entirely. Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD strictly rather than carrying the Windows report-only carve-out. The carve-out now lives solely on test-conformance, whose matrix does include windows. Repairs seven pre-existing assertions the category removal invalidated, preserving each case's purpose rather than deleting coverage, and converts the last hardcoded tier bounds to live-derived ratios -- including the macOS sanity range that was still a magic [100, 350]. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): keep the confinement test on a real OS via a documented allowlist A security review found tests/external-descriptor-confinement.test.cjs had dropped out of the Windows tier. It must stay in, and no content signal can express why: it exercises isPathConfined (src/external-descriptor-trust.cts), which uses the AMBIENT path module -- path.resolve(root, target) and path.sep -- with no injection. Its win32 semantics (drive letters, UNC, separator) are only reachable by actually running on Windows, and it is a security-relevant write-confinement gate. A content classifier cannot see 'this module reads the ambient path module', so no regex belongs here. Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows tier only. A Map rather than a list so an entry without a reason is impossible by construction, and tests assert every entry names a file that exists on disk so a stale entry fails loudly instead of rotting. This is the centrally- enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's portability-vocab.cjs already models -- deliberately not a heuristic. Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is untouched and byte-identical: the win32 concern does not apply to a POSIX runner, and a test asserts the allowlist does not leak into that tier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): inject the path impl into isPathConfined and correct the ADR count Two review findings, both fixed rather than dispositioned. A security review found tests/external-descriptor-confinement.test.cjs had left real-OS execution. The allowlist pinned it back, but that only restored INCIDENTAL coverage: isPathConfined used the ambient path module, and its test carried POSIX-only literals, so a win32 confinement escape was unverified on every platform including Windows. isPathConfined now takes an optional third parameter carrying the path implementation, defaulting to the ambient module. Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change is purely additive and every existing two-argument caller is byte-identical. Tests now inject path.win32 and path.posix, covering a different drive letter, a cross-drive absolute, backslash and forward-slash traversal, UNC, and the startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators. Proved load-bearing: dropping the + p.sep from the prefix check fails exactly the two boundary cases and nothing else. Callers' suites 149/149. The spec review caught an off-by-one: the ADR narrated a 264-file tier while the committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264 -> 265 (28.5%). Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated classify() as zeroing windows_tests, a key this change removes -- kept as historical narration but labelled as such. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#131): make the unwritable-HOME test actually test something Found by sweeping for the root-bypass class after fixing commit-files-deletion. This one is the silent variant, and it was broken twice over. First, the condition: the test made a fake HOME unwritable with chmod 0o500. The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed writable and the hostile condition never existed. Replaced with a HOME whose PARENT is a regular file, so every write under it fails ENOTDIR at the VFS layer for every uid -- no permission check is involved at all. Second, and more fundamental: the probe was npm --version, which on npm 11.19.0 performs zero filesystem I/O against HOME. Proven rather than assumed -- neutralizing runNpm()'s isolation turned the sibling test red while this one stayed green, so its assertion could never detect the regression it guards, on any uid, with or without the condition fix. npm config get cache was tried next and proved vacuous the same way (it only string-resolves the path). The probe is now npm cache verify, which really does mkdir _cacache under HOME. Re-proved load-bearing after the change: with isolation neutralized the test now fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was restored and verified diff-clean; suite 13/13. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the net drop-out figure in ADR-4641 The Consequences section still said 292 files leave real-OS Windows execution. That was the count before the narrow shell-interpreter-spawn replacement restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary categories rather than only raw-child-process, and clarifies that the 14-file filename audit was against the 292 initially dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the rejected concentration ceiling and its measurement Applying Goodhart's own question to the new ceiling -- how would you make this metric look good without improving what it represents -- surfaces a real weakness: a ratio can be satisfied by inflating the denominator, so adding OS-agnostic tests loosens it without narrowing the tier. The obvious companion gate was a sole-signal concentration ceiling, since the original defect was one detector carrying half the tier. Measured and rejected: peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the original defect; any threshold below it fails on a legitimate category. The discriminator is whether a signal is platform-meaningful, which no threshold encodes. Weakness disclosed rather than covered by a gate that does not bind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4641): add the changeset fragment for the confinement-check change changeset-lint failed on PR #4643: the PR touches user-facing paths and carried no fragment. The earlier no-changeset call matched #4604's CI-only precedent and was correct then; it was not revisited once the PR grew a src/ change, which is my miss. The fragment describes the real user-visible improvement: the external-descriptor write-confinement check's Windows semantics are now verified deterministically rather than only when the suite happened to run on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the tier count in TESTING-SUITES.md Said the tier narrowed from 546 to 254. The final committed list is 265 of 931 eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored 9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in the ADR, in a live reference page rather than a dated record, so it states the current truth rather than carrying an amendment note. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the measured aggregate from real CI job lists Epic #4589's closeout asserted its reduction from a static count; #4641's acceptance criterion asks for a figure read off a real run. Recorded here: test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4. Also states the caveat that a PR's total CHECK count is not a clean before/after comparison, since many gates are path-scoped and this change touches a broader path set -- the like-for-like figure is the test.yml job count. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): compare job totals the same way on both sides The measured-aggregate table put #4640's COMPLETED run total (21) against this run's count at matrix-expansion time (15). Those are not the same measurement: the completed total includes the post-test Coverage gate and baseline-publisher jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux jobs 7 -> 4, was correct and is unchanged. Called out in the table rather than silently corrected -- comparing two differently-derived numbers is exactly the error class this ADR is about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record measured conformance wall-clock and date the stale counterfactual Adds the per-job durations from both runs. The honest read is that this is a correctness win more than a speed one: file count fell 52% but wall-clock only 9-29%, because what was removed were the cheap static tests and what remains is concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects a future narrowing to buy time proportional to file count. The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about -- pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its tier was UNCHANGED at 197 files, which fixes that as runner variance and is noted as a caution against reading a single duration as signal. Also dates the symlink-keyword counterfactual, which cited a 254-file tier from before the replacement category and allowlist took it to its final 265. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero next gained #4644 mid-flight, so every absolute count shifted. Re-measured on the tree this actually ships against (932 eligible): 548 -> 257 by detector removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed, 9 restored. macOS 198, unchanged by this PR. The percentages did not move across three rebases (58.8% -> 28.5%), which is the whole argument for expressing the ceilings as ratios rather than counts -- noted in the ADR since it is now evidence rather than assertion. Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own win32 test cases introduced the literal win32 into the pinned file, so it classifies in on content via win32-darwin-literal. The entry stays and the reason is written down, because the file's real-OS need is a property of the code under test (isPathConfined reads the ambient path module), not of the test's text -- the text that currently saves it is incidental and could be refactored away silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1e47560e34 |
feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the same 546-file, Windows-oriented conformance-tier list as windows-latest -- built from signals like windows-shell-token/windows-env-var that have nothing to do with macOS. Issue #4593 asked for macOS coverage sized to its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin- specific behavior) instead. Issue #4593 was filed before Phase 5 (#4603) existed and referenced updating test-full's macOS legs -- that job is gone. Corrected the issue's body before any code was touched: the "shrink from full replay" half of the original ask was already done by Phase 5; what remained was narrowing the still-Windows-oriented tier macOS was inheriting. Two design assumptions were measured and rejected before accepting a design (documented in docs/adr/4593-macos-conformance-tier-architecture.md): - Reusing the general tier's signals minus its 3 Windows-specific categories barely narrows anything (546 -> 424, 78% retained) -- most files match multiple signals and only need one to survive exclusion. - A standalone CRLF/autocrlf signal, despite the issue naming "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit 143/930 files. Root cause: CRLF is primarily a Windows checkout concern in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST- PORTABILITY), so the signal was really re-selecting Windows-relevant files already covered by the general tier, not narrowing macOS specifically. Built 5 new, genuinely macOS-specific signals instead: darwin-literal (darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch, case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim from the general tier (genuinely Unix-relevant, not Windows-motivated). Measured against the real tree: 196 of 930 eligible unit-suite files (21%), versus the general tier's 546 (59%) -- a real, evidence-backed narrowing. scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/ classifyMacosTree/renderMacosGeneratedFile and a --target windows (default, unchanged)/--target macos CLI flag, so the same generator produces two independent, gated outputs rather than needing a second script. New committed output: scripts/lib/macos-conformance-tier. generated.cjs. .github/workflows/test.yml's test-conformance job: only the macos-latest leg's file-list source changes; windows-latest is byte-for-byte untouched. New shipped-file ripples handled proactively (19 install-tree fixtures regenerated, bin/install.js registered). An isolated code-review pass found one real defect: the ADR's per- category count table had drifted by 1 (zsh-dispatch, case-sensitivity) because the new test file's own fixture strings joined the tree it classifies after the table was authored -- fixed, with the union total (196, what CI actually gates on) confirmed unaffected. An isolated security-review pass found no qualifying findings. The ADR also records an explicit requirement for any future widening proposal: check whether the motivating regression is already covered by Phase 1's no-rendered-text-length-assert lint rule (#4590) before re-proposing full macOS/Linux parity, since that is exactly what #4421's root cause was (a rendered-text-length assertion, not a real behavioral divergence). Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2cefa5a5ac |
enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes (#4587)
* enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes ADR-4139's final phase. workflow.compact_content already defaulted to false (Phase 1's buildNewProjectConfig hardcoded default), but nothing surfaced it: /gsd-new-project never asked, and /gsd-settings/config had no toggle path for an already-initialized project — config-set/config-get were the only route. new-project.md gains a fourth question in the existing Round 2 AskUserQuestion array (grouped with the other general-workflow-behavior toggles, not the per-agent capability questions above it) and threads compact_content into the config-new-project CLI JSON literal. settings.md mirrors the exact pattern every other non-capability workflow.* key already follows: read_current bullet, question block, update_config write, the safe-merge non-capability-keys list, save_as_defaults, and the confirm summary table — seven edits, zero new src/*.cts code, since Phase 1's merge logic is a generic passthrough. Its success_criteria question-count ("24 settings") is bumped to 25 to match the now-25-entry main AskUserQuestion batch. settings-advanced.md deliberately does NOT get a duplicate question: no other boolean toggle in this repo is asked in both settings.md and settings-advanced.md, and there's no reason to start with this one. docs/CONFIGURATION.md, docs/USER-GUIDE.md, and a new docs/features/4139-compact- content.md fragment (regenerated into docs/FEATURES.md) document the toggle. ADR-4139 itself: Status flips Proposed -> Accepted, the acceptance-criteria section becomes a guard ledger — a 13-row table covering all 12 of #4139's original checkboxes plus the shipped-content guard criterion, each with real evidence (the merged PR that satisfied it, fetched via `gh issue view --json closedByPullRequestsReferences` rather than asserted from phase numbers) — and both "Open questions for the implementation phases" are resolved rather than left dangling: discuss-phase was never converted to spine+detail shape (verified: no detail/ subdir exists) — a genuine gap, not a reasoned decline; the disjointness check is confirmed line-based by reading compact-content-split.cjs's normalizeNonTrivialLines directly. Orthogonal review (isolated Standards/Spec code-review + security-review sub-agents) found and this fixes two real defects: the changeset fragment's body didn't match CONTRIBUTING.md's single em-dash-sentence format (was multi-sentence prose naming implementation file paths); and settings.md's own success_criteria still said "24 settings" after the new question pushed the main batch to 25. Also fixed, found by the Spec pass while confirming commands/gsd/settings.md correctly needed no sync edit: that file and its skills/gsd-settings/SKILL.md twin both still described "Interactive 5-question prompt (model, research, plan_check, verifier, branching)", stale since long before this phase (the batch has had far more than 5 questions for a while) — replaced with a description that names the current set without hardcoding a count that will drift again. gsd-test (real run, sha 1da78fe2) caught a third real regression the local sweep missed: new-project.md is a registered spine+detail split for Phase 4's token-reduction benchmark (scripts/benchmark-compact-content.cjs), and the new question's +167 tokens drifted the committed baseline (tests/fixtures/compact-content-benchmark-baseline.json). The benchmark itself is designed never to fail CI on drift, but the test asserting the COMMITTED baseline is currently non-drifted correctly caught it. Regenerated via `node scripts/benchmark-compact-content.cjs --write`; re-verified --check now reports "up to date" and the test file passes 27/27. Closes #4408. Closes #4139. Emitted-Drift-Ack-Growth: new-project.md — new 4th Round-2 AskUserQuestion entry (Compact Content, #4139) plus the config-new-project CLI JSON field and explanatory sentence; a new opt-in toggle needs new prose. Emitted-Drift-Ack-Growth: settings.md — new workflow.compact_content read_current bullet, question block, update_config write, safe-merge key, save_as_defaults field, and confirm summary row (the same seven-edit pattern every other non-capability workflow.* toggle already follows), plus the 24->25 success_criteria count fix found in review. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4408): backfill changeset PR number pr:0 -> pr:4587 now that gh pr create has returned the real number. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9770258558 |
chore(#4590): add no-rendered-text-length-assert ESLint rule (#4595)
* test(#4590): add no-rendered-text-length-assert ESLint rule Enforces ADR-456's typed-surface mandate for one specific bug shape: a test assertion whose pass/fail depends on the length/substring content of a template literal that interpolates an OS-derived path (os.tmpdir(), os.homedir(), path.join/resolve/..., or a PATH_RETURNING_FNS resolver). Because macOS's default tmpdir prefix is longer than Linux's, such an assertion can pass on one runner and fail on another -- the defect class behind #4421's incident (git show 4e75b836e9), already fixed there by pinning to a typed field per ADR-456 Sec(c) before this rule existed to catch a recurrence. Two repo-wide sweeps against the real tests/ tree narrowed the rule to a sound scope: an initial design that traced call arguments (to approximate the historical incident's cross-file render-function shape) produced false positives on ordinary fs.readFileSync(path.join(...)) + assert.match patterns; a second design that matched any bare direct path-returning call produced 45 false positives on path suffix/prefix/non-emptiness checks. The shipped rule matches only a path-returning expression interpolated into a template literal, directly or via one identifier hop -- disclosed in the rule's own "Known boundaries" as not covering the literal cross-file incident shape, which would require tracing into a callee's body. Phase 1 of epic #4589 (CI test-matrix Linux-primary migration) -- Phase 2's safety argument depends on this class of OS-dependent test assertion being enforced going forward, not merely fixed once. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4590): address code-review findings on no-rendered-text-length-assert Reletter the "Known boundaries" doc-comment list (a)-(e), fixing a gap left by an earlier edit pass and every stale cross-reference to it. Collapse isDirectPathTaint/isTaintedInterpolation's duplicated TemplateLiteral-walk into one recursive relationship (isTaintedInterpolation now delegates a nested-template-literal case back to isDirectPathTaint instead of re-implementing the .some() traversal) -- behavior unchanged, confirmed by re-running the repo-wide sweep (still zero false positives). Found by the Standards-axis /code-review pass on this PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
18c899def5 |
enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
38e4ce5f62 |
fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
aad04f4e9a |
docs(#4400): ADR-4139 — the compact-content seam (#4410)
* docs(#4400): ADR-4139 — the compact-content seam Phase 0 of epic #4139. Locks the design before any code lands. #4139's stated mechanism cannot reach the stream it exists for: 58 of 72 shipped commands deliver their whole workflow file through an eager @-include, which the host expands before any project config is in context. An in-content gate is evaluated after those bytes are already paid. The ADR declines the obvious fix (convert the 58 execution_context blocks to runtime Reads) because that removes the host guarantee for every user, not only opted-in ones — a global install shares one skill tree, so an @-include cannot be conditional. It instead keeps every @-include exactly where it is and splits what sits behind them: the canonical path becomes a runnable spine, elaborations move to a sibling detail file read at runtime. A missed Read then degrades to "runs correctly with fewer tokens", never to "runs with no instructions". Also records: the rename to workflow.compact_content, the re-pitch onto ADR-1610's context-rot argument rather than the cost argument ADR-1610 discounts, partition-not-duplication (which dissolves the dual-maintenance cost the Feature Review called disqualifying), the guard-scope and NEW_FILE_CAP mapping for the new subtree, and an argued reconciliation of the acceptance criteria this design does not meet literally. Corrects ADR-3646 §Context: it cites #3647 as open; #3647 closed 2026-09-01 as a duplicate of #3606. ADR-3646's Decision is unaffected — it explicitly disclaimed any dependence on #3647's state. The residual prose-dispatch variance named in #3647's own closure thread is unresolved, and this ADR routes around it rather than assuming it away. Refs #4139 Closes #4400 * docs(#4400): fold the two orthogonal review findings into ADR-4139 Code review (isolated context) and security review (isolated context) both returned findings. Fixed here rather than carried. Critical, from code review: the ADR repeated earlier research's claim that discuss-phase, manager and pause-work all reach a workflow by runtime Read. manager and pause-work carry plain eager @-includes and are inside the 58, not outside. discuss-phase is the only precedent, and it is one file. The Open Questions section is corrected with it. The NEW_FILE_CAP mapping was wrong in a way that changes the layout. It lives at tests/helpers/emitted-diff.cjs:96, not in workflow-size-budget, and its own doc comment records that it is a hard cap, not ack-able, and NOT tier-exemptible -- the pre-#2724 test-file version was. So a single detail.md holding plan-phase.md's elaborations is blocked outright with no exemption path. Detail content is now one or more parts under workflows/<name>/detail/, each below the cap, named by the spine in the dispatch-table shape discuss-phase.md already uses. commit-files-pathspec is in scope and earlier research called it irrelevant. Per CONTRIBUTING.md:1164-1170 it sweeps every .md under gsd-core/workflows/ for unscoped commit-seam invocations. Added to the guard table. Two byte figures were inherited rather than measured, against this ADR's own evidence note. Templates and agents re-measured; the table now carries the method and the exact numbers. From security review: "a spine that has shed a protected-content marker fails" never defined what a marker was, leaving the strongest check in the set resting on a prose-category judgment. Protection is now a literal greppable sentinel in the existing gsd: comment namespace, and the guard rule has no discretion in it. Also added: an explicit statement that the detail path is never user- or project-supplied and cannot be shadowed by a project-local file, and an exact-version pin commitment for gpt-tokenizer. Code review also found a real hole in the central fail-safe argument: spine sufficiency is verified once at split time and never again, so load-bearing procedural text carrying no sentinel could later drift into a detail part with every check green. Sufficiency is not machine-decidable, so a fifth ongoing check is added -- a spine that loses lines which reappear in its parts fails unless the PR declares the boundary move. The ADR now states plainly that this is authoring discipline with a forced checkpoint, not a structural invariant, and that the partition relocates the Feature Review's cost rather than fully eliminating it. Refs #4139 Refs #4400 --------- Co-authored-by: sim <sim@local> |
||
|
|
1db726ebbf |
feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md, and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only supersession rule. These fail until the canon and its two references are added. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(#3806): canonize the Review Dispositions Ledger contract Promote the existing planner-reviews.md Step 4 return-payload tables (Review Feedback Addressed/Deferred) into a canonical `## Review Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md and referenced (not restated) from plan-phase.md's <review_incorporation_contract> and gsd-plan-checker.md's Review Incorporation dimension. Adds round-scoping (`### Round {N} — {REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md reference survives the file being rewritten each round, and an append-only supersession rule. Scoped to part 1 only per the maintainer's approved-feature verdict — the deterministic lint/check verb (part 2) is explicitly deferred to a follow-up. Also: ADR-3806 recording the decision, a docs/features/ fragment (FEATURES.md is generated), and a changeset fragment. Closes #3806 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fenced-example count bug and lint findings from review - tests/plan-review-convergence.test.cjs: the "heading exactly once" test counted the canonical heading text globally, so it also matched the illustrative fenced-code example in planner-reviews.md that shows the same heading as sample content, always failing 2 !== 1. Rewritten as a bounded line scanner that skips fenced blocks (found by an isolated adversarial review pass). Also bounded an unbounded regex quantifier over readFileSync content flagged by local/no-unbounded-quantifier. - docs/features/review-dispositions-ledger.md: match house fragment style (bold-lead paragraphs, not #### headings) per the Standards-axis review; regenerated docs/FEATURES.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fit reference-cite fix within size hard caps; ack growth Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a single short clause pointing at gsd-core/references/planner-reviews.md (also fixes the bare `references/planner-reviews.md` cite the #3576 shipped-reference-cites gate rejects), bringing both files back under their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline. Both files still grow slightly versus origin/next, acknowledged below per ADR-2719's emitted-drift-ack contract. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block The previous commit's two Emitted-Drift-Ack-Growth trailers were separated from the Co-Authored-By trailer by a blank line, so git's own trailer parser (which tests/helpers/emitted-runtime.cjs reads via `%(trailers:key=...)`) only recognized the last contiguous block (Co-Authored-By) and treated the Ack-Growth lines as ordinary body text — invisible to the emitted-attribution gate, not malformed data. Restating them here as one contiguous trailer block, git log over the PR range aggregates trailers from every commit, so this is additive. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): isolate the ack-trailer paragraph as its own trailer block Git's trailer parser requires the trailer paragraph to be the message's final paragraph, preceded by a blank line, and to contain nothing but trailer-shaped lines. The prior commit's blank line before the trailer lines was missing, which folded the leading Emitted-Drift-Ack-Growth lines into an ordinary prose paragraph. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3806): backfill PR #4345 into changeset and ADR Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
17e163f15c |
docs(#4333): document the ADR Amends/Amended-by convention (#4334)
* docs(#4333): document the ADR Amends/Amended-by convention Two patterns for amending an accepted ADR are established practice — an in-place `## Amendment (YYYY-MM-DD)` section, and a separate ADR that declares `Amends` with a reciprocal `Amended by` back-link — but only the first was ever written down. #4030 shows the cost: a contributor concluded no ADR owned a contract that ADR-857 already covers, because nothing said the second pattern (used by ADR-1244 and ADR-2782 to extend ADR-857 itself) existed. Document both patterns in docs/contributor-standards.md, note the Amends/Amended-by reciprocity rule in docs/adr/README.md alongside the existing Supersedes/Subsumes rule (and that it isn't yet gated by scripts/gen-adr-index.cjs the way those are), and point CONTRIBUTING.md's new-ADR process at the amendment path for revisiting an existing one. Closes #4333 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4333): fix imprecise Amends/Amended-by precedent citations Orthogonal review caught two inaccuracies: PR #1643 doesn't match the in-place dated-section pattern (it rewrites the original Decision text rather than appending an untouched dated section), and ADR-1244's relationship to ADR-857 is prose ("extended by"), not the structured Amends/Amended-by header field. ADR-2782 is the verified precedent for the structured field pair — its one Amends field names four targets (857, 894, 1016, 1244), all four carrying the reciprocal back-link. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
86b745b48b |
fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1fe85cd43e |
chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk Repo-wide sweep (ahead of adding lint rules for these exact bug classes) found both incident patterns still live and unfixed on `next`: - scripts/run-tests.cjs's sweepProtectSet walk stopped on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. win32 dirname('D:\') is a fixed point (length 3, never satisfies `> 1`... wait, it does satisfy length>1), so a selected file living outside runTempRoot (the common case) spins the walk forever on Windows. Extracted a pure, exported computeSweepProtectSet helper that terminates on dirname(cur) === cur instead, with in-process RuleTester-style coverage for both win32 and posix paths. - tests/run-tests-temp-root.test.cjs's own #4020 regression test set only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never reads TMPDIR on Windows (only TEMP, then TMP), so the redirect silently no-oped there — masked because Windows CI died in the dirname-walk hang above before ever reaching this test. - tests/config-schema.property.test.cjs's fallow config-set test had the same TMPDIR-only pattern, direct process.env assignment this time, restored in its own finally block. Origin: #4220 and its shared root cause #4020. * feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for either shape. - local/require-full-tmpdir-triad: flags a TMPDIR environment override (direct process.env.TMPDIR assignment, or a TMPDIR property in a spawn-like call's env: object literal) not accompanied by TEMP and TMP in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows. Registered on tests/**/*.cjs, matching the require-userprofile-with-home precedent. - local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning from dirname() with no fixed-point termination guard (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a no-op at the platform root, but the value differs by platform (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped length/equality bound never fires on Windows. Registered on BOTH tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in scripts/run-tests.cjs, not tests/. Both rules join the zero-escape-hatch discipline already established for this catalog (no bespoke comment marker; PROTECTED_RULES in tests/portability-rule-disable-ban.test.cjs independently bans eslint-disable of either). ADR-1703 and its two companion contributing docs get an amendment documenting the mechanism, code examples, and the repo-wide sweep (three live instances found and fixed in the prior commit; no others found). CI test-scope selection updated so an edit to either rule or to scripts/run-tests.cjs re-runs the right suites. * fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too checkWhile bailed out early unless node.test was a LogicalExpression, so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } -- was silently skipped and never reported. That is the EXACT minimal shape of the original #4020/#4220 bug, and it is literally the shape used by this rule's own shipped RuleTester fixtures (the "equality-only bound" invalid cases), which were failing (0 errors reported, 1 expected) until this fix -- confirmed by running RuleTester directly against both fixtures, not just via a passing test-runner exit code. The conjunct-collection helper already handled a non-LogicalExpression test correctly (it pushes a single node as the sole conjunct); only the early-return gate needed to stop requiring a compound && / || test. Verified: RuleTester run directly against both previously-broken fixtures plus two new sanity cases (a guarded single-condition loop stays valid; an unrelated single-condition loop stays silent), and a fresh `npx eslint .` across the whole repo remains clean (no other single-condition dirname-walk shape exists in the tree). * fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call isSpawnLikeCallee only recognized a MemberExpression callee (child_process.spawnSync(...)) or a bare identifier in ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare -- const { spawnSync } = require('child_process'); spawnSync(...) -- has an Identifier callee named "spawnSync", which matched neither branch, so the whole env-literal check was skipped. gsd-test caught this: both "invalid: child_process.spawnSync with TMPDIR-only env" cases in tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors reported, 1 expected). Widened the bare-identifier branch to also match any of the known ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same lightweight convention this repo's other eslint-rules/*.cjs use (e.g. no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow tracing. Verified: RuleTester run directly against all 11 cases in tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that were failing), all pass; a fresh npx eslint . and npm run lint:ci across the whole repo remain clean. * fix(#4244): correct a stale escape-hatch reference in a test comment The comment on the "length comparison against another expression's length" case referenced a "// allow-dirname-walk marker" that doesn't exist -- the rule has zero comment-based escape hatches by design (ADR-1703), and an earlier draft's marker mechanism was removed before this branch's first commit. Spec-axis review caught the stale reference. No behavior change; comment-only. * chore(#4244): backfill changeset PR number (pr:0 -> pr:4246) --------- Co-authored-by: sim <sim@local> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
7ca5479f97 |
docs(#3672): amend ADR-1239 for quick-batch concurrency and single-writer contract (#4146)
Locks the architecture gate the maintainer required before any /gsd:quick-batch implementation PR (epic #3344, condition #2): canonical terminology, the foreground-coordinator single-writer invariant over BATCH.json/STATE.md/roadmap metadata, the dispatch.maxConcurrency sub-field (fail-closed to 1), the GSD_DISPATCH_MAX_CONCURRENCY live-capacity transport and its precedence over the descriptor value, the gsd-tools query dispatch-capacity contract, the dispatch.isolation interaction, and backpressure/recovery/deterministic-merge rules. Docs-only; closes its own Phase 0 sub-issue, not the epic. Co-authored-by: sim <sim@local> |
||
|
|
1a358ce0fd |
feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
80de48c319 |
enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs
Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.
local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.
Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.
The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.
Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.
Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.
Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.
Verification runs on the remote runner.
Refs #3914
* docs(#3914): the epic added three guards, it did not remove one
The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.
An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.
Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.
Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.
Refs #3914
* fix(#3914): make the successor actually supersede before retiring the predecessor
An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:
const exit = 'exit';
process[exit](1);
n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.
This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.
The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.
The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.
The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the two exit rules are complementary — keep both
Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.
I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:
function f(exit) { process[exit](1); } 0 vs 1
let exit='exit'; exit='exit'; process[exit]() 0 vs 1
const { exit } = ...; process[exit](1) 0 vs 1
plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.
The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.
So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.
What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.
The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.
Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the exemption term was a no-op — the net is +4
Fourth correction to this ledger, and a fourth error of the same kind.
Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in
|
||
|
|
ab69b9ce56 |
enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from |
||
|
|
12f9d1d9a0 |
enhance(#3913): docs, and the guards come down (#3994)
ADR-3889 terminal phase. Generated docs/reference/exit-codes.md from the exit-code declaration with a --check drift arm; deleted the inert soft-error-exit-zero oracle; promoted untyped-success from SMELL to VIOLATION so it can fail a build; pruned all 5 smell-baseline entries. Fixed inline: two mis-scoped oracles (routing-validity, value-hygiene), a second source behind the band table, unescaped declaration strings reaching Markdown, and a pre-existing Windows 8.3 short-name path-comparison defect. Guard ledger corrected from a claimed net -4 to a measured net -1. Closes #3913 |
||
|
|
b351c83e03 |
docs(adr): ADR-3646 — per-task external-tracker content-resolution seam (#3991)
Resolves the four conditions on the approved-feature verdict for #3646: new execute:task granularity tier below wave, hard-halt enforced via a code-side resolver seam (Lens B) rather than prose dispatch (since #3647's dispatch-reliability defect is still open), registration/validation requirements for loop-hook-dispatch.md + capability-validator.cjs, and autonomous-mode behavior via a new non-gate kind. Closes #3969 Co-authored-by: sim <sim@local> |
||
|
|
3a6c0412a9 |
enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) Extends ADR-1703's portability rule catalog with a second production-runtime rule: it flags an exact-case read of a Windows case-varying environment variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any receiver that is not process.env itself, matched via an env-shaped-receiver check to avoid colliding with ordinary `.path`-named properties elsewhere in the tree. Exports the seam's private `_envGet` as `envGet` so the rule's remediation message names a real helper, and fixes the one pre-existing violation the tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read). Closes #3624 * fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case Review findings from the code-review + isolated-adversarial passes: - extractStaticName only matched non-computed Identifier keys, so a destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8 acceptance case) silently evaded the rule. Widened to accept a Literal key regardless of computed, which is safe for MemberExpression too (its non-computed property is always an Identifier by grammar). - Added the missing RuleTester valid case for "a case-insensitive accessor call" (envGet(env, 'PATH')) from the issue's Done-when checklist. * docs: backfill changeset PR number for #3624 (PR #3976) --------- Co-authored-by: sim <sim@local> |
||
|
|
bb4f3073c0 |
docs(#3984): give every §8 rule its delivered status and its measured executor (#3986)
Surfaced by an /adr-phase-coverage audit of epic #3473 after its last sub-issue merged. §8 is the epic's declared source of truth - "where this section and the code disagree, the code is the defect" - and line 136 defines a per-rule status vocabulary. Not one of the nine rules was ever advanced past its pre-implementation status. The word Enforced appeared exactly ONCE in the document: in the sentence that defines it. Three rules read "Required - phase unassigned" for work that Phases 5, 7 and 8 demonstrably delivered, so a reader auditing coverage today would conclude three of the nine contract rules had no owner. That is the orphan shape this epic exists to eliminate, sitting in the epic's own contract, and it is the same defect class as B6's guard ledger: a status column is a claim, and it was false. Flipping all nine to Enforced would have repeated the mistake inverted. Decision 2 says a declared policy with no executor is a loud failure, so Enforced asserts an executor EXISTS - and measurement found the nine are not equally enforced. The vocabulary is now three values: Enforced a standing guard wired into lint:ci or CI fails the build; a new wrong call site cannot land Enforced (structural) the wrong call site is unrepresentable Shipped - test-covered delivered and regression-tested, but NO standing guard; a regression in covered code is caught, a new wrong call site elsewhere is not Measured per rule, with the executor cited: 8.1 Enforced lint-vendored-deps + no-external-require-in-bin 8.2 Enforced (partial) phase-enumeration-drift covers the sentinel arm; the #3357 resolver arm is test-covered only 8.3 Shipped NO guard for slug or marker re-derivation 8.4 Shipped structural for argv; no guard for the general rule 8.5 Shipped the delivering change added zero guard scripts 8.6 Enforced (structural) the transaction type; state-write-path-drift 8.7 Shipped structural via reconcileReportedFields; no guard 8.8 Enforced gen-state-md-docs --check in lint:generated-sync 8.9 Enforced (prospective) lint-fix-has-regression-tests fires on NEW fix commits; it does not assert the 19 children are covered. 17 of 19 have a test citing their number; #3364 and #3812 have none - a text match, not proof the behavior is uncovered, but not evidence it is covered either The third status value is not a euphemism for done. It names precisely where this epic's own thesis - make the wrong call site unrepresentable - is NOT yet achieved. §8.3 and §8.5 are the thinnest: nothing today stops a twelfth inline slug copy or a second silent swallow. Also corrects a claim this ADR made three days ago and #3977 has since overtaken. The ledger amendment concluded #3426/#3239 "need new detectors". The measurement behind that was right - the glob widening structurally could not reach them - but the prescription was not the best available: #3977 rerouted the test's parsers onto the ADR-2143 seam, so there is no ad-hoc parse left for a detector to find. Removing the violation beat teaching the rule to see it. Both issues are CLOSED; the roster row and the amendment now say so, and say why the earlier conclusion should not be built on. Closes #3984 Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d24e22b156 |
enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1 ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what remained was the declaration — and the pin that makes it invisible today. The census corrected two documented figures before any code changed. ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier note of mine claiming 23 was wrong and is corrected). And output({error}) is **64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2. That matters because this phase's criterion demands the pin be asserted over the enumerated population rather than sampled; asserting over a stale 60 would leave four sites unpinned while claiming full coverage, which is the shape of failure this epic exists to remove. The issue does not state the fact that shapes the design: output() never touches the exit code. Confirmed by reading it — it writes fd 1 and returns. So a declared outcome for those 64 sites had nowhere to be READ. The mapping was never the work; wiring somewhere for the declaration to land was. The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol cells, each because the module is emitted to three locations and a module-level `let` would let instances disagree, and runMain already maps a code returned by main(). A third cell inherits that solution. output() records DEGRADED for any {error} payload — key-order agnostic, which is exactly why the "42 sites" figure undercounts — and runMain projects the cell only when main() returns nothing, so an explicit return still wins. error() maps its reason through a table over the closed 25-member enum, leaving all 278 call sites untouched; 226 of them pass no reason at all. The version gate lives in error(), NOT in projectOutcome: registered names are version-invariant there, so mapping a reason straight through would make USAGE project to 64 under v1 and break the pin on its first line. projectOutcome is left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included. Proven rather than asserted. v1 is byte-identical across three real CLI paths — config-get plain, config-get --json-errors, and an output({error}) path — matching exit code and exact bytes against the pre-change build. Under GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND -> NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason, because without it a mapping where everything projects to 1 under both versions would satisfy every other assertion and the declaration would be theatre. A1 iterates all 25 enum members and A3 asserts over the measured 64-site population, so a 26th reason or a 65th site fails until it is given a mapping — the drift guard this phase needs, given ADR-2980's own count had drifted +4 unnoticed. Verification runs on the remote runner. Refs #3912 * fix(#3912): the outcome cell must never lower an exit code The remote run caught a fail-open that this phase introduced, in the phase whose entire purpose is removing fail-opens. `state validate --strict` on a missing STATE.md exited **0** where it must exit 1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set non-zero by the command was clobbered down to success. Confirmed live against a fixture, before and after. This refutes a review conclusion recorded earlier in this phase, that the cell was "fail-closed and can never mask a failure as success". It could, and did. Recording that plainly so the assumption is not repeated: the cell's danger was never only that it might add a failure — it was that projecting it unconditionally overwrites whatever decision came before. Projection is now guarded: it may set a code only when none is set, and an already-non-zero exit code always wins. The full precedence — explicit `main()` return, then an existing non-zero exitCode, then the declared outcome — is written at the projection site. A regression test drives a void return with a pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix build. The second failure was my test encoding the wrong contract, not a code defect. It asserted `output({found:false, error: undefined})` records DEGRADED because the KEY is present. `JSON.stringify` drops undefined, so the payload the user receives is `{"found":false}` — carrying no error at all, and calling that degraded would hand back exit 80 under v2 for output that reads as clean. The discriminator is a serializable error VALUE, not key presence. The test now pins `{error: undefined}` as explicitly NOT degraded, and the design doc's wording is tightened to match. Verification runs on the remote runner. Refs #3912 * docs(#3912): the versioned exit contract, and a flag defect the docs found Diataxis pass for Phase 8, plus a real fix that only surfaced because writing the how-to meant running its own examples. The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned projection this phase provides, so it gets an amendment naming #3912 / ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects DEGRADED to 80. The amendment also records the count drift rather than restating a stale figure — the ADR ratified 60 output({error}) sites in 9 modules; the AST-measured population is 64 across the same 9 (frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the enumerated 64. json-errors.md gains the outcome-declaration reference, including the precedence order a review pass got wrong and the suite refuted: an explicit main() return, then an already-set non-zero process.exitCode, then the declared outcome. Projection may only ever set a code, never lower one. A how-to is owed here and is written, not skipped. Under v1 nothing changes, so the audience is an operator opting into v2 and needing to know what the codes mean for a CI gate — a migration, which is how-to shaped. It covers turning v2 on, the code table, why 80 is "ran and reported a condition" rather than a crash, and how to split a gate that treats any non-zero as fatal. No tutorial: there is no new entry point to learn, and under the default contract a reader would be walked through observing nothing. The defect. Running the how-to's own Step 1 example returned $ gsd-tools --exit-contract=v2 state validate --strict Error: Unknown command: --exit-contract=v2 (exit 64) while the same flag trailing the subcommand worked and exited 80. The flag half-worked, by argv position. resolveContractVersion scans argv non-destructively, so the token survived into the dispatcher, which treats argv[2] as the command name. --json-errors had already solved precisely this at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv splice must happen here too, otherwise the dispatcher below sees --json-errors as an unknown command." The later flag never got the same treatment. Fixed rather than documented around: the version is resolved first — which memoizes the cell and makes an invalid value throw early — and then every --exit-contract= token is spliced out of the dispatcher's argv copy. --exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The regression test pins leading position, trailing position, agreement between the two, and a loud failure on v3 rather than a silent fall back to v1. Neither review engine would have caught this: the defect is invisible in the diff, because the diff does not touch argv handling. It surfaced only from running the documentation's own example. Writing a how-to is an execution pass. Verification runs on the remote runner. Refs #3912 * fix(#3912): the flag splice has to run before the run-with-timeout return An isolated review of the previous commit found that the fix did not deliver what it claimed, and that two of its own tests were weak. All three findings reproduced by execution before any change was made. The fix was placed below a return. main() intercepts `run-with-timeout` at gsd-tools.cjs:4436 and returns from there — above both the --json-errors block and the --exit-contract splice added in the previous commit. So the flag still died in leading position for that one command: $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..." Error: Unknown command: run-with-timeout (exit 64, child never ran) The previous commit message and the test's describe-block both claimed position-independence unconditionally. That was an overclaim, not a gap left open, and it is the part worth naming: the fix was verified by hand on the commands I happened to think of, and `run-with-timeout` returns before the code I was verifying. Both global-flag blocks now run above the interception, with a comment naming it so a later edit cannot slide them back down. Moving --json-errors up fixes the identical pre-existing bug for that flag, verified failing beforehand (exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect, same block, and a known-broken twin next to a fixed one is not a resting state. Two tests were not pulling their weight. The invalid-value test was vacuous — it passed against the pre-fix build, because `--exit-contract=v3` already exited 1 there and already printed the resolve error lazily through error() -> getContractVersion. Both its assertions held before the fix, so it pinned nothing. The real discriminator is that the pre-fix build emits BOTH "Unknown command: --exit-contract=v3" and the resolve error, while the fixed build emits only the latter; the test now asserts that absence. The leading-position and leading==trailing tests asserted proxies — "not 64", "no Unknown command", "the two agree" — none of which pin a value, and all of which would survive both positions being identically broken. With a .planning directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0 under v1 in both positions. Those numbers are pinned now. The multi-token case the descending splice loop exists for is covered too, and run-with-timeout has regression tests for both flags. The lesson is narrower than "test more". Hand-verifying the production behavior does not verify that the test would have caught its absence. The pre-fix binary has to be run against the test's own assertions. Investigated and deliberately not changed: splicing before --cwd parsing degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd: <path>", but that is pre-existing — verified on the pre-fix build via --json-errors, which already did it. This change joins the pattern rather than creating it, and both forms exit 64 on malformed input either way. Verification runs on the remote runner. Refs #3912 * chore(#3912): backfill changeset pr numbers to 3983 * test(#3912): pin the reason-table invariant as set equality, not a count A graph-backed review flagged the unchecked lookup in expectedErrorCode3912. Investigated by execution: the drift guard DOES hold — for an unmapped reason under v2 the production error() yields 1 while the table yields undefined, so the assertion fails. Not a correctness defect, and deliberately NOT made tolerant, since a tolerant lookup would destroy the guard. Two real problems remained. The guard asserted the wrong invariant: it counted the TABLE's keys at 25 rather than checking they match the ENUM's values, so a renamed member keeps the count at 25 and slips past, and a 26th member leaves the table at 25 and slips past too. Both were then caught only indirectly, by an undefined mismatch producing 'must exit undefined'. It is now a sorted set equality, so the failure names the specific missing or extra reason. And the comment above it described a '?? FAIL' fallback that does not exist anywhere in the function. It now states what the code actually does, verified by running it rather than by reading it. Refs #3912 --------- Co-authored-by: sim <sim@local> |
||
|
|
355c943b08 |
enhance(#3626): make CONTEXT.md seam claims checkable via a SEAM.*.enforced-by gate (#3975)
* feat(#3626): make CONTEXT.md seam claims checkable via SEAM.*.enforced-by gate Adds SEAM.<id>.owns / SEAM.<id>.enforced-by=lint-rule:<name>|test:<path> predicates to CONTEXT.md, generalizing the existing WORKTREE.SEAM.* shape, plus scripts/lint-seam-enforcement.cjs (wired into lint:ci) which fails when a declared single-owner seam names no existing, registered enforcement mechanism. Backs all six current module-level single-seam/ single-canonical-owner claims found in CONTEXT.md, including the Shell Command Projection Module's Windows-binary-resolution claim via #3619's local/no-private-binary-resolution rule. Scope is resolves-only per maintainer decision: the gate proves an enforcement pointer exists and is registered, not that its surface covers every file the seam claims. See docs/adr/3626-context-md-seam-claim-gate.md. Closes #3626 * fix(#3626): back the Package Identity Module's seam claim too Isolated adversarial review caught a miss in the "no grandfather list" sweep: the Package Identity Module also declares itself "Single seam owning GSD's published-package coordinates" and already names its real enforcement (scripts/lint-package-identity-drift.cjs). Backs it with SEAM.package-identity.owns/enforced-by=test:tests/package-identity.test.cjs, bringing the total to 7 backed seams. Also makes explicit, in the design doc and ADR, that function-level "single owner" sentences inside already-covered modules (STATE.md Document Module, etc.) are deliberately out of scope — a seam claim is about a module's boundary, not every function inside it. * docs(#3626): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
15af0f5536 |
enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fa41bfec5c |
enhance(#3942): the emitted-drift ack is PR-lifetime data — move it to a commit trailer (#3954)
* test(#3942): failing-first suite for the emitted-drift ack commit trailer Binds 37 input classes from the phase test matrix to the behavior ADR-3942 specifies, before any of it exists. Stubs return benign empty values rather than throwing, deliberately: several rows assert that something DOES throw (cap overflow, uncomputable commit range), and a throwing stub would turn those green for the wrong reason and destroy the red. The two rows that carry the design's load: - merge-base semantics. The range is $(git merge-base base HEAD)..HEAD, not base..HEAD, because changedPaths comes from `git diff base...HEAD` (three dot). Two-dot would let the ack set and the change set disagree about which commits are this PR's. The fixture forks a topic branch, puts a trailer on each side, and asserts only the topic-side trailer is in range. - fail-closed on an uncomputable range. With fragments a depth-1 checkout passes VACUOUSLY, every fragment reading as brand-new. With trailers the range cannot be computed at all, and returning an empty set would silently disarm the gate, so it must throw. The fixture builds a genuine shallow clone rather than simulating one. Also covers the self-inflicted case: this change's own documentation quotes the trailer syntax, so an example landing at the end of a commit message would arm a live acknowledgment keyed on the literal placeholder text. Keys carrying angle brackets or whitespace are rejected. Authored per the phase artifacts 40-design.md and 50-test-matrix.md. Not yet run on the remote runner — this commit exists to be tested. Refs #3942 * chore(#3942): move the emitted-drift ack to a commit trailer Implements ADR-3942, superseding ADR-2719 section 3 and its #2789 amendment. Sections 1, 2 and 4-7 are retained: the conservation law is unchanged, only the storage of its escape hatch moved off the working tree. An acknowledgment explains one PR's ripple, and the moment that PR merges the ripple is in the base, so it can never clear anything again. It was stored in permanent shared state anyway, and every consequence of that mismatch had to be built and then maintained. The chain is #2789 -> #2914 -> #3078 -> #3842 -> #3823 -> #3875, each fix generating the next defect, ending in a scheduled sweeper whose own first PR could not merge itself. Added parseAckTrailers + renderAckTrailer (pure) and readAckTrailers (IO shell), reading Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: trailers over the merge-base range. tests/emitted-ack-trailer.test.cjs, 37 cases, written failing-first and confirmed red before any of this existed. Changed diffEmitted takes two structurally distinct key-space maps instead of one shared paths map. That closes a latent defect: the spaces were separated by convention only, so a growth key satisfied a hash lookup by naming coincidence. staleAcks now reports which space a key was declared in. REMEDIATION teaches the trailer, per space, with its example rendered through renderAckTrailer so the taught grammar cannot drift from what the parser accepts. Removed the sweep workflow, the guard-no-ack-on-next job, the standalone linter and its lint:ci entry, the fragment directory and its three spent fragments, the legacy single-file union, and the baseAck/spentAcks mechanism -- spentness is now structural, not computed. Two range properties carry the design and are pinned by tests rather than asserted: the range is merge-base scoped, matching git diff base...HEAD, so an already-merged trailer is out of range by construction; and an uncomputable range throws instead of reading as zero acknowledgments, which is the inverse of the fragment guard's vacuous pass. Three deliberate observable changes, each disclosed in the changeset: the unread runtime field is gone, the legacy file is no longer read, and cross-space excusal no longer works. Ten open PRs carry fragments and will meet a modify/delete conflict. Measured before landing and accepted deliberately; the one-line migration is in the PR body. Verified: lint:ci exit 0. Remote runner to follow on this exact sha. Refs #3942 * fix(#3942): silent trailer collapse, lost coverage, and an unbounded cap Six findings from the orthogonal review round, all fixed in place. BLOCKER -- two trailers of the same name on one commit collapsed silently. readAckTrailers built `separator=1d` where git needs `separator=%x1d`: the `separator=` value inside a %(trailers:...) placeholder is itself a pretty-format string, so the bare hex was emitted as two literal characters and the split on \x1d never matched. Two same-name trailers therefore joined into one value with errors empty -- the first reason absorbing the second entry's key. Silent truncation, the exact class MAX_ACK_TRAILERS throws to prevent. Confirmed with od -c against real git output before and after. The failing-first matrix did not catch it because its "both spaces coexist" row uses Hash plus Growth -- different trailer NAMES -- so the value separator was never exercised. Two regression tests now cover same-name trailers directly. Coverage recovered: normalizeAckReason and INVISIBLE stayed on the live path via parseAckTrailers but lost every test when the old suite was pruned. Back under test against the current surface -- all six invisible codepoints individually, whitespace collapse, trim, CRLF, and two seeded fast-check properties. Dropping any single codepoint now fails. MAX_ACK_TRAILERS counted raw trailers before de-duplication, so one trailer carried forward across rebased commits counted once per commit and could throw on a legitimate branch. Now counts distinct entries; 100 identical repeats dedupe to one. diffEmitted validated baseline, current and changedPaths but not the new ackHash/ackGrowth, so a bad shape raised an unhandled TypeError instead of an error verdict -- the same defect shape this file documents for #2778. Docs: CONTRIBUTING and TESTING-SUITES were rewritten only in their first sections; the later passages still taught fragments, git rm and the deleted guard, contradicting the new text directly above them. Finished. Also extends lint-removed-but-needed to exempt docs/adr and docs/research. That gate fails on any docs mention of a file deleted in the same diff, which makes it impossible to document a deletion in the PR performing it -- an ADR's whole job is naming what it retired. Exemption is narrow and comes with a test proving the gate still fires for a live consumer elsewhere under docs/. A guard that cannot fail is worse than no guard. Maintainer-approved. CONTEXT.md names the retired machinery by role rather than by filename: its generated projection lands in docs/, which that gate does scan. Adds docs/how-to/acknowledge-emitted-drift.md. The required docs set is Reference and Explanation, so the task quadrant can be empty with every gate green -- and this change has a real multi-step journey, including the fragment migration ten open PRs now need. lint:ci exit 0. Refs #3942 * docs(#3942): correct the duplicate-trailer rule in CONTRIBUTING Both axes of the code review independently flagged the same passage, without seeing each other's output. It claimed two declarations of the same key are always "a hard, loudly-reported error, not a silent last-wins". That is only half true, and the missing half is the one contributors hit: identical declarations -- same key, same reason -- dedupe silently, because a trailer legitimately survives a rebase and reappears on every rebased commit. Failing there would red a branch for doing nothing wrong, which is exactly why the dedup exists. Only a same-key/different-reason pair errors, and that one is a genuine ambiguity about which explanation holds. As written, the paragraph told a contributor that a rebase-carried trailer breaks the gate -- the opposite of the behavior. CONTEXT.md's parallel entry already stated it correctly; this brings CONTRIBUTING into line. Doc-only, root-level markdown. Refs #3942 * chore(#3942): backfill changeset PR number to 3954 --------- Co-authored-by: sim <sim@local> |
||
|
|
9410f7e6e6 |
enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4 ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the other three RED before any fix. Rung 2 — the install marker has four readers and resolveRuntime is not one. resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule. Fixtures and seam names mined from PR #3382 rather than re-derived; it implemented this rung and was closed "not on the merits". Rung 3 — the sandbox map, and the fallback that was the real defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements — the map carries nothing the contract does not - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit So the map is redundant and the silent fallback is the defect. The maintainer chose to derive but hold those 16 at read-only pending the question of whether Codex enforces sandbox_mode or merely advises; HALT.md records it. T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not in aggregate — an aggregate passes while one role silently widens, which is the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a stale hold, so the hold list cannot rot into the subset map being deleted. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId` returns only documentation commits. That was the wrong instrument. Direct inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and the tests match that code rather than a guess at its semantics — including first-write-wins on a duplicate short form. T43 asserts at the consumer's output: the emitted `waves` map from the real CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is dropped. A unit assertion on resolveDependencyId would have passed throughout this defect's life. Observed RED, this tree: rung 2 11/11 fail — no marker rung, no seams rung 3 T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML whose sandbox_mode disagrees — it checks presence only) rung 4 T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1 Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's unresolvable-token warning and wave-verdict suppression, and #3785's display-mapping passthrough. If the third tier over-reaches, those go red — that is their job. Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed) cannot be isolated behaviorally, because planMap always masks it. It is a non-crash boundary pin, weaker than the other rows, and is recorded as such rather than presented as equivalent. Design: .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md Decision: .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the other three. Rung 2 — the install marker had four readers, and resolveRuntime was not one. resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no marker, while bin/install.js writes one (#2297) and four hand-rolled readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with seams), hooks/gsd-agent-isolation-guard.js, and twice in hooks/gsd-cursor-subagent-start.js. model-resolver's was already the house idiom, so it was promoted rather than replaced: src/runtime-slash.cts now owns it, and model-resolver plus both hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam lint-hooks-runtime-build-seam enforces. No import cycle existed - checked both directions before moving anything. The marker is the THIRD rung: env > project config > marker > 'claude'. N1 was checked rather than assumed, and my first reading of it was wrong. A marker holding an unknown name comes back essentially verbatim, which looked like a validation gap. Measured against the env rung with the same inputs - including "../../etc/passwd" and "claude;rm -rf /" - the two are identical, because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that, and it is met. The residual (the shared normalizer normalizes shape, it does not validate against the known-runtime set) is pre-existing on the env rung and plausibly deliberate, since a new runtime should not need a code change. The marker also does not widen the trust boundary in any real sense: it lives inside the install tree beside the code, so anyone who can write it can write runtime-slash.cjs itself. Rung 3 — the map was redundant; the silent fallback was the defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements. The map carried nothing the contract did not already have, so it is DELETED rather than reconciled. What was actually broken is `|| 'read-only'`, which silently under-granted 24 of 35 roles. 16 of those 24 declare Write or Edit and would widen under derivation. Per the maintainer's decision (HALT.md), they are held at read-only pending the question of whether Codex enforces sandbox_mode or merely advises. Emitted TOML is therefore byte-identical for all 35 roles - asserted per role, not in aggregate, because an aggregate passes while one role silently widens. The hold list self-invalidates. A hold whose role no longer derives broader fails, and so does a hold naming a role with no agents/<name>.md. Without that it would rot into exactly the hand-maintained subset map being deleted, and this commit's own ledger claim would become false over time. Both cases were proved by injecting them and watching them throw. Two committed tests asserted the deleted map's existence and contents. They were pinning the thing being removed, so the tests moved rather than the production code: the 11 role-value pairs survive as a test-local PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real derivation against real agents/*.md. The coverage is preserved; only its source moved out of production code. validate agents gains checkCodexSandboxPosture, mirroring the existing checkCodexModelPosture: each installed TOML's sandbox_mode must equal the role's expected value, failing with role, expected and found. It previously checked file presence and manifest completeness only, so a TOML whose sandbox_mode disagreed passed. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: git log -S returns only documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^ carries five occurrences, and the implementation here matches that code rather than a guess at its semantics - including first-write-wins on a duplicate short form, deterministic from the sorted plan order. It resolves the bare plan number: depends_on: ["01"] now reaches 26-01-auth-hardening. That is a control-flow change, not a diagnostic one - plans that silently collapsed into a single wave 1 now execute in their declared waves, and execute-phase.md consumes those wave values. In-phase only, by construction: the map is built from this phase's rawPlans, so a same-named short form in another phase does not resolve. #3785's display-mapping passthrough and #3885's unresolvable-token warning and wave-verdict suppression are untouched and stay green. If the third tier had over-reached, those are what would have caught it. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open I introduced, and wire the posture check to its command Two blockers from review. Both are mine, and one is a security regression my own change created. 1. A held role could escape its hold by editing its own frontmatter. The Codex install loop set the sandbox identity from the agent's frontmatter `name:` field rather than from its filename, so the hold lookup keyed off a value the file itself declares: deriveCodexSandboxMode('gsd-doc-writer', <real file>) -> read-only deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write deriveCodexSandboxMode('GSD-Doc-Writer', <same file, name: recased>) -> workspace-write What makes this a blocker rather than a nit is the DIRECTION. The deleted CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an allowlist: an unmatched key fell back to read-only, which is safe. The new scheme derives workspace-write from the tool contract and uses the hold as a subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk into a fail-open one and did not notice; the isolated reviewer proved it by execution. Neither safety net caught it. validateCodexSandboxHolds only checks that <key>.md exists, never that a file's derived identity matches its key. checkCodexSandboxPosture looks the canonical source up by the installed TOML's filename, finds nothing for a renamed agent, and treats it as a custom non-roster agent — silently no violation. The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds already validates and what an attacker editing frontmatter cannot change without renaming the file — at which point the existing validator catches it. The lookup is case-insensitive so a recase does not slip past either. The frontmatter name still drives the TOML body and filename, unchanged; only the sandbox identity moved. All 35 roster files were checked: name matches filename stem everywhere, so a stricter "they must agree or throw" invariant would have been safe against real content. It is deliberately NOT added — it would abort an install on a tampered file where emitting a correctly-derived read-only TOML is the safer outcome. Recorded as a fork rather than decided silently. 2. checkCodexSandboxPosture was exported and never called. cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and checkCodexModelPosture only; grep for the sandbox check in that file returned nothing. So criterion 3 — "validate agents fails on semantic drift, not only on missing files" — was unmet, and `validate agents` behaved exactly as before. That is ADR-3473 Decision 2's named shape: a declared policy with no executor. It also meant the T28 test asserted at the helper's return value while the COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to close, committed by me while enforcing it elsewhere in the same epic. Now wired as an additive `sandbox_posture` field beside `codex_posture`, following the sibling precedent exactly. Drift is report-only, not a non-zero exit, because that is what checkCodexModelPosture does — two sibling posture checks disagreeing about whether a violation is fatal would be its own defect. The choice is recorded in a comment rather than left implicit. A consumer-output test now drives the real CLI and asserts on the emitted JSON, and was shown failing before the wiring and passing after. Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3 was not in this deliverable, written while it was halted and false once the maintainer unblocked it. Verified after both fixes: the three bypass probes all return read-only, the per-role table is 35/35 byte-identical, and both hold self-invalidation cases still throw. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture field; docs/reference/plan-md.md documents that depends_on accepts the bare plan number. Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a concurrent PR hand-allocating a section number, regenerated into FEATURES.md. ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style, recording what was measured and built against the section's 2026-08-26 correction - including the qualification that checkAgentsInstalled itself still checks presence only, and the semantic assertion lives in a sibling wired into validate agents rather than folded into it. No how-to. Both user-visible changes are zero-step: a non-Claude install resolving its own runtime, and plans executing in their declared waves, both happen without the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does not touch, and was deliberately left alone rather than edited by association. No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it documents Claude-side tools frontmatter, never Codex sandbox_mode, and the emitted tools contract did not change. The prompt layer documents depends_on only by example, not by schema, so nothing there needed editing - and few-shot-examples/plan-checker.md already showed depends_on: ['01'], which now actually resolves. Translated copies of plan-md.md are untouched; the project treats translations as community-maintained. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser The full suite came back with 26 failures across four files. Three distinct causes, mapped individually rather than assuming the first explained the rest. A. Requiring bin/install.js printed the GSD banner to stdout and corrupted `validate agents` JSON. Unexpected token '', "[36m ██"... is not valid JSON checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring bin/install.js, whose module load prints the ASCII banner. So the command emitted banner bytes before its JSON and every JSON consumer broke, including ten tests that predate this branch. src/ reaching into bin/ was backwards layering that happened to also be loud. The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML domain module, no new module and no six-gate ripple - and both bin/install.js and src/agent-install-check.cts import it. One owner, which is §8.3's rule applied to the fix for §8.3. B. The stale-hold throw fired on a legitimate partial source dir, and masked a security assertion. validateCodexSandboxHolds treated "this hold's .md is absent from the install SOURCE dir" as a stale hold and threw. A test fixture, or any partial install source, legitimately contains a couple of agents. Worse, it threw BEFORE the path-escape check, so a test asserting that a `../../evil` frontmatter name is rejected got my unrelated error instead of the traversal rejection it was written for. A fail-closed check of mine was hiding a real security check. The "no stale holds, shrink-only" invariant is a property of the repo's canonical agents/ roster, not of whatever directory an install happens to read. It is off the runtime path and enforced where it belongs, in the tests that already existed for it. A partial source dir now installs cleanly, and the evil-name case throws with its own escapes-configHome message again. C. T8 depended on ambient process.env state. The marker/env parity assertion round-tripped through live process.env. It now compares against resolveExplicitRuntime's already-exported dependency-injection parameter - deterministic and hermetic, same claim. Proven still falsifiable rather than assumed: with the marker rung's normalization temporarily bypassed the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd") and the assertion fails, then passes again once reverted. One correction folded in along the way. The first version of the move added private _extractFrontmatterAndBody/_extractFrontmatterField helpers to codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893). Adding a third inside the epic whose thesis is one implementation per rule is not defensible. deriveCodexSandboxMode no longer parses anything: it takes (identity, toolsValue) and each caller supplies the tools value using the extractor it already has. Both helpers are deleted. The identity argument is still the filename stem, so the fail-open fix is untouched. Verified after all three: `validate agents --raw` emits parseable JSON with no banner and both posture fields; the four hold-bypass probes still return read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9 workspace-write; the hold list is still 16. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test Suite down to 7 failures from 26. Three more causes, mapped individually. A. My extractor import dragged in a script that does not exist in an installed tree. Cannot find module '../../../scripts/fix-slash-commands.cjs' Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs, which requires command-roster.cjs, whose line 36 requires ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not in an install, so every test exercising a synthetic install dir died at module load. I picked that extractor for convenience without checking what it pulls in - the same mistake that produced the banner bug, one layer further out. agent-install-check now uses a single-purpose extractToolsLine on codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser: we deleted those helpers a commit ago for good reason, and this reads one line. Verified from outside the repo root that requiring either module prints nothing and does not throw. B. A test pinned the deleted name-based fallback. 'defaults unknown agents to read-only' called generateCodexAgentToml with a fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent with a writing contract correctly derives workspace-write - design row S6, a new writing role gets the contract, not the pin. The behavior it asserted was the silent fallback this rung deleted; identity no longer decides the sandbox. Replaced with two rows rather than a flipped string: no tools declared -> read-only (absence is not a grant), and Write/Edit declared -> workspace-write. Strictly more coverage than the row it replaces. C. The stale-hold check still threw per derivation call. Last commit took the roster-existence check off the install path, but deriveCodexSandboxMode itself still threw when a hold's role did not derive broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture for a held role. The throw is gone, and it cost nothing: if a held role's content does not derive broader, the hold pins read-only and derivation returns read-only anyway, so the hold is a no-op and there is nothing to fail about. The staleness invariant is a property of the real agents/ roster, and validateCodexSandboxHolds still enforces it there - confirmed against the real roster after the change, not assumed. deriveCodexSandboxMode is now total: every (identity, toolsValue) including undefined and null returns read-only or workspace-write, never throws. Verified: validate agents emits parseable JSON; the four hold-bypass probes return read-only; the per-role table is 35/35 at 26 read-only / 9 workspace-write. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path The ADR entry and the feature fragment both ended their rung-3 explanation with "see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at read-only was reachable only from the machine that produced it. A reader of the ADR got a pointer to nothing. Both now carry the reasoning inline: the criterion asks both that the sandbox derive from the declared tool contract and that no role gain a broader sandbox, and those cannot both hold, because a faithful derivation widens 16 roles the deleted map never listed and that fell through its silent read-only default. The resolution is derive-and-hold - the derivation owns the rule now, each hold is released as its enforcement question is answered, and a hold is reversible where a widened sandbox that turns out to be enforced is not. Checked before assuming this was a defect class: CONTEXT.md cites .gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight module entries, and four other shipped docs do the same. Citing a phase artifact is an established convention here, so those are left alone. What was wrong was specific to these two: they put load-bearing rationale behind the pointer instead of provenance. docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs rather than hand-edited. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration Two orthogonal reviews on the shipped sha. Three of the findings are the same failure class this epic exists to close, committed inside it. 1. BLOCKER - the sandbox was decided for one identity and applied to another. bin/install.js derived sandbox_mode for the filename stem and then wrote the result to `${name}.toml`, where name comes from the file's own frontmatter. Make the two disagree and a HELD role's artifact goes wide: rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml add any gsd-*.md whose frontmatter name: is a held role -> clobbers that role's toml with workspace-write Both emit read-only on origin/next, because the deleted map was an allowlist and a miss fell back safe. This is a regression my change introduced. The previous review round moved the HOLD KEY off frontmatter to the filename stem and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985 calls that value attacker-editable, four lines above the line that uses it as the filename. The decision is now made over BOTH candidate identities, most-restrictive wins: if either the stem or the emitted name is held, the mode is read-only. 2. MAJOR - hold matching was toLowerCase() only, so confusables escaped. Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline, ./ and ../agents/ all slipped the hold and emitted workspace-write. Identities are now basenamed, trimmed of NBSP/zero-width/control characters, NFKC-normalized and lowercased - and anything still carrying a character outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not enumerate confusables; every shipped roster file is ASCII, so refusing to widen on an identity we cannot recognize is fail-closed with no false positives on real content. 3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY. shortFormToId keyed on the last dash-segment of any canonical id with no constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the worst shape in the epic: the unresolvable-token warning fires on a DROPPED token, so a MIS-RESOLVED one is invisible and the tool reports a confident wave assignment built from a wrong edge. A wrong edge is worse than a missing one. The segment must now match /^\d+$/, which is exactly the contract docs/reference/plan-md.md already documents. This tier was recovered verbatim from the retired SDK lineage, which carried the same defect; we are deliberately NOT preserving it bug-for-bug, and the comment says so, so the next reader does not "restore" it. 4. MAJOR - the derivation was reading a declaration as an absence. extractToolsLine read one line, so a YAML list-form tools: block returned only its first item. Two roster files use list form, and gsd-nyquist-auditor declares Write and Edit there - parsed as "- Read", found no write tool, and emitted read-only. Rung 3's headline claim is that sandbox_mode derives from the declared tool contract; that claim was false for 2 of 35 roles and materially wrong for 1. Reading a declaration as an absence is the silent-drop class this epic exists to close. Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per the standing derive-and-hold decision - so emitted TOML stays byte-identical at 26 read-only / 9 workspace-write while the hold list finally records every role that would widen. A previous pass declined this fix because it moved the count; that inverts the priority. Byte-identity is preserved THROUGH the hold, not by leaving a parser broken. Divergence check, because this is where that bug hides: both paths feeding sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now route through the one extractor. The tools readers in runtime-artifact-conversion and install.js's other frontmatter call sites serve Claude-side emission and do not feed sandbox derivation. Also fixed, each real: the posture check's `found` used a naive whole-file regex where its own sibling uses the block-aware scanner, so prose inside developer_instructions produced a false violation; `found` skipped truncatePostureValue and leaked a 300-char value into validate agents output; deriveCodexSandboxMode's absolute never-throws claim was false for an object with a throwing toString; T49 could not falsify cross-phase leakage (its target phase had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded table and pinned the FIXTURE size, so a 36th agent would be silently unchecked; three tests reimplemented the code they were testing instead of importing it; and T2-T4 deleted GSD_RUNTIME without restoring it. Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer resolves while ["01"] still does, both identity-bypass cases and every confusable vector emit read-only. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the hold list is 17, and the reason the 17th was missing The count read 16 because the derivation could not read the declaration it claimed to derive from: the tools reader was single-line, so a YAML list-form tools: block returned only its first item and gsd-nyquist-auditor's declared Write and Edit were read as an absence. Both the ADR entry and the feature fragment now carry the corrected count and the reason for it, rather than a silently updated number. Deriving from a declaration you cannot parse is not deriving, and a flattering count is worse than a wrong one because it looks settled. docs/FEATURES.md regenerated from the fragment. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3897): backfill changeset pr number Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
34399eed70 |
docs(#3942): ADR — the emitted-drift ack belongs in a commit trailer (#3943)
The acknowledgment explains one PR's unattributable emitted-artifact delta. The moment that PR merges the delta is in the base, so the acknowledgment can never clear anything again — its useful life is exactly the PR's open window. It is stored in permanent, shared, merge-path state. That single lifetime mismatch is the origin of six consecutive rounds of defect-and-fix (#2789, #2914, #3078, #3823, #3842, #3875), each fix generating the next defect, ending in a scheduled sweeper whose own PR (#3927) could not merge itself. ADR-3942 supersedes ADR-2719 section 3 only, and its #2789 Amendment. Sections 1, 2 and 4-7 are retained and depended upon: the conservation law itself is not in question, only where its escape hatch is stored. Also lands the root-cause research note behind the decision, and regenerates the ADR index. Two corrections to earlier framing, recorded here rather than dropped: - Per-commit trailers do NOT reliably survive squash-merge; the repo's own ship.md:306 reconstructs a gate trail precisely because squash discards it. The ADR therefore claims no durable audit record on next. Non-survival is the property being bought, not a cost. - The test job checks out at depth 1 (test.yml:107-110, no fetch-depth key). A depth-1 checkout cannot see the PR's commit range and fails VACUOUSLY rather than erroring, so fetch-depth: 0 is a named decision in the ADR and the acceptance criterion requires a test that fails under depth 1. Docs-only; no code, no changeset. Refs #3942 Co-authored-by: sim <sim@local> |
||
|
|
6b7df61938 |
enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises ADR-3473 §8.1 carries a blocking open question with a forcing function: it must be answered before any implementation PR for the rule opens. Answered here as (a), a string-coercing adapter, with the measurement that settles it. The sequencing note bet that §8.8's schema would make (b) tractable. Measured against merged reality it does not: only 33 of extractFrontmatter's 78 non-test call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists survive real types, leaving ~31 lines across 3 call sites as the actual prize. Also corrects three claims verified false while answering it. §8.1's justifying sentence names #3349 and #3360 as defects a real parser would fix; both are already fixed on next, confirmed by executing the compiled parser rather than reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an expected casualty of this rule, but it guards shell grep idioms in workflow bash fences and never touches our parser. The same roster calls lint-vendored-deps.cjs reusable as-is; it is hardcoded to re2js throughout. The last two were caught by applying the rule this amendment records -- a factual claim in this ADR is a hypothesis until the implementing phase executes it -- on its first use. Refs #3881 * docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable An adversarial pass on the Phase 4 design established by execution that extractFrontmatter is not a YAML parser but a line-oriented scanner whose output is a function of raw source text. Four spellings of the same value collapse to one js-yaml tree but produce four distinct legacy strings, one of them mangled. No adapter over a tree can choose among outputs the tree does not distinguish, so fork (a) -- keep a string-coercing adapter so the existing contract holds -- cannot be built. For any document with a non-scalar value, (a) collapses into (b); about 26 percent of frontmatter-carrying documents have one. Also records three design defects and one new attack surface, all confirmed by execution: catching a parse failure and returning {} would delete the frontmatter block on the next write at eight call sites that conflate empty with unparseable; an empty value yields null where legacy yields {}, and reconstructFrontmatter omits null-valued keys, so the shipped state template's empty progress key would vanish; the #1882 truncation probe is parseYamlRegion itself rather than a pre-parse heuristic, so it cannot both stay unchanged and survive that deletion; and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB. The rule is not deferred. The measurement is the deliverable and the re-scoping is recorded as an open question with a forcing function, per section 8's own rule. Refs #3881 * test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed. Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states. B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today. B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today. B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today. Refs #3881 * chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496). gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own. src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare. scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored. docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds. Refs #3881 * feat(#3881): parse .planning frontmatter with the vendored js-yaml ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and parseQuotedScalar are deleted (not patched); parsing now goes through the vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA, json: true }. Everything js-yaml does not do is layered on top, in one place, carrying the seven design-doc consequences: 1. Empty value: a null js-yaml value is coerced to {} (matching legacy's own empty-value contract) so reconstructFrontmatter — which omits null-valued keys — still round-trips a bare `key:` line instead of deleting it. Verified live: progress: with no value survives parse -> reconstruct -> re-parse. 2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS channel, is carried on the {} returned for malformed/refused YAML. Invisible to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that never inspect it are unaffected; wiring the 8 hasFrontmatter sites to consult it is a separate change, not done here. 3. Non-scalar object-list items (the four spellings of `- test: a b` that js-yaml collapses into one tree shape) are rendered as a canonical `key: value[, key2: value2]` string per item, keeping the existing array-of-strings value SHAPE. A full corpus differential over all 1702 tracked markdown files found 11 residual divergences from the legacy parser (enumerated in the PR/report), most of them the parser now being MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key defect, a dropped Unicode key). 4. The #1882 truncation probe still runs the one real parser, but derives its key count from js-yaml's own thrown error and mark.line when the whole region doesn't parse cleanly (the dominant real truncation shape: fence opened, well-formed keys, no closing fence). Verified against both the clean-parse and the exception-fallback path. 5. The #3257 comment channel now attributes each pending column-0 comment against js-yaml's own parsed top-level key list (matched by literal key text, in document order) instead of the legacy ASCII-only key regex, so a comment above a Unicode key attaches correctly. 6. Anchors, aliases and merge keys are refused outright (a raw-text pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences today: zero. A 7-line billion-laughs fixture is verified refused rather than expanded. 7. A literal U+0000 is swapped for a private-use sentinel before the parse and restored in every resulting string afterward, since js-yaml rejects NUL unconditionally under every schema. escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump() (forced double-quoted style), with control-char hex escapes lowercased to keep serialized output byte-stable (#1779 emitted lowercase); it keeps its exported name and signature for its two other call sites (commands.cts, runtime-artifact-conversion.cts), which need no change. frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments, regenerateFrontmatterKey's guard, noOpObjectListSetError and parseMustHavesBlock are all unchanged — retiring them is fork (b) and is not this phase. Refs #3881 * fix(#3881): quote template placeholders and preserve unparseable frontmatter SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax under the vendored js-yaml parser, not the literal placeholder text intended. Quote them so they parse as strings. Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the 8 call sites in state.cts/state-transition.cts that compute hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and reassemble the document without a frontmatter block when false. That check conflated 'no frontmatter' with 'unparseable frontmatter' (both parse to {}), so a document with a merge-conflict marker or refused alias in its frontmatter had that block silently dropped on write. Each site now preserves the exact raw bytes stripFrontmatter removed when the marker is set, leaving the genuinely-empty case unchanged. Refs #3881 * test(#3881): consequence and boundary coverage for the js-yaml migration Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix. Refs #3881 * docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to Refs #3881 * docs(#3881): correct the frontmatter glossary entry Two errors in the entry as first written: it named parseYamlRegion as part of the read path when that function is deleted, and it recorded the eight hasFrontmatter call sites as unwired follow-on work when they were wired in e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the frontmatter block independently, so the marker binds at the transform layer. Refs #3881 * docs(#3881): record the semantic-migration decision and the counted guard ledger The maintainer chose the full semantic migration over splitting the rule into its own epic or patching the scanner, so section 8.1 is answered as "the fork was ill-posed and the migration is semantic" rather than as (a) or (b). Also replaces the pre-implementation guess that this phase would shrink the guard surface with the counted result: excluding vendored third-party lines the hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines despite four functions being deleted, because the compatibility layer over js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is therefore not delivered as written; what improved is the kind of code maintained, not the amount. Decision 6 requires recording that rather than netting it away. Refs #3881 * chore(#3881): changeset for the vendored YAML parser migration Refs #3881 * test(#3881): golden parity, round-trip property and packaging coverage Refs #3881 * fix(#3881): refuse anchors structurally and fold in review findings ADR-3473 §8.1 review findings, addressed inline: Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping ({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor mechanics while never matching that line shape, so the exact expansion the guard exists to stop went straight through unrefused (a 303-byte quoted-key bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener` callback, which reports `state.anchor` for every event belonging to an anchored node in every spelling, and throws from inside the callback to abort before any expansion (~1-2ms vs full expand-then-discard). A merge key with an alias is still refused (merge always requires a previously anchored node, so the alias itself trips the listener); a bare merge key with NO alias is no longer separately refused, documented as intentional: FAILSAFE_SCHEMA never resolves `!!merge`, so it carries no expansion risk. Table-driven tests added for all four bypass spellings + merge key, plus a quoted-key-spelled billion-laughs fixture registered in the adversarial matrix and README. Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/ aliases were "simply UNREACHABLE from typed code" through the twin. Corrected to state the truth: anchor/alias resolution is document-level `load` mechanics reachable through exactly the declared surface, and refusal is enforced at RUNTIME (Finding 1's listener), not by the type surface. Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree back to NUL, including one the document author legitimately wrote, silently corrupting it. Now refuses outright whenever the raw region already contains U+E000 (consistent with the existing anchor/merge-key refusal path), making the substitution provably injective. Tests added for a real NUL alone (preserved), a pre-existing U+E000 alone (refused, not corrupted), and both together (refused, not merged into one byte). Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead for a hand-authored row (only read inside the upstream-verbatim branch) — exactly how Finding 2's stale docblock drifted unnoticed. Added checkHandAuthoredTwin: every value-level export the twin DECLARES must be an actual own property of the vendored runtime module at require-time. Tests added, including a sensor that a declared-but-nonexistent export IS caught. Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/ unparseableFm/reassemble preamble, copy-pasted at 7 sites in state-transition.cts plus a sixth hand-inlined copy in state.cts's cmdStateCompletePhase, is now one exported helper (beginFrontmatterReassembly) every site routes through, including the hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore) keep a literal `body = stripFrontmatter(content)` assignment alongside the helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward scan (which does not chase aliases) still sees the strip; stripFrontmatter is pure/idempotent so the extra call changes nothing observable. Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a separate change" claim (the 8 call sites are wired on this branch) and the changeset's backlink from (#3473) to (#3881). Finding 7: fixed the lint:ci failures blocking the gate — an @typescript-eslint/only-throw-error violation from throwing a bare Symbol as the anchor-detected signal (now a real Error subclass), unused-var warnings left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by two migration-specific test files (allowlisted with justification), and the lint-state-write-path-drift false positive from Finding 5's helper (fixed above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already carried an explicit timeout; no change was needed there. Golden fixture: added a golden entry for the new anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line scanner would also produce, since it independently dropped every quoted top-level key). No other corpus document diverges: real .planning/ documents carry zero anchors/aliases/merge keys/U+E000 today. Refs #3881 * fix(#3881): fold in second-round review findings Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration ASCII-only key regex for the Unicode fixture; updated to require the 相 key's value now that js-yaml has no such restriction. Audited the rest of the file for other pre-migration pins (block scalars, quoted keys, flattened values, empty values, duplicate keys, unclosed blocks, null bytes) by execution against real fixtures; found none regressed. Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts so no function still answers to the deleted hand-rolled scanner's name (ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three external call sites (src/commands.cts, src/runtime-artifact-conversion.cts) updated in the same change — a mechanical rename, not an ADR-amendment matter. Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found by execution. A top-level key named constructor/__proto__/toString/ valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read resolving an inherited Object.prototype member); a key literally named __proto__ was silently DROPPED entirely (bracket assignment on an ordinary {} invoked the inherited __proto__ setter instead of creating a data property). Fixed by building every parsed Frontmatter object with Object.create(null), and replacing an `in` check with hasOwnProperty.call in propagateCommentChannel. Added round-trip tests for all five hostile keys, each with its own leading comment. Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed full byte-stability across the migration. Verified by execution: BEL/NUL/ NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw- literal forms. Proved round-trip equivalence (each escape re-parses to the exact source codepoint) and corrected the docstring. Found and fixed a related real defect while verifying: a lone UTF-16 surrogate was emitted BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely unparseable YAML that silently collapsed to {} on re-read — extended scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped path. Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real truncation shapes (unquoted colon, open flow collection, mis-indented sibling key, refused anchor). Root cause: the mark-based prefix recovery excluded the very line whose key needed counting, and a mark-less refusal never entered the recovery branch at all. Fixed by taking the max of two lower bounds: the longest parser-verified line-prefix, and a raw-text count of key-shaped lines (reusing the same key-shape pattern this file already uses for isFrontmatterShaped). Extended test-matrix row A5 table-driven over all 4 regressed shapes. Finding 6: the design doc's claim that no test owned the #3594 adversarial fixture corpus was false — consolidation epic #1969 had already folded it into tests/frontmatter.test.cjs. An earlier commit on this branch re-created a standalone duplicate under that false premise; folded its genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures, block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the ADR's §8.1 note, including the roadmap-sibling claim (no such file exists). Finding 7: the golden serializer sorted object keys, making it structurally blind to the key-order-parity invariant ADR-3473 §8.1 actually claims. Made it order-preserving and regenerated the golden fixture from a standalone compile of the legacy (pre-#3881) parser at ddde001af; the current parser matches it with zero undocumented divergences, confirming key-order parity genuinely holds. Extended row A2 table-driven across 6 of the remaining 7 transitionCore kinds (all pass) plus documented, by execution, a newly-discovered 8th-site regression: state.cts's cmdStateCompletePhase calls the same preservation helper but its result is clobbered by a later unconditional resync — filed as a distinct finding rather than fixed here (touches syncAndPreserveStateMd, outside this change's verified scope). Refs #3881 * fix(#3881): preserve unparseable frontmatter through the CLI write path Characterization (executed, before/after shown): case (b), not (a). The frontmatter FENCE survives — `state complete-phase` on a conflict-marked STATE.md returns success and a well-formed, freshly-derived frontmatter block, not a document with no frontmatter at all. But the block's actual content (the merge-conflict markers, and with them any signal to a human that the document was in conflict) is silently discarded and replaced. Root cause was two clobber sites, not one: 1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved `transformedContent` from readModifyWriteStateMd, finds {} + the FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh frontmatter block from the body anyway. 2. Even after (1) is fixed, applyPostSyncPreservation's own postFm/applyStatePreservation/authoritativeFm-reassertion machinery re-extracts frontmatter from syncedContent, restores curated fields from the pre-write snapshot, and reconstructs a NEW block again — confirmed live via `state begin-phase`, which still lost the markers after fixing (1) alone. Both are now guarded by the same predicate (isUnparseableFrontmatter, checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not parse and the caller is not on ADR-3408 §8.3's closed "body wins" list, both functions return their input content unchanged rather than re-deriving over it. The closed list (cmdStateSync #905, /gsd-health --repair's REGENERATE_STATE, both routed only through writeStateMd, which never reaches applyPostSyncPreservation and passes sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is untouched — neither widened nor narrowed; verified by execution that `state sync` still overwrites the conflict-marked block exactly as before. Other verbs sharing the same readModifyWriteStateMd path were checked and were equally affected before this fix: state update, query state.patch, and state begin-phase all lost the conflict markers (RED, shown by execution), and all three now preserve them (GREEN). Covered table-driven in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe block, which drives the real CLI verbs via runGsdTools — not just the pure transitionCore layer the earlier A2 rows exercised — plus a control asserting state sync's body-wins contract is unchanged. Refs #3881 * fix(#3881): restore the parse surface's prototype and fix remote-runner failures Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched. Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write. Refs #3881 * chore(#3881): backfill changeset PR number Refs #3881 * test(#3881): make golden parity resistant to unrelated tree churn A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus. Refs #3881 * test(#3881): make the parser golden hermetic instead of tree-keyed This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md, docs/*.md) the prior golden pinned by tracked path. Any PR editing one of those files' frontmatter for reasons unrelated to the parser (an argument-hint addition, an allowed-tools tweak) turned the suite red, and the reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to catch a real regression. Excluding .changeset/** was not enough; the design itself was wrong: a regression fixture must not be keyed to mutable repo paths, and a single 376-entry JSON every such PR touches is also a guaranteed merge-conflict surface. Rebuilt the fixture to carry its own documents: each of 51 entries stores a stable id, literal documentText (shrunk from a real ddde001af-era corpus document), and an expectedParse captured independently from the pre-migration legacy parser (git show ddde001af:src/frontmatter.cts, compiled standalone against its byte-identical sibling modules). The test reads no tracked path, shells out to no git command, and enumerates no tree -- a PR editing commands/gsd/help.md cannot affect it. Every entry's reconstruction was verified at capture time to reproduce both the current and legacy parser's output on the original document; 0 of 51 candidates were dropped by that check (1, the deliberately-unterminated unclosed-block.md adversarial fixture, has no closing fence to truncate at and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now diverges:true entries) and the D2 order-preserving structural serializer that keeps the comparison from passing vacuously; dropped the tree-enumeration helpers, the coverage floor, the post-capture-skip logic, and the vanished-file check -- all artifacts of the path-keyed design. Refs #3881 * fix(#3881): resolve vendored-deps paths independently of cwd shape Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on windows-latest CI: the test passed absolute scratch-file paths into compareFiles()/checkRow(), whose helpers joined every input onto ROOT via path.join(ROOT, rel), producing garbage when the input was already absolute. It surfaced on windows-latest specifically because GitHub's Windows runners checkout the repo on a different drive than TEMP, so path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged (no relative traversal is representable across drives) rather than the relative form the test assumed. The remote gsd-test runner this repo gates pushes on is Linux-only and could never have caught this; GitHub CI's windows-latest job is the only signal that does, and it did. Fixed the helper itself (scripts/lint-vendored-deps.cjs's new resolvePath()) to treat an already-absolute input as absolute-in, absolute-out instead of silently mis-joining it, and updated the test to pass the scratch file's absolute path directly rather than relying on a relative conversion that is not always representable. Kept every mutation-sensor assertion intact and added coverage proving resolvePath is a no-op for relative inputs and correctly passes absolute ones through unchanged. Refs #3881 * fix(#3881): warn when state sync regenerates over unparseable frontmatter state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly overwrites an unparseable frontmatter block per its 'body wins' contract — that overwrite behavior is unchanged here. The defect was the silence: synced:true/exit 0 gave no signal that the existing block (including git merge-conflict markers) could not be parsed and was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be reported as authoritative when the derivation dropped input it could not resolve') and §8.4 ('failure is a value'). Adds a gsd: warning — ... (#3881) line on stderr, matching the existing #3573 precedent, and surfaces the same disclosure in the JSON result's existing changes[] array so a machine consumer sees it too. Exit code and synced:true are left unchanged — sync did what its contract says. REGENERATE_STATE (/gsd-health --repair's sibling on the same sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally refused by applyRepairs's dispatcher before runRepairAction ever runs (src/health-diagnostic.cts), so it is not a live path today and is not in scope for this fix. Refs #3881 * fix(#3881): exit non-zero when a state command returns an error Refs #3881 * chore(#3881): changeset for the state exit-code fix Refs #3881 * fix(#3881): honor the documented --project-dir flag Refs #3881 * revert(#3881): restore exit-0 result envelopes for state errors Reverts 9638f2936 and its changeset. The change was wrong and the revert is the correction. This repo distinguishes two error mechanisms deliberately. error() in src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure path. output({error: ...}) writes a JSON result envelope to stdout and returns normally with exit 0. The reverted commit converted 23 result-envelope sites into hard failures, which is a different contract, not a bug fix. tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope contract directly -- a failing command exits 0 with a JSON error envelope and must not publish state.json -- and the remote matrix run caught it along with four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that the original commit rewrote were encoding that real contract, not the bug it claimed; they are restored. Whether an error envelope on stdout with exit 0 is the right CLI design is a genuine question, and it is section 8.4's rule ('failure is a value') with its own phase. It is not something to flip inside this PR. Refs #3881 * chore(#3881): backfill changeset PR number for the project-dir fix Refs #3881 * test(#3881): keep the frontmatter mutation shard inside its time budget The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard cap. Root cause is NOT row-level spawn overhead (contrast the #2790/ core-utils precedent): the three shard test files' own logic runs in ~413ms total (356+30+27ms) with all 392 assertions passing. Instead, src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to the vendored YAML parser, proportionally growing the mutant count Stryker generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner bills the full 'node --test <3 files>' invocation once per mutant, and node:test's default per-file process isolation forks a child process for each of the three files on every one of those invocations — pure fork overhead multiplied by a much larger mutant population. Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares isolation: 'none', and .github/workflows/mutation.yml passes --test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e. unchanged behavior — for the other 8 shards, which were not individually audited for cross-file state leakage under shared-process execution). Measured locally via node:test's run() API on the exact 3-file set: isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same 392 passing assertions. The true CI-shard number can only be confirmed on the GitHub Actions run (Stryker cannot run locally, and 'node --test' is hard-blocked in this environment). Refs #3881 * test(#3881): register the vendored-parser tests in the frontmatter mutation shard stryker.config.mjs's own rule ("Keep this list in sync with the tests arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew src/frontmatter.cts from ~825 to 1496 lines but its new tests (tests/feat-3881-yaml-parser-consequences.test.cjs, tests/frontmatter-golden-parity.test.cjs, tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in tests/frontmatter.test.cjs) were never added to the frontmatter shard's tests array, so Stryker's mutants in the new vendored-js-yaml adapter had nothing constraining them. PR #3888 measured 55.8% against the 65 floor (748 killed / 593 survived / 17 timeout) and the shard was separately cancelled at 15m04s against the 15-minute per-shard cap. Registers all four files (each earns its slot on evidence of a unique constraining assertion, documented inline), gives the shard a measured/projected 180-minute budget via a new per-module timeoutMinutes field threaded through mutation.yml's job-level timeout-minutes the same way isolation is threaded, and removes the prior isolation:'none' override (re-measured at this file-set size, its savings are within run-to-run noise, not worth the unaudited cross-file-state-leakage risk). Refs #3881 * feat(#3881): derive the mutation test list and ratchet the score floor Refs #3881 * test(#3881): ratchet five stale mutation floors and close the frontmatter gap Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1): config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78, context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs RATCHET_BASELINE in the same diff per the ratchet's own contract. Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter), repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via extractFrontmatter), and the null-byte sentinel round-trip surviving at region offset 1. Did not lower minScore. Refs #3881 * test(#3881): decouple the ratchet test from real module floors The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour. Refs #3881 * refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency Refs #3881 * fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives. Closes #2571 Refs #3881 --------- Co-authored-by: sim <sim@local> |
||
|
|
c1f755d75d |
docs(#3889): ADR-3889 — the process-exit contract, nothing fails with success (#3890)
* docs(#3889): ADR-3889 — the process-exit contract Records the decision for epic #3889: one Outcome vocabulary, one pure projection to an integer, two terminators sharing both. Doc-only. Regenerates docs/adr/README.md via gen-adr-index. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3889): rewrite ADR-3889 as an exit-code registry, on measured evidence The first draft's census was grep-derived and roughly 2x inflated: it counted comment prose and also matched process.exitCode, the correct pattern. An AST census (@typescript-eslint/parser, CallExpression on process.exit) gives 128 deduped call sites, 71% of them in hooks/, against 88 process.exitCode assignments already in use. scripts/ is effectively migrated (39:1). Corrects the framing accordingly: the seam (src/cli-exit.cts) already exists and is already adopted; what is missing is an allocator and a hooks adapter. Replaces the fixed six-value enum with a registry: 0 and 1 free, 2 reserved to the Claude Code hook protocol, 3-13 forbidden (Node reserves them), 64+ allocated with one meaning and one owner per code. Domain codes permitted. Adds three findings obtained by reading and executing the code: - ui-safety-gate/api-coverage report empty stdin as an authoritative negative verdict (exit 1) — reproduced with a control - four modules hand-roll the same three-outcome convention, documented only in a comment at src/ui-safety-gate.cts:137 - two capability fragments fabricate {"detected":false} on probe failure, while the honest {"skipped":true,"reason":"sdk-failed"} form already exists in-tree Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
389bc86e02 |
enhance(#3883): route every slug call site through its canonical owner (#3896)
* docs(#3883): correct section 8.3 against the tree
I wrote this section in Phase 0 stating rules I had not executed against the
tree. Measured at
|
||
|
|
ddde001af6 |
enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md reference carries a Status lifecycle section that is missing from all four translations — the section documenting the status enum whose clobbering is #3853. The test derives the heading set rather than hard-coding the missing one, and names the locale and the heading when it fails. Two tripwires that must pass today and after. The field-drift guard still catches a re-derived fallback ladder: §8.8 instructs deleting that script, and that instruction rests on a wrong premise about what it guards, so the test stops a future reader from deleting it on the ADR's word. And last_activity's label resolution is pinned to what ships today, because it is declared in one of the two tables this phase consolidates and not the other — the consolidation must not silently pick a side. The locale test buckets under docs rather than state, which is what it tests; that bucket is allowlisted with justification rather than folded into an unrelated docs suite. It reads only markdown, so it carries no allow-test-rule marker — a marker there would suppress nothing and would grow the unverified pool against its ceiling. Refs #3873 * feat(#3873): one schema owns the STATE.md key set, three tables become projections ADR-3473 §8.8. The key set was declared in four places that had to agree by hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE, FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One frozen null-prototype schema now declares each key's type, enum, cardinality, source, preservation, body source, body label, accepted parse shapes and whether it is emitted unconditionally; the three tables are derived from it at module load. The projections are byte-identical to the literals they replace, key order included, and the parity tests compare against verbatim copies of today's tables rather than re-deriving both sides from the schema — a parity test fed from one source proves nothing, which is how a consolidation ships a changed policy under a green test. last_activity was the live disagreement: present in one table, absent from the other. The schema declares what ships today rather than the tidier answer, and a test pins it. The schema is a leaf module and owns the four field-policy types, re-exported from state-transition so existing importers are untouched — the same split health-diagnostic-types made to break a CJS require cycle. Refs #3873 * feat(#3873): generate the schema-derived regions, parity-check the prose tables ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in the shipped template and all five reference docs, follows gen-features.cjs's fail-closed contract, and is wired into regen:derived and lint:generated-sync. The Status lifecycle section was missing from all four translations — the section documenting the status enum behind #3853 — and is now generated into every locale. Field cardinality is a new generated table: pure schema data, no prose, so nothing to lose. The Field-reference and Status-values tables are parity-CHECKED rather than generated. Their Purpose, When-populated and Matched-text columns are genuinely hand-translated per locale, and §8.8 itself says prose stays hand-translated; generating them from an English registry would overwrite four locales' translations on every write. The row set is checked against the schema instead, so a key added to one and not the other fails, which is what field drift actually means. Building that check found last_activity_desc undocumented in all five tables. Three keys the docs describe are absent from the schema — active_phase, next_action, next_phases. They are grandfathered by name, not by wildcard, so a fourth fails: a declared gap with a forcing function rather than a silent one. Refs #3873 * fix(#3873): declare what the parsers do, and close the shape-parity gap Two declarations in the new schema described intended behavior rather than actual — the defect class this epic exists to end, committed inside the epic. Both were caught by executing the parsers instead of reading their docstrings. current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid shape errors; the path that looks like support is parseInt truncating '2 of 5' to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands this row must widen, and the shape test will go red until it does — the schema and the parser cannot drift apart quietly, which is what §8.8's checked-not- generated rule is for. STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold. normalizeStateStatus passes unrecognized prose through unchanged, so it is not closed at runtime. The seven members are the canonical values it maps onto; the docstring now says that and the test asserts the real lenient contract. Closes the acceptance item that a test asserts the parsers accept exactly the declared shapes: the check is table-driven over every row carrying acceptedShapes, guarded against passing vacuously on an empty set, and fails loudly if a future row has no registered driver. Adds the unwired-label throw and the fast-check property that every projection agrees with its schema row. Refs #3873 * fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail The remote matrix caught 12 failures with one cause. Making the template's frontmatter a generated region wrapped it in its own yaml fence ahead of the markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which both match the single markdown block — found the heading first, not the frontmatter. That breaks the contract every new project's STATE.md is created from: bug #21 and epic #1969 B8 pin that the File Template block starts with frontmatter and carries gsd_state_version. The markers now sit inside the single markdown fence, so the fence opens before the frontmatter and the region still ends ahead of the heading. Same layout as before this phase, with markers embedded rather than a second fence. Row 27 existed to catch exactly this and did not, because it was writer-seeded: it asserted against the generator's own output shape, so it passed on the broken template. It now parses the fence the way production does and was verified to fail against the broken shape before being trusted against the fixed one. A test that would not have caught the bug it exists to prevent is worse than no test. The emitted-attribution failure was separate and the fragment was the wrong remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy identity rule, so a diff touching it needs no acknowledgment. Fragment deleted rather than left explaining nothing. Refs #3873 * docs(#3873): how to change the STATE.md schema The phase gate was right and my docs artifact was wrong. I listed lint:generated-sync as the second enablement step, which is a verification command dressed as one, and then claimed a one-step sequence owed no how-to. The real sequence is build:lib then regen:derived, and the ordering is a trap: the generator reads the COMPILED schema, so regenerating before building regenerates against the previous schema and commits artifacts that look plausible while disagreeing with the code just written. A reference table cannot carry an ordering dependency; that is what the how-to test is for. The page covers adding, changing and removing a key, every reason code the check emits and what to do about each, what is generated versus hand-translated and why the two prose-bearing tables are parity-checked instead of generated, adding a language, and the three grandfathered keys. Indexed from docs/README.md. Refs #3873 * chore(#3873): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
1863f5569c |
enhance(#3871): the state transaction — mandatory snapshot, open()/rebuild() (#3874)
* test(#3871): failing-first regressions for the dropped curated progress block Pins ADR-3473 §8.6 / #3756 at the consumer's output: state record-session and state add-decision on an archived-milestone project drop the curated progress frontmatter entirely, exit 0, and report nothing. Reproduced against the real CLI before writing the tests, not inferred from the issue text. Also adds the unit-level probe that applyStatePreservation's preserve-always row is inert on a resyncing write, and an over-preservation guard that an empty project is never inflated. Refs #3871 * feat(#3871): make the STATE.md pre-write snapshot mandatory via open()/rebuild() ADR-3473 §8.6. StatePreservationInput's nullable preFm and the always-present preFmSnapshot were the same extractFrontmatter call, one of them nulled on resync — a policy flag baked into a snapshot. Both collapse into a single StateTransaction whose snapshot cannot be absent: openStateTransaction() applies preservation, rebuildStateTransaction() does not, and both carry the snapshot because the reporting phase needs it either way. An absent snapshot is now a construction failure; an empty one stays legal, because that is what a document with no parseable frontmatter honestly has. writeStateMd requires a rebuild transaction, which types ADR-3408 §8.3's closed exception list at both call sites (state sync, health --repair) instead of matching them as strings in a ratcheted baseline. Fixes the dropped curated progress block: an all-zero or absent derived total set is an unmeasured scan, not a measurement, so the curated block stands. Also fixes two defects surfaced while building — preserve-always reported a mutation even when it restored an identical value, and it re-entered the curated object by reference, which would alias the snapshot the next phase diffs against. Refs #3871 * fix(#3871): close the three remaining subsumed defects and restore the arm the type does not replace Review of the first two commits found four things. The guard shrink deleted the seam-bypass axis whole, but only its writeStateMd( arm became redundant. Its other arm catches a call site re-assembling syncStateFrontmatter + applyPostSyncPreservation instead of the owned composition, which the transaction type does not make unrepresentable and which #3469 found live. Restored as findCompositionBypasses, terminal rather than ratcheted. Three of the four issues this phase claims were untouched. All three are the epic's own shape and are fixed at the seam: current_phase_name is reasserted from the curated value when the caller names none, and cmdStateJson stops carrying a hand-maintained list parallel to FIELD_CLASSIFICATION and projects it instead. The construction failure that is the point of this phase had no test. Every enumerated matrix row now has one, including the measured-versus-unmeasured coercion boundary and a seeded property that no curated key is ever dropped. ADR-3473 §8.6 said the guard 'keeps only its raw-write check'. Verified against next: there was no raw-write check, and four other checks it does not name. Amended in place with the evidence. ARCHITECTURE.md separately advertised a preservation policy the code had deleted. Refs #3871 * fix(#3871): do not let the unmeasured-scan rule block an explicitly-requested resync The remote matrix caught over-preservation, the failure this phase's own negative space says must not happen. state update Progress re-derives the block from the body the caller just rewrote; on a project with no phase dirs the derivation yields zero totals, the unmeasured rule read that as 'the scan measured nothing', and the stale curated percent was restored over the resync the user asked for. preserve-always already said what the missing condition was: never overwrite unless the caller explicitly names this field. explicitProgressField carries it and is derived from shouldResyncStateProgress, not set by hand at a call site, so it cannot drift from what the caller asked for. Two defects found in the same mechanism and fixed with it. readModifyWriteStateMd enumerates its option keys, so a new option was silently dropped rather than rejected. And the raw-write axis captured its first argument up to the first comma, which lands inside a nested path.join, so a write to a STATE.md literal was invisible to it — the prove-it-can-fail test caught that one immediately. No test assertion was weakened; all three frontmatter rows encode #3242, #1969 B3 and #1972 and stand unchanged. Refs #3871 * docs(#3871): record why the raw-write check is kept, not why it was named The amendment justified findRawStateWrites as 'written because §8.6 requires it to exist', which is cargo-culting the contract and would have been the wrong reason to keep anything. The real reason is that writeStateMd acquires the STATE.md lockfile and a raw fs.writeFileSync acquires nothing, so this is a lock bypass and lost-update is the #500/#905/#1230 family — and after this phase it is the one reachable path into the file that nothing else covers. Also records why ADR-3408 §8.6's deletion of the 'clear' policy is not the precedent it looks like: 'clear' was dead vocabulary in a closed enum, this is coverage of a reachable path. Refs #3871 * chore(#3871): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
7bbbe495d7 |
docs(#3473): ADR-3473 — enforcement by construction, one owner per invariant (Phase 0) (#3870)
Phase 0 of epic #3473 ships this ADR alone. Every rule in §8 carries a status of Enforced or Required — Phase N. §8.1–§8.5 state the epic's B1–B5 criteria as contract rules with no phase assigned yet. §8.6–§8.8 assign Phases 1–3 to the STATE.md family that ADR-3180 and ADR-3408 left between them: the write seam's unvalidated input, its classification-excluded reporting, and the schema transcribed by hand into nine artifacts. Derivation authority is explicitly NOT in scope here — it lands as ADR-3180 Amendment 8, generalizing §7.5's already-locked sentence. Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
004e9dd741 |
fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability RED by construction. Binds to behavior renderEffortForRuntime does not yet have: an optional third `model` argument, a per-model advertised-level table, `max` passing through instead of clamping to `xhigh`, `minimal` clamping to `low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/ `reason`) so a downgrade is legible from resolver output rather than silent. Two of these pin defects that exist on next today: - `max` is discarded. Both Codex models whose catalog entries are retrievable (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing. - `minimal` is emitted to a model that refuses it. providerPresets.openai. haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's advertised floor is `low`. GSD is sending a value into a document Codex itself validates. The parity test is what pins that fixed, and it names the offending path/model/effort when it trips. Also corrects tests/model-resolver.test.cjs:351, which asserted renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as though it were a contract. ADR-443 recorded "Codex has no max" as fact and it was true when written; Codex has since added both `max` and `ultra`. That is a stale premise, so the assertion is corrected here rather than worked around. The property test asserts the invariant the whole change exists for: a rendered effort is always a level the target model actually advertises, or an explicit rejection. There is no third outcome. * fix(#3007): resolve Codex effort per model, and make every clamp visible Codex declares supported_reasoning_levels per MODEL and validates against it, so a single per-runtime capability set cannot be right for all of them. GSD's was wrong in both directions at once. `max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped max -> xhigh on that basis; it was accurate when written, and Codex has since added both `max` and `ultra`. Every Codex model whose catalog entry is retrievable advertises `max`, so the clamp was discarding a level the provider supports, silently, on the most-used path. `minimal` stops reaching Codex. No Codex model advertises it -- both retrievable entries floor at `low` -- yet providerPresets.openai.haiku.low paired gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the receiver validates and refuses into a file the receiver reads. Being unconservative in what you send is the half of Postel's rule with no defensible reading, so that preset is corrected and a parity test pins it. `ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum reasoning with automatic task delegation": at ultra, effective_multi_agent_mode returns Proactive and Codex spawns sub-agents on its own initiative, underneath GSD's orchestration rather than inside it (#2167). It is a mode switch, not a reasoning depth, so it is not added to the universal ladder -- which stays provider-agnostic by ADR-443's design -- and it is rejected even for gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered and rejected: that silently discards what the user actually asked for. Clamping is now visible. RenderedEffort carries requested/clamped/reason and resolve-execution surfaces them. The previous table clamped correctly but invisibly, so a user asking for `max` on Codex had no way to find out they were getting `xhigh` -- exactly the failure mode the robustness principle's modern critique warns about, and why "be liberal" has to mean "liberal and loud". Also closes a latent trap found while reviewing the implementation: the clamp-up loop walks the ladder upward, and for a future model advertising `ultra` but not `max` it would have selected `ultra` as the clamp target -- re-entering by the back door the mode the rejection above exists to keep out. A clamp may never produce a value that a direct request for that value would refuse. Unreachable with today's catalog, which is why no test caught it; a test now asserts the invariant directly. Signature stability is preserved: the third `model` argument is optional and the two-argument form still resolves, against the family baseline. That form's BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old answer would have fixed the defect only where a model happened to be threaded through and left it live everywhere else. tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract and is corrected here rather than worked around. * fix(#3007): close every review finding on the Codex effort alignment Two isolated reviewers, correctness and security. Both found the same two blockers, and the per-model work was inert on every surface that matters until this commit. BLOCKER — resolve-execution never passed the model and discarded the clamp. cmdResolveExecution called the two-argument form and emitted only effort_rendered/effort_param/effort_propagation, so the per-model table was unreachable from production code (tests were its only caller) and requested/ clamped/reason were computed and thrown away. Requested outcome 3 names "the effective rendered effort in resolver output" specifically, so the feature was unmet on the exact surface the issue asks for. Now passes the resolved model and emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the existing key convention rather than introducing a nested object. BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed a nested {"effort": ...} sample; the real result is flat and those keys were absent entirely. A reference doc asserting a JSON path a reader can copy is worse than no doc. Corrected against the actual emitted key set. MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex kept minimal in its supported set and still clamped max down to xhigh, so the invocation-time and install-time channels disagreed about the same runtime's capability: --host codex with max emitted xhigh while the generated TOML said max. This is the repo's documented generative-fix-divergence class, so both tables now cross-reference each other and a parity test fails if they ever diverge again. MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null _baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback never fired and every effort rendered as null. And a non-array value made the Set constructor throw at module load — model-catalog.cjs is required across the whole CLI, so one bad JSON value killed every command, not just codex effort. Guarded on size and filtered to array values; both degrade to the hardcoded baseline. MAJOR — value widened to a nullable string with two consumers left behind. runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a null effort key in generated frontmatter); install-effort-resolver still declared a non-nullable return, a structural lie that silently defeated null checking. Both corrected, both omitting the key on null — the same posture as 'inherit', where omission means "follow the host default". MAJOR — the per-model table is inert today, and the docs now say so. All three shipped models advertise the same usable range and ultra (sol's only differentiator) is rejected for every model, so no observable output differs by model. The table stays because Codex declares capability per model and the sets are free to diverge — a single per-runtime assumption is precisely what went stale and produced this issue — but overselling it as a visible per-model feature would have been the same class of error as the doc blocker above. Tests: three passed under a full revert and are strengthened rather than deleted, since each guards a real contract (#3533's inherit rule, the undeclared-host rule, off-ladder handling) — they now also assert the clamp-visibility fields, which only exist after this change. The fast-check property is kept for its shrinking, and a deterministic nested loop over the full cross-product now sits beside it so coverage is exhaustive rather than sampled. Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg form and would have written a literal null reasoning effort on the ultra path; CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT. The installer defect was found by the co-change gate, not by a reviewer — install.js is a historical co-change partner of model-catalog.cts that this diff had not touched. * test(#3007): correct assertions that pinned Codex's stale effort premise Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the shipped commit. Every one is a stale pin, not a defect: each was probed against the built module before its expectation was changed, and none failed for a reason other than this premise correction. Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale by a production change must not ride inside another commit, because the release-sdk hotfix cherry-pick filter routes by subject prefix and a correction buried under the wrong prefix ships a half-state (v1.42.3, #3621). The most valuable one was tests/model-resolver.test.cjs's cross-provider validity invariant, which hardcoded the Codex enum as `minimal|low|medium|high|xhigh` and failed with "real API would 400". That message is now false in both directions: Codex accepts `max`, and rejects `minimal`, which no model advertises. The enum is corrected to `low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the "would the real API refuse this" check worth having, and it was right to fail here. It simply carried the stale fact in its own fixture. Test NAMES were corrected alongside their assertions wherever the name asserted the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal passthrough". A renamed test that still claims the old thing is worse than a failing one, and a green test whose name states a falsehood is how the next reader inherits the wrong premise. Both channels are covered: install-time (renderEffortForRuntime, and the generated .toml in install-runtime-artifacts) and invocation-time argv (effort-surface-axis). They were deliberately brought into agreement in this change, so their assertions had to move together. Each site carries a #3007 comment recording that Codex gained max/ultra and that capability is declared per model, so a future reader can tell this was a deliberate premise correction rather than a test bent to fit an implementation. * test(#3007): separate the effort-precedence case from the clamp case The previous stale-assertion pass over-corrected one test. It saw `effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`, assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to "max passes through" and changed the expectation to `max`. The remote runner disagreed. Reproduced against the real CLI: with that config and `gsd-planner`, the resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`, `effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE — gsd-planner is heavy/opus tier and its routing-tier default outranks `effort.default` — so `max` never reaches the renderer at all. The test says nothing about clamping and never did; it only looked like a clamp pin because both mechanisms happened to yield the same string. Restored to `xhigh` and renamed to say what it actually tests. It now also asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is what makes it impossible to mistake for a clamp pin again: those two fields prove the value is what the resolver produced rather than something the renderer downgraded. Before #3007 there was no way to tell the two apart from the output — which is precisely why the previous pass could not tell them apart either. Added the test that was actually missing: `effort.agent_overrides`, which outranks the tier default, so the requested level genuinely reaches the renderer and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified by probe before asserting. One test now pins the precedence rule and the other pins the #3007 behavior, and neither can be read as the other. That the clamp-visibility fields are what resolved this is a small argument for having added them. * chore(#3007): backfill changeset pr number to 3765 * test(#3007): put model-catalog under the mutation gate The Stryker shard showed as `skipping` on this PR despite the diff rewriting model-catalog's effort logic. That was legitimate, not a detection bug: `model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the whole module — including everything #3007 touches — sat entirely outside mutation scoring with has_work "false". Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no temp dirs. That shape is not stylistic — it is the #2790 precedent this file already documents. Stryker's command runner treats a whole `node --test <file>` invocation as ONE test costing whatever its slowest case costs, and re-runs it per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses runGsdTools throughout) would reproduce exactly the 15-minute shard-cap cancellation #2790 hit. The integration file is unaffected and keeps running in full in the normal test job. Coverage spans the module rather than only the diff, because the score is measured over the whole file: effort rendering across every model and ladder level in both channels, the prototype-chain host guard, the exported enums and maps, isAnthropicFlavoredModel's provider namespacings, the profile projections, nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are worth naming — every uncovered exported function is score given away, and mergeEffortTierDefaults turned out to have a genuinely interesting contract (#3531: a partial override merges over the built-ins rather than replacing them, and isValid gates the VALUE, not the tier name, so an unknown tier key is still merged in). Every expectation was probed against the built module before being asserted. minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment. Floors in this repo are measured, not chosen — the existing entries sit at 94, 75 and 56 — and they can only be measured in CI, because mutation shards run `node --test`, which is hard-blocked locally. The first CI run on this branch reports the real number and the floor gets ratcheted to it before merge. A placeholder of 1 reaching `next` would make the gate decorative: it would pass whether or not a single mutant is ever killed. Note the target is "never regress from measured", not a fixed 80 — planning-inspect sits at 56 and is documented as an accepted ratchet candidate. * test(#3007): bootstrap model-catalog's mutation floor legally The placeholder floor was structurally illegal and the remote run said so. tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module must carry a matching RATCHET_BASELINE entry in the same diff, minScore must EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three. That is the ratchet working exactly as intended — a floor nobody can satisfy accidentally is the point of it. Bootstrapped at 50 in both places. Fifty is not a measured score and the comment says so plainly: it is the minimum the guard permits, and it coincides with Stryker's own configured `break` threshold, so it is the lowest legal starting point for a module that has never been measured. It still must be ratcheted to floor(measured) - 1 before this PR merges. Also corrected a real defect in the file's own instructions. "HOW TO UPDATE" step 1 read "Run the per-module Stryker shard locally" — which cannot be done here, and which the same file contradicts eighty lines further down, where the #2790 scores are recorded as "not a local run; mutation shards run `node --test`, hard-blocked in this repo's local environment". stryker.config.mjs confirms the command runner invokes `node --test` once per mutant, and .claude/hooks/block-local-node-test.sh denies exactly that. So the documented first step sends the next contributor at a wall. Rewritten to describe the path that works — push, read the measured score off the CI shard, then set the floor and its baseline together in one diff — and to say why local measurement is not available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched. The two-step is inherent to the environment rather than a shortcut: a floor cannot be measured before the first CI run exists, and the guard rightly refuses to accept an unmeasured one below its minimum. * test(#3007): ratchet model-catalog's mutation floor to its measured score The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this file's own rule, minScore = floor(measured) - 1, which is the same arithmetic every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94. Both halves moved together, because the ratchet guard asserts minScore equals its RATCHET_BASELINE entry and would reject them drifting apart. The spawn-free unit surface is vindicated by the clock: 57 seconds, against a 15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing the shard at tests/model-resolver.test.cjs — #2790 recorded shards being CANCELLED at that cap when they targeted a runGsdTools-heavy integration file. The registry comment is rewritten rather than deleted. It previously warned that the floor was provisional and must not ship that way; leaving that text next to a measured floor would make the file lie in the other direction. It now records the measurement the way the sibling entries do, including that 59.62 sits below TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 — comfortably clear of its own floor with real room to grow. Raise it as the tests improve; never lower it. Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is an honest floor for a module that had NO mutation coverage at all an hour ago, and it is now pinned so it cannot silently regress. --------- Co-authored-by: sim <sim@local> |
||
|
|
8526bd46f8 |
enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes The guard matched only `resolve-execution ... --effort\s`, but the CLI also accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written with the equals form was a live invocation-override caller the guard passed silently, along with `--effort` at end-of-input. Lift the matcher to a shared predicate and assert it directly against every shape the CLI accepts, plus the decoys it must not fire on (--effortless, a bare --effort with no resolve-execution, item 6's --attempt caller, a call and flag split across lines). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator invocation override needed a caller in shipped orchestration. Per the maintainer's ruling, take unblock path (b) for item 1 only -- record that the override is an operator-facing CLI surface, deliberately not driven by shipped orchestration, and ratify. The ADR's own path (b) wording is not adopted verbatim: it says the scope is limited to static install-time propagation, which is false on both counts -- item 6 has a live caller (#2296) and #2481 delivered a live invocation-time argv channel. Only one precedence step is narrowed. No consumer was invented to clear the gate: nobody has asked for a per-run effort override, and #2475's actual complaint is already closed by the cascade-to-argv path. The amendment states explicitly that --effort remains supported and is not deprecated, so the scoping is not read as dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run` There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run and gsd-mcp-server. Both sites presented a command that cannot run as written. One is in this branch's own new ADR-443 amendment; the other is a pre-existing error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than deferred. No translated copy carries either line, so no i18n drift is created. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2475): record the guard/CLI divergence risk on the item-1 matcher The predicate independently models gsd-tools.cjs's argument parser rather than sharing a constant with it, so a third --effort spelling would leave the guard reporting green while ADR-443's ratifying invariant silently stopped holding. Name that risk where the next editor will meet it. Also restores the bounded-prose rationale that was attached to the eslint directive removed in 39793079c -- the directive went unused once the regex moved to a const, but the reasoning it carried is still worth having. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9e4f0e99ad |
fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash
bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.
Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.
The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.
Refs #3631
* fix(3631): exclude derived bytecode caches from the consent digest
RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.
collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.
Three properties were preserved deliberately, each pinned by a test:
- The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
first would have turned the exclusion into a way to smuggle a symlink past the check;
a symlink named x.pyc still throws.
- Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
caps guard the WALK; the digest answers a different question, and exclusion must not
become an unbounded-bytes hole.
- An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
directory marker was the sharper half of this bug: an empty __pycache__ flipped the
hash before any .pyc existed, so a *.pyc-only filter would have left it live.
The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.
Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.
Fixes #3631
* fix(3631): narrow the digest exclusion after two isolated security reviews
The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.
HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.
FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.
Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.
Narrowed to what is actually defensible:
- a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
the walk still recurses and hashes every non-excluded child.
- .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
- a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
- declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
basename are now rejected in both validator copies — a file named .pyc can contain
perfectly valid JavaScript, so the exclusion must not be reachable from a declared
surface.
Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.
Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.
Refs #3631
* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind
Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.
HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.
The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.
The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.
Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.
Changeset rewritten: it still described the rejected wholesale-exclusion semantics.
Refs #3631
* chore(3631): backfill changeset PR number (#3650)
---------
Co-authored-by: sim <sim@local>
|
||
|
|
2972da4c9d |
enhance(#3619): ratchet the platform seam with local/no-private-binary-resolution (epic #3411 Phase 3) (#3636)
* chore(#3619): ratchet the platform seam with local/no-private-binary-resolution
Epic #3411 Phase 3, the ratchet. Scope revised with maintainer approval and
recorded on the issue: the epic's literal ask was a rule rejecting a bare-name
spawn outside the seam. Surveyed at
|
||
|
|
924f649f87 |
docs(#3625): record the spawn-library evaluation as ADR-3625 (#3632)
Spike outcome for #3625: evaluate cross-spawn / nano-spawn / execa against the hand-rolled Windows binary resolution and cmd.exe mediation that epic #3411 Phase 1 (PR #3621) is landing in the platform seam. Verdict: stay hand-rolled, with revisit-if conditions recorded so the call is not re-litigated in a future PR. Evidence, per the issue's "Done when" list: - Sync/async verdict per candidate. nano-spawn is async-only — settled by its own README, which lists "synchronous execution" among the features execa has and it does not. cross-spawn exposes `.sync`. execa exposes execaSync, which its own docs discourage. - CVE-2024-27980 escaping verdict per candidate. None uses shell:true. cross-spawn independently arrives at the SAME mechanism the seam uses: cmd.exe /d /s /c with a pre-escaped line and windowsVerbatimArguments. That validates the seam's approach rather than superseding it. - Maturity axes scored with measured data (registry metadata 2026-08-18, transitive footprint measured by install). Two premises in the issue did not survive measurement, both recorded: - The CRITICAL 167-symbol/53-file blast radius is a `direction:both` measurement, inflated by downstream callees. A call-shape change ripples to CALLERS: upstream at depth 15 is 14 symbols / 7 files, MEDIUM, and it terminates at depth 4. The decision does not rest on the CRITICAL figure. - The cited vendoring precedent path does not exist; the real one is gsd-core/bin/lib/vendor/re2js.cjs. Decisive against the only structurally-eligible candidate (cross-spawn): it resolves process.cwd() first on Windows even when an explicit PATH is supplied, calls process.chdir() during resolution, and keys escape depth on a node_modules/.bin/*.cmd path regex of the same shape #3411 was filed to delete. Adoption would also break execTool's observable not-found contract across 53 dependent files. Doc-only: docs/adr/** plus a root-level CONTEXT.md pointer. The CONTEXT.md addition is a new line rather than an edit to the seam's glossary paragraph, which PR #3621 rewrites wholesale — same-line edits would conflict on merge. ADR index regenerated via scripts/gen-adr-index.cjs --write. lint:ci exit 0; lint-docs-required and changeset/lint both run with GITHUB_BASE_REF=next and report ok_no_user_facing_changes. Closes #3625 Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cc3fd4548d |
docs(#2866): reconcile ADR-3574 against the shipped environment (#3622)
The ADR was written mid-epic and describes a tree that no longer exists. There are now two choreographies, not three: phase 6 deleted bin/install.js's agent-staging loop and its _DESCRIPTOR_AGENTS_RUNTIMES gate outright, so the refusal in decision 1 governs a two-way divergence. The ADR's own revisit condition has been met. It said to reopen the unification question when the applySurface descriptor-agents migration completed. It has. Revisited on evidence: the shapes did not converge on the axis that mattered, because phases 6 and 7 touched agents and exports, not the prune. applySurface still prunes by allow-list so it cannot delete a user file by construction; installRuntimeArtifacts still wipes and restores a snapshot. Decision 1 stands, for a narrower and better reason than when it was written. Records delivery status per decision, including that decision 3 was already satisfied when measured and decision 4 was the hardest part rather than the independent one the ADR predicted. Notes that #2875's AC1 is now doubly stale - deliberately unmet, and naming three call sites where two remain - so anyone reconciling the tracker treats this ADR as governing rather than the criterion as outstanding. Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
682eaae3f0 |
enh(#2876): retire the dead and pass-through exports from bin/install.js (#3615)
* enh(#2876): retire the dead and pass-through exports from bin/install.js The installer exported 197 names and had zero production consumers - every non-test require of it repo-wide sits inside a comment. Its interface was shaped by test access, not by callers. Removes 9 dead exports and 61 pass-throughs, repointing their tests onto the extracted modules' own interfaces. 197 down to 127. Every count in the issue was wrong: 197 exports not 188, 9 dead not 12, 61 pass-throughs not 49, 44 test files not 42 - and the audit itself then missed 7 more consumer files. restoreUserArtifacts was on the dead list but ceased to exist in phase 6, and two _GSD_EFFORT_MANIFEST_* names listed as dead are now genuinely asserted, so acting on that list would have deleted live exports. 7 of the 9 dead names collide with an independent declaration that install.js delegates TO. Each removal was justified by which declaration a reference resolves to, never by whether the name appears somewhere. Coverage parity was the gate rather than test greenness: per-file counts were captured before any edit and diffed after. 44 of 45 files are byte-identical; the single delta is one added assertion, not a loss. The sweep for scattered require sites found two forms static grep misses - require(VARIABLE) and multi-line require() - plus tests asserting that install.js re-exports the SAME object, which now assert retirement instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2876): close review findings — restore the duplicate-body guard, sweep orphaned code Both review engines found real defects in the first cut. The DEFECT.GENERATIVE-FIX single-owner guard from #1511 had been repointed from a reference-identity check to install.X === undefined. Those are not equivalent: the guard exists to catch a duplicate function body reintroduced into install.js, and the replacement passes cleanly if that duplicate is used internally and never exported. It now walks bin/install.js's real top-level bindings, so it catches a duplicate under either shape, exported or not - strictly stronger than the check it replaced. Proved by injecting a duplicate and watching it go red. That weakening survived the coverage-parity gate because the assertion count never moved. The gate compares counts, so an assertion that changes meaning rather than number is invisible to it. Removing the exports had orphaned their wrapper bodies: 14 dead wrappers, 9 consts and 9 destructure entries, several pre-existing and found by the same sweep. Dead code left in the file this phase exists to shrink. Three more comments claimed re-exports this phase removed, and tests were reading Cursor and Windsurf hook constants from install.js's local copy while calling functions from the hooks surface - equal today, with nothing holding them equal. The local consts now reference the owning module. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2876): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8a56595700 |
docs(#3574): record ADR-3574 install materialization primitives (#3575)
* docs(#3574): record ADR-3574 install materialization primitives Epic #2866 phase 6 was scoped on the premise that materializing a layout is implemented three times and should become one module. Measured against the tree, the premise does not hold: the three sites overlap in shape and diverge in mechanism. applySurface prunes by allow-list precisely so it structurally cannot delete a user's files. installRuntimeArtifacts wipes a prefix-scoped set and restores a snapshot. A single writer has to pick one, and picking either trades a working guarantee for a different one. So the ADR declines phase 6's first acceptance criterion and says why, because the next reader who notices three similar loops should find this file rather than rediscover the conflict. What is extracted instead is the genuinely shared part: durable user-artifact staging for #1874-F19, reusing the migration primitive that copies strictly before delete and never dereferences a symlink, plus the retired-kind prune both callers already share. The agents bypass closes on its own terms. Two of the issue's premises were also stale: ten runtimes have already migrated off the inline agent dispatch, and the duplication comment's deliberate-until condition is partly met. Closes #3574 * chore(#3574): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |