From 7c116b1c17f4831f446cf5bff70ccd4b66712c4d Mon Sep 17 00:00:00 2001 From: 0xdhx Date: Tue, 1 Sep 2026 13:36:57 -0500 Subject: [PATCH] fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#3697): warn when the Requirements line under-selects REQ-IDs `cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is correct for the canonical comma list the template ships, and silently wrong for every other form: `RANGE-01 … RANGE-05` -> the two ENDPOINTS only; the interior IDs are never considered, yet `requirements_updated` reports true with zero warnings `RANGE-01…05` -> ZERO IDs; the whole line is inert The silence is structural: the only cross-check, `ghostReqIds`, is itself `citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it by construction — and to `traceabilityWriteMisses` and `requirements_updated` with it. Warn on both paths. This does not add range support: the selected set is unchanged, so no existing ledger write changes. The trigger is ID-SHAPED EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a range operator joining two IDs — with parenthetical citations and HTML comments stripped before the scan, so the #2334/#2339 over-warning on `None`, on the shipped `` template comment, and on annotated lines cannot return. Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs (10 cases: 4 defect, 2 canonical controls, 4 negative-space controls). Fixes #3697 * fix(#3697): rework under-selection detection onto tokens, not a free-text scan Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb. That review refuted 5 of 9 claims; three were false-positive classes in exactly the category #2334/#2339 had to REMOVE: `RANGE-01, RANGE-02 - 3 points` the bare-hyphen alternative read `RANGE-02 - 3` as a range `REQ-01, REQ-02 — locked per ADR-7.` the trailing period kept `ADR-7.` out of the anchored filter, so the unanchored substring scan reported it as unparsed `REQ-01, REQ-02 (see (ADR-7), then ADR-8)` nested parens left `ADR-8)` behind Replaces the free-text substring scan + loose range regex with three narrow, token-based rules (R1 range-shaped token, R2 pure range operator flanked by two selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks nothing — it now says "selected". Side effect: the two false NEGATIVES the same review found are now covered — `RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`. NOT YET DONE (see the handoff prompt): regression tests for the four false positives, the two new true positives, and the #3697-4 tightening the review's CLAIM 8 asked for (it currently filters on the warning's phrasing rather than asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests green, and a 20-case standalone harness covering every case above. * test(#3697): pin the v2 token-detector boundary end-to-end Six new cases + two hardenings for the review findings against v1: - #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) — the operator set's `to|thru|through` arm was previously untested. - #3697-5 (new): a tight range hidden inside balanced parentheses (`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token, and ticks exactly RANGE-01 — the paren shave must not hide it. - #3697-4 gains the four false-positive classes a free-text detector produced: numeric estimate (`- 3 points`), date annotation, em-dash citation with trailing period (`— locked per ADR-7.`), and nested parenthetical citations. - #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty, not that one phrase is absent — a re-worded over-warning cannot pass. Negative control: against the merge-base with its lib rebuilt, all 6 defect tests fail and all 10 controls pass. * fix(#3697): close round-2 review findings — annotation false positives Round 2 of the adversarial review (against 822a72a04) refuted five claims; this closes the false-positive class and the cheap misses: - R1's bare-hyphen arm now demands a full ID on BOTH sides (`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation (`FY-2026-08`) and a sub-numbered ID, and warning on those is the expensive class. Tight hyphen shorthand with a live selection is the disclosed false negative; at zero selection R3 still catches it. - R2 requires the endpoint pair to imply an INTERIOR (same prefix, gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop nothing, so an annotation hyphen between adjacent IDs stays silent. - R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty line citing its rationale, not unparsed residue. - Token shave: quotes/backticks now shaved from alphanumeric tokens (`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a bracket-only shave so `(..)` surfaces its operator. - 256-char token cap bounds the quadratic unanchored substring test. - Warning text mentions range expansion only when a range rule fired. Tests: 6 new cases (22 total). Negative control against the merge-base: 8 defect tests fail, 14 controls pass. * fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers Round 3 of the adversarial review (against 2eb92dd0e) refuted four claims; this closes them: - Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the tokenizer before R1's `\s*` can see them and under-selected silently. A glued-fragment rule warns when an operator is glued to a full ID with an ID-shaped neighbour on the open side and the endpoint pair implies an interior. - Cross-prefix pairs around a separator no longer read as ranges: `REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and real ranges are same-prefix by nature. `impliesInterior` now returns false on prefix mismatch and computes the gap with BigInt (parseInt lost precision past 2^53). - The token shave now removes markdown emphasis markers and curly quotes, so `**None** (per ADR-7)` reaches the placeholder gate and `**RANGE-02..RANGE-05**` reaches R1. - The unanchored-substring cap rises to 2048 (a markdown-link range with a long URL cleared 256); the anchored range regexes scan linearly and drop their cap. Tests: 6 new cases (28 total; 377/377 file-wide). Negative control against the merge-base: 11 defect tests fail, 17 controls pass. * fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule `TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment rule read it as `to` + `REQ-05`, warning on the canonical two-ID list `REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only (`..`+, ellipsis, dashes): a word operator glued to an ID is an ID, not a range spelling. Tests: word-operator-prefixed ID control (misparse channel silent; the fixture's ghost-ID warning legitimately fires, so the whole-channel assertion stays with the registered controls) and an underscore-wrapped tight-range defect case. 30 targeted cases; 379/379 file-wide; negative control: 12 defect tests fail on the merge-base, 18 controls pass. * fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording - The glued-fragment TRAILING arm takes the word operators back: an ID must end in digits, so `REQ-01through` can never be an ID — the round-4 TOREQ collision was leading-arm-only, and symbol-only on both arms lost the `REQ-01through REQ-05` typo class. - A trailing run of 2+ dots survives the punctuation shave: `REQ-01..` is a glued range operator, not sentence punctuation, and the shave was silently eating the `REQ-01.. REQ-05` form. - The warning now says the line "could not be parsed as" a comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list — the selector just cannot parse decorated tokens — and a warning that misstates the input teaches readers to distrust it. Tests: two new trailing-glue defect cases (32 targeted; 381/381 file-wide). Negative control: 14 defect tests fail on the merge-base, 18 controls pass. * test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset CONTRIBUTING bans try/finally inside test bodies (it masks failures); the approved shape is `t.after(() => cleanup(tmpDir))`. All seven converted tests are this PR's own additions; the file's pre-existing instances are untouched. * chore(#3697): add changeset fragment for the Requirements-line under-selection warning changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per the house format, with the (#3697) backlink. * refactor(#3697): extract the Requirements-line detector to a testable surface Round-3 review Blocker 1 requires a fast-check property test over this detector (`RULESET.TESTS.property-based-testing`: modules implementing parsing contracts must include at least one), and Blocker 2 requires limit-1/limit/limit+1 fixtures on its 2048-char token cap (`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test reaches it by spawning the CLI, and a property test cannot pay a subprocess per generated case. So the selector and the three detection rules move to module scope as `analyzeRequirementsLine` (pure, exported) plus `formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and the pre-round suite passes against it unmodified (403/403). Two things the move makes explicit rather than incidental. The selector and the detector tokenize the SAME line DIFFERENTLY — the selector strips only `[` and `]`, the detector also shaves quotes, emphasis and trailing sentence punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected while `ADR-7` is still nameable in a warning. They now sit adjacent with the reason written down, so they cannot drift apart silently. And the stale citations in the moved comment are corrected. It pointed at src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is `gsd-core/templates/roadmap.md:32`. Both are now anchored by content. * fix(#3697): stop the warning claiming a misparse that did not happen Round-3 review Major 3 and Minor 4. Both are the same defect: the warning asserted more than the evidence supported. MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05 deferred` warned "could not be parsed ... Range forms are not expanded; rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05, i.e. every ID written on the line. Nothing was dropped, and the pinned control only stayed silent because its pair was ADJACENT (gap == 1), so the control was passing by accident of the fixture rather than by the rule. The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a range and as an annotation separator are textually identical, and no token-level rule separates them; staying quiet re-opens the exact silent under-selection #3697 is about. Deciding the ambiguity by assertion in either direction is wrong. So it is DISCLOSED: the warning now has two channels, chosen by whether any ID-shaped token was actually left unselected (`droppedIdShaped`). * something was dropped (tight range, glued fragment, inert residue) -> "could not be parsed as a comma-separated REQ-ID list", as before. * nothing was dropped (only the spaced-operator rule fired) -> "contains what reads as a range between two cited REQ-IDs", stating both readings and saying explicitly that an annotation separator means the line is already correct. This retires the "could not be parsed" wording for the four #3697-1 spaced cases too, and that is a deliberate expectation change rather than a fix counted twice: those lines never failed to parse either. They still warn, still name the selected IDs, and still assert the endpoint-only marking is unchanged; #3697-1 now also asserts the misparse channel stays SILENT. MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a citation as requirement content it had failed to read. The trigger is correct and stays: #3697's acceptance criterion asks for a warning "when it selects zero IDs from a line that is non-empty and is not the `TBD` placeholder", and inferring placeholder-ness from arbitrary prose is the free-text heuristic this detector exists to avoid. What was wrong is the wording, so the non-range arm now says "ID-shaped text that was not selected" and names the escape the author actually has (`TBD` / `None`). Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse, must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A (tracked in ADR-12)` — must warn, must not say "Unparsed text", must not diagnose a range, must name the placeholder escape). Reversion control: reverting the ambiguous channel fails #3697-9 (3 named tests); reverting the R3 wording fails #3697-10 (2 named tests). * fix(#3697): cap every token predicate, complete the dash set, cover the boundary Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three are about the detector's own predicates, so they land together. BLOCKER 2 — the 2048-char budget had no boundary coverage. `RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit / limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 / 2048 / 2049 against BOTH predicate families the cap guards, and each asserts its fixture's exact length before asserting behaviour, so a mis-built fixture fails loudly rather than passing at the wrong size. Clause (d) of that rule — an input pushed within reserve-distance of the limit — has no referent here: this is a hard cap with no reserve constant beside it, and the test comment says so rather than leaving the omission to be re-derived. NIT 6 — the cap guarded only the unanchored ID-substring regex. The anchored range regexes were left uncapped, justified by a comment asserting they scan linearly. The finding is right that this is informational (they are anchored; the input is a local ROADMAP.md), but an asserted property is cheaper to enforce than to defend, so all three predicates now share one `short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must now decline to classify. SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this code fixes at author time over a domain that grows without it, so the round owes a census of what the enumeration reaches. reached: `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure dash, U+2015 horizontal bar, U+2212 minus sign consequence: a range spelled with any of those is SILENTLY under-selected — #3697's own defect, in the code that exists to fix it Those five close. They are the same operator at a different codepoint and carry none of the ASCII hyphen's collision risk, because they are not the REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is held to is untouched, and #3697-12 pins that. Still NOT reached, declined with reason rather than left unstated: `→`, `~`, `..=`, `..<`, `until`, and `up to` (two tokens, so never one operator token). Each is a symbol or word with an independent non-range use between two REQ-IDs — the over-warning class #2334 cost three rounds. Reversion control: reverting the uniform cap fails #3697-B2 (limit+1); reverting the dash set fails #3697-11 (5 named tests). * test(#3697): add the fast-check property coverage the parser rule requires Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md) requires modules implementing parsing contracts to carry at least one fast-check property test asserting a domain invariant, and the round-2 diff had zero occurrences of `fc.` across its +311 test lines. Five properties, 1,900 generated cases: P1 soundness of silence (boundary containment) — for ANY canonical comma list of well-formed REQ-IDs, the selected set EQUALS the written set and nothing warns. This is the #2334 over-warning invariant and the #3697 under-warning invariant asserted as one statement, over generated IDs rather than hand-picked ones. It generalises #3697-4b: a prefix beginning with a word operator (`TORANGE-05`) is an ID, and P1 covers that class rather than the single example. P2 completeness — a same-prefix pair with an interior between them, separated by any of the nine spaced operators, ALWAYS warns. P3 the #2334 invariant — an ADJACENT pair around a separator can drop nothing, so it stays silent however it is annotated. P4 totality + idempotency — total over arbitrary strings, deterministic, and the formatter agrees with the analysis on whether there is anything to say (a warn with no text, or text with no warn, is a channel that can go silent or noisy on its own). P5 containment — every selected ID is ID-shaped and appears verbatim in the input. Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and P5 hold against the round-2 code as well as this one — they are regression guards, not bug-finders, and their value is that the invariants are now stated and generatively checked rather than implied by examples. P2 is the one that would have failed before the dash enumeration was completed. fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from `fc.array(...).map(join)` with the alphabet pinned to the selector's own `[A-Z0-9]` class. These live in tests/phase.test.cjs rather than a new `phase.property.test.cjs`: `lint-test-file-count` caps a production module at 2 test files and phase.cts is already at its allowlisted entry, so a new file would trade one gate for another. * docs(#3697): document the ROADMAP Requirements-line grammar Round-3 review Minor 5 — the change adds net-new user-visible warning output for a grammar constraint documented nowhere under docs/. `type: Fixed` is docs-exempt so this does not block, but a warning about a rule the reader cannot look up is not actionable, and that is worth fixing whether or not a gate demands it. Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the existing SUMMARY artifact-check advisory it is a sibling of: the supported comma-list form, why ranges are deliberately not expanded, that `TBD` and `None` are the entire placeholder vocabulary, and what each of the two warning voices means — including that the range/annotation one may be reporting a line that is already correct. Existing file rather than a new one, deliberately: docs/ carries generated indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity or index gate this change has no reason to touch. * fix(#3697): rule-scope the warning-channel discriminator Self-found at the round's pre-push review, against the Major 3 fix two commits back. That fix chose the channel from a LINE-GLOBAL question — "was any ID-shaped token left unselected?" — while the rules that produce the warning are not line-global. The two disagree as soon as the line carries an ID-shaped token no rule fired on: `RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)` `(ADR-7)` survives the selector's bracket strip, so the global test called it a drop and sent the line to the assertive channel — putting the false "could not be parsed ... rewrite the line" claim back on a correct line. That is review finding Major 3 returning through a side door, and it directly contradicts #3697-4, which pins a parenthetical citation as NOT unparsed residue. The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the ambiguous channel requires that R2 fired, that no other rule did, and that every endpoint R2 fired on was actually selected. The last conjunct is not redundant — the detector shaves brackets and the selector does not, so R2 can fire on a `(RANGE-02)` that was never selected, and that IS a drop: `RANGE-01 (RANGE-02) — RANGE-05` -> assertive, correctly `droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the channel can ask about the endpoints the rule fired on) and the `rangeReadingOnly` verdict. Reversion control: against the line-global rule, #3697-9b fails. #3697-9c passes under both rules — there the dropped token IS the R2 endpoint, so the two agree; it is a regression guard, not a bug-finder, and is recorded as such rather than counted as a second control. * fix(#3697): hold every dash to the strict range shape, not just ASCII Self-found at the round's pre-push review, and it CORRECTS a claim made two commits back. That commit widened the range-operator set by five Unicode dashes and asserted they "carry none of the ASCII hyphen's collision risk, because they are not the REQ-ID separator". That reasoning was wrong. The collision is a property of the SHAPE — `PREFIX-\d+ \d+` is also a date (`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not care which dash sits in the operator slot, because the ID's own separator is still ASCII either side of it. Measured: RANGE-01 (target FY-2026-08) silent <- pinned by #3697-4 RANGE-01 (target FY-2026‐08) WARNED <- same line, U+2010 So the widening reintroduced the #2334 over-warning class on a date annotation. It also exposed that the inconsistency PREDATES this PR: U+2013 and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash forms of that same date annotation warned before round 3 ever ran. One rule for every dash: a tight range spelled with any of the eight must carry a FULL ID on both sides, exactly as the bare hyphen already had to. `..`, `…` and the word operators stay loose — no date or sub-number reading exists between two numbers, so the strict shape would cost them coverage for nothing. The cost is a false negative, and it is one the design already accepts: `RANGE-01, RANGE-02-05` is silent today, deliberately, and now `RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so R3 catches it. Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight dashes), #3697-13b (full-ID tight range still warns for all eight), #3697-13c (loose operators keep their numeric endpoint), #3697-13d (the accepted false negative is symmetric, and the bare zero-selection line still warns). * fix(#3697): close the round's own pre-push review findings An adversarial cross-AI review of this round refuted 4 of its 10 claims. All four were real. Every fix below is to code THIS round introduced. 1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1). `REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told the author "the line is already correct and nothing needs to change" — while `(REQ-02)` had been dropped by the selector, which does not strip parentheses. The channel choice is still right, and deliberately so: `(ADR-7)` and `(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts the false "could not be parsed" claim back on a line carrying a citation — the misroute fixed two commits ago. No rule can adjudicate this; the author can. So the voice stops asserting the line is correct (it now speaks about the SEPARATOR, which is all it has evidence about), and BOTH voices gained a factual clause naming ID-shaped text the selector skipped, with the reason (brackets are not stripped) and no verdict attached. 2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2). A 2049-char range token warned before this round and went silent after it: the "uniform cap" commit bounded the predicate and, with it, the warning. That is #3697's own defect, introduced by the fix for a nit. The cap bounds the WORK, not the warning. An over-cap token carrying `-` is now recorded as unclassified (a linear `includes`, never the unanchored regex the cap exists to keep off it) and gets its own voice: "could not be checked ... the REQ-ID selection on this line is unverified". Unclassified is reported, never treated as clean. 3. THE CAP WAS NOT UNIFORM (review MISSED finding). R2 capped the operator token but not its neighbours, so `<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over both endpoints unbounded. The glued rule had the same hole. Every participant is capped now. 4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5). It drew from a bare `fc.string()`, which over 500 samples produced max length 10 and ZERO inputs containing a REQ-ID — the loop body never executed an assertion. A containment property that never contains anything is a green test measuring nothing. The generator now interleaves real IDs with noise and the property ASSERTS it saw them (>50/500), so it can never silently go vacuous again. The free-form coverage it was actually providing survives, honestly labelled, as #3697-P6. The same finding refuted this round's claim that P2 distinguishes pre-round behaviour: every operator P2 uses was already in the pre-round operator set. P2 is a regression guard, and its comment now says so. Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review MISSED finding) and is corrected with the code. Tests: #3697-9d (soft voice names the skipped ID, never claims the line is correct), #3697-9e (over-cap token reported as unclassified, still warns), #3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the boundary on the PREDICATE's verdict with `warn` asserted true at every length — asserting `warn === false` at limit+1 was itself finding 2. * docs(#3697): describe the third voice and the dash rule Follow-on to the review-findings commit: that commit corrected the docs' claim about the soft voice but left two things the code now does undescribed. - There are THREE voices, not two. The over-cap voice ("could not be checked ... unverified") arrived with the fix for the review's CLAIM 2 and had no entry. - Dash spellings require a full ID on both sides, and `..` / `…` / the word operators do not. That asymmetry is deliberate and load-bearing — `PREFIX-` is date- and sub-number-shaped — so a reader hitting `REQ-01-05` and getting silence has no way to find out why. The accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still reported) is stated rather than left to be discovered. Documentation only; no behaviour change. * fix(#3697): close the continuation review's findings A continuation of the same adversarial reviewer, run against the reworked round, refuted 6 of 7 claims. Four were real defects in this round's own work and are fixed here; the other two are answered rather than changed, below. 1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A). The previous commit's message said both voices gained it. Only the soft return appended it. The assertive voice now carries it too — and, because that voice already names range tokens and inert residue under its own clauses, the note is filtered to what those did not already name. A warning that says the same token twice is one readers learn to skim. 2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A). It read "brackets and parentheses are not stripped". Square brackets ARE stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md, which had inherited the same error. 3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B). `oversizedTokens` filtered on `includes('-')`, which misses an over-cap OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined to classify it once capped, and nothing reported it. That is the exact regression the field was added to close, one input over. Any token past the cap now counts — what it contains is irrelevant when we could not read it. 4. AND THEN OVER-REPORTED ONE (review MISSED finding). With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line is unverified" — a contradiction inside one warning. A token the selector took was examined end to end, so it is excluded. Two findings are answered, not changed: CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and deliberate: this round does not touch what gets MARKED, and the pattern is anchored at both ends with no nested quantifier, so it is linear. The claim that "all predicate paths are capped" was too broad; the DETECTOR's are. CLAIM E — `REQ-01, REQ-0205` is silent for every dash. That is the documented, deliberate cost of holding dashes to the strict shape, and it is symmetric with ASCII, which behaved that way before this PR. The reviewer is right that "without losing a range spelling that should be detected" was too strong; a bare `REQ-0205` still warns. Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a selected over-cap ID is never called unverified). #3697-9f is rebuilt — its first version used the SAME id twice, so R2 could not have fired even uncapped and it proved nothing; it now uses endpoints with a gap and fails when the neighbour cap is removed. Reversion control: all four fixes fail a named test when reverted in isolation (#3697-9h, #3697-9i, #3697-9g, #3697-9f). * fix(#3697): scope the over-cap exemption to what could actually pair A second continuation of the same reviewer, against the reworked round, confirmed the two claims that matter most and refuted three. This closes the one real defect; the other two are answered below. CLAIM J / CLAIM K (one defect, found from both directions). The previous commit exempted EVERY selector-accepted token from `oversizedTokens`, on the reasoning that the selector is uncapped and anchored so it examined the whole token. True of that token's SELECTION — and not the same as "no rule was suppressed by it". Two over-cap valid IDs either side of `..` are both selected, so both were exempted, and R2 is capped: a line that warned before this round went silent. That is the third appearance of one class in this round — the cap suppresses a check, and the suppression is not reported. Each fix for it over-corrected in the opposite direction, which is why the rule is now stated in terms of what was actually suppressed rather than in terms of the token: an over-cap token is exempt only when it was selected AND nothing beside it could have paired with it into a range (no range operator, no glued fragment, no second over-cap token). Everything else is unexaminable and says so. Two findings are answered, not changed: CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01` silent: recognise `PREFIX-ab` only when another SELECTED id on the line shares that prefix. That refutes this round's claim that the strict-dash trade was FORCED, and the claim is withdrawn — it is a design choice. The choice stands for this PR: the conservative rule is what ASCII already did before #3697, adopting a new contextual heuristic unreviewed at the end of a round is how the last three defects in this round were made, and #3697 asks for a warning rather than better range inference. Named here so the alternative is on the record rather than lost. Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not classified at all" and warns. Selection is not bounded; only range detection is. Corrected. Confirmed by the same pass, and worth recording because they are the PR's load-bearing promises: a 20,000-input comparison of the pre-extraction selector against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has changed), and the uncapped selector regex was measured linear from 100k to 800k characters. Tests: #3697-9f now asserts the range case is reported rather than silent, and #3697-9j pins the exemption's scope in both directions. Reversion control: restoring the blanket exemption fails both. * fix(#3697): warn on zero selection, as the acceptance criterion asks `Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`, the `placeholderLed` census comment, and the advice string the command emits to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned, because the citation supplied the ID-shaped residue R3 required, while bare `Deferred` did not. The claim was written into three places and never executed once. This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length === 0` while the raw capture is non-empty and not `TBD`". R3b keys on the SELECTION being empty, never on what the prose means, so it adds no free-text heuristic. It is deliberately not gated on ID-shaped residue the way R3 is, and the negative space is what settles that: all fifteen #2334/#2339 fixtures are held silent by non-zero selection or by `placeholderLed`, and not one of them by the ID-shape gate — measured, not argued. The gate was buying no negative space while costing the acceptance criterion. `tokens.length > 0` keeps an empty line and a comment-only line silent: the tokenizer strips `` before splitting, so the shipped template's own comment cannot reach the rule. Selection behavior is unchanged. This warns; it never invents an ID. Also extracts `warn` to a named const (round 3 review Minor 3) — this commit adds a disjunct to exactly that predicate, and in the return literal a later reordering would be a TDZ ReferenceError rather than a reader-visible error. Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the warning names the TBD/None escape), #3697-14b (five placeholder spellings stay whole-channel silent), #3697-14c (comment-only line stays silent). Fail-first controls: all six #3697-14 cases fail against the pre-fix tree; -14b and -14c pass at both ends, which is correct — they pin silence the widening must preserve. * fix(#3697): name the REQ-ID a glued delimiter dropped `RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with `requirements_updated: true` — #3697's own half-success failure mode, reached by one wrong delimiter, and silent before this rule. It is the issue's AC-1a ("a warning whenever the line contains ID-shaped content that the tokenizer did NOT select") at the shape most likely to be typed by accident. Round 4 review rated this Major rather than Blocker on the ground that the case is indistinguishable from a parenthesised citation, since `(ADR-7)` also shaves down to a bare ID. At the RAW token level it is distinguishable, and that is what makes the rule shippable: `REQ-01;` is shaved of a trailing DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and requires the token to sit outside any parenthetical. Measured before implementing: 0 false positives and 0 false negatives across 21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first cut without the parenthetical test scored 3 false positives — every one of them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that test is the rule's boundary rather than an optimisation. Adds the delimiter census the module did not have. The range-operator domain was already censused; the comma-substitute domain was not. Swept 26 spellings: exactly two produce a silent under-selection, `; ` and `: `. Every other spelling either selects both IDs or selects none and already warns. The review hand-listed the semicolon; the colon is the sibling that sweep found, and it fails identically. `rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing was dropped, and must not speak for a line where something demonstrably was. Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick exactly the unchanged selection), #3697-15b (three citation forms stay whole-channel silent). Fail-first control: all four #3697-15 cases fail against the previous commit's tree; -15b passes at both ends, pinning the boundary the widening must not cross. * fix(#3697): give the Requirements-line warning a stable machine kind The warning's kind existed only in the prose of its message, so every consumer and every test had to regex an English sentence — and rewording a message silently un-asserted the tests that pinned it. Round 4 review Major 3. The repo already had the settled seam for exactly these semantics. `CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a truncated scan, which is precisely this module's third voice; and `WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same reason. ADR-3473 Decision 3 ("failure is a value") points the same way. `formatRequirementsLineWarning` now returns `{ code, message }` instead of a bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a legitimate value, but the success arm is no longer a naked string one field away from the shape the ADR standardises on. The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is a documented `string[]` in `phase complete`'s JSON output, rendered by execute-phase.md's "If has_warnings is true" step, so re-typing its elements would be a breaking output-contract change for a shipped command. The code is emitted as its own additive `requirements_line_warning` field, absent entirely when the line is clean. Vocabulary, exported so tests key on it rather than on string literals: `req-line-misparse`, `req-line-range-reading`, `req-line-unverified`. Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code; message-content assertions stay where the user-visible wording is itself under test. #3697-16 pins the code end-to-end through the CLI's JSON for four line shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean line emits no kind at all, because a field present on every run carries no information. #3697-P4 holds kind-and-message-appear-together and kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added later cannot ship without one. * test(#3697): pin the divergence against the second parser of the same line CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when sharing constants/arrays/parsers between parallel surfaces, add a parity assertion test that fails if they diverge." Round 4 review Major 2. `normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP `**Requirements:**` value — its own docblock says callers "may pass the roadmap value through verbatim" — and diverges on four axes. Measured, not inferred: line phase complete gap-checker RANGE-01..RANGE-05 [] 5 IDs None (per ADR-7) [] ["ADR-7"] (REQ-02) [] ["REQ-02"] REQ-01a [] ["REQ-01a"] REQ-01, REQ-02 both both This pins the divergence rather than removing it, which is the review's second option and the correct one here: unifying the two would change what `phase complete` MARKS, and "the ledger-writing set is byte-identical to base" is the one invariant this PR holds fixed. Every axis is now asserted in BOTH directions, so drift on either side fails here instead of widening silently. The range axis is a DELIBERATE disagreement and is labelled as such — #3697 declines range expansion in terms ("I am not asking for range syntax to be supported") while gap analysis adopted it under #1269. The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a declared-empty line to `phase complete`, which reads the lead token, and a one-requirement line to gap-checker, which strips parentheses first so the citation survives its ID-shape filter. That is a citation being reported as a requirement. #3697-17b states the cost concretely: one line, five requirements in scope to gap analysis and zero to phase complete. This PR is what makes that contradiction visible, by finally giving the silent side a voice. * fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap Two round 4 review minors, both about a warning saying something it cannot support. MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so `FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`. `REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no warning at all — but whenever some OTHER rule fired on a line that also carried a date annotation, the rider told the author to "check whether any of it is a requirement" about a date. Not a false warning, since the line was warning anyway; false CONTENT, in the #2334 voice, through the side door. Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped` stays a faithful record of what the selector skipped — it is documented as a fact that never routes — while the user-facing clause declines to assert requirement-ness about a shape the design already ruled unadjudicable. #3697-18b is the other half, so the filter cannot become a silencer: a genuinely dropped REQ-ID is still named. MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent, so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel would have passed. Whole-channel silence is not available on that fixture (the pre-existing ghost-ID warning legitimately fires on the unregistered `TORANGE-05`), so the precise assertion is that no Requirements-line warning of ANY kind was emitted. The machine code added earlier in this round is what makes that statable; before it, "both channels" could only have meant a second prose regex. * docs(#3697): record the Requirements-line seam in CONTEXT.md CLAUDE.md names the CONTEXT.md glossary as a PR gate, and `get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25 co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a named seam, three warning kinds, a bound, a rule taxonomy and a deliberate cross-parser divergence, and recorded none of it. Round 4 review Major 4. The precedent is explicit rather than inferred: the directly analogous seam is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound and its boundary obligation — and that entry is the one this module's third voice was modelled on. Eight predicates, in the machine-oriented section beside it: .module the two exported functions and the code vocabulary .selector-identity citedReqIds is byte-identical to base and is the only thing reaching the ledger — a change to what phase.complete MARKS is outside this contract .rules R1 / R2 / R2' / R3 / R3b / R4 / over-cap .kinds the three codes, and why they ride beside warnings[] rather than inside it .cap 2048, neighbours included, and the boundary rule .placeholder the gate that actually holds the negative space .census-domains both open domains with their NOT-reached members .gap-checker-divergence the four axes, pinned not unified The changeset type is `Fixed`, which exempts this PR from the docs/ co-change requirement — but the glossary gate is separate from that exemption, and the 2048 cap in particular is a machine-canon-shaped fact that until now existed only inside a source comment. `docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync confirms all six targets in sync. * docs(#3697): document what the command now does, in one changeset sentence DOCS. The grammar section predated this round's two new rules, so it under-described the behaviour it exists to make lookup-able: - The placeholder paragraph enumerated three words; the rule is a DEFAULT. Any wording that selects no REQ-IDs warns, and the placeholders are matched as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty too. The comment-only line is called out, because "any other wording" would otherwise read as covering the shipped template's own ``. - The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is the quietest way to lose a requirement on this line — `requirements_updated` reads `true` either way — so it gets its own paragraph, with the parenthetical exemption stated beside it. - The machine kind is documented where a consumer would look for it, with the instruction to key on the kind rather than the wording. - The skipped-text note no longer implies it reports date shapes; it deliberately does not, and silently omitting that left the doc promising the behaviour this round removed. CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is `**** — .` and both canonical examples are one sentence; this fragment ran three. Now one, and covering what the round actually delivers rather than only the range shape it started from. * test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/` references in exempt test files and fails when a new one appears, so that comment turned four green gates red — `ci-docs-guard-registry` and the registration lint — for a file that reads no documentation at all. Caught by diffing the full suite's failing-name set against the same suite run at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base (install / config-home / shadowing tests under the sandbox HOME), and exactly these 4 did not. Rephrased rather than baselined. Adding the path to DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely starts READING a new docs path — the violation text asks the author to re-confirm the exemption still holds. Nothing here reads documentation; the guard matched prose. Baselining would have recorded a coupling that does not exist and made the next reader wonder what phase.test.cjs does with CLI-TOOLS.md. The comment still names where the contract is written, just without planting a path string. * fix(#3697): close four defects this round's own pre-push review drove An adversarial cross-AI review of this round, run before the push, returned 6 CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree, and every one was a shape the author had not probed — the rules were correct across the probe set and wrong just outside it. (1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired: `ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does not reach a BARE citation. The FP probe that scored this rule 0/0 only ever tested the parenthesised form. Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already demands an agreeing prefix for the same reason. Cost, stated in the census: a dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays silent. Same trade the strict-dash rule takes — under-report a rare shape rather than over-report a common one. Pinned as a declared blind spot by (2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped REQ-01 silently: the selector strips square brackets and R4's raw scanner did not. The bracket spelling the shipped template recommends was the one shape the rule could not see. (3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal requirement id — gap-checker's `parseRequirements` accepts it from REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement behind a rule meant only to hide dates. Narrowed to a four-digit year segment. The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED, and the removal is recorded in place: its premise was refuted, it was not inconvenient. (4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while reading as empty to the author, so R3b fired with nothing on screen to explain it. Zero-width and format characters are now stripped — stripped rather than treated as delimiters, because splitting on one would fabricate two fragments out of one ID. Also corrects the documentation the same review found overstated: the line is split on commas AND whitespace, and the ID shape is matched case-insensitively, so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was pre-existing selector behaviour; this round is the one that asserted the docs were true of it. Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is stripped, not split on), -19c (three citation forms), -19d (both bracket spellings), -19e (both halves of the rider boundary), -19f (the declared blind spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0 failures; lint:ci clean. * fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop The pre-push review's continuation refuted six of seven follow-up claims. The first one is the one that mattered: the invisible-character fix committed in c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from the detector wholesale made `REQ-01, REQ-02` go SILENT — the selector really does drop REQ-01, and the strip removed the only evidence of it. The test written alongside asserted the tokens and the empty R4 result and never asserted `warn`, so it DOCUMENTED the bug rather than catching it; that omission was the reviewer's own MISSED finding. An invisible is two different questions about one character, and the fix is to stop conflating them: absence-of-content for the empty test, DECORATION on a token for the drop rule. Neither is a reason to delete it from the line. R4 is generalised accordingly, because the continuation drove four more shapes a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`, `REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus `**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one class: decoration on a token the selector then cannot take. One rule, not four patches; patching them individually is how a list stays short and wrong. PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was never there and broke #3697-9d's channel routing with it. A parenthesis is this rule's citation marker. The rider stops adjudicating an undecidable shape. `API-2-01` is a legal requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15` are dates — no regex separates them, and both filters this round tried scored a miss in each direction. It now NAMES the token and states the ambiguity, which is the same thing the two warning voices already do about a range separator. Filtering hides a real dropped requirement; reporting it bare asks the author whether a date is a requirement; saying "this may equally be a date" does neither. The census and the docs are corrected to what the code does, including the part that is NOT complete: the prefix gate does not stop a citation that SHARES a selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level separates that from a real drop. A prose heuristic on "see" is the free-text detector this module exists to avoid, so the honest move is to say so. CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a 20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is identical to upstream/next's for every input, and marking is untouched. 513 tests in phase.test.cjs, 0 failures; lint:ci clean. * fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line Third pass of the round's own pre-push review, scoped to regression-hunting rather than further polish. Three findings, all driven, all mine. R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed a dropped requirement: the previous cut treated any shaved decoration as evidence, and emphasis is not evidence. Nothing separates that line from `**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to the skipped-text rider, which names the id without asserting a drop, exactly as `(REQ-02)` is handled. That is the #2334 class caught one cut before shipping. R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper sitting behind a delimiter. Shaves to a stable point now. The range OPERATOR lost its invisibles handling. `REQ-01 .. REQ-05` went silent, because the previous commit removed the invisible strip from BOTH the tokenizer and R4 when only R4's was wrong. An invisible is two questions about one character: for the classification rules it is noise and is stripped from the token; for the drop rule it is the evidence and must survive on the raw line. Stripping in both places hid a dropped id; stripping in neither hid a range. The reviewer's MISSED finding named the missing control — regression tests covered invisibles inside ids and not beside operators — and #3697-19i is that control. UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03` reported nothing: a running-depth counter left the unclosed `(` open through end-of-line, so every genuine drop after it inherited citation immunity. A parenthesis confers that immunity only as part of a MATCHED span now — an unmatched one is a typo, not a citation. CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode: `citedReqIds` identical to upstream/next, marking untouched, warnings appended. 522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero head-only failures against a probe worktree at upstream/next. * fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright Fourth and final pass of the round's own pre-push review. Three findings, and they shared one root cause, so this is a narrower rule rather than a longer list of shapes. TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single whitespace token, so a matched parenthetical inside it conferred immunity on the `REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans are now deleted from the line outright — which states what is actually meant, that for this rule a citation is not on the line — and an UNMATCHED paren is a typo that confers nothing. That also retires the running-depth counter whose previous bug was the mirror image: an unclosed `(` swallowing the rest of the line. DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was reported as a dropped requirement. It is a citation with sentence punctuation. The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify; `**REQ-01**;` does not, because outside the styling that character is punctuation. An invisible needs no adjacency test — nobody types one on purpose, so anywhere in the token it is corruption rather than intent. `**REQ-01**; REQ-02` therefore goes silent, and the test row asserting otherwise is inverted rather than deleted quietly: it was added one commit ago on the reasoning this pass refuted, and nothing distinguishes it from `see **REQ-7**; next topic`. Worth recording plainly: three successive cuts of this rule fired on a citation, and each fix was a narrower definition of EVIDENCE, never a longer list of shapes. The list-lengthening instinct is what produced the bug each time. The review's last MISSED finding named the missing control — the paren tests all surrounded matched spans with whitespace, so none covered a span sharing a token with an id outside it. #3697-19j carries both directions now. 526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero head-only failures against a probe at upstream/next; the invariant that `citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary Unicode inputs. * fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop Two defects, both found by this round's own pre-publication body claim-audit. 1. A DEMONSTRATED drop was discarded by the unverified voice. On `REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in delimiterDroppedIds and the formatter then reported `req-line-unverified`, whose message never mentions it — the one actionable finding masked by the token beside it. The over-cap channel now excludes a line carrying an R4 hit, exactly as rangeReadingOnly already did and for the same reason: that voice's whole claim is that nothing could be checked, and R4 has already checked something. The assertive channel still carries the over-cap rider, so nothing about the cap is traded away. Pinned by #3697-19l, which fails against the pre-fix build and nothing else does. 2. Three shipped artifacts asserted behaviour the code does not have — the same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules predicate, the CLI tools reference, and the warning's own advice string all listed markdown emphasis as an R4 trigger. It is not: styling is shaved BEFORE the test and tolerated around an id, never a trigger on its own, so `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger is exactly a glued `;`/`:` or an embedded invisible. The census predicate was wrong in a second way. Its 26-spelling separator sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and therefore biased: one-sided attachment drops silently for every punctuation outside the set — `/ | & + . > \` and the full-width and non-ASCII forms `; , ؛` all measured silent. The domain is wide open and R4 covers two characters of it. Said plainly in all three places rather than widened here: every previous widening of this rule first fired on a citation, so it is not done blind at the end of a round. Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n one-sided separators) so the documents and the code cannot drift apart again — which is what the round-4 blocker asked for. tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0. Both CONTEXT-INDEX consumers regenerated. * fix(#3697): re-sweep the separator census properly, and say what it really found The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `" from a 26-spelling sweep. That conclusion was forced by how the sweep was built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for every other separator. Different members of the domain were tested in different shapes, so no other answer was reachable. Caught by this round's pre-publication claim-audit of the response comment, reading the census comment against its own swept list. Re-swept fully crossed and driven through the built artifact: 21 separators x {bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All 34 are one shape — a separator glued to exactly one of the two ids, e.g. `REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and the `;`/`:` that R4 covers. So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census comment, the CONTEXT.md census-domains predicate and the CLI tools reference all say. The set is deliberately not widened here: three successive cuts of this rule fired on a citation, and a fourth at the end of a round with no adversarial pass is how each of those got in. Second false passage in the same block: styling-only decoration was described as "left to the skipped-text rider, which names the id". A rider only exists inside a message, and a message only exists once some rule sets `warn` — so on a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent. Describing it as handled reads as coverage. #3697-19m already pins the silence. so the test matches the documented claim. tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0. Both CONTEXT-INDEX consumers regenerated. * chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next Rebased onto next @ f4fefb0be. The two generated indexes conflicted on the replay and were resolved by regenerating, not by hand-merging: `npm run gen:context-index` for docs/CONTEXT-INDEX.json and `node examples/dynamic-context-management/gen-context-index.cjs --write` for the example's copy. Against next, each now differs only in the eight PHASE.REQ-LINE.SEAM.* predicates this PR adds (plus the PHASE class and the count); every base-side change (the SEAM.* predicates, the ADR-3942 value rewrites) is carried. `lint-example-parser-parity` and `gen-context-index --check` both pass. * fix(#3697): an over-cap token outranks the ambiguous range voice Round 7 review, Minor 1. `rangeReadingOnly`'s guard conjunction checked `hasGluedRangeFragment`, `inertIdShaped` and `delimiterDroppedIds` but not `oversizedTokens`, so a line carrying a clean, fully-selected spaced range *and* an unrelated token past the 2048-char scan cap was coded `req-line-range-reading` — a code CONTEXT.md's PHASE.REQ-LINE.SEAM.kinds predicate documents as "nothing was dropped" — over a token no rule (R1-R4 all skip over-cap tokens) had ever examined. Since the PR tells machine consumers to branch on `.code` rather than parse prose, that is a false "nothing to verify further" signal. The review's suggested fix was to add `oversizedTokens.length === 0` to the conjunction. That clause is right and is here, but on its own it routes the line to the ASSERTIVE channel: `req-line-misparse`, whose message says the line "could not be parsed as a comma-separated REQ-ID list" on a line where every ID present was in fact selected. That is the #2334 over-warning class this module's own channel-selection docblock exists to prevent — a false-clean code traded for a false-assertion one. So the correct destination is `req-line-unverified`: the line was not CHECKED, which is not the same as clean, and nothing on it demonstrably failed to parse either. Both non-assertive voices were carrying their own inline copy of the same "nothing was demonstrably dropped" conjunction, and that duplication is what let them drift: `rangeReadingOnly` omitted the cap, while the over-cap channel excluded a spaced range wholesale via `!hasSpacedRange`. Extract it once as `nothingDemonstrablyDropped` and have both read it. The new predicate is a strict superset of the old `!hasSpacedRange` guard — `spacedRangePairs.every()` is vacuously true when no spaced range fired — so the over-cap channel's behaviour on every line without a spaced range is unchanged, and R2 firing on an endpoint the selector did not take still routes to the assertive channel. Exactly one input class changes routing: a clean fully-selected spaced range beside an unexamined over-cap token, which moves from `req-line-range-reading` to `req-line-unverified`. Pinned by `#3697-19n` (the twin of `#3697-19l` on the other side of the boundary — there a demonstrated drop outranks the unverified voice, here the cap outranks the ambiguous one), with `#3697-19o` as the negative control asserting a clean range with nothing over the cap is still a range reading. * docs(#3697): state the warning-code precedence, and what the cap condition actually is Two surfaces, one point. The round 7 finding cited CONTEXT.md's PHASE.REQ-LINE.SEAM.kinds predicate as the documentation of what `req-line-range-reading` claims, and it was right to: that predicate said "R2 alone fired on selected endpoints; nothing was dropped" with no mention of the scan cap, describing a line the code could not distinguish from one carrying a token it never examined. The range reading is the weakest of the three claims and yields to the other two: a demonstrated drop makes the line a misparse, and a token the cap left unclassified makes it unverified. `docs/CLI-TOOLS.md` gains that sentence; the CONTEXT.md predicate gains it plus two precision points that this round's own pre-push adversarial review extracted over three passes, each with a driven counterexample I reproduced before acting on it: * "nothing was dropped" overstates the rule. The discriminator is RULE-scoped by design (round 3, Major 3), so that a parenthesised ID-shaped token does not re-open the #2334 over-warning class — `(REQ-02)` is indistinguishable from `(ADR-7)` at token level and is carried by the skipped-text rider, never by this code. `REQ-01, (REQ-02), REQ-03 - REQ-05` selects three, names REQ-02 as skipped, and is still a range reading. The predicate now says "no rule named a dropped ID", which is what the code tests. * The deferral condition is `oversizedTokens` being non-empty, NOT the presence of a token past the cap. Those differ: a long token the selector itself took can be exempt, because selection is uncapped and anchored and such a token was therefore examined. `RANGE-01 - RANGE-05, R-<2047 sevens>` yields `oversizedTokens=[]` and stays a range reading. Two earlier attempts to characterise WHEN the exemption applies were both refuted — the neighbour test admits any non-short neighbour, not just a range operator — so the predicate now states the condition and defers the exemption's own rule to SEAM.cap rather than paraphrasing it a third time. Verified after the edit: deferral holds if and only if the cap left a token unclassified, across all three counterexamples plus controls, and an exempt selector-taken long token is exhibited. No behaviour change — predicate text, one CLI-reference paragraph, and the two regenerated indexes. Predicate count unchanged at 285 across 22 classes, 0 duplicate ids; lint-example-parser-parity and both `--check` generators pass. * test(#3697): rename the round 7 regression pair — 19n was already taken Self-found immediately after the push, before anything was published to the review thread. The two tests added this round were named `#3697-19n` and `#3697-19o`, and `#3697-19n` was already in use: it is the declared-blind-spot case for one-sided separators in both attachment directions, generated inside a loop with a computed label, which is why a grep for a literal `test('#3697-19n'` did not find it. The PR body's own *Declared blind spots* list already refers to `#3697-19n` with that meaning. Nothing failed. Duplicate test names do not error, and the suite stayed green at 567/567 — which is the argument for fixing it rather than against. Two concrete costs: any TAP name-set differential collapses same-named tests under `sort -u`, so one of the two becomes invisible to exactly the did-not-run and pass->fail checks that name-set comparison exists to perform; and a reviewer reading "#3697-19n" in the body now gets a different test than the one the body means. Renamed to `#3697-19p` (over-cap beside a clean range) and `#3697-19q` (its negative control), the next free ids in the series; the pre-existing `#3697-19n` is untouched at its original 4 occurrences. Cross-references inside the renamed block were updated with them. Negative control re-run under the new ids: 19p still fails against pre-fix source and passes after, 19q passes at both ends. --------- Co-authored-by: Tom Boucher --- .changeset/zesty-wasps-fly.md | 5 + CONTEXT.md | 8 + docs/CLI-TOOLS.md | 130 ++ docs/CONTEXT-INDEX.json | 43 +- .../CONTEXT-INDEX.json | 597 +++--- src/phase.cts | 844 +++++++- tests/phase.test.cjs | 1834 +++++++++++++++++ 7 files changed, 3165 insertions(+), 296 deletions(-) create mode 100644 .changeset/zesty-wasps-fly.md diff --git a/.changeset/zesty-wasps-fly.md b/.changeset/zesty-wasps-fly.md new file mode 100644 index 000000000..bd143b71a --- /dev/null +++ b/.changeset/zesty-wasps-fly.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 3744 +--- +**`phase complete` now warns when the ROADMAP `**Requirements**:` line under-selects REQ-IDs** — a range (`REQ-01 … REQ-05`), a glued `;` or `:` delimiter (`REQ-01; REQ-02`), and any non-placeholder wording that selects nothing (`Deferred`, `N/A`) all marked fewer requirements than the line names while still reporting `requirements_updated: true` with zero warnings, and each now emits a warning naming what was selected and what was skipped, carrying a machine-readable kind, without expanding ranges or changing which IDs get marked. (#3697) diff --git a/CONTEXT.md b/CONTEXT.md index 56d52e9f8..fade9b4ea 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -721,6 +721,14 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr `LIVE-CONFIG.GUARD.SEAM.truncation=MAX_ENTRIES/MAX_DEPTH bound the walk; a bound hit sets truncated and diffLiveConfig emits kind:'unverified' — a truncated scan MUST NOT read as clean; boundary covered at {limit-1,limit,limit+1} via newestMtime's injected budget plus fast-check monotonicity, per RULESET.TESTS.boundary-coverage + RULESET.TESTS.property-based-testing` `LIVE-CONFIG.GUARD.SEAM.severity=reports by default locally; CI wires GSD_STRICT_LIVE_CONFIG_GUARD=1 on Linux/macOS lanes (test.yml, all three test jobs) so a suite-produced leak FAILS those runs; Windows lanes stay report-only pending the documented pre-existing USERPROFILE sweep (~190 test sites sandbox HOME alone) — promote once that lands; skipped by GSD_SKIP_LIVE_CONFIG_GUARD=1` `LIVE-CONFIG.GUARD.SEAM.ci-blind=the AMBIENT-ENV half stays CI-blind — CI never has these vars set, so green CI is not evidence for it; what strict mode catches in CI is the suite's own default-root leaks (HOME/USERPROFILE-derived), the guard remains the only loud signal for ambient-var escapes` +`PHASE.REQ-LINE.SEAM.module=src/phase.cts owns the ROADMAP **Requirements**: line seam as TWO module-scope functions, extracted so the parser is directly testable (a closure inside cmdPhaseComplete is reachable only by spawning the CLI, which no fast-check property can do): analyzeRequirementsLine(rawLine) -> RequirementsLineAnalysis, formatRequirementsLineWarning(phaseNum,rawLine,analysis) -> {code,message}|null; both exported, plus REQ_LINE_WARNING_CODE` +`PHASE.REQ-LINE.SEAM.selector-identity=citedReqIds is BYTE-IDENTICAL to the pre-extraction expression and is the ONLY thing that reaches the ledger; every rule below adds to warnings[] and NOTHING else — a change that alters what phase.complete MARKS is out of this seam's contract, not a refinement of it` +`PHASE.REQ-LINE.SEAM.rules=R1 whole-token range | R2 spaced operator between two selected interior-implying endpoints | R2' operator glued to one endpoint | R3 zero selection with ID-shaped residue | R3b zero selection on any non-placeholder non-empty line (#3697 AC-1b/AC-4) | R4 an ID the selector dropped to DECORATION — the TRIGGER is exactly a glued ;/: at either end OR an embedded invisible, never styling: quotes/backticks/emphasis are TOLERATED around the id (shaved before the test) but do NOT fire on their own, so `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are BOTH silent — plus outside any MATCHED parenthetical AND sharing a prefix with a SELECTED id (square brackets stripped as the selector strips them; parentheses deliberately NOT, they are the citation marker) | over-cap unclassified; warn is their disjunction, named rather than inlined in the return literal` +`PHASE.REQ-LINE.SEAM.kinds=three, carried as a machine code BESIDE the prose, never instead of it — req-line-misparse (ID-shaped content demonstrably not selected) | req-line-range-reading (R2 alone fired on selected endpoints and NO RULE NAMED A DROPPED ID — rule-scoped, never line-global: an unselected parenthetical is carried by the skipped-text rider, not by this code — so the voice must NOT claim a parse failure; it defers to req-line-unverified when the cap left a token UNCLASSIFIED, which is narrower than 'a token past the cap' — a long token the SELECTOR ITSELF took can be exempt, so the condition is oversizedTokens being non-empty, never the mere presence of a long token — see SEAM.cap for the exemption's own rule) | req-line-unverified (a token past the cap: the line was not classified, which is not the same as clean); emitted as the additive result field requirements_line_warning, because warnings[] is a documented string[] rendered by execute-phase.md and re-typing its elements is a breaking output-contract change; ABSENT entirely on a clean line` +`PHASE.REQ-LINE.SEAM.cap=REQ_TOKEN_SCAN_LIMIT=2048 bounds every predicate including each range participant's NEIGHBOURS, not just the operator; a bound hit sets oversizedTokens and takes the unverified kind — an unexaminable token MUST NOT read as clean, and the cap bounds the WORK, never the warning; boundary covered at {2047,2048,2049} across BOTH capped predicate families per RULESET.TESTS.boundary-coverage` +`PHASE.REQ-LINE.SEAM.placeholder=TBD/NONE as the LEAD token only, and R3b's non-empty test keys on VISIBLE content (a line of only zero-width/bidi/variation-selector codepoints is empty; an invisible INSIDE a token is decoration and R4 reports the drop); that gate — not the ID-shape gate — is what holds the whole #2334/#2339 negative space silent under R3b (measured: 15 of 15 fixtures held by non-zero selection or placeholderLed, 0 by ID shape)` +`PHASE.REQ-LINE.SEAM.census-domains=TWO open domains, each censused in-source with its NOT-reached consequence: range-operator spellings (reached: ..+, seven Unicode dashes plus ASCII -, …, to/thru/through; not reached: →, ~, ..=, ..<, until, up to) and comma-SUBSTITUTE separators (round 4's 26-spelling sweep concluded 'exactly ; and : ' because it swept the ONE-SIDED form for ;/: and only the BARE and SYMMETRIC forms for every other separator — different members tested in different shapes, so the answer was forced; re-swept round 5 FULLY CROSSED at 21 separators x {bare,trailing-space,leading-space,both} = 84, driven: 26 select both, 24 already warn, 34 UNDER-SELECT SILENTLY and all 34 are one-sided attachment — | / + & \ > . ! ? • · ؛ ; , - ~ and/plus — so R4 covers TWO CHARACTERS of a WIDE-OPEN domain, never the whole of it); R4's own NOT-reached set is therefore styling-only decoration, every non-;/: attachment, anything inside a MATCHED parenthetical, and any decorated id whose prefix is on NO selected id (REQ-01, FOO-02: x stays silent even when FOO-02 is real); the prefix gate is NOT complete in the other direction either — a citation SHARING a selected prefix (REQ-01, see REQ-7: sec 3) still fires and nothing at token level separates it from a real drop` +`PHASE.REQ-LINE.SEAM.gap-checker-divergence=normalizePhaseReqIds (src/gap-checker.cts) is a SECOND parser of the same ROADMAP value and DIVERGES on four axes — ranges (expanded there, never here, deliberately per #3697) | placeholder vocabulary (whole trimmed value after stripping parens there, LEAD token here) | parentheses (stripped there, not here) | ID shape (PHASE_REQ_ID_SHAPE_RE is wider) — pinned in BOTH directions by #3697-17 rather than unified, because unifying would change what phase.complete MARKS; consequence is user-visible: RANGE-01..RANGE-05 reports 5 requirements to gap analysis and 0 to phase complete` --- diff --git a/docs/CLI-TOOLS.md b/docs/CLI-TOOLS.md index c1eb43893..4fb30fd1e 100644 --- a/docs/CLI-TOOLS.md +++ b/docs/CLI-TOOLS.md @@ -283,6 +283,136 @@ consumer using `wave` for scheduling (`WAVE_FILTER`, the wave-safety check) is still working from the degraded assignment; only the diagnostic surfaces the loss. +### The ROADMAP `**Requirements**:` line grammar + +`phase complete` reads each phase's `**Requirements**:` line to decide which +REQ-IDs to mark. The grammar is deliberately small, and it is documented here +because the command now warns about it (#3697) — a warning about a rule that +cannot be looked up is not actionable. + +**The canonical form is a comma-separated list of REQ-IDs.** Square brackets are +optional; a REQ-ID is `PREFIX-N`, where the prefix is a letter followed by +letters or digits. Two tolerances are worth knowing because the warnings below +do **not** fire on them: the line is split on commas *and whitespace*, so +`REQ-01 REQ-02` selects both; and the ID shape is matched case-insensitively, +so `req-01` is selected and marked. Write the comma list anyway — it is what +every template and every example uses — but neither spelling is an error: + +``` +**Requirements**: REQ-01, REQ-02, REQ-03 +**Requirements**: [REQ-01, REQ-02, REQ-03] +``` + +**Ranges are not expanded** — `REQ-01 … REQ-05` selects the two endpoints and +nothing between them, and `REQ-01..REQ-05` selects nothing at all, because the +whole token fails the ID shape. This is a deliberate non-feature, not an +oversight: the line is a traceability record, and silently inventing IDs that +appear nowhere in `REQUIREMENTS.md` is worse than declining to. + +**A deliberately empty line is written `TBD` or `None`.** Those two words are +the placeholder vocabulary, and they are matched as the line's leading token, so +`None (per ADR-7)` and `**None**` are declared-empty too. **Any other wording +that selects no REQ-IDs warns** — `Deferred`, `N/A`, `Pending`, `TBA`, a bare +`-`, or free prose — because from the command's side an unrecognised word is +indistinguishable from a line that was meant to cite requirements and failed +to. A line whose only content is an HTML comment is not "other wording" and +stays silent, so the shipped template's own `` does not warn. + +**Write each requirement as a bare ID, separated by a comma.** The line is +split on commas and whitespace only, and nothing else is stripped, so anything +attached to an ID takes the ID with it. `REQ-01; REQ-02` marks **only** +`REQ-02`; so do `REQ-01 ;REQ-02`, `**REQ-01;** REQ-02` and an ID carrying a +stray invisible character. Those are the cases the warning names — it reports +the dropped ID, because this is the quietest way to lose a requirement here +(the command reports `requirements_updated: true` either way). + +**What the warning does NOT reach.** Stated so its silence is not read as a +clean bill, and enumerated rather than summarised, because each of these is a +requirement that goes missing without a word. The check fires only when the +thing touching the ID is a `;`, a `:`, or an invisible character: + +- **Markdown styling on its own is silent.** `**REQ-01**, REQ-02` drops + `REQ-01` and says nothing at all. Styling is not evidence that a separator + was meant. `**REQ-01**; REQ-02` is silent for the same reason — the `**` + sits between the ID and the `;`, so nothing is touching the ID. +- **Any other attached punctuation is silent.** `REQ-01/ REQ-02`, + `REQ-01| REQ-02`, `REQ-01. REQ-02`, `REQ-01+ REQ-02`, `REQ-01> REQ-02` and + their full-width and non-ASCII equivalents (`;`, `,`, `؛`) each mark only + `REQ-02`. A fully crossed sweep — 21 separators against bare, leading-space, + trailing-space and both-spaces spellings, 84 combinations — found **34** + silent under-selections, every one of them a separator glued to exactly one + of the two IDs. Only `;` and `:` are in the set. Widening it is a live + option — say the word — but each past widening of this check first fired on + a citation, so it is not done blind. +- **A dropped ID whose prefix matches nothing selected is silent.** + `REQ-01, REQ-02: login` is reported; `REQ-01, FOO-02: x` is not. A citation + is textually identical to a dropped requirement — `REQ-01, see ADR-7: + section 3` carries `ADR-7:` in exactly the shape `REQ-01;` has — and prefix + agreement is the only thing that separates them without guessing at prose. + Anything inside matched parentheses is left alone for the same reason: + `(see ADR-7: section 3)` is a citation. + +The prefix gate is not complete in the other direction either: a citation that +*shares* a selected prefix — `REQ-01, see REQ-7: sec 3` — does warn, naming +`REQ-7` as dropped. Nothing at the token level separates that from a real +drop. + +**Dash spellings need a full ID on both sides.** `REQ-01-REQ-05` reads as a +range; `REQ-01-05` does not, and neither does any of its typographic variants +(en dash, em dash, minus sign, and the rest). The reason is that +`PREFIX-` is also a date (`FY-2026-08`) and a sub-numbered +ID (`API-2-01`), so warning on it would be noise on lines that are perfectly +correct. `..`, `…`, `to`, `thru` and `through` have no such reading and do +accept a bare numeric endpoint (`REQ-01..05`). The cost is that +`REQ-01, REQ-02-05` is not reported; a bare `REQ-02-05` still is, because it +selects nothing. + +**What warns, and in which of the three voices.** All go to `warnings[]` and +none blocks completion: + +- *"could not be parsed as a comma-separated REQ-ID list"* — the line did not + yield the requirements it appears to name. That covers two cases, and the + message says which one it is: ID-shaped text on the line was **not** + selected, so something was demonstrably dropped; **or** the line selected + nothing at all while not being a `TBD`/`None` placeholder, in which case + there may be no ID-shaped text on it whatsoever (`Deferred` takes this + voice). Either way nothing was marked and the line needs fixing. +- *"contains what reads as a range between two cited REQ-IDs"* — the separator + is the only thing in question: a separator between two cited IDs could equally + be a range or an annotation, and the command cannot tell them apart, so it + states both readings rather than asserting a failure that may not have + happened. It speaks about the **separator**, not about the whole line. It is + also the *weakest* of the three claims, so it yields to the other two: a + demonstrated drop elsewhere on the line makes it a misparse, and an + unexamined over-cap token makes the line unverified — a voice whose claim is + that nothing was dropped cannot speak over a token no rule read. + +**Each warning carries a machine-readable kind.** The prose goes to +`warnings[]` as before — that field is unchanged and is still an array of +strings — and the kind is emitted beside it as `requirements_line_warning`, +one of `req-line-misparse`, `req-line-range-reading` or `req-line-unverified`. +The field is absent entirely when the line is clean. Key on the kind rather +than on the wording; the wording is free to improve. + +Either of those two voices may add a factual note naming **ID-shaped text on the +line that was not selected**. Square brackets *are* stripped — `[REQ-01, REQ-02]` +is the documented form — but parentheses are not, so `(REQ-02)` is not marked. +The command cannot tell that from `(ADR-7)`, which is a citation and correctly +ignored, so it names what it skipped and leaves the judgement to you. Where the +skipped text is `PREFIX--` — the shape the dash rule above +declines to adjudicate, because `FY-2026-08` is a date and `API-2-01` is a +legal requirement id and no rule separates them — it is still named, with that +ambiguity stated alongside it. Naming it and saying why it is ambiguous beats +both alternatives: filtering it hides a real dropped requirement, and reporting +it bare asks you to check whether a date is a requirement. +The third voice is for input the command could not examine: *"could not be +checked ... the REQ-ID selection on this line is unverified"*. Range detection +is bounded at 2,048 characters per token, so a longer token is not classified — +and unclassified is reported, never treated as clean. Selection itself is *not* +bounded, so a valid REQ-ID longer than that is still selected and marked +normally; it only triggers this voice if something beside it could have formed a +range with it, which is the case where the bound actually suppressed a check. + --- ## Roadmap Commands diff --git a/docs/CONTEXT-INDEX.json b/docs/CONTEXT-INDEX.json index 569c88743..d0eb0cf0f 100644 --- a/docs/CONTEXT-INDEX.json +++ b/docs/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 277, + "count": 285, "classes": { "ARCH": 1, "CI": 2, @@ -10,6 +10,7 @@ "LEARNING": 1, "LIVE-CONFIG": 6, "META": 4, + "PHASE": 8, "PLANNING": 3, "PR": 2, "PRED": 68, @@ -190,6 +191,46 @@ "klass": "META", "value": "read CONTRIBUTING.md sections \"Pull Request Guidelines\" + \"CHANGELOG Entries\" before EVERY agent dispatch" }, + { + "id": "PHASE.REQ-LINE.SEAM.cap", + "klass": "PHASE", + "value": "REQ_TOKEN_SCAN_LIMIT=2048 bounds every predicate including each range participant's NEIGHBOURS, not just the operator; a bound hit sets oversizedTokens and takes the unverified kind — an unexaminable token MUST NOT read as clean, and the cap bounds the WORK, never the warning; boundary covered at {2047,2048,2049} across BOTH capped predicate families per RULESET.TESTS.boundary-coverage" + }, + { + "id": "PHASE.REQ-LINE.SEAM.census-domains", + "klass": "PHASE", + "value": "TWO open domains, each censused in-source with its NOT-reached consequence: range-operator spellings (reached: ..+, seven Unicode dashes plus ASCII -, …, to/thru/through; not reached: →, ~, ..=, ..<, until, up to) and comma-SUBSTITUTE separators (round 4's 26-spelling sweep concluded 'exactly ; and : ' because it swept the ONE-SIDED form for ;/: and only the BARE and SYMMETRIC forms for every other separator — different members tested in different shapes, so the answer was forced; re-swept round 5 FULLY CROSSED at 21 separators x {bare,trailing-space,leading-space,both} = 84, driven: 26 select both, 24 already warn, 34 UNDER-SELECT SILENTLY and all 34 are one-sided attachment — | / + & \\ > . ! ? • · ؛ ; , - ~ and/plus — so R4 covers TWO CHARACTERS of a WIDE-OPEN domain, never the whole of it); R4's own NOT-reached set is therefore styling-only decoration, every non-;/: attachment, anything inside a MATCHED parenthetical, and any decorated id whose prefix is on NO selected id (REQ-01, FOO-02: x stays silent even when FOO-02 is real); the prefix gate is NOT complete in the other direction either — a citation SHARING a selected prefix (REQ-01, see REQ-7: sec 3) still fires and nothing at token level separates it from a real drop" + }, + { + "id": "PHASE.REQ-LINE.SEAM.gap-checker-divergence", + "klass": "PHASE", + "value": "normalizePhaseReqIds (src/gap-checker.cts) is a SECOND parser of the same ROADMAP value and DIVERGES on four axes — ranges (expanded there, never here, deliberately per #3697) | placeholder vocabulary (whole trimmed value after stripping parens there, LEAD token here) | parentheses (stripped there, not here) | ID shape (PHASE_REQ_ID_SHAPE_RE is wider) — pinned in BOTH directions by #3697-17 rather than unified, because unifying would change what phase.complete MARKS; consequence is user-visible: RANGE-01..RANGE-05 reports 5 requirements to gap analysis and 0 to phase complete" + }, + { + "id": "PHASE.REQ-LINE.SEAM.kinds", + "klass": "PHASE", + "value": "three, carried as a machine code BESIDE the prose, never instead of it — req-line-misparse (ID-shaped content demonstrably not selected) | req-line-range-reading (R2 alone fired on selected endpoints and NO RULE NAMED A DROPPED ID — rule-scoped, never line-global: an unselected parenthetical is carried by the skipped-text rider, not by this code — so the voice must NOT claim a parse failure; it defers to req-line-unverified when the cap left a token UNCLASSIFIED, which is narrower than 'a token past the cap' — a long token the SELECTOR ITSELF took can be exempt, so the condition is oversizedTokens being non-empty, never the mere presence of a long token — see SEAM.cap for the exemption's own rule) | req-line-unverified (a token past the cap: the line was not classified, which is not the same as clean); emitted as the additive result field requirements_line_warning, because warnings[] is a documented string[] rendered by execute-phase.md and re-typing its elements is a breaking output-contract change; ABSENT entirely on a clean line" + }, + { + "id": "PHASE.REQ-LINE.SEAM.module", + "klass": "PHASE", + "value": "src/phase.cts owns the ROADMAP **Requirements**: line seam as TWO module-scope functions, extracted so the parser is directly testable (a closure inside cmdPhaseComplete is reachable only by spawning the CLI, which no fast-check property can do): analyzeRequirementsLine(rawLine) -> RequirementsLineAnalysis, formatRequirementsLineWarning(phaseNum,rawLine,analysis) -> {code,message}|null; both exported, plus REQ_LINE_WARNING_CODE" + }, + { + "id": "PHASE.REQ-LINE.SEAM.placeholder", + "klass": "PHASE", + "value": "TBD/NONE as the LEAD token only, and R3b's non-empty test keys on VISIBLE content (a line of only zero-width/bidi/variation-selector codepoints is empty; an invisible INSIDE a token is decoration and R4 reports the drop); that gate — not the ID-shape gate — is what holds the whole #2334/#2339 negative space silent under R3b (measured: 15 of 15 fixtures held by non-zero selection or placeholderLed, 0 by ID shape)" + }, + { + "id": "PHASE.REQ-LINE.SEAM.rules", + "klass": "PHASE", + "value": "R1 whole-token range | R2 spaced operator between two selected interior-implying endpoints | R2' operator glued to one endpoint | R3 zero selection with ID-shaped residue | R3b zero selection on any non-placeholder non-empty line (#3697 AC-1b/AC-4) | R4 an ID the selector dropped to DECORATION — the TRIGGER is exactly a glued ;/: at either end OR an embedded invisible, never styling: quotes/backticks/emphasis are TOLERATED around the id (shaved before the test) but do NOT fire on their own, so `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are BOTH silent — plus outside any MATCHED parenthetical AND sharing a prefix with a SELECTED id (square brackets stripped as the selector strips them; parentheses deliberately NOT, they are the citation marker) | over-cap unclassified; warn is their disjunction, named rather than inlined in the return literal" + }, + { + "id": "PHASE.REQ-LINE.SEAM.selector-identity", + "klass": "PHASE", + "value": "citedReqIds is BYTE-IDENTICAL to the pre-extraction expression and is the ONLY thing that reaches the ledger; every rule below adds to warnings[] and NOTHING else — a change that alters what phase.complete MARKS is out of this seam's contract, not a refinement of it" + }, { "id": "PLANNING.PATH.PARITY.project-scope", "klass": "PLANNING", diff --git a/examples/dynamic-context-management/CONTEXT-INDEX.json b/examples/dynamic-context-management/CONTEXT-INDEX.json index fdd29e60d..69d0ae331 100644 --- a/examples/dynamic-context-management/CONTEXT-INDEX.json +++ b/examples/dynamic-context-management/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 277, + "count": 285, "classes": { "ARCH": 1, "CI": 2, @@ -10,6 +10,7 @@ "LEARNING": 1, "LIVE-CONFIG": 6, "META": 4, + "PHASE": 8, "PLANNING": 3, "PR": 2, "PRED": 68, @@ -29,1429 +30,1477 @@ "id": "ARCH.SKILL.improve-codebase.next-candidates", "klass": "ARCH", "value": "[Workstream Progress Projection Module]", - "line": 694 + "line": 697 }, { "id": "CI.GATE.changeset-lint", "klass": "CI", "value": "hard-fail for user-facing code diffs unless .changeset/* or PR has no-changelog label", - "line": 678 + "line": 681 }, { "id": "CI.GATE.issue-link-required", "klass": "CI", "value": "hard-fail if PR body lacks closes/fixes/resolves #", - "line": 677 + "line": 680 }, { "id": "CONFIG.LOCATION.SEAM.in-process-scrub", "klass": "CONFIG", "value": "TEST_ENV_BASE reaches CHILD env only; a test calling install() IN-PROCESS must additionally use helpers.scrubConfigLocationEnv() in beforeEach + its restorer in afterEach — HOME/USERPROFILE sandboxing is NOT sufficient because getGlobalConfigDir is env-FIRST", - "line": 714 + "line": 717 }, { "id": "CONFIG.LOCATION.SEAM.kimi-two-homes", "klass": "CONFIG", "value": "kimi declares TWO config-location vars: KIMI_CONFIG_DIR (registry, generic Agent-Skills root via resolveKimiGlobalDir) and KIMI_SHARE_DIR (KIMI_HOOKS_TOML_DESCRIPTOR, kimi's OWN native config.toml carrying GSD's [[hooks]] block via resolveKimiHooksTomlDir); a registry-only derivation covers the first and silently misses the second", - "line": 713 + "line": 716 }, { "id": "CONFIG.LOCATION.SEAM.scrub-set", "klass": "CONFIG", "value": "tests/helpers.cjs CONFIG_LOCATION_ENV_KEYS is DERIVED from five sources rather than maintained as one hand-written list (source 4 IS a literal residue list, for vars that fit no other rung — what is never hand-listed is the SET): capability-registry runtimes[].runtime.configHome.env AND [].configHome.skillsHome.env + runtime-homes NON_REGISTRY_CONFIG_HOME_DESCRIPTORS[].env AND [].skillsHome.env (a descriptor is a descriptor — BOTH descriptor rungs walk skillsHome, which resolves independently via resolveSkillsBaseFromDescriptor) + runtime-homes GSD_LOCATION_ENV_KEYS + a residue list (GROK_AGENTS_HOME, GSD_RUNTIME, GSD_PROJECT, GSD_WORKSTREAM) + WRITE_ESCAPE_PERMISSION_ENV_KEYS (GSD_ALLOW_SYMLINKED_DEST — a permission, not a location: it names no path but disarms the symlink-escape guard, so blanking it makes the guard STRICTER, never looser); adding a config-location var means making it ENUMERABLE at one of those sources, not appending a literal", - "line": 711 + "line": 714 }, { "id": "CONFIG.LOCATION.SEAM.two-families", "klass": "CONFIG", "value": "runtime configHomes (where a third-party runtime keeps config, registry- or descriptor-declared) and GSD's OWN location vars (GSD_HOME -> $GSD_HOME/.gsd store, GSD_AGENTS_DIR -> getAgentsDir priority 1) are DISTINCT families; no registry derivation reaches the second, and treating a miss there as a registry gap is what produced review round 2", - "line": 712 + "line": 715 }, { "id": "CONFIG.SEAM.loadConfig-context", "klass": "CONFIG", "value": "loadConfig(cwd,{workstream}) replaces env-mutation fallback; no temporary process.env GSD_WORKSTREAM rewrites", - "line": 710 + "line": 713 }, { "id": "EXEC.CLASSIFY.classes", "klass": "EXEC", "value": "{class:'quota-exceeded'|'classify-handoff-bug'|'unknown-failure', sentinel?, retryAfterSeconds?}", - "line": 932 + "line": 943 }, { "id": "EXEC.CLASSIFY.cross-runtime", "klass": "EXEC", "value": "Anthropic/CC: usage limit|rate limit|quota|429|retry-after; Copilot CLI: rate_limit (stem); Codex CLI: 429|usage_limit_reached|too many requests", - "line": 934 + "line": 945 }, { "id": "EXEC.CLASSIFY.handler", "klass": "EXEC", "value": "gsd-core/bin/lib/agent-command-router.cjs:classifyAgentFailure (registered via command-aliases.cjs; mutation:false outputMode:json)", - "line": 930 + "line": 941 }, { "id": "EXEC.CLASSIFY.precedence", "klass": "EXEC", "value": "quota sentinel wins over classifyHandoffIfNeeded bug when both appear", - "line": 935 + "line": 946 }, { "id": "EXEC.CLASSIFY.proactive-signal-not-usable", "klass": "EXEC", "value": "Anthropic exposes anthropic-ratelimit-* headers + Agent SDK RateLimitEvent; Claude Code subprocess does NOT forward to hooks/statusline today (upstream #33820, #22407, #32796)", - "line": 937 + "line": 948 }, { "id": "EXEC.CLASSIFY.retry-after-parser", "klass": "EXEC", "value": "\\bretry[-_ ]after[:\\s]+(\\d+)\\b avoids embedded-word false matches like noretry-after", - "line": 936 + "line": 947 }, { "id": "EXEC.CLASSIFY.sentinel-order", "klass": "EXEC", "value": "most specific first: 429 beats too-many-requests; resource_exhausted beats quota (array order in src/agent-command-router.cts QUOTA_SENTINELS checks resource_exhausted before quota); case-insensitive; canonical sentinel value is lower-cased form", - "line": 933 + "line": 944 }, { "id": "EXEC.CLASSIFY.workflow", "klass": "EXEC", "value": "gsd-core/workflows/execute-phase.md step 7; class-distinct prompts (quota-to-wait-for-reset; classify-handoff-bug-to-spot-check; unknown-to-continue/stop)", - "line": 931 + "line": 942 }, { "id": "GSD-RESEARCH.CONTEXT-DISCIPLINE", "klass": "GSD-RESEARCH", "value": "less-context levers: subagent isolation + compact provider output + fetches-to-disk + cache-returns-digest; API clear_tool_uses/memory tool are the conceptual model, not a Claude Code harness knob", - "line": 460 + "line": 463 }, { "id": "GSD-RESEARCH.INTEGRATION.L2-hybrid", "klass": "GSD-RESEARCH", "value": "code owns cache+legitimacy+confidence+provider-pick (gsd-tools query research-plan/research-store/package-legitimacy); MCP owns the fetch; agent returns RESEARCH.md path, never raw fetches", - "line": 458 + "line": 461 }, { "id": "GSD-RESEARCH.MODULE.package-legitimacy", "klass": "GSD-RESEARCH", "value": "registry-API verdicts (npm/PyPI/crates.io injectable adapters) computed from thresholds {minAgeDays:30,minWeeklyDownloads:1000,requireRepo:true}; verdict OK|SUS|SLOP per package; slopcheck=optional adapter that can only escalate, never the install-or-degrade gate", - "line": 457 + "line": 460 }, { "id": "GSD-RESEARCH.MODULE.research-provider", "klass": "GSD-RESEARCH", "value": "single source of truth PROVIDER_WATERFALL (docs Context7->Ref->Jina->websearch; web Exa->Tavily->Perplexity->Brave->websearch; scrape Firecrawl->Jina); planResearch returns cache-hits+fetch-plan; classifyConfidence stamps HIGH|MEDIUM|LOW by provider AUTHORITY + verification EVIDENCE (HIGH requires code-computed ground-truth corroboration e.g. legitimacyVerdict OK; provider authority alone caps at MEDIUM; SLOP caps at LOW); Firecrawl is scrape-only (not in the docs or web legs)", - "line": 456 + "line": 459 }, { "id": "GSD-RESEARCH.MODULE.research-store", "klass": "GSD-RESEARCH", "value": "content-addressed cache; key=sha256(ecosystem+library+version+query+kind); getResearch->{hit,stale} never throws (mirrors graphify staleness); ttlForSource curated HIGH 30d|MED 7d|web LOW 1d; tiers: curated-doc kinds -> ~/.gsd/research-cache (cross-project), web/synthesis -> project .planning/research/.cache", - "line": 455 + "line": 458 }, { "id": "GSD-RESEARCH.PROVIDER.availability", "klass": "GSD-RESEARCH", "value": "config flags brave_search/exa_search/firecrawl/tavily_search/ref_search/perplexity/jina (env _API_KEY or ~/.gsd/_api_key); context7/jina/websearch always available; planResearch falls through waterfall to websearch terminal", - "line": 459 + "line": 462 }, { "id": "LEARNING.prompt-budget.boundary-gap", "klass": "LEARNING", "value": "PR #3708 commit 2df566ed reserved NOTE_RESERVE_TOKENS in pressure-threshold AND in minSet pre-check; both buggy paths only fire when baseTokens ∈ (effectiveBudget - NOTE_RESERVE_TOKENS, effectiveBudget]; original test suite used budgets far from that band so neither path was exercised; fix bde1ae8f confines NOTE_RESERVE accounting to post-trim assembly path only; future budget/limit code MUST add boundary fixtures per RULESET.TESTS.boundary-coverage.fixtures", - "line": 623 + "line": 626 }, { "id": "LIVE-CONFIG.GUARD.SEAM.ci-blind", "klass": "LIVE-CONFIG", "value": "the AMBIENT-ENV half stays CI-blind — CI never has these vars set, so green CI is not evidence for it; what strict mode catches in CI is the suite's own default-root leaks (HOME/USERPROFILE-derived), the guard remains the only loud signal for ambient-var escapes", - "line": 720 + "line": 723 }, { "id": "LIVE-CONFIG.GUARD.SEAM.module", "klass": "LIVE-CONFIG", "value": "scripts/live-config-guard.cjs (deliberately NOT scripts/lib/, which the installer copies to users wholesale while uninstall removes only an allowlist; excluded from the npm tarball via package.json files[] together with its whole require chain run-tests.cjs/affected-tests-lib.cjs/run-affected-tests.cjs — a partial exclusion trips the #2858 shipped-requires-only-shipped gate); exports [resolveLiveConfigRoots, resolveExtraWatchTargets, snapshotLiveConfig, diffLiveConfig, formatViolations, newestMtime]; driven by scripts/run-tests.cjs pre/post suite", - "line": 715 + "line": 718 }, { "id": "LIVE-CONFIG.GUARD.SEAM.non-root-targets", "klass": "LIVE-CONFIG", "value": "resolveExtraWatchTargets covers THREE live write surfaces that are not runtime config ROOTS (skills bases are a DELIBERATE non-target — the config-root layout misfires beneath them, so they need their own layout): $GSD_HOME/.gsd watched WHOLESALE (exclusively GSD-owned, so the shared-root trap does not apply) plus ONE config.toml per NON_REGISTRY_CONFIG_HOME_DESCRIPTORS entry, each watched as a SINGLE FILE (those roots belong to their products) — today three targets, since #2755 split Kimi CLI (~/.kimi, KIMI_SHARE_DIR) from Kimi Code (~/.kimi-code, KIMI_CODE_HOME); the targets are DERIVED by iterating that array, never by calling a named resolver, so a further descriptor is picked up without editing the guard PROVIDED it owns the same NON_REGISTRY_OWNED_FILE ('config.toml') — one that owns a different filename needs a per-descriptor mapping, the named residual the guard states at its own definition. SECOND RESIDUAL: config.toml is not all GSD writes into those roots — installSharedHooksBundle also populates /hooks/, which is UNWATCHED; closing it is a layout decision, like skills bases; passed to snapshotLiveConfig explicitly so a fixture-root caller cannot pull the real ~/.gsd into its snapshot", - "line": 717 + "line": 720 }, { "id": "LIVE-CONFIG.GUARD.SEAM.scope", "klass": "LIVE-CONFIG", "value": "ownership-based, never whole-root: GSD_OWNED_ENTRIES top-level footprint + children whose name startsWith GSD_ARTIFACT_PREFIX ('gsd-') under GSD_PREFIXED_PARENTS (dirs shared with the host agent); watching a shared root wholesale false-positives on the host's own writes and a guard that cries wolf gets disabled", - "line": 716 + "line": 719 }, { "id": "LIVE-CONFIG.GUARD.SEAM.severity", "klass": "LIVE-CONFIG", "value": "reports by default locally; CI wires GSD_STRICT_LIVE_CONFIG_GUARD=1 on Linux/macOS lanes (test.yml, all three test jobs) so a suite-produced leak FAILS those runs; Windows lanes stay report-only pending the documented pre-existing USERPROFILE sweep (~190 test sites sandbox HOME alone) — promote once that lands; skipped by GSD_SKIP_LIVE_CONFIG_GUARD=1", - "line": 719 + "line": 722 }, { "id": "LIVE-CONFIG.GUARD.SEAM.truncation", "klass": "LIVE-CONFIG", "value": "MAX_ENTRIES/MAX_DEPTH bound the walk; a bound hit sets truncated and diffLiveConfig emits kind:'unverified' — a truncated scan MUST NOT read as clean; boundary covered at {limit-1,limit,limit+1} via newestMtime's injected budget plus fast-check monotonicity, per RULESET.TESTS.boundary-coverage + RULESET.TESTS.property-based-testing", - "line": 718 + "line": 721 }, { "id": "META.RULE.brief-must-cite-doc", "klass": "META", "value": "agent prompts MUST quote the canonical doc line being applied; paraphrasing from predicate memory drifts and produces violations", - "line": 771 + "line": 782 }, { "id": "META.RULE.brief-no-paraphrase", "klass": "META", "value": "writing \"k040 — never leave changelog box unchecked\" caused 5 of 8 agents to edit CHANGELOG.md in violation of CONTRIBUTING.md L110", - "line": 772 + "line": 783 }, { "id": "META.RULE.canonical-source-precedence", "klass": "META", "value": "CONTRIBUTING.md > docs/adr/* > CONTEXT.md > agent memory", - "line": 769 + "line": 780 }, { "id": "META.RULE.read-contributing-first", "klass": "META", "value": "read CONTRIBUTING.md sections \"Pull Request Guidelines\" + \"CHANGELOG Entries\" before EVERY agent dispatch", - "line": 770 + "line": 781 + }, + { + "id": "PHASE.REQ-LINE.SEAM.cap", + "klass": "PHASE", + "value": "REQ_TOKEN_SCAN_LIMIT=2048 bounds every predicate including each range participant's NEIGHBOURS, not just the operator; a bound hit sets oversizedTokens and takes the unverified kind — an unexaminable token MUST NOT read as clean, and the cap bounds the WORK, never the warning; boundary covered at {2047,2048,2049} across BOTH capped predicate families per RULESET.TESTS.boundary-coverage", + "line": 728 + }, + { + "id": "PHASE.REQ-LINE.SEAM.census-domains", + "klass": "PHASE", + "value": "TWO open domains, each censused in-source with its NOT-reached consequence: range-operator spellings (reached: ..+, seven Unicode dashes plus ASCII -, …, to/thru/through; not reached: →, ~, ..=, ..<, until, up to) and comma-SUBSTITUTE separators (round 4's 26-spelling sweep concluded 'exactly ; and : ' because it swept the ONE-SIDED form for ;/: and only the BARE and SYMMETRIC forms for every other separator — different members tested in different shapes, so the answer was forced; re-swept round 5 FULLY CROSSED at 21 separators x {bare,trailing-space,leading-space,both} = 84, driven: 26 select both, 24 already warn, 34 UNDER-SELECT SILENTLY and all 34 are one-sided attachment — | / + & \\ > . ! ? • · ؛ ; , - ~ and/plus — so R4 covers TWO CHARACTERS of a WIDE-OPEN domain, never the whole of it); R4's own NOT-reached set is therefore styling-only decoration, every non-;/: attachment, anything inside a MATCHED parenthetical, and any decorated id whose prefix is on NO selected id (REQ-01, FOO-02: x stays silent even when FOO-02 is real); the prefix gate is NOT complete in the other direction either — a citation SHARING a selected prefix (REQ-01, see REQ-7: sec 3) still fires and nothing at token level separates it from a real drop", + "line": 730 + }, + { + "id": "PHASE.REQ-LINE.SEAM.gap-checker-divergence", + "klass": "PHASE", + "value": "normalizePhaseReqIds (src/gap-checker.cts) is a SECOND parser of the same ROADMAP value and DIVERGES on four axes — ranges (expanded there, never here, deliberately per #3697) | placeholder vocabulary (whole trimmed value after stripping parens there, LEAD token here) | parentheses (stripped there, not here) | ID shape (PHASE_REQ_ID_SHAPE_RE is wider) — pinned in BOTH directions by #3697-17 rather than unified, because unifying would change what phase.complete MARKS; consequence is user-visible: RANGE-01..RANGE-05 reports 5 requirements to gap analysis and 0 to phase complete", + "line": 731 + }, + { + "id": "PHASE.REQ-LINE.SEAM.kinds", + "klass": "PHASE", + "value": "three, carried as a machine code BESIDE the prose, never instead of it — req-line-misparse (ID-shaped content demonstrably not selected) | req-line-range-reading (R2 alone fired on selected endpoints and NO RULE NAMED A DROPPED ID — rule-scoped, never line-global: an unselected parenthetical is carried by the skipped-text rider, not by this code — so the voice must NOT claim a parse failure; it defers to req-line-unverified when the cap left a token UNCLASSIFIED, which is narrower than 'a token past the cap' — a long token the SELECTOR ITSELF took can be exempt, so the condition is oversizedTokens being non-empty, never the mere presence of a long token — see SEAM.cap for the exemption's own rule) | req-line-unverified (a token past the cap: the line was not classified, which is not the same as clean); emitted as the additive result field requirements_line_warning, because warnings[] is a documented string[] rendered by execute-phase.md and re-typing its elements is a breaking output-contract change; ABSENT entirely on a clean line", + "line": 727 + }, + { + "id": "PHASE.REQ-LINE.SEAM.module", + "klass": "PHASE", + "value": "src/phase.cts owns the ROADMAP **Requirements**: line seam as TWO module-scope functions, extracted so the parser is directly testable (a closure inside cmdPhaseComplete is reachable only by spawning the CLI, which no fast-check property can do): analyzeRequirementsLine(rawLine) -> RequirementsLineAnalysis, formatRequirementsLineWarning(phaseNum,rawLine,analysis) -> {code,message}|null; both exported, plus REQ_LINE_WARNING_CODE", + "line": 724 + }, + { + "id": "PHASE.REQ-LINE.SEAM.placeholder", + "klass": "PHASE", + "value": "TBD/NONE as the LEAD token only, and R3b's non-empty test keys on VISIBLE content (a line of only zero-width/bidi/variation-selector codepoints is empty; an invisible INSIDE a token is decoration and R4 reports the drop); that gate — not the ID-shape gate — is what holds the whole #2334/#2339 negative space silent under R3b (measured: 15 of 15 fixtures held by non-zero selection or placeholderLed, 0 by ID shape)", + "line": 729 + }, + { + "id": "PHASE.REQ-LINE.SEAM.rules", + "klass": "PHASE", + "value": "R1 whole-token range | R2 spaced operator between two selected interior-implying endpoints | R2' operator glued to one endpoint | R3 zero selection with ID-shaped residue | R3b zero selection on any non-placeholder non-empty line (#3697 AC-1b/AC-4) | R4 an ID the selector dropped to DECORATION — the TRIGGER is exactly a glued ;/: at either end OR an embedded invisible, never styling: quotes/backticks/emphasis are TOLERATED around the id (shaved before the test) but do NOT fire on their own, so `**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are BOTH silent — plus outside any MATCHED parenthetical AND sharing a prefix with a SELECTED id (square brackets stripped as the selector strips them; parentheses deliberately NOT, they are the citation marker) | over-cap unclassified; warn is their disjunction, named rather than inlined in the return literal", + "line": 726 + }, + { + "id": "PHASE.REQ-LINE.SEAM.selector-identity", + "klass": "PHASE", + "value": "citedReqIds is BYTE-IDENTICAL to the pre-extraction expression and is the ONLY thing that reaches the ledger; every rule below adds to warnings[] and NOTHING else — a change that alters what phase.complete MARKS is out of this seam's contract, not a refinement of it", + "line": 725 }, { "id": "PLANNING.PATH.PARITY.project-scope", "klass": "PLANNING", "value": ".planning/ (never .planning/projects/); mirror planning-workspace.cjs planningDir()", - "line": 705 + "line": 708 }, { "id": "PLANNING.PATH.SEAM.helpers", "klass": "PLANNING", "value": "helpers.planningPaths delegates to workspacePlanningPaths + resolveWorkspaceContext; precedence explicit-ws > env-ws > env-project > root", - "line": 706 + "line": 709 }, { "id": "PLANNING.PATH.SEAM.init-handlers", "klass": "PLANNING", "value": "[initExecutePhase, initPlanPhase, initPhaseOp, initMilestoneOp] consume helpers.planningPaths().planning (no direct relPlanningPath join)", - "line": 707 + "line": 710 }, { "id": "PR.3267.POSTMORTEM.recovery", "klass": "PR", "value": "[issue#3270 created, label approved-enhancement applied, PR reopened, body includes \"Closes #3270\", label no-changelog applied]", - "line": 682 + "line": 685 }, { "id": "PR.3267.POSTMORTEM.root-cause", "klass": "PR", "value": "[missing issue link, missing changeset/no-changelog]", - "line": 681 + "line": 684 }, { "id": "PRED.k320.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L193-211", - "line": 775 + "line": 786 }, { "id": "PRED.k320.ci-enforcement", "klass": "PRED", "value": "scripts/changeset/lint.cjs", - "line": 781 + "line": 792 }, { "id": "PRED.k320.ci-paths-monitored", "klass": "PRED", "value": "bin/ gsd-core/ src/ agents/ commands/ hooks/ sdk/src/ sdk/prompts/", - "line": 782 + "line": 793 }, { "id": "PRED.k320.cure", "klass": "PRED", "value": "drop .changeset/--.md fragment ONLY", - "line": 777 + "line": 788 }, { "id": "PRED.k320.evidence", "klass": "PRED", "value": "PR #3302 merge-conflict against #3308 CHANGELOG.md row 2026-05-09", - "line": 784 + "line": 795 }, { "id": "PRED.k320.opt-out-label", "klass": "PRED", "value": "no-changelog", - "line": 780 + "line": 791 }, { "id": "PRED.k320.recovery", "klass": "PRED", "value": "open Removed-typed cleanup PR deleting only the redundant row", - "line": 783 + "line": 794 }, { "id": "PRED.k320.rule", "klass": "PRED", "value": "do not edit CHANGELOG.md in feature/fix/enhancement PRs", - "line": 776 + "line": 787 }, { "id": "PRED.k320.signal", "klass": "PRED", "value": "changelog-direct-edit-forbidden", - "line": 774 + "line": 785 }, { "id": "PRED.k320.tool", "klass": "PRED", "value": "npm run changeset -- --type --pr --body \"...\"", - "line": 778 + "line": 789 }, { "id": "PRED.k320.types", "klass": "PRED", "value": "Added|Changed|Deprecated|Removed|Fixed|Security", - "line": 779 + "line": 790 }, { "id": "PRED.k321.evidence", "klass": "PRED", "value": "PRs #3304/#3305 (2026-05-09): real Minor/Major findings in body, 0 threads", - "line": 790 + "line": 801 }, { "id": "PRED.k321.poll-shape", "klass": "PRED", "value": "parse pulls//reviews body AND graphql reviewThreads", - "line": 788 + "line": 799 }, { "id": "PRED.k321.resolution", "klass": "PRED", "value": "address in code; no GraphQL resolveReviewThread needed for body-only findings", - "line": 789 + "line": 800 }, { "id": "PRED.k321.shape", "klass": "PRED", "value": "CR posts \"[!CAUTION] outside the diff\" findings in review BODY, not in reviewThreads", - "line": 787 + "line": 798 }, { "id": "PRED.k321.signal", "klass": "PRED", "value": "cr-outside-diff-range-finding", - "line": 786 + "line": 797 }, { "id": "PRED.k322.cure-1", "klass": "PRED", "value": "2nd retrigger ~10min after first ack", - "line": 795 + "line": 806 }, { "id": "PRED.k322.cure-2", "klass": "PRED", "value": "if silent at 50min, treat as silent-pass with maintainer flag in merge-commit body", - "line": 796 + "line": 807 }, { "id": "PRED.k322.distinct-from", "klass": "PRED", "value": "k080", - "line": 793 + "line": 804 }, { "id": "PRED.k322.evidence", "klass": "PRED", "value": "PR #3306 (2026-05-09): 0 reviews after 50min + 2 retriggers", - "line": 798 + "line": 809 }, { "id": "PRED.k322.merge-gate-impact", "klass": "PRED", "value": "k070 real_coderabbit_review_present unsatisfied; requires maintainer judgment", - "line": 797 + "line": 808 }, { "id": "PRED.k322.shape", "klass": "PRED", "value": "ack posted, real review never lands within [5s, 410s] cooldown after burst of N PRs <15min", - "line": 794 + "line": 805 }, { "id": "PRED.k322.signal", "klass": "PRED", "value": "cr-sustained-throttle", - "line": 792 + "line": 803 }, { "id": "PRED.k323.cure-alt", "klass": "PRED", "value": "consolidate into single PR when 2+ issues share root cause", - "line": 803 + "line": 814 }, { "id": "PRED.k323.cure-pre-dispatch", "klass": "PRED", "value": "brief one agent canonical-owner; brief others to EXCLUDE shared site", - "line": 802 + "line": 813 }, { "id": "PRED.k323.evidence", "klass": "PRED", "value": "#3300 (#3297) overlapped #3306 (#3298) on add-backlog.md hunks 2026-05-09", - "line": 805 + "line": 816 }, { "id": "PRED.k323.recovery", "klass": "PRED", "value": "close smaller PR as \"subsumed by #N\" or rebase second to drop overlap hunk", - "line": 804 + "line": 815 }, { "id": "PRED.k323.shape", "klass": "PRED", "value": "2+ open issues touch same canonical bug site; each fix's sibling-audit produces overlapping diff", - "line": 801 + "line": 812 }, { "id": "PRED.k323.signal", "klass": "PRED", "value": "sibling-audit-cross-pr-overlap", - "line": 800 + "line": 811 }, { "id": "PRED.k324.cure", "klass": "PRED", "value": "verify via gh api on every agent-completion notification; never trust narrative", - "line": 809 + "line": 820 }, { "id": "PRED.k324.evidence", "klass": "PRED", "value": "2026-05-09 session: 5+ mid-monitor terminations across PRs #3232/#3271/#3251/#3255/#3262", - "line": 811 + "line": 822 }, { "id": "PRED.k324.k095-restatement", "klass": "PRED", "value": "k095 confirmed shape: agent reports \"waiting for monitor\" / \"tests still running\" then terminates", - "line": 808 + "line": 819 }, { "id": "PRED.k324.poll-shape", "klass": "PRED", "value": "gh pr view --json mergeStateStatus,statusCheckRollup + pulls//reviews + graphql reviewThreads + issues//comments tail", - "line": 810 + "line": 821 }, { "id": "PRED.k324.signal", "klass": "PRED", "value": "agent-terminates-mid-monitor", - "line": 807 + "line": 818 }, { "id": "PRED.k325.cleanup", "klass": "PRED", "value": "git worktree remove --force for aged agent worktrees", - "line": 816 + "line": 827 }, { "id": "PRED.k325.cure", "klass": "PRED", "value": "detached-HEAD: git checkout --detach $(git ls-remote origin ); modify; commit; git push --force-with-lease=: origin HEAD:refs/heads/", - "line": 815 + "line": 826 }, { "id": "PRED.k325.evidence", "klass": "PRED", "value": "2026-05-09 CHANGELOG.md strip on PRs #3300/#3302/#3304/#3305 required detached-HEAD", - "line": 817 + "line": 828 }, { "id": "PRED.k325.shape", "klass": "PRED", "value": "git checkout errors \"already used by worktree at \"", - "line": 814 + "line": 825 }, { "id": "PRED.k325.signal", "klass": "PRED", "value": "worktree-branch-lock-on-force-push", - "line": 813 + "line": 824 }, { "id": "PRED.k326.cure", "klass": "PRED", "value": "quote canonical doc verbatim in brief; mentally simulate \"if all N agents follow this brief literally, do they violate any rule?\"", - "line": 821 + "line": 832 }, { "id": "PRED.k326.evidence", "klass": "PRED", "value": "2026-05-09 brief \"k040 — update CHANGELOG.md\" → 5 of 8 agents violated CONTRIBUTING.md L110", - "line": 822 + "line": 833 }, { "id": "PRED.k326.shape", "klass": "PRED", "value": "N parallel agents amplify a single brief-vs-doc contradiction into N violations", - "line": 820 + "line": 831 }, { "id": "PRED.k326.signal", "klass": "PRED", "value": "brief-contradicts-canonical-doc", - "line": 819 + "line": 830 }, { "id": "PRED.k327.ack-shape", "klass": "PRED", "value": "body \"✅ Actions performed - Full review triggered\"", - "line": 825 + "line": 836 }, { "id": "PRED.k327.cooldown-normal", "klass": "PRED", "value": "[5s, 410s]", - "line": 828 + "line": 839 }, { "id": "PRED.k327.cooldown-throttled", "klass": "PRED", "value": "k322", - "line": 829 + "line": 840 }, { "id": "PRED.k327.distinguish-key", "klass": "PRED", "value": "len(pulls//reviews) — ack=0, real=≥1", - "line": 827 + "line": 838 }, { "id": "PRED.k327.real-review-shape", "klass": "PRED", "value": "body starts \"Actionable comments posted: N\" OR \"[!CAUTION] Some comments are outside the diff\"", - "line": 826 + "line": 837 }, { "id": "PRED.k327.signal", "klass": "PRED", "value": "cr-ack-vs-real-review", - "line": 824 + "line": 835 }, { "id": "PRED.k328.audit-list", "klass": "PRED", "value": "[heading-matches-class, closing-keyword-present, changeset-fragment-or-no-changelog-label]", - "line": 834 + "line": 845 }, { "id": "PRED.k328.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L48,L64,L81 (template links) + .github/PULL_REQUEST_TEMPLATE/{fix,enhancement,feature}.md L1 (heading text)", - "line": 832 + "line": 843 }, { "id": "PRED.k328.k100-restatement", "klass": "PRED", "value": "heading must match issue class: bug→## Fix PR, enhancement→## Enhancement PR, feature→## Feature PR", - "line": 833 + "line": 844 }, { "id": "PRED.k328.signal", "klass": "PRED", "value": "pr-template-typed-heading-required", - "line": 831 + "line": 842 }, { "id": "PRED.k329.body", "klass": "PRED", "value": "**** — . (#)", - "line": 840 + "line": 851 }, { "id": "PRED.k329.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L196-202 + .changeset/README.md", - "line": 837 + "line": 848 }, { "id": "PRED.k329.filename", "klass": "PRED", "value": ".changeset/--.md", - "line": 838 + "line": 849 }, { "id": "PRED.k329.frontmatter", "klass": "PRED", "value": "---\\\\ntype: \\\\npr: \\\\n---", - "line": 839 + "line": 850 }, { "id": "PRED.k329.observed-clean", "klass": "PRED", "value": "#3299 sunny-ibex-wave, #3301 sturdy-rams-caper, #3306 3298-phase-dir-prefix-drift-workflows", - "line": 841 + "line": 852 }, { "id": "PRED.k329.signal", "klass": "PRED", "value": "changeset-fragment-canonical-shape", - "line": 836 + "line": 847 }, { "id": "PRED.k330.fallback", "klass": "PRED", "value": "append predicate-format findings directly to CONTEXT.md", - "line": 845 + "line": 856 }, { "id": "PRED.k330.shape", "klass": "PRED", "value": "mempalace MCP tools require explicit user call; AI cannot trigger", - "line": 844 + "line": 855 }, { "id": "PRED.k330.signal", "klass": "PRED", "value": "mempalace-diary-not-callable-by-ai", - "line": 843 + "line": 854 }, { "id": "PRED.k331.cure", "klass": "PRED", "value": "gh pr close with NO --comment flag", - "line": 850 + "line": 861 }, { "id": "PRED.k331.evidence", "klass": "PRED", "value": "2026-05-09 wave-3: violation on #3300 close, deleted within 30s", - "line": 852 + "line": 863 }, { "id": "PRED.k331.k101-restatement", "klass": "PRED", "value": "k101 includes close-time --comment flag; rationale belongs in subsuming PR's squash-merge body", - "line": 849 + "line": 860 }, { "id": "PRED.k331.recovery", "klass": "PRED", "value": "if violation lands, gh api -X DELETE repos///issues/comments/", - "line": 851 + "line": 862 }, { "id": "PRED.k331.shape", "klass": "PRED", "value": "instruction \"close with no comment (rationale)\" — parenthetical is rationale, NOT comment body", - "line": 848 + "line": 859 }, { "id": "PRED.k331.signal", "klass": "PRED", "value": "close-with-no-comment-is-literal", - "line": 847 + "line": 858 }, { "id": "PROBE.ci.surface", "klass": "PROBE", "value": "the contract (parse/validate, projection round-trip, fail-closed guards), NEVER the LLM judgment (ADR-550 D5)", - "line": 592 + "line": 595 }, { "id": "PROBE.core.seam", "klass": "PROBE", "value": "analyzeCoverage(items,resolutions?,validators) ingests ALREADY-proposed items; does NOT assume deterministic propose (ADR-550 D7b)", - "line": 585 + "line": 588 }, { "id": "PROBE.edge.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 587 + "line": 590 }, { "id": "PROBE.family", "klass": "PROBE", "value": "edge-probe(shape-axis)+prohibition-probe(must-NOT-axis)+ui-consideration-probe(UI-state-axis), shared probe-core, run as spec-phase/ui-phase soft gates (ADR-550 D7; #1867)", - "line": 583 + "line": 586 }, { "id": "PROBE.item.axes", "klass": "PROBE", "value": "status{resolved|dismissed|unresolved} x verification{|null} — orthogonal; the lifecycle enum carries no verification fact (ADR-550 D7a)", - "line": 586 + "line": 589 }, { "id": "PROBE.principle", "klass": "PROBE", "value": "verifier-reach-equals-spec-reach (a goal-backward verifier only checks assertions that exist; probes make omitted assertions exist before code) — ADR-857 verification-substrate boundary; docs/design/verifier-reach.md", - "line": 582 + "line": 585 }, { "id": "PROBE.prohib.verification", "klass": "PROBE", "value": "test|judgment", - "line": 588 + "line": 591 }, { "id": "PROBE.protocol", "klass": "PROBE", "value": "recall(adversarial over-generate)->precision(drop routine-engineering); dismissals require a non-empty reason", - "line": 584 + "line": 587 }, { "id": "PROBE.ui.axis", "klass": "PROBE", "value": "MIXED — closed compiled shape-rooted 8 (empty/loading/error/populated/partial/overflow/zero-one-many/long-text) via ui-consideration-probe adapter; open UX (real-time/a11y/i18n-RTL) prose-owned in references/domain-probes.md, NOT compiled (#1867)", - "line": 590 + "line": 593 }, { "id": "PROBE.ui.seam", "klass": "PROBE", "value": "ui-phase Step 9.5 post-verification: element-cue classify -> propose-then-confirm (partial-cue mitigation, Goodhart) -> autoResolve --auto floor (never dismiss; unclassified stays unresolved #1110) -> ## UI Considerations write-back -> plan-phase `## UI Considerations` lift rule (#1867)", - "line": 591 + "line": 594 }, { "id": "PROBE.ui.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 589 + "line": 592 }, { "id": "PROC.AGENT-DISPATCH.completion-verify", "klass": "PROC", "value": "run k324.poll-shape on every agent-completion notification", - "line": 856 + "line": 867 }, { "id": "PROC.AGENT-DISPATCH.parallel-overlap-audit", "klass": "PROC", "value": "before dispatching N sibling-audit fixers, compute file-set union and assign canonical owners", - "line": 855 + "line": 866 }, { "id": "PROC.AGENT-DISPATCH.preflight", "klass": "PROC", "value": "[read-CONTRIBUTING.md-fresh, read-relevant-ADRs, cite-specific-line-in-brief, require-closing-keyword, require-changeset-fragment, forbid-CHANGELOG.md-edit, require-isolation-worktree, forbid-self-PR-comment, mandate-trust-but-verify]", - "line": 854 + "line": 865 }, { "id": "PROC.MERGE-WAVE.changelog-strip-pattern", "klass": "PROC", "value": "detached-HEAD per k325 + git checkout main -- CHANGELOG.md + commit + force-with-lease", - "line": 860 + "line": 871 }, { "id": "PROC.MERGE-WAVE.merge-tool", "klass": "PROC", "value": "gh pr merge --squash --delete-branch", - "line": 861 + "line": 872 }, { "id": "PROC.MERGE-WAVE.merge-tool-warning", "klass": "PROC", "value": "delete-branch may fail with \"used by worktree at\" — harmless; remote branch still deleted", - "line": 862 + "line": 873 }, { "id": "PROC.MERGE-WAVE.ordering", "klass": "PROC", "value": "[wave1: isolated-files, wave2: CHANGELOG-only-overlap (better: strip per k320), wave3: same-file-overlap with explicit decision]", - "line": 858 + "line": 869 }, { "id": "PROC.MERGE-WAVE.preflight", "klass": "PROC", "value": "gh pr view --json files for every PR; identify overlap pairs; surface to maintainer", - "line": 859 + "line": 870 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.observed", "klass": "PROC", "value": "#3541 + #3542 dispatched simultaneously this session; PRs #3546 #3547 opened green; one syntax slip caught by AGENT-RETIRED-SLASH-SYNTAX-DRIFT and fixed before second PR opened", - "line": 941 + "line": 952 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.pattern", "klass": "PROC", "value": "bot triage brief → worktree per branch → parallel sub-agents do rubber-duck/RCA/TDD implementation only → top-level orchestrator owns commit + gsd-test + push + PR + changeset-pr-backfill", - "line": 939 + "line": 950 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.rationale", "klass": "PROC", "value": "long-running test runs need cross-turn notifications (orchestrator-only); CONTRIBUTING.md gh-templates-first hook requires session-scoped Read calls sub-agents wouldn't otherwise make; sequencing test runs avoids GSD-TEST-CONCURRENT-OUTPUT-COLLISION", - "line": 940 + "line": 951 }, { "id": "PROC.TRIAGE.comment-shape", "klass": "PROC", "value": "lead with \"duplicate of #NNNN, fixed by PR #MMMM, in v1.X.Y\"; show current code snippet proving bug-surface gone; give @latest and @next upgrade commands; close", - "line": 944 + "line": 955 }, { "id": "PROC.TRIAGE.no-duplicate-label", "klass": "PROC", "value": "this repo has no duplicate label; framing lives in comment text + closing the issue", - "line": 945 + "line": 956 }, { "id": "PROC.TRIAGE.routing-incoming", "klass": "PROC", "value": "stale-bug-already-fixed to close as duplicate of originating issue + cite fix PR + first stable tag; release-publish-or-backport to ready-for-human; reporter-can-self-test to awaiting-retest", - "line": 943 + "line": 954 }, { "id": "PROHIB.canon-referral", "klass": "PROHIB", "value": "OWASP/GDPR/fairness-canon are REFERRED to /gsd:secure-phase+eslint, never minted as prohibitions (ADR-550 D6)", - "line": 594 + "line": 597 }, { "id": "PROHIB.descriptor.shape", "klass": "PROHIB", "value": "5 FLAT scalars (check_kind,check_target,check_rule,check_violation_fixture,check_clean_fixture) — NEVER a nested check:{} (parseMustHavesBlock is a flat parser, src/frontmatter.cts)", - "line": 599 + "line": 602 }, { "id": "PROHIB.enforce.adr", "klass": "PROHIB", "value": "docs/adr/1606-prohibition-enforcement-verify-seam.md (verify-time enforcement seam) + docs/adr/550-spec-phase-probe-contract.md (spec-phase contract)", - "line": 602 + "line": 605 }, { "id": "PROHIB.enforce.causation", "klass": "PROHIB", "value": "clean-fixture control proves the red is content-caused not env-var-set; MANDATORY for node-test (#1906 supersedes #1346 opt-in) — absent clean-fixture ⇒ node-test un-provable/fail-closed; lint-rule needs none (its subject IS the linted file)", - "line": 598 + "line": 601 }, { "id": "PROHIB.enforce.failfirst", "klass": "PROHIB", "value": "MACHINE-PROVEN against an author-supplied violation fixture (#1279); caller failFirst attestation DEMOTED to a non-authoritative hint (FF-08)", - "line": 597 + "line": 600 }, { "id": "PROHIB.enforce.green-rule", "klass": "PROHIB", "value": "passed iff provenFailFirst===true && run.passed===true (runProhibitionEnforcement); every miss/fail/un-provable HARD-GATES both modes via dispositionForProhibition's fail-closed default", - "line": 595 + "line": 598 }, { "id": "PROHIB.enforce.kinds", "klass": "PROHIB", "value": "node-test (non-vacuous red via isNonVacuousNodeTestRed; pass-side vacuity via isNonVacuousNodeTestPass) | lint-rule (eslint --format json filtered by ruleId)", - "line": 596 + "line": 599 }, { "id": "PROHIB.judgment-tier", "klass": "PROHIB", "value": "never-silent / never-hard-halt soft gate; autonomous emits \"unverified-prohibition — human review recommended\" (exogenous grading, ADR-550 D4)", - "line": 601 + "line": 604 }, { "id": "PROHIB.rail", "klass": "PROHIB", "value": "core verify rail, non-toggleable (ADR-857 verification-substrate boundary / decision #6); the verifier<->predicate contract is NOT an off-by-default capability", - "line": 600 + "line": 603 }, { "id": "PROHIB.recall", "klass": "PROHIB", "value": "LLM-prose; no compiled prohibition-probe recall engine (only the schema/projection layer is code, ADR-550 D7b)", - "line": 593 + "line": 596 }, { "id": "RELEASE-NOTES.ANTI-PATTERN", "klass": "RELEASE-NOTES", "value": "raw \"What's Changed\" PR list as final body for hotfix or feature release; \"Full Changelog only\" body for tagged release with >0 user-facing fixes", - "line": 751 + "line": 762 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.implementation-first", "klass": "RELEASE-NOTES", "value": "do not lead bullet with file path or function name; lead with symptom/user-visible behavior", - "line": 752 + "line": 763 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.risk-commentary", "klass": "RELEASE-NOTES", "value": "do not include \"may break\", \"be careful\", \"test thoroughly\" - release notes state what changed, not hedges about what might go wrong", - "line": 753 + "line": 764 }, { "id": "RELEASE-NOTES.DEFAULT-STATE", "klass": "RELEASE-NOTES", "value": "auto-generated body is \"What's Changed\" PR list + Full Changelog link; treat as draft, not final", - "line": 727 + "line": 738 }, { "id": "RELEASE-NOTES.EXAMPLE.hotfix", "klass": "RELEASE-NOTES", "value": "v1.41.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.41.1) - 14 fixes grouped by 6 subgroups", - "line": 755 + "line": 766 }, { "id": "RELEASE-NOTES.EXAMPLE.minor-auto-acceptable", "klass": "RELEASE-NOTES", "value": "v1.41.0 - kept auto-generated body; many small fixes with clean conventional-commit titles", - "line": 757 + "line": 768 }, { "id": "RELEASE-NOTES.EXAMPLE.rc", "klass": "RELEASE-NOTES", "value": "v1.7.0-rc.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.7.0-rc.1) - intro + Added/Changed/Fixed/Documentation taxonomy", - "line": 756 + "line": 767 }, { "id": "RELEASE-NOTES.GATE.hotfix", "klass": "RELEASE-NOTES", "value": "manual edit required; auto-generated body for vX.Y.{Z>0} is \"Full Changelog only\" and must be replaced with structured body", - "line": 728 + "line": 739 }, { "id": "RELEASE-NOTES.GATE.minor", "klass": "RELEASE-NOTES", "value": "auto-generated body acceptable when PR titles are clean; promote to structured body when >20 PRs or contains feature+refactor+fix mix", - "line": 730 + "line": 741 }, { "id": "RELEASE-NOTES.GATE.rc", "klass": "RELEASE-NOTES", "value": "manual edit recommended; auto-generated PR list is acceptable for early RCs but final RC before vX.Y.0 should match standard", - "line": 729 + "line": 740 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.main-branch", "klass": "RELEASE-NOTES", "value": "next (RCs) + latest (stable); install via @next or @latest", - "line": 762 + "line": 773 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.rule", "klass": "RELEASE-NOTES", "value": "streams do not mix; do not document @next in hotfix/stable notes", - "line": 763 + "line": 774 }, { "id": "RELEASE-NOTES.SCOPE", "klass": "RELEASE-NOTES", "value": "GitHub Releases body for tags vX.Y.Z, vX.Y.Z-rc.N; not CHANGELOG.md (changeset workflow owns that)", - "line": 726 + "line": 737 }, { "id": "RELEASE-NOTES.SOURCE.changesets", "klass": "RELEASE-NOTES", "value": ".changeset/*.md (frontmatter pr: + body bullets)", - "line": 742 + "line": 753 }, { "id": "RELEASE-NOTES.SOURCE.commits", "klass": "RELEASE-NOTES", "value": "git log .. --pretty=format:'%s%n%n%b' --no-merges", - "line": 741 + "line": 752 }, { "id": "RELEASE-NOTES.SOURCE.pr-bodies", "klass": "RELEASE-NOTES", "value": "gh pr view --json title,body for fixes lacking a changeset", - "line": 743 + "line": 754 }, { "id": "RELEASE-NOTES.SOURCE.precedence", "klass": "RELEASE-NOTES", "value": "changeset body > commit body > PR body > commit subject (prefer authored content over auto-generated)", - "line": 744 + "line": 755 }, { "id": "RELEASE-NOTES.STANDARD.bullet-shape", "klass": "RELEASE-NOTES", "value": "**Bold user-visible change** — explanation of what was broken or what's new, leading with symptom not implementation. Trailing (#NNN) PR ref.", - "line": 734 + "line": 745 }, { "id": "RELEASE-NOTES.STANDARD.footer.full-changelog", "klass": "RELEASE-NOTES", "value": "**Full Changelog**: https://github.com/open-gsd/gsd-core/compare/...", - "line": 738 + "line": 749 }, { "id": "RELEASE-NOTES.STANDARD.footer.hotfix", "klass": "RELEASE-NOTES", "value": "Install/upgrade: \\`npx @opengsd/gsd-core@latest\\`", - "line": 736 + "line": 747 }, { "id": "RELEASE-NOTES.STANDARD.footer.rc", "klass": "RELEASE-NOTES", "value": "Install for testing: \\`npx @opengsd/gsd-core@next\\` (per branch->dist-tag policy)", - "line": 737 + "line": 748 }, { "id": "RELEASE-NOTES.STANDARD.heading-level", "klass": "RELEASE-NOTES", "value": "## for category, ### for subgroup (area), - for bullet", - "line": 733 + "line": 744 }, { "id": "RELEASE-NOTES.STANDARD.intro", "klass": "RELEASE-NOTES", "value": "optional one-paragraph framing for RC/feature releases; omit for pure-fix hotfixes", - "line": 739 + "line": 750 }, { "id": "RELEASE-NOTES.STANDARD.subgroups", "klass": "RELEASE-NOTES", "value": "phase-planning-state | workstream | query-dispatch-cli | code-review | install | capture | docs | architecture | security", - "line": 735 + "line": 746 }, { "id": "RELEASE-NOTES.STANDARD.taxonomy", "klass": "RELEASE-NOTES", "value": "Keep-a-Changelog 1.1.0: Added | Changed | Deprecated | Removed | Fixed | Security | Documentation", - "line": 732 + "line": 743 }, { "id": "RELEASE-NOTES.TEMPLATE.hotfix", "klass": "RELEASE-NOTES", "value": "## Fixed\\n\\n### \\n- **** — . (#)\\n\\n---\\n\\nInstall/upgrade: \\`npx @opengsd/gsd-core@latest\\`\\n\\n**Full Changelog**: ", - "line": 759 + "line": 770 }, { "id": "RELEASE-NOTES.TEMPLATE.rc", "klass": "RELEASE-NOTES", "value": "\\n\\n## Added\\n### \\n- **** — . (#)\\n\\n## Changed\\n### Architecture\\n- **** — . (#)\\n\\n## Fixed\\n### \\n- **** — . (#)\\n\\n## Documentation\\n- **** — . (#)\\n\\n---\\n\\nThis is a release candidate. Install for testing:\\n\\`\\`\\`bash\\nnpx @opengsd/gsd-core@next\\n\\`\\`\\`\\n\\n**Full Changelog**: ", - "line": 760 + "line": 771 }, { "id": "RELEASE-NOTES.WORKFLOW.edit", "klass": "RELEASE-NOTES", "value": "gh release edit --notes-file ", - "line": 746 + "line": 757 }, { "id": "RELEASE-NOTES.WORKFLOW.idempotency", "klass": "RELEASE-NOTES", "value": "gh release edit overwrites body wholesale; safe to re-run after refining", - "line": 749 + "line": 760 }, { "id": "RELEASE-NOTES.WORKFLOW.token", "klass": "RELEASE-NOTES", "value": "must use .envrc GITHUB_TOKEN per RULESET.GH.AUTH.DEFAULT (this doc); never ambient gh auth", - "line": 748 + "line": 759 }, { "id": "RELEASE-NOTES.WORKFLOW.view", "klass": "RELEASE-NOTES", "value": "gh release view --json body --jq .body", - "line": 747 + "line": 758 }, { "id": "RULESET.ADR-HEADER", "klass": "RULESET", "value": "every docs/adr/NNNN-*.md must open with - **Status:** Accepted|Proposed|Superseded (by [ADR-NNNN](file.md))|Legacy + - **Date:** YYYY-MM-DD immediately after title", - "line": 649 + "line": 652 }, { "id": "RULESET.AGENT_SIZE_BUDGET", "klass": "RULESET", "value": "agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4, same mechanism and same `Emitted-Drift-Ack-Growth:` commit trailer (ADR-3942, superseding ADR-2719 §3's fragment model) as WORKFLOW_SIZE_BUDGET, scoped to agents/gsd-*.md) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). Sizes are measured via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter (tests/helpers/emitted-runtime.cjs's currentSizes() and the guard's own tier-cap checks both import it). A grown agent fails the differential guard — ack + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes. The prior per-file baseline (tests/agent-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724", - "line": 638 + "line": 641 }, { "id": "RULESET.ALLOWED-TOOLS-FRONTMATTER", "klass": "RULESET", "value": "command's allowed-tools must cover every tool the workflow calls (including Write for file creation); thin-wrapper pattern makes this easy to miss", - "line": 645 + "line": 648 }, { "id": "RULESET.ARGUMENTS-SANITIZE", "klass": "RULESET", "value": "any workflow step constructing .planning/.../{SLUG}.md path from user input ($ARGUMENTS, parsed remainder) must sanitize inline ([a-z0-9-] only, reject ..//\\\\, max-length) — \"(already sanitized)\" must trace back to explicit guard; RESUME/fallback modes need own guards", - "line": 646 + "line": 649 }, { "id": "RULESET.AUDIT.search-source-not-generated", "klass": "RULESET", "value": "verify an invariant/validation EXISTS by searching the AUTHORED source (src/*.cts OR the scripts/gen-*.cjs generator), never the generated bin/lib/*.cjs (gitignored, ADR-457); gen-time checks live in gen-*.cjs not the .cts it consumes → search BOTH before declaring absent; read generated .cjs only for output drift. Repro: grep src/*.cts for VALID_CONVERTER_NAMES → false \"5e ConverterName unenforced\"; actually enforced in gen-capability-registry.cjs. cf RULESET.TESTS.no-source-grep", - "line": 634 + "line": 637 }, { "id": "RULESET.CAPABILITY.cutover-self-gating", "klass": "RULESET", "value": "a phase-6 per-feature cutover moves the host's phase-context detection + mode/flag logic INTO the skill (self-gating, per ADR-894); the loop hook is intentionally COARSE — \"invoke skill X at point Y when config Z\" — and carries no detection/mode. WORKED EXAMPLE: plan-phase.md §5.6 UI gate (frontend-detection via ui-safety-gate.cjs + --auto/manual branch + --skip-ui bypass) must move into gsd-ui-phase before its plan:pre hook can replace the inline call without behavior loss. Spike #1018 finding.", - "line": 408 + "line": 411 }, { "id": "RULESET.CAPABILITY.off-means-off", "klass": "RULESET", "value": "the host derives shared outputs from the ACTIVE hook set (via loop.render-hooks); a hook may ADD a labeled block or be COUNTED into a host-computed aggregate (e.g. a score denominator), but NEVER mutates host source — so a disabled capability yields the base output by construction, not by authoring discipline. Ratify in ADR-894; proven by spike #1018.", - "line": 406 + "line": 409 }, { "id": "RULESET.CAPABILITY.precedence-engine-single-owner", "klass": "RULESET", "value": "the config-key four-level precedence walk (loadConfig result → workstream config.json → root config.json → registry.configSchema default → absent) is owned solely by src/capability-activation.cts: raw-value primitive resolveConfigKey(dotKey, {config,cwd,registry}) and boolean wrapper _resolveActivationValue(dotKey,config,cwd,registry); loop-resolver.cts imports the engine (no duplicate); resolveConfigValues in loop-resolver.cts delegates to resolveConfigKey; resolveCapabilityRuntimeState does NOT return registry/config — callers import capability-registry.cjs and call loadConfig(cwd) directly.", - "line": 412 + "line": 415 }, { "id": "RULESET.CAPABILITY.step-additive-gate-blocks", "klass": "RULESET", "value": "a `step` hook is purely additive (invoke skill + produce artifacts, NEVER halts the host); host-blocking preconditions are `gate`s (blocking:true, onError:halt); runtime/mode context (auto/chain vs manual) self-gates IN THE SKILL, not via `when` (config-only). §5.6 = plan:pre step (ui-phase; skill self-gates on frontend+pipeline, auto-fires only in pipelines) + a NEW plan:pre gate (frontend-and-no-UI-SPEC → halt, when:workflow.ui_safety_gate); the loop.render-hooks dispatch template handles steps AND gates. Resolves #1022.", - "line": 410 + "line": 413 }, { "id": "RULESET.CODERABBIT.GUARD.COMPLETE", "klass": "RULESET", "value": "required_checks_green && coderabbit_check_pass && graphQL(reviewThreads.unresolved_count)==0", - "line": 671 + "line": 674 }, { "id": "RULESET.CODERABBIT.GUARD.GRAPHQL", "klass": "RULESET", "value": "reviewThreads(first:100){nodes{id isResolved comments{nodes{author body path line originalLine url}}}}; use unresolved threads as authoritative, not badge text alone", - "line": 672 + "line": 675 }, { "id": "RULESET.CODERABBIT.GUARD.OPEN_PRS", "klass": "RULESET", "value": "gh pr list --repo open-gsd/gsd-core --author @me --state open; repeat near end because open PR set can change mid-run", - "line": 670 + "line": 673 }, { "id": "RULESET.CODERABBIT.GUARD.RERUN", "klass": "RULESET", "value": "after every push wait for CodeRabbit completion, then re-query unresolved threads; CodeRabbit can add new findings after earlier threads were resolved", - "line": 673 + "line": 676 }, { "id": "RULESET.CODERABBIT.GUARD.RESOLVE", "klass": "RULESET", "value": "fix validated finding -> focused tests -> commit/push -> resolveReviewThread(threadId) -> wait CI/CodeRabbit -> final unresolved_count query", - "line": 674 + "line": 677 }, { "id": "RULESET.CODERABBIT.GUARD.SCOPE", "klass": "RULESET", "value": "if a new @me open PR appears during final list, include it in the same guard pass before declaring all-open-PRs complete", - "line": 675 + "line": 678 }, { "id": "RULESET.CONTENT-PATH-NORMALIZATION", "klass": "RULESET", "value": "filesystem paths substituted into markdown body text (@-references, workflow .md, agent .md, generated docs, command bodies) MUST be normalized to POSIX forward slashes via .replace(/\\\\/g,'/') at the production source BEFORE substitution; never push normalization to tests; cross-platform content is POSIX-only; applies to: computePathPrefix output, install-path rewrites, generated shim paths emitted into .md bodies; idempotent on POSIX so unconditional; mechanically enforced by local/normalize-path-in-content (eslint, src/**/*.cts; #1733)", - "line": 878 + "line": 889 }, { "id": "RULESET.CONTRIB.CLASSIFY.enhancement", "klass": "RULESET", "value": "requires approved-enhancement before implementation", - "line": 664 + "line": 667 }, { "id": "RULESET.CONTRIB.CLASSIFY.feature", "klass": "RULESET", "value": "requires approved-feature before implementation", - "line": 665 + "line": 668 }, { "id": "RULESET.CONTRIB.CLASSIFY.fix", "klass": "RULESET", "value": "requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)", - "line": 663 + "line": 666 }, { "id": "RULESET.CONTRIB.GATE.ORDER", "klass": "RULESET", "value": "issue-first -> approval-label -> code -> PR-link -> changeset/no-changelog", - "line": 662 + "line": 665 }, { "id": "RULESET.CR-THREAD-RESOLVE", "klass": "RULESET", "value": "after adding // allow-test-rule: to silence lint, resolve existing inline CR threads via graphql resolveReviewThread mutation before merge — open threads mislead future reviewers; pattern: gh api graphql -f query='mutation { resolveReviewThread(input:{threadId:\"PRRT_...\"}) { thread { isResolved } } }'", - "line": 656 + "line": 659 }, { "id": "RULESET.EMITTED_ATTRIBUTION", "klass": "RULESET", "value": "the emitted-artifact family (ADR-2719, epic #2719) — POST-CUTOVER (#2724, Phase 4). Historically tests/fixtures/golden-install-parity/*.json (19 path→hash manifests) + tests/workflow-size-baseline.json + tests/agent-size-baseline.json were all committed, PURE FUNCTIONS of the source tree whose correct merge was ALWAYS \"recompute\" — 140 of 143 conflicted-file instances across the open PR queue were these files. #2724 DELETES all three, the golden test (tests/golden-install-parity.test.cjs), the generator (scripts/gen-golden-install-parity-zcode.cjs), `npm run gen:golden`, `UPDATE_GOLDEN`, the merge-driver bridge (scripts/git-merge-regen-driver.cjs, `npm run setup:merge-driver`, the .gitattributes merge=gsd-regen block), and scripts/update-size-baseline.cjs (`npm run size:baseline`). The differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the SOLE gate for emitted-artifact propagation AND size growth — no committed artifact, nothing to hand-merge, nothing to regenerate. `npm run regen:derived` still exists for what remains committed and derived: build, registry, ADR index, capability matrix, inventory manifest, manifest versions, and `tests/fixtures/install-tree/*.json` (now `npm run gen:install-tree`, folded into `regen:derived`). tests/fixtures/install-tree/*.json is DELIBERATELY EXCLUDED from the cutover (ADR-2719 §7): it conflicts on 0 of 7, its diffs are readable, and it preserves \"the installer stopped shipping X\" as a hard absolute failure — capturing it would convert that absolute into an attribution-free auto-resolve. The baseline the differential compares against is now published by `scripts/gen-emitted-baseline.cjs` on every push to `next` (cached, keyed on sha) and restored in PR lanes via `GSD_EMITTED_BASELINE`/`resolveBaseline()` (tests/helpers/emitted-baseline.cjs); a cache miss falls back to an in-job build via a throwaway `git worktree` (tests/helpers/emitted-runtime.cjs's `buildBaselineAtRef`). REMEDIATION IS PART OF THE GATE (#2778): the failure output names its own remedy, because a gate that states a requirement and withholds the means of satisfying it is a maintainer round-trip, not a gate — ADR-2719 §3's \"conspicuous declaration\" only works if the contributor can discover how to make it. Both failing branches name the commit trailer to add — `Emitted-Drift-Ack-Hash:` or `Emitted-Drift-Ack-Growth:` (ADR-3942) — print its exact grammar (` — `, key and reason split on the FIRST em dash), and repeat \"do NOT regenerate anything\" — post-#2724 there is nothing left to regenerate, and hunting for a deleted baseline is the predictable wrong guess. The two branches key on DIFFERENT, now STRUCTURALLY DISTINCT trailer key spaces (separate maps since ADR-3942, closing a latent defect where a growth key could satisfy a hash lookup by naming coincidence and vice versa) and each says which: the hash pass keys on the EMITTED PATH (always contains a `/`, `Emitted-Drift-Ack-Hash:`), the size ratchet keys on the BARE FILENAME (`Emitted-Drift-Ack-Growth:`; `currentSizes` writes `sizes[entry.name]` from readdirSync over `gsd-core/workflows/` + `agents/`). A stale-ack failure additionally says to drop the trailer line (amending the commit) when removing its last entry, since a lingering unused trailer signals nothing; post-#2789 it also offers CORRECTING the reason to name the ripple actually made, which is the other honest resolution and the one a contributor usually wants. NOT ack-able and deliberately given no ack text: the `NEW_FILE_CAP` branch, whose remedy is extraction. Text is sourced from one frozen `REMEDIATION` export in tests/helpers/emitted-diff.cjs, whose example line is rendered via `renderAckTrailer` (`: — `, ADR-3942) so the taught grammar cannot drift from what `parseAckTrailers` actually accepts (a round-trip test feeds the printed line back through the parser); a key that is reserved (`__proto__`/`constructor`/`prototype`) or contains `<`, `>`, or whitespace is rejected loudly, and a doc example like ` — ` must never parse as a real declaration. Note the ADR's Consequences originally called the #2724 migration \"terminal\"; #2778 corrected that — it is terminal only for a PR that grows no shipped file. The ack was PR-lifetime data kept in permanent, shared, merge-path state, and each fix generated the next defect until ADR-3942 moved it off the tree entirely (see `### Emitted Artifact Provenance`): the single shared `tests/emitted-drift-ack.json` was a guaranteed merge-conflict cell (#2789; 5 of 6 conflicting PRs in one open queue collided on it and nothing else); #2914 replaced it with per-PR fragments under `tests/emitted-drift-acks/` — the `.changeset/` shape — ending the FILE conflict but not the KEY conflict, since two sources could never name the same path; #3078 found a fully-spent fragment left on `next` still walled off every key it owned (measured at the sweep: 45 fragments owning 403 paths, up from 13/272 at triage 19 days earlier) and added the post-merge-only `guard-no-ack-on-next` job plus a manual sweep; #3842's hand sweep handed three in-flight external PRs a `modify/delete` conflict each; #3823's hand-authored sweep, computed at branch time against a guard that evaluates at merge time, lost the race to a fragment merged mid-flight and left `next` red for 24 consecutive pushes; #3875's timed sweeper workflow automated the remedy but could not merge its own PRs (three independent, deterministic defects — bad conventional-title match, wrong CI-lane classification, no auto-merge path). ADR-3942 ends the chain: the escape hatch is now a commit trailer scoped to the PR's own commits, so there is no shared file, no shared key namespace, and nothing to sweep — the fragment directory, the next-lane guard job, the scheduled sweep workflow and the standalone ack linter are all DELETED (named by ROLE rather than by filename on purpose: a backticked path here asserts a LIVE repo path and `check-glossary-refs.cjs` fails on one that does not exist, while `lint-removed-but-needed.cjs` additionally fails on a deleted file's bare BASENAME appearing anywhere it scans — and this predicate's generated projection lands in docs/, which it does scan. ADR-3942 carries the exact paths; it sits under docs/adr/, which that guard exempts as a historical record). cf `RULESET.WORKFLOW_SIZE_BUDGET`, `RULESET.AGENT_SIZE_BUDGET`; see `### Emitted Artifact Provenance`", - "line": 639 + "line": 642 }, { "id": "RULESET.GENERATIVE-FIX", "klass": "RULESET", "value": "parallel implementations diverge silently when no parity test enforces equality at the test layer; for any new constant/array/parser shared between two parallel surfaces (two workflow surfaces, or a generated artifact and its hand-authored source), the same commit MUST add a parity assertion that fails when the two diverge; exemplar: tests/runtime-launcher-parity.test.cjs (asserts every workflow bash block uses the canonical gsd_run launcher)", - "line": 876 + "line": 887 }, { "id": "RULESET.GH.AUTH.DEFAULT", "klass": "RULESET", "value": "source .envrc GITHUB_TOKEN before gh; exception=ambient allowed only when user explicitly says machine-only fallback", - "line": 669 + "line": 672 }, { "id": "RULESET.HARNESS.test-memory-guard", "klass": "RULESET", "value": "~/.claude/hooks/test-memory-guard.sh fires on every Bash PreToolUse; if argv[0]∈{node|vitest|jest|mocha|tsx|ts-node|tap|ava|playwright|cypress} OR matches (npm|pnpm|yarn|bun) (run )?(t|test|tests|vitest|jest); blocks via hookSpecificOutput.permissionDecision=deny when sum(RSS of running matching procs, excluding tsserver|*-mcp|claude|Electron|...) ≥ 4 GiB OR when argv[0] basename matches a running process's argv[0]. Exception: node --version|-v|--help|-h|-p|-e are trivial probes and skip the check. Designed for a 24 GB Mac where prior accidental fan-out exhausted RAM", - "line": 920 + "line": 931 }, { "id": "RULESET.MANIFEST-CANONICAL-KEY", "klass": "RULESET", "value": "docs/INVENTORY-MANIFEST.json has a single top-level key: families; ALL EIGHT families.* arrays (agents/commands/workflows/references/cli_modules/hooks flat, plus workflow_modes/workflow_steps nested — #2996, epic #1671 Phase 6.5) are canonical, consumed by test suites — tests/inventory-manifest-sync.test.cjs reads all eight, edit-phase/enh-2380/enh-2430 tests read commands+workflows; the six flat families are keyed by BARE BASENAME while the two nested families are keyed by // path, deliberately, because two workflows may each own a same-named step file and a basename key would silently drop one under a JSON-equality comparison; recursion is bounded at exactly one named subdirectory, never a general walk; the family tables live ONCE in scripts/gen-inventory-manifest.cjs and are IMPORTED by the test (the test formerly redeclared them, a DEFECT.GENERATIVE-FIX divergence that let a new family be verified by nobody while still reporting green); the old generated date field and the stale top-level workflows key are both gone; regen via node scripts/gen-inventory-manifest.cjs --write, AFTER build:lib; #3762 added the ROSTER half — tests/inventory-manifest-sync.test.cjs now also asserts every manifest entry has a hand-written row in docs/INVENTORY.md, via the pure matcher in tests/helpers/inventory-roster.cjs. Scope is the SIX FLAT families only, each searched inside its own `## ` section; workflow_steps/workflow_modes are DELIBERATELY exempt because docs/INVENTORY.md §\"Workflow Sub-Files\" is a shipped decision that they carry no hand-written per-file rows. Matching is whole-CELL-exact (never substring — the rostered host-integration-adapters/imperative-hook-bus.cjs must not satisfy the separate top-level hook-bus.cjs) and section-scoped (smart-entry.md and smart-entry.cjs are different families), EXCEPT commands, which match on the row's Source-column link to ../commands/gsd/.md because the six ns-* namespace routers deliberately RENDER a name that is not their file stem (/gsd-workflow ← ns-workflow.md) — DEFECT.DISPLAY-VALUE-AS-IDENTITY. Landing the gate required backfilling 32 pre-existing unrostered surfaces on next", - "line": 650 + "line": 653 }, { "id": "RULESET.PR-FLOW.docker-before-push", "klass": "RULESET", "value": "before ANY git push of any fix to any PR, run gsd-test (docker on the remote, mirrors ubuntu CI) and confirm exit 0. macOS-local node --test is NOT a substitute — many failures are platform-specific (path separators, case sensitivity, locale, fs semantics). Watchdog with Monitor on the output log; never set a sleep/timer and walk away. Source: user feedback 2026-05-16 — \"we don't set a timer we actively watch and record results in real time as possible\". SUPERSEDED 2026-07-17: 'confirm exit 0' is a false-green trap — piping/backgrounding can report exit 0 on a failed suite; gate on the verdict-line outcome:\"passed\" for the exact HEAD sha instead. See CLAUDE.md's gsd-test rule and the gsd-test-is-ref-based-commit-first predicate for the current, correct gating contract.", - "line": 922 + "line": 933 }, { "id": "RULESET.PR-FLOW.templates-mandatory", "klass": "RULESET", "value": "every gh pr create|edit|gh issue create|edit MUST first invoke the gh-templates-first skill and Read (Read tool, not Bash cat — k321 read-tracking) the matching template in .github/. Apply ALL required sections; never write freeform bodies. Repo enforces this via gsd-pr-template-policy GitHub Action which flags any non-templated body — the bot allows the PR to stay open only because authors are contributors-or-higher, but the warning is a real complaint that must be cured. Source: user feedback 2026-05-16 (multi-message escalation) — \"the whole reason i have that github action is because you fucking blow through and ignore using the templates\"", - "line": 924 + "line": 935 }, { "id": "RULESET.PR-SCOPE.one-concern-per-pr", "klass": "RULESET", "value": "split unrelated changes into separate PRs; cherry-pick doc changes to dedicated docs/ branch immediately, then force-push original to remove the commit", - "line": 652 + "line": 655 }, { "id": "RULESET.SHARED-HELPERS-LINT-VS-TEST", "klass": "RULESET", "value": "when a lint script and test suite both implement same constant (CANONICAL_TOOLS) or parser (parseFrontmatter, executionContextRefs), extract to scripts/*-helpers.cjs required by both — silent divergence otherwise", - "line": 647 + "line": 650 }, { "id": "RULESET.TESTS.CODERABBIT_FIX", "klass": "RULESET", "value": "prefer exported-function behavioral tests over source-grep; lint-no-source-grep rejects readFileSync source assertions without allow-test-rule", - "line": 676 + "line": 679 }, { "id": "RULESET.TESTS.boundary-coverage", "klass": "RULESET", "value": "tests MUST exercise inputs at and near the threshold/limit, not only trivial-fit and trivial-overflow; pick inputs where N ∈ {limit-1, limit, limit+1} and where pre-trim/pre-check accumulators ≈ effective limit; \"very small\" and \"very large\" inputs alone do not constitute edge-case coverage and routinely miss off-by-one + reservation-accounting bugs", - "line": 619 + "line": 622 }, { "id": "RULESET.TESTS.boundary-coverage.anti-pattern", "klass": "RULESET", "value": "test suites that pair budget:1_000_000 (trivially fits) with budget:1 (trivially overflows) and skip the boundary region; failure mode that shipped PR #3708 UNNEEDED_TRIM + FALSE_HARDFAIL regressions (commit 2df566ed, fixed bde1ae8f)", - "line": 622 + "line": 625 }, { "id": "RULESET.TESTS.boundary-coverage.fixtures", "klass": "RULESET", "value": "for any code with budget/limit/quota/threshold parameter, test suite MUST include: (a) input where SUT estimate == limit exactly, (b) input where estimate == limit - 1, (c) input where estimate == limit + 1, (d) input where any internal reserve/safety constant pushes baseline within reserve-distance of limit (catches early-pressure firing)", - "line": 621 + "line": 624 }, { "id": "RULESET.TESTS.clock-seam", "klass": "RULESET", "value": "concurrency logic must accept an optional {clock=Date} parameter; tests control time via t.mock.timers.enable(['Date']) + t.mock.timers.setTime(0) + t.mock.timers.tick(N); real OS scheduler races are not a permitted test pattern after ADR 456 (2026-05-28); real-race tests are deleted once deterministic seam tests cover the same logical path; clock.cjs realClock adds nowIso() (→ new Date(this.now()).toISOString()) and today() (→ nowIso().split('T')[0]) so all date-stamping in state.cjs routes through the seam; subprocess time-pin adapter: set GSD_TEST_MODE=1 + GSD_NOW_MS= in runGsdTools env to pin the date written by the SUT without touching real wall-clock (issue #474)", - "line": 626 + "line": 629 }, { "id": "RULESET.TESTS.coderabbit-fix-prefer", "klass": "RULESET", "value": "behavioral tests (call exported fn, capture JSON, assert typed fields) over source-grep", - "line": 617 + "line": 620 }, { "id": "RULESET.TESTS.delete-bad-tests", "klass": "RULESET", "value": "pass-always / vacuous-truth / source-grep / elapsed-time / real-race / permanent-allow-test-rule tests are DELETED and replaced with compliant tests in the same PR; not skipped, not commented out, not permanently exempted; replacement must cover the same logical path via typed-surface assertion or clock-seam pattern", - "line": 631 + "line": 634 }, { "id": "RULESET.TESTS.diagnostics", "klass": "RULESET", "value": "after JSON.parse, assert output shape (Array.isArray(output.phases)) with raw-output-prefix diagnostics before .map() — prevents opaque TypeErrors when CLI output shape changes", - "line": 618 + "line": 621 }, { "id": "RULESET.TESTS.escape-regex", "klass": "RULESET", "value": "new RegExp(\"prefix${var}\") must escapeRegex(var); phase-id.cjs exports escapeRegex (core.cjs re-export spine retired in epic #1267); phase IDs like 5.1 contain . which is metacharacter", - "line": 614 + "line": 617 }, { "id": "RULESET.TESTS.eslint-harness", "klass": "RULESET", "value": "ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at eslint-rules/ (repo root, NOT scripts/eslint-rules/); replaces scripts/lint-*.cjs regex scanners (fully removed in #632); all three test-rigor rules now ship at error in tests/**/*.test.cjs scope: local/no-source-grep and local/no-magic-sleep-in-tests promoted by #3313, local/no-elapsed-assertion promoted by #3331 once #3314 delivered its ADR-456 §(a) precondition (epic #1885 was subsumed into epic #3053 and closed stale before this promotion landed)", - "line": 632 + "line": 635 }, { "id": "RULESET.TESTS.feedback-loop-convergence", "klass": "RULESET", "value": "when a feature's OUTPUT feeds back into its own INPUT (calibration, retry backoff, adaptive budgets, ratchets, any self-correcting signal), step-wise tests are NOT sufficient evidence of correctness: they assert `given X return Y` while the defect lives in the TRAJECTORY across iterations. Required: a closed-loop test that (a) drives the REAL end-to-end surface — not the pure core alone, since composition bugs live between surfaces — for N >= 2x the loop's window, (b) asserts convergence on the known-true value, (c) asserts the fixed point (an already-correct history must produce NO correction), and (d) asserts boundedness under an adversarial/oscillating history. Two defects shipped past a green ~26,800-test suite in epic #1952 for want of exactly this: calibration applied twice across two surfaces (factor^2, #2631) and calibration measured against its own corrected output so it oscillated to ~1.41 instead of converging on 2.0 (#2632). Every unit, boundary, property and round-trip test passed for both. HOW TO SPOT ONE (the detection tell, not a judgment call): the feature's own acceptance criterion carries a TEMPORAL QUANTIFIER — \"after N phases\", \"subsequent\", \"over time\", \"improves\", \"learns\", \"adapts\". That phrasing means the claim is about a TRAJECTORY, so a step-wise `given X return Y` test does not test the claim that was made. #1952's AC4 read \"After N phases, the error is computed and applied as a correction to SUBSEQUENT estimates\" — the tell was in plain sight and was still tested as a point. Survey of this repo (2026-07): estimation calibration is the ONLY true instance; size/mutation ratchets are exempt because they fail on both growth AND shrinkage (cannot self-satisfy), and retry ladders (node_repair_budget, plan_bounce_passes, provider_escalation) terminate rather than feed back. Test anchor: tests/estimate-loop-convergence.test.cjs", - "line": 620 + "line": 623 }, { "id": "RULESET.TESTS.guard-toplevel-readFileSync", "klass": "RULESET", "value": "module-level const src = readFileSync(...) throws before any test() registers — wrap in try/catch in test() or use lazy load", - "line": 616 + "line": 619 }, { "id": "RULESET.TESTS.mutation-runner", "klass": "RULESET", "value": "Stryker executes every shard through the OFFICIAL @stryker-mutator/tap-runner (testRunner:'tap'), never the built-in 'command' runner (#3915); 'command' is the one runner Stryker excludes from coverage analysis, which forced coverageAnalysis:'off' and made cost strictly linear in (mutants x whole-shard test time) — the frontmatter shard measured 1751s on run 33021042847 vs 212s for the next slowest. tap.testFiles is injected per shard via MUTATION_TEST_FILES (mutation.yml env <- matrix.tests <- scripts/mutation-matrix.cjs buildResult); resolveMutationTestFiles is the SINGLE fail-closed reader and existence-checks every entry, because the tap runner's findTestyLookingFiles resolves the list with glob() and a non-matching pattern yields an EMPTY list SILENTLY (a fast, confident, meaningless run). tap.forceBail is FALSE by measurement, not preference: 3 of 26 shard test files spawn subprocesses (config-schema.property, core-utils, feat-3881-yaml-parser-consequences) and bail fires on every KILLED mutant, so leaving it on kills processes mid-spawnSync and orphans their children; Stryker's separate disableBail still skips remaining FILES, which is most of the win. tap.nodeArgs and top-level buildCommand stay UNSET so no rebuild lands between mutation and test (ADR-457). Coverage granularity is per FILE, not per test (\"a test is always a test file\"), so the #2790 excludeTests bans on spawn-heavy integration files remain necessary and unchanged", - "line": 629 + "line": 632 }, { "id": "RULESET.TESTS.mutation-score", "klass": "RULESET", "value": "Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification", - "line": 628 + "line": 631 }, { "id": "RULESET.TESTS.mutation-score-denominator", "klass": "RULESET", "value": "the gated number is mutation-testing-metrics' mutationScore = totalDetected/totalValid, which counts NoCoverage in the denominator EXACTLY as Survived; both Stryker's own thresholds.break (core dist/src/reporters/mutation-test-report-helper.js) and scripts/check-mutation-score-ratchet.cjs read THAT field, which is what makes the #3915 coverageAnalysis 'off'->'perTest' switch score-neutral. NEVER gate on mutationScoreBasedOnCoveredCode — it EXCLUDES NoCoverage and inflates sharply under perTest (measured on a synthetic report: 8 killed/2 survived = 80 and 80; 8 killed/2 noCoverage = 80 and 100), so swapping to the better-sounding field would make every minScore floor trivially satisfiable and the gate decorative. Under the pre-#3915 coverageAnalysis:'off' the two fields were ALWAYS identical (noCoverage was structurally 0), which is why nothing had ever pinned the choice; tests/mutation-score-ratchet.test.cjs now pins it with a non-vacuity assertion that the two numbers genuinely diverge", - "line": 630 + "line": 633 }, { "id": "RULESET.TESTS.no-dead-regex-in-includes", "klass": "RULESET", "value": "src.includes(\"foo.*bar\") is always false — .* is regex metacharacter not wildcard; use new RegExp(...).test(src) or delete", - "line": 615 + "line": 618 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker", "klass": "RULESET", "value": "local/no-duplicate-fold-marker ESLint AST rule (eslint-rules/no-duplicate-fold-marker.cjs, #3271) reports the 2nd and every later __foldDescribe(\"folded: ...\") call carrying a marker already seen in the SAME file, naming the first occurrence's line; error in tests/**/*.cjs. The key is the WHITESPACE-delimited token after folded:, NOT a [a-z0-9-]* slice — a slice truncates at \".\" and collides feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode (two distinct suites coexisting in tests/model-resolver.test.cjs), and NOT the whole title, so a re-fold under a different batch label (\"B1 #1970\" vs \"B5 #1975\") is still caught. Deliberately silent on: a __foldDescribe title with no folded: prefix (the alias is reused for one ordinary describe in tests/review-default-reviewers-workflow.test.cjs), a plain describe(), a non-literal title, and the same marker in two DIFFERENT files (the defect class is intra-file).", - "line": 611 + "line": 614 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker.why", "klass": "RULESET", "value": "consolidation epic #1969 folds are self-contained blocks, so a second verbatim copy parses, registers and PASSES twice — nothing reports it; #3271 found 25 such copies (~5,800 lines) in tests/install.test.cjs (18), tests/install-minimal-hooks.test.cjs (5) and tests/install-write-confinement.test.cjs (2), all from one stale-base re-application in 6d072435d (#1975 re-applying #1970's hunks, 2026-07-03). Ref DEFECT.GENERATIVE-FIX: the two copies drift apart silently when a contributor fixes one and leaves the other asserting the old behavior, with the suite still green.", - "line": 612 + "line": 615 }, { "id": "RULESET.TESTS.no-source-grep", "klass": "RULESET", "value": "local/no-source-grep ESLint AST rule (eslint-rules/no-source-grep.cjs) rejects readFileSync of a source .cjs/.js/.ts path bound to a var later hit with .includes()/.match()/.startsWith()/.endsWith()/.indexOf()/.search(); error in tests/**/*.test.cjs, warn in gsd-core/bin/**/*.cjs + scripts/**/*.cjs (ADR 452 retired the old regex script, removed for good in #632)", - "line": 608 + "line": 611 }, { "id": "RULESET.TESTS.no-source-grep.exemption", "klass": "RULESET", "value": "// allow-test-rule: with one-line justification; reserved for tests where the file content IS the product surface (STATE.md, config.toml, hooks.json, agent .md). Migration to typed-IR parser tracked in #2974.", - "line": 609 + "line": 612 }, { "id": "RULESET.TESTS.no-source-grep.tmp-file-traps", "klass": "RULESET", "value": "reading tmp files written by the SUT in tests still trips lint; round-trip through CLI (e.g. frontmatter get) instead of readFileSync+.includes()", - "line": 610 + "line": 613 }, { "id": "RULESET.TESTS.no-timing-assertion", "klass": "RULESET", "value": "do not assert on wall-clock elapsed time (Date.now() delta, performance.now(), process.hrtime() comparison); such assertions test the host machine not the SUT and flake on loaded CI runners; enforcement: local/no-elapsed-assertion ESLint rule, error (promoted by #3331 once #3314 delivered the ADR-456 §(a) reachability rule + deterministic backfill precondition); canonical replacement: clock-seam pattern with node:test mock.timers", - "line": 625 + "line": 628 }, { "id": "RULESET.TESTS.property-based-testing", "klass": "RULESET", "value": "modules implementing parsing / transformation / budget-limit / bijective contracts must include at least one fast-check (fc) property test asserting a domain invariant; invariant categories: round-trip, monotonicity, boundary-containment, idempotency; property tests live in *.test.cjs alongside unit tests; CI signal: Stryker mutation score below 80% blocks merge", - "line": 627 + "line": 630 }, { "id": "RULESET.TRIAGE-EXISTING-WORK", "klass": "RULESET", "value": "before writing agent brief for confirmed bug, check (1) local branches git branch -a | grep , (2) untracked/modified files on that branch, (3) stash, (4) open PRs with matching head branch — recover existing work rather than re-implement", - "line": 654 + "line": 657 }, { "id": "RULESET.WORKFLOW.COVERAGE-METADATA", "klass": "RULESET", "value": "#1602 SUMMARY frontmatter `coverage:` block (list of {id,description,requirement?,verification:[{kind∈unit|integration|e2e|automated_ui|manual_procedural|other, ref, status∈pass|fail|unknown}],human_judgment:bool,rationale?}) is the per-deliverable RTM consumed DETERMINISTICALLY by verify-work extract_tests via `gsd-tools uat classify-coverage --summary ` (src/coverage.cts → bin/lib/coverage.cjs). AUTHORING: execute-plan create_summary populates it from task results; every deliverable MUST be classified; fail-safe default = human_judgment:true + rationale. CLASSIFY CONTRACT: auto-pass (skip human) ONLY when human_judgment===false (strict boolean) AND verification non-empty AND every status==='pass' AND zero validation errors — else PRESENT to human. mode:legacy (no block) ⇒ byte-identical prose `## Accomplishments` fall-through; `coverage: []` ⇒ mode:coverage, zero entries (single-confirmation). Frozen IR: MODE/PRESENT_REASON/ERROR_CODE enums locked by tests/coverage-metadata-parser.test.cjs. extractFrontmatter CANNOT parse it (scalars-only `-` items) → dedicated parser, sibling of parseMustHavesBlock. Asymmetry by design: false-negative=redundant prompt (status quo); false-positive=shipped bug UAT existed to catch", - "line": 643 + "line": 646 }, { "id": "RULESET.WORKFLOW_EXECUTE_END_TO_END", "klass": "RULESET", "value": "standard for single-workflow commands is \"Execute end-to-end.\" (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses \"execute the X workflow end-to-end.\" in routing bullets — convention verified live across ~20 commands/gsd/*.md files; no ADR currently documents this specific phrasing rule (ADR-0002 covers the adjacent but distinct command-contract/@-ref-resolution seam, not this convention)", - "line": 642 + "line": 645 }, { "id": "RULESET.WORKFLOW_EXECUTION_CONTEXT", "klass": "RULESET", "value": "@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/docs-update.test.cjs (folds former \\`bug-3135-capture-backlog-workflow\\`, consolidation epic #1969); INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; \"Invoked by\" attribution must move when a flag absorbs a micro-skill", - "line": 641 + "line": 644 }, { "id": "RULESET.WORKFLOW_FILE_NAMES", "klass": "RULESET", "value": "workflow files use hyphens; XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name", - "line": 640 + "line": 643 }, { "id": "RULESET.WORKFLOW_MARKDOWN.FENCES", "klass": "RULESET", "value": "preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)", - "line": 636 + "line": 639 }, { "id": "RULESET.WORKFLOW_SIZE_BUDGET", "klass": "RULESET", "value": "workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4: tests/emitted-attribution.test.cjs's real-tree test reports growth in any gsd-core/workflows/*.md with its exact byte delta vs `next`, no committed snapshot, requires an `Emitted-Drift-Ack-Growth:` commit trailer on the PR's own commits (ADR-3942, superseding ADR-2719 §3's fragment model — key is the bare filename, reason follows ` — `)) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + discuss-phase<32000; a file that grew fails the differential guard — add an ack entry naming the file and reason, justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump. The prior per-file baseline (tests/workflow-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724. Its new-file cap (ADR-1610 Decision point 3, un-baselined files <=32768, the Codex anchor) is REVIVED inside the differential's size ratchet itself (`NEW_FILE_CAP` in tests/helpers/emitted-diff.cjs) rather than lost: \"not yet baselined\" is exactly \"present in sizeCurrent, absent from sizeBaseline\", a signal the ratchet already computes for its own reasons. NOT ack-able — same as the tier hard caps, the fix is extraction. Narrower than the original: this check cannot see XL/LARGE tiering (tests/workflow-size-budget.test.cjs's classification, invisible to the pure differential module), so a legitimately large NEW file must extract rather than tier in, one release earlier than an existing file would need to — a disclosed, deliberate simplification", - "line": 637 + "line": 640 }, { "id": "SEAM.capability-activation-precedence-owner.enforced-by", "klass": "SEAM", "value": "test:tests/capability-precedence-parity.test.cjs", - "line": 414 + "line": 417 }, { "id": "SEAM.capability-activation-precedence-owner.owns", "klass": "SEAM", "value": "the config-key four-level precedence walk (loadConfig result → workstream config.json → root config.json → registry.configSchema default → absent), owned solely by src/capability-activation.cts", - "line": 413 + "line": 416 }, { "id": "SEAM.git-query-readonly-seam.enforced-by", "klass": "SEAM", "value": "test:tests/git-base-branch.test.cjs", - "line": 206 + "line": 209 }, { "id": "SEAM.git-query-readonly-seam.owns", "klass": "SEAM", "value": "bounded, never-throw git repository introspection — base-branch detection, worktree-info detection, phase change-set detection", - "line": 205 + "line": 208 }, { "id": "SEAM.package-identity.enforced-by", "klass": "SEAM", "value": "test:tests/package-identity.test.cjs", - "line": 283 + "line": 286 }, { "id": "SEAM.package-identity.owns", "klass": "SEAM", "value": "GSD's published-package coordinates (packageName, binName, repoSlug, changelogRawUrl, manualInstallCommand) — single seam so a repoint/rename is a one-line change", - "line": 282 + "line": 285 }, { "id": "SEAM.phase-locator-milestone-enum.enforced-by", @@ -1469,13 +1518,13 @@ "id": "SEAM.shellcmdproj-win-binary-resolution.enforced-by", "klass": "SEAM", "value": "lint-rule:no-private-binary-resolution", - "line": 887 + "line": 898 }, { "id": "SEAM.shellcmdproj-win-binary-resolution.owns", "klass": "SEAM", "value": "Windows binary resolution (resolveExecutableBinary, projectSpawnInvocation) — which file a declared command name actually names, and cmd.exe mediation", - "line": 886 + "line": 897 }, { "id": "SEAM.verification-isphasecomplete.enforced-by", @@ -1493,199 +1542,199 @@ "id": "SEAM.worktree-safety-policy.enforced-by", "klass": "SEAM", "value": "test:tests/worktree-safety.test.cjs", - "line": 704 + "line": 707 }, { "id": "SEAM.worktree-safety-policy.owns", "klass": "SEAM", "value": "Worktree Safety Policy Module — resolveWorktreeContext, parseWorktreePorcelain, planWorktreePrune, executeWorktreePrunePlan, W017 classification (see WORKTREE.SEAM.* above for full interface/invariant detail)", - "line": 703 + "line": 706 }, { "id": "SESSION.2026-05-05", "klass": "SESSION", "value": "[PRED.k320..k331 introduced; DEFECT.SOURCE-GREP-IN-NEW-TESTS, DEFECT.CHANGESET-PR-FIELD-DRIFT, DEFECT.PHASE-DIR-PREFIX-DRIFT, DEFECT.PROMPT-INJECTION-SCAN-COLLISION; ADR-0002 thin-wrapper pattern findings folded into RULESET.WORKFLOW_*]", - "line": 910 + "line": 921 }, { "id": "SESSION.2026-05-05.sdk-bridge", "klass": "SESSION", "value": "PR #3158 SDK Runtime Bridge — observability isolation rule; strict-mode dispatchMode reporting invariant; transport decision ordering (guard before event emission); folded into Dispatch Policy Module glossary", - "line": 911 + "line": 922 }, { "id": "SESSION.2026-05-09", "klass": "SESSION", "value": "[8-PR triage wave, 7 merged + 1 subsumed; META.RULE.* introduced; WAVE.LESSON.* captured; k320/k322/k323/k326/k331 evidence; AI Ops Memory predicate format established]", - "line": 912 + "line": 923 }, { "id": "SESSION.2026-05-10", "klass": "SESSION", "value": "[ai-ops memory consolidation; release-notes standard taxonomy + templates; RELEASE-NOTES.* predicates introduced]", - "line": 913 + "line": 924 }, { "id": "SESSION.2026-05-13", "klass": "SESSION", "value": "[Shell Command Projection Module expansion (#3465-#3468); ADR-0009 superseded; new exports for subprocess dispatch and platform file I/O; phase-gated migration plan; PR #3464 three-gate invariant CI+CR+unresolved=0; PR #3470 stash-include-untracked rebase pattern]", - "line": 914 + "line": 925 }, { "id": "SESSION.2026-05-14", "klass": "SESSION", "value": "[#3095/PR #3490 EXEC.CLASSIFY.* introduced (Anthropic/Copilot/Codex/Gemini [runtime removed #1928] cross-runtime rate-limit sentinel coverage); #3489/PR #3499 DEFECT.STATE-TRAMPLE.idempotency-oracle (STATE.md current_phase field is oracle for state.complete-phase); #3488/PR #3501 DAG resolver same-phase short-form depends_on (shortFormToId index added to sdk/src/query/phase.ts); #3491/PR #3502 DEFECT.NESTED-GIT-INIT (gitWorktreeInfoInternal helper); #3493/PR #3500 extractCurrentMilestone generic Phase Details continuation past planned-milestone siblings; #3503/PR #3504 DEFECT.PATH-SUBSTRING-CHECK (trailing-slash anchor for homedir checks); #3346/PR #3505 codex AoT TOML leaf-key via extractFlatHookEventName; #3506/PR #3507 label-scoped stale-bot sub-job pattern; multi-PR triage operational lessons folded into PROC.TRIAGE.*; #3508 DEFECT.AGENT-ISOLATION-SILENT-FAIL; gsd-test image-missing auto-build (locally-built image via embedded heredoc Dockerfile); refined PRED.k322 threshold to 3 PRs/<10min]", - "line": 915 + "line": 926 }, { "id": "SESSION.2026-05-15", "klass": "SESSION", "value": "[#3537/PR #3538 DEFECT.PHASE-REGEX-FANOUT — phaseMarkdownRegexSource promoted to core.cjs and wired to 7 sites; parity-style regression test established as DEFECT.GENERATIVE-FIX exemplar; trek-e/gsd-test-runner#1 filed for DEFECT.GSD-TEST-MIRROR-POISONED — chown-back-before-exec legacy gap (poisoned holodeck mirror unstuck via authorized docker chown to remote 1000:1000); RULESET.PR-FLOW.* codified from project CLAUDE.md load-bearing rule; first dispatch under run-tests-before-create held cleanly (PR #3520 worker stopped on Docker exit 12 infra failure, orchestrator opened PR after unblock); CONTEXT.md refactored from 882 lines of mixed prose+predicates into ~500 lines of pure-predicate format with chronological session log]", - "line": 916 + "line": 927 }, { "id": "SESSION.2026-05-15.parallel-fix-dispatch", "klass": "SESSION", "value": "[#3542/PR #3546 prohibit git stash family in executor agents (shared refs/stash across worktrees); #3541/PR #3547 non-TTY resolution for installer prompt-user actions (default remove for SDK build artifacts, keep for skills/gsd-*/SKILL.md); #3545 filed for gsd-test-summary concurrent /tmp output collision; new predicates DEFECT.HOOK-OVER-ENFORCEMENT.read-tool-tracking, DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION, DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL, DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT, PROC.PARALLEL-FIX-DISPATCH; agent-trust-but-verify caught /gsd-update retired-syntax comment slip in #3541 implementation before PR open]", - "line": 917 + "line": 928 }, { "id": "SESSION.2026-05-16", "klass": "SESSION", "value": "[multi-PR triage wave (#3577/3581/3640/3641/3642/3648/3649/3637/3639). Established global PreToolUse hook ~/.claude/hooks/test-memory-guard.sh denying new node/test spawns when sum(RSS of node|vitest|jest|...) >= 4 GiB on the 24 GB Mac OR when a same-runner process is already in argv[0] — hard deny via hookSpecificOutput.permissionDecision=deny. PR #3577 fix: revert config-ensure-section dispatch to CJS cmdConfigEnsureSection (SDK author wrote single-section semantics under a name whose legacy callers expect full-default config init); plus 3 SDK parity carve-outs (configNewProject defaults align with sdk/shared/config-defaults.manifest.json, return relative .planning/config.json path, drop quotes from Unknown config key, lead malformed-JSON error with \"Failed to read config.json:\"). PR #3649 fix: chunk node --test spawn at 28K argv ceiling (Windows CreateProcess lpCommandLine cap 32,767 was instantly aborting unchunked spawn of 546 paths). Chunking fix surfaced 14 pre-existing Windows-only test bugs (4010 pass / 14 fail; vs 0/0 before — entire suite was un-runnable on Windows). PRs #3639 + #3637 confirmed unable to stand alone (legitimately depend on Phase 6 scaffolding only present on feat/3575-enforcement-hardening) — user decision: cherry-pick into #3577 and close. Five other PRs each had ≤1 unresolved CR thread of the changeset-pr-number / null-vs-throw / implicit-Claude-runtime / docs-stale-guidance / hardcoded-tests-path family — all quick wins. New predicates: DEFECT.SDK-PORT-NAME-COLLISION, DEFECT.WINDOWS-ARGV-OVERFLOW, DEFECT.STACKED-PR-CANNOT-STAND-ALONE, DEFECT.CANARY-VERSION-LEAK, DEFECT.GSD-TEST-HOST-MID-RUN-DEATH, RULESET.HARNESS.test-memory-guard, RULESET.PR-FLOW.docker-before-push, RULESET.PR-FLOW.templates-mandatory]", - "line": 918 + "line": 929 }, { "id": "WAVE.LESSON.agent-narrative-unreliable", "klass": "WAVE", "value": "k095/k324 confirmed at scale: 5 of 8 agents terminated mid-monitor with stale claims requiring direct verification", - "line": 869 + "line": 880 }, { "id": "WAVE.LESSON.changelog-policy-violation-multiplier", "klass": "WAVE", "value": "brief contradicting CONTRIBUTING.md's changelog-fragment policy (\"CHANGELOG Entries — Drop a Fragment\" section) produced violations on 5 of 8 PRs (#3300, #3302, #3304, #3305, #3308); k326 + k320 capture", - "line": 866 + "line": 877 }, { "id": "WAVE.LESSON.cr-throttle-burst-correlation", "klass": "WAVE", "value": "8 PRs in <15min triggered k322 sustained-throttle on multiple PRs (#3306 worst case)", - "line": 867 + "line": 878 }, { "id": "WAVE.LESSON.k101-still-trips", "klass": "WAVE", "value": "even after CONTEXT.md k101 reinforcement, agent of record posted self-PR comment on close; k331 adds explicit close-time literal-instruction guard", - "line": 870 + "line": 881 }, { "id": "WAVE.LESSON.sibling-audit-overlap", "klass": "WAVE", "value": "k015-family parallel dispatch on #3297 + #3298 produced k323 add-backlog.md cross-PR overlap", - "line": 868 + "line": 879 }, { "id": "WORKSTREAM.INVARIANT.migrate-name", "klass": "WORKSTREAM", "value": "must normalize through canonical slug policy", - "line": 690 + "line": 693 }, { "id": "WORKSTREAM.INVARIANT.slug-contract", "klass": "WORKSTREAM", "value": "all .planning/workstreams/ must be addressable by set/get/status/complete", - "line": 691 + "line": 694 }, { "id": "WORKSTREAM.NAME.POLICY.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/workstream-name-policy.cjs owns toWorkstreamSlug + active-name/path-segment validation", - "line": 708 + "line": 711 }, { "id": "WORKSTREAM.POINTER.SEAM.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/active-workstream-store.cjs owns read/write self-heal for .planning/active-workstream", - "line": 709 + "line": 712 }, { "id": "WORKSTREAM.REGRESSION.test-anchor", "klass": "WORKSTREAM", "value": "tests/workstream.test.cjs::normalizes --migrate-name to a valid workstream slug", - "line": 692 + "line": 695 }, { "id": "WORKTREE.SEAM.caller-rule", "klass": "WORKTREE", "value": "verify.cjs must consume inspectWorktreeHealth for W017 classification; no ad-hoc porcelain parsing in callers", - "line": 700 + "line": 703 }, { "id": "WORKTREE.SEAM.current", "klass": "WORKTREE", "value": "Worktree Safety Policy Module", - "line": 684 + "line": 687 }, { "id": "WORKTREE.SEAM.decision-1", "klass": "WORKTREE", "value": "retain non-destructive default; destructive path only as explicit future opt-in scaffold", - "line": 688 + "line": 691 }, { "id": "WORKTREE.SEAM.default-prune-policy", "klass": "WORKTREE", "value": "metadata_prune_only (non-destructive)", - "line": 687 + "line": 690 }, { "id": "WORKTREE.SEAM.files", "klass": "WORKTREE", "value": "[gsd-core/bin/lib/worktree-safety.cjs]", - "line": 685 + "line": 688 }, { "id": "WORKTREE.SEAM.interface", "klass": "WORKTREE", "value": "[resolveWorktreeContext, parseWorktreePorcelain, planWorktreePrune, executeWorktreePrunePlan, planWorktreeRecordAgent, cmdWorktreeRecordAgent]", - "line": 686 + "line": 689 }, { "id": "WORKTREE.SEAM.invariant", "klass": "WORKTREE", "value": "parser failure must degrade to metadata_prune_only and never escalate to destructive removal", - "line": 698 + "line": 701 }, { "id": "WORKTREE.SEAM.inventory-interface", "klass": "WORKTREE", "value": "[listLinkedWorktreePaths, inspectWorktreeHealth]", - "line": 699 + "line": 702 }, { "id": "WORKTREE.SEAM.inventory-snapshot", "klass": "WORKTREE", "value": "snapshotWorktreeInventory(repoRoot,{staleAfterMs,nowMs}) is canonical linked-worktree health snapshot for callers", - "line": 702 + "line": 705 }, { "id": "WORKTREE.SEAM.test-anchor-w017", "klass": "WORKTREE", "value": "tests/orphan-worktree-detection.test.cjs + tests/worktree-safety.test.cjs", - "line": 701 + "line": 704 }, { "id": "WORKTREE.SEAM.test-anchors", "klass": "WORKTREE", "value": "[resolveWorktreeContext:has_local_planning|linked_worktree|not_git_repo|main_worktree, planWorktreePrune:git_list_failed|worktrees_present|no_worktrees|parser_throw_fallback, executeWorktreePrunePlan:missing_plan|skip_passthrough|unsupported_action|metadata_prune_only]", - "line": 697 + "line": 700 }, { "id": "WORKTREE.SEAM.test-policy", "klass": "WORKTREE", "value": "cover all decision branches in policy module before changing prune behavior", - "line": 696 + "line": 699 } ], "duplicates": [] diff --git a/src/phase.cts b/src/phase.cts index 734b9da2f..c1df0821c 100644 --- a/src/phase.cts +++ b/src/phase.cts @@ -2425,6 +2425,797 @@ function phaseDisplayNameFromSlug(slug: string | null): string | null { return name || null; } +// ─── #3697: the `**Requirements**:` line under-selection detector ──────────── +// +// EXTRACTED from cmdPhaseComplete (round 3, review finding Blocker 1). The +// detection logic below is a parser, so `RULESET.TESTS.property-based-testing` +// requires at least one fast-check property test over it — and that is not +// reachable while the logic is a closure inside a command that only a +// subprocess can invoke (every #3697 test spawns the CLI; 100 fc runs cannot). +// Extraction is therefore load-bearing, not tidying: it is what makes the +// property test and the 2048-boundary fixtures (Blocker 2) expressible at all. +// +// BEHAVIOUR IS UNCHANGED BY THE MOVE. The two tokenizations below stay +// deliberately DIFFERENT and are co-located so they cannot drift apart: +// * the SELECTOR strips `[` and `]` only, then splits on `[,\s]+`. Its output +// IS `citedReqIds` — the ledger-writing set — so widening it would change +// what phase-complete marks, which #3697 explicitly does not do. +// * the DETECTOR additionally shaves brackets/quotes/emphasis and trailing +// sentence punctuation, so it can see an operator or an ID that the +// selector's stricter shape filter rejects. +// The gap between them is not a defect: it is why `ADR-7)` is not selected +// while `ADR-7` is still nameable in a warning. +// +// The `**Requirements**: TBD` placeholder is what phase.add / -batch / -insert +// seed (three sites in this file — locate them by the literal +// `Requirements**: TBD`, never by line number: an earlier revision of this +// comment cited 833/920/1078, which had drifted to 1132/1237/1413 by round 3). +// The shipped comma-list template is `gsd-core/templates/roadmap.md:32`. +type RequirementsLineAnalysis = { + /** The ledger-writing set — byte-identical to the pre-extraction selector. */ + citedReqIds: string[]; + /** The detector's shaved tokens (see the tokenization note above). */ + tokens: string[]; + /** R1 — tokens that are THEMSELVES a range (`RANGE-01..RANGE-05`). */ + rangeTokens: string[]; + /** R2 — a bare operator with a selected, interior-implying ID either side. */ + hasSpacedRange: boolean; + /** R2' — an operator GLUED to one endpoint (`RANGE-01 -RANGE-05`). */ + hasGluedRangeFragment: boolean; + /** R3 — zero selection on a non-placeholder line, with ID-shaped residue. */ + inertIdShaped: string[]; + /** + * R3b — zero selection on a non-placeholder line that carries ANY content. + * + * This is #3697's AC-1b/AC-4 verbatim ("warn when `citedReqIds.length === 0` + * while the raw capture is non-empty and not `TBD`"), and it is deliberately + * NOT gated on ID-shaped residue the way R3 is. Round 4 measured the reason: + * every one of the fifteen #2334/#2339 negative-space fixtures is held + * silent by non-zero SELECTION or by `placeholderLed`, and not one of them by + * the ID-shape gate — so the gate was buying no negative space while costing + * the acceptance criterion. `Deferred`, `N/A`, `Pending`, `TBA` and `-` were + * silent because of it, while the docs, this census and the advice string all + * said they warned. + */ + zeroSelectionInert: boolean; + /** + * R4 — REQ-IDs the SELECTOR dropped because a delimiter was glued to them. + * + * `REQ-01; REQ-02` selects only `REQ-02`: the selector splits on `[,\s]+`, + * so `REQ-01;` keeps its semicolon and fails the anchored ID shape. This is + * #3697's own half-success failure mode — `requirements_updated: true` with + * a silently unmarked requirement — reached by one wrong delimiter. + * + * Round 4 review called this indistinguishable from a parenthesised + * citation, because `(ADR-7)` also shaves to a bare ID. At the RAW token + * level they are not: `REQ-01;` is shaved of a trailing DELIMITER, + * `ADR-7)` of a citation wrapper. This rule keys on that shave class and + * requires the token to sit outside any parenthetical, which is what keeps + * `(see ADR-7: section 3)` silent. + */ + delimiterDroppedIds: string[]; + /** Tokens past the scan cap that could carry an ID — reported, never dropped. */ + oversizedTokens: string[]; + /** ID-shaped tokens the selector did not take. Reported as a fact; never routes. */ + unselectedIdShaped: string[]; + /** The line leads with `TBD` / `None`. */ + placeholderLed: boolean; + /** R2's hits, as `[left, right]` endpoint pairs, so the channel below can ask + * about the endpoints the rule actually fired on. */ + spacedRangePairs: Array<[string, string]>; + /** + * Nothing on the line was DEMONSTRABLY dropped: no rule that names a specific + * unselected ID fired, and any spaced range fired on endpoints the selector + * actually took. + * + * This is the shared precondition of both NON-assertive voices — the + * ambiguous range reading and the over-cap "not classified" report — and it + * is named once because they had drifted apart. Round 7 review, Minor 1: + * `rangeReadingOnly` carried the conjunction inline and omitted the cap, + * while the over-cap channel carried its own copy that excluded a spaced + * range wholesale. A line with a clean, fully-selected range beside an + * unexamined over-cap token satisfied neither guard as intended and reached + * the ambiguous voice. + */ + nothingDemonstrablyDropped: boolean; + /** + * The round-3 channel discriminator (review finding Major 3). True when the + * ONLY thing to report is a range *reading*: R2 fired, no other rule did, and + * every endpoint R2 fired on was actually selected. Nothing was dropped, so + * the line did not fail to parse and the warning must not claim it did. + * + * This is deliberately RULE-SCOPED rather than line-global. A line-global + * "was anything ID-shaped left unselected?" test reads correctly on the + * motivating example and misroutes as soon as the line carries an unrelated + * parenthesised citation: `RANGE-01, RANGE-02 — RANGE-05 deferred per + * (ADR-7)` has `(ADR-7)` outside the selector's bracket strip, so a global + * test calls it a drop and sends the line back to the assertive channel — + * reinstating exactly the false "could not be parsed" claim Major 3 is + * about, and contradicting #3697-4, which pins a parenthetical citation as + * NOT unparsed residue. Only the rules that fired may speak. + */ + rangeReadingOnly: boolean; + /** Any rule fired — the line warrants a warning. */ + warn: boolean; +}; + +// A range operator, enumerated. CENSUS (round 3): the domain is "separator +// spellings an author can put between two REQ-IDs", which is open, so the +// enumeration draws a boundary rather than covering it. Reached: ASCII `..`+, +// the seven Unicode dashes that are the SAME operator at different codepoints +// (U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure dash, U+2013 en, +// U+2014 em, U+2015 horizontal bar, U+2212 minus) plus ASCII `-`, U+2026 +// ellipsis, and the words `to`/`thru`/`through`. NOT reached, and the +// consequence is a silent under-selection — #3697's own defect — for that +// spelling: `→`, `~`, `..=`, `..<`, `until`, and `up to` (two tokens, so it +// cannot be one operator token at all). Those stay out deliberately: each is a +// symbol or word with an independent non-range use between two IDs, which is +// the over-warning class #2334 cost three rounds. The Unicode dashes DO carry +// the ASCII hyphen's date/sub-number collision — an earlier round-3 commit +// claimed they did not, and was wrong — so they take the strict arm with it; +// see the rule below. +const REQ_RANGE_DASHES = '\\u2010\\u2011\\u2012\\u2013\\u2014\\u2015\\u2212'; +// EVERY DASH IS STRICT — one rule, whatever the codepoint. `PREFIX-\d+ +// \d+` is also a date (`FY-2026-08`) and a sub-numbered ID (`API-2-01`), and +// that ambiguity is a property of the SHAPE, not of which dash key was pressed. +// The design already chose strictness for ASCII `-` on exactly this trade: a +// bare-hyphen tight range must carry a full ID on BOTH sides. Until round 3 the +// other dashes sat in the loose arm, so `RANGE-01 (target FY-202608)` +// warned while its all-ASCII twin — pinned silent by #3697-4 — did not. That +// inconsistency predates this PR for U+2013/U+2014; round 3 briefly widened it +// to five more codepoints before this commit closed it for all seven. +// The cost is symmetric and already accepted: `RANGE-01, RANGE-0205` +// goes silent, exactly as `RANGE-01, RANGE-02-05` already does today. A bare +// `RANGE-0205` still warns — it selects nothing, so R3 catches it. +// LOOSE stays loose: `..`, `…` and the word operators have no date or +// sub-number reading between two numbers, so they keep the numeric endpoint. +const REQ_RANGE_OP = `(?:\\.{2,}|\\u2026|[${REQ_RANGE_DASHES}]|-|to|thru|through)`; +const REQ_RANGE_OP_LOOSE = `(?:\\.{2,}|\\u2026|to|thru|through)`; +const REQ_RANGE_OP_SYMBOL = `(?:\\.{2,}|\\u2026|[${REQ_RANGE_DASHES}]|-)`; +const REQ_RANGE_TOKEN_RE = new RegExp( + `^([A-Z][A-Z0-9]*)-(?:\\d+)\\s*(?:${REQ_RANGE_OP_LOOSE}\\s*(?:\\1-)?|[-${REQ_RANGE_DASHES}]\\s*\\1-)\\d+$`, + 'i', +); +const REQ_PURE_RANGE_OP_RE = new RegExp(`^${REQ_RANGE_OP}$`, 'i'); +const REQ_GLUED_RANGE_LEAD_RE = new RegExp(`^${REQ_RANGE_OP_SYMBOL}([A-Z][A-Z0-9]*-\\d+)$`, 'i'); +const REQ_GLUED_RANGE_TRAIL_RE = new RegExp(`^([A-Z][A-Z0-9]*-\\d+)${REQ_RANGE_OP}$`, 'i'); +const REQ_ID_SUBSTRING_RE = /[A-Z][A-Z0-9]*-\d+/i; +const REQ_ID_SHAPE_RE = /^[A-Z][A-Z0-9]*-\d+$/i; +const REQ_ID_PARTS_RE = /^([A-Z][A-Z0-9]*)-(\d+)$/i; +// `LETTERS-\d+-\d+` — a date (`FY-2026-08`) or a sub-numbered ID (`API-2-01`). +// REQ_RANGE_TOKEN_RE's strict-dash arm exists precisely to keep this shape +// silent, because nothing at token level can tell the three readings apart. +// Round 4 review Minor 2: the skipped-text rider re-reported it through the +// side door — `REQ_ID_SUBSTRING_RE` is unanchored, so `FY-2026-08` matches as +// `FY-2026` and landed in `unselectedIdShaped`. Whenever any OTHER rule fired +// on a line carrying a date annotation, the warning then told the author to +// "check whether any of it is a requirement" about a date. Not a false +// warning — the line was warning anyway — but false CONTENT, and it is the +// #2334 voice. +// `PREFIX--` — the shape the strict-dash range rule refuses to +// act on because it is equally a date (`FY-2026-08`) and a sub-numbered id +// (`API-2-01`). NO regex separates those: `API-2026-08` is a legal requirement +// id and `FY-26-08` is a date, and both filters that tried scored a miss in +// each direction under the pre-push review's continuation. +// +// So the rider stops adjudicating and starts DISCLOSING. Round 4 Minor 2's +// real complaint was that the rider told the author to check whether a DATE +// was a requirement; the fix is to name the ambiguity rather than to guess at +// it — which is the same thing the two warning voices already do about a +// range separator. +const REQ_AMBIGUOUS_NUMERIC_RE = /^[A-Z][A-Z0-9]*(?:-\d+){2,}$/i; + +// The token-length cap. It bounds REQ_ID_SUBSTRING_RE, the one UNANCHORED +// regex here, which backtracks quadratically on a pathological token. Round 3 +// review Nit 6 objected that the anchored regexes were left uncapped on the +// strength of a comment asserting they scan linearly; they are applied through +// the same cap now, so the claim is enforced rather than asserted. No real +// REQ-ID-carrying token approaches this bound. +const REQ_TOKEN_SCAN_LIMIT = 2048; + +/** + * The Requirements-line warning KINDS, as a stable machine vocabulary (round 4 + * review Major 3). + * + * Before this, the kind existed only in the prose of the message, so every + * consumer and every test had to regex an English sentence — and rewording a + * message silently un-asserted the tests that pinned it. The repo already had + * the settled seam for exactly these semantics: `diffLiveConfig` emits + * `kind:'unverified'` for a truncated scan (`CONTEXT.md`), and + * `WAVE_CLEANUP_WARNING` carries codes in `src/worktree-safety.cts`. + * + * Carried ALONGSIDE the prose, never instead of it. `warnings[]` is a + * documented `string[]` in `phase complete`'s JSON output, rendered by + * execute-phase.md's "If has_warnings is true" step, so changing its element + * shape would be a breaking output-contract change for a shipped command. The + * code is emitted as its own additive `requirements_line_warning` field. + */ +const REQ_LINE_WARNING_CODE = { + /** ID-shaped content was demonstrably not selected — the line failed to parse. */ + misparse: 'req-line-misparse', + /** A range READING is at stake; every endpoint the rule fired on was selected. */ + rangeReading: 'req-line-range-reading', + /** A token past the scan cap means the line was not classified — never that it is clean. */ + unverified: 'req-line-unverified', +} as const; + +type ReqLineWarningCode = (typeof REQ_LINE_WARNING_CODE)[keyof typeof REQ_LINE_WARNING_CODE]; + +/** The formatter's result. `null` still means CLEAN, which is a value, not a failure. */ +type ReqLineWarning = { code: ReqLineWarningCode; message: string }; + +// R4 — a full ID with a trailing statement delimiter glued to it. ANCHORED on +// both ends, so it is linear and needs no cap of its own beyond the token +// length guard its caller applies. +// Zero-width, bidi-control, joiner and variation-selector codepoints. INVISIBLE +// to the author, and the pre-push review's continuation drove the consequence +// from both sides: a line of only these warned with nothing on screen to +// explain it, AND stripping them wholesale from the detector made +// `REQ-01, REQ-02` go SILENT while the selector really did drop REQ-01 — +// #3697's own defect, introduced by the fix for its mirror image. So they are +// never stripped from the line: they are DECORATION on a token (R4 below) and +// absence-of-content for the empty test (visibleContent), which are two +// different questions about the same character. +const REQ_INVISIBLE_RE = /[\u00AD\u200B-\u200F\u2060-\u2064\u2066-\u2069\uFE0F\uFEFF]/g; +// The wrappers R4 shaves. Emphasis, quotes and backticks, because the SELECTOR +// shaves none of them — `**REQ-01**` is genuinely not selected and is a real, +// silent drop. +// +// PARENTHESES ARE DELIBERATELY ABSENT, and this is load-bearing. A parenthesis +// is this rule's citation MARKER, not decoration to shave: `(REQ-02)` and +// `(ADR-7)` are the same shape and the rule declines both. Including them here +// made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was +// never there, and broke #3697-9d's channel routing with it — caught by the +// suite immediately after the widening. +const REQ_WRAPPER_RE = /^["'`*_~“”‘’]+|["'`*_~“”‘’]+$/g; +// An id with a list delimiter glued to EITHER end, once styling is removed. +// The capture is the bare id; a match means the delimiter was ADJACENT to it. +const REQ_DELIMITED_ID_RE = /^[;:]*([A-Z][A-Z0-9]*-\d+)[;:]*$/i; + + + +/** + * CENSUS (round 4): the domain is "separators an author writes between two + * REQ-IDs INSTEAD of a comma" — distinct from the range-operator domain + * censused above, and it had no census at all before this round. + * + * ROUND 4'S CENSUS WAS WRONG, AND THE WAY IT WAS WRONG IS THE LESSON. It swept + * 26 spellings and concluded "exactly two — `; ` and `: `". It reached that + * answer because it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the + * semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for + * every other separator. Different members of the domain were tested in + * different shapes, so the conclusion could not have come out any other way. + * + * Re-swept round 5, fully crossed: 21 separators x {bare, trailing-space, + * leading-space, both-spaces} = 84 combinations, driven through the built + * artifact. 26 select both IDs, 24 under-select and already warn, and + * 34 UNDER-SELECT SILENTLY. All 34 are the same shape — a separator glued to + * exactly ONE of the two IDs, e.g. `REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for + * every punctuation except `,` (the real delimiter) and `;` / `:` (R4). + * Measured silent: | / + & \ > . ! ? • · ؛ ; , - ~ and the word operators + * `and` / `plus` in trailing-space form. + * + * So the honest statement is that R4 covers TWO CHARACTERS of a domain that is + * wide open, not that the domain has two members. The round-4 review + * hand-listed the semicolon; the colon is its sibling and fails identically; + * everything else in that list is disclosed here and NOT caught. Widening the + * delimiter class is a small change and deliberately not made at the end of a + * round: three successive cuts of this rule fired on a citation. + * + * THE GATE IS ADJACENCY, and it is the part to read. Styling is stripped, then + * the delimiter must be touching the id: `REQ-01;`, `;REQ-02`, `**REQ-01;**` + * and the backticked form all qualify. `**REQ-01**;` does NOT — outside the + * styling a `;` is sentence punctuation, which is why `REQ-01, see **REQ-7**; + * next topic` is a citation and not a drop. An INVISIBLE anywhere in the token + * qualifies without an adjacency test, because nobody types one on purpose, so + * it is corruption rather than intent. + * + * Markdown styling on its own is NOT a trigger and NOT reported. It reaches + * the skipped-text rider, which names the id without asserting a drop — but a + * rider only exists inside a MESSAGE, and a message only exists when some rule + * set `warn`. On a line where nothing else fires, `REQ-01, **REQ-02**` is + * wholly silent. Saying it is "left to the rider" reads as coverage and is + * not; #3697-19m pins the silence so this comment cannot drift back. + * + * NOT reached, stated rather than fixed, and the second member is WIDER than + * this comment first claimed: + * - anything inside a parenthetical. A parenthesis is this rule's citation + * MARKER, never decoration to shave — `(REQ-02)` and `(ADR-7)` are the + * same shape and the rule declines both. + * - a decorated id whose prefix is on NO selected id: `REQ-01, FOO-02: x` + * stays silent even when FOO-02 is real. Prefix agreement is what + * separates a drop from a bare citation — `REQ-01, see ADR-7: section 3` + * carries `ADR-7:` in exactly `REQ-01;`'s shape — and it is the module's + * own idiom, not a new heuristic (reqEndpointsImplyInterior already + * requires an agreeing prefix). The gate is NOT complete: a citation that + * DOES share a selected prefix (`ADR-01, see ADR-7: sec 3`) still fires, + * and nothing at token level separates that from a real drop. Saying so is + * the honest position; a prose heuristic on "see" is exactly the free-text + * detector this module exists to avoid. + * The trade, plainly: an under-report on a rare shape over an over-report on a + * common one — the same call the strict-dash rule makes. + */ +function reqDelimiterDroppedIds(rawLine: string, selected: Set, cap: number): string[] { + // MATCHED parenthetical spans are removed OUTRIGHT, not tracked as a depth. + // + // Two bugs died here. A running depth counter let an unbalanced `(` stay open + // to end-of-line and swallow every real drop after it. Promoting a whole + // token to immune because it CONTAINED a matched character then leaked the + // other way: `REQ-01, REQ-02;(note) REQ-03` is one whitespace token, so the + // parenthetical conferred immunity on the `REQ-02;` sitting outside it. + // Deleting the span states what is actually meant — for this rule a citation + // is not on the line — while an UNMATCHED paren is a typo and confers + // nothing. + // + // Square brackets go too, exactly as the SELECTOR strips them: `[REQ-01; + // REQ-02]` is the documented form and was silently dropping REQ-01. + // + // INVISIBLES STAY. They are the evidence this rule reads; the tokenizer + // strips them for the classification rules, and the two sites answer two + // different questions about the same character. + const chars = [...String(rawLine).replace(//g, ' ')]; + const openStack: number[] = []; + for (let i = 0; i < chars.length; i += 1) { + if (chars[i] === '(') openStack.push(i); + else if (chars[i] === ')' && openStack.length > 0) { + const open = openStack.pop() as number; + for (let j = open; j <= i; j += 1) chars[j] = ' '; + } + } + const line = chars.join('').replace(/[[\]]/g, ''); + + // The prefixes actually SELECTED on this line. A dropped id must agree with + // one of them — that is what separates a delimiter typo from a citation, + // since `REQ-01, see ADR-7: sec 3` carries `ADR-7:` in exactly `REQ-01;`'s + // shape. Same-prefix agreement is the module's own idiom, not a new + // heuristic (see reqEndpointsImplyInterior). + const selectedPrefixes = new Set(); + for (const id of selected) { + const m = REQ_ID_PARTS_RE.exec(id); + if (m) selectedPrefixes.add(m[1].toUpperCase()); + } + + const hits: string[] = []; + for (const raw of line.split(/[,\s]+/)) { + if (!raw || raw.length > cap) continue; + // Strip STYLING only. What survives is the id plus whatever was glued + // directly to it. + const core = raw.replace(REQ_INVISIBLE_RE, '').replace(REQ_WRAPPER_RE, ''); + const m = REQ_DELIMITED_ID_RE.exec(core); + if (!m) continue; + const bare = m[1]; + // ADJACENCY IS THE WHOLE RULE. A `;`/`:` touching the id is a list + // separator someone meant; the same character OUTSIDE the styling is + // sentence punctuation — `see **REQ-7**; next topic` cites a requirement + // while `**REQ-01;** REQ-02` fails to list one, and only the delimiter's + // POSITION separates them. An INVISIBLE needs no adjacency test: nobody + // types one on purpose, so anywhere in the token it is corruption rather + // than intent. + const hadAdjacentDelimiter = core !== bare; + REQ_INVISIBLE_RE.lastIndex = 0; + const hadInvisible = REQ_INVISIBLE_RE.test(raw); + REQ_INVISIBLE_RE.lastIndex = 0; + if (!hadAdjacentDelimiter && !hadInvisible) continue; + if (selected.has(bare.toUpperCase())) continue; + const parts = REQ_ID_PARTS_RE.exec(bare); + if (parts && selectedPrefixes.has(parts[1].toUpperCase())) hits.push(bare); + } + return [...new Set(hits)]; +} + +/** Endpoints imply a dropped interior only on an AGREEING prefix and a gap > 1. */ +function reqEndpointsImplyInterior(a: string, b: string): boolean { + const ma = REQ_ID_PARTS_RE.exec(a); + const mb = REQ_ID_PARTS_RE.exec(b); + if (!ma || !mb) return false; + if (ma[1].toUpperCase() !== mb[1].toUpperCase()) return false; + // BigInt keeps the gap exact for numbers past 2^53. + const gap = BigInt(mb[2]) - BigInt(ma[2]); + return gap > 1n || gap < -1n; +} + +function analyzeRequirementsLine(rawLine: string): RequirementsLineAnalysis { + const line = typeof rawLine === 'string' ? rawLine : ''; + // SELECTOR — byte-identical to the pre-extraction expression. + const citedReqIds = line + .replace(/[\[\]]/g, '') + .split(/[,\s]+/) + .map((r) => r.trim()) + .filter(Boolean) + .filter((r) => REQ_ID_SHAPE_RE.test(r)); + + // DETECTOR tokenization. A token with NO alphanumerics is shaved of brackets + // ONLY, so `(..)` surfaces its operator while a bare `..` is not shaved to + // nothing by the punctuation classes. A trailing run of 2+ dots is a glued + // range operator (`REQ-01.. REQ-05`), not sentence punctuation — keep it. + const tokens = line + .replace(//g, ' ') + // Invisibles are removed HERE, for the classification rules — an operator + // spelled `..` is still the range operator, and a line of only + // invisibles yields no tokens at all. R4 works on the RAW line and does + // NOT strip them, because there they are the evidence of a dropped id. + // Removing them in both places is what made `REQ-01, REQ-02` silent; + // removing them in neither is what made `REQ-01 .. REQ-05` + // silent. The two questions have two different answers. + .replace(REQ_INVISIBLE_RE, '') + .split(/[,\s]+/) + .map((t) => { + const trimmed = t.trim(); + if (!/[A-Za-z0-9]/.test(trimmed)) { + return trimmed.replace(/^[[({]+/, '').replace(/[\])}]+$/, ''); + } + if (/\.{2,}$/.test(trimmed)) { + return trimmed.replace(/^[[({"'`*_~“”‘’]+/, ''); + } + return trimmed.replace(/^[[({"'`*_~“”‘’]+/, '').replace(/[\])}.;:"'`*_~“”‘’]+$/, ''); + }) + .filter(Boolean); + + // Every predicate below is applied through the scan limit (Nit 6): a token + // past the bound is not classified at all rather than classified expensively. + const short = (t: string): boolean => t.length <= REQ_TOKEN_SCAN_LIMIT; + const rangeTokens = tokens.filter((t) => short(t) && REQ_RANGE_TOKEN_RE.test(t)); + const spacedRangePairs: Array<[string, string]> = []; + tokens.forEach((t, i) => { + const left = tokens[i - 1] ?? ''; + const right = tokens[i + 1] ?? ''; + if ( + // EVERY participant is capped, not just the operator. Capping the operator + // alone left `<2049-char ID> .. <2049-char ID>` running REQ_ID_SHAPE_RE and + // BigInt over both neighbours unbounded — the cap read as uniform and was + // not (found by the round's pre-push review). + short(t) && + short(left) && + short(right) && + REQ_PURE_RANGE_OP_RE.test(t) && + i > 0 && + i < tokens.length - 1 && + REQ_ID_SHAPE_RE.test(left) && + REQ_ID_SHAPE_RE.test(right) && + reqEndpointsImplyInterior(left, right) + ) { + spacedRangePairs.push([left, right]); + } + }); + const hasSpacedRange = spacedRangePairs.length > 0; + // A half-spaced range splits at the tokenizer, so R1's own `\s*` never sees + // it. SYMBOL operators only on the LEAD arm: a word operator glued to an ID + // is an ID — `TORANGE-05` is a valid prefix-agnostic REQ-ID. The TRAIL arm + // keeps the word operators, because a valid ID must end in digits, so + // `REQ-01through` can only be a glued typo. + const hasGluedRangeFragment = tokens.some((t, i) => { + // Neighbours capped for the same reason as R2 above. + if (!short(t)) return false; + const before = tokens[i - 1] ?? ''; + const after = tokens[i + 1] ?? ''; + const lead = REQ_GLUED_RANGE_LEAD_RE.exec(t); + if ( + lead && + i > 0 && + short(before) && + REQ_ID_SHAPE_RE.test(before) && + reqEndpointsImplyInterior(before, lead[1]) + ) { + return true; + } + const trail = REQ_GLUED_RANGE_TRAIL_RE.exec(t); + return Boolean( + trail && + i < tokens.length - 1 && + short(after) && + REQ_ID_SHAPE_RE.test(after) && + reqEndpointsImplyInterior(trail[1], after), + ); + }); + const leadToken = (tokens[0] ?? '').toUpperCase(); + // CENSUS (round 3, review finding Minor 4): the placeholder domain is what + // GSD itself seeds plus what an author writes for "deliberately empty". + // Reached: `TBD` — the ONLY machine-written seed, at the three phase.add / + // -batch / -insert sites — and `None`, the author convention. NOT reached: + // `N/A`, `Deferred`, `Pending`, `TBA`, `-`. Consequence, and it is now + // ENFORCED rather than asserted: such a line selects zero IDs and warns + // through R3b below, which is what #3697's acceptance criterion asks for + // ("when it selects zero IDs from a line that is non-empty and is not the + // `TBD` placeholder"). Round 3 shipped this same paragraph while R3's + // ID-shape gate made it false for all five words — bare `Deferred` was + // silent, `Deferred (see ADR-7)` warned — and the claim sat in three + // artifacts with no test in either direction. Inferring placeholder-ness + // from arbitrary prose is still the free-text heuristic this detector + // avoids: R3b keys on the SELECTION being empty, never on what the prose + // means. + const placeholderLed = leadToken === 'TBD' || leadToken === 'NONE'; + const inertIdShaped = + citedReqIds.length === 0 && !placeholderLed + ? tokens.filter((t) => short(t) && t.includes('-') && REQ_ID_SUBSTRING_RE.test(t)) + : []; + // R3b — the acceptance criterion's own narrow form. `tokens.length > 0` is + // what keeps an empty line and a comment-only line silent: the tokenizer + // strips `` before splitting, so `` yields no + // tokens and cannot reach this rule. Every other zero-selection, + // non-placeholder line warns. + const zeroSelectionInert = citedReqIds.length === 0 && !placeholderLed && tokens.length > 0; + + // R2 is the ONLY ambiguous rule — a tight range, a glued fragment and R3 + // residue each implicate ID-shaped text the selector demonstrably did not + // take, so any of them means the line really did fail to parse. R2 is + // ambiguous only when its OWN endpoints were selected: the detector shaves + // brackets and the selector does not, so R2 can fire on a `(RANGE-02)` that + // was never selected — a real drop, and the assertive channel is right there. + // A token past the cap is NOT classified — and must therefore not be + // silently discarded. Round 3's first cut of the uniform cap did exactly + // that: a 2049-char range token warned before the round and went silent + // after it, which is #3697's own defect introduced by the fix for a nit + // (found by the round's pre-push review). The cap bounds the WORK, not the + // warning — so an over-cap token that could carry an ID is reported as + // unclassified. The test is `includes('-')`, a linear scan, never the + // unanchored regex the cap exists to keep off these tokens. + // ANY over-cap token, not just one carrying `-`. The first cut filtered on + // `includes('-')` and therefore missed an over-cap OPERATOR: + // `REQ-01 <2049 dots> REQ-05` warned before this round (R2 was uncapped) and + // went silent after it. A token we could not examine makes the line + // unverified whatever characters it happens to contain. Computed below, + // where the selected set is available. + + // ID-shaped tokens the selector did not take, ANYWHERE on the line. This is + // reported as a fact, never used to pick the channel: `(ADR-7)` and + // `(REQ-02)` are indistinguishable by shape, so routing on it would put the + // false "could not be parsed" claim back on a line carrying a citation. + // Naming them lets the author see what the tokenizer skipped without the + // warning asserting a verdict it cannot support in either direction. + const selected = new Set(citedReqIds.map((id) => id.toUpperCase())); + // A token the SELECTOR took has had its own SELECTION verified — the selector + // is uncapped and anchored, so it examined the whole token. That is not the + // same as "no rule was suppressed by it", and conflating the two was the + // second continuation review's CLAIM J/K: two over-cap valid IDs either side + // of `..` are both selected, both exempted, and R2 is capped — so a line that + // warned before this round went silent, which is the very regression the + // field exists to close, arriving through the fix for its own over-report. + // + // The exemption therefore applies only when nothing could have been + // suppressed: an over-cap token that was selected AND has no neighbour that + // could pair with it into a range. Everything else is unexaminable and is + // reported as such. + const couldPairIntoRange = (i: number): boolean => { + for (const n of [tokens[i - 1], tokens[i + 1]]) { + if (n === undefined) continue; + if (!short(n)) return true; + if (REQ_PURE_RANGE_OP_RE.test(n)) return true; + if (REQ_GLUED_RANGE_LEAD_RE.test(n) || REQ_GLUED_RANGE_TRAIL_RE.test(n)) return true; + } + return false; + }; + const oversizedTokens = tokens.filter( + (t, i) => !short(t) && (!selected.has(t.toUpperCase()) || couldPairIntoRange(i)), + ); + const unselectedIdShaped = tokens.filter( + (t) => short(t) && REQ_ID_SUBSTRING_RE.test(t) && !selected.has(t.toUpperCase()), + ); + // R4 runs on the RAW line, not on `tokens`: the shave that makes `REQ-01;` + // look like a clean `REQ-01` is exactly the evidence this rule needs, so it + // has to see the character the tokenizer removed. + const delimiterDroppedIds = reqDelimiterDroppedIds(rawLine, selected, REQ_TOKEN_SCAN_LIMIT); + const nothingDemonstrablyDropped = + rangeTokens.length === 0 && + !hasGluedRangeFragment && + inertIdShaped.length === 0 && + // R4 is a DEMONSTRATED drop, so neither non-assertive voice — one claiming + // nothing was dropped, the other that nothing could be checked — may speak + // for a line carrying one. + delimiterDroppedIds.length === 0 && + // R2 firing on an endpoint the selector did NOT take is itself a + // demonstrated drop, and the assertive channel is right there. Vacuously + // true when no spaced range fired, which is what makes this a strict + // superset of the `!hasSpacedRange` guard the over-cap channel used to + // carry — that channel's behaviour on a line with no spaced range is + // unchanged, byte for byte. + spacedRangePairs.every(([a, b]) => selected.has(a.toUpperCase()) && selected.has(b.toUpperCase())); + const rangeReadingOnly = + hasSpacedRange && + nothingDemonstrablyDropped && + // The cap bounds the WORK, never the warning. An over-cap token is not + // classified by ANY rule (R1-R4 all skip it), so the voice whose entire + // claim is that nothing was dropped has no basis to speak for this line. + // It falls to the over-cap channel below instead — `unverified`, because + // the line was not CHECKED; not `misparse`, because nothing on it + // demonstrably failed to parse either. Round 7 review, Minor 1. + oversizedTokens.length === 0; + + // Named rather than inlined into the return literal (round 3 review Minor 3): + // this disjunction is the module's single most important predicate, and in + // the literal a later edit that reordered a local below the `return` would be + // a TDZ ReferenceError at runtime rather than an error at the reader's eye + // level. R3b joins it here — see its field docs above for why it is not + // gated on ID shape. + const warn = + rangeTokens.length > 0 || + hasSpacedRange || + hasGluedRangeFragment || + inertIdShaped.length > 0 || + zeroSelectionInert || + delimiterDroppedIds.length > 0 || + oversizedTokens.length > 0; + + return { + citedReqIds, + tokens, + rangeTokens, + hasSpacedRange, + hasGluedRangeFragment, + inertIdShaped, + zeroSelectionInert, + placeholderLed, + spacedRangePairs, + nothingDemonstrablyDropped, + rangeReadingOnly, + delimiterDroppedIds, + oversizedTokens, + unselectedIdShaped, + warn, + }; +} + +/** + * Render the warning, or null when the line is clean. + * + * TWO CHANNELS, and the split is round 3's fix for review finding Major 3. The + * detector cannot distinguish `RANGE-02 — RANGE-05` meaning a range from the + * same text meaning an annotation separator; they are textually identical and + * no token-level rule separates them. What the old single-channel message did + * was resolve that ambiguity by ASSERTION — it told the author the line "could + * not be parsed" and to rewrite it, on a line where every ID present had in + * fact been selected and nothing had been dropped. That is a false statement + * under the annotation reading and the #2334 over-warning class. + * + * Going silent instead is not available: the range reading is equally live, and + * staying quiet on it re-opens the exact silent under-selection #3697 is about. + * So the ambiguity is DISCLOSED rather than decided — + * + * * any rule other than R2 fired, or R2 fired on an endpoint that was not + * selected → something ID-shaped was demonstrably NOT taken. The line did + * fail to parse; say so plainly, as before. + * * R2 alone fired and both its endpoints were selected → nothing was + * dropped. State both readings and let the author pick; never claim a parse + * failure that did not occur. + */ +function formatRequirementsLineWarning( + phaseNum: string, + rawLine: string, + analysis: RequirementsLineAnalysis, +): ReqLineWarning | null { + if (!analysis.warn) return null; + const shown = String(rawLine).trim(); + const rangeRuleFired = + analysis.rangeTokens.length > 0 || analysis.hasSpacedRange || analysis.hasGluedRangeFragment; + + // Tokens the selector skipped, stated as a fact in EITHER channel. `(ADR-7)` + // and `(REQ-02)` are the same shape, so no rule can say which one matters — + // but the author can, and only if the warning tells them. Round 3's first + // cut instead let this drive the channel, which put the false "could not be + // parsed" claim back on a line carrying a citation. + // Names only what the rule-specific clauses did NOT already name, so the + // assertive voice can carry it too without repeating itself. + const alreadyNamed = new Set( + [...analysis.rangeTokens, ...analysis.inertIdShaped, ...analysis.delimiterDroppedIds].map((t) => + t.toUpperCase(), + ), + ); + const skippedNames = analysis.unselectedIdShaped.filter((t) => !alreadyNamed.has(t.toUpperCase())); + // Named, then qualified. The `PREFIX-N-N` shape is the one the range rules + // deliberately decline to act on, so the rider says WHY it might not be a + // requirement instead of silently deciding it is not. + const ambiguousNamed = skippedNames.filter((t) => REQ_AMBIGUOUS_NUMERIC_RE.test(t)); + const skipped = + skippedNames.length > 0 + ? ` ID-shaped text on the line that was NOT selected: ${skippedNames.join(', ')}` + + ` (parentheses are not stripped, unlike square brackets) — check whether any of it is a` + + ` requirement.` + + (ambiguousNamed.length > 0 + ? ` ${ambiguousNamed.join(', ')} may equally be a date or a sub-numbered id, which is` + + ` why the range rules do not act on that shape.` + : '') + : ''; + // R4's clause. Named separately from the generic skipped-text rider because + // this one is not a "check whether any of it is a requirement" hedge — the + // token IS an ID, the selector demonstrably did not take it, and the cause + // is nameable. + const delimiterDropped = + analysis.delimiterDroppedIds.length > 0 + ? ` ${analysis.delimiterDroppedIds.join(', ')} ${analysis.delimiterDroppedIds.length === 1 ? 'was' : 'were'}` + + ` NOT selected: a \`;\` or \`:\` is glued to the ID, or it carries an invisible character, and` + + ` the line is split on commas and whitespace only. Write each requirement as a bare ID` + + ` separated by a comma.` + : ''; + const oversized = + analysis.oversizedTokens.length > 0 + ? ` One or more tokens exceed the ${REQ_TOKEN_SCAN_LIMIT}-character scan limit and were NOT` + + ` classified, so this line may carry more than is reported here.` + : ''; + + if (analysis.rangeReadingOnly) { + // AMBIGUOUS channel — the RANGE reading is what is at stake, not a parse + // failure: every endpoint the range rule fired on was selected. + // + // What this voice must NOT do is claim the whole LINE is correct. It has + // no basis for that: an unrelated `(REQ-02)` elsewhere on the line is + // dropped by the selector and invisible to every rule, so "nothing needs + // to change" is an affirmative false statement on exactly the input the + // rule-scoped discriminator was built to reach. It speaks about the + // SEPARATOR, and defers the rest to the skipped-text clause above. + return { + code: REQ_LINE_WARNING_CODE.rangeReading, + message: + `ROADMAP Phase ${phaseNum} **Requirements** line (\`${shown}\`) contains what reads as a range ` + + `between two cited REQ-IDs. Range forms are not expanded, so no interior IDs were selected; ` + + `the line selected: ${analysis.citedReqIds.join(', ')}. If a range was intended, rewrite it ` + + `naming every requirement explicitly (e.g. \`REQ-01, REQ-02, REQ-03\`); if that separator is ` + + `an annotation rather than a range, it selected nothing to expand and needs no change.` + + delimiterDropped + + skipped + + oversized, + }; + } + + if (analysis.oversizedTokens.length > 0 && analysis.nothingDemonstrablyDropped) { + // A DEMONSTRATED drop outranks this voice, whose whole claim is that + // NOTHING could be checked — both cannot be true at once. `REQ-01, + // REQ-02: ` names REQ-02 in `delimiterDroppedIds` and + // then reported `req-line-unverified`, whose message never mentions it: + // the concrete, actionable finding masked by the token beside it. That + // exclusion now lives in `nothingDemonstrablyDropped`, shared verbatim + // with `rangeReadingOnly` above rather than duplicated here — the + // duplication is what let the two drift (round 7 review, Minor 1). The + // assertive channel already appends the over-cap rider, so routing a + // demonstrated drop there loses nothing about the cap. + // OVER-CAP channel — no rule could run, so no rule may be diagnosed. Say + // exactly that: the line was not classified, rather than not a problem. + return { + code: REQ_LINE_WARNING_CODE.unverified, + message: + `ROADMAP Phase ${phaseNum} **Requirements** line (\`${shown.slice(0, 200)}…\`) could not be ` + + `checked: one or more tokens exceed the ${REQ_TOKEN_SCAN_LIMIT}-character scan limit, so the ` + + `REQ-ID selection on this line is unverified. Rewrite it as a comma-separated list ` + + `(e.g. \`REQ-01, REQ-02, REQ-03\`).`, + }; + } + + // ASSERTIVE channel — ID-shaped content was demonstrably not selected. + // Deliberately says "selected", NOT "marked complete": a range whose + // endpoints are themselves unregistered selects them and marks nothing, and a + // warning that overclaims the write is a warning the reader learns to + // distrust. + const selectedDesc = + analysis.citedReqIds.length > 0 + ? `the only REQ-ID(s) selected from it were: ${analysis.citedReqIds.join(', ')}` + : 'it selected NO REQ-IDs at all, so nothing was marked'; + const unparsed = [...new Set([...analysis.rangeTokens, ...analysis.inertIdShaped])]; + // Only diagnose "range" when a range rule actually fired — an R3 warning on + // non-range ID text must not claim one was written. And on the R3 path the + // residue is ID-SHAPED TEXT, which is not the same claim as "a requirement we + // failed to parse" (round 3 review finding Minor 4: `Deferred (see ADR-7)` + // reported `ADR-7` as missed requirement content when it is a citation). Name + // what it is, and name the placeholder escape the author actually has. + const advice = rangeRuleFired + ? ' Range forms are not expanded; rewrite the line naming every requirement explicitly ' + + '(e.g. `REQ-01, REQ-02, REQ-03`).' + : ' If these are requirements, name them explicitly (e.g. `REQ-01, REQ-02, REQ-03`); if the line ' + + 'is deliberately empty, write `TBD` or `None` — any other wording selects nothing and warns.'; + return { + code: REQ_LINE_WARNING_CODE.misparse, + message: + `ROADMAP Phase ${phaseNum} **Requirements** line could not be parsed as a comma-separated REQ-ID list ` + + `(\`${shown}\`) - ${selectedDesc}.` + + (unparsed.length > 0 + ? rangeRuleFired + ? ` Unparsed text: ${unparsed.join(', ')}.` + : ` ID-shaped text that was not selected: ${unparsed.join(', ')}.` + : '') + + advice + + delimiterDropped + + skipped + + oversized, + }; +} + function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { if (!phaseNum) { error('phase number required for phase complete'); @@ -2493,6 +3284,11 @@ function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { let stateUpdated = false; const warnings: string[] = []; + // The machine kind of the Requirements-line warning, carried out to the JSON + // result as its own field (round 4 review Major 3). Declared HERE, in the + // same scope as `warnings[]`, because the assignment happens inside + // withPlanningLock and the emission happens after it. + let reqLineWarningCode: ReqLineWarningCode | undefined; // ADR-3408 §8.5 / D2 (#3374): "liberal but visible" — when the write-seam // composition's preservation stage restores a curated frontmatter value // over a disagreeing derived one, that divergence is surfaced here rather @@ -2957,27 +3753,26 @@ function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { const traceabilityWriteMisses: string[] = []; if (reqMatch) { - // #2334 HIGH 3: filter the tokenized capture to the REQ-ID SHAPE — - // the SAME shape bodyReqIds (`\*\*([A-Z][A-Z0-9]*-\d+)\*\*`, below) - // and tableReqIds (`([A-Z][A-Z0-9]*-\d+)`, below) already require — - // so the ghost-ID / unregistered comparisons stay shape-symmetric. - // Without this, `[^\n]+` split on `[,\s]+` turned EVERY word after - // the ID list into a "cited REQ-ID": the shipped - // `templates/roadmap.md:32` line - // `**Requirements**: [REQ-01, REQ-02] ` - // warned to register ``, etc., and - // `**Requirements:** None` warned to register the literal word - // `None`. This subsumes the `TBD` placeholder special-case (`TBD` - // does not match the REQ-ID shape either); `isPlaceholderReqId` is - // kept below as a defensive no-op for any caller that still hands - // it a raw token. - const REQ_ID_SHAPE_RE = /^[A-Z][A-Z0-9]*-\d+$/i; - citedReqIds = reqMatch[1] - .replace(/[\[\]]/g, '') - .split(/[,\s]+/) - .map((r) => r.trim()) - .filter(Boolean) - .filter((r) => REQ_ID_SHAPE_RE.test(r)); + // #2334 HIGH 3 + #3697: selection and under-selection detection both + // live in `analyzeRequirementsLine` (module scope, above), extracted in + // round 3 so the parser is directly testable — a closure in here is + // reachable only by spawning the CLI, which no fast-check property test + // can do. `citedReqIds` is byte-identical to the expression that stood + // here; nothing about what phase-complete MARKS has changed. + const reqLineAnalysis = analyzeRequirementsLine(reqMatch[1]); + citedReqIds = reqLineAnalysis.citedReqIds; + const reqLineWarning = formatRequirementsLineWarning( + phaseNum, + reqMatch[1], + reqLineAnalysis, + ); + if (reqLineWarning) { + warnings.push(reqLineWarning.message); + // Carried out to the JSON result as its own field — see + // REQ_LINE_WARNING_CODE for why it is not folded into + // `warnings[]`. + reqLineWarningCode = reqLineWarning.code; + } for (const reqId of citedReqIds) { const reqEscaped = escapeRegex(reqId); @@ -3644,6 +4439,10 @@ function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { auto_pruned: autoPruned, warnings, has_warnings: warnings.length > 0, + // ADDITIVE, never a change to `warnings[]`'s element shape — that array is + // a documented string[] consumed by execute-phase.md, so re-typing it + // would break a shipped output contract. Absent when the line is clean. + ...(reqLineWarningCode ? { requirements_line_warning: { code: reqLineWarningCode } } : {}), verification_stale_check_indeterminate: staleCheckIndeterminate, milestone_conflict: milestoneConflict, preservation_warnings: preservationWarnings, @@ -3724,6 +4523,9 @@ export = { cmdPhaseInsert, cmdPhaseRemove, cmdPhaseComplete, + analyzeRequirementsLine, + formatRequirementsLineWarning, + REQ_LINE_WARNING_CODE, cmdPhaseUatPassed, cmdPhaseListPlans, computeDependencyLevels, diff --git a/tests/phase.test.cjs b/tests/phase.test.cjs index 149c2b445..fe4996cf8 100644 --- a/tests/phase.test.cjs +++ b/tests/phase.test.cjs @@ -11771,6 +11771,1840 @@ describe('issue #2334: ghost-REQ-ID classification must probe write surfaces, no ); }); +// ───────────────────────────────────────────────────────────────────────────── +// Regressions: issue #3697 — the `**Requirements**:` tokenizer UNDER-selects +// silently. #2339 fixed OVER-selection (the shape filter) and added the +// ghost-ID cross-check; neither covers an ID the tokenizer DROPPED. A spaced +// range (`RANGE-01 … RANGE-05`) survives the `[,\s]+` split as its two +// endpoints and marks only those; a tight range (`RANGE-01…05`) survives as one +// token, fails the anchored shape filter, and marks nothing. Both paths report +// success with `warnings: []`, because `ghostReqIds` is itself +// `citedReqIds.filter(...)` and so cannot see an ID that was never selected. +// +// The fix warns; it does not parse. The selected set is unchanged (range syntax +// stays unsupported), and the trigger is ID-SHAPED EVIDENCE only — the #2334 / +// #2339 over-warning on `None`, on the shipped `` template comment, +// and on parenthetical annotations must not return. +// ───────────────────────────────────────────────────────────────────────────── + +const REQ_LINE_MISPARSE_RE = /could not be parsed as a comma-separated REQ-ID list/i; +// The AMBIGUOUS channel, added in round 3. A spaced separator between two +// selected, interior-implying REQ-IDs cannot be told from a range by any +// token-level rule, and every ID on such a line IS selected — so the warning +// discloses both readings instead of asserting a parse failure that did not +// happen. See `formatRequirementsLineWarning` in src/phase.cts. +const REQ_LINE_RANGE_READING_RE = /contains what reads as a range between two cited REQ-IDs/i; + +function build3697RangeFixture(roadmapRequirementsLine) { + return build2334GhostSurfaceFixture({ + reqBody: [ + '## Functional Requirements', + '', + '- [ ] **RANGE-01**: synthetic fixture requirement 1.', + '- [ ] **RANGE-02**: synthetic fixture requirement 2.', + '- [ ] **RANGE-03**: synthetic fixture requirement 3.', + '- [ ] **RANGE-04**: synthetic fixture requirement 4.', + '- [ ] **RANGE-05**: synthetic fixture requirement 5.', + ], + roadmapRequirementsLine, + traceabilityRows: [ + '| RANGE-01 | Phase 01 | Pending |', + '| RANGE-02 | Phase 01 | Pending |', + '| RANGE-03 | Phase 01 | Pending |', + '| RANGE-04 | Phase 01 | Pending |', + '| RANGE-05 | Phase 01 | Pending |', + ], + }); +} + +const tickedReqIds = (reqContent) => + [...reqContent.matchAll(/-\s*\[x\]\s*\*\*(RANGE-\d+)\*\*/gi)].map((m) => m[1].toUpperCase()); + +describe('issue #3697: phase complete must warn when the Requirements line under-selects', () => { + for (const [label, line] of [ + ['ellipsis', 'RANGE-01 … RANGE-05'], + ['hyphen', 'RANGE-01 - RANGE-05'], + ['worded', 'RANGE-01 through RANGE-05'], + ['parenthesized operator', 'RANGE-01 (..) RANGE-05'], + ]) { + test( + `#3697-1 (${label} spaced range): a range that survives the split as its two endpoints must warn — ` + + 'and the endpoint-only marking behavior itself is UNCHANGED', + (t) => { + const tmpDir = build3697RangeFixture(line); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_RANGE_READING_RE.test(w)), + `#3697-1 FAILED (${label}): a spaced range marks ONLY its endpoints, so it must warn. ` + + `Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + assert.strictEqual( + parsed.has_warnings, true, + `#3697-1 FAILED (${label}): has_warnings must be true, got: ${JSON.stringify(parsed)}`, + ); + // The warning must name what WAS selected, so the author can see the gap. + assert.ok( + warnings.some( + (w) => REQ_LINE_RANGE_READING_RE.test(w) && /RANGE-01/.test(w) && /RANGE-05/.test(w), + ), + `#3697-1 FAILED (${label}): the warning must name the IDs actually selected ` + + `(RANGE-01, RANGE-05), got: ${JSON.stringify(warnings)}`, + ); + // Round 3 (review finding Major 3): nothing on this line was dropped — + // both endpoints were selected — so the warning must NOT assert that + // the line failed to parse. It offers the range reading and the + // annotation reading and lets the author choose. + assert.ok( + !warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-1 FAILED (${label}): a spaced range drops nothing the tokenizer could have taken, ` + + `so the misparse channel must stay silent, got: ${JSON.stringify(warnings)}`, + ); + // Behavior guard: this is a warning, NOT range support. Exactly the two + // endpoints stay ticked; the interior IDs are still not expanded. + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent).sort(), ['RANGE-01', 'RANGE-05'], + `#3697-1 FAILED (${label}): the selected set must be UNCHANGED (endpoints only — ranges are ` + + `deliberately not expanded).\nREQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + for (const [label, line] of [ + ['ellipsis', 'RANGE-01…05'], + ['double-dot', 'RANGE-01..RANGE-05'], + ]) { + test( + `#3697-2 (${label} tight range): a line that selects ZERO IDs while being non-empty and not TBD ` + + 'must warn, and must write nothing to the ledger', + (t) => { + const tmpDir = build3697RangeFixture(line); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-2 FAILED (${label}): a zero-selection line is completely inert and must warn. ` + + `Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w) && /RANGE-01/.test(w)), + `#3697-2 FAILED (${label}): the warning must name the unparsed ID-shaped text, ` + + `got: ${JSON.stringify(warnings)}`, + ); + assert.strictEqual( + parsed.requirements_updated, false, + `#3697-2 FAILED (${label}): nothing was selected, so requirements_updated must be false, ` + + `got: ${JSON.stringify(parsed)}`, + ); + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent), [], + `#3697-2 FAILED (${label}): no checkbox may be ticked.\nREQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + for (const [label, line] of [ + ['bare comma list', 'RANGE-01, RANGE-02, RANGE-03, RANGE-04, RANGE-05'], + ['bracketed comma list', '[RANGE-01, RANGE-02, RANGE-03, RANGE-04, RANGE-05]'], + ]) { + test( + `#3697-3 (control, ${label}): the canonical form must mark every ID and must not warn`, + (t) => { + const tmpDir = build3697RangeFixture(line); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + // Whole-channel assertion: the fixture registers every cited ID, so + // there is NO legitimate warning here. Filtering by the current + // phrase would let a re-worded over-warning slip through (review + // claim 8) — assert actual silence, not absence of one wording. + assert.deepStrictEqual( + warnings, [], + `#3697-3 FAILED (${label}): the canonical comma list must never warn, ` + + `got: ${JSON.stringify(warnings)}`, + ); + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent).sort(), + ['RANGE-01', 'RANGE-02', 'RANGE-03', 'RANGE-04', 'RANGE-05'], + `#3697-3 FAILED (${label}): all five requirements must be ticked.\n` + + `REQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + // The #2334/#2339 over-warning must not return. The trigger is ID-shaped + // evidence — NOT "the line has residue after the selected IDs are removed", + // which is exactly the check that once warned to register `'], + ['parenthetical citation', 'RANGE-01, RANGE-02 (locked per ADR-7)'], + // The next four are the false-positive classes a free-text detector + // produced (the #2334/#2339 regression shape): a hyphen after the list is + // an ANNOTATION separator, not a range operator, whenever what follows is + // not itself a REQ-ID — and an ID-shaped citation on a correctly-parsed + // line is a citation, not unparsed residue. + ['numeric estimate annotation', 'RANGE-01, RANGE-02 - 3 points'], + ['date annotation', 'RANGE-01, RANGE-02 - 2026-08-21 target'], + ['em-dash citation with trailing period', 'RANGE-01, RANGE-02 — locked per ADR-7.'], + ['nested parenthetical citations', 'RANGE-01, RANGE-02 (see (ADR-7), then ADR-8)'], + // A hyphen between two ADJACENT selected IDs can drop nothing (there is no + // interior), so it reads as an annotation separator, not a range. + ['adjacent-ID annotation hyphen', 'RANGE-01, RANGE-02 - RANGE-03 deferred'], + // Cross-prefix pairs around a separator are annotations, not ranges — real + // ranges are same-prefix by nature. + ['hyphen citation annotation', 'RANGE-01, RANGE-02 - (ADR-7)'], + ['ellipsis-of-omission with citation', 'RANGE-01, RANGE-02 (...) (ADR-7)'], + // Markdown emphasis around a placeholder must not defeat the placeholder + // gate. + ['bold placeholder with citation', '**None** (per ADR-7)'], + // `LETTERS-\d+-\d+` is also a date: the bare-hyphen tight-range arm demands + // a full ID on both sides precisely so this stays silent. + ['date-like parenthetical annotation', 'RANGE-01 (target FY-2026-08)'], + // A declared-empty line citing its rationale — zero selection with ID-shaped + // text, but placeholder-led. The #2334 class R3 must not recreate. + ['None with citation', 'None (per ADR-7)'], + ['TBD with citation', 'TBD (see ADR-7)'], + ]) { + test( + `#3697-4 (negative space, ${label}): must stay silent — the historical over-warning must not return`, + (t) => { + const tmpDir = build3697RangeFixture(line); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + // Whole-channel assertion, same rationale as #3697-3: every ID these + // fixtures cite is registered, so nothing here may warn at all. A + // phrase-filtered check pins wording, not silence (review claim 8). + assert.deepStrictEqual( + warnings, [], + `#3697-4 FAILED (${label}): ${JSON.stringify(line)} must not produce ANY warning, ` + + `got: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + }, + ); + } + + // A range hidden INSIDE balanced parentheses is stripped by the token shave + // and must still warn (review claim 2's first reproduction: the selector + // sees only `RANGE-01`, and v1's scan saw nothing at all). Partial + // selection: exactly RANGE-01 is ticked; the range token is reported as + // unparsed; nothing is expanded. + for (const [label5, line5] of [ + ['parenthesized', 'RANGE-01 (plus RANGE-02..RANGE-05)'], + ['backticked', 'RANGE-01, `RANGE-02..RANGE-05`'], + ['bold-wrapped', 'RANGE-01, **RANGE-02..RANGE-05**'], + ['underscore-wrapped', 'RANGE-01, _RANGE-02..RANGE-05_'], + ]) { + test( + `#3697-5 (${label5} tight range): a wrapped range must still warn and stay unexpanded`, + (t) => { + const tmpDir = build3697RangeFixture(line5); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-5 FAILED (${label5}): the wrapped range selects only RANGE-01, so it must warn. ` + + `Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w) && /RANGE-02\.\.RANGE-05/.test(w)), + `#3697-5 FAILED (${label5}): the warning must name the unparsed range token ` + + `(RANGE-02..RANGE-05), got: ${JSON.stringify(warnings)}`, + ); + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent), ['RANGE-01'], + `#3697-5 FAILED (${label5}): exactly RANGE-01 must be ticked (no expansion, no extra writes).\n` + + `REQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + // A valid prefix-agnostic ID that HAPPENS to start with a word operator + // (`TORANGE-05` reads as `to` + `RANGE-05`) is an ID, not a glued range — + // the glued-fragment rule is symbol-operator-only for exactly this reason. + // The ID is unregistered in this fixture, so the pre-existing ghost-ID + // warning legitimately fires; only the misparse channel must stay silent. + test( + '#3697-4b (word-operator-prefixed ID): a canonical list must not read as a glued range', + (t) => { + const tmpDir = build3697RangeFixture('RANGE-01, TORANGE-05'); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + // Round 4 review Minor 1: this asserted only that the ASSERTIVE channel + // stayed silent, so a regression routing the line into the AMBIGUOUS + // channel would have passed. Whole-channel silence is not available here + // — the pre-existing ghost-ID warning legitimately fires on the + // unregistered `TORANGE-05` — so the precise assertion is that NO + // Requirements-line warning of ANY kind was emitted. The machine code + // added in this round is what makes that statable at all. + assert.strictEqual( + parsed.requirements_line_warning, undefined, + `#3697-4b FAILED: "RANGE-01, TORANGE-05" is a comma list of two valid IDs and must not ` + + `produce a Requirements-line warning in ANY channel, got kind ` + + `${JSON.stringify(parsed.requirements_line_warning)} / ${JSON.stringify(warnings)}\n` + + `Full output: ${output}`, + ); + }, + ); + + // ========================================================================== + // #3697-14 — AC-1b / AC-4: zero selection on a non-placeholder line. + // + // The issue's narrow clause verbatim: "warn when `citedReqIds.length === 0` + // while the raw capture is non-empty and not `TBD`". Before R3b these five + // words selected nothing and stayed SILENT, while the CLI-TOOLS reference's + // Requirements-line grammar section, the `placeholderLed` census comment and + // the advice string all stated they warned — a claim written into three + // artifacts and never executed once. (Named without spelling the doc's path: + // the docs-guard exemption ratchet fingerprints literal docs/ references in + // exempt test files, and this test reads no documentation.) + // The asymmetry was the tell: `Deferred (see ADR-7)` warned (the citation + // supplied ID-shaped residue) while bare `Deferred` did not. + // ========================================================================== + for (const [label14, line14] of [ + ['Deferred', 'Deferred'], + ['N/A', 'N/A'], + ['Pending', 'Pending'], + ['TBA', 'TBA'], + ['bare dash', '-'], + // Prose with no ID-shaped token anywhere — the class R3's ID-shape gate + // could never reach, whatever the wording. + ['free prose', 'to be scoped after the spike'], + ]) { + test( + `#3697-14 (zero selection, ${label14}): a non-empty, non-placeholder line that selects NO ` + + 'REQ-IDs must warn — #3697 AC-1b/AC-4', + (t) => { + const tmpDir = build3697RangeFixture(line14); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-14 FAILED (${label14}): ${JSON.stringify(line14)} is non-empty, is not a ` + + `placeholder and selects zero REQ-IDs, so it must warn. ` + + `Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + // The warning must name the escape the author actually has, or it + // reports a problem with no remedy. + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w) && /`TBD` or `None`/.test(w)), + `#3697-14 FAILED (${label14}): the warning must name the TBD/None placeholder escape, ` + + `got: ${JSON.stringify(warnings)}`, + ); + // Selection behavior is UNCHANGED — this rule warns, it never invents IDs. + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent), [], + `#3697-14 FAILED (${label14}): a zero-selection line must tick NOTHING.\n` + + `REQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + // The placeholder gate is what holds the whole #2334 negative space silent + // under R3b, so it is pinned in every spelling an author actually writes. + // Whole-channel silence, same rationale as #3697-4: a phrase filter pins + // wording, not silence. + for (const [label14b, line14b] of [ + ['lowercase tbd', 'tbd'], + ['lowercase none', 'none'], + ['bold None only', '**None**'], + ['backticked TBD', '`TBD`'], + ['TBD with trailing note', 'TBD '], + ]) { + test( + `#3697-14b (placeholder spelling, ${label14b}): the placeholder gate must hold R3b off`, + (t) => { + const tmpDir = build3697RangeFixture(line14b); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.deepStrictEqual( + warnings, [], + `#3697-14b FAILED (${label14b}): ${JSON.stringify(line14b)} is a declared-empty line and ` + + `must produce NO warning at all, got: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + }, + ); + } + + // R3b's `tokens.length > 0` guard. The tokenizer strips `` + // BEFORE splitting, so a comment-only line yields no tokens and cannot + // reach the rule — without this guard the shipped template's own comment + // line would warn, which is #2334's opening move. + test( + '#3697-14c (comment-only line): a line whose only content is a template comment stays silent', + (t) => { + const tmpDir = build3697RangeFixture(''); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.deepStrictEqual( + warnings, [], + `#3697-14c FAILED: a comment-only Requirements line has no content tokens and must not ` + + `warn, got: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + }, + ); + + // ========================================================================== + // #3697-15 — AC-1a: a REQ-ID the selector dropped to a glued delimiter. + // + // `RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with + // `requirements_updated: true` — #3697's own half-success failure mode, + // reached by one wrong delimiter, and silent before this rule. + // + // Round 4 review rated this Major rather than Blocker on the ground that it + // is indistinguishable from a parenthesised citation (`(ADR-7)` shaves to a + // bare ID too). At the RAW token level it is distinguishable: the shave + // class differs, and -15b pins the citation half. + // ========================================================================== + for (const [label15, line15, expectTicked15, expectNamed15] of [ + ['semicolon', 'RANGE-01; RANGE-02', ['RANGE-02'], ['RANGE-01']], + ['colon', 'RANGE-01: RANGE-02', ['RANGE-02'], ['RANGE-01']], + // Every dropped ID is named, not just the first — the chain is what + // catches a clause that reports only one and reads as complete. + ['semicolon chain', 'RANGE-01; RANGE-02; RANGE-03', ['RANGE-03'], ['RANGE-01', 'RANGE-02']], + // The drop can be the LAST id as easily as the first. + ['trailing colon', 'RANGE-01, RANGE-02:', ['RANGE-01'], ['RANGE-02']], + ]) { + test( + `#3697-15 (delimiter-dropped ID, ${label15}): an ID the selector dropped to a glued ` + + '`;`/`:` must warn and be named — selection is UNCHANGED', + (t) => { + const tmpDir = build3697RangeFixture(line15); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-15 FAILED (${label15}): ${JSON.stringify(line15)} silently drops an ID, so it ` + + `must warn. Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + // Naming the dropped ID is the whole point — a warning that says + // "something was dropped" without saying WHAT is not actionable. + const expectedClause15 = + `${expectNamed15.join(', ')} ${expectNamed15.length === 1 ? 'was' : 'were'} NOT selected`; + assert.ok( + warnings.some((w) => w.includes(expectedClause15)), + `#3697-15 FAILED (${label15}): the warning must name EVERY dropped ID — expected the ` + + `clause ${JSON.stringify(expectedClause15)}, got: ${JSON.stringify(warnings)}`, + ); + // The selector is untouched by this PR — the rule warns, it never + // widens what gets marked. + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent), expectTicked15, + `#3697-15 FAILED (${label15}): exactly ${JSON.stringify(expectTicked15)} must be ticked — ` + + `the warning must not change the selection.\nREQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } + + // The rule's boundary, and the reason it is safe to ship. A delimiter glued + // to an ID-shaped token INSIDE a parenthetical is a citation, not a dropped + // requirement, and reporting it is the #2334 over-warning class. Whole + // channel, because there is nothing on these lines to warn about at all. + for (const [label15b, line15b] of [ + ['colon in citation', 'RANGE-01, RANGE-02 (see ADR-7: section 3)'], + ['semicolon in citation', 'RANGE-01, RANGE-02 (see ADR-7; also ADR-9)'], + ['nested colon note', 'RANGE-01, RANGE-02 (blocked: ADR-7: sec 3)'], + ]) { + test( + `#3697-15b (citation boundary, ${label15b}): a delimiter inside a parenthetical is a ` + + 'citation, not a dropped requirement', + (t) => { + const tmpDir = build3697RangeFixture(line15b); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.deepStrictEqual( + warnings, [], + `#3697-15b FAILED (${label15b}): ${JSON.stringify(line15b)} is a correctly-parsed comma ` + + `list carrying a citation and must produce NO warning, got: ${JSON.stringify(warnings)}\n` + + `Full output: ${output}`, + ); + }, + ); + } + + // ========================================================================== + // #3697-16 — the warning's machine kind reaches the JSON output (round 4 + // review Major 3). + // + // `warnings[]` stays a string[] — it is a documented output field rendered by + // execute-phase.md, so re-typing its elements would break a shipped + // contract. The kind is emitted as its own additive field, and THAT is what + // a consumer and a test key on. Rewording a message must not silently + // un-assert anything. + // ========================================================================== + for (const [label16, line16, expectCode16] of [ + ['assertive / misparse', 'RANGE-01..RANGE-05', 'req-line-misparse'], + ['ambiguous / range reading', 'RANGE-01 … RANGE-05', 'req-line-range-reading'], + ['zero selection', 'Deferred', 'req-line-misparse'], + ['delimiter drop', 'RANGE-01; RANGE-02', 'req-line-misparse'], + ]) { + test( + `#3697-16 (warning kind, ${label16}): the JSON result carries a stable machine code`, + (t) => { + const tmpDir = build3697RangeFixture(line16); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + assert.ok( + parsed.requirements_line_warning, + `#3697-16 FAILED (${label16}): a warning fired, so the result must carry ` + + `requirements_line_warning.\nFull output: ${output}`, + ); + assert.strictEqual( + parsed.requirements_line_warning.code, expectCode16, + `#3697-16 FAILED (${label16}): wrong kind for ${JSON.stringify(line16)}.\n` + + `Full output: ${output}`, + ); + // The prose channel is UNCHANGED — this field is additive, and a + // consumer reading warnings[] must see exactly what it saw before. + assert.ok( + Array.isArray(parsed.warnings) && parsed.warnings.every((w) => typeof w === 'string'), + `#3697-16 FAILED (${label16}): warnings[] must remain a string[]; ` + + `got ${JSON.stringify(parsed.warnings)}`, + ); + }, + ); + } + + // The absent half. A field that is present on every run carries no + // information, and a consumer keying on its presence would be wrong forever. + test( + '#3697-16b (clean line): no warning kind is emitted when the line parses', + (t) => { + const tmpDir = build3697RangeFixture('RANGE-01, RANGE-02, RANGE-03, RANGE-04, RANGE-05'); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + assert.strictEqual( + parsed.requirements_line_warning, undefined, + `#3697-16b FAILED: a clean Requirements line must emit no warning kind at all, got ` + + `${JSON.stringify(parsed.requirements_line_warning)}\nFull output: ${output}`, + ); + }, + ); + + // ========================================================================== + // #3697-17 — PARITY PIN against the SECOND parser of the same ROADMAP value. + // + // CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when + // sharing constants/arrays/parsers between parallel surfaces, add a parity + // assertion test that fails if they diverge." Round 4 review Major 2. + // + // `normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME + // `**Requirements:**` value — its own docblock says callers "may pass the + // roadmap value through verbatim". The two parsers do NOT agree today, and + // this PR does not make them agree: `phase complete` writes a ledger, gap + // analysis reports coverage, and unifying them would change what + // `phase complete` MARKS, which is the one invariant this PR holds fixed. + // + // So this is the pin, not the fix: every axis of disagreement is asserted + // explicitly, in BOTH directions, so drift on either side fails here rather + // than silently widening. The consequence is already user-visible — see + // -17b — and the pin is what makes it a known quantity instead of a slow + // leak. + // ========================================================================== + const { normalizePhaseReqIds } = require('../gsd-core/bin/lib/gap-checker.cjs'); + const phaseSelects = (line) => analyzeRequirementsLine(line).citedReqIds; + + for (const [axis, line17, expectPhase, expectGap, why] of [ + [ + 'ranges', + 'RANGE-01..RANGE-05', + [], + ['RANGE-01', 'RANGE-02', 'RANGE-03', 'RANGE-04', 'RANGE-05'], + 'DELIBERATE: #3697 explicitly declines range expansion ("I am not asking for range syntax ' + + 'to be supported"); gap analysis expanded ranges under #1269. This PR warns on the shape ' + + 'rather than adopting it.', + ], + [ + 'placeholder vocabulary', + 'None (per ADR-7)', + [], + ['ADR-7'], + 'DIVERGENT: phase complete reads the LEAD token, so a declared-empty line citing its ' + + 'rationale is empty. gap-checker strips parentheses first, so the whole value is no longer ' + + 'the bare placeholder and the citation survives its ID-shape filter as a requirement.', + ], + [ + 'parentheses', + '(REQ-02)', + [], + ['REQ-02'], + 'DIVERGENT: the selector strips square brackets only; gap-checker strips quotes, brackets ' + + 'AND parentheses.', + ], + [ + 'ID shape', + 'REQ-01a', + [], + ['REQ-01a'], + 'DIVERGENT: the selector requires the token to END in digits; PHASE_REQ_ID_SHAPE_RE is ' + + 'wider and admits a trailing suffix.', + ], + [ + 'the canonical form', + 'REQ-01, REQ-02', + ['REQ-01', 'REQ-02'], + ['REQ-01', 'REQ-02'], + 'AGREEMENT: on the shipped template form the two parsers agree exactly, which is what ' + + 'makes the divergences above a boundary rather than chaos.', + ], + ]) { + test(`#3697-17 (parser parity, ${axis}): the divergence is pinned in both directions`, () => { + assert.deepStrictEqual( + phaseSelects(line17), expectPhase, + `#3697-17 FAILED (${axis}): phase complete's selection drifted for ${JSON.stringify(line17)}.\n${why}`, + ); + assert.deepStrictEqual( + normalizePhaseReqIds(line17), expectGap, + `#3697-17 FAILED (${axis}): gap-checker's normalization drifted for ${JSON.stringify(line17)}.\n${why}`, + ); + }); + } + + // The consequence, made concrete. This is what the divergence COSTS a user, + // and it is the reason the pin above is worth its cost: on one line, two + // commands report contradictory scopes, and this PR is what makes the + // contradiction visible by finally giving `phase complete` a voice. + test( + '#3697-17b (user-visible contradiction): a range line reports five requirements to gap ' + + 'analysis and zero to phase complete', + () => { + const line = 'RANGE-01..RANGE-05'; + const a = analyzeRequirementsLine(line); + assert.deepStrictEqual(a.citedReqIds, [], '#3697-17b: phase complete selects nothing'); + assert.strictEqual(a.warn, true, '#3697-17b: and now says so, rather than failing silently'); + assert.strictEqual( + normalizePhaseReqIds(line).length, 5, + '#3697-17b: while gap analysis reports five requirements in scope for the same line', + ); + }, + ); + + // ========================================================================== + // #3697-18 — the skipped-text rider must not tell the author to check + // whether a DATE is a requirement (round 4 review Minor 2). + // + // This round first answered it with a FILTER, and the pre-push review's + // continuation broke that in both directions: `API-2-01` is a legal + // requirement id (gap-checker's parseRequirements accepts it) so the filter + // hid a real drop, while `FY-26-08` and `FY-2026-08-15` leaked through. No + // regex separates a date from a sub-numbered id — they are the same shape, + // which is exactly why the strict-dash range rule refuses to act on it. + // + // So the rider DISCLOSES instead of adjudicating: it names the token and + // says why it may not be a requirement. Same move the two warning voices + // already make about an ambiguous separator. + // ========================================================================== + for (const [label18, line18, expectNamed18] of [ + ['date beside a dotted range', 'RANGE-01 .. RANGE-05 (target FY-2026-08)', 'FY-2026-08'], + ['multi-segment date', 'RANGE-01 .. RANGE-05 (target FY-2026-08-15)', 'FY-2026-08-15'], + ['short-year date', 'RANGE-01 … RANGE-05 due FY-26-08', 'FY-26-08'], + ['sub-numbered requirement', 'RANGE-01 .. RANGE-05, API-2-01', 'API-2-01'], + ['four-digit sub-number', 'RANGE-01 .. RANGE-05, API-2026-08', 'API-2026-08'], + ]) { + test( + `#3697-18 (ambiguous numeric shape, ${label18}): named, and qualified rather than adjudicated`, + () => { + const a = analyzeRequirementsLine(line18); + const w = reqLineText(formatRequirementsLineWarning('1', line18, a)); + assert.ok(w, `#3697-18 (${label18}): the range rule fired, so the line warns`); + // NAMED — suppressing it hid a genuinely dropped requirement. + assert.ok( + w.includes(expectNamed18), + `#3697-18 FAILED (${label18}): ${expectNamed18} must be named — a filter here hid a real ` + + `dropped requirement: ${w}`, + ); + // QUALIFIED — the author is told why it may not be a requirement, + // instead of the warning deciding for them in either direction. + assert.match( + w, + /may equally be a date or a sub-numbered id/i, + `#3697-18 FAILED (${label18}): the shape is undecidable and the rider must say so: ${w}`, + ); + }, + ); + } + + // The other half: an UNAMBIGUOUS dropped id is named with no hedge attached. + test('#3697-18b (unambiguous skip): a plain dropped REQ-ID is named without the date caveat', () => { + const line = 'REQ-01, (REQ-02), REQ-03 — REQ-05'; + const a = analyzeRequirementsLine(line); + const w = reqLineText(formatRequirementsLineWarning('1', line, a)); + assert.match( + w, /ID-shaped text on the line that was NOT selected: REQ-02/i, + `#3697-18b FAILED: REQ-02 is a real skipped ID and must be named: ${w}`, + ); + assert.doesNotMatch( + w, /may equally be a date/i, + `#3697-18b FAILED: REQ-02 is not the ambiguous shape and must carry no caveat: ${w}`, + ); + }); + + // ========================================================================== + // #3697-19 — findings from this round's OWN pre-push adversarial review. + // + // Four claims this round made were driven and refuted before the push. Each + // is pinned here, because every one of them was a shape the author had not + // probed — the rules were correct across the probe set and wrong just + // outside it. + // ========================================================================== + + // (1) An INVISIBLE line is an empty line. A lone U+200B carried a token to + // the parser while reading as empty to the author, so R3b warned and nothing + // on screen explained why. + for (const [label19, line19] of [ + ['zero-width space', '\u200B'], + ['two zero-width spaces', '\u200B\u200B'], + ['word joiner', '\u2060'], + ['BOM', '\uFEFF'], + // All four below were driven by the pre-push review's continuation against + // a strip set that covered only the first four. + ['soft hyphen', '\u00AD'], + ['left-to-right mark', '\u200E'], + ['left-to-right isolate', '\u2066'], + ['variation selector 16', '\uFE0F'], + ]) { + test(`#3697-19 (invisible content, ${label19}): a line the author cannot see is not content`, () => { + const a = analyzeRequirementsLine(line19); + assert.strictEqual( + a.warn, false, + `#3697-19 FAILED (${label19}): an invisible-only line must not warn — no author could act on it`, + ); + }); + } + + // A zero-width character INSIDE an otherwise valid line must be stripped, not + // treated as a delimiter: splitting on it would fabricate two fragments from + // one ID and invent a drop that never happened. + // An invisible INSIDE a token breaks it for the SELECTOR, so the id is + // genuinely not marked — #3697's own defect, in its most undetectable form. + // + // The first cut of this fix stripped invisibles from the detector wholesale + // and made exactly this case go SILENT, and the test written for it asserted + // the tokens and the empty R4 result while never asserting `warn` — it + // DOCUMENTED the bug instead of catching it. That omission was the pre-push + // review's `MISSED:` finding. The assertion below is the one that was absent. + for (const [label19b, line19b] of [ + ['zero-width space', 'REQ-01\u200B, REQ-02'], + ['soft hyphen', 'REQ-01\u00AD, REQ-02'], + ['zero-width joiner', 'REQ-01\u200D, REQ-02'], + ]) { + test(`#3697-19b (embedded invisible, ${label19b}): an unmarkable id must never be silent`, () => { + const a = analyzeRequirementsLine(line19b); + assert.deepStrictEqual( + a.citedReqIds, ['REQ-02'], + `#3697-19b (${label19b}): the selector really does drop REQ-01 — selection is unchanged by this round`, + ); + assert.strictEqual( + a.warn, true, + `#3697-19b FAILED (${label19b}): an id dropped to an INVISIBLE character is the least ` + + `detectable form of #3697's own defect and must not stay silent`, + ); + assert.deepStrictEqual( + a.delimiterDroppedIds, ['REQ-01'], + `#3697-19b (${label19b}): and it must be NAMED — the author cannot see the character`, + ); + }); + } + + // (2) R4's false positive. A BARE citation carrying a colon is the same shave + // class as a real delimiter drop, and the parenthetical test does not reach + // it. Same-prefix agreement with a SELECTED id is what separates them. + for (const [label19c, line19c] of [ + ['bare citation', 'REQ-01, see ADR-7: section 3'], + ['bare citation, no verb', 'REQ-01, ADR-7: section 3'], + ['parenthesised citation', 'RANGE-01, RANGE-02 (see ADR-7: sec 3)'], + ]) { + test(`#3697-19c (citation boundary, ${label19c}): a foreign-prefix citation is not a dropped requirement`, () => { + const a = analyzeRequirementsLine(line19c); + assert.deepStrictEqual( + a.delimiterDroppedIds, [], + `#3697-19c FAILED (${label19c}): ${JSON.stringify(line19c)} cites a foreign prefix; ` + + `reporting it is the #2334 over-warning class`, + ); + }); + } + + // (3) R4's false negative, on the DOCUMENTED form. The selector strips square + // brackets; R4's raw scanner did not, so the bracket spelling the template + // recommends silently dropped an ID with no warning at all. + for (const [label19d, line19d, expectDropped19] of [ + ['bracketed semicolon', '[REQ-01; REQ-02]', ['REQ-01']], + ['bracketed colon', '[REQ-01, REQ-02: login]', ['REQ-02']], + // A trailing-only regex missed every one of these; they are one class + // (decoration on a token the selector then could not take), so they take + // one rule rather than four patches. + ['leading semicolon', 'REQ-01 ;REQ-02', ['REQ-02']], + ['leading colon', 'REQ-01 :REQ-02', ['REQ-02']], + ['bold-wrapped', '**REQ-01;** REQ-02', ['REQ-01']], + ['backtick-wrapped', '`REQ-01;` REQ-02', ['REQ-01']], + ]) { + test(`#3697-19d (bracket form, ${label19d}): the documented spelling must not defeat R4`, () => { + const a = analyzeRequirementsLine(line19d); + assert.deepStrictEqual( + a.delimiterDroppedIds, expectDropped19, + `#3697-19d FAILED (${label19d}): square brackets are the documented form and must be ` + + `stripped for R4 exactly as the selector strips them`, + ); + assert.strictEqual(a.warn, true, `#3697-19d FAILED (${label19d}): and the line must warn`); + }); + } + + // (4) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a + // legal requirement id — gap-checker's own parseRequirements accepts it — so + // only a DATE shape (four-digit year segment) may be filtered. + // The residual the same-prefix gate BUYS its safety with, pinned so it is a + // known quantity rather than a surprise. A genuinely dropped id whose prefix + // appears on no selected id stays silent — the same trade the strict-dash + // rule takes: under-report a rare shape rather than over-report a common one. + // EMPHASIS ALONE IS NOT EVIDENCE, and this pins the boundary in both + // directions. An earlier cut of this round fired on any wrapper, which made + // `REQ-01, see **REQ-7** for context` warn about a citation — the #2334 + // class again. Nothing separates that from `**REQ-01**, REQ-02` meaning to + // list one, so R4 requires the positive signal: a glued `;`/`:` (a list + // separator was intended) or an invisible (the token is corrupted; no author + // types one on purpose). Markdown styling is authorial and is left to the + // skipped-text rider, which names the id without asserting a drop. + for (const [label19h, line19h] of [ + ['emphasised citation', 'REQ-01, see **REQ-7** for context'], + ['backticked citation', 'REQ-01, see `REQ-7` for context'], + ['emphasis with no delimiter', '**REQ-01**, REQ-02'], + // The delimiter's POSITION is the rule. Outside the styling it is sentence + // punctuation, and an earlier cut of this round fired on exactly this — + // `see **REQ-7**; next topic` was reported as a dropped requirement. + ['punctuated emphasised citation', 'REQ-01, see **REQ-7**; next topic'], + ['punctuated backticked citation', 'REQ-01, see `REQ-7`; next'], + ['delimiter outside the wrapper', '**REQ-01**; REQ-02'], + ]) { + test(`#3697-19h (styling is not evidence, ${label19h}): a wrapper alone must not fire R4`, () => { + const a = analyzeRequirementsLine(line19h); + assert.deepStrictEqual( + a.delimiterDroppedIds, [], + `#3697-19h FAILED (${label19h}): markdown styling carries no list-separator intent, so ` + + `claiming a drop here is the #2334 over-warning class`, + ); + }); + } + + // An invisible attached to the range OPERATOR, not to an id. The regression + // tests covered invisibles inside ids and not this, and that gap is exactly + // what let a fix for one direction break the other — the pre-push review's + // own MISSED finding. + for (const [label19i, line19i] of [ + ['invisible around a dotted operator', 'REQ-01 \u200B..\u200B REQ-05'], + ['invisible before a glued dash', 'REQ-01 \u200B-REQ-05'], + ]) { + test(`#3697-19i (invisible on the operator, ${label19i}): the operator is still an operator`, () => { + assert.strictEqual( + analyzeRequirementsLine(line19i).warn, true, + `#3697-19i FAILED (${label19i}): an invisible beside the range operator must not hide it`, + ); + }); + } + + // An UNBALANCED parenthesis is a typo, not a citation, and must not confer + // citation immunity on the rest of the line. A running-depth counter let one + // stay open to end-of-line and swallow every real drop after it. + for (const [label19j, line19j, expectDropped19j] of [ + ['unclosed open paren', 'REQ-01, (note REQ-02; REQ-03', ['REQ-02']], + ['stray close paren', 'REQ-01, REQ-02; REQ-03)', ['REQ-02']], + // A matched span sharing a whitespace token with an id OUTSIDE it. The + // first cut promoted the whole token to immune because it CONTAINED a + // matched character, so the drop next to the citation went silent — the + // pre-push review's last MISSED finding, which named exactly the control + // these two rows add. + ['citation glued after the drop', 'REQ-01, REQ-02;(note) REQ-03', ['REQ-02']], + ['citation glued before the drop', 'REQ-01, (note)REQ-02; REQ-03', ['REQ-02']], + ]) { + test(`#3697-19j (unbalanced parens, ${label19j}): only a MATCHED span is a citation`, () => { + const a = analyzeRequirementsLine(line19j); + assert.deepStrictEqual( + a.delimiterDroppedIds, expectDropped19j, + `#3697-19j FAILED (${label19j}): an unmatched paren must not swallow the rest of the line`, + ); + }); + } + + // And the matched forms still are citations. + for (const [label19k, line19k] of [ + ['matched span', 'REQ-01, (see ADR-7: sec 3)'], + ['nested matched spans', 'RANGE-01, RANGE-02 (see (ADR-7), then ADR-8)'], + ]) { + test(`#3697-19k (matched parens, ${label19k}): a real citation is still immune`, () => { + assert.deepStrictEqual( + analyzeRequirementsLine(line19k).delimiterDroppedIds, [], + `#3697-19k FAILED (${label19k}): a matched parenthetical is a citation`, + ); + }); + } + + test('#3697-19f (declared blind spot): a dropped id with an unshared prefix stays silent', () => { + const a = analyzeRequirementsLine('REQ-01, FOO-02: x'); + assert.deepStrictEqual(a.citedReqIds, ['REQ-01'], '#3697-19f: FOO-02 really is dropped'); + assert.deepStrictEqual( + a.delimiterDroppedIds, [], + '#3697-19f: and R4 deliberately does not claim it — textually identical to a foreign citation', + ); + }); + + // Round 5 body claim-audit, MISSED item 1. A DEMONSTRATED R4 drop and an + // over-cap token on the same line: the over-cap voice's whole claim is that + // nothing could be checked, which is false the moment R4 has named an id. + // Before the fix `REQ-01, REQ-02: <2049 chars>` reported req-line-unverified + // and never mentioned REQ-02 — the actionable finding masked by the token + // beside it. Same exclusion, same reason, as rangeReadingOnly's. + test('#3697-19l (over-cap beside a drop): a demonstrated drop outranks the unverified voice', () => { + // 2049 = the 2048 cap + 1, written as a literal and asserted, in the same + // style as the #3697-B1/-B2 boundary fixtures below. + const overCap = 'x'.repeat(2049); + assert.strictEqual(overCap.length, 2049, '#3697-19l fixture must be one past the cap'); + const line = `REQ-01, REQ-02: ${overCap}`; + const a = analyzeRequirementsLine(line); + assert.deepStrictEqual(a.delimiterDroppedIds, ['REQ-02'], '#3697-19l: R4 does name the drop'); + assert.ok(a.oversizedTokens.length > 0, '#3697-19l: and the over-cap token is present'); + const w = formatRequirementsLineWarning('1', line, a); + assert.strictEqual( + reqLineCode(w), REQ_LINE_WARNING_CODE.misparse, + '#3697-19l: a line with a demonstrated drop is a misparse, not an unverified line', + ); + assert.match( + reqLineText(w), /is glued to the ID/, + '#3697-19l: and the drop diagnosis must survive — it is the only actionable part', + ); + assert.match( + reqLineText(w), /exceed the 2048-character scan limit/, + '#3697-19l: while the over-cap disclosure is still carried, not traded away', + ); + }); + + test('#3697-19p (over-cap beside a CLEAN range): the cap outranks the ambiguous voice', () => { + // The twin of #3697-19l on the other side of the boundary. There, a + // DEMONSTRATED drop outranks the unverified voice; here nothing was + // demonstrated, so the cap does — and what must not happen is the line + // reading as clean. `req-line-range-reading` is documented (in its own code + // comment and in CONTEXT.md's PHASE.REQ-LINE.SEAM.kinds predicate) to mean + // nothing was dropped, and this line carries a token no rule ever examined. + // Round 7 review, Minor 1. + const overCap = 'x'.repeat(2049); + assert.strictEqual(overCap.length, 2049, '#3697-19p fixture must be one past the cap'); + const line = `RANGE-01 - RANGE-05, ${overCap}`; + const a = analyzeRequirementsLine(line); + // The range half is genuinely clean: R2 fired, both endpoints were taken. + assert.deepStrictEqual( + a.citedReqIds, ['RANGE-01', 'RANGE-05'], + '#3697-19p: both endpoints really are selected — this is the CLEAN range shape', + ); + assert.ok(a.hasSpacedRange, '#3697-19p: and R2 really did fire'); + assert.strictEqual(a.delimiterDroppedIds.length, 0, '#3697-19p: nothing was demonstrably dropped'); + assert.ok(a.oversizedTokens.length > 0, '#3697-19p: while an over-cap token is present'); + // The predicate itself, pinned: the ambiguous voice must stand down. + assert.strictEqual( + a.rangeReadingOnly, false, + '#3697-19p: the ambiguous voice claims nothing was dropped — it may not speak over an unexamined token', + ); + const w = formatRequirementsLineWarning('1', line, a); + assert.strictEqual( + reqLineCode(w), REQ_LINE_WARNING_CODE.unverified, + '#3697-19p: an unexamined token makes the line UNVERIFIED, not clean', + ); + // And specifically NOT the assertive voice: nothing on this line failed to + // parse, so `misparse` would be the #2334 over-warning class returning + // through the fix for its own false-clean. + assert.notStrictEqual( + reqLineCode(w), REQ_LINE_WARNING_CODE.misparse, + '#3697-19p: nothing demonstrably failed to parse — asserting a misparse here is the #2334 class', + ); + assert.match( + reqLineText(w), /exceed the 2048-character scan limit/, + '#3697-19p: and the cap is named, so the reader knows what was not looked at', + ); + }); + + test('#3697-19q (control for -19p): the same range WITHOUT an over-cap token still reads as a range', () => { + // Negative control. -19p must not be satisfiable by routing every spaced + // range to `unverified`; the ambiguous voice is still correct when there + // is nothing unexamined on the line. + const line = 'RANGE-01 - RANGE-05'; + const a = analyzeRequirementsLine(line); + assert.strictEqual(a.oversizedTokens.length, 0, '#3697-19q: nothing over the cap here'); + assert.strictEqual(a.rangeReadingOnly, true, '#3697-19q: so the ambiguous voice is the right one'); + assert.strictEqual( + reqLineCode(formatRequirementsLineWarning('1', line, a)), REQ_LINE_WARNING_CODE.rangeReading, + '#3697-19q: unchanged by the -19p fix — a clean range is still a range reading', + ); + }); + + // DECLARED BLIND SPOTS, pinned so the docs and the code cannot drift apart + // again — that drift IS the round 4 blocker. R4's trigger is a glued `;`/`:` + // or an embedded invisible. Styling is TOLERATED around an id, never a + // trigger on its own, so every line below silently drops an id. Each is + // documented as silent in the CLI tools reference; if one of these ever + // starts warning, that document is wrong and this test says so first. + for (const [label19m, line19m, expectSel19m] of [ + ['bold', 'REQ-01, **REQ-02**', ['REQ-01']], + ['quotes', 'REQ-01, "REQ-02"', ['REQ-01']], + ['backticks', 'REQ-01, `REQ-02`', ['REQ-01']], + ['underscore', 'REQ-01, _REQ-02_', ['REQ-01']], + // The `**` sits BETWEEN the id and the `;`, so nothing is touching the id. + ['styling between id and delimiter', 'REQ-01, **REQ-02**;', ['REQ-01']], + ]) { + test(`#3697-19m (declared blind spot, styling-only ${label19m}): silent, and documented as silent`, () => { + const a = analyzeRequirementsLine(line19m); + assert.deepStrictEqual(a.citedReqIds, expectSel19m, `#3697-19m (${label19m}): the id really is dropped`); + assert.deepStrictEqual(a.delimiterDroppedIds, [], `#3697-19m (${label19m}): R4 does not claim it`); + assert.strictEqual(a.warn, false, `#3697-19m (${label19m}): and the line is silent`); + }); + } + + // The other half of the same boundary: only `;` and `:` are in the set. + // Round 4's separator census concluded "exactly those two" because it swept + // the ONE-SIDED form for `;`/`:` and only the bare and symmetric forms for + // every other separator — different members tested in different shapes, so + // the answer was forced. A fully crossed re-sweep (21 separators x 4 + // spellings = 84) found 34 silent under-selections, every one of them a + // separator glued to exactly ONE of the two ids. Pinned here as the + // documented COST, never asserted as coverage. + for (const sep19n of ['/', '|', '&', '+', '.', '>', '\\', ';', ',', '؛']) { + for (const [dir19n, line19n, keep19n] of [ + ['trailing', `REQ-01${sep19n} REQ-02`, 'REQ-02'], + ['leading', `REQ-01 ${sep19n}REQ-02`, 'REQ-01'], + ]) { + test(`#3697-19n (declared blind spot, ${dir19n} "${sep19n}"): silent, and documented as silent`, () => { + const a = analyzeRequirementsLine(line19n); + assert.deepStrictEqual( + a.citedReqIds, [keep19n], + `#3697-19n (${dir19n} "${sep19n}"): exactly one id survives the selector`, + ); + assert.deepStrictEqual( + a.delimiterDroppedIds, [], + `#3697-19n (${dir19n} "${sep19n}"): R4 does not reach it`, + ); + assert.strictEqual(a.warn, false, `#3697-19n (${dir19n} "${sep19n}"): and the line is silent`); + }); + } + } + + // The grammar tolerances the documentation was corrected to state. These are + // pre-existing selector behaviour, not new; the round 4 pre-push review + // caught the DOCS asserting a stricter rule than the code enforces. + for (const [label19g, line19g, expectSel19] of [ + ['whitespace-separated', 'REQ-01 REQ-02', ['REQ-01', 'REQ-02']], + ['lowercase ids', 'req-01, req-02', ['req-01', 'req-02']], + ]) { + test(`#3697-19g (documented tolerance, ${label19g}): selected and silent, as the docs now say`, () => { + const a = analyzeRequirementsLine(line19g); + assert.deepStrictEqual(a.citedReqIds, expectSel19, `#3697-19g (${label19g}): both ids are selected`); + assert.strictEqual(a.warn, false, `#3697-19g (${label19g}): and nothing warns about it`); + }); + } + + // A HALF-SPACED range splits at the tokenizer before R1's own `\s*` can see + // it: `RANGE-01 -RANGE-05` tokenizes as `RANGE-01`, `-RANGE-05` and selects + // only the well-formed side. The glued-fragment rule must warn, and the + // selection stays exactly what the tokenizer produced. + for (const [label6, line6, expectTicked] of [ + ['leading-glue', 'RANGE-01 -RANGE-05', ['RANGE-01']], + ['trailing-glue', 'RANGE-01- RANGE-05', ['RANGE-05']], + // A word operator can glue only TRAILING (an ID must end in digits, so + // `RANGE-01through` cannot be an ID — but `TORANGE-05` can, which is why + // the leading arm is symbol-only). + ['worded-glue', 'RANGE-01through RANGE-05', ['RANGE-05']], + // The trailing shave must not eat a glued `..` as sentence punctuation. + ['double-dot-glue', 'RANGE-01.. RANGE-05', ['RANGE-05']], + ]) { + test( + `#3697-6 (${label6} half-spaced range): a range glued to one endpoint must warn — ` + + 'selection is unchanged', + (t) => { + const tmpDir = build3697RangeFixture(line6); + t.after(() => cleanup(tmpDir)); + const { output } = runVerifiedPhaseComplete(['phase', 'complete', '1'], tmpDir); + const parsed = JSON.parse(output); + const warnings = parsed.warnings || []; + assert.ok( + warnings.some((w) => REQ_LINE_MISPARSE_RE.test(w)), + `#3697-6 FAILED (${label6}): a half-spaced range under-selects, so it must warn. ` + + `Got warnings: ${JSON.stringify(warnings)}\nFull output: ${output}`, + ); + const reqContent = fs.readFileSync(path.join(tmpDir, '.planning', 'REQUIREMENTS.md'), 'utf-8'); + assert.deepStrictEqual( + tickedReqIds(reqContent), expectTicked, + `#3697-6 FAILED (${label6}): exactly ${JSON.stringify(expectTicked)} must be ticked ` + + `(no expansion).\nREQUIREMENTS.md:\n${reqContent}`, + ); + }, + ); + } +}); + + + +// ───────────────────────────────────────────────────────────────────────────── +// #3697 round 3 — unit + property coverage of the EXTRACTED detector. +// +// The end-to-end block above drives the CLI, one subprocess per case. That is +// the right shape for wiring, and the wrong shape for the two rules round 3's +// review blocked on: `RULESET.TESTS.property-based-testing` wants a fast-check +// property over a parsing module (100+ runs), and +// `RULESET.TESTS.boundary-coverage.fixtures` wants limit-1 / limit / limit+1 on +// the 2048-char token cap. Both are expressible only against a callable +// surface, which is why `analyzeRequirementsLine` was extracted from +// `cmdPhaseComplete` in the same round. +// ───────────────────────────────────────────────────────────────────────────── + +const fc = require('fast-check'); +const { + analyzeRequirementsLine, + formatRequirementsLineWarning, + REQ_LINE_WARNING_CODE, +} = require('../gsd-core/bin/lib/phase.cjs'); + +// Round 4 review Major 3: the formatter returns `{ code, message }`, so channel +// IDENTITY is asserted on the stable code and never on the English sentence. +// The two channel regexes below survive for the assertions where the +// USER-VISIBLE wording is itself the thing under test. +const reqLineText = (w) => (w === null ? null : w.message); +const reqLineCode = (w) => (w === null ? null : w.code); + +describe('#3697 round 3: Requirements-line detector — properties (RULESET.TESTS.property-based-testing)', () => { + // fc arbitraries for a well-formed REQ-ID. The shape is the selector's own: + // `[A-Z][A-Z0-9]*-\d+`. Word range operators are excluded from the prefix + // alphabet nowhere — deliberately: `TORANGE-05` IS a valid ID, and property + // (a) asserting silence over it is what pins the #3697-4b behaviour + // generatively rather than at one hand-picked example. + const reqPrefix = fc + .tuple( + fc.constantFrom(...'ABCDEFGHIJKLMNOPQRSTUVWXYZ'.split('')), + // fast-check v4 removed `fc.stringOf`; build the tail from an array so + // the alphabet stays pinned to the selector's own `[A-Z0-9]` class. + fc + .array(fc.constantFrom(...'ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'.split('')), { + minLength: 0, + maxLength: 6, + }) + .map((cs) => cs.join('')), + ) + .map(([head, tail]) => head + tail); + const reqNum = fc.integer({ min: 0, max: 9999 }); + const reqId = fc.tuple(reqPrefix, reqNum).map(([p, n]) => `${p}-${n}`); + + test( + '#3697-P1 (soundness of silence): a canonical comma list of well-formed REQ-IDs NEVER warns, ' + + 'and selects exactly the IDs it lists', + () => { + fc.assert( + fc.property(fc.array(reqId, { minLength: 1, maxLength: 8 }), (ids) => { + const line = ids.join(', '); + const a = analyzeRequirementsLine(line); + // Domain invariant (boundary containment): the selected set IS the + // written set — no over-selection (#2334) and no under-selection + // (#3697) on the canonical form, whatever IDs it carries. + assert.deepStrictEqual(a.citedReqIds, ids); + assert.strictEqual( + a.warn, + false, + `#3697-P1: canonical list ${JSON.stringify(line)} must not warn; got ${JSON.stringify( + formatRequirementsLineWarning('1', line, a), + )}`, + ); + }), + { numRuns: 300 }, + ); + }, + ); + + test( + '#3697-P2 (completeness): a same-prefix pair separated by a spaced range operator, with an ' + + 'interior between them, ALWAYS warns', + // NOT a bug-finder: every operator below was already in the pre-round + // operator set, so this property holds against the pre-round code too. An + // earlier round-3 commit claimed otherwise; the review refuted it. It is a + // regression guard over the spaced-range rule, which is what it is worth. + () => { + const ops = ['..', '...', '…', '—', '–', '-', 'to', 'thru', 'through']; + fc.assert( + fc.property( + reqPrefix, + fc.integer({ min: 0, max: 400 }), + fc.integer({ min: 2, max: 400 }), + fc.constantFrom(...ops), + (prefix, lo, delta, op) => { + const line = `${prefix}-${lo} ${op} ${prefix}-${lo + delta}`; + const a = analyzeRequirementsLine(line); + assert.strictEqual( + a.warn, + true, + `#3697-P2: ${JSON.stringify(line)} implies a dropped interior and must warn`, + ); + }, + ), + { numRuns: 300 }, + ); + }, + ); + + test( + '#3697-P3 (the #2334 invariant): an ADJACENT same-prefix pair around a separator can drop ' + + 'nothing, so it never warns however it is annotated', + () => { + const ops = ['-', '—', '–', '..']; + fc.assert( + fc.property( + reqPrefix, + fc.integer({ min: 0, max: 4000 }), + fc.integer({ min: 0, max: 1 }), + fc.constantFrom(...ops), + fc.constantFrom('deferred', 'blocked', 'per ADR-7', 'see notes', ''), + (prefix, lo, delta, op, tail) => { + const line = `${prefix}-${lo}, ${prefix}-${lo} ${op} ${prefix}-${lo + delta}${ + tail ? ' ' + tail : '' + }`; + const a = analyzeRequirementsLine(line); + assert.strictEqual( + a.warn, + false, + `#3697-P3: ${JSON.stringify(line)} has no interior to drop and must stay silent; got ` + + JSON.stringify(formatRequirementsLineWarning('1', line, a)), + ); + }, + ), + { numRuns: 300 }, + ); + }, + ); + + test( + '#3697-P4 (totality + idempotency): the detector is total over arbitrary input and returns ' + + 'the same analysis every time', + () => { + fc.assert( + fc.property(fc.string({ maxLength: 300 }), (s) => { + const a = analyzeRequirementsLine(s); + const b = analyzeRequirementsLine(s); + assert.strictEqual(typeof a.warn, 'boolean'); + assert.ok(Array.isArray(a.citedReqIds) && Array.isArray(a.tokens)); + assert.deepStrictEqual(a, b, '#3697-P4: analysis must be deterministic'); + const wr = formatRequirementsLineWarning('1', s, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + // Round 4 review Major 3: a message without a kind, or a kind that is + // not in the declared vocabulary, is a channel no consumer can route + // on. Held over ARBITRARY input, so a new channel added later cannot + // ship without one. + assert.strictEqual( + wc === null, + w === null, + `#3697-P4: kind and message must appear together; got ${JSON.stringify(wr)}`, + ); + if (wc !== null) { + assert.ok( + Object.values(REQ_LINE_WARNING_CODE).includes(wc), + `#3697-P4: ${JSON.stringify(wc)} is not a declared warning kind`, + ); + } + // The formatter and the analysis must agree on whether there is + // anything to say — a warn with no text, or text with no warn, is a + // channel that can go silent or noisy on its own. + assert.strictEqual( + w === null, + a.warn === false, + `#3697-P4: warn=${a.warn} but message=${JSON.stringify(w)} for ${JSON.stringify(s)}`, + ); + }), + { numRuns: 500 }, + ); + }, + ); + + test( + '#3697-P5 (containment): every selected ID is REQ-ID-shaped and appears verbatim in the line', + () => { + // The first cut of this property drew from a bare `fc.string()` and was + // VACUOUS — measured over 500 samples it produced max length 10 and ZERO + // inputs containing a REQ-ID, so the loop body never executed a single + // assertion. Found by the round's pre-push review. The generator now + // interleaves real IDs with noise, and the property ASSERTS that it saw + // some: a containment property that never contains anything is a green + // test measuring nothing. + let sawIds = 0; + fc.assert( + fc.property( + fc.array(fc.oneof(reqId, fc.constantFrom('(', ')', '[', ']', ',', '—', '..', 'to', 'TBD', 'None', 'per', 'ADR-7')), { + minLength: 1, + maxLength: 12, + }), + (parts) => { + const line = parts.join(' '); + const cited = analyzeRequirementsLine(line).citedReqIds; + if (cited.length > 0) sawIds += 1; + for (const id of cited) { + assert.match(id, /^[A-Z][A-Z0-9]*-\d+$/i, `#3697-P5: ${JSON.stringify(id)} is not ID-shaped`); + assert.ok(line.includes(id), `#3697-P5: ${JSON.stringify(id)} is not present in the input`); + } + }, + ), + { numRuns: 500 }, + ); + assert.ok(sawIds > 50, `#3697-P5 is VACUOUS: only ${sawIds}/500 generated lines selected any ID`); + }, + ); + + test( + '#3697-P6 (totality over arbitrary text): the detector never throws on free-form input', + () => { + // What the old P5 generator was actually covering. Kept as its own + // property, honestly labelled, rather than left masquerading as + // containment coverage. + fc.assert( + fc.property(fc.string({ maxLength: 300, size: 'max' }), (str) => { + const a = analyzeRequirementsLine(str); + assert.strictEqual(typeof a.warn, 'boolean'); + formatRequirementsLineWarning('1', str, a); + }), + { numRuns: 500 }, + ); + }, + ); +}); + +describe('#3697 round 3: the 2048-char token scan cap (RULESET.TESTS.boundary-coverage)', () => { + // `REQ_TOKEN_SCAN_LIMIT` is a hard cap with NO reserve or safety constant + // beside it, so clause (d) of RULESET.TESTS.boundary-coverage.fixtures — an + // input pushed within reserve-distance of the limit — has no referent here. + // (a)/(b)/(c) are the whole obligation, and they are exercised through BOTH + // predicate families the cap now guards: the unanchored ID-substring scan + // (R3) and the anchored range-token scan (R1, capped in round 3 per Nit 6). + const inertOfLength = (n) => `REQ-01${'X'.repeat(n - 'REQ-01'.length)}`; + const rangeTokenOfLength = (n) => { + const suffix = '..RANGE-05'; + const digits = n - 'RANGE-'.length - suffix.length; + return `RANGE-${'0'.repeat(digits - 1)}1${suffix}`; + }; + + // `expectClassified` is the PREDICATE's verdict, which is what the cap + // governs. `warn` is deliberately NOT the boundary variable: past the cap the + // token is unclassified, and unclassified is reported, never treated as + // clean — asserting `warn === false` at limit+1 is precisely the silent + // regression the round's pre-push review refuted (CLAIM 2). + for (const [label, n, expectClassified] of [ + ['limit-1 (2047)', 2047, true], + ['limit (2048)', 2048, true], + ['limit+1 (2049)', 2049, false], + ]) { + test(`#3697-B1 (${label}): the UNANCHORED ID-substring scan is applied at and below the cap only`, () => { + const tok = inertOfLength(n); + assert.strictEqual(tok.length, n, `#3697-B1 fixture is ${tok.length} chars, expected ${n}`); + const a = analyzeRequirementsLine(tok); + assert.deepStrictEqual(a.citedReqIds, [], '#3697-B1: the padded token is not itself an ID'); + assert.strictEqual( + a.inertIdShaped.length > 0, + expectClassified, + `#3697-B1 (${label}): inertIdShaped membership must be ${expectClassified} at length ${n}`, + ); + assert.strictEqual( + a.oversizedTokens.length > 0, + !expectClassified, + `#3697-B1 (${label}): past the cap the token must be recorded as unclassified`, + ); + // Warns at every length — below the cap because a rule classified it, + // above it because "not classified" is itself reportable. + assert.strictEqual(a.warn, true, `#3697-B1 (${label}): unclassified is not clean`); + }); + + test(`#3697-B2 (${label}): the ANCHORED range-token scan takes the SAME cap (round-3 Nit 6)`, () => { + const tok = rangeTokenOfLength(n); + assert.strictEqual(tok.length, n, `#3697-B2 fixture is ${tok.length} chars, expected ${n}`); + const a = analyzeRequirementsLine(tok); + assert.strictEqual( + a.rangeTokens.length > 0, + expectClassified, + `#3697-B2 (${label}): rangeTokens membership must be ${expectClassified} at length ${n}`, + ); + assert.strictEqual( + a.oversizedTokens.length > 0, + !expectClassified, + `#3697-B2 (${label}): past the cap the token must be recorded as unclassified`, + ); + assert.strictEqual(a.warn, true, `#3697-B2 (${label}): unclassified is not clean`); + }); + } +}); + +describe('#3697 round 3: the two warning channels', () => { + // ── Major 3 ──────────────────────────────────────────────────────────────── + // An annotation separator between two NON-ADJACENT same-prefix IDs is + // textually identical to a range, and no token-level rule separates them. + // Before round 3 the detector resolved that ambiguity by assertion: it told + // the author the line "could not be parsed" and to rewrite it, on a line + // where every ID present HAD been selected and nothing had been dropped. + // Going silent instead is not available — the range reading is equally live, + // and silence is the #3697 defect itself. So the ambiguity is disclosed. + for (const [label, line, expectSelected] of [ + ['em-dash', 'RANGE-01, RANGE-02 — RANGE-05 deferred', ['RANGE-01', 'RANGE-02', 'RANGE-05']], + ['hyphen', 'RANGE-01, RANGE-02 - RANGE-05 deferred', ['RANGE-01', 'RANGE-02', 'RANGE-05']], + ['bare range', 'RANGE-01 … RANGE-05', ['RANGE-01', 'RANGE-05']], + ]) { + test( + `#3697-9 (${label}): a spaced separator warns through the AMBIGUOUS channel and must NOT ` + + 'claim the line failed to parse', + () => { + const a = analyzeRequirementsLine(line); + assert.deepStrictEqual( + a.citedReqIds, + expectSelected, + `#3697-9 (${label}): every ID written on the line must still be selected`, + ); + assert.strictEqual( + a.rangeReadingOnly, + true, + `#3697-9 (${label}): only the spaced-range rule fired and both endpoints were selected — ` + + 'that is why this channel exists', + ); + const wr = formatRequirementsLineWarning('1', line, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.ok(w, `#3697-9 (${label}): the range reading is live, so the line must still warn`); + assert.strictEqual( + wc, + REQ_LINE_WARNING_CODE.rangeReading, + `#3697-9 (${label}): wrong channel: ${wc} / ${w}`, + ); + assert.notStrictEqual( + wc, + REQ_LINE_WARNING_CODE.misparse, + `#3697-9 (${label}): nothing was dropped, so the warning must not assert a parse failure: ${w}`, + ); + // Both readings must be offered — the author is the only one who can + // resolve the ambiguity, and a warning that hides half of it is the + // over-warning wearing better manners. + assert.match(w, /annotation rather than a range/i, `#3697-9 (${label}): ${w}`); + assert.match(w, /needs no change/i, `#3697-9 (${label}): ${w}`); + // The soft voice speaks about the SEPARATOR, never about the whole + // line: it has no basis for the latter (see #3697-9d). + assert.doesNotMatch( + w, + /the line is already correct/i, + `#3697-9 (${label}): the soft voice must not claim the whole line is correct: ${w}`, + ); + }, + ); + } + + // The channel discriminator must be RULE-SCOPED, not line-global. Both of + // these were misrouted by the first cut of the Major 3 fix, and the first is + // the damaging direction: it puts the false "could not be parsed" claim back + // on a correct line, which is the finding itself returning through a side + // door. + test( + '#3697-9b (unrelated parenthetical citation): a citation elsewhere on the line must not flip ' + + 'the channel back to a misparse claim', + () => { + const line = 'RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)'; + const a = analyzeRequirementsLine(line); + // `(ADR-7)` survives the selector's bracket strip and so is not selected — + // but #3697-4 already pins a parenthetical citation as NOT unparsed + // residue, and no rule fires on it. Only the rules that fired may speak. + assert.strictEqual(a.rangeReadingOnly, true, '#3697-9b: R2 alone fired, on selected endpoints'); + const wr = formatRequirementsLineWarning('1', line, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual(wc, REQ_LINE_WARNING_CODE.rangeReading, `#3697-9b: wrong channel: ${wc} / ${w}`); + assert.notStrictEqual(wc, REQ_LINE_WARNING_CODE.misparse, `#3697-9b: ${wc} / ${w}`); + }, + ); + + test( + '#3697-9c (unselected endpoint): a spaced range whose own endpoint was never selected IS a ' + + 'drop, and takes the assertive channel', + () => { + const line = 'RANGE-01 (RANGE-02) — RANGE-05'; + const a = analyzeRequirementsLine(line); + // The detector shaves brackets and the selector does not, so R2 fires on + // a `RANGE-02` that was never selected. That is a genuine under-selection. + assert.deepStrictEqual(a.citedReqIds, ['RANGE-01', 'RANGE-05'], '#3697-9c: RANGE-02 is not selected'); + assert.strictEqual(a.hasSpacedRange, true, '#3697-9c: R2 still fires'); + assert.strictEqual(a.rangeReadingOnly, false, '#3697-9c: an unselected endpoint is a drop'); + const wr = formatRequirementsLineWarning('1', line, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual(wc, REQ_LINE_WARNING_CODE.misparse, `#3697-9c: wrong channel: ${wc} / ${w}`); + }, + ); + + test( + '#3697-9d (dropped ID elsewhere on the line): the soft voice must not claim the LINE is ' + + 'correct, and must name what the selector skipped', + () => { + // The pre-push review's CLAIM 1 counterexample. `(REQ-02)` survives the + // selector's bracket strip and no rule fires on it, so the channel is + // still the soft one — correctly, because `(ADR-7)` is the same shape and + // routing on it puts the false misparse claim back on a citation. What + // was wrong was the soft voice ASSERTING "the line is already correct and + // nothing needs to change" over a line that dropped a requirement. + const line = 'REQ-01, (REQ-02), REQ-03 — REQ-05'; + const a = analyzeRequirementsLine(line); + assert.deepStrictEqual(a.citedReqIds, ['REQ-01', 'REQ-03', 'REQ-05'], '#3697-9d: REQ-02 is dropped'); + assert.strictEqual(a.rangeReadingOnly, true, '#3697-9d: R2 alone fired, on selected endpoints'); + assert.deepStrictEqual(a.unselectedIdShaped, ['REQ-02'], '#3697-9d: the skip is recorded as a fact'); + const wr = formatRequirementsLineWarning('1', line, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual( + wc, REQ_LINE_WARNING_CODE.rangeReading, + `#3697-9d: R2 alone fired on selected endpoints, so this is the range-reading kind: ${wc}`, + ); + assert.doesNotMatch(w, /the line is already correct/i, `#3697-9d: ${w}`); + assert.match(w, /ID-shaped text on the line that was NOT selected: REQ-02/i, `#3697-9d: ${w}`); + assert.match(w, /check whether any of it is a requirement/i, `#3697-9d: ${w}`); + }, + ); + + test( + '#3697-9e (over-cap token): a token past the scan limit is reported as UNCLASSIFIED, never ' + + 'silently dropped', + () => { + // The review's CLAIM 2. Round 3's first cut of the uniform cap made a + // 2049-char range token silent — it warned before the round. The cap + // bounds the WORK; it must not bound the warning. + const suffix = '..RANGE-05'; + const big = `RANGE-${'0'.repeat(2049 - 'RANGE-'.length - suffix.length - 1)}1${suffix}`; + assert.strictEqual(big.length, 2049, `#3697-9e fixture is ${big.length} chars, expected 2049`); + const a = analyzeRequirementsLine(big); + assert.strictEqual(a.rangeTokens.length, 0, '#3697-9e: past the cap, no predicate classifies it'); + assert.deepStrictEqual(a.oversizedTokens, [big], '#3697-9e: but it IS recorded as unclassified'); + assert.strictEqual(a.warn, true, '#3697-9e: and the line still warns'); + const wr = formatRequirementsLineWarning('1', big, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual( + wc, REQ_LINE_WARNING_CODE.unverified, + `#3697-9e: an unclassifiable line is the 'unverified' kind — the same seam CONTEXT.md ` + + `already records for a truncated scan: ${wc}`, + ); + assert.match(w, /could not be checked/i, `#3697-9e: ${w}`); + assert.match(w, /2048-character scan limit/i, `#3697-9e: ${w}`); + assert.match(w, /unverified/i, `#3697-9e: ${w}`); + }, + ); + + test('#3697-9f (cap uniformity): R2 and the glued rule cap their NEIGHBOURS, not just the operator', () => { + // The first review's MISSED finding. Capping the operator alone left + // ` .. ` running REQ_ID_SHAPE_RE and BigInt over + // both neighbours unbounded. + // + // The endpoints must DIFFER by more than 1, or R2 could not fire even + // uncapped and the fixture would prove nothing — the first cut of this test + // used the same ID twice and was exactly that vacuous. + const pad = '0'.repeat(2049 - 'RANGE-'.length - 1); + const lo = `RANGE-${pad}1`; + const hi = `RANGE-${pad}9`; + assert.strictEqual(lo.length, 2049, `#3697-9f fixture is ${lo.length} chars, expected 2049`); + assert.strictEqual(hi.length, 2049, `#3697-9f fixture is ${hi.length} chars, expected 2049`); + const a = analyzeRequirementsLine(`${lo} .. ${hi}`); + assert.strictEqual(a.hasSpacedRange, false, '#3697-9f: an over-cap endpoint must not be classified'); + assert.deepStrictEqual(a.citedReqIds, [lo, hi], '#3697-9f: both endpoints are selected'); + // Selected does NOT mean nothing was suppressed. The second continuation + // review's CLAIM J/K: an earlier cut exempted every selector-accepted token + // from `oversizedTokens`, so this line — which WARNED before the round, + // when R2 was uncapped — went silent. Both endpoints are unexaminable and + // sit either side of a range operator, so the line is reported unverified. + assert.deepStrictEqual(a.oversizedTokens, [lo, hi], '#3697-9f: both endpoints are unexaminable'); + assert.strictEqual(a.warn, true, '#3697-9f: a suppressed classification is not clean'); + }); + + test( + '#3697-9j (over-cap exemption scope): a selected over-cap ID is exempt ONLY when nothing could ' + + 'have paired with it', + () => { + // The other half of CLAIM J/K. The exemption is what keeps #3697-9i + // silent; it must not extend to a token whose neighbour could have formed + // a range with it, because that is exactly where the cap suppressed a + // rule rather than merely declining to classify a lone token. + const big = `R-${'1'.repeat(2047)}`; + assert.strictEqual(big.length, 2049, `#3697-9j fixture is ${big.length} chars, expected 2049`); + // No neighbour that could pair -> exempt, silent. + assert.strictEqual(analyzeRequirementsLine(`REQ-01, ${big}`).warn, false, '#3697-9j: no pairing neighbour'); + // A range operator beside it -> R2 was suppressed, so report. + assert.strictEqual(analyzeRequirementsLine(`${big} .. REQ-05`).warn, true, '#3697-9j: operator neighbour'); + assert.strictEqual(analyzeRequirementsLine(`REQ-01 .. ${big}`).warn, true, '#3697-9j: operator neighbour'); + }, + ); + + test( + '#3697-9g (assertive voice): the skipped-text clause is on BOTH voices, and does not repeat ' + + 'what the rule-specific clause already named', + () => { + // The continuation review found the clause wired into the soft return + // only, while the round claimed both. It also found the wording wrong: + // square brackets ARE stripped by the selector, only parentheses are not. + const line = 'REQ-01, (REQ-02), REQ-03..REQ-05'; + const wr = formatRequirementsLineWarning('1', line, analyzeRequirementsLine(line)); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual(wc, REQ_LINE_WARNING_CODE.misparse, `#3697-9g: assertive voice expected: ${wc} / ${w}`); + assert.match(w, /ID-shaped text on the line that was NOT selected: REQ-02/i, `#3697-9g: ${w}`); + assert.match(w, /parentheses are not stripped, unlike square brackets/i, `#3697-9g: ${w}`); + // `REQ-03..REQ-05` is already named by "Unparsed text"; naming it twice + // is noise, and a warning that repeats itself is one readers skim. + // Count OUTSIDE the echoed line — the warning quotes the whole + // Requirements line first, so the raw echo is one legitimate occurrence. + const afterEcho = w.slice(w.indexOf('`)') + 2); + assert.strictEqual( + (afterEcho.match(/REQ-03\.\.REQ-05/g) || []).length, + 1, + `#3697-9g: the range token must be diagnosed exactly once: ${w}`, + ); + // And the claim must be TRUE: a bracketed ID is selected, so it can never + // appear in the skipped clause. + assert.deepStrictEqual( + analyzeRequirementsLine('[REQ-01, REQ-02]').citedReqIds, + ['REQ-01', 'REQ-02'], + '#3697-9g: square brackets are stripped by the selector', + ); + }, + ); + + test( + '#3697-9h (over-cap token with no hyphen): an unexaminable OPERATOR must not silence the line', + () => { + // The continuation review's CLAIM B. `oversizedTokens` first filtered on + // `includes('-')`, so a 2049-character run of dots between two IDs was + // missed: R2 declined to classify it (capped) and nothing reported it, so + // a line that warned before the round went silent after it. + const line = `REQ-01 ${'.'.repeat(2049)} REQ-05`; + const a = analyzeRequirementsLine(line); + assert.strictEqual(a.hasSpacedRange, false, '#3697-9h: the operator is past the cap'); + assert.strictEqual(a.oversizedTokens.length, 1, '#3697-9h: and is recorded as unexaminable'); + assert.strictEqual(a.warn, true, '#3697-9h: so the line is not silent'); + }, + ); + + test( + '#3697-9i (over-cap token the SELECTOR took): a selected ID is verified, never "unverified"', + () => { + // The continuation review's MISSED finding. The selector is uncapped and + // fully anchored, so a token it accepted was examined end to end. + // Reporting it unverified because the secondary detector declined to + // classify it is a contradiction inside one warning. + const bigId = `R-${'1'.repeat(2047)}`; + assert.strictEqual(bigId.length, 2049, `#3697-9i fixture is ${bigId.length} chars, expected 2049`); + const a = analyzeRequirementsLine(bigId); + assert.deepStrictEqual(a.citedReqIds, [bigId], '#3697-9i: it is a valid canonical REQ-ID'); + assert.deepStrictEqual(a.oversizedTokens, [], '#3697-9i: selected means verified'); + assert.strictEqual(a.warn, false, '#3697-9i: nothing to report'); + }, + ); + + // ── Minor 4 ──────────────────────────────────────────────────────────────── + // A zero-selection non-placeholder line MUST warn — #3697's own acceptance + // criterion says so in as many words. What was wrong is that it reported an + // ADR citation as "Unparsed text", i.e. as requirement content it had failed + // to read. Name what the residue actually is, and name the escape hatch. + for (const [label, line] of [ + ['Deferred (see ADR-7)', 'Deferred (see ADR-7)'], + ['N/A with citation', 'N/A (tracked in ADR-12)'], + ]) { + test( + `#3697-10 (${label}): a zero-selection line still warns, but names ID-shaped TEXT rather ` + + 'than missed requirements, and points at the placeholder escape', + () => { + const a = analyzeRequirementsLine(line); + assert.deepStrictEqual(a.citedReqIds, [], `#3697-10 (${label}): nothing is selected here`); + const wr = formatRequirementsLineWarning('1', line, a); + const w = reqLineText(wr); + const wc = reqLineCode(wr); + assert.strictEqual( + wc, REQ_LINE_WARNING_CODE.misparse, + `#3697-10 (${label}): zero selection is a demonstrated parse failure: ${wc}`, + ); + assert.ok(w, `#3697-10 (${label}): a non-empty non-placeholder line selecting zero must warn`); + assert.match(w, /ID-shaped text that was not selected/i, `#3697-10 (${label}): ${w}`); + assert.doesNotMatch( + w, + /Unparsed text/i, + `#3697-10 (${label}): the citation is not unparsed requirement content: ${w}`, + ); + assert.doesNotMatch( + w, + /Range forms are not expanded/i, + `#3697-10 (${label}): no range rule fired, so no range may be diagnosed: ${w}`, + ); + assert.match(w, /write `TBD` or `None`/i, `#3697-10 (${label}): ${w}`); + }, + ); + } + + // ── Census closure ───────────────────────────────────────────────────────── + // The range-operator enumeration is a set the code fixes at author time over + // a domain (separator spellings) that grows without it. Round 3's census + // found the ASCII and typographic dashes split: U+2013/U+2014 were reached, + // the five other Unicode dashes were not, and each miss is a SILENT + // under-selection — #3697's own defect. They carry no collision risk because + // they are not the REQ-ID separator (that is ASCII `-`), so they close. + for (const [label, cp] of [ + ['U+2010 hyphen', '‐'], + ['U+2011 non-breaking hyphen', '‑'], + ['U+2012 figure dash', '‒'], + ['U+2015 horizontal bar', '―'], + ['U+2212 minus sign', '−'], + ]) { + test(`#3697-11 (${label}): a typographic dash is the same range operator at a different codepoint`, () => { + const spaced = `RANGE-01 ${cp} RANGE-05`; + assert.strictEqual( + analyzeRequirementsLine(spaced).warn, + true, + `#3697-11 (${label}): spaced form must warn`, + ); + const tight = `RANGE-01${cp}RANGE-05`; + assert.strictEqual( + analyzeRequirementsLine(tight).warn, + true, + `#3697-11 (${label}): tight form must warn`, + ); + // And the collision the ASCII hyphen has must NOT arrive with them: a + // date-shaped annotation stays silent because its own separators are + // ASCII, so it never reaches the ID shape at all. + assert.strictEqual( + analyzeRequirementsLine(`RANGE-01 (target FY-2026-08)`).warn, + false, + `#3697-11 (${label}): the date-annotation control must stay silent`, + ); + }); + } + + // Every dash takes the STRICT shape, whatever its codepoint. `PREFIX-\d+ + // \d+` is also a date and a sub-numbered ID, and that ambiguity is a + // property of the shape rather than of which key was pressed. #3697-4 pins + // the ASCII date annotation silent; these pin its seven typographic twins to + // the same verdict. Two of them — U+2013 and U+2014 — warned BEFORE this PR, + // so this arm fixes a pre-existing inconsistency as well as the one an + // earlier round-3 commit briefly introduced for the other five. + const ALL_DASHES = [ + ['ASCII hyphen-minus', '-'], + ['U+2010 hyphen', '‐'], + ['U+2011 non-breaking hyphen', '‑'], + ['U+2012 figure dash', '‒'], + ['U+2013 en dash', '–'], + ['U+2014 em dash', '—'], + ['U+2015 horizontal bar', '―'], + ['U+2212 minus sign', '−'], + ]; + + for (const [label, d] of ALL_DASHES) { + test(`#3697-13 (${label}): a date-shaped annotation must stay silent`, () => { + const line = `RANGE-01 (target FY-2026${d}08)`; + assert.strictEqual( + analyzeRequirementsLine(line).warn, + false, + `#3697-13 (${label}): ${JSON.stringify(line)} is a date annotation, not a range`, + ); + // The sub-numbered-ID reading of the same shape, beside a selected ID. + const sub = `RANGE-01, API-2${d}01`; + assert.strictEqual( + analyzeRequirementsLine(sub).warn, + false, + `#3697-13 (${label}): ${JSON.stringify(sub)} is a sub-numbered ID, not a range`, + ); + }); + + test(`#3697-13b (${label}): a tight range with a FULL ID on both sides must still warn`, () => { + const line = `RANGE-01${d}RANGE-05`; + assert.strictEqual( + analyzeRequirementsLine(line).rangeTokens.length, + 1, + `#3697-13b (${label}): ${JSON.stringify(line)} is unambiguously a range`, + ); + }); + } + + test('#3697-13c: the LOOSE operators keep their numeric endpoint', () => { + // `..`, `…` and the word operators have no date or sub-number reading + // between two numbers, so the strict shape would cost them coverage for + // nothing. They are deliberately not moved. + for (const line of ['RANGE-01…05', 'RANGE-01..05', 'RANGE-01through05']) { + assert.strictEqual( + analyzeRequirementsLine(line).rangeTokens.length, + 1, + `#3697-13c: ${JSON.stringify(line)} must still read as a tight range`, + ); + } + }); + + test('#3697-13d: the accepted false negative is now symmetric across dashes', () => { + // `RANGE-01, RANGE-02-05` is silent in the shipped design — the strict + // shape accepts that, deliberately, for the dash people actually type. + // Every other dash now accepts it identically; the inconsistency, not the + // gap, is what round 3 removed. A BARE `RANGE-0205` still warns, + // because it selects nothing and R3 catches it. + for (const [label, d] of ALL_DASHES) { + assert.strictEqual( + analyzeRequirementsLine(`RANGE-01, RANGE-02${d}05`).warn, + false, + `#3697-13d (${label}): mixed-list numeric endpoint is the accepted false negative`, + ); + assert.strictEqual( + analyzeRequirementsLine(`RANGE-02${d}05`).warn, + true, + `#3697-13d (${label}): a bare zero-selection line must still warn via R3`, + ); + } + }); + + test('#3697-12: the ASCII-hyphen strict shape is unchanged by the dash widening', () => { + // `LETTERS-\d+-\d+` is also a date and a sub-numbered ID, which is why the + // bare-hyphen tight arm demands a full ID on both sides. Widening the + // NOHYPHEN arm must not relax that. + for (const line of ['RANGE-01-05', 'FY-2026-08', 'API-2-01']) { + assert.strictEqual( + analyzeRequirementsLine(line).rangeTokens.length, + 0, + `#3697-12: ${JSON.stringify(line)} must not read as a tight range`, + ); + } + assert.strictEqual( + analyzeRequirementsLine('RANGE-01-RANGE-05').rangeTokens.length, + 1, + '#3697-12: the full-ID-both-sides spelling must still read as a range', + ); + }); +}); + // ─── #2572: phase-SUMMARY artifact↔disk advisory at phase completion ───────── // // A SUMMARY asserts "I created these files". Until #2572 nothing checked that