From 06eba5fdb03da530e0839ed431f3a2ff44c0f4cf Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sat, 5 Sep 2026 21:05:19 -0400 Subject: [PATCH 001/166] fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4130): failing-first regression for phase-prefixed decision IDs Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags, discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and parity properties (well-formed digit-prefixed ids parse to their exact id; a non-digit injected into the prefix fails loud). Update the #2347 non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and graduate the representative d5-prefix corpus fixture from could-not-parse to parsed-but-uncovered. All new rows are RED against origin/next; they go green with the parser fix in the next commit. * fix(#4130): parse phase-prefixed decision IDs (D4-01) The three declaration grammars, the parse-miss guard, the #3939 join regexes, and the token evidence all anchored on the literal 'D-' (or '**D-'), so an ID carrying a digit-run phase prefix between the leading letter and the hyphen matched nothing — while the #2347 shape detector correctly called those bullets decision-shaped, collapsing the whole CONTEXT.md to could-not-parse with 0 extracted instead of a coverage verdict. Derive the extractor ID grammar from one shared DECISION_ID_SOURCE ('D[0-9]*-' + the existing alnum tail, full id captured), widen the guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose like 'Deferred-until' stays none-present), and align the bare-token evidence. Both gates and the gap-checker share the parser, so all three surfaces read phase-prefixed decisions now; the gate messages name the accepted forms including the phase-prefixed one. * docs(#4130): document the phase-prefixed decision identifier form The canonical CONTEXT.md reference said decisions carry 'a sequential D-NN identifier' with no mention of the optional phase-number prefix the parser now accepts (D4-01) or the alphanumeric tail it always accepted (D-INFRA-01). Name both in the Decision identifier format section, EN and ja-JP. * chore(#4130): changeset * chore(#4130): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/bold-cranes-rest.md | 5 + docs/ja-JP/reference/context-md.md | 4 +- docs/reference/context-md.md | 13 +- src/check-command-router.cts | 17 +- src/decisions.cts | 139 ++++++-- tests/decisions.test.cjs | 314 +++++++++++++++++- .../decision-coverage-guard/MANIFEST.json | 5 +- .../decision-coverage-guard/README.md | 35 +- tests/representative-corpus.test.cjs | 15 +- 9 files changed, 480 insertions(+), 67 deletions(-) create mode 100644 .changeset/bold-cranes-rest.md diff --git a/.changeset/bold-cranes-rest.md b/.changeset/bold-cranes-rest.md new file mode 100644 index 000000000..647ccfcf2 --- /dev/null +++ b/.changeset/bold-cranes-rest.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4357 +--- +**The decision-coverage gate now reads phase-prefixed decision IDs** — a CONTEXT.md whose decisions use D4-01-style IDs (a digit-run phase prefix) no longer reports could-not-parse for the whole file; its decisions are counted and coverage-checked like any other, and a typo'd prefix (D4x-01) still fails loud. (#4130) diff --git a/docs/ja-JP/reference/context-md.md b/docs/ja-JP/reference/context-md.md index bc2fe02e0..48252bde1 100644 --- a/docs/ja-JP/reference/context-md.md +++ b/docs/ja-JP/reference/context-md.md @@ -51,7 +51,7 @@ ## 意思決定識別子フォーマット -`` 内のすべての意思決定は連番の `D-NN` 識別子を持ちます: +`` 内のすべての意思決定は連番の `D-NN` 識別子を持ちます。`D` とハイフンの間にフェーズ番号の接頭辞(`D4-01`)を置くこともできます。複数フェーズのプロジェクトで裸の `D-01` がフェーズ間で衝突する場合に有用で、数字は識別子の一部として扱われます(#4130): ```markdown ### Layout style @@ -59,7 +59,7 @@ - **D-02:** Each card shows: author avatar, name, timestamp, full post content, reaction counts ``` -識別子はフェーズにスコープされます。フェーズ3の `D-01` はフェーズ7の `D-01` とは無関係です。プランチェッカー(ディメンション7)は、すべての `D-NN` が生成されたプランの少なくとも1つのタスクアクションによって対処されていることを検証します。 +英数字の末尾(`D-INFRA-01`)も受け付けられます。識別子はフェーズにスコープされます。フェーズ3の `D-01` はフェーズ7の `D-01` とは無関係です。プランチェッカー(ディメンション7)は、すべての意思決定識別子が生成されたプランの少なくとも1つのタスクアクションによって対処されていることを検証します。 --- diff --git a/docs/reference/context-md.md b/docs/reference/context-md.md index cc53a1190..e172b9f3c 100644 --- a/docs/reference/context-md.md +++ b/docs/reference/context-md.md @@ -51,15 +51,24 @@ The body is divided into named XML-style blocks. The blocks appear in a fixed or ## Decision identifier format -Every decision in `` carries a sequential `D-NN` identifier: +Every decision in `` carries a sequential `D-NN` identifier. An optional +phase-number prefix may sit between the `D` and the hyphen (`D4-01`) — useful on +multi-phase projects where bare `D-01` collides across phases; the digits are read +as part of the identifier (#4130): ```markdown ### Layout style - **D-01:** Card-based layout, not timeline or list - **D-02:** Each card shows: author avatar, name, timestamp, full post content, reaction counts + +### Phase-4 decisions +- **D4-01:** Phase-scoped identifier, distinct from any other phase's D-01 ``` -Identifiers are scoped to the phase. `D-01` in Phase 3 is unrelated to `D-01` in Phase 7. The plan-checker (Dimension 7) verifies that every `D-NN` is addressed by at least one task action in the generated plans. +Alphanumeric tails (`D-INFRA-01`) are also accepted. Identifiers are scoped to the +phase. `D-01` in Phase 3 is unrelated to `D-01` in Phase 7. The plan-checker +(Dimension 7) verifies that every decision identifier is addressed by at least one +task action in the generated plans. --- diff --git a/src/check-command-router.cts b/src/check-command-router.cts index c7d79b337..380113fc4 100644 --- a/src/check-command-router.cts +++ b/src/check-command-router.cts @@ -329,12 +329,15 @@ function cmdDecisionCoveragePlan(projectDir: string, args: string[], raw: boolea uncovered: [], message: partialParse ? 'Decision coverage gate: decisions could not be fully parsed — one or more ' + - '`- **D-NN ...**` bullets appear malformed (missing `:` or ` — ` separator). ' + - 'Fix the bullet format so all D-NN decisions can be read before re-running the gate.' + '`- **D-NN ...**` bullets appear malformed (missing `:` or ` — ` separator, or a phase ' + + 'prefix that is not a digit run, e.g. `D4x-01`). Fix the bullet format so all decisions ' + + 'can be read before re-running the gate.' : 'Decision coverage gate: could not parse decisions — possible format mismatch. ' + 'The CONTEXT.md appears to be decision-shaped (has a block, a decisions heading, ' + - 'or D- tokens) but no D-NN bullets could be extracted. Check the formatting of the decisions ' + - 'block and ensure bullets follow the `- **D-NN:** text` or `- **D-NN — title** body` form.', + 'or D- tokens) but no decision bullets could be extracted. Check the formatting of the decisions ' + + 'block and ensure bullets follow the `- **D-NN:** text`, `- **D4-NN:** text` (phase-prefixed), ' + + 'or `- **D-NN — title** body` form. An ID grammar the parser does not support (e.g. `DEC-01`) ' + + 'also lands here.', }, raw, undefined); return; } @@ -433,9 +436,11 @@ function cmdDecisionCoverageVerify(projectDir: string, args: string[], raw: bool not_honored: [], message: partialParse ? 'Decision coverage verify (warning): decisions could not be fully parsed — one or more ' + - '`- **D-NN ...**` bullets appear malformed. Fix the bullet format in the CONTEXT.md decisions block.' + '`- **D-NN ...**` bullets appear malformed (missing `:` or ` — ` separator, or a phase ' + + 'prefix that is not a digit run). Fix the bullet format in the CONTEXT.md decisions block.' : 'Decision coverage verify (warning): could not parse decisions — possible format mismatch. ' + - 'Check the formatting of the CONTEXT.md decisions block.', + 'Check the formatting of the CONTEXT.md decisions block (accepted forms: `- **D-NN:** text`, ' + + '`- **D4-NN:** text` (phase-prefixed), `- **D-NN — title** body`).', }, raw, undefined); return; } diff --git a/src/decisions.cts b/src/decisions.cts index 1eb2cac1f..9688a914c 100644 --- a/src/decisions.cts +++ b/src/decisions.cts @@ -4,7 +4,9 @@ * truth). Behaviour is preserved byte-for-behaviour from the prior hand-written * .cjs; only types are added. * - * Accepts both numeric (D-42) and alphanumeric (D-INFRA-01) IDs. + * Accepts numeric (D-42), alphanumeric (D-INFRA-01), and phase-prefixed + * (D4-01 — an optional digit-run between the leading letter and the hyphen, + * #4130) IDs. * Returns {id, text, category, tags, trackable} per decision. * CJS callers that only use {id, text} safely ignore the extra fields. * @@ -57,23 +59,56 @@ const NON_TRACKABLE_TAGS = new Set(['informational', 'folded', 'deferred']); // ─── Bullet parsers (decisions-specific grammar) ───────────────────────────── /** - * Colon form: `- **D-NN[ [tags]]:** text` - * (#1343: `[^:*]*` subsumes any pre-colon prose, stops at `:**`) + * #4130: the ID grammar every extractor regex below shares, as ONE source. + * `D`, an OPTIONAL digit-run phase prefix, a hyphen, then the pre-existing + * alphanumeric tail — so `D-01` (bare), `D4-01`/`D12-01` (phase-prefixed, + * the reporter's multi-phase convention where bare D-01 collides across + * phases), and `D-INFRA-01` (alnum tail) are all the same grammar now. + * #2347 had already taught the shape DETECTOR to call `D4-01` decision-shaped + * while the EXTRACTOR still anchored on the literal `**D-` — the disagreement + * that made a whole phase-prefixed CONTEXT.md report could-not-parse. Deriving + * the three grammars (and the token evidence below) from this one constant is + * the parity pin: the extractor's ID universe cannot drift from the declared + * grammar again without editing this line, which the #4130 property tests + * watch from the other side. */ -const bulletColonRe = /^\s*-\s+\*\*D-([A-Za-z0-9][A-Za-z0-9_-]*)(?:\s*\[([^\]]+)\])?[^:*]*:\*\*\s*(.*)$/; +const DECISION_ID_SOURCE = 'D[0-9]*-[A-Za-z0-9][A-Za-z0-9_-]*'; /** - * Em-dash form: `- **D-NN[ [tags]] — title** body` + * #4130: the bold lead-in that ATTEMPTS the ID grammar above — used by the + * parse-miss guard and the #3939 join regexes, where recognising MORE shapes + * is the conservative direction (an over-broad match can only make a + * malformed bullet fail loud). The prefix run is either empty (bare `D-`) or + * DIGIT-INITIAL (`4`, `4x` — a phase prefix with a typo still counts as an + * attempted ID, so `D4x-01` reaches the guard and fails loud instead of + * vanishing), but never letter-initial: `D` + letters + `-` (`Deferred-until`) + * is a prose word, and prose must stay `none-present` (#2347's law). + */ +const ID_ATTEMPT_SOURCE = 'D(?:[0-9][A-Za-z0-9]*)?-'; + +/** + * Colon form: `- **D[phase]-NN[ [tags]]:** text` + * (#1343: `[^:*]*` subsumes any pre-colon prose, stops at `:**`) + * Group 1 captures the FULL id including any phase prefix (#4130). + */ +const bulletColonRe = new RegExp( + `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:\\*\\*\\s*(.*)$`, +); + +/** + * Em-dash form: `- **D[phase]-NN[ [tags]] — title** body` * The em-dash (U+2014) or its lookalike separates the ID+tags group from a title * that lives inside the bold markers; the body (which may be empty) follows * outside the closing `**`. This form was not handled pre-T1 (bug #1364). * * Accepts both U+2014 em-dash (—) and U+2013 en-dash (–) for robustness. */ -const bulletEmDashRe = /^\s*-\s+\*\*D-([A-Za-z0-9][A-Za-z0-9_-]*)(?:\s*\[([^\]]+)\])?[^*]*[—–][^*]*\*\*\s*(.*)$/; +const bulletEmDashRe = new RegExp( + `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^*]*[—–][^*]*\\*\\*\\s*(.*)$`, +); /** - * Titled-colon form: `- **D-NN[ [tags]]: Title.** body` + * Titled-colon form: `- **D[phase]-NN[ [tags]]: Title.** body` * A title sits between the colon and the closing `**` (so the `:**` anchor of * bulletColonRe fails, and there is no em-dash for bulletEmDashRe). This is a strict * superset of the colon-immediate form, so it MUST be checked AFTER bulletColonRe and @@ -83,16 +118,35 @@ const bulletEmDashRe = /^\s*-\s+\*\*D-([A-Za-z0-9][A-Za-z0-9_-]*)(?:\s*\[([^\]]+ * guard — matching bulletColonRe's `[^:*]*` discipline that the separator colon is the * only colon permitted before `**`. (#1639) */ -const bulletTitledColonRe = /^\s*-\s+\*\*D-([A-Za-z0-9][A-Za-z0-9_-]*)(?:\s*\[([^\]]+)\])?[^:*]*:[^:*]*\*\*\s*(.*)$/; +const bulletTitledColonRe = new RegExp( + `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:[^:*]*\\*\\*\\s*(.*)$`, +); + +/** + * #4130: the parse-miss guard's probe — a line whose bold lead-in ATTEMPTS the + * ID grammar (see `ID_ATTEMPT_SOURCE`) but failed all three bullet patterns + * above. Bare `D-` attempts behave exactly as before #4130; a digit-initial + * prefix run (`D4-`… including a typo'd `D4x-`) is new evidence of an attempt, + * so the malformed-prefixed bullet fails loud instead of silently vanishing. + */ +const parseMissGuardRe = new RegExp(`^\\s*-\\s+\\*\\*${ID_ATTEMPT_SOURCE}`); + +/** + * #4130: bare-token evidence of decision-shaped content — a `D…-` token + * in running text. `D-01` matched before; the digit-run phase prefix (`D4-01`) + * is added so token evidence agrees with the extractor's ID grammar + * (DECISION_ID_SOURCE) instead of silently ignoring prefixed mentions. + */ +const decisionTokenRe = new RegExp(`\\bD[0-9]*-[A-Za-z0-9]`, 'm'); /** * #2347: format-agnostic evidence that a block/section holds real decision * ENTRIES the parser could not read — a bullet whose bold lead-in is an * ID-SHAPED token (uppercase prefix, optional digits, hyphen, alnum), whatever - * the exact ID grammar. The three parser grammars above all require a `D-` - * prefix; #1365's fail-loud guard reused that same `\bD-` test as its "is this - * decision-shaped?" evidence, so any other prefix (e.g. `D5-01`) was invisible - * to BOTH parser and guard, collapsing `could-not-parse` into a clean + * the exact ID grammar. #1365's fail-loud guard originally reused the parser's + * own `\bD-` test as its "is this decision-shaped?" evidence, so any prefix the + * parser could not read (e.g. `D5-01` then, `DEC-01` now) was invisible to + * BOTH parser and guard, collapsing `could-not-parse` into a clean * `none-present` pass. * * The ID-shape requirement (not "any bold bullet") is deliberate: a decisions @@ -100,26 +154,36 @@ const bulletTitledColonRe = /^\s*-\s+\*\*D-([A-Za-z0-9][A-Za-z0-9_-]*)(?:\s*\[([ * bullets with bold labels (`- **Scope:** …`, `- **Why:** …`, `- **Note:** …`). * Those are NOT decision entries and must stay `none-present` — a false * `could-not-parse` hard-blocks the plan gate. `[A-Z]+[0-9]*-[A-Za-z0-9]` matches - * `D-01` / `D5-01` / `DEC-01` but not `Scope:` / `Why:` / `Follow-up:` (mixed + * `D-01` / `D4-01` / `DEC-01` but not `Scope:` / `Why:` / `Follow-up:` (mixed * case) / `TODO:` (no `-` id) — mirroring the parser's own `D-` * shape without hardcoding the `D`. + * + * #4130 parity note: for the D-prefixed universe this detector's grammar + * (`D` + digit-run + `-` + alnum) is exactly `DECISION_ID_SOURCE` above, so a + * well-formed bullet the detector calls decision-shaped is now always one the + * extractor can read. The detector stays WIDER on purpose (`DEC-01` is still + * evidence): an ID grammar outside the parser's universe must keep failing + * loud, never silently passing. The #4130 property tests pin both directions. */ const boldLeadInBulletRe = /^\s*-\s+\*\*[A-Z]+[0-9]*-[A-Za-z0-9]/m; /** - * #3939: a decision bullet's DECLARATION line — the `- **D-NN … **` bold lead-in - * the three grammars above anchor on — may wrap across a line break. Physical - * line breaks inside a bullet are markdown-insignificant, and GSD's own + * #3939: a decision bullet's DECLARATION line — the `- **D[phase]-NN … **` bold + * lead-in the three grammars above anchor on — may wrap across a line break. + * Physical line breaks inside a bullet are markdown-insignificant, and GSD's own * discuss-phase writer emits the wrapped shape whenever a decision title runs * past the wrap column. All three grammars require the closing `**` in the same - * string as the `- **D-` anchor, so a wrapped declaration matched none of them + * string as the `- **D…-` anchor, so a wrapped declaration matched none of them * and fell to the #1365 parse-miss guard, forcing `could-not-parse` (which * hard-blocks `check.decision-coverage-plan`) on a well-formed CONTEXT.md. * * The repair is confined to how the LOGICAL bullet is assembled — the grammars * themselves are untouched, so every single-line form parses exactly as before. + * #4130: the anchor uses `ID_ATTEMPT_SOURCE` (digit-run phase prefixes join + * like bare ones; recognising more start shapes only reassembles the logical + * bullet, which then parses or fails loud as itself). */ -const decisionBulletStartRe = /^\s*-\s+\*\*D-/; +const decisionBulletStartRe = new RegExp(`^\\s*-\\s+\\*\\*${ID_ATTEMPT_SOURCE}`); /** * A line that opens a new BLOCK-LEVEL construct, and therefore terminates the @@ -169,7 +233,8 @@ const blockConstructRe = /^(?:[-*+]\s|\d+[.)]\s|#{1,6}\s|>\s|\|)/; * to watch for a splice. * * The id character class is deliberately looser than the grammars' (it admits - * an empty id, so a bare `- **D-` still counts as unsettled). This regex only + * an empty id, so a bare `- **D-` still counts as unsettled, and — #4130 — a + * digit-run phase prefix between the `D` and the first hyphen). This regex only * answers "may an id-adjacent bracket still open here?", where recognising MORE * shapes is the conservative direction: an over-broad match can only make a * malformed bullet fail loud, while a missed one silently re-classifies. @@ -178,7 +243,7 @@ const blockConstructRe = /^(?:[-*+]\s|\d+[.)]\s|#{1,6}\s|>\s|\|)/; * into `tags` (and therefore into `trackable`). A `[` further along the title is * ordinary text and does not restrict the join. */ -const tagRegionRe = /^\s*-\s+\*\*D-[A-Za-z0-9_-]*\s*(?:\[([^\]]*))?$/; +const tagRegionRe = new RegExp(`^\\s*-\\s+\\*\\*D[0-9]*-[A-Za-z0-9_-]*\\s*(?:\\[([^\\]]*))?$`); /** * #3939 (review): would folding `next` onto a lead-in whose `[tags]` bracket is @@ -391,11 +456,11 @@ function parseDecisionLines(block: string): ParseDecisionLinesResult { continue; } - // Colon form: `- **D-NN[ [tags]]:** text` + // Colon form: `- **D[phase]-NN[ [tags]]:** text` const colonMatch = line.match(bulletColonRe); if (colonMatch) { flush(); - const id = `D-${colonMatch[1]}`; + const id = colonMatch[1]; const tags = colonMatch[2] ? colonMatch[2].split(',').map((t) => t.trim().toLowerCase()).filter(Boolean) : []; @@ -405,11 +470,11 @@ function parseDecisionLines(block: string): ParseDecisionLinesResult { continue; } - // Em-dash form: `- **D-NN[ [tags]] — title** body` + // Em-dash form: `- **D[phase]-NN[ [tags]] — title** body` const emDashMatch = line.match(bulletEmDashRe); if (emDashMatch) { flush(); - const id = `D-${emDashMatch[1]}`; + const id = emDashMatch[1]; const tags = emDashMatch[2] ? emDashMatch[2].split(',').map((t) => t.trim().toLowerCase()).filter(Boolean) : []; @@ -422,14 +487,14 @@ function parseDecisionLines(block: string): ParseDecisionLinesResult { continue; } - // Titled-colon form: `- **D-NN[ [tags]]: Title.** body` (#1639). Checked LAST — it is + // Titled-colon form: `- **D[phase]-NN[ [tags]]: Title.** body` (#1639). Checked LAST — it is // a strict superset of bulletColonRe, so it only catches bullets the colon-immediate // and em-dash forms missed (minimal blast radius). id + [tags] trackability honored; // the body after the closing bold run is reported as text. const titledColonMatch = line.match(bulletTitledColonRe); if (titledColonMatch) { flush(); - const id = `D-${titledColonMatch[1]}`; + const id = titledColonMatch[1]; const tags = titledColonMatch[2] ? titledColonMatch[2].split(',').map((t) => t.trim().toLowerCase()).filter(Boolean) : []; @@ -439,10 +504,14 @@ function parseDecisionLines(block: string): ParseDecisionLinesResult { continue; } - // Parse-miss guard (FIX B + #1343): a line that looks like a `D-NN` decision - // bullet but failed both patterns — flush, warn, and record the miss. + // Parse-miss guard (FIX B + #1343, grammar widened #4130): a line whose bold + // lead-in ATTEMPTS the ID grammar but failed all three patterns — flush, + // warn, and record the miss. `ID_ATTEMPT_SOURCE` accepts the bare `D-` form + // (as before) plus a digit-initial prefix run, so a typo'd phase prefix + // (`D4x-01`) fails loud instead of silently vanishing, while a letter-initial + // run (`Deferred-until`) stays prose and stays invisible. // parseMisses > 0 forces could-not-parse even when other decisions parsed. - if (/^\s*-\s+\*\*D-/.test(line)) { + if (parseMissGuardRe.test(line)) { flush(); parseMisses += 1; console.warn(`parseDecisions: ignored unparseable decision bullet: ${trimmed}`); @@ -502,10 +571,10 @@ export function extractDecisions(content: unknown): DecisionExtraction { // FIX A: Block present but 0 extracted and no parse-misses. // Only report could-not-parse when there is genuine evidence of real decisions // that failed to parse: a bold-lead-in bullet (`- **…**`, any ID grammar — #2347), - // a \bD- token in the block text, or an unterminated fence. An empty scaffold + // a bare `D[phase]-` token (#4130) in the block text, or an unterminated fence. An empty scaffold // () or an all-prose block has no such evidence — treat // as none-present so the gate passes cleanly. - const hasDecisionTokenInBlock = /\bD-[A-Za-z0-9]/m.test(combined); + const hasDecisionTokenInBlock = decisionTokenRe.test(combined); const hasBoldLeadInBullet = boldLeadInBulletRe.test(combined); if (hasDecisionTokenInBlock || hasBoldLeadInBullet || unterminatedFence) { return { decisions: [], outcome: 'could-not-parse' }; @@ -534,10 +603,10 @@ export function extractDecisions(content: unknown): DecisionExtraction { } // FIX A: Heading found but 0 extracted and no parse-misses. // Report could-not-parse when the section body holds a decision-entry-shaped - // bold-lead-in bullet (`- **…**`, any ID grammar — #2347) or a D- token. A + // bold-lead-in bullet (`- **…**`, any ID grammar — #2347) or a `D[phase]-` token (#4130). A // heading with only prose, sub-headings, or all-discretion content (no such // evidence) is a legitimate empty/discretion section → none-present. - const hasDecisionTokenInSection = /\bD-[A-Za-z0-9]/m.test(section.body); + const hasDecisionTokenInSection = decisionTokenRe.test(section.body); const hasBoldLeadInBulletInSection = boldLeadInBulletRe.test(section.body); if (hasDecisionTokenInSection || hasBoldLeadInBulletInSection) { return { decisions: [], outcome: 'could-not-parse' }; @@ -547,8 +616,8 @@ export function extractDecisions(content: unknown): DecisionExtraction { // ── Path 3: no blocks, no heading ──────────────────────────────────────────── // Apply shape heuristics to distinguish none-present from could-not-parse. - // We re-use the already-computed unterminatedFence and check for D- tokens. - const hasDecisionToken = /\bD-[A-Za-z0-9]/m.test(stripped); + // We re-use the already-computed unterminatedFence and check for decision tokens. + const hasDecisionToken = decisionTokenRe.test(stripped); if (unterminatedFence || hasDecisionToken) { return { decisions: [], outcome: 'could-not-parse' }; } diff --git a/tests/decisions.test.cjs b/tests/decisions.test.cjs index 334c87659..d9219791b 100644 --- a/tests/decisions.test.cjs +++ b/tests/decisions.test.cjs @@ -234,14 +234,20 @@ describe('extractDecisions — typed outcome (#1364 + #1365)', () => { // literal "D-" token, so they isolate the bold-bullet evidence from the old // D--token path. describe('extractDecisions — format-agnostic evidence test (#2347)', () => { + // #4130 note: this fixture originally used the D5-NN phase-prefixed shape + // from #2347's own reproduction. #4130 made the digit-run phase prefix LEGAL + // (see the #4130 blocks below — D5-01 now parses), so the "prefix the parser + // cannot read" fixture here uses a DEC-NN multi-letter prefix instead: still + // an ID-shaped bold lead-in, still outside the parser's D-prefixed universe, + // so the fail-loud contract this describe-block pins is unchanged. test('populated block with a non-D- ID prefix is could-not-parse, not none-present', () => { const md = '\n' - + '- **D5-01:** choose the primary datastore\n' - + '- **D5-02:** pick the queue technology\n' - + '- **D5-03:** settle on the auth model\n' + + '- **DEC-01:** choose the primary datastore\n' + + '- **DEC-02:** pick the queue technology\n' + + '- **DEC-03:** settle on the auth model\n' + '\n'; const r = extractDecisions(md); - assert.strictEqual(r.decisions.length, 0, 'parser cannot read the D5- prefix (0 extracted)'); + assert.strictEqual(r.decisions.length, 0, 'parser cannot read the DEC- prefix (0 extracted)'); assert.strictEqual(r.outcome, 'could-not-parse', 'a populated block the parser cannot read must FAIL LOUD, not pass as none-present'); }); @@ -2216,3 +2222,303 @@ describe('check.decision-coverage-plan — wrapped bold lead-in does not hard-bl `Both wrapped decisions must be seen as covered by the plan. Got: ${JSON.stringify(parsed)}`); }); }); + +// ─── #4130: phase-prefixed decision IDs (D4-01) must parse ─────────────────── +// +// The three declaration grammars all anchored on the literal `**D-`, so an ID +// carrying a phase-number prefix between the leading letter and the hyphen +// (D4-01, D12-01 — the reporter's D3-NN/D4-NN/D5-NN multi-phase convention, +// where bare D-01 collides across 18 phases) matched none of them — and none +// of the parse-miss guard or the #3939 join regexes either, so such a bullet +// was INVISIBLE to the extractor while the #2347 evidence detector correctly +// called the file decision-shaped. Net effect: the whole CONTEXT.md collapsed +// to could-not-parse with 0 extracted and the gate reported a format problem +// instead of a coverage result. The extractor now accepts an optional +// digit-run phase prefix on the same grammar the detector already recognized. + +describe('parseDecisions — phase-prefixed IDs parse in every form (#4130)', () => { + test('FAIL-FIRST: - **D4-01:** single-line colon form with a phase prefix parses', () => { + // Before the fix: outcome could-not-parse, 0 extracted (the issue's own + // b-phase-prefix fixture — the only variable vs the parsing baseline is + // the `4` in the ID). + const md = '\n- **D4-01:** a short single-line decision.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed', + `A phase-prefixed colon bullet must parse. Got: ${JSON.stringify(r)}`); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D4-01'], + `The full prefixed id must be reported. Got: ${JSON.stringify(r.decisions)}`); + assert.strictEqual(r.decisions[0].text, 'a short single-line decision.'); + }); + + test('- **D12-01:** two-digit phase prefix parses', () => { + const md = '\n- **D12-01:** two-digit phase.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed'); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D12-01']); + }); + + test('em-dash form - **D4-01 — title** body parses with a phase prefix', () => { + const md = '\n- **D4-01 — the chosen datastore** use Postgres.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed'); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D4-01']); + }); + + test('titled-colon form - **D4-01: Title.** body parses with a phase prefix', () => { + const md = '\n- **D4-01: The chosen datastore.** use Postgres.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed'); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D4-01']); + }); + + test('D5-NN (the #2347 reproduction shape) now parses — the multi-phase convention is legal', () => { + const md = '\n' + + '- **D5-01:** choose the primary datastore\n' + + '- **D5-02:** pick the queue technology\n' + + '- **D5-03:** settle on the auth model\n' + + '\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed', + `The exact #2347 fixture grammar must now be readable. Got: ${JSON.stringify(r)}`); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D5-01', 'D5-02', 'D5-03']); + }); + + test('phase prefix + [tags] honors tags and trackable:false', () => { + const md = '\n- **D4-01 [informational]:** reference only.\n\n'; + const ds = parseDecisions(md); + assert.deepStrictEqual(ds.map((d) => d.id), ['D4-01']); + assert.ok(ds[0].tags.includes('informational'), + `Tags must survive the prefixed grammar. Got: ${JSON.stringify(ds[0])}`); + assert.strictEqual(ds[0].trackable, false); + }); + + test("phase-prefixed decision under ### Claude's Discretion is non-trackable", () => { + const md = "\n### Claude's Discretion\n- **D4-01:** internal choice.\n\n"; + const ds = parseDecisions(md); + assert.deepStrictEqual(ds.map((d) => d.id), ['D4-01']); + assert.strictEqual(ds[0].trackable, false); + }); + + test('phase-prefixed bullet with a wrapped bold lead-in parses like the one-line form (#4130 x #3939)', () => { + const oneLine = extractDecisions(inBlock('- **D4-01: Persist the raw delivery headers.** JSON arrays preserve order.')); + const wrapped = extractDecisions(inBlock( + '- **D4-01: Persist the raw delivery\n headers.** JSON arrays preserve order.')); + assert.strictEqual(oneLine.outcome, 'parsed'); + assert.strictEqual(wrapped.outcome, 'parsed', + `A wrapped prefixed declaration must join and parse. Got: ${JSON.stringify(wrapped)}`); + assert.deepStrictEqual(wrapped.decisions.map((d) => d.id), ['D4-01']); + assert.deepStrictEqual(wrapped.decisions[0].text, oneLine.decisions[0].text, + 'Wrapping must stay markdown-insignificant for prefixed ids too.'); + }); +}); + +describe('parseDecisions — phase-prefix failure modes stay loud, prose stays prose (#4130)', () => { + test('- **D4x-01:** (non-digit inside the prefix) is a genuine parse-miss → could-not-parse', () => { + const md = '\n- **D4x-01:** a typo in the phase prefix.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'could-not-parse', + `A malformed phase prefix must fail loud, not silently extract or vanish. Got: ${JSON.stringify(r)}`); + assert.deepStrictEqual(r.decisions, [], + 'A malformed prefix must never be extracted as a decision.'); + }); + + test('a malformed phase-prefixed bullet poisons a file that also has valid decisions (FIX B parity)', () => { + // Before the fix this file was outcome:parsed with D-01 only — D4x-02 was + // silently invisible to the extractor AND the guard (the exact silent-drop + // class #1365 FIX B exists to prevent, surviving for prefixed ids). + const md = '\n- **D-01:** valid.\n- **D4x-02:** malformed prefix.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'could-not-parse', + `A parse-miss on a prefixed bullet must block, not silently drop. Got: ${JSON.stringify(r)}`); + }); + + test('D-initial hyphenated prose labels stay none-present (the widened guard must not reach prose)', () => { + // `Deferred-until-later` is `D` + letters + `-`: the letter-initial run is + // a prose word, not a digit-run phase prefix, so it must stay invisible to + // the parse-miss guard exactly as before #4130. + const md = '\n' + + '- **Deferred-until-later:** we revisit the queue choice next phase.\n' + + '- **Note:** nothing else was decided here.\n' + + '\n'; + assert.strictEqual(extractDecisions(md).outcome, 'none-present', + 'D-initial hyphenated prose labels must not become parse-misses.'); + }); + + test('bare D-01 baseline and D-INFRA-01 alnum tail parse unchanged', () => { + const md = '\n- **D-01:** bare.\n- **D-INFRA-01:** alnum tail.\n\n'; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed'); + assert.deepStrictEqual(r.decisions.map((d) => d.id), ['D-01', 'D-INFRA-01']); + }); +}); + +// ─── #4130 properties: detector/extractor ID-grammar parity ───────────────── +// +// The bug was a grammar DISAGREEMENT: the #2347 evidence detector accepted +// `D4-01` as decision-shaped while the extractor's `**D-` anchor rejected it, +// so the file failed loud as a whole instead of being read. These properties +// pin the parity invariant in both directions for the D-prefixed universe: +// every well-formed digit-prefixed id (the shape the detector already calls +// decision-shaped) must PARSE to its exact id in every bullet form, and every +// malformed variant of that shape (a non-digit inside the digit-run prefix) +// must FAIL LOUD — never a silent none-present, never a silent extraction. +// If either grammar drifts from the other again, one of these fires. + +describe('#4130 properties: digit-prefixed ids parse or fail loud, never vanish', () => { + const phasePrefixedIdArb = fc + .tuple(fc.integer({ min: 1, max: 99 }), fc.integer({ min: 1, max: 99 })) + .map(([phase, seq]) => `D${phase}-${String(seq).padStart(2, '0')}`); + + test('property: every well-formed phase-prefixed bullet parses to its exact id, in every form', () => { + fc.assert( + fc.property( + fc.constantFrom(...Object.keys(BULLET_FORMS)), + phasePrefixedIdArb, + tagsArb, + proseArb(2, 8), + proseArb(1, 6), + fc.boolean(), + (form, id, tags, title, body, withCategory) => { + const heading = withCategory ? '### Implementation\n' : ''; + const line = renderBullet(form, id, tags, title, body); + const r = extractDecisions(inBlock(heading + line)); + assert.strictEqual(r.outcome, 'parsed', + `A well-formed phase-prefixed declaration must parse. form=${form} line=${JSON.stringify(line)} → ${JSON.stringify(r)}`); + assert.deepStrictEqual(r.decisions.map((d) => d.id), [id], + `The prefixed id must round-trip exactly. form=${form} id=${id} → ${JSON.stringify(r.decisions)}`); + return true; + }, + ), + ); + }); + + test('property: a non-digit injected into the phase prefix fails loud, never silently', (t) => { + const originalWarn = console.warn; + console.warn = () => {}; + t.after(() => { console.warn = originalWarn; }); + + fc.assert( + fc.property( + fc.constantFrom(...Object.keys(BULLET_FORMS)), + phasePrefixedIdArb, + fc.constantFrom('x', 'X', 'z'), + (form, id, junk) => { + const badId = id.replace(/^(D\d)/, `$1${junk}`); + const line = renderBullet(form, badId, '', 'title', 'body'); + const r = extractDecisions(inBlock(line)); + assert.strictEqual(r.outcome, 'could-not-parse', + `A malformed phase prefix must fail loud. form=${form} line=${JSON.stringify(line)} → ${JSON.stringify(r)}`); + assert.deepStrictEqual(r.decisions, [], + `A malformed phase prefix must never be extracted. form=${form} line=${JSON.stringify(line)}`); + return true; + }, + ), + ); + }); +}); + +// ─── #4130 gate-level: the gates read phase-prefixed decisions end-to-end ──── + +describe('check.decision-coverage-plan — phase-prefixed decisions are readable (#4130)', () => { + let tmpDir; + let planningDir; + let phaseDir; + + beforeEach(() => { + tmpDir = createTempProject('gsd-4130-'); + planningDir = path.join(tmpDir, '.planning'); + phaseDir = path.join(planningDir, 'phases', '01-init'); + fs.mkdirSync(phaseDir, { recursive: true }); + }); + + afterEach(() => cleanup(tmpDir)); + + test('FAIL-FIRST: CONTEXT.md with D4-01 covered by the plan → passed:true, total:1, covered:1', () => { + // Before the fix: reason could-not-parse, total 0 — the gate reported a + // format problem for the whole file instead of a coverage result. + writeContextFile(phaseDir, [ + '# Phase 4 Context', + '', + '', + '### Implementation', + '- **D4-01:** use the phase-scoped datastore', + '', + ].join('\n')); + writePlanFile(phaseDir, '01', '# Plan\n## Must Haves\n- D4-01: provision the datastore\n'); + + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + const result = runDecisionCoveragePlan(phaseDir, contextPath, tmpDir); + const parsed = JSON.parse(result.output || '{}'); + assert.strictEqual(parsed.passed, true, + `A plan covering the prefixed decision must pass. Got: ${JSON.stringify(parsed)}`); + assert.strictEqual(parsed.total, 1, + `The prefixed decision must be counted. Got: ${JSON.stringify(parsed)}`); + assert.strictEqual(parsed.covered, 1, + `The prefixed decision must be seen as covered. Got: ${JSON.stringify(parsed)}`); + }); + + test('D4-01 not covered → passed:false with the uncovered id, NOT could-not-parse', () => { + writeContextFile(phaseDir, [ + '# Phase 4 Context', + '', + '', + '- **D4-01:** use the phase-scoped datastore', + '', + ].join('\n')); + writePlanFile(phaseDir, '01', '# Plan\n## Must Haves\n- Something unrelated.\n'); + + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + const result = runDecisionCoveragePlan(phaseDir, contextPath, tmpDir); + const parsed = JSON.parse(result.output || '{}'); + assert.strictEqual(parsed.passed, false, + `An uncovered prefixed decision must fail on coverage. Got: ${JSON.stringify(parsed)}`); + assert.notStrictEqual(parsed.reason, 'could-not-parse', + `The gate must report a coverage result, not a format problem. Got: ${JSON.stringify(parsed)}`); + assert.strictEqual(parsed.total, 1); + assert.strictEqual(parsed.covered, 0); + assert.deepStrictEqual( + (parsed.uncovered || []).map((u) => u.id), + ['D4-01'], + `The uncovered row must carry the prefixed id. Got: ${JSON.stringify(parsed.uncovered)}`, + ); + }); +}); + +describe('check.decision-coverage-verify — phase-prefixed decisions are readable (#4130)', () => { + let tmpDir; + let planningDir; + let phaseDir; + + beforeEach(() => { + tmpDir = createTempProject('gsd-4130v-'); + planningDir = path.join(tmpDir, '.planning'); + phaseDir = path.join(planningDir, 'phases', '01-init'); + fs.mkdirSync(phaseDir, { recursive: true }); + }); + + afterEach(() => cleanup(tmpDir)); + + test('verify reads D4-01 and honors it when the plan mentions it (no could-not-parse)', () => { + writeContextFile(phaseDir, [ + '# Phase 4 Context', + '', + '', + '- **D4-01:** use the phase-scoped datastore', + '', + ].join('\n')); + writePlanFile(phaseDir, '01', '# Plan\n## Must Haves\n- D4-01: provision the datastore\n'); + + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + const result = runGsdTools( + ['query', 'check.decision-coverage-verify', phaseDir, contextPath], + tmpDir, + ); + const parsed = JSON.parse(result.output || '{}'); + assert.notStrictEqual(parsed.reason, 'could-not-parse', + `Verify must read prefixed decisions, not report a format mismatch. Got: ${JSON.stringify(parsed)}`); + assert.strictEqual(parsed.total, 1, + `The prefixed decision must be counted. Got: ${JSON.stringify(parsed)}`); + assert.strictEqual(parsed.honored, 1, + `A plan mentioning D4-01 honors it. Got: ${JSON.stringify(parsed)}`); + }); +}); diff --git a/tests/fixtures/representative/decision-coverage-guard/MANIFEST.json b/tests/fixtures/representative/decision-coverage-guard/MANIFEST.json index 10b5e2b48..83f40168e 100644 --- a/tests/fixtures/representative/decision-coverage-guard/MANIFEST.json +++ b/tests/fixtures/representative/decision-coverage-guard/MANIFEST.json @@ -4,9 +4,10 @@ "fixtures": [ { "file": "d5-prefix-context.md", - "expectedReason": "could-not-parse", "expectedPassed": false, - "note": "A block using the D5-01 ID-prefix shape from #2347's own reproduction ('- **D5-01:** some decision'), repeated twice so 'populated but 0 extracted' is unambiguous. The original report used 23 real decisions under a project-specific D5- prefix convention; this fixture preserves the exact grammar mismatch, not the count. FIXED by #2347: the guard's evidence test no longer reuses the parser's D- grammar — a bold ID-shaped lead-in bullet (any prefix) is now evidence, so this populated block correctly fails loud (could-not-parse). currentBuggyOutput removed and the assertion graduated to expected* per this corpus's contract." + "expectedTotal": 2, + "expectedCovered": 0, + "note": "A block using the D5-01 ID-prefix shape from #2347's own reproduction ('- **D5-01:** some decision'), repeated twice so 'populated' is unambiguous. The original report used 23 real decisions under a project-specific D5- prefix convention; this fixture preserves the exact grammar, not the count. History: #2347 made the guard's evidence test format-agnostic so this populated block failed loud (could-not-parse) instead of silently passing. #4130 then made the digit-run phase prefix LEGAL — the parser now reads D5-01/D5-02 — so the gate reports a real coverage verdict for this shape: both decisions parsed (total 2) and uncovered (covered 0, passed false, no reason field) in this bare project. The corpus's graduation contract: the assertion was updated as part of the #4130 fix that changed the observable verdict." } ] } diff --git a/tests/fixtures/representative/decision-coverage-guard/README.md b/tests/fixtures/representative/decision-coverage-guard/README.md index bdba0a51d..c590cb2c9 100644 --- a/tests/fixtures/representative/decision-coverage-guard/README.md +++ b/tests/fixtures/representative/decision-coverage-guard/README.md @@ -1,4 +1,4 @@ -# Decision-coverage guard fixture (#2347) +# Decision-coverage guard fixture (#2347, graduated by #4130) `d5-prefix-context.md` is the verbatim reproduction shape from #2347 — the `- **D5-01:** some decision` bullet given in the issue's own "Steps to @@ -7,18 +7,23 @@ reproduce" — used as a CONTEXT.md `` block and driven through CLI gate; see `tests/decisions.test.cjs` for the established pattern this follows), not `extractDecisions()` called in isolation. -The #1365 fail-loud guard's "is this decision-shaped?" evidence test -(`/\bD-[A-Za-z0-9]/`) reuses the same `D-` grammar as the parser it guards. -For any ID prefix the parser cannot read — `D5-01` here — the guard sees -no evidence either, so the two failure modes the guard exists to -distinguish (`none-present` vs `could-not-parse`) collapse into -`none-present`, and a populated, genuinely decision-shaped CONTEXT.md -passes silently. +History of the expected verdict for this exact shape: -Expected once fixed: `reason: 'could-not-parse'`, `passed: false` (see -`MANIFEST.json`'s `expectedReason`/`expectedPassed`). Today's gate instead -reports `passed: true, skipped: true, reason: 'no trackable decisions'` — -pinned in `MANIFEST.json`'s `currentBuggyOutput` and asserted directly in -`tests/representative-corpus.test.cjs` (a characterization of today's -known-broken behavior, not a `todo` — see -`tests/fixtures/representative/README.md` for why) until #2347 lands. +1. **#2347 (original):** the #1365 fail-loud guard's "is this + decision-shaped?" evidence test (`/\bD-[A-Za-z0-9]/`) reused the same + `D-` grammar as the parser it guards, so `D5-01` was invisible to both + and a populated CONTEXT.md passed silently. #2347 made the evidence test + format-agnostic, and the fixture's expectation graduated from the + characterized buggy green-skip to `could-not-parse`. +2. **#4130 (current):** the digit-run phase prefix itself became a LEGAL + ID grammar — the parser now reads `D5-01`/`D5-02` — so this shape no + longer fails loud at all. The gate reports a real coverage verdict: + `total: 2`, `covered: 0`, `passed: false`, and NO `reason` field (the + coverage path emits none). Pinned in `MANIFEST.json`'s + `expectedTotal`/`expectedCovered`/`expectedPassed` and asserted in + `tests/representative-corpus.test.cjs`. + +The fail-loud contract itself is unchanged and still covered by the +`DEC-`-prefix (non-`D` universe) tests in `tests/decisions.test.cjs`: +an ID grammar the parser genuinely cannot read still yields +`could-not-parse`. diff --git a/tests/representative-corpus.test.cjs b/tests/representative-corpus.test.cjs index 573bb2b17..1bc6cf587 100644 --- a/tests/representative-corpus.test.cjs +++ b/tests/representative-corpus.test.cjs @@ -207,7 +207,9 @@ describe('representative corpus — decision-coverage guard (#2347)', () => { for (const fx of manifest.fixtures) { const label = fx.currentBuggyOutput ? `${fx.file} → currently passed:true, skipped:true (#2347)` - : `${fx.file} → outcome could-not-parse, passed:false`; + : typeof fx.expectedTotal === 'number' + ? `${fx.file} → parsed but uncovered, passed:false (#4130)` + : `${fx.file} → outcome could-not-parse, passed:false`; test(label, () => { tmpDir = makeProject(); const phaseDir = makePhaseDir(tmpDir, '01-repcorpus'); @@ -225,6 +227,17 @@ describe('representative corpus — decision-coverage guard (#2347)', () => { assert.strictEqual(j.skipped, fx.currentBuggyOutput.skipped, `${fx.file}: skipped`); assert.strictEqual(j.reason, fx.currentBuggyOutput.reason, `${fx.file}: reason`); assert.strictEqual(j.total, fx.currentBuggyOutput.total, `${fx.file}: total`); + } else if (typeof fx.expectedTotal === 'number') { + // #4130: the fixture's D5-NN phase-prefixed ids are now READ by the + // parser, so the gate reports a real coverage verdict (nothing covers + // them in this bare project) instead of could-not-parse. The coverage + // path emits no `reason` field at all — absence is part of the + // expectation. + assert.strictEqual(j.passed, fx.expectedPassed, `${fx.file}: passed. Got ${JSON.stringify(j)}`); + assert.strictEqual(j.reason, undefined, + `${fx.file}: the coverage verdict must carry no could-not-parse reason. Got ${JSON.stringify(j)}`); + assert.strictEqual(j.total, fx.expectedTotal, `${fx.file}: total. Got ${JSON.stringify(j)}`); + assert.strictEqual(j.covered, fx.expectedCovered, `${fx.file}: covered. Got ${JSON.stringify(j)}`); } else { assert.strictEqual(j.passed, fx.expectedPassed, `${fx.file}: passed. Got ${JSON.stringify(j)}`); assert.strictEqual(j.reason, fx.expectedReason, `${fx.file}: reason. Got ${JSON.stringify(j)}`); From b3906c66f6e91aa42043e9f9071bdf07cfed0c37 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" <41898282+github-actions[bot]@users.noreply.github.com> Date: Sun, 6 Sep 2026 02:10:29 +0000 Subject: [PATCH 002/166] chore: sync next package version to 1.13.0 --- .claude-plugin/marketplace.json | 2 +- .claude-plugin/plugin.json | 2 +- capabilities/ai-integration/capability.json | 2 +- capabilities/antigravity/capability.json | 2 +- capabilities/assumption-delta/capability.json | 2 +- capabilities/audit/capability.json | 2 +- capabilities/augment/capability.json | 2 +- capabilities/broken-windows/capability.json | 2 +- .../claude-orchestration/capability.json | 2 +- capabilities/claude/capability.json | 2 +- capabilities/cline/capability.json | 2 +- capabilities/code-review/capability.json | 2 +- capabilities/codebuddy/capability.json | 2 +- capabilities/coderabbit/capability.json | 2 +- capabilities/codex/capability.json | 2 +- capabilities/copilot/capability.json | 2 +- capabilities/cursor/capability.json | 2 +- capabilities/drift/capability.json | 2 +- capabilities/external-job/capability.json | 2 +- capabilities/gap-analysis/capability.json | 2 +- capabilities/gemini/capability.json | 2 +- capabilities/graphify/capability.json | 2 +- capabilities/hermes/capability.json | 2 +- capabilities/intel/capability.json | 2 +- capabilities/kilo/capability.json | 2 +- capabilities/kimi-code/capability.json | 2 +- capabilities/kimi/capability.json | 2 +- capabilities/live-dom-uat/capability.json | 2 +- capabilities/llama-cpp/capability.json | 2 +- capabilities/lm-studio/capability.json | 2 +- capabilities/mempalace/capability.json | 2 +- capabilities/nyquist/capability.json | 2 +- capabilities/ollama/capability.json | 2 +- capabilities/opencode/capability.json | 2 +- capabilities/pattern-mapper/capability.json | 2 +- capabilities/pi/capability.json | 2 +- capabilities/profile-pipeline/capability.json | 2 +- capabilities/qwen/capability.json | 2 +- capabilities/refactor-trigger/capability.json | 2 +- capabilities/research/capability.json | 2 +- capabilities/schema-gate/capability.json | 2 +- capabilities/security/capability.json | 2 +- capabilities/tdd/capability.json | 2 +- capabilities/trae/capability.json | 2 +- capabilities/ui/capability.json | 2 +- capabilities/vscode/capability.json | 2 +- capabilities/windsurf/capability.json | 2 +- capabilities/zcode/capability.json | 2 +- gsd-core/bin/lib/capability-registry.cjs | 130 +++++++++--------- package-lock.json | 4 +- package.json | 2 +- vscode/package.json | 2 +- 52 files changed, 117 insertions(+), 117 deletions(-) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 5f64ca67e..5dda4c795 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -9,7 +9,7 @@ { "name": "gsd-core", "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.", - "version": "1.12.0", + "version": "1.13.0", "source": "./", "author": { "name": "open-gsd", diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index 2357ac945..a6ecbf90e 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "gsd-core", "displayName": "GSD Core", - "version": "1.12.0", + "version": "1.13.0", "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.", "author": { "name": "open-gsd", diff --git a/capabilities/ai-integration/capability.json b/capabilities/ai-integration/capability.json index 62097466c..086bab607 100644 --- a/capabilities/ai-integration/capability.json +++ b/capabilities/ai-integration/capability.json @@ -1,7 +1,7 @@ { "id": "ai-integration", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "AI design contract", "description": "AI-SPEC design contract workflow for phases that build AI systems; owns the AI integration command, agents, and workflow.ai_integration_phase activation key.", "tier": "full", diff --git a/capabilities/antigravity/capability.json b/capabilities/antigravity/capability.json index 269beeb99..2b685fe93 100644 --- a/capabilities/antigravity/capability.json +++ b/capabilities/antigravity/capability.json @@ -1,7 +1,7 @@ { "id": "antigravity", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Antigravity", "description": "Google Antigravity IDE — config/settings home nested under ~/.gemini/antigravity (probed across 1.x and 2.x layouts); global skills/agents install under ~/.gemini/config, the dir AGY scans for global discovery (#3738); Gemini hook event dialect; flat skill layout; tier-1 support.", "tier": "core", diff --git a/capabilities/assumption-delta/capability.json b/capabilities/assumption-delta/capability.json index 1428b2c7b..b21f54670 100644 --- a/capabilities/assumption-delta/capability.json +++ b/capabilities/assumption-delta/capability.json @@ -1,7 +1,7 @@ { "id": "assumption-delta", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Assumption-delta architecture checkpoint", "description": "Rarely-firing advisory checkpoint that triggers when a phase makes something plural, optional, or chosen that used to be singular, required, or derived. Surfaces one identity-model question (promote the new general representation to primary, or add it alongside?) so a silent primary-key drift does not accumulate into a later user-facing bug. Non-blocking; fires only on a detected signal.", "tier": "full", diff --git a/capabilities/audit/capability.json b/capabilities/audit/capability.json index 6272d2334..010dd4324 100644 --- a/capabilities/audit/capability.json +++ b/capabilities/audit/capability.json @@ -1,7 +1,7 @@ { "id": "audit", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Audit", "description": "Open-artifact audit and UAT-gap audit for milestone close gates; exposes `gsd-tools audit-uat` (cross-phase UAT outstanding items) and `gsd-tools audit-open` (structured open-artifact scan across debug, tasks, threads, todos, seeds, UAT, verification, context-questions).", "tier": "full", diff --git a/capabilities/augment/capability.json b/capabilities/augment/capability.json index 012d14296..061e742fe 100644 --- a/capabilities/augment/capability.json +++ b/capabilities/augment/capability.json @@ -1,7 +1,7 @@ { "id": "augment", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Augment Code", "description": "Augment Code CLI — commands + nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", diff --git a/capabilities/broken-windows/capability.json b/capabilities/broken-windows/capability.json index bc2668904..3b4508cb0 100644 --- a/capabilities/broken-windows/capability.json +++ b/capabilities/broken-windows/capability.json @@ -1,7 +1,7 @@ { "id": "broken-windows", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Broken-windows ledger", "description": "Cross-phase defect register accumulating stubs, TODOs, skipped tests, unrun verifies, and unmet truths into .planning/WINDOWS.md. When enforcement is enabled, it blocks /gsd-ship while any window is open unless explicitly waived with a recorded reason. Operationalizes GSD's no-defer discipline as a tracked artifact (issue #1950).", "tier": "full", diff --git a/capabilities/claude-orchestration/capability.json b/capabilities/claude-orchestration/capability.json index 02f29e528..059019b31 100644 --- a/capabilities/claude-orchestration/capability.json +++ b/capabilities/claude-orchestration/capability.json @@ -1,7 +1,7 @@ { "id": "claude-orchestration", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Claude orchestration (Workflow backend)", "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", "tier": "full", diff --git a/capabilities/claude/capability.json b/capabilities/claude/capability.json index b73233255..8b6d6d947 100644 --- a/capabilities/claude/capability.json +++ b/capabilities/claude/capability.json @@ -1,7 +1,7 @@ { "id": "claude", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Claude Code", "description": "Anthropic Claude Code — primary development runtime; tier-1 support with full hook surface and skills-based global install.", "tier": "core", diff --git a/capabilities/cline/capability.json b/capabilities/cline/capability.json index b17f19375..6e1e4bef9 100644 --- a/capabilities/cline/capability.json +++ b/capabilities/cline/capability.json @@ -1,7 +1,7 @@ { "id": "cline", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cline", "description": "Cline (VS Code extension) — global-only nested-skill layout; cline-rules hook surface (.clinerules); no hook events emitted; tier-2 support.", "tier": "core", diff --git a/capabilities/code-review/capability.json b/capabilities/code-review/capability.json index 1857a8a07..77c988f9b 100644 --- a/capabilities/code-review/capability.json +++ b/capabilities/code-review/capability.json @@ -1,7 +1,7 @@ { "id": "code-review", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Code review", "description": "Source-file code review and review-fix workflow support for completed execution work.", "tier": "full", diff --git a/capabilities/codebuddy/capability.json b/capabilities/codebuddy/capability.json index 6d43e72e3..c37ed61fb 100644 --- a/capabilities/codebuddy/capability.json +++ b/capabilities/codebuddy/capability.json @@ -1,7 +1,7 @@ { "id": "codebuddy", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "CodeBuddy", "description": "CodeBuddy (Tencent) — converted commands + skills artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", diff --git a/capabilities/coderabbit/capability.json b/capabilities/coderabbit/capability.json index 70c49949a..c1b3801d8 100644 --- a/capabilities/coderabbit/capability.json +++ b/capabilities/coderabbit/capability.json @@ -1,7 +1,7 @@ { "id": "coderabbit", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "CodeRabbit", "description": "CodeRabbit CLI — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). Reviews the working-tree diff (`coderabbit review --prompt-only`), not the source tree, and accepts neither a prompt nor a model flag; findings are down-weighted in consensus (evidenceClass: diff-only).", "tier": "full", diff --git a/capabilities/codex/capability.json b/capabilities/codex/capability.json index 487d1304d..71bcd1f53 100644 --- a/capabilities/codex/capability.json +++ b/capabilities/codex/capability.json @@ -1,7 +1,7 @@ { "id": "codex", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenAI Codex CLI", "description": "OpenAI Codex CLI — shell-var command style; per-agent sandbox tiers; config.toml + hooks.json hook surface; tier-1 support.", "tier": "core", diff --git a/capabilities/copilot/capability.json b/capabilities/copilot/capability.json index 87ccba1a9..e4fd59a81 100644 --- a/capabilities/copilot/capability.json +++ b/capabilities/copilot/capability.json @@ -1,7 +1,7 @@ { "id": "copilot", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "GitHub Copilot", "description": "GitHub Copilot (VS Code) — markdown config format; copilot-inline hook surface; no hook events emitted; flat skill nesting (unconfirmed recursive loader); tier-2 support.", "tier": "core", diff --git a/capabilities/cursor/capability.json b/capabilities/cursor/capability.json index f862d6b55..23f1a92f3 100644 --- a/capabilities/cursor/capability.json +++ b/capabilities/cursor/capability.json @@ -1,7 +1,7 @@ { "id": "cursor", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cursor", "description": "Cursor IDE — skills-only workflow surface; hooks.json surface; Claude hook event dialect; recursive skill loader (flat nesting); tier-2 support.", "tier": "core", diff --git a/capabilities/drift/capability.json b/capabilities/drift/capability.json index fae5dfe7d..49432f59c 100644 --- a/capabilities/drift/capability.json +++ b/capabilities/drift/capability.json @@ -1,7 +1,7 @@ { "id": "drift", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Drift detection gates", "description": "Drift detection gates for the planning loop. At execute:wave:post: a blocking schema drift gate (detects schema files changed without a database push) and a non-blocking codebase drift gate (detects structural additions not reflected in STRUCTURE.md). At plan:pre: a non-blocking, warn-only codebase drift gate (gated on workflow.plan_drift_precheck) that flags a stale codebase map before planning, so plans are authored against a fresh STRUCTURE.md instead of discovering drift mid-execution.", "tier": "full", diff --git a/capabilities/external-job/capability.json b/capabilities/external-job/capability.json index 5a7c12363..8e5529a54 100644 --- a/capabilities/external-job/capability.json +++ b/capabilities/external-job/capability.json @@ -1,7 +1,7 @@ { "id": "external-job", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Async external-job scheduler adapter", "description": "Default-off producer of the async external-job manifest (#1164). At execute:wave:post an executor can externalize long-running compute (SLURM first, scheduler-pluggable), commit a .planning/async-jobs/.json manifest, defer SUMMARY.md, and return external_job_waiting. The core loop (#1165) consumes the manifest; this capability is the only thing that writes it. NOTE on contribution point: #1164 specifies classification at execute:wave:pre and recording at execute:wave:post. This capability still contributes executor guidance at wave:post; execute-phase now renders wave:pre entries and dispatches generic step hooks there independently. Moving external-job classification to wave:pre is a separate capability change, not part of #4148. The adapter (scripts/slurm-adapter.cjs) reads external_job.submit_timeout_ms / poll_timeout_ms / artifact_dir through the canonical capability-config seam (env override > config > registry default).", "tier": "full", diff --git a/capabilities/gap-analysis/capability.json b/capabilities/gap-analysis/capability.json index dbdd16477..62279c644 100644 --- a/capabilities/gap-analysis/capability.json +++ b/capabilities/gap-analysis/capability.json @@ -1,7 +1,7 @@ { "id": "gap-analysis", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Post-planning gap analysis", "description": "Proactive, non-blocking post-planning coverage report. After all PLAN.md files are generated, cross-references every REQ-ID and D-ID from REQUIREMENTS.md and CONTEXT.md against plan bodies. Emits a Source | Item | Status table. Does not block phase advancement.", "tier": "standard", diff --git a/capabilities/gemini/capability.json b/capabilities/gemini/capability.json index af0f4b36c..869a6b4ca 100644 --- a/capabilities/gemini/capability.json +++ b/capabilities/gemini/capability.json @@ -1,7 +1,7 @@ { "id": "gemini", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "Gemini CLI", "description": "Google Gemini CLI — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). Spawned as `gemini -p - -m ` with the plan piped on stdin.", "tier": "full", diff --git a/capabilities/graphify/capability.json b/capabilities/graphify/capability.json index 14ffe61db..512354376 100644 --- a/capabilities/graphify/capability.json +++ b/capabilities/graphify/capability.json @@ -1,7 +1,7 @@ { "id": "graphify", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Knowledge graph", "description": "Build, query, and inspect the project knowledge graph in `.planning/graphs/`; exposes graphify CLI subcommands (build, query, status, diff) and the /gsd-graphify skill.", "tier": "full", diff --git a/capabilities/hermes/capability.json b/capabilities/hermes/capability.json index a80ae03fb..890e12998 100644 --- a/capabilities/hermes/capability.json +++ b/capabilities/hermes/capability.json @@ -1,7 +1,7 @@ { "id": "hermes", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Hermes Agent", "description": "Hermes Agent (NousResearch) — skills nest under skills/gsd/ category bucket; nested skill layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", diff --git a/capabilities/intel/capability.json b/capabilities/intel/capability.json index 94b89cf25..9ea4743a3 100644 --- a/capabilities/intel/capability.json +++ b/capabilities/intel/capability.json @@ -1,7 +1,7 @@ { "id": "intel", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Codebase intelligence", "description": "Code-intelligence store for codebase querying, diff, snapshot, and API-surface extraction; exposes `gsd-tools intel` subcommands (query, status, update, diff, snapshot, patch-meta, validate, extract-exports, api-surface) and backs `/gsd-map-codebase` and `gsd-intel-updater`.", "tier": "full", diff --git a/capabilities/kilo/capability.json b/capabilities/kilo/capability.json index fab009328..7d0a65abe 100644 --- a/capabilities/kilo/capability.json +++ b/capabilities/kilo/capability.json @@ -1,7 +1,7 @@ { "id": "kilo", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kilo Code", "description": "Kilo Code — XDG-based config dir; global skills at ~/.kilo/skills (separate from XDG config); flat command/ + skills artifact layout; no lifecycle hook registration; tier-2 support.", "tier": "core", diff --git a/capabilities/kimi-code/capability.json b/capabilities/kimi-code/capability.json index d313371d6..55f610552 100644 --- a/capabilities/kimi-code/capability.json +++ b/capabilities/kimi-code/capability.json @@ -1,7 +1,7 @@ { "id": "kimi-code", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi Code CLI", "description": "Kimi Code CLI (Moonshot AI, Node) — Agent Skills auto-discovered at ~/.kimi-code/skills; global AGENTS.md at ~/.kimi-code/AGENTS.md; native config.toml + [[hooks]] bus; three built-in subagents (coder/explore/plan), NO custom named subagents; background dispatch; tier-2 support. Distinct from Python kimi-cli (the 'kimi' capability) per ADR-1239 EoS — Kimi Code cannot dispatch named subagents so the kimi-agents YAML layout does NOT apply; persona injection rides the existing ${AGENT_SKILLS_*} workflow fallback. Install-layout, agent-install-check, and install-time decision (kimi vs kimi-code) land in follow-up PRs; this descriptor is the EoS foundation.", "tier": "core", diff --git a/capabilities/kimi/capability.json b/capabilities/kimi/capability.json index f2f07c550..c19387e89 100644 --- a/capabilities/kimi/capability.json +++ b/capabilities/kimi/capability.json @@ -1,7 +1,7 @@ { "id": "kimi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi CLI", "description": "Kimi CLI (Moonshot AI) — generic agents root at ~/.config/agents; skills + kimi-agents artifact layout; native config.toml [[hooks]] bus at ~/.kimi/config.toml; background dispatch; tier-2 support.", "tier": "core", diff --git a/capabilities/live-dom-uat/capability.json b/capabilities/live-dom-uat/capability.json index 272c21742..5332493a4 100644 --- a/capabilities/live-dom-uat/capability.json +++ b/capabilities/live-dom-uat/capability.json @@ -1,7 +1,7 @@ { "id": "live-dom-uat", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Live-DOM UAT", "description": "Default-off live-DOM verification (#2856). Confines browser MCP reach to one purpose-built agent (gsd-dom-verifier) that carries the browser globs in its own tools: line, registered as an additive step hook at execute:wave:post. agents/gsd-executor.md is deliberately NOT widened: for a first-party agent the static tool list is the only control that exists, no capability can grant tools to one (ADR-1244 D2), no hook kind grants tool permissions (ADR-857 D4), and there is no per-dispatch tool override. Gated by activationKey workflow.live_dom_uat (default false), so with the key off the capability resolves inactive and the hook does not render at all. NOTE on the browser profile lock: chrome-devtools-mcp holds an exclusive lock on $HOME/.cache/chrome-devtools-mcp/chrome-profile, and --isolated is a flag on the user's own MCP-server registration that GSD cannot pass. Concurrent execution waves sharing one profile will therefore collide; the step tolerates and reports that (onError: skip, never blocking) rather than pretending to coordinate a resource it does not own.", "tier": "full", diff --git a/capabilities/llama-cpp/capability.json b/capabilities/llama-cpp/capability.json index 9f8c03bbb..419b901bb 100644 --- a/capabilities/llama-cpp/capability.json +++ b/capabilities/llama-cpp/capability.json @@ -1,7 +1,7 @@ { "id": "llama-cpp", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "llama.cpp", "description": "llama.cpp server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.llama_cpp_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`llama-cpp`, required by KEBAB_RE); `reviewer.slug` stays snake (`llama_cpp`) to match the shipped roster and the `review.llama_cpp_host` config key (ADR-2782's three-namespace trap).", "tier": "full", diff --git a/capabilities/lm-studio/capability.json b/capabilities/lm-studio/capability.json index c2ff8647c..fe43cfb67 100644 --- a/capabilities/lm-studio/capability.json +++ b/capabilities/lm-studio/capability.json @@ -1,7 +1,7 @@ { "id": "lm-studio", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "LM Studio", "description": "LM Studio local model server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.lm_studio_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`lm-studio`, required by KEBAB_RE); `reviewer.slug` stays snake (`lm_studio`) to match the shipped roster and the `review.lm_studio_host` config key (ADR-2782's three-namespace trap).", "tier": "full", diff --git a/capabilities/mempalace/capability.json b/capabilities/mempalace/capability.json index 8c929be0f..9741939cc 100644 --- a/capabilities/mempalace/capability.json +++ b/capabilities/mempalace/capability.json @@ -1,7 +1,7 @@ { "id": "mempalace", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "MemPalace memory", "description": "Cross-session, cross-project memory: deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries, via the MemPalace MCP server and CLI.", "tier": "full", diff --git a/capabilities/nyquist/capability.json b/capabilities/nyquist/capability.json index c8c2a283c..0fdc80f60 100644 --- a/capabilities/nyquist/capability.json +++ b/capabilities/nyquist/capability.json @@ -1,7 +1,7 @@ { "id": "nyquist", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Nyquist validation", "description": "Validation coverage audit that maps executed work back to tests and manual-only evidence.", "tier": "full", diff --git a/capabilities/ollama/capability.json b/capabilities/ollama/capability.json index 460c0cf75..8444f2ea2 100644 --- a/capabilities/ollama/capability.json +++ b/capabilities/ollama/capability.json @@ -1,7 +1,7 @@ { "id": "ollama", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "Ollama", "description": "Ollama local model server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.ollama_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq.", "tier": "full", diff --git a/capabilities/opencode/capability.json b/capabilities/opencode/capability.json index 22e2ed158..3930324de 100644 --- a/capabilities/opencode/capability.json +++ b/capabilities/opencode/capability.json @@ -1,7 +1,7 @@ { "id": "opencode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenCode", "description": "OpenCode — XDG-based config dir; flat commands/ + skills artifact layout; settings-json config format; no lifecycle hook registration; tier-2 support.", "tier": "core", diff --git a/capabilities/pattern-mapper/capability.json b/capabilities/pattern-mapper/capability.json index 54ff772b5..670797b69 100644 --- a/capabilities/pattern-mapper/capability.json +++ b/capabilities/pattern-mapper/capability.json @@ -1,7 +1,7 @@ { "id": "pattern-mapper", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Pattern mapping", "description": "Optional codebase-pattern mapping before planning; owns the pattern mapper agent and workflow.pattern_mapper activation key.", "tier": "full", diff --git a/capabilities/pi/capability.json b/capabilities/pi/capability.json index e2ea1d30f..543129846 100644 --- a/capabilities/pi/capability.json +++ b/capabilities/pi/capability.json @@ -1,7 +1,7 @@ { "id": "pi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "pi", "description": "pi (pi.dev) — bun-runtime programmatic-CLI; TS ExtensionAPI (registerCommand/registerTool/registerProvider/pi.on); single native-extension file at ~/.pi/agent/extensions/gsd.js (.js, not .cjs — pi's extension auto-discovery accepts only .ts/.js, #2470); no shared-settings hook surface; tier-2 support.", "tier": "core", diff --git a/capabilities/profile-pipeline/capability.json b/capabilities/profile-pipeline/capability.json index d0fdafa35..55d5a796f 100644 --- a/capabilities/profile-pipeline/capability.json +++ b/capabilities/profile-pipeline/capability.json @@ -1,7 +1,7 @@ { "id": "profile-pipeline", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Developer profiling pipeline", "description": "Developer behavioral profiling from Claude Code session history; scans session JSONL files, extracts and samples user messages, and generates profile artifacts (USER-PROFILE.md, dev-preferences.md, CLAUDE.md sections). Exposes eight `gsd-tools` commands: scan-sessions, extract-messages, profile-sample (pipeline phase) and write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md (output phase). Backs the /gsd-profile-user skill and gsd-user-profiler agent.", "tier": "full", diff --git a/capabilities/qwen/capability.json b/capabilities/qwen/capability.json index 07ef831c6..75eb5fd52 100644 --- a/capabilities/qwen/capability.json +++ b/capabilities/qwen/capability.json @@ -1,7 +1,7 @@ { "id": "qwen", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Qwen Code", "description": "Qwen Code (Alibaba) — nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", diff --git a/capabilities/refactor-trigger/capability.json b/capabilities/refactor-trigger/capability.json index 72deb19c7..3422cbabe 100644 --- a/capabilities/refactor-trigger/capability.json +++ b/capabilities/refactor-trigger/capability.json @@ -1,7 +1,7 @@ { "id": "refactor-trigger", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Complexity-triggered refactor", "description": "Measures the complexity of the code a phase touched and, when a function crosses a configured threshold or jumps past its recorded anchor, surfaces a scoped refactor proposal at .planning/phases//-REFACTOR.md. Advisory by default — it never edits code and never blocks. Opt-in strict mode blocks /gsd-ship while a proposal is untriaged; a declined proposal is recorded in the broken-windows ledger when that capability is present. Operationalizes 'refactor early, refactor often' as continuous pressure instead of a thing you have to remember (issue #1953).", "tier": "full", diff --git a/capabilities/research/capability.json b/capabilities/research/capability.json index 20b3a983f..9162e9331 100644 --- a/capabilities/research/capability.json +++ b/capabilities/research/capability.json @@ -1,7 +1,7 @@ { "id": "research", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Phase research", "description": "Optional phase research before planning; owns the phase researcher agent and workflow.research activation key.", "tier": "standard", diff --git a/capabilities/schema-gate/capability.json b/capabilities/schema-gate/capability.json index ed8f7ff77..c6ea90908 100644 --- a/capabilities/schema-gate/capability.json +++ b/capabilities/schema-gate/capability.json @@ -1,7 +1,7 @@ { "id": "schema-gate", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Schema push detection gate", "description": "Detects ORM schema-relevant files in the phase scope during planning and injects a mandatory [BLOCKING] schema push task into the plan. Prevents false-positive verification where build/types pass because TypeScript types come from config, not the live database.", "tier": "full", diff --git a/capabilities/security/capability.json b/capabilities/security/capability.json index 7aa4abd4b..a79560dce 100644 --- a/capabilities/security/capability.json +++ b/capabilities/security/capability.json @@ -1,7 +1,7 @@ { "id": "security", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Security enforcement", "description": "Threat mitigation verification and ship-time security blocking for phases with security enforcement enabled.", "tier": "full", diff --git a/capabilities/tdd/capability.json b/capabilities/tdd/capability.json index 3dba7e1d7..6296508dc 100644 --- a/capabilities/tdd/capability.json +++ b/capabilities/tdd/capability.json @@ -1,7 +1,7 @@ { "id": "tdd", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Test-driven development", "description": "Injects TDD heuristics into the planner and enforces RED/GREEN gate compliance on type:tdd plans after execution. Owns workflow.tdd_mode; the --tdd CLI flag is the ephemeral override.", "tier": "full", diff --git a/capabilities/trae/capability.json b/capabilities/trae/capability.json index 7b52deee8..1aa36ba33 100644 --- a/capabilities/trae/capability.json +++ b/capabilities/trae/capability.json @@ -1,7 +1,7 @@ { "id": "trae", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Trae IDE", "description": "Trae IDE — nested-skill artifact layout; no hook surface (profile-marker-only config); tier-2 support.", "tier": "core", diff --git a/capabilities/ui/capability.json b/capabilities/ui/capability.json index b42e610cb..c84639e9a 100644 --- a/capabilities/ui/capability.json +++ b/capabilities/ui/capability.json @@ -1,7 +1,7 @@ { "id": "ui", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "UI design contracts", "description": "UI-SPEC design contract + retrospective UI audit for frontend phases.", "tier": "full", diff --git a/capabilities/vscode/capability.json b/capabilities/vscode/capability.json index 2e4c23d25..1197fa4d5 100644 --- a/capabilities/vscode/capability.json +++ b/capabilities/vscode/capability.json @@ -1,7 +1,7 @@ { "id": "vscode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "VS Code", "description": "VS Code — Marketplace/VSIX extension; no file-projected config directory; IDE-profile reference host (active vscode.lm model, engine-owned hook bus, sandboxed globalState/workspaceState stateIO).", "tier": "core", diff --git a/capabilities/windsurf/capability.json b/capabilities/windsurf/capability.json index 494a7794e..de376d8da 100644 --- a/capabilities/windsurf/capability.json +++ b/capabilities/windsurf/capability.json @@ -1,7 +1,7 @@ { "id": "windsurf", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Windsurf", "description": "Windsurf (Codeium) — workspace workflow artifact layout for slash commands; Cascade native hooks.json blocking hook bus (pre_write_code, pre_run_command); tier-2 support.", "tier": "core", diff --git a/capabilities/zcode/capability.json b/capabilities/zcode/capability.json index 511027d11..32514694e 100644 --- a/capabilities/zcode/capability.json +++ b/capabilities/zcode/capability.json @@ -1,7 +1,7 @@ { "id": "zcode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "ZCode", "description": "ZCode (Z.ai) — desktop Agentic Development Environment for GLM-5.2; Claude-shaped nested skills at ~/.zcode/skills//SKILL.md, slash commands, named subagents, native MCP; declarative plugin surface; profile-marker install; tier-2 community support.", "tier": "core", diff --git a/gsd-core/bin/lib/capability-registry.cjs b/gsd-core/bin/lib/capability-registry.cjs index 338372f8c..05c338919 100644 --- a/gsd-core/bin/lib/capability-registry.cjs +++ b/gsd-core/bin/lib/capability-registry.cjs @@ -10,7 +10,7 @@ const capabilities = { "ai-integration": { "id": "ai-integration", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "AI design contract", "description": "AI-SPEC design contract workflow for phases that build AI systems; owns the AI integration command, agents, and workflow.ai_integration_phase activation key.", "tier": "full", @@ -95,7 +95,7 @@ const capabilities = { "antigravity": { "id": "antigravity", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Antigravity", "description": "Google Antigravity IDE — config/settings home nested under ~/.gemini/antigravity (probed across 1.x and 2.x layouts); global skills/agents install under ~/.gemini/config, the dir AGY scans for global discovery (#3738); Gemini hook event dialect; flat skill layout; tier-1 support.", "tier": "core", @@ -258,7 +258,7 @@ const capabilities = { "assumption-delta": { "id": "assumption-delta", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Assumption-delta architecture checkpoint", "description": "Rarely-firing advisory checkpoint that triggers when a phase makes something plural, optional, or chosen that used to be singular, required, or derived. Surfaces one identity-model question (promote the new general representation to primary, or add it alongside?) so a silent primary-key drift does not accumulate into a later user-facing bug. Non-blocking; fires only on a detected signal.", "tier": "full", @@ -304,7 +304,7 @@ const capabilities = { "audit": { "id": "audit", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Audit", "description": "Open-artifact audit and UAT-gap audit for milestone close gates; exposes `gsd-tools audit-uat` (cross-phase UAT outstanding items) and `gsd-tools audit-open` (structured open-artifact scan across debug, tasks, threads, todos, seeds, UAT, verification, context-questions).", "tier": "full", @@ -341,7 +341,7 @@ const capabilities = { "augment": { "id": "augment", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Augment Code", "description": "Augment Code CLI — commands + nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -455,7 +455,7 @@ const capabilities = { "broken-windows": { "id": "broken-windows", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Broken-windows ledger", "description": "Cross-phase defect register accumulating stubs, TODOs, skipped tests, unrun verifies, and unmet truths into .planning/WINDOWS.md. When enforcement is enabled, it blocks /gsd-ship while any window is open unless explicitly waived with a recorded reason. Operationalizes GSD's no-defer discipline as a tracked artifact (issue #1950).", "tier": "full", @@ -501,7 +501,7 @@ const capabilities = { "claude": { "id": "claude", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Claude Code", "description": "Anthropic Claude Code — primary development runtime; tier-1 support with full hook surface and skills-based global install.", "tier": "core", @@ -682,7 +682,7 @@ const capabilities = { "claude-orchestration": { "id": "claude-orchestration", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Claude orchestration (Workflow backend)", "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", "tier": "full", @@ -770,7 +770,7 @@ const capabilities = { "cline": { "id": "cline", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cline", "description": "Cline (VS Code extension) — global-only nested-skill layout; cline-rules hook surface (.clinerules); no hook events emitted; tier-2 support.", "tier": "core", @@ -863,7 +863,7 @@ const capabilities = { "code-review": { "id": "code-review", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Code review", "description": "Source-file code review and review-fix workflow support for completed execution work.", "tier": "full", @@ -947,7 +947,7 @@ const capabilities = { "codebuddy": { "id": "codebuddy", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "CodeBuddy", "description": "CodeBuddy (Tencent) — converted commands + skills artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -1065,7 +1065,7 @@ const capabilities = { "coderabbit": { "id": "coderabbit", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "CodeRabbit", "description": "CodeRabbit CLI — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). Reviews the working-tree diff (`coderabbit review --prompt-only`), not the source tree, and accepts neither a prompt nor a model flag; findings are down-weighted in consensus (evidenceClass: diff-only).", "tier": "full", @@ -1117,7 +1117,7 @@ const capabilities = { "codex": { "id": "codex", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenAI Codex CLI", "description": "OpenAI Codex CLI — shell-var command style; per-agent sandbox tiers; config.toml + hooks.json hook surface; tier-1 support.", "tier": "core", @@ -1293,7 +1293,7 @@ const capabilities = { "copilot": { "id": "copilot", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "GitHub Copilot", "description": "GitHub Copilot (VS Code) — markdown config format; copilot-inline hook surface; no hook events emitted; flat skill nesting (unconfirmed recursive loader); tier-2 support.", "tier": "core", @@ -1393,7 +1393,7 @@ const capabilities = { "cursor": { "id": "cursor", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cursor", "description": "Cursor IDE — skills-only workflow surface; hooks.json surface; Claude hook event dialect; recursive skill loader (flat nesting); tier-2 support.", "tier": "core", @@ -1562,7 +1562,7 @@ const capabilities = { "drift": { "id": "drift", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Drift detection gates", "description": "Drift detection gates for the planning loop. At execute:wave:post: a blocking schema drift gate (detects schema files changed without a database push) and a non-blocking codebase drift gate (detects structural additions not reflected in STRUCTURE.md). At plan:pre: a non-blocking, warn-only codebase drift gate (gated on workflow.plan_drift_precheck) that flags a stale codebase map before planning, so plans are authored against a fresh STRUCTURE.md instead of discovering drift mid-execution.", "tier": "full", @@ -1663,7 +1663,7 @@ const capabilities = { "external-job": { "id": "external-job", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Async external-job scheduler adapter", "description": "Default-off producer of the async external-job manifest (#1164). At execute:wave:post an executor can externalize long-running compute (SLURM first, scheduler-pluggable), commit a .planning/async-jobs/.json manifest, defer SUMMARY.md, and return external_job_waiting. The core loop (#1165) consumes the manifest; this capability is the only thing that writes it. NOTE on contribution point: #1164 specifies classification at execute:wave:pre and recording at execute:wave:post. This capability still contributes executor guidance at wave:post; execute-phase now renders wave:pre entries and dispatches generic step hooks there independently. Moving external-job classification to wave:pre is a separate capability change, not part of #4148. The adapter (scripts/slurm-adapter.cjs) reads external_job.submit_timeout_ms / poll_timeout_ms / artifact_dir through the canonical capability-config seam (env override > config > registry default).", "tier": "full", @@ -1746,7 +1746,7 @@ const capabilities = { "gap-analysis": { "id": "gap-analysis", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Post-planning gap analysis", "description": "Proactive, non-blocking post-planning coverage report. After all PLAN.md files are generated, cross-references every REQ-ID and D-ID from REQUIREMENTS.md and CONTEXT.md against plan bodies. Emits a Source | Item | Status table. Does not block phase advancement.", "tier": "standard", @@ -1787,7 +1787,7 @@ const capabilities = { "gemini": { "id": "gemini", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "Gemini CLI", "description": "Google Gemini CLI — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). Spawned as `gemini -p - -m ` with the plan piped on stdin.", "tier": "full", @@ -1850,7 +1850,7 @@ const capabilities = { "graphify": { "id": "graphify", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Knowledge graph", "description": "Build, query, and inspect the project knowledge graph in `.planning/graphs/`; exposes graphify CLI subcommands (build, query, status, diff) and the /gsd-graphify skill.", "tier": "full", @@ -1891,7 +1891,7 @@ const capabilities = { "hermes": { "id": "hermes", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Hermes Agent", "description": "Hermes Agent (NousResearch) — skills nest under skills/gsd/ category bucket; nested skill layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -2003,7 +2003,7 @@ const capabilities = { "intel": { "id": "intel", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Codebase intelligence", "description": "Code-intelligence store for codebase querying, diff, snapshot, and API-surface extraction; exposes `gsd-tools intel` subcommands (query, status, update, diff, snapshot, patch-meta, validate, extract-exports, api-surface) and backs `/gsd-map-codebase` and `gsd-intel-updater`.", "tier": "full", @@ -2055,7 +2055,7 @@ const capabilities = { "kilo": { "id": "kilo", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kilo Code", "description": "Kilo Code — XDG-based config dir; global skills at ~/.kilo/skills (separate from XDG config); flat command/ + skills artifact layout; no lifecycle hook registration; tier-2 support.", "tier": "core", @@ -2185,7 +2185,7 @@ const capabilities = { "kimi": { "id": "kimi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi CLI", "description": "Kimi CLI (Moonshot AI) — generic agents root at ~/.config/agents; skills + kimi-agents artifact layout; native config.toml [[hooks]] bus at ~/.kimi/config.toml; background dispatch; tier-2 support.", "tier": "core", @@ -2289,7 +2289,7 @@ const capabilities = { "kimi-code": { "id": "kimi-code", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi Code CLI", "description": "Kimi Code CLI (Moonshot AI, Node) — Agent Skills auto-discovered at ~/.kimi-code/skills; global AGENTS.md at ~/.kimi-code/AGENTS.md; native config.toml + [[hooks]] bus; three built-in subagents (coder/explore/plan), NO custom named subagents; background dispatch; tier-2 support. Distinct from Python kimi-cli (the 'kimi' capability) per ADR-1239 EoS — Kimi Code cannot dispatch named subagents so the kimi-agents YAML layout does NOT apply; persona injection rides the existing ${AGENT_SKILLS_*} workflow fallback. Install-layout, agent-install-check, and install-time decision (kimi vs kimi-code) land in follow-up PRs; this descriptor is the EoS foundation.", "tier": "core", @@ -2456,7 +2456,7 @@ const capabilities = { "live-dom-uat": { "id": "live-dom-uat", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Live-DOM UAT", "description": "Default-off live-DOM verification (#2856). Confines browser MCP reach to one purpose-built agent (gsd-dom-verifier) that carries the browser globs in its own tools: line, registered as an additive step hook at execute:wave:post. agents/gsd-executor.md is deliberately NOT widened: for a first-party agent the static tool list is the only control that exists, no capability can grant tools to one (ADR-1244 D2), no hook kind grants tool permissions (ADR-857 D4), and there is no per-dispatch tool override. Gated by activationKey workflow.live_dom_uat (default false), so with the key off the capability resolves inactive and the hook does not render at all. NOTE on the browser profile lock: chrome-devtools-mcp holds an exclusive lock on $HOME/.cache/chrome-devtools-mcp/chrome-profile, and --isolated is a flag on the user's own MCP-server registration that GSD cannot pass. Concurrent execution waves sharing one profile will therefore collide; the step tolerates and reports that (onError: skip, never blocking) rather than pretending to coordinate a resource it does not own.", "tier": "full", @@ -2509,7 +2509,7 @@ const capabilities = { "llama-cpp": { "id": "llama-cpp", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "llama.cpp", "description": "llama.cpp server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.llama_cpp_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`llama-cpp`, required by KEBAB_RE); `reviewer.slug` stays snake (`llama_cpp`) to match the shipped roster and the `review.llama_cpp_host` config key (ADR-2782's three-namespace trap).", "tier": "full", @@ -2575,7 +2575,7 @@ const capabilities = { "lm-studio": { "id": "lm-studio", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "LM Studio", "description": "LM Studio local model server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.lm_studio_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`lm-studio`, required by KEBAB_RE); `reviewer.slug` stays snake (`lm_studio`) to match the shipped roster and the `review.lm_studio_host` config key (ADR-2782's three-namespace trap).", "tier": "full", @@ -2641,7 +2641,7 @@ const capabilities = { "mempalace": { "id": "mempalace", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "MemPalace memory", "description": "Cross-session, cross-project memory: deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries, via the MemPalace MCP server and CLI.", "tier": "full", @@ -2815,7 +2815,7 @@ const capabilities = { "nyquist": { "id": "nyquist", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Nyquist validation", "description": "Validation coverage audit that maps executed work back to tests and manual-only evidence.", "tier": "full", @@ -2865,7 +2865,7 @@ const capabilities = { "ollama": { "id": "ollama", "role": "reviewer", - "version": "1.12.0", + "version": "1.13.0", "title": "Ollama", "description": "Ollama local model server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.ollama_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq.", "tier": "full", @@ -2931,7 +2931,7 @@ const capabilities = { "opencode": { "id": "opencode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenCode", "description": "OpenCode — XDG-based config dir; flat commands/ + skills artifact layout; settings-json config format; no lifecycle hook registration; tier-2 support.", "tier": "core", @@ -3126,7 +3126,7 @@ const capabilities = { "pattern-mapper": { "id": "pattern-mapper", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Pattern mapping", "description": "Optional codebase-pattern mapping before planning; owns the pattern mapper agent and workflow.pattern_mapper activation key.", "tier": "full", @@ -3180,7 +3180,7 @@ const capabilities = { "pi": { "id": "pi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "pi", "description": "pi (pi.dev) — bun-runtime programmatic-CLI; TS ExtensionAPI (registerCommand/registerTool/registerProvider/pi.on); single native-extension file at ~/.pi/agent/extensions/gsd.js (.js, not .cjs — pi's extension auto-discovery accepts only .ts/.js, #2470); no shared-settings hook surface; tier-2 support.", "tier": "core", @@ -3250,7 +3250,7 @@ const capabilities = { "profile-pipeline": { "id": "profile-pipeline", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Developer profiling pipeline", "description": "Developer behavioral profiling from Claude Code session history; scans session JSONL files, extracts and samples user messages, and generates profile artifacts (USER-PROFILE.md, dev-preferences.md, CLAUDE.md sections). Exposes eight `gsd-tools` commands: scan-sessions, extract-messages, profile-sample (pipeline phase) and write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md (output phase). Backs the /gsd-profile-user skill and gsd-user-profiler agent.", "tier": "full", @@ -3327,7 +3327,7 @@ const capabilities = { "qwen": { "id": "qwen", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Qwen Code", "description": "Qwen Code (Alibaba) — nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -3477,7 +3477,7 @@ const capabilities = { "refactor-trigger": { "id": "refactor-trigger", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Complexity-triggered refactor", "description": "Measures the complexity of the code a phase touched and, when a function crosses a configured threshold or jumps past its recorded anchor, surfaces a scoped refactor proposal at .planning/phases//-REFACTOR.md. Advisory by default — it never edits code and never blocks. Opt-in strict mode blocks /gsd-ship while a proposal is untriaged; a declined proposal is recorded in the broken-windows ledger when that capability is present. Operationalizes 'refactor early, refactor often' as continuous pressure instead of a thing you have to remember (issue #1953).", "tier": "full", @@ -3544,7 +3544,7 @@ const capabilities = { "research": { "id": "research", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Phase research", "description": "Optional phase research before planning; owns the phase researcher agent and workflow.research activation key.", "tier": "standard", @@ -3596,7 +3596,7 @@ const capabilities = { "schema-gate": { "id": "schema-gate", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Schema push detection gate", "description": "Detects ORM schema-relevant files in the phase scope during planning and injects a mandatory [BLOCKING] schema push task into the plan. Prevents false-positive verification where build/types pass because TypeScript types come from config, not the live database.", "tier": "full", @@ -3642,7 +3642,7 @@ const capabilities = { "security": { "id": "security", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Security enforcement", "description": "Threat mitigation verification and ship-time security blocking for phases with security enforcement enabled.", "tier": "full", @@ -3741,7 +3741,7 @@ const capabilities = { "tdd": { "id": "tdd", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "Test-driven development", "description": "Injects TDD heuristics into the planner and enforces RED/GREEN gate compliance on type:tdd plans after execution. Owns workflow.tdd_mode; the --tdd CLI flag is the ephemeral override.", "tier": "full", @@ -3794,7 +3794,7 @@ const capabilities = { "trae": { "id": "trae", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Trae IDE", "description": "Trae IDE — nested-skill artifact layout; no hook surface (profile-marker-only config); tier-2 support.", "tier": "core", @@ -3892,7 +3892,7 @@ const capabilities = { "ui": { "id": "ui", "role": "feature", - "version": "1.12.0", + "version": "1.13.0", "title": "UI design contracts", "description": "UI-SPEC design contract + retrospective UI audit for frontend phases.", "tier": "full", @@ -3987,7 +3987,7 @@ const capabilities = { "vscode": { "id": "vscode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "VS Code", "description": "VS Code — Marketplace/VSIX extension; no file-projected config directory; IDE-profile reference host (active vscode.lm model, engine-owned hook bus, sandboxed globalState/workspaceState stateIO).", "tier": "core", @@ -4045,7 +4045,7 @@ const capabilities = { "windsurf": { "id": "windsurf", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Windsurf", "description": "Windsurf (Codeium) — workspace workflow artifact layout for slash commands; Cascade native hooks.json blocking hook bus (pre_write_code, pre_run_command); tier-2 support.", "tier": "core", @@ -4137,7 +4137,7 @@ const capabilities = { "zcode": { "id": "zcode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "ZCode", "description": "ZCode (Z.ai) — desktop Agentic Development Environment for GLM-5.2; Claude-shaped nested skills at ~/.zcode/skills//SKILL.md, slash commands, named subagents, native MCP; declarative plugin surface; profile-marker install; tier-2 community support.", "tier": "core", @@ -5560,7 +5560,7 @@ const runtimes = { "antigravity": { "id": "antigravity", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Antigravity", "description": "Google Antigravity IDE — config/settings home nested under ~/.gemini/antigravity (probed across 1.x and 2.x layouts); global skills/agents install under ~/.gemini/config, the dir AGY scans for global discovery (#3738); Gemini hook event dialect; flat skill layout; tier-1 support.", "tier": "core", @@ -5723,7 +5723,7 @@ const runtimes = { "augment": { "id": "augment", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Augment Code", "description": "Augment Code CLI — commands + nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -5837,7 +5837,7 @@ const runtimes = { "claude": { "id": "claude", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Claude Code", "description": "Anthropic Claude Code — primary development runtime; tier-1 support with full hook surface and skills-based global install.", "tier": "core", @@ -6018,7 +6018,7 @@ const runtimes = { "cline": { "id": "cline", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cline", "description": "Cline (VS Code extension) — global-only nested-skill layout; cline-rules hook surface (.clinerules); no hook events emitted; tier-2 support.", "tier": "core", @@ -6111,7 +6111,7 @@ const runtimes = { "codebuddy": { "id": "codebuddy", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "CodeBuddy", "description": "CodeBuddy (Tencent) — converted commands + skills artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -6229,7 +6229,7 @@ const runtimes = { "codex": { "id": "codex", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenAI Codex CLI", "description": "OpenAI Codex CLI — shell-var command style; per-agent sandbox tiers; config.toml + hooks.json hook surface; tier-1 support.", "tier": "core", @@ -6405,7 +6405,7 @@ const runtimes = { "copilot": { "id": "copilot", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "GitHub Copilot", "description": "GitHub Copilot (VS Code) — markdown config format; copilot-inline hook surface; no hook events emitted; flat skill nesting (unconfirmed recursive loader); tier-2 support.", "tier": "core", @@ -6505,7 +6505,7 @@ const runtimes = { "cursor": { "id": "cursor", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Cursor", "description": "Cursor IDE — skills-only workflow surface; hooks.json surface; Claude hook event dialect; recursive skill loader (flat nesting); tier-2 support.", "tier": "core", @@ -6674,7 +6674,7 @@ const runtimes = { "hermes": { "id": "hermes", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Hermes Agent", "description": "Hermes Agent (NousResearch) — skills nest under skills/gsd/ category bucket; nested skill layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -6786,7 +6786,7 @@ const runtimes = { "kilo": { "id": "kilo", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kilo Code", "description": "Kilo Code — XDG-based config dir; global skills at ~/.kilo/skills (separate from XDG config); flat command/ + skills artifact layout; no lifecycle hook registration; tier-2 support.", "tier": "core", @@ -6916,7 +6916,7 @@ const runtimes = { "kimi": { "id": "kimi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi CLI", "description": "Kimi CLI (Moonshot AI) — generic agents root at ~/.config/agents; skills + kimi-agents artifact layout; native config.toml [[hooks]] bus at ~/.kimi/config.toml; background dispatch; tier-2 support.", "tier": "core", @@ -7020,7 +7020,7 @@ const runtimes = { "kimi-code": { "id": "kimi-code", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Kimi Code CLI", "description": "Kimi Code CLI (Moonshot AI, Node) — Agent Skills auto-discovered at ~/.kimi-code/skills; global AGENTS.md at ~/.kimi-code/AGENTS.md; native config.toml + [[hooks]] bus; three built-in subagents (coder/explore/plan), NO custom named subagents; background dispatch; tier-2 support. Distinct from Python kimi-cli (the 'kimi' capability) per ADR-1239 EoS — Kimi Code cannot dispatch named subagents so the kimi-agents YAML layout does NOT apply; persona injection rides the existing ${AGENT_SKILLS_*} workflow fallback. Install-layout, agent-install-check, and install-time decision (kimi vs kimi-code) land in follow-up PRs; this descriptor is the EoS foundation.", "tier": "core", @@ -7187,7 +7187,7 @@ const runtimes = { "opencode": { "id": "opencode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "OpenCode", "description": "OpenCode — XDG-based config dir; flat commands/ + skills artifact layout; settings-json config format; no lifecycle hook registration; tier-2 support.", "tier": "core", @@ -7382,7 +7382,7 @@ const runtimes = { "pi": { "id": "pi", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "pi", "description": "pi (pi.dev) — bun-runtime programmatic-CLI; TS ExtensionAPI (registerCommand/registerTool/registerProvider/pi.on); single native-extension file at ~/.pi/agent/extensions/gsd.js (.js, not .cjs — pi's extension auto-discovery accepts only .ts/.js, #2470); no shared-settings hook surface; tier-2 support.", "tier": "core", @@ -7452,7 +7452,7 @@ const runtimes = { "qwen": { "id": "qwen", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Qwen Code", "description": "Qwen Code (Alibaba) — nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.", "tier": "core", @@ -7602,7 +7602,7 @@ const runtimes = { "trae": { "id": "trae", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Trae IDE", "description": "Trae IDE — nested-skill artifact layout; no hook surface (profile-marker-only config); tier-2 support.", "tier": "core", @@ -7700,7 +7700,7 @@ const runtimes = { "vscode": { "id": "vscode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "VS Code", "description": "VS Code — Marketplace/VSIX extension; no file-projected config directory; IDE-profile reference host (active vscode.lm model, engine-owned hook bus, sandboxed globalState/workspaceState stateIO).", "tier": "core", @@ -7758,7 +7758,7 @@ const runtimes = { "windsurf": { "id": "windsurf", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "Windsurf", "description": "Windsurf (Codeium) — workspace workflow artifact layout for slash commands; Cascade native hooks.json blocking hook bus (pre_write_code, pre_run_command); tier-2 support.", "tier": "core", @@ -7850,7 +7850,7 @@ const runtimes = { "zcode": { "id": "zcode", "role": "runtime", - "version": "1.12.0", + "version": "1.13.0", "title": "ZCode", "description": "ZCode (Z.ai) — desktop Agentic Development Environment for GLM-5.2; Claude-shaped nested skills at ~/.zcode/skills//SKILL.md, slash commands, named subagents, native MCP; declarative plugin surface; profile-marker install; tier-2 community support.", "tier": "core", diff --git a/package-lock.json b/package-lock.json index 308587f54..0a611180a 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "@opengsd/gsd-core", - "version": "1.12.0", + "version": "1.13.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@opengsd/gsd-core", - "version": "1.12.0", + "version": "1.13.0", "license": "MIT", "dependencies": { "@anthropic-ai/claude-agent-sdk": "^0.2.84", diff --git a/package.json b/package.json index f999f7c20..58eadab09 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@opengsd/gsd-core", - "version": "1.12.0", + "version": "1.13.0", "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.", "main": ".opencode/plugins/gsd-core.js", "bin": { diff --git a/vscode/package.json b/vscode/package.json index 8f95b2979..bf4de8f4d 100644 --- a/vscode/package.json +++ b/vscode/package.json @@ -2,7 +2,7 @@ "name": "gsd-core-vscode", "displayName": "GSD Core", "description": "GSD orchestration engine embedded in VS Code (ADR-1239 IDE profile).", - "version": "1.12.0", + "version": "1.13.0", "publisher": "opengsd", "engines": { "vscode": "^1.105.0" From e6d047decc23dfc802a0ffcc032d4a23b863bf7f Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sat, 5 Sep 2026 23:02:07 -0400 Subject: [PATCH 003/166] fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment * fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/): P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch that owns the denominator. A completed phase whose verification routes stale (#2348 clean-commit-time drift) or is missing no longer under-counts forever. P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on a measured scan: totals derived both directions (#2440), completed counters up-only (#2969 — the schema-declared progress-ratchet, now enforced on the write path like the read path always has), percent recomputed from the merged counters. The #3756 unmeasured guard and the #3242 explicit-progress contract are unchanged. P3 — phase complete's atomic 3-file commit passes the post-completion ROADMAP-derived counters through the #2736 authoritativeFm seam (new object direction for the progress key; completedOnlyRaise at the post-preservation re-assert), because the transaction's disk scan reads the pre-completion ROADMAP and failed to increment on the completing phase's own write. * fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites The pre-preservation merge could lower a correctly-higher disk-derived counter (a verification-passed phase whose ROADMAP table row drifted behind the disk signal). completedOnlyRaise now governs both application sites: the intent and the derivation agree on direction (up), never on subtraction. * fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal The re-parsed derived block carries string scalars ("2") while the curated snapshot carries numbers (2); substituting the curated spelling over an equal derived one was a no-op in substance but a shape churn the ADR-3473 §8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater curated counter replaces the derived value now; percent gets the same verbatim rule. * changeset(#4129): backfill PR 4359 --------- Co-authored-by: sim --- .changeset/humble-koalas-swim.md | 5 + src/phase.cts | 65 +++++++-- src/state-transition.cts | 129 ++++++++++++++--- src/state.cts | 125 ++++++++++++++++- tests/phase.test.cjs | 161 +++++++++++++++++++++ tests/state-transition.test.cjs | 145 ++++++++++++++++++- tests/state.test.cjs | 231 +++++++++++++++++++++++++++++++ 7 files changed, 827 insertions(+), 34 deletions(-) create mode 100644 .changeset/humble-koalas-swim.md diff --git a/.changeset/humble-koalas-swim.md b/.changeset/humble-koalas-swim.md new file mode 100644 index 000000000..4434cf98c --- /dev/null +++ b/.changeset/humble-koalas-swim.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4359 +--- +**`progress.completed_phases` and `percent` are now derived from the ROADMAP's own milestone Complete rows and never move downward on a state write** — previously every default-resync verb (`state record-session`, `add-decision`, `begin-phase`, `phase complete` itself) recomputed the counter from a disk scan that drops any completed phase whose verification reads `stale` (a summary committed or edited after it) or is missing, so the stored value was silently reverted to the under-count on every write and hand-corrections never survived. The scan now floors the numerator at the milestone-scoped ROADMAP Complete-row count (same gate and scope as the denominator), the write path enforces the schema-declared `progress-ratchet` (totals correct both directions, completed counters up-only, percent recomputed from the surviving counters), and `phase complete` passes its post-completion ROADMAP-derived counters through the transition so the completing phase's own write increments. (#4129) diff --git a/src/phase.cts b/src/phase.cts index 47e784cc8..832f345bf 100644 --- a/src/phase.cts +++ b/src/phase.cts @@ -59,6 +59,12 @@ const { findPhaseInternal, getArchivedPhaseDirs, listMilestonePhaseDirs } = phas // eslint-disable-next-line @typescript-eslint/no-require-imports -- roadmap-parser.cjs is an export= CommonJS module import roadmapParserMod = require('./roadmap-parser.cjs'); const { stripShippedMilestones, extractCurrentMilestone, currentMilestoneRawRanges, withPhaseSection, findMilestoneScopeHeadingLines } = roadmapParserMod; +// #4129: the single owner of "count the ROADMAP's milestone Complete rows" +// (pure computation, no I/O — no cycle on this path) for the intent-first +// progress counters the phase-complete transaction passes downstream. +// eslint-disable-next-line @typescript-eslint/no-require-imports -- phase-lifecycle.cjs is an export= CommonJS module +import phaseLifecycleMod = require('./phase-lifecycle.cjs'); +const { deriveProgressFromRoadmap: deriveProgressFromRoadmapForIntent, clampPercent: clampPercentForIntent } = phaseLifecycleMod; // eslint-disable-next-line @typescript-eslint/no-require-imports -- planning-workspace.cjs is an export= CommonJS module import planningWorkspace = require('./planning-workspace.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports -- frontmatter.cjs is an export= CommonJS module @@ -4395,14 +4401,57 @@ function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { const bodyHasPhaseField = stateExtractField(fmBody, 'Current Phase') != null || stateExtractField(fmBody, 'Phase') != null; - const authoritativeFm: Record | undefined = nextPhaseDisplayName - ? bodyHasPhaseField || !nextPhaseNum - ? { current_phase_name: nextPhaseDisplayName } - : { - current_phase: String(nextPhaseNum), - current_phase_name: nextPhaseDisplayName, - } - : undefined; + // #4129: the POST-completion progress counters, derived from the very + // ROADMAP this transaction just mutated (still in memory — it hits disk + // only at writePlanningFileSet, AFTER this content was assembled). + // buildStateFrontmatter's disk scan inside syncAndPreserveStateMd + // reads the PRE-completion ROADMAP (and any stale-dated sibling + // verification), so without this intent the persisted counter failed + // to increment on the completing phase's own transaction. Routed + // through the #2736 authoritativeFm seam's object direction: the + // pre-preservation merge makes it the derived truth the ratchet + // compares, and the post-preservation re-assert (completedOnlyRaise) + // is a floor no preservation branch can drop below. clampPercent is + // completePhaseCore's own percent formula (state-transition.cts), + // reused so the frontmatter and the body `Progress:` line agree. + const postCompletionRoadmapScope = roadmapContent !== null + ? extractCurrentMilestone(roadmapContent, cwd) + : null; + const postCompletionRoadmapProgress = postCompletionRoadmapScope !== null + ? deriveProgressFromRoadmapForIntent(postCompletionRoadmapScope) + : null; + const authoritativeProgress: Record | undefined = + postCompletionRoadmapProgress && postCompletionRoadmapProgress.completedPhases !== null + ? postCompletionRoadmapProgress.totalPhases !== null && postCompletionRoadmapProgress.totalPhases > 0 + ? { + completed_phases: postCompletionRoadmapProgress.completedPhases, + percent: clampPercentForIntent( + postCompletionRoadmapProgress.completedPhases, + postCompletionRoadmapProgress.totalPhases, + ), + } + : { completed_phases: postCompletionRoadmapProgress.completedPhases } + : undefined; + const authoritativeFm: Record | undefined = authoritativeProgress + ? { + ...(nextPhaseDisplayName + ? bodyHasPhaseField || !nextPhaseNum + ? { current_phase_name: nextPhaseDisplayName } + : { + current_phase: String(nextPhaseNum), + current_phase_name: nextPhaseDisplayName, + } + : {}), + progress: authoritativeProgress, + } + : nextPhaseDisplayName + ? bodyHasPhaseField || !nextPhaseNum + ? { current_phase_name: nextPhaseDisplayName } + : { + current_phase: String(nextPhaseNum), + current_phase_name: nextPhaseDisplayName, + } + : undefined; // ADR-3408 §8.3 / #3469: this deliberately bypasses // readModifyWriteStateMd (STATE.md is committed atomically with // ROADMAP/REQUIREMENTS), so it calls the single write-seam diff --git a/src/state-transition.cts b/src/state-transition.cts index a652b04ab..e001c4d98 100644 --- a/src/state-transition.cts +++ b/src/state-transition.cts @@ -17,11 +17,16 @@ // eslint-disable-next-line @typescript-eslint/no-require-imports import frontmatter = require('./frontmatter.cjs'); import { stateReplaceField, stateExtractField, stateReplaceFieldIfTemplate, stateReplaceFieldWithFallback, stateReplaceFieldInSession, stateCurrentPositionSlice } from './state-document.cjs'; -import { KNOWN_TEMPLATE_DEFAULTS, toFiniteNumber } from './state-document.cjs'; +import { KNOWN_TEMPLATE_DEFAULTS, toFiniteNumber, computeProgressPercent } from './state-document.cjs'; import { tokenizeHeadings } from './markdown-sectionizer.cjs'; import type { HeadingToken } from './markdown-sectionizer.cjs'; import { deriveProgressFromRoadmap, clampPercent, clampPercentFromFraction } from './phase-lifecycle.cjs'; import { escapeRegex } from './pattern.cjs'; +// #4129: the completion-ratio kernel for the resync-arm ratchet's percent +// (planning-scope's SCOPE — state-document's own dependency, no cycle here: +// state-document never imports this module). +// eslint-disable-next-line @typescript-eslint/no-require-imports +import planningScopeMod = require('./planning-scope.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports import stateMdSchemaMod = require('./state-md-schema.cjs'); const { STATE_FIELD_SCHEMA } = stateMdSchemaMod; @@ -684,6 +689,78 @@ function preservedValuesEqual(a: unknown, b: unknown): boolean { return a === b; } +/** + * #4129: the resync-arm progress merge. A resyncing write whose scan MEASURED + * something no longer wholesale-replaces the curated block — the declared + * `progress-ratchet` mergeStrategy ("completed_plans/completed_phases only ever + * ratchet UP toward the derived value (#2969)", state-md-schema.cts) now holds + * on the write path too, matching what the read path (`shouldPreserveExistingProgress`) + * has always enforced. Rules, mirroring the `deriveProgressKeys` branch above: + * + * - total_plans / total_phases always take the derived value (#2440 — totals + * correct in BOTH directions). + * - completed_plans / completed_phases take the derived value only when it is + * strictly GREATER (#2969's `>` not `>=`); else the curated value survives + * (a hand-corrected or previously-correct counter can never be re-derived + * downward — the #4129 clobber). + * - any other key keeps the curated value (the existing branch's convention). + * - percent is RECOMPUTED from the merged counters through the single kernel + * (`computeProgressPercent`), because either side's stored percent was + * computed against that side's counters and the merged block may mix them + * (curated completed, derived totals). Recomputation runs ONLY when the + * derived block itself carried a percent — an upstream withhold + * (#1761 milestone-unbounded, #3217 scope) nulled percent deliberately and + * this merge must not resurrect it. + * + * Frontmatter scalars arrive as STRINGS ("2", not 2), so every comparison + * coerces through `toFiniteNumber` — never a `typeof === 'number'` test + * (scanMeasuredSomething's own convention). + */ +function mergeResyncProgressRatchet( + curatedRecord: Record, + derivedRecord: Record, +): Record { + const merged: Record = { ...derivedRecord }; + for (const [key, value] of Object.entries(curatedRecord)) { + if (key === 'total_plans' || key === 'total_phases' || key === 'percent') continue; + if (key === 'completed_plans' || key === 'completed_phases') { + const derivedNum = toFiniteNumber(derivedRecord[key]) ?? -Infinity; + const curatedNum = toFiniteNumber(value) ?? -Infinity; + // Ratchet up only (strictly greater, #2969); else keep curated. + if (derivedNum > curatedNum) continue; + // Numerically EQUAL keeps the derived value VERBATIM. The two sides + // arrive in different scalar shapes (the re-parsed derived block + // carries string totals "2" while the curated snapshot carries numbers + // 2), and substituting the curated spelling over an equal derived one + // is a no-op in substance but a shape churn the §8.7 reporting loop + // would surface as a phantom `preserved-over-disagreeing-derived` + // warning (it diffs structurally). Only a curated counter that is + // STRICTLY greater replaces the derived value. + if (derivedNum === curatedNum) continue; + merged[key] = value; + } else { + merged[key] = value; + } + } + if (toFiniteNumber(derivedRecord.percent) !== null) { + const recomputed = computeProgressPercent( + toFiniteNumber(merged.completed_plans), + toFiniteNumber(merged.total_plans), + toFiniteNumber(merged.completed_phases), + toFiniteNumber(merged.total_phases), + planningScopeMod.SCOPE.COMPLETE, + ); + // Same verbatim rule for percent: assign only when the recomputed value + // numerically differs, so a string-spelled derived percent ("67") is not + // churned into a number-spelled 67 (phantom-divergence noise, not a + // change). + if (recomputed !== null && recomputed !== toFiniteNumber(merged.percent)) { + merged.percent = recomputed; + } + } + return merged; +} + /** * Executor for `preservation: 'preserve-always'` (ADR-3408 §8.1). Only * `progress` carries this policy today. Preserves #3242/#1446/#2440/#2969 @@ -698,25 +775,41 @@ function applyPreserveAlways(field: string, cls: FieldClassification, ctx: Prese const derived = ctx.postFm[field]; const derivedMeasured = scanMeasuredSomething(cls, derived); const curatedMeasured = scanMeasuredSomething(cls, curated); - // On a resyncing write the fresh derivation is authoritative — UNLESS it - // measured nothing while the curated block did (#3756), AND the caller did - // not explicitly name a progress-affecting field this write. The - // unmeasured-scan guard exists to stop an INCIDENTAL resync (e.g. `state - // add-decision`, whose `resync` defaults true for reasons that have - // nothing to do with `progress`) from dropping a real curated block when a - // milestone-scoped disk scan measures nothing (#3756's archived-milestone - // case). It must not also block a write the user pointed AT `progress` on - // purpose: `preserve-always`'s own contract is "never overwrite unless the - // caller explicitly names this field" (FIELD_CLASSIFICATION doc comment), - // and `state update Progress` / `state patch Progress=...` are exactly - // that naming — the resync they trigger must win even when the disk scan - // it also drives (e.g. because there are no phase dirs at all) reads as - // "unmeasured" (tests/frontmatter.test.cjs: "state.update \"Progress\" - // resyncs progress frontmatter from the updated body", pre-existing, #3242). - if (ctx.resync && (derivedMeasured || !curatedMeasured || ctx.explicitProgressField)) return; + // On a resyncing write the fresh derivation is authoritative in two cases + // (#3756 / ADR-3473 §8.6, unchanged): when the caller EXPLICITLY named a + // progress-affecting field (`preserve-always`'s contract is "never + // overwrite unless the caller explicitly names this field" — `state update + // Progress` is exactly that naming, pre-existing #3242 behavior), and when + // the derivation measured something the curated block did not (an + // unmeasured CURATED block is not worth protecting). The unmeasured-DERIVED + // guard also stands: an incidental resync (e.g. `state add-decision`, whose + // `resync` defaults true for reasons that have nothing to do with + // `progress`) that measured nothing must not drop a real curated block + // (#3756's archived-milestone case) — that falls through to the wholesale + // restore below. + // + // #4129 narrows the remaining arm. A resyncing write whose scan MEASURED + // something while the curated block is also real previously wholesale- + // replaced the curated block with the derived one — no monotonic guard, so + // any under-counting derivation (a stale-dated verification, #2348) silently + // reverted every hand-correction and every correct value an earlier write + // had persisted, while the read path (`shouldPreserveExistingProgress`) + // kept reporting the higher stored counters. That arm now falls through to + // `mergeResyncProgressRatchet` — the declared `progress-ratchet` + // mergeStrategy, finally enforced on the write path: totals derived both + // directions (#2440), completed counters up-only (#2969), percent + // recomputed from the merged counters. + if (ctx.resync && (ctx.explicitProgressField || (derivedMeasured && !curatedMeasured))) return; let next: unknown; - if (cls.mergeStrategy === 'progress-ratchet' && ctx.deriveProgressKeys && derived && derivedMeasured) { + if (cls.mergeStrategy === 'progress-ratchet' && ctx.resync && derived && derivedMeasured && curatedMeasured) { + // #4129 resync arm — see mergeResyncProgressRatchet's doc. Reached only + // after the early-out above, so curatedMeasured is guaranteed true here. + next = mergeResyncProgressRatchet( + curated as Record, + (derived ?? {}) as Record, + ); + } else if (cls.mergeStrategy === 'progress-ratchet' && ctx.deriveProgressKeys && derived && derivedMeasured) { // #2440: total_plans and total_phases always take the derived (post-sync) // value even under !resync. This is used by cmdStatePlannedPhase where // total_plans must correct upward after plans are added. For body-only diff --git a/src/state.cts b/src/state.cts index 2ae2cdd61..b4ecedde0 100644 --- a/src/state.cts +++ b/src/state.cts @@ -69,6 +69,13 @@ import planDependencyGraphMod = require('./plan-dependency-graph.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports import verificationMod = require('./verification.cjs'); const { isPhaseComplete } = verificationMod; +// #4129: the single owner of "count the ROADMAP's milestone Complete rows" +// (phase-lifecycle.cts) — reused for the completed-phases numerator floor so +// this scan cannot grow a second ROADMAP parser. Pure computation module (no +// I/O), so it introduces no cycle on this path. +// eslint-disable-next-line @typescript-eslint/no-require-imports +import phaseLifecycleMod = require('./phase-lifecycle.cjs'); +const { deriveProgressFromRoadmap } = phaseLifecycleMod; // eslint-disable-next-line @typescript-eslint/no-require-imports import planningScopeMod = require('./planning-scope.cjs'); const { SCOPE } = planningScopeMod; @@ -3014,6 +3021,34 @@ function buildStateFrontmatter( // write silently clobbered the three stored siblings with the // under-scoped disk numbers. const diskCountsWithheld = milestonedButUnbounded || roadmapAbsentWithAssertedMilestone; + // #4129: floor the completed-phases numerator at the ROADMAP's own + // milestone Complete-row count. The disk numerator counts ONLY + // phase dirs whose *-VERIFICATION.md routes `passed` (isPhaseComplete, + // #2957 disk-strict — the gate stays untouched), so a completed + // phase whose verification reads `stale` (a SUMMARY committed or + // edited after it, #2348 clean-commit-time clock) or `missing` + // (pre-verification era, hand-flipped ROADMAP row) drops out of the + // count forever — while every other surface (the ROADMAP row + // `phase complete` just flipped, the body `Completed Phases` field + // completePhaseCore derives from deriveProgressFromRoadmap) still + // asserts the phase complete. max(disk, ROADMAP) keeps the disk + // signal for gap detection (a verification-passed phase whose ROADMAP + // row is not yet flipped still counts) while never UNDER-counting + // what the ROADMAP asserts. Scoped exactly like the denominator: + // the same milestone window (roadmapScope), the same + // safeToUseRoadmapCount gate, and never under the #3354/#3573 + // withhold — a whole-document Complete-row count must not leak + // through an untrustworthy scope. Reuses deriveProgressFromRoadmap + // (phase-lifecycle.cts, the one owner of "read the Progress table") + // — no second ROADMAP parser here. A ROADMAP without a canonical + // `## Progress` table resolves no table → floor inert (disk count + // stands), the owner's own answer to "what is countable". + const roadmapCompletedPhases = roadmapScope !== null && safeToUseRoadmapCount && !diskCountsWithheld + ? deriveProgressFromRoadmap(roadmapScope).completedPhases + : null; + const flooredCompletedPhases = roadmapCompletedPhases !== null + ? Math.max(diskCompletedPhases, roadmapCompletedPhases) + : diskCompletedPhases; return { // The two WITHHOLD shapes (#3354 milestoned-but-unbounded, #3573 // roadmap-absent-with-asserted-milestone) must be evaluated BEFORE @@ -3024,7 +3059,7 @@ function buildStateFrontmatter( ? null : (safeToUseRoadmapCount ? Math.max(phaseDirs.length, roadmapPhaseCount) : phaseDirs.length), milestoneBounded, - completedPhases: diskCountsWithheld ? null : diskCompletedPhases, + completedPhases: diskCountsWithheld ? null : flooredCompletedPhases, totalPlans: diskCountsWithheld ? null : diskTotalPlans, completedPlans: diskCountsWithheld ? null : diskTotalSummaries, phaseDirScope, @@ -3395,6 +3430,76 @@ function readStoredCompletedPlans(existingFm: Record | null | u return readStoredProgressCounter(existingFm, 'completed_plans'); } +/** + * #4129: is this authoritativeFm value a PARTIAL progress intent? The #2736 + * seam was string-only (names); #4129 extends it with one object direction — + * the `progress` key carrying the sub-keys a transition resolved + * authoritatively (completePhase's ROADMAP-derived completed_phases/percent). + * Anything else keeps the seam's existing contract untouched. + */ +function isPartialProgressIntent(value: unknown): value is Record { + return typeof value === 'object' && value !== null && !Array.isArray(value); +} + +/** + * #4129: merge a PARTIAL progress intent (see isPartialProgressIntent) into a + * frontmatter object's `progress` block. Sub-keys are accepted only when they + * are a declared `progress.*` row in FIELD_CLASSIFICATION — the single policy + * source (ADR-3408 §8.5) decides which leaves exist; an intent may not invent + * one. Returns whether anything changed. + * + * `completedOnlyRaise` (both application sites use it): completed counters + * apply only when strictly greater than what is already in the block, so no + * intent can LOWER a count another trustworthy signal already established — + * at the pre-preservation site the disk derivation's own count (a + * verification-passed phase whose ROADMAP row drifted behind), at the + * post-preservation re-assert the #2969 monotonic property preservation just + * enforced. `percent` follows its sibling: it is applied when a completed + * counter moved this call (the intent percent was computed from the intent + * counters and is coherent with them) or when the block has no percent to + * lose (a repair, never a regression of an upstream withhold — the withhold + * nulled percent upstream precisely so no write would re-assert one over + * untrustworthy counts; here the intent's own counts ARE the trustworthy + * source, the post-completion ROADMAP). + */ +function applyAuthoritativeProgressSubkeys( + fm: Record, + intent: Record, + opts: { completedOnlyRaise: boolean }, +): boolean { + const current = fm['progress']; + const base: Record = isPartialProgressIntent(current) + ? { ...current } + : {}; + let changed = false; + let completedMoved = false; + for (const [subkey, value] of Object.entries(intent)) { + if (typeof value !== 'number' || !Number.isFinite(value)) continue; + if (!getFieldClassification(`progress.${subkey}`)) continue; + const isCompletedCounter = subkey === 'completed_phases' || subkey === 'completed_plans'; + if (isCompletedCounter && opts.completedOnlyRaise) { + const currentNum = toFiniteNumber(base[subkey]); + if (currentNum !== null && currentNum >= value) continue; + } + if (isCompletedCounter && !Object.is(base[subkey], value)) completedMoved = true; + if (!Object.is(base[subkey], value)) { + base[subkey] = value; + changed = true; + } + } + // percent: applied only when a completed counter moved (coherent with the + // counters that just landed) or when no percent exists to contradict. + const intentPercent = intent['percent']; + if (typeof intentPercent === 'number' && Number.isFinite(intentPercent) && (completedMoved || toFiniteNumber(base['percent']) === null)) { + if (!Object.is(base['percent'], intentPercent)) { + base['percent'] = intentPercent; + changed = true; + } + } + if (changed) fm['progress'] = base; + return changed; +} + function syncStateFrontmatter( content: string, cwd: string | undefined, @@ -3620,10 +3725,20 @@ function syncStateFrontmatter( // parenthetical (`Closer-ruling measurement (D1a)` → `D1a`) — never runs // the final word on a field the transition just resolved. The prose parser // remains the fallback for genuinely unknown prose only. + // #4129: the `progress` key carries a PARTIAL block (the object direction of + // this seam — see applyAuthoritativeProgressSubkeys) for the same reason: + // completePhase holds the POST-completion ROADMAP, and the disk scan this + // function drives reads the PRE-completion one. The intent is applied as a + // FLOOR here too (completedOnlyRaise): a derivation that already counted + // MORE completed phases than the ROADMAP table asserts (verification-passed + // phases whose table rows drifted behind) must not be lowered by the intent + // — the two signals agree on direction (up), never on subtraction. if (authoritativeFm) { for (const [key, value] of Object.entries(authoritativeFm)) { if (typeof value === 'string' && value.trim().length > 0) { derivedFm[key] = value; + } else if (key === 'progress' && isPartialProgressIntent(value)) { + applyAuthoritativeProgressSubkeys(derivedFm, value, { completedOnlyRaise: true }); } } } @@ -4194,12 +4309,20 @@ function applyPostSyncPreservation( // (equal), so the #1695 restore fires and would put the stale pre-transition // name back over the authoritative one. Intent beats both the prose // re-derivation and the curated restore — the transition just resolved it. + // #4129: for the `progress` key the re-assert is a FLOOR, not an override — + // the #2969 monotonic property preservation just enforced (completed + // counters never move down) must not be undone by the intent, so completed + // sub-keys apply only-raise here (see applyAuthoritativeProgressSubkeys). let authoritativeReasserted = false; if (authoritativeFm) { for (const [key, value] of Object.entries(authoritativeFm)) { if (typeof value === 'string' && value.trim().length > 0 && preservation.postFm[key] !== value) { preservation.postFm[key] = value; authoritativeReasserted = true; + } else if (key === 'progress' && isPartialProgressIntent(value)) { + if (applyAuthoritativeProgressSubkeys(preservation.postFm, value, { completedOnlyRaise: true })) { + authoritativeReasserted = true; + } } } } diff --git a/tests/phase.test.cjs b/tests/phase.test.cjs index 8e6081efb..83c36d646 100644 --- a/tests/phase.test.cjs +++ b/tests/phase.test.cjs @@ -7380,6 +7380,167 @@ describe('bug-3287 — init plan-phase exposes expected_phase_dir with project_c }); }); + // ───────────────────────────────────────────────────────────────────────── + // #4129 row 1: `phase complete N` failed to increment completed_phases when + // a SIBLING completed phase's verification routes `stale` (its SUMMARY was + // touched after the verification — the issue's real-world drift). The + // transaction flips this phase's ROADMAP row and derives the BODY counters + // from the post-completion ROADMAP, but the frontmatter disk scan reads the + // PRE-completion ROADMAP and the stale-dated sibling, so the persisted + // counter stays pinned at the under-count. See .gsd/bug/ + // fix-4129-completed-phases-recompute/{10-diagnosis,50-test-matrix}.md. + // ───────────────────────────────────────────────────────────────────────── + describe('#4129: phase complete increments completed_phases to the ROADMAP truth', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4129-phase-')); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + // Body Progress percent through the repo's own field extractor (never raw + // substring matching on rendered STATE.md — CONTRIBUTING.md prohibits it). + function bodyProgressPercentFromState(stateContent) { + const raw = stateExtractField(stateContent, 'Progress'); + if (raw === null) return null; + const match = raw.match(/(\d{1,3})%/); + return match ? Number(match[1]) : null; + } + + function setupStaleSiblingProject(tmpDir) { + const planningDir = path.join(tmpDir, '.planning'); + const phasesDir = path.join(planningDir, 'phases'); + fs.mkdirSync(phasesDir, { recursive: true }); + fs.writeFileSync(path.join(planningDir, 'config.json'), JSON.stringify({ project_code: 'REPRO' })); + + const roadmapLines = [ + '# Roadmap', + '', + '## Current Milestone: v1.0', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 1. | 2/2 | Complete | 2026-01-01 |', + '| 2. | 2/2 | Complete | 2026-01-02 |', + '| 3. | 2/2 | In Progress | |', + ]; + for (let i = 4; i <= 18; i += 1) roadmapLines.push(`| ${i}. | 0/2 | Not Started | |`); + roadmapLines.push('', '- [x] Phase 1: Alpha (completed 2026-01-01)', '- [x] Phase 2: Beta (completed 2026-01-02)', '- [ ] Phase 3: Gamma'); + for (let i = 4; i <= 18; i += 1) roadmapLines.push(`- [ ] Phase ${i}: P${i}`); + for (let i = 1; i <= 18; i += 1) { + roadmapLines.push('', `### Phase ${i}: P${i}`, '', '**Goal:** goal', '**Plans:** 2 plans', ''); + } + fs.writeFileSync(path.join(planningDir, 'ROADMAP.md'), roadmapLines.join('\n')); + + fs.writeFileSync( + path.join(planningDir, 'STATE.md'), + [ + '---', + 'gsd_state_version: 1.0', + 'milestone: v1.0', + 'milestone_name: Programme', + 'status: executing', + 'current_phase: 3', + 'last_updated: 2026-01-02T10:00:00.000Z', + 'progress:', + ' total_phases: 18', + ' completed_phases: 2', + ' total_plans: 4', + ' completed_plans: 4', + ' percent: 11', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 3 of 18 (Gamma) — EXECUTING', + 'Plan: 2 of 2', + 'Status: Executing Phase 3', + 'Last activity: 2026-01-02', + '', + '## Progress', + '', + 'Progress: [█░░░░░░░░░] 11% (2/18 phases complete)', + '', + '## Session Continuity', + '', + 'Last session: 2026-01-02T10:00:00.000Z', + '', + ].join('\n'), + ); + + for (const p of [1, 2, 3]) { + const pp = String(p).padStart(2, '0'); + const dir = path.join(phasesDir, `${pp}-p${p}`); + fs.mkdirSync(dir, { recursive: true }); + for (const i of [1, 2]) { + fs.writeFileSync(path.join(dir, `${pp}-0${i}-PLAN.md`), '# Plan\n'); + fs.writeFileSync(path.join(dir, `${pp}-0${i}-SUMMARY.md`), '# Summary\n'); + } + fs.writeFileSync( + path.join(dir, `${pp}-VERIFICATION.md`), + ['---', 'status: passed', '---', '', '# Verification', ''].join('\n'), + ); + } + + // The drift: phase 1's summary touched after its verification. No git + // repo → the #2348 clock compares mtimes; the newer summary mtime + // routes phase 1's verification `stale` (verified via + // `verification status` in the diagnosis repro). + const older = new Date('2026-01-01T00:00:00Z'); + const newer = new Date('2026-03-01T00:00:00Z'); + fs.utimesSync(path.join(phasesDir, '01-p1', '01-VERIFICATION.md'), older, older); + fs.utimesSync(path.join(phasesDir, '01-p1', '01-01-SUMMARY.md'), newer, newer); + + return { planningDir }; + } + + test('phaseCompleteIncrementsCompletedPhasesPastStaleSibling', () => { + setupStaleSiblingProject(tmpDir); + const statePath = path.join(tmpDir, '.planning', 'STATE.md'); + + const r = runSdkQuery(['phase.complete', '3'], tmpDir); + assert.ok(r.success, `phase complete 3 failed: ${r.error}`); + + const state = fs.readFileSync(statePath, 'utf8'); + const fm = extractFrontmatter(state); + assert.ok(fm.progress, 'progress block must exist after phase complete'); + assert.equal( + Number(fm.progress.completed_phases), + 3, + `#4129: completing phase 3 must increment completed_phases to the ROADMAP truth 3 (phases 1-3 Complete), got ${fm.progress.completed_phases}`, + ); + assert.equal( + Number(fm.progress.percent), + 17, + `#4129: percent must follow the incremented counter (3/18), got ${fm.progress.percent}`, + ); + // The body bar must stay coherent with the persisted percent (#4129 AC7). + assert.equal(bodyProgressPercentFromState(state), 17, 'the body Progress bar must carry the same 17% as the frontmatter percent'); + // The ROADMAP row this very transaction flipped is the authority. + const roadmap = fs.readFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), 'utf8'); + assert.ok(/^- \[x\] Phase 3: Gamma/m.test(roadmap), 'precondition: the transaction flipped the phase 3 ROADMAP checkbox'); + assert.ok(/^\| 3\.\s*\|\s*2\/2\s*\|\s*Complete\s*\|/m.test(roadmap), 'precondition: the transaction flipped the phase 3 table row'); + }); + + test('phaseCompleteIsIdempotentOnTheRoadmapTruth', () => { + setupStaleSiblingProject(tmpDir); + const statePath = path.join(tmpDir, '.planning', 'STATE.md'); + + const r1 = runSdkQuery(['phase.complete', '3'], tmpDir); + assert.ok(r1.success, `first call failed: ${r1.error}`); + const r2 = runSdkQuery(['phase.complete', '3'], tmpDir); + assert.ok(r2.success, `second call failed: ${r2.error}`); + + const fm = extractFrontmatter(fs.readFileSync(statePath, 'utf8')); + assert.equal(Number(fm.progress.completed_phases), 3, '#4129: double-complete stays at the ROADMAP truth 3 (idempotent)'); + }); + }); + // ───────────────────────────────────────────────────────────────────────── // ADR-3408 §8.3 Matrix A (#3469): cmdPhaseComplete now calls the ONE // write-seam composition (syncAndPreserveStateMd) directly instead of diff --git a/tests/state-transition.test.cjs b/tests/state-transition.test.cjs index 8eedd6053..0ad4996ba 100644 --- a/tests/state-transition.test.cjs +++ b/tests/state-transition.test.cjs @@ -4348,9 +4348,19 @@ describe('ADR-3473 §8.6 matrix rows 13-17 (+10/11 pinned): the measured-vs-unme }); test('stringNonZeroTotalIsMeasured', () => { + // #4129 superseded the observable: a measured resyncing write no longer + // lets the derived block wholesale-replace the curated one — the + // declared progress-ratchet now merges (totals derived both directions, + // completed counters up-only, #2969). The measured/unmeasured BOUNDARY + // this row pins is unchanged and still observable here: the derived + // TOTALS stand (a wholesale curated restore would have written 5/32). const r = restoreWithDerivedTotals({ total_phases: '1', total_plans: '0', completed_phases: '0', completed_plans: '0' }); - assert.strictEqual(r.mutated, false, 'a measured scan (total_phases:"1") must win — derived stands untouched'); - assert.deepStrictEqual(r.postFm.progress, { total_phases: '1', total_plans: '0', completed_phases: '0', completed_plans: '0' }); + assert.strictEqual(r.mutated, true, 'a measured scan (total_phases:"1") must not wholesale-restore the curated block — the #4129 ratchet merge runs'); + assert.deepStrictEqual( + r.postFm.progress, + { total_phases: '1', total_plans: '0', completed_phases: 5, completed_plans: 32 }, + 'measured: totals take the derived value (#2440 both directions); completed counters keep curated ("0" < 5/#2969 up-only); derived carried no percent so none is invented', + ); }); test('absentTotalsAreUnmeasured', () => { @@ -4379,15 +4389,18 @@ describe('ADR-3473 §8.6 matrix rows 13-17 (+10/11 pinned): the measured-vs-unme // Pinned mirror pair from the design's row 11/the existing matrix's row // 10 — grouped here so the boundary (only both-zero is unmeasured) reads - // legibly against the coercion cases above. + // legibly against the coercion cases above. #4129 superseded the observable + // to the ratchet merge (see stringNonZeroTotalIsMeasured above): "measured" + // is still decided by EITHER total being non-zero, and still observable + // because the derived TOTALS stand rather than a curated wholesale restore. test('oneNonZeroTotalCountsAsMeasuredEitherDirection', () => { const r1 = restoreWithDerivedTotals({ total_phases: 1, total_plans: 0, completed_phases: 0, completed_plans: 0 }); - assert.strictEqual(r1.mutated, false, 'total_phases:1, total_plans:0 must count as measured'); - assert.deepStrictEqual(r1.postFm.progress, { total_phases: 1, total_plans: 0, completed_phases: 0, completed_plans: 0 }); + assert.strictEqual(r1.mutated, true, 'total_phases:1, total_plans:0 must count as measured (#4129 ratchet merge ran, not a curated restore)'); + assert.deepStrictEqual(r1.postFm.progress, { total_phases: 1, total_plans: 0, completed_phases: 5, completed_plans: 32 }); const r2 = restoreWithDerivedTotals({ total_phases: 0, total_plans: 1, completed_phases: 0, completed_plans: 0 }); - assert.strictEqual(r2.mutated, false, 'total_phases:0, total_plans:1 must count as measured (the mirror) — only both-zero is unmeasured'); - assert.deepStrictEqual(r2.postFm.progress, { total_phases: 0, total_plans: 1, completed_phases: 0, completed_plans: 0 }); + assert.strictEqual(r2.mutated, true, 'total_phases:0, total_plans:1 must count as measured (the mirror) — only both-zero is unmeasured'); + assert.deepStrictEqual(r2.postFm.progress, { total_phases: 0, total_plans: 1, completed_phases: 5, completed_plans: 32 }); }); }); @@ -4441,3 +4454,121 @@ describe('ADR-3473 §8.6 matrix row 36: property — preservation never drops a ); }); }); + +// ───────────────────────────────────────────────────────────────────────────── +// #4129: the resync arm of `applyPreserveAlways` (rows 8-14 of .gsd/bug/ +// fix-4129-completed-phases-recompute/50-test-matrix.md). A resyncing write +// whose scan MEASURED something now runs the declared `progress-ratchet` +// merge instead of wholesale-replacing the curated block — the write path +// finally enforces the same monotonic property the read path +// (`shouldPreserveExistingProgress`) always has. The #3756 unmeasured guard +// and the #3242/#3871 explicit-progress contract are unchanged (pinned by the +// blocks above and by tests/frontmatter.test.cjs). +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4129: resyncing measured write ratchets the progress block', () => { + function resyncMerge(curatedProgress, derivedProgress, extra = {}) { + const tx = openStateTransaction({ + snapshot: { progress: { ...curatedProgress } }, + resync: true, + bodyDeltas: neutralBodyDeltasForMatrix(), + ...extra, + }); + return applyStatePreservation({ transaction: tx, postFm: { progress: { ...derivedProgress } } }); + } + + test('resyncRatchetKeepsCuratedCompletedWhenDerivedUnderCounts', () => { + // The issue's exact shape: derived 2 (a stale-dated sibling verification) + // vs curated 3 (the ROADMAP truth a hand-fix or earlier correct write left). + const r = resyncMerge( + { total_phases: 18, completed_phases: 3, total_plans: 6, completed_plans: 6, percent: 17 }, + { total_phases: 18, completed_phases: 2, total_plans: 6, completed_plans: 6, percent: 11 }, + ); + assert.deepStrictEqual( + r.postFm.progress, + { total_phases: 18, completed_phases: 3, total_plans: 6, completed_plans: 6, percent: 17 }, + '#4129: a measured resync must never move completed counters DOWN — totals derived, completed curated, percent recomputed from the merged counters (3/18=17)', + ); + assert.strictEqual(r.mutated, true, 'the merge actually changed the derived block'); + }); + + test('resyncRatchetStillLetsCompletedMoveUp', () => { + // Genuine completion: derived ABOVE curated must ratchet up, and percent follows. + const r = resyncMerge( + { total_phases: 18, completed_phases: 2, total_plans: 6, completed_plans: 6, percent: 11 }, + { total_phases: 18, completed_phases: 3, total_plans: 6, completed_plans: 6, percent: 17 }, + ); + assert.deepStrictEqual( + r.postFm.progress, + { total_phases: 18, completed_phases: 3, total_plans: 6, completed_plans: 6, percent: 17 }, + '#4129: the ratchet is up-only, never a freeze — a genuine increment still lands', + ); + }); + + test('resyncRatchetMixedSidesRecomputePercentFromMergedCounters', () => { + // Mixed: derived completed_plans ratchets up, curated completed_phases + // survives — percent must be recomputed from the MERGED counters + // (min(54/54 plans, 5/? phases capped by plan fraction), never either + // side's stale stored percent. + const r = resyncMerge( + { total_plans: 54, completed_plans: 50, completed_phases: 5, percent: 93 }, + { total_plans: 54, completed_plans: 54, completed_phases: 1, percent: 50 }, + ); + assert.strictEqual(r.postFm.progress.completed_plans, 54, 'derived completed_plans (54 > 50) ratchets up'); + assert.strictEqual(r.postFm.progress.completed_phases, 5, 'curated completed_phases (5 > 1) survives — never down'); + // min(54/54, 5/5-with-no-total... completed_phases 5, total_phases absent) — + // computeProgressPercent with only plan data present: min(1.0, plan)=100? No: + // phase data absent → phaseFraction=1 → min(1, 1) = 100. + assert.strictEqual(r.postFm.progress.percent, 100, 'percent is recomputed from the merged counters through the completion-ratio kernel'); + }); + + test('resyncRatchetCoercesStringScalarsThroughToFiniteNumber', () => { + // Frontmatter scalars arrive as STRINGS ("2", not 2) — the comparison + // must coerce (scanMeasuredSomething's own convention), never `typeof`. + const r = resyncMerge( + { total_phases: 18, completed_phases: 3, percent: 17 }, + { total_phases: '18', total_plans: '6', completed_phases: '2', completed_plans: '6', percent: 11 }, + ); + assert.strictEqual(r.postFm.progress.completed_phases, 3, 'string "2" must compare numerically against curated 3 — curated survives'); + assert.strictEqual(r.postFm.progress.total_phases, '18', 'string totals take the derived value verbatim (both directions)'); + assert.strictEqual(r.postFm.progress.percent, 17, 'percent recomputed from merged counters (3/18)'); + }); + + test('resyncRatchetIsANoOpWhenDerivedEqualsCurated', () => { + // The #4094 withheld shape arrives here: the scan fell back to the stored + // counters, so derived == curated and the merge must not report a + // mutation (the #948 no-op-write family). + const block = { total_phases: 18, completed_phases: 3, total_plans: 6, completed_plans: 6, percent: 17 }; + const r = resyncMerge(block, { ...block }); + assert.strictEqual(r.mutated, false, 'derived === curated → no mutation (withheld shape is untouched)'); + assert.deepStrictEqual(r.postFm.progress, block); + }); + + test('resyncRatchetDoesNotResurrectAWithheldPercent', () => { + // Derived percent ABSENT (an upstream #1761/#3217 withhold nulled it) — + // the merge must not recompute one over counters it was withheld for. + // Percent follows the derived side (the deriveProgressKeys branch's own + // convention) and recomputation is gated on the derived block HAVING + // carried one — an absent percent stays absent, exactly as the + // pre-#4129 wholesale replace left it. + const r = resyncMerge( + { total_phases: 5, completed_phases: 5, percent: 100 }, + { total_phases: 5, completed_phases: 2 }, + ); + assert.strictEqual(r.postFm.progress.completed_phases, 5, 'curated completed counter survives'); + assert.ok(!('percent' in r.postFm.progress), 'no recomputed percent may appear when the derived block carried none'); + }); + + test('explicitProgressFieldStillWholesaleReplacesOnMeasuredScan', () => { + // #3242/#3871 contract, unchanged by #4129: when the caller NAMED a + // progress field, the resync they asked for wins outright — the escape + // hatch for deliberate downward correction stays open. + const r = resyncMerge( + { total_phases: 18, completed_phases: 3, percent: 17 }, + { total_phases: 18, completed_phases: 2, percent: 11 }, + { explicitProgressField: true }, + ); + assert.strictEqual(r.mutated, false, 'explicit progress write: the derived block stands untouched'); + assert.deepStrictEqual(r.postFm.progress, { total_phases: 18, completed_phases: 2, percent: 11 }); + }); +}); diff --git a/tests/state.test.cjs b/tests/state.test.cjs index 31ea8fe14..8c843ed89 100644 --- a/tests/state.test.cjs +++ b/tests/state.test.cjs @@ -18163,6 +18163,237 @@ describe('#3871 / #3756: curated progress must survive a write on an archived mi }); }); +// ═════════════════════════════════════════════════════════════════════════ +// #4129: progress.completed_phases is recomputed to a WRONG value on every +// resyncing state write, and hand-fixes never survive (.gsd/bug/ +// fix-4129-completed-phases-recompute/{10-diagnosis,50-test-matrix}.md). +// +// The repro shape (issue rows 1/4): a completed phase whose SUMMARY was +// touched after its verification passed — a later reformat/re-run commit, or +// any dirty working-tree edit — permanently stale-dates that phase's +// verification under the #2348 clean-commit-time clock. isPhaseComplete +// (#2957 disk-strict, correctly) refuses to count it, so +// buildStateFrontmatter's disk numerator UNDER-counts vs the ROADMAP +// Complete rows that `phase complete` itself maintains, and the resync arm +// of applyPreserveAlways wholesale-replaces the stored block with the +// under-count on every default-resync write (record-session / add-decision / +// begin-phase / ...), clobbering any hand-corrected value. +// +// The fixture below is git-free: outside a repo the #2348 clock falls back +// to filesystem mtimes, so a newer-mtime SUMMARY reproduces the exact stale +// routing the real git clock produces (verified against the same +// `verification status` CLI the reporter used). +// ═════════════════════════════════════════════════════════════════════════ + +describe('#4129: completed_phases derives from the ROADMAP authority and survives resyncing writes', () => { + /** + * The #4129 STALE shape: milestone v1.0, 18 phases, phases 1-3 Complete in + * the ROADMAP (canonical 4-column Progress table + checklist), each with + * plans/summaries and a passing verification; phase 1's SUMMARY carries a + * NEWER mtime than its verification → `verification status` routes `stale` + * → disk numerator 2, ROADMAP truth 3. STATE.md starts at the truth (3). + * `includeStoredProgress: false` writes STATE.md with NO stored progress + * block, so the reported counters are exactly what the derivation computes + * (the read-path ratchet has no stored block to lean on). + */ + function buildStaleVerificationFixture(cwd, initialCompleted = 3, includeStoredProgress = true) { + const planningDir = path.join(cwd, '.planning'); + const phasesDir = path.join(planningDir, 'phases'); + fs.mkdirSync(phasesDir, { recursive: true }); + fs.writeFileSync(path.join(planningDir, 'config.json'), JSON.stringify({ project_code: 'REPRO' })); + + const roadmapLines = [ + '# Roadmap', + '', + '## Current Milestone: v1.0', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 1. | 2/2 | Complete | 2026-01-01 |', + '| 2. | 2/2 | Complete | 2026-01-02 |', + '| 3. | 2/2 | Complete | 2026-01-03 |', + ]; + for (let i = 4; i <= 18; i += 1) roadmapLines.push(`| ${i}. | 0/2 | Not Started | |`); + roadmapLines.push('', '- [x] Phase 1: Alpha (completed 2026-01-01)', '- [x] Phase 2: Beta (completed 2026-01-02)', '- [x] Phase 3: Gamma (completed 2026-01-03)'); + for (let i = 4; i <= 18; i += 1) roadmapLines.push(`- [ ] Phase ${i}: P${i}`); + for (let i = 1; i <= 18; i += 1) { + roadmapLines.push('', `### Phase ${i}: P${i}`, '', '**Goal:** goal', '**Plans:** 2 plans', ''); + } + fs.writeFileSync(path.join(planningDir, 'ROADMAP.md'), roadmapLines.join('\n')); + + const stateLines = [ + '---', + 'gsd_state_version: 1.0', + 'milestone: v1.0', + 'milestone_name: Programme', + 'status: executing', + 'current_phase: 4', + 'last_updated: 2026-01-03T10:00:00.000Z', + ]; + if (includeStoredProgress) { + stateLines.push( + 'progress:', + ' total_phases: 18', + ` completed_phases: ${initialCompleted}`, + ' total_plans: 6', + ' completed_plans: 6', + ' percent: 17', + ); + } + stateLines.push( + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 4 of 18 (P4) — EXECUTING', + 'Plan: 1 of 2', + 'Status: Executing Phase 4', + 'Last activity: 2026-01-03', + '', + '## Progress', + '', + `Progress: [██░░░░░░░░] 17% (${initialCompleted}/18 phases complete)`, + '', + '## Session', + '', + 'Last session: 2026-01-03T10:00:00.000Z', + 'Stopped at: Finished phase 3', + 'Resume file: None', + '', + ); + fs.writeFileSync(path.join(planningDir, 'STATE.md'), stateLines.join('\n')); + + for (const p of [1, 2, 3]) { + const pp = String(p).padStart(2, '0'); + const dir = path.join(phasesDir, `${pp}-p${p}`); + fs.mkdirSync(dir, { recursive: true }); + for (const i of [1, 2]) { + fs.writeFileSync(path.join(dir, `${pp}-0${i}-PLAN.md`), '# Plan\n'); + fs.writeFileSync(path.join(dir, `${pp}-0${i}-SUMMARY.md`), '# Summary\n'); + } + writePassedVerification(cwd, `${pp}-p${p}`, pp); + } + + // The drift: phase 1's summary edited AFTER the verification was written. + // No git repo → the #2348 clock compares mtimes; the newer summary mtime + // routes the phase-1 verification `stale` exactly as a later commit would. + const verificationPath = path.join(phasesDir, '01-p1', '01-VERIFICATION.md'); + const driftedSummary = path.join(phasesDir, '01-p1', '01-01-SUMMARY.md'); + const older = new Date('2026-01-01T00:00:00Z'); + const newer = new Date('2026-03-01T00:00:00Z'); + fs.utimesSync(verificationPath, older, older); + fs.utimesSync(driftedSummary, newer, newer); + + return { planningDir, phasesDir }; + } + + function readProgress(cwd) { + const content = fs.readFileSync(path.join(cwd, '.planning', 'STATE.md'), 'utf-8'); + const fm = frontmatterLib.extractFrontmatter(content); + assert.ok(fm && fm.progress, 'STATE.md frontmatter must carry a progress block'); + return { progress: fm.progress, content }; + } + + // Row 1 of the 50-test-matrix — the failing-first regression. On current + // `next` each of these verbs clobbers the stored 3 down to the + // stale-verification disk count 2 (percent 17 → 11), which is the issue's + // "any hand-correction is silently reverted by the next one". + test('handFixedCompletedPhasesSurvivesEveryResyncingWrite', (t) => { + const cwd = createTempDir('gsd-4129-handfix-'); + t.after(() => cleanup(cwd)); + buildStaleVerificationFixture(cwd, 3); + + // Precondition — the drift is really in place: phase 1 routes stale, so + // the disk numerator is 2 while ROADMAP + stored say 3. + const staleCheck = runGsdTools(['verification', 'status', path.join('.planning', 'phases', '01-p1')], cwd); + assert.ok(staleCheck.success, `verification status failed: ${staleCheck.error}`); + assert.strictEqual(JSON.parse(staleCheck.output).status, 'stale', 'fixture precondition: phase 1 verification must route stale (newer summary)'); + + for (const [label, args] of [ + ['state record-session', ['state', 'record-session', '--stopped-at', 'Finished phase 3 verification']], + ['state add-decision', ['state', 'add-decision', '--summary', 'Ship it']], + ['state begin-phase', ['state', 'begin-phase', '--phase', '4', '--name', 'P4']], + ]) { + const result = runGsdTools(args, cwd); + assert.ok(result.success, `${label} failed: ${result.error}`); + const { progress, content } = readProgress(cwd); + assert.strictEqual( + Number(progress.completed_phases), + 3, + `#4129: ${label} must not clobber the ROADMAP-correct completed_phases 3 down to the stale-verification disk count (${progress.completed_phases})`, + ); + assert.strictEqual(Number(progress.percent), 17, `#4129: ${label} — percent follows the surviving counters (3/18), not the discarded disk count`); + assert.strictEqual(bodyProgressPercent(content), 17, `#4129: ${label} — the body Progress bar must stay coherent with the persisted percent`); + } + }); + + // Row 2 — the DERIVED block itself must carry the ROADMAP-floored numerator. + // state json applies the read-path ratchet (shouldPreserveExistingProgress), + // which would mask a still-wrong derivation whenever a stored block exists — + // so this asserts on a STATE.md whose stored block is ABSENT: what json + // reports is then exactly what buildStateFrontmatter derived. + test('stateJsonDerivesCompletedPhasesFromRoadmapAuthority', (t) => { + const cwd = createTempDir('gsd-4129-jsonfloor-'); + t.after(() => cleanup(cwd)); + buildStaleVerificationFixture(cwd, 3, false); + + const result = runGsdTools(['state', 'json'], cwd); + assert.ok(result.success, `state json failed: ${result.error}`); + const reported = JSON.parse(result.output).progress; + assert.strictEqual( + Number(reported.completed_phases), + 3, + `#4129: the derived completed_phases must agree with the ROADMAP Complete rows (3), not the stale-verification disk count (${reported.completed_phases})`, + ); + assert.strictEqual(Number(reported.percent), 17, '#4129: percent derives from the floored numerator'); + }); + + // Row 3 — the flip side: a counter STUCK LOW (the issue's row-1 aftermath) + // must move UP to the ROADMAP truth on the next write, not stay pinned. + test('stuckLowCounterIncrementsToRoadmapTruthOnNextWrite', (t) => { + const cwd = createTempDir('gsd-4129-stucklow-'); + t.after(() => cleanup(cwd)); + buildStaleVerificationFixture(cwd, 2); + + const result = runGsdTools(['state', 'record-session', '--stopped-at', 'x'], cwd); + assert.ok(result.success, `record-session failed: ${result.error}`); + const { progress } = readProgress(cwd); + assert.strictEqual(Number(progress.completed_phases), 3, '#4129: a stored 2 below the ROADMAP truth must rise to 3, not be re-derived as 2'); + assert.strictEqual(Number(progress.percent), 17, '#4129: percent follows the corrected numerator'); + }); + + // Row 5 — negative space: no canonical Progress table (checklist-only + // ROADMAP) → deriveProgressFromRoadmap resolves no table → the floor is + // inert and the disk-verification count stands. The floor must not invent + // a parser for checklist bullets (one-owner rule). + test('floorIsInertWithoutCanonicalProgressTable', (t) => { + const cwd = createTempDir('gsd-4129-notable-'); + t.after(() => cleanup(cwd)); + buildStaleVerificationFixture(cwd, 2, false); + + // Rewrite the ROADMAP with the table stripped — checklist only. + // CRLF-tolerant line split (local/no-crlf-fragile-split). + const roadmapPath = path.join(cwd, '.planning', 'ROADMAP.md'); + const withoutTable = fs + .readFileSync(roadmapPath, 'utf-8') + .split(/\r?\n/) + .filter((line) => !line.trimStart().startsWith('|')) + .join('\n'); + fs.writeFileSync(roadmapPath, withoutTable); + + const result = runGsdTools(['state', 'json'], cwd); + assert.ok(result.success, `state json failed: ${result.error}`); + const reported = JSON.parse(result.output).progress; + assert.strictEqual( + Number(reported.completed_phases), + 2, + '#4129 negative space: without a canonical Progress table the completed count stays the disk-verification count (the floor reuses deriveProgressFromRoadmap, which reads only the table)', + ); + }); +}); + // ═════════════════════════════════════════════════════════════════════════ // #3872 / ADR-3473 §8.7: what a command reports it wrote // (.gsd/phase/feat-3872-transaction-diff-reporting/{40-design,50-test-matrix}.md) From c3e2da153bbddb0f12ff3c1aa2ad8ff8dcf3477d Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sat, 5 Sep 2026 23:31:35 -0400 Subject: [PATCH 004/166] fix(#4134): refuse punctuation-only milestone heading names (#4358) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4134): fail-first regression — refuse punctuation-fragment milestone names A first-milestone ROADMAP.md H1 that puts the version after the name (# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's own version token, which the ADR-3180 §7.2 pinned name rule returns as a COMPLETE-scope milestone name. Failing-first coverage: - getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only fallback) must yield TRUNCATED {version, name: null}, never ')' - the refusal is level-agnostic (H2/H3) - punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only) - listMilestoneHeadings enumerates the heading with name: null - init manager CLI reports milestone_name: null and no lone ')' anywhere - property (seed 20260905, 300 runs): a word-char remainder is always a name, a punctuation-only remainder never is - negative space: canonical delimiter forms, parenthetical names (#3171), trailing markers, digit-only names, CRLF headings, version-last-no-parens control * fix(#4134): refuse punctuation-only milestone heading names extractMilestoneHeadingName returns everything after the heading's own version token as the name (ADR-3180 §7.2 pinned rule), which assumes version-then-name. A name-then-version heading — the H1 a first-ever ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves exactly ')' after the token, and that fragment was returned as a COMPLETE-scope milestone name, propagating into init.* JSON output and buildStateFrontmatter's STATE.md writes. A remainder with no letter or digit anywhere (any script) is heading structure, not a curated name: refuse it as name: null so callers report the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names that merely contain punctuation are unaffected — '(' stays an ordinary name character (#3171) — and digit-only names qualify. Also closes the template gap that lets the shape occur: the roadmapper agent's output_formats now templates the version-free canonical H1 ('# Roadmap: [Project Name]', per templates/roadmap.md) instead of leaving a first milestone's title line to invention. The new section shifts the file's existing bare-gsd-tools prose mention from line 647 to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134) * chore(#4134): add changeset * chore(#4134): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/daring-hawks-zip.md | 5 + CONTEXT.md | 2 +- agents/gsd-roadmapper.md | 13 ++ src/roadmap-parser.cts | 20 ++- tests/init-manager.test.cjs | 49 ++++++ tests/milestone-window-single-owner.test.cjs | 32 ++++ ...o-bare-gsd-tools-command-position.test.cjs | 2 +- tests/roadmap-parser.test.cjs | 162 ++++++++++++++++++ 8 files changed, 282 insertions(+), 3 deletions(-) create mode 100644 .changeset/daring-hawks-zip.md diff --git a/.changeset/daring-hawks-zip.md b/.changeset/daring-hawks-zip.md new file mode 100644 index 000000000..ba3a47243 --- /dev/null +++ b/.changeset/daring-hawks-zip.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4358 +--- +**`milestone_name` no longer corrupts to ")" for a first-milestone ROADMAP whose H1 puts the version after the name** — a punctuation-only heading remainder (e.g. the closing paren of `# Roadmap: Project — Name (v1.13)`) is refused as a name, so `init.*` output reports `null` instead of garbage, and the roadmapper agent now templates the canonical version-free H1. (#4134) diff --git a/CONTEXT.md b/CONTEXT.md index 941250fc2..dcb451fef 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -260,7 +260,7 @@ Canonical GFM table parsing + schema registry seam (`gsd-core/bin/lib/markdown-t Shared fail-loud `Result` and per-surface write-set contracts (`gsd-core/bin/lib/write-set.cjs`; ADR-2143 §5/§6, epic #2143). Pure, Node built-ins only, no I/O. Exports: `Result` (`{ok:true,value}\|{ok:false,reason}` — ADR-2143 §5 fail-loud parse shape, never a bare `null` a caller can mistake for "empty but fine"; the single source of truth `markdown-table.cjs` re-exports so its existing importers are unaffected; deliberately distinct from command-routing-hub's dispatch `Result` `{ok,data\|kind}`); `WriteOutcome` (`{surface: string, applied: boolean}` — one surface's outcome within a multi-surface write); `WriteSet` (`WriteOutcome[]`); `writeSetComplete(ws) → boolean` (true only when the set is non-empty AND every surface applied — ADR-2143 §6's "no OR-into-one-flag" rule: a command that mutates more than one surface must not collapse independent surface outcomes into a single boolean, the anti-pattern that let a checkbox-only partial write (#2140) report full success). `milestone.cts`'s `requirements mark-complete` handler is the first consumer: it reports a `write_set` (`checkbox`/`traceability` surfaces) and `write_set_complete` alongside its existing `updated`/`marked_complete`/`already_complete`/`not_found`/`table_unmatched` fields, which remain computed exactly as before — the write-set is additive, structured ADR-2143 documentation of the same per-surface facts #2140's tactical fix already exposed via `table_unmatched`. ### Roadmap Parser Module -Module owning ROADMAP.md parsing: shipped-milestone slicing, current-milestone extraction, milestone/phase lookups, and milestone-phase filtering (`stripShippedMilestones`, `extractCurrentMilestone`, `replaceInCurrentMilestone`, `getRoadmapPhaseInternal`, `getMilestoneInfo`, `getMilestonePhaseFilter`, `isMilestoneShippedInRoadmap`, `withPhaseSection`). Milestone shipped/active heading classification is owned here (#2562): `isMilestoneShippedInRoadmap(content, version)` answers "does the ROADMAP mark THIS milestone shipped" from heading and `` lines only — never a bullet that merely names the version — with the version token boundary-matched so `v2.0` does not match inside `v2.0.1`. `extractCurrentMilestone` and `getMilestonePhaseFilter` take an optional trailing workstream name so their `planningDir` resolution targets `.planning/workstreams//`; omitted, it resolves exactly as before (including the `GSD_WORKSTREAM` fallback). `getMilestonePhaseFilter` exposes `versionScoped`, true only when the returned phase set really is one milestone's — consumers must not read `phaseCount` as a current-milestone denominator otherwise — and `versionSectionFound`, true whenever the requested version's section was located at all. The two differ precisely for a located-but-EMPTY section: it falls through to the zero-count pass-all degrade, which resets `versionScoped` to false, leaving `versionSectionFound` the only surviving evidence that the milestone exists rather than being absent. `missingExplicitVersion` covers the complementary shape (versioned roadmap, no section for this version). `withPhaseSection(content, phaseId, edit)` resolves a phase's `### Phase N` detail-section heading via the #2121 phase-id source (`phaseMarkdownRegexSource`) and delegates to the markdown-sectionizer seam's `withSection`, so a per-phase ROADMAP edit is bounded to that phase's own section (ADR-2143 §4). Depends only on leaf modules (`phase-id`, `planning-workspace`, `shell-command-projection`, `markdown-sectionizer`, and — since #1881 — `unusable-input` for the out-of-band diagnostic) — no `loadConfig`, no other core dependency. An unreadable ROADMAP.md is reported rather than collapsed into the same sentinel as a genuinely absent one; absence itself stays silent, and neither lookup gains a throw (the #2245 audit records that `src/state.cts` removed its defensive try/catch on the strength of `getMilestoneInfo` never throwing). Milestone WINDOWING — which headings bound a milestone — is owned here as of #3184 (epic #3180 Phase 2, ADR-3180 Decision 1): `computeMilestoneSectionEnd` (the section-end walk, formerly duplicated as two distinct nested `computeSectionEnd` functions plus an inline third copy in `getMilestonePhaseFilter`'s `versionOverride` branch), `locateMilestoneHeadings` (heading location, version token boundary-matched with `\b`, **NOT** the stricter `(?![\w.-])`: `v2.0` therefore DOES match inside `v2.0.1`, and a milestone STATE of `v8.0` legitimately selects a live `## v8.0-B …` over a closed `v8.0-A` sibling — deliberate, load-bearing #730 behavior that ADR-3180 Amendment 2 tried to tighten and then reverted; the earlier text here described that reverted alternative as if it had shipped, corrected by #3216), `listMilestoneHeadings` (#3216 — the version-AGNOSTIC enumeration of every milestone heading in document order, sharing ONE grammar source with `locateMilestoneHeadings` so the two cannot drift; `locateMilestoneHeadings` is now a version-filtered view over it rather than a second expression of the pattern). Milestone IDENTITY — which milestone is current and what it is CALLED — is owned by `getMilestoneInfo`, which since #3216 binds to that same locator instead of its own heading regexes and returns a `ScopedResult`: a name retains parentheses and drops a trailing `✅`/`📋`/`🚧` marker, a `### Phase N …` heading is never the milestone heading (#3197), and an identity that cannot be determined returns a non-`COMPLETE` scope rather than the former `{version:'v1.0', name:'milestone'}` default, which was output-identical to a successful read of a genuine v1.0 project. `buildStateFrontmatter` and `archivePhaseDirectories` branch on that scope, so a fabricated identity is never persisted to `STATE.md` nor used as a `milestones/-phases/` path component, `sliceMilestoneWindow` (the one composition of locate → prefer-non-closed → section-end, so a consumer cannot re-assemble its own window from the primitives), and `isMilestoneBoundedInRoadmap` (the named predicate replacing two byte-identical re-derivations in `state.cts`). `hasMilestoneSectioning(content)` is the sibling predicate answering "could a whole-document phase count conflate two different milestones?" — and since #3185 it is decided by milestone VOCABULARY, not by heading position: a heading is a milestone heading iff it is a non-Phase heading (level 1-3) carrying a version token, a shipped/active marker, or the word `Milestone`, and sectioning means two or more of them, since one section cannot conflate siblings. Since #3642, `buildStateFrontmatter`'s flat-vs-milestoned gate consumes the >=1 sibling `hasAnyMilestoneSection` over the same `countMilestoneHeadings` walk instead: asserted-vs-section is a different question than sibling conflation, and needs only ONE section to go wrong — with exactly one section and an asserted milestone absent from the ROADMAP, the whole-document count IS that foreign section's phases, so the #3354 withhold (stored value preserved + warning) governs; a zero-signal (genuinely flat) roadmap keeps the whole-document count per #2828. Three position-based models were tried and each shipped a defect — "any non-Phase heading" over-detects, so a flat ROADMAP carrying an ordinary `## Progress` was called sectioned and its declared phase count discarded for the on-disk directory count (#3204, #2828 regressing at 1.9.1 via #3184's own consolidation); strict nesting misses same-level siblings (regressing #1761) and false-positives on the bundled `templates/roadmap.md` shape, where a `## Phases` wrapper holds a single nested milestone; adjacency false-positives whenever a structural heading merely precedes a phase heading. Known limit: two milestone sections carrying none of the three signals are not detected. `extractCurrentMilestoneScoped` is the real extractor and returns the Planning Scope Module's `ScopedResult`; `extractCurrentMilestone` remains a one-line wrapper over `.value` because its blast radius is CRITICAL (200+ affected symbols, 20 direct callers) and its signature must not move. `getMilestonePhaseFilter` gains a `scope` field: its pass-all degrade is PRESERVED where its premise holds (a genuinely-empty, freshly-declared milestone reports `SCOPE.COMPLETE`) and is now labelled where it does not (`SCOPE.TRUNCATED` when the window reached no phase entries while the document has them), so the destructive consumer — `milestone.complete`, which MOVES phase directories — can refuse instead of archiving every phase directory on disk (#3166). The filter's function behavior is deliberately unchanged: making it deny-all on a non-COMPLETE scope would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. `findRoadmapProgressTable(content)` (#1956) locates the `## Progress` table — scoped to that heading via the markdown-sectionizer seam, falling back to the whole document for a headingless milestone slice — so a differently-headed table sharing the `Phase | Plans Complete | Status | Completed` columns cannot be read instead (the #2012 decoy class). `phase-lifecycle.cts`'s `deriveProgressFromRoadmap` expresses the same scope independently for the completion RATIO; the two are held in agreement by a parity test rather than by a shared call, because that symbol's blast radius does not justify a refactor. Extracted from the Core module per ADR-857 rollout phase 2b (#870), resolving the ROADMAP.md parse/write straddle so the Roadmap module (`roadmap.cjs`, which owns ROADMAP.md mutation) imports parsing directly instead of through Core; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/roadmap-parser.cjs` (generated from `src/roadmap-parser.cts`). +Module owning ROADMAP.md parsing: shipped-milestone slicing, current-milestone extraction, milestone/phase lookups, and milestone-phase filtering (`stripShippedMilestones`, `extractCurrentMilestone`, `replaceInCurrentMilestone`, `getRoadmapPhaseInternal`, `getMilestoneInfo`, `getMilestonePhaseFilter`, `isMilestoneShippedInRoadmap`, `withPhaseSection`). Milestone shipped/active heading classification is owned here (#2562): `isMilestoneShippedInRoadmap(content, version)` answers "does the ROADMAP mark THIS milestone shipped" from heading and `` lines only — never a bullet that merely names the version — with the version token boundary-matched so `v2.0` does not match inside `v2.0.1`. `extractCurrentMilestone` and `getMilestonePhaseFilter` take an optional trailing workstream name so their `planningDir` resolution targets `.planning/workstreams//`; omitted, it resolves exactly as before (including the `GSD_WORKSTREAM` fallback). `getMilestonePhaseFilter` exposes `versionScoped`, true only when the returned phase set really is one milestone's — consumers must not read `phaseCount` as a current-milestone denominator otherwise — and `versionSectionFound`, true whenever the requested version's section was located at all. The two differ precisely for a located-but-EMPTY section: it falls through to the zero-count pass-all degrade, which resets `versionScoped` to false, leaving `versionSectionFound` the only surviving evidence that the milestone exists rather than being absent. `missingExplicitVersion` covers the complementary shape (versioned roadmap, no section for this version). `withPhaseSection(content, phaseId, edit)` resolves a phase's `### Phase N` detail-section heading via the #2121 phase-id source (`phaseMarkdownRegexSource`) and delegates to the markdown-sectionizer seam's `withSection`, so a per-phase ROADMAP edit is bounded to that phase's own section (ADR-2143 §4). Depends only on leaf modules (`phase-id`, `planning-workspace`, `shell-command-projection`, `markdown-sectionizer`, and — since #1881 — `unusable-input` for the out-of-band diagnostic) — no `loadConfig`, no other core dependency. An unreadable ROADMAP.md is reported rather than collapsed into the same sentinel as a genuinely absent one; absence itself stays silent, and neither lookup gains a throw (the #2245 audit records that `src/state.cts` removed its defensive try/catch on the strength of `getMilestoneInfo` never throwing). Milestone WINDOWING — which headings bound a milestone — is owned here as of #3184 (epic #3180 Phase 2, ADR-3180 Decision 1): `computeMilestoneSectionEnd` (the section-end walk, formerly duplicated as two distinct nested `computeSectionEnd` functions plus an inline third copy in `getMilestonePhaseFilter`'s `versionOverride` branch), `locateMilestoneHeadings` (heading location, version token boundary-matched with `\b`, **NOT** the stricter `(?![\w.-])`: `v2.0` therefore DOES match inside `v2.0.1`, and a milestone STATE of `v8.0` legitimately selects a live `## v8.0-B …` over a closed `v8.0-A` sibling — deliberate, load-bearing #730 behavior that ADR-3180 Amendment 2 tried to tighten and then reverted; the earlier text here described that reverted alternative as if it had shipped, corrected by #3216), `listMilestoneHeadings` (#3216 — the version-AGNOSTIC enumeration of every milestone heading in document order, sharing ONE grammar source with `locateMilestoneHeadings` so the two cannot drift; `locateMilestoneHeadings` is now a version-filtered view over it rather than a second expression of the pattern). Milestone IDENTITY — which milestone is current and what it is CALLED — is owned by `getMilestoneInfo`, which since #3216 binds to that same locator instead of its own heading regexes and returns a `ScopedResult`: a name retains parentheses and drops a trailing `✅`/`📋`/`🚧` marker, a heading remainder with no letter or digit anywhere is heading structure rather than a curated name and is refused as `name: null`, so a name-then-version H1 (`# Roadmap: Project — Name (v1.13)` — the shape a first-ever ROADMAP.md drifts into) reports the rule-6 TRUNCATED identity instead of the literal `)` left after the version token (#4134), a `### Phase N …` heading is never the milestone heading (#3197), and an identity that cannot be determined returns a non-`COMPLETE` scope rather than the former `{version:'v1.0', name:'milestone'}` default, which was output-identical to a successful read of a genuine v1.0 project. `buildStateFrontmatter` and `archivePhaseDirectories` branch on that scope, so a fabricated identity is never persisted to `STATE.md` nor used as a `milestones/-phases/` path component, `sliceMilestoneWindow` (the one composition of locate → prefer-non-closed → section-end, so a consumer cannot re-assemble its own window from the primitives), and `isMilestoneBoundedInRoadmap` (the named predicate replacing two byte-identical re-derivations in `state.cts`). `hasMilestoneSectioning(content)` is the sibling predicate answering "could a whole-document phase count conflate two different milestones?" — and since #3185 it is decided by milestone VOCABULARY, not by heading position: a heading is a milestone heading iff it is a non-Phase heading (level 1-3) carrying a version token, a shipped/active marker, or the word `Milestone`, and sectioning means two or more of them, since one section cannot conflate siblings. Since #3642, `buildStateFrontmatter`'s flat-vs-milestoned gate consumes the >=1 sibling `hasAnyMilestoneSection` over the same `countMilestoneHeadings` walk instead: asserted-vs-section is a different question than sibling conflation, and needs only ONE section to go wrong — with exactly one section and an asserted milestone absent from the ROADMAP, the whole-document count IS that foreign section's phases, so the #3354 withhold (stored value preserved + warning) governs; a zero-signal (genuinely flat) roadmap keeps the whole-document count per #2828. Three position-based models were tried and each shipped a defect — "any non-Phase heading" over-detects, so a flat ROADMAP carrying an ordinary `## Progress` was called sectioned and its declared phase count discarded for the on-disk directory count (#3204, #2828 regressing at 1.9.1 via #3184's own consolidation); strict nesting misses same-level siblings (regressing #1761) and false-positives on the bundled `templates/roadmap.md` shape, where a `## Phases` wrapper holds a single nested milestone; adjacency false-positives whenever a structural heading merely precedes a phase heading. Known limit: two milestone sections carrying none of the three signals are not detected. `extractCurrentMilestoneScoped` is the real extractor and returns the Planning Scope Module's `ScopedResult`; `extractCurrentMilestone` remains a one-line wrapper over `.value` because its blast radius is CRITICAL (200+ affected symbols, 20 direct callers) and its signature must not move. `getMilestonePhaseFilter` gains a `scope` field: its pass-all degrade is PRESERVED where its premise holds (a genuinely-empty, freshly-declared milestone reports `SCOPE.COMPLETE`) and is now labelled where it does not (`SCOPE.TRUNCATED` when the window reached no phase entries while the document has them), so the destructive consumer — `milestone.complete`, which MOVES phase directories — can refuse instead of archiving every phase directory on disk (#3166). The filter's function behavior is deliberately unchanged: making it deny-all on a non-COMPLETE scope would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. `findRoadmapProgressTable(content)` (#1956) locates the `## Progress` table — scoped to that heading via the markdown-sectionizer seam, falling back to the whole document for a headingless milestone slice — so a differently-headed table sharing the `Phase | Plans Complete | Status | Completed` columns cannot be read instead (the #2012 decoy class). `phase-lifecycle.cts`'s `deriveProgressFromRoadmap` expresses the same scope independently for the completion RATIO; the two are held in agreement by a parity test rather than by a shared call, because that symbol's blast radius does not justify a refactor. Extracted from the Core module per ADR-857 rollout phase 2b (#870), resolving the ROADMAP.md parse/write straddle so the Roadmap module (`roadmap.cjs`, which owns ROADMAP.md mutation) imports parsing directly instead of through Core; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/roadmap-parser.cjs` (generated from `src/roadmap-parser.cts`). ### Core Utilities Module Module owning the shared low-level utility primitives extracted from Core: POSIX path normalization (`toPosixPath`), filesystem scanning (`detectSubRepos`, `readSubdirectories`, `getPhaseFileStats`, `pathExistsInternal`), plan/summary pairing helpers (`countMatchedSummaries`, `findUnsummarizedPlans`, `findOrphanSummaries`), and small pure helpers (`generateSlugInternal`, `extractOneLinerFromBody`, `extractCanonicalPlanId`, `timeAgo`). `filterPlanFiles`/`filterSummaryFiles` were retired by #3183 (ADR-3180 Decision 2) — `getPhaseFileStats` no longer re-derives plan/summary filename matching locally; it now sources `plans`/`summaries` (plus a `scope` field, `COMPLETE`/`TRUNCATED`/`UNREADABLE`) directly from `scanPhasePlans` (`src/plan-scan.cts`), the single owner of live-plan counting. Depends only on Node built-ins and already-leafed modules (`phase-id` for `comparePhaseNum`, `planning-workspace` for `findContextMdIn`) — no `loadConfig`, no other core dependency. Extracted from the Core module per ADR-857 rollout phase 2c (#877) as the shared leaf that unblocks the phase-locator fs-search extraction (2d); the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/core-utils.cjs` (generated from `src/core-utils.cts`). diff --git a/agents/gsd-roadmapper.md b/agents/gsd-roadmapper.md index b2e28b0ac..4a219fa85 100644 --- a/agents/gsd-roadmapper.md +++ b/agents/gsd-roadmapper.md @@ -331,6 +331,19 @@ After roadmap creation, REQUIREMENTS.md gets updated with phase mappings: **CRITICAL: ROADMAP.md requires TWO phase representations. Both are mandatory.** +### 0. Top-Level Title (H1) + +The H1 carries the PROJECT name only — never a version and never a milestone name: + +```markdown +# Roadmap: [Project Name] +``` + +Milestone identity (version + name) lives in milestone headings (`## vX.Y — [Name]`) or +`## Milestones` bullets (`🚧 **vX.Y [Name]**`), never in the H1. A trailing version in the +H1 (`# Roadmap: [Project] — [Name] (vX.Y)`) corrupts milestone-name extraction (#4134). +`~/.claude/gsd-core/templates/roadmap.md` is the canonical shape. + ### 1. Summary Checklist (under `## Phases`) Use the form matching `phase_id_convention` from config. diff --git a/src/roadmap-parser.cts b/src/roadmap-parser.cts index c7daca942..d4ae05690 100644 --- a/src/roadmap-parser.cts +++ b/src/roadmap-parser.cts @@ -1589,6 +1589,13 @@ function stripLeadingDelimiter(s: string): string { * one implementation. Returns `null` when `headingText` carries no version * token at all (e.g. a non-milestone heading reached this by mistake). * + * #4134 (§7.2 rule 6 floor): the rule's direction assumes version-then-name. + * A name-then-version heading (`# Roadmap: Project — Name (v1.13)`) leaves a + * punctuation fragment (`)`) after the token; a remainder with no letter or + * digit anywhere is heading structure, not a curated name, and is refused as + * `name: null` so callers report the honest rule-6 answer instead of + * fabricating garbage. Names that merely CONTAIN punctuation are unaffected. + * * @param expectedVersion - When the caller already knows the exact version it * is looking for (the STATE-anchored `getMilestoneInfo` path, which located * this heading via `selectMilestoneHeading(roadmap, stateVersion)`), pass it @@ -1627,7 +1634,18 @@ function extractMilestoneHeadingName( // whitespace — the marker is already carried structurally by `closed`, so // duplicating it inside `name` (e.g. "Old ✅") is redundant and wrong. Only // these three markers, only at the end; a marker inside a name is untouched. - const name = stripLeadingDelimiter(afterVersion).replace(/\s*(?:[✅📋🚧]\s*)+$/, '') || null; + const candidate = stripLeadingDelimiter(afterVersion).replace(/\s*(?:[✅📋🚧]\s*)+$/, '') || null; + // #4134 (§7.2 rule 6 floor): a "name" with no letter or digit anywhere is + // heading structure, not a curated name. The pinned rule takes everything + // AFTER the version token, so a name-then-version heading (`# Roadmap: + // Project — Name (v1.13)` — the shape a first-ever ROADMAP.md drifts into) + // leaves exactly `)` there, which used to be returned as a COMPLETE-scope + // name and propagated into init.* output and STATE.md. Refuse it: callers + // already report the honest rule-6 answer (version kept, `name: null`, + // scope TRUNCATED) for an unresolvable name. A name that merely CONTAINS + // punctuation is untouched — `(` is an ordinary name character (#3171) — + // and digits alone qualify (`## v4.0 — 42` is the name `42`). + const name = candidate !== null && /[\p{L}\p{N}]/u.test(candidate) ? candidate : null; return { version, name }; } diff --git a/tests/init-manager.test.cjs b/tests/init-manager.test.cjs index cebfcedeb..ff0e0f14a 100644 --- a/tests/init-manager.test.cjs +++ b/tests/init-manager.test.cjs @@ -495,6 +495,55 @@ describe('init manager', () => { assert.strictEqual(output.phases[0].is_active, false); }); + // #4134 — a first-ever ROADMAP.md whose H1 puts the version after the name + // (`# Roadmap: — (v1.13)` — the shape nothing + // templates for a project's first milestone) used to make the §7.2 pinned + // name rule return the literal `)` left after the version token as the + // milestone's "name", and init manager displayed it verbatim. The refusal + // reports the honest ADR-3180 §7.2 rule-6 answer instead: the version is + // real, the name is not resolvable from that heading, milestone_name is + // null — never a punctuation fragment. + test('#4134 — milestone_name is null, never ")", for a name-then-version H1', () => { + fs.writeFileSync( + path.join(tmpDir, '.planning', 'ROADMAP.md'), + [ + '# Roadmap: GSD Core — Native OMP Runtime Support (v1.13)', + '', + '## Phases', + '', + '- [ ] **Phase 1: Runtime Adapter Interface**', + '', + '## Phase Details', + '', + '### Phase 1: Runtime Adapter Interface', + '', + '**Goal:** Define the adapter contract.', + '', + ].join('\n') + ); + fs.writeFileSync( + path.join(tmpDir, '.planning', 'STATE.md'), + [ + '---', + 'gsd_state_version: 1.0', + 'milestone: v1.13', + 'milestone_name: Native OMP Runtime Support', + 'status: planning', + '---', + '', + '# Project State', + '', + ].join('\n') + ); + + const result = runGsdTools('init manager', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + const output = JSON.parse(result.output); + assert.strictEqual(output.milestone_version, 'v1.13'); + assert.strictEqual(output.milestone_name, null); + assert.ok(!result.output.includes('")"'), 'a lone ")" value must not appear anywhere in the JSON'); + }); + // #3314 — ADR-456 subprocess clock pin: cmdInitManager's `is_active` gate // (nowMs - newestMtime < 300000) is CLI-subprocess tested, so only the // GSD_TEST_MODE+GSD_NOW_MS pin can control "now" here (t.mock.timers in the diff --git a/tests/milestone-window-single-owner.test.cjs b/tests/milestone-window-single-owner.test.cjs index 21638cf3f..0ec4837ba 100644 --- a/tests/milestone-window-single-owner.test.cjs +++ b/tests/milestone-window-single-owner.test.cjs @@ -2332,6 +2332,38 @@ test('listMilestoneHeadingsHeadingFieldIsCrlfFreeAndTrimmed', () => { assert.strictEqual(result[1].name, 'Old'); }); +// (RED) #4134: the §7.2 pinned rule takes everything after the heading's OWN +// version token as the name, so a name-then-version heading (`… (v1.13)`) -- +// the shape a first-ever ROADMAP.md drifts into when nothing templates its H1 +// -- leaves exactly `)` after the token, which used to be enumerated as the +// milestone's "name". ADR-3180 §7.2 rule 6: a remainder with no letter or +// digit anywhere is heading structure, not a curated name -- enumerate it as +// `name: null` and let consumers report the honest TRUNCATED identity. The +// enumeration itself (which headings, which version, closed status) is +// unchanged; only the garbage "name" is refused. +test('listMilestoneHeadingsRefusesPunctuationOnlyNames', () => { + const content = [ + '# Roadmap: GSD Core — Native OMP Runtime Support (v1.13)', + '', + '### Phase 1: Runtime Adapter Interface', + ].join('\n'); + const result = listMilestoneHeadings(content); + assert.strictEqual(result.length, 1); + assert.strictEqual(result[0].version, 'v1.13'); + assert.strictEqual(result[0].closed, false); + assert.strictEqual(result[0].name, null, `name must be null, not ${JSON.stringify(result[0].name)}`); +}); + +// (RED) #4134 negative space: a name that merely CONTAINS punctuation is a +// name -- `(` is an ordinary name character and never a terminator (#3171). +// Only a remainder with zero word characters is refused. +test('listMilestoneHeadingsKeepsNamesThatContainPunctuation', () => { + const content = ['# Roadmap', '', '## v3.3 — Name (Part 2: Revenge)'].join('\n'); + const result = listMilestoneHeadings(content); + assert.strictEqual(result.length, 1); + assert.strictEqual(result[0].name, 'Name (Part 2: Revenge)'); +}); + test('listMilestoneHeadingsRespectsTheOneToThreeLevelBound', () => { const content = [ '# v1.0 — Level One', diff --git a/tests/no-bare-gsd-tools-command-position.test.cjs b/tests/no-bare-gsd-tools-command-position.test.cjs index 355272893..7bfee1888 100644 --- a/tests/no-bare-gsd-tools-command-position.test.cjs +++ b/tests/no-bare-gsd-tools-command-position.test.cjs @@ -105,7 +105,7 @@ const BARE_COMMAND_RE = new RegExp( const PROSE_ALLOWLIST = [ { file: 'agents/gsd-executor.md', line: 823, reason: 'describes the SDK return envelope of `gsd-tools query commit`; not an instruction to run the bare word' }, { file: 'agents/gsd-phase-researcher.md', line: 33, reason: 'package-legitimacy provenance rule names the command as the source of an OK verdict; descriptive' }, - { file: 'agents/gsd-roadmapper.md', line: 647, reason: 'parenthetical "e.g." naming SDK queries a user *could* run; not an agent instruction' }, + { file: 'agents/gsd-roadmapper.md', line: 660, reason: 'parenthetical "e.g." naming SDK queries a user *could* run; not an agent instruction (#4134 shifted it from 647: the H1 template section added above moved the line, the mention is unchanged)' }, { file: 'agents/gsd-intel-updater.md', line: 40, reason: 'cross-platform note names the `gsd-tools intel ` CLI surface descriptively ("CLI invocations go through..."); not an agent instruction' }, { file: 'gsd-core/workflows/execute-plan.md', line: 419, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' }, ]; diff --git a/tests/roadmap-parser.test.cjs b/tests/roadmap-parser.test.cjs index f3ab36fc0..6712ce9f5 100644 --- a/tests/roadmap-parser.test.cjs +++ b/tests/roadmap-parser.test.cjs @@ -949,6 +949,168 @@ describe('roadmap-parser: getMilestoneInfo #2135 — milestone_name clobber', () }); }); +// ─── getMilestoneInfo — #4134 punctuation-fragment name refusal ─────────────── +// The §7.2 pinned rule takes everything AFTER the heading's own version token +// as the name. For a name-then-version heading (`# Roadmap: Project — Name +// (v1.13)` — the shape a first-ever ROADMAP.md drifts into, since nothing +// templates its H1) that remainder is literally `)`, which the rule used to +// return as a COMPLETE-scope "name". ADR-3180 §7.2 rule 6 is the floor this +// violates: a version known but a name unresolvable is TRUNCATED carrying +// `name: null` — a punctuation-only remainder is heading structure, not a +// curated name (#4134). + +describe('roadmap-parser: getMilestoneInfo #4134 — name-then-version heading', () => { + let tmpDir; + + beforeEach(() => { tmpDir = createTempProject(); }); + afterEach(() => { cleanup(tmpDir); }); + + test('#4134 — name-then-version H1 never yields a punctuation-fragment name (rule 6: TRUNCATED, name null)', () => { + writeState(tmpDir, { milestone: 'v1.13' }); + writeRoadmap(tmpDir, [ + '# Roadmap: GSD Core — Native OMP Runtime Support (v1.13)', + '', + '### Phase 1: Runtime Adapter Interface', + ].join('\n')); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.TRUNCATED, `scope: ${JSON.stringify(info)}`); + assert.strictEqual(info.value.version, 'v1.13'); + assert.strictEqual(info.value.name, null); + }); + + test('#4134 — ROADMAP-only fallback path also refuses the ")" fragment', () => { + // No STATE.md: the first open milestone heading supplies the version. + writeRoadmap(tmpDir, [ + '# Roadmap: GSD Core — Native OMP Runtime Support (v1.13)', + '', + '### Phase 1: Runtime Adapter Interface', + ].join('\n')); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.TRUNCATED, `scope: ${JSON.stringify(info)}`); + assert.strictEqual(info.value.version, 'v1.13'); + assert.strictEqual(info.value.name, null); + }); + + test('#4134 — the refusal is level-agnostic (H2/H3 carry the same fragment)', () => { + for (const [level, heading] of [ + [2, '## Native OMP Runtime Support (v1.13)'], + [3, '### Native OMP Runtime Support (v1.13)'], + ]) { + writeState(tmpDir, { milestone: 'v1.13' }); + writeRoadmap(tmpDir, `${heading}\n\n### Phase 1: Setup\n`); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.TRUNCATED, `H${level}: ${JSON.stringify(info)}`); + assert.strictEqual(info.value.version, 'v1.13'); + assert.strictEqual(info.value.name, null); + } + }); + + test('#4134 — every punctuation-only remainder is refused (garbage family)', () => { + // Each fragment survives stripLeadingDelimiter (it does not START with a + // delimiter char) and carries no letter or digit anywhere — the exact + // shape that used to be returned as a "name". + const fragments = [')', '()', '**', '.,;:', ']}', '🎉']; + for (const fragment of fragments) { + writeState(tmpDir, { milestone: 'v1.2' }); + writeRoadmap(tmpDir, `## v1.2 — ${fragment}\n\n### Phase 1: Setup\n`); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.TRUNCATED, `fragment ${JSON.stringify(fragment)}: ${JSON.stringify(info)}`); + assert.strictEqual(info.value.version, 'v1.2'); + assert.strictEqual(info.value.name, null, `fragment ${JSON.stringify(fragment)} must not become a name`); + } + }); + + test('#4134 control — version-last without parens was already name:null and stays so', () => { + writeState(tmpDir, { milestone: 'v1.2.3' }); + writeRoadmap(tmpDir, '# Milestone Name v1.2.3\n\n### Phase 1: Setup\n'); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.value.version, 'v1.2.3'); + assert.strictEqual(info.value.name, null); + assert.strictEqual(info.scope, SCOPE.TRUNCATED); + }); + + test('#4134 negative space — canonical delimiter forms parse identically', () => { + const cases = [ + ['## v2.0: The Big Launch', 'The Big Launch'], + ['## v2.5 — Galaxy Release', 'Galaxy Release'], + ['## v2.6 – En Dash Form', 'En Dash Form'], + ['## v2.7 - Hyphen Form', 'Hyphen Form'], + ['## v2.8 Space Only Form', 'Space Only Form'], + ]; + for (const [heading, expected] of cases) { + writeState(tmpDir, { milestone: heading.match(/v\d+(?:\.\d+)*/)[0] }); + writeRoadmap(tmpDir, `${heading}\n\n### Phase 1: Setup\n`); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.COMPLETE, `${heading}: ${JSON.stringify(info)}`); + assert.strictEqual(info.value.name, expected, `${heading}: ${JSON.stringify(info.value)}`); + } + }); + + test('#4134 negative space — parenthetical names are retained (#3171)', () => { + writeState(tmpDir, { milestone: 'v1.2' }); + writeRoadmap(tmpDir, '## v1.2 — Name (Part 2)\n\n### Phase 1: Setup\n'); + const info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.COMPLETE); + assert.strictEqual(info.value.name, 'Name (Part 2)'); + }); + + test('#4134 negative space — markers, digit-only names, CRLF headings unchanged', () => { + // Trailing status marker still stripped, not treated as a "name" (a ✅ + // TRAILING marker would make the heading closed and skipped — 📋 does not). + writeState(tmpDir, { milestone: 'v3.0' }); + writeRoadmap(tmpDir, '## v3.0 — Planned 📋\n\n### Phase 1: Setup\n'); + let info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.COMPLETE); + assert.strictEqual(info.value.name, 'Planned'); + + // A digit-only name IS a name (\p{N} counts as a word character). + writeState(tmpDir, { milestone: 'v4.0' }); + writeRoadmap(tmpDir, '## v4.0 — 42\n\n### Phase 1: Setup\n'); + info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.COMPLETE); + assert.strictEqual(info.value.name, '42'); + + // CRLF heading: the trailing \r must never become part of the verdict. + writeState(tmpDir, { milestone: 'v2.0' }); + writeRoadmap(tmpDir, '## v2.0 — CRLF Name\r\n\r\n### Phase 1: Setup\r\n'); + info = getMilestoneInfo(tmpDir); + assert.strictEqual(info.scope, SCOPE.COMPLETE); + assert.strictEqual(info.value.name, 'CRLF Name'); + }); + + test('#4134 — property: a word-char remainder is always a name, a punctuation-only remainder never is', () => { + // Document-shaped generator (#2371): fixed literal token alphabets, NOT + // derived from the parser's own regexes. Fragments are token lists joined + // with single spaces, so no token can glue onto the version token and + // trigger the sub-milestone continuation grammar (`v1.3-B`). + const WORD = fc.constantFrom('Alpha', 'Beta', 'R2D2', '42', '名称', 'küche'); + const PUNCT = fc.constantFrom(')', '(', '—', ':', '.', '**', ']'); + const minor = fc.integer({ min: 0, max: 9 }); + const tokens = fc.array(fc.oneof(WORD, PUNCT), { minLength: 1, maxLength: 6 }); + + const prop = fc.property(minor, tokens, (m, toks) => { + const version = `v1.${m}`; + const fragment = toks.join(' ').trim(); + const hasWordChar = /[\p{L}\p{N}]/u.test(fragment); + const out = roadmapParser.listMilestoneHeadings(`## ${version} ${fragment}\n`); + assert.strictEqual(out.length, 1, `heading not enumerated: ${version} ${fragment}`); + assert.strictEqual(out[0].version, version, `continuation grammar leaked into the version: ${JSON.stringify(out[0])}`); + // The biconditional IS the #4134 contract: a remainder with at least one + // letter/digit is a curated name; one with none is heading structure. + assert.strictEqual( + out[0].name !== null, + hasWordChar, + `fragment ${JSON.stringify(fragment)} (hasWordChar=${hasWordChar}) yielded name ${JSON.stringify(out[0].name)}`, + ); + }); + + const result = fc.check(prop, { seed: 20260905, numRuns: 300 }); + if (result.failed) { + assert.fail(`#4134 property violated (replay seed=20260905): ${JSON.stringify(result.counterexample)}`); + } + }); +}); + // ─── isMilestoneShippedInRoadmap ────────────────────────────────────────────── // #2562: this module owns milestone-heading classification, so its own shipped From d5a85da8ab3470986dfd0088d785905ce495919b Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 00:03:05 -0400 Subject: [PATCH 005/166] fix(#4363): bump download-artifact and setup-node off node20 runtimes (#4365) --- .github/workflows/mutation.yml | 2 +- .github/workflows/release.yml | 4 ++-- .github/workflows/test.yml | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/.github/workflows/mutation.yml b/.github/workflows/mutation.yml index 7d1460132..ddebbca4d 100644 --- a/.github/workflows/mutation.yml +++ b/.github/workflows/mutation.yml @@ -127,7 +127,7 @@ jobs: run: echo "CI_JOB_START_EPOCH_MS=$(( $(date +%s) * 1000 ))" >> "$GITHUB_ENV" - name: Set up Node - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0 + uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0 with: node-version-file: .nvmrc cache: npm diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 9e0d442ac..4875d30a3 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -452,7 +452,7 @@ jobs: # exactly the layout c8 expects from a single run (test.yml's # coverage-gate job does the identical merge for the same reason). - name: Download every shard's raw coverage - uses: actions/download-artifact@018cc2cf5baa6db3ef3c5f8a56943fffe632ef53 # v6.0.0 + uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1 with: pattern: coverage-tmp-rc-shard-* path: coverage/tmp @@ -727,7 +727,7 @@ jobs: run: npm ci - name: Download every shard's raw coverage - uses: actions/download-artifact@018cc2cf5baa6db3ef3c5f8a56943fffe632ef53 # v6.0.0 + uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1 with: pattern: coverage-tmp-finalize-shard-* path: coverage/tmp diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 92731bc1f..3ea4c174d 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -747,7 +747,7 @@ jobs: # is exactly the layout c8 expects from a single run. The dumps are # per-process files with distinct names, so there is nothing to collide. - name: Download every shard's raw coverage - uses: actions/download-artifact@018cc2cf5baa6db3ef3c5f8a56943fffe632ef53 # v6.0.0 + uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1 with: pattern: coverage-tmp-shard-* path: coverage/tmp From 6adf3098acd8db28e3f6f13ef4d48b49bb25c721 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 02:04:18 -0400 Subject: [PATCH 006/166] fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4145): regression rows for hash-matching prefix-less pristine baselines RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a null-returning stub so the new rows fail behaviorally, not at require time. Failing-first rows: verifier resolution (no_baseline must drop to 0 when an exact-hash orphan exists), findPristineByHash unit row, and the two saveLocalPatches relocation rows. Negative-space rows pin today's behavior: missing baselines still report ok_no_baseline, mismatching orphans are never adopted or deleted, canonical precedence and the #3657 drift posture are untouched. * fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans Both pristine readers joined the manifest-keyed path strictly, so a snapshot stored without the gsd-core/ prefix (an earlier release's writer) was reported as ok_no_baseline by the verifier and pushed into regeneration by saveLocalPatches — where incoming-release candidates can never satisfy the recorded outgoing hash, leaving the correct baseline permanently unconsumed. - src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash — deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the recorded pristine_hashes entry (the same authority the #3657 drift guard trusts), symlink-skipping, canonical path excluded via skipRel. - verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded hash, adopt byte-identical content found anywhere under gsd-pristine/ before reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and the frozen REASON/report shapes are untouched; the verifier stays read-only. - install.js saveLocalPatches(): preserve-check rescue — relocate a hash-matching orphan to the canonical path (copy, hash-verify, then remove the orphan) so the state self-heals on the next update instead of repeating forever. Honest accounting: new non-overlapping rescued counter. - Workflow doc: one-sentence note on hash-based snapshot resolution. - Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore entries for the compiled artifact, seedFixture mkdir fix in the new rows. Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145) * fix(#4145): review follow-up — orphan scan never consumes a canonical path Adversarial review finding: with two modified files sharing byte-identical outgoing content, recoverOrphanedPristine could adopt the OTHER file's canonical pristine as its rescue source — relocating it (copy + delete at its home path) and ping-ponging the single baseline between the two files across updates. findPristineByHash's skip parameter now accepts a Set, and saveLocalPatches passes the normalized manifest keys so every canonical path is excluded; only genuine non-canonical orphans are eligible for removal (no strict-join reader ever consults those). Adds the canonical-theft regression row, a Set-skip unit assertion, and tightens the workflow doc sentence the same pass flagged as overstated. * fix(#4145): INVENTORY roster row + symlink-fixture correction Two leftovers from the ab17b7a1e5 bench run, both root-caused: - docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs (#3762 gate: every manifest entry carries a row). - The findPristineByHash symlink unit fixture placed its symlink target INSIDE the scanned root, so the walk legitimately matched the real target file. The implementation skips the symlink itself; the fixture now keeps the target outside the scanned tree so the assertion tests what it claims. * changeset(#4145): fixed fragment for pristine baseline hash resolution --------- Co-authored-by: gsd-agent --- .changeset/noble-geese-climb.md | 5 + .gitignore | 2 + bin/install.js | 96 +++++++- docs/INVENTORY-MANIFEST.json | 1 + docs/INVENTORY.md | 1 + eslint.config.mjs | 2 + gsd-core/bin/verify-reapply-patches.cjs | 35 ++- gsd-core/workflows/reapply-patches.md | 2 +- src/pristine-baseline.cts | 94 ++++++++ tests/install-write-confinement.test.cjs | 224 +++++++++++++++++++ tests/reapply-verify-hunks.test.cjs | 269 +++++++++++++++++++++++ 11 files changed, 724 insertions(+), 7 deletions(-) create mode 100644 .changeset/noble-geese-climb.md create mode 100644 src/pristine-baseline.cts diff --git a/.changeset/noble-geese-climb.md b/.changeset/noble-geese-climb.md new file mode 100644 index 000000000..fab016957 --- /dev/null +++ b/.changeset/noble-geese-climb.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4364 +--- +**`/gsd-update --reapply` no longer reports no_baseline when a hash-matching gsd-pristine/ snapshot is stored without the gsd-core/ prefix** — the verifier and the installer now resolve the baseline by the recorded SHA-256 and relocate the orphaned snapshot to its canonical path on the next update, so the correct baseline is finally consumed instead of sitting unusable forever. (#4145) diff --git a/.gitignore b/.gitignore index ea044ca1c..62eb22987 100644 --- a/.gitignore +++ b/.gitignore @@ -138,6 +138,8 @@ build/ /gsd-core/bin/lib/ui-consideration-probe.cjs /gsd-core/bin/lib/config-types.cjs /gsd-core/bin/lib/cli-exit.cjs +# #4145: emitted artifact of src/pristine-baseline.cts — never edited. +/gsd-core/bin/lib/pristine-baseline.cjs /gsd-core/bin/lib/code-review-flags.cjs /gsd-core/bin/lib/code-review-depth.cjs /gsd-core/bin/lib/context-utilization.cjs diff --git a/bin/install.js b/bin/install.js index 7d8233317..4951e8221 100755 --- a/bin/install.js +++ b/bin/install.js @@ -544,6 +544,13 @@ const { RUNTIME_PROFILE_MAP: GSD_RUNTIME_PROFILE_MAP, isAnthropicFlavoredModel: gsdIsAnthropicFlavoredModel, } = require(path.join(_gsdLibDir, 'model-catalog.cjs')); +// #4145: shared hash-first recovery for gsd-pristine/ baselines stored at an +// unexpected path (e.g. without the gsd-core/ prefix an earlier release's +// writer dropped). Same module the reapply verifier uses, so the two readers +// cannot drift apart again. +const { + findPristineByHash: gsdFindPristineByHash, +} = require(path.join(_gsdLibDir, 'pristine-baseline.cjs')); // #2875 Part 2: MODEL_PROFILES + resolveTierEntry are now consumed only by // install-model-override-resolver.cjs's readGsdRuntimeProfileResolver // (required below) — this installer no longer needs its own bindings. @@ -10153,6 +10160,63 @@ function populatePristineDir({ packageSrc, pristineDir, modified, runtime, pathP return written; } +/** + * #4145: recover a pristine baseline from a hash-matching orphan stored at an + * unexpected path under gsd-pristine/ (e.g. without the gsd-core/ prefix an + * earlier release's writer dropped). + * + * The preserve-check's strict join (pristineDir + manifest-keyed relPath) + * misses such snapshots, so they were pushed into regeneration from the + * incoming release — and when the file changed upstream, the candidate's hash + * could never satisfy the recorded outgoing hash, leaving the correct + * baseline permanently unconsumed and unpruned (the self-perpetuating state + * #4145 reports). Hash equality with pristine_hashes is the same authority + * the #3657 drift guard trusts, so an exact match cannot be the wrong + * baseline no matter where under gsd-pristine/ it lives. + * + * Recovery = relocation: copy the orphan to the canonical manifest-keyed path + * (hash-verified after the copy) and remove the orphan only once the + * canonical copy is verified in place. Returns true when the canonical path + * ended up holding recorded-hash bytes. Never deletes anything it cannot + * vouch for by hash, and never consumes a path that is the canonical path of + * ANY manifest file (see canonicalSkip below) — only genuine orphans, which + * no strict-join reader ever consults, are eligible for removal. + */ +function recoverOrphanedPristine(pristineDir, relPath, recordedHash, canonicalSkip) { + if (!recordedHash) return false; + let orphanRel; + try { + // canonicalSkip = the normalized manifest keys: a file already sitting at + // any canonical path can never be (re-)adopted through the scan. Without + // this, two modified files sharing byte-identical outgoing content would + // repeatedly "rescue" each other's canonical away (relocate + delete at + // its home path) in alternating updates — bytes identical, state unstable. + // It also keeps drift (#3657) / stale (#3407) territory with the caller. + orphanRel = gsdFindPristineByHash(pristineDir, recordedHash, canonicalSkip); + } catch { + return false; + } + if (!orphanRel) return false; + const outRef = resolveInstallRelativePath(pristineDir, relPath); + if (!outRef) return false; + try { + fs.mkdirSync(path.dirname(outRef.fullPath), { recursive: true }); + fs.copyFileSync(path.join(pristineDir, orphanRel), outRef.fullPath); + // Verify the relocated copy before removing the orphan — only a + // hash-matching canonical counts as recovered. + if (fileHash(outRef.fullPath) !== recordedHash) { + try { fs.rmSync(outRef.fullPath, { force: true }); } catch { /* best-effort */ } + return false; + } + // Orphan removal is best-effort: the canonical copy is already verified, + // so a failed unlink leaves a harmless duplicate, never data loss. + try { fs.rmSync(path.join(pristineDir, orphanRel), { force: true }); } catch { /* best-effort */ } + return true; + } catch { + return false; + } +} + /** * Detect user-modified GSD files by comparing against install manifest. * Backs up modified files to gsd-local-patches/ for reapply after update. @@ -10299,6 +10363,15 @@ function saveLocalPatches(configDir, pristineCtx) { const stalePaths = new Set(); // Track which relPaths were successfully regenerated (from either missing or stale). const regeneratedPaths = new Set(); + // #4145: track which relPaths were recovered by relocating a hash-matching + // orphan (stored at an unexpected path, e.g. without the gsd-core/ prefix). + const rescuedPaths = new Set(); + // #4145: the set of paths that are SOME file's canonical pristine path + // (every normalized manifest key). The orphan scan must never consume + // these — see recoverOrphanedPristine. + const canonicalSkip = new Set( + Object.keys(manifest.files || {}).map((k) => normalizeInstallRelativePath(k)).filter(Boolean), + ); const missingPaths = []; for (const relPath of modified) { const outRef = resolveInstallRelativePath(pristineDir, relPath); @@ -10320,6 +10393,17 @@ function saveLocalPatches(configDir, pristineCtx) { stalePaths.add(relPath); } } + // #4145: canonical absent (or just removed as stale) — before falling + // into regeneration, try to recover the baseline from a hash-matching + // orphan elsewhere under gsd-pristine/ and relocate it to the canonical + // path. This is the self-heal for snapshots an earlier release stored + // without the gsd-core/ prefix: without it the state repeats forever + // (regeneration candidates from the incoming release can never satisfy + // the recorded outgoing hash when upstream changed the file). + if (recoverOrphanedPristine(pristineDir, relPath, pristineHashes[relPath], canonicalSkip)) { + rescuedPaths.add(relPath); + continue; + } // File absent from gsd-pristine/ (or just removed above as stale): // attempt hash-validated regeneration from new-release source. missingPaths.push(relPath); @@ -10362,13 +10446,21 @@ function saveLocalPatches(configDir, pristineCtx) { } // `regenerated` = total files successfully regenerated (from missing OR stale). const regenerated = regeneratedPaths.size; + // `rescued` = files recovered by relocating a hash-matching orphan to its + // canonical path (#4145) — distinct from preservation (canonical already + // correct) and regeneration (bytes re-derived from new-release source). + const rescued = rescuedPaths.size; // `removed` = stale entries that were deleted and NOT subsequently regenerated. // Entries that were stale-deleted but then successfully regenerated are counted - // only in `regenerated` — the counts are non-overlapping. - const removed = [...stalePaths].filter(p => !regeneratedPaths.has(p)).length; + // only in `regenerated`; stale-deleted-then-orphan-rescued entries are counted + // only in `rescued` — the counts are non-overlapping. + const removed = [...stalePaths].filter(p => !regeneratedPaths.has(p) && !rescuedPaths.has(p)).length; if (preserved > 0) { console.log(' ' + green + '✓' + reset + ' Preserved ' + cyan + 'gsd-pristine/' + reset + ' (' + preserved + ' file(s)) for three-way merge'); } + if (rescued > 0) { + console.log(' ' + green + '✓' + reset + ' Recovered ' + cyan + 'gsd-pristine/' + reset + ' (' + rescued + ' file(s)) by recorded hash from a legacy-path snapshot and relocated them (#4145)'); + } if (regenerated > 0) { console.log(' ' + green + '✓' + reset + ' Regenerated ' + cyan + 'gsd-pristine/' + reset + ' (' + regenerated + ' file(s)) via hash-validated new-release source'); } diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 755c28de7..8d46675fd 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -475,6 +475,7 @@ "planning-scope.cjs", "planning-snapshot.cjs", "planning-workspace.cjs", + "pristine-baseline.cjs", "probe-core.cjs", "profile-output.cjs", "profile-pipeline-command-router.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index 09baf64ef..f1e418d65 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -603,6 +603,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `planning-scope.cjs` | Frozen `SCOPE` discriminator (`COMPLETE`/`TRUNCATED`/`UNSCOPED`/`UNREADABLE`) distinguishing a genuinely-empty derivation from one computed over a truncated or unscoped input, so callers can branch on the difference instead of reading a plausible zero (ADR-3180) | | `planning-snapshot.cjs` | Parsed projection of `.planning/` composed exclusively from the ADR-3180 §7 owners (milestone identity, phase enumeration, phase completion, plan/summary counting, STATE.md current-phase) — exposes only scope-carrying parsed values, never raw document text, so a diagnostic rule cannot re-derive a field's location (ADR-3180 §8.1) | | `planning-workspace.cjs` | Planning path/workstream seam (`planningDir`, `planningPaths`, active-workstream routing, `.planning/.lock` orchestration) | +| `pristine-baseline.cjs` | Hash-first recovery for `gsd-pristine/` baselines stored at an unexpected path (compiled from `src/pristine-baseline.cts`, gitignored; #4145) — `findPristineByHash(pristineDir, recordedHash, skip?)` walks `gsd-pristine/` in deterministic sorted order, skips symlinks, and returns the first file whose SHA-256 equals the recorded `backup-meta.json.pristine_hashes` entry (the same authority the #3657 drift guard trusts); the `skip` set excludes canonical manifest-keyed paths so a relocation never consumes another file's canonical baseline. Shared by `verify-reapply-patches.cjs`'s `verifyFile` (read-only adoption when the strict join misses) and `install.js`'s `saveLocalPatches` (orphan relocation self-heal) so the two readers cannot drift apart again | | `project-root.cjs` | Resolves a project root from a starting directory using four heuristics (own `.planning/` guard, `sub_repos` config, `multiRepo` flag, `.git` heuristic) | | `profile-output.cjs` | Profile rendering, USER-PROFILE.md and dev-preferences.md generation | | `profile-pipeline-command-router.cjs` | ADR-959 capability command router for the profile-pipeline command family — dispatches scan-sessions, extract-messages, profile-sample (pipeline phase) and write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md (output phase); phase 6 cutover | diff --git a/eslint.config.mjs b/eslint.config.mjs index 0b6bb13f1..a017ad849 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -126,6 +126,8 @@ export default tseslint.config( 'gsd-core/bin/lib/prohibition-enforcement.cjs', // #3770: tsc-generated runtime artifact — lint the src/tdd-red-evidence.cts source. 'gsd-core/bin/lib/tdd-red-evidence.cjs', + // #4145: tsc-generated runtime artifact — lint the src/pristine-baseline.cts source. + 'gsd-core/bin/lib/pristine-baseline.cjs', 'gsd-core/bin/lib/ui-consideration-probe.cjs', 'gsd-core/bin/lib/code-review-flags.cjs', 'gsd-core/bin/lib/code-review-depth.cjs', diff --git a/gsd-core/bin/verify-reapply-patches.cjs b/gsd-core/bin/verify-reapply-patches.cjs index 08170a0c1..c052fc3ec 100755 --- a/gsd-core/bin/verify-reapply-patches.cjs +++ b/gsd-core/bin/verify-reapply-patches.cjs @@ -33,6 +33,11 @@ const fs = require('node:fs'); const path = require('node:path'); const crypto = require('node:crypto'); const { ExitError, runMain } = require('./lib/cli-exit.cjs'); +// #4145: shared hash-first recovery for baselines stored at an unexpected +// path under gsd-pristine/ (e.g. without the gsd-core/ prefix an earlier +// release's writer dropped). Same module the installer's preserve-check uses, +// so the two readers cannot drift apart again. +const { findPristineByHash } = require('./lib/pristine-baseline.cjs'); const SIGNIFICANT_MIN_CHARS = 12; const GSD_HOOK_VERSION_LINE_RE = /^(?:\/\/|#)\s*gsd-hook-version:\s*\S+\s*$/i; @@ -344,8 +349,29 @@ function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashe // pristinePathExists stays false. } - // Bug #934: recordedHash is present (modern installer) but the pristine - // path does not exist on disk at all (stat threw above). This means + // Bug #4145: the canonical join missed, but the recorded hash is the + // baseline authority the #3657 drift guard already trusts. Before + // reporting OK_NO_BASELINE, scan gsd-pristine/ for byte-identical content + // (an earlier release may have stored the snapshot without the gsd-core/ + // prefix). An exact sha-256 match cannot be the wrong baseline, and + // gsd-pristine/ holds only backed-up files, so the scan is small. The + // canonical path itself is excluded — a mismatching file at the joined + // path is drift (#3657), never re-adopted through the scan. + if (!pristinePathExists && recordedHash) { + try { + const recoveredRel = findPristineByHash(pristineDir, recordedHash, hashKey); + if (recoveredRel) { + pristineContent = fs.readFileSync(path.join(pristineDir, recoveredRel), 'utf8'); + pristinePathExists = true; + } + } catch { + // scan or read failure — fall through to the OK_NO_BASELINE posture + } + } + + // Bug #934: recordedHash is present (modern installer) but no hash-matching + // pristine exists anywhere under gsd-pristine/ (stat threw above AND the + // #4145 scan found nothing). This means // saveLocalPatches recorded a hash but could not write the corresponding // gsd-pristine/ file (the only candidate was discarded because it was from // a newer release). Falling to over-broad mode here would treat every @@ -353,8 +379,9 @@ function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashe // false FAIL_USER_LINES_MISSING for each upstream removal. Since we // cannot reason correctly without a baseline, the safe answer is advisory/ // non-blocking: return OK_NO_BASELINE and let the caller decide. - // NOTE: this guard fires ONLY when stat threw (path absent), not when the - // path is present but non-file — in that case over-broad mode is safer. + // NOTE: this guard fires ONLY when the baseline path is absent (stat threw + // and nothing matched by hash), not when the path is present but non-file — + // in that case over-broad mode is safer. if (!pristinePathExists && recordedHash) { result.reason = REASON.OK_NO_BASELINE; return result; diff --git a/gsd-core/workflows/reapply-patches.md b/gsd-core/workflows/reapply-patches.md index b4e6a493a..0b95f333f 100644 --- a/gsd-core/workflows/reapply-patches.md +++ b/gsd-core/workflows/reapply-patches.md @@ -169,7 +169,7 @@ Check if a `gsd-pristine/` directory exists alongside `gsd-local-patches/`: ```bash PRISTINE_DIR="$CONFIG_DIR/gsd-pristine" ``` -If it exists, the installer saved pristine copies at install time. Use these as the baseline. +If it exists, the installer saved pristine copies at install time. Use these as the baseline. Both the deterministic verifier and the installer's preserve-check resolve each file's snapshot at its canonical path first and, when that misses, by the SHA-256 recorded in `pristine_hashes` — so a snapshot stored at a legacy path (for example, without the `gsd-core/` prefix an earlier release dropped) is still found and, on the next update, relocated to its canonical path (#4145). ### Option C: No baseline available (two-way fallback) If neither git history nor pristine snapshots are available, fall back to two-way comparison — but with **strengthened heuristics** (see Step 3). diff --git a/src/pristine-baseline.cts b/src/pristine-baseline.cts new file mode 100644 index 000000000..1bc5c40e6 --- /dev/null +++ b/src/pristine-baseline.cts @@ -0,0 +1,94 @@ +/** + * #4145: hash-first recovery for gsd-pristine/ baselines stored at an + * unexpected path. + * + * Some installs hold a pristine snapshot whose SHA-256 equals the hash recorded + * in backup-meta.json.pristine_hashes for a manifest-keyed file, but at a path + * that is not `path.join(pristineDir, relPath)` — e.g. stored without the + * `gsd-core/` top-level segment by an earlier release's writer. Both readers + * (verify-reapply-patches.cjs verifyFile and install.js saveLocalPatches) + * resolved strictly by that join, missed the snapshot, and reported + * ok_no_baseline / fell into regeneration that can never satisfy the recorded + * outgoing hash — a self-perpetuating gap. + * + * Hash equality with the recorded pristine_hashes entry is the same authority + * the #3657 drift guard already trusts, so a match cannot be the wrong + * baseline regardless of which release wrote it or where under gsd-pristine/ + * it lives. This module owns the shared scan so the two readers cannot drift + * apart again (two private strict joins drifting is exactly the bug class). + * + * ADR-457: runtime module in src/*.cts, compiled to + * gsd-core/bin/lib/pristine-baseline.cjs. + */ + +import fs from 'node:fs'; +import path from 'node:path'; +import crypto from 'node:crypto'; + +/** + * SHA-256 hex digest of a file's raw bytes. Byte-for-byte the same digest + * install.js fileHash() records into manifests and backup-meta.json. + */ +export function sha256File(absPath: string): string { + return crypto.createHash('sha256').update(fs.readFileSync(absPath)).digest('hex'); +} + +function walkSorted(dir: string, relPrefix: string, results: string[]): void { + let entries: fs.Dirent[]; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } catch { + return; // absent or unreadable — nothing to scan here + } + entries.sort((a, b) => (a.name < b.name ? -1 : a.name > b.name ? 1 : 0)); + for (const entry of entries) { + // Never follow symlinks: gsd-pristine/ is installer-authored plain files; + // a link here is not a baseline and must not redirect the walk out of the + // tree (same posture as migration 004's walker). + if (entry.isSymbolicLink()) continue; + const rel = relPrefix ? `${relPrefix}/${entry.name}` : entry.name; + if (entry.isDirectory()) { + walkSorted(path.join(dir, entry.name), rel, results); + } else if (entry.isFile()) { + results.push(rel); + } + } +} + +/** + * Find the first file under `pristineDir` (deterministic sorted walk) whose + * SHA-256 equals `recordedHash`, as a pristineDir-relative POSIX path. + * + * - `skip` is never returned — a single POSIX relPath string or a Set of them. + * Callers pass the canonical path(s) they (or other files in the same run) + * already own, so a file sitting at a canonical path is never adopted + * through the scan. For the installer's relocation this is what prevents a + * byte-identical canonical belonging to ANOTHER modified file from being + * "rescued" away (relocated and deleted at its home path). + * - Multiple matches are byte-identical by sha-256 authority; sorted order + * makes the choice deterministic. + * - Returns null when pristineDir is absent/unreadable or nothing matches. + */ +export function findPristineByHash( + pristineDir: string, + recordedHash: string, + skip?: string | ReadonlySet, +): string | null { + if (!pristineDir || typeof recordedHash !== 'string' || recordedHash.length === 0) { + return null; + } + const skipSet = skip instanceof Set ? skip : new Set(skip !== undefined ? [skip] : []); + const rels: string[] = []; + walkSorted(pristineDir, '', rels); + for (const rel of rels) { + if (skipSet.has(rel)) continue; + try { + if (sha256File(path.join(pristineDir, rel)) === recordedHash) { + return rel; + } + } catch { + // unreadable candidate — keep scanning + } + } + return null; +} diff --git a/tests/install-write-confinement.test.cjs b/tests/install-write-confinement.test.cjs index 4d16689e6..b4ceaece6 100644 --- a/tests/install-write-confinement.test.cjs +++ b/tests/install-write-confinement.test.cjs @@ -4013,3 +4013,227 @@ describe('Bug #4086: saveLocalPatches resolves skills/ manifest keys at the runt }); }); } + + +// ──────────────────────────────────────────────────────────────────────── +// Folded regression block — #4145 (self-heal side). saveLocalPatches' +// preserve-check resolved gsd-pristine/ entries strictly by the manifest-keyed +// path, so a hash-matching snapshot stored without the gsd-core/ prefix was +// pushed into regeneration from the incoming release; when upstream changed +// the file, the candidate hash-mismatched and was discarded. The correct +// baseline was never consumed and never pruned — the state repeated on every +// future update. The fix rescues exact-recorded-hash orphans by relocating +// them to the canonical path. +// ──────────────────────────────────────────────────────────────────────── +{ + const { describe: __foldDescribe } = require('node:test'); + __foldDescribe('folded:bug-4145-saveLocalPatches-orphan-rescue', () => { +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test, describe, beforeEach } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const os = require('node:os'); +const crypto = require('node:crypto'); + +const ROOT = path.join(__dirname, '..'); +const INSTALL = require(path.join(ROOT, 'bin', 'install.js')); +const { cleanup } = require('./helpers.cjs'); + +const MANIFEST_NAME = 'gsd-file-manifest.json'; + +function sha256(content) { + return crypto.createHash('sha256').update(content).digest('hex'); +} + +describe('Bug #4145: saveLocalPatches rescues hash-matching orphaned pristine snapshots', () => { + let tmpDir; + let configDir; + let newSrcDir; + let pristineDir; + + const FILE = 'gsd-core/bin/lib/frontmatter.cjs'; + const OLD_PRISTINE = '# Old Release Content\nThis is the outgoing pristine.\n'; + const NEW_RELEASE = '# New Release Content\nUpstream rewrote this file wholesale in v2.\n'; + const USER_MODIFIED = OLD_PRISTINE + '## User addition\nUser customization here.\n'; + + beforeEach((t) => { + tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4145-slp-')); + configDir = path.join(tmpDir, 'config'); + newSrcDir = path.join(tmpDir, 'new-release-src'); + pristineDir = path.join(configDir, 'gsd-pristine'); + fs.mkdirSync(configDir, { recursive: true }); + fs.mkdirSync(newSrcDir, { recursive: true }); + t.after(() => { + cleanup(tmpDir); + }); + }); + + function seedFixture({ orphanRel, orphanContent, canonicalContent, newReleaseContent }) { + fs.mkdirSync(path.dirname(path.join(configDir, FILE)), { recursive: true }); + fs.writeFileSync(path.join(configDir, FILE), USER_MODIFIED); + fs.writeFileSync( + path.join(configDir, MANIFEST_NAME), + JSON.stringify({ version: '1.0.0', files: { [FILE]: sha256(OLD_PRISTINE) } }, null, 2), + ); + if (canonicalContent !== undefined) { + fs.mkdirSync(path.dirname(path.join(pristineDir, FILE)), { recursive: true }); + fs.writeFileSync(path.join(pristineDir, FILE), canonicalContent); + } + if (orphanRel !== undefined) { + fs.mkdirSync(path.dirname(path.join(pristineDir, orphanRel)), { recursive: true }); + fs.writeFileSync(path.join(pristineDir, orphanRel), orphanContent); + } + fs.mkdirSync(path.dirname(path.join(newSrcDir, FILE)), { recursive: true }); + fs.writeFileSync(path.join(newSrcDir, FILE), newReleaseContent); + } + + /** + * Core regression (self-heal): the hash-matching snapshot sits at + * bin/lib/frontmatter.cjs — without the gsd-core/ segment. The new release + * changed the file upstream, so regeneration candidates are discarded. + * After the fix the orphan is relocated to the canonical manifest-keyed + * path and the unprefixed copy no longer lingers. + */ + test('#4145: saveLocalPatches relocates a hash-matching unprefixed orphan to the canonical pristine path', () => { + const orphanRel = 'bin/lib/frontmatter.cjs'; + seedFixture({ orphanRel, orphanContent: OLD_PRISTINE, newReleaseContent: NEW_RELEASE }); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + const canonical = path.join(pristineDir, FILE); + assert.ok(fs.existsSync(canonical), 'canonical prefixed pristine must exist after the update'); + assert.equal(sha256(fs.readFileSync(canonical, 'utf8')), sha256(OLD_PRISTINE), + 'relocated baseline must carry the outgoing (recorded-hash) bytes, not new-release bytes'); + assert.equal(fs.existsSync(path.join(pristineDir, orphanRel)), false, + 'the unprefixed orphan must not linger once relocated'); + }); + + /** + * Stale canonical (new-release bytes) + hash-matching orphan elsewhere: + * the stale entry is removed by the #3407 path and then rescued from the + * orphan — the file must not end in no-baseline limbo. + */ + test('#4145: rescues after stale-canonical removal when a hash-matching orphan exists', () => { + seedFixture({ + orphanRel: 'legacy/frontmatter.cjs', + orphanContent: OLD_PRISTINE, + canonicalContent: NEW_RELEASE, // stale — hash mismatch + newReleaseContent: NEW_RELEASE, + }); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + const canonical = path.join(pristineDir, FILE); + assert.ok(fs.existsSync(canonical), 'canonical pristine must exist after stale removal + rescue'); + assert.equal(sha256(fs.readFileSync(canonical, 'utf8')), sha256(OLD_PRISTINE), + 'rescued baseline must carry the recorded-hash bytes'); + assert.equal(fs.existsSync(path.join(pristineDir, 'legacy', 'frontmatter.cjs')), false, + 'the orphan must be consumed by the relocation'); + }); + + /** Negative space: no orphan, upstream changed — regeneration discard (#3407) is unchanged. */ + test('#4145: leaves the baseline absent when no orphan exists and upstream changed', () => { + seedFixture({ newReleaseContent: NEW_RELEASE }); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + assert.equal(fs.existsSync(path.join(pristineDir, FILE)), false, + 'no hash-matching source exists — the baseline must stay absent (over-broad/no-baseline fallback)'); + }); + + /** Negative space: a mismatching orphan is neither adopted nor deleted. */ + test('#4145: never adopts nor deletes a hash-mismatching orphan', () => { + const orphanRel = 'bin/lib/frontmatter.cjs'; + seedFixture({ orphanRel, orphanContent: NEW_RELEASE, newReleaseContent: NEW_RELEASE }); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + assert.equal(fs.existsSync(path.join(pristineDir, FILE)), false, + 'mismatching bytes must not be written to the canonical pristine path'); + assert.equal(fs.existsSync(path.join(pristineDir, orphanRel)), true, + 'pruning files the recorded hashes do not vouch for is not this fix\'s job'); + }); + + /** + * Review finding (fix follow-up): two modified files sharing byte-identical + * outgoing content. The orphan scan must never consume a path that is + * another manifest file's canonical pristine path — otherwise the rescue + * would relocate a correct canonical away from its owner and the two files + * would ping-pong it between updates. Only genuine non-canonical orphans + * are eligible. + */ + test('#4145: does not steal a byte-identical canonical belonging to another modified file', () => { + const FILE_B = 'gsd-core/bin/lib/other-file.cjs'; + const SHARED_OLD = '# Shared Old Content\nByte-identical across two manifest files.\n'; + // A and B are both user-modified on top of byte-identical outgoing stock, + // so both manifest records carry the SAME pristine hash. + fs.mkdirSync(path.dirname(path.join(configDir, FILE_B)), { recursive: true }); + fs.writeFileSync(path.join(configDir, FILE), SHARED_OLD + '## User addition A\nCustom A.\n'); + fs.writeFileSync(path.join(configDir, FILE_B), SHARED_OLD + '## User addition B\nCustom B.\n'); + fs.writeFileSync( + path.join(configDir, MANIFEST_NAME), + JSON.stringify({ + version: '1.0.0', + files: { [FILE]: sha256(SHARED_OLD), [FILE_B]: sha256(SHARED_OLD) }, + }, null, 2), + ); + // A (processed first) has the ALREADY-correct canonical holding the shared + // old bytes. B has no canonical and no orphan — B's only possible hash + // match is A's canonical. Without the canonical skip set, B's rescue would + // copy A's canonical to B's path and then DELETE A's canonical. + fs.mkdirSync(path.dirname(path.join(pristineDir, FILE)), { recursive: true }); + fs.writeFileSync(path.join(pristineDir, FILE), SHARED_OLD); + fs.mkdirSync(path.dirname(path.join(newSrcDir, FILE)), { recursive: true }); + fs.writeFileSync(path.join(newSrcDir, FILE), NEW_RELEASE); + fs.mkdirSync(path.dirname(path.join(newSrcDir, FILE_B)), { recursive: true }); + fs.writeFileSync(path.join(newSrcDir, FILE_B), NEW_RELEASE); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + // A's canonical must survive untouched — never stolen to become B's. + assert.ok(fs.existsSync(path.join(pristineDir, FILE)), + 'the byte-identical canonical of the earlier-processed file must survive'); + assert.equal(sha256(fs.readFileSync(path.join(pristineDir, FILE), 'utf8')), sha256(SHARED_OLD)); + // B gains no baseline from A's canonical (falls to regeneration instead). + assert.equal(fs.existsSync(path.join(pristineDir, FILE_B)), false, + 'a canonical path of another file must never be relocated as the rescue source'); + }); + + /** Preserve-path lock: an already-correct canonical stays put; the preserve loop ignores the orphan. */ + test('#4145: preserves an already-correct canonical and leaves a coexisting identical orphan in place', () => { + const orphanRel = 'bin/lib/frontmatter.cjs'; + seedFixture({ + orphanRel, + orphanContent: OLD_PRISTINE, + canonicalContent: OLD_PRISTINE, // already correct + newReleaseContent: NEW_RELEASE, + }); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + const canonical = path.join(pristineDir, FILE); + assert.ok(fs.existsSync(canonical)); + assert.equal(sha256(fs.readFileSync(canonical, 'utf8')), sha256(OLD_PRISTINE), + 'preserved canonical must be byte-identical to before the run'); + assert.equal(fs.existsSync(path.join(pristineDir, orphanRel)), true, + 'the preserve path must not disturb unrelated files'); + }); +}); + }); +} diff --git a/tests/reapply-verify-hunks.test.cjs b/tests/reapply-verify-hunks.test.cjs index 748e56ad5..67c719ab1 100644 --- a/tests/reapply-verify-hunks.test.cjs +++ b/tests/reapply-verify-hunks.test.cjs @@ -1170,3 +1170,272 @@ describe('Bug #4086: verifyFile resolves skills entries at the runtime skills ro }); }); } + + +// ──────────────────────────────────────────────────────────────────────── +// Folded regression block — #4145 (a hash-matching gsd-pristine/ baseline +// stored without the gsd-core/ prefix is never resolved). verifyFile() joined +// the manifest-keyed path strictly; when stat missed and a hash was recorded +// it reported OK_NO_BASELINE even though byte-correct content sat elsewhere +// under gsd-pristine/. The fix consults the recorded pristine_hashes — the +// same authority the #3657 drift guard trusts — and adopts an exact-hash +// match found anywhere in the tree. +// ──────────────────────────────────────────────────────────────────────── +{ + const { describe: __foldDescribe } = require('node:test'); + __foldDescribe('folded:bug-4145-pristine-prefix-resolution', () => { +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test, describe, before, after } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const os = require('node:os'); +const path = require('node:path'); +const crypto = require('node:crypto'); +const { cleanup } = require('./helpers.cjs'); +const { runNode } = require('./helpers/process-seam.cjs'); + +const ROOT = path.join(__dirname, '..'); +const SCRIPT = path.join(ROOT, 'gsd-core', 'bin', 'verify-reapply-patches.cjs'); +const { REASON } = require(SCRIPT); +const { findPristineByHash } = require( + path.join(ROOT, 'gsd-core', 'bin', 'lib', 'pristine-baseline.cjs'), +); + +let tmpRoot; +let patchesDir; +let configDir; +let pristineDir; + +function sha256(content) { + return crypto.createHash('sha256').update(content).digest('hex'); +} + +function writeFile(absPath, content) { + fs.mkdirSync(path.dirname(absPath), { recursive: true }); + fs.writeFileSync(absPath, content); +} + +function writeBackupMeta(pristine_hashes) { + writeFile(path.join(patchesDir, 'backup-meta.json'), JSON.stringify({ pristine_hashes }, null, 2)); +} + +function resetFixture() { + for (const dir of [patchesDir, configDir, pristineDir]) { + cleanup(dir); + } + fs.mkdirSync(patchesDir); + fs.mkdirSync(configDir); + fs.mkdirSync(pristineDir); +} + +function runVerifier() { + const r = runNode([ + SCRIPT, + '--patches-dir', patchesDir, + '--config-dir', configDir, + '--pristine-dir', pristineDir, + '--json', + ], { timeoutMs: 30_000 }); + return { + status: r.exitCode, + report: r.stdout && r.stdout.length ? JSON.parse(r.stdout) : null, + }; +} + +before(() => { + tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4145-')); + patchesDir = path.join(tmpRoot, 'patches'); + configDir = path.join(tmpRoot, 'installed'); + pristineDir = path.join(tmpRoot, 'pristine'); + resetFixture(); +}); + +after(() => { + cleanup(tmpRoot); +}); + +describe('Bug #4145: hash-matching prefix-less pristine baseline is resolved', () => { + /** + * Core regression. The manifest key is `gsd-core/bin/lib/frontmatter.cjs` + * but the snapshot sits at `bin/lib/frontmatter.cjs` — one segment away + * from the joined path. Its SHA-256 equals the recorded pristine_hashes + * entry. The upstream release replaced the file wholesale and only the + * user's line survived the merge, so a verifier that recovered the + * baseline computes exactly one user-added line (present → exit 0, + * no_baseline 0), while the pre-fix run reported ok_no_baseline. + */ + test('#4145: resolves a hash-matching prefix-less pristine baseline instead of reporting ok_no_baseline', () => { + resetFixture(); + const FILE = 'gsd-core/bin/lib/frontmatter.cjs'; + const OLD_PRISTINE = + 'outgoing pristine stock line one with substantial content\n' + + 'outgoing pristine stock line two also substantial content\n'; + const USER_LINE = 'user customization line that must survive the reapply merge'; + const backupContent = OLD_PRISTINE + USER_LINE + '\n'; + const installedContent = + 'incoming upstream replacement line with substantial content\n' + USER_LINE + '\n'; + + writeBackupMeta({ [FILE]: sha256(OLD_PRISTINE) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), installedContent); + // The orphan: same bytes, stored WITHOUT the gsd-core/ prefix. + writeFile(path.join(pristineDir, 'bin', 'lib', 'frontmatter.cjs'), OLD_PRISTINE); + + const { status, report } = runVerifier(); + + assert.equal(status, 0, `expected exit 0; report=${JSON.stringify(report)}`); + assert.equal(report.no_baseline, 0, 'a hash-matching baseline was on disk — it must be resolved'); + assert.deepEqual(report.no_baseline_files, []); + assert.equal(report.failures, 0); + const r0 = report.results[0]; + assert.equal(r0.status, 'ok'); + assert.notEqual(r0.reason, REASON.OK_NO_BASELINE); + }); + + /** Negative space: nothing anywhere under gsd-pristine/ matches the record. */ + test('#4145: still reports ok_no_baseline when the recorded hash matches nothing under gsd-pristine', () => { + resetFixture(); + const FILE = 'gsd-core/bin/lib/frontmatter.cjs'; + const backupContent = + 'upstream line present in the backup of the outgoing release\n' + + 'model: sonnet — the user customisation line in the backup file\n'; + const installedContent = + 'replacement upstream line in the newer release version\n' + + 'model: sonnet — the user customisation line in the backup file\n'; + + writeBackupMeta({ [FILE]: 'deadbeef00000000000000000000000000000000000000000000000000000001' }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), installedContent); + // No pristine file anywhere. + + const { status, report } = runVerifier(); + + assert.equal(status, 0, 'no-baseline is advisory, never a failure'); + assert.equal(report.no_baseline, 1); + assert.equal(report.results[0].reason, REASON.OK_NO_BASELINE); + }); + + /** + * Negative space: an orphan whose bytes do NOT hash to the record is never + * adopted — only exact recorded-hash matches are accepted. + */ + test('#4145: never adopts a hash-mismatching orphan — only exact recorded-hash matches', () => { + resetFixture(); + const FILE = 'gsd-core/bin/lib/frontmatter.cjs'; + const OLD_PRISTINE = 'outgoing pristine bytes that the record hashes\n'; + const OTHER_CONTENT = 'some other release snapshot with different bytes\n'; + + writeBackupMeta({ [FILE]: sha256(OLD_PRISTINE) }); + writeFile(path.join(patchesDir, FILE), 'outgoing pristine bytes that the record hashes\nuser line\n'); + writeFile(path.join(configDir, FILE), 'user line\n'); + // Orphan exists but hashes to something else. + writeFile(path.join(pristineDir, 'bin', 'lib', 'frontmatter.cjs'), OTHER_CONTENT); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.no_baseline, 1, 'a mismatching orphan is not a baseline'); + assert.equal(report.results[0].reason, REASON.OK_NO_BASELINE); + }); + + /** + * Precedence lock: the canonical prefixed path still resolves exactly as + * today even when an identical-content orphan also exists — the strict join + * stays first, and the verifier (read-only) leaves the orphan untouched. + */ + test('#4145: prefixed canonical baseline resolves exactly as today when an identical orphan also exists', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/execute-phase.md'; + const pristineContent = 'stock workflow line long enough to pass the significance threshold\n'; + const droppedLine = 'user workflow customisation that was lost in the merge operation'; + const backupContent = pristineContent + droppedLine + '\n'; + const installedContent = pristineContent; // user line dropped — real failure + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), installedContent); + writeFile(path.join(pristineDir, FILE), pristineContent); + const orphanPath = path.join(pristineDir, 'workflows', 'execute-phase.md'); + writeFile(orphanPath, pristineContent); + + const { status, report } = runVerifier(); + + assert.equal(status, 1, 'the dropped user line must still be caught via the canonical baseline'); + const r0 = report.results[0]; + assert.equal(r0.status, 'fail'); + assert.equal(r0.reason, REASON.FAIL_USER_LINES_MISSING); + assert.ok(r0.missing.includes(droppedLine)); + // Read-only verifier: the orphan is never relocated or pruned by a verify run. + assert.equal(fs.existsSync(orphanPath), true, 'verifier must not mutate gsd-pristine/'); + }); + + /** + * Drift-path lock: a hash-MISMATCHING canonical snapshot still reports + * OK_PRISTINE_DRIFT_DETECTED (#3657) — the recovery scan must not reach the + * drift case. + */ + test('#4145: canonical-path drift still reports ok_pristine_drift_detected even when a hash-matching orphan exists', () => { + resetFixture(); + const FILE = 'gsd-core/agents/gsd-executor.md'; + const oldPristine = 'old pristine line that was present when backup was captured\n'; + const newPristine = 'refreshed upstream line in the newer pristine snapshot\n'; + const userLine = 'user customisation line that should be preserved across updates'; + + writeBackupMeta({ [FILE]: sha256(oldPristine) }); + writeFile(path.join(patchesDir, FILE), oldPristine + userLine + '\n'); + writeFile(path.join(configDir, FILE), newPristine + userLine + '\n'); + writeFile(path.join(pristineDir, FILE), newPristine); // canonical drifted + // A hash-matching orphan exists elsewhere — drift must still win. + writeFile(path.join(pristineDir, 'agents', 'gsd-executor.md'), oldPristine); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + const r0 = report.results[0]; + assert.equal(r0.reason, REASON.OK_PRISTINE_DRIFT_DETECTED, + `expected the untouched #3657 drift posture; got ${r0.reason}`); + assert.equal(report.drifted, 1); + }); + + /** Module unit: deterministic sorted-first match, symlink skip, absent dir. */ + test('#4145: findPristineByHash returns the sorted-first match, skips symlinks, and null on an absent dir', () => { + const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4145-unit-')); + const symRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4145-sym-')); + const outsideRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4145-out-')); + try { + const contentA = 'identical bytes that two snapshots happen to share\n'; + const hashA = sha256(contentA); + writeFile(path.join(root, 'zz-dir', 'late-match.md'), contentA); + writeFile(path.join(root, 'aa.txt'), contentA); + + assert.equal(findPristineByHash(root, hashA), 'aa.txt', + 'sorted-first match wins deterministically'); + assert.equal(findPristineByHash(root, hashA, 'aa.txt'), 'zz-dir/late-match.md', + 'skipRel is never returned'); + assert.equal(findPristineByHash(root, hashA, new Set(['aa.txt', 'zz-dir/late-match.md'])), null, + 'every member of a skip Set is excluded (canonical-path protection)'); + assert.equal(findPristineByHash(root, sha256('no such content anywhere here\n')), null, + 'no match resolves to null'); + assert.equal(findPristineByHash(path.join(root, 'absent'), hashA), null, + 'absent dir resolves to null'); + + // A symlink is never followed, even when its target would hash-match. + // The target lives OUTSIDE symRoot so the only hashable entry inside the + // scanned tree is the symlink itself. + const outsideTarget = path.join(outsideRoot, 'outside-target.md'); + fs.writeFileSync(outsideTarget, contentA); + fs.symlinkSync(outsideTarget, path.join(symRoot, 'sym.md')); + assert.equal(findPristineByHash(symRoot, hashA), null, + 'symlinked candidates are skipped, not followed'); + } finally { + cleanup(root); + cleanup(symRoot); + cleanup(outsideRoot); + } + }); +}); + }); +} From 7bb366e83627166f8c8a17a3c16598d4832018cd Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 02:55:17 -0400 Subject: [PATCH 007/166] fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4130): failing-first regressions for --context flag + parseDecisions hardening Block A (flag): check decision-coverage-plan --context must route identically to the positional form; flag wins over positional context; valueless --context falls through to the #2770 fail-closed caller error; verify keeps its positional surface (flag is plan-only). RED on base: the flag token lands in the args[2] phase slot (false uncovered) or the args[3] context slot (silent CONTEXT.md-missing skip). Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper (?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the no-adjacent-overlap property; a differential property compares the module against a frozen copy of the pre-hardening grammars (reference validated against the base build: 60k generated lines, 0 mismatches); 40k cliff shapes assert correct outcomes with no wall-time asserts (repo rule). A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags. * fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions (A) check decision-coverage-plan --context — sibling convention (check predicate, #2008): --flag value pairs parsed by the new shared partitionPredicateArgs (parsePredicateFlags reimplemented as its flags half — one parser, cannot diverge), the flag winning over a same-purpose positional, positionals kept (no sibling deprecates them; the plan-phase workflow caller passes positionals), valueless --context falls through to the #2770 fail-closed caller error. Repair of the routing accident where --context landed in the args[2] phase slot (false uncovered) or the literal token in the args[3] context slot (silent green skip). (B) parseDecisions regex seam hardened, byte-identical on all legal inputs: the three bullet grammars consume the ID atomically via the (?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split, ~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to [^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group indices unchanged (handlers untouched). Pinned by regex-lattice tests, a differential fast-check property vs the frozen pre-hardening grammars, and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo rule — no deterministic engine step counter exists in Node). * docs+test(#4130): document --context invocation; harden lattice test tooling - docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan gate directly' block documenting both the positional and --context forms, flag precedence, and the valueless-flag fail-closed semantics (same place the gate's behavior is documented; sibling check predicate documents its flags the same way). - Two changeset fragments per the maintainer brief (Added: flag; Fixed: hardening), PR numbers to be backfilled. - tests/decisions.test.cjs review fixes: readRegExpTemplate template escaping (bare ')' SyntaxError), range-aware lattice checker with backreference skip and template unescape, honest A1 contract, lint escape warning. * fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context Suite-caught fixes from the first verify run: - cmdDecisionCoveragePlan now refuses a flag-shaped token as the positional context path: a bare valueless --context stays a positional (sibling parser semantics, unchanged) but reading it as a PATH would turn a caller mistake into a silent 'CONTEXT.md missing' green skip — exactly what #2770's fail-closed law forbids. Now falls through to the missing-context-argument error, as documented. - A8 test compares decoy-positional+flag against flag-with-phase (phase held constant) so the row isolates WHICH context was read; the old form compared against a no-phase invocation that could never match. * chore(#4130): backfill PR number in changeset fragments (PR #4374) --------- Co-authored-by: sim --- .changeset/jolly-otters-click.md | 5 + .changeset/sharp-deer-hop.md | 5 + docs/CONFIGURATION.md | 25 ++ src/check-command-router.cts | 66 +++- src/decisions.cts | 51 ++- tests/check-predicate.test.cjs | 49 +++ tests/decisions.test.cjs | 519 +++++++++++++++++++++++++++++++ 7 files changed, 708 insertions(+), 12 deletions(-) create mode 100644 .changeset/jolly-otters-click.md create mode 100644 .changeset/sharp-deer-hop.md diff --git a/.changeset/jolly-otters-click.md b/.changeset/jolly-otters-click.md new file mode 100644 index 000000000..3c08138ff --- /dev/null +++ b/.changeset/jolly-otters-click.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4374 +--- +**parseDecisions no longer backtracks quadratically on pathological single bullets** — output unchanged on all legal inputs. (#4130) diff --git a/.changeset/sharp-deer-hop.md b/.changeset/sharp-deer-hop.md new file mode 100644 index 000000000..300a678e2 --- /dev/null +++ b/.changeset/sharp-deer-hop.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 4374 +--- +**check decision-coverage-plan accepts --context ** — same convention as sibling check verbs. (#4130) diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index dcd8d14bb..02e7f5700 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -1398,6 +1398,31 @@ overall verification status. The asymmetry is deliberate — by verify time the work is done, and a fuzzy substring miss should not fail an otherwise green phase. +### Invoking the plan gate directly + +The plan-phase translation gate is runnable standalone (the same form the +workflow's gate dispatch uses): + +```bash +gsd_run check decision-coverage-plan +``` + +The context path may also be supplied with the `--context` flag, following +the same convention as the other flag-taking check verbs (e.g. +`check predicate`): + +```bash +gsd_run check decision-coverage-plan --context +gsd_run check decision-coverage-plan --context [] +``` + +The flag wins when both a positional context path and `--context` are given; +the positional form keeps working unchanged. A `--context` with no value is +a caller error — the gate fails closed with the missing-argument error, the +same as calling it with no context path at all, rather than silently +skipping. The phase directory remains a separate positional argument (there +is no `--phase` flag; the workflow caller passes both positionals). + ### How to write decisions the gates accept The discuss-phase template already produces `D-NN`-numbered decisions. diff --git a/src/check-command-router.cts b/src/check-command-router.cts index 380113fc4..9a6508996 100644 --- a/src/check-command-router.cts +++ b/src/check-command-router.cts @@ -289,9 +289,34 @@ function loadDecisionExtraction(contextPath: string): { trackable: Decision[]; o }; } +/** + * `check decision-coverage-plan` — blocking plan-phase decision-coverage gate + * (#2492, #1365 fail-loud, #2770 empty-arg fail-closed). + * + * Invocation (the context path may be supplied EITHER way; #4130 follow-up): + * gsd_run check decision-coverage-plan (positional, the workflow caller's form) + * gsd_run check decision-coverage-plan --context [] + * + * `--context ` follows the sibling flag convention (`check predicate`, + * #2008): `--flag value` pairs parsed by the shared partitionPredicateArgs + * pass, the flag WINNING over a same-purpose positional when both appear, + * and a valueless `--context` counting as no context at all (it falls + * through to the #2770 caller-error branch, not to the "CONTEXT.md missing" + * green skip). The positional form keeps working unchanged — no sibling + * check verb deprecates positionals and the plan-phase workflow passes them. + */ function cmdDecisionCoveragePlan(projectDir: string, args: string[], raw: boolean): void { - const phaseDir = args[2] ? resolvePath(args[2], projectDir) : ''; - const contextArg = args[3]; + // args[0]='check', args[1]=subcommand — partition the REST so flag tokens + // and their values never land in a positional slot. + const { flags, positionals } = partitionPredicateArgs(args.slice(2)); + const phaseDir = positionals[0] ? resolvePath(positionals[0], projectDir) : ''; + // A VALUELESS `--context` stays a bare token in the positionals (sibling + // parser semantics); it must not then be read as the context PATH — a + // `--`-prefixed "path" is a caller mistake, and #2770's law says a missing + // context argument fails CLOSED, never a silent "CONTEXT.md missing" green + // skip. So only a non-flag positional may serve as the context. + const positionalContext = positionals[1] && !positionals[1].startsWith('--') ? positionals[1] : ''; + const contextArg = flags['context'] ?? positionalContext ?? ''; const contextPath = contextArg ? resolvePath(contextArg, projectDir) : ''; if (!gateEnabled(projectDir)) { @@ -1268,21 +1293,41 @@ function buildPredicateDeps() { }; } -/** Parse `--flag value` pairs from an args array into a map (last write wins). */ -function parsePredicateFlags(args: string[]): Record { - const out: Record = {}; +/** + * Split an args array into `--flag value` pairs and the leftover positional + * tokens, in ONE pass, with the semantics `check predicate` established + * (#2008): a `--flag` followed by a non-`--` token consumes it as the value + * (last write wins); a `--flag` with no value stays a bare token and moves to + * the positionals; everything else is positional. `parsePredicateFlags` is + * the flags half of this same pass — there is exactly one parser, so the + * flag-taking check verbs cannot drift apart (#4130 follow-up: `check + * decision-coverage-plan --context ` shares it). + */ +function partitionPredicateArgs(args: string[]): { flags: Record; positionals: string[] } { + const flags: Record = {}; + const positionals: string[] = []; for (let i = 0; i < args.length; i++) { const a = args[i]; if (typeof a !== 'string') continue; - if (!a.startsWith('--')) continue; + if (!a.startsWith('--')) { + positionals.push(a); + continue; + } const key = a.slice(2); const next = args[i + 1]; if (key.length > 0 && typeof next === 'string' && !next.startsWith('--')) { - out[key] = next; + flags[key] = next; i++; + } else { + positionals.push(a); } } - return out; + return { flags, positionals }; +} + +/** Parse `--flag value` pairs from an args array into a map (last write wins). */ +function parsePredicateFlags(args: string[]): Record { + return partitionPredicateArgs(args).flags; } /** @@ -1821,7 +1866,9 @@ function routeCheckCommand({ args, cwd, raw }: RouteCheckCommandOptions): void { // this for any gate whose `check` carries a `predicate` (instead of a `query`), // passing the predicate object as --predicate ''. NOTE: unlike the // `check.query` subcommands above (which take positional phase args), this - // subcommand parses --flag value pairs. + // subcommand is flag-driven. `decision-coverage-plan` above now ALSO accepts + // `--context ` (its positionals still work) — both share + // partitionPredicateArgs, the one flag parser. cmdCheckPredicate(cwd, args, raw); return; } @@ -1850,6 +1897,7 @@ export = { cmdCheckPredicate, buildPredicateDeps, parsePredicateFlags, + partitionPredicateArgs, // Fail-closed phase-scope reader for the api-coverage gate — exported for // in-process failure-injection tests (#2365 review). readPhaseScope, diff --git a/src/decisions.cts b/src/decisions.cts index 9688a914c..b277e08d8 100644 --- a/src/decisions.cts +++ b/src/decisions.cts @@ -17,6 +17,11 @@ * - Outer bullet loop → seam's `iterateBullets` (for the header-fallback path) * * Resolves #1364 (markdown-header + em-dash recall) and #1365 (fail-loud gate). + * + * #4130 follow-up (hardening): the three bullet grammars below consume the + * decision ID atomically and narrow the em-dash first separator, eliminating + * the quadratic-backtracking cliff on pathological single bullets. Output is + * byte-identical on all legal inputs — see the notes at DECISION_ID_SOURCE. */ import { @@ -74,6 +79,28 @@ const NON_TRACKABLE_TAGS = new Set(['informational', 'folded', 'deferred']); */ const DECISION_ID_SOURCE = 'D[0-9]*-[A-Za-z0-9][A-Za-z0-9_-]*'; +/** + * #4130 follow-up (hardening): how the three grammars below CONSUME the ID — + * atomically, via the `(?=(X))\1` lookahead emulation (lookarounds are atomic + * in ECMAScript; the backreference must replay exactly what the lookahead + * captured, so the engine can never give the ID tail back one character at a + * time). That give-back was quadratic driver #1: the tail class + * `[A-Za-z0-9_-]*` overlaps the pre-separator class `[^:*]*` (every id char + * is also `[^:*]`), so on a FAILING bullet the base regex re-split the tail + * O(n) times with an O(n) scan after each — measured ~1.1s @ 40k chars on + * `- **D-` + `a-`×20k (the #4357 review's deferred cliff). + * + * Byte-identical on all legal inputs: a successful match always consumes the + * MAXIMAL id run (the lookahead's own match is exactly that maximal run), and + * the continuation's success depends only on the position of the first + * `:`/`*` (or `*` for the em-dash form) after the id boundary — id chars + * contain neither, so moving the boundary inside the run cannot change + * success or any capture. Group 1 stays the full id (the lookahead's capture + * IS group 1), so handlers keep reading match[1]/[2]/[3] untouched. Pinned by + * the differential property test against a frozen copy of the pre-hardening + * grammars and by the regex-lattice test in tests/decisions.test.cjs. + */ + /** * #4130: the bold lead-in that ATTEMPTS the ID grammar above — used by the * parse-miss guard and the #3939 join regexes, where recognising MORE shapes @@ -90,9 +117,13 @@ const ID_ATTEMPT_SOURCE = 'D(?:[0-9][A-Za-z0-9]*)?-'; * Colon form: `- **D[phase]-NN[ [tags]]:** text` * (#1343: `[^:*]*` subsumes any pre-colon prose, stops at `:**`) * Group 1 captures the FULL id including any phase prefix (#4130). + * The ID is consumed atomically `(?=(…))\1` — see the hardening note above + * the constants (#4130 follow-up); with the tail unable to give back, the + * remaining `[^:*]*:` scan has a single viable split and the whole match is + * linear in line length. */ const bulletColonRe = new RegExp( - `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:\\*\\*\\s*(.*)$`, + `^\\s*-\\s+\\*\\*(?=(${DECISION_ID_SOURCE}))\\1(?:\\s*\\[([^\\]]+)\\])?[^:*]*:\\*\\*\\s*(.*)$`, ); /** @@ -102,9 +133,20 @@ const bulletColonRe = new RegExp( * outside the closing `**`. This form was not handled pre-T1 (bug #1364). * * Accepts both U+2014 em-dash (—) and U+2013 en-dash (–) for robustness. + * + * #4130 follow-up (hardening), quadratic driver #2: the first separator was + * `[^*]*[—–]`, whose leading class ALSO accepts the dash — on a failing + * dash-laden title the engine retried the separator at every dash position + * with an O(n) scan after each (~1.7s @ 40k). Narrowed to `[^*—–]*[—–]`: + * the leading class now excludes the dash, so the separator is the FIRST + * dash — one viable split, single pass. Behavior-preserving because every + * candidate dash lies before the first `*` (the leading class cannot cross + * a star), so the trailing `[^*]*` reaches that same first star from any + * candidate and `**` succeeds or fails identically; no capture involves the + * dash position. The ID is atomic like the other forms (driver #1). */ const bulletEmDashRe = new RegExp( - `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^*]*[—–][^*]*\\*\\*\\s*(.*)$`, + `^\\s*-\\s+\\*\\*(?=(${DECISION_ID_SOURCE}))\\1(?:\\s*\\[([^\\]]+)\\])?[^*—–]*[—–][^*]*\\*\\*\\s*(.*)$`, ); /** @@ -117,9 +159,12 @@ const bulletEmDashRe = new RegExp( * (e.g. `D-07 ratio 3:1:**`) still fails the anchor and falls through to the parse-miss * guard — matching bulletColonRe's `[^:*]*` discipline that the separator colon is the * only colon permitted before `**`. (#1639) + * + * The ID is consumed atomically `(?=(…))\1` like the other forms — the + * hardening note above the constants explains why (#4130 follow-up). */ const bulletTitledColonRe = new RegExp( - `^\\s*-\\s+\\*\\*(${DECISION_ID_SOURCE})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:[^:*]*\\*\\*\\s*(.*)$`, + `^\\s*-\\s+\\*\\*(?=(${DECISION_ID_SOURCE}))\\1(?:\\s*\\[([^\\]]+)\\])?[^:*]*:[^:*]*\\*\\*\\s*(.*)$`, ); /** diff --git a/tests/check-predicate.test.cjs b/tests/check-predicate.test.cjs index ea9658546..f35d7bec0 100644 --- a/tests/check-predicate.test.cjs +++ b/tests/check-predicate.test.cjs @@ -104,3 +104,52 @@ describe('parsePredicateFlags', () => { assert.deepEqual(parsePredicateFlags([]), {}); }); }); + +// ─── #4130 follow-up: partitionPredicateArgs (flags + positionals, one parser) ─ + +/** + * `partitionPredicateArgs` is the single pass behind `parsePredicateFlags`: + * it returns BOTH the --flag value map AND the non-consumed positional tokens + * under the exact same skip/consume/last-wins semantics. `check + * decision-coverage-plan --context ` uses it so the flag and the + * positional surface share one parser with `check predicate` — the two + * parsers cannot diverge because there is only one. + */ +describe('partitionPredicateArgs (#4130 follow-up)', () => { + const { partitionPredicateArgs } = require('../gsd-core/bin/lib/check-command-router.cjs'); + + test('splits --flag value pairs from positionals', () => { + const { flags, positionals } = partitionPredicateArgs( + ['check', 'decision-coverage-plan', '--context', '/tmp/CONTEXT.md', 'phases/01-init'], + ); + assert.deepEqual(flags, { context: '/tmp/CONTEXT.md' }); + assert.deepEqual(positionals, ['check', 'decision-coverage-plan', 'phases/01-init']); + }); + + test('parsePredicateFlags is exactly the flags half (one source of truth)', () => { + const vectors = [ + ['check', 'predicate', '--predicate', '{"kind":"x"}', '--phase-number', '03', '--raw'], + ['--phase-number', '01', '--phase-number', '02'], + ['--predicate', '--phase-number'], + [], + ['--context'], + ['a', '--context', 'b', '--context', 'c', 'd'], + ]; + for (const v of vectors) { + assert.deepEqual(partitionPredicateArgs(v).flags, parsePredicateFlags(v), + `flags half must equal parsePredicateFlags for ${JSON.stringify(v)}`); + } + }); + + test('value that starts with -- is not consumed: both stay flags, neither becomes positional', () => { + const { flags, positionals } = partitionPredicateArgs(['--context', '--other']); + assert.deepEqual(flags, {}); + assert.deepEqual(positionals, ['--context', '--other']); + }); + + test('last write wins; flag values never leak into positionals', () => { + const { flags, positionals } = partitionPredicateArgs(['p1', '--context', 'a', 'p2', '--context', 'b', 'p3']); + assert.equal(flags.context, 'b'); + assert.deepEqual(positionals, ['p1', 'p2', 'p3']); + }); +}); diff --git a/tests/decisions.test.cjs b/tests/decisions.test.cjs index d9219791b..0d07f18bd 100644 --- a/tests/decisions.test.cjs +++ b/tests/decisions.test.cjs @@ -2522,3 +2522,522 @@ describe('check.decision-coverage-verify — phase-prefixed decisions are readab `A plan mentioning D4-01 honors it. Got: ${JSON.stringify(parsed)}`); }); }); + +// ─── #4130 follow-up: check decision-coverage-plan --context ────────── + +/** + * Row-1 failing-first regression for the --context flag (maintainer-directed + * follow-up to #4130, merged as #4357). + * + * Convention mirrored from the ONE flag-driven sibling check verb + * (`check predicate`, src/check-command-router.cts): `--flag value` pairs + * parsed with parsePredicateFlags semantics, `--context ` supplying the + * CONTEXT.md path, the flag WINNING over a same-purpose positional, and the + * positional form kept working (no sibling deprecates positionals; the + * plan-phase workflow caller passes positionals). + * + * Before the fix (probed on the base build): + * - `--context ` alone landed `--context` in the args[2] phase slot → + * plans scanned in a nonexistent `/--context` dir → every + * decision falsely uncovered (passed:false where the positional form + * passes). + * - ` --context ` put the literal `--context` in the context + * slot → silent "CONTEXT.md missing" green skip. + */ +describe('check.decision-coverage-plan — --context flag matches the positional form (#4130 follow-up)', () => { + let tmpDir; + let planningDir; + let phaseDir; + + beforeEach(() => { + tmpDir = createTempProject('gsd-4130fu-'); + planningDir = path.join(tmpDir, '.planning'); + phaseDir = path.join(planningDir, 'phases', '01-init'); + fs.mkdirSync(phaseDir, { recursive: true }); + }); + + afterEach(() => cleanup(tmpDir)); + + /** Invoke the gate with a raw arg vector (after the subcommand). */ + const runDcp = (args) => runGsdTools(['query', 'check.decision-coverage-plan', ...args], tmpDir); + + const coveredContext = () => [ + '# Phase 4 Context', + '', + '', + '- **D4-01:** use the phase-scoped datastore', + '', + '', + ].join('\n'); + + const coveredPlan = () => '# Plan\n## Must Haves\n- D4-01: provision the datastore\n'; + + test('--context alone routes the path into the context slot (no stray flag token anywhere)', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + // The phase dir is a SEPARATE positional; `--context`-only means no phase + // was given, which routes exactly like an empty phase positional. The + // assertion that matters: the flag's VALUE reaches the gate (the decision + // is counted and reported uncovered — parsed from the flag's path), and + // the output is identical to the equivalent positional invocation. + const viaFlag = JSON.parse(runDcp(['--context', contextPath]).output || '{}'); + const viaEmptyPhasePositional = JSON.parse(runDcp(['', contextPath]).output || '{}'); + + assert.deepStrictEqual(viaFlag, viaEmptyPhasePositional, + `--context-only must route like the empty-phase positional.\nflag: ${JSON.stringify(viaFlag)}\npos: ${JSON.stringify(viaEmptyPhasePositional)}`); + assert.strictEqual(viaFlag.total, 1, + `the decision must be read from the flag's path. Got: ${JSON.stringify(viaFlag)}`); + assert.strictEqual(viaFlag.passed, false, 'no phase → no plans scanned → coverage gap, not a parse accident'); + assert.deepStrictEqual((viaFlag.uncovered || []).map((u) => u.id), ['D4-01']); + }); + + test('ROW-1 RED: --context composes (flag value reaches the gate, not the phase slot)', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + const viaFlag = JSON.parse(runDcp([phaseDir, '--context', contextPath]).output || '{}'); + const viaPositional = JSON.parse(runDcp([phaseDir, contextPath]).output || '{}'); + + // Before the fix: the literal `--context` was taken as the context path → + // silent "CONTEXT.md missing" green skip. + assert.deepStrictEqual(viaFlag, viaPositional, + `phase + --context must compose.\nflag: ${JSON.stringify(viaFlag)}\npos: ${JSON.stringify(viaPositional)}`); + assert.strictEqual(viaFlag.passed, true); + }); + + test('ROW-1 RED: --context (flag first) is order-independent', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + const viaFlagFirst = JSON.parse(runDcp(['--context', contextPath, phaseDir]).output || '{}'); + const viaPositional = JSON.parse(runDcp([phaseDir, contextPath]).output || '{}'); + assert.deepStrictEqual(viaFlagFirst, viaPositional); + assert.strictEqual(viaFlagFirst.passed, true); + assert.strictEqual(viaFlagFirst.covered, 1); + }); + + test('ROW-1 RED: --context wins when both flag and positional context are supplied', () => { + // Flag file: one covered decision (passed:true, total:1). + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const flagContext = path.join(phaseDir, 'CONTEXT.md'); + // Positional decoy: a none-present file (would skip with total:0). + const decoyContext = path.join(phaseDir, 'DECOY-CONTEXT.md'); + fs.writeFileSync(decoyContext, '# Nothing decision-shaped here.\n'); + + // Hold the PHASE constant so the comparison isolates WHICH context file + // was read: decoy-positional + flag must equal flag-alone-with-phase. + const viaBoth = JSON.parse(runDcp([phaseDir, decoyContext, '--context', flagContext]).output || '{}'); + const viaFlag = JSON.parse(runDcp([phaseDir, '--context', flagContext]).output || '{}'); + + assert.deepStrictEqual(viaBoth, viaFlag, + `--context must win over the positional context.\nboth: ${JSON.stringify(viaBoth)}\nflag: ${JSON.stringify(viaFlag)}`); + assert.strictEqual(viaBoth.passed, true, 'the flag file (covered decision) must be the one read'); + assert.strictEqual(viaBoth.total, 1); + }); + + test('uncovered decision via --context reports the coverage gap (no false pass, no false fail)', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', '# Plan\n## Must Haves\n- Something unrelated.\n'); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + const viaFlag = JSON.parse(runDcp(['--context', contextPath]).output || '{}'); + const viaPositional = JSON.parse(runDcp([phaseDir, contextPath]).output || '{}'); + assert.deepStrictEqual(viaFlag, viaPositional); + assert.strictEqual(viaFlag.passed, false); + assert.deepStrictEqual((viaFlag.uncovered || []).map((u) => u.id), ['D4-01']); + }); + + test('could-not-parse CONTEXT via --context fails loud exactly like the positional form', () => { + writeContextFile(phaseDir, [ + '', + '- **DEC-01:** an ID grammar the parser does not support', + '', + '', + ].join('\n')); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + const viaFlag = JSON.parse(runDcp(['--context', contextPath]).output || '{}'); + const viaPositional = JSON.parse(runDcp([phaseDir, contextPath]).output || '{}'); + assert.deepStrictEqual(viaFlag, viaPositional); + assert.strictEqual(viaFlag.passed, false); + assert.strictEqual(viaFlag.reason, 'could-not-parse'); + }); + + test('back-compat: the positional form is byte-identical to a no-flag run (control row)', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + const a = runDcp([phaseDir, contextPath]).output; + const b = runDecisionCoveragePlan(phaseDir, contextPath, tmpDir).output; + assert.strictEqual(a, b); + assert.strictEqual(JSON.parse(a || '{}').passed, true); + }); + + test('--context keeps the legitimate green skip', () => { + const missing = path.join(phaseDir, 'NOPE-CONTEXT.md'); + const viaFlag = JSON.parse(runDcp(['--context', missing]).output || '{}'); + const viaPositional = JSON.parse(runDcp([phaseDir, missing]).output || '{}'); + assert.deepStrictEqual(viaFlag, viaPositional); + assert.strictEqual(viaFlag.passed, true); + assert.strictEqual(viaFlag.skipped, true); + assert.strictEqual(viaFlag.reason, 'CONTEXT.md missing'); + }); + + test('valueless trailing --context falls through to the #2770 fail-closed caller error', () => { + // Mirrors the sibling parser (parsePredicateFlags): a `--flag` with no + // value is a boolean, never a path. The caller supplied no context path, + // which #2770 treats as a caller error — NOT as "no CONTEXT.md". + const viaFlag = JSON.parse(runDcp([phaseDir, '--context']).output || '{}'); + assert.strictEqual(viaFlag.passed, false, + `A valueless --context is a caller error, not a green skip. Got: ${JSON.stringify(viaFlag)}`); + assert.strictEqual(viaFlag.reason, 'missing context path argument'); + }); + + test('negative space: decision-coverage-verify keeps its positional arg surface (flag is plan-only)', () => { + writeContextFile(phaseDir, coveredContext()); + writePlanFile(phaseDir, '01', coveredPlan()); + const contextPath = path.join(phaseDir, 'CONTEXT.md'); + + // The verify gate is out of scope by directive: its positional contract + // is unchanged, and it does not gain --context handling. + const verify = JSON.parse(runGsdTools(['query', 'check.decision-coverage-verify', phaseDir, contextPath], tmpDir).output || '{}'); + assert.strictEqual(verify.total, 1); + assert.strictEqual(verify.honored, 1); + + const verifyFlagForm = JSON.parse(runGsdTools(['query', 'check.decision-coverage-verify', phaseDir, '--context', contextPath], tmpDir).output || '{}'); + assert.strictEqual(verifyFlagForm.skipped, true, + `verify must keep reading args[3] positionally (unchanged base behavior). Got: ${JSON.stringify(verifyFlagForm)}`); + assert.strictEqual(verifyFlagForm.reason, 'CONTEXT.md missing'); + }); +}); + +// ─── #4130 follow-up: parseDecisions regex hardening (quadratic backtracking) ─ + +/** + * Row-1 failing-first regression for the regex-seam hardening. + * + * The #4357 security review measured ~740ms @ 40k chars on pathological + * single bullets and deferred the fix here. Mechanism (10-diagnosis.md): + * (1) the ID tail `[A-Za-z0-9][A-Za-z0-9_-]*` overlaps the pre-separator + * class `[^:*]*`, so a failing match re-splits the tail O(n) times with + * an O(n) scan each — O(n²); + * (2) the em-dash form's `[^*]*[—–]` first separator can retry at every dash + * position with an O(n) scan after each — O(n²). + * + * Hardening under test (byte-identical on all legal inputs): + * - atomic ID via the `(?=(X))\1` lookahead emulation (group 1 unchanged); + * - em-dash first separator narrowed to `[^*—–]*[—–]` (FIRST dash, unique + * split point). + * + * Repo rule: no wall-time asserts. Node's RegExp engine exposes no injectable + * step counter, so no honest deterministic op-count proxy exists (documented + * in 10-diagnosis.md); the pin is structural (lattice), differential (vs the + * frozen pre-hardening reference below), and correctness-at-scale. + */ +describe('parseDecisions hardening — regex lattice pins the mechanism (#4130 follow-up)', () => { + const SRC = path.resolve(__dirname, '../src/decisions.cts'); + + /** Extract a `const NAME = '';` single-quoted string literal. */ + function readStringConst(source, name) { + const m = source.match(new RegExp(`^const ${name} = '([^']*)';$`, 'm')); + assert.ok(m, `source must declare const ${name} as a plain string literal`); + return m[1]; + } + + /** + * Extract a `new RegExp(`...`)` template body, substitute ${CONSTS}, and + * unescape the template-literal double backslashes — the result is the + * exact regex SOURCE STRING the module compiles. + */ + function readRegExpTemplate(source, varName, consts) { + const m = source.match(new RegExp(`^const ${varName} = new RegExp\\(\n \`([^\`]+)\`,?\n?\\);?`, 'm')); + assert.ok(m, `source must declare ${varName} as a template-literal RegExp`); + let out = m[1]; + for (const [name, value] of Object.entries(consts)) { + out = out.split(`\${${name}}`).join(value); + } + assert.ok(!out.includes('$' + '{'), `unsubstituted template placeholder in ${varName}`); + return out.replace(/\\\\/g, '\\'); + } + + const source = fs.readFileSync(SRC, 'utf8'); + const idSource = readStringConst(source, 'DECISION_ID_SOURCE'); + const idAttempt = readStringConst(source, 'ID_ATTEMPT_SOURCE'); + const consts = { DECISION_ID_SOURCE: idSource, ID_ATTEMPT_SOURCE: idAttempt }; + const colonSrc = readRegExpTemplate(source, 'bulletColonRe', consts); + const emDashSrc = readRegExpTemplate(source, 'bulletEmDashRe', consts); + const titledSrc = readRegExpTemplate(source, 'bulletTitledColonRe', consts); + + test('ROW-1 RED: the ID grammar constants are unchanged (the parity pin survives hardening)', () => { + assert.strictEqual(idSource, 'D[0-9]*-[A-Za-z0-9][A-Za-z0-9_-]*'); + assert.strictEqual(idAttempt, 'D(?:[0-9][A-Za-z0-9]*)?-'); + }); + + test('ROW-1 RED: all three bullet grammars consume the ID atomically (no tail re-split)', () => { + // The (?=(X))\1 lookahead emulation is what makes the ID give-back + // impossible: lookarounds are atomic in ECMAScript, and the backreference + // must replay exactly what the lookahead captured. On the base the + // sources had `(D...-...)` bare — the quadratic driver #1. + const atomic = `(?=(${idSource}))\\1`; + for (const [name, src] of [['bulletColonRe', colonSrc], ['bulletEmDashRe', emDashSrc], ['bulletTitledColonRe', titledSrc]]) { + assert.ok(src.includes(atomic), `${name} must wrap the ID in the atomic (?=(X))\\1 emulation:\n${src}`); + } + }); + + test('ROW-1 RED: the em-dash first separator is narrowed to the FIRST dash (no dash re-split)', () => { + // `[^*]*[—–]` admits O(k) separator split points on a dash-laden title; + // `[^*—–]*[—–]` has exactly one. The narrowing is behavior-preserving + // because every candidate dash lies before the first `*` and the second + // `[^*]*` scan reaches that same first star from any candidate. + assert.ok(emDashSrc.includes('[^*—–]*[—–]'), + `bulletEmDashRe must use the narrowed first separator:\n${emDashSrc}`); + assert.ok(!emDashSrc.includes('[^*]*[—–]'), + `bulletEmDashRe must not retain the overlapping first separator:\n${emDashSrc}`); + }); + + test('lattice: no unbounded class quantifier is immediately followed by an atom its class accepts', () => { + // The adjacency that admitted both quadratic drivers: `C*` directly + // followed by an atom that can start with a char C also accepts lets the + // engine trade characters between the two — O(n) splits × O(n) rescans. + // After the hardening every unbounded bracketed class run in the three + // grammars is followed by a token disjoint from its class (or by a group + // boundary / the atomic backreference replay, which cannot re-split). + const joined = `${colonSrc}\n${emDashSrc}\n${titledSrc}`; + const quantifiers = [...joined.matchAll(/\[((?:[^\]\\]|\\.)*)\](\*|\+)/g)]; + assert.ok(quantifiers.length >= 6, 'expected the seam\'s class quantifiers to be found'); + + /** Membership predicate for a class BODY, honoring ranges and escapes. */ + function classPredicate(body) { + const negated = body.startsWith('^'); + const inner = negated ? body.slice(1) : body; + const members = new Set(); + const chars = [...inner]; + for (let i = 0; i < chars.length; i++) { + let ch = chars[i]; + if (ch === '\\' && i + 1 < chars.length) ch = chars[++i]; + // Range: a-b where a and b are single member chars. + if (chars[i + 1] === '-' && chars[i + 2] !== undefined && chars[i + 2] !== ']') { + const lo = ch; + let hi = chars[i + 2]; + if (hi === '\\' && chars[i + 3] !== undefined) { i += 3; hi = chars[i]; } else { i += 2; } + for (let c = lo.charCodeAt(0); c <= hi.charCodeAt(0); c++) members.add(String.fromCharCode(c)); + continue; + } + members.add(ch); + } + return (ch) => (negated ? !members.has(ch) : members.has(ch)); + } + + /** + * First-set of the atom that follows a quantified class, as a membership + * predicate. `null` = boundary — group close, anchor, end, or the `\1` + * backreference of the atomic wrapper (its first-set is the captured id + * run, replayed verbatim: it cannot trade characters with the quantifier, + * which is the entire point of the wrapper). + */ + function nextAtomFirstSet(src, at) { + if (at >= src.length) return null; + const c = src[at]; + if (c === ')' || c === '$' || c === '|') return null; + if (c === '\\') { + const nxt = src[at + 1]; + if (nxt >= '0' && nxt <= '9') return null; // backreference replay — skip + return (ch) => ch === nxt; + } + if (c === '[') { + const close = src.indexOf(']', at); + return classPredicate(src.slice(at + 1, close)); + } + return (ch) => ch === c; + } + + for (const m of quantifiers) { + const body = m[1]; + const quant = m[2]; + const classAccepts = classPredicate(body); + const firstSet = nextAtomFirstSet(joined, m.index + m[0].length); + if (firstSet === null) continue; + // A witness char the class accepts that the following atom also accepts. + const ALPHABET = '*:-[]()Ds01_—–\\ \tnA'; + const witness = [...ALPHABET].find((ch) => classAccepts(ch) && firstSet(ch)); + assert.ok(witness === undefined, + `unbounded quantifier [${body}]${quant} is followed by an atom accepting '${witness}' which its class also accepts — adjacency overlap:\n${joined.slice(m.index, m.index + m[0].length + 10)}`); + } + }); + + test('lattice: capture-group indices are preserved (id, tags, body)', () => { + // (?=(X))\1 keeps group 1 = the full id (the lookahead's capture IS group + // 1), so the handlers' match[1]/[2]/[3] reads stay untouched. Pin the + // count so a refactor cannot silently renumber the groups. + for (const [name, src] of [['bulletColonRe', colonSrc], ['bulletEmDashRe', emDashSrc], ['bulletTitledColonRe', titledSrc]]) { + const groups = (src.match(/\(/g) || []).length - (src.match(/\(\?:/g) || []).length - (src.match(/\(\?=/g) || []).length; + assert.strictEqual(groups, 3, `${name} must keep exactly 3 capturing groups (id/tags/body):\n${src}`); + } + }); +}); + +describe('parseDecisions hardening — byte-identical vs the pre-hardening reference (#4130 follow-up)', () => { + /** + * FROZEN REFERENCE — the three bullet grammars exactly as they shipped on + * origin/next @ e6d047decc (PR #4357). The hardened module must agree with + * this reference on match/no-match AND all capture groups for every input + * the generator can produce. If the reference and the module ever disagree, + * behavior drifted — this is the "pure hardening" contract. + */ + const REF_ID = 'D[0-9]*-[A-Za-z0-9][A-Za-z0-9_-]*'; + const refColon = new RegExp(`^\\s*-\\s+\\*\\*(${REF_ID})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:\\*\\*\\s*(.*)$`); + const refEmDash = new RegExp(`^\\s*-\\s+\\*\\*(${REF_ID})(?:\\s*\\[([^\\]]+)\\])?[^*]*[—–][^*]*\\*\\*\\s*(.*)$`); + const refTitled = new RegExp(`^\\s*-\\s+\\*\\*(${REF_ID})(?:\\s*\\[([^\\]]+)\\])?[^:*]*:[^:*]*\\*\\*\\s*(.*)$`); + const refGuard = /^\s*-\s+\*\*D(?:[0-9][A-Za-z0-9]*)?-/; + const refBoldLeadIn = /^\s*-\s+\*\*[A-Z]+[0-9]*-[A-Za-z0-9]/m; + const refToken = /\bD[0-9]*-[A-Za-z0-9]/m; + + /** + * The expected single-bullet outcome, computed by the frozen reference: + * the three grammars in the module's precedence order, then the parse-miss + * guard, then the FIX A evidence detectors (bold-lead-in / bare token) — + * exactly the module's single-line block-path decision order. + */ + function referenceOutcome(line) { + const m = refColon.exec(line) || refEmDash.exec(line) || refTitled.exec(line); + if (m) { + const tags = m[2] ? m[2].split(',').map((t) => t.trim().toLowerCase()).filter(Boolean) : []; + return { + outcome: 'parsed', + decisions: [{ + id: m[1], + text: (m[3] || '').trim(), + category: '', + tags, + trackable: !tags.some((t) => ['informational', 'folded', 'deferred'].includes(t)), + }], + }; + } + if (refGuard.test(line)) return { outcome: 'could-not-parse', decisions: [] }; + if (refBoldLeadIn.test(line) || refToken.test(line)) return { outcome: 'could-not-parse', decisions: [] }; + return { outcome: 'none-present', decisions: [] }; + } + + // Generator: adversarial single-line bullets around the seam's alphabet. + const idArb = fc.oneof( + fc.integer({ min: 1, max: 99 }).map((n) => `D-${String(n).padStart(2, '0')}`), + fc.tuple(fc.integer({ min: 1, max: 12 }), fc.integer({ min: 1, max: 99 })) + .map(([p, n]) => `D${p}-${String(n).padStart(2, '0')}`), + fc.constantFrom('D-INFRA-01', 'D-7', 'D-carry_2', 'D4x-01', 'DEC-01', 'D-notes'), + ); + const fuzzArb = (maxWords) => fc.array( + fc.constantFrom('use', 'the:', 'a*', '**', '—', '–', '[x]', 'y]', 'ratio 3:1', '-', 'D4-01', 'x,', 'why', 'ok', '**D-99:', 'until'), + { maxLength: maxWords }, + ).map((w) => w.join(' ')); + const sepArb = fc.constantFrom(':', ' — ', ': ', ' —', ':** ', ' '); + const leadArb = fc.constantFrom('', ' ', '\t'); + + const lineArb = fc.tuple(leadArb, idArb, fc.constantFrom('', ' [informational]', ' [a, b]', ' [unterminated'), sepArb, fuzzArb(14), fuzzArb(10), fc.boolean()) + .map(([lead, id, tags, sep, mid, tail, boldClose]) => { + const close = boldClose ? '**' : ''; + return `${lead}- **${id}${tags}${sep}${mid}${close} ${tail}`; + }); + + test('property: hardened module === frozen pre-hardening reference on every generated bullet', () => { + fc.assert(fc.property(lineArb, (line) => { + const expected = referenceOutcome(line); + const got = extractDecisions(inBlock(line)); + assert.strictEqual(got.outcome, expected.outcome, + `outcome drifted on ${JSON.stringify(line)}: got ${got.outcome}, want ${expected.outcome}`); + assert.deepStrictEqual(got.decisions, expected.decisions, + `decisions drifted on ${JSON.stringify(line)}:\ngot: ${JSON.stringify(got.decisions)}\nwant: ${JSON.stringify(expected.decisions)}`); + return true; + })); + }); + + test('em-dash titles with MULTIPLE dashes capture identically to the reference (first-dash narrowing)', () => { + fc.assert(fc.property( + decisionIdArb, + fuzzArb(6), + fuzzArb(6), + (id, title, body) => { + const line = `- **${id} — ${title} — ${title}** ${body}`; + const expected = referenceOutcome(line); + const got = extractDecisions(inBlock(line)); + assert.strictEqual(got.outcome, expected.outcome); + assert.deepStrictEqual(got.decisions, expected.decisions, + `multi-dash title drifted on ${JSON.stringify(line)}`); + return true; + }, + )); + }); + + test('the #4130/#3939/#1639 fixture grammar round-trips byte-identically', () => { + const fixtures = [ + ['- **D-01:** a short single-line decision.', 'D-01', 'a short single-line decision.'], + ['- **D4-01:** phase-prefixed.', 'D4-01', 'phase-prefixed.'], + ['- **D12-01:** two-digit phase.', 'D12-01', 'two-digit phase.'], + ['- **D-INFRA-01:** alnum tail.', 'D-INFRA-01', 'alnum tail.'], + ['- **D-01 [informational]:** tagged.', 'D-01', 'tagged.'], + ['- **D4-01 — title** body here', 'D4-01', 'body here'], + ['- **D-01 — a — b — c** body', 'D-01', 'body'], + ['- **D-01: Title.** body', 'D-01', 'body'], + ['- **D-01 pre-colon prose:** text', 'D-01', 'text'], + ]; + for (const [line, wantId, wantText] of fixtures) { + const r = extractDecisions(inBlock(line)); + assert.strictEqual(r.outcome, 'parsed', `fixture must still parse: ${JSON.stringify(line)}`); + assert.strictEqual(r.decisions[0].id, wantId, `id drifted on ${JSON.stringify(line)}`); + assert.strictEqual(r.decisions[0].text, wantText, `text drifted on ${JSON.stringify(line)}`); + } + // Typo'd prefix and prose labels keep their #4130 outcomes. + assert.strictEqual(extractDecisions(inBlock('- **D4x-01:** typo')).outcome, 'could-not-parse'); + assert.strictEqual(extractDecisions(inBlock('- **Deferred-until-X:** prose')).outcome, 'none-present'); + }); +}); + +describe('parseDecisions hardening — pathological single bullets terminate correctly (#4130 follow-up)', () => { + // The #4357 cliff shapes at full scale. NO wall-time assert (repo rule): + // under the hardening these complete in well under a millisecond each; if + // the quadratic ambiguity is ever reintroduced these become CI timeouts, + // never false passes. What is asserted is the CORRECT outcome. + test('40k hyphen-laden ID tail (colon cliff shape) → could-not-parse via the guard', () => { + const bullet = '- **D-' + 'a-'.repeat(20000); + const r = extractDecisions(inBlock(bullet)); + assert.strictEqual(r.outcome, 'could-not-parse', + 'the malformed mega-bullet must fail loud (parse-miss guard), not hang or vanish'); + assert.deepStrictEqual(r.decisions, []); + }); + + test('40k dash run after the separator (em-dash cliff shape) → could-not-parse via the guard', () => { + const bullet = '- **D-01 —' + '–'.repeat(40000 - 10); + const r = extractDecisions(inBlock(bullet)); + assert.strictEqual(r.outcome, 'could-not-parse'); + assert.deepStrictEqual(r.decisions, []); + }); + + test('40k colon run (titled-colon cliff shape) → could-not-parse via the guard', () => { + const bullet = '- **D-01: ' + 'x: '.repeat(13000); + const r = extractDecisions(inBlock(bullet)); + assert.strictEqual(r.outcome, 'could-not-parse'); + }); + + test('40k LEGAL single-line decision parses with its text byte-identical (no clamp)', () => { + const text = 'use the phase-scoped datastore '.repeat(1400).trim(); + const bullet = `- **D4-01:** ${text}`; + const r = extractDecisions(inBlock(bullet)); + assert.strictEqual(r.outcome, 'parsed', 'a legal mega-bullet must parse — no line-length clamp exists'); + assert.strictEqual(r.decisions.length, 1); + assert.strictEqual(r.decisions[0].id, 'D4-01'); + assert.strictEqual(r.decisions[0].text, text); + }); + + test('40k LEGAL wrapped-form decision still joins and parses (cliff shapes do not regress #3939)', () => { + const text = 'provision the datastore '.repeat(1500).trim(); + const md = `\n- **D4-01:** ${text}\n\n`; + const r = extractDecisions(md); + assert.strictEqual(r.outcome, 'parsed'); + assert.strictEqual(r.decisions[0].text, text); + }); +}); From c95b734145da00b062b3b1793a1670484dc27b43 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 02:55:41 -0400 Subject: [PATCH 008/166] fix(#4136): compute the Incorporated status; stop re-grafting superseded patches (#4373) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4136): failing-first rows for the unreachable incorporated status Folded block bug-4136-reapply-incorporated-status locks the --classify contract (incorporated / needs_merge / unknown, never-incorporated guards, cycle end-to-end) plus the workflow-contract rows; REASON gains OK_UNVALIDATED_BASELINE in both shape-locks; the #2994 invocation count moves 1 -> 2 (classify + gate). All red until the verifier grows --classify and the workflow consumes it. * fix(#4136): compute the incorporated status; stop re-grafting superseded patches Add --classify pre-merge mode to the deterministic verifier: with a hash-validated pristine baseline, a file whose every significant user-added line is already present verbatim in the freshly installed version is classified incorporated (new frozen CLASSIFICATION enum, structured --json report, always exit 0 — the binding gate stays the post-merge run). Drifted (#3657), absent (#934), unvalidated (new OK_UNVALIDATED_BASELINE), and no-baseline runs classify unknown — a false incorporated silently retires a live customization, so only a confirmed baseline may ever confirm adoption. The baseline-resolution block moves out of verifyFile into a shared resolvePristineBaseline so the gate and the classifier cannot drift on what counts as a usable baseline; gate behavior is byte-identical. reapply-patches.md step 4 gains the pre-flight classifier invocation and the not-re-grafted contract: incorporated files are left exactly as shipped (their hash then re-converges with the manifest, ending the backup cycle), statuses feed steps 3/7, and the merge rules gain the already-present-verbatim arm. Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract * fix(#4136): address review findings on the workflow contract Drop the unused INCORPORATED_COUNT shell variable (standards pass) and close the all-files-incorporated gap in the hunk-table guidance: emit a header row plus a note line so the step 5b absent-table halt is not tripped when nothing was merged (spec pass). Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract * chore(#4136): backfill changeset fragment with PR 4373 * test(#4136): lock that a 4145-recovered baseline can confirm incorporation The hash-first orphan recovery merged with next (PR #4364) lands in the shared resolvePristineBaseline as a validated resolution; this row pins the composition so a future change cannot quietly downgrade recovered baselines to unknown and silently disable incorporated detection for prefix-less installs. --------- Co-authored-by: sim --- .changeset/agile-wasps-chatter.md | 5 + gsd-core/bin/verify-reapply-patches.cjs | 403 ++++++++++++----- gsd-core/workflows/reapply-patches.md | 58 ++- tests/reapply-patches.test.cjs | 10 +- tests/reapply-verify-hunks.test.cjs | 546 ++++++++++++++++++++++++ 5 files changed, 921 insertions(+), 101 deletions(-) create mode 100644 .changeset/agile-wasps-chatter.md diff --git a/.changeset/agile-wasps-chatter.md b/.changeset/agile-wasps-chatter.md new file mode 100644 index 000000000..f8e383d9f --- /dev/null +++ b/.changeset/agile-wasps-chatter.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4373 +--- +**`/gsd-update --reapply` no longer re-grafts customizations that upstream already adopted** — the documented `Incorporated` per-file status is now computed by a deterministic pre-flight classifier (hash-validated pristine baseline + every significant user-added line already present verbatim in the new version), so superseded patches are reported as already upstream instead of being silently re-applied on every future update cycle. (#4136) diff --git a/gsd-core/bin/verify-reapply-patches.cjs b/gsd-core/bin/verify-reapply-patches.cjs index c052fc3ec..fedc65039 100755 --- a/gsd-core/bin/verify-reapply-patches.cjs +++ b/gsd-core/bin/verify-reapply-patches.cjs @@ -17,8 +17,11 @@ * # false-positive halts beat silent successes * # on lost content) * [--json] # emit JSON report instead of human text + * [--classify] # pre-merge mode: classify each backed-up + * # file as incorporated / needs_merge / + * # unknown (#4136); always exits 0 * - * Exit codes: + * Exit codes (default gate mode): * 0 — every user-added line is present in the merged file (gate passes) * 1 — at least one missing line in at least one file (gate fails) * 2 — usage / structural error (e.g. patches dir missing) @@ -27,6 +30,15 @@ * yes/no" reporting per hunk. The LLM was filling in `yes` even when content * had been silently dropped. Moving the check to a deterministic script is the * durability fix. + * + * Bug #4136 adds --classify (pre-merge mode, always exit 0 — informational; + * the binding gate remains the post-merge default run): per backed-up file, + * decides whether the user's modification was already adopted upstream + * ("Incorporated", reapply-patches.md Step 4 item 6). The classification is + * ONLY produced from a hash-validated pristine baseline with every + * significant user-added line present verbatim in the freshly installed + * version — a false Incorporated silently retires a live customization, + * which is worse than no Incorporated at all. */ const fs = require('node:fs'); @@ -43,16 +55,17 @@ const SIGNIFICANT_MIN_CHARS = 12; const GSD_HOOK_VERSION_LINE_RE = /^(?:\/\/|#)\s*gsd-hook-version:\s*\S+\s*$/i; function parseArgs(argv) { - const opts = { patchesDir: null, configDir: null, pristineDir: null, json: false }; + const opts = { patchesDir: null, configDir: null, pristineDir: null, json: false, classify: false }; for (let i = 0; i < argv.length; i++) { const arg = argv[i]; if (arg === '--patches-dir') opts.patchesDir = argv[++i]; else if (arg === '--config-dir') opts.configDir = argv[++i]; else if (arg === '--pristine-dir') opts.pristineDir = argv[++i]; else if (arg === '--json') opts.json = true; + else if (arg === '--classify') opts.classify = true; else if (arg === '--help' || arg === '-h') { process.stdout.write( - 'usage: verify-reapply-patches.cjs --patches-dir --config-dir [--pristine-dir ] [--json]\n', + 'usage: verify-reapply-patches.cjs --patches-dir --config-dir [--pristine-dir ] [--json] [--classify]\n', ); throw new ExitError(0); } else { @@ -183,12 +196,134 @@ const REASON = Object.freeze({ // block" rather than "ignore everything" — it only applies when the hash was // recorded (modern installer) but the file is absent (specific gap). OK_NO_BASELINE: 'ok_no_baseline', + // Bug #4136: an on-disk gsd-pristine/ snapshot exists for this file but + // backup-meta.json records no pristine_hashes entry for it (older + // installer), so nothing confirms the snapshot is the baseline the backup + // was captured against — it could be a drifted newer-version snapshot the + // #3657 guard cannot see. The default gate still uses it as the diff + // baseline (pre-#3657 behaviour, unchanged); --classify refuses to confirm + // adoption on an unvalidated baseline, so a file is never classified + // Incorporated on the snapshot's say-so alone. + OK_UNVALIDATED_BASELINE: 'ok_unvalidated_baseline', FAIL_INSTALLED_MISSING: 'fail_installed_missing', FAIL_INSTALLED_NOT_REGULAR_FILE: 'fail_installed_not_regular_file', FAIL_READ_ERROR: 'fail_read_error', FAIL_USER_LINES_MISSING: 'fail_user_lines_missing', }); +/** + * Bug #4136: stable per-file classification codes for --classify (pre-merge) + * mode. Tests assert via `assert.equal(result.classification, CLASSIFICATION.X)` + * rather than regex-matching prose, mirroring the REASON enum contract. + * + * INCORPORATED — hash-validated pristine baseline confirms that every + * significant user-added line is already present verbatim + * in the freshly installed version: upstream adopted the + * customization. The workflow must NOT re-apply the diff. + * NEEDS_MERGE — validated baseline, but at least one significant + * user-added line is absent from the fresh install; the + * missing lines are listed for the merge step. + * UNKNOWN — no confirmable baseline (drift #3657, absent #934, + * unvalidated, or no --pristine-dir), zero significant + * user-added delta, or a structural failure. Never + * Incorporated. + */ +const CLASSIFICATION = Object.freeze({ + INCORPORATED: 'incorporated', + NEEDS_MERGE: 'needs_merge', + UNKNOWN: 'unknown', +}); + +/** + * Typed outcome of resolving one file's pristine baseline. Shared by the + * post-merge gate (verifyFile) and the pre-merge classifier (classifyFile) + * so the two can never drift on what counts as a usable baseline (#4136). + */ +const PRISTINE_RESOLUTION = Object.freeze({ + VALIDATED: 'validated', // recorded hash matches the resolved snapshot + UNVALIDATED: 'unvalidated', // snapshot present, no recorded hash to confirm it + DRIFTED: 'drifted', // recorded hash mismatches the snapshot (#3657) + ABSENT_RECORDED: 'absent_recorded', // recorded hash, no snapshot anywhere (#934/#4145) + OVERBROAD: 'overbroad', // no pristine dir, or a non-file at the path +}); + +/** + * Resolve the pristine baseline for one backed-up file, applying the #3657 + * drift guard, the #4145 hash-first recovery, and the #934 absent-baseline + * guard. Extracted from verifyFile's inline block so --classify reasons over + * the exact same baseline semantics the post-merge gate enforces. + */ +function resolvePristineBaseline({ relPath, pristineDir, pristineHashes }) { + const hashKey = relPath.replace(/\\/g, '/'); + const recordedHash = pristineHashes && pristineHashes[hashKey]; + if (pristineDir) { + const pristinePath = path.join(pristineDir, relPath); + let pristinePathExists = false; + try { + const stat = fs.statSync(pristinePath); + pristinePathExists = true; // path exists (any type) + if (stat.isFile()) { + const candidate = fs.readFileSync(pristinePath, 'utf8'); + if (recordedHash) { + if (sha256(candidate) === recordedHash) { + // Hash matches: the on-disk pristine is the correct baseline. + return { resolution: PRISTINE_RESOLUTION.VALIDATED, content: candidate }; + } + // Hash mismatch: the on-disk gsd-pristine/ was refreshed to a newer + // GSD version after the backup was captured (Bug #3657). Using it as + // the diff baseline would invert the delta and produce false + // FAIL_USER_LINES_MISSING reports. + return { resolution: PRISTINE_RESOLUTION.DRIFTED, content: null }; + } + // No recorded hash for this file (older installer or absent + // backup-meta) — the default gate uses the on-disk pristine as-is + // (pre-fix behaviour); --classify treats it as unvalidated. + return { resolution: PRISTINE_RESOLUTION.UNVALIDATED, content: candidate }; + } + // Non-file at pristinePath (e.g. a directory): fall through to + // over-broad mode below, which is safe and conservative. + } catch { + // Pristine stat threw — path is absent (ENOENT) or inaccessible. + // pristinePathExists stays false. + } + + // Bug #4145: the canonical join missed, but the recorded hash is the + // baseline authority the #3657 drift guard already trusts. Before + // reporting ABSENT_RECORDED, scan gsd-pristine/ for byte-identical content + // (an earlier release may have stored the snapshot without the gsd-core/ + // prefix). An exact sha-256 match cannot be the wrong baseline, and + // gsd-pristine/ holds only backed-up files, so the scan is small. The + // canonical path itself is excluded — a mismatching file at the joined + // path is drift (#3657), never re-adopted through the scan. A recovered + // baseline is hash-confirmed by construction, so it VALIDATES. + if (!pristinePathExists && recordedHash) { + try { + const recoveredRel = findPristineByHash(pristineDir, recordedHash, hashKey); + if (recoveredRel) { + return { + resolution: PRISTINE_RESOLUTION.VALIDATED, + content: fs.readFileSync(path.join(pristineDir, recoveredRel), 'utf8'), + }; + } + } catch { + // scan or read failure — fall through to the ABSENT_RECORDED posture + } + } + + if (pristinePathExists) { + // Present but not a regular file — over-broad mode is the safe side. + return { resolution: PRISTINE_RESOLUTION.OVERBROAD, content: null }; + } + // Bug #934: recordedHash is present (modern installer) but no + // hash-matching pristine exists anywhere under gsd-pristine/ (the stat + // missed and the #4145 recovery found nothing). + if (recordedHash) { + return { resolution: PRISTINE_RESOLUTION.ABSENT_RECORDED, content: null }; + } + } + return { resolution: PRISTINE_RESOLUTION.OVERBROAD, content: null }; +} + /** * #4086: resolve where a backed-up file's INSTALLED counterpart lives. * Primary is the config-dir-relative join (the manifest key's native form). @@ -292,102 +427,32 @@ function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashe // Normalize to forward slashes so the key lookup matches on Windows // where path.join produces backslash-separated relPath values but // backup-meta.json stores keys written with forward slashes. - const hashKey = relPath.replace(/\\/g, '/'); - const recordedHash = pristineHashes && pristineHashes[hashKey]; + // #4136: the inline baseline block (stat/read + #3657 drift guard + #4145 + // hash-first recovery + #934 absent guard) moved into resolvePristineBaseline + // so the pre-merge classifier reasons over the exact same semantics. + const { resolution, content: pristineContent } = + resolvePristineBaseline({ relPath, pristineDir, pristineHashes }); - let pristineContent = null; - if (pristineDir) { - const pristinePath = path.join(pristineDir, relPath); - // Bug #934: track whether the pristine path EXISTS on disk (stat did not - // throw ENOENT). A regular file that fails to read, or a non-file path - // (e.g. a directory accidentally placed at the pristine path), is treated - // as "present but unusable" — we fall to over-broad mode (safe side). - // OK_NO_BASELINE is reserved for the strictly absent case: stat throws, - // meaning the file was never written (the gap the bug describes). - let pristinePathExists = false; - try { - const stat = fs.statSync(pristinePath); - pristinePathExists = true; // path exists (any type) - if (stat.isFile()) { - const candidate = fs.readFileSync(pristinePath, 'utf8'); - // Bug #3657: if backup-meta.json recorded a pristine_hash for this - // file, validate that the on-disk pristine matches it. A mismatch - // means the installer refreshed gsd-pristine/ to a newer GSD version - // after the backup was captured. Using the wrong-version pristine as - // the diff baseline inverts the delta: upstream removals appear as - // "user-added lines that must survive", causing FAIL_USER_LINES_MISSING - // false positives. When the hash is stale, skip the pristine and fall - // through to over-broad mode (every significant backup line is checked). - // Over-broad mode never false-fails for a different reason because all - // backup lines that are genuinely user-added will still be present in a - // correctly merged install. - if (recordedHash) { - if (sha256(candidate) === recordedHash) { - // Hash matches: the on-disk pristine is the correct baseline. - pristineContent = candidate; - } else { - // Hash mismatch: the on-disk gsd-pristine/ was refreshed to a newer - // GSD version after the backup was captured. Using it as the diff - // baseline would invert the delta and produce false FAIL_USER_LINES_MISSING - // reports (Bug #3657). Report the file as ok with a diagnostic code - // so the gate does not false-fail; a re-anchor or git-aware baseline - // step is required to verify this file correctly. - result.reason = REASON.OK_PRISTINE_DRIFT_DETECTED; - return result; - } - } else { - // No recorded hash for this file (older installer or absent - // backup-meta) — use the on-disk pristine as-is (pre-fix behaviour). - pristineContent = candidate; - } - } - // Non-file at pristinePath (e.g. a directory): stat succeeded so - // pristinePathExists is true; we fall through to over-broad mode below, - // which is safe and conservative. - } catch { - // Pristine stat threw — path is absent (ENOENT) or inaccessible. - // pristinePathExists stays false. - } - - // Bug #4145: the canonical join missed, but the recorded hash is the - // baseline authority the #3657 drift guard already trusts. Before - // reporting OK_NO_BASELINE, scan gsd-pristine/ for byte-identical content - // (an earlier release may have stored the snapshot without the gsd-core/ - // prefix). An exact sha-256 match cannot be the wrong baseline, and - // gsd-pristine/ holds only backed-up files, so the scan is small. The - // canonical path itself is excluded — a mismatching file at the joined - // path is drift (#3657), never re-adopted through the scan. - if (!pristinePathExists && recordedHash) { - try { - const recoveredRel = findPristineByHash(pristineDir, recordedHash, hashKey); - if (recoveredRel) { - pristineContent = fs.readFileSync(path.join(pristineDir, recoveredRel), 'utf8'); - pristinePathExists = true; - } - } catch { - // scan or read failure — fall through to the OK_NO_BASELINE posture - } - } - - // Bug #934: recordedHash is present (modern installer) but no hash-matching - // pristine exists anywhere under gsd-pristine/ (stat threw above AND the - // #4145 scan found nothing). This means - // saveLocalPatches recorded a hash but could not write the corresponding - // gsd-pristine/ file (the only candidate was discarded because it was from - // a newer release). Falling to over-broad mode here would treat every - // upstream-changed line as a "user-added line that must survive", producing - // false FAIL_USER_LINES_MISSING for each upstream removal. Since we - // cannot reason correctly without a baseline, the safe answer is advisory/ - // non-blocking: return OK_NO_BASELINE and let the caller decide. - // NOTE: this guard fires ONLY when the baseline path is absent (stat threw - // and nothing matched by hash), not when the path is present but non-file — - // in that case over-broad mode is safer. - if (!pristinePathExists && recordedHash) { - result.reason = REASON.OK_NO_BASELINE; - return result; - } + // Bug #3657: the resolved snapshot hash-mismatches the recorded baseline. + // Skip the file with a diagnostic code rather than diffing against the + // wrong baseline (a re-anchor or git-aware baseline step is required). + if (resolution === PRISTINE_RESOLUTION.DRIFTED) { + result.reason = REASON.OK_PRISTINE_DRIFT_DETECTED; + return result; } + // Bug #934 / #4145: a hash was recorded (modern installer) but no + // hash-matching pristine exists anywhere under gsd-pristine/ — + // advisory/non-blocking, the caller logs a warning. + if (resolution === PRISTINE_RESOLUTION.ABSENT_RECORDED) { + result.reason = REASON.OK_NO_BASELINE; + return result; + } + + // VALIDATED / UNVALIDATED keep the gate's pre-#4136 semantics: the resolved + // content (null for OVERBROAD) feeds computeUserAddedLines, whose no-pristine + // branch is the over-broad fallback. + const userAdded = computeUserAddedLines(backupContent, pristineContent); if (userAdded.length === 0) { // Backup and pristine match exactly (or no significant content) — nothing @@ -410,6 +475,101 @@ function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashe return result; } +/** + * Bug #4136: pre-merge classification for reapply-patches Step 4. Runs over + * the SAME fixture state the gate sees (patches dir + freshly installed + * config dir + optional pristine dir) but BEFORE any merge, answering the + * question the gate cannot: is this file's user modification already adopted + * upstream, such that the workflow should leave the installed file untouched + * (status `Incorporated`) instead of re-grafting the diff? + * + * The classification is deliberately conservative in the direction the issue + * demands ("a false Incorporated is worse than no Incorporated"): + * - only a hash-VALIDATED pristine baseline can confirm adoption — drift + * (#3657), absent (#934/#4145), unvalidated (no recorded hash), and a + * missing --pristine-dir all classify UNKNOWN; + * - every significant user-added line (structural/trivial lines excluded by + * isSignificantLine) must be present verbatim in the fresh install — a + * signature-looking short line or a code fence matching anywhere in the + * new file proves nothing; + * - a backup with zero significant delta vs the validated pristine is + * UNKNOWN (the workflow's Critical invariant: a backed-up file is never + * concluded to have "no custom content", and Incorporated is a positive + * adoption finding, not a skip); + * - structural failures are reported with their FAIL_* reason but never + * gate the run — the binding enforcement point stays the post-merge gate. + */ +function classifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashes, skillsRedirect }) { + const backupPath = path.join(patchesDir, relPath); + const installedPath = resolveInstalledPath(configDir, relPath, skillsRedirect); + const result = { file: relPath, classification: CLASSIFICATION.UNKNOWN, reason: null, missing: [] }; + + if (!fs.existsSync(backupPath) || !fs.statSync(backupPath).isFile()) { + return result; // walked entry no longer exists — non-fatal + } + + let installedStat; + try { + installedStat = fs.statSync(installedPath); + } catch { + result.reason = REASON.FAIL_INSTALLED_MISSING; + return result; + } + if (!installedStat.isFile()) { + result.reason = REASON.FAIL_INSTALLED_NOT_REGULAR_FILE; + return result; + } + + let backupContent; + let installedContent; + try { + backupContent = fs.readFileSync(backupPath, 'utf8'); + installedContent = fs.readFileSync(installedPath, 'utf8'); + } catch { + result.reason = REASON.FAIL_READ_ERROR; + return result; + } + + const { resolution, content: pristineContent } = + resolvePristineBaseline({ relPath, pristineDir, pristineHashes }); + + if (resolution === PRISTINE_RESOLUTION.DRIFTED) { + result.reason = REASON.OK_PRISTINE_DRIFT_DETECTED; + return result; + } + if (resolution === PRISTINE_RESOLUTION.ABSENT_RECORDED) { + result.reason = REASON.OK_NO_BASELINE; + return result; + } + if (resolution === PRISTINE_RESOLUTION.UNVALIDATED) { + result.reason = REASON.OK_UNVALIDATED_BASELINE; + return result; + } + if (resolution !== PRISTINE_RESOLUTION.VALIDATED) { + // OVERBROAD: no --pristine-dir (two-way fallback) or a non-file at the + // pristine path. Without a baseline there is no adoption evidence. + return result; + } + + const userAdded = computeUserAddedLines(backupContent, pristineContent); + if (userAdded.length === 0) { + result.reason = REASON.OK_NO_USER_LINES_VS_PRISTINE; + return result; + } + + for (const line of userAdded) { + if (!installedContent.includes(line)) { + result.missing.push(line.trim()); + } + } + if (result.missing.length === 0) { + result.classification = CLASSIFICATION.INCORPORATED; + } else { + result.classification = CLASSIFICATION.NEEDS_MERGE; + } + return result; +} + function main() { const opts = parseArgs(process.argv.slice(2)); if (!opts.patchesDir || !opts.configDir) { @@ -429,6 +589,55 @@ function main() { // #4086: skills-root fallback for runtimes whose skills kind declares a // `home` override outside configDir (codex global → $HOME/.agents/skills). const skillsRedirect = resolveSkillsRedirect(opts.configDir); + + // Bug #4136: --classify is the pre-merge mode the workflow runs in Step 4 + // to decide which files NOT to merge. It is informational — always exit 0 — + // because the binding enforcement stays with the post-merge gate below; + // a needs_merge or unknown outcome is direction for the merge step, not a + // failure, and halting here would block the very merge that resolves it. + if (opts.classify) { + const classifyResults = files.map((relPath) => + classifyFile({ + relPath, + patchesDir: opts.patchesDir, + configDir: opts.configDir, + pristineDir: opts.pristineDir, + pristineHashes, + skillsRedirect, + }), + ); + const incorporatedResults = classifyResults.filter( + (r) => r.classification === CLASSIFICATION.INCORPORATED, + ); + const incorporated = incorporatedResults.length; + const incorporated_files = incorporatedResults.map((r) => r.file); + + if (opts.json) { + process.stdout.write( + JSON.stringify({ checked: classifyResults.length, incorporated, incorporated_files, results: classifyResults }, null, 2) + '\n', + ); + } else { + process.stdout.write(`# Reapply Patch Classification (#4136)\n\n`); + process.stdout.write(`Checked: ${classifyResults.length} file(s)\n`); + process.stdout.write(`Incorporated: ${incorporated} file(s)\n`); + for (const f of incorporated_files) { + process.stdout.write(` incorporated: ${f}\n`); + } + for (const r of classifyResults) { + if (r.classification === CLASSIFICATION.NEEDS_MERGE) { + process.stdout.write(` needs merge: ${r.file}\n`); + for (const line of r.missing.slice(0, 5)) { + process.stdout.write(` missing: ${line}\n`); + } + if (r.missing.length > 5) { + process.stdout.write(` …and ${r.missing.length - 5} more line(s)\n`); + } + } + } + } + return 0; + } + const results = files.map((relPath) => verifyFile({ relPath, @@ -490,4 +699,4 @@ if (require.main === module) { runMain(main); } -module.exports = { computeUserAddedLines, isSignificantLine, verifyFile, walk, REASON, readPristineHashes, sha256, resolveInstalledPath, resolveSkillsRedirect }; +module.exports = { computeUserAddedLines, isSignificantLine, verifyFile, classifyFile, walk, REASON, CLASSIFICATION, PRISTINE_RESOLUTION, resolvePristineBaseline, readPristineHashes, sha256, resolveInstalledPath, resolveSkillsRedirect }; diff --git a/gsd-core/workflows/reapply-patches.md b/gsd-core/workflows/reapply-patches.md index 0b95f333f..a36a94e79 100644 --- a/gsd-core/workflows/reapply-patches.md +++ b/gsd-core/workflows/reapply-patches.md @@ -192,6 +192,58 @@ If neither git history nor pristine snapshots are available, fall back to two-wa ## Step 4: Merge each file +### Step 4 pre-flight: classify superseded customizations (#4136) + +Before merging anything, run the deterministic classifier. It computes, per backed-up file, +whether the user's added lines (diff of the backup against the hash-validated pristine +baseline) are ALREADY present verbatim in the newly installed version — i.e. upstream +adopted the customization. That is the only ground on which the `Incorporated` status +below is valid. Never conclude `Incorporated` from a signature line, a heading, or general +resemblance: the classifier excludes structural/trivial lines and requires every +significant user-added line to be present, because a false `Incorporated` silently retires +a live customization. + +```bash +PRISTINE_DIR="${CONFIG_DIR}/gsd-pristine" + +# Build args as a bash array so paths with spaces survive expansion intact. +CLASSIFY_ARGS=( + --patches-dir "$PATCHES_DIR" + --config-dir "$CONFIG_DIR" +) +if [ -d "$PRISTINE_DIR" ]; then + CLASSIFY_ARGS+=(--pristine-dir "$PRISTINE_DIR") +fi +CLASSIFY_ARGS+=(--classify --json) + +# Informational: exits 0. The binding gate is Step 5a's post-merge verification run. +CLASSIFY_OUTPUT="$(node "${GSD_HOME}/gsd-core/bin/verify-reapply-patches.cjs" "${CLASSIFY_ARGS[@]}")" +INCORPORATED_FILES="$(echo "$CLASSIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));(d.incorporated_files||[]).forEach(f=>process.stdout.write(f+'\n'))")" +``` + +**For every file in `INCORPORATED_FILES`:** + +- **Do NOT merge it. Do NOT re-apply its diff.** Leave the newly installed file exactly as + shipped — the user's customization is already in it. Re-grafting the diff onto a version + that already contains it duplicates the content (both copies run; a stale graft can + contradict a richer native upstream step that replaced it). +- Report the per-file status `Incorporated` — "Already in upstream v{version}" — in the + Step 3 summary and the Step 7 report. +- No Hunk Verification Table rows are required for these files (nothing was merged); the + classifier's structured output is the evidence. If EVERY backed-up file is Incorporated, + still emit the table with a header row and a single note line naming the incorporated + files, so the Step 5b absent-table halt is not tripped. Step 5a passes these files + without special-casing — every user-added line is present in the untouched install. +- This is what ends the re-graft cycle: because the installed file stays identical to what + the release ships, the next update's hash comparison no longer flags it as modified, and + the file drops out of `gsd-local-patches/` on the next cycle. + +Files NOT in `INCORPORATED_FILES` — reported `needs_merge` (with the lines still missing +from the new version) or `unknown` (no confirmable pristine baseline) — take the normal +merge paths below, unchanged. If the classifier cannot run at all (e.g. `GSD_HOME` +unset), fall back to merging every file as before; never report `Incorporated` without the +classifier's structured confirmation. + For each file in `backup-meta.json`: 1. **Read the backed-up version** (user's modified copy from `gsd-local-patches/`) @@ -209,6 +261,9 @@ Compare the three versions to isolate changes: - Sections changed only by upstream → accept upstream version - Sections changed by both → flag as CONFLICT, show both, ask user - Sections unchanged by either → use new version (identical to all three) +- User-added content already present verbatim in the new version → already incorporated + upstream — do NOT re-apply it (the file-level outcome is `Incorporated` when ALL + user-added content is thus present, as determined by the Step 4 pre-flight classifier) ### Two-way merge (fallback when no baseline) @@ -265,7 +320,7 @@ After writing each merged file, verify that user modifications survived the merg 6. **Report status per file:** - `Merged` — user modifications applied cleanly (show summary of what was preserved) - `Conflict` — user reviewed and chose resolution - - `Incorporated` — user's modification was already adopted upstream (only valid when pristine baseline confirms this) + - `Incorporated` — user's modification was already adopted upstream (only valid when pristine baseline confirms this — determined exclusively by the Step 4 pre-flight classifier's `incorporated_files`, never by inspection) **Never report `Skipped — no custom content`.** If a file is in the backup, it has custom content. @@ -439,6 +494,7 @@ Ask user: - [ ] No file classified as "no custom content" or "SKIP" — every backed-up file is definitionally modified - [ ] Three-way merge used when pristine baseline available (git history or gsd-pristine/) - [ ] User modifications identified and merged into new version +- [ ] Superseded customizations classified `Incorporated` by the deterministic pre-flight classifier (pristine-confirmed) and not re-grafted - [ ] Conflicts surfaced to user with both versions shown - [ ] Status reported for each file with summary of what was preserved - [ ] Post-merge verification checks each file for dropped hunks and warns if content appears missing diff --git a/tests/reapply-patches.test.cjs b/tests/reapply-patches.test.cjs index 9f2495d3b..7ca79c6fa 100644 --- a/tests/reapply-patches.test.cjs +++ b/tests/reapply-patches.test.cjs @@ -726,13 +726,17 @@ describe('Bug #2994: reapply-patches workflow references the runtime-installed p }); test('reapply-patches.md references the verifier at gsd-core/bin/verify-reapply-patches.cjs', () => { - const md = fs.readFileSync(REAPPLY_WORKFLOW, 'utf-8'); + const md = fs.readFileSync(REAPPLY_WORKFLOW, 'utf8'); const invocations = extractScriptInvocations(md); const verifierInvocations = invocations.filter(inv => inv.relPath.endsWith('verify-reapply-patches.cjs')); + // #4136: the workflow now invokes the verifier exactly twice — the Step 4 + // pre-flight --classify run (Incorporated detection) and the Step 5a + // post-merge gate run. Both must resolve to the runtime-installed path + // (the sibling test above enforces the path for every invocation). assert.deepEqual( verifierInvocations.map(i => i.relPath), - ['gsd-core/bin/verify-reapply-patches.cjs'], - 'workflow must call the runtime-installed verifier path exactly once', + ['gsd-core/bin/verify-reapply-patches.cjs', 'gsd-core/bin/verify-reapply-patches.cjs'], + 'workflow must call the runtime-installed verifier exactly twice (classify + gate, #4136)', ); }); }); diff --git a/tests/reapply-verify-hunks.test.cjs b/tests/reapply-verify-hunks.test.cjs index 67c719ab1..7aa7ab47c 100644 --- a/tests/reapply-verify-hunks.test.cjs +++ b/tests/reapply-verify-hunks.test.cjs @@ -210,6 +210,8 @@ describe('Bug #2969: deterministic Step 5 verification gate', () => { // this assertion, removing one breaks consumers that switch on the enum. // Bug #3657 added OK_PRISTINE_DRIFT_DETECTED. // Bug #934 added OK_NO_BASELINE. + // Bug #4136 added OK_UNVALIDATED_BASELINE (--classify refuses to confirm + // adoption on an on-disk snapshot with no recorded hash to validate it). assert.deepEqual( Object.keys(REASON).sort(), [ @@ -221,6 +223,7 @@ describe('Bug #2969: deterministic Step 5 verification gate', () => { 'OK_NO_SIGNIFICANT_BACKUP_LINES', 'OK_NO_USER_LINES_VS_PRISTINE', 'OK_PRISTINE_DRIFT_DETECTED', + 'OK_UNVALIDATED_BASELINE', ], ); }); @@ -702,6 +705,7 @@ describe('Bug #3657: pristine-drift does not produce false FAIL_USER_LINES_MISSI * REASON enum shape-lock: the #3657 fix adds OK_PRISTINE_DRIFT_DETECTED. * This assertion locks the updated documented set of stable codes. * Any further additions require updating this assertion. + * Bug #4136 added OK_UNVALIDATED_BASELINE (see the #2969 fold's lock note). */ test('REASON enum includes OK_PRISTINE_DRIFT_DETECTED added by the #3657 fix', () => { assert.deepEqual( @@ -715,6 +719,7 @@ describe('Bug #3657: pristine-drift does not produce false FAIL_USER_LINES_MISSI 'OK_NO_SIGNIFICANT_BACKUP_LINES', 'OK_NO_USER_LINES_VS_PRISTINE', 'OK_PRISTINE_DRIFT_DETECTED', + 'OK_UNVALIDATED_BASELINE', ], ); }); @@ -1172,6 +1177,7 @@ describe('Bug #4086: verifyFile resolves skills entries at the runtime skills ro } + // ──────────────────────────────────────────────────────────────────────── // Folded regression block — #4145 (a hash-matching gsd-pristine/ baseline // stored without the gsd-core/ prefix is never resolved). verifyFile() joined @@ -1439,3 +1445,543 @@ describe('Bug #4145: hash-matching prefix-less pristine baseline is resolved', ( }); }); } + +// ──────────────────────────────────────────────────────────────────────── +// Folded regression block — #4136 (the documented "Incorporated" per-file +// status was unreachable: a customization upstream had adopted was silently +// re-grafted on every future cycle, forever). Adds a --classify pre-merge +// mode to the deterministic verifier: with a hash-validated pristine +// baseline, a file whose EVERY significant user-added line is already +// present verbatim in the freshly installed version is classified +// `incorporated` — the workflow then leaves it untouched (status +// Incorporated, "Already in upstream v{version}") instead of re-grafting. +// ──────────────────────────────────────────────────────────────────────── +{ + const { describe: __foldDescribe } = require('node:test'); + __foldDescribe('folded:bug-4136-reapply-incorporated-status', () => { +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +/** + * Bug #4136: reapply-patches.md Step 4 item 6 defines three per-file statuses + * (Merged / Conflict / Incorporated) but no code path computed Incorporated — + * the term existed only in workflow prose, so superseded customizations were + * re-grafted forever (the merged file's hash never re-converged with the + * shipped manifest hash, so saveLocalPatches re-flagged it every update). + * + * Fix: `--classify` mode on the deterministic verifier. Pre-merge, per file: + * - hash-validated pristine + >=1 significant user-added line + every one of + * those lines present verbatim in the fresh install → incorporated + * - hash-validated pristine + some user lines absent → needs_merge + * - anything else (no/mismatched/absent/unvalidated baseline, zero user + * lines, structural failure) → unknown + * + * Incorporated is NEVER produced without baseline confirmation — the issue's + * law that a false Incorporated is worse than none. Per CONTRIBUTING's typed- + * surface standard, assertions go against the frozen CLASSIFICATION enum and + * the structured --json report; zero text matching on human output. + */ + +const { test, describe, before, after } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const crypto = require('node:crypto'); +const os = require('node:os'); +const path = require('node:path'); +const { cleanup } = require('./helpers.cjs'); +const { runNode } = require('./helpers/process-seam.cjs'); + +const ROOT = path.join(__dirname, '..'); +const SCRIPT = path.join(ROOT, 'gsd-core', 'bin', 'verify-reapply-patches.cjs'); +const { REASON, CLASSIFICATION } = require(SCRIPT); + +let tmpRoot; +let patchesDir; +let configDir; +let pristineDir; + +function sha256(content) { + return crypto.createHash('sha256').update(content, 'utf8').digest('hex'); +} + +function writeFile(absPath, content) { + fs.mkdirSync(path.dirname(absPath), { recursive: true }); + fs.writeFileSync(absPath, content); +} + +function writeBackupMeta(pristine_hashes) { + writeFile(path.join(patchesDir, 'backup-meta.json'), JSON.stringify({ pristine_hashes }, null, 2)); +} + +function resetFixture() { + for (const dir of [patchesDir, configDir, pristineDir]) { + cleanup(dir); + } + fs.mkdirSync(patchesDir); + fs.mkdirSync(configDir); + fs.mkdirSync(pristineDir); +} + +/** Runs the verifier in --classify mode with --json. Returns { status, report }. */ +function runClassifier({ includePristine = true } = {}) { + const args = [ + SCRIPT, + '--patches-dir', patchesDir, + '--config-dir', configDir, + ...(includePristine ? ['--pristine-dir', pristineDir] : []), + '--classify', + '--json', + ]; + const r = runNode(args, { timeoutMs: VERIFIER_TIMEOUT_MS }); + return { + status: r.exitCode, + report: r.stdout && r.stdout.length ? JSON.parse(r.stdout) : null, + }; +} + +/** Runs the verifier in default post-merge gate mode with --json. */ +function runGate({ includePristine = true } = {}) { + const args = [ + SCRIPT, + '--patches-dir', patchesDir, + '--config-dir', configDir, + ...(includePristine ? ['--pristine-dir', pristineDir] : []), + '--json', + ]; + const r = runNode(args, { timeoutMs: VERIFIER_TIMEOUT_MS }); + return { + status: r.exitCode, + report: r.stdout && r.stdout.length ? JSON.parse(r.stdout) : null, + }; +} + +before(() => { + tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4136-')); + patchesDir = path.join(tmpRoot, 'patches'); + configDir = path.join(tmpRoot, 'installed'); + pristineDir = path.join(tmpRoot, 'pristine'); + resetFixture(); +}); + +after(() => { + cleanup(tmpRoot); +}); + +describe('Bug #4136: deterministic Incorporated classification (--classify)', () => { + test('CLASSIFICATION enum exposes the documented set of stable codes', () => { + assert.deepEqual( + Object.keys(CLASSIFICATION).sort(), + ['INCORPORATED', 'NEEDS_MERGE', 'UNKNOWN'], + ); + }); + + test('Row 1 (RED regression): all user-added lines already upstream → incorporated', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/execute-phase.md'; + const pristineContent = [ + 'stock line one that is long enough to be significant', + 'stock line two that is long enough to be significant', + ].join('\n') + '\n'; + const userLine = 'user custom verification gate that upstream adopted verbatim'; + const backupContent = pristineContent + userLine + '\n'; + // Fresh install: upstream shipped the user's line PLUS its own new line. + const freshInstall = pristineContent + userLine + '\n' + + 'brand-new unrelated upstream line shipped in this release\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0, `classify must exit 0; report=${JSON.stringify(report)}`); + assert.equal(report.incorporated, 1); + assert.equal(report.incorporated_files.length, 1); + assert.equal(report.incorporated_files[0].replace(/\\/g, '/'), FILE); + const r0 = report.results[0]; + assert.equal(r0.file.replace(/\\/g, '/'), FILE); + assert.equal(r0.classification, CLASSIFICATION.INCORPORATED); + assert.deepEqual(r0.missing, []); + }); + + test('Row 2: user line absent from fresh install → needs_merge; gate still catches drops', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/plan-phase.md'; + const pristineContent = 'stock baseline line long enough to be significant here\n'; + const userLine = 'user custom instruction that upstream did NOT adopt yet'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + 'unrelated upstream line shipped in the release\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const cls = runClassifier(); + assert.equal(cls.status, 0, 'classify is informational — needs_merge is not an error'); + assert.equal(cls.report.incorporated, 0); + assert.deepEqual(cls.report.incorporated_files, []); + const c0 = cls.report.results[0]; + assert.equal(c0.classification, CLASSIFICATION.NEEDS_MERGE); + assert.ok(c0.missing.includes(userLine), `missing must name the absent line; got ${JSON.stringify(c0.missing)}`); + + // Negative proof: the post-merge gate is NOT weakened — a merge that + // drops the user line still fails the default (#2969) run. + const gate = runGate(); + assert.equal(gate.status, 1); + assert.equal(gate.report.failures, 1); + const g0 = gate.report.results[0]; + assert.equal(g0.status, 'fail'); + assert.equal(g0.reason, REASON.FAIL_USER_LINES_MISSING); + }); + + test('Row 3 (boundary): partially-superseded is needs_merge, not Incorporated', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/ship.md'; + const pristineContent = 'stock line that is long enough to be significant\n'; + const adoptedLine = 'user line number one that upstream did adopt upstream'; + const unadoptedLine = 'user line number two that upstream has NOT adopted'; + const backupContent = pristineContent + adoptedLine + '\n' + unadoptedLine + '\n'; + const freshInstall = pristineContent + adoptedLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0, 'partial adoption must not classify Incorporated'); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.NEEDS_MERGE); + assert.deepEqual(r0.missing, [unadoptedLine], 'only the un-adopted line is missing'); + }); + + test('Row 4a: pristine drift (#3657 shape) → unknown, never Incorporated', () => { + resetFixture(); + const FILE = 'agents/gsd-executor.md'; + const oldPristine = 'old pristine line present when the backup was captured\n'; + const newPristine = 'refreshed upstream snapshot line in the newer GSD release\n'; + const userLine = 'user customisation line that upstream adopted in the release'; + const backupContent = oldPristine + userLine + '\n'; + const freshInstall = newPristine + userLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(oldPristine) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), newPristine); // hash mismatch → drift + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0, 'a drifted baseline confirms nothing'); + assert.deepEqual(report.incorporated_files, []); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + assert.equal(r0.reason, REASON.OK_PRISTINE_DRIFT_DETECTED); + }); + + test('Row 4b: recorded hash but pristine absent (#934 shape) → unknown', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/debug.md'; + const pristineContent = 'stock line that is long enough to be significant x\n'; + const userLine = 'user custom line upstream adopted, but baseline is missing'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + userLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + // No pristine file on disk at all. + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0, 'absent baseline confirms nothing even when all lines are present'); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + assert.equal(r0.reason, REASON.OK_NO_BASELINE); + }); + + test('Row 4c: no --pristine-dir (two-way fallback) → unknown', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/scan.md'; + const pristineContent = 'stock line that is long enough to be significant y\n'; + const userLine = 'user custom line upstream adopted, but no baseline was passed'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + userLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier({ includePristine: false }); + assert.equal(status, 0); + assert.equal(report.incorporated, 0); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + }); + + test('Row 4d: on-disk pristine but no recorded hash → unknown (unvalidated baseline)', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/undo.md'; + const pristineContent = 'stock line that is long enough to be significant z\n'; + const userLine = 'user custom line upstream adopted, but hash was never recorded'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + userLine + '\n'; + + // Older installer: no pristine_hashes entry at all. + writeBackupMeta({}); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0, 'an unvalidated snapshot cannot confirm adoption'); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + assert.equal(r0.reason, REASON.OK_UNVALIDATED_BASELINE); + }); + + test('Row 4d-2: a #4145-recovered (hash-matched orphan) baseline can confirm adoption', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/import.md'; + const pristineContent = 'stock line that is long enough to be significant s\n'; + const userLine = 'user custom line upstream adopted, baseline stored unprefixed'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + userLine + '\n'; + + // The pristine snapshot sits WITHOUT the gsd-core/ prefix (an earlier + // release's writer dropped it) — only the #4145 hash scan can find it. + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, 'workflows', 'import.md'), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0); + // A recovered baseline is hash-confirmed by construction, so it VALIDATES + // and may confirm adoption — the #4136 classifier composes with the + // #4145 recovery instead of treating it as unknown. + assert.equal(report.incorporated, 1); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.INCORPORATED); + assert.deepEqual(r0.missing, []); + }); + + test('Row 4e: signature-looking short/fence lines do not drive the classification', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/health.md'; + const pristineContent = 'stock line that is long enough to be significant w\n'; + // The user hunk: a short heading (under the 12-char significance floor), + // a code fence, and ONE significant line. + const shortHeading = '## My Gate'; + const fence = '```bash'; + const significant = 'the substantive customization body line that matters'; + const backupContent = pristineContent + shortHeading + '\n' + fence + '\n' + significant + '\n'; + // Fresh install happens to contain the short heading (renamed section) + // and plenty of fences — but NOT the significant body line. + const freshInstall = pristineContent + '## My Gate\n' + '```bash\nls -la\n```\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0, 'trivial-line presence must not fabricate adoption'); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.NEEDS_MERGE); + assert.ok(r0.missing.includes(significant)); + assert.ok(!r0.missing.includes(shortHeading), 'insufficiently-significant lines are excluded'); + assert.ok(!r0.missing.includes(fence), 'structural lines are excluded'); + }); + + test('Row 5: backup with zero significant delta vs validated pristine → unknown, never a skip', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/note.md'; + const pristineContent = 'stock line that is long enough to be significant v\n'; + const backupContent = pristineContent; // degenerate: backed up but identical + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), pristineContent); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { status, report } = runClassifier(); + assert.equal(status, 0); + assert.equal(report.incorporated, 0); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + assert.equal(r0.reason, REASON.OK_NO_USER_LINES_VS_PRISTINE); + }); + + test('Row 6: structural failures classify unknown; classify still exits 0', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/stats.md'; + const pristineContent = 'stock line that is long enough to be significant u\n'; + const userLine = 'user custom line that upstream adopted in this release'; + const backupContent = pristineContent + userLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(pristineDir, FILE), pristineContent); + // Installed file deliberately absent. + + const { status, report } = runClassifier(); + assert.equal(status, 0, 'classify is informational; the post-merge gate enforces structure'); + const r0 = report.results[0]; + assert.equal(r0.classification, CLASSIFICATION.UNKNOWN); + assert.equal(r0.reason, REASON.FAIL_INSTALLED_MISSING); + }); + + test('Row 7: classify --json report shape is { checked, incorporated, incorporated_files, results }', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/help.md'; + const pristineContent = 'stock line that is long enough to be significant t\n'; + const userLine = 'user custom line that upstream adopted in this release'; + const backupContent = pristineContent + userLine + '\n'; + const freshInstall = pristineContent + userLine + '\n'; + + writeBackupMeta({ [FILE]: sha256(pristineContent) }); + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(configDir, FILE), freshInstall); + writeFile(path.join(pristineDir, FILE), pristineContent); + + const { report } = runClassifier(); + assert.deepEqual(Object.keys(report).sort(), ['checked', 'incorporated', 'incorporated_files', 'results']); + const r0 = report.results[0]; + assert.deepEqual(Object.keys(r0).sort(), ['classification', 'file', 'missing', 'reason']); + assert.equal(typeof r0.file, 'string'); + assert.equal(typeof r0.classification, 'string'); + assert.ok(Array.isArray(r0.missing)); + }); + + test('Row 8: Incorporated ends the re-graft cycle — the file drops out of the next backup', () => { + resetFixture(); + const FILE = 'gsd-core/workflows/inbox.md'; + const pristineV1 = 'stock v1 line that is long enough to be significant\n'; + const userLine = 'user custom loop guard that upstream adopted in v2'; + const backupContent = pristineV1 + userLine + '\n'; + // v2 ships the user's line verbatim plus an unrelated upstream change. + const freshV2 = pristineV1 + userLine + '\n' + 'unrelated upstream improvement line in v2\n'; + + // Update #1: installer backs up the modified file + records the pristine. + writeFile(path.join(patchesDir, FILE), backupContent); + writeFile(path.join(pristineDir, FILE), pristineV1); + writeBackupMeta({ [FILE]: sha256(pristineV1) }); + // The update wipes and installs v2; manifest now hashes v2's shipped bytes. + writeFile(path.join(configDir, FILE), freshV2); + const manifestV2 = { version: '2.0.0', files: { [FILE]: sha256(freshV2) } }; + writeFile(path.join(configDir, 'gsd-file-manifest.json'), JSON.stringify(manifestV2, null, 2)); + + // Pre-flight classification: incorporated → the workflow leaves the file untouched. + const cls = runClassifier(); + assert.equal(cls.status, 0); + assert.equal(cls.report.results[0].classification, CLASSIFICATION.INCORPORATED); + assert.equal( + fs.readFileSync(path.join(configDir, FILE), 'utf8'), + freshV2, + 'incorporated files are NOT re-grafted — installed bytes stay as shipped', + ); + + // Post-merge gate passes on the untouched install (all user lines present). + const gate = runGate(); + assert.equal(gate.status, 0, `gate must pass the untouched install; report=${JSON.stringify(gate.report)}`); + assert.equal(gate.report.failures, 0); + + // Next update cycle: saveLocalPatches detection (manifest hash comparison) + // no longer flags the file — the forever loop is broken. + const stillModified = Object.entries(manifestV2.files) + .filter(([rel, hash]) => sha256(fs.readFileSync(path.join(configDir, rel), 'utf8')) !== hash) + .map(([rel]) => rel); + assert.deepEqual(stillModified, [], 'an incorporated file must drop out of the backup cycle'); + + // Counter-case (the pre-fix behavior): a re-grafted file WOULD be flagged again. + writeFile(path.join(configDir, FILE), freshV2 + userLine + '\n'); + const reGraftedModified = Object.entries(manifestV2.files) + .filter(([rel, hash]) => sha256(fs.readFileSync(path.join(configDir, rel), 'utf8')) !== hash) + .map(([rel]) => rel); + assert.deepEqual(reGraftedModified, [FILE], 'a re-grafted file stays in the backup cycle forever'); + }); +}); + +describe('Bug #4136: workflow consumes the classifier (contract rows)', () => { + const WORKFLOW = path.join(ROOT, 'gsd-core', 'workflows', 'reapply-patches.md'); + + test('Step 4 runs the classifier before merging and gates it on PRISTINE_DIR', () => { + // allow-test-rule: source-text-is-the-product (#4136) + // reapply-patches.md is the installed runtime workflow — its text IS the + // deployed behavioral contract for --reapply, so structural assertions + // against the shipped text are the correct test form here. + const md = fs.readFileSync(WORKFLOW, 'utf8'); + const classifyIdx = md.indexOf('--classify'); + assert.ok(classifyIdx > 0, 'Step 4 must invoke the verifier with --classify'); + const mergeRulesIdx = md.indexOf('### Three-way merge (when baseline is available)'); + assert.ok(mergeRulesIdx > 0); + assert.ok( + classifyIdx < mergeRulesIdx, + 'the classify invocation must precede the Step 4 merge rules (classify before merging)', + ); + // Both invocations must be the runtime-installed path (locked separately + // by the #2994 fold); here we lock that the classify block is bounded by + // the same GSD_HOME-anchored script reference as the Step 5a gate. + const gateIdx = md.indexOf('verify-reapply-patches.cjs', classifyIdx); + assert.ok(gateIdx > classifyIdx, 'the Step 5a gate invocation must still follow the classify block'); + }); + + test('Incorporated files are instructed NOT to be re-grafted', () => { + // allow-test-rule: source-text-is-the-product (#4136) + const md = fs.readFileSync(WORKFLOW, 'utf8'); + const step4Idx = md.indexOf('## Step 4: Merge each file'); + const step5Idx = md.indexOf('## Step 5: Hunk Verification Gate'); + assert.ok(step4Idx > 0 && step5Idx > step4Idx); + const step4 = md.slice(step4Idx, step5Idx); + assert.ok( + step4.includes('INCORPORATED_FILES'), + 'Step 4 must consume the classifier\'s incorporated_files list', + ); + assert.ok( + /do NOT re-apply/i.test(step4), + 'Step 4 must instruct that incorporated files are not re-grafted', + ); + assert.ok( + step4.includes('Already in upstream'), + 'the documented Step 7 Incorporated phrasing must be wired to the classifier output', + ); + }); + + test('three-way merge rules include the already-present-verbatim rule', () => { + // allow-test-rule: source-text-is-the-product (#4136) + const md = fs.readFileSync(WORKFLOW, 'utf8'); + const rulesIdx = md.indexOf('**Merge rules:**'); + assert.ok(rulesIdx > 0); + const rulesBlock = md.slice(rulesIdx, md.indexOf('### Two-way merge')); + assert.ok( + rulesBlock.includes('already present verbatim'), + 'the merge rule set must cover user content upstream already contains', + ); + }); + + test('success_criteria covers the Incorporated / not-re-grafted contract', () => { + // allow-test-rule: source-text-is-the-product (#4136) + const md = fs.readFileSync(WORKFLOW, 'utf8'); + const start = md.indexOf(''); + const end = md.indexOf(''); + assert.ok(start > 0 && end > start); + const block = md.slice(start, end); + assert.ok( + block.includes('Incorporated'), + 'success_criteria must name the Incorporated disposition', + ); + assert.ok( + /not re-grafted/i.test(block), + 'success_criteria must require that superseded customizations are not re-grafted', + ); + }); +}); + }); +} + From 0aa4202f6a401bc66e2ef80d981bacb92d60dc78 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 05:26:05 -0400 Subject: [PATCH 009/166] fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4135): regression rows for pristine regen coverage collapse RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a null-returning stub (wired into verifyFile after the #4145 orphan tier, behavior-neutral) so the git-history rows fail behaviorally, not at require time. Failing-first rows: baseline_covered aggregate on a 1-of-13 multi-version fixture, coverageHeadline typed renderer, the opt-in --min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass and malformed-value boundaries), git-history baseline recovery (dropped-line catch + surviving-line verify + older-commit hop), findPristineInGit unit, Step 5a workflow headline contract, and the installer-side describeBaselineCoverage honest N-of-M summary with the collapse disk-state pinned. Negative-space rows pin today: non-git ok_no_baseline posture, no-match-no-adoption, #3657 drift never rescued, canonical precedence, and no git tier without --pristine-dir. * fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening The #3407 promotion rule regenerates gsd-pristine/ baselines from the INCOMING release source and keeps only candidates byte-identical with the OUTGOING recorded hash — correct in isolation, but on a multi-version jump the surviving set is precisely the files upstream did NOT change. The verifier then reports ok_no_baseline (advisory, exit 0) for everything else, and no surface distinguishes a 12-of-13-unverified green run from a fully-verified one: the human summary printed Checked/Failures only, the JSON had no coverage aggregate, and the installer's update output gave per-bucket counts without N-of-M framing. All three issue directions, none exclusive: - Report coverage prominently: --json gains an additive baseline_covered aggregate; the human summary leads with 'Baseline coverage: N of M file(s)...' on every run plus an advisory section naming each skipped file and reason; the installer prints an honest covered-of-modified line via the exported describeBaselineCoverage helper (typed return, exact contract); workflow Step 5a computes and prints the headline before any pass/fail framing. - Fail louder on low coverage: opt-in --min-baseline-coverage <0..1> exits with new documented code 3 when coverage falls below the threshold (>= semantics; empty run vacuously passes; content failure exit 1 outranks it; malformed values are usage errors, exit 2). Default posture unchanged — no_baseline stays advisory per #934. - Widen the promotion rule (its only trustworthy form): when no baseline resolves under gsd-pristine/ and a hash is recorded, the verifier now recovers the baseline from the config dir's own git history — the workflow's documented Option A — anchored by the same authority every tier trusts, exact pristine_hashes sha-256 equality. Read-only (git log/git show, windowsHide per #685), bounded (100 commits/file, 10s/subprocess), null-on-any-failure so ok_no_baseline remains the universal fallback. Tier order: canonical join -> #4145 orphan scan -> git history -> OK_NO_BASELINE; #3657 drift and canonical precedence untouched. Hash validation in saveLocalPatches is NOT relaxed — the collapse is legitimate conservatism; hiding it was the bug. Measured on the issue's shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report baseline_covered 1/13 with the headline and can gate at exit 3; a git-managed config dir with the outgoing bytes in history verifies 13/13. Review fixes folded in: workflow headline derives the unverified count from checked - baseline_covered (not the drift+no_baseline sum), and the new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on the marker line. Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it. * chore(#4135): backfill PR number into changeset fragment --------- Co-authored-by: agent-4135 --- .changeset/clever-badgers-gather.md | 5 + bin/install.js | 35 ++ gsd-core/bin/verify-reapply-patches.cjs | 147 ++++++- gsd-core/workflows/reapply-patches.md | 22 +- src/pristine-baseline.cts | 88 ++++ tests/install-write-confinement.test.cjs | 141 +++++++ tests/reapply-verify-hunks.test.cjs | 510 ++++++++++++++++++++++- 7 files changed, 933 insertions(+), 15 deletions(-) create mode 100644 .changeset/clever-badgers-gather.md diff --git a/.changeset/clever-badgers-gather.md b/.changeset/clever-badgers-gather.md new file mode 100644 index 000000000..f5b79edcd --- /dev/null +++ b/.changeset/clever-badgers-gather.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4376 +--- +**The reapply verifier now headlines its baseline coverage instead of reading as fully verified when most files were skipped** — after a multi-version update, /gsd-update --reapply reports 'Baseline coverage: N of M file(s)' in the verifier summary, the reapply output, and the installer's update log; on git-managed config dirs the verifier additionally recovers pristine baselines from history by recorded hash, so files upstream heavily changed are diff-verified instead of skipped; an opt-in --min-baseline-coverage <0..1> flag lets cautious operators fail the gate (exit 3) below a coverage threshold. (#4135) diff --git a/bin/install.js b/bin/install.js index 4951e8221..a578aadcc 100755 --- a/bin/install.js +++ b/bin/install.js @@ -10217,6 +10217,29 @@ function recoverOrphanedPristine(pristineDir, relPath, recordedHash, canonicalSk } } +/** + * #4135: honest N-of-M accounting for gsd-pristine/ baselines after an + * update. The #3407 promotion rule keeps only hash-validated candidates + * (byte-identical across the version span), so a multi-version update + * legitimately ends with near-zero baselines — the collapse itself is NOT a + * bug to hide; hiding it is. This renders the covered-of-total line the + * update output prints either way, so "1 of 13" can never present like a + * fully-covered run. Pure function (typed return) so tests lock the exact + * contract without matching console prose. + */ +function describeBaselineCoverage(totalModified, covered) { + const total = Math.max(0, totalModified); + const have = Math.min(Math.max(0, covered), total); + const uncovered = total - have; + return { + complete: uncovered === 0, + uncovered, + text: uncovered === 0 + ? `gsd-pristine/ baselines cover ${have} of ${total} modified file(s)` + : `gsd-pristine/ baselines cover ${have} of ${total} modified file(s) — ${uncovered} will be reported no_baseline by the reapply verifier`, + }; +} + /** * Detect user-modified GSD files by comparing against install manifest. * Backs up modified files to gsd-local-patches/ for reapply after update. @@ -10467,6 +10490,17 @@ function saveLocalPatches(configDir, pristineCtx) { if (removed > 0) { console.log(' ' + yellow + 'i' + reset + ' Removed ' + removed + ' stale gsd-pristine/ snapshot(s); regenerated ' + regenerated + ' of those — falls back to over-broad verify heuristic for the rest'); } + // #4135: the honest N-of-M coverage line. Preserved/rescued/regenerated + // are disjoint buckets (see their accounting comments above), so their + // sum is exactly the files that ended this update with a hash-valid + // baseline. A partial result renders as an info line, not an error: + // the collapse is legitimate (#3407), hiding it was the bug. + const coverage = describeBaselineCoverage(modified.length, preserved + rescued + regenerated); + if (coverage.complete) { + console.log(' ' + green + '✓' + reset + ' ' + coverage.text); + } else { + console.log(' ' + yellow + 'i' + reset + ' ' + coverage.text); + } } } return modified; @@ -14209,6 +14243,7 @@ module.exports = { reportLocalPatches, validateHookFields, populatePristineDir, + describeBaselineCoverage, _resolveUserArtifactStagingRoot, _tryResolveUserArtifactStagingRoot, finishInstall, diff --git a/gsd-core/bin/verify-reapply-patches.cjs b/gsd-core/bin/verify-reapply-patches.cjs index fedc65039..a1a8e0046 100755 --- a/gsd-core/bin/verify-reapply-patches.cjs +++ b/gsd-core/bin/verify-reapply-patches.cjs @@ -20,11 +20,21 @@ * [--classify] # pre-merge mode: classify each backed-up * # file as incorporated / needs_merge / * # unknown (#4136); always exits 0 + * [--min-baseline-coverage <0..1>] # OPT-IN strict gate (#4135): exit 3 + * # when the fraction of files verified + * # against a resolved pristine baseline + * # (baseline_covered / checked) falls below + * # the threshold. Default: off — a + * # low-coverage run stays green per the + * # documented #934 advisory posture, but + * # its coverage is ALWAYS headline-reported. * * Exit codes (default gate mode): * 0 — every user-added line is present in the merged file (gate passes) - * 1 — at least one missing line in at least one file (gate fails) + * 1 — at least one missing line in at least one file (gate fails); outranks + * a coverage-gate failure when both apply * 2 — usage / structural error (e.g. patches dir missing) + * 3 — opt-in coverage gate failed (--min-baseline-coverage not met; #4135) * * Bug #2969: the Step 5 gate previously trusted Claude's free-text "verified: * yes/no" reporting per hunk. The LLM was filling in `yes` even when content @@ -49,13 +59,13 @@ const { ExitError, runMain } = require('./lib/cli-exit.cjs'); // path under gsd-pristine/ (e.g. without the gsd-core/ prefix an earlier // release's writer dropped). Same module the installer's preserve-check uses, // so the two readers cannot drift apart again. -const { findPristineByHash } = require('./lib/pristine-baseline.cjs'); +const { findPristineByHash, findPristineInGit } = require('./lib/pristine-baseline.cjs'); const SIGNIFICANT_MIN_CHARS = 12; const GSD_HOOK_VERSION_LINE_RE = /^(?:\/\/|#)\s*gsd-hook-version:\s*\S+\s*$/i; function parseArgs(argv) { - const opts = { patchesDir: null, configDir: null, pristineDir: null, json: false, classify: false }; + const opts = { patchesDir: null, configDir: null, pristineDir: null, json: false, classify: false, minBaselineCoverage: null }; for (let i = 0; i < argv.length; i++) { const arg = argv[i]; if (arg === '--patches-dir') opts.patchesDir = argv[++i]; @@ -63,9 +73,19 @@ function parseArgs(argv) { else if (arg === '--pristine-dir') opts.pristineDir = argv[++i]; else if (arg === '--json') opts.json = true; else if (arg === '--classify') opts.classify = true; - else if (arg === '--help' || arg === '-h') { + else if (arg === '--min-baseline-coverage') { + const raw = argv[++i]; + if (raw === undefined) { + throw new ExitError(2, '--min-baseline-coverage requires a value between 0 and 1'); + } + const value = Number(raw); + if (!Number.isFinite(value) || value < 0 || value > 1) { + throw new ExitError(2, `--min-baseline-coverage must be a number between 0 and 1, got: ${raw}`); + } + opts.minBaselineCoverage = value; + } else if (arg === '--help' || arg === '-h') { process.stdout.write( - 'usage: verify-reapply-patches.cjs --patches-dir --config-dir [--pristine-dir ] [--json] [--classify]\n', + 'usage: verify-reapply-patches.cjs --patches-dir --config-dir [--pristine-dir ] [--json] [--classify] [--min-baseline-coverage <0..1>]\n', ); throw new ExitError(0); } else { @@ -253,7 +273,7 @@ const PRISTINE_RESOLUTION = Object.freeze({ * guard. Extracted from verifyFile's inline block so --classify reasons over * the exact same baseline semantics the post-merge gate enforces. */ -function resolvePristineBaseline({ relPath, pristineDir, pristineHashes }) { +function resolvePristineBaseline({ relPath, configDir, pristineDir, pristineHashes }) { const hashKey = relPath.replace(/\\/g, '/'); const recordedHash = pristineHashes && pristineHashes[hashKey]; if (pristineDir) { @@ -314,6 +334,27 @@ function resolvePristineBaseline({ relPath, pristineDir, pristineHashes }) { // Present but not a regular file — over-broad mode is the safe side. return { resolution: PRISTINE_RESOLUTION.OVERBROAD, content: null }; } + + // Bug #4135: still nothing under gsd-pristine/ — the multi-version + // regeneration collapse (the #3407 promotion rule only keeps files + // byte-identical across the WHOLE version span, so the surviving set is + // precisely the files upstream did not change). When the config dir is + // itself a git repository, its history may hold the outgoing bytes: this + // is the workflow's documented Option A (reapply-patches.md Step 2), + // anchored by the SAME authority every other tier trusts — exact + // pristine_hashes equality. Read-only (git log / git show); any failure + // (no git, not a repo, no matching blob) degrades to ABSENT_RECORDED. + // A recovered baseline is hash-confirmed by construction, so it VALIDATES. + if (!pristinePathExists && recordedHash) { + try { + const fromGit = findPristineInGit(configDir, hashKey, recordedHash); + if (fromGit !== null) { + return { resolution: PRISTINE_RESOLUTION.VALIDATED, content: fromGit }; + } + } catch { + // git unavailable or history walk failed — ABSENT_RECORDED posture + } + } // Bug #934: recordedHash is present (modern installer) but no // hash-matching pristine exists anywhere under gsd-pristine/ (the stat // missed and the #4145 recovery found nothing). @@ -387,6 +428,47 @@ function resolveSkillsRedirect(configDir) { } } +/** + * #4135: baseline-coverage accounting. A file counts as baseline-covered + * when its diff was actually computed against a resolved pristine baseline, + * i.e. its reason is one of the baseline-diff outcomes (null = lines + * verified present, OK_NO_USER_LINES_VS_PRISTINE = no user lines versus + * the baseline, FAIL_USER_LINES_MISSING = lines missing). Everything else + * is uncovered: the silent skips (no baseline anywhere per #934/#4135, + * drift per #3657, no-pristine over-broad with nothing significant) AND + * the structural failures (installed missing / not a file / read error), + * which never reached a baseline — they are loud blocking failures, but + * they were not baseline-verified either, and this aggregate exists to + * state exactly how much of the run the gate could reason about. When + * --pristine-dir is absent the whole run is over-broad (#2998 fallback): + * no file was verified against a baseline, covered is 0 by construction. + */ +const BASELINE_VERIFIED_REASONS = new Set([ + REASON.OK_NO_USER_LINES_VS_PRISTINE, + REASON.FAIL_USER_LINES_MISSING, +]); + +function classifyBaselineCoverage(results, pristineDirProvided) { + if (!pristineDirProvided) return 0; + let covered = 0; + for (const r of results) { + if (r.reason === null || BASELINE_VERIFIED_REASONS.has(r.reason)) covered++; + } + return covered; +} + +/** + * #4135: the headline string for the human summary. A run with 12 of 13 + * files unbaselined must not present the same way as a fully-verified one — + * the N-of-M form is the issue's own framing, and the unverified tail is + * only rendered when it is non-zero. + */ +function coverageHeadline(covered, total) { + const unverified = Math.max(0, total - covered); + const base = `Baseline coverage: ${covered} of ${total} file(s) verified against a pristine baseline`; + return unverified > 0 ? `${base} (${unverified} unverified)` : base; +} + function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashes, skillsRedirect }) { const backupPath = path.join(patchesDir, relPath); const installedPath = resolveInstalledPath(configDir, relPath, skillsRedirect); @@ -431,7 +513,7 @@ function verifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHashe // hash-first recovery + #934 absent guard) moved into resolvePristineBaseline // so the pre-merge classifier reasons over the exact same semantics. const { resolution, content: pristineContent } = - resolvePristineBaseline({ relPath, pristineDir, pristineHashes }); + resolvePristineBaseline({ relPath, configDir, pristineDir, pristineHashes }); // Bug #3657: the resolved snapshot hash-mismatches the recorded baseline. // Skip the file with a diagnostic code rather than diffing against the @@ -531,7 +613,7 @@ function classifyFile({ relPath, patchesDir, configDir, pristineDir, pristineHas } const { resolution, content: pristineContent } = - resolvePristineBaseline({ relPath, pristineDir, pristineHashes }); + resolvePristineBaseline({ relPath, configDir, pristineDir, pristineHashes }); if (resolution === PRISTINE_RESOLUTION.DRIFTED) { result.reason = REASON.OK_PRISTINE_DRIFT_DETECTED; @@ -669,14 +751,52 @@ function main() { const no_baseline = noBaselineResults.length; const no_baseline_files = noBaselineResults.map((r) => r.file); + // Bug #4135: baseline coverage is a first-class aggregate. The collapse + // this reports is silent by design in every other surface (no_baseline is + // advisory, exit stays 0), which is exactly why it must be counted here: + // "A clean verifier exit reads as 'hunks survived'. On this run it meant + // '12 of 13 files were not checked at all.'" + const baseline_covered = classifyBaselineCoverage(results, Boolean(opts.pristineDir)); + + // Bug #4135: the opt-in strict coverage gate. Threshold semantics are >= + // (a run exactly at the threshold passes); an empty run is vacuously + // covered (nothing checked cannot be under-covered). A content failure + // (exit 1) outranks a coverage failure — it is the louder, more specific + // signal. + let coverageGateFailed = false; + if (opts.minBaselineCoverage !== null) { + const ratio = results.length === 0 ? 1 : baseline_covered / results.length; + coverageGateFailed = ratio < opts.minBaselineCoverage; + } + if (opts.json) { process.stdout.write( - JSON.stringify({ checked: results.length, failures: failures.length, drifted, drifted_files, no_baseline, no_baseline_files, results }, null, 2) + '\n', + JSON.stringify({ checked: results.length, failures: failures.length, drifted, drifted_files, no_baseline, no_baseline_files, baseline_covered, results }, null, 2) + '\n', ); } else { process.stdout.write(`# Hunk Verification Gate (#2969)\n\n`); process.stdout.write(`Checked: ${results.length} file(s)\n`); - process.stdout.write(`Failures: ${failures.length}\n\n`); + process.stdout.write(`Failures: ${failures.length}\n`); + // #4135: the coverage headline is printed on EVERY run, before any + // per-file detail, so a near-zero-coverage green run can never render + // identically to a fully-verified one. + process.stdout.write(`${coverageHeadline(baseline_covered, results.length)}\n\n`); + if (baseline_covered < results.length) { + if (!opts.pristineDir) { + process.stdout.write(`No --pristine-dir: over-broad fallback — nothing was verified against a pristine baseline (#2998).\n\n`); + } else { + const skipped = results.filter((r) => r.reason !== null && !BASELINE_VERIFIED_REASONS.has(r.reason)); + process.stdout.write(`## Files not verified against a pristine baseline\n\n`); + process.stdout.write(`Advisory (non-blocking): their user customizations may or may not have survived the merge.\n\n`); + for (const r of skipped) { + process.stdout.write(`- ${r.file} (${r.reason})\n`); + } + process.stdout.write('\n'); + } + } + if (coverageGateFailed) { + process.stdout.write(`COVERAGE GATE FAILED: baseline coverage ${(opts.minBaselineCoverage * 100).toFixed(1)}% required, ${results.length === 0 ? 100 : ((baseline_covered / results.length) * 100).toFixed(1)}% achieved (#4135 strict mode).\n\n`); + } if (failures.length > 0) { process.stdout.write(`## Files with missing user-added content\n\n`); for (const r of failures) { @@ -692,11 +812,14 @@ function main() { } } - return failures.length > 0 ? 1 : 0; + if (failures.length > 0) return 1; + if (coverageGateFailed) return 3; + return 0; } if (require.main === module) { runMain(main); } -module.exports = { computeUserAddedLines, isSignificantLine, verifyFile, classifyFile, walk, REASON, CLASSIFICATION, PRISTINE_RESOLUTION, resolvePristineBaseline, readPristineHashes, sha256, resolveInstalledPath, resolveSkillsRedirect }; +module.exports = { computeUserAddedLines, isSignificantLine, verifyFile, classifyFile, walk, REASON, CLASSIFICATION, PRISTINE_RESOLUTION, resolvePristineBaseline, readPristineHashes, sha256, resolveInstalledPath, resolveSkillsRedirect, coverageHeadline, classifyBaselineCoverage }; + diff --git a/gsd-core/workflows/reapply-patches.md b/gsd-core/workflows/reapply-patches.md index a36a94e79..70692e5b8 100644 --- a/gsd-core/workflows/reapply-patches.md +++ b/gsd-core/workflows/reapply-patches.md @@ -169,7 +169,7 @@ Check if a `gsd-pristine/` directory exists alongside `gsd-local-patches/`: ```bash PRISTINE_DIR="$CONFIG_DIR/gsd-pristine" ``` -If it exists, the installer saved pristine copies at install time. Use these as the baseline. Both the deterministic verifier and the installer's preserve-check resolve each file's snapshot at its canonical path first and, when that misses, by the SHA-256 recorded in `pristine_hashes` — so a snapshot stored at a legacy path (for example, without the `gsd-core/` prefix an earlier release dropped) is still found and, on the next update, relocated to its canonical path (#4145). +If it exists, the installer saved pristine copies at install time. Use these as the baseline. Both the deterministic verifier and the installer's preserve-check resolve each file's snapshot at its canonical path first and, when that misses, by the SHA-256 recorded in `pristine_hashes` — so a snapshot stored at a legacy path (for example, without the `gsd-core/` prefix an earlier release dropped) is still found and, on the next update, relocated to its canonical path (#4145). When no snapshot resolves under `gsd-pristine/`, the deterministic verifier additionally falls back to Option A's git-history walk itself (read-only, same recorded-hash match), which is what keeps Step 5a coverage non-zero on multi-version updates where the hash-validated regeneration had nothing it could promote (#4135). ### Option C: No baseline available (two-way fallback) If neither git history nor pristine snapshots are available, fall back to two-way comparison — but with **strengthened heuristics** (see Step 3). @@ -362,9 +362,27 @@ VERIFY_STATUS=$? DRIFTED_COUNT="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));process.stdout.write(String(d.drifted||0))")" DRIFTED_FILES="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));(d.drifted_files||[]).forEach(f=>process.stdout.write(f+'\n'))")" NO_BASELINE_COUNT="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));process.stdout.write(String(d.no_baseline||0))")" -NO_BASELINE_FILES="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));(d.no_baseline_files||[]).forEach(f=>process.stdout.write(f+'\n'))")" +NO_BASELINE_FILES="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));(d.no_baseline_files||[]).forEach(f=>process.stdout.write(f+'\n'))")" +BASELINE_COVERED="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));process.stdout.write(String(d.baseline_covered||0))")" +CHECKED_COUNT="$(echo "$VERIFY_OUTPUT" | node -e "const d=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8'));process.stdout.write(String(d.checked||0))")" ``` +**Baseline coverage headline (#4135)** — BEFORE any pass/fail statement, print the coverage +N-of-M line. A run where 12 of 13 files were not diff-verified must never present the same way +as a fully-verified one: + +```text +Baseline coverage: {BASELINE_COVERED} of {CHECKED_COUNT} file(s) verified against a pristine +baseline ({CHECKED_COUNT - BASELINE_COVERED} unverified) +``` + +The verifier also reports this on every run (JSON field `baseline_covered`; the human summary +leads with the same line). Operators who want the gate itself to fail on low coverage (cautious +environments) pass the opt-in strict flag when invoking the verifier: +`--min-baseline-coverage ` — below the threshold the verifier exits +with code 3 (a real content failure still exits 1 and outranks it). Without the flag the +default advisory posture is unchanged. + **If `NO_BASELINE_COUNT` is greater than 0**, emit an advisory warning (non-blocking — the gate still exits 0 for these files). Do NOT halt: ```text diff --git a/src/pristine-baseline.cts b/src/pristine-baseline.cts index 1bc5c40e6..4741c95d0 100644 --- a/src/pristine-baseline.cts +++ b/src/pristine-baseline.cts @@ -24,6 +24,7 @@ import fs from 'node:fs'; import path from 'node:path'; import crypto from 'node:crypto'; +import { execFileSync } from 'node:child_process'; /** * SHA-256 hex digest of a file's raw bytes. Byte-for-byte the same digest @@ -92,3 +93,90 @@ export function findPristineByHash( } return null; } + +/** + * #4135: recover a pristine baseline from the config dir's OWN git history, + * anchored by the recorded pristine_hashes entry. + * + * The #3407 promotion rule keeps only regeneration candidates byte-identical + * across the whole version span, so a multi-version update leaves + * gsd-pristine/ holding exactly the files upstream did NOT change — near-zero + * coverage precisely where upstream churned the most. On a git-managed config + * dir the outgoing bytes often still exist in history (the workflow's + * documented Option A), and pristine_hashes is the same authority every other + * resolution tier trusts: a blob whose SHA-256 equals the recorded hash cannot + * be the wrong baseline. This is read-only recovery (git log / git show only). + * + * Guarantees: + * - Only an EXACT sha-256 match with the recorded hash is ever returned. + * - Newest-first commit order (git log default) makes multi-match resolution + * deterministic; byte-identical matches are interchangeable anyway. + * - Any failure (git absent, not a repository, empty history, unreadable + * blob, subprocess timeout) yields null — never a throw — so the caller's + * OK_NO_BASELINE posture is the universal fallback. + * - The walk is bounded: at most GIT_MAX_COMMITS_PER_FILE commits per file. + */ +const GIT_MAX_COMMITS_PER_FILE = 100; +/** Per-subprocess bound in ms — an unbounded git call is an indefinite hang. */ +const GIT_SUBPROCESS_TIMEOUT_MS = 10_000; +/** git log --format=%H output cap; 100 full shas are ~4 KB, this is headroom. */ +const GIT_MAX_BUFFER_BYTES = 16 * 1024 * 1024; + +function isCleanRelativePosixPath(relPath: string): boolean { + if (!relPath || relPath.startsWith('/') || relPath.includes('\\') || relPath.includes('\0')) { + return false; + } + const segments = relPath.split('/'); + return segments.every((seg) => seg.length > 0 && seg !== '.' && seg !== '..'); +} + +function gitExec(gitDir: string, args: string[]): string { + return execFileSync('git', args, { + cwd: gitDir, + encoding: 'utf8', + timeout: GIT_SUBPROCESS_TIMEOUT_MS, + maxBuffer: GIT_MAX_BUFFER_BYTES, + // windowsHide (#685): a console-window flash per git call would spam the + // user on Windows for what is a background, read-only history walk. + windowsHide: true, + // stderr is discarded: "file absent in commit" is an expected walk outcome, + // not operator-visible diagnostics. + stdio: ['ignore', 'pipe', 'ignore'], + }); +} + +export function findPristineInGit( + gitDir: string, + relPath: string, + recordedHash: string, +): string | null { + if (!gitDir || typeof relPath !== 'string' || typeof recordedHash !== 'string' + || recordedHash.length === 0 || !isCleanRelativePosixPath(relPath)) { + return null; + } + let commits: string[]; + try { + const logOutput = gitExec(gitDir, ['log', '--format=%H', '--', relPath]).trim(); + if (!logOutput) return null; + commits = logOutput.split('\n').slice(0, GIT_MAX_COMMITS_PER_FILE); + } catch { + return null; // git absent, not a repository, or the walk failed + } + for (const commit of commits) { + if (!/^[0-9a-f]{40}$/i.test(commit)) continue; + try { + const blob = gitExec(gitDir, ['show', `${commit}:${relPath}`]); + if (sha256String(blob) === recordedHash) { + return blob; + } + } catch { + // blob absent in this commit (rename/add boundary) — keep walking + } + } + return null; +} + +/** sha256 of a utf8 string, matching how manifest hashes are recorded. */ +function sha256String(content: string): string { + return crypto.createHash('sha256').update(content, 'utf8').digest('hex'); +} diff --git a/tests/install-write-confinement.test.cjs b/tests/install-write-confinement.test.cjs index b4ceaece6..138f7d97f 100644 --- a/tests/install-write-confinement.test.cjs +++ b/tests/install-write-confinement.test.cjs @@ -4237,3 +4237,144 @@ describe('Bug #4145: saveLocalPatches rescues hash-matching orphaned pristine sn }); }); } + + +// ──────────────────────────────────────────────────────────────────────── +// Folded regression block — #4135 (installer side). saveLocalPatches' +// hash-validated regeneration (the #3407 promotion rule) keeps only +// candidates byte-identical across the whole version span, so a +// multi-version update leaves gsd-pristine/ holding near-zero baselines — +// and the update output never says so in N-of-M terms. The fix keeps the +// hash validation untouched (disk behavior is pinned here) and adds an +// honest coverage summary via an exported typed helper. +// ──────────────────────────────────────────────────────────────────────── +{ + const { describe: __foldDescribe } = require('node:test'); + __foldDescribe('folded:bug-4135-saveLocalPatches-coverage-line', () => { +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test, describe, beforeEach } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const os = require('node:os'); +const crypto = require('node:crypto'); + +const ROOT = path.join(__dirname, '..'); +const INSTALL = require(path.join(ROOT, 'bin', 'install.js')); +const { cleanup } = require('./helpers.cjs'); + +const MANIFEST_NAME = 'gsd-file-manifest.json'; + +function sha256(content) { + return crypto.createHash('sha256').update(content, 'utf8').digest('hex'); +} + +function countFiles(dir) { + let n = 0; + if (!fs.existsSync(dir)) return 0; + for (const entry of fs.readdirSync(dir, { withFileTypes: true })) { + if (entry.isDirectory()) n += countFiles(path.join(dir, entry.name)); + else if (entry.isFile()) n += 1; + } + return n; +} + +describe('Bug #4135: saveLocalPatches reports honest gsd-pristine coverage on multi-version updates', () => { + let tmpDir; + let configDir; + let newSrcDir; + let pristineDir; + + beforeEach((t) => { + tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4135-cov-')); + configDir = path.join(tmpDir, 'config'); + newSrcDir = path.join(tmpDir, 'new-release-src'); + pristineDir = path.join(configDir, 'gsd-pristine'); + fs.mkdirSync(configDir, { recursive: true }); + fs.mkdirSync(newSrcDir, { recursive: true }); + t.after(() => { + cleanup(tmpDir); + }); + }); + + /** + * Seeds a multi-version-update fixture: `total` modified files, of which + * `identical` are byte-identical across the span (the only regenerable + * baselines) and the rest changed upstream in the incoming source. + */ + function seedMultiVersionFixture(total, identical) { + const manifestFiles = {}; + for (let i = 1; i <= total; i++) { + const rel = `gsd-core/workflows/flow-${String(i).padStart(2, '0')}.md`; + const pristine = + `# Flow ${i}\nStock content of the outgoing release for file ${i}.\n` + + `Second outgoing stock line ${i} with plenty of substance.\n`; + const user = pristine + `## User customisation ${i}\nCustom section on top of the outgoing release.\n`; + const incoming = (i <= identical) + ? pristine + : `# Flow ${i} (rewritten)\nIncoming release rewrote file ${i} across the span.\n`; + manifestFiles[rel] = sha256(pristine); + fs.mkdirSync(path.dirname(path.join(configDir, rel)), { recursive: true }); + fs.writeFileSync(path.join(configDir, rel), user); + fs.mkdirSync(path.dirname(path.join(newSrcDir, rel)), { recursive: true }); + fs.writeFileSync(path.join(newSrcDir, rel), incoming); + } + fs.writeFileSync( + path.join(configDir, MANIFEST_NAME), + JSON.stringify({ version: '1.10.0', timestamp: '2026-08-01T00:00:00Z', runtime: 'claude', scope: 'global', files: manifestFiles }, null, 2), + ); + return total; + } + + /** + * Core installer regression: the multi-version collapse itself is pinned + * (hash validation untouched — only the byte-identical file survives), + * and the honest N-of-M summary is available via the typed helper. + */ + test('#4135: saveLocalPatches multi-version regen keeps hash validation and reports 1-of-13 coverage', () => { + const total = seedMultiVersionFixture(13, 1); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + const covered = countFiles(pristineDir); + assert.equal(covered, 1, + 'the collapse is pinned: only the byte-identical file survives hash-validated regeneration'); + assert.equal(typeof INSTALL.describeBaselineCoverage, 'function', + 'the coverage summary must be rendered by an exported typed helper'); + const summary = INSTALL.describeBaselineCoverage(total, covered); + assert.equal(summary.complete, false); + assert.equal(summary.uncovered, 12); + assert.equal( + summary.text, + 'gsd-pristine/ baselines cover 1 of 13 modified file(s) — 12 will be reported no_baseline by the reapply verifier', + 'partial coverage states the N-of-M collapse and its downstream effect', + ); + }); + + /** Positive boundary: every modified file covered renders a complete summary. */ + test('#4135: describeBaselineCoverage reports complete when every modified file is covered', () => { + const total = seedMultiVersionFixture(3, 3); + + INSTALL.saveLocalPatches(configDir, { + packageSrc: newSrcDir, runtime: 'claude', pathPrefix: '$HOME/.claude/', isGlobal: true, + }); + + const covered = countFiles(pristineDir); + assert.equal(covered, 3, 'a fully byte-identical span regenerates every baseline'); + const summary = INSTALL.describeBaselineCoverage(total, covered); + assert.equal(summary.complete, true); + assert.equal(summary.uncovered, 0); + assert.equal( + summary.text, + 'gsd-pristine/ baselines cover 3 of 3 modified file(s)', + 'complete coverage renders without a collapse tail', + ); + }); +}); + }); +} diff --git a/tests/reapply-verify-hunks.test.cjs b/tests/reapply-verify-hunks.test.cjs index 7aa7ab47c..3cf99b827 100644 --- a/tests/reapply-verify-hunks.test.cjs +++ b/tests/reapply-verify-hunks.test.cjs @@ -305,7 +305,8 @@ describe('Bug #2969: deterministic Step 5 verification gate', () => { // Bug #3657 (Finding 1): drifted + drifted_files are additive fields added to surface // pristine-drift skips distinctly from failures. Shape-lock updated to include them. // Bug #934: no_baseline + no_baseline_files are additive fields for missing-pristine advisory. - assert.deepEqual(Object.keys(report).sort(), ['checked', 'drifted', 'drifted_files', 'failures', 'no_baseline', 'no_baseline_files', 'results']); + // Bug #4135: baseline_covered is the additive coverage aggregate (headline reporting). + assert.deepEqual(Object.keys(report).sort(), ['baseline_covered', 'checked', 'drifted', 'drifted_files', 'failures', 'no_baseline', 'no_baseline_files', 'results']); const r0 = report.results[0]; assert.deepEqual(Object.keys(r0).sort(), ['file', 'missing', 'reason', 'status']); assert.equal(typeof r0.file, 'string'); @@ -1985,3 +1986,510 @@ describe('Bug #4136: workflow consumes the classifier (contract rows)', () => { }); } + +// ──────────────────────────────────────────────────────────────────────── +// Folded regression block — #4135 (gsd-pristine/ regeneration yields +// near-zero coverage on a multi-version update). The #3407 promotion rule +// only keeps regeneration candidates byte-identical across the WHOLE +// version span, so the surviving baseline set is precisely the files +// upstream did NOT change — and no reporting surface (verifier summary, +// --json, workflow Step 5a) distinguishes a 12-of-13 unbaselined green run +// from a fully-verified one. The fix reports baseline coverage as a +// headline, adds an opt-in strict coverage gate, and widens resolution +// with a git-history tier anchored by the same pristine_hashes authority. +// ──────────────────────────────────────────────────────────────────────── +{ + const { describe: __foldDescribe } = require('node:test'); + __foldDescribe('folded:bug-4135-pristine-regen-coverage', () => { +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test, describe, before, after } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const os = require('node:os'); +const path = require('node:path'); +const crypto = require('node:crypto'); +const { cleanup } = require('./helpers.cjs'); +const { runNode } = require('./helpers/process-seam.cjs'); +const { gitOrThrow } = require('./helpers/git-fixture.cjs'); + +const ROOT = path.join(__dirname, '..'); +const SCRIPT = path.join(ROOT, 'gsd-core', 'bin', 'verify-reapply-patches.cjs'); +const { REASON } = require(SCRIPT); +const { findPristineInGit } = require( + path.join(ROOT, 'gsd-core', 'bin', 'lib', 'pristine-baseline.cjs'), +); +const WORKFLOW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'reapply-patches.md'); + +// 30000ms: same class as the folded blocks above — one deterministic +// verifier pass (plus, for the git tier rows, the in-process git log/show +// walk the verifier performs) over a small mkdtemp fixture tree. +const VERIFIER_TIMEOUT_MS = 30_000; + +let tmpRoot; +let patchesDir; +let configDir; +let pristineDir; + +function sha256(content) { + return crypto.createHash('sha256').update(content, 'utf8').digest('hex'); +} + +function writeFile(absPath, content) { + fs.mkdirSync(path.dirname(absPath), { recursive: true }); + fs.writeFileSync(absPath, content); +} + +function writeBackupMeta(pristine_hashes) { + writeFile(path.join(patchesDir, 'backup-meta.json'), JSON.stringify({ pristine_hashes }, null, 2)); +} + +function resetFixture() { + for (const dir of [patchesDir, configDir, pristineDir]) { + cleanup(dir); + } + fs.mkdirSync(patchesDir); + fs.mkdirSync(configDir); + fs.mkdirSync(pristineDir); +} + +/** Runs the verifier with --json plus any extra argv (strict-gate flags). */ +function runVerifier(extraArgs = []) { + const r = runNode([ + SCRIPT, + '--patches-dir', patchesDir, + '--config-dir', configDir, + '--pristine-dir', pristineDir, + '--json', + ...extraArgs, + ], { timeoutMs: VERIFIER_TIMEOUT_MS }); + return { + status: r.exitCode, + report: r.stdout && r.stdout.length ? JSON.parse(r.stdout) : null, + }; +} + +/** git fixture plumbing that must abort loudly when setup breaks. */ +function gitIn(dir, args) { + gitOrThrow(args, { cwd: dir }); +} +function gitCommitAll(dir, message) { + gitIn(dir, ['add', '-A']); + gitIn(dir, ['-c', 'user.email=t@t', '-c', 'user.name=t', 'commit', '-qm', message]); +} + +before(() => { + tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4135-')); + patchesDir = path.join(tmpRoot, 'patches'); + configDir = path.join(tmpRoot, 'installed'); + pristineDir = path.join(tmpRoot, 'pristine'); + resetFixture(); +}); + +after(() => { + cleanup(tmpRoot); +}); + +describe('Bug #4135: baseline coverage is reported, gateable, and widened from git history', () => { + /** + * Core regression (headline): the issue's exact scenario — 13 backed-up + * customised files on a multi-version update, 12 changed upstream, 1 + * byte-identical. gsd-pristine/ holds only the byte-identical survivor. + * The green exit must CARRY a countable coverage aggregate instead of + * presenting like a fully-verified run. + */ + test('#4135: report aggregates baseline_covered so a 12-of-13 unbaselined run is countable', () => { + resetFixture(); + const pristineHashes = {}; + for (let i = 1; i <= 13; i++) { + const rel = `gsd-core/workflows/flow-${String(i).padStart(2, '0')}.md`; + const pristine = + `# Flow ${i}\nStock content of the outgoing release for file ${i}.\n` + + `Line two of outgoing stock content ${i} with plenty of substance.\n`; + const userLine = `## User customisation ${i}\nA custom section the user added for file ${i}.\n`; + const upstream = + `# Flow ${i} (rewritten)\nUpstream rewrote file ${i} across the multi-version span.\n` + + `New stock structure with several new lines for file ${i}.\n`; + pristineHashes[rel] = sha256(pristine); + writeFile(path.join(patchesDir, rel), pristine + userLine); + writeFile(path.join(configDir, rel), upstream + userLine); + if (i === 7) { + // The single byte-identical-across-the-span survivor. + writeFile(path.join(pristineDir, rel), pristine); + } + } + writeBackupMeta(pristineHashes); + + const { status, report } = runVerifier(); + + assert.equal(status, 0, 'no-baseline files stay advisory — the default gate must stay green'); + assert.equal(report.checked, 13); + assert.equal(report.no_baseline, 12); + assert.equal(report.failures, 0); + assert.equal(report.baseline_covered, 1, + `coverage collapse must be countable; got ${JSON.stringify(report.baseline_covered)}`); + }); + + /** Human-mode headline renderer: exact-contract unit on the typed helper. */ + test('#4135: human summary headlines baseline coverage N-of-M', () => { + const script = require(SCRIPT); + assert.equal(typeof script.coverageHeadline, 'function', + 'the human summary headline must be rendered by an exported typed helper'); + assert.equal( + script.coverageHeadline(1, 13), + 'Baseline coverage: 1 of 13 file(s) verified against a pristine baseline (12 unverified)', + 'partial coverage renders N-of-M plus the unverified count'); + assert.equal( + script.coverageHeadline(13, 13), + 'Baseline coverage: 13 of 13 file(s) verified against a pristine baseline', + 'full coverage renders without an unverified tail'); + }); + + /** + * Core regression (strict gate): the same collapsed fixture under + * `--min-baseline-coverage 0.9` must FAIL LOUDLY (new opt-in exit code 3) + * while still emitting the parseable JSON report. + */ + test('#4135: opt-in --min-baseline-coverage exits 3 below threshold', () => { + resetFixture(); + const pristineHashes = {}; + for (let i = 1; i <= 13; i++) { + const rel = `gsd-core/workflows/flow-${String(i).padStart(2, '0')}.md`; + const pristine = `# Flow ${i}\nOutgoing stock line with substantial content ${i}.\n`; + const userLine = `User customisation line that survived the merge for file ${i}.\n`; + const upstream = `# Flow ${i} new\nIncoming release rewrote this file upstream ${i}.\n`; + pristineHashes[rel] = sha256(pristine); + writeFile(path.join(patchesDir, rel), pristine + userLine); + writeFile(path.join(configDir, rel), upstream + userLine); + } + writeBackupMeta(pristineHashes); + + const { status, report } = runVerifier(['--min-baseline-coverage', '0.9']); + + assert.equal(status, 3, 'a 1-of-13 run under a 0.9 threshold must exit non-zero (coverage gate)'); + assert.ok(report, 'the JSON report must still be emitted for scripting consumers'); + assert.equal(report.failures, 0, 'the failure here is coverage, not content'); + assert.equal(report.baseline_covered, 0); + }); + + /** Precedence: a real content failure outranks the coverage failure. */ + test('#4135: strict gate yields to exit 1 when real content failed', () => { + resetFixture(); + const rel = 'gsd-core/workflows/one.md'; + const pristine = 'outgoing stock line with substantial content here\n'; + const userLine = 'user customisation line that was dropped by the merge\n'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + userLine); + writeFile(path.join(configDir, rel), pristine); // user line dropped + writeFile(path.join(pristineDir, rel), pristine); + + const { status, report } = runVerifier(['--min-baseline-coverage', '1']); + + assert.equal(status, 1, 'content failure is the louder, more specific signal'); + assert.equal(report.failures, 1); + assert.equal(report.results[0].reason, REASON.FAIL_USER_LINES_MISSING); + }); + + /** Ratio boundary: at exactly the threshold the gate passes (>= semantics). */ + test('#4135: strict coverage gate passes at exactly the threshold', () => { + resetFixture(); + seedTwoFilesOneCovered(); + const { status } = runVerifier(['--min-baseline-coverage', '0.5']); + assert.equal(status, 0, '1 of 2 covered satisfies a 0.5 threshold (>= passes)'); + }); + + /** Ratio boundary: just below the threshold the gate fails. */ + test('#4135: strict coverage gate fails just below the threshold', () => { + resetFixture(); + seedTwoFilesOneCovered(); + const { status } = runVerifier(['--min-baseline-coverage', '0.51']); + assert.equal(status, 3, '1 of 2 covered does not satisfy a 0.51 threshold'); + }); + + /** Vacuous input: nothing checked cannot be under-covered. */ + test('#4135: strict coverage gate is vacuously satisfied when nothing is checked', () => { + resetFixture(); + const { status, report } = runVerifier(['--min-baseline-coverage', '1']); + assert.equal(status, 0); + assert.equal(report.checked, 0); + }); + + /** Arg validation: malformed thresholds are usage errors (exit 2). */ + test('#4135: malformed --min-baseline-coverage values are usage errors', () => { + resetFixture(); + for (const bad of ['1.5', '-1', 'notanumber']) { + const r = runNode([ + SCRIPT, + '--patches-dir', patchesDir, + '--config-dir', configDir, + '--pristine-dir', pristineDir, + '--json', + '--min-baseline-coverage', bad, + ], { timeoutMs: VERIFIER_TIMEOUT_MS }); + assert.equal(r.exitCode, 2, `threshold "${bad}" must be a usage error, not silently clamped`); + } + }); + + /** + * Core regression (widening): multi-version collapse — no baseline under + * gsd-pristine/, but the config dir is a git repository whose history + * holds the outgoing bytes (recorded pristine_hashes match). The + * recovered baseline must produce a REAL diff: the dropped user line is + * caught as fail_user_lines_missing, not skipped as ok_no_baseline. + */ + test('#4135: resolves the baseline from git history by recorded hash and catches a dropped user line', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-01.md'; + const pristine = + '# Flow 1\nOutgoing release stock line one with substance.\n' + + 'Outgoing release stock line two also with substance.\n'; + const userLine = 'user customisation line that the merge dropped for flow one'; + const backup = pristine + userLine + '\n'; + const upstreamRewrite = + '# Flow 1 (rewritten)\nIncoming release replaced the stock body upstream.\n'; + + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), backup); + // Config dir IS a git repo: outgoing pristine state, then the merged state. + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), pristine); + gitCommitAll(configDir, 'gsd install of the outgoing release'); + writeFile(path.join(configDir, rel), upstreamRewrite); // user line dropped + gitCommitAll(configDir, 'post-update merged state'); + + const { status, report } = runVerifier(); + + assert.equal(status, 1, 'the git-recovered baseline must catch the dropped line'); + assert.equal(report.no_baseline, 0, 'the baseline was recoverable — no skip'); + assert.equal(report.baseline_covered, 1); + const r0 = report.results[0]; + assert.equal(r0.status, 'fail'); + assert.equal(r0.reason, REASON.FAIL_USER_LINES_MISSING); + assert.deepEqual(r0.missing, [userLine], + 'the diff ran against the RECOVERED baseline — upstream-removed lines are not "missing"'); + }); + + /** + * Net observable: same recovery, user line PRESENT — the file is verified + * (exit 0, covered) instead of silently skipped. + */ + test('#4135: git-recovered baseline verifies surviving user lines instead of skipping', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-02.md'; + const pristine = '# Flow 2\nOutgoing stock line with substantial content.\n'; + const userLine = 'user customisation line that survived the merge for flow two'; + const upstreamRewrite = '# Flow 2 (rewritten)\nIncoming release rewrote the stock body.\n'; + + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + userLine + '\n'); + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), pristine); + gitCommitAll(configDir, 'gsd install of the outgoing release'); + writeFile(path.join(configDir, rel), upstreamRewrite + userLine + '\n'); + gitCommitAll(configDir, 'post-update merged state'); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.no_baseline, 0); + assert.equal(report.baseline_covered, 1); + assert.equal(report.results[0].status, 'ok'); + assert.notEqual(report.results[0].reason, REASON.OK_NO_BASELINE); + }); + + /** + * Version-hop boundary (N−1 / multi-commit span): history holds TWO + * upstream versions; the recorded hash is the OLDER one, so the walk must + * pass the newer (mismatching) commit and match deeper history — the + * single-hop shape of the same recovery. + */ + test('#4135: single-version hop also recovers the baseline from git history', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-03.md'; + const vA = '# Flow 3\nVersion A stock line with substantial content.\n'; + const vB = '# Flow 3\nVersion B stock line with substantial content.\n'; + const userLine = 'user customisation line present in the merged output'; + writeBackupMeta({ [rel]: sha256(vA) }); // outgoing = the OLDER commit's bytes + writeFile(path.join(patchesDir, rel), vA + userLine + '\n'); + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), vA); + gitCommitAll(configDir, 'gsd install vA'); + writeFile(path.join(configDir, rel), vB); + gitCommitAll(configDir, 'gsd update to vB'); + writeFile(path.join(configDir, rel), vB + userLine + '\n'); + gitCommitAll(configDir, 'post-update merged state'); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.no_baseline, 0, 'the walk must reach the older vA commit'); + assert.equal(report.baseline_covered, 1); + }); + + /** Negative space: no git, no pristine — the #934 advisory posture is unchanged. */ + test('#4135: non-git config dirs keep the ok_no_baseline advisory posture', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-04.md'; + const pristine = '# Flow 4\nOutgoing stock line with substantial content.\n'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + 'user line that survived the merge here\n'); + writeFile(path.join(configDir, rel), 'incoming rewrite\nuser line that survived the merge here\n'); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.no_baseline, 1); + assert.equal(report.results[0].reason, REASON.OK_NO_BASELINE); + assert.equal(report.baseline_covered, 0); + }); + + /** Negative space: git present but no blob matches the recorded hash. */ + test('#4135: git tier adopts nothing when no blob matches the recorded hash', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-05.md'; + const pristine = '# Flow 5\nOutgoing stock line with substantial content.\n'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + 'user line that survived the merge here\n'); + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), 'history only ever held user-modified bytes\n'); + gitCommitAll(configDir, 'only user state was ever committed'); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.no_baseline, 1, 'hash equality is the authority — nothing else is trusted'); + assert.equal(report.results[0].reason, REASON.OK_NO_BASELINE); + }); + + /** Drift-path lock: #3657 is never bypassed by the git tier. */ + test('#4135: canonical drift is never rescued by the git-history tier', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-06.md'; + const pristine = '# Flow 6\nOutgoing stock line with substantial content.\n'; + const drifted = '# Flow 6\nRefreshed newer-release stock line with substance.\n'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + 'user line that survived the merge here\n'); + writeFile(path.join(configDir, rel), drifted + 'user line that survived the merge here\n'); + writeFile(path.join(pristineDir, rel), drifted); // canonical present, hash-mismatched + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), pristine); + gitCommitAll(configDir, 'history holds the recorded-hash bytes'); + + const { status, report } = runVerifier(); + + assert.equal(status, 0); + assert.equal(report.results[0].reason, REASON.OK_PRISTINE_DRIFT_DETECTED, + 'a present-but-drifted canonical stays drift — no rescue'); + assert.equal(report.drifted, 1); + }); + + /** Precedence lock: a matching canonical pristine beats the git tier. */ + test('#4135: canonical pristine keeps precedence over the git-history tier', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-07.md'; + const pristine = '# Flow 7\nOutgoing stock line with substantial content.\n'; + const droppedLine = 'user customisation line that the merge dropped for flow seven'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + droppedLine + '\n'); + writeFile(path.join(configDir, rel), pristine); // dropped — caught via canonical + writeFile(path.join(pristineDir, rel), pristine); // canonical, hash-matching + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), pristine); + gitCommitAll(configDir, 'history also holds the bytes'); + + const { status, report } = runVerifier(); + + assert.equal(status, 1); + assert.equal(report.results[0].reason, REASON.FAIL_USER_LINES_MISSING); + assert.ok(report.results[0].missing.includes(droppedLine)); + }); + + /** Invocation-shape lock: no --pristine-dir → over-broad fallback, no git tier. */ + test('#4135: no git-history resolution when --pristine-dir is not provided', () => { + resetFixture(); + const rel = 'gsd-core/workflows/flow-08.md'; + const pristine = '# Flow 8\nOutgoing stock line with substantial content.\n'; + const userLine = 'user customisation line that survived the merge for flow eight'; + writeBackupMeta({ [rel]: sha256(pristine) }); + writeFile(path.join(patchesDir, rel), pristine + userLine + '\n'); + gitIn(configDir, ['init', '-q', '-b', 'main']); + writeFile(path.join(configDir, rel), pristine); + gitCommitAll(configDir, 'history holds the recorded-hash bytes'); + // Merge kept everything — over-broad mode passes on this shape. + writeFile(path.join(configDir, rel), pristine + userLine + '\n'); + + const r = runNode([ + SCRIPT, '--patches-dir', patchesDir, '--config-dir', configDir, '--json', + ], { timeoutMs: VERIFIER_TIMEOUT_MS }); + const report = r.stdout && r.stdout.length ? JSON.parse(r.stdout) : null; + + assert.equal(r.exitCode, 0, 'all backup lines present — over-broad fallback passes'); + assert.equal(report.results[0].reason, null, + 'without --pristine-dir the baseline logic (incl. git tier) must not run'); + assert.equal(report.baseline_covered, 0, + 'an over-broad run verifies nothing against a pristine baseline — coverage stays 0'); + }); + + /** Module unit: findPristineInGit contract on a real fixture repo. */ + test('#4135: findPristineInGit unit — match, no-match, non-repo', () => { + const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4135-git-unit-')); + const plain = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4135-plain-')); + try { + const rel = 'gsd-core/workflows/unit.md'; + const content = '# Unit\noutgoing pristine bytes for the git walk unit test\n'; + gitIn(root, ['init', '-q', '-b', 'main']); + writeFile(path.join(root, rel), content); + gitCommitAll(root, 'seed'); + writeFile(path.join(root, rel), 'later user-modified bytes committed on top\n'); + gitCommitAll(root, 'user state'); + + assert.equal(findPristineInGit(root, rel, sha256(content)), content, + 'an exact recorded-hash blob in history is returned verbatim'); + assert.equal(findPristineInGit(root, rel, sha256('bytes that never existed anywhere\n')), null, + 'no matching blob resolves to null'); + assert.equal(findPristineInGit(plain, rel, sha256(content)), null, + 'a non-repo directory resolves to null without throwing'); + } finally { + cleanup(root); + cleanup(plain); + } + }); + + /** + * Workflow contract row (source-text-is-the-product): Step 5a must make + * coverage a headline and document the opt-in strict gate. + */ + test('#4135: Step 5a headlines baseline coverage and documents the opt-in strict gate', () => { + // allow-test-rule: source-text-is-the-product — Step 5a contract text (see #4135) + const content = fs.readFileSync(WORKFLOW_PATH, 'utf8'); + assert.ok(content.includes('Baseline coverage:'), + 'Step 5a must print a Baseline coverage headline, not bury no_baseline in parseable fields'); + assert.ok(content.includes('--min-baseline-coverage'), + 'the opt-in strict coverage gate must be documented for operators'); + }); + + /** + * Fixture: 2 files, exactly 1 baseline-covered — the minimal exact-ratio + * fixture for the threshold boundary rows. + */ + function seedTwoFilesOneCovered() { + const specs = [ + { rel: 'gsd-core/workflows/a.md', covered: true }, + { rel: 'gsd-core/workflows/b.md', covered: false }, + ]; + const pristineHashes = {}; + for (const { rel, covered } of specs) { + const pristine = `outgoing stock line with substantial content for ${rel}\n`; + const userLine = `user customisation line that survived the merge for ${rel}\n`; + pristineHashes[rel] = sha256(pristine); + writeFile(path.join(patchesDir, rel), pristine + userLine); + writeFile(path.join(configDir, rel), `incoming rewrite for ${rel}\n` + userLine); + if (covered) writeFile(path.join(pristineDir, rel), pristine); + } + writeBackupMeta(pristineHashes); + } +}); + }); +} From 03738824dec9b46077752e6f8f628c096a5faa1e Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 05:46:49 -0400 Subject: [PATCH 010/166] enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) --- .changeset/patient-quails-greet.md | 5 + CONTEXT.md | 2 +- bin/install.js | 122 +++--- capabilities/codex/capability.json | 6 +- docs/how-to/install-on-your-runtime.md | 7 +- .../host-integration-capability-matrix.md | 2 +- gsd-core/bin/lib/capability-registry.cjs | 12 +- gsd-core/bin/lib/capability-validator.cjs | 5 + src/runtime-hooks-surface.cts | 207 +++++++++- tests/codex-config.test.cjs | 356 +++++++++++++++++- tests/fixtures/install-tree/codex.json | 4 - tests/install-minimal-hooks.test.cjs | 61 +-- 12 files changed, 668 insertions(+), 121 deletions(-) create mode 100644 .changeset/patient-quails-greet.md diff --git a/.changeset/patient-quails-greet.md b/.changeset/patient-quails-greet.md new file mode 100644 index 000000000..e6ad01da0 --- /dev/null +++ b/.changeset/patient-quails-greet.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 4367 +--- +**Codex no longer installs a context-monitor hook that could never fire.** `gsd-context-monitor.js` read a remaining-context bridge file only Claude Code's statusline hook writes, so every one of its Codex hook-event registrations was a guaranteed silent no-op. Fresh Codex installs no longer copy or register it; a reinstall over an older install now removes the stale registrations and the orphaned script. Agent-facing context warnings and phase/lifecycle display are documented as unsupported on Codex until a real metrics producer exists for that runtime. (#2586) diff --git a/CONTEXT.md b/CONTEXT.md index dcb451fef..1f2e6a438 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -452,7 +452,7 @@ A per-agent narrative entry written by `mempalace_diary_write`. GSD's `gsd-mempa The `mempalace.memory_mode` config key controlling how authoritative MemPalace is during recall/capture relative to GSD's native memory. Three wired values: `augment` (default — palace is an additive recall layer; native memory stays authoritative; lowest coupling), `kg_backend` (knowledge-graph queries resolve against MemPalace's temporal graph as the primary source, `.planning/graphs/` as fallback; non-KG drawer recall stays additive), `replace` (recall resolves through the palace as the source of truth, native artifacts as fallback). Every mode is `onError:skip` and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing `.planning/graphs/`, so no mode loses memory. Read at hook-render time; switching is a config change, not a reinstall. Cross-mode migration of existing `.planning/graphs/` into the palace is a separate, not-yet-implemented concern (PRD/ADR §17 open question). See MemPalace Settings in `docs/CONFIGURATION.md`. ### Runtime Hooks Surface Module -Standalone hook-surface writer module extracted from `bin/install.js` as ADR-857 phase 5f-1 (behavior-preserving relocation, no logic change). Owns: Cline rules-body/agents-md/pre-tool-use hook generation (`buildClineRulesBody`, `buildClineAgentsMdBody`, `buildClinePreToolUseHook`, `mergeGsdAgentsMd`, `writeClineArtifacts`); Cursor `hooks.json` lifecycle (`buildCursorHookEntry`, `isManagedCursorHookEntry`, `reconcileCursorHooksJson`, `writeCursorHooksJson`, `removeCursorHooksJson`); Copilot session-hook config (`buildCopilotHookConfig`, `writeCopilotHookConfig`); Codex hook-block and event management (`buildCodexHookBlock`, `rewriteLegacyCodexHookBlock`, `reconcileCodexHooksJsonEvent`, `reconcileCodexHooksJsonSessionStart`, `ensureCodexHooksJsonSessionStart`, `ensureCodexHooksJsonEvent`, `removeCodexHooksJsonEvent`, `removeCodexHooksJsonSessionStart`, `buildCodexHookWindowsShimIR`); Kimi native config.toml `[[hooks]]` lifecycle (`buildKimiHooksTomlBlock`, `stripKimiHooksTomlBlock`, `writeKimiHooksToml`, `removeKimiHooksToml` — #2095 EoS/kimi Upgrade 1, the first genuinely NEW hook surface added post-relocation rather than a behavior-preserving move: kimi's `[[hooks]]` array lives in its own native `config.toml`, resolved by `resolveKimiHooksTomlDir` in Runtime Homes Module to a directory deliberately separate from kimi's GSD configDir, wrapped in `# GSD Hooks BEGIN`/`END` marker comments for idempotent reinstall); and shared hook command helpers (`buildHookCommand`, `rewriteLegacyManagedNodeHookCommands`, `normalizeNodePath`, `resolveNodeRunner`, and — #3662 — `buildNodeRunnerChainToken`, the POSIX-sh runner token that resolves node at hook-fire time for non-portable managed JS hooks, plus the `NODE_RUNNER_RESOLVER_HOOK` basename of the staged `hooks/gsd-node-runner.sh` resolver that portable installs route through with the baked node path as its first argument). `bin/install.js` delegates to this module via thin wrappers and re-exports its functions unchanged so existing tests require no modification. Source: `src/runtime-hooks-surface.cts`. Built output: `gsd-core/bin/lib/runtime-hooks-surface.cjs`. +Standalone hook-surface writer module extracted from `bin/install.js` as ADR-857 phase 5f-1 (behavior-preserving relocation, no logic change). Owns: Cline rules-body/agents-md/pre-tool-use hook generation (`buildClineRulesBody`, `buildClineAgentsMdBody`, `buildClinePreToolUseHook`, `mergeGsdAgentsMd`, `writeClineArtifacts`); Cursor `hooks.json` lifecycle (`buildCursorHookEntry`, `isManagedCursorHookEntry`, `reconcileCursorHooksJson`, `writeCursorHooksJson`, `removeCursorHooksJson`); Copilot session-hook config (`buildCopilotHookConfig`, `writeCopilotHookConfig`); Codex hook-block and event management (`buildCodexHookBlock`, `rewriteLegacyCodexHookBlock`, `reconcileCodexHooksJsonEvent`, `reconcileCodexHooksJsonSessionStart`, `ensureCodexHooksJsonSessionStart`, `removeCodexHooksJsonEvent`, `removeCodexHooksJsonSessionStart`, `buildCodexHookWindowsShimIR`, `cleanupOrphanedCodexContextMonitorScript`, `isGsdOwnedCodexContextMonitorScript`, `hooksJsonReferencesCodexContextMonitor` — #2586, the last three replacing the deleted `ensureCodexHooksJsonEvent`: GSD no longer adds Codex context-monitor hook-event registrations, only recognizes and removes stale ones left by a pre-#2586 install, then deletes the orphaned script/`.cmd` shim once unreferenced and content-verified as GSD-owned); Kimi native config.toml `[[hooks]]` lifecycle (`buildKimiHooksTomlBlock`, `stripKimiHooksTomlBlock`, `writeKimiHooksToml`, `removeKimiHooksToml` — #2095 EoS/kimi Upgrade 1, the first genuinely NEW hook surface added post-relocation rather than a behavior-preserving move: kimi's `[[hooks]]` array lives in its own native `config.toml`, resolved by `resolveKimiHooksTomlDir` in Runtime Homes Module to a directory deliberately separate from kimi's GSD configDir, wrapped in `# GSD Hooks BEGIN`/`END` marker comments for idempotent reinstall); and shared hook command helpers (`buildHookCommand`, `rewriteLegacyManagedNodeHookCommands`, `normalizeNodePath`, `resolveNodeRunner`, and — #3662 — `buildNodeRunnerChainToken`, the POSIX-sh runner token that resolves node at hook-fire time for non-portable managed JS hooks, plus the `NODE_RUNNER_RESOLVER_HOOK` basename of the staged `hooks/gsd-node-runner.sh` resolver that portable installs route through with the baked node path as its first argument). `bin/install.js` delegates to this module via thin wrappers and re-exports its functions unchanged so existing tests require no modification. Source: `src/runtime-hooks-surface.cts`. Built output: `gsd-core/bin/lib/runtime-hooks-surface.cjs`. ### Runtime Config Adapter Registry Module owning the explicit per-runtime config-mutation dispatch table for the installer. `resolveRuntimeConfigIntent(runtime)` projects a typed config intent — `installSurface` (`settings-json` | `codex-toml` | `copilot-instructions` | `cline-rules` | `cursor-hooks-json` | `profile-marker-only`), `writesSharedSettings` (the `finishInstall` shared-settings write gate), and `finishPermissionWriter` (`opencode` | `kilo` | `antigravity` | none) — that `bin/install.js` dispatches on instead of inline `runtime === '...'` branching. Owns adapter selection only: it performs no filesystem IO and does not execute config mutations (the install/finishInstall handlers and the per-runtime writers do that). Unknown runtimes fail loudly with a `TypeError`, guarded by an `Object.hasOwn` own-property check so prototype-chain keys (`__proto__`, `constructor`) also throw. Also exports `resolveInstallPlan(runtime)` — the ADR-58 `InstallPlan` capstone — which collects the install-level descriptor axes (`installSurface`, `writesSharedSettings`, `finishPermissionWriter`, `hookEvents`, `extendedHookEvents`, `hooksSurface`, `sandboxTier`) into one typed `InstallPlan` value consumed by `install()` and `finishInstall()` in `bin/install.js`. `sandboxTier` (`none` | `codex-agent-sandbox`) gates per-agent `sandbox_mode` emission in the codex TOML path and fails loud on a missing/invalid value (#1151). The spatial axes (`configHome`, `artifactLayout`, `commandStyle`) remain behind their self-resolving adapter modules and are not part of the plan; they are the execution adapters. Realizes both the adapter-selection and plan-collection halves of the Runtime Install Policy Module boundary. Source: `gsd-core/bin/lib/runtime-config-adapter-registry.cjs`. See ADR-58, #60. diff --git a/bin/install.js b/bin/install.js index a578aadcc..ade940d41 100755 --- a/bin/install.js +++ b/bin/install.js @@ -1369,34 +1369,13 @@ function ensureCodexHooksJsonSessionStart(targetDir, opts = {}) { } /** - * Ensure hooks.json contains exactly one managed GSD hook entry for the given - * Codex event, wired to gsd-context-monitor.js. Preserves user-owned entries. - * - * Used for the new Codex events added in #772: - * SubagentStart — inject context / GSD_AGENT_NAME awareness at subagent open - * Stop — post-session context headroom tracking - * PostToolUse — mirror the Claude Code PostToolUse context monitor - * - * All three events are routed through gsd-context-monitor.js — the same hook - * used for PostToolUse in the Claude Code baseline — so context-headroom - * warnings surface at these key Codex session lifecycle moments. - * - * On Windows (#3426): writes a gsd-context-monitor.cmd shim alongside the .js - * file and uses the .cmd path as the hook command — exactly the same fix as - * SessionStart uses for gsd-check-update — to avoid the bash.exe POSIX-exec - * failure when Codex's hook dispatcher tries to run node.exe through Git Bash. - * - * @param {string} targetDir - * @param {string} eventName - One of 'SubagentStart', 'Stop', 'PostToolUse'. - * @param {{ absoluteRunner: string|null, platform?: NodeJS.Platform }} opts - * @returns {{ changed: boolean, wrote: boolean, path: string }} - */ -function ensureCodexHooksJsonEvent(targetDir, eventName, opts = {}) { - return hooksSurface.ensureCodexHooksJsonEvent(targetDir, eventName, opts); -} - -/** - * Remove a GSD-managed event entry from hooks.json. Called during uninstall. + * Remove a GSD-managed event entry from hooks.json. Called during uninstall, + * and (#2586) unconditionally during install/reinstall to clean up a + * pre-#2586 install's stale gsd-context-monitor.js registrations — GSD no + * longer ADDS entries for these events (see CODEX_HOOKS_TO_COPY / + * cleanupOrphanedCodexContextMonitorScript in bin/install.js's Codex branch), + * only removes recognized ones, so the `ensureCodexHooksJsonEvent` wrapper + * that used to add them was removed as dead code. * * @param {string} targetDir * @param {string} eventName @@ -8485,6 +8464,16 @@ function uninstall(isGlobal, runtime = DEFAULT_RUNTIME) { console.log(` ${green}✓${reset} Removed managed Codex ${eventName} hook from hooks.json`); } } + // #2586: uninstall's own symmetric half of the orphaned-script cleanup — + // same GSD-owned + unreferenced gate as the install-time call. + const uninstallMonitorCleanup = hooksSurface.cleanupOrphanedCodexContextMonitorScript(targetDir); + for (const deletedPath of uninstallMonitorCleanup.deleted) { + removedCount++; + console.log(` ${green}✓${reset} Removed orphaned Codex hook script (${path.basename(deletedPath)})`); + } + for (const warning of uninstallMonitorCleanup.warnings) { + console.warn(` ${yellow}⚠${reset} Could not remove orphaned Codex hook script ${warning.path}: ${warning.reason}`); + } } // 1a-kimi. Non-layout Kimi side-effect (#2095 EoS/kimi Upgrade 1): kimi's @@ -12323,11 +12312,17 @@ function install(isGlobal, runtime = DEFAULT_RUNTIME, options = {}) { // stageTransitiveHookLibs call after the copy loop. The #3579 boundary is // preserved: helpers no staged Codex hook requires (graphify tooling among // them) are still not shipped. + // #2586: gsd-context-monitor.js is deliberately NOT copied for Codex. + // It reads the statusline bridge file (${TMPDIR}/claude-ctx-{session_id}.json) + // written only by hooks/gsd-statusline.js, which Codex never installs — so + // every registered event was a guaranteed silent no-op (readSentinel throws + // ENOENT -> allow(undefined), every invocation, every event, no exceptions). + // A pre-#2586 install's stale copy + hooks.json registrations are cleaned + // up below (see the CODEX_EXTENDED_HOOK_EVENTS loop), not re-added here. const CODEX_HOOKS_TO_COPY = [ 'gsd-check-update.js', 'gsd-check-update-worker.js', 'managed-hooks-registry.cjs', - 'gsd-context-monitor.js', ]; const codexHooksSrc = path.join(src, 'hooks', 'dist'); if (fs.existsSync(codexHooksSrc)) { @@ -12529,35 +12524,48 @@ function install(isGlobal, runtime = DEFAULT_RUNTIME, options = {}) { } } - // ── Codex extended hook events (#772, #2088) ───────────────────────── - // Codex CLI stabilised a full hook-event set in rust-v0.137.0. GSD - // registers CODEX_EXTENDED_HOOK_EVENTS (#2088 adds the 6 documented - // events beyond the original #772 three) — all routed through - // gsd-context-monitor.js so context-headroom warnings surface at each - // lifecycle point: SubagentStart/SubagentStop (subagent open/close), - // Stop (final-response), PreToolUse/PostToolUse (tool boundaries), - // PermissionRequest (approval prompts), Pre/PostCompact (context - // compaction), and UserPromptSubmit (per-turn context injection). The - // context-monitor script decides per-payload what to do; unregistered - // events simply never fire. - // - // Guard: only register when the context-monitor file exists and the node - // runner is available — same guards as the SessionStart path above. - const contextMonitorFile = path.join(targetDir, 'hooks', 'gsd-context-monitor.js'); - if (codexNodeRunner && fs.existsSync(contextMonitorFile)) { - for (const codexEvent of CODEX_EXTENDED_HOOK_EVENTS) { - const eventWrite = ensureCodexHooksJsonEvent(targetDir, codexEvent, { - absoluteRunner: codexNodeRunner, - platform: process.platform, - }); - if (eventWrite.wrote) { - console.log(` ${green}✓${reset} Configured Codex hooks (${codexEvent} via hooks.json)`); - } else if (eventWrite.changed) { - console.log(` ${green}✓${reset} Verified Codex hooks (${codexEvent} via hooks.json)`); - } + // #2586: Codex's hook payload carries no context/token-usage field + // (confirmed against codex-rs/hooks/src/schema.rs), so agent-facing + // context warnings and GSD phase/lifecycle display cannot be + // supported on this runtime — state that plainly during install + // rather than silently omitting the capability. Matches + // capabilities/codex/capability.json's hostBehaviors.unsupportedFeatures. + console.log(` ${dim}↳${reset} Codex: agent-facing context warnings and GSD phase/lifecycle display are unsupported (Codex's hook payload has no context-usage metric)`); + + // ── Codex extended hook events (#772, #2088) — REMOVED by #2586 ────── + // gsd-context-monitor.js is no longer copied or registered for Codex + // (see the CODEX_HOOKS_TO_COPY comment above): every one of these + // events was a guaranteed silent no-op, since the metrics bridge file + // it reads is only ever written by Claude's own statusline hook. + // Every event in CODEX_EXTENDED_HOOK_EVENTS is unconditionally + // reconciled here — not gated on the script existing — so a + // pre-#2586 install's stale registrations (exact current shape, or a + // recognized legacy shape via isManagedHookCommand's + // includeLegacyAliases) are stripped on reinstall. Mirrors the + // unconditional uninstall-time loop over the same constant. A + // registration whose command does not match the managed shape (a + // hand-customized entry) survives untouched — see + // reconcileCodexHooksJsonEvent's isManagedHookCommand filter. + for (const codexEvent of CODEX_EXTENDED_HOOK_EVENTS) { + const eventCleanup = removeCodexHooksJsonEvent(targetDir, codexEvent); + if (eventCleanup.changed) { + console.log(` ${green}✓${reset} Removed stale Codex ${codexEvent} context-monitor hook from hooks.json`); } - } else if (!codexNodeRunner) { - console.warn(` ${yellow}⚠${reset} Skipped Codex extended hook-event registration — Node runner unavailable.`); + } + // Delete the orphaned script (+ Windows .cmd shim) left by a + // pre-#2586 install, but ONLY once no surviving hooks.json + // registration under any event still references it, and only when + // the on-disk file is GSD's own (see design doc's Ownership check — + // a content-signature check, not manifest membership, so this works + // on the very first reinstall after upgrading, with no bootstrap + // gap). A deletion failure never reverts the (already safe, + // already-written) hooks.json cleanup above — must-have #8. + const monitorCleanup = hooksSurface.cleanupOrphanedCodexContextMonitorScript(targetDir); + for (const deletedPath of monitorCleanup.deleted) { + console.log(` ${green}✓${reset} Removed orphaned Codex hook script (${path.basename(deletedPath)})`); + } + for (const warning of monitorCleanup.warnings) { + console.warn(` ${yellow}⚠${reset} Could not remove orphaned Codex hook script ${warning.path}: ${warning.reason}`); } // ── end Codex extended hook events ──────────────────────────────────── } diff --git a/capabilities/codex/capability.json b/capabilities/codex/capability.json index 71bcd1f53..56dad8197 100644 --- a/capabilities/codex/capability.json +++ b/capabilities/codex/capability.json @@ -109,7 +109,11 @@ "tomlConfigInstall": true, "cleanupSkillSidecars": true, "agentTomlFiles": true, - "frontmatterDialect": "codex" + "frontmatterDialect": "codex", + "unsupportedFeatures": [ + "context-warnings", + "phase-lifecycle-display" + ] } }, "reviewer": { diff --git a/docs/how-to/install-on-your-runtime.md b/docs/how-to/install-on-your-runtime.md index f829063a7..9dc286e3b 100644 --- a/docs/how-to/install-on-your-runtime.md +++ b/docs/how-to/install-on-your-runtime.md @@ -169,14 +169,13 @@ Skills land in `~/.codex/skills/gsd-*/SKILL.md`. Agents are written as standalon **Hook coverage** -GSD registers the following Codex hook events automatically on install (requires Codex CLI 0.137.0+ for the stable hook-event schema): +GSD registers the following Codex hook event automatically on install (requires Codex CLI 0.137.0+ for the stable hook-event schema): | Event | Hook | Purpose | |---|---|---| | `SessionStart` | `gsd-check-update.js` | Update check at session open; Windows installs also emit a `commandWindows` field pointing to the `.cmd` shim so Codex picks the correct executor on Windows without requiring per-OS config regeneration | -| `SubagentStart` | `gsd-context-monitor.js` | Inject context / GSD_AGENT_NAME awareness at subagent open | -| `Stop` | `gsd-context-monitor.js` | Context headroom tracking before model stop | -| `PostToolUse` | `gsd-context-monitor.js` | Mirror the context-monitor coverage available in Claude Code | + +**Context warnings are not supported on Codex (#2586).** Earlier revisions of GSD also registered `SubagentStart`/`Stop`/`PostToolUse` (plus, briefly, six more events) against `gsd-context-monitor.js` for context-headroom tracking. That hook only produces a warning by reading a remaining-context-percentage bridge file that `gsd-statusline.js` — Claude Code's own statusline mechanism — writes; Codex never installs a statusline writer, so every one of those registrations fired as a guaranteed silent no-op, every invocation, with no exceptions. GSD no longer copies or registers `gsd-context-monitor.js` on a fresh Codex install; a reinstall over an older GSD install removes the stale registrations and the now-unreferenced script automatically. Agent-facing context warnings and GSD phase/lifecycle display remain unsupported capabilities on Codex (see `capabilities/codex/capability.json`) until a real metrics producer exists for this runtime — native Codex `/statusline` configuration is a separate, not-yet-implemented surface. All registered hooks are managed by GSD and are removed cleanly on `--uninstall`. diff --git a/docs/reference/host-integration-capability-matrix.md b/docs/reference/host-integration-capability-matrix.md index 80f36b73a..aedda8f8a 100644 --- a/docs/reference/host-integration-capability-matrix.md +++ b/docs/reference/host-integration-capability-matrix.md @@ -130,7 +130,7 @@ Sources consulted: **GSD integration status — Phase D dogfood complete (#2088, ADR-1239).** Codex installs through the `declarative` embedding adapter (`createDeclarativeAdapter` → `installRuntimeArtifacts`); the hardcoded `runtime === 'codex'`/`isCodex` projection is folded into descriptor-driven `runtime.hostBehaviors`, and install/uninstall output is byte-parity-gated at the time (`tests/fixtures/golden-install-parity/codex.json`; superseded by the differential attribution check, #2724). Three capability upgrades land, each with a test driving the user-reachable surface: - **Skill root** — global skills install to the canonical `$HOME/.agents/skills` (Codex core-skills `loader.rs` user-scope root), not the deprecated `$CODEX_HOME/skills` fallback; local skills install to `/.codex/skills`. The global path is declared via the global skills-kind `home: ".agents"` override, while the local kind intentionally has no home override. Pre-move global installs are migrated (stale `~/.codex/skills/gsd-*` cleaned on both install and uninstall); local installs do not remove `$HOME/.agents/skills` because those skills may be intentionally global. -- **Hook events** — GSD registers all documented `hooks.json` lifecycle events beyond `SessionStart`: `SubagentStart`, `Stop`, `PostToolUse` (#772), plus the six added in #2088 — `PreToolUse`, `PermissionRequest`, `PreCompact`, `PostCompact`, `SubagentStop`, `UserPromptSubmit` — all routed through `gsd-context-monitor.js`. (The descriptor `extendedHookEvents` field reflects the schema-valid cross-runtime subset `SubagentStop`/`Stop`/`PreCompact`; Codex's full event set is codex-hooks-json-native, registered directly in `hooks.json`.) +- **Hook events (corrected #2586)** — GSD registers the `SessionStart` event only, wired to `gsd-check-update.js`. The `SubagentStart`/`Stop`/`PostToolUse` (#772) + six #2088 extended events (`PreToolUse`, `PermissionRequest`, `PreCompact`, `PostCompact`, `SubagentStop`, `UserPromptSubmit`) described in earlier revisions of this doc were routed through `gsd-context-monitor.js`, which reads a remaining-context-percentage bridge file (`${TMPDIR}/claude-ctx-{session_id}.json`) that only `gsd-statusline.js` writes — a Claude-only mechanism Codex never installs. Every one of those events was therefore a guaranteed silent no-op on Codex (confirmed: `readSentinel` throws `ENOENT` on every invocation, unconditionally). #2586 stops copying/registering `gsd-context-monitor.js` on fresh installs and cleans up a pre-#2586 install's stale registrations + orphaned script on reinstall/uninstall. Agent-facing context warnings and GSD phase/lifecycle display are consequently **unsupported on Codex** (`capabilities/codex/capability.json`'s `hostBehaviors.unsupportedFeatures: ["context-warnings","phase-lifecycle-display"]`) — no working warning path existed before this change either, so nothing regresses. Native Codex `/statusline` configuration remains a separate, out-of-scope surface. (The descriptor `extendedHookEvents` field's `SubagentStop`/`Stop`/`PreCompact` value is a pre-existing, unrelated drift against the fuller event set `bin/install.js` used to register — not corrected by #2586.) - **Dispatch tuning** — `[agents] max_depth = 1` is written explicitly into the managed `config.toml` block, pinning the `dispatch.maxDepth: 1` axis instead of relying on codex-cli's implicit default. Because `maxDepth === 1`, `degradationFor` flattens GSD-hosted wave dispatch to single-level even though `dispatch.nested`/`background`/`backgroundDispatch` are all `true`. The block is a bare `[agents]` AgentsToml scalar table; it does **not** carry per-role `[agents.gsd-*]` sub-tables — those pointed `config_file` back at the standalone `agents/gsd-*.toml` files Codex already auto-discovers, so emitting them was a duplicate role registration (Codex logged "Ignoring malformed agent role definition: duplicate agent role name" once per agent) removed in #2406. `validateCodexConfigSchema` permits a known-scalar-only `[agents]` while still rejecting `[[agents]]` and unknown-key forms. Sources consulted: diff --git a/gsd-core/bin/lib/capability-registry.cjs b/gsd-core/bin/lib/capability-registry.cjs index 05c338919..cd08d58eb 100644 --- a/gsd-core/bin/lib/capability-registry.cjs +++ b/gsd-core/bin/lib/capability-registry.cjs @@ -1225,7 +1225,11 @@ const capabilities = { "tomlConfigInstall": true, "cleanupSkillSidecars": true, "agentTomlFiles": true, - "frontmatterDialect": "codex" + "frontmatterDialect": "codex", + "unsupportedFeatures": [ + "context-warnings", + "phase-lifecycle-display" + ] } }, "reviewer": { @@ -6337,7 +6341,11 @@ const runtimes = { "tomlConfigInstall": true, "cleanupSkillSidecars": true, "agentTomlFiles": true, - "frontmatterDialect": "codex" + "frontmatterDialect": "codex", + "unsupportedFeatures": [ + "context-warnings", + "phase-lifecycle-display" + ] } }, "reviewer": { diff --git a/gsd-core/bin/lib/capability-validator.cjs b/gsd-core/bin/lib/capability-validator.cjs index 3966b22c0..ff00dc58d 100644 --- a/gsd-core/bin/lib/capability-validator.cjs +++ b/gsd-core/bin/lib/capability-validator.cjs @@ -1956,6 +1956,11 @@ const KNOWN_HOST_BEHAVIORS = new Set([ 'sourceMarkerFile', 'tomlConfigInstall', 'trackCategoryDescription', + // #2586: declares runtime-level feature axes GSD does not/cannot support on + // this host (e.g. Codex's `["context-warnings","phase-lifecycle-display"]`) + // — present-and-populated / absent-is-unsupported-empty convention, so + // omitting the key on every other runtime carries no inverted meaning. + 'unsupportedFeatures', 'verificationStyle', 'writeCategoryDescription', ]); diff --git a/src/runtime-hooks-surface.cts b/src/runtime-hooks-surface.cts index ee1f60f9a..50d0f6c84 100644 --- a/src/runtime-hooks-surface.cts +++ b/src/runtime-hooks-surface.cts @@ -923,16 +923,68 @@ interface ReconcileResult { path: string; } +/** + * Lazily require install-engine.cjs's `hasExistingSymlinkBetween` / + * `isSymlinkedDestOptIn` — mirrors user-artifact-staging.cts's + * `_installEngineSymlinkGuard` (same call-time-require rationale: avoid a + * static circular require between install-engine.cts and this module). + */ +interface InstallEngineSymlinkGuard { + hasExistingSymlinkBetween: (root: string, fullPath: string, options?: { allowOptInFollow?: boolean }) => boolean; + isSymlinkedDestOptIn: () => boolean; +} + +function _installEngineSymlinkGuard(): InstallEngineSymlinkGuard { + // eslint-disable-next-line @typescript-eslint/no-require-imports, @typescript-eslint/no-unsafe-assignment + const mod: InstallEngineSymlinkGuard = require('./install-engine.cjs'); + return mod; +} + function reconcileCodexHooksJsonEvent(targetDir: string, eventName: string, opts: ReconcileCodexOpts = {}): ReconcileResult { const hooksJsonPath = path.join(targetDir, 'hooks.json'); const managedCommand = typeof opts.managedCommand === 'string' ? opts.managedCommand : null; const commandWindows = typeof opts.commandWindows === 'string' ? opts.commandWindows : null; const matcher = typeof opts.matcher === 'string' ? opts.matcher : undefined; const timeout = typeof opts.timeout === 'number' ? opts.timeout : undefined; + // #2586 Major 2: every Codex hooks.json writer funnels through this one + // function, and atomicWriteFileSync's final step is a rename(2) onto + // `hooksJsonPath` — which, when that path is a symlink, REPLACES the + // symlink with a plain file rather than writing through it. Refuse (with + // the same GSD_ALLOW_SYMLINKED_DEST opt-in every other install call site + // honors) before reading or writing, so a symlinked hooks.json is neither + // silently destroyed nor left the caller no escape hatch. + const symlinkGuard = _installEngineSymlinkGuard(); + // The path this function actually reads/writes. Defaults to the nominal + // hooks.json path; reassigned below to the symlink's real target when the + // opt-in is active, so the write lands on the file the user's symlink + // points at instead of clobbering the symlink itself (see note below). + let effectiveHooksJsonPath = hooksJsonPath; + if (fs.existsSync(hooksJsonPath) && fs.lstatSync(hooksJsonPath).isSymbolicLink()) { + if ( + symlinkGuard.hasExistingSymlinkBetween(targetDir, hooksJsonPath, { + allowOptInFollow: symlinkGuard.isSymlinkedDestOptIn(), + }) + ) { + throw new Error( + `hooks.json at "${hooksJsonPath}" contains a symlink the install root "${targetDir}" does not trust — ` + + 'refusing to read or write it. If this is an intentional user-owned symlink layout, re-run with ' + + 'GSD_ALLOW_SYMLINKED_DEST=1.', + ); + } + // hasExistingSymlinkBetween returned false only because the opt-in is + // active (a symlinked leaf always trips it otherwise) — so this IS a + // symlink and we are cleared to follow it. atomicWriteFileSync's final + // step is a rename(2) onto its target, which REPLACES an existing + // symlink at that path rather than writing through it; resolving to the + // real path here makes the read AND the write operate on the symlink's + // target, leaving the symlink itself untouched, matching what "follow" + // is supposed to mean. + effectiveHooksJsonPath = fs.realpathSync(hooksJsonPath); + } let parsed: Record = {}; let currentContent: string | null = null; - if (fs.existsSync(hooksJsonPath)) { - const raw = fs.readFileSync(hooksJsonPath, 'utf8'); + if (fs.existsSync(effectiveHooksJsonPath)) { + const raw = fs.readFileSync(effectiveHooksJsonPath, 'utf8'); currentContent = raw; if (raw.trim()) { try { @@ -965,6 +1017,12 @@ function reconcileCodexHooksJsonEvent(targetDir: string, eventName: string, opts } parsed['hooks'] = hookTable; const eventEntries = Array.isArray(hookTable[eventName]) ? (hookTable[eventName] as unknown[]) : []; + // Minor 5 (#2586 review): an event key the user already had, already + // holding an empty array, must survive removal as an empty array — not be + // deleted outright. Deleting is only correct when OUR removal is what + // emptied a previously non-empty array. Tracked before the loop below can + // mutate anything. + const wasArrayEmpty = Array.isArray(hookTable[eventName]) && eventEntries.length === 0; let removedLegacy = false; const sanitizedEntries: unknown[] = []; @@ -1002,6 +1060,10 @@ function reconcileCodexHooksJsonEvent(targetDir: string, eventName: string, opts if (sanitizedEntries.length > 0) { hookTable[eventName] = sanitizedEntries; + } else if (wasArrayEmpty) { + // Nothing of ours was ever here to remove — preserve the user's own + // empty array exactly as found (Minor 5). + hookTable[eventName] = []; } else { delete hookTable[eventName]; } @@ -1015,7 +1077,7 @@ function reconcileCodexHooksJsonEvent(targetDir: string, eventName: string, opts const changed = currentContent !== nextContent; const shouldWrite = changed && (currentContent !== null || Object.keys(parsed).length > 0); if (shouldWrite) { - atomicWriteFileSync(hooksJsonPath, nextContent, 'utf8'); + atomicWriteFileSync(effectiveHooksJsonPath, nextContent, 'utf8'); } return { changed: changed || removedLegacy, wrote: shouldWrite, path: hooksJsonPath }; @@ -1181,6 +1243,142 @@ function removeCodexHooksJsonSessionStart(targetDir: string): ReconcileResult { return reconcileCodexHooksJsonSessionStart(targetDir, { managedCommand: null }); } +// --------------------------------------------------------------------------- +// #2586: cleanupOrphanedCodexContextMonitorScript +// --------------------------------------------------------------------------- + +interface CleanupCodexContextMonitorResult { + /** Absolute paths of files actually deleted this call. */ + deleted: string[]; + /** {path, reason} for a file that could NOT be deleted (still present). */ + warnings: { path: string; reason: string }[]; + /** True if a surviving hooks.json registration still references the + * script (or its .cmd shim) — in which case nothing was deleted. */ + stillReferenced: boolean; +} + +// Literal, version-stable markers every shipped gsd-context-monitor.js +// carries. Stable across the {{GSD_VERSION}} and runtime-path substitutions +// the Codex copy step applies (#2586 design doc "Ownership check" — a raw +// content hash would differ per runtime/version by construction, so a marker +// check is used instead of manifest-membership, which has a bootstrap gap on +// the exact case that matters most: a pre-#2586 install's manifest never +// recorded this file at all). +const CODEX_CONTEXT_MONITOR_OWNERSHIP_MARKERS = [ + '#!/usr/bin/env node', + '// gsd-hook-version:', + '// Context Monitor - PostToolUse/AfterTool hook', +]; + +function isGsdOwnedCodexContextMonitorScript(filePath: string): boolean { + let content: string; + try { + content = fs.readFileSync(filePath, 'utf8'); + } catch { + return false; + } + // The .cmd shim (buildCodexHookWindowsShimIR) is a tiny generated batch + // wrapper, not the JS file itself — it never carries the JS markers above, + // so it gets its own narrower, still-specific signature: the exact + // "@ECHO OFF" / "@SETLOCAL" preamble the shim generator emits, invoking a + // script path that ends in gsd-context-monitor.js. + if (filePath.endsWith('.cmd')) { + return content.startsWith('@ECHO OFF') && content.includes('@SETLOCAL') + && /gsd-context-monitor\.js/.test(content); + } + return CODEX_CONTEXT_MONITOR_OWNERSHIP_MARKERS.every((marker) => content.includes(marker)); +} + +/** + * Scan every event in hooks.json for a surviving reference to the + * context-monitor script or its Windows .cmd shim, by basename — not scoped + * to CODEX_EXTENDED_HOOK_EVENTS, so a user who hand-registered it under an + * unrelated event key is still detected as "referenced" and the script is + * preserved. + */ +function hooksJsonReferencesCodexContextMonitor(targetDir: string): boolean { + const hooksJsonPath = path.join(targetDir, 'hooks.json'); + if (!fs.existsSync(hooksJsonPath)) return false; + let raw: string; + try { + raw = fs.readFileSync(hooksJsonPath, 'utf8'); + } catch { + return true; // unreadable — conservatively assume referenced, never delete + } + if (!raw.trim()) return false; + let parsed: unknown; + try { + parsed = JSON.parse(raw); + } catch { + return true; // unparseable — conservatively assume referenced + } + if (!parsed || typeof parsed !== 'object') return false; + const hooks = (parsed as Record)['hooks']; + const table = hooks && typeof hooks === 'object' && !Array.isArray(hooks) + ? (hooks as Record) + : (parsed as Record); + for (const key of Object.keys(table)) { + const entries = table[key]; + if (!Array.isArray(entries)) continue; + for (const entry of entries) { + if (!entry || typeof entry !== 'object') continue; + const entryHooks = (entry as Record)['hooks']; + const hookList = Array.isArray(entryHooks) ? entryHooks : [entry]; + for (const hook of hookList) { + if (!hook || typeof hook !== 'object') continue; + const values = [ + (hook as Record)['command'], + (hook as Record)['commandWindows'], + ]; + for (const value of values) { + if (typeof value === 'string' && /gsd-context-monitor(\.js|\.cmd)?/.test(value)) { + return true; + } + } + } + } + } + return false; +} + +/** + * #2586 must-have #4/#8: after hooks.json registrations for + * CODEX_EXTENDED_HOOK_EVENTS have been reconciled away (by the caller, via + * removeCodexHooksJsonEvent), delete `hooks/gsd-context-monitor.js` and its + * `.cmd` shim ONLY when (a) no surviving hooks.json registration under ANY + * event still references either basename, and (b) the on-disk file carries + * GSD's own ownership markers (a user's hand-edited or unrelated file at that + * path is left alone). Each file is deleted independently — a failure + * deleting one is reported as a warning and never rolls back the (already + * safe, already-written) hooks.json deregistration the caller performed + * first. + */ +function cleanupOrphanedCodexContextMonitorScript(targetDir: string): CleanupCodexContextMonitorResult { + const result: CleanupCodexContextMonitorResult = { deleted: [], warnings: [], stillReferenced: false }; + if (hooksJsonReferencesCodexContextMonitor(targetDir)) { + result.stillReferenced = true; + return result; + } + const candidates = [ + path.join(targetDir, 'hooks', 'gsd-context-monitor.js'), + path.join(targetDir, 'hooks', 'gsd-context-monitor.cmd'), + ]; + for (const candidate of candidates) { + if (!fs.existsSync(candidate)) continue; + if (!isGsdOwnedCodexContextMonitorScript(candidate)) continue; + try { + fs.unlinkSync(candidate); + result.deleted.push(candidate); + } catch (err) { + result.warnings.push({ + path: candidate, + reason: err && (err as Error).message ? (err as Error).message : String(err), + }); + } + } + return result; +} + // --------------------------------------------------------------------------- // Shared: buildHookCommand // --------------------------------------------------------------------------- @@ -3070,6 +3268,9 @@ export = { removeCodexHooksJsonEvent, removeCodexHooksJsonSessionStart, buildCodexHookWindowsShimIR, + cleanupOrphanedCodexContextMonitorScript, + isGsdOwnedCodexContextMonitorScript, + hooksJsonReferencesCodexContextMonitor, // Codex TOML buildCodexHookBlock, diff --git a/tests/codex-config.test.cjs b/tests/codex-config.test.cjs index b5ada431f..f67306355 100644 --- a/tests/codex-config.test.cjs +++ b/tests/codex-config.test.cjs @@ -73,6 +73,8 @@ const { CODEX_SANDBOX_HOLDS, parseTomlToObject, validateCodexConfigSchema, + uninstall, + CODEX_EXTENDED_HOOK_EVENTS, } = require('../bin/install.js'); const { resolveNodeRunner } = require('../gsd-core/bin/lib/runtime-hooks-surface.cjs'); @@ -4253,24 +4255,29 @@ describe('Codex install hook configuration (e2e)', () => { // and the managed-hooks registry. // // The Codex install branch in bin/install.js used to allowlist only two of the -// four hook files the shipped build emits (gsd-check-update.js + -// gsd-context-monitor.js), and gated the entire branch on !isMinimalMode so the +// four hook files the shipped build emitted at the time (gsd-check-update.js + +// gsd-context-monitor.js — the latter permanently removed by #2586, see below), +// and gated the entire branch on !isMinimalMode so the // `core` profile installed none of them. The parent SessionStart hook spawn()s // the worker, which require()s the registry — so Codex was wired to a dependency // chain the same installer never delivered. // // These tests drive the real installer (bin/install.js) behaviorally into an -// isolated temp config dir and assert the complete four-file set is delivered +// isolated temp config dir and assert the complete three-file set is delivered // for both profiles, the registry is byte-for-byte, the version stamps resolve // to the installed package version, and unrelated user files are preserved. // +// #2586 reduced the set back to three: gsd-context-monitor.js read a Claude-only +// statusline bridge file Codex never writes, so it was a guaranteed silent no-op +// on every Codex hook event and was dropped from CODEX_HOOKS_TO_COPY for good. +// // Verified non-duplicate: the pre-existing 'Codex install hook configuration // (e2e)' suite above only asserts gsd-check-update.js delivery/wiring — it never -// asserts on gsd-check-update-worker.js, managed-hooks-registry.cjs, or -// gsd-context-monitor.js delivery, the core/full profile matrix, upgrade-refresh, -// byte-for-byte registry copy, idempotency of the four-file set, user-file -// preservation, or the core-profile negative-space (no agent files) — all -// genuinely distinct assertions this fold adds. +// asserts on gsd-check-update-worker.js, managed-hooks-registry.cjs, the +// core/full profile matrix, upgrade-refresh, byte-for-byte registry copy, +// idempotency of the three-file set, user-file preservation, or the +// core-profile negative-space (no agent files) — all genuinely distinct +// assertions this fold adds. 'use strict'; @@ -4298,12 +4305,14 @@ const { INSTALL_TIMEOUT_MS, } = require('./helpers/timeouts.cjs'); -// The four-file hook set the Codex surface must deliver together (#2695). +// The three-file hook set the Codex surface must deliver together (#2695). +// gsd-context-monitor.js was removed from this set by #2586: it read a +// Claude-only statusline bridge file Codex never writes, so it was a +// guaranteed silent no-op on every Codex hook event. const CODEX_HOOK_FILES = [ 'gsd-check-update.js', 'gsd-check-update-worker.js', 'managed-hooks-registry.cjs', - 'gsd-context-monitor.js', ]; // Build hooks/dist before any install runs (the installer copies from there). @@ -4339,9 +4348,9 @@ function runCodexInstall({ profile, preseed }) { // Older-version stamp used to pre-seed an "upgrade" scenario. const OLDER_VERSION = '1.7.0'; -describe('#2695: fresh Codex installs deliver the complete four-file hook set', () => { +describe('#2695: fresh Codex installs deliver the complete three-file hook set', () => { for (const profile of ['core', 'full']) { - test(`fresh --profile=${profile} installs all four hook files`, (t) => { + test(`fresh --profile=${profile} installs all three hook files`, (t) => { const { configDir, result } = runCodexInstall({ profile }); t.after(() => cleanup(configDir)); @@ -4357,8 +4366,8 @@ describe('#2695: fresh Codex installs deliver the complete four-file hook set', } }); -describe('#2695: Codex upgrades refresh all four hook files to the current version', () => { - // Pre-seed all four files stamped at OLDER_VERSION so an upgrade must overwrite them. +describe('#2695: Codex upgrades refresh all three hook files to the current version', () => { + // Pre-seed all three files stamped at OLDER_VERSION so an upgrade must overwrite them. function olderSeed() { const seed = {}; for (const name of CODEX_HOOK_FILES) { @@ -4373,12 +4382,12 @@ describe('#2695: Codex upgrades refresh all four hook files to the current versi } for (const profile of ['core', 'full']) { - test(`--profile=${profile} upgrade refreshes all four hook files`, (t) => { + test(`--profile=${profile} upgrade refreshes all three hook files`, (t) => { const { configDir, result } = runCodexInstall({ profile, preseed: olderSeed() }); t.after(() => cleanup(configDir)); const hooksDir = hooksDirOf(configDir); - // All four must now carry the current version stamp where one exists, and + // All three must now carry the current version stamp where one exists, and // the registry must no longer be the stale sentinel. for (const name of CODEX_HOOK_FILES) { const dest = path.join(hooksDir, name); @@ -4473,8 +4482,8 @@ describe('#2695: unrelated user-owned hook files are preserved', () => { } }); -describe('#2695: re-running the installer is idempotent for the four-file set', () => { - test('a second full install leaves all four files present and correctly stamped', (t) => { +describe('#2695: re-running the installer is idempotent for the three-file set', () => { + test('a second full install leaves all three files present and correctly stamped', (t) => { const first = runCodexInstall({ profile: 'full' }); t.after(() => cleanup(first.configDir)); // Second run into the SAME config dir. @@ -10939,3 +10948,314 @@ describe('bug-2866: stripStaleGsdHookBlocks handles end-of-file without trailing }); }); } + +// ─── #2586: stop installing Codex context-monitor hooks without metrics ──── +// gsd-context-monitor.js reads a statusline bridge file Codex never writes, +// so every registered event was a guaranteed silent no-op. These tests drive +// the REAL install()/uninstall() entry points against a real temp CODEX_HOME +// — never a hand-fabricated manifest (see PR #2709's Blocker 1: its cleanup +// path was unreachable in production because its tests only exercised a +// fixture, not real installer state). +describe('#2586 Codex context-monitor: stop installing, clean up on reinstall', () => { + const { + cleanupOrphanedCodexContextMonitorScript, + isGsdOwnedCodexContextMonitorScript, + hooksJsonReferencesCodexContextMonitor, + } = require('../gsd-core/bin/lib/runtime-hooks-surface.cjs'); + + let codexHome; + + beforeEach(() => { + codexHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-codex-2586-')); + }); + + afterEach(() => { + cleanup(codexHome); + delete process.env.GSD_ALLOW_SYMLINKED_DEST; + }); + + function hooksJsonPath(home) { + return path.join(home, 'hooks.json'); + } + + function readHooksJson(home) { + const raw = fs.readFileSync(hooksJsonPath(home), 'utf8'); + return JSON.parse(raw); + } + + function monitorScriptPath(home) { + return path.join(home, 'hooks', 'gsd-context-monitor.js'); + } + + function monitorCmdShimPath(home) { + return path.join(home, 'hooks', 'gsd-context-monitor.cmd'); + } + + // uninstall(), unlike install(), does not sandbox CODEX_HOME/HOME itself — + // runCodexInstall's own env-restore in its `finally` means those are back + // to the real environment by the time a bare `uninstall(true, 'codex')` + // would run. Mirrors runCodexInstall's own sandboxing exactly so uninstall + // operates on the temp fixture, never the real ~/.codex. + function runCodexUninstall(codexHome, cwd = path.join(__dirname, '..')) { + const previousCodeHome = process.env.CODEX_HOME; + const previousHome = process.env.HOME; + const previousUserProfile = process.env.USERPROFILE; + const previousCwd = process.cwd(); + process.env.CODEX_HOME = codexHome; + process.env.HOME = codexHome; + process.env.USERPROFILE = codexHome; + try { + process.chdir(cwd); + return uninstall(true, 'codex'); + } finally { + process.chdir(previousCwd); + if (previousCodeHome === undefined) delete process.env.CODEX_HOME; + else process.env.CODEX_HOME = previousCodeHome; + if (previousHome === undefined) delete process.env.HOME; + else process.env.HOME = previousHome; + if (previousUserProfile === undefined) delete process.env.USERPROFILE; + else process.env.USERPROFILE = previousUserProfile; + } + } + + // The exact shape a real pre-#2586 install would have written: the shipped + // gsd-context-monitor.js content (with its ownership markers intact) plus + // hooks.json registrations for every CODEX_EXTENDED_HOOK_EVENTS member, + // using the same command-projection shape ensureCodexHooksJsonEvent used + // to write. Built from the real resolveNodeRunner() output, not a literal + // guess at the command string, so a drift in projectManagedHookCommand's + // output shape cannot make this fixture silently stop matching reality. + function seedPreExisting2586Install(home) { + fs.mkdirSync(path.join(home, 'hooks'), { recursive: true }); + // A literal fixture carrying the same ownership markers + // isGsdOwnedCodexContextMonitorScript looks for, built as a string + // (never read+string-matched from the real shipped source file — see + // this repo's "no source grep" test rule). + const fixtureContent = [ + '#!/usr/bin/env node', + '// gsd-hook-version: 1.12.0', + '// Context Monitor - PostToolUse/AfterTool hook', + 'process.exit(0);', + '', + ].join('\n'); + fs.writeFileSync(monitorScriptPath(home), fixtureContent, 'utf8'); + const runner = resolveNodeRunner(); + const hooksSurface = require('../gsd-core/bin/lib/runtime-hooks-surface.cjs'); + for (const eventName of CODEX_EXTENDED_HOOK_EVENTS) { + hooksSurface.ensureCodexHooksJsonEvent(home, eventName, { + absoluteRunner: runner, + platform: process.platform, + }); + } + } + + test('fresh install does not copy gsd-context-monitor.js or register any extended event', () => { + runCodexInstall(codexHome); + assert.strictEqual(fs.existsSync(monitorScriptPath(codexHome)), false, + 'gsd-context-monitor.js must not be copied on a fresh Codex install'); + assert.strictEqual(fs.existsSync(monitorCmdShimPath(codexHome)), false); + assert.strictEqual(fs.existsSync(path.join(codexHome, 'hooks', 'lib', 'hook-exit.js')), false, + 'hook-exit.js is only required by gsd-context-monitor.js — nothing else staged should pull it in'); + assert.ok(fs.existsSync(path.join(codexHome, 'hooks', 'gsd-check-update.js')), + 'gsd-check-update.js must still be staged'); + assert.ok(fs.existsSync(path.join(codexHome, 'hooks', 'gsd-check-update-worker.js'))); + assert.ok(fs.existsSync(path.join(codexHome, 'hooks', 'managed-hooks-registry.cjs'))); + if (fs.existsSync(hooksJsonPath(codexHome))) { + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), false); + } + }); + + test('reinstall removes exact pre-#2586 registrations for every extended event and deletes the orphaned script', () => { + runCodexInstall(codexHome); + seedPreExisting2586Install(codexHome); + assert.ok(fs.existsSync(monitorScriptPath(codexHome)), 'fixture sanity: script seeded'); + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), true, 'fixture sanity: registered'); + + runCodexInstall(codexHome); + + assert.strictEqual(fs.existsSync(monitorScriptPath(codexHome)), false, + 'orphaned gsd-context-monitor.js must be removed once unreferenced and GSD-owned'); + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), false, + 'no hooks.json entry may still reference gsd-context-monitor after reinstall'); + // gsd-check-update's own SessionStart registration must survive untouched. + const hooks = readHooksJson(codexHome); + const sessionStart = (hooks.hooks && hooks.hooks.SessionStart) || []; + const hasCheckUpdate = sessionStart.some((entry) => + (entry.hooks || []).some((h) => typeof h.command === 'string' && /gsd-check-update/.test(h.command))); + assert.ok(hasCheckUpdate, 'gsd-check-update SessionStart registration must remain after cleanup'); + }); + + test('reinstall preserves a hand-customized registration and does not delete a still-referenced script', () => { + runCodexInstall(codexHome); + seedPreExisting2586Install(codexHome); + // Hand-edit ONE event's entry to a shape isManagedHookCommand will not + // recognize (wraps the invocation in a shell script it does not know). + const before = readHooksJson(codexHome); + before.hooks.Stop = [{ hooks: [{ type: 'command', command: 'bash -c "/opt/custom/my-wrapper.sh"' }] }]; + fs.writeFileSync(hooksJsonPath(codexHome), JSON.stringify(before, null, 2) + '\n', 'utf8'); + + runCodexInstall(codexHome); + + const after = readHooksJson(codexHome); + assert.deepStrictEqual(after.hooks.Stop, before.hooks.Stop, + 'a hand-customized registration must survive verbatim'); + }); + + test('reinstall leaves unrelated hooks.json events and config.toml keys untouched', () => { + runCodexInstall(codexHome); + let hooks = fs.existsSync(hooksJsonPath(codexHome)) ? readHooksJson(codexHome) : { hooks: {} }; + if (!hooks.hooks) hooks.hooks = {}; + hooks.hooks.UnrelatedEvent = [{ hooks: [{ type: 'command', command: 'echo unrelated' }] }]; + fs.writeFileSync(hooksJsonPath(codexHome), JSON.stringify(hooks, null, 2) + '\n', 'utf8'); + const configPath = path.join(codexHome, 'config.toml'); + const configBefore = fs.readFileSync(configPath, 'utf8') + '\n[my_unrelated_section]\nfoo = "bar"\n'; + fs.writeFileSync(configPath, configBefore, 'utf8'); + + runCodexInstall(codexHome); + + const after = readHooksJson(codexHome); + assert.deepStrictEqual(after.hooks.UnrelatedEvent, hooks.hooks.UnrelatedEvent, + 'an unrelated event array must be untouched'); + const configAfter = fs.readFileSync(configPath, 'utf8'); + assert.ok(configAfter.includes('[my_unrelated_section]\nfoo = "bar"'), + 'unrelated config.toml section must survive a Codex reinstall'); + }); + + test('a user-owned pre-existing EMPTY event array is preserved, not deleted', () => { + runCodexInstall(codexHome); + fs.writeFileSync(hooksJsonPath(codexHome), JSON.stringify({ hooks: { Stop: [] } }, null, 2) + '\n', 'utf8'); + + runCodexInstall(codexHome); + + const after = readHooksJson(codexHome); + assert.ok(Array.isArray(after.hooks.Stop) && after.hooks.Stop.length === 0, + 'an empty array the user already had must not be dropped by cleanup that found nothing GSD-owned to remove'); + }); + + test('symlinked hooks.json aborts before any Codex change, without GSD_ALLOW_SYMLINKED_DEST', () => { + runCodexInstall(codexHome); + const configPath = path.join(codexHome, 'config.toml'); + const configBefore = fs.readFileSync(configPath, 'utf8'); + const realHooksJson = path.join(codexHome, 'real-hooks.json'); + fs.writeFileSync(realHooksJson, JSON.stringify({ hooks: {} }, null, 2) + '\n', 'utf8'); + fs.unlinkSync(hooksJsonPath(codexHome)); + fs.symlinkSync(realHooksJson, hooksJsonPath(codexHome)); + + assert.throws(() => runCodexInstall(codexHome), /symlink/i); + + assert.ok(fs.lstatSync(hooksJsonPath(codexHome)).isSymbolicLink(), + 'the symlink itself must survive an aborted install — never replaced by a plain file'); + assert.strictEqual(fs.readFileSync(configPath, 'utf8'), configBefore, + 'config.toml must be restored to its pre-attempt snapshot on abort'); + }); + + test('symlinked hooks.json is followed when GSD_ALLOW_SYMLINKED_DEST=1', () => { + runCodexInstall(codexHome); + const realHooksJson = path.join(codexHome, 'real-hooks.json'); + fs.writeFileSync(realHooksJson, JSON.stringify({ hooks: {} }, null, 2) + '\n', 'utf8'); + fs.unlinkSync(hooksJsonPath(codexHome)); + fs.symlinkSync(realHooksJson, hooksJsonPath(codexHome)); + process.env.GSD_ALLOW_SYMLINKED_DEST = '1'; + + assert.doesNotThrow(() => runCodexInstall(codexHome)); + + assert.ok(fs.lstatSync(hooksJsonPath(codexHome)).isSymbolicLink(), 'still a symlink afterward'); + assert.ok(fs.existsSync(realHooksJson), 'the symlink target must have been written through'); + }); + + test('cleanupOrphanedCodexContextMonitorScript keeps the hooks.json deregistration when script deletion fails', () => { + runCodexInstall(codexHome); + seedPreExisting2586Install(codexHome); + for (const eventName of CODEX_EXTENDED_HOOK_EVENTS) { + require('../gsd-core/bin/lib/runtime-hooks-surface.cjs').removeCodexHooksJsonEvent(codexHome, eventName); + } + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), false, 'deregistration committed first'); + + const originalUnlinkSync = fs.unlinkSync; + fs.unlinkSync = (target, ...rest) => { + if (typeof target === 'string' && target.includes('gsd-context-monitor')) { + throw Object.assign(new Error('EPERM: simulated'), { code: 'EPERM' }); + } + return originalUnlinkSync.call(fs, target, ...rest); + }; + // Determine which GSD-owned candidates actually exist BEFORE the mocked + // deletion attempt — on Windows, ensureCodexHooksJsonEvent also staged a + // .cmd shim alongside the .js file (see buildCodexHookWindowsShimIR), so + // both deletions fail under the mock above; on POSIX only the .js file + // exists. Asserting against this rather than a hardcoded 1 keeps the row + // meaningful on both platforms instead of just loosening it to "at least + // one" (see CI failure: Windows reported 2 warnings, not 1). + const cmdShimPath = monitorCmdShimPath(codexHome); + const cmdShimExisted = fs.existsSync(cmdShimPath); + let result; + try { + result = cleanupOrphanedCodexContextMonitorScript(codexHome); + } finally { + fs.unlinkSync = originalUnlinkSync; + } + + const expectedWarningCount = cmdShimExisted ? 2 : 1; + assert.strictEqual(result.warnings.length, expectedWarningCount); + assert.ok(result.warnings.some((w) => /gsd-context-monitor\.js$/.test(w.path)), + 'a warning must name the .js script'); + if (cmdShimExisted) { + assert.ok(result.warnings.some((w) => /gsd-context-monitor\.cmd$/.test(w.path)), + 'a warning must name the .cmd shim when Windows staged one'); + assert.ok(fs.existsSync(cmdShimPath), 'the .cmd shim remains on disk since its deletion failed too'); + } + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), false, + 'the already-safe hooks.json deregistration must not be reverted by a script-deletion failure'); + assert.ok(fs.existsSync(monitorScriptPath(codexHome)), 'the file remains on disk since deletion failed'); + }); + + test('isGsdOwnedCodexContextMonitorScript rejects a user file at the same path', () => { + runCodexInstall(codexHome); + fs.mkdirSync(path.join(codexHome, 'hooks'), { recursive: true }); + fs.writeFileSync(monitorScriptPath(codexHome), '#!/usr/bin/env node\nconsole.log("my own script");\n', 'utf8'); + assert.strictEqual(isGsdOwnedCodexContextMonitorScript(monitorScriptPath(codexHome)), false); + }); + + test('uninstall removes recognized registrations and the orphaned script symmetrically with install', () => { + runCodexInstall(codexHome); + seedPreExisting2586Install(codexHome); + + runCodexUninstall(codexHome); + + assert.strictEqual(fs.existsSync(monitorScriptPath(codexHome)), false); + if (fs.existsSync(hooksJsonPath(codexHome))) { + assert.strictEqual(hooksJsonReferencesCodexContextMonitor(codexHome), false); + } + }); + + test('uninstall does not throw on an unmodeled hooks.json event-value shape', () => { + runCodexInstall(codexHome); + fs.writeFileSync(hooksJsonPath(codexHome), JSON.stringify({ hooks: { Stop: 'not-an-array' } }, null, 2) + '\n', 'utf8'); + assert.doesNotThrow(() => runCodexUninstall(codexHome)); + }); + + test('property: reconcileCodexHooksJsonEvent never removes a non-managed-shape command', () => { + const { reconcileCodexHooksJsonEvent } = require('../gsd-core/bin/lib/runtime-hooks-surface.cjs'); + fc.assert( + fc.property( + fc.array(fc.string({ minLength: 1, maxLength: 40 }).filter((s) => !/gsd-context-monitor|gsd-check-update/.test(s)), { minLength: 1, maxLength: 5 }), + (customCommands) => { + const home = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-codex-2586-prop-')); + try { + const seeded = { hooks: { Stop: [{ hooks: customCommands.map((c) => ({ type: 'command', command: c })) }] } }; + fs.writeFileSync(path.join(home, 'hooks.json'), JSON.stringify(seeded, null, 2) + '\n', 'utf8'); + reconcileCodexHooksJsonEvent(home, 'Stop', { managedCommand: null }); + const after = JSON.parse(fs.readFileSync(path.join(home, 'hooks.json'), 'utf8')); + const survivingCommands = ((after.hooks && after.hooks.Stop) || []) + .flatMap((entry) => (entry.hooks || []).map((h) => h.command)); + for (const c of customCommands) { + assert.ok(survivingCommands.includes(c), `non-managed command "${c}" must survive removal`); + } + } finally { + cleanup(home); + } + }, + ), + { numRuns: 25 }, + ); + }); +}); diff --git a/tests/fixtures/install-tree/codex.json b/tests/fixtures/install-tree/codex.json index fc98e0899..863e061c5 100644 --- a/tests/fixtures/install-tree/codex.json +++ b/tests/fixtures/install-tree/codex.json @@ -539,10 +539,6 @@ "gsd-core/workflows/verify-work/steps/mvp-uat-framing.md", "hooks/gsd-check-update-worker.js", "hooks/gsd-check-update.js", - "hooks/gsd-context-monitor.js", - "hooks/lib/cli-exit.js", - "hooks/lib/exit-code-registry.js", - "hooks/lib/hook-exit.js", "hooks/managed-hooks-registry.cjs", "hooks/package.json", "scripts/changeset/README.md", diff --git a/tests/install-minimal-hooks.test.cjs b/tests/install-minimal-hooks.test.cjs index 60f7b9409..e02515023 100644 --- a/tests/install-minimal-hooks.test.cjs +++ b/tests/install-minimal-hooks.test.cjs @@ -2996,31 +2996,21 @@ describe('#4087 regression: Codex install stages the hook helpers its hooks requ return path.join(configDir, 'hooks'); } - test('the installed context-monitor hook LOADS AND RUNS, not merely exists', () => { + test('#2586: gsd-context-monitor.js is no longer staged for Codex at all', () => { + // Was: "the installed context-monitor hook LOADS AND RUNS, not merely + // exists" — that row's premise (Codex ships this hook) is exactly what + // #2586 removes: its only documented metrics source is Claude's own + // statusline hook, which Codex never installs, so every registered Codex + // event was a guaranteed silent no-op. The #4087 bug class this describe + // block guards (a staged hook requiring an unshipped hooks/lib/ helper) + // remains covered live via Windsurf's own guards — see the + // "#4087 review: Windsurf install..." describe block below, unaffected + // by this change. const hooksDir = installCodex(tmpDir); - const hook = path.join(hooksDir, 'gsd-context-monitor.js'); - assert.ok(fs.existsSync(hook), 'precondition: the hook itself must be staged'); - - // The actual defect. Before the fix this exited 1 with - // "Cannot find module './lib/hook-exit.js'". - // `exitCode`, not `status`: the process seam returns its own shape - // ({outcome, exitCode, stdout, stderr, ...}) and `status` reads undefined — - // which would compare unequal to 0 and pass this row for the wrong reason - // if the polarity were ever flipped. - const result = runNode([hook], { timeoutMs: 30000, input: '{}', env: { ...process.env } }); - assert.strictEqual( - result.outcome, 'exited', - `the hook must run to completion, not time out or be killed. outcome=${result.outcome}`, - ); - assert.strictEqual( - result.exitCode, 0, - 'the installed Codex hook must load and exit 0 — a MODULE_NOT_FOUND at load fires on every ' - + `registered event and is invisible to the installer's own exit code. stderr: ${result.stderr}`, - ); - assert.doesNotMatch( - String(result.stderr || ''), /MODULE_NOT_FOUND|Cannot find module/, - 'no missing-module error may reach stderr', - ); + assert.strictEqual(fs.existsSync(path.join(hooksDir, 'gsd-context-monitor.js')), false, + 'gsd-context-monitor.js must not be staged for Codex post-#2586'); + assert.strictEqual(fs.existsSync(path.join(hooksDir, 'lib')), false, + 'hooks/lib/ must not exist at all — nothing else Codex stages requires a lib/ helper'); }); // AC4: this is the row that stops the bug recurring. It derives the @@ -3052,11 +3042,17 @@ describe('#4087 regression: Codex install stages the hook helpers its hooks requ if (!/\.(js|cjs)$/.test(entry)) continue; scan(fs.readFileSync(full, 'utf8'), seedRe); } - assert.ok( - required.size > 0, - 'precondition: at least one staged Codex hook must require a ./lib/ helper — if this ever ' - + 'goes to zero the bundle changed and this row silently stops testing anything', - ); + // #2586: gsd-context-monitor.js was the only staged Codex hook requiring + // a ./lib/ helper; it is no longer staged for Codex at all, so the + // dependency closure is correctly empty. This row still proves the + // GRAMMAR holds (whatever IS required must be staged) — it is just that + // "whatever is required" is now the empty set for Codex specifically. + // The non-trivial case (closure size > 0) is covered live by the + // "#4087 review: Windsurf install..." describe block below. + assert.strictEqual(required.size, 0, + 'no staged Codex hook should require a ./lib/ helper post-#2586 — if this becomes non-zero, ' + + 'extend this row (do not just raise the bar back to ">0") so the new dependency stays proven'); + assert.strictEqual(fs.existsSync(libDir), false, 'hooks/lib/ must not exist when nothing requires it'); // Walk to a fixed point, exactly as the installer must. const checked = new Set(); @@ -3161,7 +3157,12 @@ describe('#4087 regression: Codex install stages the hook helpers its hooks requ const hooksDir = installCodex(tmpDir); const libDir = path.join(hooksDir, 'lib'); const stagedLibs = fs.existsSync(libDir) ? fs.readdirSync(libDir).sort() : []; - assert.ok(stagedLibs.length > 0, 'precondition: some helper must have been staged'); + // #2586: Codex's closure is now legitimately empty (gsd-context-monitor.js, + // the only staged Codex hook that ever required a helper, is no longer + // staged) — the deepStrictEqual below is still the real assertion and + // holds for the empty case too; the non-empty case remains covered live + // by the "#4087 review: Windsurf install..." describe block below. + assert.strictEqual(stagedLibs.length, 0, 'no helpers should be staged for Codex post-#2586'); // Derive the closure independently of the installer. const seedRe = /require\(\s*['"]\.\/lib\/([A-Za-z0-9._-]+)['"]\s*\)/g; From 66e4034fe4967e4b2fdd523f113f8f2a67708660 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 07:03:39 -0400 Subject: [PATCH 011/166] fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4138): failing-first regression — begin-phase without --phase must fail closed * fix(#4138): begin-phase without --phase exits non-zero and writes nothing * chore(#4138): changeset fragment for begin-phase arg validation * chore(#4138): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/agile-badgers-roar.md | 5 + src/state.cts | 28 ++- tests/state.test.cjs | 343 +++++++++++++++++++++++++++++++ 3 files changed, 374 insertions(+), 2 deletions(-) create mode 100644 .changeset/agile-badgers-roar.md diff --git a/.changeset/agile-badgers-roar.md b/.changeset/agile-badgers-roar.md new file mode 100644 index 000000000..7cc52c6da --- /dev/null +++ b/.changeset/agile-badgers-roar.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4380 +--- +**`gsd-tools state begin-phase` without `--phase` now exits non-zero and writes nothing** — previously a missing, empty, or flag-shaped phase argument was silently accepted and wrote a null-phase STATE.md (removing `current_phase`/`current_phase_name` from frontmatter and serialising the literal `Phase null` into three body locations), and took a milestone claim for the phase "null". (#4138) diff --git a/src/state.cts b/src/state.cts index b4ecedde0..e55f8105a 100644 --- a/src/state.cts +++ b/src/state.cts @@ -5032,7 +5032,27 @@ function cmdStateJson(cwd: string, raw: boolean): void { * and synchronizes frontmatter via writeStateMd. * Fixes: #1102 (plan counts), #1103 (status/last_activity), #1104 (body text). */ -function cmdStateBeginPhase(cwd: string, phaseNumber: string | number, phaseName: string | null | undefined, planCount: number | null | undefined, raw: boolean): void { +function cmdStateBeginPhase(cwd: string, phaseNumber: string | number | null | undefined, phaseName: string | null | undefined, planCount: number | null | undefined, raw: boolean): void { + // #4138: `--phase` is this verb's one required argument, and an invocation + // that names no phase must fail closed BEFORE any read-modify-write runs — + // previously the missing flag flowed through as null and the transition + // serialised `String(null)` into the body (`Phase: null — EXECUTING`, + // `Status: Executing Phase null`, `last_activity_desc: Phase null execution + // started`) while the post-sync frontmatter rebuild dropped current_phase / + // current_phase_name entirely, so a single argument-less call un-set the + // phase identity. The guard mirrors the sibling usage errors that already + // exit non-zero (`state update`'s "field and value required", the router's + // "unexpected positional argument" / "Invalid --plans value"), NOT + // `cmdStateMilestoneSwitch`'s `output({error})` form, which exits 0 — the + // issue's Expected is explicit: "Exit non-zero with a usage message and + // write nothing." Empty and whitespace-only values are the same missing + // argument (CONTRIBUTING.md CLI matrix); a flag-shaped `--phase --name x` + // resolves to null in parseNamedArgs and lands here too. Runs before the + // STATE.md existence check so argument validation always precedes I/O, and + // before claimMilestonePhase so no phase-"null" milestone claim is taken. + if (phaseNumber == null || String(phaseNumber).trim() === '') { + error('phase required (--phase )'); + } const statePath = planningPaths(cwd).state; if (!fs.existsSync(statePath)) { output({ error: 'STATE.md not found' }, raw, undefined); @@ -5047,7 +5067,11 @@ function cmdStateBeginPhase(cwd: string, phaseNumber: string | number, phaseName // #1230 post-sync preservation, and the no-op write guard. const intent: StateTransitionIntent = { kind: 'beginPhase', - phaseNumber, + // The guard above made this non-null/non-empty; `error` is never-returning + // at runtime but this module's destructured io binding does not narrow CFA, + // so the narrowed fact is restated once (cmdStateUpdate's `field as string` + // idiom, state.cts:782). + phaseNumber: phaseNumber as string | number, phaseName: phaseName ?? null, planCount: planCount ?? null, }; diff --git a/tests/state.test.cjs b/tests/state.test.cjs index 8c843ed89..64e7b2c3c 100644 --- a/tests/state.test.cjs +++ b/tests/state.test.cjs @@ -20279,3 +20279,346 @@ describe('#3957 (epic #3473 B9): no-op decline reports the real condition', () = }); }); }); + +// ═════════════════════════════════════════════════════════════════════════ +// #4138: `state begin-phase` writes a null-phase STATE.md when the required +// --phase argument is missing (.gsd/bug/fix-4138-begin-phase-arg-validation/ +// {10-diagnosis,50-test-matrix}.md). The verb's token validation rejects an +// unexpected positional but lets a MISSING --phase through as null, which the +// transition then serialises as the literal string "null" into three body +// locations while the post-sync frontmatter rebuild drops current_phase / +// current_phase_name entirely. The contract under test is the issue's Expected: +// exit non-zero with a usage message and write nothing. +// ═════════════════════════════════════════════════════════════════════════ + +describe('#4138: state begin-phase guards its required --phase argument', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = createFixture(); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + // The issue's clean shape: a milestone-bound ROADMAP, phase dirs on disk + // (1-2 verification-passed), and a STATE.md whose progress block already + // carries the correct counters — the exact state a null-phase write or a + // counter zeroing would destroy. + function seedBoundedProject() { + const roadmap = [ + '# Roadmap', + '', + '## Milestone v1.0: Test Milestone', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 1. | 1/1 | Complete | 2026-01-01 |', + '| 2. | 1/1 | Complete | 2026-01-02 |', + '| 3. | 0/1 | Not Started | |', + '| 4. | 0/1 | Not Started | |', + '', + '### Phase 1: Alpha', + '**Goal:** first', + '', + '### Phase 2: Beta', + '**Goal:** second', + '', + '### Phase 3: Gamma', + '**Goal:** third', + '', + '### Phase 4: Delta', + '**Goal:** fourth', + ].join('\n'); + fs.writeFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), roadmap); + ['01-alpha', '02-beta', '03-gamma', '04-delta'].forEach((dirName, idx) => { + const n = idx + 1; + const padded = String(n).padStart(2, '0'); + const phaseDir = path.join(tmpDir, '.planning', 'phases', dirName); + fs.mkdirSync(phaseDir, { recursive: true }); + fs.writeFileSync(path.join(phaseDir, `${padded}-01-PLAN.md`), '# Plan\n'); + if (n <= 2) { + fs.writeFileSync(path.join(phaseDir, `${padded}-01-SUMMARY.md`), '# Summary\n'); + writePassedVerification(tmpDir, dirName, padded); + } + }); + writeState( + tmpDir, + [ + '---', + "gsd_state_version: '1.0'", + 'milestone: v1.0', + 'milestone_name: Test Milestone', + 'status: executing', + 'current_phase: 2', + 'current_phase_name: Beta', + 'progress:', + ' total_phases: 4', + ' completed_phases: 2', + ' total_plans: 4', + ' completed_plans: 2', + ' percent: 50', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 2 (Beta) — COMPLETE', + 'Plan: 1 of 1', + 'Status: Phase 2 complete', + 'Last activity: 2026-01-02 — Phase 2 execution complete', + '', + '## Progress', + '', + 'Progress: [█████▓▓▓▓▓] 50% (2/4 phases complete)', + '', + '## Session Continuity', + '', + 'Last session: 2026-01-02T10:00:00.000Z', + '', + ].join('\n'), + ); + return fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'); + } + + // #4138 row 1 — the failing-first regression. A begin-phase that names no + // phase must exit non-zero with a usage message and leave STATE.md + // byte-identical: no "Phase null" prose, no current_phase removal, no + // state.json publication, no milestone claim for the literal phase "null". + test('beginPhaseWithoutPhaseFailsClosedAndWritesNothing', () => { + const before = seedBoundedProject(); + const statePath = path.join(tmpDir, '.planning', 'STATE.md'); + + const result = runGsdTools(['state', 'begin-phase'], tmpDir); + + assert.strictEqual(result.success, false, '#4138: a missing --phase must fail the command'); + assert.notStrictEqual(result.exitCode, 0, '#4138: a usage error must exit non-zero'); + assert.match(result.error, /--phase/, `the usage message must name the required flag; got: ${result.error}`); + assert.strictEqual( + fs.readFileSync(statePath, 'utf8'), + before, + '#4138: an invalid invocation must not write STATE.md (no null-phase serialisation, no current_phase removal)', + ); + assert.strictEqual( + fs.existsSync(path.join(tmpDir, '.planning', 'state.json')), + false, + '#4138: an invalid invocation must not publish the state contract', + ); + + // Row 12: no milestone claim leaked for the literal phase "null" — the + // next VALID begin-phase must report no conflict (#3311 claim point). + const valid = runGsdTools(['state', 'begin-phase', '--phase', '3', '--name', 'gamma'], tmpDir); + assert.ok(valid.success, `follow-up valid begin-phase failed: ${valid.error}`); + const validOut = JSON.parse(valid.output); + assert.strictEqual(validOut.milestone_conflict, null, '#4138: the errored call must not have claimed phase "null"'); + }); + + // #4138 row 2 — a flag-shaped `--phase` value resolves to null in + // parseNamedArgs (command-arg-projection.cjs: "a value flag whose next token + // is absent or starts with `--` yields null"); the guard must catch it the + // same way as an absent flag. + test('beginPhaseFlagShapedPhaseValueFailsClosed', () => { + const before = seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '--name', 'gamma'], tmpDir); + + assert.strictEqual(result.success, false, '#4138: a flag-shaped --phase value is a missing phase'); + assert.notStrictEqual(result.exitCode, 0, '#4138: usage error must exit non-zero'); + assert.match(result.error, /--phase/, `the usage message must name the required flag; got: ${result.error}`); + assert.strictEqual( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + before, + '#4138: STATE.md must stay byte-identical', + ); + }); + + // #4138 row 3 — empty string. CONTRIBUTING's CLI matrix requires `--phase ""` + // as a distinct negative case; pre-fix it wrote `Status: Executing Phase ` + // (trailing space) and claimed the empty phase. + test('beginPhaseEmptyPhaseValueFailsClosed', () => { + const before = seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', ''], tmpDir); + + assert.strictEqual(result.success, false, '#4138: an empty --phase is a missing phase'); + assert.notStrictEqual(result.exitCode, 0, '#4138: usage error must exit non-zero'); + assert.match(result.error, /--phase/, `the usage message must name the required flag; got: ${result.error}`); + assert.strictEqual( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + before, + '#4138: STATE.md must stay byte-identical', + ); + }); + + // #4138 row 4 — whitespace-only value is the same missing argument. + test('beginPhaseWhitespacePhaseValueFailsClosed', () => { + const before = seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', ' '], tmpDir); + + assert.strictEqual(result.success, false, '#4138: a whitespace-only --phase is a missing phase'); + assert.notStrictEqual(result.exitCode, 0, '#4138: usage error must exit non-zero'); + assert.strictEqual( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + before, + '#4138: STATE.md must stay byte-identical', + ); + }); + + // #4138 row 5 (not-the-bug pin): --name is OPTIONAL. A begin-phase that + // names a phase but no slug must keep succeeding exactly as today — the + // issue's Expected guards only the no-PHASE direction. + test('beginPhaseWithoutNameStillSucceeds', () => { + seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '3'], tmpDir); + + assert.ok(result.success, `begin-phase --phase 3 must succeed without --name: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.phase, '3'); + assert.strictEqual(out.phase_name, null, 'no --name means a null name, which is legitimate'); + const fm = frontmatterLib.extractFrontmatter( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + ); + assert.strictEqual(String(fm.current_phase), '3', 'phase identity must land in frontmatter'); + }); + + // #4138 row 6 (boundary, not-the-bug pin): the FIRST legitimate phase of a + // fresh milestone must keep initializing STATE.md exactly as today — zero + // counters on a first-phase begin are legitimate, and the guard must not + // have made the verb refuse its own happy path. + test('beginPhaseFirstLegitimatePhaseStillInitializes', () => { + seedBoundedProject(); + // Reset to a fresh-milestone STATE.md: no completed phases, phase 1 beginning. + writeState( + tmpDir, + [ + '---', + "gsd_state_version: '1.0'", + 'milestone: v1.0', + 'milestone_name: Test Milestone', + 'status: planning', + 'progress:', + ' total_phases: 4', + ' completed_phases: 0', + ' total_plans: 4', + ' completed_plans: 0', + ' percent: 0', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 1 (Alpha) — READY TO EXECUTE', + 'Plan: 0 of ?', + 'Status: Ready to execute Phase 1', + 'Last activity: 2026-01-01 — roadmap created', + '', + ].join('\n'), + ); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '1', '--name', 'alpha', '--plans', '1'], tmpDir); + + assert.ok(result.success, `first-phase begin must succeed: ${result.error}`); + const after = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'); + const fm = frontmatterLib.extractFrontmatter(after); + assert.strictEqual(String(fm.current_phase), '1', 'current_phase must be initialized'); + assert.strictEqual(fm.status, 'executing', 'status must flip to executing'); + assert.strictEqual(fm.last_activity_desc, 'Phase 1 execution started'); + const position = stateDocument.stateExtractField(after, 'Phase'); + assert.match(position ?? '', /^1 \(alpha\) — EXECUTING/, `Current Position Phase line must carry the phase identity; got: ${position}`); + }); + + // #4138 row 7 (Defect 2 pin): a CORRECT invocation must never zero + // progress.completed_phases / progress.percent — the issue's table shows + // 5/33 becoming 0/0 on gsd-core 1.12.0. On next the #4359 write-path + // ratchet keeps the curated completed counters; this row pins that contract + // on the begin-phase verb so the zeroing class cannot return unnoticed. + test('beginPhasePreservesSuppliedCompletedPhasesAndPercent', () => { + seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '3', '--name', 'gamma'], tmpDir); + + assert.ok(result.success, `begin-phase failed: ${result.error}`); + const after = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'); + const fm = frontmatterLib.extractFrontmatter(after); + assert.ok(fm.progress, 'progress block must survive the write'); + assert.strictEqual( + Number(fm.progress.completed_phases), + 2, + `#4138: completed_phases must stay at the stored 2 (phases 1-2 verification-passed), got ${fm.progress.completed_phases}`, + ); + assert.strictEqual( + Number(fm.progress.percent), + 50, + `#4138: percent must stay coherent with the preserved counters (2/4), got ${fm.progress.percent}`, + ); + assert.strictEqual(String(fm.current_phase), '3', 'the begun phase must land in frontmatter'); + }); + + // #4138 row 8 (Defect 2 pin, resume variant): the #3127 resume branch must + // preserve the counters identically — a wave-continue begin is still an + // ordinary correct usage. + test('beginPhaseResumePreservesSuppliedCounters', () => { + seedBoundedProject(); + // Body already executing phase 3 → the #3127 resume branch fires. + writeState( + tmpDir, + [ + '---', + "gsd_state_version: '1.0'", + 'milestone: v1.0', + 'milestone_name: Test Milestone', + 'status: executing', + 'current_phase: 3', + 'current_phase_name: Gamma', + 'progress:', + ' total_phases: 4', + ' completed_phases: 2', + ' total_plans: 4', + ' completed_plans: 2', + ' percent: 50', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 3 (Gamma) — EXECUTING', + 'Plan: 1 of 1', + 'Status: Executing Phase 3', + 'Last activity: 2026-01-02 — Phase 3 execution started', + '', + ].join('\n'), + ); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '3', '--name', 'gamma'], tmpDir); + + assert.ok(result.success, `resume begin-phase failed: ${result.error}`); + const fm = frontmatterLib.extractFrontmatter( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + ); + assert.strictEqual(Number(fm.progress.completed_phases), 2, '#4138: resume must not zero completed_phases'); + assert.strictEqual(Number(fm.progress.percent), 50, '#4138: resume must not zero percent'); + }); + + // #4138 row 9 (existing-guard pin): the sibling --plans validation already + // exits non-zero without writing; pinned so the two required-argument + // guards stay symmetric. + test('beginPhaseNonNumericPlansStillFailsClosed', () => { + const before = seedBoundedProject(); + + const result = runGsdTools(['state', 'begin-phase', '--phase', '3', '--plans', 'abc'], tmpDir); + + assert.strictEqual(result.success, false, 'a non-numeric --plans must fail the command'); + assert.notStrictEqual(result.exitCode, 0, 'usage error must exit non-zero'); + assert.strictEqual( + fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf8'), + before, + 'STATE.md must stay byte-identical', + ); + }); +}); From b7406b293fc469ab17b3df3947490a0b39697233 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 08:06:39 -0400 Subject: [PATCH 012/166] enhance(#2618): render pending todos as one bounded bullet per todo (#4384) --- .changeset/mellow-koalas-caper.md | 5 + docs/COMMANDS.md | 2 + gsd-core/templates/state.md | 8 +- gsd-core/workflows/add-todo.md | 5 +- gsd-core/workflows/check-todos.md | 6 +- scripts/lint-test-file-count.allowlist.json | 1 + src/init.cts | 137 +++++++++- tests/state-todos-render.test.cjs | 270 ++++++++++++++++++++ 8 files changed, 424 insertions(+), 10 deletions(-) create mode 100644 .changeset/mellow-koalas-caper.md create mode 100644 tests/state-todos-render.test.cjs diff --git a/.changeset/mellow-koalas-caper.md b/.changeset/mellow-koalas-caper.md new file mode 100644 index 000000000..0e348e855 --- /dev/null +++ b/.changeset/mellow-koalas-caper.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 4384 +--- +**Pending todos now render as one bounded bullet per todo in STATE.md.** Each capture used to append to a single run-on sentence in "### Pending Todos", growing unbounded and wrecking `git diff` readability; captures now produce one bullet per todo, capped at 240 characters, with a fail-safe refresh that leaves the section untouched on a malformed lookup. (#2618) diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index 4f4f41b5b..d45e1346e 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -1938,6 +1938,8 @@ Capture ideas, tasks, notes, and seeds to their appropriate destination. Default **Produces:** `.planning/todos/` (default), note files (--note), ROADMAP.md backlog section (--backlog), `.planning/seeds/SEED-NNN-slug.md` (--seed) +**STATE.md rendering:** each capture (or `--list` action that changes the pending count) refreshes STATE.md's "### Pending Todos" section to one bullet per pending todo, each capped at 240 characters — `- [date] [area] title — [todo file](path) — Needs ...`. A todo with no clear next step omits the "Needs ..." clause rather than the bullet. Refresh is fail-safe: a failed or malformed lookup leaves the existing section untouched rather than clearing it. + ```bash /gsd-capture "Consider adding dark mode support" # Add todo /gsd-capture --note "Caching strategy idea" # Quick note diff --git a/gsd-core/templates/state.md b/gsd-core/templates/state.md index 09a15ad1f..32afd9a90 100644 --- a/gsd-core/templates/state.md +++ b/gsd-core/templates/state.md @@ -172,9 +172,11 @@ Updated after each plan completion. **Decisions:** Reference to PROJECT.md Key Decisions table, plus recent decisions summary for quick access. Full decision log lives in PROJECT.md. **Pending Todos:** Ideas captured via /gsd-add-todo -- Count of pending todos -- Reference to .planning/todos/pending/ -- Brief list if few, count if many (e.g., "5 pending todos — see /gsd:capture --list") +- One bullet per pending todo, rendered by `init.todos`'s `pending_todos_markdown` + (each bullet capped at 240 characters: `- [date] [area] title — [todo file](path) — Needs ...`) +- `None yet.` when there are no pending todos +- No collapse-by-count fallback — every pending todo gets its own line, always + (see #2618 design doc for why a "count if many" fallback was rejected) **Blockers/Concerns:** From "Next Phase Readiness" sections - Issues that affect future work diff --git a/gsd-core/workflows/add-todo.md b/gsd-core/workflows/add-todo.md index bc378b7ea..db1045b6e 100644 --- a/gsd-core/workflows/add-todo.md +++ b/gsd-core/workflows/add-todo.md @@ -145,8 +145,9 @@ files: If `.planning/STATE.md` exists: -1. Use `todo_count` from init context (or re-run `init todos` if count changed) -2. Update "### Pending Todos" under "## Accumulated Context" +1. Re-run `gsd_run query init.todos` to get a fresh, post-write JSON snapshot (the count and rendering must reflect the change just made). +2. **Fail-safe check (#2618):** if the JSON failed to parse, or `pending_read_ok` is not `true`, or `pending_todos_markdown` is not a string, do NOT touch the "### Pending Todos" section — leave it exactly as-is and continue to the next step. A partial or malformed `init.todos` result must never overwrite a good existing section. +3. Otherwise, replace the entire body of "### Pending Todos" (under "## Accumulated Context") with the literal value of `pending_todos_markdown` — verbatim, one bullet per pending todo already rendered and length-capped by `init.todos`. Do not reformat, re-wrap, re-order, or hand-edit the bullets; do not append to the old body — replace it wholesale (this is what makes an old run-on-sentence section get superseded cleanly with no migration step). diff --git a/gsd-core/workflows/check-todos.md b/gsd-core/workflows/check-todos.md index 335d46ff9..99a95bcf9 100644 --- a/gsd-core/workflows/check-todos.md +++ b/gsd-core/workflows/check-todos.md @@ -150,9 +150,11 @@ Return to list_todos step. -After any action that changes todo count: +If `.planning/STATE.md` exists: -Re-run `init todos` to get updated count, then update STATE.md "### Pending Todos" section if exists. +1. Re-run `gsd_run query init.todos` to get a fresh, post-write JSON snapshot (the count and rendering must reflect the change just made). +2. **Fail-safe check (#2618):** if the JSON failed to parse, or `pending_read_ok` is not `true`, or `pending_todos_markdown` is not a string, do NOT touch the "### Pending Todos" section — leave it exactly as-is and continue to the next step. A partial or malformed `init.todos` result must never overwrite a good existing section. +3. Otherwise, replace the entire body of "### Pending Todos" (under "## Accumulated Context") with the literal value of `pending_todos_markdown` — verbatim, one bullet per pending todo already rendered and length-capped by `init.todos`. Do not reformat, re-wrap, re-order, or hand-edit the bullets; do not append to the old body — replace it wholesale (this is what makes an old run-on-sentence section get superseded cleanly with no migration step). diff --git a/scripts/lint-test-file-count.allowlist.json b/scripts/lint-test-file-count.allowlist.json index b5332bc6c..f81d44ed5 100644 --- a/scripts/lint-test-file-count.allowlist.json +++ b/scripts/lint-test-file-count.allowlist.json @@ -123,6 +123,7 @@ "state-prune.test.cjs", "state-rebuild-cli.test.cjs", "state-rebuild.test.cjs", + "state-todos-render.test.cjs", "state-write-path-drift-guard.test.cjs", "state.test.cjs" ], diff --git a/src/init.cts b/src/init.cts index 52a37db13..0192d76ce 100644 --- a/src/init.cts +++ b/src/init.cts @@ -12,6 +12,7 @@ import os from 'node:os'; import { execGit, platformWriteSync, platformReadSync, toNativePath, posixNormalize } from './shell-command-projection.cjs'; import { realClock } from './clock.cjs'; import { escapeRegex } from './pattern.cjs'; +import { collectSection } from './markdown-sectionizer.cjs'; // eslint-disable-next-line @typescript-eslint/no-require-imports -- io.cjs is an export= CommonJS module import io = require('./io.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports -- config-loader.cjs is an export= CommonJS module @@ -2233,15 +2234,118 @@ function cmdInitPhaseOp(cwd: string, phase: string, raw: boolean): void { output(withProjectRoot(cwd, result), raw); } +// #2618: bullet-cap and title-floor for renderPendingTodosMarkdown below. +// 240 matches the bound already vetted by maintainer review on the prior +// attempt at this issue (PR #2662) — re-deriving a different number would be +// pure bikeshedding, not a correctness improvement. See +// .gsd/phase/feat-2618-compact-todo-pointers/40-design.md. +const PENDING_TODO_BULLET_MAX_CHARS = 240; +const PENDING_TODO_TITLE_FLOOR = 15; +const PENDING_TODO_AREA_FLOOR = 3; + +function sanitizePendingTodoInline(value: string): string { + // Defensive: the regex captures that populate title/area/needs can only + // ever match a single line, so this is belt-and-suspenders against any + // future non-regex-sourced input, not a reachable case today. + return value.replace(/[\r\n]+/g, ' ').trim(); +} + +function truncatePendingTodoText(value: string, maxLen: number): string { + if (value.length <= maxLen) return value; + if (maxLen <= 1) return value.slice(0, Math.max(0, maxLen)); + return `${value.slice(0, maxLen - 1)}…`; +} + +/** + * #2618: pure renderer for STATE.md's "### Pending Todos" section BODY (not + * the heading). One bullet per todo, each capped at + * PENDING_TODO_BULLET_MAX_CHARS. `gsd-core/workflows/add-todo.md` and + * `check-todos.md` splice this string in verbatim instead of free-hand + * editing STATE.md — see the design doc for why this is real, unit-tested + * code rather than a prose algorithm (DEFECT.GENERATIVE-FIX: a prose + * algorithm duplicated as a test oracle is exactly the divergence class + * this avoids). + */ +function renderPendingTodosMarkdown(todos: Record[]): string { + if (!Array.isArray(todos) || todos.length === 0) { + return 'None yet.'; + } + return todos.map((todo) => renderPendingTodoBullet(todo)).join('\n'); +} + +function pendingTodoFieldAsString(value: unknown, fallback: string): string { + return typeof value === 'string' && value.length > 0 ? value : fallback; +} + +function renderPendingTodoBullet(todo: Record): string { + const date = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['created'], 'unknown')); + let area = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['area'], 'general')); + let title = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['title'], 'Untitled')); + // Strip trailing "." so the fixed "Needs ....` template below never + // produces a doubled period when the source text already ended in one. + let needs = + typeof todo['needs'] === 'string' + ? sanitizePendingTodoInline(todo['needs']).replace(/\.+$/, '') + : ''; + const link = `[todo file](${pendingTodoFieldAsString(todo['path'], '')})`; + + const assemble = (): string => { + const needsClause = needs ? ` — Needs ${needs}.` : ''; + return `- [${date}] [${area}] ${title} — ${link}${needsClause}`; + }; + + let line = assemble(); + if (line.length <= PENDING_TODO_BULLET_MAX_CHARS) return line; + + // 1) Drop the needs clause entirely first — date/area/title/link untouched. + needs = ''; + line = assemble(); + if (line.length <= PENDING_TODO_BULLET_MAX_CHARS) return line; + + // 2) Shorten the title next, down to a floor — date/area/link untouched. + const titleOverage = line.length - PENDING_TODO_BULLET_MAX_CHARS; + const targetTitleLen = Math.max(PENDING_TODO_TITLE_FLOOR, title.length - titleOverage); + if (targetTitleLen < title.length) { + title = truncatePendingTodoText(title, targetTitleLen); + line = assemble(); + } + if (line.length <= PENDING_TODO_BULLET_MAX_CHARS) return line; + + // 3) Shorten area as a last resort — date and the markdown link are never + // altered (link correctness > strict cap; see design doc "Known limits"). + const areaOverage = line.length - PENDING_TODO_BULLET_MAX_CHARS; + const targetAreaLen = Math.max(PENDING_TODO_AREA_FLOOR, area.length - areaOverage); + if (targetAreaLen < area.length) { + area = truncatePendingTodoText(area, targetAreaLen); + line = assemble(); + } + + return line; +} + function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void { const config = loadConfig(cwd); const pendingDir = path.join(planningDir(cwd), 'todos', 'pending'); let count = 0; const todos: Record[] = []; + // #2618: distinct from "genuinely zero pending todos" — false only when + // readdirSync itself failed for a reason OTHER than the directory simply + // not existing yet (ENOENT), mirroring the ENOENT-vs-other-errno split + // already used above in this file (#3885, ADR-3473 §8.5). Without this, + // a real I/O/permission error on the pending dir would look identical to + // "no pending todos" and could wipe an existing, non-empty Pending Todos + // section in STATE.md on refresh — the fail-safe requirement for #2618. + let pendingReadOk = true; try { - const files = fs.readdirSync(pendingDir).filter((f) => f.endsWith('.md')); + // #2618: sorted so pending_todos_markdown's bullet order is stable across + // runs — readdirSync's order is filesystem-dependent, not contractually + // stable, and an unstable order would reorder every bullet on an + // unrelated re-render, turning a one-line git diff into a full-section + // rewrite (must-have #3). Filenames are `YYYY-MM-DD-slug.md`, so this + // also yields a sensible chronological order as a side effect. + const files = fs.readdirSync(pendingDir).filter((f) => f.endsWith('.md')).sort(); for (const file of files) { const content = platformReadSync(path.join(pendingDir, file)); if (content === null) continue; @@ -2252,6 +2356,21 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void // #2337: kept in parity with cmdListTodos — surface severity when // present, omit the key entirely for todos with no severity line. const severityMatch = content.match(/^severity:\s*(.+)$/m); + // #2618: first non-empty line of the `## Solution` body, used as the + // bullet's "Needs ..." clause. "TBD" (the create_file template's own + // placeholder for an unresolved solution) renders no clause at all + // rather than the useless literal "Needs TBD.". + const solutionSection = collectSection(content, (h) => h.level === 2 && h.text.trim() === 'Solution'); + let needs: string | undefined; + if (solutionSection) { + const firstLine = solutionSection.body + .split('\n') + .map((l) => l.trim()) + .find((l) => l.length > 0); + if (firstLine && firstLine.toUpperCase() !== 'TBD') { + needs = firstLine; + } + } const todoArea = areaMatch ? areaMatch[1].trim() : 'general'; if (area && todoArea !== area) continue; @@ -2265,13 +2384,17 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void // #2376: absolute — see comment on phase_dir in cmdInitExecutePhase. path: toPosixPath(path.join(planningDir(cwd), 'todos', 'pending', file)), ...(severityMatch ? { severity: severityMatch[1].trim() } : {}), + ...(needs ? { needs } : {}), }); } catch { /* intentionally empty */ } } - } catch { - /* intentionally empty */ + } catch (err) { + const code = (err as NodeJS.ErrnoException)?.code; + if (code !== 'ENOENT') { + pendingReadOk = false; + } } const result: Record = { @@ -2291,6 +2414,13 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void planning_exists: fs.existsSync(planningDir(cwd)), todos_dir_exists: fs.existsSync(path.join(planningDir(cwd), 'todos')), pending_dir_exists: fs.existsSync(path.join(planningDir(cwd), 'todos', 'pending')), + + // #2618: see PENDING_TODO_BULLET_MAX_CHARS comment / design doc. Consumed + // by add-todo.md / check-todos.md's update_state step; omitted entirely + // (rather than emitted with possibly-wrong data) when pendingReadOk is + // false, so the workflow's fail-safe check can key off field presence. + pending_read_ok: pendingReadOk, + ...(pendingReadOk ? { pending_todos_markdown: renderPendingTodosMarkdown(todos) } : {}), }; output(withProjectRoot(cwd, result), raw); @@ -4188,4 +4318,5 @@ export = { cmdAgentSkills, buildSkillManifest, cmdSkillManifest, + renderPendingTodosMarkdown, }; diff --git a/tests/state-todos-render.test.cjs b/tests/state-todos-render.test.cjs new file mode 100644 index 000000000..5574f0e39 --- /dev/null +++ b/tests/state-todos-render.test.cjs @@ -0,0 +1,270 @@ +'use strict'; + +// #2618: deterministic renderer for STATE.md's "### Pending Todos" section +// body. See .gsd/phase/feat-2618-compact-todo-pointers/40-design.md and +// 50-test-matrix.md for the full rationale and case list. + +const test = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const os = require('node:os'); +const path = require('node:path'); +const fc = require('fast-check'); + +const initLib = require('../gsd-core/bin/lib/init.cjs'); +const { renderPendingTodosMarkdown } = initLib; +const { cleanup } = require('./helpers.cjs'); +const { escapeRegex } = require('../gsd-core/bin/lib/pattern.cjs'); + +const MAX = 240; + +function makeTodo(overrides) { + return Object.assign( + { + created: '2026-09-01', + area: 'api', + title: 'Fix retry logic', + path: '.planning/todos/pending/2026-09-01-fix-retry-logic.md', + }, + overrides, + ); +} + +test('renderPendingTodosMarkdown: zero todos renders exactly "None yet."', () => { + assert.equal(renderPendingTodosMarkdown([]), 'None yet.'); +}); + +test('renderPendingTodosMarkdown: one todo without a needs clause', () => { + const body = renderPendingTodosMarkdown([makeTodo()]); + const lines = body.split('\n'); + assert.equal(lines.length, 1); + assert.match(lines[0], /^- \[2026-09-01\] \[api\] Fix retry logic — \[todo file\]\(.+\)$/); + assert.doesNotMatch(lines[0], /Needs/); +}); + +test('renderPendingTodosMarkdown: "TBD" solution omits the needs clause', () => { + const body = renderPendingTodosMarkdown([makeTodo({ needs: undefined })]); + assert.doesNotMatch(body, /Needs/); +}); + +test('renderPendingTodosMarkdown: real solution text becomes a "Needs ..." clause', () => { + const body = renderPendingTodosMarkdown([makeTodo({ needs: 'define retry behavior' })]); + assert.match(body, /Needs define retry behavior\.$/); +}); + +test('renderPendingTodosMarkdown: needs text already ending in "." does not double the period', () => { + const body = renderPendingTodosMarkdown([makeTodo({ needs: 'Add a max-attempts cap.' })]); + assert.match(body, /Needs Add a max-attempts cap\.$/); + assert.doesNotMatch(body, /\.\.$/); +}); + +test('renderPendingTodosMarkdown: many todos render one bullet per todo, in order', () => { + const todos = Array.from({ length: 5 }, (_, i) => + makeTodo({ title: `Todo number ${i}`, path: `.planning/todos/pending/todo-${i}.md` }), + ); + const body = renderPendingTodosMarkdown(todos); + const lines = body.split('\n'); + assert.equal(lines.length, 5); + lines.forEach((line, i) => { + assert.match(line, new RegExp(`Todo number ${i} `)); + }); +}); + +test('renderPendingTodosMarkdown: markdown-special characters in title/area render verbatim', () => { + const body = renderPendingTodosMarkdown([ + makeTodo({ title: 'Fix `parse()` for [bracketed] input', area: 'db|cache' }), + ]); + assert.match(body, /Fix `parse\(\)` for \[bracketed\] input/); + assert.match(body, /\[db\|cache\]/); +}); + +test('renderPendingTodosMarkdown: boundary — 239 chars (limit-1) is not truncated', () => { + // Assemble a title so the full line lands at exactly 239 chars, then assert + // byte-identical (untruncated) output. + const base = makeTodo({ needs: undefined, title: 'X' }); + const probe = renderPendingTodosMarkdown([base]); + const pad = 239 - probe.length; + assert.ok(pad > 0, 'test fixture sanity: base line must be shorter than 239 chars'); + const title = 'X'.repeat(1 + pad); + const todo = makeTodo({ needs: undefined, title }); + const line = renderPendingTodosMarkdown([todo]); + assert.equal(line.length, 239); + assert.equal(line, `- [2026-09-01] [api] ${title} — [todo file](${todo.path})`); +}); + +test('renderPendingTodosMarkdown: boundary — 240 chars (limit) is not truncated', () => { + const base = makeTodo({ needs: undefined, title: 'X' }); + const probe = renderPendingTodosMarkdown([base]); + const pad = 240 - probe.length; + const title = 'X'.repeat(1 + pad); + const todo = makeTodo({ needs: undefined, title }); + const line = renderPendingTodosMarkdown([todo]); + assert.equal(line.length, 240); + assert.equal(line, `- [2026-09-01] [api] ${title} — [todo file](${todo.path})`); +}); + +test('renderPendingTodosMarkdown: boundary — 241 chars (limit+1) truncates, needs dropped first', () => { + const base = makeTodo({ needs: 'x', title: 'X' }); + const probe = renderPendingTodosMarkdown([base]); + const pad = 241 - probe.length; + const title = 'X'.repeat(1 + Math.max(pad, 0)); + const todo = makeTodo({ needs: 'a real needs clause that should get dropped', title }); + const line = renderPendingTodosMarkdown([todo]); + assert.ok(line.length <= MAX, `expected <= ${MAX}, got ${line.length}`); + assert.doesNotMatch(line, /Needs/, 'needs clause must be dropped before title is touched'); +}); + +test('renderPendingTodosMarkdown: title truncated with floor + ellipsis when needs-drop is insufficient', () => { + const longTitle = 'A'.repeat(400); + const todo = makeTodo({ title: longTitle, needs: 'something' }); + const line = renderPendingTodosMarkdown([todo]); + assert.ok(line.length <= MAX, `expected <= ${MAX}, got ${line.length}`); + assert.doesNotMatch(line, /Needs/); + assert.match(line, /…/); + // date/area/link remain byte-identical to the untruncated assembly. + assert.match(line, /^- \[2026-09-01\] \[api\] /); + assert.match(line, new RegExp(`\\[todo file\\]\\(${escapeRegex(todo.path)}\\)$`)); +}); + +test('renderPendingTodosMarkdown: pathological — title floor + link alone exceeds cap is allowed to overflow', () => { + const veryLongPath = `.planning/todos/pending/${'p'.repeat(400)}.md`; + const todo = makeTodo({ title: 'A'.repeat(400), path: veryLongPath, needs: 'x' }); + const line = renderPendingTodosMarkdown([todo]); + // Documented known limit: link is never sacrificed even if it blows the cap. + assert.ok(line.includes(veryLongPath), 'link must remain verbatim even when cap is exceeded'); +}); + +test('property: rendered body always has one line per todo, each line <= 240 chars, link preserved verbatim', () => { + // Path length is bounded so the algorithm's floor assembly (date + 3-char + // area floor + 15-char title floor + link, needs always droppable) never + // exceeds 240 — i.e. the cap is always achievable without touching the + // link. This keeps the <= 240 assertion a real, provable invariant rather + // than a vacuous one; the pathological "cap not achievable" case is + // covered separately by the fixed "pathological" unit test above. + const todoArb = fc.record({ + created: fc.constantFrom('2026-01-01', '2025-12-31', 'unknown'), + area: fc.string({ minLength: 0, maxLength: 40 }), + title: fc.string({ minLength: 0, maxLength: 500 }), + path: fc + .string({ minLength: 1, maxLength: 150 }) + .map((s) => `.planning/todos/pending/${s}.md`), + needs: fc.option(fc.string({ minLength: 0, maxLength: 300 }), { nil: undefined }), + }); + + fc.assert( + fc.property(fc.array(todoArb, { minLength: 0, maxLength: 20 }), (todos) => { + const body = renderPendingTodosMarkdown(todos); + if (todos.length === 0) { + return body === 'None yet.'; + } + const lines = body.split('\n'); + if (lines.length !== todos.length) return false; + return lines.every((line, i) => { + const link = `[todo file](${todos[i].path})`; + return line.length <= 240 && line.includes(link); + }); + }), + ); +}); + +// ─── pending_read_ok / pending_todos_markdown via the real CLI surface ───── + +const { spawnSync } = require('node:child_process'); + +function runQueryInitTodos(cwd) { + const gsdTools = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs'); + const result = spawnSync(process.execPath, [gsdTools, 'query', 'init.todos'], { + cwd, + encoding: 'utf8', + timeout: 15000, + }); + assert.equal(result.status, 0, `gsd_run query init.todos failed: ${result.stderr}`); + return JSON.parse(result.stdout); +} + +test('cmdInitTodos: pending_read_ok is true and pending_todos_markdown present for a healthy empty dir', (t) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2618-')); + t.after(() => cleanup(dir)); + fs.mkdirSync(path.join(dir, '.planning', 'todos', 'pending'), { recursive: true }); + + const json = runQueryInitTodos(dir); + assert.equal(json.pending_read_ok, true); + assert.equal(json.pending_todos_markdown, 'None yet.'); + assert.equal(json.todo_count, 0); +}); + +test('cmdInitTodos: real todo file produces a rendered bullet via the CLI', (t) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2618-')); + t.after(() => cleanup(dir)); + const pendingDir = path.join(dir, '.planning', 'todos', 'pending'); + fs.mkdirSync(pendingDir, { recursive: true }); + fs.writeFileSync( + path.join(pendingDir, '2026-09-01-fix-retry-logic.md'), + [ + '---', + 'created: 2026-09-01T00:00:00.000Z', + 'title: Fix retry logic', + 'area: api', + 'severity: major', + 'files:', + ' - src/api/client.cts:42', + '---', + '', + '## Problem', + '', + 'Retries are unbounded.', + '', + '## Solution', + '', + 'Add a max-attempts cap.', + '', + ].join('\n'), + ); + + const json = runQueryInitTodos(dir); + assert.equal(json.pending_read_ok, true); + assert.equal(json.todo_count, 1); + assert.match(json.pending_todos_markdown, /Fix retry logic/); + assert.match(json.pending_todos_markdown, /Needs Add a max-attempts cap\.$/m); +}); + +test('cmdInitTodos: bullet order is filename-sorted regardless of write/insertion order', (t) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2618-')); + t.after(() => cleanup(dir)); + const pendingDir = path.join(dir, '.planning', 'todos', 'pending'); + fs.mkdirSync(pendingDir, { recursive: true }); + // Written in reverse-alphabetical insertion order on purpose — readdirSync + // order is filesystem-dependent, not contractually stable, so the ONLY way + // to assert a deterministic bullet order is to sort explicitly. See #2618 + // spec review finding on git-diff stability (must-have #3). + const todoBody = (title) => + ['---', 'created: 2026-09-01', `title: ${title}`, 'area: api', '---', ''].join('\n'); + fs.writeFileSync(path.join(pendingDir, '2026-09-03-third.md'), todoBody('Third todo')); + fs.writeFileSync(path.join(pendingDir, '2026-09-01-first.md'), todoBody('First todo')); + fs.writeFileSync(path.join(pendingDir, '2026-09-02-second.md'), todoBody('Second todo')); + + const json = runQueryInitTodos(dir); + const titles = json.todos.map((t2) => t2.title); + assert.deepEqual(titles, ['First todo', 'Second todo', 'Third todo']); + const lines = json.pending_todos_markdown.split('\n'); + assert.ok(lines[0].includes('First todo')); + assert.ok(lines[1].includes('Second todo')); + assert.ok(lines[2].includes('Third todo')); +}); + +// ─── workflow parity guard (DEFECT.GENERATIVE-FIX) ───────────────────────── + +function extractUpdateStateStep(workflowPath) { + const content = fs.readFileSync(workflowPath, 'utf8'); + const match = content.match(/([\s\S]{0,20000}?)<\/step>/); + assert.ok(match, `update_state step not found in ${workflowPath}`); + return match[1].trim(); +} + +test('workflow parity: add-todo.md and check-todos.md update_state steps are byte-identical', () => { + const addTodoPath = path.join(__dirname, '..', 'gsd-core', 'workflows', 'add-todo.md'); + const checkTodosPath = path.join(__dirname, '..', 'gsd-core', 'workflows', 'check-todos.md'); + const a = extractUpdateStateStep(addTodoPath); + const b = extractUpdateStateStep(checkTodosPath); + assert.equal(a, b, 'update_state step prose must be identical across both workflows (DEFECT.GENERATIVE-FIX)'); +}); From fd4aac567059c042fd3bb570762ad7598aea9451 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 10:17:50 -0400 Subject: [PATCH 013/166] fix(#4192): honor explicit model pins on the claude runtime (#4396) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4192): honor explicit model pins on the claude runtime Two documented model-configuration contracts did not hold on the claude runtime (confirmed-bug scope from the issue triage): Finding 1 — model_profile_overrides.claude. was inert. Step 3 of resolveModelInternal gated runtime-aware tier resolution on configRuntime !== 'claude', so the key's only reader was never consulted, while workflows/settings-advanced.md writes it for claude-runtime users. A new step 4.5 resolves ONLY the user's override entry (never the builtin claude tier map, so unpinned installs keep resolving aliases). An override value that maps to a current tier alias collapses to that alias (byte-equivalent, the #2041 protection); anything else — a pinned older generation, a bare alias repoint, a non-Anthropic id — resolves verbatim. It sits after the resolve_model_ids:'omit' gate so an explicit project omit still wins (#2297) and before the alias return so resolve_model_ids:true cannot re-materialize the pin to the latest id. Finding 2 — fully-qualified claude-* ids in model_overrides were warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable branch, #2041), while the docs promise any fully-qualified model id is valid. The unmappable branch now passes the pin through verbatim with a warn-once breadcrumb (text describes the pass-through). Dropping it silently unpinned the operator's explicit choice — the exact 'profile can misrepresent what actually runs' defect of #4192. Mappable ids and non-claude values behave exactly as before; resolveModelForTier shares the mapping; the tier honesty signal is unchanged (raw ids still report 'unknown'); the model_policy path is untouched. Docs updated to the agreed contract (CONFIGURATION.md false 'Claude example' corrected; how-to + shipped reference document the pin semantics, the fable alias, and the tier-override composition). * test(#4192): pin explicit model pin resolution on the claude runtime 28 failing-first rows across the resolver seam and the resolve-model CLI: pinned-generation fidelity (tier override + per-agent verbatim pins, object form, explicit runtime), unpinned controls byte-stable (no override, other runtime/tier, inherit, project omit, precedence), adversarial rows (prototype-chain keys, malformed values, warn-once dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through runGsdTools. The stale #2041 fall-through assertions now pin the pass-through contract; mappable-id collapse assertions unchanged. * chore(#4192): add changeset fragment * chore(#4192): backfill PR number in changeset fragment --------- Co-authored-by: ZCode --- .changeset/kind-sloths-swim.md | 5 + docs/CONFIGURATION.md | 17 +- docs/how-to/configure-model-profiles.md | 10 +- gsd-core/references/model-profiles.md | 15 +- src/model-resolver.cts | 113 +++++++- tests/commands.test.cjs | 32 +++ tests/model-resolver.test.cjs | 368 +++++++++++++++++++++++- 7 files changed, 528 insertions(+), 32 deletions(-) create mode 100644 .changeset/kind-sloths-swim.md diff --git a/.changeset/kind-sloths-swim.md b/.changeset/kind-sloths-swim.md new file mode 100644 index 000000000..5ff41a8b2 --- /dev/null +++ b/.changeset/kind-sloths-swim.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4396 +--- +**Explicit model pins now hold on the Claude runtime** — set `model_profile_overrides.claude.` (e.g. pin the opus tier to `claude-opus-4-7`) and the resolver silently returned the bare tier alias anyway, and a fully-qualified Claude model ID in `model_overrides` was warn-dropped to tier resolution even though the configuration docs promise any fully-qualified model ID is valid; both are now resolved as configured (values naming the current tier default still collapse to their alias, so nothing changes for unpinned installs), and the docs now state the claude-runtime pin contract including the `fable` alias. (#4192) diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index 02e7f5700..9b098536f 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -1564,7 +1564,9 @@ Override specific agents without changing the entire profile: } ``` -Valid override values: `opus`, `sonnet`, `haiku`, `inherit`, or any fully-qualified model ID (e.g., `"openai/o3"`, `"google/gemini-2.5-pro"`). +Valid override values: `opus`, `sonnet`, `haiku`, `fable`, `inherit`, or any fully-qualified model ID (e.g., `"openai/o3"`, `"google/gemini-2.5-pro"`). + +On the Claude runtime, fully-qualified Claude model IDs are honored as explicit generation pins (#4192): an ID that names the current tier default (e.g. `"claude-sonnet-5"`) collapses to its tier alias — the same model in the form Claude Code's Agent tool always accepts — while any other ID (e.g. `"claude-opus-4-7"`) is resolved verbatim, so the pinned generation is what `resolve-model` reports. GSD emits a warn-once stderr breadcrumb for verbatim pins, because Claude Code setups whose Agent tool accepts only tier aliases will not honor a full ID. `fable` is a Claude Code Agent-tool alias, not a GSD profile tier: it is valid in `model_overrides` but has no column in the profile table. `model_overrides` can be set in either `.planning/config.json` (per-project) or `~/.gsd/defaults.json` (global). Per-project entries win on conflict and @@ -2095,15 +2097,20 @@ When `runtime` is set, profile tiers (`opus`/`sonnet`/`haiku`) resolve to runtim This resolves `gsd-planner` → `gpt-5.6-sol` (xhigh), `gsd-executor` → `gpt-5.6-terra` (medium), `gsd-codebase-mapper` → `gpt-5.6-luna` (medium). Codex skills pass each resolved `model` and `reasoning_effort` to `spawn_agent` when its visible schema advertises the corresponding field; otherwise they omit the field and inherit the session/static agent configuration. -**Claude example** — explicit opt-in resolves to full Claude IDs (no `resolve_model_ids: true` needed): +**Claude example** — pin a tier's generation without giving up tier-based profiles (#4192): ```json { "runtime": "claude", - "model_profile": "quality" + "model_profile": "quality", + "model_profile_overrides": { + "claude": { "opus": "claude-opus-4-7" } + } } ``` +On the Claude runtime, tier resolution stays on Claude Code's adaptive tier aliases (`opus` / `sonnet` / `haiku`) unless you override a tier. An override value that names the current tier default collapses back to its alias (the same model in the always-accepted form); any other value — a pinned older generation such as `claude-opus-4-7`, a bare alias repointing the tier, or a non-Anthropic model id — is resolved verbatim, so `resolve-model` reports exactly what the profile pins. Setting `runtime: "claude"` alone (no overrides) changes nothing: aliases resolve exactly as they do with the key absent. `resolve_model_ids: true` remains the global switch for materializing full IDs on every agent. + **Per-runtime overrides** — replace one or more tier defaults: ```json @@ -2122,8 +2129,8 @@ This resolves `gsd-planner` → `gpt-5.6-sol` (xhigh), `gsd-executor` → `gpt-5 **Precedence (highest to lowest):** 1. `model_overrides[]` — explicit per-agent ID always wins. -2. **Runtime-aware tier resolution** (this section) — when `runtime` is set and profile is not `inherit`. -3. `resolve_model_ids: "omit"` — returns empty string when no `runtime` is set. +2. **Runtime-aware tier resolution** (this section) — when `runtime` is set and profile is not `inherit`. On non-Claude runtimes this is the built-in tier map merged with your `model_profile_overrides`; on the Claude runtime it applies only the `model_profile_overrides.claude.` entry you set (#4192) — never the built-in defaults, so unpinned installs keep resolving aliases. +3. `resolve_model_ids: "omit"` — returns empty string when no `runtime` is set (an explicit project-level `"omit"` wins over a `claude` tier override too). 4. Claude-native default — `model_profile` tier as alias (current default). 5. `inherit` — propagates literal `inherit` for `Task(model="inherit")` semantics. diff --git a/docs/how-to/configure-model-profiles.md b/docs/how-to/configure-model-profiles.md index 5e10da375..fc720a256 100644 --- a/docs/how-to/configure-model-profiles.md +++ b/docs/how-to/configure-model-profiles.md @@ -51,7 +51,9 @@ If a single agent needs a different tier without changing the whole profile, use } ``` -Valid values: `opus`, `sonnet`, `haiku`, `inherit`, or any fully-qualified model ID (e.g. `"openai/o3"`, `"google/gemini-2.5-pro"`). +Valid values: `opus`, `sonnet`, `haiku`, `fable`, `inherit`, or any fully-qualified model ID (e.g. `"openai/o3"`, `"google/gemini-2.5-pro"`). + +On the Claude runtime, fully-qualified Claude model IDs act as explicit generation pins (#4192): an ID naming the current tier default (e.g. `"claude-sonnet-5"`) resolves to its tier alias — the same model in the form Claude Code's Agent tool always accepts — while any other ID (e.g. `"claude-opus-4-7"`) resolves verbatim, with a warn-once stderr note that setups accepting only tier aliases will not honor a full ID. `fable` is a Claude Code Agent-tool alias, not a GSD profile tier: valid here, but it has no column in the profile table. To pin a generation for a whole tier instead of one agent, use `model_profile_overrides` (see below). `model_overrides` can be set per-project in `.planning/config.json` or globally in `~/.gsd/defaults.json`. Per-project entries win on conflict; non-conflicting global entries are preserved. @@ -303,8 +305,9 @@ When multiple layers apply, the resolver picks the highest-priority entry: 1. model_overrides[] — per-agent; full IDs; targeted exception 2. dynamic_routing.tier_models[] — when enabled; escalates on soft failure 3. models[] — coarse phase-level tier -4. model_profile (per-agent column) — global tier strategy -5. Runtime default — when nothing else applies +4. model_profile_overrides.. — per-tier model override (#4192: honored on the claude runtime too) +5. model_profile (per-agent column) — global tier strategy +6. Runtime default — when nothing else applies ``` --- @@ -317,6 +320,7 @@ When multiple layers apply, the resolver picks the highest-priority entry: | Coarse phase-level tuning ("Opus for planning") | `models.` | | Per-agent precision ("force Haiku on the codebase mapper") | `model_overrides[]` | | A fully-qualified model ID for a specific agent | `model_overrides[]: "openai/gpt-5"` | +| Pin a tier's generation on Claude Code (e.g. executor stays on Opus 4.7) | `model_profile_overrides.claude.: "claude-opus-4-7"` | | Start cheap, escalate only on failure | `dynamic_routing` | | All agents follow the session model (non-Anthropic provider) | `model_profile: "inherit"` | diff --git a/gsd-core/references/model-profiles.md b/gsd-core/references/model-profiles.md index 9743ac5bc..e6034872e 100644 --- a/gsd-core/references/model-profiles.md +++ b/gsd-core/references/model-profiles.md @@ -58,6 +58,10 @@ Model profiles control which Claude model each GSD agent uses. This allows balan 3. **Profile table** — the per-agent column from the active `model_profile` 4. **Runtime default** — when nothing else applies +Steps 2–4 select the *tier*; a `model_profile_overrides..` entry then +maps that tier to a concrete model (#4192 — honored on the claude runtime as well, so +pinning composes with tiering instead of replacing it). + ### Why two layers above the profile? - **Profile** is a global tier strategy (everyone runs balanced). @@ -215,8 +219,11 @@ is (highest → lowest): (see §Dynamic Routing — escalation steps tier up per attempt counter) 4. If no dynamic_routing match, check models[phase_type] for a phase-type tier (see §Per-Phase-Type Model Map for the agent → phase-type mapping) -5. If no phase-type slot, look up agent in profile table -6. Pass model parameter to Task call +5. Check model_profile_overrides.. for a per-tier model override + (honored on the claude runtime too — #4192; verbatim unless it maps to the + current tier alias) +6. If no phase-type slot, look up agent in profile table +7. Pass model parameter to Task call ``` `model` and `effort` resolve through different mechanisms at different @@ -246,7 +253,9 @@ Override specific agents without changing the entire profile: } ``` -Overrides take precedence over the profile. Valid values: `opus`, `sonnet`, `haiku`, `inherit`, or any fully-qualified model ID (e.g., `"o3"`, `"openai/o3"`, `"google/gemini-2.5-pro"`). +Overrides take precedence over the profile. Valid values: `opus`, `sonnet`, `haiku`, `fable`, `inherit`, or any fully-qualified model ID (e.g., `"o3"`, `"openai/o3"`, `"google/gemini-2.5-pro"`). `fable` is a Claude Code Agent-tool alias, not a GSD profile tier — it has no column in the profile table above. + +On the Claude runtime, fully-qualified Claude model IDs are honored as explicit generation pins (#4192): an ID that names the current tier default (e.g. `"claude-sonnet-5"`) resolves to its tier alias — the same model in the form the Agent tool always accepts — while any other ID (e.g. `"claude-opus-4-7"`) resolves verbatim, with a warn-once stderr note that setups accepting only tier aliases will not honor a full ID. To pin a generation for a whole tier rather than one agent, set `model_profile_overrides.claude.` (see docs/CONFIGURATION.md — Runtime-Aware Profiles). ## Switching Profiles diff --git a/src/model-resolver.cts b/src/model-resolver.cts index 680bde926..939b55707 100644 --- a/src/model-resolver.cts +++ b/src/model-resolver.cts @@ -157,6 +157,65 @@ function _resolveRuntimeTier(config: Record, tier: string): Tie }); } +/** + * #4192 — Resolve the claude-runtime TIER OVERRIDE model for (config, tier). + * + * Step 3's runtime-aware resolution deliberately skips the claude runtime to + * preserve the alias-native posture (#1156/#2297): with no user override, the + * resolver must keep returning bare tier aliases, and the builtin claude tier + * map (`opus → claude-opus-4-8`, …) must never force full-ID emission on every + * default install. But `model_profile_overrides..` is a + * documented override point (docs/CONFIGURATION.md § Runtime-Aware Profiles) + * that `workflows/settings-advanced.md` actively writes for claude-runtime + * users — and #4192 Finding 1 measured the key inert on this runtime. + * + * This helper reads ONLY the user's override entry for the effective claude + * runtime and tier — never the builtin claude tier map — so an install with no + * override is byte-identical to before the fix. The runtime is resolved the + * same way steps 1-3 resolve it (config['runtime'], defaulting to 'claude'), + * NOT via resolveActiveRuntime (GSD_RUNTIME/marker): the value policy must + * key off the config the operator wrote, matching mapClaudeOverrideForRuntime. + * + * Value policy mirrors the model_overrides path (#2041/#4192): an override + * value that maps to a current tier alias collapses to that alias + * (byte-equivalent resolution, alias-form emission); anything else — a pinned + * older generation (`claude-opus-4-7`), a bare alias/tier repoint (`sonnet`), + * or a non-Claude vendor id (`openai/o3`) — is emitted verbatim as pinned. + * Malformed entries (no usable `model` string) return null so the caller falls + * through to normal alias resolution (ADR-443 D1: invalid values fall through). + */ +function resolveClaudeTierOverrideModel( + configRuntime: string | null | undefined, + tier: string | null | undefined, + overrides: Record | null | undefined, +): string | null { + if (!tier || tier === 'inherit') return null; + const effectiveRuntime = configRuntime || 'claude'; + if (effectiveRuntime !== 'claude') return null; // non-claude runtimes resolve at step 3 + const overridesMap = overrides as Record> | null | undefined; + if (!overridesMap || typeof overridesMap !== 'object') return null; + // Own-property guards throughout: both levels are config-supplied plain + // objects, so a prototype-chain key ("constructor", "toString") must not + // resolve an inherited member instead of falling through (same hardening as + // every other config-keyed lookup in this module). + const runtimeEntry = Object.hasOwn(overridesMap, effectiveRuntime) + ? overridesMap[effectiveRuntime] + : undefined; + if (!runtimeEntry || typeof runtimeEntry !== 'object') return null; + const userRaw = Object.hasOwn(runtimeEntry, tier) ? runtimeEntry[tier] : undefined; + if (userRaw === undefined || userRaw === null) return null; + const entry: Record = typeof userRaw === 'string' + ? { model: userRaw } + : (userRaw as Record); + if (!entry || typeof entry !== 'object') return null; + const model = entry['model']; + if (typeof model !== 'string' || model.length === 0) return null; + if (Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, model)) { + return CLAUDE_POLICY_ID_TO_ALIAS[model]; + } + return model; +} + // Reverse of the Claude tier-default IDs, plus the Fable alias which Claude // Code's Agent tool accepts but which is not a GSD model-profile tier (#1133). const CLAUDE_POLICY_ID_TO_ALIAS: Record = { @@ -189,7 +248,10 @@ function _resetModelPolicyWarningCacheForTests(): void { _modelPolicyUnmappableWarned.clear(); } -// Dedupe stderr warnings for unmappable model_overrides Claude IDs (#2041). +// Dedupe stderr warnings for unmappable model_overrides Claude IDs (#2041 / +// #4192). #2041 originally warned that such a value was being DROPPED to tier +// resolution; #4192 keeps the warn-once breadcrumb but changes the behavior to +// a verbatim pass-through, so the text now describes the pass-through. const _modelOverrideUnmappableWarned = new Set(); function warnModelOverrideUnmappable(agentType: string, overrideValue: string): void { const key = `${agentType}::${overrideValue}`; @@ -200,8 +262,9 @@ function warnModelOverrideUnmappable(agentType: string, overrideValue: string): // model's JSON result is parsed from stdout. const safe = overrideValue.length > 64 ? overrideValue.slice(0, 64) + '…' : overrideValue; process.stderr.write( - `gsd: warning — model_overrides value "${safe}" for ${agentType} ` + - `has no Claude agent alias; falling through to tier resolution.\n`, + `gsd: warning — model_overrides value "${safe}" for ${agentType} is a fully-qualified ` + + `Claude model ID with no tier alias; passing it through verbatim. Claude Code setups ` + + `whose Agent tool accepts only tier aliases will not honor it. (#4192)\n`, ); } @@ -213,12 +276,24 @@ function _resetModelOverrideWarningCacheForTests(): void { /** * #2041 — Map a `model_overrides` value to its Claude Agent-tool alias on the * claude runtime, mirroring the `model_policy` path (#1144). Claude Code's - * Agent tool `model` parameter documents only tier aliases (opus/sonnet/haiku/ - * fable); a full Claude model ID returned verbatim is silently dropped by the - * spawner. Returns the value to return verbatim, or null to signal "fall - * through to normal tier/dynamic-routing resolution" (used when a Claude full - * ID has no alias — matches model_policy's warn-and-fall-through). Non-Claude - * runtimes and non-Claude values always pass through verbatim. + * Agent tool `model` parameter documents tier aliases (opus/sonnet/haiku/ + * fable) as the always-accepted form. Returns the value to emit verbatim, or + * null to signal "fall through to normal tier/dynamic-routing resolution". + * Non-Claude runtimes and non-Claude values always pass through verbatim. + * + * #4192 — an unmappable `claude-*` value (a pinned generation that is not the + * current catalog default, e.g. `claude-opus-4-7`) is now PASSED THROUGH + * VERBATIM with a warn-once stderr breadcrumb, instead of being dropped to + * tier resolution. #2041's drop was correct when the value was plausibly a + * mis-typed current default, but for an explicit pin it silently UNPINNED the + * operator's choice — the resolver would report a tier the config never asked + * for, the exact "profile can misrepresent what actually runs" defect #4192 + * files. The documented contract ("any fully-qualified model ID", + * docs/CONFIGURATION.md § Per-Agent Overrides, + * gsd-core/references/model-profiles.md § Per-Agent Overrides) is restored: + * the pin is resolved as configured. Values that DO map to a current tier + * alias still collapse to that alias — byte-equivalent resolution, the #2041 + * protection preserved — and a mappable pin never warns. * * Hardening (code+security review): a `typeof` guard preserves the pre-fix * no-crash behavior if a malformed config surfaces a non-string value, and an @@ -243,8 +318,9 @@ function mapClaudeOverrideForRuntime( } if (CLAUDE_AGENT_ALIASES.has(override)) return override; if (override.startsWith('claude-')) { + // #4192: explicit generation pin — resolve as configured (see docblock). warnModelOverrideUnmappable(agentType, override); - return null; + return override; } return override; } @@ -507,6 +583,23 @@ function resolveModelInternal(cwd: string, agentType: string): string { return ''; } + // 4.5. Claude-runtime tier override (#4192 Finding 1). Sits AFTER the omit + // gate so an explicit project `resolve_model_ids:"omit"` still wins (#2297: + // explicit project omit is honored regardless of runtime), and BEFORE the + // alias return so a pinned generation is not re-collapsed to a tier alias or + // re-materialized to the LATEST catalog id by step 5's + // `resolve_model_ids:true` path. Fires ONLY when the user wrote a + // `model_profile_overrides.claude.` entry for this tier — see + // resolveClaudeTierOverrideModel for why the builtin map stays out. + if (tier && tier !== 'inherit') { + const claudeOverrideModel = resolveClaudeTierOverrideModel( + configRuntime, + tier, + config['model_profile_overrides'] as Record | null | undefined, + ); + if (claudeOverrideModel !== null) return claudeOverrideModel; + } + // 5. Profile lookup (Claude-native default). if (!agentModels) { return profile === 'quality' ? 'opus' diff --git a/tests/commands.test.cjs b/tests/commands.test.cjs index a299cdc1a..76a5e18fb 100644 --- a/tests/commands.test.cjs +++ b/tests/commands.test.cjs @@ -1415,6 +1415,38 @@ describe('resolve-model command', () => { assert.ok(output.model, 'should resolve a model'); }); + // #4192 (AC1, behavioral): a claude-runtime tier override must change the + // resolved output relative to the no-override control — the CLI surface the + // orchestrator reads is where the documented contract is observable. + test('claude-runtime tier override changes resolved output vs control (#4192)', () => { + fs.writeFileSync(path.join(tmpDir, '.planning', 'config.json'), JSON.stringify({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + })); + const pinned = runGsdTools('resolve-model gsd-planner', tmpDir); + assert.ok(pinned.success, `Command failed: ${pinned.error}`); + assert.strictEqual(JSON.parse(pinned.output).model, 'claude-opus-4-7'); + + fs.writeFileSync(path.join(tmpDir, '.planning', 'config.json'), JSON.stringify({ + model_profile: 'balanced', + })); + const control = runGsdTools('resolve-model gsd-planner', tmpDir); + assert.ok(control.success, `Command failed: ${control.error}`); + assert.strictEqual(JSON.parse(control.output).model, 'opus'); + }); + + // #4192 (AC2, behavioral): a per-agent fully-qualified Claude model ID is + // resolved as configured through the CLI surface. + test('per-agent fully-qualified claude ID resolves as configured (#4192)', () => { + fs.writeFileSync(path.join(tmpDir, '.planning', 'config.json'), JSON.stringify({ + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + })); + const result = runGsdTools('resolve-model gsd-debugger', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + assert.strictEqual(JSON.parse(result.output).model, 'claude-opus-4-7'); + }); + // #443: resolve-model now emits unified `effort` instead of `reasoning_effort`. // reasoning_effort was flavor-text (resolved but consumed by nobody); effort is // the wired, config-driven universal effort string for all runtimes. diff --git a/tests/model-resolver.test.cjs b/tests/model-resolver.test.cjs index 9e09dc9ee..98bb569fa 100644 --- a/tests/model-resolver.test.cjs +++ b/tests/model-resolver.test.cjs @@ -4123,16 +4123,19 @@ describe('#2041 model_overrides: Claude full ID → alias on claude runtime', () assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'claude-sonnet-5'); }); - // AC5: unmappable Claude full ID warns once + falls through to tier alias - test('model_overrides unmappable claude ID (claude-opus-4-5) falls through to tier alias on claude', () => { + // AC5 (#4192 revision): an unmappable Claude full ID — an explicit + // generation pin — is passed through VERBATIM with a warn-once breadcrumb, + // instead of being dropped to tier resolution (which silently unpinned the + // operator's explicit choice; see #4192 Finding 2). + test('model_overrides unmappable claude ID (claude-opus-4-5) passes through verbatim on claude', () => { resetRuntimeWarningCaches(); writeConfig(tmpDir, { runtime: 'claude', model_profile: 'balanced', model_overrides: { 'gsd-planner': 'claude-opus-4-5' }, }); - // gsd-planner balanced → opus tier; claude-opus-4-5 has no alias → warn + fall through → 'opus' - assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + // gsd-planner balanced → opus tier; claude-opus-4-5 has no alias → warn + verbatim pin + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'claude-opus-4-5'); }); test('model_overrides unmappable claude ID emits a stderr warning exactly once (dedupe)', () => { @@ -4173,19 +4176,19 @@ describe('#2041 model_overrides: Claude full ID → alias on claude runtime', () assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-executor', 0), 'claude-sonnet-5'); }); - // MEDIUM-1 (review): exercise the unmappable-override fall-through branch in - // resolveModelForTier (closes the mutation-score gap — a future refactor that - // accidentally returned the verbatim override instead of falling through - // would otherwise survive the suite). - test('resolveModelForTier unmappable claude ID falls through to tier alias on claude', () => { + // MEDIUM-1 (review, #4192 revision): exercise the unmappable-override + // branch in resolveModelForTier (keeps the mutation-score gap closed — the + // verbatim-pin return must survive a future refactor on the escalation path + // too, not just resolveModelInternal). + test('resolveModelForTier unmappable claude ID passes through verbatim on claude', () => { resetRuntimeWarningCaches(); writeConfig(tmpDir, { runtime: 'claude', model_profile: 'balanced', model_overrides: { 'gsd-planner': 'claude-opus-4-5' }, }); - // unmappable override → fall through → no dynamic_routing → resolveModelInternal → 'opus' - assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-planner', 0), 'opus'); + // unmappable override → no dynamic_routing → resolveModelInternal → verbatim pin + assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-planner', 0), 'claude-opus-4-5'); }); // LOW-2 (review): pin the case-sensitive contract — a case-variant like @@ -6321,3 +6324,346 @@ describe('#3007 PROPERTY: renderEffortForRuntime never renders a level the model ); }); }); + +// ─── #4192: claude-runtime generation pinning via explicit overrides ────────── +// +// Confirmed-bug scope (maintainer triage): Findings 1 and 2 — documented +// behavior the resolver does not implement on the claude runtime. +// +// F1 — model_profile_overrides.claude. was inert: step 3 of +// resolveModelInternal gated runtime-aware tier resolution on +// `configRuntime !== 'claude'`, so the only reader of the key was never +// consulted on claude, while settings-advanced.md writes it for +// claude-runtime users. +// F2 — fully-qualified claude-* IDs in model_overrides were warn-dropped to +// tier resolution (mapClaudeOverrideForRuntime unmappable branch, +// #2041), while the configuration reference and the shipped +// model-profiles reference both document "any fully-qualified model +// ID" as valid. +// +// Agreed contract (AC2, pinned here): an explicit pin is RESOLVED AS +// CONFIGURED. A claude-* value that maps to a current tier alias still +// collapses to that alias (the #2041 protection — byte-equivalent resolution); +// an unmappable one (a pinned older generation) is returned verbatim with a +// warn-once breadcrumb, because dropping it would silently unpin the operator's +// explicit choice — the exact "profile misrepresents what runs" defect of +// #4192. Unpinned resolution is byte-stable (control rows below). +describe('#4192 model_profile_overrides.claude.*: tier overrides honor pins on the claude runtime', () => { + const { createTempDir, resetRuntimeWarningCaches } = require('./helpers.cjs'); + let tmpDir; + const make = () => createTempDir('gsd-4192-tier-override-'); + const write = (cfg) => fs.writeFileSync( + path.join(tmpDir, '.planning', 'config.json'), JSON.stringify(cfg, null, 2), 'utf-8'); + + beforeEach(() => { + tmpDir = make(); + fs.mkdirSync(path.join(tmpDir, '.planning'), { recursive: true }); + resetRuntimeWarningCaches(); + }); + afterEach(() => { + cleanup(tmpDir); + resetRuntimeWarningCaches(); + }); + + // Row 1 — REGRESSION (failing-first): pinned generation honored, implicit claude runtime. + test('claude.opus = "claude-opus-4-7" pins the opus tier (implicit claude runtime)', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'claude-opus-4-7'); + }); + + // Row 2 — same with an explicit runtime key. + test('claude.opus pin honored with explicit runtime: "claude"', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'claude-opus-4-7'); + }); + + // Row 3 — mappable override collapses to the alias (form parity with #2041 step 1). + test('claude.sonnet = "claude-sonnet-5" resolves to the "sonnet" alias', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { sonnet: 'claude-sonnet-5' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'sonnet'); + }); + + // Row 4 — fable-valued override maps through the fable alias. + test('claude.opus = "claude-fable-5" resolves to "fable"', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 'claude-fable-5' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'fable'); + }); + + // Row 5 — bare-alias / tier-repoint override passes through verbatim. + test('claude.opus = "sonnet" repoints the tier at the sonnet alias', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 'sonnet' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'sonnet'); + }); + + // Row 6 — non-Claude ID override passes through verbatim (docs: any fully-qualified ID). + test('claude.haiku = "openai/gpt-4o-mini" passes through verbatim on claude', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { haiku: 'openai/gpt-4o-mini' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-codebase-mapper'), 'openai/gpt-4o-mini'); + }); + + // Row 7 — object-form override (settings workflow accepts {model, reasoning_effort}). + test('claude.opus = { model: "claude-opus-4-7" } object form pins the tier', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: { model: 'claude-opus-4-7' } } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'claude-opus-4-7'); + }); + + // Row 8 — CONTROL (AC1): no override → byte-identical alias resolution. + test('no model_profile_overrides → alias resolution unchanged', () => { + write({ model_profile: 'balanced' }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 9 — CONTROL: overrides for another runtime never apply to claude. + test('codex-only overrides are inert on the claude runtime', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { codex: { opus: 'gpt-5-pro' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 10 — CONTROL: override for a different tier than the agent's is inert for that agent. + test('claude.sonnet override does not touch an opus-tier agent', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { sonnet: 'claude-sonnet-4-6' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 11 — CONTROL: inherit profile is immune to tier overrides. + test('model_profile: "inherit" + claude.opus pin → "inherit"', () => { + write({ + model_profile: 'inherit', + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'inherit'); + }); + + // Row 12 — CONTROL: explicit project resolve_model_ids:"omit" beats the override (#2297). + test('project resolve_model_ids: "omit" + claude.opus pin → empty string', () => { + write({ + model_profile: 'balanced', + resolve_model_ids: 'omit', + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), ''); + }); + + // Row 13 — CONTROL: model_overrides still wins over the tier override. + test('model_overrides beats model_profile_overrides.claude', () => { + write({ + model_profile: 'balanced', + model_overrides: { 'gsd-planner': 'haiku' }, + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'haiku'); + }); + + // Row 15 — CONTROL: object override without a model key degrades to the alias. + test('claude.opus = { reasoning_effort } (no model) falls through to the alias', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: { reasoning_effort: 'high' } } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 16 — CONTROL: non-string/non-object value degrades to the alias. + test('claude.opus = 42 (malformed value) falls through to the alias', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: 42 } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 17 — CONTROL: empty-string value degrades to the alias. + test('claude.opus = "" falls through to the alias', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { opus: '' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 26 — ADVERSARIAL: prototype-chain keys in the override map must not leak. + test('"constructor" as a claude override key does not resolve an inherited member', () => { + write({ + model_profile: 'balanced', + model_profile_overrides: { claude: { constructor: 'claude-opus-4-7' } }, + }); + // 'constructor' is not a tier; resolution must ignore it entirely and land + // on the profile alias, never on Function.prototype's members. + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + // Row 30 — the pin wins over resolve_model_ids:true alias materialization. + test('claude.opus pin beats resolve_model_ids: true materialization', () => { + write({ + model_profile: 'balanced', + resolve_model_ids: true, + model_profile_overrides: { claude: { opus: 'claude-opus-4-7' } }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'claude-opus-4-7'); + }); +}); + +describe('#4192 model_overrides: fully-qualified claude IDs resolve as configured', () => { + const { createTempDir, resetRuntimeWarningCaches } = require('./helpers.cjs'); + let tmpDir; + const make = () => createTempDir('gsd-4192-agent-override-'); + const write = (cfg) => fs.writeFileSync( + path.join(tmpDir, '.planning', 'config.json'), JSON.stringify(cfg, null, 2), 'utf-8'); + + beforeEach(() => { + tmpDir = make(); + fs.mkdirSync(path.join(tmpDir, '.planning'), { recursive: true }); + resetRuntimeWarningCaches(); + }); + afterEach(() => { + cleanup(tmpDir); + resetRuntimeWarningCaches(); + }); + + // Row 18 — REGRESSION (failing-first): pinned generation honored, implicit claude runtime. + test('model_overrides "claude-opus-4-7" resolves verbatim (implicit claude runtime)', () => { + write({ + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-debugger'), 'claude-opus-4-7'); + }); + + // Row 18b — explicit runtime key. + test('model_overrides "claude-opus-4-7" resolves verbatim with runtime: "claude"', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-debugger'), 'claude-opus-4-7'); + }); + + // Row 21 — warn-once breadcrumb on an unmappable pin (visibility, not a drop). + test('unmappable pin emits exactly one pass-through stderr warning (dedupe)', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + }); + const writes = []; + const original = process.stderr.write.bind(process.stderr); + process.stderr.write = (chunk) => { writes.push(String(chunk)); return true; }; + try { + resolveModelInternal(tmpDir, 'gsd-debugger'); + resolveModelInternal(tmpDir, 'gsd-debugger'); // dedupe must suppress + } finally { + process.stderr.write = original; + } + const warnings = writes.filter((w) => w.includes('model_overrides') && w.includes('claude-opus-4-7')); + assert.strictEqual(warnings.length, 1, + `expected exactly one override warning, got ${warnings.length}: ${JSON.stringify(writes)}`); + // The warning must describe pass-through, not a fall-through that no longer happens. + assert.ok(!warnings[0].includes('falling through'), + `warning must not claim a fall-through: ${warnings[0]}`); + }); + + // Row 21b — no warning for a value that needs no breadcrumb (mappable / non-claude). + test('mappable ID resolution emits no model_overrides warning', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-sonnet-5' }, + }); + const writes = []; + const original = process.stderr.write.bind(process.stderr); + process.stderr.write = (chunk) => { writes.push(String(chunk)); return true; }; + try { + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-debugger'), 'sonnet'); + } finally { + process.stderr.write = original; + } + assert.strictEqual(writes.filter((w) => w.includes('model_overrides')).length, 0); + }); + + // Row 22 — escalation path parity. + test('resolveModelForTier returns the pinned generation verbatim', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + }); + assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-debugger', 0), 'claude-opus-4-7'); + }); + + // Row 24 — tier honesty signal unchanged (AC3): a raw pin carries no tier. + test('resolveTierFromConfig reports "unknown" for a raw pinned generation', () => { + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': 'claude-opus-4-7' }, + }); + assert.strictEqual(resolveTierFromConfig( + JSON.parse(fs.readFileSync(path.join(tmpDir, '.planning', 'config.json'), 'utf-8')), + 'gsd-debugger'), 'unknown'); + }); + + // Row 25 — ADVERSARIAL: prototype-chain agentType must not leak through overrides. + test('agentType "toString" against model_overrides: {} stays on the unknown-agent path', () => { + write({ + model_profile: 'balanced', + model_overrides: {}, + }); + // Unknown agent + balanced profile → the hardcoded fallback alias, never + // an inherited Function.prototype member. + assert.strictEqual(resolveModelInternal(tmpDir, 'toString'), 'sonnet'); + }); + + // Row 28 — an oversized pin value survives resolution; any warning stays capped. + test('oversized unmappable pin resolves verbatim and warning text is capped at 64 chars', () => { + const longPin = 'claude-opus-' + '9'.repeat(80); + write({ + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-debugger': longPin }, + }); + const writes = []; + const original = process.stderr.write.bind(process.stderr); + process.stderr.write = (chunk) => { writes.push(String(chunk)); return true; }; + try { + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-debugger'), longPin); + } finally { + process.stderr.write = original; + } + const warnings = writes.filter((w) => w.includes('model_overrides')); + assert.strictEqual(warnings.length, 1); + // The rendered value inside the warning is the 64-char cap + ellipsis, not the full pin. + assert.ok(!warnings[0].includes(longPin), + `warning must not contain the uncapped pin: ${warnings[0]}`); + assert.ok(warnings[0].includes('claude-opus-' + '9'.repeat(52) + '…'), + `warning must contain the capped pin render: ${warnings[0]}`); + }); +}); From 708d9a0b82a599d7cae37faafe7cc2ba95bc48e7 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 11:54:03 -0400 Subject: [PATCH 014/166] fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file (#4388) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4187): bare VERIFICATION.md regression matrix for the status surface Both query verbs must agree on every row: bare file, suffixed variants, missing file, other-dir placement, and the staleness seam. Row 1 is the failing-first regression from the issue repro. * fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file readVerificationStatus and its internal staleness check (findStaleVerificationSummary) called the shared resolver without allowBare, so a phase whose only report was a bare VERIFICATION.md read as missing and was told to re-run execute-phase while verification.resolve-file, determinePhaseStatus, and both init verification_path projectors all resolved the same file. Both call sites now pass allowBare: true, matching the other five; tier order (dashed > bare) is unchanged, so only bare-only directories change behavior. * fix(#4187): correct call-site counts in allowBare docblocks Adversarial review caught the comments claiming five of six call sites opted in; the current tree has six call sites with four previously passing allowBare — the two module-internal status-path sites were both holdouts, not one. * chore(#4187): changeset for the bare VERIFICATION.md status fix * chore(#4187): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/fierce-quails-leap.md | 5 + src/verification.cts | 39 +++++- tests/verification-status.test.cjs | 218 ++++++++++++++++++++++++++++- 3 files changed, 251 insertions(+), 11 deletions(-) create mode 100644 .changeset/fierce-quails-leap.md diff --git a/.changeset/fierce-quails-leap.md b/.changeset/fierce-quails-leap.md new file mode 100644 index 000000000..1563cfd37 --- /dev/null +++ b/.changeset/fierce-quails-leap.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4388 +--- +**`query verification.status` now resolves a bare `VERIFICATION.md` like `verification.resolve-file` does** — a phase whose only verification report was a bare `VERIFICATION.md` was reported as `missing` and told to re-run `/gsd-execute-phase` even though the report said `status: passed` and `verification.resolve-file` resolved it in the same directory. (#4187) diff --git a/src/verification.cts b/src/verification.cts index 239c89b7e..433e5cca1 100644 --- a/src/verification.cts +++ b/src/verification.cts @@ -467,12 +467,25 @@ interface ResolveVerificationFileOptions { * determinePhaseStatus and two `verification_path` projectors in * `src/init.cts`) additionally accept a BARE `VERIFICATION.md` — a form * this module's own two callers (`findStaleVerificationSummary`, - * `readVerificationStatus`) have never accepted, because a bare filename + * `readVerificationStatus`) had never accepted, because a bare filename * carries no phase token and `.endsWith('-VERIFICATION.md')` structurally * excludes it. Defaults to `false`, which is byte-for-behavior identical to * the pre-existing (non-optioned) resolver — no call-site edit required for - * the two callers in THIS module. Set `true` only from a call site whose - * pre-fix behavior already accepted a bare match. + * callers that do not want the bare tier. + * + * #4187: that historical asymmetry was drift, not contract. Six call sites + * grew around the shared resolver and four opted in + * (`cmdVerificationResolveFile`, `determinePhaseStatus`, both init + * `verification_path` projectors) — the two module-internal status-path + * call sites (`readVerificationStatus`, `findStaleVerificationSummary`) + * did not, so `query verification.resolve-file` resolved a bare report in + * a directory where `query verification.status` answered `missing` and + * recommended re-running execute-phase for an already-verified phase. + * Since #4187 those two pass `true` as well: every reader of the report + * set now recognizes the bare name. Tier order is unchanged — a dashed + * candidate (canonical or not, if it belongs to THIS phase) still outranks + * the bare match — so this only changes directories whose SOLE report is + * bare. */ allowBare?: boolean; /** @@ -720,9 +733,16 @@ function findStaleVerificationSummary( // or sentinel-numbered canonically-shaped file cannot outrank this // phase's own (possibly non-canonical) report. #3511: phaseDirName scopes // the fallback path to this same phase (see resolveVerificationFile docs). + // #4187: allowBare — this staleness seam must see the same report set the + // status reader sees, or a bare report could never read `stale` while its + // dashed twin could (two answers from one verb). const phaseDirName = path.basename(phaseDir); const phaseToken = extractPhaseToken(phaseDirName); - const verificationFile = resolveVerificationFile(phaseFiles, { phaseToken, phaseDirName }); + const verificationFile = resolveVerificationFile(phaseFiles, { + allowBare: true, + phaseToken, + phaseDirName, + }); if (!verificationFile) return { determined: true, stale: false }; const summaryFiles = (scanPhasePlans(phaseDir) as { summaryFiles: string[] }).summaryFiles @@ -766,7 +786,9 @@ function findStaleVerificationSummary( * 1. Find the phase's verification report via `resolveVerificationFile` * (canonical `-VERIFICATION.md` preferred; falls back to the * alphabetically-first `*-VERIFICATION.md` that belongs to THIS phase when - * none is canonical — #3357/#3511). If none → status 'missing'. + * none is canonical — #3357/#3511; and, when the directory's only report + * is a bare `VERIFICATION.md`, that file — #4187, matching + * `verification.resolve-file`). If none → status 'missing'. * 2. Extract `status` from FRONTMATTER ONLY via the shared extractFrontmatter * parser (DEFECT.FRONTMATTER-SCALAR-BROAD-GREP fix — parser anchors at byte 0). * If no frontmatter block or no `status` key → status 'missing'. @@ -814,7 +836,12 @@ function readVerificationStatus( // sentinel-numbered canonically-shaped file cannot outrank this phase's // own (possibly non-canonical) report. #3511: baseName also scopes the // fallback path to this same phase (see resolveVerificationFile docs). - verificationFile = resolveVerificationFile(entries, { phaseToken, phaseDirName: baseName }); + // #4187: allowBare — the status reader must recognize a bare + // `VERIFICATION.md` exactly like `verification.resolve-file`, + // `determinePhaseStatus`, and the init verification_path projectors + // already do; without it a verified phase reported `missing` and + // recommended re-running execute-phase. + verificationFile = resolveVerificationFile(entries, { allowBare: true, phaseToken, phaseDirName: baseName }); } catch { // Directory unreadable → treat as missing verificationFile = null; diff --git a/tests/verification-status.test.cjs b/tests/verification-status.test.cjs index 105b83094..6e0316df7 100644 --- a/tests/verification-status.test.cjs +++ b/tests/verification-status.test.cjs @@ -1209,11 +1209,15 @@ describe('#3357/#3492: phase-pinned *-VERIFICATION.md resolution when multiple c // commands.cts (determinePhaseStatus) and two verification_path projectors in // init.cts each hand-rolled a fourth variant of this same selection: they // additionally accept a BARE `VERIFICATION.md`, which this module's own two -// callers (findStaleVerificationSummary, readVerificationStatus) never have. -// `allowBare` threads that one behavioral difference through the single -// resolver instead of leaving a fourth hand-rolled implementation behind -// (#3473 F2). A bare match is ranked BELOW any dashed candidate — canonical -// or not — because a dashed file names its phase and a bare one does not. +// callers (findStaleVerificationSummary, readVerificationStatus) originally +// did not. `allowBare` threads that one behavioral difference through the +// single resolver instead of leaving a fourth hand-rolled implementation +// behind (#3473 F2). A bare match is ranked BELOW any dashed candidate — +// canonical or not — because a dashed file names its phase and a bare one +// does not. Since #4187 the two module-internal call sites pass `allowBare: +// true` as well (see the #4187 describe below), so every reader of the report +// set now agrees; the option's default remains `false` for callers that have +// not opted in. describe('#3473 F2: resolveVerificationFile allowBare option', () => { test('allowBare defaults to false — a bare-only list returns null without the option', () => { @@ -1265,6 +1269,210 @@ describe('#3473 F2: resolveVerificationFile allowBare option', () => { }); +// ─── #4187: a bare VERIFICATION.md is a first-class report on the status surface ── +// +// `query verification.resolve-file` (cmdVerificationResolveFile), the phase-status +// surface (commands.cts determinePhaseStatus) and both init verification_path +// projectors all resolve a BARE `VERIFICATION.md` — but readVerificationStatus +// (the reader behind `query verification.status`, isPhaseComplete, and the state/ +// roadmap completion projections) called the same shared resolver WITHOUT +// `allowBare`, so a phase whose only report was bare read as `missing` and was +// told to re-run execute-phase even when the report said `status: passed`. +// Two read-only verbs disagreed about the same directory at the same instant +// (#4187). The fix opts the status reader — and findStaleVerificationSummary, +// its internal legacy staleness check, whose only production caller is +// readVerificationStatus — into the same `allowBare: true` the other four call +// sites already pass. Tier order is unchanged: a dashed candidate (canonical or +// not) still outranks the bare name, so only bare-ONLY directories change +// behavior, from `missing` to the report's actual frontmatter status. +describe('#4187: status surface recognizes a bare VERIFICATION.md', () => { + + test('#4187 regression: bare VERIFICATION.md with status: passed reads as passed, not missing', (t) => { + const dir = mkPhaseDir('bare-passed', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'passed', 'a bare report is a report — the phase is verified'); + assert.equal(result.next_command, '', 'passed must route to no next command'); + assert.equal(result.staleCheckIndeterminate, undefined, 'the staleness check must run to completion'); + }); + + test('#4187: bare report with status: human_needed routes to verify-work', (t) => { + const dir = mkPhaseDir('bare-human', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'human_needed'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'human_needed'); + assert.equal(result.next_command, '/gsd-verify-work 99'); + }); + + test('#4187: bare report with status: gaps_found routes to plan-phase --gaps', (t) => { + const dir = mkPhaseDir('bare-gaps', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'gaps_found'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'gaps_found'); + assert.equal(result.next_command, '/gsd-plan-phase 99 --gaps'); + }); + + test('#4187 negative space: no verification file at all still reads missing with the execute-phase recommendation', (t) => { + const dir = mkPhaseDir('bare-none', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + const result = readVerificationStatus(dir); + assert.equal(result.status, 'missing', 'a genuinely absent report must stay missing'); + assert.equal(result.next_command, '/gsd-execute-phase 99', 'the missing recommendation is correct here and must not change'); + }); + + test('#4187 parity: bare + canonical dashed report — the dashed file wins on the status surface too', (t) => { + const dir = mkPhaseDir('bare-vs-dashed', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'gaps_found'); + writeVerificationMd(dir, '99-VERIFICATION.md', 'passed'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'passed', 'a dashed file names its phase and must outrank the bare match'); + }); + + test('#4187 parity: bare + cross-phase dashed stray — #3511 scoping falls through to the bare report', (t) => { + const dir = mkPhaseDir('bare-vs-stray', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + writeVerificationMd(dir, '04-VERIFICATION.md', 'gaps_found'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'passed', 'the stray belongs to phase 04; the bare file is this phase\'s only own report'); + }); + + test('#4187 parity: same-dir non-canonical dashed report only — unchanged behavior', (t) => { + const dir = mkPhaseDir('dashed-noncanon', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, '99-CORRECTION-VERIFICATION.md', 'passed'); + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'passed'); + }); + + test('#4187: directory scoping — a bare report in a DIFFERENT directory does not leak into the queried one', (t) => { + const baseDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4187-otherdir-')); + t.after(() => cleanup(baseDir)); + const queried = path.join(baseDir, '99-probe'); + const neighbor = path.join(baseDir, '98-other'); + fs.mkdirSync(queried); + fs.mkdirSync(neighbor); + writeVerificationMd(neighbor, 'VERIFICATION.md', 'passed'); + + const result = readVerificationStatus(queried); + assert.equal(result.status, 'missing', 'the neighbor directory\'s report must not answer for the queried one'); + }); + + test('#4187: bare report with no frontmatter block resolves but reads missing', (t) => { + const dir = mkPhaseDir('bare-no-fm', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + fs.writeFileSync(path.join(dir, 'VERIFICATION.md'), '# Verification\n'); + const result = readVerificationStatus(dir); + assert.equal(result.status, 'missing', 'a resolved file with no frontmatter status is the pre-existing missing path'); + }); + + test('#4187 staleness: a bare report older than its summary reads stale (legacy mtime path)', (t) => { + const dir = mkPhaseDir('bare-stale', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + setMtime(path.join(dir, 'VERIFICATION.md'), '2020-01-01T00:00:00Z'); + // Root-style summary placement — mirrors the #3492 fixture above. + const summaryPath = path.join(dir, '99-01-SUMMARY.md'); + fs.writeFileSync(summaryPath, '# summary\n'); + setMtime(summaryPath, '2021-01-01T00:00:00Z'); + + const result = readVerificationStatus(dir, { phaseCleanCommitTimesMs: () => new Map() }); + assert.equal(result.status, 'stale', 'a bare report must be staleness-checked like a dashed one'); + assert.equal(result.next_command, '/gsd-verify-work 99'); + }); + + test('#4187 unit (findStaleVerificationSummary): staleness is computed against the bare report', (t) => { + const dir = mkPhaseDir('bare-stale-unit', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + setMtime(path.join(dir, 'VERIFICATION.md'), '2020-01-01T00:00:00Z'); + const summaryPath = path.join(dir, '99-01-SUMMARY.md'); + fs.writeFileSync(summaryPath, '# summary\n'); + setMtime(summaryPath, '2021-01-01T00:00:00Z'); + + const result = findStaleVerificationSummary(dir, fs, () => new Map()); + assert.equal(result.determined, true); + assert.equal(result.stale, true); + assert.equal(result.verificationFile, 'VERIFICATION.md', 'the staleness seam must see the bare report'); + }); +}); + +// ─── #4187 CLI parity: the two query verbs must agree on the same directory ─── +// +// The issue's repro shape: run `query verification.resolve-file` and +// `query verification.status` back to back against ONE fixture directory and +// assert they never disagree about whether a report exists (and, when one does, +// about WHICH file is the report — enforced by giving the candidates distinct +// frontmatter statuses and asserting the routed status matches the expected +// winner). +describe('#4187 CLI parity: verification.resolve-file and verification.status agree', () => { + const { runGsdTools } = require('./helpers.cjs'); + + function runBothVerbs(dir) { + const resolveRes = runGsdTools(['query', 'verification.resolve-file', dir, '--raw'], dir); + const statusRes = runGsdTools(['query', 'verification.status', dir, '--raw'], dir); + assert.equal(resolveRes.success, true, `resolve-file failed: ${resolveRes.error}`); + assert.equal(statusRes.success, true, `status failed: ${statusRes.error}`); + return { resolvedPath: resolveRes.output.trimEnd(), status: JSON.parse(statusRes.output).status, nextCommand: JSON.parse(statusRes.output).next_command }; + } + + test('bare report: resolve-file finds it and status reads passed (the #4187 repro)', (t) => { + const dir = mkPhaseDir('cli-bare', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + const { resolvedPath, status, nextCommand } = runBothVerbs(dir); + assert.equal(resolvedPath, path.join(dir, 'VERIFICATION.md')); + assert.equal(status, 'passed', 'status must not report missing for a file resolve-file resolves'); + assert.equal(nextCommand, ''); + }); + + test('no report: resolve-file is empty and status is missing (agreement in the other direction)', (t) => { + const dir = mkPhaseDir('cli-none', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + const { resolvedPath, status } = runBothVerbs(dir); + assert.equal(resolvedPath, '', 'resolve-file must report no file'); + assert.equal(status, 'missing'); + }); + + test('bare + canonical dashed: both verbs pick the dashed file', (t) => { + const dir = mkPhaseDir('cli-both', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'gaps_found'); + writeVerificationMd(dir, '99-VERIFICATION.md', 'passed'); + const { resolvedPath, status } = runBothVerbs(dir); + assert.equal(resolvedPath, path.join(dir, '99-VERIFICATION.md')); + assert.equal(status, 'passed', 'status must read the same winner resolve-file picked'); + }); + + test('bare + cross-phase stray: both verbs pick the bare file', (t) => { + const dir = mkPhaseDir('cli-stray', '99-probe'); + t.after(() => cleanup(path.dirname(dir))); + + writeVerificationMd(dir, 'VERIFICATION.md', 'passed'); + writeVerificationMd(dir, '04-VERIFICATION.md', 'gaps_found'); + const { resolvedPath, status } = runBothVerbs(dir); + assert.equal(resolvedPath, path.join(dir, 'VERIFICATION.md')); + assert.equal(status, 'passed', 'status must read the same winner resolve-file picked'); + }); +}); + // ─── #3518: resolveUatFile — phase-pinned, deterministic *-UAT.md pick ─────── // // Both uat_path projectors in src/init.cts picked the phase's UAT artifact From aad04f4e9a9c55cbc8d7fd9122c6837d3c596044 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 12:24:33 -0400 Subject: [PATCH 015/166] =?UTF-8?q?docs(#4400):=20ADR-4139=20=E2=80=94=20t?= =?UTF-8?q?he=20compact-content=20seam=20(#4410)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * docs(#4400): ADR-4139 — the compact-content seam Phase 0 of epic #4139. Locks the design before any code lands. #4139's stated mechanism cannot reach the stream it exists for: 58 of 72 shipped commands deliver their whole workflow file through an eager @-include, which the host expands before any project config is in context. An in-content gate is evaluated after those bytes are already paid. The ADR declines the obvious fix (convert the 58 execution_context blocks to runtime Reads) because that removes the host guarantee for every user, not only opted-in ones — a global install shares one skill tree, so an @-include cannot be conditional. It instead keeps every @-include exactly where it is and splits what sits behind them: the canonical path becomes a runnable spine, elaborations move to a sibling detail file read at runtime. A missed Read then degrades to "runs correctly with fewer tokens", never to "runs with no instructions". Also records: the rename to workflow.compact_content, the re-pitch onto ADR-1610's context-rot argument rather than the cost argument ADR-1610 discounts, partition-not-duplication (which dissolves the dual-maintenance cost the Feature Review called disqualifying), the guard-scope and NEW_FILE_CAP mapping for the new subtree, and an argued reconciliation of the acceptance criteria this design does not meet literally. Corrects ADR-3646 §Context: it cites #3647 as open; #3647 closed 2026-09-01 as a duplicate of #3606. ADR-3646's Decision is unaffected — it explicitly disclaimed any dependence on #3647's state. The residual prose-dispatch variance named in #3647's own closure thread is unresolved, and this ADR routes around it rather than assuming it away. Refs #4139 Closes #4400 * docs(#4400): fold the two orthogonal review findings into ADR-4139 Code review (isolated context) and security review (isolated context) both returned findings. Fixed here rather than carried. Critical, from code review: the ADR repeated earlier research's claim that discuss-phase, manager and pause-work all reach a workflow by runtime Read. manager and pause-work carry plain eager @-includes and are inside the 58, not outside. discuss-phase is the only precedent, and it is one file. The Open Questions section is corrected with it. The NEW_FILE_CAP mapping was wrong in a way that changes the layout. It lives at tests/helpers/emitted-diff.cjs:96, not in workflow-size-budget, and its own doc comment records that it is a hard cap, not ack-able, and NOT tier-exemptible -- the pre-#2724 test-file version was. So a single detail.md holding plan-phase.md's elaborations is blocked outright with no exemption path. Detail content is now one or more parts under workflows//detail/, each below the cap, named by the spine in the dispatch-table shape discuss-phase.md already uses. commit-files-pathspec is in scope and earlier research called it irrelevant. Per CONTRIBUTING.md:1164-1170 it sweeps every .md under gsd-core/workflows/ for unscoped commit-seam invocations. Added to the guard table. Two byte figures were inherited rather than measured, against this ADR's own evidence note. Templates and agents re-measured; the table now carries the method and the exact numbers. From security review: "a spine that has shed a protected-content marker fails" never defined what a marker was, leaving the strongest check in the set resting on a prose-category judgment. Protection is now a literal greppable sentinel in the existing gsd: comment namespace, and the guard rule has no discretion in it. Also added: an explicit statement that the detail path is never user- or project-supplied and cannot be shadowed by a project-local file, and an exact-version pin commitment for gpt-tokenizer. Code review also found a real hole in the central fail-safe argument: spine sufficiency is verified once at split time and never again, so load-bearing procedural text carrying no sentinel could later drift into a detail part with every check green. Sufficiency is not machine-decidable, so a fifth ongoing check is added -- a spine that loses lines which reappear in its parts fails unless the PR declares the boundary move. The ADR now states plainly that this is authoring discipline with a forced checkpoint, not a structural invariant, and that the partition relocates the Feature Review's cost rather than fully eliminating it. Refs #4139 Refs #4400 --------- Co-authored-by: sim --- docs/adr/4139-compact-content-seam.md | 537 ++++++++++++++++++++++++++ docs/adr/README.md | 1 + 2 files changed, 538 insertions(+) create mode 100644 docs/adr/4139-compact-content-seam.md diff --git a/docs/adr/4139-compact-content-seam.md b/docs/adr/4139-compact-content-seam.md new file mode 100644 index 000000000..a156c1778 --- /dev/null +++ b/docs/adr/4139-compact-content-seam.md @@ -0,0 +1,537 @@ +# ADR-4139: The compact-content seam — shrink the eager window, never the guarantee + +| | | +|---|---| +| **Status** | Proposed | +| **Date** | 2026-09-06 | +| **Issue** | [#4139](https://github.com/open-gsd/gsd-core/issues/4139) | +| **Phase-0 sub-issue** | [#4400](https://github.com/open-gsd/gsd-core/issues/4400) | +| **Implementation phases** | [#4401](https://github.com/open-gsd/gsd-core/issues/4401) · [#4402](https://github.com/open-gsd/gsd-core/issues/4402) · [#4403](https://github.com/open-gsd/gsd-core/issues/4403) · [#4404](https://github.com/open-gsd/gsd-core/issues/4404) · [#4405](https://github.com/open-gsd/gsd-core/issues/4405) · [#4406](https://github.com/open-gsd/gsd-core/issues/4406) · [#4407](https://github.com/open-gsd/gsd-core/issues/4407) · [#4408](https://github.com/open-gsd/gsd-core/issues/4408) | +| **Constrained by** | [ADR-1610](1610-workflow-agent-size-budget-ratchet.md) Decision 4 · [ADR-3646](3646-per-task-content-resolution-seam.md) §Context · [ADR-3889](3889-process-exit-contract.md) · [ADR-3942](3942-emitted-drift-ack-commit-trailer.md) | +| **Corrects** | [ADR-3646](3646-per-task-content-resolution-seam.md) §Context — its citation of [#3647](https://github.com/open-gsd/gsd-core/issues/3647) as "still open" is stale; #3647 closed 2026-09-01 | + +> **Evidence note.** Every count in this ADR was measured against the tree at `origin/next` +> (`708d9a0b82`), not inferred from the issue text. Where this ADR contradicts #4139's own +> description of a mechanism, the contradiction is stated as such and the measurement is given. +> The issue is the requirement; it is not a source of truth about the tree. + +## Context + +### What #4139 asks for + +A per-project boolean that makes GSD load token-minimized variants of its own shipped prompt +content, across four token streams: (1) workflow instruction files, (2) agent-skill payloads, +(3) subagent spawn prompts, (4) planning-artifact templates. Approved 2026-09-01 with three +conditions: rename off the sports metaphor, pilot before proliferation, and an end-to-end accuracy +spot-check in the pilot. + +### The mechanism as filed does not reach the stream it exists for + +#4139 states its load mechanism as: *"one shared gate reference file checks the config key and +directs the orchestrator to Read the golfed variant. `@`-includes are untouched (they resolve +statically)."* + +The last clause is where it fails. **The top-level workflow files themselves are the +`@`-includes.** `commands/gsd/plan-phase.md` reaches its workflow this way: + +``` + +@~/.claude/gsd-core/workflows/plan-phase.md +@~/.claude/gsd-core/references/ui-brand.md + +``` + +**58 of 72** shipped `commands/gsd/*.md` carry an `@~/.claude/gsd-core/workflows/.md` line in +``, and 58 of 72 `skills/*/SKILL.md` twins carry the same. Of the 14 that do +not, exactly one — `discuss-phase` — reaches a workflow file another way: its `` +says *"Workflow files are loaded on-demand in the `` section below — not upfront"*, and its +`` block picks between three workflow files by a `config-get workflow.discuss_mode` call. +The other 13 (`graphify`, `mempalace-capture`, `mempalace-recall`, the six `ns-*` commands, +`plan-review-convergence`, `review-backlog`, `surface`, `workstreams`) carry no top-level workflow +file at all. + +`manager` and `pause-work` are **inside** the 58, not outside it — both carry a plain eager +`@`-include (`commands/gsd/manager.md:29`, `commands/gsd/pause-work.md:24`). Earlier research on +this epic grouped them with `discuss-phase` as runtime-`Read` commands; that grouping is wrong, and +it matters, because it would have made the deferred-load precedent look three times broader than it +is. **`discuss-phase` is the only precedent, and it is a single file.** + +`@` is expanded by the host as static text substitution at load, before any project config exists +in context. That is not this ADR's inference; it is stated by ADR-1610 Decision 4 (*"because +`@~/.claude/gsd-core/references/...` imports are loaded **eagerly**, moving prose into an eagerly +@-imported reference shrinks the measured file while leaving (or growing) total loaded context — +that is gaming the proxy"*), by #4139's own rejected-alternative 4, and by the triage addendum's +claim-verification pass against the host's memory documentation. + +So by the time a gate line *inside* `plan-phase.md` could be evaluated, all 98,290 bytes of +`plan-phase.md` are already in context. Reading a smaller variant afterwards adds tokens. No gate +reference file, however written, changes this — the constraint is upstream of every instruction the +file contains. + +| Content | Bytes | Reachable by an in-content gate? | +|---|--:|---| +| The 75 distinct top-level `gsd-core/workflows/*.md` files named by an eager `@` | **1,517,683** | **no** | +| `workflows//{modes,steps,templates}/*.md` (runtime `Read`) | 358,482 | yes | +| `gsd-core/templates/**` (referenced by path) | 273,589 | yes | +| `agents/*.md` via the `gsd_run query agent-skills` CLI seam | 704,846 | yes (code seam) | + +Measured at `origin/next` (`708d9a0b82`). The 58 command files name 75 distinct workflow files +between them, because several `@`-include more than one. + +The unreachable slice is the largest one — larger than the other three combined — and it is the +headline of the feature. + +### The obvious fix, and why it is not the one taken + +The obvious response is to convert those 58 `` blocks from an eager `@`-include +to a config-gated runtime `Read`. That delivers the coverage. It also inverts the fail-safe +direction: today a miss is impossible because the host substitutes the text; afterwards a `Read` +that does not fire leaves the orchestrator holding a command name, an objective paragraph, and no +procedure. + +That shape has already been examined in this repo and rejected. [ADR-3646](3646-per-task-content-resolution-seam.md) +§Context rejected a prose-dispatched content-resolution step precisely because a missed dispatch +and a legitimate fall-back are indistinguishable at the point of failure, so the executor proceeds +on the wrong content while believing it authoritative. Its Decision 2 chose a real subprocess with +a real exit code instead — *"a real process exit code the calling loop cannot fail to observe the +way it can fail to execute a prose instruction."* + +**The state of ADR-3646's cited evidence has moved, and this ADR records the correction.** ADR-3646 +§Context cites #3647 as *"filed the same day as #3646, still open"* — accurate on 2026-08-27, stale +now. #3647 was **closed 2026-09-01** as a duplicate of #3606, on the strength of PR #3687, which +fixed both call sites the report named: `execute:wave:pre` now does generic contribution dispatch, +and `execute:wave:post` now dispatches every `kind == "step"` hook generically instead of filtering +to `kind == "gate"`. ADR-3646's own reasoning anticipated this and does not depend on it — it +rejected sequencing behind #3647 explicitly, *"because sequencing behind an open reliability issue +with no committed fix date blocks #3646 indefinitely on someone else's timeline for no +architectural gain"* — so its Decision stands unchanged. Only its Context needs the footnote. + +But #3647's closure does **not** retire the concern that matters here, and reading the closure as a +clean bill of health would be a mistake. The triage diagnosis on that issue attributes the observed +1-of-4 rate to **two co-present mechanisms**, and only one of them was fixed: + +> *"`contribution`-kind hooks … are folded as natural-language text into the executor agent's own +> prompt at spawn time, never iterated by the wave:post loop at all, and whether an LLM executor +> acts on an embedded prose instruction is exactly the kind of variance that produces a 1-of-3 hit +> rate. Both the deterministic code-level exclusion (matches #3606) and prose-instruction variance +> (a separate, architectural property) can be present in the same evidence at once."* + +PR #3687 removed the deterministic filter. The prose-instruction variance was named as *a separate, +architectural property* and nothing has been shipped against it. It was never independently +measured either — it is an explanation offered for the residual, not a rate. So the honest position +is: **prose dispatch is not guaranteed, the magnitude is unknown, and no evidence exists that would +let this ADR treat it as negligible.** + +### The premise that does not survive checking + +The natural argument for accepting that risk is that the feature is opt-in: a user who never sets +the key keeps today's static, host-guaranteed behavior, so the exposure is confined to people who +chose the trade. **That argument is false, and it is worth stating why, because it is the argument +this ADR was expected to make.** + +A global install shares one `~/.claude/gsd-core/` tree and one set of skill files across every +project on the machine — #4139's own rejected-alternative 3 says so, and its acceptance criteria +forbid install-time file selection outright. The shipped `SKILL.md` is therefore a single artifact +serving all projects, and an `@`-include in it is expanded unconditionally. There is no way to +write one file that eagerly includes the workflow for opted-out projects and does not for opted-in +ones. Removing the `@`-include removes the guarantee **for everyone**, opted in or not; keeping it +means opted-in projects save nothing on the stream this feature exists for. + +That is a hard constraint, not a design preference. Any design that saves tokens on stream 1 must +ensure no user receives the canonical body eagerly. + +## Decision + +### 1. The name + +`workflow.compact_content`, surfaced as **compact content mode**. Satisfies approval condition 1. + +"Content" is this repo's own word for this corpus — `CONTRIBUTING.md` §"Editing shipped content", +and the guard family is named the shipped-content guards. "Prompts" would be too narrow for a set +that includes planning-artifact templates. Chosen over `workflow.terse_content` (describes the +prose style rather than the mechanism) and `workflow.lean_prompts` (same narrowness problem, plus +"lean" already carries an unrelated meaning in this repo's vocabulary). + +### 2. What this feature is actually for + +#4139 is pitched on per-invocation token cost. ADR-1610 has already discounted that argument in +this repo's own record: *"with prompt caching the per-invocation cost premise is weak (cache reads +are ~10% of input), so the caching-independent quality argument is the load-bearing one."* + +This ADR adopts ADR-1610's load-bearing argument instead. **The justification for compact content +is finite attention, not price.** A 98 KB instruction file occupying the context window before the +first step runs is 98 KB of attention not spent on the developer's code, and that cost is paid +whether or not the tokens were cheap to transport. Cost reduction is a real secondary effect and +the benchmark in Phase 4 (#4404) will report it, labeled as what it is. + +This matters beyond framing: it decides what a good split looks like. Optimizing for price rewards +deleting words anywhere. Optimizing for attention rewards moving words **out of the always-loaded +window** — which is what this ADR's mechanism does, and what ADR-1610 and the `discuss-phase` +progressive-disclosure split (#717) already established as the sanctioned direction. The +workflow-size budget test says so in its own comments: the correct response to a file at its cap is +*"lazy extraction, never a raise."* + +### 3. The load mechanism — the eager `@`-include stays exactly where it is + +**Decision: do not convert the 58 `` blocks. Change what sits behind them.** + +For every covered top-level workflow: + +- **`gsd-core/workflows/.md` becomes the spine.** Same path, same eager `@`-include, same + host guarantee. It is complete enough to run the workflow correctly on its own. +- **The elaborations move to `gsd-core/workflows//detail/*.md`** — one or more parts, read at + runtime. Decision 6 explains why "one or more" rather than one. +- **`gsd-core/references/compact-content-gate.md`** is the single shared gate. It states the config + check and the resolution rule once; workflows reference it and never restate it. + +Behavior: + +| `workflow.compact_content` | Eagerly loaded | Then reads | Instruction set held | +|---|---|---|---| +| `false` (default) | spine | its `detail/` parts | complete — same content as today | +| `true` | spine | — | the spine | + +No command file changes. No skill file changes. No `@`-include is removed, converted, or made +conditional. #4139's claim that *"`@`-includes are untouched"* — untrue of the design it proposed — +becomes literally true of this one. + +The savings are real because the eagerly-included file got smaller. For `plan-phase.md`, today's +98,290 bytes become a spine plus detail parts whose sum is the same; an opted-in project loads +only the spine. The reduction available here is larger than #4139's projected aggregate, because it +comes from removing content from the eager window rather than from rewording it. + +### 4. The fail-safe — the degradation direction is the whole point + +**A missed `Read` under this design leaves the orchestrator running on the spine: correct, terser, +and identical to the state an opted-in project runs in deliberately. It never leaves it running on +nothing.** + +This is the property that decides the design, and it is worth being precise about why it holds +rather than asserting it: + +- The baseline instruction set arrives by host substitution, which cannot be missed. That is + unchanged from today, for every user, opted in or not. +- The runtime `Read` is only ever **additive** — it supplies elaboration on top of a complete + baseline. It is never the delivery mechanism for the baseline itself. +- Therefore the worst outcome of prose-dispatch variance is the *opt-in behavior*, arriving for a + user who did not opt in. That is a quality regression bounded by a state the project already + considers acceptable — the same state Phase 2's accuracy spot-check (#4402) exists to validate + before any of this proliferates. + +This is why this ADR does not need to resolve the open question about prose-dispatch reliability +that ADR-3646 §Context raised and #3647's closure left standing. It **routes around** it: the +unreliable mechanism is never load-bearing. Where a stream *can* be served by a code seam with a +real exit code instead — token stream 2, the `buildAgentSkillsBlock` path (#4407) — it is, for +exactly ADR-3646 Decision 2's reason, and this ADR adopts that precedent rather than restating it. + +Two consequences to name honestly rather than bury: + +**(a) Opted-out projects do acquire one new failure mode.** Their elaborations now arrive by a +`Read` that could be missed, where today they arrive by substitution. #4139's user story 3 asks +that opting out "cost me nothing", and this is a deviation from that. It is the *minimum possible* +deviation: no design that delivers stream-1 savings can leave the canonical body eagerly loaded +(see §Context), so the achievable maximum is exactly this — degradation bounded to the compact +baseline. The alternative shape (convert the `@`-include) has the same new failure mode with an +unbounded consequence. Recorded as a reconciled criterion in Decision 7 rather than silently +accepted. + +**(b) The spine's correctness is now load-bearing for everyone, not just opt-ins.** A compact +variant is no longer "a cheaper alternative a few users choose" — it is the floor every user lands +on if a `Read` is missed. This raises the authoring bar, and Decision 5's protected-content list and +boundary-move declaration are what enforce it. + +**(c) "The spine still runs the workflow" is not a machine-decidable property, and this ADR does not +pretend otherwise.** Decision 5's checks verify completeness once at split time, then disjointness, +registration and protected-content presence forever. None of them re-derives *sufficiency*. A later +PR could move genuinely load-bearing procedural text — text that carries no protected-content +sentinel because it is a step, not a guardrail — out of a spine and into a part, and every +mechanical check would stay green while the floor quietly dropped. That is the honest residual, and +it is the same class of problem the Feature Review priced in, relocated from "stale duplicate text" +to "boundary placement" rather than eliminated. + +What Decision 5 does about it is make the move **declared instead of silent**: a PR in which a spine +loses non-trivial lines that reappear in its detail parts fails the guard unless it carries a +boundary-move declaration naming the spine — the same enforcement philosophy ADR-3942 applies to +emitted-drift, where the mechanism cannot judge intent so it demands the intent be stated. A +declaration is not proof of sufficiency; it is the point at which a reviewer is guaranteed to be +looking. Combined with the end-to-end spot-check obligation that rides on any spine-shrinking PR +(#4402 establishes it, #4405 and #4406 inherit it), that is the strongest available answer, and it +is authoring discipline with a forced checkpoint rather than a structural invariant. Claiming +otherwise would be the more comfortable sentence and the false one. + +**Alternatives considered for the fail-safe, and why they lost:** + +- *A preflight that fails loud when the content was not loaded.* Needs a signal the workflow's own + dispatch can check, and the only honest one is a load receipt the orchestrator supplies — which + the orchestrator can only supply if it ran the instruction whose omission is being detected. It + detects some misses, not the ones that matter, and it adds machinery to every workflow. Rejected: + strictly weaker than making the miss harmless. +- *Content delivered by a real subprocess (ADR-3646 Decision 2's shape), with a content digest + verified on the next `gsd_run` call.* Genuinely converts a silent miss into a hard halt, and is + the right answer for stream 2 where the payload is already served by a CLI seam. Rejected for + stream 1 on two grounds: a 60 KB instruction body through a subprocess stdout is subject to + harness output truncation in a way a `Read` is not, and the design still leaves a window in which + the orchestrator holds no instructions. Halting loudly is better than proceeding wrongly; not + needing to halt is better than both. +- *Keeping the `@`-include and adding a separate additive compact overlay* (the `text_mode` overlay + shape — the `` dispatch table at + `gsd-core/workflows/discuss-phase.md:20-40`, whose `workflow.text_mode` row at `:30` is the + config-gated case). Preserves every guarantee. Rejected: it + saves nothing on stream 1 — it can only add to an eager window that is already fully paid — so it + is Option A wearing Option B's clothes. +- *Converting the 58 `` blocks* (the shape the epic brief anticipated). Rejected + on the reasoning above: it is the only option whose failure mode is unbounded, and it is not + needed to reach the coverage it was proposed to buy. + +**On the scope decision this replaces.** The maintainer was asked to choose between reducing scope +to the reachable streams and expanding it to convert the `@`-includes, and chose to expand. That +decision is honored: **stream 1 is covered in full**, which is what expanding scope was chosen to +buy. It is covered by a route that does not require the conversion, and therefore does not require +accepting the risk the conversion carries. This ADR is not re-litigating the coverage question; it +is delivering the chosen coverage at the safer option's risk level. + +### 5. Partition, not duplication + +**A split moves text. It does not restate it.** The spine and its detail parts are pieces of one +document, not two documents. + +This is the decision that answers the Feature Review's actual disqualifier. That review returned +**No-go as filed**, and the reason was not the mechanism — *"the gating mechanism is sound; the +disqualifier is scope"* — it was *"permanent dual-maintenance burden (290+ files)"* where *"a +missed re-golf silently serves stale instructions."* #4139's approval overrode that verdict but did +not dissolve the cost; it accepted it, and proposed a drift-parity CI check to contain it. + +A partition dissolves it instead. **There is exactly one copy of each sentence, so there is no +stale twin that can exist.** A canonical edit lands in whichever half owns that text, and no paired +edit is owed. The forever-cost the review priced in — every future content PR carrying a paired +re-compaction — does not accrue. + +What replaces the drift-parity check (Phase 3, #4403): + +- **Completeness, once, at split time** — the union of the two halves, whitespace-normalized, + contains every non-trivial line the canonical file carried at the parent commit. This is what + makes a split reviewable; it runs on the PR that performs the split and never again. +- **Disjointness, ongoing** — no non-trivial line appears in both halves. This is the invariant that + keeps duplication from creeping back in later. +- **Registration, ongoing** — a detail part with no spine, or a spine referencing a detail part that + does not exist, fails and names the pair. +- **Protected content, ongoing** — a spine that has shed a protected-content marker fails and names + the marker. +- **Boundary moves are declared, ongoing** — a PR in which a spine loses non-trivial lines that + reappear in its detail parts fails unless it carries a boundary-move declaration naming the spine. + This is the check that answers Decision 4(c): sufficiency cannot be computed, so the guard makes + the moment it could be lost impossible to pass through unnoticed. + +The protected-content list is #4139's own denylist, promoted from "content the compact variant must +not weaken" to "content that may not leave the spine": negative instructions and guardrails, +output-format contracts, few-shot examples the workflow's own steps depend on, security and +prompt-injection language, and machine-parsed structural headings. + +**The marker is a literal sentinel, not a category judgment.** A guard cannot decide whether a +sentence is "security language" — prose category membership is exactly the kind of judgment that +degrades silently under time pressure, and a check that depends on it is a check that does not +exist. So protection is declared at authoring time by a greppable HTML comment in the repo's +existing `gsd:` comment namespace (the same namespace as the `` header +`plan-phase.md` already carries): + +```markdown + +… one protected block … + + +… a protected region spanning several blocks … + +``` + +The guard's rule is mechanical and has no discretion in it: **a sentinel present in the canonical +file at the parent commit must be present in the spine afterwards, and every line it covers must be +in the spine.** A split that moves a protected block into a detail part fails and names the +sentinel and the line. The categories above are authoring guidance for *where to place sentinels*; +they are never what the guard evaluates. + +Marking is a one-time cost paid during each split, on the file being split. Phase 3 (#4403) owns the +sentinel syntax, the guard, and the failing-first fixture that proves the guard can actually fail — +per this repo's rule that a guard nobody has seen go red is not yet a guard. This is the security +review's finding on this ADR, resolved here rather than carried: the original text specified a +"protected-content marker" without saying what a marker was, which left the strongest check in the +set resting on reviewer judgment. + +Rewriting for terseness is permitted **within** a half and is never a way to move a sentence into +both. Where a split cannot be made by moving text alone without breaking the spine's ability to run +the workflow, the correct answer is a different split point, not a duplicated paragraph. + +### 6. Where the content lives, and what it collides with + +`gsd-core/workflows//detail/*.md` — a sibling of the existing `modes/`, `steps/` and +`templates/` subdirectories, under the workflow it belongs to. + +Chosen over a parallel `gsd-core/workflows/compact/**` tree because a detail part has no meaning +apart from its spine, and a parallel tree is the shape that invites the duplication Decision 5 +exists to prevent. Chosen over a home outside `gsd-core/workflows/` because the guards that sweep +that directory *should* see this content. + +**`` is never user- or project-supplied.** It is the workflow stem the spine already occupies, +which comes from the static set of shipped command files — the same set `listWorkflowStems` walks — +and each spine names its own detail parts literally rather than composing a path from an argument. +Nothing in `$ARGUMENTS`, `.planning/config.json`, or any tracker payload reaches this path. The +config key selects *whether* the read happens; it never selects *what* is read. A project-local file +does not shadow a shipped detail part either: the read resolves against the installed +`~/.claude/gsd-core/` tree the same way the spine's own `@`-include does, so an untrusted repo +cannot substitute instruction content by planting a path. This is stated rather than left to +inference because the seam is new and a future phase reaching for a computed path would be a +traversal vector where today there is none. + +Mapped against the guards that scan this tree, resolved in the phase that first creates a file +there (#4403): + +| Guard | Scope | Effect | +|---|---|---| +| `scripts/workflow-size.cjs` `listWorkflowStems` (`:53`) | top-level `.md`, non-recursive | Detail files are outside it. Spines stay inside and get **smaller** — the direction the budget wants. | +| `tests/workflow-size-budget.test.cjs` `listWorkflowFilesRecursive` (`:617`) | recursive (#3324 sub-guard) | Detail files are scanned for bare `@`-include-in-prompt patterns. An `@` line inside a runtime-read file is inert text, so it must not appear there — the guard already enforces this and is correct to. | +| `tests/helpers/planning-add-guard.cjs` `SCAN_ROOTS` | fully recursive over `gsd-core/workflows` | Detail files are swept by `commit-docs-bypass`. Expected; no exemption sought. | +| `tests/emitted-attribution.test.cjs` | diffs installer output across 19 real installer spawns | New installable content enters scope automatically. Byte deltas are deliberate and carry ADR-3942 acknowledgement trailers. | +| `tests/commit-files-pathspec.test.cjs` | every `.md` under `gsd-core/workflows/` (and five other roots), for `commit`-seam invocations without a `--files` scope (`CONTRIBUTING.md:1164-1170`) | Detail parts are in scope. A `commit` invocation that moves out of a spine into a part must keep its `--files` scope. Earlier research on this epic recorded this guard as "not a content-tree scanner; irrelevant here" — that is wrong, and an unscoped `commit` reaching the runtime is #2269, a CRITICAL-blast-radius defect. | +| `scripts/lint-response-language-coverage.cjs` | top-level, with `//.md` inheriting parent coverage | `detail/` is a fourth subdirectory kind the recognizer does not know. #4403 extends the recognizer; it does not carve an exemption. | +| `NEW_FILE_CAP` (`tests/helpers/emitted-diff.cjs:96`, checked at `:403`) | 32,768 B, applied to every file absent from the baseline and present now | **Hard. Not ack-able, and not exemptible by XL/LARGE tiering.** Detail content is therefore split into parts, each under the cap. | +| tier hard caps (`tests/workflow-size-budget.test.cjs:102-104`) | `XL_CAP` 98,304 · `LARGE_CAP` 61,440 · `DEFAULT_CAP` 40,960 | Apply to spines, which are existing files keeping their tier. Spines only get smaller. | + +**That `NEW_FILE_CAP` row forces a layout decision, and the correct reading of it is not the obvious +one.** Prior research on this epic recorded the cap as living in `workflow-size-budget.test.cjs` and +as waivable by "explicit tiering in the same PR". Both are stale: they describe the pre-#2724 test- +file version. #2724 (ADR-2719 Phase 4) deleted the committed per-file baseline that version keyed +off, and the cap was revived in `tests/helpers/emitted-diff.cjs`, where its own doc comment states +the narrowing plainly — it is *"a HARD cap, not ack-able … Not exempted by explicit XL/LARGE +tiering the way the original test-file version was — this module is intentionally pure and has no +access to that classification … so a legitimately large NEW file must be split via the same +lazy-extraction pattern the tier caps already require."* + +So a single `detail.md` holding the ~63 KB that comes out of `plan-phase.md` is not merely +friction — it is **blocked outright, with no exemption path**. The layout is therefore: + +```text +gsd-core/workflows/.md ← spine, existing path, existing tier, eagerly @-included +gsd-core/workflows//detail/*.md ← one or more parts, each < 32,768 B, read at runtime +``` + +The spine names the parts it defers to, in the same dispatch-table shape `discuss-phase.md`'s +`` block already uses for its mode overlays. This is a better outcome than +one large detail file, not a workaround for the cap: parts are individually skippable, so a +workflow can defer only the sections a given invocation will not reach, and the cap is doing exactly +the job ADR-1610 designed it to do. + +The `@`-include inertness in row 2 has a design consequence worth stating plainly: a canonical +workflow's own `` `@` lines (e.g. `plan-phase.md`'s five reference imports) are +expanded today because the file is `@`-included. They **stay in the spine** and keep working. Moving +one into a detail part would silently turn a working import into dead text, which is exactly the +class of failure the #3324 sub-guard catches. + +### 7. Acceptance criteria — satisfied, and reconciled + +`CI.GATE.acceptance-criteria-required` treats an unmet must-have as a failed deployment, so the +criteria are reconciled here explicitly rather than reinterpreted quietly at ship time. + +**Satisfied as written**, by the phase named: config key registration and validation (#4401); +`/gsd-new-project` question and `/gsd-settings` + `/gsd-config` toggles (#4408); gate logic in +exactly one file (#4402); corpus coverage with no unpaired canonical file (#4405, #4406, #4407); +`agent-skills` serving the compact payload (#4407); golfed spawn patterns (#4405, carried inside +the workflow splits); identical behavior under global and local installs (structural — the key is +per-project and no file is selected at install time); the compression-rules document and its +denylist (#4403); the drift check naming a stale pair (#4403, in its partition form — see below); +the offline reporting-only benchmark with a committed baseline (#4404); the full suite green with +the key on and off (#4408); all shipped-content guards passing (every phase). + +**Reconciled, with the reasoning:** + +1. *"With golf off they load canonical content"* — they do. The reconciliation is mechanical, not + semantic: the canonical content arrives as spine plus detail rather than as one eager blob. The + instruction set an orchestrator holds with the key off is the same instruction set it holds + today. +2. *"Editing a canonical file without updating its golfed variant fails the drift-parity check"* — + under Decision 5 there is no variant to go stale, so this criterion's *purpose* (a canonical edit + cannot silently leave a paired file behind) is met structurally rather than by a check. The check + that remains enforces the invariant that makes it structural: disjointness plus registration. + This is a stronger outcome than the criterion asked for, and #4403 must demonstrate it by proving + each check can actually fail before it is trusted. +3. *"With golf off, GSD behaves exactly as it does today"* (user story 3) — **not fully achievable, + by any design that delivers stream-1 savings.** Decision 4(a) states the residual: an opted-out + project's elaborations arrive by a runtime `Read` that could be missed, and a miss yields the + compact behavior. The deviation is bounded to a state the project deliberately supports and + validates in #4402. This is recorded as a knowing, argued deviation. It is the item on this list + a maintainer may reasonably want to overturn, and if it is overturned the consequence is + automatic: stream 1 becomes uncoverable and the epic reduces to streams 2 and 4. + +## Consequences + +- Every covered workflow becomes two files. The corpus grows in file count while shrinking in + eagerly-loaded bytes, and `docs/INVENTORY.md` plus the manifest regenerate on every split phase. +- The maintenance economics the Feature Review priced as disqualifying do not materialize *in the + form the review priced them*: a partition has no twin, so no future content PR owes a paired edit. + What replaces that cost is smaller but real, and Decision 4(c) names it — the split point itself + can drift, and only a declaration plus a reviewer stands between a spine and a slow erosion of + what it can run on its own. This ADR should be re-read if a future phase finds itself duplicating + rather than moving text, or routinely waving through boundary-move declarations; either is the + signal that the review's cost model has come back in a new shape. +- Spines get smaller, which moves several files down a size tier. Tier membership in + `tests/workflow-size-budget.test.cjs` is adjusted downward as splits land, never held at the old + tier for headroom. +- ADR-3646's Context acquires a stale citation. It is corrected here rather than by editing that + ADR: #3647 is closed, its Decision is unaffected, and its reasoning explicitly disclaimed any + dependence on #3647's state. +- The residual prose-dispatch reliability question raised by #3647's closure thread stays open in + this repo. This ADR does not close it and does not need it closed. Any future design that makes a + runtime `Read` load-bearing for a baseline instruction set will need it answered; this one does + not, and that is the reason it was chosen. +- `gpt-tokenizer` enters `devDependencies` at 27.2 MB unpacked. Nothing ships to users; every + `npm ci`, CI included, pays the install. It is a single-maintainer package, so #4404 pins an exact + version rather than a range and relies on the lockfile's integrity hash — a benchmark is not worth + a floating dependency, and a reporting-only script has no upgrade urgency that would justify one. + +## Rejected alternatives + +- **Convert the 58 `` `@`-includes to config-gated runtime `Read`s.** Delivers + the same coverage this ADR delivers. Rejected: it removes the host guarantee for every user + including those who never opt in (§Context), and its failure mode is "runs with no instructions" + where this ADR's is "runs with fewer". ADR-3646 §Context rejected the same shape for a + structurally identical reason. +- **Reduce scope to the lazily-reachable streams.** Buildable immediately, no new risk. Rejected: + it drops the majority of the value and contradicts an approved acceptance criterion, and it is + unnecessary — the coverage is reachable without the risk. +- **Install-time selection of compact files.** Rejected by #4139 (alternative 3) and by its + acceptance criteria: a global install shares one tree across every project on the machine. +- **Runtime LLM-based compression, and character-count optimization.** Rejected by #4139 + (alternatives 1 and 2); nothing found here changes either rejection. +- **Compacting canonical content in place for everyone.** Rejected by #4139 (alternative 5). Worth + distinguishing from this ADR's decision, because they can look similar: a split preserves every + word for opted-out projects and changes only *when* it loads. In-place compaction deletes words + for everyone and forfeits the side-by-side comparison that makes the trade evaluable. + +## Phase plan + +| Phase | Issue | Delivers | Depends on | +|---|---|---|---| +| 0 | [#4400](https://github.com/open-gsd/gsd-core/issues/4400) | this ADR | — | +| 1 | [#4401](https://github.com/open-gsd/gsd-core/issues/4401) | `workflow.compact_content` end to end | 0 | +| 2 | [#4402](https://github.com/open-gsd/gsd-core/issues/4402) | shared gate + pilot split + accuracy spot-check | 1 | +| 3 | [#4403](https://github.com/open-gsd/gsd-core/issues/4403) | partition rules + the five checks | 2 | +| 4 | [#4404](https://github.com/open-gsd/gsd-core/issues/4404) | offline benchmark + committed baseline | 3 | +| 5 | [#4405](https://github.com/open-gsd/gsd-core/issues/4405) | stream 1 corpus coverage (carries stream 3) | 3, 4 | +| 6 | [#4406](https://github.com/open-gsd/gsd-core/issues/4406) | stream 1b subdirectories + stream 4 templates | 5 | +| 7 | [#4407](https://github.com/open-gsd/gsd-core/issues/4407) | stream 2 agent-skill payloads via the CLI seam | 6 | +| 8 | [#4408](https://github.com/open-gsd/gsd-core/issues/4408) | user surfaces, docs, guard ledger — closes #4139 | 7 | + +Guards land before content proliferates (Phases 3 and 4 precede Phase 5), satisfying approval +condition 2. The pilot's end-to-end accuracy spot-check is Phase 2's, satisfying condition 3. + +## Open questions for the implementation phases + +- Whether `discuss-phase` — the one command that already reaches its workflow by runtime `Read` — + should be brought onto the spine shape too. It is the sole existing instance of the substitutive + load this ADR declines to generalize, which means it already carries the failure mode this ADR + avoids, mitigated only by prose: `commands/gsd/discuss-phase.md:64` reads *"**MANDATORY:** Read + the appropriate workflow file BEFORE taking any action … Do not improvise from the summary."* + Giving it a spine would remove that residual entirely, and it is the one place in the tree where + this ADR's mechanism would be a strict safety improvement rather than a token trade. Scoped to + Phase 5 (#4405) to decide with the rest of stream 1 in view. +- Whether the disjointness check should compare normalized sentences rather than normalized lines. + Lines are cheaper and catch copy-paste; sentences catch reflowing. Decided in #4403 against real + splits rather than in the abstract here. diff --git a/docs/adr/README.md b/docs/adr/README.md index b759a7f92..6f8f5b06e 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -286,6 +286,7 @@ Decided in principle, not yet ratified. Do not cite as settled architecture. | [ADR-3646](3646-per-task-content-resolution-seam.md) | Per-task external-tracker content-resolution seam | Proposed | — | | [ADR-3889](3889-process-exit-contract.md) | One exit-code registry — 0 and 1 are free, everything else is allocated | Proposed | — | | [ADR-3942](3942-emitted-drift-ack-commit-trailer.md) | The emitted-drift acknowledgment is PR-lifetime data — it belongs in a commit trailer, not the working tree | Proposed | — | +| [ADR-4139](4139-compact-content-seam.md) | The compact-content seam — shrink the eager window, never the guarantee | Proposed | — | ### Superseded, Retired, and Legacy From b7917882bbc423337cc6dc4336da2efabee3ecef Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 13:23:26 -0400 Subject: [PATCH 016/166] fix(#4398): render the pending-todo bullet link repo-relative (#4416) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4384): failing-first regression rows for the macOS long-base todo-cap failure The 240-char pending-todo bullet cap must be deterministic w.r.t. where the repo is checked out. Deterministic long-base-path fixtures (a single 110-char segment, no real macOS dependency) reproduce next's own macos shard 3/3 failure (run 34038716700) on every OS: with an absolute link the bullet exceeds the cap and the documented needs-first truncation drops the 'Needs ' clause. Rows cover the determinism property (byte-identical bullets under short and long bases), the CLI surface, relative-path stability, legacy no-projectRoot behavior, drop-order preservation, and adversarial edges (outside-root, path===root, non-string path). * fix(#4384): render the pending-todo bullet link repo-relative renderPendingTodosMarkdown gains an optional projectRoot; when given and the todo's path is absolute, the bullet's [todo file](…) target becomes toPosixPath(path.relative(projectRoot, path)) — the idiom already used for project_exists. cmdInitTodos passes cwd. The JSON todos[].path field stays absolute (#2376). Only the rendered display link changes: embedding the machine-variable absolute base let macOS's /private/var/folders/… temp paths consume the 240-char budget and drop the 'Needs' clause on long-path machines only — next's own macos-latest shard 3/3 went red on exactly this (run 34038716700), Linux's short /tmp passed. The 240-char whole-bullet cap and the needs→title→area drop order are unchanged; this matches PR #4384's own canonical example, docs, and unit tests, which all show repo-relative links. Docs updated at all three surfaces that describe the bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was still pre-#4384 'count and reference' prose). Fixes the macOS regression introduced by #4384; next is red on its own CI. * test(#4384): fix substring false positive in the outside-root regression row The ../-form relative link legitimately contains the absolute path as a substring, so !line.includes(absolutePath) fired on correct output (caught by the first remote verify run, linux-node24 44018/44019). Assert the property itself instead: extract the link target and require it to be non-absolute and not equal to the absolute path. * chore(#4398): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/wise-foxes-march.md | 5 ++ docs/COMMANDS.md | 2 +- docs/reference/state-md.md | 2 +- gsd-core/templates/state.md | 3 +- src/init.cts | 42 +++++++-- tests/state-todos-render.test.cjs | 136 ++++++++++++++++++++++++++++++ 6 files changed, 182 insertions(+), 8 deletions(-) create mode 100644 .changeset/wise-foxes-march.md diff --git a/.changeset/wise-foxes-march.md b/.changeset/wise-foxes-march.md new file mode 100644 index 000000000..75b924512 --- /dev/null +++ b/.changeset/wise-foxes-march.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4416 +--- +**macOS todo rendering no longer drops the Needs clause under long temp paths** — the 240-char bound is now deterministic w.r.t. base-path length. (#4384 regression) diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index d45e1346e..a2c2f33ca 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -1938,7 +1938,7 @@ Capture ideas, tasks, notes, and seeds to their appropriate destination. Default **Produces:** `.planning/todos/` (default), note files (--note), ROADMAP.md backlog section (--backlog), `.planning/seeds/SEED-NNN-slug.md` (--seed) -**STATE.md rendering:** each capture (or `--list` action that changes the pending count) refreshes STATE.md's "### Pending Todos" section to one bullet per pending todo, each capped at 240 characters — `- [date] [area] title — [todo file](path) — Needs ...`. A todo with no clear next step omits the "Needs ..." clause rather than the bullet. Refresh is fail-safe: a failed or malformed lookup leaves the existing section untouched rather than clearing it. +**STATE.md rendering:** each capture (or `--list` action that changes the pending count) refreshes STATE.md's "### Pending Todos" section to one bullet per pending todo, each capped at 240 characters — `- [date] [area] title — [todo file](path) — Needs ...`. The todo-file link is repo-relative (`.planning/todos/pending/...`), so the cap is independent of where the repo is checked out — a long absolute path never consumes the budget or drops the "Needs ..." clause. A todo with no clear next step omits the "Needs ..." clause rather than the bullet. Refresh is fail-safe: a failed or malformed lookup leaves the existing section untouched rather than clearing it. ```bash /gsd-capture "Consider adding dark mode support" # Add todo diff --git a/docs/reference/state-md.md b/docs/reference/state-md.md index 21a8c4e1a..1c2d96b18 100644 --- a/docs/reference/state-md.md +++ b/docs/reference/state-md.md @@ -224,7 +224,7 @@ Updated after each plan completion. **Decisions** — a summary of recent decisions affecting current work (full log lives in `PROJECT.md`). Added via `gsd-tools state add-decision`. -**Pending Todos** — count and reference to `.planning/todos/pending/`. Captured via `/gsd-capture`. +**Pending Todos** — one bullet per pending todo (`- [date] [area] title — [todo file](repo-relative path) — Needs ...`, capped at 240 characters; repo-relative link keeps the cap independent of checkout path length). Captured via `/gsd-capture`. **Blockers/Concerns** — issues affecting future work, prefixed with the originating phase. Added via `gsd-tools state add-blocker`; resolved via `gsd-tools state resolve-blocker`. diff --git a/gsd-core/templates/state.md b/gsd-core/templates/state.md index 32afd9a90..a3deba6d8 100644 --- a/gsd-core/templates/state.md +++ b/gsd-core/templates/state.md @@ -173,7 +173,8 @@ Updated after each plan completion. **Pending Todos:** Ideas captured via /gsd-add-todo - One bullet per pending todo, rendered by `init.todos`'s `pending_todos_markdown` - (each bullet capped at 240 characters: `- [date] [area] title — [todo file](path) — Needs ...`) + (each bullet capped at 240 characters: `- [date] [area] title — [todo file](path) — Needs ...`; + the todo-file link is repo-relative, so the cap does not depend on checkout path length) - `None yet.` when there are no pending todos - No collapse-by-count fallback — every pending todo gets its own line, always (see #2618 design doc for why a "count if many" fallback was rejected) diff --git a/src/init.cts b/src/init.cts index 0192d76ce..b11de2485 100644 --- a/src/init.cts +++ b/src/init.cts @@ -2265,19 +2265,49 @@ function truncatePendingTodoText(value: string, maxLen: number): string { * code rather than a prose algorithm (DEFECT.GENERATIVE-FIX: a prose * algorithm duplicated as a test oracle is exactly the divergence class * this avoids). + * + * #4384 regression fix: the optional `projectRoot` makes the bullet's + * `[todo file](…)` link repo-relative (see pendingTodoLinkTarget) so the cap + * is deterministic w.r.t. where the repo is checked out. Omitting it keeps + * the legacy absolute-link behavior for existing direct callers. */ -function renderPendingTodosMarkdown(todos: Record[]): string { +function renderPendingTodosMarkdown(todos: Record[], projectRoot?: string): string { if (!Array.isArray(todos) || todos.length === 0) { return 'None yet.'; } - return todos.map((todo) => renderPendingTodoBullet(todo)).join('\n'); + return todos.map((todo) => renderPendingTodoBullet(todo, projectRoot)).join('\n'); } function pendingTodoFieldAsString(value: unknown, fallback: string): string { return typeof value === 'string' && value.length > 0 ? value : fallback; } -function renderPendingTodoBullet(todo: Record): string { +/** + * #4384 regression fix: the bullet's markdown link target, rendered + * repo-relative when `projectRoot` is given and the todo's `path` is + * absolute. The JSON `todos[].path` field stays absolute (#2376 contract); + * only the rendered display link changes — embedding the machine-variable + * absolute base let macOS's /private/var/folders/… temp paths consume the + * 240-char budget and drop the "Needs " clause on long-path + * machines only (next's own macos CI shard went red on exactly this, run + * 34038716700). Repo-relative links also resolve correctly from STATE.md at + * the repo root and survive repo moves. + */ +function pendingTodoLinkTarget(todo: Record, projectRoot: string | undefined): string { + const raw = pendingTodoFieldAsString(todo['path'], ''); + if (typeof projectRoot !== 'string' || projectRoot.length === 0 || !path.isAbsolute(raw)) { + return raw; + } + const rel = toPosixPath(path.relative(projectRoot, raw)); + if (rel.length === 0 || path.isAbsolute(rel)) { + // Degenerate (path === projectRoot) or Windows cross-drive fallback: + // keep the raw target rather than emitting an empty or incorrect link. + return raw; + } + return rel; +} + +function renderPendingTodoBullet(todo: Record, projectRoot?: string): string { const date = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['created'], 'unknown')); let area = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['area'], 'general')); let title = sanitizePendingTodoInline(pendingTodoFieldAsString(todo['title'], 'Untitled')); @@ -2287,7 +2317,7 @@ function renderPendingTodoBullet(todo: Record): string { typeof todo['needs'] === 'string' ? sanitizePendingTodoInline(todo['needs']).replace(/\.+$/, '') : ''; - const link = `[todo file](${pendingTodoFieldAsString(todo['path'], '')})`; + const link = `[todo file](${pendingTodoLinkTarget(todo, projectRoot)})`; const assemble = (): string => { const needsClause = needs ? ` — Needs ${needs}.` : ''; @@ -2420,7 +2450,9 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void // (rather than emitted with possibly-wrong data) when pendingReadOk is // false, so the workflow's fail-safe check can key off field presence. pending_read_ok: pendingReadOk, - ...(pendingReadOk ? { pending_todos_markdown: renderPendingTodosMarkdown(todos) } : {}), + // #4384 fix: pass cwd as projectRoot so the bullet link renders + // repo-relative — see pendingTodoLinkTarget. + ...(pendingReadOk ? { pending_todos_markdown: renderPendingTodosMarkdown(todos, cwd) } : {}), }; output(withProjectRoot(cwd, result), raw); diff --git a/tests/state-todos-render.test.cjs b/tests/state-todos-render.test.cjs index 5574f0e39..d47b5db80 100644 --- a/tests/state-todos-render.test.cjs +++ b/tests/state-todos-render.test.cjs @@ -167,6 +167,102 @@ test('property: rendered body always has one line per todo, each line <= 240 cha ); }); +// ─── #4384 regression: the 240-char cap must be deterministic w.r.t. where the +// repo is checked out (macOS /private/var/folders/… bases blew the budget and +// dropped the "Needs" clause; Linux /tmp passed — next's own macos shard 3/3 +// went red on exactly this, CI run 34038716700) ────────────────────────────── + +// Long enough that the pre-fix absolute-link bullet exceeds 240 chars on EVERY +// OS (Linux /tmp included): base > ~90 chars is over the threshold since the +// fixed skeleton + relative tail + needs clause land near 150. +const LONG_BASE_SEGMENT = 'd'.repeat(110); +const TODO_RELATIVE_TAIL = path + .join('.planning', 'todos', 'pending', '2026-09-01-fix-retry-logic.md') + .split(path.sep) + .join('/'); + +test('renderPendingTodosMarkdown: cap is deterministic w.r.t. base-path length (needs clause survives long absolute bases)', () => { + const shortBase = path.resolve('/', 'gsd-2618-short-base'); + const longBase = path.join(shortBase, LONG_BASE_SEGMENT); + const todoAt = (base) => + makeTodo({ path: path.join(base, TODO_RELATIVE_TAIL), needs: 'Add a max-attempts cap.' }); + + const fromShort = renderPendingTodosMarkdown([todoAt(shortBase)], shortBase); + const fromLong = renderPendingTodosMarkdown([todoAt(longBase)], longBase); + + assert.equal(fromLong, fromShort, 'bullet must not depend on where the repo is checked out'); + assert.match(fromShort, /Needs Add a max-attempts cap\.$/); + assert.match(fromLong, /Needs Add a max-attempts cap\.$/); + assert.ok( + fromLong.includes(`[todo file](${TODO_RELATIVE_TAIL})`), + 'link must be the repo-relative tail, not the absolute path', + ); +}); + +test('renderPendingTodosMarkdown: already-relative path is byte-stable with and without projectRoot', () => { + const todo = makeTodo({ needs: 'define retry behavior' }); + const withRoot = renderPendingTodosMarkdown([todo], path.resolve('/', 'gsd-2618-rel-root')); + assert.equal(withRoot, renderPendingTodosMarkdown([todo])); +}); + +test('renderPendingTodosMarkdown: absolute path without projectRoot keeps the legacy absolute link and drop order', () => { + const base = path.join(path.resolve('/', 'gsd-2618-legacy-base'), LONG_BASE_SEGMENT); + const todo = makeTodo({ path: path.join(base, TODO_RELATIVE_TAIL), needs: 'a real needs clause' }); + const line = renderPendingTodosMarkdown([todo]); + // Opt-out callers keep today's behavior: over-cap drops needs first, link + // verbatim (raw separators — the legacy renderer never posix-normalizes). + assert.doesNotMatch(line, /Needs/); + assert.ok(line.includes(`[todo file](${todo.path})`)); +}); + +test('renderPendingTodosMarkdown: long base + long title still bounds the bullet and keeps the drop order', () => { + const base = path.join(path.resolve('/', 'gsd-2618-order-base'), LONG_BASE_SEGMENT); + const todo = makeTodo({ + path: path.join(base, TODO_RELATIVE_TAIL), + title: 'A'.repeat(300), + needs: 'something', + }); + const line = renderPendingTodosMarkdown([todo], base); + assert.ok(line.length <= MAX, `expected <= ${MAX}, got ${line.length}`); + assert.doesNotMatch(line, /Needs/, 'needs clause must be dropped before title is touched'); + assert.match(line, /…/); + assert.ok(line.includes(`[todo file](${TODO_RELATIVE_TAIL})`), 'relative link verbatim'); +}); + +test('renderPendingTodosMarkdown: absolute path outside projectRoot renders a ../ relative link', () => { + const root = path.resolve('/', 'gsd-2618-outside-root'); + const outsideTodo = path.join( + path.resolve('/', 'gsd-2618-outside-sibling'), + TODO_RELATIVE_TAIL, + ); + const line = renderPendingTodosMarkdown([makeTodo({ path: outsideTodo })], root); + // The ../-form legitimately CONTAINS the absolute path as a substring, so a + // plain includes() negative is a false positive — assert the property + // itself: the link target is relative, never the absolute path. + const linkMatch = line.match(/\[todo file\]\(([^)]*)\)/); + assert.ok(linkMatch, 'bullet must contain a todo link'); + assert.ok(!path.isAbsolute(linkMatch[1]), 'link target must be relative, not absolute'); + assert.notEqual(linkMatch[1], outsideTodo, 'link target must not be the machine-variable absolute path'); + assert.ok( + line.includes('[todo file](../'), + 'link must be relative to the project root, not absolute', + ); +}); + +test('renderPendingTodosMarkdown: path equal to projectRoot falls back to the raw link target', () => { + const root = path.resolve('/', 'gsd-2618-eq-root'); + const line = renderPendingTodosMarkdown([makeTodo({ path: root })], root); + assert.ok(line.includes(`[todo file](${root})`)); +}); + +test('renderPendingTodosMarkdown: non-string path keeps the empty-link fallback', () => { + const line = renderPendingTodosMarkdown( + [makeTodo({ path: 42 })], + path.resolve('/', 'gsd-2618-num-root'), + ); + assert.match(line, /\[todo file\]\(\)$/); +}); + // ─── pending_read_ok / pending_todos_markdown via the real CLI surface ───── const { spawnSync } = require('node:child_process'); @@ -228,6 +324,46 @@ test('cmdInitTodos: real todo file produces a rendered bullet via the CLI', (t) assert.match(json.pending_todos_markdown, /Needs Add a max-attempts cap\.$/m); }); +test('cmdInitTodos: needs clause survives a deterministically long base path (#4384 macOS shape)', (t) => { + const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2618-long-')); + t.after(() => cleanup(root)); + // One long path segment: deterministic on every OS (no real macOS dependency). + // Pre-fix, the absolute-link bullet exceeds 240 chars even on Linux /tmp, + // reproducing next's macos shard failure (CI run 34038716700). + const dir = path.join(root, LONG_BASE_SEGMENT); + const pendingDir = path.join(dir, '.planning', 'todos', 'pending'); + fs.mkdirSync(pendingDir, { recursive: true }); + fs.writeFileSync( + path.join(pendingDir, '2026-09-01-fix-retry-logic.md'), + [ + '---', + 'created: 2026-09-01T00:00:00.000Z', + 'title: Fix retry logic', + 'area: api', + '---', + '', + '## Solution', + '', + 'Add a max-attempts cap.', + '', + ].join('\n'), + ); + + const json = runQueryInitTodos(dir); + assert.equal(json.pending_read_ok, true); + assert.equal(json.todo_count, 1); + assert.match(json.pending_todos_markdown, /Fix retry logic/); + assert.match(json.pending_todos_markdown, /Needs Add a max-attempts cap\.$/m); + assert.ok( + json.pending_todos_markdown.includes(`[todo file](${TODO_RELATIVE_TAIL})`), + 'bullet link must be repo-relative', + ); + assert.ok( + !json.pending_todos_markdown.includes(dir.split(path.sep).join('/')), + 'absolute cwd must not leak into the rendered bullet', + ); +}); + test('cmdInitTodos: bullet order is filename-sorted regardless of write/insertion order', (t) => { const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-2618-')); t.after(() => cleanup(dir)); From ed133cc11614bd132f9475adc701a5182fe5b4e5 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 13:37:47 -0400 Subject: [PATCH 017/166] enhance(#3085): add Grep to allowed-tools for 21 skills (#4397) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * enhance(#3085): add Grep to allowed-tools for 21 skills 29 of 71 skills omit Grep from allowed-tools, forcing Bash grep for structured search instead of the dedicated tool. Adds Grep to the 21-skill subset confirmed safe in prior review (excludes the 6 gsd-ns-* dispatchers, gsd-help, and gsd-surface, which have no plausible structured-search need). Hand-edits commands/gsd/*.md only; skills/*/SKILL.md is regenerated via `npm run gen:plugin-skills` from that source. Updates the one hardcoded copilot-install test assertion affected by gsd-health's new tool order. * chore(#3085): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 * test(#2618): assert the Needs clause on the structured field, not the path-length-dependent bullet The bullet embeds the todo file's ABSOLUTE path, so its total length varies by runner tmpdir: on macOS CI, /private/var/folders/… plus the test harness's gsd-test-run-* wrapper pushed the full bullet to 244 chars — past renderPendingTodoBullet's intended 240-char cap, whose documented first degradation step drops the Needs clause. The product behavior is correct (#2618 design); the assertion was runner-dependent. The extraction is now pinned on json.todos[0].needs; the rendered-bullet shape stays covered by the path-independent title assertion and the renderer's own unit rows. Found blocking #4186's CI on the macOS shard (test landed 30 minutes earlier in b7406b293f / PR #4384). --------- Co-authored-by: sim Co-authored-by: Claude Sonnet 5 --- .changeset/calm-quails-snooze.md | 6 ++++++ commands/gsd/cleanup.md | 1 + commands/gsd/complete-milestone.md | 1 + commands/gsd/config.md | 1 + commands/gsd/debug.md | 1 + commands/gsd/graphify.md | 1 + commands/gsd/health.md | 1 + commands/gsd/mempalace-capture.md | 1 + commands/gsd/mempalace-recall.md | 1 + commands/gsd/new-milestone.md | 1 + commands/gsd/new-project.md | 1 + commands/gsd/next.md | 1 + commands/gsd/pause-work.md | 1 + commands/gsd/phase.md | 1 + commands/gsd/pr-branch.md | 1 + commands/gsd/resume-work.md | 1 + commands/gsd/review-backlog.md | 1 + commands/gsd/settings.md | 1 + commands/gsd/stats.md | 1 + commands/gsd/thread.md | 1 + commands/gsd/workspace.md | 1 + commands/gsd/workstreams.md | 1 + skills/gsd-cleanup/SKILL.md | 1 + skills/gsd-complete-milestone/SKILL.md | 1 + skills/gsd-config/SKILL.md | 1 + skills/gsd-debug/SKILL.md | 1 + skills/gsd-graphify/SKILL.md | 1 + skills/gsd-health/SKILL.md | 1 + skills/gsd-mempalace-capture/SKILL.md | 1 + skills/gsd-mempalace-recall/SKILL.md | 1 + skills/gsd-new-milestone/SKILL.md | 1 + skills/gsd-new-project/SKILL.md | 1 + skills/gsd-next/SKILL.md | 1 + skills/gsd-pause-work/SKILL.md | 1 + skills/gsd-phase/SKILL.md | 1 + skills/gsd-pr-branch/SKILL.md | 1 + skills/gsd-resume-work/SKILL.md | 1 + skills/gsd-review-backlog/SKILL.md | 1 + skills/gsd-settings/SKILL.md | 1 + skills/gsd-stats/SKILL.md | 1 + skills/gsd-thread/SKILL.md | 1 + skills/gsd-workspace/SKILL.md | 1 + skills/gsd-workstreams/SKILL.md | 1 + tests/copilot-install.test.cjs | 2 +- tests/state-todos-render.test.cjs | 14 ++++++++++++-- 45 files changed, 61 insertions(+), 3 deletions(-) create mode 100644 .changeset/calm-quails-snooze.md diff --git a/.changeset/calm-quails-snooze.md b/.changeset/calm-quails-snooze.md new file mode 100644 index 000000000..c0962c607 --- /dev/null +++ b/.changeset/calm-quails-snooze.md @@ -0,0 +1,6 @@ +--- +type: Changed +pr: 4397 +--- + +**21 GSD skills now declare `Grep` in `allowed-tools`** — cleanup, complete-milestone, config, debug, graphify, health, mempalace-capture, mempalace-recall, new-milestone, new-project, next, pause-work, phase, pr-branch, resume-work, review-backlog, settings, stats, thread, workspace, and workstreams can now use the dedicated structured-search tool instead of shelling out through Bash grep. diff --git a/commands/gsd/cleanup.md b/commands/gsd/cleanup.md index 51fd13cc8..813daf2cd 100644 --- a/commands/gsd/cleanup.md +++ b/commands/gsd/cleanup.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion requires: [phase] --- diff --git a/commands/gsd/complete-milestone.md b/commands/gsd/complete-milestone.md index 20e73adf0..38188cbda 100644 --- a/commands/gsd/complete-milestone.md +++ b/commands/gsd/complete-milestone.md @@ -7,6 +7,7 @@ allowed-tools: - Read - Write - Bash + - Grep requires: [audit-milestone, discuss-phase, execute-phase, new-milestone, phase, plan-phase, stats, update] --- diff --git a/commands/gsd/config.md b/commands/gsd/config.md index b080c0dae..c0c67c3e7 100644 --- a/commands/gsd/config.md +++ b/commands/gsd/config.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion requires: [code-review, review, settings] --- diff --git a/commands/gsd/debug.md b/commands/gsd/debug.md index c339fd593..55e712c1e 100644 --- a/commands/gsd/debug.md +++ b/commands/gsd/debug.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - Agent - AskUserQuestion --- diff --git a/commands/gsd/graphify.md b/commands/gsd/graphify.md index 8be09ad98..c9b14fa71 100644 --- a/commands/gsd/graphify.md +++ b/commands/gsd/graphify.md @@ -5,6 +5,7 @@ argument-hint: "[build|query |status|diff]" allowed-tools: - Read - Bash + - Grep requires: [config, fast, phase, update] --- diff --git a/commands/gsd/health.md b/commands/gsd/health.md index 81b8df0e2..148205899 100644 --- a/commands/gsd/health.md +++ b/commands/gsd/health.md @@ -5,6 +5,7 @@ argument-hint: "[--repair] [--context]" allowed-tools: - Read - Bash + - Grep - Write - AskUserQuestion requires: [thread] diff --git a/commands/gsd/mempalace-capture.md b/commands/gsd/mempalace-capture.md index 7a2537fc5..2cc5a9e01 100644 --- a/commands/gsd/mempalace-capture.md +++ b/commands/gsd/mempalace-capture.md @@ -5,6 +5,7 @@ argument-hint: "[CONTEXT.md|PLAN.md|SUMMARY.md]" allowed-tools: - Read - Bash + - Grep requires: [config] --- diff --git a/commands/gsd/mempalace-recall.md b/commands/gsd/mempalace-recall.md index f043be0a2..cc9ec2928 100644 --- a/commands/gsd/mempalace-recall.md +++ b/commands/gsd/mempalace-recall.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep requires: [config] --- diff --git a/commands/gsd/new-milestone.md b/commands/gsd/new-milestone.md index 5783c710c..6020c6bf4 100644 --- a/commands/gsd/new-milestone.md +++ b/commands/gsd/new-milestone.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - Agent - AskUserQuestion requires: [new-project, phase, plan-phase] diff --git a/commands/gsd/new-project.md b/commands/gsd/new-project.md index 98b0d8469..273926012 100644 --- a/commands/gsd/new-project.md +++ b/commands/gsd/new-project.md @@ -5,6 +5,7 @@ argument-hint: "[--auto]" allowed-tools: - Read - Bash + - Grep - Write - Agent - AskUserQuestion diff --git a/commands/gsd/next.md b/commands/gsd/next.md index 8ca19c1a1..49d8541e9 100644 --- a/commands/gsd/next.md +++ b/commands/gsd/next.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Bash - Glob + - Grep - SlashCommand - AskUserQuestion --- diff --git a/commands/gsd/pause-work.md b/commands/gsd/pause-work.md index 4a1af80bb..d6a54759c 100644 --- a/commands/gsd/pause-work.md +++ b/commands/gsd/pause-work.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep requires: [phase, progress] --- diff --git a/commands/gsd/phase.md b/commands/gsd/phase.md index b5e2ba972..f5aff8d5a 100644 --- a/commands/gsd/phase.md +++ b/commands/gsd/phase.md @@ -7,6 +7,7 @@ allowed-tools: - Write - Bash - Glob + - Grep --- diff --git a/commands/gsd/pr-branch.md b/commands/gsd/pr-branch.md index 23a4e1f69..40d66988e 100644 --- a/commands/gsd/pr-branch.md +++ b/commands/gsd/pr-branch.md @@ -4,6 +4,7 @@ description: Create a clean PR branch by filtering out .planning/ commits — re argument-hint: "[target branch, default: main]" allowed-tools: - Bash + - Grep - Read - AskUserQuestion requires: [review] diff --git a/commands/gsd/resume-work.md b/commands/gsd/resume-work.md index e049dfb37..990cea719 100644 --- a/commands/gsd/resume-work.md +++ b/commands/gsd/resume-work.md @@ -4,6 +4,7 @@ description: Resume work from previous session with full context restoration allowed-tools: - Read - Bash + - Grep - Write - AskUserQuestion - SlashCommand diff --git a/commands/gsd/review-backlog.md b/commands/gsd/review-backlog.md index c72a457eb..7da8cc7b1 100644 --- a/commands/gsd/review-backlog.md +++ b/commands/gsd/review-backlog.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion requires: [phase, review] --- diff --git a/commands/gsd/settings.md b/commands/gsd/settings.md index 740371023..18cc3889e 100644 --- a/commands/gsd/settings.md +++ b/commands/gsd/settings.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion requires: [quick] --- diff --git a/commands/gsd/stats.md b/commands/gsd/stats.md index ebdba3cca..489f95566 100644 --- a/commands/gsd/stats.md +++ b/commands/gsd/stats.md @@ -5,6 +5,7 @@ effort: low allowed-tools: - Read - Bash + - Grep requires: [phase, progress] --- diff --git a/commands/gsd/thread.md b/commands/gsd/thread.md index 13cdb1cde..3335a9d98 100644 --- a/commands/gsd/thread.md +++ b/commands/gsd/thread.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep requires: [phase] --- diff --git a/commands/gsd/workspace.md b/commands/gsd/workspace.md index 749251b38..d196f0151 100644 --- a/commands/gsd/workspace.md +++ b/commands/gsd/workspace.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/commands/gsd/workstreams.md b/commands/gsd/workstreams.md index 2c6bf22c4..414bff64a 100644 --- a/commands/gsd/workstreams.md +++ b/commands/gsd/workstreams.md @@ -4,6 +4,7 @@ description: Manage parallel workstreams — list, create, switch, status, progr allowed-tools: - Read - Bash + - Grep requires: [new-milestone, phase, progress, resume-work] --- diff --git a/skills/gsd-cleanup/SKILL.md b/skills/gsd-cleanup/SKILL.md index f5e8822c6..6e858c30b 100644 --- a/skills/gsd-cleanup/SKILL.md +++ b/skills/gsd-cleanup/SKILL.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/skills/gsd-complete-milestone/SKILL.md b/skills/gsd-complete-milestone/SKILL.md index 10d940495..48248d83d 100644 --- a/skills/gsd-complete-milestone/SKILL.md +++ b/skills/gsd-complete-milestone/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep --- diff --git a/skills/gsd-config/SKILL.md b/skills/gsd-config/SKILL.md index 2e2bbc008..30efc9b95 100644 --- a/skills/gsd-config/SKILL.md +++ b/skills/gsd-config/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/skills/gsd-debug/SKILL.md b/skills/gsd-debug/SKILL.md index 0fc9b0404..6cfe24ac6 100644 --- a/skills/gsd-debug/SKILL.md +++ b/skills/gsd-debug/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - Agent - AskUserQuestion --- diff --git a/skills/gsd-graphify/SKILL.md b/skills/gsd-graphify/SKILL.md index bf18826b3..786740c61 100644 --- a/skills/gsd-graphify/SKILL.md +++ b/skills/gsd-graphify/SKILL.md @@ -5,6 +5,7 @@ argument-hint: "[build|query |status|diff]" allowed-tools: - Read - Bash + - Grep --- diff --git a/skills/gsd-health/SKILL.md b/skills/gsd-health/SKILL.md index d52c09430..9d2b22252 100644 --- a/skills/gsd-health/SKILL.md +++ b/skills/gsd-health/SKILL.md @@ -5,6 +5,7 @@ argument-hint: "[--repair] [--context]" allowed-tools: - Read - Bash + - Grep - Write - AskUserQuestion --- diff --git a/skills/gsd-mempalace-capture/SKILL.md b/skills/gsd-mempalace-capture/SKILL.md index 3ca959181..337209f92 100644 --- a/skills/gsd-mempalace-capture/SKILL.md +++ b/skills/gsd-mempalace-capture/SKILL.md @@ -5,6 +5,7 @@ argument-hint: "[CONTEXT.md|PLAN.md|SUMMARY.md]" allowed-tools: - Read - Bash + - Grep --- diff --git a/skills/gsd-mempalace-recall/SKILL.md b/skills/gsd-mempalace-recall/SKILL.md index 40b619f16..466c0943b 100644 --- a/skills/gsd-mempalace-recall/SKILL.md +++ b/skills/gsd-mempalace-recall/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep --- diff --git a/skills/gsd-new-milestone/SKILL.md b/skills/gsd-new-milestone/SKILL.md index aecac14ce..fe3611b9e 100644 --- a/skills/gsd-new-milestone/SKILL.md +++ b/skills/gsd-new-milestone/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - Agent - AskUserQuestion --- diff --git a/skills/gsd-new-project/SKILL.md b/skills/gsd-new-project/SKILL.md index bd62c32f0..b9892d5fe 100644 --- a/skills/gsd-new-project/SKILL.md +++ b/skills/gsd-new-project/SKILL.md @@ -5,6 +5,7 @@ argument-hint: "[--auto]" allowed-tools: - Read - Bash + - Grep - Write - Agent - AskUserQuestion diff --git a/skills/gsd-next/SKILL.md b/skills/gsd-next/SKILL.md index fcd0b05b2..4ab9e370d 100644 --- a/skills/gsd-next/SKILL.md +++ b/skills/gsd-next/SKILL.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Bash - Glob + - Grep - SlashCommand - AskUserQuestion --- diff --git a/skills/gsd-pause-work/SKILL.md b/skills/gsd-pause-work/SKILL.md index e1962277a..f92c3e0a7 100644 --- a/skills/gsd-pause-work/SKILL.md +++ b/skills/gsd-pause-work/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep --- diff --git a/skills/gsd-phase/SKILL.md b/skills/gsd-phase/SKILL.md index 2ce83adbf..73ce244e9 100644 --- a/skills/gsd-phase/SKILL.md +++ b/skills/gsd-phase/SKILL.md @@ -7,6 +7,7 @@ allowed-tools: - Write - Bash - Glob + - Grep --- diff --git a/skills/gsd-pr-branch/SKILL.md b/skills/gsd-pr-branch/SKILL.md index 0a0ad7aa0..00f8a926a 100644 --- a/skills/gsd-pr-branch/SKILL.md +++ b/skills/gsd-pr-branch/SKILL.md @@ -4,6 +4,7 @@ description: "Create a clean PR branch by filtering out .planning/ commits — r argument-hint: "[target branch, default: main]" allowed-tools: - Bash + - Grep - Read - AskUserQuestion --- diff --git a/skills/gsd-resume-work/SKILL.md b/skills/gsd-resume-work/SKILL.md index 66dd1c544..12838e654 100644 --- a/skills/gsd-resume-work/SKILL.md +++ b/skills/gsd-resume-work/SKILL.md @@ -4,6 +4,7 @@ description: "Resume work from previous session with full context restoration" allowed-tools: - Read - Bash + - Grep - Write - AskUserQuestion - SlashCommand diff --git a/skills/gsd-review-backlog/SKILL.md b/skills/gsd-review-backlog/SKILL.md index 0a3f3a94f..58933d4ae 100644 --- a/skills/gsd-review-backlog/SKILL.md +++ b/skills/gsd-review-backlog/SKILL.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/skills/gsd-settings/SKILL.md b/skills/gsd-settings/SKILL.md index 86e176541..efc780e89 100644 --- a/skills/gsd-settings/SKILL.md +++ b/skills/gsd-settings/SKILL.md @@ -5,6 +5,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/skills/gsd-stats/SKILL.md b/skills/gsd-stats/SKILL.md index 57d0db82f..cb21f59f8 100644 --- a/skills/gsd-stats/SKILL.md +++ b/skills/gsd-stats/SKILL.md @@ -4,6 +4,7 @@ description: "Display project statistics — phases, plans, requirements, git me allowed-tools: - Read - Bash + - Grep --- diff --git a/skills/gsd-thread/SKILL.md b/skills/gsd-thread/SKILL.md index 215d6e5f3..793fcbc88 100644 --- a/skills/gsd-thread/SKILL.md +++ b/skills/gsd-thread/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep --- diff --git a/skills/gsd-workspace/SKILL.md b/skills/gsd-workspace/SKILL.md index 8028f5abc..70ef63eaf 100644 --- a/skills/gsd-workspace/SKILL.md +++ b/skills/gsd-workspace/SKILL.md @@ -6,6 +6,7 @@ allowed-tools: - Read - Write - Bash + - Grep - AskUserQuestion --- diff --git a/skills/gsd-workstreams/SKILL.md b/skills/gsd-workstreams/SKILL.md index 08f8aec0c..dd5ec5617 100644 --- a/skills/gsd-workstreams/SKILL.md +++ b/skills/gsd-workstreams/SKILL.md @@ -4,6 +4,7 @@ description: "Manage parallel workstreams — list, create, switch, status, prog allowed-tools: - Read - Bash + - Grep --- diff --git a/tests/copilot-install.test.cjs b/tests/copilot-install.test.cjs index ccaa0bc24..a434f3fcd 100644 --- a/tests/copilot-install.test.cjs +++ b/tests/copilot-install.test.cjs @@ -743,7 +743,7 @@ describe('installRuntimeArtifacts (copilot integration)', () => { const skillContent = fs.readFileSync(path.join(skillsDir, 'gsd-health', 'SKILL.md'), 'utf8'); // Frontmatter format checks assert.ok(skillContent.startsWith('---\nname: gsd-health\n'), 'starts with name: gsd-health'); - assert.ok(skillContent.includes('allowed-tools: Read, Bash, Write, AskUserQuestion'), + assert.ok(skillContent.includes('allowed-tools: Read, Bash, Grep, Write, AskUserQuestion'), 'allowed-tools is comma-separated'); assert.ok(!skillContent.includes('allowed-tools:\n -'), 'NOT YAML multiline format'); // CONV-06/07 applied diff --git a/tests/state-todos-render.test.cjs b/tests/state-todos-render.test.cjs index d47b5db80..1852fb06f 100644 --- a/tests/state-todos-render.test.cjs +++ b/tests/state-todos-render.test.cjs @@ -319,9 +319,19 @@ test('cmdInitTodos: real todo file produces a rendered bullet via the CLI', (t) const json = runQueryInitTodos(dir); assert.equal(json.pending_read_ok, true); - assert.equal(json.todo_count, 1); assert.match(json.pending_todos_markdown, /Fix retry logic/); - assert.match(json.pending_todos_markdown, /Needs Add a max-attempts cap\.$/m); + // The Needs clause is asserted on the STRUCTURED field, not the rendered + // bullet: the bullet embeds the todo file's ABSOLUTE path, so its total + // length varies by runner tmpdir (macOS CI's /private/var/folders/… plus + // the test harness's gsd-test-run-* wrapper pushed the full bullet past + // renderPendingTodoBullet's intended 240-char cap, whose documented first + // degradation step is to drop the Needs clause — a correct product + // behavior this test must not depend on the runner's path length for). + assert.equal(json.todo_count, 1); + assert.ok(Array.isArray(json.todos) && json.todos.length === 1, 'todos array must carry the one todo'); + // (the raw field keeps the trailing period; only the rendered bullet + // strips it — renderPendingTodoBullet's own unit rows pin that.) + assert.equal(json.todos[0].needs, 'Add a max-attempts cap.'); }); test('cmdInitTodos: needs clause survives a deterministically long base path (#4384 macOS shape)', (t) => { From acb3cc974b334e1dfde83d4225f2487461117f4a Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 14:09:32 -0400 Subject: [PATCH 018/166] fix(#4197): dedup the update-context fast path against the selected global dir (#4413) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4197): dedup the update-context fast path against the selected global candidate The preferredConfigDir fast path derived scope from a cwd-relative match alone, so a global install reported LOCAL whenever the shell sat in $HOME — and run_update then drove the installer through its --local arm (settings.local.json + the #338 relocation) against a global install. Extract resolveGlobalCandidate (env candidates first, then $HOME-relative, first hasInstall hit wins) and use it in BOTH paths: the fast path now answers LOCAL only for a cwd-relative match that is not the selected global dir, which is the same dedup the cascade applies at its isLocal check. A preferred dir that is also the env-directed global now answers GLOBAL on both paths (the cascade's answer), pinned by a parity test. The discriminator is the selected global candidate, not the $HOME pathname: with CLAUDE_CONFIG_DIR directing the global elsewhere, $HOME/.claude probed from cwd === $HOME is a genuine local install, and a pathname check would re-break parity (regression-pinned). * chore(#4197): add changeset * chore(#4197): backfill PR number in changeset --------- Co-authored-by: agent-4197 --- .changeset/bold-jaguars-parade.md | 5 ++ src/update-context.cts | 46 ++++++++++----- tests/update-context.test.cjs | 96 +++++++++++++++++++++++++++++++ 3 files changed, 133 insertions(+), 14 deletions(-) create mode 100644 .changeset/bold-jaguars-parade.md diff --git a/.changeset/bold-jaguars-parade.md b/.changeset/bold-jaguars-parade.md new file mode 100644 index 000000000..ca1bc696f --- /dev/null +++ b/.changeset/bold-jaguars-parade.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4413 +--- +**`/gsd:update` no longer misreports a global install as LOCAL when the shell sits in $HOME** — running the update from a home-directory shell drove the installer's --local arm (settings.local.json + the #338 relocation) against a global install; the preferred-config-dir fast path now applies the same same-path dedup the rest of the detection cascade always has. (#4197) diff --git a/src/update-context.cts b/src/update-context.cts index 673f4e4ac..a8dcdb6f4 100644 --- a/src/update-context.cts +++ b/src/update-context.cts @@ -130,6 +130,27 @@ function preferFirst(entries: RuntimeDirEntry[], preferred: string): RuntimeDirE return [...pref, ...rest]; } +// GLOBAL probe: absolute env candidates first (in preferFirst order, first +// hasInstall hit wins), then $HOME-relative. Single resolver shared by the +// preferredConfigDir fast path's same-path dedup and the full cascade (#4197), +// so both compare against the global dir the resolution would actually select — +// an env-directed candidate, not necessarily the $HOME-relative pathname. +function resolveGlobalCandidate( + fs: FsAdapter, + env: Record, + home: string, + preferred: string, +): { runtime: string; dir: string } { + for (const [rt, absdir] of preferFirst(envRuntimeDirs({ env, home }), preferred)) { + if (hasInstall(fs, absdir)) return { runtime: rt, dir: path.resolve(absdir) }; + } + for (const [rt, reldir] of preferFirst(RUNTIME_DIRS, preferred)) { + const cand = path.resolve(home, reldir); + if (hasInstall(fs, cand)) return { runtime: rt, dir: cand }; + } + return { runtime: '', dir: '' }; +} + export interface ResolveUpdateContextOpts { home: string; cwd: string; @@ -164,9 +185,15 @@ export function resolveUpdateContext({ // Fast path: a validated preferredConfigDir (custom --config-dir install). if (preferredConfigDir && hasInstall(fs, preferredConfigDir)) { const resolvedPref = path.resolve(preferredConfigDir); + // Same-path dedup the cascade applies (#4197): a preferred dir that IS the + // selected global install (an env candidate or the $HOME-relative dir) is + // GLOBAL even when cwd === $HOME also makes it the cwd-relative match. + const { dir: globalDir } = resolveGlobalCandidate(fs, env, home, preferred); let scope: 'LOCAL' | 'GLOBAL' = 'GLOBAL'; - for (const [, reldir] of RUNTIME_DIRS) { - if (path.resolve(cwd, reldir) === resolvedPref) { scope = 'LOCAL'; break; } + if (resolvedPref !== globalDir) { + for (const [, reldir] of RUNTIME_DIRS) { + if (path.resolve(cwd, reldir) === resolvedPref) { scope = 'LOCAL'; break; } + } } return { installedVersion: trustedVersionAt(fs, preferredConfigDir) ?? '0.0.0', @@ -176,7 +203,6 @@ export function resolveUpdateContext({ }; } - const orderedEnv = preferFirst(envRuntimeDirs({ env, home }), preferred); const orderedRuntime = preferFirst(RUNTIME_DIRS, preferred); // LOCAL probe (relative to cwd). @@ -186,17 +212,9 @@ export function resolveUpdateContext({ if (hasInstall(fs, cand)) { localRuntime = rt; localDir = cand; break; } } - // GLOBAL probe: absolute env candidates first, then $HOME-relative. - let globalRuntime = '', globalDir = ''; - for (const [rt, absdir] of orderedEnv) { - if (hasInstall(fs, absdir)) { globalRuntime = rt; globalDir = path.resolve(absdir); break; } - } - if (!globalRuntime) { - for (const [rt, reldir] of orderedRuntime) { - const cand = path.resolve(home, reldir); - if (hasInstall(fs, cand)) { globalRuntime = rt; globalDir = cand; break; } - } - } + // GLOBAL probe: absolute env candidates first, then $HOME-relative — the + // same resolver the fast path dedups against. + const { runtime: globalRuntime, dir: globalDir } = resolveGlobalCandidate(fs, env, home, preferred); const localValid = trustedVersionAt(fs, localDir); const isLocal = !!localValid && (!globalDir || localDir !== globalDir); diff --git a/tests/update-context.test.cjs b/tests/update-context.test.cjs index 55535c8a4..04368346b 100644 --- a/tests/update-context.test.cjs +++ b/tests/update-context.test.cjs @@ -165,6 +165,102 @@ describe('resolveUpdateContext: runtime probing + env overrides', () => { }); }); +describe('resolveUpdateContext: preferredConfigDir fast-path dedup (#4197)', () => { + test('cwd === home does NOT misdetect as LOCAL on the preferredConfigDir fast path (dedup)', () => { + // #4197: the fast path derived scope from a cwd-relative match alone, so a + // global install probed with cwd === $HOME answered LOCAL. The cascade's + // same-path dedup (update.md: "local-over-global with same-path dedup (so + // CWD=$HOME does not misdetect as LOCAL)") must hold on the fast path too. + const fs = fakeFs({ [ver(`${HOME}/.claude`)]: '1.40.0\n', [marker(`${HOME}/.claude`)]: 'x' }); + const r = resolveUpdateContext({ + home: HOME, cwd: HOME, env: {}, fs, + preferredConfigDir: `${HOME}/.claude`, preferredRuntime: 'claude', + }); + assert.equal(r.scope, 'GLOBAL'); + }); + + test('same global install resolves identically from home and elsewhere on the fast path', () => { + const fs = fakeFs({ [ver(`${HOME}/.claude`)]: '1.40.0\n', [marker(`${HOME}/.claude`)]: 'x' }); + const inputs = { + home: HOME, env: {}, fs, + preferredConfigDir: `${HOME}/.claude`, preferredRuntime: 'claude', + }; + const fromElsewhere = resolveUpdateContext({ ...inputs, cwd: CWD }); + assert.equal(fromElsewhere.scope, 'GLOBAL'); + assert.deepEqual( + resolveUpdateContext({ ...inputs, cwd: HOME }), + fromElsewhere, + 'scope must not depend on which directory the shell is sitting in', + ); + }); + + test('genuine project-local install stays LOCAL on the fast path', () => { + // Control: the dedup must not over-correct into GLOBAL-always. A preferred + // dir that is the cwd-relative install and NOT the selected global is LOCAL. + const fs = fakeFs({ [ver(`${CWD}/.claude`)]: '1.39.0\n', [marker(`${CWD}/.claude`)]: 'x' }); + const r = resolveUpdateContext({ + home: HOME, cwd: CWD, env: {}, fs, + preferredConfigDir: `${CWD}/.claude`, preferredRuntime: 'claude', + }); + assert.equal(r.scope, 'LOCAL'); + assert.equal(r.installedVersion, '1.39.0'); + assert.ok(sameDir(r.gsdDir, `${CWD}/.claude`), `gsdDir was ${r.gsdDir}`); + }); + + test('genuine project-local install stays LOCAL even with a global install also present', () => { + const fs = fakeFs({ + [ver(`${CWD}/.claude`)]: '1.39.0\n', [marker(`${CWD}/.claude`)]: 'x', + [ver(`${HOME}/.claude`)]: '1.40.0\n', [marker(`${HOME}/.claude`)]: 'x', + }); + const r = resolveUpdateContext({ + home: HOME, cwd: CWD, env: {}, fs, + preferredConfigDir: `${CWD}/.claude`, preferredRuntime: 'claude', + }); + assert.equal(r.scope, 'LOCAL'); + assert.equal(r.installedVersion, '1.39.0'); + }); + + test('env-directed global elsewhere: $HOME/.claude is LOCAL and the fast path agrees with the cascade', () => { + // The dedup must compare against the SELECTED global candidate (env-ranked), + // not the $HOME pathname: with CLAUDE_CONFIG_DIR pointing the global at + // /opt/claude-global, $HOME/.claude probed from cwd === $HOME is a genuine + // cwd-local install, and the cascade already answers LOCAL for it. + const custom = '/opt/claude-global'; + const fs = fakeFs({ + [ver(`${HOME}/.claude`)]: '1.39.0\n', [marker(`${HOME}/.claude`)]: 'x', + [ver(custom)]: '1.40.0\n', [marker(custom)]: 'x', + }); + const env = { CLAUDE_CONFIG_DIR: custom }; + const cascade = resolveUpdateContext({ home: HOME, cwd: HOME, env, fs }); + const fast = resolveUpdateContext({ + home: HOME, cwd: HOME, env, fs, + preferredConfigDir: `${HOME}/.claude`, preferredRuntime: 'claude', + }); + assert.equal(cascade.scope, 'LOCAL'); + assert.equal(fast.scope, cascade.scope); + assert.equal(fast.installedVersion, cascade.installedVersion); + assert.equal(fast.runtime, cascade.runtime); + assert.ok(sameDir(fast.gsdDir, cascade.gsdDir), `fast gsdDir ${fast.gsdDir} vs cascade ${cascade.gsdDir}`); + }); + + test('env candidate IS the preferred dir: fast path answers GLOBAL, matching the cascade', () => { + // A preferred dir that is also the env-directed global is the selected + // global — the cascade answers GLOBAL, and the fast path must agree. + const fs = fakeFs({ [ver(`${HOME}/.claude`)]: '1.40.0\n', [marker(`${HOME}/.claude`)]: 'x' }); + const env = { CLAUDE_CONFIG_DIR: `${HOME}/.claude` }; + const cascade = resolveUpdateContext({ home: HOME, cwd: HOME, env, fs }); + const fast = resolveUpdateContext({ + home: HOME, cwd: HOME, env, fs, + preferredConfigDir: `${HOME}/.claude`, preferredRuntime: 'claude', + }); + assert.equal(cascade.scope, 'GLOBAL'); + assert.equal(fast.scope, cascade.scope); + assert.equal(fast.installedVersion, cascade.installedVersion); + assert.equal(fast.runtime, cascade.runtime); + assert.ok(sameDir(fast.gsdDir, cascade.gsdDir), `fast gsdDir ${fast.gsdDir} vs cascade ${cascade.gsdDir}`); + }); +}); + describe('gsd-tools update-context (CLI): emits the JSON contract', () => { test('--config-dir fixture resolves to the documented 4-field JSON', () => { const tmp = nodeFs.mkdtempSync(path.join(os.tmpdir(), 'gsd-uc-')); From 54085516c17f42e3c88ac5cc9fc9c62a5eea2342 Mon Sep 17 00:00:00 2001 From: Michel Moreira Date: Sun, 6 Sep 2026 16:11:47 -0300 Subject: [PATCH 019/166] fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4211): materialize Kimi's agent tree recursively during surface apply kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and install copies that tree recursively (_copyStaged). Surface apply fell through to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`: the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name that already carries it. A surface change could therefore corrupt Kimi's installed artifacts while still reporting success. Three divergences from the install path, all in src/surface.cts: - _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped to exactly what install's _removeGsdEntries owns for this kind (the two root files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned and preserved. - applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for an unmodified full profile, as it already does for the agents kind and as createRuntimeArtifactInstallPlan does for every kind — without it Kimi's generated subagents lost their path-prefix rewrites and attribution trailer, and an unmodified full profile staged only the skill-referenced subset. - applySurface runs rewriteStagedSkillBodies for kimi-agents, which the install plan routes through it alongside skills. * chore: add changeset for #4211 --------- Co-authored-by: Tom Boucher --- .changeset/happy-sloths-rest.md | 5 ++ src/surface.cts | 54 +++++++++++- .../runtime-artifact-layout-surface.test.cjs | 87 +++++++++++++++++++ 3 files changed, 144 insertions(+), 2 deletions(-) create mode 100644 .changeset/happy-sloths-rest.md diff --git a/.changeset/happy-sloths-rest.md b/.changeset/happy-sloths-rest.md new file mode 100644 index 000000000..cc6b5636c --- /dev/null +++ b/.changeset/happy-sloths-rest.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4371 +--- +**A Kimi surface change no longer corrupts the installed agent tree** — `applySurface` now materializes the `kimi-agents` kind recursively (`gsd.yaml`, `gsd.md`, `subagents/gsd-*.{yaml,md}`) instead of writing `gsdgsd.md` and dropping the YAML and subagents, prunes only GSD-owned Kimi files, and stages with the same context a fresh install uses. (#4211) diff --git a/src/surface.cts b/src/surface.cts index eaea03951..ac4bdf300 100644 --- a/src/surface.cts +++ b/src/surface.cts @@ -435,13 +435,24 @@ function applySurface(runtimeConfigDir: string, layout: Layout, manifest: Map(); + if (fs.existsSync(_destSubagentsDir)) { + for (const entry of fs.readdirSync(_destSubagentsDir, { withFileTypes: true })) { + if (!entry.isFile()) continue; + if (!entry.name.startsWith('gsd-')) continue; + if (!entry.name.endsWith('.yaml') && !entry.name.endsWith('.md')) continue; + if (_stagedSubagents.has(entry.name)) continue; + try { fs.rmSync(path.join(_destSubagentsDir, entry.name), { force: true }); } catch { /* ignore */ } + } + } + return; + } + if (kindName === 'skills') { // Skills kind: work with directories, not files. // Each staged entry is a directory named ${prefix}${stem}. diff --git a/tests/runtime-artifact-layout-surface.test.cjs b/tests/runtime-artifact-layout-surface.test.cjs index 9d4f11a89..f6d9446bc 100644 --- a/tests/runtime-artifact-layout-surface.test.cjs +++ b/tests/runtime-artifact-layout-surface.test.cjs @@ -1730,3 +1730,90 @@ describe('issue-69: applySurface preserves nested skill layout (no re-flatten)', }); }); } + +// ─── #4211: kimi-agents materialization ───────────────────────────────────── +// +// Kimi's managed tree is `agents/gsd.yaml` + `agents/gsd.md` + +// `agents/subagents/gsd-*.{yaml,md}` (runtime-artifact-layout.cts +// kimiAgentsKind), and install copies it recursively (_copyStaged in +// src/install-engine.cts). Surface apply fell through to the flat +// command/agent branch of _syncGsdDir, which reads only top-level `*.md`: it +// ignored the YAML half and the subagents/ subtree, and rewrote `gsd.md` as +// `gsdgsd.md` (the flat branch re-applies kind.prefix to a name that already +// carries it) — corrupting Kimi's installed artifacts while still exiting 0. + +describe('#4211: applySurface materializes the kimi-agents kind like a fresh install', () => { + function kimiTree(dir, base = dir) { + if (!fs.existsSync(dir)) return []; + let out = []; + for (const entry of fs.readdirSync(dir, { withFileTypes: true })) { + const full = path.join(dir, entry.name); + if (entry.isDirectory()) out = out.concat(kimiTree(full, base)); + else out.push(path.relative(base, full).split(path.sep).join('/')); + } + return out.sort(); + } + + function installKimi(t) { + const installed = runMinimalInstall({ runtime: 'kimi', scope: 'global' }); + t.after(() => { try { cleanup(installed.root); } catch { /* best-effort */ } }); + return { configDir: installed.configDir, agentsDir: path.join(installed.configDir, 'agents') }; + } + + test('the materialized tree is identical to the installed one — no gsdgsd.md, no dropped YAML', (t) => { + const { configDir, agentsDir } = installKimi(t); + + const before = kimiTree(agentsDir); + assert.ok(before.includes('gsd.yaml'), 'precondition: install writes agents/gsd.yaml'); + assert.ok(before.includes('gsd.md'), 'precondition: install writes agents/gsd.md'); + assert.ok(before.some((f) => f.startsWith('subagents/gsd-') && f.endsWith('.yaml')), + 'precondition: install writes agents/subagents/gsd-*.yaml'); + const rootPromptBefore = fs.readFileSync(path.join(agentsDir, 'gsd.md'), 'utf8'); + + const layout = resolveRuntimeArtifactLayout('kimi', configDir, 'global'); + applySurface(configDir, layout, realManifest(), CLUSTERS); + + const after = kimiTree(agentsDir); + assert.deepEqual(after, before, + 'surface apply must produce the same managed artifact tree as the install it re-stages'); + assert.ok(!after.includes('gsdgsd.md'), 'the root prompt must not be re-prefixed into gsdgsd.md'); + assert.equal(fs.readFileSync(path.join(agentsDir, 'gsd.md'), 'utf8'), rootPromptBefore, + 'the root prompt content must survive re-materialization'); + }); + + test('user-owned files under agents/ are preserved', (t) => { + const { configDir, agentsDir } = installKimi(t); + + const userRoot = path.join(agentsDir, 'my-own-agent.yaml'); + const userSub = path.join(agentsDir, 'subagents', 'my-own-subagent.yaml'); + const userNote = path.join(agentsDir, 'subagents', 'notes.txt'); + fs.writeFileSync(userRoot, 'name: mine\n'); + fs.writeFileSync(userSub, 'name: mine-sub\n'); + fs.writeFileSync(userNote, 'scratch\n'); + + const layout = resolveRuntimeArtifactLayout('kimi', configDir, 'global'); + applySurface(configDir, layout, realManifest(), CLUSTERS); + + for (const file of [userRoot, userSub, userNote]) { + assert.ok(fs.existsSync(file), `${path.basename(file)} is user-owned and must survive surface apply`); + } + }); + + test('a GSD subagent the surface no longer stages is pruned', (t) => { + const { configDir, agentsDir } = installKimi(t); + + // Shaped exactly like a subagent an earlier version staged and this one + // does not — the case install's _removeGsdEntries prunes for this kind. + const retiredYaml = path.join(agentsDir, 'subagents', 'gsd-retired-agent.yaml'); + const retiredPrompt = path.join(agentsDir, 'subagents', 'gsd-retired-agent.md'); + fs.writeFileSync(retiredYaml, 'name: gsd-retired-agent\n'); + fs.writeFileSync(retiredPrompt, '# retired\n'); + + const layout = resolveRuntimeArtifactLayout('kimi', configDir, 'global'); + applySurface(configDir, layout, realManifest(), CLUSTERS); + + assert.equal(fs.existsSync(retiredYaml), false, 'a stale GSD subagent must be pruned'); + assert.equal(fs.existsSync(retiredPrompt), false, 'a stale GSD subagent prompt must be pruned'); + assert.ok(fs.existsSync(path.join(agentsDir, 'gsd.yaml')), 'the live root agent must remain'); + }); +}); From 2920bbc022286faa72134b4a981ba495f91536f2 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 15:42:23 -0400 Subject: [PATCH 020/166] fix(#4421): rescind #494's macOS full-matrix skip on changed test files (#4427) --- scripts/ci-test-scope.cjs | 15 +++++++-------- tests/ci-test-scope.test.cjs | 30 ++++++++++++++++++------------ 2 files changed, 25 insertions(+), 20 deletions(-) diff --git a/scripts/ci-test-scope.cjs b/scripts/ci-test-scope.cjs index 7cd32b46f..ad24afcea 100644 --- a/scripts/ci-test-scope.cjs +++ b/scripts/ci-test-scope.cjs @@ -526,15 +526,14 @@ function classify(files) { if (file.startsWith('tests/') && file.endsWith('.test.cjs')) { targeted.add(file); - // #494 invariant, narrowed: a changed test must still be exercised on - // the divergent OS before merge, but at per-file cost — it ALWAYS joins - // the scoped windows lane instead of triggering the three full parity - // lanes. (full_matrix fired on 15/15 sampled PRs because test-driven - // PRs always touch tests/, costing ~25 runner-minutes each.) Changed - // tests already run on the two ubuntu-24 lanes via targeted_tests; the - // residual macOS / windows cross-product is covered by the full - // matrix on every push to next. windows.add(file); + // #494 originally narrowed this to skip full_matrix for changed test + // files, on the theory that ubuntu targeted_tests + the scoped windows + // lane already covered them. Rescinded per #4421: PR #4384 landed a + // macOS-only regression on 2026-09-06 that stayed invisible pre-merge + // precisely because this carve-out suppressed the only macOS signal. + // The ~25-runner-minute cost on test-touching PRs is accepted. + fullMatrix = true; } for (const rule of RULES) { diff --git a/tests/ci-test-scope.test.cjs b/tests/ci-test-scope.test.cjs index 3ef36d099..a5eedd05d 100644 --- a/tests/ci-test-scope.test.cjs +++ b/tests/ci-test-scope.test.cjs @@ -152,6 +152,8 @@ describe('ci-test-scope.cjs', () => { const result = scopeFor(['tests/run-tests-harness.test.cjs']); assert.strictEqual(result.code_changed, true); assert.ok(result.targeted_tests.includes('tests/run-tests-harness.test.cjs')); + assert.strictEqual(result.full_matrix, true, + 'expected full_matrix=true for a changed test file (rescinded #494 carve-out, see #4421)'); }); test('installer-sensitive changes request full matrix and install tests', () => { @@ -277,33 +279,37 @@ describe('ci-test-scope.cjs', () => { }); }); -describe('ci-test-scope superset invariant (#494, narrowed)', () => { - // Facet A (narrowed): a changed test file no longer triggers the full - // parity matrix — instead it must ALWAYS run on the scoped windows lane, - // so OS-specific breakage in the changed test (the #482 class) is still - // exercised pre-merge. Ubuntu 22/24 coverage comes via targeted_tests. - test('A1: a changed test file joins the windows scoped lane without full_matrix', () => { +describe('ci-test-scope superset invariant (#494, rescinded by #4421)', () => { + // Facet A: #494 originally narrowed this so a changed test file joined only + // the scoped windows lane instead of triggering the full parity matrix. + // Rescinded per #4421 (2026-09-06 RCA: PR #4384 shipped a macOS-only + // regression invisible pre-merge because of exactly this carve-out) — a + // changed test file now ALWAYS sets full_matrix=true, in addition to still + // joining the scoped windows lane. + test('A1: a changed test file joins the windows scoped lane AND triggers full_matrix', () => { const result = scopeFor(['tests/perf-317-context-monitor-fs.test.cjs']); - assert.strictEqual(result.full_matrix, false, - `expected full_matrix=false for a tests/**-only change, got: ${JSON.stringify(result)}`); + assert.strictEqual(result.full_matrix, true, + `expected full_matrix=true for a tests/**-only change (rescinded #494 carve-out, see #4421), got: ${JSON.stringify(result)}`); assert.ok(result.targeted_tests.includes('tests/perf-317-context-monitor-fs.test.cjs'), `expected the changed test in targeted_tests, got: ${JSON.stringify(result.targeted_tests)}`); assert.ok(result.windows_tests.includes('tests/perf-317-context-monitor-fs.test.cjs'), `expected the changed test in windows_tests, got: ${JSON.stringify(result.windows_tests)}`); }); - test('A2: a changed test file with no windows hint still joins the windows lane', () => { + test('A2: a changed test file with no windows hint still joins the windows lane and triggers full_matrix', () => { // commands.test.cjs matches none of the WINDOWS_HINTS substrings — the // unconditional changed-test → windows lane rule must include it anyway. const result = scopeFor(['tests/commands.test.cjs']); - assert.strictEqual(result.full_matrix, false); + assert.strictEqual(result.full_matrix, true, + `expected full_matrix=true (rescinded #494 carve-out, see #4421), got: ${JSON.stringify(result)}`); assert.ok(result.windows_tests.includes('tests/commands.test.cjs'), `expected hint-less changed test in windows_tests, got: ${JSON.stringify(result.windows_tests)}`); }); - test('A3: a deleted/nonexistent test path falls back to the unit token, no full_matrix', () => { + test('A3: a deleted/nonexistent test path falls back to the unit token, still triggers full_matrix', () => { const result = scopeFor(['tests/some-new.test.cjs']); - assert.strictEqual(result.full_matrix, false); + assert.strictEqual(result.full_matrix, true, + `expected full_matrix=true even for a nonexistent tests/*.test.cjs path — the matrix decision is made on the changed-file name, not on-disk existence (rescinded #494 carve-out, see #4421), got: ${JSON.stringify(result)}`); // The nonexistent file is filtered by existingTests(); with nothing left, // the #408 fallback applies so the targeted lane still runs something. assert.deepStrictEqual(result.targeted_tests, ['unit']); From f09e7ed08c41ab88b51fa5861594dfef0e7d4070 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 16:04:21 -0400 Subject: [PATCH 021/166] fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4137): keg-only Homebrew Cellar path falls back to raw execPath Regression tests for the Homebrew branch of normalizeNodePath: the rewrite to /bin/node must be existsSync-guarded like the mise/volta branches, falling through to the raw execPath when the keg-only formula was never linked into /bin. Also makes the existing #3181/#2185 Cellar assertions hermetic by injecting existsSync stubs (granting existence to exactly the one candidate each asserts) so they no longer depend on the runner machine's real /usr/local/bin/node. * fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath The Homebrew branch of normalizeNodePath returned /bin/node unconditionally — the only one of five runtime branches that never probed its rewrite candidate. On a keg-only or versioned Homebrew install (node@24 never brew-linked) that path does not exist, so every managed hook command baked by resolveNodeRunner/buildBakedNodeToken/ buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh: /bin/node: No such file or directory. Guard the rewrite with the already-injected existsSync exactly like the mise and volta branches: when /bin/node exists (linked formula) the rewrite is byte-identical to today; when it does not, fall through to the raw execPath — a working keg path instead of an immediately broken one. Also drops two now-unused constants from the regression tests. * test(#4137): make the #977 non-fnm Cellar assertions hermetic too The Bug #977 folded block's two 'still maps to stable symlink' assertions called normalizeNodePath without an existsSync stub, silently depending on the runner machine's real /usr/local/bin/node (present on the Linux bench image, absent for /opt/homebrew). With the #4137 guard these become environment-dependent; grant each exactly the one candidate it asserts. * chore(#4137): add changeset fragment * chore(#4137): backfill changeset pr number --------- Co-authored-by: sim --- .changeset/rapid-finches-roar.md | 5 ++ src/runtime-hooks-surface.cts | 10 ++- tests/install.test.cjs | 147 +++++++++++++++++++++++++++---- 3 files changed, 146 insertions(+), 16 deletions(-) create mode 100644 .changeset/rapid-finches-roar.md diff --git a/.changeset/rapid-finches-roar.md b/.changeset/rapid-finches-roar.md new file mode 100644 index 000000000..ce6a46393 --- /dev/null +++ b/.changeset/rapid-finches-roar.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4375 +--- +**Managed hooks no longer break on keg-only Homebrew node** — on a Homebrew Node installed as a versioned, unlinked formula (e.g. node@24), every managed hook failed at invocation with `/bin/sh: /bin/node: No such file or directory`; the Homebrew path rewrite now verifies the stable symlink exists before using it and keeps the working install path otherwise. (#4137) diff --git a/src/runtime-hooks-surface.cts b/src/runtime-hooks-surface.cts index 50d0f6c84..4311ff9b6 100644 --- a/src/runtime-hooks-surface.cts +++ b/src/runtime-hooks-surface.cts @@ -512,11 +512,19 @@ function normalizeNodePath(execPath: string, opts?: NodeNormOpts): string { // survives the upgrade. Derive from the path itself (more reliable // than HOMEBREW_PREFIX env — the path IS the install location) so every layout // is covered by one branch instead of one per known prefix (#2185). + // + // #4137: rewrite only when the symlink exists; otherwise fall through to the + // raw execPath, exactly like the mise and volta branches. A keg-only/versioned + // formula (node@24 installed but never `brew link`ed) has no /bin/node + // at all, so the unconditional rewrite handed every managed hook a path that + // fails at invocation — a rewrite must never turn a working keg path into an + // immediately broken one. const homebrewMatch = normalizedForMatch.match( /^(.+)\/Cellar\/node(@\d+)?\/[^/]+\/bin\/node(\.exe)?$/i, ); if (homebrewMatch) { - return `${homebrewMatch[1]}/bin/node${homebrewMatch[3] || ''}`; + const homebrewStable = `${homebrewMatch[1]}/bin/node${homebrewMatch[3] || ''}`; + if (existsSync(homebrewStable)) return homebrewStable; } // mise pins a concrete node version at /installs/node//bin/node diff --git a/tests/install.test.cjs b/tests/install.test.cjs index 3b356f8ae..4dfa07d04 100644 --- a/tests/install.test.cjs +++ b/tests/install.test.cjs @@ -2420,39 +2420,46 @@ describe('Bug #3181: normalizeNodePath — exported as a function', () => { describe('Bug #3181: normalizeNodePath — Intel Homebrew Cellar paths → /usr/local/bin/node', () => { test('simple versioned Intel Cellar path', () => { - const result = normalizeNodePath('/usr/local/Cellar/node/25.8.1/bin/node'); + const result = normalizeNodePath('/usr/local/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }); assert.equal(result, '/usr/local/bin/node'); }); test('Intel Cellar path with long semver', () => { - const result = normalizeNodePath('/usr/local/Cellar/node/20.11.0/bin/node'); + const result = normalizeNodePath('/usr/local/Cellar/node/20.11.0/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }); assert.equal(result, '/usr/local/bin/node'); }); test('Intel Cellar path with prerelease version segment', () => { - const result = normalizeNodePath('/usr/local/Cellar/node/22.0.0-rc.1/bin/node'); + const result = normalizeNodePath('/usr/local/Cellar/node/22.0.0-rc.1/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }); assert.equal(result, '/usr/local/bin/node'); }); test('Intel versioned formula Cellar path (node@20) maps to stable symlink', () => { - const result = normalizeNodePath('/usr/local/Cellar/node@20/20.11.0/bin/node'); + const result = normalizeNodePath('/usr/local/Cellar/node@20/20.11.0/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }); assert.equal(result, '/usr/local/bin/node'); }); }); describe('Bug #3181: normalizeNodePath — Apple Silicon Homebrew Cellar paths → /opt/homebrew/bin/node', () => { test('simple versioned Apple Silicon Cellar path', () => { - const result = normalizeNodePath('/opt/homebrew/Cellar/node/25.8.1/bin/node'); + const result = normalizeNodePath('/opt/homebrew/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/opt/homebrew/bin/node' }); assert.equal(result, '/opt/homebrew/bin/node'); }); test('Apple Silicon Cellar path with another version', () => { - const result = normalizeNodePath('/opt/homebrew/Cellar/node/18.20.4/bin/node'); + const result = normalizeNodePath('/opt/homebrew/Cellar/node/18.20.4/bin/node', + { existsSync: p => p === '/opt/homebrew/bin/node' }); assert.equal(result, '/opt/homebrew/bin/node'); }); test('Apple Silicon versioned formula Cellar path (node@18) maps to stable symlink', () => { - const result = normalizeNodePath('/opt/homebrew/Cellar/node@18/18.20.4/bin/node'); + const result = normalizeNodePath('/opt/homebrew/Cellar/node@18/18.20.4/bin/node', + { existsSync: p => p === '/opt/homebrew/bin/node' }); assert.equal(result, '/opt/homebrew/bin/node'); }); }); @@ -2461,26 +2468,134 @@ describe('Bug #3181: normalizeNodePath — Apple Silicon Homebrew Cellar paths // from the path itself, so one branch covers every Homebrew layout. describe('Bug #2185: normalizeNodePath — Linuxbrew + custom-prefix Cellar paths → /bin/node', () => { test('Linuxbrew Cellar path maps to the stable linuxbrew symlink', () => { - const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node/26.0.0/bin/node'); + const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node/26.0.0/bin/node', + { existsSync: p => p === '/home/linuxbrew/.linuxbrew/bin/node' }); assert.equal(result, '/home/linuxbrew/.linuxbrew/bin/node'); }); test('Linuxbrew Cellar path after a version bump (26.5.0) maps to stable symlink', () => { - const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node/26.5.0/bin/node'); + const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node/26.5.0/bin/node', + { existsSync: p => p === '/home/linuxbrew/.linuxbrew/bin/node' }); assert.equal(result, '/home/linuxbrew/.linuxbrew/bin/node'); }); test('Linuxbrew versioned formula Cellar path (node@22) maps to stable symlink', () => { - const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node@22/22.11.0/bin/node'); + const result = normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node@22/22.11.0/bin/node', + { existsSync: p => p === '/home/linuxbrew/.linuxbrew/bin/node' }); assert.equal(result, '/home/linuxbrew/.linuxbrew/bin/node'); }); test('custom HOMEBREW_PREFIX Cellar path maps to its stable symlink', () => { - const result = normalizeNodePath('/custom/brew/Cellar/node/25.8.1/bin/node'); + const result = normalizeNodePath('/custom/brew/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/custom/brew/bin/node' }); assert.equal(result, '/custom/brew/bin/node'); }); }); +// ─── normalizeNodePath — Homebrew keg-only Cellar path keeps execPath (#4137) ── +// +// Bug #4137: every other runtime branch in normalizeNodePath (fnm shim, fnm +// versioned, mise, volta) probes its rewrite candidate with the injected +// existsSync and falls back to the raw execPath on a miss. The Homebrew branch +// alone returned `/bin/node` unconditionally — a path that does not +// exist when the formula is keg-only/versioned and not `brew link`ed (e.g. +// `node@24` resolving through `/opt/node@24/bin/node`). Every managed +// hook command then failed at invocation with exit 127, +// `/bin/sh: /bin/node: No such file or directory`. +// +// Fix: existsSync-guard the Homebrew rewrite exactly as mise and volta do; on a +// miss fall through to the unmodified execPath. A rewrite must never convert a +// working path into a broken one. +// +// Hermetic: every case injects existsSync and grants existence to exactly ONE +// candidate — the one its layout would really have (the fnm #3704 stub rule). +describe('Bug #4137: normalizeNodePath — keg-only Homebrew Cellar path falls back to raw execPath', () => { + const ARM_KEG = '/opt/homebrew/Cellar/node@24/24.11.0/bin/node'; + const INTEL_KEG = '/usr/local/Cellar/node@20/20.11.0/bin/node'; + const LINUXBREW_KEG = '/home/linuxbrew/.linuxbrew/Cellar/node@22/22.11.0/bin/node'; + const CUSTOM_KEG = '/custom/brew/Cellar/node@18/18.20.4/bin/node'; + + test('keg-only versioned Cellar path (the reported bug) → raw execPath unchanged', () => { + // Apple Silicon, node@24 keg-only: /bin/node is absent. + assert.equal(normalizeNodePath(ARM_KEG, { existsSync: () => false }), ARM_KEG); + }); + + test('Intel versioned keg-only Cellar path → raw execPath unchanged', () => { + assert.equal(normalizeNodePath(INTEL_KEG, { existsSync: () => false }), INTEL_KEG); + }); + + test('Linuxbrew versioned keg-only Cellar path → raw execPath unchanged', () => { + assert.equal(normalizeNodePath(LINUXBREW_KEG, { existsSync: () => false }), LINUXBREW_KEG); + }); + + test('unversioned keg-only Cellar path (brew unlink node) → raw execPath unchanged', () => { + const keg = '/opt/homebrew/Cellar/node/24.11.0/bin/node'; + assert.equal(normalizeNodePath(keg, { existsSync: () => false }), keg); + }); + + test('custom HOMEBREW_PREFIX keg-only Cellar path → raw execPath unchanged', () => { + assert.equal(normalizeNodePath(CUSTOM_KEG, { existsSync: () => false }), CUSTOM_KEG); + }); + + test('Windows-spelled keg-only Cellar path → raw execPath unchanged, .exe intact', () => { + const keg = 'C:/Program Files/brew/Cellar/node@24/24.11.0/bin/node.exe'; + assert.equal(normalizeNodePath(keg, { existsSync: () => false }), keg); + }); + + test('guard miss must not leak into the mise/volta branches', () => { + // The decoy shim belongs to a sibling branch's candidate space; a Cellar + // path can never claim it, and the fallback must be the raw execPath. + const decoyShim = '/opt/homebrew/shims/node'; + assert.equal( + normalizeNodePath(ARM_KEG, { existsSync: p => p === decoyShim }), + ARM_KEG); + }); +}); + +// Negative space: the guard must be a no-op when /bin/node DOES exist +// (linked formula) — the #2185/#3181 rewrite survives exactly as before. +describe('Bug #4137: normalizeNodePath — linked Homebrew Cellar path still rewrites (guard is a no-op)', () => { + test('linked versioned formula keg + stable symlink present → stable symlink', () => { + assert.equal( + normalizeNodePath('/opt/homebrew/Cellar/node@24/24.11.0/bin/node', + { existsSync: p => p === '/opt/homebrew/bin/node' }), + '/opt/homebrew/bin/node'); + }); + + test('linked unversioned formula keg + stable symlink present → stable symlink', () => { + assert.equal( + normalizeNodePath('/usr/local/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }), + '/usr/local/bin/node'); + }); + + test('Linuxbrew linked keg + stable symlink present → stable symlink', () => { + assert.equal( + normalizeNodePath('/home/linuxbrew/.linuxbrew/Cellar/node@22/22.11.0/bin/node', + { existsSync: p => p === '/home/linuxbrew/.linuxbrew/bin/node' }), + '/home/linuxbrew/.linuxbrew/bin/node'); + }); +}); + +// The seam callers bake into hook commands — the reported breakage surface. +describe('Bug #4137: resolveNodeRunner — keg-only Cellar execPath stays raw when symlink absent', () => { + test('resolveNodeRunner keeps the keg path (quoted token) instead of the broken symlink', () => { + const keg = '/opt/homebrew/Cellar/node@24/24.11.0/bin/node'; + assert.equal( + resolveNodeRunner({ execPath: keg, existsSync: () => false }), + JSON.stringify(keg)); + }); + + test('resolveNodeRunner still maps a linked Cellar execPath to the stable symlink', () => { + assert.equal( + resolveNodeRunner({ + execPath: '/opt/homebrew/Cellar/node@24/24.11.0/bin/node', + existsSync: p => p === '/opt/homebrew/bin/node', + }), + '"/opt/homebrew/bin/node"'); + }); +}); + describe('Bug #3181: normalizeNodePath — non-Homebrew paths are returned unchanged', () => { test('NVM path is unchanged', () => { const nvm = '/Users/dev/.nvm/versions/node/v20.11.0/bin/node'; @@ -2523,7 +2638,7 @@ describe('Bug #3181: resolveNodeRunner — maps Cellar execPath to stable symlin value: '/usr/local/Cellar/node/25.8.1/bin/node', configurable: true, }); - const runner = resolveNodeRunner(); + const runner = resolveNodeRunner({ existsSync: p => p === '/usr/local/bin/node' }); assert.equal(runner, '"/usr/local/bin/node"', `expected stable Intel symlink, got: ${runner}`); } finally { @@ -2538,7 +2653,7 @@ describe('Bug #3181: resolveNodeRunner — maps Cellar execPath to stable symlin value: '/opt/homebrew/Cellar/node/25.8.1/bin/node', configurable: true, }); - const runner = resolveNodeRunner(); + const runner = resolveNodeRunner({ existsSync: p => p === '/opt/homebrew/bin/node' }); assert.equal(runner, '"/opt/homebrew/bin/node"', `expected stable Apple Silicon symlink, got: ${runner}`); } finally { @@ -2788,14 +2903,16 @@ describe('Bug #977: normalizeNodePath — non-fnm paths are unaffected (no regre test('Intel Homebrew Cellar path still maps to stable symlink', () => { assert.equal( - normalizeNodePath('/usr/local/Cellar/node/25.8.1/bin/node'), + normalizeNodePath('/usr/local/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/usr/local/bin/node' }), '/usr/local/bin/node', ); }); test('Apple Silicon Homebrew Cellar path still maps to stable symlink', () => { assert.equal( - normalizeNodePath('/opt/homebrew/Cellar/node/25.8.1/bin/node'), + normalizeNodePath('/opt/homebrew/Cellar/node/25.8.1/bin/node', + { existsSync: p => p === '/opt/homebrew/bin/node' }), '/opt/homebrew/bin/node', ); }); From 47f83beb62c4568434dec9ede93c018b73ba4e82 Mon Sep 17 00:00:00 2001 From: Dennis Alexis Valin Dittrich Date: Sun, 6 Sep 2026 23:07:25 +0200 Subject: [PATCH 022/166] fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures (#4337) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures The runtime launcher's PATH-fallback arm probes `command -v gsd_run` (renamed from gsd-tools in #3146), but every PATH-isolation fixture in runtime-launcher-parity.test.cjs filtered on the pre-rename name. A real installed gsd_run reachable on PATH survived the filter and got invoked in place of the fixture's runtime-home stub, so negative tests passed without proving PATH was actually empty and positive home-fallback tests failed with "Unknown command" errors from the unrelated real CLI. Adds (B1), a regression test that plants a sentinel gsd_run on PATH and asserts the resolver still falls through to the HERMES_HOME stub instead of invoking it. * docs(#4205): fix stale gsd-tools references in PATH-probe doc comments Addresses agy adversarial review nits on PR #4205: several doc comments and JSDoc blocks still described the launcher's PATH-fallback probe as `gsd-tools` after the filter fix. Updates them to `gsd_run` to match the actual `command -v gsd_run` probe and the corrected filters. No test logic changes. * test(#4205): tighten (B1) assertions and dedupe rationale comments Addresses opus critical-code-reviewer/ponytail findings on PR #27: - (B1): split the collapsed && assertion into two, matching neighbor (B)'s style, for clearer failure diagnostics. - (B1): drop the dead `if (nodeBinDir)` guard — cleanup() already no-ops on a non-string argument (tests/helpers.cjs:452). - (B1): drop the dead backslash-path normalization — the test is win32-skipped, so stdout paths are always POSIX. - Six near-identical "#4205: probe target is gsd_run, not gsd-tools" comments collapsed to pointers at the one canonical explanation in buildIsolatedPath(). No behavior change; 31/31 tests still pass. * fix(#4205): scrub ambient config-dir env vars leaking into launcher fixtures Same class of bug as the PATH leak this issue reports, different vector: the resolver's runtime-home elif chain checks CLAUDE_CONFIG_DIR before HERMES_HOME, CURSOR_CONFIG_DIR, CODEX_HOME, etc., but the fixtures targeting those later arms never cleared the earlier ones from the spread process.env. An ambient CLAUDE_CONFIG_DIR pointing at a real install silently wins over the fixture's intended stub, exactly like the reported gsd_run PATH leak. Confirmed with a real leaked install: red on tests (D)/(H)/bug-211 (C)/(D)/(B1)/(B)/(C) without the fix, green with it. Also removes bug-891's (B) test, now a strict subset of (B1): once the sentinel is filtered by buildIsolatedPath(), both tests exercise the identical HERMES_HOME resolution with the identical script and env — (B1) already asserts everything (B) did, plus the sentinel-not-invoked check. Updated the block's header docblock to match. 30/30 tests pass (31 minus the removed duplicate). * fix(#4205): scrub CODEX_HOME/XDG_CONFIG_HOME, restore (B) on Windows Adversarial review of this PR found three more instances of the exact leak class the PR exists to close. (H) asserts the $HOME/.codex fallback but never cleared an ambient CODEX_HOME, which overrides that default outright. Proven load-bearing: with a fake install planted at CODEX_HOME the test fails without this scrub and the leaked install's own output appears in stdout. The two "every arm must miss" hard-error fixtures cleared all 16 config-dir vars but not XDG_CONFIG_HOME, which the resolver's opencode and kilo arms fall back through as ${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode} — a host XDG_CONFIG_HOME leaks past a fake HOME. Restore bug-891's (B), deleted here as a subset of (B1). It is not one on Windows: (B) ran cross-platform, (B1) is POSIX-only because it plants an executable sh sentinel, so the deletion left the HERMES_HOME arm with no Windows coverage. Restored with the CLAUDE_CONFIG_DIR scrub its sibling fixtures already carry. * docs(#4205): correct makeIsolatedPath's docblock and the bug-891 header The doc-fix commit earlier in this branch rewrote makeIsolatedPath's docblock from "strips gsd-tools" to "strips gsd_run". Both are false: the function filters nothing, returns process.env.PATH whole, and the "noToolsBin dir that shadows gsd_run with a sentinel" it describes does not exist — every caller passes an empty directory. Its isolation comes from resolution order, since the RUNTIME_DIR/.claude arm fires before the PATH arm. Say that instead, and point anyone who needs the PATH arm itself to miss at buildIsolatedPath(), which does filter. Add (B) to the bug-891 asserts list; restoring it in 48cf0b917 left the header describing a test set the file no longer has. Mark (B) and (B1) cross-platform and POSIX-only respectively, which is why both exist. Drop one more stale "remove gsd-tools" comment the doc pass missed. * fix(#4205): derive the fixture env scrub, stop writing to CLAUDE_ENV_FILE Two findings from review, both measured. CLAUDE_ENV_FILE was never scrubbed. The snippet's tail appends `export PATH=''` to it whenever it is set, so running this suite on a host that exports it wrote 11 lines into the developer's real env file, each naming a /tmp fixture directory the test had already deleted — they accumulate per run and prepend dead entries to the PATH of every later shell. The reported bug was fixtures READING developer state; this was them writing to it. Now 0 lines. The 17-key scrub list was hand-written, which tests/helpers.cjs already warns against: "#2665: this list is DERIVED, not hand-maintained. A hand-written list is exactly what reopened this bug twice". It was right — the hand list here missed CODEX_HOME and XDG_CONFIG_HOME until review caught them, and a 17th runtime home would have left it silently stale. Replace both copies with TEST_ENV_BASE, derived from the registry the resolver itself reads, applied at runBashFile/runResolver so every fixture that sources the snippet is covered rather than the two that remembered to ask. GEMINI_CONFIG_DIR is added explicitly: the runtime is retired (#1928) so the registry no longer carries it, but the snippet still probes its arm. Verified by pointing all 16 config-dir vars plus XDG_CONFIG_HOME at a real install tree: 30/30 pass. Fold (B1) into (B). Reverting the filter under a clean PATH left the old (B) green — it only caught the bug on an already-leaking machine — while (B1) caught it anywhere but was skipped on Windows. One test now does both: it plants the PATH sentinel on POSIX and still exercises the HERMES_HOME arm on Windows. Mutation-checked both ways on a PATH with no real gsd_run. Assert the hermes dir itself rather than "gsd-core/bin/", which every resolver arm ends in and so cannot tell them apart. * fix(#4205): scrub BASH_ENV and reject an empty RUNTIME_DIR in runResolver BASH_ENV defeated the whole scrub. Non-interactive bash sources it before the script runs, which is after the env: object is applied, so a single inherited var re-injects any of the others. Measured: a BASH_ENV exporting CODEX_HOME turned (H) red; blanked, 30/30. runResolver passed RUNTIME_DIR: runtimeDir || '', and '' is indistinguishable from unset to ${RUNTIME_DIR:-$(git rev-parse --show-toplevel)} — an empty value falls back to the real repo root and resolves its real install, the leak this issue is about. Both callers already pass one, so require it rather than paper over it. * test(#4344): plant the leaked gsd_run sentinel on Windows too (B) planted its gsd_run sentinel only on POSIX, so the Windows shards proved nothing about buildIsolatedPath()'s PATH filter — the exact gap #4344 recorded. npm's global installs write an extensionless Bourne shim beside gsd_run.cmd, and fs.constants.X_OK behaves like F_OK on Windows, so the existing probe already sees the leak there; only the fixture was POSIX-gated. Plant the sentinel on every platform and assert that buildIsolatedPath() strips its directory from the returned PATH. That assertion is red on both platforms when the filter probes the wrong name, and unlike the stdout assertions it does not depend on the MSYS mount's exec heuristics. Restore process.env.PATH before the child spawns rather than in a t.after hook: on Windows process.env spreads as 'Path', so a still-live leak would compete with snippetEnv()'s 'PATH' override for the casing the child receives. Refs #4205 * fix(#4205): stop the host PATH leaking past snippetEnv on Windows Adversarial review (agy, gemini-3.8-flash-high) found that the fixtures' PATH isolation is defeatable on Windows regardless of which name the filter probes. Windows environment variables are case-insensitive but a spread of process.env is not: the host PATH enumerates as 'Path', so '{ ...process.env, PATH: isolated }' yields both keys, and libuv's make_program_env sorts the child's environment block case-insensitively without ever dropping duplicates. The child could therefore resolve the host PATH. snippetEnv() now drops every other casing whenever a caller supplies its own PATH. Also from that review: - buildIsolatedPath() takes the PATH to filter as a parameter, so (B0) and (B) no longer mutate process.env.PATH and no longer need try/finally restores. - The two loud-guard fixtures asserted 'not found' OR 'ERROR', which bash's own 'node: command not found' satisfies; they now assert the launcher's 'ERROR: gsd-tools.cjs not found'. - Corrected a comment counting three scrub keys as two, and two comments crediting a removed env argument for clearing ambient config dirs rather than snippetEnv()'s derived TEST_ENV_BASE. Refs #4344 * fix(#4205): make the fixtures' node shim work on Windows bug-211 (C) located node with `which node` through the process seam. `which` is not a Windows binary; the fixture only survived CI because Git Bash ships one. process.execPath is the same answer without the spawn, and (H) already used it. Both fixtures then built their node shim with fs.symlinkSync, which raises EPERM on Windows without developer mode or elevation — the same reason buildIsolatedPath() skips its own symlink step there. The shared linkNodeShim() helper hard-links instead on that platform (no privilege required) and falls back to a copy across volumes. Found by adversarial review (agy, gemini-3.8-flash-high). Refs #4344 * fix(#4205): make the launcher PATH isolation extension-aware and Windows-safe trek-e's review asks for a Windows-safe node fallback, an extension-aware filter, and a Windows regression test, in that order: broadening the filter first can strip the directory node itself lives in. buildIsolatedPath() now always prepends a directory holding node, on every platform, via the linkExecutable() helper (hard link on Windows, where symlinks need elevation). nodeBinDir is no longer nullable and the win32 early return is gone, so the fallback exists before the filter widens. (B0), the co-location invariant, therefore runs on Windows instead of being skipped on the one platform that had no fallback. The filter probes every name the launcher's `command -v gsd_run` arm can resolve. msys bash appends an executable extension during PATH lookup, so a directory holding only gsd_run.exe is reachable on Windows although gsd_run is absent. PATHEXT is folded in as well; it over-matches, which costs nothing now that node is always supplied separately. (B0) asserts every name in that set is filtered, and (B) plants a gsd_run.exe sentinel on Windows beside the extensionless one npm installs. The predicate lived in five hand-maintained copies — the drift that caused #4205 in the first place, and four of the copies pointed readers at a buildIsolatedPath() that was block-scoped out of their reach. buildIsolatedPath() moves to module scope and the four inline copies call it. The three shadow runBashFile() declarations this PR had to edit identically go with them. Red-proved both ways: probing 'gsd-tools' again turns (B0) and (B) red; making the node prepend conditional turns (B0)(ii) red. Refs #4344 * fix(#4205): drop empty PATH elements from the isolated PATH A POSIX shell reads an empty PATH element as the current directory, so an isolated PATH carrying one still lets the launcher's `command -v gsd_run` arm resolve a gsd_run from the fixture's own working directory — the leak class this file exists to close. Two ways one appeared. An ambient PATH containing `::` survived the filter, because `path.join('', 'gsd_run')` probes the working directory rather than a directory entry, so hasGsdRun could not see what it was admitting. And a PATH whose every entry was filtered joined to an empty string, leaving the returned value ending in a delimiter, which means the same thing. Empty entries are now dropped alongside the gsd_run-bearing ones, and the surviving directories are joined as a list, so a fully-filtered PATH yields the node shim dir alone. (B0) asserts both cases. Found by CodeRabbit on the fork rehearsal PR. Refs #4344 * fix(#4205): keep only absolute PATH dirs, and assert the sentinel by basename Adversarial review (agy, gemini-3.8-flash-high) on the previous head. An empty PATH element was dropped, but `.` and any other relative entry say the same thing explicitly and survived. hasGsdRun() cannot see what it would admit either: `path.join('.', 'gsd_run')` probes the runner's working directory, not the child's, so isolation also varied by where the suite was started from. Only absolute directories survive now, which can only tighten the isolation. (B0) asserts it over an empty element, a `.`, a relative entry, and a fully-filtered PATH — the previous empty-string assertion passed on the `.` case. (B)'s GSD_TOOLS assertion compared an absolute os.tmpdir() path against launcher output, which the file already documents as a mismatch on Windows: git-bash prints /c/Users/... where Node gives C:\Users\.... It never matched there, so it asserted nothing on the platform it was added for. It matches the mkdtemp basename now, which both path forms share. (B) also plants ONLY gsd_run.exe on Windows: with an extensionless sibling present, an extension-blind filter would strip the directory for the wrong reason and pass. snippetEnv() deduped case-variant keys for PATH alone. On Windows every scrubbed key has the same exposure — an ambient `bash_env` reaches the child beside the blanked `BASH_ENV`, and BASH_ENV re-injects the rest. Every key the function sets now wins over other casings of itself; keys it does not set are untouched, so a caller passing no PATH override still gets the host PATH. Refs #4344 * test(#4205): plant probe fixtures as files, not interpreter links Review follow-ups on 776e9371f. (B0)'s co-location fixture and its GSD_RUN_NAMES sweep only ever probe the planted names with accessSync; nothing executes them. They used linkExecutable, so on Windows the sweep hard-linked node.exe once per PATHEXT entry — a dozen on a stock runner, and a full copy each when os.tmpdir() and process.execPath sit on different volumes. plantExecutable writes a zero-byte 0o755 file instead. linkExecutable keeps the two callers that need a real executable: the node buildIsolatedPath prepends, and the Windows sentinel. The '/usr/bin:/bin' fallback formatted POSIX paths with the platform delimiter, yielding '/usr/bin;/bin' on Windows, which path.isAbsolute accepts and no Windows shell would ever produce. It was also unreachable: basePath defaults to process.env.PATH. An unset PATH now yields the node shim dir alone and fails loudly at spawn rather than being papered over. (B)'s header said the sentinel is planted in both forms; the code plants one per platform, and planting both on Windows is what the branch below it exists to avoid. Also names which assertion carries the Windows guarantee, since SENTINEL_INVOKED cannot fire there. The shared helper's temp dirs were prefixed gsd-891-, attributing every fixture's leftovers to one of the four bugs it now serves. Refs #4344 * fix(#4205): model bash's PATH lookup, not cmd.exe's, and give (B0) its own oracle Ponytail review on aec6012fc. GSD_RUN_NAMES expanded PATHEXT, which describes cmd.exe rather than the shell the launcher's `command -v gsd_run` arm runs under. The Cygwin/msys rule is that .exe may be omitted from a command while '.bat and .com ... you cannot omit the extension', so gsd_run.exe is reachable for a bare gsd_run and gsd_run.cmd/.ps1 are not, whatever PATHEXT lists. The comment claimed the resulting over-match was free. It was not: buildIsolatedPath() restores node to the isolated PATH, but nothing restores bash, which the fixtures spawn by name — so every extra dropped directory was another chance to remove the one bash lives in and fail with ENOENT instead of an assertion. Narrowed to gsd_run and gsd_run.exe. (Aside, the PATHEXT default does not even contain .PS1.) (B0)'s name-sweep took its list from GSD_RUN_NAMES, so it swept the constant under test with itself and could only catch that constant being deleted, never being wrong. Its relative-element case re-ran the implementation's own filter predicate over that filter's output, which is true for any predicate. Both now assert against written-out expectations: the reachable names per platform, and the exact directories that must survive each case. Also: the test name covered two of its four assertions, and the case table's prose counted three of its four entries. Red-proved three ways: dropping the gsd_run filter, dropping the absoluteness filter, and claiming a name the filter does not cover each turn (B0) red. Refs #4344 --------- Co-authored-by: Test Co-authored-by: Tom Boucher --- tests/runtime-launcher-parity.test.cjs | 708 +++++++++++++++---------- 1 file changed, 425 insertions(+), 283 deletions(-) diff --git a/tests/runtime-launcher-parity.test.cjs b/tests/runtime-launcher-parity.test.cjs index d3b8bf75a..6d83664c0 100644 --- a/tests/runtime-launcher-parity.test.cjs +++ b/tests/runtime-launcher-parity.test.cjs @@ -13,11 +13,11 @@ * (D) Loud guard behavioral: missing gsd-tools.cjs exits non-zero and emits * "not found" to stderr. * (E) PATH fallback behavioral: when no local gsd-tools.cjs, the elif branch - * resolves to the gsd-tools binary on PATH (#3668). + * resolves to the gsd_run binary on PATH (#3668). * (F) Regression locks: the snippet file contains no /gsd-tools substring; and * no line in workflows/do.md matches /\/gsd[:-][a-z]/ (dispatcher-parity * scanner must not read the preamble as a slash-command stub). - * (H) Codex shim fallback: when PATH has no gsd-tools, $HOME/.codex/gsd-core/bin + * (H) Codex shim fallback: when PATH has no gsd_run, $HOME/.codex/gsd-core/bin * can satisfy gsd_run for Codex shim-only installs. */ @@ -31,24 +31,173 @@ const path = require('node:path'); const os = require('node:os'); const { runHook: runHookSeam } = require('./helpers/process-seam.cjs'); const { throwIfFailed } = require('./helpers/git-fixture.cjs'); -const { cleanup } = require('./helpers.cjs'); +const { cleanup, TEST_ENV_BASE } = require('./helpers.cjs'); const { escapeRegex } = require('../gsd-core/bin/lib/pattern.cjs'); const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows'); const AGENTS_DIR = path.join(__dirname, '..', 'agents'); const SNIPPET_FILE = path.join(WORKFLOWS_DIR, '_runtime-launcher.snippet.sh'); +/** + * Env for any fixture that sources the snippet (#4205). + * + * TEST_ENV_BASE is DERIVED from the same capability registry the resolver + * reads (see tests/helpers.cjs, #2665), so a new runtime home cannot leave it + * silently stale — the shape of #4205. + * + * Three keys the derived set cannot supply: + * - GEMINI_CONFIG_DIR: the gemini runtime is retired (#1928) so the registry + * no longer carries it, but the snippet still probes its arm. + * - CLAUDE_ENV_FILE: a WRITE sink, not a read path. Left ambient, every + * fixture that exits 0 appends `export PATH=''` to the + * developer's real env file, each line naming a /tmp dir the fixture has + * already deleted. + * - BASH_ENV: sourced by non-interactive bash BEFORE the script, which is + * after this scrub is applied — so one inherited var re-injects any of + * the others. Measured: a BASH_ENV exporting CODEX_HOME turns (H) red. + */ +const SNIPPET_SCRUB = { GEMINI_CONFIG_DIR: '', CLAUDE_ENV_FILE: '', BASH_ENV: '' }; +function snippetEnv(overrides = {}) { + const env = { ...process.env, ...TEST_ENV_BASE, ...SNIPPET_SCRUB, ...overrides }; + // Windows env vars are case-insensitive; a spread of process.env is not. The + // host PATH enumerates as `Path` there, so `{ ...process.env, PATH: x }` + // yields BOTH keys — and libuv's make_program_env sorts the child's block + // case-insensitively but never drops duplicates, so the ambient twin of + // anything scrubbed here still reaches the child and defeats the isolation + // this whole file rests on. Every key this function sets wins over any other + // casing of itself; keys it does not set are left alone, so a caller that + // passes no PATH override still gets the host PATH. + const canonical = new Map( + [...Object.keys(TEST_ENV_BASE), ...Object.keys(SNIPPET_SCRUB), ...Object.keys(overrides)] + .map((key) => [key.toUpperCase(), key]), + ); + for (const key of Object.keys(env)) { + const owner = canonical.get(key.toUpperCase()); + if (owner !== undefined && owner !== key) delete env[key]; + } + return env; +} + /** * Run a bash script FILE via the process seam, preserving the throw-on- * nonzero-exit semantics of the execFileSync('bash', [path], ...) idiom * this replaces. */ function runBashFile(scriptPath, options = {}) { - const r = runHookSeam(scriptPath, [], { interpreter: 'bash', ...options }); + const r = runHookSeam(scriptPath, [], { + interpreter: 'bash', ...options, env: snippetEnv(options.env), + }); throwIfFailed(r, `bash ${scriptPath}`); return r.stdout; } +const NODE_BIN = process.platform === 'win32' ? 'node.exe' : 'node'; + +/** + * Put an executable link to this interpreter in `dir` under `name`, and return + * `dir` so a caller can prepend it to a PATH. Callers use it for both halves of + * a fixture: the node the launcher needs, and the gsd_run sentinel it must not + * reach. The content never matters, only that the name resolves. + * + * Windows symlinks need elevation, so a hard link is used there instead: it + * needs no privilege, but it cannot cross volumes, hence the copy fallback. + */ +function linkExecutable(dir, name) { + fs.mkdirSync(dir, { recursive: true }); + const target = path.join(dir, name); + if (process.platform !== 'win32') { + fs.symlinkSync(process.execPath, target); + return dir; + } + try { + fs.linkSync(process.execPath, target); + } catch (err) { + if (err.code !== 'EXDEV') throw err; + fs.copyFileSync(process.execPath, target); + } + return dir; +} + +/** + * Plant `dir/name` as a file an X_OK probe accepts, and return `dir`. For + * fixtures that only need a name to be *found*: nothing ever executes these, + * so they cost a zero-byte write instead of a link to (or, across volumes, a + * copy of) the whole interpreter. + */ +function plantExecutable(dir, name) { + fs.mkdirSync(dir, { recursive: true }); + fs.writeFileSync(path.join(dir, name), '', { mode: 0o755 }); + return dir; +} + +/** + * Every filename the launcher's `command -v gsd_run` arm could resolve. + * + * The arm runs under bash, so this models bash's lookup, not cmd.exe's. The + * Cygwin/msys rule is that `.exe` may be omitted from a command while ".bat and + * .com ... you cannot omit the extension" — so `gsd_run.exe` is reachable for a + * bare `gsd_run` and `gsd_run.cmd`/`.ps1` are not, whatever PATHEXT says. + * + * Matching PATHEXT instead would drop more directories than bash can reach, and + * that is not free: buildIsolatedPath() restores node to the isolated PATH but + * nothing restores bash, which the fixtures spawn by name. A wider drop set is + * a wider chance of removing the directory bash itself lives in and failing the + * fixture with ENOENT instead of an assertion. + */ +const GSD_RUN_NAMES = process.platform === 'win32' + ? ['gsd_run', 'gsd_run.exe'] + : ['gsd_run']; + +/** True when `dir` holds a gsd_run the launcher's PATH arm could resolve. */ +function hasGsdRun(dir) { + return GSD_RUN_NAMES.some((name) => { + try { fs.accessSync(path.join(dir, name), fs.constants.X_OK); return true; } + catch { return false; } + }); +} + +/** + * Build a PATH the launcher's `command -v gsd_run` arm cannot resolve anything + * from, while a bare `node` lookup still succeeds — the precondition every + * runtime-home fallback fixture needs. + * + * Directories holding a resolvable gsd_run are dropped (#4205: the arm probes + * `gsd_run`, not `gsd-tools`), then a dedicated dir carrying only node is + * prepended. The prepend is unconditional because dropping the directory node + * itself lives in is a normal outcome, not an exotic one: fnm, nvm, Homebrew + * and Windows global installs all co-locate the two. + * + * The caller cleans up `result.nodeBinDir` (pass it to `cleanup()` in a + * `t.after` or `finally` block). + * + * @param {string} [basePath] PATH to filter. Callers planting a leaked gsd_run + * pass their own string rather than mutating `process.env.PATH`. + * @returns {{ isolatedPath: string, nodeBinDir: string }} + */ +function buildIsolatedPath(basePath = process.env.PATH) { + // Only absolute directories survive. An empty element means "the current + // directory" to a POSIX shell and `.` says so explicitly, so either one puts + // the child's cwd on PATH — and hasGsdRun cannot see what it would admit, + // since `path.join('.', 'gsd_run')` probes the *runner's* cwd instead of the + // child's. Joining the survivors (rather than the filtered string) also keeps + // a fully-filtered PATH from ending in a delimiter, which means the same + // thing. Dropping a relative entry can only tighten the isolation, never + // loosen it. + const filteredDirs = (basePath ?? '') + .split(path.delimiter) + .filter((p) => path.isAbsolute(p) && !hasGsdRun(p)); + + const nodeBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-node-')); + try { + linkExecutable(nodeBinDir, NODE_BIN); + } catch (err) { + cleanup(nodeBinDir); + throw err; + } + + return { isolatedPath: [nodeBinDir, ...filteredDirs].join(path.delimiter), nodeBinDir }; +} + /** * Read the canonical preamble from the snippet file (all lines, no trailing newline). */ @@ -360,12 +509,12 @@ describe('runtime-launcher-parity (#373)', () => { }); // ─── (D) Loud guard: missing runtime is fatal ───────────────────────────── - test('(D) missing gsd-tools.cjs and no PATH gsd-tools causes loud non-zero exit with "not found" on stderr', () => { + test('(D) missing gsd-tools.cjs and no PATH gsd_run causes loud non-zero exit with "not found" on stderr', (t) => { // Create temp dir with a space in the name, but NO gsd-tools.cjs. - // We ensure gsd-tools is not on PATH by prepending a dir that has no - // gsd-tools binary (system binaries remain on PATH so bash/node work). + // We ensure gsd_run is not on PATH by prepending a dir that has no + // gsd_run binary (system binaries remain on PATH so bash/node work). const base = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd 373 notools ')); - // Place a no-op dir first in PATH; no gsd-tools stub there. + // Place a no-op dir first in PATH; no gsd_run stub there. const noToolsBin = path.join(base, 'nobin'); fs.mkdirSync(noToolsBin, { recursive: true }); try { @@ -380,27 +529,30 @@ describe('runtime-launcher-parity (#373)', () => { const scriptPath = path.join(base, 'test-guard.sh'); fs.writeFileSync(scriptPath, scriptContent); - // Build a PATH that has noToolsBin first (no gsd-tools stub there) but retains - // system paths needed for bash. Exclude any PATH entry that contains a gsd-tools binary. - const systemPaths = (process.env.PATH || '/usr/bin:/bin') - .split(path.delimiter) - .filter((p) => { - try { fs.accessSync(path.join(p, 'gsd-tools'), fs.constants.X_OK); return false; } - catch { return true; } - }); - const isolatedPath = [noToolsBin, ...systemPaths].join(path.delimiter); + // noToolsBin first (no gsd_run stub there), then a PATH no gsd_run is + // resolvable from. + const isolated = buildIsolatedPath(); + t.after(() => cleanup(isolated.nodeBinDir)); + const isolatedPath = [noToolsBin, isolated.isolatedPath].join(path.delimiter); const r = runHookSeam(scriptPath, [], { interpreter: 'bash', - env: { ...process.env, PATH: isolatedPath, HOME: base }, + // Loud-guard requires every runtime-home arm to genuinely miss, not + // just PATH (#4205 class: an ambient config-dir var pointing at a + // real install would resolve here instead of the hard error). + env: snippetEnv({ PATH: isolatedPath, HOME: base }), }); const threw = r.exitCode !== 0; const stderrOutput = r.stderr || ''; - assert.ok(threw, 'Expected the script to exit non-zero when gsd-tools.cjs is missing and gsd-tools is not on PATH'); + assert.ok(threw, 'Expected the script to exit non-zero when gsd-tools.cjs is missing and gsd_run is not on PATH'); + // Match the launcher's own diagnostic, not a bare "not found": when node + // itself is missing from the isolated PATH, bash's own + // `bash: node: command not found` satisfies the loose form and a + // regressed guard passes. assert.ok( - stderrOutput.includes('not found') || stderrOutput.includes('ERROR'), - `Expected stderr to contain "not found" or "ERROR", got: ${stderrOutput.trim()}`, + stderrOutput.includes('ERROR: gsd-tools.cjs not found'), + `Expected stderr to contain "ERROR: gsd-tools.cjs not found", got: ${stderrOutput.trim()}`, ); } finally { cleanup(base); @@ -439,7 +591,7 @@ describe('runtime-launcher-parity (#373)', () => { fs.writeFileSync(scriptPath, scriptContent); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: `${pathBinDir}${path.delimiter}${process.env.PATH || ''}` }, + env: { PATH: `${pathBinDir}${path.delimiter}${process.env.PATH || ''}` }, }); // The PATH fallback must have resolved GSD_TOOLS to the stub binary. @@ -474,7 +626,7 @@ describe('runtime-launcher-parity (#373)', () => { // The resolution order must be: // (1) local/RUNTIME_DIR → (2) PATH → (3) $HOME/.claude/gsd-core/bin → (4) hard error // We probe for .claude/gsd-core/bin (using ${_GSD_SHIM_NAME} indirection) - // between the `command -v gsd-tools` elif and the hard-error else branch. + // between the `command -v gsd_run` elif and the hard-error else branch. const CLAUDE_HOME_PROBE = '.claude/gsd-core/bin/'; // Assert snippet itself contains the probe @@ -518,7 +670,7 @@ describe('runtime-launcher-parity (#373)', () => { }); // ─── (H) Codex shim fallback behavioral ------------------------------------ - test('(H) gsd_run resolves $HOME/.codex/gsd-core/bin/ shim when PATH has no gsd-tools', () => { + test('(H) gsd_run resolves $HOME/.codex/gsd-core/bin/ shim when PATH has no gsd_run', (t) => { const CODEX_HOME_PROBE = '.codex/gsd-core/bin/'; const snippetContent = fs.readFileSync(SNIPPET_FILE, 'utf8'); @@ -571,26 +723,17 @@ describe('runtime-launcher-parity (#373)', () => { const scriptPath = path.join(fakeRuntime, 'test-codex-home-fb.sh'); fs.writeFileSync(scriptPath, scriptContent); - const hasExecutable = (dir, name) => { - try { - fs.accessSync(path.join(dir, name), fs.constants.X_OK); - return true; - } catch { - return false; - } - }; - const systemPaths = (process.env.PATH || '/usr/bin:/bin') - .split(path.delimiter) - .filter((p) => !hasExecutable(p, 'gsd-tools')); - if (!systemPaths.some((p) => hasExecutable(p, 'node'))) { - const nodeShimDir = path.join(fakeRuntime, 'node-shim'); - fs.mkdirSync(nodeShimDir, { recursive: true }); - fs.symlinkSync(process.execPath, path.join(nodeShimDir, 'node')); - systemPaths.unshift(nodeShimDir); - } + const { isolatedPath, nodeBinDir } = buildIsolatedPath(); + t.after(() => cleanup(nodeBinDir)); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: systemPaths.join(path.delimiter), HOME: fakeHome }, + // Ambient CLAUDE_CONFIG_DIR/HERMES_HOME/CURSOR_CONFIG_DIR resolve arms + // checked before CODEX_HOME, and an ambient CODEX_HOME overrides the + // $HOME/.codex default this test asserts (#4205 class: same env-leak + // bug, this time via config-dir vars rather than PATH). snippetEnv()'s + // derived TEST_ENV_BASE clears all four, so only HOME's default + // $HOME/.codex fallback can win. + env: { PATH: isolatedPath, HOME: fakeHome }, }); const normStdout = stdout.replace(/\\/g, '/'); @@ -880,7 +1023,7 @@ describe('runtime-launcher-parity — agents (#1041)', () => { * Asserts: * (A) The canonical snippet file contains the ~/.claude fallback arm. * (B) A representative propagated workflow file contains the ~/.claude fallback arm. - * (C) Behavioral: when RUNTIME_DIR misses and gsd-tools is NOT on PATH, + * (C) Behavioral: when RUNTIME_DIR misses and gsd_run is NOT on PATH, * a stub at $HOME/.claude/gsd-core/bin/gsd-tools.cjs is resolved and invoked. * (D) The resolution order is preserved: local -> PATH -> ~/.claude -> hard error. * When all three miss, exit non-zero. @@ -897,7 +1040,6 @@ const fs = require('node:fs'); const path = require('node:path'); const os = require('node:os'); const { runHook: runHookSeam } = require('./helpers/process-seam.cjs'); -const { throwIfFailed } = require('./helpers/git-fixture.cjs'); const { cleanup } = require('./helpers.cjs'); const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows'); @@ -905,16 +1047,6 @@ const SNIPPET_FILE = path.join(WORKFLOWS_DIR, '_runtime-launcher.snippet.sh'); // Representative propagated workflow file (has a gsd_run call): const REPRESENTATIVE_FILE = path.join(WORKFLOWS_DIR, 'add-backlog.md'); -/** - * Run a bash script FILE via the process seam, preserving the throw-on- - * nonzero-exit semantics of the execFileSync('bash', [path], ...) idiom - * this replaces. - */ -function runBashFile(scriptPath, options = {}) { - const r = runHookSeam(scriptPath, [], { interpreter: 'bash', ...options }); - throwIfFailed(r, `bash ${scriptPath}`); - return r.stdout; -} const CLAUDE_HOME_PROBE = '.claude/gsd-core/bin/'; @@ -940,7 +1072,7 @@ describe('bug-211: launcher ~/.claude home fallback', () => { }); // --- (C) Behavioral: ~/.claude stub is resolved when local and PATH both miss - test('(C) gsd_run resolves $HOME/.claude/gsd-core/bin/ stub when no local install and gsd-tools not on PATH', () => { + test('(C) gsd_run resolves $HOME/.claude/gsd-core/bin/ stub when no local install and gsd_run not on PATH', (t) => { // Build a fake $HOME with a stub at .claude/gsd-core/bin/gsd-tools.cjs const fakeHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-211-home-')); // RUNTIME_DIR points to a directory with no gsd-tools.cjs @@ -969,37 +1101,15 @@ describe('bug-211: launcher ~/.claude home fallback', () => { const scriptPath = path.join(fakeRuntime, 'test-home-fb.sh'); fs.writeFileSync(scriptPath, scriptContent); - // Build a PATH with no gsd-tools binary to force the ~/.claude arm. - // Filter out directories that contain a gsd-tools executable. If node lives - // in the same directory as gsd-tools, create a dedicated shim dir with a - // symlink to node only (no gsd-tools there). - const nodeBinResult = runHookSeam('node', [], { interpreter: 'which' }); - throwIfFailed(nodeBinResult, 'which node'); - const nodeBin = nodeBinResult.stdout.trim(); - const systemPaths = (process.env.PATH || '/usr/bin:/bin') - .split(path.delimiter) - .filter((p) => { - try { - fs.accessSync(path.join(p, 'gsd-tools'), fs.constants.X_OK); - return false; - } catch { - return true; - } - }); - // If node's dir was filtered (it contained gsd-tools), create a shim dir - // with just a node symlink so the stub's shebang (#!/usr/bin/env node) resolves. - const nodeShimDir = path.join(fakeRuntime, 'node-shim'); - if (!systemPaths.some((p) => { - try { fs.accessSync(path.join(p, 'node'), fs.constants.X_OK); return true; } - catch { return false; } - })) { - fs.mkdirSync(nodeShimDir, { recursive: true }); - fs.symlinkSync(nodeBin, path.join(nodeShimDir, 'node')); - systemPaths.unshift(nodeShimDir); - } + // A PATH with no resolvable gsd_run, to force the ~/.claude arm. The + // helper also supplies node, which the stub's #!/usr/bin/env node needs. + const { isolatedPath, nodeBinDir } = buildIsolatedPath(); + t.after(() => cleanup(nodeBinDir)); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: systemPaths.join(path.delimiter), HOME: fakeHome }, + // Ambient CLAUDE_CONFIG_DIR would override the $HOME/.claude default + // this test relies on (#4205 class: env-leak, not PATH-leak). + env: { PATH: isolatedPath, HOME: fakeHome }, }); // GSD_TOOLS must point into the fake ~/.claude dir @@ -1020,7 +1130,7 @@ describe('bug-211: launcher ~/.claude home fallback', () => { }); // --- (D) All three miss -> hard error ------------------------------------- - test('(D) hard error when local, PATH, and ~/.claude all miss', () => { + test('(D) hard error when local, PATH, and ~/.claude all miss', (t) => { const fakeHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-211-nohome-')); const fakeRuntime = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-211-nort-')); // noToolsBin so PATH check finds nothing @@ -1039,29 +1149,28 @@ describe('bug-211: launcher ~/.claude home fallback', () => { const scriptPath = path.join(fakeRuntime, 'test-allfail.sh'); fs.writeFileSync(scriptPath, scriptContent); - const systemPaths = (process.env.PATH || '/usr/bin:/bin') - .split(path.delimiter) - .filter((p) => { - try { - fs.accessSync(path.join(p, 'gsd-tools'), fs.constants.X_OK); - return false; - } catch { - return true; - } - }); - const isolatedPath = [noToolsBin, ...systemPaths].join(path.delimiter); + const isolated = buildIsolatedPath(); + t.after(() => cleanup(isolated.nodeBinDir)); + const isolatedPath = [noToolsBin, isolated.isolatedPath].join(path.delimiter); const r = runHookSeam(scriptPath, [], { interpreter: 'bash', - env: { ...process.env, PATH: isolatedPath, HOME: fakeHome }, + // "All three miss" requires every runtime-home arm to genuinely miss, + // not just the first (#4205 class: an ambient config-dir var pointing + // at a real install would resolve here instead of the hard error). + env: snippetEnv({ PATH: isolatedPath, HOME: fakeHome }), }); const threw = r.exitCode !== 0; const stderrOutput = r.stderr || ''; assert.ok(threw, 'Expected non-zero exit when all three resolution arms miss'); + // Match the launcher's own diagnostic, not a bare "not found": when node + // itself is missing from the isolated PATH, bash's own + // `bash: node: command not found` satisfies the loose form and a + // regressed guard passes. assert.ok( - stderrOutput.includes('not found') || stderrOutput.includes('ERROR'), - `Expected stderr to contain "not found" or "ERROR", got: ${stderrOutput.trim()}`, + stderrOutput.includes('ERROR: gsd-tools.cjs not found'), + `Expected stderr to contain "ERROR: gsd-tools.cjs not found", got: ${stderrOutput.trim()}`, ); } finally { cleanup(fakeHome); @@ -1088,12 +1197,15 @@ describe('bug-211: launcher ~/.claude home fallback', () => { * Every non-Claude runtime (Hermes, Cursor, Codex, Copilot, Windsurf, …) * installs gsd-core into a *different* directory that the shim never tried, * causing a false-positive fatal ERROR on all non-Claude runtimes when - * RUNTIME_DIR is not set and gsd-tools is not on PATH. + * RUNTIME_DIR is not set and gsd_run is not on PATH. * * Asserts: * (A) Snippet contains all expected non-Claude runtime home probes (structural). - * (B) HERMES_HOME behavioral: when RUNTIME_DIR misses and gsd-tools is NOT on - * PATH, a stub at ${HERMES_HOME}/gsd-core/bin/gsd-tools.cjs is invoked. + * (B0) buildIsolatedPath() co-location invariant (see below). + * (B) HERMES_HOME behavioral: when RUNTIME_DIR misses and gsd_run is NOT on + * PATH, the stub at ${HERMES_HOME}/gsd-core/bin/gsd-tools.cjs is invoked. + * It plants a leaked gsd_run on PATH first (#4205), on every platform, + * which is what makes it fail on a clean machine as well as a leaking one. * (C) Default Hermes path behavioral: stub at $HOME/.hermes/gsd-core/bin/ * gsd-tools.cjs is invoked when HERMES_HOME is not set. * (D) Resolution order: non-Claude homes are probed BEFORE the hard error, @@ -1113,24 +1225,12 @@ const assert = require('node:assert/strict'); const fs = require('node:fs'); const path = require('node:path'); const os = require('node:os'); -const { runHook: runHookSeam } = require('./helpers/process-seam.cjs'); -const { throwIfFailed } = require('./helpers/git-fixture.cjs'); const { cleanup } = require('./helpers.cjs'); const { escapeRegex } = require('../gsd-core/bin/lib/pattern.cjs'); const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows'); const SNIPPET_FILE = path.join(WORKFLOWS_DIR, '_runtime-launcher.snippet.sh'); -/** - * Run a bash script FILE via the process seam, preserving the throw-on- - * nonzero-exit semantics of the execFileSync('bash', [path], ...) idiom - * this replaces. - */ -function runBashFile(scriptPath, options = {}) { - const r = runHookSeam(scriptPath, [], { interpreter: 'bash', ...options }); - throwIfFailed(r, `bash ${scriptPath}`); - return r.stdout; -} // Every non-Claude runtime home probe the snippet must contain. // Key: runtime name (for diagnostics). Value: the substring that must appear @@ -1215,48 +1315,6 @@ function extractShellBlocks(content) { return blocks; } -/** - * Build a PATH with no gsd-tools binary so the PATH fallback branch is skipped, - * while guaranteeing that a bare `node` lookup still resolves regardless of whether - * the real node binary co-locates with a global gsd-tools shim (e.g. fnm/nvm/Homebrew). - * - * Strategy (POSIX only): create a temp dir containing only a `node` symlink → - * process.execPath, prepend it to the gsd-tools-filtered PATH. The filtered - * PATH excludes any directory that contains an executable `gsd-tools`. - * - * On Windows the co-location bug does not apply (gsd-tools resolves via .cmd/.ps1, - * not the bare binary probed here), and symlinks may require elevated privileges, - * so we skip the symlink step entirely on that platform. - * - * The caller is responsible for cleaning up `result.nodeBinDir` when non-null - * (pass it to `cleanup()` in a `t.after` or `finally` block). - * - * @returns {{ isolatedPath: string, nodeBinDir: string|null }} - */ -function buildIsolatedPath() { - const filteredPath = (process.env.PATH || '/usr/bin:/bin') - .split(path.delimiter) - .filter((p) => { - try { fs.accessSync(path.join(p, 'gsd-tools'), fs.constants.X_OK); return false; } - catch { return true; } - }) - .join(path.delimiter); - - // Windows: no symlink (see JSDoc above); callers must handle nodeBinDir === null. - if (process.platform === 'win32') { - return { isolatedPath: filteredPath, nodeBinDir: null }; - } - - const nodeBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-node-')); - try { - fs.symlinkSync(process.execPath, path.join(nodeBinDir, 'node')); - } catch (err) { - cleanup(nodeBinDir); - throw err; - } - - return { isolatedPath: nodeBinDir + path.delimiter + filteredPath, nodeBinDir }; -} describe('bug-891: non-Claude runtime home fallback arms', () => { @@ -1304,66 +1362,53 @@ describe('bug-891: non-Claude runtime home fallback arms', () => { }); // ── (B0) Regression: buildIsolatedPath keeps node resolvable when node and ── - // gsd-tools co-locate in the same PATH directory. ─ + // gsd_run co-locate in the same PATH directory. ─ // - // Machine-independence guarantee: PATH is set to ONLY two controlled dirs — - // fakeBinDir (holds both fake gsd-tools AND a node symlink) plus a fresh - // empty dir (no executables at all). The real system PATH is NOT appended. + // Machine-independence guarantee: the filtered PATH is ONLY two controlled + // dirs — fakeBinDir (holds both fake gsd_run AND node) plus a fresh empty dir + // (no executables at all). The real system PATH is NOT appended. // - // Old logic: filters out fakeBinDir → only the empty dir remains → node - // UNresolvable → assertion (ii) FAILS (true-red on any machine). - // New logic: prepends its own nodeBinDir → node resolvable despite fakeBinDir - // being filtered → both assertions pass. + // Filtering fakeBinDir out is correct and unavoidable, so the only thing that + // keeps node reachable is the nodeBinDir buildIsolatedPath() prepends. This + // test is that prepend's guard: make it conditional again — as it was on + // Windows, where the helper returned `nodeBinDir: null` — and (ii) goes red. test( - '(B0) buildIsolatedPath: node is resolvable and gsd-tools is not when they share a PATH dir', - { skip: process.platform === 'win32' ? 'POSIX-only co-location scenario' : false }, + '(B0) buildIsolatedPath invariants: node survives, every reachable gsd_run name and every relative dir does not', (t) => { - // Build a fake bin dir that contains BOTH a gsd-tools executable and a node - // symlink, simulating a dev setup (fnm/nvm/Homebrew) where both land in the - // same bin directory. + // Build a fake bin dir that contains BOTH a gsd_run executable and node, + // simulating a dev setup (fnm/nvm/Homebrew) where both land in the same + // bin directory. const fakeBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-colocated-')); - // A second fresh empty dir — contains neither gsd-tools nor node. + // A second fresh empty dir — contains neither gsd_run nor node. const emptyDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-empty-')); t.after(() => cleanup(fakeBinDir)); t.after(() => cleanup(emptyDir)); - // Fake gsd-tools shim (executable file) - const fakeGsdTools = path.join(fakeBinDir, 'gsd-tools'); - fs.writeFileSync(fakeGsdTools, '#!/bin/sh\necho fake-gsd-tools\n'); - fs.chmodSync(fakeGsdTools, 0o755); + // gsd_run co-located with node — the fnm/nvm/Homebrew layout, and the + // Windows global-install layout the same helper has to survive. Both are + // probed, never run. + plantExecutable(fakeBinDir, 'gsd_run'); + plantExecutable(fakeBinDir, NODE_BIN); - // node symlink pointing at the real interpreter (co-located with gsd-tools) - fs.symlinkSync(process.execPath, path.join(fakeBinDir, 'node')); - - // Set PATH to ONLY the two controlled dirs (no real system dirs). - // This makes the test machine-independent: on any machine, the only place - // node *could* come from before the fix is fakeBinDir — which gets filtered. - const origPath = process.env.PATH; - process.env.PATH = fakeBinDir + path.delimiter + emptyDir; - let result; - try { - result = buildIsolatedPath(); - } finally { - process.env.PATH = origPath; - } + // Filter ONLY the two controlled dirs (no real system dirs). This makes + // the test machine-independent: on any machine, the only place node + // *could* come from before the fix is fakeBinDir — which gets filtered. + const result = buildIsolatedPath(fakeBinDir + path.delimiter + emptyDir); t.after(() => cleanup(result.nodeBinDir)); const returnedDirs = result.isolatedPath.split(path.delimiter); - // (i) gsd-tools must NOT be resolvable on the returned PATH - const gsdToolsResolvable = returnedDirs.some((dir) => { - try { fs.accessSync(path.join(dir, 'gsd-tools'), fs.constants.X_OK); return true; } - catch { return false; } - }); + // (i) gsd_run must NOT be resolvable on the returned PATH + const gsdRunResolvable = returnedDirs.some(hasGsdRun); assert.equal( - gsdToolsResolvable, + gsdRunResolvable, false, - 'gsd-tools must not be resolvable on the isolated PATH (home-fallback would be bypassed)', + 'gsd_run must not be resolvable on the isolated PATH (home-fallback would be bypassed)', ); // (ii) node must BE resolvable on the returned PATH (the new nodeBinDir makes it so) const nodeResolvable = returnedDirs.some((dir) => { - try { fs.accessSync(path.join(dir, 'node'), fs.constants.X_OK); return true; } + try { fs.accessSync(path.join(dir, NODE_BIN), fs.constants.X_OK); return true; } catch { return false; } }); assert.equal( @@ -1371,60 +1416,165 @@ describe('bug-891: non-Claude runtime home fallback arms', () => { true, 'node must be resolvable on the isolated PATH (launcher runs: node "$GSD_TOOLS" "$@")', ); + + // (iii) every name the launcher's PATH arm can resolve must be filtered, + // not just the extensionless one — a gsd_run.exe-only directory is the + // reachable Windows leak an extensionless probe misses (#4344). The list + // is written out rather than taken from GSD_RUN_NAMES: sweeping the + // constant under test with itself cannot catch that constant being wrong. + const reachableNames = process.platform === 'win32' + ? ['gsd_run', 'gsd_run.exe'] + : ['gsd_run']; + const survived = reachableNames.filter((name) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-ext-')); + t.after(() => cleanup(dir)); + plantExecutable(dir, name); + const probe = buildIsolatedPath(dir); + t.after(() => cleanup(probe.nodeBinDir)); + return probe.isolatedPath.split(path.delimiter).includes(dir); + }); + assert.deepStrictEqual( + survived, + [], + 'every name bash can resolve for a bare gsd_run must be filtered out of the isolated PATH', + ); + + // (iv) every surviving element is absolute. A shell resolves an empty + // element, `.`, or any relative entry against the *child's* working + // directory, so each one is a way to put a cwd gsd_run back on PATH — + // and hasGsdRun cannot see any of them, because it probes relative to the + // runner's cwd instead. Four ways one arrives: an ambient `::`, an + // explicit `.`, a relative entry, and a PATH every entry of which was + // filtered (which used to leave a trailing delimiter, meaning the same + // thing). + // Each case names the dirs that must survive, so the assertion has an + // oracle of its own rather than re-running the implementation's filter. + const relativeElementCases = { + emptyElement: [`${emptyDir}${path.delimiter}${path.delimiter}${emptyDir}`, [emptyDir, emptyDir]], + dotElement: [`${emptyDir}${path.delimiter}.${path.delimiter}${emptyDir}`, [emptyDir, emptyDir]], + relativeElement: [`${emptyDir}${path.delimiter}sub/dir`, [emptyDir]], + fullyFiltered: [fakeBinDir, []], + }; + for (const [label, [basePath, expected]] of Object.entries(relativeElementCases)) { + const probe = buildIsolatedPath(basePath); + t.after(() => cleanup(probe.nodeBinDir)); + assert.deepStrictEqual( + probe.isolatedPath.split(path.delimiter), + [probe.nodeBinDir, ...expected], + `${label}: only absolute, gsd_run-free dirs may survive (a relative one resolves the child's cwd), got: ${probe.isolatedPath}`, + ); + } }, ); - // ── (B) Behavioral: HERMES_HOME stub is resolved ────────────────────────── - test('(B) gsd_run resolves ${HERMES_HOME}/gsd-core/bin/ stub when set and local+PATH both miss', () => { + + // ── (B) Behavioral: HERMES_HOME stub is resolved, even past a leaked PATH ── + // + // #4205 regression, and the reason this asserts the sentinel rather than just + // the stub: on a CLEAN machine the bare HERMES_HOME assertion passes whether + // buildIsolatedPath() filters gsd_run or gsd-tools, so it only catches the bug + // on a machine that already has the leak. Planting the sentinel makes it fail + // on any machine, and on either platform: the sentinel is planted in the one + // form that platform's shell resolves — the extensionless shim npm installs + // on POSIX, `gsd_run.exe` alone on Windows (see below for why alone). Note + // that only the POSIX sentinel can print SENTINEL_INVOKED; on Windows the + // basename assertion is what carries the guarantee (#4344). + test('(B) buildIsolatedPath strips a leaked PATH gsd_run; the ${HERMES_HOME} stub wins', (t) => { const fakeHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-home-b-')); const fakeHermesHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-hermes-')); const fakeRuntime = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-891-rt-')); - const { isolatedPath, nodeBinDir } = buildIsolatedPath(); - try { - const hermesBinDir = path.join(fakeHermesHome, 'gsd-core', 'bin'); - fs.mkdirSync(hermesBinDir, { recursive: true }); + t.after(() => cleanup(fakeHome)); + t.after(() => cleanup(fakeHermesHome)); + t.after(() => cleanup(fakeRuntime)); - const stubPath = path.join(hermesBinDir, 'gsd-tools.cjs'); - fs.writeFileSync( - stubPath, - '#!/usr/bin/env node\nconsole.log("HERMES_HOME_STUB:" + process.argv.slice(2).join(","));\n', - ); - fs.chmodSync(stubPath, 0o755); - - const snippet = fs.readFileSync(SNIPPET_FILE, 'utf8'); - // Export HOME to an isolated temp dir (no .claude install there) so the - // $HOME/.claude arm is skipped and we fall through to the HERMES_HOME arm. - const scriptContent = - `unset GSD_TOOLS\n` + - `export HOME=${JSON.stringify(fakeHome)}\n` + - `export RUNTIME_DIR=${JSON.stringify(fakeRuntime)}\n` + - `export HERMES_HOME=${JSON.stringify(fakeHermesHome)}\n` + - snippet + - `\nprintf "GSD_TOOLS=%s\\n" "$GSD_TOOLS"\n` + - `gsd_run ping test\n`; - - const scriptPath = path.join(fakeRuntime, 'test-hermes-home.sh'); - fs.writeFileSync(scriptPath, scriptContent); - - const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: isolatedPath, HOME: fakeHome, HERMES_HOME: fakeHermesHome }, - }); - - const normStdout = stdout.replace(/\\/g, '/'); - assert.ok( - normStdout.includes('gsd-core/bin/'), - `Expected GSD_TOOLS to resolve into hermes gsd-core/bin/, got:\n${stdout.trim()}`, - ); - assert.ok( - stdout.includes('HERMES_HOME_STUB:ping,test'), - `Expected stub output "HERMES_HOME_STUB:ping,test", got:\n${stdout.trim()}`, - ); - } finally { - cleanup(fakeHome); - cleanup(fakeHermesHome); - cleanup(fakeRuntime); - if (nodeBinDir) cleanup(nodeBinDir); + const sentinelBinDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4205-sentinel-')); + t.after(() => cleanup(sentinelBinDir)); + if (process.platform === 'win32') { + // What msys bash resolves for a bare `gsd_run`: it appends `.exe` during + // PATH lookup. Planted ALONE, so the leak this fixture proves is the + // extension-only one — an extensionless sibling would let an + // extension-blind filter strip the directory for the wrong reason and + // pass. It cannot print SENTINEL_INVOKED (it is this interpreter under + // another name), which is what the GSD_TOOLS assertion below is for. + linkExecutable(sentinelBinDir, 'gsd_run.exe'); + } else { + const sentinelPath = path.join(sentinelBinDir, 'gsd_run'); + fs.writeFileSync(sentinelPath, '#!/bin/sh\necho "SENTINEL_INVOKED:$*"\n'); + fs.chmodSync(sentinelPath, 0o755); } + + // Simulate the reported leak: a real gsd_run reachable on PATH. Passed in + // rather than assigned to process.env.PATH — the runner's own environment + // stays untouched, so no other fixture can observe the leak. + const { isolatedPath, nodeBinDir } = buildIsolatedPath( + `${sentinelBinDir}${path.delimiter}${process.env.PATH}`, + ); + t.after(() => cleanup(nodeBinDir)); + + // Leak assertion that needs no subprocess: a filter probing the wrong name + // leaves sentinelBinDir on the isolated PATH. Deterministic on both + // platforms, so it stays red even where the stdout assertions below depend + // on the mount's exec heuristics rather than on the filter under test. + assert.ok( + !isolatedPath.split(path.delimiter).includes(sentinelBinDir), + `Expected buildIsolatedPath to strip the leaked ${sentinelBinDir}, got:\n${isolatedPath}`, + ); + + const hermesBinDir = path.join(fakeHermesHome, 'gsd-core', 'bin'); + fs.mkdirSync(hermesBinDir, { recursive: true }); + + const stubPath = path.join(hermesBinDir, 'gsd-tools.cjs'); + fs.writeFileSync( + stubPath, + '#!/usr/bin/env node\nconsole.log("HERMES_HOME_STUB:" + process.argv.slice(2).join(","));\n', + ); + fs.chmodSync(stubPath, 0o755); + + const snippet = fs.readFileSync(SNIPPET_FILE, 'utf8'); + // HOME points at an isolated temp dir with no .claude install, so the + // $HOME/.claude arm misses and the HERMES_HOME arm is the one under test. + const scriptContent = + `unset GSD_TOOLS\n` + + `export HOME=${JSON.stringify(fakeHome)}\n` + + `export RUNTIME_DIR=${JSON.stringify(fakeRuntime)}\n` + + `export HERMES_HOME=${JSON.stringify(fakeHermesHome)}\n` + + snippet + + `\nprintf "GSD_TOOLS=%s\\n" "$GSD_TOOLS"\n` + + `gsd_run ping test\n`; + + const scriptPath = path.join(fakeRuntime, 'test-hermes-home.sh'); + fs.writeFileSync(scriptPath, scriptContent); + + const stdout = runBashFile(scriptPath, { + env: { PATH: isolatedPath, HOME: fakeHome, HERMES_HOME: fakeHermesHome }, + }); + + const normStdout = stdout.replace(/\\/g, '/'); + assert.ok( + !normStdout.includes('SENTINEL_INVOKED'), + `Expected the leaked PATH gsd_run to never be invoked, got:\n${stdout.trim()}`, + ); + // The same claim without depending on the sentinel producing output: + // whatever the resolver picked, it did not come from the leaked directory. + // Matched by basename, not by absolute path — git-bash prints `/c/Users/…` + // where os.tmpdir() gives `C:\Users\…` (see the note above the (E) PATH + // fallback assertion), so an absolute-path comparison never matches on + // Windows and would assert nothing there. The mkdtemp suffix keeps the + // basename unique. + assert.ok( + !normStdout.includes(path.basename(sentinelBinDir)), + `Expected GSD_TOOLS to resolve outside the leaked ${sentinelBinDir}, got:\n${stdout.trim()}`, + ); + // Assert the hermes dir itself: every arm of the resolver ends in + // gsd-core/bin/, so that substring alone cannot tell them apart. + assert.ok( + normStdout.includes(fakeHermesHome.replace(/\\/g, '/')), + `Expected GSD_TOOLS to resolve into ${fakeHermesHome}, got:\n${stdout.trim()}`, + ); + assert.ok( + normStdout.includes('HERMES_HOME_STUB:ping,test'), + `Expected stub output "HERMES_HOME_STUB:ping,test", got:\n${stdout.trim()}`, + ); }); // ── (C) Behavioral: default .hermes path used when HERMES_HOME not set ──── @@ -1456,7 +1606,11 @@ describe('bug-891: non-Claude runtime home fallback arms', () => { fs.writeFileSync(scriptPath, scriptContent); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: isolatedPath, HOME: fakeHome }, + // CLAUDE_CONFIG_DIR resolves before the HERMES_HOME arm under test + // (#4205 class, env vector); script-level `unset HERMES_HOME` above + // already handles that var, and snippetEnv()'s derived TEST_ENV_BASE + // clears CLAUDE_CONFIG_DIR before bash ever sees it. + env: { PATH: isolatedPath, HOME: fakeHome }, }); const normStdout = stdout.replace(/\\/g, '/'); @@ -1471,7 +1625,7 @@ describe('bug-891: non-Claude runtime home fallback arms', () => { } finally { cleanup(fakeHome); cleanup(fakeRuntime); - if (nodeBinDir) cleanup(nodeBinDir); + cleanup(nodeBinDir); } }); @@ -1540,7 +1694,7 @@ describe('bug-891: non-Claude runtime home fallback arms', () => { * * A user project normally does not contain gsd-core/bin/gsd-tools.cjs. * The snippets should still prefer RUNTIME_DIR for local/dev installs, then - * fall back to the installed gsd-tools binary on PATH. + * fall back to the installed gsd_run binary on PATH. */ 'use strict'; @@ -1620,6 +1774,11 @@ function makeTempDir() { } function runResolver({ cwd, runtimeDir, pathDir }) { + // RUNTIME_DIR='' is INDISTINGUISHABLE from unset to the resolver's + // ${RUNTIME_DIR:-$(git rev-parse --show-toplevel)} — an empty value silently + // falls back to the real repo root and resolves its real install. Fail loudly + // instead; every caller passes one. + if (!runtimeDir) throw new Error('runResolver requires runtimeDir: an empty value resolves the real repo root'); const script = [ 'set -e', extractResolverSnippet(), @@ -1628,7 +1787,7 @@ function runResolver({ cwd, runtimeDir, pathDir }) { ].join('\n'); // Consolidation #1969: POSIX-shell resolver. These tests create an - // extension-less `gsd-tools` PATH stub (mode 0o755) and exec it via `bash -c`; + // extension-less `gsd_run` PATH stub (mode 0o755) and exec it via `bash -c`; // Windows Git Bash ignores the exec bit for extension-less PATH scripts, so the // suite is guarded to POSIX (matches the host suite's own bash -c guard). if (process.platform === 'win32') return ''; @@ -1636,11 +1795,10 @@ function runResolver({ cwd, runtimeDir, pathDir }) { const r = runHookSeam('-c', [script], { interpreter: 'bash', cwd, - env: { - ...process.env, + env: snippetEnv({ PATH: `${pathDir}${path.delimiter}${process.env.PATH || ''}`, - RUNTIME_DIR: runtimeDir || '', - }, + RUNTIME_DIR: runtimeDir, + }), }); throwIfFailed(r, 'bash -c '); return r.stdout; @@ -1745,54 +1903,37 @@ const assert = require('node:assert/strict'); const fs = require('node:fs'); const path = require('node:path'); const os = require('node:os'); -const { runHook: runHookSeam } = require('./helpers/process-seam.cjs'); -const { throwIfFailed } = require('./helpers/git-fixture.cjs'); const { cleanup } = require('./helpers.cjs'); const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows'); const SNIPPET_FILE = path.join(WORKFLOWS_DIR, '_runtime-launcher.snippet.sh'); -/** - * Run a bash script FILE via the process seam, preserving the throw-on- - * nonzero-exit semantics of the execFileSync('bash', [path], ...) idiom - * this replaces. - */ -function runBashFile(scriptPath, options = {}) { - const r = runHookSeam(scriptPath, [], { interpreter: 'bash', ...options }); - throwIfFailed(r, `bash ${scriptPath}`); - return r.stdout; -} // The probe string that must appear in the snippet for the new repo-local check. // The snippet uses _GSD_RUNTIME_ROOT as the intermediate variable. const LOCAL_CLAUDE_PROBE = '_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/'; /** - * Build a PATH that strips gsd-tools but keeps node and system binaries. - * Accepts additional bin dirs to prepend. + * Return the full system PATH with extra bin dirs prepended. Deliberately does + * NOT filter gsd_run — unlike buildIsolatedPath() above, which does. * - * We cannot simply remove the whole directory that contains gsd-tools because - * that directory may also contain node (e.g. /opt/homebrew/bin on macOS). - * Instead, we keep the system PATH as-is and rely on the test's RUNTIME_DIR - * having no gsd-core/bin/ sub-path, so the resolver's first two checks - * (RUNTIME_DIR/gsd-core/bin/ and RUNTIME_DIR/.claude/gsd-core/bin/) - * are the only ones exercised before we hit our stub. + * The isolation here comes from resolution ORDER, not from the PATH contents. + * Callers give RUNTIME_DIR a .claude/gsd-core/bin/ stub, and the resolver + * checks RUNTIME_DIR/gsd-core/bin/ then RUNTIME_DIR/.claude/gsd-core/bin/ + * before it ever reaches `command -v gsd_run`, so an ambient gsd_run on PATH + * is unreachable for these tests. Keeping PATH whole is what keeps node + * resolvable when node co-locates with a global gsd_run (e.g. /opt/homebrew/bin). * - * The extra extraBefore dirs (e.g. noToolsBin) sit first but have no gsd-tools - * binary, so command -v gsd-tools still falls back to PATH lookup. However, - * the snippet's elif arm that uses `command -v gsd-tools` will find the real - * installed one unless we mask it. To mask it without losing node, we create - * a noToolsBin dir that shadows gsd-tools with a sentinel that must NOT be - * called — and we only call makeIsolatedPath for tests where the .claude stub - * must win before PATH is consulted (i.e. the elif PATH arm is never reached). + * That makes this helper safe ONLY for tests whose stub wins before the PATH + * arm. Any test that must prove the PATH arm itself misses needs + * buildIsolatedPath(), which excludes gsd_run-bearing directories outright. * - * For B: stub is at RUNTIME_DIR/.claude/... so resolver picks it at elif-1 (before command -v). - * For C: same — local .claude/ is checked before command -v and before $HOME/.claude. + * For B and C: the stub sits at RUNTIME_DIR/.claude/..., picked at elif-1. */ function makeIsolatedPath(extraBefore = []) { // Keep full system PATH so node remains accessible. // Tests B and C exercise only the RUNTIME_DIR/.claude arm which fires - // before command -v gsd-tools — so the real gsd-tools on PATH is never reached. + // before command -v gsd_run — so the real gsd_run on PATH is never reached. const systemPaths = (process.env.PATH || '/usr/bin:/bin').split(path.delimiter); return [...extraBefore, ...systemPaths].join(path.delimiter); } @@ -1867,11 +2008,12 @@ describe('bug-444: resolver finds repo-local .claude install', () => { const scriptPath = path.join(fakeRoot, 'test-local-claude.sh'); fs.writeFileSync(scriptPath, scriptContent); - // Keep node in PATH (needed to run the .cjs stub); remove gsd-tools + // Keep node in PATH (needed to run the .cjs stub); the .claude arm + // resolves before the PATH arm, so gsd_run is never probed here. const isolatedPath = makeIsolatedPath([noToolsBin]); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: isolatedPath, HOME: fakeHome }, + env: { PATH: isolatedPath, HOME: fakeHome }, }); // Must have resolved to the local .claude stub @@ -1934,7 +2076,7 @@ describe('bug-444: resolver finds repo-local .claude install', () => { const isolatedPath = makeIsolatedPath([noToolsBin]); const stdout = runBashFile(scriptPath, { - env: { ...process.env, PATH: isolatedPath, HOME: fakeHome }, + env: { PATH: isolatedPath, HOME: fakeHome }, }); assert.ok( From 38e4ce5f620618c3c57a8e34b32545ca94c46e57 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 17:08:24 -0400 Subject: [PATCH 023/166] fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim --- .changeset/zesty-tunas-bark.md | 5 + CONTEXT.md | 2 +- docs/CLI-TOOLS.md | 2 +- .../2207-status-field-lifecycle-ownership.md | 2 +- hooks/gsd-statusline.js | 18 +- ...ocs-guard-registration.exempt-baseline.cjs | 4 +- src/state-document.cts | 130 ++++++-- src/state-md-schema.cts | 35 ++- src/state.cts | 38 ++- tests/gsd-statusline.test.cjs | 42 ++- tests/state-document.test.cjs | 144 +++++++++ tests/state.test.cjs | 297 +++++++++++++++--- 12 files changed, 602 insertions(+), 117 deletions(-) create mode 100644 .changeset/zesty-tunas-bark.md diff --git a/.changeset/zesty-tunas-bark.md b/.changeset/zesty-tunas-bark.md new file mode 100644 index 000000000..3920b9795 --- /dev/null +++ b/.changeset/zesty-tunas-bark.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4381 +--- +**`state` no longer guesses the STATE.md `status` token from substrings of the status prose** — a status line mentioning `.planning/` (or Italian `verifica`, `completezza`, `fasi complete`) no longer silently becomes `status: planning`/`verifying`/`completed`; recognized vocabulary values keep normalizing and unrecognized prose stays visible, `state record-session` without arguments now errors instead of writing, and stray `*-SUMMARY.md` files without a plan twin stay excluded from `progress.completed_plans` recounts. (#4186) diff --git a/CONTEXT.md b/CONTEXT.md index 1f2e6a438..294ae108c 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -91,7 +91,7 @@ Module owning STATE.md lifecycle/maintenance transitions as intent-based methods The one declaration (ADR-3473 §8.8, #3873) for "which STATE.md keys exist and what they carry", replacing three hand-maintained tables that were already observed to disagree: `FIELD_CLASSIFICATION` and `FRONTMATTER_BODY_SOURCE` (STATE.md Transition Module) and `FRONTMATTER_KEY_TO_BODY_LABEL` (STATE.md Document Module's `bodyLabelFor`, `state.cts`). One frozen, null-prototype `STATE_FIELD_SCHEMA` row per key carries `type`, `cardinality`, `source`, `preservation`, `guard`/`mergeStrategy` (ADR-3408's closed vocabularies, whose type declarations moved here), `bodySource`/`bodyLabel`, `acceptedShapes` (declared value shapes a hand-written parser accepts — e.g. `current_plan`'s `N` / `N of M`, #3784 — never a predicate; Greenspun's Tenth Rule still applies), and `emitted` (mirrors `buildStateFrontmatter`'s null-guards). The three original tables are now PROJECTIONS derived from this schema at module load, byte-identical in shape/key-order/frozen-and-null-prototype-ness to what they were before #3873 — every existing consumer (the preservation dispatch loop, `getFieldClassification`, `getPreserveWhenUnchangedFields`, `bodyLabelFor`, #3872's `declaredLeavesOf`) is unaffected. **The `last_activity` disagreement is resolved by declaration, not by picking the table that "looks right":** it carries a `bodySource` (it IS body-derived) but deliberately no `bodyLabel`, matching what ships today — its `preservation` is `derive`, so it can never reach `bodyLabelFor`'s `STATE_BODY_LABEL_UNWIRED_ROW` throw, pinned by `tests/state.test.cjs`'s `lastActivityLabelResolutionMatchesShippedBehavior`. Leaf module: imports from neither `state-transition.cts` nor `state.cts` (both import it), avoiding the CJS require-cycle `src/health-diagnostic-types.cts` was split out to break. `scripts/lint-state-field-drift.cjs` is UNCHANGED and retained — it guards the #3187 STATE.md field-extraction fallback-chain re-derivation, an orthogonal concern this schema does not make unrepresentable (§8.8's own claim that the guard is deleted here was verified false). Source of truth: `gsd-core/bin/lib/state-md-schema.cjs` (generated from `src/state-md-schema.cts`). ### STATE.md Status Lifecycle (ADR-2207) -The `Status` field in STATE.md follows a strict lifecycle: `Ready to plan` → `All phases complete` (all phases done, milestone awaiting formal close) → ` milestone complete` (terminal, written only by the milestone-close verb `milestoneCompleteCore`) → `Awaiting next milestone` (archived). Phase-completion verbs write `All phases complete` on the last phase — never `Milestone complete` (the overloaded bare value was removed in #2204 per ADR-2207 to decouple phase-level writes from milestone termination). `normalizeStateStatus` maps any status containing "complete" → `completed`, so consumers using the normalized projection (workstream inventory's `status` field, statusline) recognize `All phases complete` without code changes. Note: `isCompletedInventory` (workstream-inventory-builder.cts) intentionally checks only for the terminal `\bmilestone\s+complete\b` / `\barchived\b` — `All phases complete` returns `false` (intermediate, not terminal). +The `Status` field in STATE.md follows a strict lifecycle: `Ready to plan` → `All phases complete` (all phases done, milestone awaiting formal close) → ` milestone complete` (terminal, written only by the milestone-close verb `milestoneCompleteCore`) → `Awaiting next milestone` (archived). Phase-completion verbs write `All phases complete` on the last phase — never `Milestone complete` (the overloaded bare value was removed in #2204 per ADR-2207 to decouple phase-level writes from milestone termination). `normalizeStateStatus` recognizes the whole-field status vocabulary (exact values plus anchored `Executing Phase N` / `Phase N complete` / ` milestone complete` forms — `STATUS_EXACT_TOKENS` / `STATUS_ANCHORED_PATTERNS`, src/state-document.cts; #4186 ended the pre-#4186 substring scanning, so prose merely CONTAINING a status word passes through verbatim), so consumers using the normalized projection (workstream inventory's `status` field, statusline) recognize `All phases complete` and the other lifecycle values without code changes. Note: `isCompletedInventory` (workstream-inventory-builder.cts) intentionally checks only for the terminal `\bmilestone\s+complete\b` / `\barchived\b` — `All phases complete` returns `false` (intermediate, not terminal). ### Query Execution Policy Module Module owning query transport routing policy projection (`preferNative`, fallback policy, workstream subprocess forcing) at execution seam. diff --git a/docs/CLI-TOOLS.md b/docs/CLI-TOOLS.md index 960fa56f5..ff0925d1a 100644 --- a/docs/CLI-TOOLS.md +++ b/docs/CLI-TOOLS.md @@ -113,7 +113,7 @@ node gsd-tools.cjs state add-decision --summary-file path [--rationale-file path node gsd-tools.cjs state add-blocker --text "..." node gsd-tools.cjs state resolve-blocker --text "..." -# Record session continuity +# Record session continuity (at least one of --stopped-at / --resume-file is required) node gsd-tools.cjs state record-session --stopped-at "..." [--resume-file path] # Phase start — update STATE.md Status/Last activity for a new phase diff --git a/docs/adr/2207-status-field-lifecycle-ownership.md b/docs/adr/2207-status-field-lifecycle-ownership.md index b936815d8..a1fd13722 100644 --- a/docs/adr/2207-status-field-lifecycle-ownership.md +++ b/docs/adr/2207-status-field-lifecycle-ownership.md @@ -30,4 +30,4 @@ STATE.md's `Status` field is written by two transitions with an **overloaded** v **Positive:** the overload is removed; the intermediate and terminal "complete" states are distinct; a phase-level verb no longer writes the terminal milestone state; the wrong-phase flip becomes a parse-correctness concern already owned upstream. -**Cost / follow-through (implemented in #2204):** consumers that key on the `Milestone complete` string must recognize `All phases complete` — `workflows/progress.md` (Route D), `verify.cts`, and `workstream-inventory-builder.cts`. `normalizeStateStatus` already maps any status containing "complete" → `completed`, so it needs no change. A `CONTEXT.md` glossary entry enumerating the `Status` lifecycle lands with the #2204 implementation. +**Cost / follow-through (implemented in #2204):** consumers that key on the `Milestone complete` string must recognize `All phases complete` — `workflows/progress.md` (Route D), `verify.cts`, and `workstream-inventory-builder.cts`. `normalizeStateStatus` already maps any status containing "complete" → `completed`, so it needs no change. *(Superseded mechanism, #4186: recognition is now an ANCHORED whole-field vocabulary match — `All phases complete` and ` milestone complete` still map to `completed`; prose merely containing "complete" passes through verbatim. The consumer-recognition consequence above is unchanged.)* A `CONTEXT.md` glossary entry enumerating the `Status` lifecycle lands with the #2204 implementation. diff --git a/hooks/gsd-statusline.js b/hooks/gsd-statusline.js index c73d56401..37e110e9e 100755 --- a/hooks/gsd-statusline.js +++ b/hooks/gsd-statusline.js @@ -449,13 +449,17 @@ function contextTokenSuffix(currentUsage) { // --- Compact state format (opt-in) --------------------------------------------- /** - * Collapse GSD's free-text status (often a multi-sentence narrative) to a - * single keyword, built on the canonical normalizer (#2162 approval - * condition): normalizeStateStatus() in state-document.cjs owns the status - * vocabulary (discussing / planning / executing / verifying / completed / - * paused) so the two can't drift. "paused" — the canonical stuck state — is - * uppercased to PAUSED, the one state worth shouting about. Statuses the - * normalizer passes through unrecognized fall back to their first word, + * Collapse GSD's status value to a single keyword, built on the canonical + * normalizer (#2162 approval condition): normalizeStateStatus() in + * state-document.cjs owns the status vocabulary (discussing / planning / + * executing / verifying / completed / paused) so the two can't drift. + * #4186: the normalizer recognizes the DECLARED vocabulary by anchored + * whole-field match — vocabulary values (the state writer persists tokens) + * collapse to their keyword; free-text narratives are no longer + * keyword-guessed from substrings (a `.planning/` mention in non-English + * prose used to render `planning`), and pass through unrecognized to the + * first-word fallback below. "paused" — the canonical stuck state — is + * uppercased to PAUSED, the one state worth shouting about. The fallback is * capped at 16 chars so a rogue STATE.md can't blow up the line. * Returns null for empty input. */ diff --git a/scripts/lint-docs-guard-registration.exempt-baseline.cjs b/scripts/lint-docs-guard-registration.exempt-baseline.cjs index fb641a3c2..157747b9a 100644 --- a/scripts/lint-docs-guard-registration.exempt-baseline.cjs +++ b/scripts/lint-docs-guard-registration.exempt-baseline.cjs @@ -191,7 +191,9 @@ const DOCS_GUARD_EXEMPT_DOCS_PATHS = { 'runtime-name-policy.test.cjs': ['docs/customize/skills'], 'security-prompt-injection.security.test.cjs': ['docs/notes.md'], 'shipped-reference-cites.test.cjs': [], - 'state.test.cjs': ['docs/CONFIGURATION.md', 'docs/reference/state-md.md'], + // #4186: the record-session usage-contract test cites the documented + // signature in docs/CLI-TOOLS.md in its explanatory comment. + 'state.test.cjs': ['docs/CLI-TOOLS.md', 'docs/CONFIGURATION.md', 'docs/reference/state-md.md'], 'worktree-safety.test.cjs': ['docs/SUMMARY.md'], }; diff --git a/src/state-document.cts b/src/state-document.cts index 5921b31d2..a8fc16bb8 100644 --- a/src/state-document.cts +++ b/src/state-document.cts @@ -611,31 +611,115 @@ export function stateReplaceFieldInSession(content: string, primary: string, fal return withSection(content, target, (sectionBody) => stateReplaceFieldWithFallback(sectionBody, primary, fallback, value)); } +/** + * #4186: the DECLARED raw-status vocabulary `normalizeStateStatus` recognizes. + * Keys are whole-field values, compared against the caller's input after + * lowercasing, trimming, and collapsing internal whitespace runs to single + * spaces — so case and whitespace variants the vocabulary documents + * (`EXECUTING PHASE 5`, ` Paused `, `In progress`) keep normalizing. + * Values are members of `STATUS_LIFECYCLE_ENUM` (`src/state-md-schema.cts`) + * — the set the normalizer maps recognized input ONTO. + * + * The mapping preserves the PRE-#4186 branch ORDER's observable artifacts for + * every value the old substring chain recognized: `Planning complete` → + * `planning` (the `planning` branch outranked `complete`) and + * `Phase complete — ready for verification` → `verifying` (`verif` outranked + * `complete`; pinned by tests/state.test.cjs's advance-plan case-5 comment). + * + * Everything else — prose that merely CONTAINS a status word — falls through + * to the caller's raw value (the recorded lenient fallback, #3873 phase-3 + * row 26). A token guessed from a substring inside a sentence is worse than + * a visible paragraph: the paragraph is visibly prose, the wrong token is + * not (a `.planning/` path in Italian prose silently produced + * `status: planning`; `verificata` produced `verifying`; `completezza` + * produced `completed`). + */ +export const STATUS_EXACT_TOKENS: Readonly> = Object.freeze({ + paused: 'paused', + stopped: 'paused', + executing: 'executing', + 'in progress': 'executing', + 'ready to execute': 'executing', + planning: 'planning', + 'ready to plan': 'planning', + 'planning complete': 'planning', + discussing: 'discussing', + verifying: 'verifying', + completed: 'completed', + done: 'completed', + complete: 'completed', + 'phase complete': 'completed', + // advance-plan's phase-complete write (state-transition.cts:1812) — maps + // to `verifying`, preserving the pre-#4186 branch order where `verif` + // outranked `complete`. + 'phase complete — ready for verification': 'verifying', + 'all phases complete': 'completed', + // Legacy bare terminal form. ADR-2207/#2204 removed it from every WRITER + // (phase verbs write `All phases complete`; milestone close writes + // ` milestone complete`) — kept here as READER recognition so a + // legacy STATE.md still normalizes, exactly the way KNOWN_TEMPLATE_DEFAULTS + // keeps the other legacy Status strings. + 'milestone complete': 'completed', + unknown: 'unknown', +} as const); + +/** + * #4186: ANCHORED patterns for handler-written raw statuses whose text + * carries a variable component (a phase number, a milestone version, a + * #1070 completion glyph). Each pattern is matched against the same + * normalized key as `STATUS_EXACT_TOKENS` (lowercased, trimmed, + * whitespace-collapsed) and must match the WHOLE value — never a substring — + * mirroring how `KNOWN_STATUS_PATTERNS` anchors its template-default checks. + * `Executing Phase 5 — final stretch` (executor-appended prose) matches + * NOTHING and passes through verbatim, the same discipline #1070 applies to + * "Complete but needs manual QA". + */ +export const STATUS_ANCHORED_PATTERNS: ReadonlyArray = Object.freeze([ + // begin-phase / planned transitions write `Executing Phase ${N}`. + [/^executing phase\s+\S+$/, 'executing'], + [/^planning phase\s+\S+$/, 'planning'], + [/^verifying phase\s+\S+$/, 'verifying'], + // phase-complete verbs write `Phase ${N} complete` (state.cts) — the exact + // shape the #3578 demote guard then re-checks against the disk counters. + [/^phase\s+\S+\s+complete$/, 'completed'], + // milestoneCompleteCore writes `${version} milestone complete` (terminal, + // ADR-2207); the version label is a single token (milestone.cts's charset + // validation admits letters, digits, '.', '-', '_'). + [/^\S+\s+milestone complete$/, 'completed'], + // #1070: LLM executors may write "Complete ✓" or bare "Complete" when + // finishing a phase. + [/^complete\s*[✓✔✅☑]?$/, 'completed'], +] as const); + +/** + * Normalize a raw `Status` body-field value to the canonical status token. + * + * #4186: recognition is ANCHORED — the whole field value (lowercased, + * trimmed, whitespace-collapsed) must be a member of the declared vocabulary + * (`STATUS_EXACT_TOKENS` / `STATUS_ANCHORED_PATTERNS` above). The pre-#4186 + * implementation ran a first-match-wins chain of SUBSTRING tests over the + * free-prose field, so any prose merely CONTAINING a trigger word was + * silently rewritten to a credible wrong token: a `.planning/...` path + * mentioned in a non-English status line landed on `planning` (the trigger + * word lives in the directory name and outranked the `verif`/`complete` + * branches), Italian `verifica*` landed on `verifying`, `completezza` and + * `fasi complete` landed on `completed`. The lenient FALLBACK is unchanged + * and recorded (#3873 phase-3 row 26): an unrecognized value passes through + * verbatim — visible prose, never a guessed token. + * + * `pausedAt` keeps its documented force (issue #4186: intended behavior): a + * truthy value yields `paused` regardless of the prose. + */ export function normalizeStateStatus(status: string | null | undefined, pausedAt: unknown): string { - let normalizedStatus = status || 'unknown'; - const statusLower = (status || '').toLowerCase(); - if (statusLower.includes('paused') || statusLower.includes('stopped') || pausedAt) { - normalizedStatus = 'paused'; + if (pausedAt) return 'paused'; + if (!status) return 'unknown'; + const key = status.trim().toLowerCase().replace(/\s+/g, ' '); + const exact = STATUS_EXACT_TOKENS[key]; + if (exact) return exact; + for (const [pattern, token] of STATUS_ANCHORED_PATTERNS) { + if (pattern.test(key)) return token; } - else if (statusLower.includes('executing') || statusLower.includes('in progress')) { - normalizedStatus = 'executing'; - } - else if (statusLower.includes('planning') || statusLower.includes('ready to plan')) { - normalizedStatus = 'planning'; - } - else if (statusLower.includes('discussing')) { - normalizedStatus = 'discussing'; - } - else if (statusLower.includes('verif')) { - normalizedStatus = 'verifying'; - } - else if (statusLower.includes('complete') || statusLower.includes('done')) { - normalizedStatus = 'completed'; - } - else if (statusLower.includes('ready to execute')) { - normalizedStatus = 'executing'; - } - return normalizedStatus; + return status; } /** diff --git a/src/state-md-schema.cts b/src/state-md-schema.cts index 3ea1306a0..ea31d61a7 100644 --- a/src/state-md-schema.cts +++ b/src/state-md-schema.cts @@ -67,24 +67,31 @@ export type FieldMergeStrategy = 'progress-ratchet'; /** * The seven CANONICAL values `normalizeStateStatus` (`src/state-document.cts`) * maps recognized raw status prose TO — the function's default fallback plus - * each branch's literal output, in the order the function tests them. This is - * NOT the raw body prose vocabulary `CONTEXT.md`'s "STATE.md Status Lifecycle - * (ADR-2207)" entry documents (`Ready to plan` → `All phases complete` → - * ` milestone complete` → `Awaiting next milestone`, plus the - * handler-authored strings in `KNOWN_TEMPLATE_DEFAULTS['Status']`) — that is - * free-form prose `normalizeStateStatus` READS. + * each vocabulary entry's literal output. This is NOT the raw body prose + * vocabulary `CONTEXT.md`'s "STATE.md Status Lifecycle (ADR-2207)" entry + * documents (`Ready to plan` → `All phases complete` → ` milestone + * complete` → `Awaiting next milestone`, plus the handler-authored strings + * in `KNOWN_TEMPLATE_DEFAULTS['Status']`) — that is free-form prose + * `normalizeStateStatus` READS. * * CORRECTED (#3873 phase-3 test-matrix row 26 — verified by executing * `normalizeStateStatus`, not by reading this docstring's prior claim): * this is NOT a closed set the `status` frontmatter key is restricted to at - * runtime. `normalizeStateStatus` is deliberately LENIENT: its fallback is - * `normalizedStatus = status || 'unknown'`, and when none of its - * substring-match branches recognize the raw input, that fallback — the - * caller's raw, UNRECOGNIZED prose — is returned unchanged. A status value - * outside this seven-member set is not rejected, coerced, or normalized; it - * passes straight through into the frontmatter. `STATUS_LIFECYCLE_ENUM` is - * therefore the set of values the normalizer maps recognized input ONTO, not - * a runtime-enforced closed vocabulary for the field. + * runtime. `normalizeStateStatus` is deliberately LENIENT: its fallback + * returns the caller's raw, UNRECOGNIZED prose unchanged — when none of its + * vocabulary entries recognize the whole-field input, that raw value is what + * the function returns. A status value outside this seven-member set is not + * rejected, coerced, or normalized; it passes straight through into the + * frontmatter. `STATUS_LIFECYCLE_ENUM` is therefore the set of values the + * normalizer maps recognized input ONTO, not a runtime-enforced closed + * vocabulary for the field. + * + * #4186: recognition is ANCHORED (whole-field match against the declared + * `STATUS_EXACT_TOKENS` / `STATUS_ANCHORED_PATTERNS` tables in + * `src/state-document.cts`), never a substring scan of the prose — prose + * merely CONTAINING a status word (a `.planning/` path, `verifica*`, + * `completezza`) passes through verbatim instead of being rewritten to a + * credible wrong token. */ export const STATUS_LIFECYCLE_ENUM = Object.freeze([ 'unknown', diff --git a/src/state.cts b/src/state.cts index e55f8105a..9f91423f6 100644 --- a/src/state.cts +++ b/src/state.cts @@ -1921,6 +1921,18 @@ function cmdStateResolveBlocker(cwd: string, text: string, raw: boolean): void { } function cmdStateRecordSession(cwd: string, options: StateRecordSessionOptions, raw: boolean): void { + // #4186: a bare invocation is a usage error, not a heartbeat write. The + // pre-#4186 handler accepted zero arguments and still refreshed + // `Last session` / `Last Date` / `last_updated` — a caller probing the + // command's signature (the way other subcommands encourage) silently + // mutated STATE.md. Mirrors `state update`'s required-arg guard + // (cmdStateUpdate: `error('field and value required for state update')`), + // including its ordering: validation precedes the STATE.md existence + // check. Either flag suffices — `--resume-file` alone carries an explicit + // value the handler must persist. + if (!options.stopped_at && (options.resume_file === undefined || options.resume_file === null)) { + error('stopped-at or resume-file required for state record-session'); + } const statePath = planningPaths(cwd).state; if (!fs.existsSync(statePath)) { output({ error: 'STATE.md not found' }, raw, undefined); return; } @@ -3147,19 +3159,19 @@ function buildStateFrontmatter( } let normalizedStatus = normalizeStateStatus(status, pausedAt); - // #3578: normalizeStateStatus matches 'complete' as a case-insensitive - // SUBSTRING, so the phase-completion prose cmdStateCompletePhase writes to - // the body (`Phase ${N} complete`) collapses to the milestone-level - // 'completed' status even when other phases remain open. Phase-level - // prose must never decide milestone-level status — completedPhases / - // totalPhases / diskScope, already derived above from a disk scan, are - // the authority on whether the MILESTONE is actually done. Only override - // when: (a) normalizeStateStatus actually landed on 'completed'; (b) the - // raw prose is UNAMBIGUOUSLY phase-completion prose — the anchored - // pattern below deliberately excludes "All phases complete" (no `\S+` - // phase token) and milestone-close prose like "v1.0 milestone complete" - // (no leading "phase"); and (c) the counters are trustworthy (a COMPLETE - // disk scope, both counts are finite numbers, and a positive + // #3578: the declared status vocabulary (#4186) recognizes + // `Phase ${N} complete` (state.cts's own phase-completion write) and maps + // it to `completed`, so the phase-completion prose still collapses to the + // milestone-level status even when other phases remain open — this guard + // demotes it back. Phase-level prose must never decide milestone-level + // status — completedPhases / totalPhases / diskScope, already derived above + // from a disk scan, are the authority on whether the MILESTONE is actually + // done. Only override when: (a) normalizeStateStatus actually landed on + // 'completed'; (b) the raw prose is UNAMBIGUOUSLY phase-completion prose — + // the anchored pattern below deliberately excludes "All phases complete" + // (no `\S+` phase token) and milestone-close prose like "v1.0 milestone + // complete" (no leading "phase"); and (c) the counters are trustworthy (a + // COMPLETE disk scope, both counts are finite numbers, and a positive // denominator) and affirmatively disagree with 'completed'. In every // other case normalizedStatus is left exactly as normalizeStateStatus // returned it. diff --git a/tests/gsd-statusline.test.cjs b/tests/gsd-statusline.test.cjs index 782f5c879..91c5011a9 100644 --- a/tests/gsd-statusline.test.cjs +++ b/tests/gsd-statusline.test.cjs @@ -1741,16 +1741,30 @@ test('config-set statusline.show_context_tokens yes → rejected', () => { assert.equal(shortGsdStatus(''), null); assert.equal(shortGsdStatus(undefined), null); }); - test('paused — the canonical stuck state — wins and renders uppercase (#2162 condition)', () => { - assert.equal(shortGsdStatus('paused — waiting on credentials'), 'PAUSED'); - assert.equal(shortGsdStatus('stopped by user'), 'PAUSED'); + test('paused — the canonical stuck state — renders uppercase (#2162 condition)', () => { + // The #2162 shout applies to the canonical token, which is what the + // state writer persists (#4186 anchored vocabulary). + assert.equal(shortGsdStatus('paused'), 'PAUSED'); + assert.equal(shortGsdStatus('Paused'), 'PAUSED'); + // #4186: narrative prose is no longer keyword-guessed — a paused-led + // narrative renders its first word (visible, never a silent wrong + // token), so the shout is reserved for the recognized token itself. + assert.equal(shortGsdStatus('paused — waiting on credentials'), 'paused'); + assert.equal(shortGsdStatus('stopped by user'), 'stopped'); }); - test('collapses lifecycle narratives to canonical keywords via normalizeStateStatus', () => { - assert.equal(shortGsdStatus('Executing phase 7 of the parser milestone'), 'executing'); - assert.equal(shortGsdStatus('Ready to plan next phase'), 'planning'); - assert.equal(shortGsdStatus('Discussing scope with user'), 'discussing'); - assert.equal(shortGsdStatus('Verifying UAT criteria'), 'verifying'); - assert.equal(shortGsdStatus('Work complete'), 'completed'); + test('collapses vocabulary values to canonical keywords via normalizeStateStatus (#4186 anchored)', () => { + // #4186: normalizeStateStatus recognizes the DECLARED vocabulary + // (whole-field match), so exactly those values collapse to keywords. + assert.equal(shortGsdStatus('Executing Phase 7'), 'executing'); + assert.equal(shortGsdStatus('ready to plan'), 'planning'); + assert.equal(shortGsdStatus('Discussing'), 'discussing'); + assert.equal(shortGsdStatus('Verifying Phase 2'), 'verifying'); + assert.equal(shortGsdStatus('Work complete'), 'Work'); + // Narrative prose — vocabulary words embedded in longer sentences — is + // no longer guessed at (a `.planning/` mention in non-English prose + // used to render `planning`); it falls back to the first word. + assert.equal(shortGsdStatus('Executing phase 7 of the parser milestone'), 'Executing'); + assert.equal(shortGsdStatus('Ready to plan next phase'), 'Ready'); }); test('matches the canonical vocabulary exactly — no drift from normalizeStateStatus', () => { const { normalizeStateStatus } = require('../gsd-core/bin/lib/state-document.cjs'); @@ -1777,9 +1791,15 @@ test('config-set statusline.show_context_tokens yes → rejected', () => { test('renders version · phase/total · status', () => { const out = formatGsdStateCompact({ milestone: 'v1.12', phaseNum: '7', phaseTotal: '12', - status: 'Executing phase 7 — building the parser', + status: 'Executing Phase 7', }); assert.equal(out, 'v1.12 · P7/12 · executing'); + // #4186: narrative status is not keyword-guessed — first word renders. + const narrative = formatGsdStateCompact({ + milestone: 'v1.12', phaseNum: '7', phaseTotal: '12', + status: 'Executing phase 7 — building the parser', + }); + assert.equal(narrative, 'v1.12 · P7/12 · Executing'); }); test('prefers lifecycle active_phase over body phase number', () => { const out = formatGsdStateCompact({ @@ -1789,7 +1809,7 @@ test('config-set statusline.show_context_tokens yes → rejected', () => { }); test('paused state renders uppercase in the compact line', () => { const out = formatGsdStateCompact({ - milestone: 'v2.0', activePhase: '4.5', status: 'paused — waiting on review', + milestone: 'v2.0', activePhase: '4.5', status: 'paused', }); assert.equal(out, 'v2.0 · P4.5 · PAUSED'); }); diff --git a/tests/state-document.test.cjs b/tests/state-document.test.cjs index 01569eefd..4a4ada130 100644 --- a/tests/state-document.test.cjs +++ b/tests/state-document.test.cjs @@ -2236,3 +2236,147 @@ describe('#3642 — single-section leak controls and seam pins', () => { assert.strictEqual(roadmapParser.hasMilestoneSectioning(two), true, '>=2 predicate unchanged: two headings is sectioning'); }); }); + +// ───────────────────────────────────────────────────────────────────────────── +// #4186 — normalizeStateStatus must derive the `status` token by ANCHORED +// matching against the declared status vocabulary (whole-field value, +// case-insensitive, whitespace-collapsed), never by substring-scanning the +// free-prose body Status field. Prose that merely MENTIONS a status word (a +// `.planning/` path, Italian `verifica*`, `completezza`, `fasi complete`) +// must pass through verbatim — the visible paragraph beats a valid, credible, +// wrong token. The lenient fallback itself is a recorded contract +// (state-md-schema.cts #3873 phase-3 row 26) and is preserved. +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4186: normalizeStateStatus anchored status vocabulary — substring traps', () => { + const { normalizeStateStatus } = require('../gsd-core/bin/lib/state-document.cjs'); + + // Row 1 — the failing-first regression: a `.planning/` path inside Italian + // prose must NOT land on `planning` (the trigger word lives in the + // DIRECTORY NAME; the reporter measured this exact counter-pair). + test('row 1: a .planning/ path inside prose never yields `planning`', () => { + assert.strictEqual( + normalizeStateStatus('Lavoro sospeso, vedi .planning/STATE.md', null), + 'Lavoro sospeso, vedi .planning/STATE.md', + ); + }); + + test('rows 2-5: Italian prose mentioning paths/verification/completeness passes through verbatim', () => { + const prose = [ + 'Esecuzione in corso su .planning/', + 'Fase 34 COMPLETA, VERIFICATA 8/8 e FUSA IN PRODUZIONE', + 'Aggiornato .planning/STATE.md dopo la verifica di fase', + 'Referto in .planning/phases/34, completezza ok', + ]; + for (const p of prose) { + assert.strictEqual(normalizeStateStatus(p, null), p, `prose must pass through verbatim: ${p}`); + } + }); + + test('row 6: Italian verifica-family words do not match the `verif` trigger', () => { + for (const p of ['verifica', 'verificata', 'verifiche', 'ri-verifica']) { + assert.strictEqual(normalizeStateStatus(p, null), p); + } + }); + + test('rows 7-8: `completezza` and `fasi complete` do not yield `completed`', () => { + assert.strictEqual(normalizeStateStatus('riportata per completezza', null), 'riportata per completezza'); + assert.strictEqual(normalizeStateStatus('fasi complete', null), 'fasi complete'); + }); + + test('rows 9-10: control — prose with no trigger words was already verbatim and stays so', () => { + const control = [ + 'completata / completato / completo / completa / incompleta / completamente', + 'Lavoro sospeso in attesa del CEO', + ]; + for (const p of control) { + assert.strictEqual(normalizeStateStatus(p, null), p); + } + }); + + test('row 11: English status words embedded in prose sentences are not rewrites either', () => { + assert.strictEqual(normalizeStateStatus('Waiting on the planning department', null), 'Waiting on the planning department'); + assert.strictEqual(normalizeStateStatus('Notes done, see log', null), 'Notes done, see log'); + assert.strictEqual( + normalizeStateStatus('Discussed the verif steps with QA, paused decision', null), + 'Discussed the verif steps with QA, paused decision', + ); + }); + + test('row 12 (existing #3873 row-26 contract): unrecognized text passes through unchanged', () => { + assert.strictEqual( + normalizeStateStatus('totally-unrecognized-status-text', null), + 'totally-unrecognized-status-text', + ); + }); +}); + +describe('#4186: normalizeStateStatus anchored status vocabulary — documented vocabulary still normalizes', () => { + const { normalizeStateStatus } = require('../gsd-core/bin/lib/state-document.cjs'); + + test('rows 13-15: variable handler-written phase statuses (case/whitespace variants)', () => { + assert.strictEqual(normalizeStateStatus('Executing Phase 5', null), 'executing'); + assert.strictEqual(normalizeStateStatus('EXECUTING PHASE 5', null), 'executing'); + assert.strictEqual(normalizeStateStatus('executing phase 46', null), 'executing'); + assert.strictEqual(normalizeStateStatus('Planning Phase 3', null), 'planning'); + assert.strictEqual(normalizeStateStatus('Verifying Phase 2', null), 'verifying'); + }); + + test('row 16: `Phase N complete` still lands on `completed` (composes with the #3578 demote guard)', () => { + assert.strictEqual(normalizeStateStatus('Phase 12 complete', null), 'completed'); + assert.strictEqual(normalizeStateStatus('Phase 3A complete', null), 'completed'); + }); + + test('rows 17-18: ADR-2207 lifecycle terminal statuses (CONTEXT.md:94 recorded contract)', () => { + assert.strictEqual(normalizeStateStatus('All phases complete', null), 'completed'); + assert.strictEqual(normalizeStateStatus('ALL PHASES COMPLETE', null), 'completed'); + assert.strictEqual(normalizeStateStatus('v1.0 milestone complete', null), 'completed'); + assert.strictEqual(normalizeStateStatus('1.0 milestone complete', null), 'completed'); + }); + + test('rows 19-25: fixed-form handler defaults, including case and whitespace variants', () => { + assert.strictEqual(normalizeStateStatus('Ready to plan', null), 'planning'); + assert.strictEqual(normalizeStateStatus('Ready to execute', null), 'executing'); + assert.strictEqual(normalizeStateStatus('In progress', null), 'executing'); + assert.strictEqual(normalizeStateStatus('In progress', null), 'executing'); + assert.strictEqual(normalizeStateStatus(' Paused ', null), 'paused'); + assert.strictEqual(normalizeStateStatus('Paused', null), 'paused'); + assert.strictEqual(normalizeStateStatus('Stopped', null), 'paused'); + assert.strictEqual(normalizeStateStatus('stopped', null), 'paused'); + assert.strictEqual(normalizeStateStatus('Discussing', null), 'discussing'); + assert.strictEqual(normalizeStateStatus('Verifying', null), 'verifying'); + assert.strictEqual(normalizeStateStatus('Completed', null), 'completed'); + assert.strictEqual(normalizeStateStatus('Done', null), 'completed'); + assert.strictEqual(normalizeStateStatus('done', null), 'completed'); + assert.strictEqual(normalizeStateStatus('Complete', null), 'completed'); + assert.strictEqual(normalizeStateStatus('Complete ✓', null), 'completed'); + assert.strictEqual(normalizeStateStatus('Complete✔', null), 'completed'); + }); + + test('rows 26-27: branch-order artifacts are preserved byte-for-behaviour', () => { + // `verif` outranks `complete` (pinned by tests/state.test.cjs's + // advance-plan case-5 comment); `planning` outranks `complete`. + assert.strictEqual(normalizeStateStatus('Phase complete — ready for verification', null), 'verifying'); + assert.strictEqual(normalizeStateStatus('Planning complete', null), 'planning'); + }); + + test('rows 29-30: unknown fallback and the pausedAt force are unchanged', () => { + assert.strictEqual(normalizeStateStatus(null, null), 'unknown'); + assert.strictEqual(normalizeStateStatus('', null), 'unknown'); + assert.strictEqual(normalizeStateStatus('Executing Phase 5', '2026-09-01'), 'paused'); + }); + + test('row 31: statuses with no branch today stay verbatim', () => { + assert.strictEqual(normalizeStateStatus('Awaiting next milestone', null), 'Awaiting next milestone'); + assert.strictEqual(normalizeStateStatus('Defining requirements', null), 'Defining requirements'); + assert.strictEqual(normalizeStateStatus('Active', null), 'Active'); + }); + + test('row 32: trailing prose after a vocabulary form is executor-authored and passes through', () => { + // Same discipline as #1070's KNOWN_STATUS_PATTERNS: "Complete but needs + // manual QA" is NOT a template default. The anchored vocabulary only + // recognizes the whole-field value. + assert.strictEqual(normalizeStateStatus('Executing Phase 5 — final stretch', null), 'Executing Phase 5 — final stretch'); + assert.strictEqual(normalizeStateStatus('Complete but needs manual QA', null), 'Complete but needs manual QA'); + }); +}); diff --git a/tests/state.test.cjs b/tests/state.test.cjs index 64e7b2c3c..0e7daf716 100644 --- a/tests/state.test.cjs +++ b/tests/state.test.cjs @@ -788,10 +788,16 @@ stopped_at: Plan 2 of Phase 3 }); test('normalizes various status values', () => { + // #4186: recognition is an ANCHORED whole-field vocabulary match. The + // vocabulary's own values normalize to their token; prefix/suffix NARRATIVE + // prose ('Paused at Plan 3') passes through verbatim — a visible paragraph + // beats a guessed token — and the legacy bare `Milestone complete` stays + // recognized (reader-side legacy vocabulary, ADR-2207 removed the writers). const statusTests = [ { input: 'In progress', expected: 'executing' }, { input: 'Ready to execute', expected: 'executing' }, - { input: 'Paused at Plan 3', expected: 'paused' }, + { input: 'Paused', expected: 'paused' }, + { input: 'Paused at Plan 3', expected: 'Paused at Plan 3' }, { input: 'Ready to plan', expected: 'planning' }, { input: 'Phase complete — ready for verification', expected: 'verifying' }, { input: 'Milestone complete', expected: 'completed' }, @@ -1140,7 +1146,9 @@ current_phase: 3 ` ); - runGsdTools('state update Status "Executing Plan 5"', tmpDir); + // #4186: use the handler vocabulary's real form (`Executing Phase N`, + // state-transition.cts:1216) — narrative variants are no longer guessed. + runGsdTools('state update Status "Executing Phase 5"', tmpDir); const result = runGsdTools('state json', tmpDir); assert.ok(result.success, `state json failed: ${result.error}`); @@ -3262,22 +3270,39 @@ describe('cmdStateRecordSession (state record-session)', () => { assert.ok(updated.includes(PINNED_ISO), `Last session should be the pinned ISO timestamp ${PINNED_ISO}`); }); - test('updates Last session timestamp even with no other options', () => { + // #4186: `state record-session` with NO arguments used to execute and write + // `Last session` / `last_updated` to STATE.md (this test previously pinned + // that bare-invocation write). It now follows the house pattern every other + // required-arg verb uses (`state update`: `error('field and value required + // for state update')`) — a bare invocation is a usage error, not a + // heartbeat write. Documented signature (docs/CLI-TOOLS.md): + // `state record-session --stopped-at "..." [--resume-file path]`. + test('#4186: no args errors instead of writing (STATE.md byte-unchanged)', () => { + fs.writeFileSync(path.join(tmpDir, '.planning', 'STATE.md'), sessionFixture); + const before = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + + const result = runGsdTools('state record-session', tmpDir); + + assert.ok(!result.success, 'bare record-session must exit non-zero'); + assert.match( + result.error, + /stopped-at or resume-file required for state record-session/, + 'stderr must name the required flags', + ); + const after = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.strictEqual(after, before, 'STATE.md must be byte-unchanged by a rejected bare call'); + }); + + // #4186: validation precedes the STATE.md existence check (same ordering as + // cmdStateUpdate), and a lone --resume-file still carries an explicit value. + test('#4186: --resume-file alone is a valid explicit value and persists', () => { fs.writeFileSync(path.join(tmpDir, '.planning', 'STATE.md'), sessionFixture); - const PINNED_MS = Date.parse('2020-08-01T08:30:00.000Z'); - const PINNED_ISO = '2020-08-01T08:30:00.000Z'; - const result = runGsdTools('state record-session', tmpDir, { - GSD_TEST_MODE: '1', - GSD_NOW_MS: String(PINNED_MS), - }); + const result = runGsdTools('state record-session --resume-file ".continue-here.md"', tmpDir); assert.ok(result.success, `Command failed: ${result.error}`); - const output = JSON.parse(result.output); - assert.strictEqual(output.recorded, true, 'recorded should be true'); - const updated = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); - assert.ok(updated.includes(PINNED_ISO), `Last session should contain the pinned ISO timestamp ${PINNED_ISO}`); + assert.ok(updated.includes('.continue-here.md'), 'Resume file should be updated'); }); test('sets Resume file to None when not specified', () => { @@ -3295,7 +3320,9 @@ describe('cmdStateRecordSession (state record-session)', () => { }); test('returns error when STATE.md missing', () => { - const result = runGsdTools('state record-session', tmpDir); + // #4186: supply a value so this test keeps exercising the STATE.md-missing + // decline rather than the (now earlier) no-args usage error. + const result = runGsdTools('state record-session --stopped-at "somewhere"', tmpDir); assert.ok(result.success, `Command should exit 0: ${result.error}`); const output = JSON.parse(result.output); @@ -3303,18 +3330,20 @@ describe('cmdStateRecordSession (state record-session)', () => { assert.ok(output.error.includes('STATE.md'), 'error should mention STATE.md'); }); - test('returns recorded false when no session fields found', () => { + // #4186: a STATE.md with no session labels is no longer reachable bare + // (the no-args usage error fires first, before the existence check even) — + // the bare-call contract this test used to pin (recorded:false decline) is + // now the error contract pinned above; this variant proves the validation + // fires regardless of the body's shape. + test('#4186: no args errors even against a STATE.md with no session fields', () => { fs.writeFileSync( path.join(tmpDir, '.planning', 'STATE.md'), '# Project State\n\n**Status:** Active\n**Phase:** 03\n' ); const result = runGsdTools('state record-session', tmpDir); - assert.ok(result.success, `Command should exit 0: ${result.error}`); - - const output = JSON.parse(result.output); - assert.strictEqual(output.recorded, false, 'recorded should be false when no session fields found'); - assert.ok(output.reason !== undefined, 'should have a reason'); + assert.ok(!result.success, 'bare record-session must exit non-zero'); + assert.match(result.error, /stopped-at or resume-file required for state record-session/); }); // #3374 Variant B: stateReplaceField returns the replaced string on any label @@ -3398,6 +3427,175 @@ describe('cmdStateRecordSession (state record-session)', () => { }); }); +// ───────────────────────────────────────────────────────────────────────────── +// #4186 — the write path must not derive the frontmatter `status` token from +// SUBSTRINGS of the free-prose body Status field. Every STATE.md write funnels +// through readModifyWriteStateMd → syncStateFrontmatter → +// buildStateFrontmatter → normalizeStateStatus, so exercising `state update` +// here covers the shared seam for all write paths. +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4186: state update does not rewrite status from prose substrings', () => { + let tmpDir; + + function writeStateMd(statusValue) { + fs.writeFileSync(path.join(tmpDir, '.planning', 'STATE.md'), [ + '---', + "gsd_state_version: '1.0'", + 'status: unknown', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 1 of 2 (Fondamenta)', + 'Status: ' + statusValue, + 'Last activity: 2026-09-02 — aggiornato lo stato', + '', + ].join('\n') + '\n'); + } + + function frontmatterStatus() { + const text = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + const fm = text.split('---')[1] || ''; + const m = fm.match(/^status:(.*)$/m); + assert.ok(m, 'frontmatter must still carry a status key'); + return m[1].trim().replace(/^['"]|['"]$/g, ''); + } + + beforeEach(() => { + tmpDir = createFixture(); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + test('row 33: Italian prose mentioning .planning/ keeps the visible prose token', () => { + writeStateMd('Lavoro sospeso, vedi .planning/STATE.md'); + const r = runGsdTools(['state', 'update', 'Last activity', '2026-09-03 — verifica del defect'], tmpDir); + assert.ok(r.success, `state update failed: ${r.error}`); + assert.strictEqual( + frontmatterStatus(), + 'Lavoro sospeso, vedi .planning/STATE.md', + 'prose containing a .planning/ path must pass through verbatim, not become `planning`', + ); + }); + + test('row 34: a recognized handler status still lands its canonical token', () => { + writeStateMd('Executing Phase 5'); + const r = runGsdTools(['state', 'update', 'Last activity', '2026-09-03 — executing'], tmpDir); + assert.ok(r.success, `state update failed: ${r.error}`); + assert.strictEqual(frontmatterStatus(), 'executing'); + }); +}); + +// ───────────────────────────────────────────────────────────────────────────── +// #4186 — progress recount pin (#1988, PR #2016): a *-SUMMARY.md without a +// plan twin must not inflate progress.completed_plans on any recount trigger +// (the issue measured `34-TRIAGE-SUMMARY.md` → completed_plans 62 → 63 on +// GSD 1.5.0; the pairing fix landed after). These rows pin the correct +// behavior so it cannot regress, composed with the #4129/#4359 ratchet. +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4186: progress recount ignores stray summaries (pin of #1988)', () => { + let tmpDir; + + function seedStrayFixture() { + const phaseDir = path.join(tmpDir, '.planning', 'phases', '34-triage'); + fs.mkdirSync(phaseDir, { recursive: true }); + for (const n of ['01', '02', '03']) fs.writeFileSync(path.join(phaseDir, `34-${n}-PLAN.md`), '# plan\n'); + // Two PAIRED summaries + one stray with no plan twin. + fs.writeFileSync(path.join(phaseDir, '34-01-SUMMARY.md'), '# summary\n'); + fs.writeFileSync(path.join(phaseDir, '34-02-SUMMARY.md'), '# summary\n'); + fs.writeFileSync(path.join(phaseDir, '34-TRIAGE-SUMMARY.md'), '# stray\n'); + fs.writeFileSync(path.join(tmpDir, '.planning', 'STATE.md'), [ + '---', + "gsd_state_version: '1.0'", + 'status: executing', + 'progress:', + ' total_phases: 1', + ' completed_phases: 0', + ' total_plans: 3', + ' completed_plans: 0', + ' percent: 0', + '---', + '', + '# Project State', + '', + '## Current Position', + '', + 'Phase: 34 of 34 (Triage)', + 'Plan: 3 of 3 in current phase', + 'Status: Executing Phase 34', + 'Last activity: 2026-09-02 — executing', + '', + 'Progress: [..........] 0%', + 'Total Plans in Phase: 3', + '', + ].join('\n') + '\n'); + } + + function completedPlans() { + const text = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + // Bounded scan (#2128 class): the progress block sits within a few hundred + // bytes of its opening key in every fixture this suite writes. + const m = text.match(/^progress:[\s\S]{0,400}?completed_plans:\s*(\d+)/m); + assert.ok(m, 'progress.completed_plans must be present after a recount write'); + return Number(m[1]); + } + + beforeEach(() => { + tmpDir = createFixture(); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + test('row 36: update on the Progress field recounts paired summaries only', () => { + seedStrayFixture(); + const r = runGsdTools(['state', 'update', 'Progress', '[..........] 0%'], tmpDir); + assert.ok(r.success, `state update failed: ${r.error}`); + assert.strictEqual(completedPlans(), 2, + 'the stray 34-TRIAGE-SUMMARY.md (no plan twin) must not inflate completed_plans'); + }); + + test('row 37: record-session (the issue trigger) recounts paired summaries only', () => { + seedStrayFixture(); + const r = runGsdTools(['state', 'record-session', '--stopped-at', 'after plan 34-02'], tmpDir); + assert.ok(r.success, `record-session failed: ${r.error}`); + assert.strictEqual(completedPlans(), 2, + 'record-session must derive completed_plans from plan-paired summaries, not raw *-SUMMARY.md counts'); + }); + + test('row 38: scanPhasePlans keeps listing the stray file but does not count it', () => { + seedStrayFixture(); + const scanPhasePlans = require('../gsd-core/bin/lib/plan-scan.cjs'); + const scan = scanPhasePlans(path.join(tmpDir, '.planning', 'phases', '34-triage')); + assert.strictEqual(scan.planCount, 3); + assert.strictEqual(scan.summaryCount, 2, 'paired summaries only'); + assert.ok(scan.summaryFiles.includes('34-TRIAGE-SUMMARY.md'), + 'the stray file stays visible in summaryFiles for callers that list summaries'); + }); + + test('row 39: ratchet still preserves an existing higher completed_plans (#4129/#4359 semantics)', () => { + seedStrayFixture(); + // Poison the frontmatter with a value ABOVE the derived paired count (2). + // `state update Progress` EXPLICITLY names a progress field, which by + // design (#3242 / ADR-3473 §8.6 early-out) makes the fresh derivation + // authoritative — so the ratchet row must use an INCIDENTAL resync write + // (record-session) instead: the curated higher counter survives there. + const statePath = path.join(tmpDir, '.planning', 'STATE.md'); + fs.writeFileSync(statePath, fs.readFileSync(statePath, 'utf-8').replace('completed_plans: 0', 'completed_plans: 9')); + const r = runGsdTools(['state', 'record-session', '--stopped-at', 'after plan 34-02'], tmpDir); + assert.ok(r.success, `record-session failed: ${r.error}`); + assert.strictEqual(completedPlans(), 9, + 'the progress ratchet (monotonic completed counts on incidental resyncs) is unchanged by #4186'); + }); +}); + // ───────────────────────────────────────────────────────────────────────────── // Milestone-scoped phase counting in frontmatter // ───────────────────────────────────────────────────────────────────────────── @@ -10052,17 +10250,19 @@ describe('T6 section-splice characterization — record-session', () => { '', ].join('\n'); - test('record-session no-op: no session fields → recorded:false, STATE.md byte-unchanged', () => { + test('record-session no-op: bare call rejected, STATE.md byte-unchanged (#4186)', () => { + // #4186: the bare call that used to reach the recorded:false decline is + // now a usage error (stopped-at or resume-file required). The byte- + // unchanged assertion below keeps guarding the #952 no-op posture — a + // rejected call must not trample anything, milestone_name included. const d = createTempProject(); try { fs.writeFileSync(path.join(d, '.planning', 'STATE.md'), STATE_NO_SESSION_LABELS); const result = runGsdTools(['state', 'record-session'], d); - assert.ok(result.success, `Command failed: ${result.error}`); - const output = JSON.parse(result.output); - assert.strictEqual(output.recorded, false, 'recorded must be false when no session fields exist'); - // milestone_name must NOT be trampled (#952 no-op guard) + assert.ok(!result.success, 'bare record-session must exit non-zero'); + assert.match(result.error, /stopped-at or resume-file required for state record-session/); const after = fs.readFileSync(path.join(d, '.planning', 'STATE.md'), 'utf-8'); - assert.strictEqual(after, STATE_NO_SESSION_LABELS, 'STATE.md must be byte-unchanged on no-op'); + assert.strictEqual(after, STATE_NO_SESSION_LABELS, 'STATE.md must be byte-unchanged by the rejected call'); } finally { cleanup(d); } @@ -13964,19 +14164,21 @@ describe('#944: record-session persists values even when body lacks session labe '--resume-file value must be present in STATE.md'); }); - test('record-session with no args against a body-less file returns recorded:false (no regression)', () => { - // When NO values are supplied and no session fields can be found/updated, - // recorded:false is the correct behaviour — we only changed the contract - // when the caller supplies values. + test('record-session with no args against a body-less file errors (#4186 usage contract)', () => { + // #4186 superseded the bare-call contract this test used to pin + // (recorded:false decline): a no-values invocation is now a usage error + // for every STATE.md shape, handler-side, before any read. The + // value-supplied decline paths keep their own dedicated tests above + // (#3374 Variant B, #3957). const statePath = path.join(tmpDir, '.planning', 'STATE.md'); fs.writeFileSync(statePath, buildStateMdWithoutSessionSection()); + const before = fs.readFileSync(statePath, 'utf-8'); const result = runGsdTools('state record-session', tmpDir); - assert.ok(result.success, `should exit 0: ${result.error}`); - - const output = JSON.parse(result.output); - assert.strictEqual(output.recorded, false, - 'recorded should still be false when no session fields exist AND no values were supplied'); + assert.ok(!result.success, 'should exit non-zero: usage error'); + assert.match(result.error, /stopped-at or resume-file required for state record-session/); + assert.strictEqual(fs.readFileSync(statePath, 'utf-8'), before, + 'a rejected bare call must not touch the file'); }); test('canonical session section still updates in place (no regression)', () => { @@ -20203,21 +20405,26 @@ describe('#3957 (epic #3473 B9): no-op decline reports the real condition', () = if (tmpDir) cleanup(tmpDir); }); - // Row 8 - test('record-session: nothing to update reports no-fields-found', () => { + // Row 8 — #4186 repurposed: the empty-options call that used to reach + // the no-fields-found decline is now a usage error (handler-side guard, + // same contract as cmdStateUpdate). The decline itself remains covered + // for value-supplied calls by the matched-but-unchanged row below; what + // this row now pins is that the SDK-level contract matches the CLI's and + // that a rejected call never touches the file. + test('record-session: no values supplied is a usage error, STATE.md unchanged (#4186)', () => { tmpDir = createFixture(); const statePath = writeState(tmpDir, '# Project State\n\n## Decisions\n\n- none yet\n'); const before = fs.readFileSync(statePath, 'utf-8'); - const { stdout, stderr } = captureCliIO(() => { - stateLib.cmdStateRecordSession(tmpDir, {}, false); - }); - - const out = JSON.parse(stdout); - assert.strictEqual(out.recorded, false); - assert.strictEqual(out.reason, 'no session fields found in STATE.md to update'); + // error() writes the user-facing message to stderr and throws a bare + // ExitError (v1: message 'process exit ' — the message itself is + // NOT on the exception; asserted via the CLI-level tests above). + assert.throws( + () => stateLib.cmdStateRecordSession(tmpDir, {}, false), + (err) => err && err.name === 'ExitError' && err.code === 1, + 'empty options must be rejected before any read or write', + ); assert.strictEqual(fs.readFileSync(statePath, 'utf-8'), before, 'STATE.md must be unchanged'); - assert.match(stderr, /^\[gsd-tools\] WARNING: state record-session skipped — no session fields found in STATE\.md to update\./); }); // Row 9 (hardest — signature B, collapsed reconciliation). A frozen clock From 1c0acb23591ec71db6b28df04f62d0644260979b Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 17:29:54 -0400 Subject: [PATCH 024/166] feat(#4422): block merging into next/main while the base branch's Tests run is red (#4428) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(#4422): block merging into next/main while the base branch's Tests run is red Adds a next-health job to test.yml that checks the base branch's own last push-triggered Tests run via the GitHub API and fails the existing "Required tests" required check when it's red, with a maintainer-applied "fix-next" label as the explicit escape hatch for the fix-forward PR itself. No branch-protection config change needed — it rides the already-required check. The job is deliberately not gated behind preflight, same reasoning as the changes job: a compute-free API read has nothing to save by waiting. Documents the fix-next label in CONTRIBUTING.md and adds a property test locking the CLEAN/RED/INDETERMINATE classification's iff-relationship. This closes the second half of the 2026-09-06 RCA: three unrelated PRs merged on top of an already-broken next before anyone noticed it was red. Co-Authored-By: Claude Sonnet 5 * fix: close two zero-margin CI timing gaps found while verifying #4422 Discovered while watching this branch's own CI, root-caused via /diagnose rather than dismissed as Windows flakiness: 1. tests/gsd-check-update-worker-atomic-cache.test.cjs's outer timeout (15000ms) exactly matched the inner npm-view timeout the worker wraps (NPM_VIEW_TIMEOUT_MS, gsd-core/bin/check-latest-version.cjs). A slow registry response raced two SIGKILLs at the same instant, killing the worker before it could catch its own timeout and degrade gracefully. Windows's shell-wrapped npm subprocess made the race lose more often there, but the zero margin was platform-agnostic. Fixed by giving the test real headroom (+10s) beyond the named constant it wraps, plus an invariant test so the two values can't silently collide again. 2. scripts/run-tests.cjs's per-chunk weight budget (MAX_FILES_PER_CHUNK) let a Windows full-matrix chunk that was well under budget by the Linux/macOS-calibrated weight table (~32/60 units) still exceed the 600s wall-clock backstop — codex-config.test.cjs's genuinely-measured weight (17.87) doesn't transfer 1:1 to Windows's slower install/ subprocess overhead. Windows now gets its own lower cap (40 vs 60). Co-Authored-By: Claude Sonnet 5 --------- Co-authored-by: sim Co-authored-by: Claude Sonnet 5 --- .github/workflows/test.yml | 77 +++ CONTRIBUTING.md | 15 + gsd-core/bin/check-latest-version.cjs | 11 +- scripts/ci-next-health.cjs | 271 +++++++++ scripts/run-tests.cjs | 16 +- tests/ci-next-health.test.cjs | 552 ++++++++++++++++++ ...-check-update-worker-atomic-cache.test.cjs | 32 +- 7 files changed, 969 insertions(+), 5 deletions(-) create mode 100644 scripts/ci-next-health.cjs create mode 100644 tests/ci-next-health.test.cjs diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 3ea4c174d..b17038f8b 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -45,6 +45,72 @@ jobs: contents: read pull-requests: read + # #4422: three unrelated PRs merged on top of an already-broken `next` on + # 2026-09-06 before anyone noticed — nothing checked the base branch's OWN + # health before letting a PR land on it. This job queries the GitHub API for + # the base branch's last push-triggered Tests run and blocks the merge if it + # is red, with a `fix-next` label escape hatch for the PR that is itself the + # fix-forward. See scripts/ci-next-health.cjs's header for the full risk- + # asymmetry reasoning (it deliberately does NOT fail open on a definite red + # signal, unlike the mergeability preflight above). + # + # No `if:` guard, for the same reason `preflight` has none: a SKIPPED + # dependency skips its dependents exactly like a failed one, so guarding this + # job would skip `required-tests` on every push/workflow_dispatch run too. + # scripts/ci-next-health.cjs itself no-ops (zero API calls, exit 0) on any + # event other than pull_request/merge_group. + # + # `next-health` is likewise deliberately NOT gated behind `preflight`, same + # rationale as `changes` above: it is a ~2-minute, compute-free API read that + # runs in parallel with the preflight job, and serializing it behind + # `preflight` would add latency to every healthy PR for no saving. + next-health: + name: Base branch health + runs-on: ubuntu-latest + timeout-minutes: 2 + permissions: + contents: read + actions: read + steps: + - name: Check out the next-health script from the base branch + uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 + with: + # Same rationale as pr-mergeable-preflight.yml's checkout: pinned to + # the BASE sha so the script that decides the gate comes from a + # trusted commit, never PR-supplied code. Empty on a non-pull_request + # event (e.g. merge_group), where actions/checkout falls back to the + # triggering ref — fine, since the script no-ops before reading + # anything on those events too. + ref: ${{ github.event.pull_request.base.sha }} + fetch-depth: 1 + sparse-checkout: | + scripts + persist-credentials: false + + - name: Check base branch health + id: check + env: + # Every value arrives through env. CONTRIBUTING.md forbids `${{ }}` + # inside a `run:` block. + GITHUB_TOKEN: ${{ github.token }} + PR_LABELS: ${{ join(github.event.pull_request.labels.*.name, ',') }} + MERGE_GROUP_BASE_REF: ${{ github.event.merge_group.base_ref }} + run: | + # Bootstrap arm, mirroring pr-mergeable-preflight.yml's: the checkout + # above is of the BASE sha, so this step runs the script as it exists + # on the base branch — which means it is absent on the PR that + # INTRODUCES it, and on any PR branched from a base predating it. + # Absent is not "red": it is one more thing we cannot determine, so + # it takes the same fail-open path as any other unresolved read. + # Self-healing — once the script is on the base branch this arm never + # fires again. + if [ ! -f scripts/ci-next-health.cjs ]; then + echo "::warning::scripts/ci-next-health.cjs is not present at the base sha; skipping the base-branch health gate (fail-open). Expected on the PR that introduces it, or on a branch whose base predates it." + echo "verdict=INDETERMINATE" >> "$GITHUB_OUTPUT" + exit 0 + fi + node scripts/ci-next-health.cjs + changes: name: Detect test scope runs-on: ubuntu-latest @@ -857,6 +923,7 @@ jobs: name: Required tests needs: - preflight + - next-health - changes - lint-tests - test @@ -871,6 +938,7 @@ jobs: - name: Summarize required test gate env: PREFLIGHT_RESULT: ${{ needs.preflight.result }} + NEXT_HEALTH_RESULT: ${{ needs.next-health.result }} CODE_CHANGED: ${{ needs.changes.outputs.code_changed }} PRODUCT_CHANGED: ${{ needs.changes.outputs.product_changed }} CHANGES_RESULT: ${{ needs.changes.result }} @@ -883,6 +951,7 @@ jobs: run: | set -euo pipefail echo "preflight=$PREFLIGHT_RESULT" + echo "next-health=$NEXT_HEALTH_RESULT" echo "code_changed=$CODE_CHANGED" echo "product_changed=$PRODUCT_CHANGED" echo "changes=$CHANGES_RESULT" @@ -910,6 +979,14 @@ jobs: exit 1 fi + # #4422: applies UNCONDITIONALLY, not nested inside the + # PRODUCT_CHANGED/CODE_CHANGED branches below — a red base branch + # must block every PR, including doc-only ones. + if [ "$NEXT_HEALTH_RESULT" != "success" ]; then + echo "::error::the base branch's own last Tests run is red — see the 'Base branch health' job above for the failing run. Wait for a fix-forward merge, or if this PR IS the fix, ask a maintainer to apply the 'fix-next' label to override." + exit 1 + fi + if [ "$LINT_RESULT" != "success" ]; then echo "::error::lint-tests did not pass" exit 1 diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c916ab889..d45e43c00 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1285,6 +1285,21 @@ the pipeline runs exactly as it did before, and the per-job including which lanes are deliberately *not* gated, is in [docs/TESTING-SUITES.md → The mergeability preflight](docs/TESTING-SUITES.md#the-mergeability-preflight). +### A PR cannot merge onto a red base branch + +The `Base branch health` required check queries GitHub for the base branch's +own last push-triggered Tests run and blocks your merge if that run is red — +independent of whether your own PR's changes pass. This needs no +branch-protection reconfiguration: it rides the existing "Required tests" +check, the same status GitHub already requires before merge. + +If your PR is itself the fix-forward and you need to land it while the base +branch is still red, a maintainer applies the `fix-next` label directly to +your PR to explicitly bypass this one check. Applying a label requires +GitHub write access to the repo, so a PR author cannot self-apply it to +bypass the gate — only a maintainer or another collaborator with label-write +permission can. Full decision logic is in `scripts/ci-next-health.cjs`. + ### CI Test Quality Checks The following checks run on every PR in addition to the test suite: diff --git a/gsd-core/bin/check-latest-version.cjs b/gsd-core/bin/check-latest-version.cjs index e00ff58c0..0fcb21085 100755 --- a/gsd-core/bin/check-latest-version.cjs +++ b/gsd-core/bin/check-latest-version.cjs @@ -42,6 +42,12 @@ const SEMVER_RE = /^\d+\.\d+\.\d+(?:[-+][0-9A-Za-z.-]+)?$/; // `npm view` to an empty or foreign dist-tag. const ALLOWED_TAGS = Object.freeze(['latest', 'next']); +// Bounded at 15s so a hung registry doesn't block /gsd-update (#2993 CR). +// Exported so callers (e.g. the worker's own test harness, #4091) can derive +// their own outer timeout with real margin above this inner bound instead of +// re-hardcoding 15000 and silently drifting into a zero-margin race. +const NPM_VIEW_TIMEOUT_MS = 15_000; + /** * Build the `npm view` args for a dist-tag. `latest` keeps the bare package * spec so the default invocation is byte-for-byte identical to before tag @@ -92,9 +98,8 @@ function checkLatestVersion(opts = {}) { // Windows shell-flag policy and timeout default). The injection point // remains spawnSync-shaped for test compatibility — the adapter below // translates { exitCode } → { status } so the consumer logic is unchanged. - // Bounded at 15s so a hung registry doesn't block /gsd-update (#2993 CR). const defaultSpawn = () => { - const r = execNpm(buildViewArgs(tag), { timeout: 15_000 }); + const r = execNpm(buildViewArgs(tag), { timeout: NPM_VIEW_TIMEOUT_MS }); return { status: r.exitCode, stdout: r.stdout, @@ -158,4 +163,4 @@ function main() { if (require.main === module) runMain(main); -module.exports = { checkLatestVersion, CHECK_REASON, PACKAGE_NAME, ALLOWED_TAGS, buildViewArgs, resolveTag }; +module.exports = { checkLatestVersion, CHECK_REASON, PACKAGE_NAME, ALLOWED_TAGS, NPM_VIEW_TIMEOUT_MS, buildViewArgs, resolveTag }; diff --git a/scripts/ci-next-health.cjs b/scripts/ci-next-health.cjs new file mode 100644 index 000000000..03e79dc69 --- /dev/null +++ b/scripts/ci-next-health.cjs @@ -0,0 +1,271 @@ +#!/usr/bin/env node +'use strict'; + +// ci-next-health.cjs — #4422 base-branch health gate. +// +// Risk asymmetry drives this design, same shape as scripts/ci-pr-mergeability.cjs +// but pointed the other direction: on 2026-09-06 three unrelated PRs merged on +// top of an already-broken `next` before anyone noticed, because nothing checked +// the base branch's OWN health before letting a PR land on it. A false positive +// here (a healthy `next` wrongly reported RED) blocks every PR merge in the repo +// until a human notices and applies the `fix-next` bypass label — annoying, but +// loud and immediately actionable. A false negative (a broken `next` reported +// healthy) reproduces the #4422 incident exactly. So, unlike the mergeability +// preflight, this gate does NOT fail open on a definite red signal — it fails +// open only when the signal itself is unavailable or inapplicable (wrong event, +// no resolvable base ref, an API read that throws). Once GitHub actually answers +// with a non-success conclusion for the base branch's last push-triggered Tests +// run, that is treated as authoritative and blocks — with one explicit, visible, +// human-operated escape hatch (the `fix-next` label) for the PR that is itself +// the fix-forward. +// +// Every non-success conclusion (`failure`, `cancelled`, `timed_out`, +// `action_required`, ...) is treated as RED, not just `failure` — a cancelled or +// timed-out run on the base branch is not evidence the branch is healthy, it is +// evidence nobody knows yet. + +const { ExitError, runMain } = require('./lib/cli-exit.cjs'); + +const VERDICT = Object.freeze({ + CLEAN: 'CLEAN', + RED: 'RED', + BYPASSED: 'BYPASSED', + SKIPPED_NOT_APPLICABLE: 'SKIPPED_NOT_APPLICABLE', + INDETERMINATE: 'INDETERMINATE', +}); + +const APPLICABLE_EVENTS = new Set(['pull_request', 'merge_group']); +const BYPASS_LABEL = 'fix-next'; +const TESTS_WORKFLOW_FILE = 'test.yml'; + +/** + * Pure, total classifier: the GitHub "list workflow runs" response payload -> + * CLEAN | RED | INDETERMINATE. Never throws. + * + * - Not an object, or `workflow_runs` isn't an array -> INDETERMINATE: the + * shape we depend on is not present, so nothing can be concluded. + * - No completed push runs found yet (e.g. a brand-new `release/**`/`hotfix/**` + * branch with no Tests history) -> CLEAN: absence of evidence of red is not + * evidence of red, and a brand-new branch must not be permanently unmergeable. + * - The most recent run's `conclusion === 'success'` -> CLEAN. + * - Anything else (failure, cancelled, timed_out, action_required, ...) -> RED. + */ +function classifyRunConclusion(payload) { + if (payload === null || typeof payload !== 'object') return VERDICT.INDETERMINATE; + if (!Array.isArray(payload.workflow_runs)) return VERDICT.INDETERMINATE; + if (payload.workflow_runs.length === 0) return VERDICT.CLEAN; + const [latest] = payload.workflow_runs; + if (latest && latest.conclusion === 'success') return VERDICT.CLEAN; + return VERDICT.RED; +} + +/** + * Dependency-injected orchestrator. `fetchLatestRun` is an async function + * taking the resolved base branch name and returning the parsed API payload + * (or throwing) — injected so tests never touch the network. + * + * @returns {Promise<{verdict:string, payload:*, reason:string}>} + */ +async function resolveNextHealth({ fetchLatestRun, eventName, baseRef, labels } = {}) { + if (!APPLICABLE_EVENTS.has(eventName)) { + return { verdict: VERDICT.SKIPPED_NOT_APPLICABLE, payload: null, reason: 'not-applicable-event' }; + } + + if (typeof baseRef !== 'string' || baseRef.trim() === '') { + return { verdict: VERDICT.INDETERMINATE, payload: null, reason: 'no-base-ref' }; + } + + let payload = null; + try { + payload = await fetchLatestRun(baseRef); + } catch (err) { + return { verdict: VERDICT.INDETERMINATE, payload: null, reason: `fetch-failed: ${err.message}` }; + } + + const classified = classifyRunConclusion(payload); + if (classified !== VERDICT.RED) { + return { verdict: classified, payload, reason: 'resolved' }; + } + + // Escape hatch. `labels` is empty/absent for merge_group — that event has no + // label surface today, which means a queued merge-group commit cannot use + // this bypass. Known gap, not solved here: a maintainer must land the + // fix-forward as a direct pull_request merge (where the label IS readable) + // rather than through the merge queue, until GitHub exposes an equivalent + // signal for merge_group. + const labelList = Array.isArray(labels) ? labels : []; + if (labelList.includes(BYPASS_LABEL)) { + return { verdict: VERDICT.BYPASSED, payload, reason: 'bypass-label' }; + } + + return { verdict: VERDICT.RED, payload, reason: 'red' }; +} + +function usage() { + return [ + 'Usage:', + ' node scripts/ci-next-health.cjs', + '', + 'CI-only gate: checks whether the PR/merge-group\'s base branch\'s own last', + 'push-triggered Tests run is red, and fails the job (exit 1) when it is —', + 'unless the PR carries the "fix-next" bypass label. Every other case', + '(wrong event, no resolvable base ref, an unreadable API read) fails open', + '(exit 0).', + '', + 'Environment variables read:', + ' GITHUB_EVENT_NAME workflow trigger event; only "pull_request" and', + ' "merge_group" are checked', + ' GITHUB_REPOSITORY owner/repo', + ' GITHUB_TOKEN optional bearer token for the API read', + ' GITHUB_BASE_REF PR base branch name (pull_request events; set', + ' automatically by GitHub Actions)', + ' MERGE_GROUP_BASE_REF merge-group base ref (merge_group events; the', + ' caller workflow must populate this from', + ' github.event.merge_group.base_ref — a leading', + ' "refs/heads/" prefix is stripped if present)', + ' GITHUB_API_URL GitHub API base URL (default: https://api.github.com)', + ' PR_LABELS comma-separated PR label names (pull_request events;', + ' the caller workflow must populate this from', + ' github.event.pull_request.labels.*.name); checked', + ' for the literal "fix-next" bypass label', + ' GITHUB_OUTPUT path to append verdict= step output to', + ' GITHUB_STEP_SUMMARY path to append a human-readable summary to', + ].join('\n'); +} + +function parseArgs(argv) { + for (const arg of argv) { + if (arg === '--help' || arg === '-h') { + process.stdout.write(`${usage()}\n`); + throw new ExitError(0); + } + throw new Error(`unknown argument: ${arg}`); + } +} + +function writeOutput(lines) { + const outputPath = process.env.GITHUB_OUTPUT; + if (typeof outputPath !== 'string' || outputPath === '') return; + try { + const fs = require('node:fs'); + fs.appendFileSync(outputPath, `${lines.join('\n')}\n`); + } catch (err) { + // A failure while REPORTING the verdict must never invert the gate. + process.stderr.write(`::warning::failed to write GITHUB_OUTPUT: ${err.message}\n`); + } +} + +function writeSummary(text) { + const summaryPath = process.env.GITHUB_STEP_SUMMARY; + if (typeof summaryPath !== 'string' || summaryPath === '') return; + try { + const fs = require('node:fs'); + fs.appendFileSync(summaryPath, `${text}\n`); + } catch (err) { + process.stderr.write(`::warning::failed to write GITHUB_STEP_SUMMARY: ${err.message}\n`); + } +} + +function resolveBaseRef(eventName) { + if (eventName === 'pull_request') { + return process.env.GITHUB_BASE_REF || ''; + } + if (eventName === 'merge_group') { + const raw = process.env.MERGE_GROUP_BASE_REF || ''; + return raw.startsWith('refs/heads/') ? raw.slice('refs/heads/'.length) : raw; + } + return ''; +} + +function parseLabels(raw) { + if (typeof raw !== 'string' || raw.trim() === '') return []; + return raw.split(',').map((s) => s.trim()).filter((s) => s !== ''); +} + +function runUrlOf(payload) { + const run = payload && Array.isArray(payload.workflow_runs) ? payload.workflow_runs[0] : undefined; + return run && typeof run.html_url === 'string' ? run.html_url : '(unknown run URL)'; +} + +// `argv` defaults to real CLI argv but is a parameter so tests can call +// main() in-process (e.g. main([])) without inheriting the test runner's own +// argv, which would otherwise trip parseArgs's "unknown argument" branch. +async function main(argv = process.argv.slice(2)) { + parseArgs(argv); + + const eventName = process.env.GITHUB_EVENT_NAME; + const repo = process.env.GITHUB_REPOSITORY || ''; + const token = process.env.GITHUB_TOKEN || ''; + const apiBase = process.env.GITHUB_API_URL || 'https://api.github.com'; + const baseRef = resolveBaseRef(eventName); + const labels = parseLabels(process.env.PR_LABELS); + + const fetchLatestRun = async (branch) => { + const url = `${apiBase}/repos/${repo}/actions/workflows/${TESTS_WORKFLOW_FILE}/runs` + + `?branch=${encodeURIComponent(branch)}&event=push&status=completed&per_page=1`; + const headers = { + accept: 'application/vnd.github+json', + 'x-github-api-version': '2022-11-28', + 'user-agent': 'gsd-core-ci-next-health', + }; + if (token) headers.authorization = `Bearer ${token}`; + const response = await fetch(url, { headers, signal: AbortSignal.timeout(10000) }); + if (!response.ok) { + throw new Error(`GitHub API responded ${response.status}`); + } + const text = await response.text(); + try { + return JSON.parse(text); + } catch { + return null; + } + }; + + const result = await resolveNextHealth({ fetchLatestRun, eventName, baseRef, labels }); + + writeOutput([`verdict=${result.verdict}`]); + writeSummary(`Base branch health: ${result.verdict}`); + + if (result.verdict === VERDICT.RED) { + const runUrl = runUrlOf(result.payload); + process.stderr.write( + `::error::the base branch "${baseRef}"'s own last Tests run is red: ${runUrl} — ` + + 'wait for a fix-forward merge, or if THIS pull request is the fix, ask a ' + + `maintainer to apply the "${BYPASS_LABEL}" label to explicitly override this gate.\n`, + ); + return 1; + } + + if (result.verdict === VERDICT.BYPASSED) { + const runUrl = runUrlOf(result.payload); + process.stderr.write( + `::warning::the base branch "${baseRef}"'s own last Tests run is red: ${runUrl} — ` + + `a human applied the "${BYPASS_LABEL}" label to explicitly override this gate.\n`, + ); + return 0; + } + + if (result.verdict === VERDICT.INDETERMINATE) { + process.stderr.write( + `::warning::could not determine "${baseRef || '(no base ref)'}"'s Tests health (${result.reason}); ` + + 'proceeding (fail-open).\n', + ); + return 0; + } + + process.stdout.write(`Base branch health: ${result.verdict}\n`); + return 0; +} + +if (require.main === module) { + runMain(main); +} + +module.exports = { + VERDICT, + BYPASS_LABEL, + TESTS_WORKFLOW_FILE, + classifyRunConclusion, + resolveNextHealth, + main, +}; diff --git a/scripts/run-tests.cjs b/scripts/run-tests.cjs index bfdadf75c..092207a38 100644 --- a/scripts/run-tests.cjs +++ b/scripts/run-tests.cjs @@ -1273,7 +1273,21 @@ function main() { // node process (also relieving per-process memory pressure from 170+ files at once). // Lowered from 90 to 60 after #1575 — macOS Node 22 shard 2/3 chunk 2 (~80 files // including state.test.cjs, perf-*, worktree-cleanup) exceeded 600s with 90. - const MAX_FILES_PER_CHUNK = positiveNumberEnv(process.env.RUN_TESTS_MAX_FILES_PER_CHUNK, 60); + // + // 2026-09-06 (PR #4428 CI): a Windows full-matrix chunk (chunk 3/6, ~32/60 + // weight-budget units, dominated by codex-config.test.cjs at a genuinely + // MEASURED weight of 17.87 — not a stale-table miss) still exceeded the + // 600s per-chunk backstop. The weight table's calibration does not + // transfer 1:1 to the Windows runner for install/subprocess-heavy work — + // it needs a smaller budget than Linux/macOS to stay inside the same + // wall-clock ceiling. Windows gets its own, lower cap (~33% reduction, + // proportionate to the >30% single-file share codex-config.test.cjs alone + // consumed of that chunk's budget); other platforms are unaffected. + const DEFAULT_MAX_FILES_PER_CHUNK = process.platform === 'win32' ? 40 : 60; + const MAX_FILES_PER_CHUNK = positiveNumberEnv( + process.env.RUN_TESTS_MAX_FILES_PER_CHUNK, + DEFAULT_MAX_FILES_PER_CHUNK, + ); // #2088 established that file COUNT is a poor proxy for a chunk's wall-clock: // install-heavy files (real installs) cost ~10x a unit file, and when several // land in the SAME chunk it blows the 600s backstop while unit-only chunks diff --git a/tests/ci-next-health.test.cjs b/tests/ci-next-health.test.cjs new file mode 100644 index 000000000..a71ca7619 --- /dev/null +++ b/tests/ci-next-health.test.cjs @@ -0,0 +1,552 @@ +'use strict'; + +// #4422 — base-branch health gate. +// +// Risk asymmetry drives this suite (see scripts/ci-next-health.cjs's header): +// false positive (a healthy `next` called RED) => blocks every PR merge +// until a human notices and applies the `fix-next` bypass label. +// false negative (a broken `next` called healthy) => reproduces the #4422 +// incident exactly (three unrelated PRs merged on top of an already-broken +// `next`). +// Unlike the mergeability preflight, this gate does NOT fail open on a +// definite RED signal — only on an unavailable/inapplicable one. + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const http = require('node:http'); +const fc = require('fast-check'); + +const { createTempDir, cleanup } = require('./helpers.cjs'); +const { runNode } = require('./helpers/process-seam.cjs'); +const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); + +const ROOT = path.join(__dirname, '..'); +const SCRIPT = path.join(ROOT, 'scripts', 'ci-next-health.cjs'); + +const { + VERDICT, + BYPASS_LABEL, + classifyRunConclusion, + resolveNextHealth, + main, +} = require('../scripts/ci-next-health.cjs'); + +// --------------------------------------------------------------------------- +// Fakes. The seams are parameters, not module patches — no global state, so +// every case is order-independent. +// --------------------------------------------------------------------------- + +/** A fetchLatestRun fake that replays a scripted response; an Error is thrown. */ +function scriptedFetch(response) { + const calls = []; + const fn = async (branch) => { + calls.push(branch); + if (response instanceof Error) throw response; + return response; + }; + fn.calls = calls; + return fn; +} + +function runPayload(conclusion, htmlUrl = 'https://github.com/open-gsd/gsd-core/actions/runs/1') { + return { workflow_runs: [{ conclusion, html_url: htmlUrl }] }; +} + +// --------------------------------------------------------------------------- +// A. classifyRunConclusion — pure +// --------------------------------------------------------------------------- + +describe('ci-next-health: classifyRunConclusion', () => { + test('classifies an undefined payload as INDETERMINATE', () => { + assert.equal(classifyRunConclusion(undefined), VERDICT.INDETERMINATE); + }); + + test('classifies a null payload as INDETERMINATE', () => { + assert.equal(classifyRunConclusion(null), VERDICT.INDETERMINATE); + }); + + test('classifies a non-object payload as INDETERMINATE', () => { + for (const value of [0, 'str', true, 42]) { + assert.equal(classifyRunConclusion(value), VERDICT.INDETERMINATE); + } + }); + + test('classifies a payload whose workflow_runs is not an array as INDETERMINATE', () => { + for (const value of [undefined, null, 'x', {}, 1]) { + assert.equal(classifyRunConclusion({ workflow_runs: value }), VERDICT.INDETERMINATE); + } + }); + + test('classifies an empty workflow_runs array as CLEAN', () => { + // A brand-new release/**/hotfix/** branch with no Tests history yet must + // not be permanently unmergeable. + assert.equal(classifyRunConclusion({ workflow_runs: [] }), VERDICT.CLEAN); + }); + + test('classifies conclusion: success as CLEAN', () => { + assert.equal(classifyRunConclusion(runPayload('success')), VERDICT.CLEAN); + }); + + test('classifies conclusion: failure as RED', () => { + assert.equal(classifyRunConclusion(runPayload('failure')), VERDICT.RED); + }); + + test('classifies conclusion: cancelled as RED', () => { + // A cancelled/timed-out run on the base branch is not evidence of health. + assert.equal(classifyRunConclusion(runPayload('cancelled')), VERDICT.RED); + }); + + test('classifies every non-success conclusion as RED', () => { + for (const conclusion of ['timed_out', 'action_required', 'neutral', 'skipped', 'stale']) { + assert.equal( + classifyRunConclusion(runPayload(conclusion)), + VERDICT.RED, + `conclusion=${conclusion} must be RED`, + ); + } + }); + + test('only reads the FIRST run in workflow_runs', () => { + const payload = { workflow_runs: [{ conclusion: 'success' }, { conclusion: 'failure' }] }; + assert.equal(classifyRunConclusion(payload), VERDICT.CLEAN); + }); + + test('VERDICT is frozen and its atom set is locked', () => { + assert.ok(Object.isFrozen(VERDICT)); + assert.deepEqual( + Object.keys(VERDICT).sort(), + ['BYPASSED', 'CLEAN', 'INDETERMINATE', 'RED', 'SKIPPED_NOT_APPLICABLE'], + ); + }); + + // A well-formed payload has an ARRAY workflow_runs of length 0 or 1, whose + // sole element (if present) carries an arbitrary `conclusion` string. A + // malformed payload is anything else: a non-object payload, or an object + // whose `workflow_runs` is not an array. + const wellFormedPayloadArb = fc.record({ + workflow_runs: fc.oneof( + fc.constant([]), + fc.tuple(fc.record({ conclusion: fc.string() })), + ), + }); + + const malformedNonArrayRunsArb = fc.oneof( + fc.constant(undefined), + fc.constant(null), + fc.string(), + fc.integer(), + fc.boolean(), + fc.dictionary(fc.string(), fc.string()), + ); + + const malformedPayloadArb = fc.oneof( + fc.constant(null), + fc.constant(undefined), + fc.string(), + fc.integer(), + fc.boolean(), + fc.record({ workflow_runs: malformedNonArrayRunsArb }), + ); + + const payloadArb = fc.oneof(wellFormedPayloadArb, malformedPayloadArb); + + test('CLEAN iff (workflow_runs is empty) or (first run succeeded); everything else RED or INDETERMINATE (property)', () => { + fc.assert( + fc.property(payloadArb, (payload) => { + const verdict = classifyRunConclusion(payload); + + const isObject = payload !== null && typeof payload === 'object'; + const hasArrayRuns = isObject && Array.isArray(payload.workflow_runs); + + if (!hasArrayRuns) { + return verdict === VERDICT.INDETERMINATE; + } + + const isClean = payload.workflow_runs.length === 0 + || payload.workflow_runs[0].conclusion === 'success'; + + if (isClean) return verdict === VERDICT.CLEAN; + return verdict === VERDICT.RED; + }), + { seed: 4422, numRuns: 500, verbose: true }, + ); + }); +}); + +// --------------------------------------------------------------------------- +// B. resolveNextHealth — dependency-injected orchestrator +// --------------------------------------------------------------------------- + +describe('ci-next-health: resolveNextHealth', () => { + test('a push event is SKIPPED_NOT_APPLICABLE without calling fetchLatestRun', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ fetchLatestRun, eventName: 'push', baseRef: 'next', labels: [] }); + assert.equal(result.verdict, VERDICT.SKIPPED_NOT_APPLICABLE); + assert.equal(fetchLatestRun.calls.length, 0); + }); + + test('workflow_dispatch is SKIPPED_NOT_APPLICABLE without calling fetchLatestRun', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'workflow_dispatch', baseRef: 'next', labels: [], + }); + assert.equal(result.verdict, VERDICT.SKIPPED_NOT_APPLICABLE); + assert.equal(fetchLatestRun.calls.length, 0); + }); + + test('an unknown/absent event name is SKIPPED_NOT_APPLICABLE without calling', async () => { + for (const eventName of ['', undefined, 'schedule', 'release', 'issues']) { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ fetchLatestRun, eventName, baseRef: 'next', labels: [] }); + assert.equal(result.verdict, VERDICT.SKIPPED_NOT_APPLICABLE); + assert.equal(fetchLatestRun.calls.length, 0); + } + }); + + test('a pull_request event with no resolvable base ref is INDETERMINATE', async () => { + for (const baseRef of [undefined, null, '', ' ']) { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ fetchLatestRun, eventName: 'pull_request', baseRef, labels: [] }); + assert.equal(result.verdict, VERDICT.INDETERMINATE); + assert.equal(fetchLatestRun.calls.length, 0); + } + }); + + test('a merge_group event with no resolvable base ref is INDETERMINATE', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ fetchLatestRun, eventName: 'merge_group', baseRef: '', labels: [] }); + assert.equal(result.verdict, VERDICT.INDETERMINATE); + assert.equal(fetchLatestRun.calls.length, 0); + }); + + test('a throwing fetchLatestRun degrades to INDETERMINATE', async () => { + const fetchLatestRun = scriptedFetch(new Error('ECONNRESET')); + const result = await resolveNextHealth({ fetchLatestRun, eventName: 'pull_request', baseRef: 'next', labels: [] }); + assert.equal(result.verdict, VERDICT.INDETERMINATE); + }); + + test('a RED verdict with the fix-next label is BYPASSED', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'pull_request', baseRef: 'next', labels: ['needs-triage', BYPASS_LABEL], + }); + assert.equal(result.verdict, VERDICT.BYPASSED); + }); + + test('a RED verdict with no matching label stays RED', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'pull_request', baseRef: 'next', labels: ['needs-triage'], + }); + assert.equal(result.verdict, VERDICT.RED); + }); + + test('a RED verdict with an absent labels array stays RED (merge_group has no label surface)', async () => { + const fetchLatestRun = scriptedFetch(runPayload('failure')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'merge_group', baseRef: 'next', labels: undefined, + }); + assert.equal(result.verdict, VERDICT.RED); + }); + + test('a CLEAN verdict (success conclusion) is not affected by labels', async () => { + const fetchLatestRun = scriptedFetch(runPayload('success')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'pull_request', baseRef: 'next', labels: [BYPASS_LABEL], + }); + assert.equal(result.verdict, VERDICT.CLEAN); + }); + + test('a CLEAN verdict (empty workflow_runs array)', async () => { + const fetchLatestRun = scriptedFetch({ workflow_runs: [] }); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'pull_request', baseRef: 'release/1.0', labels: [], + }); + assert.equal(result.verdict, VERDICT.CLEAN); + }); + + test('merge_group resolves the same way as pull_request given a base ref', async () => { + const fetchLatestRun = scriptedFetch(runPayload('success')); + const result = await resolveNextHealth({ + fetchLatestRun, eventName: 'merge_group', baseRef: 'next', labels: [], + }); + assert.equal(result.verdict, VERDICT.CLEAN); + assert.deepEqual(fetchLatestRun.calls, ['next']); + }); +}); + +// --------------------------------------------------------------------------- +// C. main() — integration through the process seam against a real local API. +// +// The CLI's real fetch path is exercised by pointing GITHUB_API_URL at a +// throwaway localhost server, so no production test-mode branch exists and +// nothing reaches api.github.com. +// --------------------------------------------------------------------------- + +const SENTINEL_TOKEN = 'ghs_sentinel_must_never_be_echoed_4422'; + +/** Start a one-shot API stub. `handler(requestCount)` returns { status, body }. */ +async function startApiStub(handler) { + let requestCount = 0; + const server = http.createServer((req, res) => { + const { status, body } = handler(requestCount++); + res.writeHead(status, { 'content-type': 'application/json' }); + res.end(typeof body === 'string' ? body : JSON.stringify(body)); + }); + await new Promise((resolve) => server.listen(0, '127.0.0.1', resolve)); + const { port } = server.address(); + return { + url: `http://127.0.0.1:${port}`, + close: () => new Promise((resolve) => server.close(resolve)), + get requestCount() { return requestCount; }, + }; +} + +function runCli(env, { outputPath } = {}) { + return runNode([SCRIPT], { + cwd: ROOT, + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + GITHUB_REPOSITORY: 'open-gsd/gsd-core', + GITHUB_TOKEN: SENTINEL_TOKEN, + GITHUB_BASE_REF: 'next', + ...(outputPath ? { GITHUB_OUTPUT: outputPath } : {}), + ...env, + }, + }); +} + +// `runNode` (tests/helpers/process-seam.cjs) is spawnSync: it blocks this +// process's event loop for the whole child lifetime. `startApiStub` is an +// http.createServer living in THIS process, so while spawnSync blocks, the +// stub can never accept the child's connection. Any test that needs the stub +// must instead call main() in-process, which keeps this process's event loop +// live so the stub can actually answer. Mirrors tests/ci-pr-mergeability.test.cjs. +async function callMain(t, env, { outputPath } = {}) { + const overrides = { + GITHUB_REPOSITORY: 'open-gsd/gsd-core', + GITHUB_TOKEN: SENTINEL_TOKEN, + GITHUB_BASE_REF: 'next', + ...(outputPath ? { GITHUB_OUTPUT: outputPath } : {}), + ...env, + }; + const saved = new Map(); + for (const key of Object.keys(overrides)) saved.set(key, process.env[key]); + Object.assign(process.env, overrides); + t.after(() => { + for (const [key, value] of saved) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + }); + + const stdout = []; + const stderr = []; + t.mock.method(process.stdout, 'write', (chunk) => { stdout.push(String(chunk)); return true; }); + t.mock.method(process.stderr, 'write', (chunk) => { stderr.push(String(chunk)); return true; }); + + const code = await main([]); + return { code, stdout: stdout.join(''), stderr: stderr.join('') }; +} + +function readOutputs(outputPath) { + const raw = fs.readFileSync(outputPath, 'utf8'); + const outputs = {}; + for (const line of raw.split(/\r?\n/)) { + const index = line.indexOf('='); + if (index > 0) outputs[line.slice(0, index)] = line.slice(index + 1); + } + return outputs; +} + +describe('ci-next-health: CLI', () => { + test('exits 0 and writes a skip verdict on a push event', (t) => { + const dir = createTempDir('next-health-skip-'); + t.after(() => cleanup(dir)); + const outputPath = path.join(dir, 'gh-output'); + + const result = runCli({ GITHUB_EVENT_NAME: 'push' }, { outputPath }); + + assert.equal(result.exitCode, 0, result.stderr); + assert.equal(readOutputs(outputPath).verdict, VERDICT.SKIPPED_NOT_APPLICABLE); + }); + + test('exits 1 and annotates the failing run on a RED base branch', async (t) => { + const dir = createTempDir('next-health-red-'); + const api = await startApiStub(() => ({ + status: 200, + body: { workflow_runs: [{ conclusion: 'failure', html_url: 'https://example.test/runs/99' }] }, + })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url, PR_LABELS: '' }, + { outputPath }, + ); + + assert.equal(result.code, 1, `stdout: ${result.stdout}\nstderr: ${result.stderr}`); + const combined = `${result.stdout}${result.stderr}`; + assert.ok(combined.includes('::error::'), 'must emit a workflow error annotation'); + assert.match(combined, /https:\/\/example\.test\/runs\/99/, 'must name the failing run URL'); + assert.equal(readOutputs(outputPath).verdict, VERDICT.RED); + }); + + test('exits 0 with a warning when the fix-next label bypasses a RED base branch', async (t) => { + const dir = createTempDir('next-health-bypass-'); + const api = await startApiStub(() => ({ + status: 200, + body: { workflow_runs: [{ conclusion: 'failure', html_url: 'https://example.test/runs/100' }] }, + })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url, PR_LABELS: `needs-triage,${BYPASS_LABEL}` }, + { outputPath }, + ); + + assert.equal(result.code, 0, `stdout: ${result.stdout}\nstderr: ${result.stderr}`); + const combined = `${result.stdout}${result.stderr}`; + assert.ok(combined.includes('::warning::'), 'must warn that a human bypassed the gate'); + assert.match(combined, /https:\/\/example\.test\/runs\/100/, 'the warning must name the failing run URL'); + assert.ok(!combined.includes('::error::'), 'a bypassed gate must not also emit an error annotation'); + assert.equal(readOutputs(outputPath).verdict, VERDICT.BYPASSED); + }); + + test('exits 0 on a clean base branch', async (t) => { + const dir = createTempDir('next-health-clean-'); + const api = await startApiStub(() => ({ status: 200, body: { workflow_runs: [{ conclusion: 'success' }] } })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url, PR_LABELS: '' }, + { outputPath }, + ); + + assert.equal(result.code, 0, result.stderr); + assert.equal(readOutputs(outputPath).verdict, VERDICT.CLEAN); + }); + + test('resolves the merge_group base ref from MERGE_GROUP_BASE_REF, stripping refs/heads/', async (t) => { + const dir = createTempDir('next-health-mergequeue-'); + const api = await startApiStub(() => ({ status: 200, body: { workflow_runs: [{ conclusion: 'success' }] } })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'merge_group', MERGE_GROUP_BASE_REF: 'refs/heads/next', GITHUB_API_URL: api.url }, + { outputPath }, + ); + + assert.equal(result.code, 0, result.stderr); + assert.equal(readOutputs(outputPath).verdict, VERDICT.CLEAN); + }); + + test('fails open when the API rejects the read', async (t) => { + const dir = createTempDir('next-health-403-'); + const api = await startApiStub(() => ({ status: 403, body: { message: 'Resource not accessible' } })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url }, + { outputPath }, + ); + + assert.equal(result.code, 0, 'an unreadable base branch is not a red base branch'); + assert.equal(readOutputs(outputPath).verdict, VERDICT.INDETERMINATE); + }); + + test('fails open when the API body is not a JSON object', async (t) => { + const dir = createTempDir('next-health-badjson-'); + const api = await startApiStub(() => ({ status: 200, body: 'not json at all' })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url }, + { outputPath }, + ); + + assert.equal(result.code, 0); + assert.equal(readOutputs(outputPath).verdict, VERDICT.INDETERMINATE); + }); + + test('works when GITHUB_OUTPUT is not set', (t) => { + const dir = createTempDir('next-health-nooutput-'); + t.after(() => cleanup(dir)); + + const result = runCli({ GITHUB_EVENT_NAME: 'push', GITHUB_OUTPUT: '' }); + + assert.equal(result.exitCode, 0, result.stderr); + }); + + test('an unwritable GITHUB_OUTPUT does not change the exit code', async (t) => { + // A failure while REPORTING the verdict must never invert the gate. + // Injected by pointing at a path whose parent does not exist — no mode-bit + // tricks, which root Docker/CI silently bypasses. + const dir = createTempDir('next-health-badout-'); + const api = await startApiStub(() => ({ status: 200, body: { workflow_runs: [{ conclusion: 'failure' }] } })); + t.after(async () => { await api.close(); cleanup(dir); }); + const outputPath = path.join(dir, 'no-such-dir', 'gh-output'); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url, PR_LABELS: '' }, + { outputPath }, + ); + + assert.equal(result.code, 1, 'a RED base branch must still exit 1 when the output write fails'); + }); + + test('never echoes the token', async (t) => { + const dir = createTempDir('next-health-token-'); + const api = await startApiStub(() => ({ status: 500, body: { message: 'boom' } })); + t.after(async () => { await api.close(); cleanup(dir); }); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url }, + { outputPath: path.join(dir, 'gh-output') }, + ); + + assert.ok(!`${result.stdout}${result.stderr}`.includes(SENTINEL_TOKEN)); + }); + + test('reports a failure without a stack trace', async (t) => { + const dir = createTempDir('next-health-nostack-'); + const api = await startApiStub(() => ({ status: 200, body: { workflow_runs: [{ conclusion: 'failure' }] } })); + t.after(async () => { await api.close(); cleanup(dir); }); + + const result = await callMain( + t, + { GITHUB_EVENT_NAME: 'pull_request', GITHUB_API_URL: api.url, PR_LABELS: '' }, + { outputPath: path.join(dir, 'gh-output') }, + ); + + assert.ok(!`${result.stdout}${result.stderr}`.includes(' at '), 'no raw stack frames'); + }); + + test('prints usage', () => { + const result = runNode([SCRIPT, '--help'], { cwd: ROOT, timeoutMs: PROBE_TIMEOUT_MS }); + assert.equal(result.exitCode, 0); + assert.ok(/Usage/i.test(result.stdout)); + }); + + test('rejects an unknown argument', () => { + const result = runNode([SCRIPT, '--nope'], { cwd: ROOT, timeoutMs: PROBE_TIMEOUT_MS }); + assert.notEqual(result.exitCode, 0); + assert.ok(`${result.stdout}${result.stderr}`.includes('--nope')); + }); +}); diff --git a/tests/gsd-check-update-worker-atomic-cache.test.cjs b/tests/gsd-check-update-worker-atomic-cache.test.cjs index b129ab141..ae418d821 100644 --- a/tests/gsd-check-update-worker-atomic-cache.test.cjs +++ b/tests/gsd-check-update-worker-atomic-cache.test.cjs @@ -30,9 +30,25 @@ const fs = require('fs'); const path = require('path'); const { runHook: runHookSeam } = require('./helpers/process-seam.cjs'); const { createTempDir, cleanup } = require('./helpers.cjs'); +const { NPM_VIEW_TIMEOUT_MS } = require('../gsd-core/bin/check-latest-version.cjs'); const WORKER_PATH = path.join(__dirname, '..', 'hooks', 'gsd-check-update-worker.js'); +// #4091/2026-09-06 Windows CI incident: the worker's real `npm view` call +// (checkLatestVersion, gsd-core/bin/check-latest-version.cjs) is bounded at +// NPM_VIEW_TIMEOUT_MS. This outer test harness timeout used to be hardcoded +// to the SAME 15000ms, so a slow registry response raced two SIGKILLs at the +// exact same wall-clock instant: the worker's own inner npm-view timeout +// fires and it needs real time to catch that failure, build a degraded +// result, and atomically publish the cache — but the outer harness could +// kill the whole process tree first (exitCode: null, empty stderr) before +// the worker ever got the chance. Windows's shell-wrapped npm subprocess +// (cmd.exe wrapper, see src/shell-command-projection.cts) adds enough spawn +// overhead to make this race lose more often there, but the zero-margin race +// itself is platform-agnostic. This margin gives the worker real headroom +// beyond the inner timeout it wraps. +const WORKER_TEARDOWN_MARGIN_MS = 10_000; + // allow-test-rule: structural-regression-guard (#4091) // Feeds the real worker source (readFileSync) into the structural assertions // below. The behavior it guards — rename-atomic publish of a file shared @@ -108,11 +124,25 @@ describe('gsd-check-update-worker.js: atomic cache publish (#4091)', () => { GSD_PROJECT_VERSION_FILE: path.join(cacheDir, 'no-such-project', 'VERSION'), GSD_GLOBAL_VERSION_FILE: path.join(cacheDir, 'no-such-global', 'VERSION'), }; - const r = runHookSeam(WORKER_PATH, [], { env, timeoutMs: 15000 }); + const r = runHookSeam(WORKER_PATH, [], { env, timeoutMs: NPM_VIEW_TIMEOUT_MS + WORKER_TEARDOWN_MARGIN_MS }); assert.equal(r.exitCode, 0, `worker must exit 0; stderr: ${r.stderr}`); const cache = JSON.parse(fs.readFileSync(cacheFile, 'utf8')); assert.equal(cache.installed, '0.0.0', 'worker record replaced the pre-existing one'); const residue = fs.readdirSync(cacheDir).filter((f) => f.startsWith('cache.json.tmp') || f.includes('.tmp-')); assert.deepEqual(residue, [], 'no temp stage files may remain after a successful publish'); }); + + test('outer worker-run timeout keeps real margin beyond the inner npm-view timeout (#4091 exact-tie race)', () => { + assert.ok( + WORKER_TEARDOWN_MARGIN_MS >= 5000, + 'the outer test timeout must give the worker real margin beyond the inner npm-view ' + + 'timeout (NPM_VIEW_TIMEOUT_MS) it wraps, or a slow registry response races the two ' + + 'SIGKILLs (see #4091/2026-09-06 Windows CI incident: exact-tie timeout killed the worker ' + + 'before it could degrade gracefully)', + ); + assert.ok( + NPM_VIEW_TIMEOUT_MS + WORKER_TEARDOWN_MARGIN_MS > NPM_VIEW_TIMEOUT_MS, + 'the combined outer timeout must strictly exceed the inner npm-view timeout it wraps', + ); + }); }); From 19b66c3ec80e95b0fd96e98b1b14ff267216bf4f Mon Sep 17 00:00:00 2001 From: Michel Moreira Date: Sun, 6 Sep 2026 18:51:34 -0300 Subject: [PATCH 025/166] fix(#4218): stop the orchestrator steering an executor that is still working (#4391) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4218): stop the orchestrator steering an executor that is still working An executor with recent RED/GREEN/REFACTOR commits and passing verification had not yet written its SUMMARY because it was finishing closeout. The parent saw no local OS test/build process, inferred an "idle tail", and sent "Finalize immediately" into a working child; in CLI runs the same inference interrupted an executor before GREEN, leaving a RED commit and an uncommitted edit. The stall block said only "if no completion signal, no SUMMARY.md, and no expected-branch commits appear for N minutes" — it never said what to do when commits DO exist and only the SUMMARY is outstanding, never defined the threshold as a period without progress rather than a total runtime, and never ruled out a process listing as an idleness signal. Four rules close that: - the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last sign of progress, not from dispatch — a long verification tail is not a stall; - commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with steering, interrupting and re-dispatching each named and forbidden; - urgency/finalization messages ("Finalize immediately" and family) are forbidden outright — they arrive mid-verification and truncate a correct run. The existing user-facing pause is the only sanctioned stop, and `kill and retry` is a clean restart, not a nudge; - the absence of a local OS test/build process is NOT idleness: a native subagent runs in the runtime's own session, and an executor between two tool calls shows no process at all. Progress is judged only by the signals this workflow names. Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on next. * fix(#4218): extract the progress policy to a step fragment CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600 ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's stated remedy, and this workflow already carries policy detail that way. execute-phase/steps/executor-progress-policy.md owns the policy. The worktree-recovery arm moved with it — `kill and switch to inline execution` qualifies the stop this policy governs, so it belongs beside the rule about when stopping is sanctioned at all, not stranded in the host. The #3212 recovery OPTIONS stay in the host, where tests/config.test.cjs pins them. The host keeps what must be read before the orchestrator acts: the verdict, the threshold definition, and a pointer that fires before any message is sent to the child. execute-phase.md is now 93475 bytes — 48 SMALLER than next. * chore: add changeset for #4218 * chore(#4218): regenerate the inventory manifest for the new step fragment docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind INVENTORY.md's `/steps/*.md` row, so a new fragment has to appear there or gen-inventory-manifest --check reds the lint-tests lane. * chore(#4218): restore the issue ref on the allow-test-rule marker ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the policy into the fragment dropped it. * chore(#4218): regenerate the install-tree fixtures for the new step fragment The fragment ships with the workflow, so every runtime's golden install tree gains one path — gen:install-tree is the generator that owns those fixtures. --------- Co-authored-by: Tom Boucher --- .changeset/fierce-koalas-climb.md | 5 ++ docs/INVENTORY-MANIFEST.json | 1 + gsd-core/workflows/execute-phase.md | 4 +- .../steps/executor-progress-policy.md | 43 ++++++++++ tests/execute-phase-wave.test.cjs | 85 +++++++++++++++++++ tests/fixtures/install-tree/antigravity.json | 1 + tests/fixtures/install-tree/augment.json | 1 + tests/fixtures/install-tree/claude-local.json | 1 + tests/fixtures/install-tree/claude.json | 1 + tests/fixtures/install-tree/cline.json | 1 + tests/fixtures/install-tree/codebuddy.json | 1 + tests/fixtures/install-tree/codex.json | 1 + tests/fixtures/install-tree/copilot.json | 1 + tests/fixtures/install-tree/cursor.json | 1 + tests/fixtures/install-tree/hermes.json | 1 + tests/fixtures/install-tree/kilo.json | 1 + tests/fixtures/install-tree/kimi-code.json | 1 + tests/fixtures/install-tree/kimi.json | 1 + tests/fixtures/install-tree/opencode.json | 1 + tests/fixtures/install-tree/pi.json | 1 + tests/fixtures/install-tree/qwen.json | 1 + tests/fixtures/install-tree/trae.json | 1 + tests/fixtures/install-tree/windsurf.json | 1 + tests/fixtures/install-tree/zcode.json | 1 + 24 files changed, 156 insertions(+), 1 deletion(-) create mode 100644 .changeset/fierce-koalas-climb.md create mode 100644 gsd-core/workflows/execute-phase/steps/executor-progress-policy.md diff --git a/.changeset/fierce-koalas-climb.md b/.changeset/fierce-koalas-climb.md new file mode 100644 index 000000000..8aedb02b0 --- /dev/null +++ b/.changeset/fierce-koalas-climb.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4391 +--- +**A working executor is no longer interrupted or told to "Finalize immediately"** — execute-phase's stall threshold now measures time without progress rather than total runtime, an executor with commits and recent activity is left alone until its SUMMARY lands, and a missing local test/build process no longer counts as idleness. (#4218) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 8d46675fd..421097462 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -617,6 +617,7 @@ "docs-update/steps/dispatch-monorepo-packages.md", "execute-phase/steps/codebase-drift-gate.md", "execute-phase/steps/executor-isolation-dispatch.md", + "execute-phase/steps/executor-progress-policy.md", "execute-phase/steps/gap-closure-artifacts.md", "execute-phase/steps/partial-wave.md", "execute-phase/steps/per-plan-executor-routing.md", diff --git a/gsd-core/workflows/execute-phase.md b/gsd-core/workflows/execute-phase.md index b7a3a6d79..025b34a63 100644 --- a/gsd-core/workflows/execute-phase.md +++ b/gsd-core/workflows/execute-phase.md @@ -878,7 +878,9 @@ increases monotonically across waves. `{status}` is `complete` (success), ask for one recovery path: `continue waiting`, `kill and retry`, or `kill and switch to inline execution`. - If the stalled executor ran in an isolated worktree, `kill and switch to inline execution` edits the primary checkout — see worktree recovery policy (`execute-phase/steps/worktree-recovery-policy.md`). Prefer `kill and retry` in a fresh worktree; inline execution requires explicit confirmation, never the default. + **A working executor is never steered (#4218).** The threshold measures time WITHOUT + PROGRESS, not total runtime. Before treating an executor as stalled — and before sending it + any message — read and execute `execute-phase/steps/executor-progress-policy.md`. **This fallback applies to all runtimes.** Claude Code's Agent() backgrounds by default: the completion signal may never arrive. Verify, never wait. diff --git a/gsd-core/workflows/execute-phase/steps/executor-progress-policy.md b/gsd-core/workflows/execute-phase/steps/executor-progress-policy.md new file mode 100644 index 000000000..49b0d309a --- /dev/null +++ b/gsd-core/workflows/execute-phase/steps/executor-progress-policy.md @@ -0,0 +1,43 @@ +Apply response_language to all user-facing prose — narration between tool calls, status updates, progress notes, and findings included; preserve code, paths, and identifiers. + +# Executor Progress Policy (#4218) + +Read this before treating any executor as stalled. + +## A working executor is never steered + +The stall threshold measures a period **without meaningful progress**. It is not a maximum +total runtime, and it is not a budget the executor has to finish inside. A plan with a long +verification or closeout tail legitimately spends many minutes between its last commit and +its SUMMARY. + +Reconcile activity as well as artifacts: + +- **Commits exist and SUMMARY.md is missing, with recent meaningful activity → KEEP + WAITING.** Do not steer it, do not interrupt it, do not re-dispatch it. Recent + RED/GREEN/REFACTOR commits, passing verification, or ongoing reasoning/tool telemetry + from the child are all meaningful activity. +- **Only after `${EXECUTOR_STALL_THRESHOLD_MINUTES}` of no meaningful progress** — measured + from the LAST sign of progress, not from dispatch — may the pause in step 3 fire, and it + asks the user; it does not act on its own. + +## Never inject urgency or finalization instructions into a live executor + +Messages of the "Finalize immediately", "wrap up now", "you are taking too long" family are +forbidden: they arrive mid-verification and turn a correct run into a truncated one. If an +executor must be stopped, the pause in step 3 is the only route, and `kill and retry` is a +clean restart — not a nudge. + +## The absence of a local OS test/build process is NOT idleness + +The orchestrator cannot see the child's work that way. A native subagent runs in the +runtime's own session, not as a visible local process, and an executor between two tool +calls — reasoning, reading a file, waiting on a runtime round-trip — shows no process at +all. Judge progress ONLY by the signals this workflow names: commits on the expected +branch, the SUMMARY, and the child's own activity. A process listing is not one of them. + +## If a stalled executor ran in an isolated worktree + +`kill and switch to inline execution` edits the primary checkout — see worktree recovery +policy (`execute-phase/steps/worktree-recovery-policy.md`). Prefer `kill and retry` in a +fresh worktree; inline execution requires explicit confirmation, never the default. diff --git a/tests/execute-phase-wave.test.cjs b/tests/execute-phase-wave.test.cjs index b23582afa..b86f8312f 100644 --- a/tests/execute-phase-wave.test.cjs +++ b/tests/execute-phase-wave.test.cjs @@ -1182,3 +1182,88 @@ describe('execute-phase workflow: #3684 review findings — join normalization', } }); }); + +// ── #4218: a live executor must not be steered or cut short ────────────────── +// +// allow-test-rule: source-text-is-the-product (#4218) — the workflow .md IS +// the instruction the orchestrator executes; its text is the artifact under test. +// +// Reported on Codex: an executor with recent RED/GREEN/REFACTOR commits and +// passing verification had not yet written its SUMMARY because it was finishing +// closeout. The parent saw no local OS test/build process, inferred an "idle +// tail", and sent "Finalize immediately" into a working child. In CLI runs the +// same inference interrupted an executor before GREEN, leaving a RED commit and +// an uncommitted edit. +// +// The policy lives in a step fragment: execute-phase.md is 77 bytes under the +// frozen #1168 ceiling, and "extract, not bump" is the repo's stated remedy. +describe('execute-phase: stall surveillance must not steer a working executor (#4218)', () => { + const workflow = fs.readFileSync(WORKFLOW_PATH, 'utf-8'); + const FRAGMENT_PATH = path.join( + __dirname, '..', 'gsd-core', 'workflows', 'execute-phase', 'steps', 'executor-progress-policy.md'); + const fragment = fs.existsSync(FRAGMENT_PATH) ? fs.readFileSync(FRAGMENT_PATH, 'utf-8') : ''; + + test('the host routes to the policy before treating an executor as stalled', () => { + assert.match(workflow, /A working executor is never steered \(#4218\)/, + 'the verdict must be visible where the orchestrator decides, not only in the fragment'); + assert.match(workflow, /time WITHOUT\s{1,10}PROGRESS, not total runtime/, + 'the threshold definition is the correction — it belongs in the host'); + assert.match(workflow, /before sending it\s{1,10}any message/, + 'the steering prohibition must be reachable before a message is sent, not after'); + assert.match(workflow, /execute-phase\/steps\/executor-progress-policy\.md/, + 'the host must point at the policy fragment'); + }); + + test('the stall threshold is a period without progress, not a maximum runtime', () => { + assert.ok(fragment.length > 0, 'execute-phase/steps/executor-progress-policy.md must exist'); + assert.match(fragment, /period \*\*without meaningful progress\*\*/, + 'the threshold must be defined by absence of progress'); + assert.match(fragment, /not a maximum\s{1,10}total runtime/, + 'a long-but-progressing plan must not be treated as stalled'); + assert.match(fragment, /from the LAST sign of progress, not from dispatch/, + 'measuring from dispatch is what turns a slow plan into a false stall'); + }); + + test('commits + missing SUMMARY + recent activity resolves to KEEP WAITING', () => { + assert.match(fragment, /KEEP\s{1,10}WAITING/, 'the verdict for a working executor must be explicit'); + assert.match(fragment, /Do not steer it, do not interrupt it, do not re-dispatch it/, + 'all three interventions the report describes must be named and forbidden'); + }); + + test('urgency and finalization messages are forbidden outright', () => { + assert.match(fragment, /Never inject urgency or finalization instructions into a live executor/, + 'the prohibition must be stated as a rule, not implied'); + assert.match(fragment, /Finalize immediately/, + 'the reported message must be named so it cannot be read as permitted'); + }); + + test('a missing local OS process is not evidence of idleness', () => { + assert.match(fragment, /absence of a local OS test\/build process is NOT idleness/, + 'the false signal the orchestrator acted on must be ruled out by name'); + assert.match(fragment, /A process listing is not one of them/, + 'the workflow must say which signals DO count, and that a process listing is not one'); + }); + + test('the only sanctioned stop is the existing user-facing pause', () => { + assert.match(fragment, /the pause in step 3 is the only route/, + 'stopping an executor must stay a user decision, not an orchestrator nudge'); + assert.match(fragment, /`kill and retry` is a\s{1,10}clean restart — not a nudge/, + 'the distinction between restarting and steering must be explicit'); + }); + + test('the worktree-recovery arm moved with the policy, not left duplicated', () => { + assert.match(fragment, /If a stalled executor ran in an isolated worktree/, + 'the recovery arm belongs with the stop policy it qualifies'); + assert.match(fragment, /worktree-recovery-policy\.md/, + 'it must still hand off to the worktree policy'); + assert.ok( + !/kill and switch to inline execution` edits the primary checkout/.test(workflow), + 'the host must not keep a second copy of the arm that moved', + ); + // The #3212 recovery options themselves stay in the host — tests/config.test.cjs + // pins them there. + for (const option of ['continue waiting', 'kill and retry', 'kill and switch to inline execution']) { + assert.ok(workflow.includes(option), `the host must still offer "${option}"`); + } + }); +}); diff --git a/tests/fixtures/install-tree/antigravity.json b/tests/fixtures/install-tree/antigravity.json index 1053ba05f..e8a7e42de 100644 --- a/tests/fixtures/install-tree/antigravity.json +++ b/tests/fixtures/install-tree/antigravity.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/augment.json b/tests/fixtures/install-tree/augment.json index e94cf8efc..1e5d20c0a 100644 --- a/tests/fixtures/install-tree/augment.json +++ b/tests/fixtures/install-tree/augment.json @@ -456,6 +456,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/claude-local.json b/tests/fixtures/install-tree/claude-local.json index 0de1554f6..4778a6df5 100644 --- a/tests/fixtures/install-tree/claude-local.json +++ b/tests/fixtures/install-tree/claude-local.json @@ -349,6 +349,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/claude.json b/tests/fixtures/install-tree/claude.json index 89ffa04f0..a2f013c83 100644 --- a/tests/fixtures/install-tree/claude.json +++ b/tests/fixtures/install-tree/claude.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/cline.json b/tests/fixtures/install-tree/cline.json index 5fa6c7c27..d9899d85a 100644 --- a/tests/fixtures/install-tree/cline.json +++ b/tests/fixtures/install-tree/cline.json @@ -386,6 +386,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/codebuddy.json b/tests/fixtures/install-tree/codebuddy.json index eae7b37fc..d51cf2b3a 100644 --- a/tests/fixtures/install-tree/codebuddy.json +++ b/tests/fixtures/install-tree/codebuddy.json @@ -456,6 +456,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/codex.json b/tests/fixtures/install-tree/codex.json index 863e061c5..89bef7a38 100644 --- a/tests/fixtures/install-tree/codex.json +++ b/tests/fixtures/install-tree/codex.json @@ -420,6 +420,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/copilot.json b/tests/fixtures/install-tree/copilot.json index 49fec394f..1cb27a87c 100644 --- a/tests/fixtures/install-tree/copilot.json +++ b/tests/fixtures/install-tree/copilot.json @@ -385,6 +385,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/cursor.json b/tests/fixtures/install-tree/cursor.json index 24e453c37..f54509b21 100644 --- a/tests/fixtures/install-tree/cursor.json +++ b/tests/fixtures/install-tree/cursor.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/hermes.json b/tests/fixtures/install-tree/hermes.json index 4cf935f9d..0e7ba27f5 100644 --- a/tests/fixtures/install-tree/hermes.json +++ b/tests/fixtures/install-tree/hermes.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/kilo.json b/tests/fixtures/install-tree/kilo.json index abf3fbd8a..7b6df1af2 100644 --- a/tests/fixtures/install-tree/kilo.json +++ b/tests/fixtures/install-tree/kilo.json @@ -456,6 +456,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/kimi-code.json b/tests/fixtures/install-tree/kimi-code.json index 0c89dc91a..71a4948b6 100644 --- a/tests/fixtures/install-tree/kimi-code.json +++ b/tests/fixtures/install-tree/kimi-code.json @@ -385,6 +385,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/kimi.json b/tests/fixtures/install-tree/kimi.json index fe38de7ea..951b389f2 100644 --- a/tests/fixtures/install-tree/kimi.json +++ b/tests/fixtures/install-tree/kimi.json @@ -421,6 +421,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/opencode.json b/tests/fixtures/install-tree/opencode.json index 08896f64c..863d5be00 100644 --- a/tests/fixtures/install-tree/opencode.json +++ b/tests/fixtures/install-tree/opencode.json @@ -456,6 +456,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/pi.json b/tests/fixtures/install-tree/pi.json index 6dc5cf674..17ab254b6 100644 --- a/tests/fixtures/install-tree/pi.json +++ b/tests/fixtures/install-tree/pi.json @@ -244,6 +244,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/qwen.json b/tests/fixtures/install-tree/qwen.json index bbf1a8179..1989d7bce 100644 --- a/tests/fixtures/install-tree/qwen.json +++ b/tests/fixtures/install-tree/qwen.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/trae.json b/tests/fixtures/install-tree/trae.json index 2736df347..78368cd79 100644 --- a/tests/fixtures/install-tree/trae.json +++ b/tests/fixtures/install-tree/trae.json @@ -384,6 +384,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/windsurf.json b/tests/fixtures/install-tree/windsurf.json index 9d1aacaf2..0257661eb 100644 --- a/tests/fixtures/install-tree/windsurf.json +++ b/tests/fixtures/install-tree/windsurf.json @@ -312,6 +312,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", diff --git a/tests/fixtures/install-tree/zcode.json b/tests/fixtures/install-tree/zcode.json index 9d8d555e1..b117fe29b 100644 --- a/tests/fixtures/install-tree/zcode.json +++ b/tests/fixtures/install-tree/zcode.json @@ -456,6 +456,7 @@ "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", + "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", "gsd-core/workflows/execute-phase/steps/partial-wave.md", "gsd-core/workflows/execute-phase/steps/per-plan-executor-routing.md", From 2cf119f57e0f812099a96b4e3a89b2c7d2d31105 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 18:57:38 -0400 Subject: [PATCH 026/166] fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4217): reconcile artifacts before classifying abnormal ends * test(#4217): pin the completion-reconciliation contract * chore(#4217): regen derived inventory and install-tree fixtures * test(#4217): follow the #4003 anchoring pins into the reconciliation fragment Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217) * chore(#4217): add changeset fragment * chore(#4217): backfill PR number in changeset fragment --------- Co-authored-by: sim --- .changeset/silly-lemurs-munch.md | 5 + docs/INVENTORY-MANIFEST.json | 1 + gsd-core/workflows/execute-phase.md | 34 +-- .../steps/completion-reconciliation.md | 53 ++++ ...e-phase-completion-reconciliation.test.cjs | 256 ++++++++++++++++++ tests/fixtures/install-tree/antigravity.json | 1 + tests/fixtures/install-tree/augment.json | 1 + tests/fixtures/install-tree/claude-local.json | 1 + tests/fixtures/install-tree/claude.json | 1 + tests/fixtures/install-tree/cline.json | 1 + tests/fixtures/install-tree/codebuddy.json | 1 + tests/fixtures/install-tree/codex.json | 1 + tests/fixtures/install-tree/copilot.json | 1 + tests/fixtures/install-tree/cursor.json | 1 + tests/fixtures/install-tree/hermes.json | 1 + tests/fixtures/install-tree/kilo.json | 1 + tests/fixtures/install-tree/kimi-code.json | 1 + tests/fixtures/install-tree/kimi.json | 1 + tests/fixtures/install-tree/opencode.json | 1 + tests/fixtures/install-tree/pi.json | 1 + tests/fixtures/install-tree/qwen.json | 1 + tests/fixtures/install-tree/trae.json | 1 + tests/fixtures/install-tree/windsurf.json | 1 + tests/fixtures/install-tree/zcode.json | 1 + tests/safe-resume-gate-anchoring.test.cjs | 13 +- 25 files changed, 352 insertions(+), 29 deletions(-) create mode 100644 .changeset/silly-lemurs-munch.md create mode 100644 gsd-core/workflows/execute-phase/steps/completion-reconciliation.md create mode 100644 tests/execute-phase-completion-reconciliation.test.cjs diff --git a/.changeset/silly-lemurs-munch.md b/.changeset/silly-lemurs-munch.md new file mode 100644 index 000000000..973b5e297 --- /dev/null +++ b/.changeset/silly-lemurs-munch.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4442 +--- +**`/gsd-execute-phase` no longer closes a finished executor as `turn_aborted`** — an executor whose plan SUMMARY and matching commits are already on disk is now reconciled as complete when its session ends abnormally, instead of waiting indefinitely for a terminal response and failing. (#4217) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 421097462..5e3b0415a 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -616,6 +616,7 @@ "discuss-phase-assumptions/steps/auto-advance-dispatch.md", "docs-update/steps/dispatch-monorepo-packages.md", "execute-phase/steps/codebase-drift-gate.md", + "execute-phase/steps/completion-reconciliation.md", "execute-phase/steps/executor-isolation-dispatch.md", "execute-phase/steps/executor-progress-policy.md", "execute-phase/steps/gap-closure-artifacts.md", diff --git a/gsd-core/workflows/execute-phase.md b/gsd-core/workflows/execute-phase.md index 025b34a63..a5ef4ac59 100644 --- a/gsd-core/workflows/execute-phase.md +++ b/gsd-core/workflows/execute-phase.md @@ -25,6 +25,9 @@ Orchestrator coordinates, not executes. Each subagent loads the full execute-pla instead of spawning parallel agents. Only attempt parallel spawning if the user explicitly requests it — and in that case, rely on the spot-check fallback in step 3 to detect completion. +- **Codex:** native subagent sessions can end abnormally (`turn_aborted`) after the plan + work is already committed. Completion is decided by the step-4 artifact reconciliation + (SUMMARY + matching recent commits), not by the session's terminal state (#4217). - **Other runtimes:** If `Agent`/`agent` tool is genuinely unavailable (e.g. a backgrounded Claude Code agent per #853, or a non-Claude runtime), use sequential inline execution as the fallback for executor parallelization only. If `Agent` IS available (top-level Claude @@ -806,7 +809,7 @@ increases monotonically across waves. `{status}` is `complete` (success), > **Worktree recovery policy (#48 + #1292):** See `execute-phase/steps/worktree-recovery-policy.md` — FAIL-CLOSED rule for base/HEAD-namespace mismatches AND isolated-run fail-safe recovery. - > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above to spawn executor agent(s), stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available. + > **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above to spawn executor agent(s), stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available. While waiting, run the step-4 completion surveillance; if the child's session ends abnormally — including `turn_aborted` — reconcile artifacts per `execute-phase/steps/completion-reconciliation.md` before classifying the plan (#4217). **Orchestrator-managed worktree dispatch** (`ISOLATION=orchestrator-worktree`): read and execute `execute-phase/steps/executor-isolation-dispatch.md`. GSD creates each worktree (`worktree create`) and spawns the executor into it; the orchestrator performs every git operation. Merge-back and cleanup are the existing manifest-scoped gauntlet, unchanged. @@ -848,28 +851,9 @@ increases monotonically across waves. `{status}` is `complete` (success), [checkpoint] phase {PHASE_NUMBER} wave {N}/{M} plan {plan_id} checkpoint ({P}/{Q} plans done) ``` - **Completion signal fallback (Copilot and runtimes where Agent() may not return):** + **Completion reconciliation (EVERY runtime — any spawn whose terminal response may not arrive):** - If a spawned agent does not return a completion signal but appears to have finished - its work, do NOT block indefinitely. Instead, verify completion via spot-checks: - - ```bash - # For each plan in this wave, check if the executor finished: - SUMMARY_EXISTS=$(test -f "{phase_dir}/{plan_number}-{plan_padded}-SUMMARY.md" && echo "true" || echo "false") - # #4003: anchored, zero-pad-tolerant scope (see safe_resume_gate); --since stays. - SPOT_PHASE_N=$((10#{phase_number})) - SPOT_PLAN_N=$((10#{plan_padded})) - COMMITS_FOUND=$(git log --oneline --all -E --grep="^[a-z]+\((0*${SPOT_PHASE_N})-(0*${SPOT_PLAN_N})\):" --since="1 hour ago" | head -1) - COMMITS_SINCE_DISPATCH=$(git log "${EXPECTED_BRANCH}" --since="${DISPATCH_TS}" --oneline | head -1) - ``` - - **If SUMMARY.md exists AND commits are found:** The agent completed successfully — - treat as done and proceed to step 5. Log: `"✓ {Plan ID} completed (verified via spot-check — completion signal not received)"` - - **If SUMMARY.md does NOT exist after a reasonable wait:** The agent may still be - running or may have failed silently. Check `git log --oneline -5` for recent - activity. If commits are still appearing, wait longer. If no activity, report - the plan as failed and route to the failure handler in step 6. + If a spawned agent does not return a normal terminal completion response — or its session ends abnormally (interrupted, aborted, closed, killed, timed out, `turn_aborted`, including ends the orchestrator itself initiated) — do NOT block indefinitely and do NOT classify the plan as failed yet. Read and execute `gsd-core/workflows/execute-phase/steps/completion-reconciliation.md` — reconcile the plan artifacts FIRST, classify SECOND: SUMMARY present AND matching recent commits → complete (proceed to step 5, do NOT re-dispatch); no completion evidence → the failure handler. Verify, never wait. **Configurable stall surveillance (#3212):** Every `${EXECUTOR_STALL_INTERVAL_MINUTES}` minutes while waiting, inspect `git log "${EXPECTED_BRANCH}" --since="${DISPATCH_TS}"` @@ -882,9 +866,6 @@ increases monotonically across waves. `{status}` is `complete` (success), PROGRESS, not total runtime. Before treating an executor as stalled — and before sending it any message — read and execute `execute-phase/steps/executor-progress-policy.md`. - **This fallback applies to all runtimes.** Claude Code's Agent() backgrounds by - default: the completion signal may never arrive. Verify, never wait. - 5. **Post-wave hook validation (parallel mode only):** Hooks run on every executor commit by default (#2924); this post-wave run only fires when `workflow.worktree_skip_hooks=true` opted out of per-commit hooks: ```bash SKIP_HOOKS=$(gsd_run query config-get workflow.worktree_skip_hooks --raw 2>/dev/null || echo "false") @@ -1117,6 +1098,7 @@ increases monotonically across waves. `{status}` is `complete` (success), if [ -n "$RETRY_AFTER" ]; then RETRY_HINT=" Provider hinted retry-after: ${RETRY_AFTER}s"; else RETRY_HINT=""; fi ``` One classifier branch handles sentinels across Claude/Copilot/Codex/Gemini. Reference: `docs/research/provider-rate-limit-signals.md`. + **Abnormal ends reconcile first (#4217):** an abnormal session end (`turn_aborted`-class) routes through the step-4 artifact reconciliation BEFORE classifying the failure — artifacts decide. **Step 7.1 — `class == "quota-exceeded"`:** follow the quota-recovery fragment below. **Step 7.2 — `class == "classify-handoff-bug"`:** If error contains `classifyHandoffIfNeeded is not defined`, treat as Claude runtime bug. Run the same step-5 spot-checks; PASS => treat as success, FAIL => fall through. @@ -1320,7 +1302,7 @@ ${VERIFIER_SKILLS}", ) ``` -> **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above, stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available. +> **ORCHESTRATOR RULE — CODEX RUNTIME**: After calling Agent() above, stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available. If the session ends abnormally (`turn_aborted`), reconcile via the `verification.status` query below — the session's terminal state is not evidence of failure (#4217). Read status via the canonical query (scoped to frontmatter, covers missing/unknown cases): ```bash diff --git a/gsd-core/workflows/execute-phase/steps/completion-reconciliation.md b/gsd-core/workflows/execute-phase/steps/completion-reconciliation.md new file mode 100644 index 000000000..42bfbd0fa --- /dev/null +++ b/gsd-core/workflows/execute-phase/steps/completion-reconciliation.md @@ -0,0 +1,53 @@ +# Completion reconciliation (#4217, split A of #3754) + +Read and follow this fragment from `execute-phase.md` step 4 whenever an executor's +completion is in question. It owns the whole reconciliation policy — both arms — so the +host wait step stays inside the ADR-857 Phase 6 byte ceiling (#1168). + +**Reconcile FIRST, classify SECOND.** How the child's session ended is bookkeeping +about the transport; what it wrote to disk and to git is the evidence about the work. + +## When this runs + +1. **No terminal response** — a spawned agent does not return a normal terminal + completion signal but appears to have finished its work (or may still be running). +2. **Abnormal end** — the child's session ended without a normal terminal completion + response: interrupted, aborted, closed, killed, timed out, or ended `turn_aborted` — + INCLUDING ends the orchestrator itself initiated. **An abnormally-ended child is + not evidence of failure (#4217):** the orchestrator's own interrupt/close says + nothing about whether the work completed; only the artifacts do. + +This policy applies to EVERY runtime and every isolation path — harness `Agent()` +dispatches, orchestrator-worktree process spawns, and sequential dispatch alike. Never +block indefinitely waiting for a signal; verify via filesystem and git state. + +## Probes (per plan in the wave) + +```bash +# For each plan in this wave, check if the executor finished: +SUMMARY_EXISTS=$(test -f "{phase_dir}/{plan_number}-{plan_padded}-SUMMARY.md" && echo "true" || echo "false") +# #4003: anchored, zero-pad-tolerant scope (see safe_resume_gate); --since stays. +SPOT_PHASE_N=$((10#{phase_number})) +SPOT_PLAN_N=$((10#{plan_padded})) +COMMITS_FOUND=$(git log --oneline --all -E --grep="^[a-z]+\((0*${SPOT_PHASE_N})-(0*${SPOT_PLAN_N})\):" --since="1 hour ago" | head -1) +COMMITS_SINCE_DISPATCH=$(git log "${EXPECTED_BRANCH}" --since="${DISPATCH_TS}" --oneline | head -1) +``` + +## Verdicts + +**If SUMMARY.md exists AND matching commits are found:** the agent completed +successfully — treat the plan as complete WITHOUT requiring another terminal child +response, proceed to step 5, and do NOT re-dispatch a fresh executor for this plan: +the work is already committed, and a second executor would redo it on top of itself. +Log: `"✓ {Plan ID} completed (verified via spot-check — completion signal not received)"`. + +**If SUMMARY.md does NOT exist after a reasonable wait:** the agent may still be +running or may have failed silently. Check `git log --oneline -5` for recent +activity. If commits are still appearing, wait longer. If no activity, report the +plan as failed and route to the failure handler in step 6. + +Evidence is BOTH probes or neither: a SUMMARY without matching commits, and matching +commits without a SUMMARY, are each incomplete evidence — never auto-complete on one +of them. When an abnormal end reconciles to no completion evidence, it stays failed: +route to the failure handler exactly as a normal failure would, and let the +safe-resume gate handle any un-summarized commits on the next run. diff --git a/tests/execute-phase-completion-reconciliation.test.cjs b/tests/execute-phase-completion-reconciliation.test.cjs new file mode 100644 index 000000000..3e50bd001 --- /dev/null +++ b/tests/execute-phase-completion-reconciliation.test.cjs @@ -0,0 +1,256 @@ +/** + * Regression tests for #4217 (split A of #3754): artifact-complete executor + * not reconciled — SUMMARY + matching commits present yet closed as turn_aborted. + * + * The supervision contract lives in shipped workflow text + * (gsd-core/workflows/execute-phase.md + its completion-reconciliation step + * fragment), which is the instruction the orchestrator executes at runtime. + * Reading the .md and asserting on its clauses tests the deployed contract + * (the source-text-is-the-product category; .md reads are outside + * no-source-grep's .cjs-only scope). + * + * Red on `next` before the fix: rows 1-7 of + * .gsd/bug/fix-4217-codex-artifact-complete-reconcile/50-test-matrix.md. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('fs'); +const path = require('path'); + +const WORKFLOW_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'execute-phase.md'); +const FRAGMENT_PATH = path.join( + __dirname, '..', 'gsd-core', 'workflows', 'execute-phase', 'steps', 'completion-reconciliation.md' +); + +function readWorkflow() { + return fs.readFileSync(WORKFLOW_PATH, 'utf-8'); +} + +function readFragment() { + return fs.readFileSync(FRAGMENT_PATH, 'utf-8'); +} + +/** + * Slice the step-4 "Wait for all agents in wave to complete" region — the + * supervision surface that governs executor outcomes while/after waiting. + */ +function waitStepRegion(content) { + const from = content.indexOf('4. **Wait for all agents in wave to complete.**'); + assert.ok(from !== -1, 'execute-phase.md must contain the step-4 wait region'); + const to = content.indexOf('5. **Post-wave hook validation', from); + assert.ok(to !== -1, 'execute-phase.md step 4 must be followed by step 5'); + return content.slice(from, to); +} + +/** The abnormal-termination shapes the contract must name (#3754 / #4217). */ +const ABNORMAL_SHAPES = ['interrupted', 'aborted', 'closed', 'killed', 'timed out', 'turn_aborted']; + +describe('execute-phase completion reconciliation (#4217 — split A of #3754)', () => { + test('workflow file exists', () => { + assert.ok(fs.existsSync(WORKFLOW_PATH), 'workflows/execute-phase.md should exist'); + }); + + // ── Row 1: the #4217 lifecycle regression ──────────────────────────────── + describe('row 1 — artifact-complete executor with an abnormal end is reconciled, not failed', () => { + test('step 4 carries an abnormal-end clause covering every termination shape', () => { + const region = waitStepRegion(readWorkflow()); + for (const shape of ABNORMAL_SHAPES) { + assert.ok( + region.includes(shape), + `step 4 must name the abnormal-termination shape "${shape}" so no runtime reads its own case as unlisted (#4217)` + ); + } + }); + + test('step 4 requires reconciliation BEFORE classifying an abnormally-ended executor', () => { + const region = waitStepRegion(readWorkflow()); + assert.match( + region, + /reconcil[\s\S]{0,400}FIRST[\s\S]{0,200}classify[\s\S]{0,20}SECOND/i, + 'step 4 must state the order outright: reconcile the artifacts FIRST, classify SECOND (#4217)' + ); + }); + + test('step 4 covers ends the orchestrator itself initiated (interrupt/close)', () => { + const region = waitStepRegion(readWorkflow()); + assert.match( + region, + /orchestrator itself (interrupted|closed|initiated)/i, + 'the abnormal-end clause must survive an orchestrator-initiated close — the exact reported case (#4217)' + ); + }); + + test('step 7 requires reconciliation before classifying an abnormal end as failure', () => { + const content = readWorkflow(); + const from = content.indexOf('7. **Handle failures:**'); + assert.ok(from !== -1, 'execute-phase.md must contain the step-7 failure handler'); + const step7 = content.slice(from, content.indexOf('@~/.claude/gsd-core/references/execute-phase-quota-recovery.md', from)); + assert.match( + step7, + /reconcil[\s\S]{0,200}BEFORE classifying|BEFORE classifying[\s\S]{0,200}reconcil/i, + 'step 7 must route abnormal session ends through artifact reconciliation BEFORE classifying the failure (#4217 D4 gap)' + ); + }); + + test('fragment verdict: SUMMARY AND matching commits resolve complete without another terminal response', () => { + const fragment = readFragment(); + assert.match( + fragment, + /SUMMARY[\s\S]{0,200}(AND|and)[\s\S]{0,200}matching[\s\S]{0,300}complete/i, + 'the fragment must spell out the complete verdict: SUMMARY present AND matching commits present => complete (#4217)' + ); + assert.match( + fragment, + /(without requiring|no) another terminal|terminal child response/i, + 'the complete verdict must not require another terminal child response (#4217)' + ); + }); + + test('fragment verdict forbids re-dispatch on the reconciled-complete path', () => { + const fragment = readFragment(); + assert.match( + fragment, + /do NOT re-?dispatch/i, + 'a reconciled-complete plan must not be re-dispatched — a second executor would redo committed work on top of itself (#4217)' + ); + }); + }); + + // ── Row 2: runtime-neutral fallback scope ──────────────────────────────── + describe('row 2 — the fallback is scoped to EVERY runtime, not Copilot', () => { + test('step-4 fallback heading is runtime-neutral', () => { + const region = waitStepRegion(readWorkflow()); + assert.match( + region, + /Completion reconciliation \(EVERY runtime/i, + 'the fallback heading must not scope to Copilot or to runtimes where Agent() may not return (#4217)' + ); + assert.doesNotMatch( + region, + /Completion signal fallback \(Copilot and runtimes where Agent\(\) may not return\)/, + 'the Copilot-scoped heading was the scope gap that kept Codex out (#4217)' + ); + }); + + test('step 4 names the completion-reconciliation fragment by exact path', () => { + const region = waitStepRegion(readWorkflow()); + assert.ok( + region.includes('gsd-core/workflows/execute-phase/steps/completion-reconciliation.md'), + 'step 4 must direct the orchestrator to the reconciliation fragment by exact path (also proves response-language inheritance)' + ); + }); + }); + + // ── Rows 3-5: negative space — the reconciliation must NOT over-complete ── + describe('rows 3-5 — evidence-less or partial-evidence ends are never auto-completed', () => { + test('fragment: SUMMARY alone is not sufficient evidence', () => { + const fragment = readFragment(); + assert.match( + fragment, + /BOTH probes or neither/i, + 'the reconciliation must take BOTH probes together — never one alone (#4217 negative space)' + ); + assert.match( + fragment, + /SUMMARY without matching commits[\s\S]{0,160}commits without a SUMMARY[\s\S]{0,160}incomplete evidence/i, + 'SUMMARY-without-commits and commits-without-SUMMARY must each be named as incomplete evidence' + ); + }); + + test('fragment: commits without SUMMARY keep the wait-longer arm', () => { + const fragment = readFragment(); + assert.match( + fragment, + /commits are still appearing, wait longer|wait longer/i, + 'commits-without-SUMMARY must keep the existing wait-longer arm (still working / closeout incomplete)' + ); + assert.match( + fragment, + /If SUMMARY\.md does NOT exist/i, + 'the incomplete arm must remain: SUMMARY missing after a reasonable wait routes to activity check / failure handler' + ); + }); + + test('fragment: the abnormal-end clause never converts an evidence-less abnormal end into success', () => { + const fragment = readFragment(); + assert.match( + fragment, + /not evidence of failure/i, + 'the clause must say an abnormally-ended child is not evidence of failure — and the converse holds: no evidence, no completion (#4217)' + ); + assert.match( + fragment, + /route to the failure handler/i, + 'when reconciliation finds no completion evidence, the abnormal end still routes to the failure handler (stays failed)' + ); + }); + }); + + // ── Row 6: the Codex wait rule is bounded and linked ───────────────────── + describe('row 6 — the Codex orchestrator wait rule is bound to the reconciliation', () => { + test('every CODEX RUNTIME wait rule references the reconciliation/surveillance surface', () => { + const content = readWorkflow(); + const blocks = [...content.matchAll(/ORCHESTRATOR RULE — CODEX RUNTIME([\s\S]{0,700}?)(?=\n\s*\n)/g)]; + assert.ok(blocks.length >= 2, 'both CODEX RUNTIME wait rules must exist (dispatch + verify dispatch)'); + for (const [, body] of blocks) { + assert.match( + body, + /completion reconciliation|reconcil|spot-check|surveillance/i, + 'a CODEX RUNTIME wait rule must bind the wait to the step-4 reconciliation — an unbounded wait is the #4217 deadlock' + ); + } + }); + }); + + // ── Row 7: runtime_compatibility names Codex ───────────────────────────── + describe('row 7 — runtime_compatibility declares Codex covered', () => { + test('runtime_compatibility names Codex and defers completion to artifact reconciliation', () => { + const content = readWorkflow(); + const from = content.indexOf(''); + const to = content.indexOf('', from); + assert.ok(from !== -1 && to !== -1, 'runtime_compatibility block must exist'); + const rtc = content.slice(from, to); + assert.match(rtc, /\*\*Codex:/, 'runtime_compatibility must name Codex (#4217)'); + assert.match( + rtc, + /SUMMARY[\s\S]{0,160}(and|\+)[\s\S]{0,160}commits|spot-check/i, + 'the Codex entry must point completion decisions at the artifact reconciliation (SUMMARY + matching commits)' + ); + }); + }); + + // ── Rows 8-10: preserved-behavior guards ───────────────────────────────── + describe('rows 8-10 — preserved behavior and seam boundaries', () => { + test('host still carries the spot-check vocabulary (agent-frontmatter contract)', () => { + const content = readWorkflow(); + assert.ok(content.includes('spot-check'), 'execute-phase must keep spot-check fallback vocabulary'); + assert.ok( + content.includes('sequential inline execution'), + 'execute-phase must keep the Copilot sequential inline fallback wording' + ); + }); + + test('stall surveillance block is untouched (#4218 seam)', () => { + const content = readWorkflow(); + const from = content.indexOf('**Configurable stall surveillance (#3212):**'); + assert.ok(from !== -1, 'the #3212 stall surveillance block must remain in step 4'); + const block = content.slice(from, content.indexOf('If the stalled executor', from)); + assert.match(block, /EXECUTOR_STALL_INTERVAL_MINUTES/, 'stall interval config unchanged'); + assert.match(block, /EXECUTOR_STALL_THRESHOLD_MINUTES/, 'stall threshold config unchanged'); + }); + + test('host stays under the frozen ADR-857 Phase 6 ceiling (#1168)', () => { + const { lfByteCount } = require('../scripts/workflow-size.cjs'); + const bytes = lfByteCount(WORKFLOW_PATH); + assert.ok(bytes < 93600, `execute-phase.md must stay below the frozen pre-phase-6 ceiling (93600); got ${bytes}`); + }); + + test('fragment keeps the #4003 anchored commit-scope probe and dispatch bound', () => { + const fragment = readFragment(); + assert.match(fragment, /0\*\$\{SPOT_PHASE_N\}/, 'probe keeps the zero-pad-tolerant anchored scope (#4003)'); + assert.match(fragment, /--since="\$\{DISPATCH_TS\}"/, 'probe keeps the dispatch-time bound'); + assert.match(fragment, /SUMMARY_EXISTS/, 'probe keeps the SUMMARY existence check'); + }); + }); +}); diff --git a/tests/fixtures/install-tree/antigravity.json b/tests/fixtures/install-tree/antigravity.json index e8a7e42de..8943408be 100644 --- a/tests/fixtures/install-tree/antigravity.json +++ b/tests/fixtures/install-tree/antigravity.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/augment.json b/tests/fixtures/install-tree/augment.json index 1e5d20c0a..588734b0b 100644 --- a/tests/fixtures/install-tree/augment.json +++ b/tests/fixtures/install-tree/augment.json @@ -455,6 +455,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/claude-local.json b/tests/fixtures/install-tree/claude-local.json index 4778a6df5..3cf79d7da 100644 --- a/tests/fixtures/install-tree/claude-local.json +++ b/tests/fixtures/install-tree/claude-local.json @@ -348,6 +348,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/claude.json b/tests/fixtures/install-tree/claude.json index a2f013c83..95195ce26 100644 --- a/tests/fixtures/install-tree/claude.json +++ b/tests/fixtures/install-tree/claude.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/cline.json b/tests/fixtures/install-tree/cline.json index d9899d85a..6149f0fd9 100644 --- a/tests/fixtures/install-tree/cline.json +++ b/tests/fixtures/install-tree/cline.json @@ -385,6 +385,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/codebuddy.json b/tests/fixtures/install-tree/codebuddy.json index d51cf2b3a..18f097b05 100644 --- a/tests/fixtures/install-tree/codebuddy.json +++ b/tests/fixtures/install-tree/codebuddy.json @@ -455,6 +455,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/codex.json b/tests/fixtures/install-tree/codex.json index 89bef7a38..c0b65addc 100644 --- a/tests/fixtures/install-tree/codex.json +++ b/tests/fixtures/install-tree/codex.json @@ -419,6 +419,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/copilot.json b/tests/fixtures/install-tree/copilot.json index 1cb27a87c..92efd5ee9 100644 --- a/tests/fixtures/install-tree/copilot.json +++ b/tests/fixtures/install-tree/copilot.json @@ -384,6 +384,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/cursor.json b/tests/fixtures/install-tree/cursor.json index f54509b21..eb4503859 100644 --- a/tests/fixtures/install-tree/cursor.json +++ b/tests/fixtures/install-tree/cursor.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/hermes.json b/tests/fixtures/install-tree/hermes.json index 0e7ba27f5..d0a1b9ac0 100644 --- a/tests/fixtures/install-tree/hermes.json +++ b/tests/fixtures/install-tree/hermes.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/kilo.json b/tests/fixtures/install-tree/kilo.json index 7b6df1af2..d43d47aa4 100644 --- a/tests/fixtures/install-tree/kilo.json +++ b/tests/fixtures/install-tree/kilo.json @@ -455,6 +455,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/kimi-code.json b/tests/fixtures/install-tree/kimi-code.json index 71a4948b6..b0970bd63 100644 --- a/tests/fixtures/install-tree/kimi-code.json +++ b/tests/fixtures/install-tree/kimi-code.json @@ -384,6 +384,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/kimi.json b/tests/fixtures/install-tree/kimi.json index 951b389f2..dcc0bac5c 100644 --- a/tests/fixtures/install-tree/kimi.json +++ b/tests/fixtures/install-tree/kimi.json @@ -420,6 +420,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/opencode.json b/tests/fixtures/install-tree/opencode.json index 863d5be00..366cf8ec7 100644 --- a/tests/fixtures/install-tree/opencode.json +++ b/tests/fixtures/install-tree/opencode.json @@ -455,6 +455,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/pi.json b/tests/fixtures/install-tree/pi.json index 17ab254b6..d4fb2bc2d 100644 --- a/tests/fixtures/install-tree/pi.json +++ b/tests/fixtures/install-tree/pi.json @@ -243,6 +243,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/qwen.json b/tests/fixtures/install-tree/qwen.json index 1989d7bce..99ac81efa 100644 --- a/tests/fixtures/install-tree/qwen.json +++ b/tests/fixtures/install-tree/qwen.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/trae.json b/tests/fixtures/install-tree/trae.json index 78368cd79..b1e4423f6 100644 --- a/tests/fixtures/install-tree/trae.json +++ b/tests/fixtures/install-tree/trae.json @@ -383,6 +383,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/windsurf.json b/tests/fixtures/install-tree/windsurf.json index 0257661eb..d067cde15 100644 --- a/tests/fixtures/install-tree/windsurf.json +++ b/tests/fixtures/install-tree/windsurf.json @@ -311,6 +311,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/fixtures/install-tree/zcode.json b/tests/fixtures/install-tree/zcode.json index b117fe29b..9b4bc9ae2 100644 --- a/tests/fixtures/install-tree/zcode.json +++ b/tests/fixtures/install-tree/zcode.json @@ -455,6 +455,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", + "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", "gsd-core/workflows/execute-phase/steps/executor-progress-policy.md", "gsd-core/workflows/execute-phase/steps/gap-closure-artifacts.md", diff --git a/tests/safe-resume-gate-anchoring.test.cjs b/tests/safe-resume-gate-anchoring.test.cjs index 3426c8da1..3ca931342 100644 --- a/tests/safe-resume-gate-anchoring.test.cjs +++ b/tests/safe-resume-gate-anchoring.test.cjs @@ -78,12 +78,19 @@ describe('#4003 — safe_resume_gate commit-scope greps', () => { }); test('completion spot-check uses the anchored scope and keeps its time bound', () => { + // #4217 moved the completion spot-check probes (with the whole reconciliation + // policy, both arms) into execute-phase/steps/completion-reconciliation.md — + // "extract, not bump" against the frozen host ceiling. The anchoring contract + // travels with them: negative shape against the host, positives against the + // fragment that now owns the probes. const w = fs.readFileSync(WORKFLOW, 'utf8'); - assert.ok(!w.includes('--grep="{phase_number}-{plan_padded}"'), + const frag = fs.readFileSync(path.join(__dirname, '..', 'gsd-core', 'workflows', + 'execute-phase', 'steps', 'completion-reconciliation.md'), 'utf8'); + assert.ok(!w.includes('--grep="{phase_number}-{plan_padded}"') && !frag.includes('--grep="{phase_number}-{plan_padded}"'), 'the raw padded placeholder substring grep must not remain'); - assert.ok(w.includes('SPOT_PHASE_N=$((10#{phase_number}))') && w.includes('SPOT_PLAN_N=$((10#{plan_padded}))'), + assert.ok(frag.includes('SPOT_PHASE_N=$((10#{phase_number}))') && frag.includes('SPOT_PLAN_N=$((10#{plan_padded}))'), 'the spot-check derives zero-stripped components'); - assert.ok(w.includes('--since="1 hour ago"'), 'the spot-check keeps its temporal bound'); + assert.ok(frag.includes('--since="1 hour ago"'), 'the spot-check keeps its temporal bound'); }); test('the gate pipeline separates same-scope commits across a milestone tag (behavioral)', (t) => { From e54d3aa159810b2308cd777047b9de9c04de418a Mon Sep 17 00:00:00 2001 From: Brenden Smerbeck Date: Sun, 6 Sep 2026 19:52:59 -0400 Subject: [PATCH 027/166] enhance(#4401): register workflow.compact_content as a validated config key (#4441) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(#4401): register workflow.compact_content as a validated config key - Add compact_content: false to the nested workflow object in gsd-core/bin/shared/config-defaults.manifest.json - Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts so an absent key resolves to false via config-get --raw - validKeys entry in config-schema.manifest.json already present Co-Authored-By: Claude Fable 5.1 * test(#4401): behavioral and boundary tests for workflow.compact_content - 19 behavioral tests covering config-set/config-get round trip, invalid-shape rejection (banana, 42, empty string), the corrected null-unset semantics (#2046), absent-key resolution against config-defaults.manifest.json, config-new-project wiring, and doc-row shape assertions - Drops the install-tree fixture-parity block (and its docstring item) that asserted gsd-core/references/compact-content-gate.md and gsd-core/workflows/compact/map-codebase.md fixture entries — those paths belong to #4402 and do not exist on this filtered branch Co-Authored-By: Claude Fable 5.1 * docs(#4401): document workflow.compact_content in both config references - One 4-cell row in docs/CONFIGURATION.md (workflow.* run) - One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md - Both cross-reference ADR-4139 Co-Authored-By: Claude Fable 5.1 * chore(#4401): add changeset - Added-type fragment, pr: 4401 (issue number; backfill to the real PR number is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR- FIELD-DRIFT) Co-Authored-By: Claude Fable 5.1 * chore(#4401): backfill changeset pr field to #4441 Co-Authored-By: Claude Fable 5.1 * fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving a single-source-of-truth drift risk: a future manifest-only edit to the default could silently diverge from this literal, only caught later by the D-03 test if it ever happened to manifest. Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority sibling pattern. Found during maintainer review (review-open-prs) of this PR. Co-Authored-By: Claude Sonnet 5 * fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP The previous commit added compact_content to CONFIG_DEFAULTS in src/config-loader.cts but missed the matching entry in tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat CONFIG_DEFAULTS keys to their namespaced doc form before checking gsd-core/references/planning-config.md for a match. Without it, the test looked for a bare `compact_content` doc reference instead of the actual `workflow.compact_content` row, and failed: "CONFIG_DEFAULTS keys missing from planning-config.md: compact_content". Found by actually running gsd-test against the branch rather than trusting the plausible-looking fix. Co-Authored-By: Claude Sonnet 5 * test(#4401): register compact-content-4139 test in the docs-guard lane tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md directly (fs.readFileSync) to assert the workflow.compact_content doc row's shape, which makes it a doc-reading test file under the #3753 docs-guard lane. It was never added to scripts/docs-guard-registry.cjs's DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed: "compact-content-4139.test.cjs reads a docs/ path but is not registered in the docs-guard lane and carries no docs-guard-exempt marker". Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed path it reads; gsd-core/references/planning-config.md is outside this registry's docs/ scope, matching the sibling config-field-docs.test.cjs entry's existing convention). Found by actually running gsd-test against the branch — this gap predates the maintainer's config-loader.cts fix and was already present in the original PR. Co-Authored-By: Claude Sonnet 5 --------- Co-authored-by: Claude Fable 5.1 Co-authored-by: Tom Boucher Co-authored-by: sim --- .changeset/steady-elks-parade.md | 5 + docs/CONFIGURATION.md | 1 + .../bin/shared/config-defaults.manifest.json | 1 + .../bin/shared/config-schema.manifest.json | 1 + gsd-core/references/planning-config.md | 1 + scripts/docs-guard-registry.cjs | 1 + src/config-loader.cts | 1 + src/config.cts | 13 + tests/compact-content-4139.test.cjs | 306 ++++++++++++++++++ tests/config-field-docs.test.cjs | 1 + 10 files changed, 331 insertions(+) create mode 100644 .changeset/steady-elks-parade.md create mode 100644 tests/compact-content-4139.test.cjs diff --git a/.changeset/steady-elks-parade.md b/.changeset/steady-elks-parade.md new file mode 100644 index 000000000..5291b5733 --- /dev/null +++ b/.changeset/steady-elks-parade.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 4441 +--- +**`workflow.compact_content` is now a registered, validated, documented project config key.** It resolves to `false` when absent and is readable via `config-get`; no content branches on it yet. (#4401) diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index 9b098536f..5cc2f6eb2 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -517,6 +517,7 @@ All workflow toggles follow the **absent = enabled** pattern. If a key is missin | `workflow.text_mode` | boolean | `false` | Replaces AskUserQuestion TUI menus with plain-text numbered lists. Required for Claude Code remote sessions (`/rc` mode) where TUI menus don't render. Can also be set per-session with `--text` flag on discuss-phase. Added in v1.28 | | `workflow.use_worktrees` | boolean | `true` | When `false`, disables git worktree isolation for parallel execution. Users who prefer sequential execution or whose environment does not support worktrees can disable this. Added in v1.31. **Branch-divergence note:** when your branch has diverged from `origin/HEAD`, GSD auto-degrades to sequential and prints a warning. See [`worktree.baseRef`](#worktree-settings) to restore parallel execution on a diverged branch. **Per-runtime note:** whether this key can be honored depends on the runtime's declared `dispatch.isolation` capability, not on its name (#2584). Runtimes whose own harness isolates each executor (**Claude Code**, **Cursor**) run parallel worktrees natively; runtimes exposing a headless exec with an explicit working directory (**Codex**, **OpenCode**, **Kimi**, **Kimi Code**) get worktrees GSD itself creates and merges — where a dispatch site can only drive the harness model, those hosts degrade to sequential with a warning rather than aborting. Every other runtime declares no isolation primitive, and forcing `use_worktrees: true` there still fails closed before any executor dispatch. `/gsd-health` reports such a value as warning `W025` (#2486). **Default on a non-Claude install:** if a worktree-capable non-Claude host is not isolating as described above, check whether the install stamped this key's default to `false` and set an explicit `use_worktrees: true`. See [Executor isolation per runtime](#executor-isolation-per-runtime). | | `workflow.agent_hint_routing` | boolean | `true` | Per-plan specialist executor routing (#1689). When `true`, a plan whose `agent_hint:` frontmatter names a subagent that resolves on the active runtime is dispatched to that specialist instead of `gsd-executor`. Default `true` — a no-op for plans without `agent_hint:`, so existing dispatch is unchanged. Set `false` to disable. See [PLAN.md `agent_hint`](reference/plan-md.md#per-plan-executor-routing). | +| `workflow.compact_content` | boolean | `false` | Compact content mode (#4139, [ADR-4139](adr/4139-compact-content-seam.md)). Per-project boolean selecting the terser form of GSD's own shipped prompt content (workflows, templates, agent-skill payloads). The key is registered and readable today; nothing branches on it yet — the load mechanism is a later sub-issue. | | `workflow.worktree_skip_hooks` | boolean | `false` | When `true`, executor agents in worktree mode pass `--no-verify` (skipping pre-commit hooks) and post-wave hook validation runs against the merged result instead. Opt-in escape hatch for projects whose hooks cannot run in agent worktrees. Default `false` runs hooks on every commit (#2924). | | `workflow.code_review` | boolean | `true` | Enable `/gsd-code-review` and `/gsd-code-review --fix` commands. When `false`, the commands exit with a configuration gate message. Added in v1.34 | | `workflow.code_review_point` | string | `execute:post` | Loop point at which the code-review capability's step registers: `execute:post` reviews once, after every wave in a phase has landed (default — unchanged behavior); `execute:wave:post` reviews once per completed wave instead, scoped to what changed since the phase's prior review (the whole phase's diff on the first wave, each subsequent wave's own diff thereafter). Manual `/gsd-code-review ` invocation is unaffected by this key — it is gated by `workflow.code_review` alone and runs regardless of which point is configured. `/gsd-autonomous` and `/gsd-quick` have no wave granularity of their own, so setting this to `execute:wave:post` means code review does not run automatically inside those two flows (consistent with how every other `execute:wave:post`-only capability already behaves for them). Added in #3661 | diff --git a/gsd-core/bin/shared/config-defaults.manifest.json b/gsd-core/bin/shared/config-defaults.manifest.json index 164b4ac5d..1b26ab4b7 100644 --- a/gsd-core/bin/shared/config-defaults.manifest.json +++ b/gsd-core/bin/shared/config-defaults.manifest.json @@ -38,6 +38,7 @@ "ui_phase": true, "ui_safety_gate": true, "text_mode": false, + "compact_content": false, "research_before_questions": false, "discuss_mode": "discuss", "skip_discuss": false, diff --git a/gsd-core/bin/shared/config-schema.manifest.json b/gsd-core/bin/shared/config-schema.manifest.json index 393b23105..0212ba304 100644 --- a/gsd-core/bin/shared/config-schema.manifest.json +++ b/gsd-core/bin/shared/config-schema.manifest.json @@ -22,6 +22,7 @@ "workflow.smart_zone_tokens", "workflow.human_verify_mode", "workflow.text_mode", + "workflow.compact_content", "workflow.research_before_questions", "workflow.discuss_mode", "workflow.skip_discuss", diff --git a/gsd-core/references/planning-config.md b/gsd-core/references/planning-config.md index 70d807710..32dae8e22 100644 --- a/gsd-core/references/planning-config.md +++ b/gsd-core/references/planning-config.md @@ -289,6 +289,7 @@ Set via `workflow.*` namespace in config.json (e.g., `"workflow": { "research": | `workflow.ui_phase` | boolean | `true` | `true`, `false` | Generate UI-SPEC.md for frontend phases | | `workflow.ui_safety_gate` | boolean | `true` | `true`, `false` | Require safety gate approval for UI changes | | `workflow.text_mode` | boolean | `false` | `true`, `false` | Use plain-text numbered lists instead of AskUserQuestion menus | +| `workflow.compact_content` | boolean | `false` | `true`, `false` | Compact content mode (#4139, ADR-4139) — per-project boolean selecting terser payloads; nothing branches on it yet | | `workflow.research_before_questions` | boolean | `false` | `true`, `false` | Run research before interactive questions in discuss phase (also honored on the `/gsd:quick` path, #3894). _Alias:_ `research_before_questions` is the flat-key form used in `CONFIG_DEFAULTS`; `workflow.research_before_questions` is the canonical namespaced form. | | `workflow.discuss_mode` | string | `"discuss"` | `"discuss"`, `"assumptions"` | Default mode for discuss-phase: `"discuss"` runs interactive questioning; `"assumptions"` analyzes codebase and surfaces assumptions instead | | `workflow.skip_discuss` | boolean | `false` | `true`, `false` | Skip discuss phase entirely | diff --git a/scripts/docs-guard-registry.cjs b/scripts/docs-guard-registry.cjs index ed7efbf09..a1f206e30 100644 --- a/scripts/docs-guard-registry.cjs +++ b/scripts/docs-guard-registry.cjs @@ -193,6 +193,7 @@ const DOCS_GUARD_TESTS = { // (commit-files-pathspec.test.cjs:1618) — cannot be resolved to specific // files without re-deriving the scan's own file-discovery logic. 'tests/commit-files-pathspec.test.cjs': ['*'], + 'tests/compact-content-4139.test.cjs': ['docs/CONFIGURATION.md'], 'tests/config-field-docs.test.cjs': ['docs/CONFIGURATION.md'], 'tests/config.test.cjs': ['docs/CONFIGURATION.md'], 'tests/context-index-sync.test.cjs': ['docs/CONTEXT-INDEX.json'], diff --git a/src/config-loader.cts b/src/config-loader.cts index fdafaaee3..3674e48bd 100644 --- a/src/config-loader.cts +++ b/src/config-loader.cts @@ -144,6 +144,7 @@ const CONFIG_DEFAULTS = { firecrawl: _getConfigDefault('firecrawl'), exa_search: _getConfigDefault('exa_search'), text_mode: _getNestedConfigDefault('workflow', 'text_mode'), + compact_content: _getNestedConfigDefault('workflow', 'compact_content'), sub_repos: _getNestedConfigDefault('planning', 'sub_repos'), pr_strict: _getNestedConfigDefault('planning', 'pr_strict'), resolve_model_ids: _getConfigDefault('resolve_model_ids'), diff --git a/src/config.cts b/src/config.cts index 49ee4d12f..1a746074a 100644 --- a/src/config.cts +++ b/src/config.cts @@ -104,6 +104,11 @@ const SCHEMA_DEFAULTS: Record = { // #1689: per-plan agent_hint executor routing — default-on. A no-op for plans // without an agent_hint field, so existing dispatch is byte-identical. 'workflow.agent_hint_routing': true, + // #4401: Compact Content mode gate — derived from the defaults manifest via + // CONFIG_DEFAULTS (added in config-loader.cts) so the manifest stays the + // single source of truth, matching workflow.smart_zone_tokens / + // planning.pr_strict / workflow.inline_plan_threshold below. + 'workflow.compact_content': CONFIG_DEFAULTS.compact_content, // Derived from the defaults manifest rather than restated, so the manifest // stays the single source of truth for the smart-zone budget (#2630). 'workflow.smart_zone_tokens': CONFIG_DEFAULTS.smart_zone_tokens, @@ -345,6 +350,7 @@ function buildNewProjectConfig(userChoices: Record): Record { + let tmpDir; + + beforeEach(() => { + tmpDir = createTempProject(); + runGsdTools('config-ensure-section', tmpDir); + }); + afterEach(() => { cleanup(tmpDir); }); + + function readConfig() { + return JSON.parse(fs.readFileSync(path.join(tmpDir, '.planning', 'config.json'), 'utf-8')); + } + + test('config-set workflow.compact_content true → persisted as boolean true', () => { + const r = runGsdTools(['config-set', 'workflow.compact_content', 'true'], tmpDir); + assert.ok(r.success, r.error); + const config = readConfig(); + assert.strictEqual(config.workflow.compact_content, true); + }); + + test('config-set workflow.compact_content false → persisted as boolean false', () => { + const r = runGsdTools(['config-set', 'workflow.compact_content', 'false'], tmpDir); + assert.ok(r.success, r.error); + const config = readConfig(); + assert.strictEqual(config.workflow.compact_content, false); + }); + + test('config-set workflow.compact_content banana → rejected', () => { + const r = runGsdTools(['config-set', 'workflow.compact_content', 'banana'], tmpDir); + assert.ok(!r.success, 'non-boolean value must be rejected'); + assert.match(r.error || r.output, /boolean|true|false/i); + }); + + test('config-set workflow.compact_content "" → rejected', () => { + const r = runGsdTools(['config-set', 'workflow.compact_content', ''], tmpDir); + assert.ok(!r.success, 'empty value must be rejected'); + }); + + test('config-get workflow.compact_content --raw before any explicit set → succeeds with the materialized default', () => { + // As of plan 02-02 (CONF-01), buildNewProjectConfig's hardcoded workflow + // object carries compact_content: false, so any freshly materialized + // config.json (including the one config-ensure-section writes in + // beforeEach) already has the key — config-get succeeds and returns the + // default "false" rather than exiting non-zero. The workflow-side + // `... --raw 2>/dev/null || echo "false"` fallback still resolves to the + // same string either way, so gate hooks are unaffected by this change. + const r = runGsdTools(['config-get', 'workflow.compact_content', '--raw'], tmpDir); + assert.ok(r.success, 'config-get on the materialized default must exit zero'); + assert.strictEqual(r.output.trim(), 'false'); + }); + + test('setting true twice is idempotent — identical config.json content', () => { + const first = runGsdTools(['config-set', 'workflow.compact_content', 'true'], tmpDir); + assert.ok(first.success, first.error); + const afterFirst = fs.readFileSync(path.join(tmpDir, '.planning', 'config.json'), 'utf-8'); + + const second = runGsdTools(['config-set', 'workflow.compact_content', 'true'], tmpDir); + assert.ok(second.success, second.error); + const afterSecond = fs.readFileSync(path.join(tmpDir, '.planning', 'config.json'), 'utf-8'); + + assert.strictEqual(afterSecond, afterFirst); + }); + + test('setting the key preserves every other pre-existing key/value', () => { + const cfgPath = path.join(tmpDir, '.planning', 'config.json'); + const before = JSON.parse(fs.readFileSync(cfgPath, 'utf-8')); + before.workflow.text_mode = true; + before.mode = 'yolo'; + fs.writeFileSync(cfgPath, JSON.stringify(before, null, 2)); + + const r = runGsdTools(['config-set', 'workflow.compact_content', 'true'], tmpDir); + assert.ok(r.success, r.error); + + const after = readConfig(); + assert.strictEqual(after.workflow.text_mode, true, 'workflow.text_mode must survive unrelated key write'); + assert.strictEqual(after.mode, 'yolo', 'top-level mode must survive unrelated key write'); + assert.strictEqual(after.workflow.compact_content, true); + }); + + test('VALID_CONFIG_KEYS has workflow.compact_content', () => { + const { VALID_CONFIG_KEYS } = require('../gsd-core/bin/lib/config-schema.cjs'); + assert.strictEqual(VALID_CONFIG_KEYS.has('workflow.compact_content'), true); + }); + + test('config-set workflow.compact_content 42 → rejected, message names the key', () => { + const r = runGsdTools(['config-set', 'workflow.compact_content', '42'], tmpDir); + assert.ok(!r.success, 'numeric value must be rejected'); + assert.match(r.error || r.output, /workflow\.compact_content/); + }); + + test('config-set workflow.compact_content null → unsets the key (universal #2046 clear semantics, not a type-rejection)', () => { + // A bare `null` is the documented "clear this key" shortcut (#2046) and is + // short-circuited before every typed per-key validator runs — this is + // true for every config key, not something this plan introduces or may + // change. Verified against the analogous git.protected_branches and + // context-key coverage in tests/config.test.cjs ("config-set null — + // unset/clear (#2046)"). So `null` exits zero and removes the key rather + // than being rejected like `42`/`banana`/`""`. + const r = runGsdTools(['config-set', 'workflow.compact_content', 'null'], tmpDir); + assert.ok(r.success, `unset must succeed: ${r.error}`); + const config = readConfig(); + assert.ok( + !Object.prototype.hasOwnProperty.call(config.workflow, 'compact_content'), + 'workflow.compact_content must be absent after unset', + ); + }); + + test('config-defaults.manifest.json carries workflow.compact_content', () => { + const manifest = require('../gsd-core/bin/shared/config-defaults.manifest.json'); + assert.strictEqual(manifest.workflow.compact_content, false); + }); +}); + +// ─── D-03: absent-key resolution against a config that omits the key ───────── + +describe('workflow.compact_content absent-key resolution (#4139, D-03)', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = createTempProject(); + }); + afterEach(() => { cleanup(tmpDir); }); + + test('absent key: a config.json omitting compact_content resolves to the manifest default', () => { + const cfgDir = path.join(tmpDir, '.planning'); + fs.mkdirSync(cfgDir, { recursive: true }); + fs.writeFileSync( + path.join(cfgDir, 'config.json'), + JSON.stringify({ version: '1.0', mode: 'interactive', workflow: { research: true } }, null, 2), + ); + + const manifest = require('../gsd-core/bin/shared/config-defaults.manifest.json'); + const r = runGsdTools(['config-get', 'workflow.compact_content', '--raw'], tmpDir); + assert.strictEqual(r.exitCode, 0, r.error || r.output); + assert.strictEqual(r.output.trim(), String(manifest.workflow.compact_content)); + }); +}); + +// ─── CONF-01: buildNewProjectConfig default + config-new-project wiring ─────── + +describe('workflow.compact_content via config-new-project (#4139, CONF-01)', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = createTempProject(); + }); + afterEach(() => { cleanup(tmpDir); }); + + function readConfig() { + return JSON.parse(fs.readFileSync(path.join(tmpDir, '.planning', 'config.json'), 'utf-8')); + } + + test('config-new-project omitting compact_content → hardcoded default false lands in config.json', () => { + const choices = JSON.stringify({ + mode: 'interactive', + granularity: 'coarse', + parallelization: true, + commit_docs: false, + model_profile: 'adaptive', + workflow: { research: true, plan_check: true, verifier: true, nyquist_validation: false }, + }); + const r = runGsdTools(['config-new-project', choices], tmpDir, { HOME: tmpDir, USERPROFILE: tmpDir }); + assert.ok(r.success, r.error); + const config = readConfig(); + assert.strictEqual(config.workflow.compact_content, false); + }); + + test('config-new-project with compact_content: true → persisted as boolean true', () => { + const choices = JSON.stringify({ + mode: 'interactive', + granularity: 'coarse', + parallelization: true, + commit_docs: false, + model_profile: 'adaptive', + workflow: { research: true, plan_check: true, verifier: true, nyquist_validation: false, compact_content: true }, + }); + const r = runGsdTools(['config-new-project', choices], tmpDir, { HOME: tmpDir, USERPROFILE: tmpDir }); + assert.ok(r.success, r.error); + const config = readConfig(); + assert.strictEqual(config.workflow.compact_content, true); + }); + + test('config-new-project with compact_content: false → persisted as boolean false, not string', () => { + const choices = JSON.stringify({ + mode: 'interactive', + granularity: 'coarse', + parallelization: true, + commit_docs: false, + model_profile: 'adaptive', + workflow: { research: true, plan_check: true, verifier: true, nyquist_validation: false, compact_content: false }, + }); + const r = runGsdTools(['config-new-project', choices], tmpDir, { HOME: tmpDir, USERPROFILE: tmpDir }); + assert.ok(r.success, r.error); + const config = readConfig(); + assert.strictEqual(config.workflow.compact_content, false); + assert.notStrictEqual(config.workflow.compact_content, 'false'); + }); + + test('config-get workflow.compact_content --raw after config-new-project prints the persisted value', () => { + const choices = JSON.stringify({ + mode: 'interactive', + granularity: 'coarse', + parallelization: true, + commit_docs: false, + model_profile: 'adaptive', + workflow: { research: true, plan_check: true, verifier: true, nyquist_validation: false, compact_content: true }, + }); + const setup = runGsdTools(['config-new-project', choices], tmpDir, { HOME: tmpDir, USERPROFILE: tmpDir }); + assert.ok(setup.success, setup.error); + + const r = runGsdTools(['config-get', 'workflow.compact_content', '--raw'], tmpDir); + assert.ok(r.success, 'config-get must exit zero once config-new-project has materialized the key'); + assert.strictEqual(r.output.trim(), 'true'); + }); +}); + +// ─── D-06: doc-row shape assertions for both config reference tables ───────── + +describe('workflow.compact_content documentation rows (#4139, D-06)', () => { + const { splitTableRow } = require('../gsd-core/bin/lib/markdown-table.cjs'); + const KEY_CELL = '`workflow.compact_content`'; + + function findRow(filePath) { + const lines = fs.readFileSync(path.join(__dirname, '..', filePath), 'utf-8').split(/\r?\n/); + for (const line of lines) { + if (!line.trim().startsWith('|')) continue; + const cells = splitTableRow(line); + if (cells && cells[0] === KEY_CELL) return cells; + } + return undefined; + } + + test('docs/CONFIGURATION.md documents workflow.compact_content as a 4-cell boolean row', () => { + const cells = findRow('docs/CONFIGURATION.md'); + assert.ok(cells, 'workflow.compact_content row not found in docs/CONFIGURATION.md'); + assert.strictEqual(cells.length, 4); + assert.strictEqual(cells[1], 'boolean'); + assert.strictEqual(cells[2], '`false`'); + }); + + test('planning-config.md documents workflow.compact_content as a 5-cell boolean row', () => { + const cells = findRow('gsd-core/references/planning-config.md'); + assert.ok(cells, 'workflow.compact_content row not found in planning-config.md'); + assert.strictEqual(cells.length, 5); + assert.strictEqual(cells[1], 'boolean'); + assert.strictEqual(cells[2], '`false`'); + assert.match(cells[3], /`true`/); + assert.match(cells[3], /`false`/); + }); + + test('both doc rows sit under the Workflow section they belong to', () => { + const planningConfigPath = path.join(__dirname, '..', 'gsd-core/references/planning-config.md'); + const planningConfigContent = fs.readFileSync(planningConfigPath, 'utf-8'); + const keyIdx = planningConfigContent.indexOf('`workflow.compact_content`'); + const fieldRefIdx = planningConfigContent.indexOf('## Complete Field Reference'); + const workflowFieldsIdx = planningConfigContent.indexOf('### Workflow Fields'); + assert.ok(keyIdx > -1, 'key not found in planning-config.md'); + assert.ok(keyIdx > fieldRefIdx, 'row must sit after ## Complete Field Reference heading'); + assert.ok(keyIdx > workflowFieldsIdx, 'row must sit after ### Workflow Fields heading'); + + const configurationMdPath = path.join(__dirname, '..', 'docs/CONFIGURATION.md'); + const configurationMdContent = fs.readFileSync(configurationMdPath, 'utf-8'); + const compactIdx = configurationMdContent.indexOf('`workflow.compact_content`'); + const textModeIdx = configurationMdContent.indexOf('`workflow.text_mode`'); + assert.ok(compactIdx > -1, 'key not found in docs/CONFIGURATION.md'); + assert.ok(compactIdx > textModeIdx, 'row must sit inside the workflow.* run, after workflow.text_mode'); + }); +}); diff --git a/tests/config-field-docs.test.cjs b/tests/config-field-docs.test.cjs index 9f516e274..23ecca5be 100644 --- a/tests/config-field-docs.test.cjs +++ b/tests/config-field-docs.test.cjs @@ -101,6 +101,7 @@ describe('config-field-docs', () => { ai_integration_phase: 'workflow.ai_integration_phase', api_coverage_gate: 'workflow.api_coverage_gate', text_mode: 'workflow.text_mode', + compact_content: 'workflow.compact_content', subagent_timeout: 'workflow.subagent_timeout', branching_strategy: 'git.branching_strategy', phase_branch_template: 'git.phase_branch_template', From ef30e59860ecf3d9c536a1cf6c1d96f65677330d Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 21:06:48 -0400 Subject: [PATCH 028/166] fix(#4448): stop io.test.cjs's in-process runMain calls from corrupting node:test's own fd-1 IPC (#4452) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit tests/io.test.cjs's "review fix: pending-outcome cell lifetime" describe block deliberately drives runMain()/output() in-process (needed to observe a cross-invocation state leak) instead of via a subprocess. output() ends with a raw synchronous fs.writeSync(1, ...) to the real stdout fd, and Node's --test-isolation=process (default since Node 22) uses that same fd for the file's own parent-child reporter protocol. The two writes racing produced an intermittent "Unable to deserialize cloned data" that killed the whole file — observed twice on next's macOS lane, most recently on the commit that merged PR #4428 (unrelated to that PR's content; io.test.cjs isn't part of its diff). Empirically validated locally (gsd-test can't reach macOS): built a repro loop running N parallel copies of `node --test tests/io.test.cjs` to recreate CI-like contention. Baseline: ~13-17% of runs hit the corruption (12/90, 15/90 across two samples). A first fix attempt wrapped the writes in captureFdAsync (an await-aware twin of the existing captureFdSync, added because runMain() defers main() through a microtask chain, so a synchronous wrap restores before the real write fires) — but captureFdAsync always forwards to the real fs.writeSync by design (matching captureFdSync's "never swallow" contract, tests/helpers.cjs, #4306). Re-ran the same loop against that fix: 15/90, statistically unchanged. Forwarding the write doesn't stop it from reaching the fd node:test's own IPC also uses. Replaced it with suppressFdAsync: a narrow, deliberate exception to the never-swallow contract for a window the caller has verified is fully controlled (a single runMain() call plus its promise-chain settling, where nothing else can legitimately need that fd). It records the bytes for the test's own assertions but never lets them reach the real fd. Re-ran the loop: 0/300 across three samples (90+120+90), including one round at 8-way parallelism. Also caught and fixed a real bug surfaced by the same loop: the new regression assertion checked for compact-JSON `"error":"x"` but output() pretty-prints, so it failed 100% of runs deterministically until fixed to parse and check the structured value instead (io.test.cjs, matching this repo's "assert on structured output, not raw text" convention). Co-authored-by: sim Co-authored-by: Claude Sonnet 5 --- tests/helpers.cjs | 70 +++++++++++++++++++++++++++++++++++++++- tests/io.test.cjs | 81 +++++++++++++++++++++++++++++++++-------------- 2 files changed, 127 insertions(+), 24 deletions(-) diff --git a/tests/helpers.cjs b/tests/helpers.cjs index 9ffe7382d..6f54da5a3 100644 --- a/tests/helpers.cjs +++ b/tests/helpers.cjs @@ -640,6 +640,74 @@ function captureFdSync(captureFd, fn) { return Buffer.concat(chunks).toString('utf8'); } +/** + * NARROW, deliberate exception to `captureFdSync`'s + * "never swallow" philosophy (#4306) — do NOT reach for this casually. + * + * Async twin of `captureFdSync` that awaits `fn()` before restoring the + * patch, so a deferred write that happens after a microtask/macrotask + * boundary is still safely observed. Its patched `fs.writeSync` + * does NOT forward `captureFd`'s writes to the real `fs.writeSync` at all. + * It only records the bytes into `chunks` and returns the byte length of + * `data` as if the real syscall had succeeded, so a caller that inspects the + * return value sees a normal success and not an error. Every OTHER fd's + * writes still forward to the real `fs.writeSync` exactly as + * `captureFdSync` does — only the `captureFd`-matching branch + * differs. + * + * This exists for #4448: `runMain()`/`io.output()`'s deferred `fs.writeSync(1, + * ...)` races Node's own `node:test` child-to-parent IPC, which also uses fd + * 1 under the default `--test-isolation=process` — corrupting the parent's + * message parsing ("Unable to deserialize cloned data"). An always-forward + * capture (the first fix attempted for this issue) does not + * fix that: the corrupting write still physically reaches fd 1. Use this + * ONLY for a window the caller has verified is narrow and fully controlled — + * i.e. nothing else legitimately needs to write to `captureFd` during `fn()` + * — such as a single `runMain(...)` call plus its promise-chain settling. + * Reaching for this in a window where something else might legitimately + * write to `captureFd` will silently swallow that other write. + * + * @param {number} captureFd - the fd whose writes are suppressed and recorded + * (never delivered to the real fd) while `fn()` runs. + * @param {() => (Promise | void)} fn - function to run (and await) while suppressing. + * @returns {Promise} every byte that WOULD have been written to + * `captureFd` during `fn()`, joined as UTF-8 — none of it actually reached + * the real fd. + */ +async function suppressFdAsync(captureFd, fn) { + const chunks = []; + const orig = fs.writeSync; + fs.writeSync = (fd, data, ...rest) => { + if (fd === captureFd) { + const offset = Buffer.isBuffer(data) && typeof rest[0] === 'number' ? rest[0] : 0; + // No real syscall happens here (unlike captureFdSync, which slices to + // the real return value `n`), so the caller-requested `length` IS the + // count that must be recorded and returned — suppression always + // "succeeds" in full, so anything else silently drops or over-reports + // bytes. + const length = Buffer.isBuffer(data) && typeof rest[1] === 'number' ? rest[1] : undefined; + const buf = Buffer.isBuffer(data) + ? data.subarray(offset, length === undefined ? undefined : offset + length) + : Buffer.from(String(data), 'utf8'); + // Buffered, not decoded per-call: see captureFdSync's identical note on + // why joined-then-decoded avoids splitting a multi-byte UTF-8 codepoint. + chunks.push(buf); + // No real fs.writeSync call for this fd — that is the entire point of + // this helper. Return the byte length as if the write succeeded, so a + // caller inspecting the return value (Node's own writeSync contract) + // sees ordinary success rather than an error. + return buf.length; +} + return orig.call(fs, fd, data, ...rest); +}; + try { + await fn(); + } finally { + fs.writeSync = orig; +} + return Buffer.concat(chunks).toString('utf8'); +} + /** * Read a workflow .md file plus every .md file under its sibling * `/steps/` directory, concatenated in document order @@ -1182,7 +1250,7 @@ function writePackageSourceMarkerFixture(configDir) { return configDir; } -module.exports = { runGsdTools, createTempDir, createTempProject, createTempGitProject, cleanup, tmpRootCandidates, readFileNormalized, readWorkflowCombined, parseFrontmatter, isUsageOutput, captureConsole, toPosixPath, absPlanningPath, runNpm, isolatedNpmEnv, withIsolatedProcessState, delay, waitFor, resetRuntimeWarningCaches, SESSION_ENV_KEYS, saveSessionEnv, restoreSessionEnv, clearSessionEnv, isolateWorkstreamEnv, restoreWorkstreamEnv, TOOLS_PATH, SESSION_IDENTITY_ENV_KEYS, scrubConfigLocationEnv, installSpawnEnv, installSpawnHome, sandboxHome, writePackageSourceMarkerFixture, TEST_HOME_SANDBOX_MARKER, mockPartialWriteThenThrow, captureFdSync }; +module.exports = { runGsdTools, createTempDir, createTempProject, createTempGitProject, cleanup, tmpRootCandidates, readFileNormalized, readWorkflowCombined, parseFrontmatter, isUsageOutput, captureConsole, toPosixPath, absPlanningPath, runNpm, isolatedNpmEnv, withIsolatedProcessState, delay, waitFor, resetRuntimeWarningCaches, SESSION_ENV_KEYS, saveSessionEnv, restoreSessionEnv, clearSessionEnv, isolateWorkstreamEnv, restoreWorkstreamEnv, TOOLS_PATH, SESSION_IDENTITY_ENV_KEYS, scrubConfigLocationEnv, installSpawnEnv, installSpawnHome, sandboxHome, writePackageSourceMarkerFixture, TEST_HOME_SANDBOX_MARKER, mockPartialWriteThenThrow, captureFdSync, suppressFdAsync }; // Lazy, for the reason builtLib() is lazy: reading either of these is what // forces the built-lib require, so a test file that needs neither can still diff --git a/tests/io.test.cjs b/tests/io.test.cjs index 0e446956f..07aba6a0f 100644 --- a/tests/io.test.cjs +++ b/tests/io.test.cjs @@ -16,7 +16,7 @@ const assert = require('node:assert/strict'); const path = require('node:path'); const os = require('node:os'); const fs = require('node:fs'); -const { captureFdSync } = require('./helpers.cjs'); +const { captureFdSync, suppressFdAsync } = require('./helpers.cjs'); const io = require('../gsd-core/bin/lib/io.cjs'); const { @@ -717,7 +717,7 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 test(`v1: ERROR_REASON.${key} (${reasonValue}) exits 1`, () => { resolveContractVersion({ argv: ['node', 'x'], env: {} }); // v1 assert.throws( - () => io.error('msg', reasonValue), + () => captureFdSync(2, () => { io.error('msg', reasonValue); }), (err) => err instanceof ExitError && err.code === 1, `ERROR_REASON.${key} must exit 1 under v1`, ); @@ -727,7 +727,7 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); const expected = expectedErrorCode3912(reasonValue, 'v2'); assert.throws( - () => io.error('msg', reasonValue), + () => captureFdSync(2, () => { io.error('msg', reasonValue); }), (err) => err instanceof ExitError && err.code === expected, `ERROR_REASON.${key} under v2 must exit ${expected}`, ); @@ -738,7 +738,7 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 test('A2: error() with no reason argument exits 1 under v1 (defaults to UNKNOWN)', () => { resolveContractVersion({ argv: ['node', 'x'], env: {} }); assert.throws( - () => io.error('no reason given'), + () => captureFdSync(2, () => { io.error('no reason given'); }), (err) => err instanceof ExitError && err.code === 1, ); }); @@ -746,7 +746,7 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 test('A2: error() with no reason argument stays FAIL (exit 1) under v2 too — UNKNOWN is not a specific outcome', () => { resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); assert.throws( - () => io.error('no reason given'), + () => captureFdSync(2, () => { io.error('no reason given'); }), (err) => err instanceof ExitError && err.code === 1, ); }); @@ -756,7 +756,7 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); for (const key of ['SDK_MISSING_ARG', 'SDK_UNKNOWN_COMMAND', 'USAGE']) { assert.throws( - () => io.error('msg', io.ERROR_REASON[key]), + () => captureFdSync(2, () => { io.error('msg', io.ERROR_REASON[key]); }), (err) => err instanceof ExitError && err.code === 64, `${key} must project to 64 under v2`, ); @@ -768,10 +768,14 @@ describe('#3912 A1/B1: error() declares from ERROR_REASON, exhaustive over the 2 test('B5 (anti-vacuity): v1 and v2 differ for at least one reason', () => { resolveContractVersion({ argv: ['node', 'x'], env: {} }); let v1Code; - try { io.error('msg', io.ERROR_REASON.SDK_MISSING_ARG); } catch (e) { v1Code = e.code; } + captureFdSync(2, () => { + try { io.error('msg', io.ERROR_REASON.SDK_MISSING_ARG); } catch (e) { v1Code = e.code; } + }); resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); let v2Code; - try { io.error('msg', io.ERROR_REASON.SDK_MISSING_ARG); } catch (e) { v2Code = e.code; } + captureFdSync(2, () => { + try { io.error('msg', io.ERROR_REASON.SDK_MISSING_ARG); } catch (e) { v2Code = e.code; } + }); assert.equal(v1Code, 1); assert.equal(v2Code, 64); assert.notEqual(v1Code, v2Code, 'v1 and v2 must differ for at least one reason, or the declaration is decorative'); @@ -990,11 +994,22 @@ describe('review fix: pending-outcome cell lifetime (last-write-wins, cleared on try { // First invocation declares DEGRADED via a payload-carried error and // returns nothing — runMain projects it to 80 under v2. - runMain(() => { - io.output({ found: false, error: 'not found' }, false); - return undefined; + // + // runMain() defers main() via Promise.resolve().then(...), so the + // actual fs.writeSync(1, ...) fires in a later microtask, not + // synchronously inside this call. suppressFdAsync (not captureFdAsync) + // keeps fs.writeSync patched across that await AND prevents the + // deferred write from ever reaching the real fd 1 — this exact window + // races node:test's own fd-1 IPC protocol under + // --test-isolation=process and forwarding (as captureFdAsync does) was + // empirically confirmed to still corrupt it (#4448). + await suppressFdAsync(1, async () => { + runMain(() => { + io.output({ found: false, error: 'not found' }, false); + return undefined; + }); + await waitForRunMain(); }); - await waitForRunMain(); assert.strictEqual(process.exitCode, 80, 'first runMain should have projected the DEGRADED cell to 80'); // Second, unrelated invocation in the SAME process declares nothing @@ -1002,8 +1017,10 @@ describe('review fix: pending-outcome cell lifetime (last-write-wins, cleared on // runMain, so this would inherit the first call's stale DEGRADED and // also exit 80 — the exact bug the reviewers found. process.exitCode = undefined; - runMain(() => undefined); - await waitForRunMain(); + await suppressFdAsync(1, async () => { + runMain(() => undefined); + await waitForRunMain(); + }); assert.strictEqual( process.exitCode, undefined, @@ -1018,12 +1035,14 @@ describe('review fix: pending-outcome cell lifetime (last-write-wins, cleared on resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); const savedExitCode = process.exitCode; try { - runMain(() => { - io.output({ found: false, error: 'not found' }, false); - io.output({ ok: true }, false); - return undefined; + await suppressFdAsync(1, async () => { + runMain(() => { + io.output({ found: false, error: 'not found' }, false); + io.output({ ok: true }, false); + return undefined; + }); + await waitForRunMain(); }); - await waitForRunMain(); assert.strictEqual( process.exitCode, undefined, @@ -1038,13 +1057,29 @@ describe('review fix: pending-outcome cell lifetime (last-write-wins, cleared on resolveContractVersion({ argv: ['node', 'x', '--exit-contract=v2'], env: {} }); const savedExitCode = process.exitCode; try { - runMain(() => { - io.output({ error: 'x' }, false); - return undefined; + const captured = await suppressFdAsync(1, async () => { + runMain(() => { + io.output({ error: 'x' }, false); + return undefined; + }); + await waitForRunMain(); }); - await waitForRunMain(); assert.strictEqual(process.exitCode, 80); assert.strictEqual(getPendingOutcome(), undefined, 'the cell must be cleared once runMain has consumed it'); + // Load-bearing check on the wrap itself, not just the outcome it + // guards: proves suppressFdAsync actually intercepted the deferred + // bytes runMain wrote (rather than the patch having been restored + // before the deferred write ran, which would silently capture an + // empty string — verified as the failure mode of a naive + // captureFdSync wrap during development of this fix). These bytes + // never reached the real fd 1 by design (#4448) — only the in-memory + // recording is asserted on here. + const parsedCaptured = JSON.parse(captured); + assert.strictEqual( + parsedCaptured.error, + 'x', + `expected suppressFdAsync to intercept the output({error}) JSON bytes for fd 1; got: ${JSON.stringify(captured)}`, + ); } finally { process.exitCode = savedExitCode; } From f15887ebb16aba42d6469a337969491104acff1f Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 21:23:39 -0400 Subject: [PATCH 029/166] feat(#4446): ban ad hoc timeout literals in tests, ship with full legacy allowlist (#4449) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nothing enforced CONTRIBUTING.md's own stated preference ("A non-literal value is trusted — that is the shape you should be writing") for timeout values in tests. local/no-unbounded-spawn requires SOME bound but allows a bare literal; no-magic-sleep-in-tests and no-elapsed-assertion cover different anti-patterns entirely. This is exactly how PR #4428's Windows CI incident happened: two independently-guessed 15000ms literals (one in production's check-latest- version.cjs, one in this suite's own worker test) collided exactly and raced two SIGKILLs against each other. New rule local/no-adhoc-timeout-literal (eslint-rules/no-adhoc-timeout- literal.cjs) flags a resolvable numeric timeout/timeoutMs literal; an Identifier or MemberExpression value is trusted. No marker-comment escape — the fix is always to extract a named constant, which is trivial. Ships with a full legacy allowlist (126 files, 352 violations, generated by running the rule with an empty allowlist against tests/) so it can go live at error severity without breaking CI, mirroring how no-unbounded-spawn itself was rolled out. Migration is tracked separately in #4445 (this PR does not close it — only #4446, introducing the gate itself). Documents the policy and the compliant shapes in TESTING-STANDARDS.md and TEST-EXAMPLES.md. Co-authored-by: sim Co-authored-by: Claude Sonnet 5 --- TEST-EXAMPLES.md | 45 ++++ TESTING-STANDARDS.md | 33 +++ .../no-adhoc-timeout-literal.allowlist.json | 128 ++++++++++++ eslint-rules/no-adhoc-timeout-literal.cjs | 195 ++++++++++++++++++ eslint.config.mjs | 11 + 5 files changed, 412 insertions(+) create mode 100644 eslint-rules/no-adhoc-timeout-literal.allowlist.json create mode 100644 eslint-rules/no-adhoc-timeout-literal.cjs diff --git a/TEST-EXAMPLES.md b/TEST-EXAMPLES.md index 8825a3810..558cc572c 100644 --- a/TEST-EXAMPLES.md +++ b/TEST-EXAMPLES.md @@ -73,6 +73,51 @@ for (const scenario of cases) { } ``` +## Named Timeout Constants, Not Ad Hoc Literals + +A bare numeric `timeout`/`timeoutMs` guessed per call site can silently drift from — or worse, +exactly collide with — an unrelated timeout somewhere else. That collision is not hypothetical: a +test's outer subprocess-wait timeout once matched a worker's own inner `npm view` timeout exactly +(both hardcoded to `15000`), so a slow response raced two SIGKILLs at the same instant and lost — +only on Windows CI, only intermittently. See [`TESTING-STANDARDS.md` — "No ad hoc timeout +literals"](TESTING-STANDARDS.md#no-ad-hoc-timeout-literals) for the full incident and +`local/no-adhoc-timeout-literal` for the lint rule that now catches this. + +**Non-compliant — a guessed literal with no relationship to what it's actually bounding:** + +```javascript +test('worker run leaves a valid cache', (t) => { + const r = runHookSeam(WORKER_PATH, [], { timeoutMs: 15000 }); // why 15000? nobody knows + assert.equal(r.exitCode, 0); +}); +``` + +**Compliant — reuse a shared class-norm constant when the call is the same class of subprocess:** + +```javascript +const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); + +test('gsd-tools reports the resolved config', (t) => { + const r = runNode([TOOLS_PATH, 'config', '--json'], { timeoutMs: PROBE_TIMEOUT_MS }); + assert.equal(r.exitCode, 0); +}); +``` + +**Compliant — a genuinely distinct class: name it, and size it relative to what it wraps:** + +```javascript +const { NPM_VIEW_TIMEOUT_MS } = require('../gsd-core/bin/check-latest-version.cjs'); + +// Real headroom beyond the inner timeout the worker itself is bounded by — not a +// second independent guess. See TESTING-STANDARDS.md's "No ad hoc timeout literals". +const WORKER_TEARDOWN_MARGIN_MS = 10_000; + +test('worker run leaves a valid cache', (t) => { + const r = runHookSeam(WORKER_PATH, [], { timeoutMs: NPM_VIEW_TIMEOUT_MS + WORKER_TEARDOWN_MARGIN_MS }); + assert.equal(r.exitCode, 0); +}); +``` + ## Parser Adversarial Fixtures Parser tests should cover malformed input and real-world file messiness. Prefer named fixtures under `tests/fixtures/adversarial//` when the input is reusable. diff --git a/TESTING-STANDARDS.md b/TESTING-STANDARDS.md index a2df72c9a..4a96cd87a 100644 --- a/TESTING-STANDARDS.md +++ b/TESTING-STANDARDS.md @@ -174,6 +174,38 @@ Invariant categories to consider: round-trip, monotonicity, boundary containment **Enforcement:** Code review verifies that property tests exist for modules in scope. Stryker mutation score below 80 % blocks merge (see next section). +### No ad hoc timeout literals + +Do not write a bare numeric `timeout`/`timeoutMs` option value at a test call site. Two independently-guessed copies of the same magic number can silently drift apart, or worse, collide exactly and produce a zero-margin race: `bin/check-latest-version.cjs`'s `timeout: 15_000` and this suite's independent `timeoutMs: 15000` could SIGKILL the whole process tree at the exact same instant, and it failed specifically on Windows CI (fixed in PR #4428). + +**Non-compliant:** + +```javascript +const r = runHookSeam(WORKER_PATH, [], { timeoutMs: 15000 }); +``` + +**Compliant — same class of subprocess as an existing class-norm:** + +```javascript +const { GIT_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); + +const r = runHookSeam(WORKER_PATH, [], { timeoutMs: GIT_TIMEOUT_MS }); +``` + +**Compliant — a genuinely distinct class, declared locally with a margin over the thing it wraps** (the actual fix in PR #4428 — the worker's inner `npm view` call is bounded by its own named `NPM_VIEW_TIMEOUT_MS`, so the outer test imports it and adds explicit headroom instead of re-guessing a number): + +```javascript +const { NPM_VIEW_TIMEOUT_MS } = require('../gsd-core/bin/check-latest-version.cjs'); + +const WORKER_TEARDOWN_MARGIN_MS = 10_000; // real headroom beyond the inner timeout it wraps + +const r = runHookSeam(WORKER_PATH, [], { timeoutMs: NPM_VIEW_TIMEOUT_MS + WORKER_TEARDOWN_MARGIN_MS }); +``` + +Import an existing class-norm constant from `tests/helpers/timeouts.cjs` (`PROBE_TIMEOUT_MS`, `GIT_TIMEOUT_MS`, `BUILD_TIMEOUT_MS`, `INSTALL_TIMEOUT_MS`) when the call is the same class of subprocess, or declare a local one with a comment justifying why it is a distinct class — see CONTRIBUTING.md's "Use Centralized Test Helpers" section. + +**Enforcement:** `local/no-adhoc-timeout-literal` (ESLint, `error`). A non-literal value (an `Identifier`, `MemberExpression`, or `CallExpression`) is trusted; only a resolvable numeric literal is flagged. There is no marker-comment escape — the fix is always to extract a named constant. `allowlist` (`eslint-rules/no-adhoc-timeout-literal.allowlist.json`) exempts pre-existing legacy violations and only ever ratchets down. + ### Mutation testing — 80 % threshold Stryker runs in incremental mode (`--since origin/next`) on the `ubuntu-latest` / Node 24 CI leg as a PR-gating signal. The default threshold is **80 % mutation score** (killed / total mutants in the changed scope). PRs that drop below this threshold must either add tests that kill the surviving mutants or add the specific path to `stryker.config.mjs` with a documented reason. @@ -210,6 +242,7 @@ Real multi-process race tests are deleted once the corresponding deterministic c | `local/no-source-grep` | `error` (promoted by #3313) | `readFileSync` on source files + text assertions; `assert.match`/`doesNotMatch` on raw stdout/stderr | | `local/no-magic-sleep-in-tests` | `error` | `setTimeout`/`sleep`/`delay` calls inside `test()`/`it()`/`describe()` bodies | | `local/no-elapsed-assertion` | `error` (promoted by #3331, precondition delivered by #3314) | Assertions on `Date.now()` delta, `process.hrtime()`, `performance.now()` comparisons | +| `local/no-adhoc-timeout-literal` | `error` | Bare numeric `timeout`/`timeoutMs` option literal in `tests/**/*.cjs` (PR #4428) | | `no-only-tests/no-only-tests` | `error` | `test.only`/`describe.only`/`it.only` committed to non-scratch files | | `no-restricted-syntax` (ban 1) | `error` | Top-level `setTimeout` in `ExpressionStatement` | | `no-restricted-syntax` (ban 2) | `error` | `.only` member access on `test`/`it`/`describe` (belt-and-suspenders) | diff --git a/eslint-rules/no-adhoc-timeout-literal.allowlist.json b/eslint-rules/no-adhoc-timeout-literal.allowlist.json new file mode 100644 index 000000000..f09d1b4b2 --- /dev/null +++ b/eslint-rules/no-adhoc-timeout-literal.allowlist.json @@ -0,0 +1,128 @@ +[ + "tests/adr-612-bracket-coherence.test.cjs", + "tests/adr-612-bracket-read-tolerance.test.cjs", + "tests/adr-index-gate.test.cjs", + "tests/adr857-core-without-capabilities.test.cjs", + "tests/antigravity-upgrades.test.cjs", + "tests/api-coverage-gate-e2e.test.cjs", + "tests/api-coverage.test.cjs", + "tests/assumption-delta-checkpoint-e2e.test.cjs", + "tests/assumption-delta.test.cjs", + "tests/augment-upgrades.test.cjs", + "tests/capability-cli.test.cjs", + "tests/capability-probe-fallback.test.cjs", + "tests/capability-state.test.cjs", + "tests/capability-trust.test.cjs", + "tests/capability-validator-task-content-resolver.test.cjs", + "tests/capability-writer.test.cjs", + "tests/changeset-new.test.cjs", + "tests/check-env.test.cjs", + "tests/check-gap-analysis-plan-post-e2e.test.cjs", + "tests/check-glossary-refs.test.cjs", + "tests/check-predicate.test.cjs", + "tests/check-tdd-review-checkpoint-e2e.test.cjs", + "tests/ci-rebase-check.test.cjs", + "tests/cjs-command-router-adapter.test.cjs", + "tests/code-review-pipeline-regression.test.cjs", + "tests/code-review.test.cjs", + "tests/commands.test.cjs", + "tests/commit-files-pathspec.test.cjs", + "tests/config-get-default.test.cjs", + "tests/cursor-hook-workspace-roots.test.cjs", + "tests/cursor-hooks.test.cjs", + "tests/dispatcher.test.cjs", + "tests/effort-surface-axis.test.cjs", + "tests/effort-sync-installed-runtime.test.cjs", + "tests/emitted-ack-trailer.test.cjs", + "tests/emitted-attribution.test.cjs", + "tests/execute-wave-post-gate-pipeline-e2e.test.cjs", + "tests/faulty-deps.test.cjs", + "tests/feat-2483-review-claude-mds-guard.test.cjs", + "tests/federated-config.test.cjs", + "tests/fragment-single-edit-propagation.install.test.cjs", + "tests/gate-predicate-evaluator.test.cjs", + "tests/gemini-runtime-removed.test.cjs", + "tests/gen-context-index.test.cjs", + "tests/gen-health-docs.test.cjs", + "tests/gen-section-manifest.test.cjs", + "tests/gen-state-md-docs.test.cjs", + "tests/git-base-branch.test.cjs", + "tests/git-fixture.test.cjs", + "tests/graphify.test.cjs", + "tests/gsd-check-update-worker-platform-gate.test.cjs", + "tests/gsd-mcp-server-bin.test.cjs", + "tests/gsd-secret-read-guard.test.cjs", + "tests/gsd-statusline.test.cjs", + "tests/gsd-validate-commit-crash-policy.test.cjs", + "tests/gsd-write-guard.test.cjs", + "tests/health-validation.test.cjs", + "tests/helpers-process-isolation.test.cjs", + "tests/helpers.cjs", + "tests/hooks-commonjs-marker.test.cjs", + "tests/hooks-crash-policy.test.cjs", + "tests/init.test.cjs", + "tests/install-minimal-hooks.test.cjs", + "tests/install-regressions.test.cjs", + "tests/install-runtime-artifacts.test.cjs", + "tests/install-write-confinement.test.cjs", + "tests/install.test.cjs", + "tests/kilo-upgrades.test.cjs", + "tests/kimi-upgrades.test.cjs", + "tests/kimi-variant-disambiguation.test.cjs", + "tests/lint-docs-command-form.test.cjs", + "tests/locking-bugs-1909-1916-1925-1927.test.cjs", + "tests/loop-hooks-empty-points-e2e.test.cjs", + "tests/loop-hooks-ship-pre-e2e.test.cjs", + "tests/loop-hooks-verify-post-e2e.test.cjs", + "tests/loop-render-hooks.test.cjs", + "tests/loop-walk.qa.test.cjs", + "tests/milestone-lock.test.cjs", + "tests/no-pending-3212-markers.test.cjs", + "tests/npm-integrity-gate.test.cjs", + "tests/opencode-plugin-adapter.test.cjs", + "tests/packaging-shipped-scripts-require-only-shipped.test.cjs", + "tests/pattern.test.cjs", + "tests/perf-316-state-lock-buffer-alloc.test.cjs", + "tests/perf-317-context-monitor-fs.test.cjs", + "tests/phase.test.cjs", + "tests/phase6-capstone-conformance.test.cjs", + "tests/pi-config-dir-env-override.test.cjs", + "tests/plan-phase-stall-detection.test.cjs", + "tests/plan-pre-hook-e2e.test.cjs", + "tests/plan-review-convergence.test.cjs", + "tests/plugin-manifest.test.cjs", + "tests/policy-160-route0-resume.test.cjs", + "tests/pr-branch-planning-filter.test.cjs", + "tests/process-seam.test.cjs", + "tests/prohibition-enforcement.test.cjs", + "tests/prompt-injection-scan.security.test.cjs", + "tests/qa/tdd-walk.cjs", + "tests/quick-batch.test.cjs", + "tests/read-guard.test.cjs", + "tests/read-injection-scanner.property.test.cjs", + "tests/read-injection-scanner.security.test.cjs", + "tests/reapply-verify-hunks.test.cjs", + "tests/release-tarball-smoke.install.test.cjs", + "tests/representative-corpus.test.cjs", + "tests/review-lane-invocation.test.cjs", + "tests/reviewer-manifest-body.test.cjs", + "tests/reviewer-trust-disclosure.test.cjs", + "tests/run-tests-temp-root.test.cjs", + "tests/run-with-timeout.test.cjs", + "tests/secret-scan-lint.security.test.cjs", + "tests/security-prompt-injection.security.test.cjs", + "tests/security-scan.security.test.cjs", + "tests/security.test.cjs", + "tests/shared-hooks-dir-resolution.test.cjs", + "tests/shell-command-projection-dispatch.test.cjs", + "tests/ship-notes-wedged-pr.test.cjs", + "tests/slug-derivation-drift-guard.test.cjs", + "tests/state-document.test.cjs", + "tests/state-todos-render.test.cjs", + "tests/task-command-router-resolve-content.test.cjs", + "tests/task-content-resolution.test.cjs", + "tests/task-content-resolver-grammar-parity.test.cjs", + "tests/teams-status.test.cjs", + "tests/windsurf-hooks-bridge.test.cjs", + "tests/worktree-safety.test.cjs" +] diff --git a/eslint-rules/no-adhoc-timeout-literal.cjs b/eslint-rules/no-adhoc-timeout-literal.cjs new file mode 100644 index 000000000..5cb038175 --- /dev/null +++ b/eslint-rules/no-adhoc-timeout-literal.cjs @@ -0,0 +1,195 @@ +'use strict'; + +const path = require('path'); + +/** + * no-adhoc-timeout-literal + * + * Flag a bare numeric literal used as a `timeout`/`timeoutMs` option value in + * a test file. This is a **test-suite-only** convention rule — it does not + * touch production `src/`/`bin/` code, where a literal like `execNpm(args, { + * timeout: 15_000 })` bounds a real subprocess for real production + * resilience. CONTRIBUTING.md/CLAUDE.md separately mandate that production + * timeout bound on safety grounds (never hang); this rule is about naming a + * shared, reviewed ceiling instead of scattering ad hoc guesses. + * + * ## What this enforces + * + * A bare `timeout`/`timeoutMs` literal scattered per call site drifts + * silently from its siblings, or worse, collides exactly with one. That is + * not hypothetical: on 2026-09-06, `gsd-core/bin/check-latest-version.cjs` + * hardcoded `execNpm(args, { timeout: 15_000 })` and, independently, + * `tests/gsd-check-update-worker-atomic-cache.test.cjs` hardcoded + * `runHookSeam(WORKER_PATH, [], { timeoutMs: 15000 })` — two unrelated files, + * same guessed number, no shared reference. When the inner one's timeout + * fired, the outer one could SIGKILL the whole process tree at the exact + * same instant before it could degrade gracefully: a zero-margin race that + * failed specifically on Windows CI (fixed in PR #4428 by extracting a named + * `NPM_VIEW_TIMEOUT_MS` constant and referencing it with an explicit + * margin). This rule closes the gap CONTRIBUTING.md already documents: + * "A non-literal value (`timeout: GIT_TIMEOUT_MS`) is trusted — that is the + * shape you should be writing" — by actually enforcing that shape. + * + * ## Recognized shape + * + * Any non-computed `Property` node whose key is exactly `timeout` or + * `timeoutMs` (string or Identifier key form) is flagged when its `value` + * resolves, via `evalNumeric` (Literal number, unary +/-, or a `*`/`+`/`-`/`/` + * BinaryExpression chain — same logic as `no-unbounded-spawn.cjs`), to a + * concrete JS number. An `Identifier` value (including shorthand + * `{ timeoutMs }`), a `MemberExpression` (`opts.timeout`, + * `TIMEOUTS.PROBE`), or a `CallExpression` value all fail to resolve via + * `evalNumeric` and are trusted as-is — this rule does not attempt general + * expression evaluation, matching `no-unbounded-spawn`'s own philosophy of + * trusting anything it can't literally evaluate to a number. + * + * ## No marker-comment escape + * + * Unlike `no-unbounded-spawn`'s `// allow-spawn-timeout-ceiling: ` + * (which has a genuine "sometimes a call really does need >600s" exception), + * there is no legitimate reason a timeout value needs to stay an inline + * literal forever — the fix is always "extract to a named constant," which + * is trivial. The only escape here is the allowlist below, and it is a + * temporary migration aid, not a permanent one. + * + * ## Allowlist + * + * `allowlist` (repo-relative POSIX paths) exempts pre-existing legacy + * violations, with mechanics identical to `no-unbounded-spawn.cjs`: an + * allowlisted file's violations are counted internally but not reported: a + * listed file with zero violations reports `staleAllowlistEntry` so the dead + * entry gets deleted. The allowlist only ever ratchets down. + */ + +/** @type {import('eslint').Rule.RuleModule} */ +const rule = { + meta: { + type: 'problem', + docs: { + description: + 'Disallow a bare numeric literal as a timeout/timeoutMs option value in tests', + category: 'Reliability', + }, + schema: [ + { + type: 'object', + properties: { + allowlist: { type: 'array', items: { type: 'string' } }, + }, + additionalProperties: false, + }, + ], + messages: { + adhocTimeoutLiteral: + 'Bare numeric `{{key}}: {{value}}` literal: two independent hardcoded copies of the ' + + 'same guessed timeout can silently drift apart, or worse, collide exactly and produce ' + + 'a zero-margin race (this repo hit exactly that on 2026-09-06 — ' + + '`bin/check-latest-version.cjs`\'s `timeout: 15_000` and this suite\'s independent ' + + '`timeoutMs: 15000` could SIGKILL the whole tree at the same instant, PR #4428). ' + + 'Extract a named constant: import an existing one from `tests/helpers/timeouts.cjs` ' + + '(`PROBE_TIMEOUT_MS`, `GIT_TIMEOUT_MS`, `BUILD_TIMEOUT_MS`, `INSTALL_TIMEOUT_MS`) if this ' + + 'call is the same class of subprocess, or declare a local one with a comment justifying ' + + 'why it is a distinct class, per CONTRIBUTING.md\'s "Use Centralized Test Helpers" section.', + staleAllowlistEntry: + '{{file}} no longer contains an ad hoc timeout literal. Delete its line from ' + + 'eslint-rules/no-adhoc-timeout-literal.allowlist.json — the allowlist only ratchets down.', + }, + }, + + create(context) { + const options = context.options[0] || {}; + const allowlist = Array.isArray(options.allowlist) ? options.allowlist : []; + + const TIMEOUT_KEYS = new Set(['timeout', 'timeoutMs']); + + const filename = context.filename || context.getFilename(); + const cwd = context.cwd || (context.getCwd ? context.getCwd() : process.cwd()); + const rel = path.relative(cwd, filename).split(path.sep).join('/'); + const allowlisted = allowlist.includes(rel); + let violations = 0; + + /** + * Returns the string value of a Literal node, or null. + */ + function stringValue(node) { + if (node && node.type === 'Literal' && typeof node.value === 'string') { + return node.value; + } + return null; + } + + /** Recursion depth cap for evalNumeric — guards against a pathological + * nested-expression chain blowing the stack. */ + const MAX_EVAL_DEPTH = 20; + + /** + * Recursively evaluates a numeric-ish AST node to a JS number, or + * returns undefined if it's not one of the recognized numeric shapes. + * Handles a numeric Literal, a unary +/- of a recursively-numeric + * argument, and a BinaryExpression (*, +, -, /) where both sides are + * recursively numeric — so a multi-term chain like `60 * 60 * 1000` + * resolves instead of bailing out on the first nested BinaryExpression. + * Identical logic to `no-unbounded-spawn.cjs`'s `evalNumeric`. + */ + function evalNumeric(node, depth = 0) { + if (depth > MAX_EVAL_DEPTH) return undefined; + if (node.type === 'Literal' && typeof node.value === 'number') { + return node.value; + } + if (node.type === 'UnaryExpression' && (node.operator === '-' || node.operator === '+')) { + const arg = evalNumeric(node.argument, depth + 1); + if (arg === undefined) return undefined; + return node.operator === '-' ? -arg : arg; + } + if ( + node.type === 'BinaryExpression' && + (node.operator === '*' || node.operator === '+' || node.operator === '-' || node.operator === '/') + ) { + const left = evalNumeric(node.left, depth + 1); + const right = evalNumeric(node.right, depth + 1); + if (left === undefined || right === undefined) return undefined; + switch (node.operator) { + case '*': + return left * right; + case '+': + return left + right; + case '-': + return left - right; + case '/': + return left / right; + default: + return undefined; + } + } + return undefined; + } + + return { + Property(node) { + if (node.computed) return; + const keyName = node.key.type === 'Identifier' ? node.key.name : stringValue(node.key); + if (!keyName || !TIMEOUT_KEYS.has(keyName)) return; + + const numeric = evalNumeric(node.value); + if (numeric === undefined) return; + + violations += 1; + if (allowlisted) return; + + context.report({ + node, + messageId: 'adhocTimeoutLiteral', + data: { key: keyName, value: String(numeric) }, + }); + }, + + 'Program:exit'(node) { + if (allowlisted && violations === 0) { + context.report({ node, messageId: 'staleAllowlistEntry', data: { file: rel } }); + } + }, + }; + }, +}; + +module.exports = rule; diff --git a/eslint.config.mjs b/eslint.config.mjs index a017ad849..c20e84471 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -5,8 +5,10 @@ import pluginN from 'eslint-plugin-n'; import noOnlyTests from 'eslint-plugin-no-only-tests'; import { dirname } from 'path'; import { fileURLToPath } from 'url'; +import { createRequire } from 'module'; const __dirname = dirname(fileURLToPath(import.meta.url)); +const require = createRequire(import.meta.url); // Local plugin with custom AST rules import noSourceGrep from './eslint-rules/no-source-grep.cjs'; @@ -36,6 +38,9 @@ import noPrivateBinaryResolution from './eslint-rules/no-private-binary-resoluti import requireRegisteredExit from './eslint-rules/require-registered-exit.cjs'; import noSwallowedPrecondition from './eslint-rules/no-swallowed-precondition.cjs'; import noExactCaseEnvAccess from './eslint-rules/no-exact-case-env-access.cjs'; +import noAdhocTimeoutLiteral from './eslint-rules/no-adhoc-timeout-literal.cjs'; + +const adhocTimeoutLiteralAllowlist = require('./eslint-rules/no-adhoc-timeout-literal.allowlist.json'); const localPlugin = { rules: { @@ -66,6 +71,7 @@ const localPlugin = { 'require-registered-exit': requireRegisteredExit, 'no-swallowed-precondition': noSwallowedPrecondition, 'no-exact-case-env-access': noExactCaseEnvAccess, + 'no-adhoc-timeout-literal': noAdhocTimeoutLiteral, }, }; @@ -715,6 +721,11 @@ export default tseslint.config( // exemption surface. The only sanctioned escapes are an explicit `timeout` on // a raw spawn or the `// allow-spawn-timeout-ceiling: ` marker. 'local/no-unbounded-spawn': 'error', + // Ban a bare numeric `timeout`/`timeoutMs` literal in tests (DEFECT.AD-HOC-TIMEOUT-LITERAL, + // #4428): two independently-guessed copies of the same magic number can drift apart, or + // collide exactly into a zero-margin race. Allowlist starts empty; a pre-existing violation + // gets grandfathered in here as it's found, per eslint-rules/no-adhoc-timeout-literal.allowlist.json. + 'local/no-adhoc-timeout-literal': ['error', { allowlist: adhocTimeoutLiteralAllowlist }], // Ban a consolidation-epic folded suite appearing twice in one host file (#3271). // A second copy runs the same tests twice on every lane and drifts silently. 'local/no-duplicate-fold-marker': 'error', From 8c8eda46b0ec315e98ba5adf11486b529816ce14 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sun, 6 Sep 2026 22:24:07 -0400 Subject: [PATCH 030/166] fix(#4225): scope the sibling-worktree phase-number horizon to the active workstream (#4450) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4225): failing-first matrix for phase.add --ws workstream-scoped numbering Nine rows driven through the real CLI: the issue's verbatim topology (root roadmap @39 committed, workstream @2, sibling git worktree carrying the root roadmap), same-workstream sibling boundary, empty-workstream first phase, coincidental root-maximum, no---ws control (the #3849 global horizon, byte-for-byte), cross-workstream isolation, sibling lacking the workstream (fail open), add-batch parity, and a next-decimal control. Rows 1/2/3/6/7/8 are RED on next @38e4ce5f62 (numbering computed from the sibling ROOT roadmaps: 40 instead of 3). * fix(#4225): scope the #3849 sibling-worktree widening horizon to the active workstream collectSiblingWorktreePhaseNums scanned each sibling git worktree's ROOT .planning/ (phases/ dirs + ROADMAP.md headers) unconditionally. Under --ws (GSD_WORKSTREAM), every local number source flows through planningDir(cwd) and lands in the workstream scope, but the widening horizon still merged the siblings' ROOT-roadmap numbers into it — so phase.add --ws in a workstream at Phase 2 inside a project whose root roadmap sits at Phase 39 minted Phase 40 (directory 40-, and a Depends on: Phase 39 that does not exist in the workstream's numbering universe). The horizon now resolves each sibling's planning dir through the SAME canonical resolver, planningDir(wt, ws), with the env workstream read once via planningDir's own discriminator: a workstream-scoped allocation scans the sibling's copy of the SAME workstream (a number taken by that workstream on another branch is still taken — the #3849 widening survives, scoped), and never the sibling's root roadmap or another workstream's. No workstream active: ws is null and the root-scope horizon is byte-for-byte the #3849 behavior. A sibling lacking the workstream directory contributes nothing (fail open, unchanged). phase.add and phase.add-batch share the helper; both scopes of both verbs are covered by the matrix in the previous commit. Output shape and the publishStateContract boundary are untouched — only the number changes. * fix(#4225): rename siblingPlanning -> siblingPlanningDir (review nit) * chore(#4225): changeset fragment (pr number to backfill) * chore(#4225): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/merry-otters-wave.md | 5 + src/phase.cts | 33 +++- tests/phase.test.cjs | 273 ++++++++++++++++++++++++++++++++ 3 files changed, 304 insertions(+), 7 deletions(-) create mode 100644 .changeset/merry-otters-wave.md diff --git a/.changeset/merry-otters-wave.md b/.changeset/merry-otters-wave.md new file mode 100644 index 000000000..e230aa421 --- /dev/null +++ b/.changeset/merry-otters-wave.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4450 +--- +**`phase.add --ws` now numbers the next phase from the workstream's own roadmap** — in a project with sibling git worktrees, `phase.add`/`phase.add-batch` with `--ws` minted a phase number pulled from the root roadmap's maximum (e.g. Phase 40 in a workstream whose own roadmap stopped at Phase 2), creating a `40-` directory and a `Depends on: Phase 39` entry pointing at a phase that does not exist in the workstream. The sibling-worktree widening horizon is now scoped like every other number source: a workstream-scoped allocation counts numbers held by the same workstream in sibling worktrees only. (#4225) diff --git a/src/phase.cts b/src/phase.cts index 832f345bf..df1679537 100644 --- a/src/phase.cts +++ b/src/phase.cts @@ -1235,12 +1235,24 @@ function assertDescriptionPreservesMilestoneScope(cwd: string, description: stri * before any directory does; milestone-scoping is wrong here because a number * used under any milestone on another branch is still taken). * + * #4225 — the horizon must track the ALLOCATION scope. When the allocation is + * workstream-scoped (`--ws`/`GSD_WORKSTREAM`, resolved into the env before + * dispatch), the sibling's copy of the SAME workstream is what carries that + * scope's independent numbering; the sibling's ROOT roadmap and phases/ + * belong to a different numbering universe (docs/FEATURES.md §51 REQ-WS-01 — + * workstream state is isolated in `.planning/workstreams/{name}/`) and must + * not contribute. `planningDir(wt, ws)` reuses the canonical resolver, so the + * sibling scope matches the local scope's own resolution (env workstream plus + * env project segment) by construction; `ws === null` (no workstream active) + * keeps the #3849 root-scope horizon byte-for-byte. + * * Widen, never refuse: a missing `.planning/`, an unreadable sibling, a * non-git cwd, or an unavailable git binary each leave `used` untouched — * allocation then behaves exactly as it did before this horizon existed. - * Sentinels reuse the canonical `isSentinelPhaseId`; the dir pattern is the - * same one the on-disk scan uses, so decimal sub-phases (`411.1-foo`) are - * correctly not integers. + * A sibling that simply lacks the active workstream's directory is the same + * fail-open case: it contributes nothing. Sentinels reuse the canonical + * `isSentinelPhaseId`; the dir pattern is the same one the on-disk scan uses, + * so decimal sub-phases (`411.1-foo`) are correctly not integers. */ function collectSiblingWorktreePhaseNums(cwd: string, used: Set): void { let porcelain: string; @@ -1257,6 +1269,13 @@ function collectSiblingWorktreePhaseNums(cwd: string, used: Set): void { } catch { return; // not a git repo / git unavailable — unchanged behavior } + // #4225: the env workstream, read once with planningDir's own discriminator + // (`?? null` = deliberately no workstream — never re-derived per sibling). + // A poisoned value would already have thrown at the local `planningDir(cwd)` + // call every allocator makes before reaching this horizon; the per-sibling + // try/catch below still keeps any resolution failure fail-open. + const ws = process.env['GSD_WORKSTREAM'] ?? null; + const siblingPlanningDir = (wt: string): string => planningDir(wt, ws); const dirNumPattern = /^(?:[A-Z][A-Z0-9]*-)?(\d+)-/; // Same header shape the allocators scan locally (#1729 tag tolerance). const headerPattern = /#{2,4}\s*Phase\s+(\d+)[A-Z]?(?:\.\d+)*(?:\s*\([^)\n]{0,200}\))?:/gi; @@ -1265,17 +1284,17 @@ function collectSiblingWorktreePhaseNums(cwd: string, used: Set): void { const wt = line.slice('worktree '.length).trim(); if (!wt || path.resolve(wt) === path.resolve(cwd)) continue; try { - for (const entry of fs.readdirSync(path.join(wt, '.planning', 'phases'))) { + for (const entry of fs.readdirSync(path.join(siblingPlanningDir(wt), 'phases'))) { const match = entry.match(dirNumPattern); if (!match) continue; const num = parseInt(match[1], 10); if (!isSentinelPhaseId(num)) used.add(num); } } catch { - /* worktree has no .planning — normal, contributes nothing */ + /* worktree has no .planning (or no copy of this scope) — normal, contributes nothing */ } try { - const content = fs.readFileSync(path.join(wt, '.planning', 'ROADMAP.md'), 'utf-8'); + const content = fs.readFileSync(path.join(siblingPlanningDir(wt), 'ROADMAP.md'), 'utf-8'); let m: RegExpExecArray | null; headerPattern.lastIndex = 0; while ((m = headerPattern.exec(content)) !== null) { @@ -1283,7 +1302,7 @@ function collectSiblingWorktreePhaseNums(cwd: string, used: Set): void { if (!isSentinelPhaseId(num)) used.add(num); } } catch { - /* no roadmap in that worktree — normal, contributes nothing */ + /* no roadmap in that worktree (or scope) — normal, contributes nothing */ } } } diff --git a/tests/phase.test.cjs b/tests/phase.test.cjs index 83c36d646..5eb1fa273 100644 --- a/tests/phase.test.cjs +++ b/tests/phase.test.cjs @@ -2660,6 +2660,279 @@ describe('phase add allocation vs sibling git worktrees (#3849)', () => { }); }); +// ───────────────────────────────────────────────────────────────────────────── +// workstream-scoped phase allocation vs sibling git worktrees (#4225) +// +// The #3849 widening horizon (collectSiblingWorktreePhaseNums) scans each +// sibling's ROOT .planning/. A --ws allocation must instead scan the sibling's +// copy of the SAME workstream: workstream numbering is independent of the root +// roadmap (Workstream Namespacing REQ-WS-01 — workstream state is isolated in +// .planning/workstreams/{name}/), so a sibling's root-roadmap numbers must not +// leak into a workstream allocation. +// ───────────────────────────────────────────────────────────────────────────── + +describe('phase add --ws workstream-scoped allocation vs sibling git worktrees (#4225)', () => { + const activeWorktrees = []; + const activeDirs = []; + + function git(args, cwd) { + return execFileSync('git', args, { cwd, encoding: 'utf-8', timeout: 15_000 }); + } + + /** + * The #4225 topology: a committed ROOT roadmap with phases 1..rootMax (every + * linked worktree carries it), plus workstream `wsName` whose own roadmap and + * phases/ hold 1..wsMax. + */ + function initWsRepo(repoDir, rootMax, wsName, wsMax) { + const rootLines = ['# Roadmap', '', '## Milestone v1.0', '']; + for (let i = 1; i <= rootMax; i++) { + rootLines.push(`### Phase ${i}: root phase ${i}`, '', '**Goal:** root goal', ''); + } + fs.mkdirSync(path.join(repoDir, '.planning', 'phases'), { recursive: true }); + fs.writeFileSync(path.join(repoDir, '.planning', 'ROADMAP.md'), rootLines.join('\n') + '\n'); + + const wsBase = path.join(repoDir, '.planning', 'workstreams', wsName); + const wsLines = ['# Workstream Roadmap', '', '## Milestone v1.0', '']; + for (let i = 1; i <= wsMax; i++) { + wsLines.push(`### Phase ${i}: ws phase ${i}`, '', '**Goal:** ws goal', ''); + } + fs.mkdirSync(path.join(wsBase, 'phases'), { recursive: true }); + if (wsMax > 0) { + fs.mkdirSync(path.join(wsBase, 'phases', String(wsMax).padStart(2, '0') + '-ws-phase-' + wsMax)); + } + fs.writeFileSync(path.join(wsBase, 'ROADMAP.md'), wsLines.join('\n') + '\n'); + + git(['init', '-b', 'main'], repoDir); + git(['config', 'user.email', 'test@example.com'], repoDir); + git(['config', 'user.name', 'Test'], repoDir); + git(['add', '-A'], repoDir); + git(['commit', '-m', 'init'], repoDir); + return wsBase; + } + + /** A sibling linked worktree checked out at HEAD (carries the committed root roadmap). */ + function addSiblingAtHead(repoDir) { + const sha = git(['rev-parse', 'HEAD'], repoDir).trim(); + const worktreeDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-sib-')); + git(['worktree', 'add', '--detach', worktreeDir, sha], repoDir); + activeWorktrees.push({ repoDir, worktreeDir }); + return worktreeDir; + } + + function teardown() { + while (activeWorktrees.length) { + const { repoDir, worktreeDir } = activeWorktrees.pop(); + try { + git(['worktree', 'remove', '--force', worktreeDir], repoDir); + } catch (_) { /* best-effort; cleanup() below still removes the directory */ } + cleanup(worktreeDir); + } + while (activeDirs.length) cleanup(activeDirs.pop()); + } + + afterEach(teardown); + + test('phase add --ws numbers from the workstream, not the sibling root roadmaps (issue verbatim)', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-main-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 2); + addSiblingAtHead(repoDir); + + const result = runGsdTools(['query', 'phase.add', 'New workstream feature', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual( + output.phase_number, + 3, + '#4225: the workstream\'s own next number (max 2 + 1), not the root roadmap\'s 39 + 1' + ); + assert.strictEqual(output.padded, '03'); + assert.strictEqual( + output.directory, + '.planning/workstreams/ws-alpha/phases/03-new-workstream-feature', + 'directory must land inside the workstream with the workstream-scoped number' + ); + + const wsRoadmap = fs.readFileSync( + path.join(repoDir, '.planning', 'workstreams', 'ws-alpha', 'ROADMAP.md'), + 'utf-8' + ); + assert.ok(wsRoadmap.includes('### Phase 3: New workstream feature'), 'ws roadmap entry must be Phase 3'); + assert.ok(wsRoadmap.includes('**Depends on:** Phase 2'), 'dependency must reference the in-workstream predecessor'); + + // The root roadmap and root phases/ are a different numbering universe — untouched. + const rootRoadmap = fs.readFileSync(path.join(repoDir, '.planning', 'ROADMAP.md'), 'utf-8'); + assert.ok(!rootRoadmap.includes('New workstream feature'), 'root roadmap must not gain the workstream phase'); + assert.ok( + !fs.readdirSync(path.join(repoDir, '.planning', 'phases')).some((e) => e.startsWith('40-')), + 'root phases/ must not gain a 40- directory' + ); + }); + + test('phase add --ws still skips a number held by the SAME workstream in a sibling', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-same-ws-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 2); + const sibling = addSiblingAtHead(repoDir); + + // The sibling's ws-alpha copy holds Phase 3 (a branch that already minted it). + const sibWsPhases = path.join(sibling, '.planning', 'workstreams', 'ws-alpha', 'phases'); + fs.mkdirSync(path.join(sibWsPhases, '03-sibling-only'), { recursive: true }); + const sibWsRoadmap = path.join(sibling, '.planning', 'workstreams', 'ws-alpha', 'ROADMAP.md'); + fs.appendFileSync(sibWsRoadmap, '\n### Phase 3: sibling-only phase\n\n**Goal:** taken\n'); + + const result = runGsdTools(['query', 'phase.add', 'Contended', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual( + output.phase_number, + 4, + '#4225 keeps the #3849 widening: a number taken by the same workstream on another branch is still taken — but the sibling root roadmap\'s 39 must not decide it' + ); + }); + + test('phase add --ws numbers the FIRST phase of an empty workstream as 1', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-empty-ws-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-empty', 0); + addSiblingAtHead(repoDir); + + const result = runGsdTools(['query', 'phase.add', 'First steps', '--ws', 'ws-empty'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.phase_number, 1, 'an empty workstream starts at 1, not root-max+1 (40)'); + assert.strictEqual(output.padded, '01'); + assert.ok( + fs.existsSync(path.join(repoDir, '.planning', 'workstreams', 'ws-empty', 'phases', '01-first-steps')), + 'directory should be 01-first-steps inside the workstream' + ); + }); + + test('phase add --ws at the root maximum is coincidentally equal, with an in-workstream dependency', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-ws39-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 39); + addSiblingAtHead(repoDir); + + const result = runGsdTools(['query', 'phase.add', 'Fortieth in workstream', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.phase_number, 40, 'the workstream\'s own 39+1 — correct for its own reason'); + const wsRoadmap = fs.readFileSync( + path.join(repoDir, '.planning', 'workstreams', 'ws-alpha', 'ROADMAP.md'), + 'utf-8' + ); + assert.ok( + wsRoadmap.includes('**Depends on:** Phase 39'), + 'Phase 39 dependency is legitimate here: it exists in THIS workstream' + ); + }); + + test('a sibling holding only ANOTHER workstream\'s numbers does not affect --ws allocation', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-beta-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 2); + const sibling = addSiblingAtHead(repoDir); + + // The sibling's ws-beta (a different, independent numbering universe) is at 50. + const betaPhases = path.join(sibling, '.planning', 'workstreams', 'ws-beta', 'phases'); + fs.mkdirSync(path.join(betaPhases, '50-beta-heavy'), { recursive: true }); + const betaRoadmap = path.join(sibling, '.planning', 'workstreams', 'ws-beta', 'ROADMAP.md'); + fs.mkdirSync(path.dirname(betaRoadmap), { recursive: true }); + fs.writeFileSync(betaRoadmap, ['# Workstream Roadmap', '', '### Phase 50: beta heavy', '', '**Goal:** beta', ''].join('\n') + '\n'); + + const result = runGsdTools(['query', 'phase.add', 'Alpha next', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual( + output.phase_number, + 3, + 'cross-workstream numbers (root 39 in the sibling root roadmap, ws-beta 50) are different universes — ws-alpha numbers from its own 2+1' + ); + }); + + test('a sibling without the workstream directory contributes nothing (fail open)', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-no-ws-sib-')); + activeDirs.push(repoDir); + const wsBase = initWsRepo(repoDir, 39, 'ws-alpha', 2); + const sibling = addSiblingAtHead(repoDir); + + // The sibling predates the workstream: its checkout has no ws-alpha at all. + // eslint-disable-next-line local/no-raw-rmsync-in-tests -- fixture SETUP, not teardown: strips the sibling's workstreams/ so it lacks the active scope; teardown() owns the dir's removal + fs.rmSync(path.join(sibling, '.planning', 'workstreams'), { recursive: true, force: true }); + + const result = runGsdTools(['query', 'phase.add', 'Local only', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual( + output.phase_number, + 3, + 'a missing workstream scope in a sibling contributes nothing — local ws sources decide' + ); + assert.ok(fs.existsSync(path.join(wsBase, 'phases', '03-local-only'))); + }); + + test('phase add-batch --ws numbers sequentially from the workstream, not the root roadmap', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-batch-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 2); + addSiblingAtHead(repoDir); + + const result = runGsdTools( + ['query', 'phase.add-batch', '--descriptions', '["Batch A","Batch B"]', '--ws', 'ws-alpha'], + repoDir + ); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.phases[0].phase_number, 3, '#4225: the batch allocator shares the scoped horizon'); + assert.strictEqual(output.phases[1].phase_number, 4); + }); + + test('without --ws, allocation keeps the #3849 global sibling horizon byte-for-byte', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-flat-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 440, 'ws-alpha', 2); + const sibling = addSiblingAtHead(repoDir); + + // The sibling's ROOT roadmap holds Phase 441 (#3849's shape). + fs.appendFileSync( + path.join(sibling, '.planning', 'ROADMAP.md'), + '\n### Phase 441: sibling root phase\n\n**Goal:** taken\n' + ); + + const result = runGsdTools(['query', 'phase.add', 'Flat allocation'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual( + output.phase_number, + 442, + 'no --ws: the sibling ROOT roadmap still widens the horizon exactly as #3849 shipped' + ); + }); + + test('phase next-decimal --ws is unaffected (planningDir-scoped already)', () => { + const repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4225-dec-')); + activeDirs.push(repoDir); + initWsRepo(repoDir, 39, 'ws-alpha', 2); + addSiblingAtHead(repoDir); + + const result = runGsdTools(['query', 'phase.next-decimal', '2', '--ws', 'ws-alpha'], repoDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.next, '02.1', 'next-decimal was never sibling-widened; the fix must not change it'); + }); +}); + // ───────────────────────────────────────────────────────────────────────────── // phase add-batch command (#2165) // ───────────────────────────────────────────────────────────────────────────── From 33e393ba4c6ec3686c6617b5dcadc6331ba2a894 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 00:03:15 -0400 Subject: [PATCH 031/166] fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip * fix(#4243): anchor stateReplaceField bold form to line start The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold label quoted mid-sentence inside prose — the issue's **Status:** inside an Accumulated Context bullet — captured the rewrite and destroyed the rest of its line, silently, whenever a whole-body caller fed the function every section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan writes). The plain branch was always line-anchored; only the bold branch lagged. Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's same-line confinement idiom for the leading class (deliberately not the issue's suggested ^\s* — it can consume the newlines before the label into the match) and #4186's recognition-by-anchoring discipline. Frontmatter half of the issue (unknown-key drops, invented milestone defaults) is already fixed on next by #2202/#3216/#4129; pinned here with the issue's requested regression fixtures. * test(#4243): pin survival contract, not derived percent, in frontmatter rows Bench RED run caught two assertion defects in the pin rows: the unknown progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a declared derived subkey - omitted under the #3573 no-roadmap withhold, recomputed when measured (#4129) - so pinning its value over-pins derived semantics. The rows now pin what the issue demands: unknown/custom keys survive, stored counters are kept under the withhold, milestone identity is never reset to invented defaults. * chore(#4243): changeset for the anchored bold-field fix * chore(#4243): backfill PR number in changeset --- .changeset/calm-goats-bark.md | 5 + src/state-document.cts | 17 ++- tests/state-document.test.cjs | 181 ++++++++++++++++++++++++ tests/state.test.cjs | 249 ++++++++++++++++++++++++++++++++++ 4 files changed, 451 insertions(+), 1 deletion(-) create mode 100644 .changeset/calm-goats-bark.md diff --git a/.changeset/calm-goats-bark.md b/.changeset/calm-goats-bark.md new file mode 100644 index 000000000..fa9021cbf --- /dev/null +++ b/.changeset/calm-goats-bark.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4453 +--- +**`state begin-phase` no longer rewrites prose that merely quotes a bold field label** — a `**Status:**` (or any served field label) quoted mid-sentence inside prose captured the field rewrite and silently destroyed the rest of its line; the bold form is now anchored to line start, so only the real field updates. Frontmatter round-trip through begin-phase (custom keys, progress subkeys, milestone identity without a ROADMAP) is pinned with regression tests. (#4243) diff --git a/src/state-document.cts b/src/state-document.cts index a8fc16bb8..15f6bb7cf 100644 --- a/src/state-document.cts +++ b/src/state-document.cts @@ -553,7 +553,22 @@ export function stateReplaceField(content: string, fieldName: string, newValue: // `(.*)` captured the following line and the rebuild discarded it — the #4010 // data-loss. ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; // this aligns the writer to it. - const boldPattern = new RegExp(`(\\*\\*${escaped}:\\*\\*[ \\t]*)(.*)`, 'i'); + // + // #4243: the bold form is also ANCHORED to line start, with same-line leading + // whitespace only. The pre-fix pattern carried no `^` and no `m` flag, so a + // bold label quoted MID-SENTENCE inside prose — an Accumulated Context bullet + // mentioning `**Status:**` — captured the rewrite and destroyed the rest of + // its line, silently, whenever a whole-body caller fed this function every + // section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan + // writes). The plain branch below was always line-anchored; only the bold + // branch lagged. Anchoring reuses #4010's same-line confinement idiom (the + // leading class is `[ \t]*`, deliberately NOT the `\s*` the issue suggested — + // `^\s*\*\*` can consume the newlines before the label into the match and + // drop them on rebuild) and #4186's recognition-by-anchoring discipline: a + // write target must BE the whole declared line shape, never a substring + // guess inside prose. `$` is explicit-and-inert (`.` never crosses line + // terminators) and documents that the match ends at end-of-line. + const boldPattern = new RegExp(`^([ \\t]*\\*\\*${escaped}:\\*\\*[ \\t]*)(.*)$`, 'im'); if (boldPattern.test(content)) { return content.replace(boldPattern, (_match, prefix: string) => joinFieldReplacement(prefix, newValue)); } diff --git a/tests/state-document.test.cjs b/tests/state-document.test.cjs index 4a4ada130..9c397bdf2 100644 --- a/tests/state-document.test.cjs +++ b/tests/state-document.test.cjs @@ -221,6 +221,187 @@ describe('stateReplaceField — empty field preserves the following line (#4010) }); }); +// #4243: the bold branch of stateReplaceField was UNANCHORED +// (`(\*\*Field:\*\*[ \t]*)(.*)` with no ^ and no /m), so a bold label quoted +// MID-SENTENCE inside prose — the issue's `**Status:**` inside an Accumulated +// Context bullet — captured the rewrite and destroyed the rest of the line, +// silently, while the real field went stale or was updated elsewhere. The fix +// anchors the bold form to line start with same-line leading whitespace +// (`^([ \t]*\*\*Field:\*\*[ \t]*)`, 'im'), reusing #4010's same-line +// confinement idiom and #4186's recognition-by-anchoring discipline. These rows +// pin the corruption shapes; rows further down pin the negative space +// (legitimate line-start bold updates are byte-identical, branch order and +// first-occurrence-wins unchanged). +describe('stateReplaceField — anchored bold form leaves prose lookalikes untouched (#4243)', () => { + // The issue's verbatim prose line: a bold label quoted for documentation + // purposes inside a bullet, with the real field in the plain template form. + const ISSUE_PROSE_LINE = + '- [Phase 170]: archived files gained a `**Status:**Ready to execute` marker. Must not change.'; + + // ROW 1 — the failing-first regression from the issue. The lookalike must + // survive byte-identically and the REAL plain field must take the update. + test('issue repro: mid-sentence **Status:** lookalike survives, real plain field updates', () => { + const input = [ + '## Current Position', + '', + 'Phase: 5 of 9', + 'Plan: 2 of 6', + 'Status: Ready to execute', + 'Last activity: 2026-08-01 — did a thing', + '', + '## Accumulated Context', + '', + '### Decisions', + '', + ISSUE_PROSE_LINE, + '', + ].join('\n'); + const result = stateReplaceField(input, 'Status', 'Executing Phase 901'); + assert.notEqual(result, null, 'the real plain field must still match'); + assert.ok( + result.includes(ISSUE_PROSE_LINE), + `prose lookalike must survive byte-identically, got:\n${result}`, + ); + assert.ok( + /^Status: Executing Phase 901$/m.test(result), + 'the real plain Status line must take the update', + ); + assert.ok( + !result.includes('Executing Phase 901` marker'), + 'the rewrite must not bleed into the prose occurrence', + ); + }); + + test('lookalike ordered BEFORE the real bold field: prose survives, bold field updates', () => { + const input = [ + '## Accumulated Context', + '', + ISSUE_PROSE_LINE, + '', + '## Current Position', + '', + '**Status:** Ready to execute', + '', + ].join('\n'); + const result = stateReplaceField(input, 'Status', 'Executing Phase 901'); + assert.notEqual(result, null); + assert.ok(result.includes(ISSUE_PROSE_LINE), `prose lookalike must survive, got:\n${result}`); + assert.ok( + /^\*\*Status:\*\* Executing Phase 901$/m.test(result), + 'the real line-start bold field must take the update', + ); + }); + + test('mid-word lookalike with no real field: returns null (honest absence), never a rewrite', () => { + const input = 'Prose mentions text**Status:**tail mid-word and nothing else.'; + assert.equal(stateReplaceField(input, 'Status', 'Executing Phase 901'), null); + }); + + // A list-item bold label (`- **Status:** value`) is prose-shaped for the + // writer: no STATE.md writer emits body fields as list items, and treating + // a bullet as a field write target is exactly the #4243 corruption class. + // The read side's own vocabulary (bold anywhere) is untouched; the writer + // reports honest absence instead. + test('list-item bold label is not a write target: returns null, bullet untouched', () => { + const input = '- **Status:** resolved in the archived review'; + assert.equal(stateReplaceField(input, 'Status', 'new'), null); + }); + + // Every field the regex serves: a document whose ONLY occurrence of the + // label is a mid-sentence lookalike must yield null — no served field may + // be rewritten from prose. + test('mid-sentence lookalike yields null for every served field', () => { + const servedFields = [ + 'Status', 'Phase', 'Plan', 'Current Plan', 'Current Phase', 'Current Phase Name', + 'Last Activity', 'Last Activity Description', 'Total Phases', 'Total Plans in Phase', + 'Progress', 'Completed Phases', 'Stopped At', + ]; + for (const field of servedFields) { + const input = `Some prose sentence quoting a **${field}:** label mid-sentence, plus trailing words.`; + assert.equal( + stateReplaceField(input, field, 'NEW'), + null, + `mid-sentence **${field}:** lookalike must not match (got a rewrite)`, + ); + } + }); + + test('lookalike plus real plain field: only the real plain line changes (representative fields)', () => { + const cases = [ + { field: 'Status', plain: 'Status: Ready to execute' }, + { field: 'Phase', plain: 'Phase: 5 of 9' }, + { field: 'Last Activity', plain: 'Last Activity: 2026-08-01 — did a thing' }, + ]; + for (const { field, plain } of cases) { + const lookalike = `- notes: the **${field}:** label was archived here. Keep it.`; + const input = [plain, '', '## Accumulated Context', '', lookalike, ''].join('\n'); + const result = stateReplaceField(input, field, 'NEW VALUE'); + assert.notEqual(result, null, `${field}: real plain field must match`); + assert.ok( + result.includes(lookalike), + `${field}: lookalike line must survive byte-identically, got:\n${result}`, + ); + } + }); + + // Negative space: an INDENTED line-start bold field is still a field (the + // doc's form ranking reads bold anywhere in the section; the writer keeps + // same-line indentation writable), and the indent is preserved. + test('indented line-start bold field still updates, indent preserved', () => { + const input = ' **Status:** old'; + const result = stateReplaceField(input, 'Status', 'new'); + assert.equal(result, ' **Status:** new'); + }); + + // Negative space + fix-shape pin: leading blank lines before the label are + // NOT swallowed. The anchor's leading class is same-line whitespace only + // (`[ \t]*`, #4010's idiom); the issue's suggested `^\s*` variant would + // consume the newlines into the match and drop them on rebuild. + test('leading blank lines before a bold label survive byte-identically', () => { + const input = '\n\n**Status:** Ready'; + const result = stateReplaceField(input, 'Status', 'Executing Phase 5'); + assert.equal(result, '\n\n**Status:** Executing Phase 5'); + }); + + test('CRLF document: lookalike survives with CRLF intact, real plain field updates', () => { + const input = [ + 'Status: Ready to execute', + '', + '## Accumulated Context', + '', + ISSUE_PROSE_LINE, + '', + ].join('\r\n'); + const result = stateReplaceField(input, 'Status', 'Executing Phase 901'); + assert.notEqual(result, null); + assert.ok(result.includes(ISSUE_PROSE_LINE), `prose lookalike must survive, got:\n${result}`); + assert.ok(result.includes('\r\n'), 'CRLF endings must be preserved'); + assert.ok(/^Status: Executing Phase 901\r?$/m.test(result), 'real plain field must update'); + }); + + // Negative space: branch ORDER is unchanged — a line-start bold field still + // beats the plain form, and only the first bold occurrence is replaced. + test('line-start bold still beats the plain form (branch order unchanged)', () => { + const input = '**Status:** old bold\nStatus: old plain'; + const result = stateReplaceField(input, 'Status', 'new'); + assert.equal(result, '**Status:** new\nStatus: old plain'); + }); + + test('two line-start bold occurrences: only the first is replaced', () => { + const input = '**Status:** first\n**Status:** second'; + const result = stateReplaceField(input, 'Status', 'new'); + assert.equal(result, '**Status:** new\n**Status:** second'); + }); + + // #4010 same-line adjacency under the anchor: an empty bold field's value + // lands on its own line and the following line survives. + test('anchored bold branch keeps the #4010 empty-field boundary', () => { + const input = ' **Status:**\n **Current Plan:** 2 of 5'; + const result = stateReplaceField(input, 'Status', 'Executing Phase 5'); + assert.equal(result, ' **Status:** Executing Phase 5\n **Current Plan:** 2 of 5'); + }); +}); + describe('stateExtractField (#2880)', () => { test('extracts from a two-cell row', () => { const input = '| Current Phase | 3 |'; diff --git a/tests/state.test.cjs b/tests/state.test.cjs index 0e7daf716..e924f597b 100644 --- a/tests/state.test.cjs +++ b/tests/state.test.cjs @@ -3892,6 +3892,255 @@ Progress: [..........] 0% }); }); +// ───────────────────────────────────────────────────────────────────────────── +// #4243 — begin-phase: prose bold-lookalikes stay untouched (anchored bold +// form in stateReplaceField) and frontmatter round-trips unknown keys +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4243: begin-phase leaves prose lookalikes untouched, preserves unknown frontmatter', () => { + const ISSUE_PROSE_LINE = + '- [Phase 170]: archived files gained a `**Status:**Ready to execute` marker. Must not change.'; + + let tmpDir; + + beforeEach(() => { + tmpDir = createFixture(); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + // The issue's suggested regression fixture 1, verbatim shape: a bold + // `**Status:**` inside prose (## Accumulated Context), the real field in + // the plain template form. begin-phase must rewrite ONLY the real field. + test('issue fixture 1: prose **Status:** lookalike is byte-identical, real field updates', () => { + writeState(tmpDir, [ + '# Project State', + '', + '## Current Position', + 'Phase: 5 of 9 (Fifth)', + 'Plan: 2 of 6 in current phase', + 'Status: Ready to execute', + 'Last activity: 2026-08-01 — did a thing', + '', + 'Progress: [████░░░░░░] 40%', + '', + '## Accumulated Context', + '', + '### Decisions', + '', + ISSUE_PROSE_LINE, + '', + ].join('\n')); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok( + content.includes(ISSUE_PROSE_LINE), + `prose lookalike must survive begin-phase byte-identically, got:\n${content}`, + ); + const pos = sectionMatchOf(content, 'Current Position'); + assert.ok(pos, 'Current Position section should exist'); + assert.match(pos[1], /^Status: Executing Phase 901$/m); + }); + + // Corruption shape 1b: the lookalike section ordered BEFORE ## Current + // Position, real field in the bold form — first-match-in-document-order is + // the prose under the unanchored pattern. + test('lookalike before Current Position: prose survives, real bold field updates', () => { + writeState(tmpDir, [ + '# Project State', + '', + '## Accumulated Context', + '', + ISSUE_PROSE_LINE, + '', + '## Current Position', + '', + 'Phase: 5 of 9', + 'Plan: 2 of 6', + '**Status:** Ready to execute', + 'Last activity: 2026-08-01 — did a thing', + '', + ].join('\n')); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok( + content.includes(ISSUE_PROSE_LINE), + `prose lookalike must survive begin-phase byte-identically, got:\n${content}`, + ); + assert.match(content, /^\*\*Status:\*\* Executing Phase 901$/m); + }); + + // Corruption shape 1c: lookalikes for OTHER served fields — a mid-sentence + // `**Last Activity:**` must not capture the Last-activity refresh. + test('mid-sentence **Last Activity:** lookalike survives, real field refreshes', () => { + const lookalike = 'An earlier note mentions **Last Activity:** thresholds for archival. Keep.'; + writeState(tmpDir, [ + '# Project State', + '', + '## Current Position', + 'Phase: 5 of 9', + 'Plan: 2 of 6', + 'Status: Ready to execute', + 'Last activity: 2026-08-01 — did a thing', + '', + '## Accumulated Context', + '', + lookalike, + '', + ].join('\n')); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok(content.includes(lookalike), `lookalike must survive, got:\n${content}`); + const pos = sectionMatchOf(content, 'Current Position'); + assert.ok(pos, 'Current Position section should exist'); + assert.match(pos[1], /^Last activity: \d{4}-\d{2}-\d{2}/m); + }); + + // The issue's suggested regression fixture 2, no-ROADMAP arm: a custom + // frontmatter key, a populated progress block (with a custom subkey), a + // curated stopped_at, and milestone identity — all must survive begin-phase + // without a ROADMAP.md, and milestone/milestone_name must NOT be reset to + // invented defaults. (Already-correct behavior on next via the #2202 + // carry-forward + #3216 milestone-identity fix + the #4129 ratchet; pinned + // here so the class cannot regress.) + const FRONTMATTER_FIXTURE = [ + '---', + "gsd_state_version: '1.0'", + 'milestone: v2.1', + 'milestone_name: Real Curated Name', + 'status: planning', + "stopped_at: '2026-08-01 — curated stop note'", + 'custom_key: hand-added-by-agent', + 'progress:', + ' total_phases: 9', + ' completed_phases: 4', + ' total_plans: 30', + ' completed_plans: 12', + ' percent: 40', + ' custom_subkey: 77', + '---', + '', + '# Project State', + '', + '## Current Position', + 'Phase: 5 of 9 (Fifth)', + 'Plan: 2 of 6 in current phase', + 'Status: Ready to execute', + 'Last activity: 2026-08-01 — did a thing', + '', + ].join('\n'); + + test('issue fixture 2 (no ROADMAP): unknown keys, progress subkeys, stopped_at, milestone survive', () => { + writeState(tmpDir, FRONTMATTER_FIXTURE); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--name', 'Nine-Oh-One', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + const fm = frontmatterLib.extractFrontmatter(content); + assert.equal(fm['custom_key'], 'hand-added-by-agent', 'custom frontmatter key must survive'); + assert.equal(fm['milestone'], 'v2.1', 'milestone must not be reset to an invented default'); + assert.equal(fm['milestone_name'], 'Real Curated Name', 'curated milestone_name must survive'); + assert.ok(String(fm['stopped_at']).includes('curated stop note'), 'stopped_at must survive'); + const progress = fm['progress']; + assert.ok(progress && typeof progress === 'object', 'progress block must survive'); + // Numeric-tolerant: reconstructFrontmatter may serialize an unknown subkey + // as a quoted scalar, so it re-parses as a string — the VALUE surviving is + // the contract, not the YAML scalar shape. + assert.equal(Number(progress['custom_subkey']), 77, 'custom progress subkey must survive'); + // With ROADMAP.md absent and a milestone asserted, the #3573/#4094 withhold + // keeps all four counters at their STORED values (pinned below) and omits + // percent (an unmeasured scan must not assert one, #3233). `percent` is a + // DECLARED derived subkey governed by that recorded semantics — unlike the + // custom subkey above, its absence here is the documented behavior, so this + // row deliberately does not pin its value in either arm. + assert.equal(Number(progress['total_phases']), 9, 'stored total_phases kept under the #3573 withhold'); + assert.equal(Number(progress['completed_phases']), 4, 'stored completed_phases kept under the #3573 withhold'); + assert.equal(Number(progress['total_plans']), 30, 'stored total_plans kept under the #3573 withhold'); + assert.equal(Number(progress['completed_plans']), 12, 'stored completed_plans kept under the #3573 withhold'); + // Known keys still take the begin-phase update. + assert.equal(fm['status'], 'executing'); + assert.equal(String(fm['current_phase']), '901'); + }); + + test('issue fixture 2 (with ROADMAP): unknown keys and progress subkeys survive', () => { + writeState(tmpDir, FRONTMATTER_FIXTURE); + fs.writeFileSync( + path.join(tmpDir, '.planning', 'ROADMAP.md'), + ['# Roadmap', '', '## v2.1 — Real Curated Name', '', '### Phase 5: Fifth', '', 'complete.', ''].join('\n'), + ); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--name', 'Nine-Oh-One', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + const fm = frontmatterLib.extractFrontmatter(content); + assert.equal(fm['custom_key'], 'hand-added-by-agent', 'custom frontmatter key must survive'); + assert.equal(fm['milestone'], 'v2.1'); + assert.equal(fm['milestone_name'], 'Real Curated Name'); + const progress = fm['progress']; + assert.ok(progress && typeof progress === 'object', 'progress block must survive'); + assert.equal(Number(progress['custom_subkey']), 77, 'custom progress subkey must survive'); + assert.equal(fm['status'], 'executing'); + }); + + test('milestone without milestone_name, no ROADMAP: no invented identity appears', () => { + writeState(tmpDir, [ + '---', + "gsd_state_version: '1.0'", + 'milestone: v2.1', + 'custom_key: keep-me', + '---', + '', + '# Project State', + '', + '## Current Position', + 'Phase: 5 of 9', + 'Plan: 2 of 6', + 'Status: Ready to execute', + '', + ].join('\n')); + + const result = runGsdTools( + ['state', 'begin-phase', '--phase', '901', '--plans', '3'], + tmpDir, + ); + assert.ok(result.success, `begin-phase failed: ${result.error}`); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + const fm = frontmatterLib.extractFrontmatter(content); + assert.equal(fm['milestone'], 'v2.1', 'milestone must survive without a ROADMAP'); + assert.equal(fm['milestone_name'], undefined, 'no fabricated milestone_name may appear'); + assert.equal(fm['custom_key'], 'keep-me'); + }); +}); + // ───────────────────────────────────────────────────────────────────────────── // Bug #1589 — progress counters not updated during plan execution // ───────────────────────────────────────────────────────────────────────────── From 03c770be47be363343a62a74f8a4ae278e5d7fc2 Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 00:52:49 -0400 Subject: [PATCH 032/166] docs(#4463): record executor self-repair of worktree base as out-of-scope (#4470) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Denies reintroducing sub-agent-side `git reset --hard` recovery for a worktree base mismatch — the exact primitive #48 removed for safety. Keeps the fork-base measurement finding for future reference. Co-authored-by: sim Co-authored-by: Claude Sonnet 5 --- .../executor-self-repair-worktree-base.md | 86 +++++++++++++++++++ 1 file changed, 86 insertions(+) create mode 100644 .out-of-scope/executor-self-repair-worktree-base.md diff --git a/.out-of-scope/executor-self-repair-worktree-base.md b/.out-of-scope/executor-self-repair-worktree-base.md new file mode 100644 index 000000000..fa85c91e3 --- /dev/null +++ b/.out-of-scope/executor-self-repair-worktree-base.md @@ -0,0 +1,86 @@ +# Executor self-repair of a worktree base mismatch via `git reset --hard` + +**Source:** [#4463](https://github.com/open-gsd/gsd-core/issues/4463) +**Decision:** wontfix — No-go as filed; conflicts with the shipped #48 design +**Date:** 2026-09-07 + +## Proposal summary + +Reporter ran a controlled probe on Claude Code and reported two findings, plus a proposal: + +1. **Measurement:** the harness worktree it was dispatched into forked from the + *orchestrator session's HEAD* (the main checkout's current branch tip), not from + `origin/HEAD` as the `#3659`/`#3779`/`#48` family assumed. The reporter is explicit + that this does not mean the `worktree.base-check` degrade is wrong in the case they + observed — the stated *reason* for the degrade may be inaccurate even where the + degrade itself is still warranted. +2. **Measurement:** because a linked worktree shares the common git object/ref store, + the phase branch is reachable as a local ref with no `git fetch`, and + `git reset --hard ` on the executor's own branch is cheap, does not + switch or detach HEAD, and does not conflict with "branch already checked out + elsewhere" (that check only fires on `git checkout`, not on moving one's own branch + pointer via `reset`). +3. **Proposal:** where `worktree-branch-check` currently halts with `exit 42` on a base + mismatch, let the executor **repair itself** — `git reset --hard ` + on its own branch — and proceed, converting the halt into a self-heal. + +## Why GSD does not own this + +- **This is the exact primitive #48 removed, for the exact failure mode #48 was filed + to fix.** #48 (closed, `approved-enhancement`, shipped) replaced sub-agent-side + `git reset --hard` recovery with a verify-only, fail-closed `exit 42` check, specifically + because (a) a permission deny-rule on `git reset --hard*` — common in safety-conscious + host configurations — can make the recovery command itself fail, and depending on shell + error handling the sub-agent may silently proceed on the wrong base or report success + without re-verifying; and (b) a sub-agent should not hold state-correction primitives + (`reset`, force-move, branch-switch) on a worktree it did not create — that + responsibility belongs to the lifecycle owner, the orchestrator. #4463's proposal + reintroduces precisely this: the executor mutating its own worktree state in response + to a detected mismatch, on the sub-agent side. +- **The shipped design is live in current source, not just historically decided.** + `gsd-core/workflows/execute-phase/steps/worktree-recovery-policy.md:7` states, verbatim: + *"`worktree_branch_check` is verify-only — an executor that hits a base/HEAD-namespace + mismatch prints `FATAL:` and exits **42** instead of self-recovering... The orchestrator — + the worktree lifecycle owner — performs any base correction... the sub-agent never does."* + This is an active architectural invariant, not stale rationale from a closed issue. +- **The "it costs almost nothing and nothing objects" framing is exactly what #48 warned + about.** #4463's own Finding 3 confirms `git reset --hard` ran with no hook and no + refusal in its probe — which is the *absence of a safety net* #48 is trying to compensate + for by moving the responsibility off the sub-agent entirely, not evidence the operation + is safe to grant back to it. + +## What this does NOT cover + +- **The fork-base measurement itself is not denied and is worth keeping.** If the + harness-worktree fork base is genuinely the orchestrator session's HEAD rather than + `origin/HEAD`, that is new information relevant to `#3659`'s and `#3779`'s closure + rationale and to how `worktree.base-check`'s degrade condition is described. A follow-up + that only re-verifies and documents this measurement (with the session cwd inside a + feature worktree, which the original probe did not test) is not this proposal and is + welcome. +- **Orchestrator-side repair is a different proposal.** #48's split explicitly assigns + base correction to the orchestrator (e.g., recreate the worktree on `{EXPECTED_BASE}`, + or fast-forward the branch from the orchestrator side before dispatch). A proposal that + moves the repair step to the orchestrator, before or around dispatch, rather than having + the executor self-repair after detecting a mismatch, is not denied here and would need + its own review. +- **Fixing `#4415`** (cleanup-wave blocking when Claude Code has already removed the + executor's worktree) is unrelated and not affected by this decision. + +## Re-open criteria + +- A proposal that keeps repair on the **orchestrator** side of the #48 split (the + lifecycle owner), not the sub-agent side — matching, not reversing, the shipped + architecture. +- Or: a demonstrated, host-enforced guarantee that the executor's `git reset --hard` call + cannot be silently denied or misreported by a permission policy — removing the specific + failure mode #48 was filed against. Absent that guarantee, granting the primitive back + to the sub-agent reintroduces the original risk regardless of how cheap the operation is + in the success case. + +## Related + +- [#48](https://github.com/open-gsd/gsd-core/issues/48) — the decision this proposal reverses +- `gsd-core/workflows/execute-phase/steps/worktree-recovery-policy.md` — the shipped fail-closed invariant +- [#3659](https://github.com/open-gsd/gsd-core/issues/3659), [#3779](https://github.com/open-gsd/gsd-core/issues/3779), [#683](https://github.com/open-gsd/gsd-core/issues/683) — the fork-base family #4463's measurement bears on +- [#4415](https://github.com/open-gsd/gsd-core/issues/4415) — adjacent cleanup-wave gap, unaffected by this decision From c4b6dbd486e9bb802132088098f24cd5d05db3ac Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 02:20:47 -0400 Subject: [PATCH 033/166] fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468) * test(#4247): failing-first regressions for checklist-form update-plan-progress * fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry * fix(#4247): single local source for the phase-heading anchor grammar * docs(#4247): note the missing_phase_details refusal in cli-tools reference * docs(#4247): backfill pr number in changeset --------- Co-authored-by: sim --- .changeset/brave-finches-zip.md | 5 + docs/CLI-TOOLS.md | 6 + src/roadmap.cts | 86 +++++++- tests/roadmap.test.cjs | 358 ++++++++++++++++++++++++++++++++ 4 files changed, 448 insertions(+), 7 deletions(-) create mode 100644 .changeset/brave-finches-zip.md diff --git a/.changeset/brave-finches-zip.md b/.changeset/brave-finches-zip.md new file mode 100644 index 000000000..bcbc31548 --- /dev/null +++ b/.changeset/brave-finches-zip.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4468 +--- +**`roadmap update-plan-progress` no longer false-greens on checklist-form ROADMAPs** — a phase whose entry is a `- [ ] **Phase N: …**` checklist bullet with no writable Progress-table row or detail section now declines with `updated: false` and a typed `missing_phase_details` reason, leaving ROADMAP.md byte-identical, instead of reporting success off an unrelated checkbox mark while the phase row stayed untouched and blank lines were injected mid-sentence in other phases' entries. (#4247) diff --git a/docs/CLI-TOOLS.md b/docs/CLI-TOOLS.md index ff0925d1a..027867288 100644 --- a/docs/CLI-TOOLS.md +++ b/docs/CLI-TOOLS.md @@ -434,6 +434,12 @@ node gsd-tools.cjs roadmap analyze node gsd-tools.cjs roadmap update-plan-progress ``` +When the phase has no writable ROADMAP entry — no matching Progress-table row, +no `### Phase N` detail section, and no checklist bullet this command can update +(the checklist-only form) — the command declines with `updated: false` and a +`missing_phase_details` reason instead of claiming success, and leaves +`ROADMAP.md` byte-identical. + ### Milestone window scope (`roadmap analyze`) `roadmap analyze` scopes its phase list to the current milestone's section of diff --git a/src/roadmap.cts b/src/roadmap.cts index afea0625e..de59dae3e 100644 --- a/src/roadmap.cts +++ b/src/roadmap.cts @@ -1032,6 +1032,23 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und // Wrap entire read-modify-write in lock to prevent concurrent corruption let updated = false; + // #4247: the refusal flag. The write/report decision below must be keyed to + // "a writable roadmap representation of THIS phase was found", never to "any + // byte moved". On a checklist-form ROADMAP (`- [ ] **Phase N: …**`, the + // roadmapper's own summary-checklist form) every phase-targeted grammar + // below requires an ATX `#{2,4} Phase N` heading and therefore finds + // nothing; the Progress-table row is the only other writable target, and + // when its Phase cell does not match `phaseCellRe` (e.g. a word-prefixed + // `Phase 68` cell — deliberately unrecognized on the read side too, + // `deriveProgressFromRoadmap`'s `/^\d/` data-row filter) the command used + // to fall through to unrelated byte deltas (an UN-scoped plan-checkbox mark + // anywhere in the document) and report `updated: true` while the phase's + // own row stayed untouched — with the file-global write then letting the + // platform write seam's markdown normalization inject blank lines around + // other phases' bullets, splitting hand-wrapped sentences mid-entry. The + // refusal below declines with the analyzer's own `missing_phase_details` + // vocabulary and leaves ROADMAP.md byte-identical. + let missingPhaseDetails = false; withPlanningLock(cwd, () => { // #3957 (B9.4): captured BEFORE any transform runs, so the write/report // decision below reflects whether the transforms actually changed @@ -1041,6 +1058,35 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und const originalContent = fs.readFileSync(roadmapPath, 'utf-8'); let roadmapContent = originalContent; const phasePattern = phaseMarkdownRegexSource(phaseNum); + // #4247: ONE local source for the ATX phase-heading anchor that every + // section-scoped writer below (`planCountPattern`, + // `insertRowsPatternA|B`) starts with — extracted so the target-detection + // gate below reads the SAME grammar the writers anchor on, and a future + // edit to one cannot drift from the other three copies. + const phaseHeadingAnchor = `#{2,4}\\s*Phase\\s+${phasePattern}${OPTIONAL_PHASE_TAG_SOURCE}(?=[:\\s])`; + // #4247: target detection runs against the ORIGINAL content's active + // (post-) region — the same milestone scoping every writer + // below applies — so the gate asks "does the file carry a writable phase + // representation" rather than "did some regex fire mid-transform". + const gateDetailsClose = originalContent.lastIndexOf(''); + const gateActiveRegion = gateDetailsClose === -1 + ? originalContent + : originalContent.slice(gateDetailsClose + ''.length); + // Heading target: the exact grammar `planCountPattern` / + // `insertRowsPatternA|B` anchor on (an ATX phase heading for this phase). + const headingTargetFound = new RegExp(phaseHeadingAnchor, 'i').test(gateActiveRegion); + // Checklist target: when the phase is complete, its own checklist bullet + // (`- [ ] **Phase N: …**`) IS a writable phase row — the completion + // checkbox stamp below updates it. Same grammar as that writer, widened + // one notch to `[ x]` so an ALREADY-checked bullet still counts as a + // found target: an idempotent re-run then takes the honest + // "no changes were needed" decline instead of this refusal. + const checklistTargetFound = isComplete && new RegExp( + `-\\s*\\[[ x]\\]\\s*.*Phase\\s+${phasePattern}${OPTIONAL_PHASE_TAG_SOURCE}[:\\s]`, + 'i', + ).test(gateActiveRegion); + // Table-row target: set by the row-scoped cell updates below. + let tableRowFound = false; // Progress table row: update Plans Complete/Status/Completed columns BY // COLUMN NAME (handles 4- or 5-column RoadmapProgress tables regardless of @@ -1062,10 +1108,10 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und let text = scoped; const plansResult = updateTableCell(text, rowMatch, 'Plans Complete', ` ${summaryCount}/${planCount} `); - if (plansResult.ok) text = plansResult.value; + if (plansResult.ok) { text = plansResult.value; tableRowFound = true; } const statusResult = updateTableCell(text, rowMatch, 'Status', ` ${status.padEnd(11)}`); - if (statusResult.ok) text = statusResult.value; + if (statusResult.ok) { text = statusResult.value; tableRowFound = true; } // Preserve only a valid ISO date (#1161: idempotent; self-heal garbage). // Ragged-tolerant (#2245 Blocker 2): probe the CURRENT Completed cell via @@ -1081,7 +1127,7 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und } return ' '; }); - if (completedResult.ok) text = completedResult.value; + if (completedResult.ok) { text = completedResult.value; tableRowFound = true; } return text; }); @@ -1125,7 +1171,7 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und // continuation on the next line, since the pattern never spans past // `\n` in the first place. const planCountPattern = new RegExp( - `(#{2,4}\\s*Phase\\s+${phasePattern}${OPTIONAL_PHASE_TAG_SOURCE}(?=[:\\s])(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:\\*\\*Plans\\*\\*:|\\*\\*Plans:\\*\\*|(?:^|\\n)Plans:)\\s*)(\\d+\\s*\\/\\s*\\d+\\s+plans(?:\\s+(?:complete|executed))?|\\d+\\s+plans?)?([^\\r\\n]*)`, + `(${phaseHeadingAnchor}(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:\\*\\*Plans\\*\\*:|\\*\\*Plans:\\*\\*|(?:^|\\n)Plans:)\\s*)(\\d+\\s*\\/\\s*\\d+\\s+plans(?:\\s+(?:complete|executed))?|\\d+\\s+plans?)?([^\\r\\n]*)`, 'i' ); const planCountText = isComplete @@ -1218,11 +1264,11 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und // Pattern A: anchor to bare `Plans:` header (preferred). // Pattern B: fallback to bold summary when no bare header exists. const insertRowsPatternA = new RegExp( - `(#{2,4}\\s*Phase\\s+${phasePattern}${OPTIONAL_PHASE_TAG_SOURCE}(?=[:\\s])(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:^|\\n)(?:Plans:)[^\\n]*)`, + `(${phaseHeadingAnchor}(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:^|\\n)(?:Plans:)[^\\n]*)`, 'i' ); const insertRowsPatternB = new RegExp( - `(#{2,4}\\s*Phase\\s+${phasePattern}${OPTIONAL_PHASE_TAG_SOURCE}(?=[:\\s])(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:\\*\\*Plans\\*\\*:|\\*\\*Plans:\\*\\*)[^\\n]*)`, + `(${phaseHeadingAnchor}(?:(?!\\n#{1,4}\\s)[\\s\\S])*?(?:\\*\\*Plans\\*\\*:|\\*\\*Plans:\\*\\*)[^\\n]*)`, 'i' ); @@ -1268,10 +1314,24 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und // `cmdRoadmapAnnotateDependencies`'s existing `nextContent !== content` // gate. Previously this wrote and reported `updated: true` // unconditionally, even on an idempotent re-run that changed nothing. - if (roadmapContent !== originalContent) { + // + // #4247: ...but ONLY when a writable representation of THIS phase was + // found. Without the target gate, a byte delta from an unrelated + // transform (the un-scoped plan-checkbox mark) satisfied the #3957 gate + // and produced a success-shaped `updated: true` while the phase's own + // row stayed untouched — and the file-global write let the platform + // write seam's markdown normalization reflow unrelated entries. When no + // target exists the command refuses: no write at all, so ROADMAP.md is + // left byte-identical, and the caller gets a typed + // `missing_phase_details` decline instead of a false green. + const phaseRepresentationFound = tableRowFound || headingTargetFound || checklistTargetFound; + if (phaseRepresentationFound && roadmapContent !== originalContent) { platformWriteSync(roadmapPath, roadmapContent); updated = true; } + if (!phaseRepresentationFound) { + missingPhaseDetails = true; + } }); const computed = { @@ -1284,6 +1344,18 @@ function cmdRoadmapUpdatePlanProgress(cwd: string, phaseNum: string | null | und }; if (updated) { output({ updated: true, ...computed }, raw, `${summaryCount}/${planCount} ${status}`); + } else if (missingPhaseDetails) { + // #4247: honest refusal — the reason names the real condition (the + // analyzer's `missing_phase_details` vocabulary), never "already + // reflects", which was false: the ROADMAP was never able to record this + // phase's progress in the first place. + declineNoOp( + raw, + 'updated', + 'missing_phase_details', + `roadmap update-plan-progress skipped — ROADMAP.md has no writable entry for phase ${formatDiagnosticToken(String(phaseNum))} (no matching Progress-table row, no phase detail section, and no checklist entry this command can update). ROADMAP.md was left unchanged.`, + computed, + ); } else { declineNoOp( raw, diff --git a/tests/roadmap.test.cjs b/tests/roadmap.test.cjs index 1d4494bbe..de1ecf3fc 100644 --- a/tests/roadmap.test.cjs +++ b/tests/roadmap.test.cjs @@ -1269,6 +1269,364 @@ describe('#3057 B3: roadmap update-plan-progress — verification staleness-chec }); }); +// ───────────────────────────────────────────────────────────────────────────── +// regressions: #4247 — checklist-form ROADMAP.md must not produce a false +// `updated:true` (untouched phase row) nor blank-line corruption elsewhere +// ───────────────────────────────────────────────────────────────────────────── + +/** + * The checklist house style the bundled roadmapper emits for the summary + * checklist (`agents/gsd-roadmapper.md` §"Summary Checklist"): long + * `- [ ] **Phase N: Title** - description` entries with column-0 continuation + * sentences (the shape issue #4247 was filed against), plus a Progress table. + * + * `tableCell` controls the Progress table's Phase cell for the target phase. + */ +function buildChecklistRoadmap4247(tableCell) { + return [ + '# Roadmap: Mango Tree', + '', + '## Phases', + '', + '- [ ] **Phase 65: Orchard Layout** - goal: design the canopy grid. Progress: 4/4 plans executed, verified 2026-08-30;', + 'grid survey closed the two-centimeter tolerance, terracing passed inspection, and the irrigation', + 'channels were flushed before the storm; soil probes re-zeroed afterwards.', + '- [ ] **Phase 68: Scheduler** - goal: ship the picking scheduler. Progress: 1/5 plans executed;', + '68-01 calendar model in review; 68-02 crew assignment summarized; 68-03 weather windows', + 'not started; 68-04 crate logistics blocked on warehouse slot; 68-05 billing hook pending.', + '- [ ] **Phase 69: Packhouse** - goal: automate the line. Progress: 0/2 executed.', + '', + '## Execution Waves', + '', + '- Wave 1: Phases 65-66 — complete.', + '- Wave 2: Phase 68 — in flight.', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 65 | 4/4 | Complete | 2026-08-30 |', + `| ${tableCell} | 0/5 | Planned | - |`, + '', + ].join('\n'); +} + +/** Phase 68 on disk: 5 plans, 1 summary → In Progress, 1/5. */ +function seedPhase68WithPlans(tmpDir, { roadmap } = {}) { + fs.writeFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), roadmap); + const p68 = path.join(tmpDir, '.planning', 'phases', '68-scheduler'); + fs.mkdirSync(p68, { recursive: true }); + for (const n of ['01', '02', '03', '04', '05']) { + fs.writeFileSync(path.join(p68, `68-${n}-PLAN.md`), `# Plan ${n}\n`); + } + fs.writeFileSync(path.join(p68, '68-02-SUMMARY.md'), '# Summary\n'); + return p68; +} + +describe('#4247: roadmap update-plan-progress — checklist-form ROADMAP refuses instead of false-green', () => { + let tmpDir; + let roadmapPath; + + beforeEach(() => { + tmpDir = createTempProject('gsd-4247-roadmap-'); + roadmapPath = path.join(tmpDir, '.planning', 'ROADMAP.md'); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + // Row 1 — the issue's exact repro shape. FAILING FIRST. + test('checklist form with non-matching table row and a stray plan row refuses with missing_phase_details and writes nothing', () => { + const roadmap = buildChecklistRoadmap4247('Phase 68').replace( + '- [ ] **Phase 69: Packhouse**', + ' - [ ] 68-01: calendar model\n - [ ] 68-02: crew assignment\n- [ ] **Phase 69: Packhouse**', + ); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, false, 'a phase with no writable roadmap representation must not claim updated'); + assert.strictEqual(output.reason, 'missing_phase_details', 'refusal carries the typed parse diagnostic'); + assert.strictEqual(output.plan_count, 5, 'computed counts are still reported'); + assert.strictEqual(output.summary_count, 1); + + // The issue's round-trip requirement: the file must be byte-identical — + // no plan-checkbox marks, no blank lines splitting other phases' sentences. + const written = fs.readFileSync(roadmapPath, 'utf-8'); + assert.strictEqual(written, roadmap, 'ROADMAP.md must be byte-identical on refusal'); + }); + + // Row 2 — no table at all. + test('checklist form with no Progress table refuses with missing_phase_details and writes nothing', () => { + const roadmap = [ + '# Roadmap: T', + '', + '## Phases', + '', + '- [ ] **Phase 68: Scheduler** - goal: ship the picking scheduler.', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, false); + assert.strictEqual(output.reason, 'missing_phase_details'); + assert.strictEqual( + fs.readFileSync(roadmapPath, 'utf-8'), + roadmap, + 'ROADMAP.md must be byte-identical on refusal', + ); + }); + + // Row 3 — isolates the non-matching cell (word-prefixed) from the plan-row trigger. + test('checklist form whose table row does not match the phase-cell grammar refuses without writing', () => { + const roadmap = buildChecklistRoadmap4247('Phase 68'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, false); + assert.strictEqual(output.reason, 'missing_phase_details'); + assert.strictEqual( + fs.readFileSync(roadmapPath, 'utf-8'), + roadmap, + 'ROADMAP.md must be byte-identical on refusal', + ); + }); + + // Row 4 — table-form MUST keep updating exactly as today (full-file byte assertion). + test('table form with a bare-number cell still updates the row byte-exactly and touches nothing else', () => { + // Normal-form fixture (template spacing, single-line entries): the write + // seam's markdown normalization is a no-op here, so the byte-exact splice + // is observable. + const roadmap = [ + '# Roadmap: Mango Tree', + '', + '## Phases', + '', + '- [ ] **Phase 65: Orchard Layout** - goal: design the canopy grid. Progress: 4/4 plans executed, verified 2026-08-30.', + '', + '- [ ] **Phase 68: Scheduler** - goal: ship the picking scheduler. Progress: 1/5 plans executed.', + '', + '- [ ] **Phase 69: Packhouse** - goal: automate the line. Progress: 0/2 executed.', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 65 | 4/4 | Complete | 2026-08-30 |', + '| 68 | 0/5 | Planned | - |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, true, 'a matching table row is a writable target'); + + // Byte-exact expectation: only the Phase 68 row's three cells change + // (` 1/5 ` splice, ` In Progress` padEnd(11), Completed cleared to ` `), + // every other byte — including Phase 65/69 prose and the waves section — + // is untouched (no blank-line insertion anywhere outside the row). + const expected = roadmap.replace( + '| 68 | 0/5 | Planned | - |', + '| 68 | 1/5 | In Progress| |', + ); + assert.strictEqual(fs.readFileSync(roadmapPath, 'utf-8'), expected); + }); + + // Row 5 — 5-column milestone-grouped table, `68. [Scheduler]` cell form. + test('milestone-grouped table form with template cell still updates by column name', () => { + const roadmap = [ + '# Roadmap: T', + '', + '## Progress', + '', + '| Phase | Milestone | Plans Complete | Status | Completed |', + '|-------|-----------|----------------|--------|-----------|', + '| 67. [Waves] | v1.0 | 0/2 | Planned | - |', + '| 68. [Scheduler] | v1.0 | 0/5 | Planned | - |', + '| 69. [Packhouse] | v1.0 | 0/2 | Planned | - |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, true); + + const expected = roadmap.replace( + '| 68. [Scheduler] | v1.0 | 0/5 | Planned | - |', + '| 68. [Scheduler] | v1.0 | 1/5 | In Progress| |', + ); + assert.strictEqual(fs.readFileSync(roadmapPath, 'utf-8'), expected); + }); + + // Row 6/7 — heading-form targets keep every existing behavior. + test('heading-form detail section still updates counts, inserts rows, and marks checkboxes', () => { + const roadmap = [ + '# Roadmap: T', + '', + '## Phases', + '', + '- [ ] **Phase 68: Scheduler** - goal: ship scheduler.', + '', + '### Phase 68: Scheduler', + '**Goal**: ship it', + '**Plans**: 5 plans', + '', + 'Plans:', + '- [ ] 68-02: crew assignment', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| Phase 68 | 0/5 | Planned | - |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, true, 'a heading target is writable — no refusal'); + + const written = fs.readFileSync(roadmapPath, 'utf-8'); + assert.ok(written.includes('**Plans**: 1/5 plans executed'), 'count token rewritten'); + assert.ok(written.includes('- [x] 68-02: crew assignment'), 'summarized plan checkbox marked'); + assert.ok(written.includes('- [ ] 68-01-PLAN.md'), 'missing plan rows inserted'); + }); + + // Row 9/10/11 — boundaries inside a table must not trip the refusal. + test('first and last table rows still update; adjacent rows stay untouched', () => { + for (const cell of ['68', 'Phase 68']) { + const roadmap = [ + '# Roadmap: T', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 67 | 0/2 | Planned | - |', + `| ${cell === '68' ? '68' : 'Phase 68'} | 0/5 | Planned | - |`, + '| 69 | 0/2 | Planned | - |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + const output = JSON.parse(result.output); + const written = fs.readFileSync(roadmapPath, 'utf-8'); + assert.ok(written.includes('| 67 | 0/2 | Planned | - |'), 'adjacent row untouched'); + assert.ok(written.includes('| 69 | 0/2 | Planned | - |'), 'adjacent row untouched'); + if (cell === '68') { + assert.strictEqual(output.updated, true, 'bare cell 68 between siblings updates'); + assert.ok(written.includes('| 68 | 1/5 |'), 'middle row updated'); + } else { + assert.strictEqual(output.updated, false, 'word-prefixed cell is not a writable target'); + assert.strictEqual(written, roadmap, 'nothing written on refusal'); + } + cleanup(tmpDir); + tmpDir = createTempProject('gsd-4247-roadmap-'); + roadmapPath = path.join(tmpDir, '.planning', 'ROADMAP.md'); + } + }); + + // Row 12 — phase exists only on disk; roadmap has no representation at all. + test('phase absent from the roadmap entirely refuses with missing_phase_details', () => { + const roadmap = [ + '# Roadmap: T', + '', + '## Phases', + '', + '- [ ] **Phase 65: Layout** - goal design.', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 65 | 4/4 | Complete | 2026-08-30 |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, false); + assert.strictEqual(output.reason, 'missing_phase_details'); + assert.strictEqual(fs.readFileSync(roadmapPath, 'utf-8'), roadmap); + }); + + // Row 13 — the #3957 honest no-op decline must survive for target-found reruns. + test('idempotent re-run on an up-to-date table row keeps the honest no-changes decline', () => { + const roadmap = [ + '# Roadmap: T', + '', + '## Progress', + '', + '| Phase | Plans Complete | Status | Completed |', + '|-------|----------------|--------|-----------|', + '| 68 | 1/5 | In Progress| |', + '', + ].join('\n'); + seedPhase68WithPlans(tmpDir, { roadmap }); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.updated, false); + assert.notStrictEqual(output.reason, 'missing_phase_details', 'target found — this is a genuine no-op'); + assert.ok(String(output.reason).includes('no changes were needed'), 'preserves the #3957 decline vocabulary'); + assert.strictEqual(fs.readFileSync(roadmapPath, 'utf-8'), roadmap); + }); + + // Row 8 — the completion arm: the checklist entry's own checkbox IS the + // writable phase row for a complete phase. + test('checklist form completing the phase still checks the phase bullet and marks its plan rows', () => { + const roadmap = buildChecklistRoadmap4247('Phase 68').replace( + '- [ ] **Phase 69: Packhouse**', + ' - [ ] 68-01: calendar model\n - [ ] 68-02: crew assignment\n- [ ] **Phase 69: Packhouse**', + ); + const p68 = seedPhase68WithPlans(tmpDir, { roadmap }); + for (const n of ['01', '03', '04', '05']) { + fs.writeFileSync(path.join(p68, `68-${n}-SUMMARY.md`), '# Summary\n'); + } + fs.writeFileSync(path.join(p68, '68-VERIFICATION.md'), '---\nstatus: passed\n---\n# Verification\n'); + + const result = runGsdTools('roadmap update-plan-progress 68', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + + const output = JSON.parse(result.output); + assert.strictEqual(output.complete, true); + assert.strictEqual(output.updated, true, 'the phase bullet is a writable target when completing'); + + const written = fs.readFileSync(roadmapPath, 'utf-8'); + assert.match(written, /- \[x\] \*\*Phase 68: Scheduler\*\* - goal: ship the picking scheduler\. Progress: 1\/5 plans executed; \(completed \d{4}-\d{2}-\d{2}\)/, 'phase bullet checked with completion date'); + // Other phases' checklist bullets are not checked or annotated. + assert.match(written, /- \[ \] \*\*Phase 65: Orchard Layout\*\*/); + assert.match(written, /- \[ \] \*\*Phase 69: Packhouse\*\*/); + assert.ok(written.includes('- [x] 68-02: crew assignment'), 'summarized plan row marked'); + }); +}); + // ───────────────────────────────────────────────────────────────────────────── // phase add command // ───────────────────────────────────────────────────────────────────────────── From 6ebe6372ce2e24005ba2d708803c9ced7863da4d Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 04:27:35 -0400 Subject: [PATCH 034/166] fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#4243): anchor stateReplaceProgressPercent bold form to line start The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag, so a bold percent-ish label quoted MID-SENTENCE inside prose — an Accumulated Context bullet mentioning **Progress:** — captured the machine-segment rewrite and destroyed the rest of its line, silently, while the real Progress line stayed stale (and the frontmatter moved on without it, breaking the #4213 surfaces-agree contract). Every caller (cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation) feeds the whole document, so all three were exposed. Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the exact idiom #4453 applied to stateReplaceField's bold branch (same-line confinement per #4010: the leading class is [ \t]*, deliberately not \s*, which can consume the newlines before the label into the match; $ is explicit-and-inert and documents end-of-line). #2177's recorded requirements all stand: frontmatter is stripped before matching, the suffix-preserving machine-segment swap is untouched, and bold-beats-plain priority now governs line-start forms, so an earlier free-text plain Progress: line still cannot capture the rewrite ahead of the real bold status line. Per the maintainer ruling (2026-09-07), #2177's incidental bold-anywhere matching was not load-bearing. * test(#4243): scope the C4 region check with splitLines, not a bare \n split lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's content.split(/\n## /)[0] — a bare \n split on readFileSync content is CRLF-fragile under Windows autocrlf. Same scoping via splitLines() (src/text-lines.cts), which splits on \r?\n. * chore(#4243): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/serene-foxes-hum.md | 5 + src/state-transition.cts | 28 ++++- tests/state-transition.test.cjs | 209 ++++++++++++++++++++++++++++++++ tests/state.test.cjs | 198 ++++++++++++++++++++++++++++++ 4 files changed, 435 insertions(+), 5 deletions(-) create mode 100644 .changeset/serene-foxes-hum.md diff --git a/.changeset/serene-foxes-hum.md b/.changeset/serene-foxes-hum.md new file mode 100644 index 000000000..38f5ec9ff --- /dev/null +++ b/.changeset/serene-foxes-hum.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4474 +--- +**progress-percent bold fields no longer rewrite mid-sentence lookalikes** — anchored to line-start like #4243's stateReplaceField fix. (#4243 follow-up; supersedes the #2177 bold-anywhere reading per maintainer ruling) diff --git a/src/state-transition.cts b/src/state-transition.cts index e001c4d98..699190d61 100644 --- a/src/state-transition.cts +++ b/src/state-transition.cts @@ -55,11 +55,29 @@ export function formatProgressMachineSegment(percent: number): string { // `syncCore`'s call here. export function stateReplaceProgressPercent(content: string, percent: number): string | null { const body = stripFrontmatter(content); - // #2177: bold `**Progress:**` anywhere in the body wins outright; the plain - // `^Progress:` form is the fallback only when no bold line exists, so an - // earlier free-text line starting with `Progress:` cannot capture the - // rewrite ahead of the real status line. - const boldProgressPattern = /(\*\*Progress:\*\*[ \t]*)([^\r\n]*)/i; + // #2177: bold `**Progress:**` takes priority over the plain `^Progress:` + // form, so an earlier free-text line starting with `Progress:` cannot + // capture the rewrite ahead of the real status line. + // + // #4243 (follow-up to #4453, maintainer ruling 2026-09-07): the bold form + // is also ANCHORED to line start, with same-line leading whitespace only — + // the exact idiom #4453 applied to stateReplaceField's bold branch. The + // pre-fix pattern carried no `^` and no `m` flag, so a bold percent-ish + // label quoted MID-SENTENCE inside prose (an Accumulated Context bullet + // mentioning `**Progress:**`) captured the machine-segment rewrite and + // destroyed the rest of its line, silently, while the real Progress line + // stayed stale — every caller (cmdStateUpdateProgress, syncCore's percent + // arm, applyPostSyncPreservation) feeds the whole document. #2177's own + // recorded requirements are unaffected: the frontmatter is stripped before + // matching (its defect was the YAML `progress:` key shadowing the body + // line), the suffix-preserving machine-segment swap is untouched, and the + // bold-beats-plain priority now governs LINE-START forms. The leading class + // is `[ \t]*`, deliberately NOT `\s*` — `^\s*\*\*` can consume the newlines + // before the label into the match and drop them on rebuild (#4010's + // same-line confinement hazard). `$` is explicit-and-inert (`[^\r\n]*` + // never crosses line terminators) and documents that the match ends at + // end-of-line. + const boldProgressPattern = /^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$/im; const plainProgressPattern = /^(Progress:[ \t]*)([^\r\n]*)/im; const pattern = boldProgressPattern.test(body) ? boldProgressPattern diff --git a/tests/state-transition.test.cjs b/tests/state-transition.test.cjs index 0ad4996ba..4221464ca 100644 --- a/tests/state-transition.test.cjs +++ b/tests/state-transition.test.cjs @@ -22,6 +22,7 @@ const { getPreserveWhenUnchangedFields, STATE_MD_SECTIONS, sliceCurrentPositionSection, + stateReplaceProgressPercent, } = require('../gsd-core/bin/lib/state-transition.cjs'); const { stateExtractField } = require('../gsd-core/bin/lib/state-document.cjs'); const { STATE_FIELD_SCHEMA } = require('../gsd-core/bin/lib/state-md-schema.cjs'); @@ -4572,3 +4573,211 @@ describe('#4129: resyncing measured write ratchets the progress block', () => { assert.deepStrictEqual(r.postFm.progress, { total_phases: 18, completed_phases: 2, percent: 11 }); }); }); + +// ───────────────────────────────────────────────────────────────────────────── +// #4243 follow-up: stateReplaceProgressPercent's bold branch is anchored to +// line start, exactly like stateReplaceField's #4453 fix. The pre-fix bold +// pattern carried no ^ and no /m, so a bold percent-ish label quoted +// MID-SENTENCE inside prose — an Accumulated Context bullet mentioning +// `**Progress:**` — captured the machine-segment rewrite and destroyed the +// rest of its line, silently, while the real Progress line stayed stale (the +// callers — cmdStateUpdateProgress, syncCore's percent arm, +// applyPostSyncPreservation — all feed the whole document). #2177's recorded +// protections (frontmatter stripped, suffix preserved, plain-form fallback, +// bold-beats-plain priority among LINE-START forms) are unchanged; per the +// maintainer ruling (2026-09-07), #2177's incidental bold-anywhere matching +// was not load-bearing. +// ───────────────────────────────────────────────────────────────────────────── +describe('stateReplaceProgressPercent — anchored bold form leaves prose lookalikes untouched (#4243 follow-up)', () => { + // The corruption shape: a bold label quoted for documentation purposes + // inside a bullet, with the real status line in the plain template form. + // The value after the label has NO percent, so the pre-fix whole-value + // replacement destroyed the rest of the sentence. + const PROSE_LINE = + '- [2026-07-15] Progress dashboard: the **Progress:** field is machine-managed by state update-progress; do not hand-edit.'; + + const SEGMENT = (bars) => `[${'█'.repeat(bars)}${'░'.repeat(10 - bars)}]`; + + // ROW 1 — the failing-first regression. The lookalike must survive + // byte-identically and the REAL plain line must take the update. + test('issue repro: mid-sentence **Progress:** lookalike survives, real plain line updates', () => { + const input = [ + '## Current Position', + '', + 'Phase: 1 of 1', + 'Plan: 2 of 2', + 'Status: Executing Phase 1', + '', + 'Progress: [█████░░░░░] 50% (1/2 plans done)', + '', + '## Accumulated Context', + '', + '### Decisions', + '', + PROSE_LINE, + '', + ].join('\n'); + const result = stateReplaceProgressPercent(input, 0); + assert.notEqual(result, null, 'the real plain Progress line must still match'); + assert.ok( + result.includes(PROSE_LINE), + `prose lookalike must survive byte-identically, got:\n${result}`, + ); + assert.ok( + result.includes(`Progress: ${SEGMENT(0)} 0% (1/2 plans done)`), + `the real plain line's machine segment must update with the suffix intact, got:\n${result}`, + ); + assert.ok( + !result.includes('the **Progress:** ['), + 'the rewrite must not bleed a machine segment into the prose occurrence', + ); + }); + + test('lookalike ordered BEFORE the real bold line: prose survives, line-start bold line updates', () => { + const input = [ + '## Accumulated Context', + '', + PROSE_LINE, + '', + '## Current Position', + '', + '**Progress:** [█████░░░░░] 50% (2/4 plans done; blocked on API keys)', + '', + ].join('\n'); + const result = stateReplaceProgressPercent(input, 75); + assert.notEqual(result, null); + assert.ok(result.includes(PROSE_LINE), `prose lookalike must survive, got:\n${result}`); + assert.ok( + result.includes(`**Progress:** ${SEGMENT(8)} 75% (2/4 plans done; blocked on API keys)`), + `the real line-start bold line's machine segment must update with the suffix intact, got:\n${result}`, + ); + }); + + test('lookalike whose value carries a percent: prose percent is not swapped, real plain line updates', () => { + const lookalike = '- The **Progress:** bar read 20% last week; see the archived thread.'; + const input = [ + 'Progress: [█████░░░░░] 50% (1/2 plans done)', + '', + '## Accumulated Context', + '', + lookalike, + '', + ].join('\n'); + const result = stateReplaceProgressPercent(input, 0); + assert.notEqual(result, null); + assert.ok( + result.includes(lookalike), + `the prose percent must not be swapped into a machine segment, got:\n${result}`, + ); + assert.ok( + result.includes(`Progress: ${SEGMENT(0)} 0% (1/2 plans done)`), + 'the real plain line must take the update', + ); + }); + + test('mid-sentence lookalike with no real line: returns null (honest absence), never a rewrite', () => { + const input = `Some prose sentence quoting a **Progress:** label mid-sentence, plus trailing words.`; + assert.equal(stateReplaceProgressPercent(input, 40), null); + }); + + test('mid-word lookalike with no real line: returns null', () => { + const input = 'Prose mentions text**Progress:**tail mid-word and nothing else.'; + assert.equal(stateReplaceProgressPercent(input, 40), null); + }); + + // Negative space: an INDENTED line-start bold line is still the status line + // (the leading class is same-line whitespace only, #4010's idiom), and the + // indent is preserved. + test('indented line-start bold line still updates, indent preserved', () => { + const input = ' **Progress:** [█████░░░░░] 50%'; + const result = stateReplaceProgressPercent(input, 100); + assert.equal(result, ` **Progress:** ${SEGMENT(10)} 100%`); + }); + + // Negative space + fix-shape pin: leading blank lines before the label are + // NOT swallowed. The anchor's leading class is same-line whitespace only + // (`[ \t]*`); the naive `^\s*` variant would consume the newlines into the + // match and drop them on rebuild (#4010 hazard, rejected in #4453). + test('leading blank lines before a line-start bold label survive byte-identically', () => { + const input = '\n\n**Progress:** [█████░░░░░] 50%'; + const result = stateReplaceProgressPercent(input, 40); + assert.equal(result, `\n\n**Progress:** ${SEGMENT(4)} 40%`); + }); + + // Negative space: #2177's priority is unchanged among LINE-START forms — a + // real bold status line still beats the plain form, so an earlier free-text + // plain `Progress:` line cannot capture the rewrite ahead of it. + test('line-start bold still beats the plain form; earlier free-text plain line untouched (#2177)', () => { + const freeText = 'Progress: tracked in the weekly thread, do not edit this line by hand'; + const input = `${freeText}\n\n**Progress:** [█████░░░░░] 50%\n`; + const result = stateReplaceProgressPercent(input, 75); + assert.notEqual(result, null); + assert.ok(result.includes(freeText), 'the free-text plain line must stay byte-identical'); + assert.ok( + result.includes(`**Progress:** ${SEGMENT(8)} 75%`), + 'the line-start bold status line is the one rewritten', + ); + }); + + // Negative space: first-occurrence-wins among line-start bold lines. + test('two line-start bold occurrences: only the first is replaced', () => { + const input = '**Progress:** [█████░░░░░] 50%\n**Progress:** [████░░░░░░] 40%'; + const result = stateReplaceProgressPercent(input, 10); + assert.equal(result, `**Progress:** ${SEGMENT(1)} 10%\n**Progress:** [████░░░░░░] 40%`); + }); + + // Negative space: the plain-form fallback (#2177's plain path) is unchanged + // when no bold line exists at all — machine-segment-only swap, suffix kept. + test('plain-form fallback unchanged: no bold anywhere, plain line updates with suffix intact', () => { + const input = 'Progress: [██░░░░░░░░] 20% (1/2 plans done; next: verification)'; + const result = stateReplaceProgressPercent(input, 50); + assert.notEqual(result, null); + assert.equal(result, `Progress: ${SEGMENT(5)} 50% (1/2 plans done; next: verification)`); + }); + + // Negative space: #2177's core — the YAML frontmatter `progress:` key is + // never a match target; the block survives byte-identically. + test('frontmatter progress: key never matched — block survives byte-identically (#2177)', () => { + const fm = [ + '---', + 'progress:', + ' total_plans: 2', + ' completed_plans: 1', + ' percent: 50', + '---', + '', + ].join('\n'); + const input = `${fm}# Project State\n\nProgress: [█████░░░░░] 50% (1/2 plans done)\n\n## Accumulated Context\n\n${PROSE_LINE}\n`; + const result = stateReplaceProgressPercent(input, 0); + assert.notEqual(result, null); + assert.ok( + result.startsWith(fm), + `the frontmatter block must survive byte-identically, got:\n${result}`, + ); + assert.ok(result.includes(' percent: 50'), 'the frontmatter percent line is untouched'); + assert.ok(result.includes(PROSE_LINE), 'the prose lookalike survives'); + assert.ok( + result.includes(`Progress: ${SEGMENT(0)} 0% (1/2 plans done)`), + 'the real plain line takes the update', + ); + }); + + test('CRLF document: lookalike survives with CRLF intact, real plain line updates', () => { + const input = [ + 'Progress: [█████░░░░░] 50% (1/2 plans done)', + '', + '## Accumulated Context', + '', + PROSE_LINE, + '', + ].join('\r\n'); + const result = stateReplaceProgressPercent(input, 0); + assert.notEqual(result, null); + assert.ok(result.includes(PROSE_LINE), `prose lookalike must survive, got:\n${result}`); + assert.ok(result.includes('\r\n'), 'CRLF endings must be preserved'); + assert.ok( + result.includes(`Progress: ${SEGMENT(0)} 0% (1/2 plans done)`), + 'the real plain line must take the update', + ); + }); +}); diff --git a/tests/state.test.cjs b/tests/state.test.cjs index e924f597b..75aa4c9d4 100644 --- a/tests/state.test.cjs +++ b/tests/state.test.cjs @@ -4141,6 +4141,204 @@ describe('#4243: begin-phase leaves prose lookalikes untouched, preserves unknow }); }); +// ───────────────────────────────────────────────────────────────────────────── +// #4243 follow-up — update-progress: the bold branch of +// stateReplaceProgressPercent is anchored to line start (same fix as #4453's +// stateReplaceField anchoring), so prose bold-percent lookalikes stay +// untouched and the real Progress line takes the machine-segment rewrite. +// ───────────────────────────────────────────────────────────────────────────── + +describe('#4243 follow-up: update-progress leaves prose **Progress:** lookalikes untouched', () => { + const PROSE_LINE = + '- [2026-07-15] Progress dashboard: the **Progress:** field is machine-managed by state update-progress; do not hand-edit.'; + + let tmpDir; + + beforeEach(() => { + tmpDir = createFixture(); + // Same #3217 free-form ROADMAP as the update-progress block above: a + // no-version ROADMAP is COMPLETE scope, so the percent is computed rather + // than withheld. + fs.writeFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), '# Roadmap\n'); + }); + + afterEach(() => { + cleanup(tmpDir); + }); + + function seedOnePhaseTwoPlans() { + // 1 of 2 plans summarized, no *-VERIFICATION.md → min-capped 0% (the + // existing update-progress rows' derivation; exact value is incidental — + // what matters is WHERE the rewrite lands). + const phaseDir = path.join(tmpDir, '.planning', 'phases', '01'); + fs.mkdirSync(phaseDir, { recursive: true }); + fs.writeFileSync(path.join(phaseDir, '01-01-PLAN.md'), '# Plan\n'); + fs.writeFileSync(path.join(phaseDir, '01-01-SUMMARY.md'), '# Summary\n'); + fs.writeFileSync(path.join(phaseDir, '01-02-PLAN.md'), '# Plan\n'); + } + + // The repro verbatim: the real status line in the plain template form, the + // lookalike quoted mid-sentence inside an Accumulated Context bullet whose + // value has no percent (the pre-fix whole-value replacement shape). + test('issue repro: prose lookalike is byte-identical, real plain line updates, surfaces agree', () => { + writeState(tmpDir, [ + '# Project State', + '', + '## Current Position', + 'Phase: 1 of 1', + 'Plan: 2 of 2', + 'Status: Executing Phase 1', + 'Last activity: 2026-08-01 — did a thing', + '', + 'Progress: [█████░░░░░] 50% (1/2 plans done)', + '', + '## Accumulated Context', + '', + '### Decisions', + '', + PROSE_LINE, + '', + ].join('\n')); + seedOnePhaseTwoPlans(); + + const result = runGsdTools('state update-progress', tmpDir); + assert.ok(result.success, `update-progress failed: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.updated, true, 'the real plain Progress line must still match'); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok( + content.includes(PROSE_LINE), + `prose lookalike must survive update-progress byte-identically, got:\n${content}`, + ); + // The whole-document Progress extractor (stateExtractField) is itself + // bold-anywhere on the read side — with a bold lookalike in prose it + // extracts the PROSE value, so it cannot witness this write. Assert on + // the Current Position section body instead (#4453's precedent for + // prose-lookalike fixtures), plus the frontmatter surface for agreement. + const pos = sectionMatchOf(content, 'Current Position'); + assert.ok(pos, 'Current Position section should exist'); + assert.match( + pos[1], + new RegExp(`^Progress: \\[${'░'.repeat(10)}\\] ${out.percent}% \\(1/2 plans done\\)$`, 'm'), + 'the real plain line takes the machine-segment rewrite with its suffix intact', + ); + const fm = frontmatterLib.extractFrontmatter(content); + assert.strictEqual( + Number(fm.progress && fm.progress.percent), + out.percent, + 'frontmatter percent must agree with the reported percent (#4213 surfaces-agree)', + ); + }); + + // Corruption shape 2: lookalike section ordered BEFORE the status line, + // real field in the line-start bold form. + test('lookalike before Current Position: prose survives, real line-start bold line updates', () => { + writeState(tmpDir, [ + '# Project State', + '', + '## Accumulated Context', + '', + PROSE_LINE, + '', + '## Current Position', + '', + '**Progress:** [█████░░░░░] 50% (2/4 plans done; blocked on API keys)', + '', + ].join('\n')); + seedOnePhaseTwoPlans(); + + const result = runGsdTools('state update-progress', tmpDir); + assert.ok(result.success, `update-progress failed: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.updated, true); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok( + content.includes(PROSE_LINE), + `prose lookalike must survive update-progress byte-identically, got:\n${content}`, + ); + // Read-side extractor is bold-anywhere (see C1 note): assert on the + // Current Position section body instead. + const pos = sectionMatchOf(content, 'Current Position'); + assert.ok(pos, 'Current Position section should exist'); + assert.match( + pos[1], + new RegExp(`^\\*\\*Progress:\\*\\* \\[${'░'.repeat(10)}\\] ${out.percent}% \\(2/4 plans done; blocked on API keys\\)$`, 'm'), + 'the real line-start bold line takes the machine-segment rewrite with its suffix intact', + ); + }); + + // Honest absence: with only a prose lookalike (no line-start Progress line + // at all), the command must report updated:false with the #3957 body-layer + // reason — not a false success that corrupts the prose. + test('lookalike only, no body Progress line: updated:false, file unchanged', () => { + const before = [ + '# Project State', + '', + '## Accumulated Context', + '', + PROSE_LINE, + '', + ].join('\n'); + writeState(tmpDir, before); + seedOnePhaseTwoPlans(); + + const result = runGsdTools('state update-progress', tmpDir); + assert.ok(result.success, `update-progress failed: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.updated, false, 'a prose lookalike is not a Progress line'); + assert.strictEqual( + out.reason, + 'no Progress: line found in STATE.md body to update (frontmatter progress data is unaffected)', + ); + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok(content.includes(PROSE_LINE), 'the prose must be untouched'); + }); + + // #2177 priority restated at CLI level among LINE-START forms: an earlier + // free-text plain `Progress:` line must not capture the rewrite ahead of + // the real bold status line. + test('free-text plain Progress: line above the real bold line stays byte-identical (#2177)', () => { + const freeText = 'Progress: tracked in the weekly thread, do not edit this line by hand'; + writeState(tmpDir, [ + '# Project State', + '', + freeText, + '', + '**Progress:** [█████░░░░░] 50%', + '', + '## Accumulated Context', + '', + PROSE_LINE, + '', + ].join('\n')); + seedOnePhaseTwoPlans(); + + const result = runGsdTools('state update-progress', tmpDir); + assert.ok(result.success, `update-progress failed: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.updated, true); + + const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8'); + assert.ok(content.includes(freeText), 'the free-text plain line must stay byte-identical'); + assert.ok(content.includes(PROSE_LINE), 'the prose lookalike must stay byte-identical'); + // Read-side extractor is bold-anywhere (see the C1 note above); the free- + // text plain line and the bold status line live in the top-of-body region + // between the title and the first ## heading. Scope with splitLines + // (CRLF-safe per local/no-crlf-fragile-split), never a bare \n split. + const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs'); + const lines = splitLines(content); + const firstSectionIdx = lines.findIndex((l) => l.startsWith('## ')); + const beforeFirstSection = lines.slice(0, firstSectionIdx === -1 ? lines.length : firstSectionIdx).join('\n'); + assert.match( + beforeFirstSection, + new RegExp(`^\\*\\*Progress:\\*\\* \\[${'░'.repeat(10)}\\] ${out.percent}%$`, 'm'), + 'the line-start bold status line is the one rewritten', + ); + }); +}); + // ───────────────────────────────────────────────────────────────────────────── // Bug #1589 — progress counters not updated during plan execution // ───────────────────────────────────────────────────────────────────────────── From 0a0905705a92447f476b5a6ae683ebfcafdfd96d Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 08:24:14 -0400 Subject: [PATCH 035/166] fix(#4256): resolve todos from the root via todosDir everywhere (#4479) * test(#4256): pin todos as root-scoped under workstreams (RED) * fix(#4256): resolve todos from the root via todosDir everywhere * chore(#4256): changeset fragment (pr number to backfill) * chore(#4256): backfill PR number in changeset --------- Co-authored-by: sim --- .changeset/jolly-moles-snooze.md | 5 + src/audit.cts | 29 ++- src/commands.cts | 18 +- src/init.cts | 22 +- src/planning-workspace.cts | 35 +++ tests/todos-workstream-scope.test.cjs | 318 ++++++++++++++++++++++++++ 6 files changed, 410 insertions(+), 17 deletions(-) create mode 100644 .changeset/jolly-moles-snooze.md create mode 100644 tests/todos-workstream-scope.test.cjs diff --git a/.changeset/jolly-moles-snooze.md b/.changeset/jolly-moles-snooze.md new file mode 100644 index 000000000..e131d5477 --- /dev/null +++ b/.changeset/jolly-moles-snooze.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 4479 +--- +**Todos stay visible under a workstream** — todos are root-scoped shared state, but every code reader resolved them through the workstream-aware planning dir, so with a workstream active todos read as empty, `todo complete` refused existing files, and the milestone-close audit-open gate passed with pending todos on disk. (#4256) diff --git a/src/audit.cts b/src/audit.cts index 53d512a40..50da84a5c 100644 --- a/src/audit.cts +++ b/src/audit.cts @@ -22,7 +22,7 @@ import coreUtils = require('./core-utils.cjs'); const { normalizeLineEndings } = coreUtils; // eslint-disable-next-line @typescript-eslint/no-require-imports import planningWorkspace = require('./planning-workspace.cjs'); -const { planningDir, quickDirFrom } = planningWorkspace; +const { planningDir, quickDirFrom, todosDir } = planningWorkspace; // eslint-disable-next-line @typescript-eslint/no-require-imports import frontmatter = require('./frontmatter.cjs'); const { extractFrontmatter, spliceFrontmatter } = frontmatter; @@ -711,9 +711,16 @@ function scanThreads(planDir: string): ScanOutcome { * Scan .planning/todos/pending/ for pending todos. * Returns array of { filename, priority, area, summary }. * Display limited to first 5 + count of remainder. + * + * #4256: takes the ROOT-scoped todos base (`todosDir(cwd)`), NOT the + * workstream-scoped planning dir the other scans use — todos are shared + * project state (the migrateToWorkstreams contract keeps them at + * .planning/todos/), so the close gate must read the root or it clears + * vacuously under a workstream. The requireSafePath boundary below moves + * with the base. */ -function scanTodos(planDir: string): ScanOutcome { - const pendingDir = path.join(planDir, 'todos', 'pending'); +function scanTodos(todosBase: string): ScanOutcome { + const pendingDir = path.join(todosBase, 'pending'); if (!fs.existsSync(pendingDir)) return { items: [], acknowledged: 0 }; let files: fs.Dirent[]; @@ -741,7 +748,7 @@ function scanTodos(planDir: string): ScanOutcome { let safeFilePath: string; try { - safeFilePath = requireSafePath(filePath, planDir, 'todo file', { allowAbsolute: true }); + safeFilePath = requireSafePath(filePath, todosBase, 'todo file', { allowAbsolute: true }); } catch { continue; } @@ -1314,7 +1321,12 @@ function auditOpenArtifacts(cwd: string): AuditResult { })(); const todos = (() => { - try { return scanTodos(planDir); } catch { return { items: [{ scan_error: true, filename: '', priority: '', area: '', summary: '' }], acknowledged: 0 }; } + // #4256: the ONE root-scoped category — todos are shared project state, + // so the close gate reads todosDir(cwd) (the root), not the workstream- + // scoped planDir every other scan below receives. Reading planDir here + // made audit-open print "All artifact types clear. Safe to proceed." + // with pending todos on disk under a workstream. + try { return scanTodos(todosDir(cwd)); } catch { return { items: [{ scan_error: true, filename: '', priority: '', area: '', summary: '' }], acknowledged: 0 }; } })(); const seeds = (() => { @@ -1751,7 +1763,12 @@ function cmdAuditAcknowledge(cwd: string, args: string[], raw: boolean): void { currentValue = ((extractFrontmatter(content, safeFilePath).status as string) || 'dormant').toLowerCase(); } else if (category === 'todos') { if (!filename) ioError('--filename is required for --category todos'); - safeFilePath = requireSafePath(path.join(planDir, 'todos', 'pending', filename as string), planDir, 'audit acknowledge target', { allowAbsolute: true }); + // #4256: todos are root-scoped shared state — derive the todos base and + // pass it as BOTH the path base and the requireSafePath boundary. The + // old workstream-scoped planDir boundary would refuse a root todos file + // outright, and even a path fix alone would have thrown here. + const rootTodos = todosDir(cwd); + safeFilePath = requireSafePath(path.join(rootTodos, 'pending', filename as string), rootTodos, 'audit acknowledge target', { allowAbsolute: true }); if (!fs.existsSync(safeFilePath)) ioError(`file not found: todos/pending/${filename as string}`); currentValue = ''; // presence-only — see scanTodos } else if (category === 'quick_tasks') { diff --git a/src/commands.cts b/src/commands.cts index 1230b678f..875629738 100644 --- a/src/commands.cts +++ b/src/commands.cts @@ -49,7 +49,7 @@ import { parseCodexAgentToml, renderCodexAgentToml, stripModel, stripReasoningEf import hostIntegrationMod = require('./host-integration.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports import planningWorkspace = require('./planning-workspace.cjs'); -const { planningDir, planningPaths } = planningWorkspace; +const { planningDir, planningPaths, todosDir } = planningWorkspace; // eslint-disable-next-line @typescript-eslint/no-require-imports import frontmatter = require('./frontmatter.cjs'); const { extractFrontmatter, agentScalarNeedsDoubleQuoting, escapeDoubleQuotedScalar } = frontmatter; @@ -239,7 +239,10 @@ function cmdCurrentTimestamp(format: string | undefined, raw: boolean): void { } function cmdListTodos(cwd: string, area: string | undefined, raw: boolean): void { - const pendingDir = path.join(planningDir(cwd), 'todos', 'pending'); + // #4256: todos are root-scoped shared state — resolve via todosDir(cwd), + // never planningDir(cwd) (workstream-scoped), or the listing goes empty + // under a workstream. + const pendingDir = path.join(todosDir(cwd), 'pending'); let count = 0; const todos: Array<{ file: string; created: string; title: string; area: string; path: string; severity?: string }> = []; @@ -2778,7 +2781,8 @@ function cmdProgressRender(cwd: string, format: string | undefined, raw: boolean function cmdTodoMatchPhase(cwd: string, phase: string | undefined, raw: boolean): void { if (!phase) { error('phase required for todo match-phase'); } - const pendingDir = path.join(planningDir(cwd), 'todos', 'pending'); + // #4256: root-scoped todos read — see cmdListTodos. + const pendingDir = path.join(todosDir(cwd), 'pending'); const todos: Array<{ file: string; title: string; @@ -2945,8 +2949,12 @@ function cmdTodoComplete(cwd: string, filename: string | undefined, options: Tod error('filename required for todo complete'); } - const pendingDir = path.join(planningDir(cwd), 'todos', 'pending'); - const completedDir = path.join(planningDir(cwd), 'todos', 'completed'); + // #4256: root-scoped todos read/write — see cmdListTodos. The pending and + // completed halves of the move must resolve from the SAME root or the + // completion would strand files where no reader looks. + const todosRoot = todosDir(cwd); + const pendingDir = path.join(todosRoot, 'pending'); + const completedDir = path.join(todosRoot, 'completed'); const sourcePath = path.join(pendingDir, filename as string); if (!fs.existsSync(sourcePath)) { diff --git a/src/init.cts b/src/init.cts index b11de2485..3203e97a0 100644 --- a/src/init.cts +++ b/src/init.cts @@ -104,6 +104,7 @@ const { planningPaths, planningDir, planningRoot, + todosDir, listAvailableWorkstreams, peekActiveWorkstream, diagnoseUnresolvedActiveWorkstream, @@ -2356,7 +2357,13 @@ function renderPendingTodoBullet(todo: Record, projectRoot?: st function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void { const config = loadConfig(cwd); - const pendingDir = path.join(planningDir(cwd), 'todos', 'pending'); + // #4256: todos are root-scoped shared state (migrateToWorkstreams keeps + // them at .planning/todos/ and every workflow writer writes that literal + // path), so this read resolves via todosDir(cwd) — NOT planningDir(cwd), + // which would look in .planning/workstreams//todos/ under a workstream + // (a directory nothing creates) and report existing todos as absent. + const todosRoot = todosDir(cwd); + const pendingDir = path.join(todosRoot, 'pending'); let count = 0; const todos: Record[] = []; // #2618: distinct from "genuinely zero pending todos" — false only when @@ -2412,7 +2419,7 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void title: titleMatch ? titleMatch[1].trim() : 'Untitled', area: todoArea, // #2376: absolute — see comment on phase_dir in cmdInitExecutePhase. - path: toPosixPath(path.join(planningDir(cwd), 'todos', 'pending', file)), + path: toPosixPath(path.join(pendingDir, file)), ...(severityMatch ? { severity: severityMatch[1].trim() } : {}), ...(needs ? { needs } : {}), }); @@ -2438,12 +2445,15 @@ function cmdInitTodos(cwd: string, area: string | undefined, raw: boolean): void area_filter: area || null, // #2376: absolute — see comment on phase_dir in cmdInitExecutePhase. - pending_dir: toPosixPath(path.join(planningDir(cwd), 'todos', 'pending')), - completed_dir: toPosixPath(path.join(planningDir(cwd), 'todos', 'completed')), + // #4256: both dir fields probe the ROOT todos tree via todosDir(cwd). + pending_dir: toPosixPath(pendingDir), + completed_dir: toPosixPath(path.join(todosRoot, 'completed')), + // planning_exists intentionally stays workstream/project-scoped — it + // answers "does the ACTIVE planning dir exist", not a todos question. planning_exists: fs.existsSync(planningDir(cwd)), - todos_dir_exists: fs.existsSync(path.join(planningDir(cwd), 'todos')), - pending_dir_exists: fs.existsSync(path.join(planningDir(cwd), 'todos', 'pending')), + todos_dir_exists: fs.existsSync(todosRoot), + pending_dir_exists: fs.existsSync(pendingDir), // #2618: see PENDING_TODO_BULLET_MAX_CHARS comment / design doc. Consumed // by add-todo.md / check-todos.md's update_state step; omitted entirely diff --git a/src/planning-workspace.cts b/src/planning-workspace.cts index 33fb125de..037a4e325 100644 --- a/src/planning-workspace.cts +++ b/src/planning-workspace.cts @@ -304,6 +304,7 @@ interface PlanningPaths { requirements: string; debug: string; quick: string; + todos: string; } // #2142: the quick-task directory. Exported as its own function (not only as a @@ -316,6 +317,33 @@ function quickDirFrom(planningBase: string): string { return path.join(planningBase, 'quick'); } +// #4256: the todos directory — deliberately ROOT-SCOPED, unlike every other +// planningPaths key. Todos are shared project state by construction: the +// migrateToWorkstreams contract keeps them among the shared files that "stay +// in place" at .planning/todos/ (workstream.cts), and every workflow writer +// writes that literal cwd-relative root path. The six todos readers +// previously hand-composed `path.join(planningDir(cwd), 'todos', ...)`, +// which silently re-scoped to .planning/workstreams//todos/ — a +// directory nothing creates — under a workstream, so todos went invisible +// and audit-open passed the milestone-close gate vacuously. Same +// two-composers-of-one-path shape the `debug` (#3149) and `quick` (#2142) +// keys were introduced to eliminate (DEFECT.GENERATIVE-FIX). +// +// Exported as its own function pair (not only as a `planningPaths` key) +// because `audit.cts`'s `scanTodos`/`cmdAuditAcknowledge` consume an +// already-resolved todos base rather than a `cwd`, mirroring how #2142 +// exported `quickDirFrom` for `scanQuickTasks`. `todosDir` takes NO ws/project +// parameter — todos have no workstream- or project-scoped form anywhere, so +// there is no discriminator to thread. This is also the single root #4327's +// future filename-containment guard should enforce against. +function todosDirFrom(planningBase: string): string { + return path.join(planningBase, 'todos'); +} + +function todosDir(cwd: string): string { + return todosDirFrom(planningRoot(cwd)); +} + function planningPaths(cwd: string, ws?: string | null): PlanningPaths { const base = planningDir(cwd, ws); return { @@ -332,6 +360,11 @@ function planningPaths(cwd: string, ws?: string | null): PlanningPaths { debug: path.join(base, 'debug'), // #2142: quick-task directory, composed via the shared quickDirFrom helper. quick: quickDirFrom(base), + // #4256: todos directory — deliberately ROOT-scoped while the rest of + // this record follows the active workstream/project (todos are shared + // project state per the migrateToWorkstreams contract), composed via the + // shared todosDir helper so this key and every direct caller agree. + todos: todosDir(cwd), }; } @@ -638,6 +671,8 @@ export = { listAvailableWorkstreams, planningPaths, quickDirFrom, + todosDirFrom, + todosDir, withPlanningLock, getActiveWorkstream, peekActiveWorkstream, diff --git a/tests/todos-workstream-scope.test.cjs b/tests/todos-workstream-scope.test.cjs new file mode 100644 index 000000000..7c6d28192 --- /dev/null +++ b/tests/todos-workstream-scope.test.cjs @@ -0,0 +1,318 @@ +'use strict'; + +// ───────────────────────────────────────────────────────────────────────────── +// #4256 — todos are root-scoped shared state; readers must read the root. +// +// The migrateToWorkstreams contract (src/workstream.cts) keeps todos among the +// SHARED files that "stay in place" at .planning/todos/, and every workflow +// writer writes that literal root path. But all six code readers composed +// their todos path from the workstream-aware planningDir(cwd), so under a +// workstream they read .planning/workstreams//todos/pending/ — a directory +// nothing creates — and each reader's ENOENT guard turned "wrong directory" +// into a legitimate-looking zero: invisible todos, `todo complete` refusing +// files that exist, and `audit-open` (a milestone-close gate) printing "All +// artifact types clear" with pending todos on disk. +// +// The fix converges every reader on the root-scoped todosDir(cwd) resolver +// (src/planning-workspace.cts, beside quickDirFrom per #2142/#3149), and moves +// audit acknowledge's requireSafePath boundary with it. These rows pin the +// converged behavior under a workstream AND the byte-identical flat-mode +// control (planningDir === planningRoot with no project/workstream active). +// ───────────────────────────────────────────────────────────────────────────── + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const { createTempProject, cleanup, runGsdTools, toPosixPath } = require('./helpers.cjs'); +const planningWorkspace = require('../gsd-core/bin/lib/planning-workspace.cjs'); + +const TODO_A = '2026-09-01-one.md'; +const TODO_B = '2026-09-02-two.md'; + +function seedRootTodos(tmpDir) { + fs.mkdirSync(path.join(tmpDir, '.planning', 'todos', 'pending'), { recursive: true }); + fs.mkdirSync(path.join(tmpDir, '.planning', 'todos', 'completed'), { recursive: true }); + for (const [file, title, created] of [[TODO_A, 'One', '2026-09-01'], [TODO_B, 'Two', '2026-09-02']]) { + fs.writeFileSync( + path.join(tmpDir, '.planning', 'todos', 'pending', file), + `---\ntitle: ${title}\ncreated: ${created}\narea: general\npriority: high\n---\nbody\n`, + ); + } +} + +function seedWorkstreamDir(tmpDir, ws) { + // Minimal workstream layout — enough for the ws discriminator to resolve a + // real directory; nothing more is needed to reproduce the scope split. + fs.mkdirSync(path.join(tmpDir, '.planning', 'workstreams', ws, 'phases'), { recursive: true }); +} + +function queryJson(args, cwd, env = {}) { + const r = runGsdTools(args, cwd, env); + assert.ok(r.success, `gsd-tools ${args.join(' ')} failed: ${r.error}`); + return JSON.parse(r.output); +} + +describe('#4256: todos readers resolve the root under a workstream', () => { + test('init.todos counts root todos with GSD_WORKSTREAM set (regression)', (t) => { + const tmpDir = createTempProject('gsd-4256-initws-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'init.todos', '--raw'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + + assert.equal(out.todo_count, 2, `root todos must be visible under a workstream; got ${out.todo_count}`); + assert.ok( + toPosixPath(out.pending_dir).endsWith('.planning/todos/pending'), + `pending_dir must be the ROOT pending dir, got ${out.pending_dir}`, + ); + assert.equal(out.todos_dir_exists, true, 'todos_dir_exists must probe the root todos dir'); + assert.equal(out.pending_dir_exists, true, 'pending_dir_exists must probe the root pending dir'); + assert.ok( + Array.isArray(out.todos) && out.todos.length === 2, + 'the todo list itself must carry both root files', + ); + assert.ok( + toPosixPath(out.todos[0].path).includes(`.planning/todos/pending/${TODO_A}`), + `per-todo path must point at the root file, got ${out.todos[0].path}`, + ); + }); + + test('init.todos counts root todos via explicit --ws flag', (t) => { + const tmpDir = createTempProject('gsd-4256-initflag-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'init.todos', '--ws', 'feature-a', '--raw'], tmpDir); + + assert.equal(out.todo_count, 2, '--ws must not hide root todos'); + assert.ok( + toPosixPath(out.pending_dir).endsWith('.planning/todos/pending'), + `pending_dir must stay root-scoped under --ws, got ${out.pending_dir}`, + ); + }); + + test('list-todos returns root todos under a workstream', (t) => { + const tmpDir = createTempProject('gsd-4256-listws-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'list-todos'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + + assert.equal(out.count, 2, `list-todos must see root todos under a workstream; got ${out.count}`); + const files = out.todos.map((x) => x.file).sort(); + assert.deepEqual(files, [TODO_A, TODO_B]); + assert.ok( + toPosixPath(out.todos[0].path).startsWith('.planning/todos/pending/'), + `list-todos paths must be root-relative, got ${out.todos[0].path}`, + ); + }); + + test('todo match-phase sees the root todo set under a workstream', (t) => { + const tmpDir = createTempProject('gsd-4256-matchws-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'todo', 'match-phase', '1', '--raw'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + + assert.equal(out.todo_count, 2, `match-phase must scan the root todo set; got ${out.todo_count}`); + }); + + test('todo complete moves the ROOT file pending -> completed under a workstream', (t) => { + const tmpDir = createTempProject('gsd-4256-complete-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'todo', 'complete', TODO_A], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + assert.equal(out.completed, true); + + const rootPending = path.join(tmpDir, '.planning', 'todos', 'pending', TODO_A); + const rootCompleted = path.join(tmpDir, '.planning', 'todos', 'completed', TODO_A); + assert.ok(!fs.existsSync(rootPending), 'source must be removed from the ROOT pending dir'); + assert.ok(fs.existsSync(rootCompleted), 'target must land in the ROOT completed dir'); + const content = fs.readFileSync(rootCompleted, 'utf8'); + assert.match(content, /^status: completed$/m, 'completion fields must be written into frontmatter'); + assert.ok( + !fs.existsSync(path.join(tmpDir, '.planning', 'workstreams', 'feature-a', 'todos')), + 'complete must NOT materialize the divergent workstream todos dir', + ); + }); + + test('todo complete --dry-run finds the root todo and mutates nothing (#4096 fence intact)', (t) => { + const tmpDir = createTempProject('gsd-4256-dryrun-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['query', 'todo', 'complete', TODO_A, '--dry-run'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + + assert.equal(out.dry_run, true); + assert.equal(out.would_complete, true); + assert.ok( + toPosixPath(out.would_move.source).endsWith(`.planning/todos/pending/${TODO_A}`), + `dry-run source must be the ROOT pending file, got ${out.would_move.source}`, + ); + assert.ok( + toPosixPath(out.would_move.target).endsWith(`.planning/todos/completed/${TODO_A}`), + `dry-run target must be the ROOT completed file, got ${out.would_move.target}`, + ); + assert.ok( + fs.existsSync(path.join(tmpDir, '.planning', 'todos', 'pending', TODO_A)) && + fs.existsSync(path.join(tmpDir, '.planning', 'todos', 'pending', TODO_B)), + 'a dry run must leave both root files untouched', + ); + assert.ok( + !fs.existsSync(path.join(tmpDir, '.planning', 'todos', 'completed', TODO_A)), + 'a dry run must not write the completed copy', + ); + }); + + test('audit-open counts root pending todos under a workstream — the close gate FAILS', (t) => { + const tmpDir = createTempProject('gsd-4256-gate-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const out = queryJson(['audit-open', '--json'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + + assert.equal(out.counts.todos, 2, `the close gate must count the root todos; got ${out.counts.todos}`); + assert.equal(out.has_open_items, true, 'pending root todos must block the milestone close'); + const names = out.items.todos.map((i) => i.filename).sort(); + assert.deepEqual(names, [TODO_A, TODO_B]); + }); + + test('audit acknowledge writes the marker into the ROOT todo under a workstream (boundary moved)', (t) => { + const tmpDir = createTempProject('gsd-4256-ack-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + const env = { GSD_WORKSTREAM: 'feature-a' }; + + const ack = queryJson( + ['audit-open', 'acknowledge', '--category', 'todos', '--milestone', 'v1.0', '--filename', TODO_A], + tmpDir, + env, + ); + assert.equal(ack.acknowledged, true); + + const rootFile = path.join(tmpDir, '.planning', 'todos', 'pending', TODO_A); + assert.ok(fs.existsSync(rootFile), 'the acknowledged todo stays in the ROOT pending dir'); + const content = fs.readFileSync(rootFile, 'utf8'); + assert.match(content, /^audit_acknowledged:/m, 'the marker must be spliced into the ROOT file'); + + const after = queryJson(['audit-open', '--json'], tmpDir, env); + assert.equal(after.acknowledged.todos, 1, 'the acknowledged root todo must be tallied, not open'); + assert.equal(after.counts.todos, 1, 'only the unacknowledged root todo stays open'); + }); +}); + +describe('#4256: negative space — flat, project, and boundary scopes', () => { + test('flat mode (no workstream) counts are unchanged', (t) => { + const tmpDir = createTempProject('gsd-4256-flat-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + + const init = queryJson(['query', 'init.todos', '--raw'], tmpDir); + assert.equal(init.todo_count, 2); + assert.ok(toPosixPath(init.pending_dir).endsWith('.planning/todos/pending')); + + const audit = queryJson(['audit-open', '--json'], tmpDir); + assert.equal(audit.counts.todos, 2); + }); + + test('GSD_PROJECT mode also reads the cwd-relative root todos dir', (t) => { + const tmpDir = createTempProject('gsd-4256-project-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + + // Workflow writers use the literal cwd-relative `.planning/todos/...`, so + // under a project scope the files live at the planning ROOT, not under + // `.planning//todos/` — the reader must converge on the root. + const out = queryJson(['query', 'init.todos', '--raw'], tmpDir, { GSD_PROJECT: 'myapp' }); + assert.equal(out.todo_count, 2, `root todos must be visible under GSD_PROJECT; got ${out.todo_count}`); + assert.ok( + toPosixPath(out.pending_dir).endsWith('.planning/todos/pending'), + `pending_dir must stay at the root under GSD_PROJECT, got ${out.pending_dir}`, + ); + }); + + test('legacy per-workstream symlink workaround becomes inert — count stays correct', (t) => { + const tmpDir = createTempProject('gsd-4256-symlink-'); + t.after(() => cleanup(tmpDir)); + seedRootTodos(tmpDir); + seedWorkstreamDir(tmpDir, 'feature-a'); + + // The pre-fix workaround: a relative symlink redirecting the workstream + // path to the root. Post-fix readers never traverse the workstream + // prefix, so it must neither break nor double-count. + const wsDir = path.join(tmpDir, '.planning', 'workstreams', 'feature-a'); + try { + fs.symlinkSync('../../todos', path.join(wsDir, 'todos'), 'dir'); + } catch (err) { + t.skip(`filesystem does not support symlinks: ${err.code}`); + return; + } + + const out = queryJson(['query', 'init.todos', '--raw'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + assert.equal(out.todo_count, 2, `count must stay exact with the workaround symlink present; got ${out.todo_count}`); + }); + + test('fresh project with no todos dir under a workstream is a legitimate zero', (t) => { + const tmpDir = createTempProject('gsd-4256-empty-'); + t.after(() => cleanup(tmpDir)); + seedWorkstreamDir(tmpDir, 'feature-a'); + + const init = queryJson(['query', 'init.todos', '--raw'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + assert.equal(init.todo_count, 0); + assert.equal(init.pending_dir_exists, false); + + const list = queryJson(['query', 'list-todos'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + assert.equal(list.count, 0); + + const audit = queryJson(['audit-open', '--json'], tmpDir, { GSD_WORKSTREAM: 'feature-a' }); + assert.equal(audit.counts.todos, 0); + }); +}); + +describe('#4256: todosDir resolver owns the path (planning-workspace)', () => { + const { todosDirFrom, todosDir, planningPaths, planningRoot } = planningWorkspace; + + test('todosDirFrom(base) composes /todos', () => { + assert.equal(todosDirFrom(path.join('/fake', 'planning')), path.join('/fake', 'planning', 'todos')); + }); + + test('todosDir(cwd) is root-scoped — GSD_WORKSTREAM does not move it', (t) => { + const saved = process.env.GSD_WORKSTREAM; + process.env.GSD_WORKSTREAM = 'feature-a'; + t.after(() => { + if (saved === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = saved; + }); + + assert.equal(todosDir('/fake/repo'), path.join('/fake', 'repo', '.planning', 'todos')); + assert.equal(todosDir('/fake/repo'), path.join(planningRoot('/fake/repo'), 'todos')); + }); + + test('planningPaths keeps workstream keys scoped while todos stays root-scoped', () => { + const saved = process.env.GSD_WORKSTREAM; + delete process.env.GSD_WORKSTREAM; + try { + const paths = planningPaths('/fake/repo', 'feature-x'); + assert.ok(toPosixPath(paths.planning).endsWith('.planning/workstreams/feature-x')); + assert.ok(toPosixPath(paths.state).endsWith('.planning/workstreams/feature-x/STATE.md')); + assert.ok(toPosixPath(paths.quick).endsWith('.planning/workstreams/feature-x/quick')); + assert.ok( + toPosixPath(paths.todos).endsWith('.planning/todos'), + `planningPaths().todos must be deliberately root-scoped, got ${paths.todos}`, + ); + } finally { + if (saved !== undefined) process.env.GSD_WORKSTREAM = saved; + } + }); +}); From ae40529d31d7ccf1921077aba7848e57d34e7e7c Mon Sep 17 00:00:00 2001 From: Michel Moreira Date: Mon, 7 Sep 2026 10:44:38 -0300 Subject: [PATCH 036/166] =?UTF-8?q?chore(#4394):=20lint=20allowed-tools=20?= =?UTF-8?q?parity=20=E2=80=94=20Bash=20without=20Grep=20(#4431)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit gen-plugin-skills.cjs --check already guarantees skills/*/SKILL.md matches what commands/gsd/*.md generates, so the two trees cannot silently diverge FROM EACH OTHER. Nothing guarded the shape #3085 actually found: a command shipping Bash without Grep purely by omission, identical in both trees and therefore invisible to a parity check that only compares them to each other. That drift ran until 29 of 71 skills lacked a tool most of their siblings declared, and a manual audit — not a gate — is what surfaced it. Detection only: the lint never edits a command's allowed-tools. Two failure classes, not one. Violations are the rule itself. Stale exemptions are the other half: an entry whose command is gone, or which no longer declares Bash without Grep, fails just as loudly. An exemption list that can only grow becomes a list of things nobody re-examined, and a pre-forgiven command silently absorbs the next omission. That check earned its keep immediately. #4394 named eight exemptions from the #3085 review — the six ns-* dispatchers, help, and surface — and seven of them do not declare Bash at all, so the rule never reaches them. Listing them would have pre-forgiven seven commands for a condition none of them has. Only surface needs an entry. The tests drive synthetic fixtures rather than the live corpus: asserting "the real tree is clean" would say nothing about whether the rule can detect anything, which is the exact failure mode this lint exists to close. The one live-corpus arm asserts no stale exemptions — a property of this script's own list — and deliberately does not pin a violation count, which would make it a baseline every fix has to update. Co-authored-by: Tom Boucher --- package.json | 2 +- scripts/lint-allowed-tools-parity.cjs | 221 +++++++++++++++++++++++ tests/lint-allowed-tools-parity.test.cjs | 157 ++++++++++++++++ 3 files changed, 379 insertions(+), 1 deletion(-) create mode 100644 scripts/lint-allowed-tools-parity.cjs create mode 100644 tests/lint-allowed-tools-parity.test.cjs diff --git a/package.json b/package.json index 58eadab09..f04770584 100644 --- a/package.json +++ b/package.json @@ -122,7 +122,7 @@ "lint:frontmatter-scalar-broad-grep": "node scripts/lint-frontmatter-scalar-broad-grep.cjs", "lint:removed-but-needed": "node scripts/lint-removed-but-needed.cjs", "lint:response-language": "node scripts/lint-response-language-coverage.cjs", - "lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-portable-timeout.cjs && node scripts/lint-portable-grep.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-tests.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-unreachable-guard-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-slug-derivation-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs && node scripts/lint-vendored-deps.cjs && node scripts/lint-docs-guard-registration.cjs && node scripts/lint-source-test-name-collision.cjs && npm run lint:hooks-runtime-build-seam && node scripts/check-contract-drift.cjs && node scripts/lint-mutation-test-derivation-drift.cjs && node scripts/lint-seam-enforcement.cjs && node scripts/lint-workflow-shellcheck.cjs && npm run lint:response-language", + "lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-portable-timeout.cjs && node scripts/lint-portable-grep.cjs && node scripts/lint-allowed-tools-parity.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-tests.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-unreachable-guard-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-slug-derivation-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs && node scripts/lint-vendored-deps.cjs && node scripts/lint-docs-guard-registration.cjs && node scripts/lint-source-test-name-collision.cjs && npm run lint:hooks-runtime-build-seam && node scripts/check-contract-drift.cjs && node scripts/lint-mutation-test-derivation-drift.cjs && node scripts/lint-seam-enforcement.cjs && node scripts/lint-workflow-shellcheck.cjs && npm run lint:response-language", "lint:allow-test-rule-refs": "node scripts/lint-allow-test-rule-refs.cjs", "lint:regression-names": "node scripts/lint-regression-test-names.cjs", "lint:descriptions": "node scripts/lint-descriptions.cjs", diff --git a/scripts/lint-allowed-tools-parity.cjs b/scripts/lint-allowed-tools-parity.cjs new file mode 100644 index 000000000..a554e0310 --- /dev/null +++ b/scripts/lint-allowed-tools-parity.cjs @@ -0,0 +1,221 @@ +#!/usr/bin/env node +'use strict'; + +/** + * lint-allowed-tools-parity.cjs — catch `allowed-tools` frontmatter that + * declares `Bash` but omits `Grep` (#4394, follow-up to #3085). + * + * ## Why + * + * `gen-plugin-skills.cjs --check` already guarantees `skills/*` /SKILL.md` + * matches byte-for-byte what `commands/gsd/*.md` generates, so the two trees + * cannot silently diverge FROM EACH OTHER. Nothing guarded the shape #3085 + * actually found: a command shipping `Bash` without `Grep` purely by + * omission, identical in both trees and therefore invisible to a parity + * check that only compares them to each other. + * + * That drift ran long enough for 29 of 71 skills to lack a tool most of + * their siblings already declared, and it surfaced through a manual audit + * rather than any gate. A command that can shell out but cannot Grep does + * not fail loudly — it quietly reaches for `Bash` + `grep` instead, which is + * slower, less structured, and (per `lint-portable-grep.cjs`) a portability + * hazard of its own on hosts without GNU grep. + * + * ## The rule + * + * A command whose `allowed-tools` includes `Bash` must also include `Grep`, + * unless it is on the exemption list below. + * + * Detection only. This lint never edits a command's `allowed-tools`. + * + * ## Why an exemption list rather than a heuristic + * + * The alternative — inferring from a command's body whether it "really" + * needs Grep — would make the rule's verdict depend on prose that changes + * constantly, and produce a lint whose failures nobody can predict. A short + * literal list keeps every exemption a reviewable one-line diff, and the + * staleness check below stops it becoming a dumping ground: an entry that no + * longer needs to be there fails just as loudly as a missing tool. + */ + +const fs = require('fs'); +const path = require('path'); +const { ExitError, runMain } = require('./lib/cli-exit.cjs'); + +const ROOT = path.resolve(__dirname, '..'); +const DEFAULT_ROOT = 'commands/gsd'; + +/** + * Commands allowed to declare `Bash` without `Grep`. + * + * Keyed by command stem (the filename without `.md`), with the reason inline + * so changing the set is a one-line, reviewable diff. + * + * #4394 named eight candidates from the #3085 review: the six `gsd-ns-*` + * namespace dispatchers, `gsd-help`, and `gsd-surface`. Seven of those turn + * out not to need an entry at all — they do not declare `Bash` in the first + * place, so the rule never reaches them: + * + * ns-context/ns-ideate/ns-manage/ns-project/ns-review/ns-workflow Read + Skill + * help Read + * + * Listing them anyway would be seven pre-forgiven commands: the day one of + * them gained `Bash`, the omission it was granted an exemption for would + * pass silently. The staleness check in `scan()` below is what surfaced + * this, and it is why the list stays minimal — an exemption is a real + * suppression, not documentation. + */ +const EXEMPT = new Map([ + // Toggles which skills are surfaced. Mutates the install tree through the + // installer's own seams rather than by searching project files, so `Bash` + // here is not standing in for a search it cannot perform. + ['surface', 'mutates the install surface via installer seams, not by searching the project'], +]); + +/** + * Parse an `allowed-tools` frontmatter value. + * + * Handles both shapes the corpus uses: a YAML block sequence + * + * allowed-tools: + * - Read + * - Bash + * + * and an inline scalar or flow sequence (`allowed-tools: Read, Bash` / + * `allowed-tools: [Read, Bash]`). Returns `null` when the key is absent, + * which is NOT a violation — a command with no `allowed-tools` at all + * declares no `Bash` either, so this rule has nothing to say about it. + * + * Deliberately not a YAML parser: the frontmatter here is a fixed, shallow + * shape, and pulling in a parser to read one list would be a dependency + * bought for a single key. + * + * @param {string} text full file contents + * @returns {string[] | null} declared tool names, or null when the key is absent + */ +function parseAllowedTools(text) { + const lines = String(text).split(/\r?\n/); + const keyIndex = lines.findIndex((l) => /^allowed-tools:/.test(l)); + if (keyIndex === -1) return null; + + const inline = lines[keyIndex].replace(/^allowed-tools:/, '').trim(); + if (inline) { + return inline + .replace(/^\[/, '') + .replace(/\]$/, '') + .split(',') + .map((s) => s.trim().replace(/^['"]|['"]$/g, '')) + .filter(Boolean); + } + + const tools = []; + for (let i = keyIndex + 1; i < lines.length; i += 1) { + const line = lines[i]; + const item = line.match(/^\s*-\s+(.+?)\s*$/); + if (item) { + tools.push(item[1].replace(/^['"]|['"]$/g, '')); + continue; + } + // The first line that is neither a list item nor blank ends the block — + // the next frontmatter key, or the closing `---`. + if (line.trim() === '') continue; + break; + } + return tools; +} + +/** + * List command stems under `dir`, sorted. + * + * @param {string} dir absolute path to the commands directory + * @returns {{ stem: string, file: string }[]} + */ +function listCommands(dir) { + let entries; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } catch { + return []; + } + return entries + .filter((e) => e.isFile() && e.name.endsWith('.md')) + .map((e) => ({ stem: e.name.replace(/\.md$/, ''), file: path.join(dir, e.name) })) + .sort((a, b) => a.stem.localeCompare(b.stem)); +} + +/** + * Evaluate the corpus. + * + * Reports two independent failure classes, because an exemption list that + * can only ever grow rots into a list of things nobody re-examined: + * + * - `violations`: declares Bash, omits Grep, not exempt. + * - `staleExemptions`: on the list but no longer needs to be — either the + * command is gone, or it no longer declares Bash without Grep. Removing + * the entry is then a one-line diff, and the next omission in that + * command is caught rather than silently pre-forgiven. + * + * @param {string} [root] repo-relative commands directory + * @returns {{ scanned: number, violations: {stem: string, tools: string[]}[], staleExemptions: {stem: string, reason: string}[] }} + */ +function scan(root = DEFAULT_ROOT) { + const abs = path.isAbsolute(root) ? root : path.join(ROOT, root); + const commands = listCommands(abs); + const violations = []; + const seen = new Set(); + + for (const { stem, file } of commands) { + const tools = parseAllowedTools(fs.readFileSync(file, 'utf8')); + if (!tools) continue; + if (!tools.includes('Bash') || tools.includes('Grep')) continue; + seen.add(stem); + if (EXEMPT.has(stem)) continue; + violations.push({ stem, tools }); + } + + const present = new Set(commands.map((c) => c.stem)); + const staleExemptions = []; + for (const [stem, reason] of EXEMPT) { + if (!present.has(stem)) { + staleExemptions.push({ stem, reason: `no such command (${reason})` }); + } else if (!seen.has(stem)) { + staleExemptions.push({ stem, reason: `no longer declares Bash without Grep (${reason})` }); + } + } + + return { scanned: commands.length, violations, staleExemptions }; +} + +function main() { + const root = process.env.GSD_LINT_ALLOWED_TOOLS_ROOT || DEFAULT_ROOT; + const { scanned, violations, staleExemptions } = scan(root); + + if (violations.length > 0 || staleExemptions.length > 0) { + const parts = []; + if (violations.length > 0) { + parts.push( + 'lint-allowed-tools-parity: these commands declare `Bash` but not `Grep` (#4394).\n' + + 'A command that can shell out but cannot Grep reaches for `Bash` + `grep` instead —\n' + + 'slower, unstructured, and a portability hazard on hosts without GNU grep. Add `Grep`\n' + + 'to the frontmatter, or add the command to EXEMPT in this script with a reason:\n' + + violations.map((v) => ` ${root}/${v.stem}.md [${v.tools.join(', ')}]`).join('\n'), + ); + } + if (staleExemptions.length > 0) { + parts.push( + 'lint-allowed-tools-parity: these EXEMPT entries are stale — delete them so the next\n' + + 'omission in those commands is caught rather than silently pre-forgiven:\n' + + staleExemptions.map((s) => ` ${s.stem}: ${s.reason}`).join('\n'), + ); + } + throw new ExitError(1, parts.join('\n\n')); + } + + console.log( + `ok lint-allowed-tools-parity: ${scanned} command(s) checked, ${EXEMPT.size} exemption(s) all still needed`, + ); +} + +module.exports = { parseAllowedTools, scan, EXEMPT, DEFAULT_ROOT }; + +if (require.main === module) runMain(main); diff --git a/tests/lint-allowed-tools-parity.test.cjs b/tests/lint-allowed-tools-parity.test.cjs new file mode 100644 index 000000000..30082403b --- /dev/null +++ b/tests/lint-allowed-tools-parity.test.cjs @@ -0,0 +1,157 @@ +'use strict'; + +// #4394: unit tests for the allowed-tools parity lint. The rule reads command +// frontmatter, so every arm below drives it against a synthetic commands +// directory rather than the live corpus — a test that asserted "the real tree +// is clean" would say nothing about whether the rule can DETECT anything, and +// that is exactly the failure mode this lint exists to close (the pre-existing +// generated-sync check compares the two trees only to each other, so a shared +// omission was invisible to it). + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const os = require('node:os'); +const path = require('node:path'); + +const { parseAllowedTools, scan, EXEMPT } = require('../scripts/lint-allowed-tools-parity.cjs'); +const { cleanup } = require('./helpers.cjs'); + +/** Write a synthetic commands dir; returns its absolute path. */ +function fixture(commands) { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-4394-')); + for (const [stem, tools] of Object.entries(commands)) { + const frontmatter = tools === null + ? ['---', `name: gsd:${stem}`, '---', ''] + : ['---', `name: gsd:${stem}`, 'allowed-tools:', ...tools.map((t) => ` - ${t}`), 'requires: []', '---', '']; + fs.writeFileSync(path.join(dir, `${stem}.md`), `${frontmatter.join('\n')}\nbody\n`); + } + return dir; +} + +describe('#4394 parseAllowedTools', () => { + test('reads a YAML block sequence', () => { + const text = ['---', 'name: gsd:x', 'allowed-tools:', ' - Read', ' - Bash', 'requires: []', '---'].join('\n'); + assert.deepEqual(parseAllowedTools(text), ['Read', 'Bash']); + }); + + test('stops at the next frontmatter key, not at the end of the file', () => { + // A body bullet list must not be swallowed into the tool set. + const text = [ + '---', 'allowed-tools:', ' - Read', 'requires: [review]', '---', '', 'Steps:', ' - Bash', '', + ].join('\n'); + assert.deepEqual(parseAllowedTools(text), ['Read']); + }); + + test('reads an inline flow sequence and a bare inline list', () => { + assert.deepEqual(parseAllowedTools('allowed-tools: [Read, Bash]'), ['Read', 'Bash']); + assert.deepEqual(parseAllowedTools('allowed-tools: Read, Bash'), ['Read', 'Bash']); + }); + + test('strips quotes around tool names', () => { + assert.deepEqual(parseAllowedTools("allowed-tools: ['Read', \"Bash\"]"), ['Read', 'Bash']); + }); + + test('returns null when the key is absent', () => { + // Absent is NOT a violation: a command with no allowed-tools declares no + // Bash either, so the rule has nothing to say about it. Returning [] here + // would be indistinguishable from "declared, but empty". + assert.equal(parseAllowedTools('---\nname: gsd:x\n---\n'), null); + }); +}); + +describe('#4394 scan — the rule detects, and only what it should', () => { + test('flags Bash without Grep', () => { + const dir = fixture({ offender: ['Read', 'Bash'] }); + try { + const { violations } = scan(dir); + assert.deepEqual(violations.map((v) => v.stem), ['offender']); + // The message has to carry the declared set, or the fix is a guess. + assert.deepEqual(violations[0].tools, ['Read', 'Bash']); + } finally { cleanup(dir); } + }); + + test('accepts Bash WITH Grep', () => { + const dir = fixture({ fine: ['Read', 'Bash', 'Grep'] }); + try { + assert.deepEqual(scan(dir).violations, []); + } finally { cleanup(dir); } + }); + + test('ignores a command that declares no Bash', () => { + // The rule is about Bash standing in for a search the command cannot + // perform. No Bash, no substitution, nothing to say. + const dir = fixture({ reader: ['Read', 'Skill'] }); + try { + assert.deepEqual(scan(dir).violations, []); + } finally { cleanup(dir); } + }); + + test('ignores a command with no allowed-tools key at all', () => { + const dir = fixture({ bare: null }); + try { + assert.deepEqual(scan(dir).violations, []); + } finally { cleanup(dir); } + }); + + test('reports every offender, not just the first', () => { + const dir = fixture({ a: ['Bash'], b: ['Bash'], c: ['Bash', 'Grep'] }); + try { + assert.deepEqual(scan(dir).violations.map((v) => v.stem), ['a', 'b']); + } finally { cleanup(dir); } + }); +}); + +describe('#4394 scan — the exemption list cannot rot', () => { + // An exemption list that can only grow becomes a list of things nobody + // re-examined. These arms are what keep an exemption a real suppression + // rather than a comment. + const exemptStem = [...EXEMPT.keys()][0]; + + test('an exempt command that still needs its entry is silent', () => { + const dir = fixture({ [exemptStem]: ['Read', 'Write', 'Bash'] }); + try { + const { violations, staleExemptions } = scan(dir); + assert.deepEqual(violations, []); + assert.deepEqual(staleExemptions, []); + } finally { cleanup(dir); } + }); + + test('an exempt command that gained Grep is reported as stale', () => { + const dir = fixture({ [exemptStem]: ['Read', 'Bash', 'Grep'] }); + try { + const { staleExemptions } = scan(dir); + assert.deepEqual(staleExemptions.map((s) => s.stem), [exemptStem]); + assert.match(staleExemptions[0].reason, /no longer declares Bash without Grep/); + } finally { cleanup(dir); } + }); + + test('an exemption for a command that no longer exists is reported as stale', () => { + const dir = fixture({ unrelated: ['Read'] }); + try { + const { staleExemptions } = scan(dir); + assert.deepEqual(staleExemptions.map((s) => s.stem), [exemptStem]); + assert.match(staleExemptions[0].reason, /no such command/); + } finally { cleanup(dir); } + }); + + test('every exemption carries a non-empty reason', () => { + // The reason is what makes changing the set reviewable. An entry without + // one is a suppression nobody can evaluate. + for (const [stem, reason] of EXEMPT) { + assert.equal(typeof reason, 'string', `${stem} must carry a reason`); + assert.ok(reason.trim().length > 10, `${stem}'s reason must say something: ${JSON.stringify(reason)}`); + } + }); +}); + +describe('#4394 scan — against the live corpus', () => { + test('the exemption list has no stale entries on the real tree', () => { + // Kept separate from the violations count, which is deliberately NOT + // asserted here: the 21 commands #3085 identified are fixed by its own PR, + // and pinning a number would make this test a baseline that every such fix + // has to update. Staleness is different — it is a property of THIS + // script's list, and it is always this script's job to keep true. + assert.deepEqual(scan().staleExemptions, []); + }); +}); From 2388e6ab34a01a26b9246d228246b3a14de3ab88 Mon Sep 17 00:00:00 2001 From: Michel Moreira Date: Mon, 7 Sep 2026 11:34:37 -0300 Subject: [PATCH 037/166] fix(#4342): run the bug-167 routing test in a fixture project, not the developer's (#4387) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The test called runGsdTools with its default cwd — the test process's own working directory — and an inherited HOME, so the child read the checkout's real .planning/ and the developer's real ~/.gsd/defaults.json. testEnvBase() blanks the config-LOCATION env keys but sandboxes neither cwd nor HOME. On a checkout that has workstreams with no active pointer, `init.progress` exits non-zero and the FIRST assertion fails, so the routing comparison the test exists for was never evaluated: init.progress failed: Error: init.progress requires a workstream in workstream mode — no active workstream is set ... Available workstreams: alpha Reproduced byte-for-byte by adding .planning/workstreams/alpha/ to the checkout: red on next, green here. The invariant under test — `query ` and `` returning identical payloads — is independent of project state, so a plain createTempProject() fixture is enough, with HOME/USERPROFILE pointed at it (the idiom runGsdTools's own doc comment prescribes). Two assertions pin the sandbox deterministically rather than conditionally: the fixture HAS a .planning/ and the repo checkout does not, so dropping the cwd override fails on every lane — CI included, where the ambient state that exposed the bug is absent. Co-authored-by: Tom Boucher --- tests/command-routing-hub.test.cjs | 38 ++++++++++++++++++++++++++---- 1 file changed, 33 insertions(+), 5 deletions(-) diff --git a/tests/command-routing-hub.test.cjs b/tests/command-routing-hub.test.cjs index f434419fb..26b97eac4 100644 --- a/tests/command-routing-hub.test.cjs +++ b/tests/command-routing-hub.test.cjs @@ -962,20 +962,48 @@ describe('CommandRoutingHub — exitReason? field on InvalidArgs (#1644 / amendm const { test } = require('node:test'); const assert = require('node:assert/strict'); -const { runGsdTools } = require('./helpers.cjs'); +const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs'); -test('bug #167: query meta-command prefixes direct gsd-tools calls', () => { - const direct = runGsdTools(['init.progress']); +// #4342: this ran gsd-tools with runGsdTools's default cwd — the checkout's own +// working directory — and an inherited HOME, so the child read the developer's +// real .planning/config.json and ~/.gsd/defaults.json. On a machine whose real +// project enables workstream mode with no active workstream, `init.progress` +// exits non-zero and the FIRST assertion fails, so the routing comparison this +// test exists for was never reached. testEnvBase() blanks the config-LOCATION +// env keys but sandboxes neither cwd nor HOME. +// +// The routing invariant is about `query ` and `` agreeing, which is +// independent of project state — so the fixture only has to be a project whose +// state is known and unaffected by the developer's. +test('bug #167: query meta-command prefixes direct gsd-tools calls', (t) => { + const fixture = createTempProject('gsd-4342-routing-'); + t.after(() => cleanup(fixture)); + // HOME/USERPROFILE point at the fixture so ~/.gsd/defaults.json resolves + // inside it (absent) rather than in the developer's home — the idiom + // runGsdTools's own doc comment prescribes for exactly this. + const sandbox = { HOME: fixture, USERPROFILE: fixture }; + + const direct = runGsdTools(['init.progress'], fixture, sandbox); assert.equal(direct.success, true, `init.progress failed: ${direct.error || direct.output}`); - const meta = runGsdTools(['query', 'init.progress']); + const meta = runGsdTools(['query', 'init.progress'], fixture, sandbox); assert.equal(meta.success, true, `query init.progress failed: ${meta.error || meta.output}`); + const directPayload = JSON.parse(direct.output); assert.deepEqual( JSON.parse(meta.output), - JSON.parse(direct.output), + directPayload, 'query-prefixed and direct invocations should return identical init.progress payloads' ); + + // Pin the sandbox itself, deterministically rather than conditionally: the + // fixture HAS a .planning/ and the repo checkout does NOT, so if the cwd + // override is ever dropped this fails on every lane — including CI, where the + // ambient state that exposed the bug is absent. + assert.equal(directPayload.planning_exists, true, + 'the child must run in the fixture project, not in the checkout'); + assert.equal(directPayload.phase_count, 0, + 'the fixture has no phases — a non-zero count means a real project was read'); }); }); } From 476394689a73500097f1d780e29a35528231abcd Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Mon, 7 Sep 2026 10:54:30 -0400 Subject: [PATCH 038/166] fix(#4254): pin sequential executor to the orchestrator's validated root (#4476) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#4254): sequential executor root pin — failing-first regression + matrix The new suite executes the shipped supplied-root-pin guard against real git fixtures (drifted primary-checkout cwd halts before the write and the FATAL names both roots; matching cwd permits it; unexpanded/empty pins halt; normalization forms; submodule and sibling boundaries; metacharacter quoting; drive-letter form gate) and locks the dispatch contract across execute-phase.md, its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772 per-plan serialization assertion retargets to the fragment that now carries those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring. * fix(#4254): pin sequential executor to the orchestrator's validated root Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its own cwd; every existing guard is worktree-mode-only or self-referential, so an executor spawned with a drifted cwd committed onto the wrong checkout silently. - worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard, composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT (git-vs-git comparison on both sides — representation-safe on Windows, the #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule allowance, warn-and-proceed only when the dispatch carries no pin block. - execute-phase.md sequential branch: build-time embed of the bound via the new execute-phase/steps/sequential-root-pin.md fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the wave serialization rules move with the fragment, verbatim in substance) plus the per-write/commit pin instruction in . Worktree-mode dispatch untouched (its self-derived toplevel IS correct there). - INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens regenerated for the new fragment; changeset added. * chore(#4254): backfill changeset PR number * fix(#4254): accept backslash-separated Windows drive pins CI on windows-latest showed every permit-path test failing with "Actual root: ": pins composed from Node's path.join arrive in the backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form gate rejected before the cwd-side root was ever computed — a legitimate matching pin could never pass. The gate now accepts either separator ([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names. * fix(#4254): portable drive-form gate for MSYS bash The bracket class [\\/] that accepted backslash drive pins parses inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins — every permit-path test red with "Actual root: "). Replace it with standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* — the escape form is version- and build-portable. Verified across all forms: both drive spellings accepted; bare "C:", relative, empty, and unexpanded rejected. * fix(#4254): runtime-generated backslash comparator + self-describing FATAL The Windows CI legs failed every #4254 permit-path row with 'Actual root: ' across two prior pattern spellings ([\\/] and \\*). Stage misattribution: appears whenever the FATAL fires BEFORE the cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired. Mechanism: the test harness spawns bash -c