diff --git a/.changeset/patient-moles-sprint.md b/.changeset/patient-moles-sprint.md new file mode 100644 index 000000000..855d43632 --- /dev/null +++ b/.changeset/patient-moles-sprint.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 3861 +--- +**`/gsd-execute-phase`'s code review gate now reports what it found and records what happened to it** — the gate parsed REVIEW.md's severity counts and discarded them, printing a message that was byte-identical for one `info` finding and for a Critical, and nothing recorded a per-finding disposition anywhere. It now states the breakdown (`23 findings — 1 critical, 9 warning, 8 info`, accepting `blocker:` as the documented tier-equivalent of `critical:`) and writes `-REVIEW-DISPOSITION.md` beside the review, one row per finding defaulting to `open`, preserving any disposition a human recorded, and carrying a row forward even after the review stops reporting it. `/gsd-code-review --fix` reconciles `fixed`/`skipped` into the same ledger once its report exists, and `--auto`'s iterations are reconciled too — the loop overwrites its fix report each pass and the re-review drops what it fixed, so the gate reads the per-iteration backups before they are removed, and a finding closed in iteration 1 is recorded as `fixed` rather than as never triaged. A decision is carried forward only while the finding ID still names the same finding — the ledger records each finding's title, so a renumbered `CR-01` cannot inherit an earlier `CR-01`'s disposition; when an ID is reused the earlier decision loses its row and the drop is reported. Severity follows the section a finding is filed under, an out-of-vocabulary disposition falls back to `open` rather than counting as a decision, and a finding whose heading the gate cannot parse is reported rather than dropped. The gate stays advisory and never blocks. (#3829) diff --git a/docs/FEATURES.md b/docs/FEATURES.md index f0be56c7d..b867d8f87 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -2382,6 +2382,8 @@ Test suite that scans all agent, workflow, and command files for embedded inject - REQ-REVIEW-06: `--auto` flag MUST enable fix + re-review iteration loop, capped at 3 iterations - REQ-REVIEW-07: Feature MUST be gated by `workflow.code_review` config flag - REQ-REVIEW-08: `workflow.code_review_point` MUST select which loop point the automatic review step registers at (`execute:post` default, or `execute:wave:post`), independent of the `workflow.code_review` on/off gate and of manual `/gsd-code-review` invocation (#3661) +- REQ-REVIEW-09: The in-phase `code_review_gate` MUST report the per-severity counts it parses from REVIEW.md, so a review with one `info` finding is distinguishable from a review with a Critical +- REQ-REVIEW-10: Each finding MUST carry a recorded disposition, so a triaged finding is distinguishable from a forgotten one **Config:** | Setting | Type | Default | Description | @@ -2415,6 +2417,143 @@ Escalation is **whole-review, not per-file**: depth is a single scalar handed to v1 supports **directory-prefix matching only, not glob syntax**: no glob engine (`minimatch`, `picomatch`, `fast-glob`) exists in this project and none was added for this feature. A path containing `*` or `?` (e.g. `src/auth/**`) is a configuration error rather than a silent near-miss, because accepting it as sugar for a prefix would make unsupported patterns look armed when they match nothing. Every use case in the issue is expressible as a directory prefix. See [Scope code review depth by path](how-to/scope-code-review-depth-by-path.md) for the resolution order, error table, and a worked example. **Optional external reviewer lanes (#4209):** `/gsd-code-review` accepts the same reviewer-lane flags as `/gsd-review` — any flag the roster declares (run `gsd_run review-lane flags` to list them for your installation, e.g. `--codex`, `--agy`). No reviewer-lane flag is the default and is byte-for-byte unchanged from before #4209: zero lane selection, plan, or invoke calls, and only the internal `gsd-code-reviewer` agent runs. Passing one or more flags asks those lanes to independently review the same already-resolved file scope alongside the internal agent, through the same shared capability-trait interpreter and `review-lane plan`/`invoke` machinery `/gsd-review` uses — no second implementation. Each lane's prompt carries only the repository root, canonical file paths, review depth, and base SHA, never source file contents, under four fixed prohibitions (no source mutation, no test execution, no background processes, no polling). External findings are unverified corroborating evidence: `gsd-code-reviewer` independently re-verifies every claim against the actual source before writing it to `REVIEW.md`, so there remains exactly one `REVIEW.md` schema regardless of how many lanes ran. An explicitly requested lane that is unavailable or fails is reported as a warning, never silently dropped and never a raw-CLI fallback. This is separate from `/gsd-review`, which reviews `PLAN.md` files before execution, not source code. +**In-phase review reporting and disposition** + +`/gsd-execute-phase`'s `code_review_gate` runs code review, then reports what it found: + +``` +Code review: 23 findings — 1 critical, 9 warning, 8 info. +Consider running: /gsd-code-review 1 --fix +``` + +Both `critical:` and its documented tier-equivalent `blocker:` are accepted. A REVIEW.md written +without a `findings:` block has no counts to report, and the gate falls back to the countless form +rather than printing a half-filled line. + +The gate then writes `-REVIEW-DISPOSITION.md` beside the review — one row per finding ID, +defaulting to `open`: + +| Finding | Severity | Disposition | Source | +|---------|----------|-------------|--------| +| CR-01 | critical | open | - | +| WR-01 | warning | fixed | 01-REVIEW-FIX.md | + +`open` means recorded but not yet triaged, and it is the only value the gate assigns on its own. + +**Where the reconciliation happens, which is not where you might expect.** The in-phase gate runs +immediately after review, and at that moment `-REVIEW-FIX.md` does not exist — the gate invokes +review with neither `--fix` nor `--auto` — so every row it writes is `open`. The `fixed` and +`skipped` outcomes are reconciled by `/gsd-code-review --fix`, which records them once the fix +report is on disk. Running the gate alone therefore tells you what was found; running `--fix` is +what records what happened to it. + +**`--auto`'s iterations are reconciled too, and that takes reading more than the final report.** +The loop overwrites `REVIEW-FIX.md` on every pass and the re-review stops reporting a finding once +it is fixed, so a finding closed in iteration 1 appears in neither final artifact. The gate +therefore also reads the per-iteration backups the loop writes (`-REVIEW-FIX.iterN.md`), newest +first, so the most recent statement about an ID wins; the backups are removed after the ledger has +read them, not before. Without this a fully successful multi-iteration run recorded its early fixes +as `open (not in the current review)` — indistinguishable from a finding that vanished for an +unrelated reason, which is the one distinction this ledger exists to make. A finding a fix report +decided but the current review no longer reports gets a row of its own, carrying that decision and +marked *(not in the current review)*. + +A converged `--auto` run also reaches the gate with a clean review and, on a direct +`/gsd-code-review` invocation, no ledger from the in-phase gate. A fix report on disk is reason +enough to record: without that, a run in which every finding was fixed and committed produced no +disposition record at all. + +**A decision belongs to a finding, not to an ID.** IDs are reused across re-reviews — `--auto` +renumbers — so the ledger records each finding's title alongside its disposition, and a recorded +decision is carried forward only while the ID still names the same finding. Without that, a prior +`CR-01 fixed` would be inherited by a brand-new `CR-01`, which is a false decision in the one +artifact whose purpose is telling triaged from forgotten. + +**The limitation that follows, stated rather than hidden.** When an ID *is* reused, the new finding +is recorded `open` (nobody decided anything about it) and the earlier decision **loses its row** — +the ledger keys rows on the finding ID, and two rows under one ID is an ambiguity, not a record. The +drop is reported on the console, naming the ID and what had been decided — for a *recorded* decision. +A prior row still sitting at `open` is replaced silently, and deliberately: `open` means nobody had +decided anything, so there is no decision to lose. Preserving a recorded decision in the file +was tried and withdrawn: it needed a second identity scheme and produced a fresh defect on each of +three review passes. Recovery is the ledger's own git history where `commit_docs` is on — which is +why the console reports the drop rather than pointing at a commit that may not exist. + +**Two residuals, since (ID, title) is not proof of identity.** A ledger written before titles were +recorded carries none, so its decisions are inherited on the ID alone — refusing there would reset +every decision in every existing ledger, which is the loss the guard exists to prevent. And two +genuinely distinct findings that share both an ID and a title are indistinguishable to this key. +Separating them needs a second identity scheme, which is the thing that was just withdrawn. + +Reconciliation applies an outcome only when the fix report names the **same** finding, because +finding IDs are reused across re-reviews and a stale report would otherwise declare a brand-new +`CR-01` already fixed. Titles are compared with runs of whitespace collapsed, so a fixer that +re-spaces a title still reconciles. A title that differs otherwise — including one **wrapped across +lines**, since a `###` heading is one line and the continuation is a separate paragraph — leaves the +row `open` and is reported, naming the ids it could not reconcile, rather than passing silently. The report does not claim to +know whether such a report is stale or merely re-titled, because it cannot tell. + +The disposition column is a closed vocabulary — `open`, `fixed`, `skipped`, `deferred`. A +hand-edited value outside it is not treated as a decision: the row falls back to `open`, so a +typo cannot quietly mark a phase as triaged. + +Severity comes from the section a finding sits under (`## Critical Issues`, `## Warnings`, +`## Info`) when the review uses those headings, and from the ID prefix otherwise. The section is +the reviewer's own statement of severity, so a Critical filed under `## Critical Issues` is +recorded critical even if its ID was mis-numbered `WR-04`. That severity is then **remembered**: +a row the current review no longer reports, or reports under no recognized heading, keeps the +severity the ledger recorded rather than having it re-inferred from the ID prefix — so the +mis-numbered `WR-04` above stays `critical` on the second run, deferred or not. The precedence is +the current review's section, then the recorded value, then the prefix; the recorded value is +inherited only while the ID still names the same finding, by the same title check the disposition +uses, so a reused ID starts from its own review. That check has the same compatibility arm the +disposition has: a ledger written before titles were recorded carries no title to compare, so its +severities — like its decisions — inherit on the ID alone. + +If the review's `total:` exceeds the number of findings whose headings the gate could parse, the +shortfall is stated — on the console and as an `unparsed:` key in the ledger's frontmatter. A +finding the gate cannot record is the one a human most needs to see, so it is never dropped +silently. The one input that produces no shortfall is a `findings:` block that disagrees with +itself: where `critical` (or `blocker`), `warning` and `info` are all present, numeric and do not +sum to `total`, the gate has no trustworthy number to reconcile against and withholds the key +rather than reporting a figure derived from one — the same input on which the console line +already withholds the severity breakdown. Counts that are merely absent, partial or non-numeric +are not a disagreement, and the shortfall is still reported from `total` alone. `deferred` is the one disposition the gate never writes: it is recorded by hand, and the +reason recorded beside it in the Source cell is preserved across re-runs, a literal `|` included +once escaped. One exception, because it cannot be resolved: a reason ending in the literal phrase +*(not in the current review)* loses that trailing phrase, since it is indistinguishable from the +carried marker the gate appends. The alternative is worse — a stored marker never leaves, so a +carried finding that later reappears would keep claiming it is absent from the review reporting it. + +Re-running the gate keeps every row it can, so a decision recorded here is not +overwritten by a later pass. A finding the current review no longer reports — `--auto` re-reviews +and rewrites REVIEW.md, so this happens routinely — is **carried** rather than dropped, marked +*(not in the current review)*. That holds whether or not it was triaged: losing a decided row would +erase the record that the finding was seen, and losing an *untriaged* one would erase the record +that it was never answered, which is the trace this ledger exists to keep. The cost is that a +renumbered finding shows under both IDs until the old row is decided; the marker makes that legible. +A run that changes nothing rewrites nothing, so a re-executed phase does not produce a docs commit +with no content. + +One residual is concurrency: the ledger is read, rebuilt and written whole, with no lock. Two +dispatchers can run this step (`code_review_gate` and `code-review-fix`'s `record_disposition`) and +a human is invited to hand-edit the file, so two writers overlapping would lose one's update. This is +the shape #3780 reported for `WINDOWS.md` under parallel executors, closed there by a cross-process +lock (#4681); the disposition step does not take that lock. No lost update has been reproduced; the +window is stated so it is not mistaken for a guarantee. + +Not to be confused with the **Review Dispositions Ledger** of reviews-mode planning +([ADR-3806](../adr/3806-review-dispositions-ledger.md), `docs/features/review-dispositions-ledger.md`): +that one is a `## Review Dispositions Ledger` section inside `PLAN.md`, append-only per round, over +`REVIEWS.md` findings. This artifact is a sibling file beside `REVIEW.md`, rewritten idempotently +with rows carried. Same word, different inputs, writers, files and durability rules; neither governs +the other. + +The record is a sibling artifact rather than a section inside REVIEW.md because `--auto`'s +re-review loop rewrites REVIEW.md on every iteration — a ledger kept inside it would not survive +the next pass — and because REVIEW.md has a single writer (`gsd-code-reviewer`) that the gate is +not. The gate remains advisory throughout: it reports and records, and never blocks phase +completion. --- diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 0711da028..e9624103f 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -651,6 +651,7 @@ "complete-milestone/steps/git-tag.md", "discuss-phase-assumptions/steps/auto-advance-dispatch.md", "docs-update/steps/dispatch-monorepo-packages.md", + "execute-phase/steps/code-review-disposition.md", "execute-phase/steps/codebase-drift-gate.md", "execute-phase/steps/completion-reconciliation.md", "execute-phase/steps/executor-isolation-dispatch.md", diff --git a/docs/adr/3806-review-dispositions-ledger.md b/docs/adr/3806-review-dispositions-ledger.md index 798484872..1f9ee68ef 100644 --- a/docs/adr/3806-review-dispositions-ledger.md +++ b/docs/adr/3806-review-dispositions-ledger.md @@ -137,3 +137,19 @@ construction the moment a new capability ships one. - #724 / #728 — established the content requirement this decision only gives a canonical shape to. - [ADR-766](766-claude-code-plugin-manifest-module.md) — precedent for an ADR and its full implementation landing in the same PR. + +## Amendment (2026-09-14): not the code-review disposition ledger + +A second artifact with an adjacent name and the opposite durability rule now exists, and neither +document mentioned the other. `-REVIEW-DISPOSITION.md` (#3829, PR #3861 — +[`docs/features/code-review-pipeline.md`](../features/code-review-pipeline.md)) records a +per-finding disposition (`open` / `fixed` / `skipped` / `deferred`) for the **code-review +pipeline's** `REVIEW.md` findings, written by `execute-phase`'s `code_review_gate` and by +`code-review-fix`'s `record_disposition`. It is a sibling file beside `REVIEW.md`, rewritten +idempotently on every run, with rows carried forward and a human hand-editing `deferred` in place. + +This ADR's ledger is the **reviews-mode planning** one: a `## Review Dispositions Ledger` section +inside `PLAN.md`, append-only per round, over `REVIEWS.md` findings. The two share the word +"disposition" and nothing else — different inputs, different writers, different files, and opposite +rules for what a later run may change. Neither governs the other; a `recall_decision` hit on either +name is not evidence about the other. This section is the cross-reference the adjacency warranted. diff --git a/docs/features/code-review-pipeline.md b/docs/features/code-review-pipeline.md index 4db8ec4d1..ecbdf3634 100644 --- a/docs/features/code-review-pipeline.md +++ b/docs/features/code-review-pipeline.md @@ -17,6 +17,8 @@ group: v1.34.0 Features - REQ-REVIEW-06: `--auto` flag MUST enable fix + re-review iteration loop, capped at 3 iterations - REQ-REVIEW-07: Feature MUST be gated by `workflow.code_review` config flag - REQ-REVIEW-08: `workflow.code_review_point` MUST select which loop point the automatic review step registers at (`execute:post` default, or `execute:wave:post`), independent of the `workflow.code_review` on/off gate and of manual `/gsd-code-review` invocation (#3661) +- REQ-REVIEW-09: The in-phase `code_review_gate` MUST report the per-severity counts it parses from REVIEW.md, so a review with one `info` finding is distinguishable from a review with a Critical +- REQ-REVIEW-10: Each finding MUST carry a recorded disposition, so a triaged finding is distinguishable from a forgotten one **Config:** | Setting | Type | Default | Description | @@ -50,3 +52,140 @@ Escalation is **whole-review, not per-file**: depth is a single scalar handed to v1 supports **directory-prefix matching only, not glob syntax**: no glob engine (`minimatch`, `picomatch`, `fast-glob`) exists in this project and none was added for this feature. A path containing `*` or `?` (e.g. `src/auth/**`) is a configuration error rather than a silent near-miss, because accepting it as sugar for a prefix would make unsupported patterns look armed when they match nothing. Every use case in the issue is expressible as a directory prefix. See [Scope code review depth by path](how-to/scope-code-review-depth-by-path.md) for the resolution order, error table, and a worked example. **Optional external reviewer lanes (#4209):** `/gsd-code-review` accepts the same reviewer-lane flags as `/gsd-review` — any flag the roster declares (run `gsd_run review-lane flags` to list them for your installation, e.g. `--codex`, `--agy`). No reviewer-lane flag is the default and is byte-for-byte unchanged from before #4209: zero lane selection, plan, or invoke calls, and only the internal `gsd-code-reviewer` agent runs. Passing one or more flags asks those lanes to independently review the same already-resolved file scope alongside the internal agent, through the same shared capability-trait interpreter and `review-lane plan`/`invoke` machinery `/gsd-review` uses — no second implementation. Each lane's prompt carries only the repository root, canonical file paths, review depth, and base SHA, never source file contents, under four fixed prohibitions (no source mutation, no test execution, no background processes, no polling). External findings are unverified corroborating evidence: `gsd-code-reviewer` independently re-verifies every claim against the actual source before writing it to `REVIEW.md`, so there remains exactly one `REVIEW.md` schema regardless of how many lanes ran. An explicitly requested lane that is unavailable or fails is reported as a warning, never silently dropped and never a raw-CLI fallback. This is separate from `/gsd-review`, which reviews `PLAN.md` files before execution, not source code. +**In-phase review reporting and disposition** + +`/gsd-execute-phase`'s `code_review_gate` runs code review, then reports what it found: + +``` +Code review: 23 findings — 1 critical, 9 warning, 8 info. +Consider running: /gsd-code-review 1 --fix +``` + +Both `critical:` and its documented tier-equivalent `blocker:` are accepted. A REVIEW.md written +without a `findings:` block has no counts to report, and the gate falls back to the countless form +rather than printing a half-filled line. + +The gate then writes `-REVIEW-DISPOSITION.md` beside the review — one row per finding ID, +defaulting to `open`: + +| Finding | Severity | Disposition | Source | +|---------|----------|-------------|--------| +| CR-01 | critical | open | - | +| WR-01 | warning | fixed | 01-REVIEW-FIX.md | + +`open` means recorded but not yet triaged, and it is the only value the gate assigns on its own. + +**Where the reconciliation happens, which is not where you might expect.** The in-phase gate runs +immediately after review, and at that moment `-REVIEW-FIX.md` does not exist — the gate invokes +review with neither `--fix` nor `--auto` — so every row it writes is `open`. The `fixed` and +`skipped` outcomes are reconciled by `/gsd-code-review --fix`, which records them once the fix +report is on disk. Running the gate alone therefore tells you what was found; running `--fix` is +what records what happened to it. + +**`--auto`'s iterations are reconciled too, and that takes reading more than the final report.** +The loop overwrites `REVIEW-FIX.md` on every pass and the re-review stops reporting a finding once +it is fixed, so a finding closed in iteration 1 appears in neither final artifact. The gate +therefore also reads the per-iteration backups the loop writes (`-REVIEW-FIX.iterN.md`), newest +first, so the most recent statement about an ID wins; the backups are removed after the ledger has +read them, not before. Without this a fully successful multi-iteration run recorded its early fixes +as `open (not in the current review)` — indistinguishable from a finding that vanished for an +unrelated reason, which is the one distinction this ledger exists to make. A finding a fix report +decided but the current review no longer reports gets a row of its own, carrying that decision and +marked *(not in the current review)*. + +A converged `--auto` run also reaches the gate with a clean review and, on a direct +`/gsd-code-review` invocation, no ledger from the in-phase gate. A fix report on disk is reason +enough to record: without that, a run in which every finding was fixed and committed produced no +disposition record at all. + +**A decision belongs to a finding, not to an ID.** IDs are reused across re-reviews — `--auto` +renumbers — so the ledger records each finding's title alongside its disposition, and a recorded +decision is carried forward only while the ID still names the same finding. Without that, a prior +`CR-01 fixed` would be inherited by a brand-new `CR-01`, which is a false decision in the one +artifact whose purpose is telling triaged from forgotten. + +**The limitation that follows, stated rather than hidden.** When an ID *is* reused, the new finding +is recorded `open` (nobody decided anything about it) and the earlier decision **loses its row** — +the ledger keys rows on the finding ID, and two rows under one ID is an ambiguity, not a record. The +drop is reported on the console, naming the ID and what had been decided — for a *recorded* decision. +A prior row still sitting at `open` is replaced silently, and deliberately: `open` means nobody had +decided anything, so there is no decision to lose. Preserving a recorded decision in the file +was tried and withdrawn: it needed a second identity scheme and produced a fresh defect on each of +three review passes. Recovery is the ledger's own git history where `commit_docs` is on — which is +why the console reports the drop rather than pointing at a commit that may not exist. + +**Two residuals, since (ID, title) is not proof of identity.** A ledger written before titles were +recorded carries none, so its decisions are inherited on the ID alone — refusing there would reset +every decision in every existing ledger, which is the loss the guard exists to prevent. And two +genuinely distinct findings that share both an ID and a title are indistinguishable to this key. +Separating them needs a second identity scheme, which is the thing that was just withdrawn. + +Reconciliation applies an outcome only when the fix report names the **same** finding, because +finding IDs are reused across re-reviews and a stale report would otherwise declare a brand-new +`CR-01` already fixed. Titles are compared with runs of whitespace collapsed, so a fixer that +re-spaces a title still reconciles. A title that differs otherwise — including one **wrapped across +lines**, since a `###` heading is one line and the continuation is a separate paragraph — leaves the +row `open` and is reported, naming the ids it could not reconcile, rather than passing silently. The report does not claim to +know whether such a report is stale or merely re-titled, because it cannot tell. + +The disposition column is a closed vocabulary — `open`, `fixed`, `skipped`, `deferred`. A +hand-edited value outside it is not treated as a decision: the row falls back to `open`, so a +typo cannot quietly mark a phase as triaged. + +Severity comes from the section a finding sits under (`## Critical Issues`, `## Warnings`, +`## Info`) when the review uses those headings, and from the ID prefix otherwise. The section is +the reviewer's own statement of severity, so a Critical filed under `## Critical Issues` is +recorded critical even if its ID was mis-numbered `WR-04`. That severity is then **remembered**: +a row the current review no longer reports, or reports under no recognized heading, keeps the +severity the ledger recorded rather than having it re-inferred from the ID prefix — so the +mis-numbered `WR-04` above stays `critical` on the second run, deferred or not. The precedence is +the current review's section, then the recorded value, then the prefix; the recorded value is +inherited only while the ID still names the same finding, by the same title check the disposition +uses, so a reused ID starts from its own review. That check has the same compatibility arm the +disposition has: a ledger written before titles were recorded carries no title to compare, so its +severities — like its decisions — inherit on the ID alone. + +If the review's `total:` exceeds the number of findings whose headings the gate could parse, the +shortfall is stated — on the console and as an `unparsed:` key in the ledger's frontmatter. A +finding the gate cannot record is the one a human most needs to see, so it is never dropped +silently. The one input that produces no shortfall is a `findings:` block that disagrees with +itself: where `critical` (or `blocker`), `warning` and `info` are all present, numeric and do not +sum to `total`, the gate has no trustworthy number to reconcile against and withholds the key +rather than reporting a figure derived from one — the same input on which the console line +already withholds the severity breakdown. Counts that are merely absent, partial or non-numeric +are not a disagreement, and the shortfall is still reported from `total` alone. `deferred` is the one disposition the gate never writes: it is recorded by hand, and the +reason recorded beside it in the Source cell is preserved across re-runs, a literal `|` included +once escaped. One exception, because it cannot be resolved: a reason ending in the literal phrase +*(not in the current review)* loses that trailing phrase, since it is indistinguishable from the +carried marker the gate appends. The alternative is worse — a stored marker never leaves, so a +carried finding that later reappears would keep claiming it is absent from the review reporting it. + +Re-running the gate keeps every row it can, so a decision recorded here is not +overwritten by a later pass. A finding the current review no longer reports — `--auto` re-reviews +and rewrites REVIEW.md, so this happens routinely — is **carried** rather than dropped, marked +*(not in the current review)*. That holds whether or not it was triaged: losing a decided row would +erase the record that the finding was seen, and losing an *untriaged* one would erase the record +that it was never answered, which is the trace this ledger exists to keep. The cost is that a +renumbered finding shows under both IDs until the old row is decided; the marker makes that legible. +A run that changes nothing rewrites nothing, so a re-executed phase does not produce a docs commit +with no content. + +One residual is concurrency: the ledger is read, rebuilt and written whole, with no lock. Two +dispatchers can run this step (`code_review_gate` and `code-review-fix`'s `record_disposition`) and +a human is invited to hand-edit the file, so two writers overlapping would lose one's update. This is +the shape #3780 reported for `WINDOWS.md` under parallel executors, closed there by a cross-process +lock (#4681); the disposition step does not take that lock. No lost update has been reproduced; the +window is stated so it is not mistaken for a guarantee. + +Not to be confused with the **Review Dispositions Ledger** of reviews-mode planning +([ADR-3806](../adr/3806-review-dispositions-ledger.md), `docs/features/review-dispositions-ledger.md`): +that one is a `## Review Dispositions Ledger` section inside `PLAN.md`, append-only per round, over +`REVIEWS.md` findings. This artifact is a sibling file beside `REVIEW.md`, rewritten idempotently +with rows carried. Same word, different inputs, writers, files and durability rules; neither governs +the other. + +The record is a sibling artifact rather than a section inside REVIEW.md because `--auto`'s +re-review loop rewrites REVIEW.md on every iteration — a ledger kept inside it would not survive +the next pass — and because REVIEW.md has a single writer (`gsd-code-reviewer`) that the gate is +not. The gate remains advisory throughout: it reports and records, and never blocks phase +completion. diff --git a/gsd-core/workflows/code-review-fix.md b/gsd-core/workflows/code-review-fix.md index 3b69339a2..5ad46303c 100644 --- a/gsd-core/workflows/code-review-fix.md +++ b/gsd-core/workflows/code-review-fix.md @@ -259,10 +259,11 @@ if [ "$AUTO_MODE" = "true" ]; then # Total fix passes = MAX_ITERATIONS. Loop uses -lt (not -le) intentionally. ITERATION=1 MAX_ITERATIONS=3 - # #3190: track whether the loop converged (re-review came back clean) vs - # degraded (hit the cap). Convergence determines whether the .iterN.md backups - # are spent scratch (cleaned below) or retained for post-mortem analysis. - CONVERGED=false + # #3190's convergence-vs-degradation distinction still governs whether the .iterN.md backups are + # spent scratch or a post-mortem trail — but the removal moved to `cleanup_iteration_backups`, + # after `record_disposition` has read them, so nothing in THIS loop consumes the flag any more. + # It is re-derived there from the final REVIEW.md's status, which is exactly the condition the + # `break` below fires on. A variable set here and read nowhere would just be a decoy. while [ $ITERATION -lt $MAX_ITERATIONS ]; do ITERATION=$((ITERATION + 1)) @@ -322,7 +323,8 @@ ${AGENT_SKILLS_REVIEWER}") " 2>/dev/null) if [ "$NEW_STATUS" = "clean" ]; then - CONVERGED=true + # Convergence: the re-review came back clean. This leaves REVIEW.md at `status: clean`, which + # is what `cleanup_iteration_backups` re-derives the decision from. echo "" echo "✓ All issues resolved after iteration ${ITERATION}." break @@ -356,28 +358,44 @@ ${AGENT_SKILLS_FIXER}") fi done - # After loop completes + # After loop completes. The iteration COUNTER alone does not distinguish degradation from success: + # a loop that converges on the final iteration exits with ITERATION == MAX_ITERATIONS and printed + # "Reached maximum iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which + # every finding was fixed — telling the operator the opposite of what happened, on the one path + # where the cap and convergence coincide. Convergence is re-derived from the review the loop left + # behind, the same signal cleanup_iteration_backups reads, so the two cannot disagree. if [ $ITERATION -ge $MAX_ITERATIONS ]; then - echo "" - echo "⚠ Reached maximum iterations (${MAX_ITERATIONS}). Remaining issues documented in REVIEW-FIX.md." + LOOP_END_STATUS=$(REVIEW_PATH="${REVIEW_PATH}" node -e " + const fs = require('fs'); + try { + const content = fs.readFileSync(process.env.REVIEW_PATH, 'utf-8'); + const match = content.replace(/\r\n/g, '\n').match(/^---\n([\s\S]*?)\n---/); + console.log(match && /status:\s*(\S+)/.test(match[1]) ? match[1].match(/status:\s*(\S+)/)[1] : 'unknown'); + } catch (e) { console.log('unknown'); } + " 2>/dev/null) + if [ "$LOOP_END_STATUS" = "clean" ]; then + # SILENT on convergence, deliberately. The loop's own break already printed "All issues + # resolved after iteration N" on its way out, and reaching the cap does not make that less + # true — adding a second success line here printed both, which is noise the first cut shipped. + # What this branch exists for is to NOT print the degradation warning; saying nothing is the + # whole behaviour. + : + else + echo "" + echo "⚠ Reached maximum iterations (${MAX_ITERATIONS}). Remaining issues documented in REVIEW-FIX.md." + fi fi - # #3190: on convergence the .iterN.md backups are spent scratch — their - # stated purpose is post-mortem analysis "if iterations degrade", and - # convergence means no degradation. Remove them so the phase directory is - # clean (no dirty REVIEW.md or backup files after the run). They are RETAINED - # when the loop degraded (hit MAX_ITERATIONS / fixer failure) so the - # post-mortem trail survives. Backup CREATION (cp … .iterN.md) is unchanged. - if [ "$CONVERGED" = "true" ]; then - rm -f "${REVIEW_PATH%.md}.iter"*.md "${FIX_REPORT_PATH%.md}.iter"*.md 2>/dev/null || true - fi + # #3190's cleanup of the .iterN.md backups now runs in `cleanup_iteration_backups`, AFTER + # `record_disposition` — see that step for why. Removing them here deleted the only record of + # what the earlier iterations fixed before anything had read it. fi ``` Key design decisions for --auto (addresses ALL review HIGH concerns): 1. **Re-review scope**: Uses REVIEW_FILES_ARRAY from original REVIEW.md frontmatter, falling back to full phase scope. Scope is NOT lost between iterations. Uses portable while-read loop (bash 3.2+ compatible, handles spaces in paths). 2. **Artifact semantics**: REVIEW.md is overwritten by each re-review (latest review state). REVIEW-FIX.md is overwritten by each fixer iteration (latest fix state with iteration count). There is ONE final version of each artifact, not per-iteration copies. - Backup files (.iterN.md) preserve history for post-mortem analysis if iterations degrade. On successful convergence (#3190) the backups are spent scratch and removed; on degradation (hit MAX_ITERATIONS / fixer failure) they are retained for post-mortem. + Backup files (.iterN.md) preserve history for post-mortem analysis if iterations degrade. On successful convergence (#3190) the backups are spent scratch and removed; on degradation (hit MAX_ITERATIONS / fixer failure) they are retained for post-mortem. The removal happens in `cleanup_iteration_backups`, after `record_disposition` has read them — they are the only surviving record of what an earlier iteration fixed, since this artifact keeps one final version rather than per-iteration copies. 3. **Commit timing**: Fix commits happen per-finding inside the agent. REVIEW-FIX.md is NOT committed until step 7 (after ALL iterations complete). Only ONE docs commit, not one per iteration. In --auto that single commit also stages the converged REVIEW.md alongside REVIEW-FIX.md (#3190), so the two committed artifacts agree — the initial code-review commit held iteration-1 REVIEW.md content, and the --auto re-review loop overwrote it each iteration. @@ -427,6 +445,73 @@ fi This commit happens ONCE at the end of the workflow, after all iterations (if --auto) complete. Not per-iteration. + +**Reconcile the per-finding disposition ledger, now that REVIEW-FIX.md exists.** Read and execute +`gsd-core/workflows/execute-phase/steps/code-review-disposition.md`. It consumes `PHASE_DIR` and +`PHASE_NUMBER` and derives everything else, and it is advisory throughout: it never blocks. + +This is the step's second and final call site, and without it REQ-REVIEW-10 is unreachable in every +shipped path. `execute-phase.md`'s `code_review_gate` runs the same step immediately after review, +where `-REVIEW-FIX.md` cannot yet exist — the gate invokes review with neither `--fix` nor +`--auto` — so every row it writes is `open` by construction. The reconciliation logic that turns +those rows into `fixed` / `skipped` is reachable only once a fix report is on disk, which is here. +Wired anywhere earlier and it reads a report that has not been written; wired into `code-review.md` +instead, it is unreachable — that workflow delegates through `code-review/steps/dispatch-fix.md`, +which exits the workflow after invoking this one, so there is no point in it that is after +REVIEW-FIX.md exists. Placing it here also covers a direct invocation of this workflow, which a +wiring in `code-review.md` would miss. + +Placed AFTER `commit_fix_report` deliberately: the fix report must be on disk and committed before +the ledger claims anything about it, and the step's own reconciliation reads +`${PHASE_DIR}/${PADDED_PHASE}-REVIEW-FIX.md` directly. + +Safe to run twice. The step is idempotent — a re-render that changes nothing reports `unchanged` and +rewrites no file — so a phase that reaches the gate and then a fix run ends with one ledger reflecting +both, not two competing ones. + + + +Only runs if AUTO_MODE is true. If AUTO_MODE is false, skip this step entirely. + +**Removes the `.iterN.md` scratch — deliberately AFTER `record_disposition`, not at the end of the +loop.** #3190's rule is unchanged: on convergence the backups are spent scratch and go; on degradation +they are retained for post-mortem. What changed is the timing, and it was load-bearing. This workflow +keeps ONE final version of `REVIEW.md` and `REVIEW-FIX.md` rather than per-iteration copies, so the +backups are the only surviving record of what an earlier iteration fixed — and the re-review drops a +finding once it is fixed, so the final review does not carry it either. Deleting them inside the loop +meant the disposition ledger reached a converged `--auto` run with every early fix already erased, and +recorded those findings as `open (not in the current review)`: indistinguishable from never triaged, +which is the one distinction #3829 exists to make. + +Convergence is re-derived here rather than carried. Shell state does not survive the loop's fence, so +a flag set there would be gone by this step — and once nothing in the loop consumed it, keeping it +would have left a variable set in one place and read in none. The loop breaks on exactly one +condition, a re-review returning clean, and that leaves `REVIEW.md` at `status: clean`: the final +review's status IS the converged case. Hitting `MAX_ITERATIONS` leaves the last re-review non-clean, +and the backups are kept. + +```bash +if [ "$AUTO_MODE" = "true" ]; then + FINAL_STATUS=$(REVIEW_PATH="${REVIEW_PATH}" node -e " + const fs = require('fs'); + try { + const content = fs.readFileSync(process.env.REVIEW_PATH, 'utf-8'); + const match = content.replace(/\r\n/g, '\n').match(/^---\n([\s\S]*?)\n---/); + console.log(match && /status:\s*(\S+)/.test(match[1]) ? match[1].match(/status:\s*(\S+)/)[1] : 'unknown'); + } catch (e) { console.log('unknown'); } + " 2>/dev/null) + # Anything but a proven-clean final review RETAINS the backups. An unreadable or unparseable + # review is not evidence of convergence, and retaining costs a few scratch files where deleting + # costs the post-mortem trail the retention rule exists for. + if [ "$FINAL_STATUS" = "clean" ]; then + rm -f "${REVIEW_PATH%.md}.iter"*.md "${FIX_REPORT_PATH%.md}.iter"*.md 2>/dev/null || true + else + echo "Iteration backups retained (final review status: ${FINAL_STATUS}) — post-mortem trail." + fi +fi +``` + + Parse REVIEW-FIX.md frontmatter and present formatted summary to user. diff --git a/gsd-core/workflows/execute-phase.md b/gsd-core/workflows/execute-phase.md index 6691a2fe5..1c4759b39 100644 --- a/gsd-core/workflows/execute-phase.md +++ b/gsd-core/workflows/execute-phase.md @@ -1162,20 +1162,11 @@ If no active code-review step hook exists: display "Code review skipped (code-re Skill(skill="gsd-${ref.skill}", args="${PHASE_NUMBER}") ``` -**Check results using deterministic path (not glob):** -```bash -# #4748: bind init's normalized id — `printf "%02d"` cannot pad a letter id -# (03A → `03`, exit 1) and reads an already-padded `08` as octal (→ `00`). -PADDED="{padded_phase}" -REVIEW_FILE="${PHASE_DIR}/${PADDED}-REVIEW.md" -REVIEW_STATUS=$(sed -n '/^---$/,/^---$/p' "$REVIEW_FILE" | grep "^status:" | head -1 | cut -d: -f2 | tr -d ' ') -``` - -If REVIEW_STATUS is not "clean" and not "skipped" and not empty, display: -``` -Code review found issues. Consider running: -/gsd:code-review ${PHASE_NUMBER} --fix -``` +**Report the review, and record what happened to each finding.** Read and execute `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`. +It parses REVIEW.md's frontmatter, states the per-severity counts, and writes +`-REVIEW-DISPOSITION.md` — one row per finding, defaulting to `open` — so a triaged finding is +distinguishable downstream from a forgotten one. It consumes `PHASE_DIR` and `PHASE_NUMBER`, and is +advisory throughout: it never blocks. **Error handling:** If the Skill invocation fails or throws, catch the error, display "Code review encountered an error (non-blocking): {error}" and proceed to gate dispatch. Review failures must never block execution. diff --git a/gsd-core/workflows/execute-phase/steps/code-review-disposition.md b/gsd-core/workflows/execute-phase/steps/code-review-disposition.md new file mode 100644 index 000000000..f4a54a20d --- /dev/null +++ b/gsd-core/workflows/execute-phase/steps/code-review-disposition.md @@ -0,0 +1,1017 @@ +# `code_review_gate` — report the review and record a per-finding disposition + +Read and executed by `execute-phase.md`'s `code_review_gate` step, immediately after code review +returns. It consumes `PHASE_DIR` and `PHASE_NUMBER` and derives everything else. + +It lives here rather than inline in the parent because `execute-phase.md` sits against two size +ceilings — the XL hard cap in `tests/workflow-size-budget.test.cjs` and the frozen ADR-857 +pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs` — and both are red lines to be kept +under, not budgets to spend. + +**What it is for.** The counts the gate prints say how many findings there were, not what happened +to any of them. Without a record, the phase directory carries no answer to *what happened to CR-01* +and a phase can reach `phase.complete` with a Critical standing and no trace it was ever seen. + +**Why a sibling artifact rather than a section inside REVIEW.md.** `--auto`'s re-review loop +rewrites REVIEW.md on every iteration, so a ledger kept inside it would not survive the next pass; +and REVIEW.md has a single writer, `gsd-code-reviewer`, which this step is not. + +**Advisory — it never blocks.** Every failure path reports and steps over. + +**Check results using deterministic path (not glob):** +```bash +# PADDED must survive a DOTTED phase number, of ANY segment count. This step is dispatched from +# exactly TWO places: `execute-phase.md` (`code_review_gate`) and `code-review-fix.md` +# (`record_disposition`). Only the second validates anything -- `code-review-fix.md`'s PADDED_PHASE validator anchors +# `^[0-9]+[A-Z]?(\.[0-9]+)*$`, an unbounded `*` widened by #4568 and a letter axis widened by #4744, so it accepts `03.1`, `23.1.2` AND `12A`. +# `execute-phase.md` applies NO shape gate at all, so this fence is not mirroring an upstream +# guarantee; it IS the guarantee. (`code-review.md`'s own PADDED_PHASE validator is identical but never +# dispatches this step. It was cited here as a caller for several rounds and is not one.) And +# `printf "%02d"` cannot format one: bash prints `invalid number` and exits 1, which under +# `set -euo pipefail` aborts this step on its FIRST line -- the loudest possible failure from +# the gate that promises never to block, and it takes the whole phase's review reporting with +# it. Pad the integer part only and carry the sub-number verbatim, so 3.1 -> 03.1, 23.1.2 -> +# 23.1.2 and 3 -> 03. The segment count is deliberately NOT bounded here: the canonical +# grammar in src/phase-id.cts (`PHASE_NUMBER_TOKEN_SOURCE`, #2128) is unbounded in segments, +# and #4568 widened the one dispatcher that validates to match it on that axis, so a guard +# narrower than that dispatcher means no ledger for a phase id the dispatcher already accepted. +# On failure NO path is built and the fence refuses by name: advisory means advisory, and it +# also means never probing a path assembled out of a value we just rejected. +# VALIDATE, THEN FORMAT -- never format and fall back on failure. `printf "%02d" abc` writes +# `00` to stdout BEFORE it fails, so a `$(printf ... || printf %s ...)` fallback CONCATENATES +# the two and yields `00abc`; `08` fails the same way as invalid octal, giving `0008.1` for a +# legitimate `08.1`. Both were driven. `${PHASE_NUMBER:-}` because an UNSET input must not trip +# `set -u` in a step that promises not to abort. Those printf failures are why this block VALIDATES +# instead of formatting; the pad itself performs NO arithmetic since round 14 (see the +# `case "${#_dig}"` line below), so it has no octal hazard to guard and needs no `10#`. `10#` +# survives in this step only where it still belongs -- on the severity COUNTS, which really are +# numbers being added. +# VALIDATE THE WHOLE VALUE, then format -- and on failure build NO path at all. +# Carrying an unusable value verbatim was the first draft and it was worse than the bug it +# replaced: PHASE_NUMBER is interpolated into a file path, so `../../etc/passwd` produced +# `${PHASE_DIR}/../../etc/passwd-REVIEW.md`, where the old `printf "%02d"` had at least +# mangled it to `00`. `code-review-fix.md`'s PADDED_PHASE validator already checks `^[0-9]+[A-Z]?(\.[0-9]+)*$` against its +# own PADDED_PHASE -- the padded form, not the raw PHASE_NUMBER this step is handed -- while +# `execute-phase.md` validates nothing at all; this step has two call sites and validates for +# itself rather than trusting either. Anything else yields an EMPTY PADDED and the blocks +# below refuse to build a path from it. +# PHASE_DIR is checked for NON-EMPTINESS ONLY. Both inputs come from the caller's init query, so +# neither is raw user input; only PHASE_NUMBER has a SHAPE (`^[0-9]+[A-Z]?(\.[0-9]+)*$`) to check +# against. A filesystem path admits `..` and symlinked parents alike, so a shape +# check here rejects working setups and proves nothing. Residual: PHASE_DIR may itself be a symlink +# and the ledger is written through it -- left alone, and not a security boundary. +_pd="${PHASE_DIR:-}" +_pn="${PHASE_NUMBER:-}" +_ok=1 +[ -n "$_pd" ] || _ok=0 +case "$_pn" in + ''|*[!0-9.A-Z]*) _ok=0 ;; # empty, or any character outside [0-9.A-Z] -- this is the traversal fence + .*|*.) _ok=0 ;; # leading or trailing dot + *..*) _ok=0 ;; # EMPTY SEGMENT. The three arms plus the letter-axis block below + # accept exactly digits[LETTER](.digits)* -- byte-congruent with the + # callers' ^[0-9]+[A-Z]?(\.[0-9]+)*$ -- rather than merely wider than + # the retired `*.*.*` arity bound, which masked `1..2` by accident. +esac +# THE LETTER AXIS. The canonical grammar (src/phase-id.cts) is digits, an OPTIONAL single uppercase +# letter, then dotted digit segments -- `12A`, `3A`, `23A.1.2`. #4744 (#4660) widened the six +# shell/markdown mirrors to it after this branch was cut, and its lint ratchet then flagged this +# step as the one letterless mirror left. The character class above admits the letter; these +# arms pin WHERE it may sit -- only as the last character of the integer part, at most once -- +# so `23a`, `A23`, `2A3`, `23AB` and `23.1A` are all refused. +if [ "$_ok" = "1" ]; then + _int="${_pn%%.*}" + case "$_pn" in *.*) _sub=".${_pn#*.}" ;; *) _sub="" ;; esac + _let="${_int##*[0-9]}" # what trails the last digit: '' or the letter + _dig="${_int%"$_let"}" + case "$_dig" in ''|*[!0-9]*) _ok=0 ;; esac # the integer part must be digits first + case "$_let" in ''|[A-Z]) ;; *) _ok=0 ;; esac # at most ONE letter, uppercase + case "$_sub" in *[!0-9.]*) _ok=0 ;; esac # no letter in any later segment +fi +# LENGTH-BOUND EACH COMPONENT SEPARATELY -- and the REASON changed at round 14, so read this rather +# than inherit it. It used to be integer overflow: the pad ran `$((10#$_int))`, bash integers wrap at +# 2^64, and a 54-digit value yielded -7908320945662590977 SILENTLY as the padded phase. The pad is a +# string pad now and converts nothing, so that overflow is unreachable and its rationale is dead. +# WHAT THE BOUND STILL DOES, stated narrowly because the obvious wider claim is FALSE: it bounds each +# SEGMENT, and NOTHING here bounds the COMPOSITE. Every segment is joined into ONE filename +# component and depth is unbounded, so a per-segment bound does not enforce a filename limit -- +# driven at round 14: thirty 8-digit segments yield a 278-character PADDED and a 294-character name +# against a NAME_MAX of 255. That is a real residual of this validator, it predates the pad change, +# and it is named here rather than papered over with a filesystem rationale the bound does not +# deliver. The bound belongs on the INTEGER PART: applied to the whole value it rejected +# `12345678.1`, whose integer part is a legal 8 digits, while accepting `1.123456` -- an accidental +# bound on the composite that was both too strict and too loose. Every later segment is bounded too, +# on the same narrow reading. +# THE LOOP IS THE POINT, and it is what makes the heading above TRUE. The earlier form bounded +# `${_pn#*.}` -- the WHOLE tail after the first dot -- which is one component only while the id +# has at most two. Once N-segment ids are accepted (see the shape arms), that form rejects +# `1.1234567.1`, whose every component is a legal 7-or-fewer digits, purely because the tail +# measures 9 characters. That is the composite bound this comment already called "too strict", +# surviving one level up. Walk the segments instead, so the rule is per-component in fact and +# not only in the heading. The loop terminates on any string the shape arms admit: each pass +# strips a leading `.`, and the no-dot pass clears $_rest. +if [ "$_ok" = "1" ]; then + _rest="$_pn" + while [ -n "$_rest" ]; do + case "$_rest" in + *.*) _seg="${_rest%%.*}"; _rest="${_rest#*.}" ;; + *) _seg="$_rest"; _rest="" ;; + esac + # The bound is on the DIGITS: a letter suffix is one character the digit bound has no + # stake in, so `12345678A` is within it exactly as `12345678` is. + case "${_seg%[A-Z]}" in ?????????*) _ok=0 ;; esac + done +fi +if [ "$_ok" = "1" ]; then + # padStart(2,'0'), EXACTLY, and as a STRING -- the canonical normalizer left-pads the digit run to a + # MINIMUM of two and otherwise preserves it, so arithmetic is the wrong tool. `printf "%02d" + # "$((10#$_dig))"` agreed on `8`/`08`/`09` and silently DISAGREED on every longer leading-zero run: + # `008` -> `08`, `0008A` -> `08A`, resolving a REVIEW.md path init never writes. Driven at round 14 + # by the property that asserts agreement with `normalizePhaseName` over generated ids; the 13-shape + # matrix that preceded it sampled no run longer than two and could not see it. Dropping the + # arithmetic also RETIRES the octal hazard `10#` existed to work around, rather than guarding it. + # The letter and any dot segments ride along verbatim, as the canonical padder does: `3A` -> `03A`. + case "${#_dig}" in 1) PADDED="0${_dig}${_let}${_sub}" ;; *) PADDED="${_dig}${_let}${_sub}" ;; esac +else + PADDED="" +fi +# REFUSE BEFORE BUILDING ANY PATH. An unusable input yields an empty PADDED above, and the +# earlier placement -- after the assignments -- meant a rejected value still had +# `${_pd}/-REVIEW.md` assembled and stat'ed before the refusal fired. Nothing is constructed +# from a value we have already rejected. +if [ -z "$PADDED" ]; then + echo "Code review reporting skipped (unusable phase number or directory: '${PHASE_NUMBER:-}')" + return 0 2>/dev/null || exit 0 +fi +REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md" +DISPOSITION_FILE="${_pd}/${PADDED}-REVIEW-DISPOSITION.md" +# Extract ONLY the leading frontmatter block: `sed -n '/^---$/,/^---$/p'` re-opens its range +# on a body `---` and runs to EOF, which leaks body lines into the scan. That leak is benign +# for a key the frontmatter always carries (the first match still wins) but NOT for an +# optional one — a review with no `findings:` block and a body `total:` line would otherwise +# report the body's number as the count. Stop at the closing delimiter instead, and strip CR +# first so a CRLF-authored review neither breaks the delimiter match nor injects a carriage +# return into the message below (DEFECT.FRONTMATTER-SCALAR-BROAD-GREP). +# Buffered, and emitted only if the CLOSING delimiter was actually seen: an unterminated +# frontmatter block would otherwise run to EOF and hand the whole review body to the reads below, +# defeating the scoping entirely. +# Guarded and `|| true`: this step is advisory, so a REVIEW.md that is missing, a directory, or +# otherwise unreadable must leave the counts empty and let execution continue — never abort the +# step under `set -e`/`pipefail`. +# REVIEW_READ records that the file was actually OPENED, separately from what it yielded. An +# absent or unreadable review and a present-but-unparseable one both leave every value below +# empty, and the reporting arm used to treat the two identically -- silence -- so a REVIEW.md +# with three criticals and an unterminated frontmatter read exactly like a clean review. A +# malformed report must not read as a clean one; the arm below tells them apart on this flag. +REVIEW_FM="" +REVIEW_READ=0 +if [ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]; then + REVIEW_READ=1 + REVIEW_FM=$(tr -d '\r' < "$REVIEW_FILE" 2>/dev/null | LC_ALL=C awk 'NR==1{if($0!="---") exit; next} /^---$/{closed=1; exit} {buf = buf $0 "\n"} END{if (closed) printf "%s", buf}' || true) +fi +# `|| true` on every read: under `pipefail` a non-matching `grep` exits 1, and an assignment +# whose command substitution fails aborts the step under `set -e`. An advisory gate must survive +# a REVIEW.md with no frontmatter at all. +# STATUS TAKES THE SAME PARSER AS THE COUNTS, and it is the read where truncation costs most. Under +# `cut -d: -f2` the valid YAML scalar `status: clean:junk` arrived as the bare `clean` -- so an +# unusable status SILENTLY took the clean arm, suppressing both the report and the ledger. `-f2-` +# keeps the whole scalar, `clean:junk` matches no arm, and the step reports. Found by the round's +# fourth adversarial pass as a sibling of the count-parser class, in the same file. +REVIEW_STATUS=$(echo "$REVIEW_FM" | LC_ALL=C grep -m1 "^status:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) +# The counts belong to the `findings:` MAPPING, not merely to the frontmatter, and the scoping now +# goes all the way there. `^[[:space:]]*total:` matches any indented key anywhere in the block, so +# a top-level key later named `total:`, `info:` or `critical:` was picked up ahead of the nested +# one — the extensive comment above is about scoping the frontmatter, and the scoping stopped one +# level short of the mapping the values actually live in. `status:` was never exposed: it is +# anchored to column 0 because it IS top-level. +# The awk selects the `findings:` block and stops at the next column-0 key, so the reads below can +# only see keys nested under it. Block 2 derives REVIEW_TOTAL through the same filter. +# `blocker:` is the documented tier-equivalent of `critical:` (gsd-code-reviewer.md § "Label +# equivalence") — accept either, exactly as code-review.md's present_results already does. +REVIEW_FINDINGS_FM=$(echo "$REVIEW_FM" | LC_ALL=C awk '/^findings:[[:space:]]*$/{f=1; next} f&&/^[^[:space:]]/{exit} f' || true) +REVIEW_CRITICAL=$(echo "$REVIEW_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*(critical|blocker):" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) +REVIEW_WARNING=$(echo "$REVIEW_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*warning:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) +REVIEW_INFO=$(echo "$REVIEW_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*info:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) +REVIEW_TOTAL=$(echo "$REVIEW_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*total:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) +# ONE PARSER FOR THE WHOLE STEP. These reads used `cut -d: -f2 | tr -d ' '`, which repairs a +# malformed scalar into a number twice over: `tr -d` deletes INTERNAL spaces (`1 0` -> `10`) and +# `-f2` keeps only the SECOND FIELD (`1: junk` -> `1`). Block 2 was tightened first, which left the +# two fences disagreeing about the same bytes -- driven: `critical: 1 0` made this block print +# `10 findings -- 10 critical` while the ledger recorded three rows and no shortfall. A console line +# and a ledger contradicting each other is the exact confusion this PR exists to remove, so the fix +# is one parser rather than a disclosed divergence. `-f2-` keeps the whole scalar; only the ends are +# trimmed. A repaired number is not a number. +# LC_ALL=C ON EVERY `grep`/`sed` IN THESE READS, and it is load-bearing rather than cosmetic: the +# POSIX classes are LOCALE-DEFINED, and glibc's C.UTF-8 puts U+2003 (and U+1680, U+2000-U+200A, +# U+205F, U+3000) in BOTH [[:space:]] and [[:blank:]], where C and en_US.UTF-8 put them in neither. +# Unpinned, `status: clean` trimmed to `clean` under one locale and stayed unusable under +# another -- the same silent suppression as the `clean:junk` truncation, reachable only on some +# machines. Pinned to C the class is exactly {space, tab, NL, VT, FF, CR}, which is what the mirror +# in tests/code-review-pipeline-regression.test.cjs spells out literally, so the two agree by +# construction rather than by coincidence of locale. +# The breakdown is reportable only when ALL FOUR counts are numbers. Deciding on REVIEW_TOTAL +# alone would still emit `6 findings — critical` for a review carrying a total and nothing else. +REVIEW_COUNTS_OK=1 +for _c in "$REVIEW_TOTAL" "$REVIEW_CRITICAL" "$REVIEW_WARNING" "$REVIEW_INFO"; do + # Length-bounded as well as digit-only: bash integers wrap at 2^64, so a 20-digit count + # arrives at the sum below as 0 and an inconsistent breakdown passes. No real review + # reports nine digits of findings. + case "$_c" in ''|*[!0-9]*) REVIEW_COUNTS_OK=0 ;; ?????????*) REVIEW_COUNTS_OK=0 ;; esac +done +# Numeric is necessary and not sufficient. `total: 0` beside `critical: 1` is four valid numbers +# that render the self-contradicting line `0 findings — 1 critical, 0 warning, 0 info`. An +# inconsistent breakdown is unavailable for the same reason a partial one is: half-true is worse +# than withheld, and the countless form is already the documented fallback. +# `10#` on every operand: bash infers the base from a leading zero, so a review reporting +# `critical: 08` makes $(( )) fail with "value too great for base". The CONSEQUENCE stated here +# used to be "takes the whole advisory step down under `set -e`", and that is WRONG: the +# arithmetic sits inside an `if` condition, a TESTED context, where `set -e` is inert. What +# actually happens is that the consistency check is SKIPPED -- which the regression suite +# already records. Skipping it is still the wrong outcome for a check whose whole job is +# refusing a half-true breakdown, so `10#` stays; only the account of what it prevents is +# corrected. The values are already digit-only by the loop above. +if [ "$REVIEW_COUNTS_OK" = "1" ] \ + && [ "$((10#$REVIEW_CRITICAL + 10#$REVIEW_WARNING + 10#$REVIEW_INFO))" -ne "$((10#$REVIEW_TOTAL))" ]; then + REVIEW_COUNTS_OK=0 +fi +# EMIT — inside the fence, on every reporting arm. Until this block existed, the fence computed six +# values and printed none of them, and the prose below then asked the agent to display four of them. +# The shell exits at the closing fence and the agent sees only stdout, so those values were +# unobtainable: the message could not be rendered, and the whole block was decorative. That is the +# rule block 2 states about itself — a prose-only gate on a value no later block can see is not a +# gate — applied to the block that is this step's primary deliverable rather than only to its +# sibling. The status arm is re-derived here, not left to the reader, for the same reason. +case "$REVIEW_STATUS" in + '') + # NO STATUS is not NO REVIEW. When the file was read and yielded no status -- unterminated + # frontmatter, no frontmatter, no `status:` key, a zero-byte file -- the review is + # UNPARSEABLE, and saying nothing would make it indistinguishable from a clean one. State it, + # without the breakdown (there is none to trust) and without the --fix suggestion (nothing + # here proves there are findings to fix). An absent or unreadable review stays silent: there + # is nothing to describe, and guessing is the failure the guard above exists to prevent. + if [ "$REVIEW_READ" = "1" ]; then + echo "Code review status unparsed: REVIEW.md is present but its frontmatter has no parseable status; severity counts unavailable." + fi + ;; + clean|skipped) ;; # nothing to report; block 2 still reconciles an existing ledger + *) + if [ "$REVIEW_COUNTS_OK" = "1" ]; then + echo "Code review: ${REVIEW_TOTAL} findings — ${REVIEW_CRITICAL} critical, ${REVIEW_WARNING} warning, ${REVIEW_INFO} info." + else + # A REVIEW.md written without a `findings:` block has no counts to report, and any count that + # is empty, non-numeric, over-long or inconsistent makes the whole breakdown unavailable + # rather than half-filled. Half-true is worse than withheld. + echo "Code review found issues." + fi + echo "Consider running: /gsd:code-review ${PHASE_NUMBER:-} --fix" + ;; +esac +``` + +**Display that block's stdout verbatim.** It prints the severity breakdown when all four counts are +numeric and mutually consistent, and the countless form otherwise; on a clean, skipped or absent +review it prints nothing and there is nothing to display. Do not re-derive any of it — a number the +shell computed and did not print is gone once the fence closes, which is precisely the defect this +arm exists to close. + +**Record a per-finding disposition.** The counts say how many findings there were, not what +happened to any of them. On the same condition as the message above — REVIEW_STATUS not "clean", +not "skipped" and not empty — write `${DISPOSITION_FILE}`: one row per finding ID, defaulting to +`open`, reconciling `fixed`/`skipped` from REVIEW-FIX.md and preserving any disposition already +recorded, its stated reason included. It is a sibling artifact because `--auto` rewrites +REVIEW.md every iteration and `gsd-code-reviewer` is its single writer. Advisory like the rest of +the step — never blocks: + +## Design notes for the embedded record-builder + +These notes document the `node -e` script in the fence below. They live here rather than +as comments inside that script because the script is passed to `node -e` as a single +command-line argument, and Windows caps a command line at 32,767 characters +(`CreateProcess`); with the commentary inline the argument reached 33,353 characters and +the step failed to launch on Windows with `ENAMETOOLONG`. Each note names the line it +precedes, so the pairing survives the move. + +**Before `(function main() {`** + +EVERYTHING BELOW RUNS INSIDE main() AND LEAVES BY return, NEVER an explicit exit call. The script +prints its one-line verdict and then ends; with an explicit exit directly after console.log, +the exit can pre-empt the write when stdout is a pipe or socket (Node documents those writes +as asynchronous on POSIX), and the caller then sees an exit 0 with NO verdict line. A +hardening against that documented hazard, not a reproduced defect: the 'unchanged' branch was +the only one that exited explicitly, and the empty stdout that first pointed at it turned out to +be a reviewing sandbox's own. A function that returns lets the event loop drain stdout before +the process ends. Same exit status either way. + +**Before `if (fs.existsSync(process.env.DISPOSITION_FILE) && !fs.lstatSync(process.env.DISPOSITION_FILE).isFile()) {`** + +AN EXISTING LEDGER THAT IS NOT A REGULAR FILE IS NOT A LEDGER. Checked FIRST, before any read +or write of that path, and the ordering is the fix rather than a tidy-up: + * writeFileSync FOLLOWS a symlink, so a planted link replaced the contents of whatever it + pointed at -- outside the phase directory, link left intact so nothing looked wrong; + * a FIFO at that path made readFileSync BLOCK FOREVER, which is the one behaviour a gate + documented as advisory and non-blocking must never have; + * and the unchanged-run fast path read the file before the check, so a symlink whose + target already matched slipped through reporting 'unchanged'. +All three driven. lstatSync does not follow the link, which is why it is the right call. +NAMED RESIDUAL, not silently accepted: this is a check-then-write, so a symlink planted +between the lstat and the write still wins. Node exposes no portable O_NOFOLLOW write, and +an attacker who can write into the phase directory mid-run already has what the check would +protect. It narrows a real accident; it is not a security boundary, and the docs do not +claim one. A hard link likewise passes isFile() by construction. + +**Before `let fence = null;`** + +The OPEN fence's marker is remembered, not just the fact of being fenced. A bare toggle +treats every fence marker as interchangeable, so a ~~~ line inside a ` ` ` block CLOSES it +and the block's real close REOPENS one — which silently swaps a fenced example for the +real findings around it. Driven: a review quoting ~~~ inside a fenced example recorded +the EXAMPLE's id and dropped the real finding entirely. Per CommonMark, a fence closes +only on the same character, at least as long as the one that opened it. + +**Before `const SECTION_SEV = [[/^##\s+Critical Issues\s*$/, 'critical'], [/^##\s+Warnings\s*$/, 'warning'], [/^##\s+Info\s*$/, 'info']];`** + +SEVERITY COMES FROM THE SECTION FIRST, the recorded ledger value second, the id prefix +third (the full precedence is at sev(), below the identity check). The section heading is the +reviewer's OWN statement of a finding's severity -- gsd-code-reviewer.md emits findings under +'## Critical Issues' / '## Warnings' / '## Info' -- and this walker already visits every line, +so the signal was in hand and discarded. Deriving from the prefix alone means a reviewer who +mis-numbers a Critical as WR-04 while filing it under '## Critical Issues' gets a ledger row +reading 'warning', which then disagrees with the review it summarizes AND with the frontmatter +count line block 1 prints from findings.critical. The Severity column is the whole basis for +triaging the ledger, so it has to agree with the document it describes. +Matched WHOLE, exactly as the fix-report sections are: a prefix match would let a heading like +'## Critical Issues Verification' re-tier everything under it. + +**Before `const declaredTotal = /^[0-9]+$/.test(process.env.REVIEW_TOTAL || '') ? Number(process.env.REVIEW_TOTAL) : null;`** + +A review that reports nothing still has to reconcile an EXISTING ledger: its decided rows +and its untriaged rows are BOTH carried, marked. Exiting here would freeze a stale ledger +showing findings as open that the review no longer reports. +A fix report with no ledger is also something to record: a converged '--auto' run has neither, +and exiting here recorded nothing for a fully fixed phase. +A review that reports findings NONE of which this parser understood is also something to +record, and it is the case with the least evidence anywhere else. The shortfall is derived +HERE, above the guard, rather than at its old site beside the render: order is final from +the heading walk above and never grows again, so the value is the same either way -- but at +the old site it was computed AFTER this return had already fired, so it could not reach the +one exit that discards it. Partial shortfalls (some findings parsed, some not) always +reported, which is exactly why the total one read as covered. + +**Before `const prior = new Map();`** + +Prior rows: keep the disposition AND its source cell — the source is where a human writes +the reason a finding was deferred, and rewriting it would discard the very thing the +'set deferred by hand, with the reason' instruction asks for. The Source cell is the LAST +column, so it is captured through to the end of the line, less an optional trailing pipe: +a bare | inside it is prose, not a column break. The previous capture admitted a pipe only +when escaped, and the whole-line match then FAILED on a bare one -- a human who wrote +'waiting on team A | team B' as a deferral reason had the row not match at all, the finding +reset to open, and the reason destroyed: a triaged Critical rendered indistinguishable from +one never seen, off an ordinary typo in the one field this ledger asks a human to hand-edit. +The render below escapes a bare pipe on the next write, so the file converges to the escaped +form either way. The trailing pipe is optional so a hand-mangled row loses no decision. + +**Before `const SEV_VOCAB = ['critical', 'warning', 'info'];`** + +SEVERITY, READ BACK. The ledger has always WRITTEN a severity for every row -- in the table's +Severity cell and in the frontmatter's 'severity:' key -- and until this map existed nothing +read either back: the row regex discarded the cell as [^|]*, the frontmatter walk collected +only titles, and a CARRIED row was rebuilt through sev() from the id prefix, because +sectionSev holds only findings the CURRENT review reports. So a WR-04 the reviewer filed under +'## Critical Issues' was recorded 'critical', a human deferred it, and the next run -- the +review no longer reporting it -- silently re-recorded it 'warning'. The one artifact whose +purpose is remembering a finding's severity lost it on the second run, in the unsafe +direction. Driven by executing the shipped script twice (round 11). +The table cell is read first (it is the human-facing surface, and the one the disposition +already comes from); the frontmatter key is the fallback for a hand-mangled cell. Both are +ENUM-validated -- a value outside critical|warning|info is not a severity and is ignored, so +the row falls through to inference rather than carrying garbage (ADR-227, the same rule the +disposition column takes). + +**Before `const m = l.match(/^\|\s*((?:CR|BL|WR|IN)-\d+)\s*\|\s*([^|]*?)\s*\|\s*(open|fixed|skipped|deferred)\s*\|\s*(.*?)\s*\|?\s*$/);`** + +The disposition column is an ENUM, not 'any lowercase token'. ADR-227 requires a trust +boundary to validate semantic SHAPE, not merely type, and to coerce a failure to the +contract's safe default -- and this ledger is a trust boundary by construction, because +the rendered instruction tells a human to hand-edit it. Under the old ([a-z]+) capture a +single transposed character ('opne') was stored as a decision: it is not 'open', so it +beat the default, was excluded from the open: headline count, and was carried forward +forever. One typo and the ledger reported a phase fully triaged. +Note the asymmetry that made this a correctness bug rather than a style point: a typo +OUTSIDE [a-z] ('Deferred') already failed to match, lost the decision and reset the row +to open -- safe. A typo INSIDE [a-z] was unsafe. The parser failed open in the one +direction that matters. A row that does not match now yields no prior entry, so the row +falls back to 'open' -- the safe default, by the same path the capital-D case took. +The Severity cell is CAPTURED, not skipped: it is the value the carry-forward below has to +preserve, and skipping it (the previous [^|]*) is how a carried row lost its tier. + +**Before `if (m) {`** + +Strip the carried marker before storing: it is rendered from the carried flag, so +leaving it on the stored value would re-append it every run — the cell grows without +bound AND the file changes on every run, defeating the unchanged-run check below. +Strip AT MOST ONE trailing marker, unconditionally. Storing the cell verbatim looked +like the way to stop the strip eating human text, and it introduced a worse defect: +once the generated marker is stored it can never leave, so a carried finding that +REAPPEARS in a later review still renders 'not in the current review' -- a ledger that +is now factually wrong about its own contents. The residual ambiguity is irreducible +(a reason ending in exactly that phrase is indistinguishable from the marker) and it +costs nothing real: on a carried row the render puts the phrase straight back, and on a +current row the phrase was self-contradictory to begin with. The unbounded quantifier is +what had to go, not the strip itself. + +**Before `const sameTitle = (a, b) => String(a === undefined ? '' : a).replace(/\s+/g, ' ').trim()`** + +TITLE COMPARISON, and its FALSE-POSITIVE mode, which was previously unacknowledged. +The strict instinct is right -- ids are reused across re-reviews, so a stale REVIEW-FIX.md +must not mark a brand-new CR-01 as already fixed -- but gsd-code-fixer.md writes +'### {finding_id}: {title}' under no contract that the title is copied byte-for-byte from +REVIEW.md. A fixer that REFLOWS a long title produced a spurious stale note, left a +genuinely-fixed row 'open', and told the reader the fix report named a different finding. +Runs of whitespace are collapsed because re-spacing carries no information. The BOUND, stated +because it is easy to over-read this: a title WRAPPED across lines is NOT reconciled. A '###' +heading is one line by definition, so the continuation is a separate paragraph the heading +parser correctly never captures, and collapsing whitespace cannot reach across that boundary. +Not widened -- absorbing whatever follows a heading into the title would swallow arbitrary +prose and make this very check meaningless. Case changes and truncation stay strict too -- +they are the shapes a genuinely different finding actually takes, and widening to them would +trade this false positive for the silent false NEGATIVE the strict match exists to prevent. +Residual, stated: a fixer that re-cases or truncates still produces a spurious note. That is +the safe direction (a visible note, not a silent wrong 'fixed'), and the note's wording below +no longer asserts which of the two it is. + +**Before `if (h.id && sect && !applied.has(h.id)) {`** + +First occurrence wins, so an id listed under BOTH sections is not decided by row order. +And the fix report must name the SAME finding: ids are reused across re-reviews, so a +stale REVIEW-FIX.md would otherwise mark a brand-new CR-01 as already fixed. +A title mismatch is the STALE-report case and must not pass silently: the id is +reused, the finding is not, and a reader who sees the row stay 'open' has no way to +tell that from 'the fix report never mentioned it'. Record it and say so below. + +**Before `const sev = (id) => sectionSev.get(id) || (priorSev.has(id) && sameFinding(id) ? priorSev.get(id) : prefixSev(id));`** + +SEVERITY PRECEDENCE: the current review's SECTION (the reviewer's own statement, this run), +then the severity this ledger RECORDED (an earlier reviewer's statement, persisted), then the +id PREFIX (an inference). A recorded value is inherited only while the id still names the +SAME finding -- the identity rule the disposition already obeys -- so a reused id starts from +its own review's section or its prefix, never from the finding it replaced. A carried row is +absent from the current review, so sameFinding() is true for it by construction and its +recorded severity is what it keeps. Defined here, below the identity check, because it +depends on it. + +**Before `const carriedIds = [];`** + +A prior finding the current review no longer reports is CARRIED, never dropped -- and that +now holds for UNTRIAGED rows too, which is the correction. Carrying only decided rows meant +an untriaged row for a dropped or renumbered finding disappeared without trace, and combined +with the reconciliation gap that left EVERY row untriaged, a re-review silently deleted the +whole ledger. The --auto loop rewrites REVIEW.md on every iteration, so it does not retain it +either: run 1 records CR-01 open, the re-review renumbers it to CR-02, and run 2's ledger +contains neither. That is #3829's complaint verbatim -- 'no trace of what happened to them' -- +reproduced by the artifact built to prevent it, and 'nothing was decided about it' is exactly +the state #3829 says must leave a trace. +The carried marker is what keeps this honest rather than merely additive: the row does not +claim the finding is live, it records that it was seen and never triaged. Stated cost, since +it is real: a RENUMBERED finding appears twice until someone triages the old row, and a +carried untriaged row persists across runs until decided. Both are bounded by the phase's own +findings, both are legible from the marker, and both are strictly better than a silent delete. +Prior rows UNION ids a fix report decided that the review no longer reports: a decision the +ledger cannot render is a decision lost. Precedence matches row() -- applied beats recorded. + +**Before `const reusedNote = reused.length ? ' (' + reused.length + ' recorded decision(s) DROPPED -- the id now names a different finding, so the decision no longer has a row: ' + reused.join(', ') + ')' : '';`** + +Surfaced, not thrown: the gate is advisory. But a fix report naming a finding whose title +no longer matches is the one case where 'open' understates what is known, so it is stated. +The wording no longer ASSERTS a stale report. Both causes reach here -- a genuinely different +finding under a reused id, and a fixer that re-titled the same one -- and the step cannot tell +them apart, so it reports the observation rather than a conclusion it has not earned. +On the console too, for a reader who never opens the ledger. + +**Before `if (rows.length === 0 && !unparsed && !fs.existsSync(process.env.DISPOSITION_FILE)) return;`** + +RECONCILE THE TWO PARSERS. The counts come from REVIEW.md's frontmatter; the rows come from +heading matches against a CLOSED CR|BL|WR|IN alternation. A finding the heading parser cannot +match -- a fifth prefix, a missing ': ' separator, a '#### ' heading -- contributed no row, no +note and no diagnostic, and the ledger then declared 'open: 3 of 3' over a set strictly +smaller than the console line reported one paragraph earlier. Two findings recorded nowhere, +and neither artifact said so. +The earlier argument for the closed alternation -- that an unlisted prefix produces no row +rather than a MIS-CLASSIFIED one -- is the wrong trade under this repo's own fail-safe rule: +a dropped finding is demoted below every finding that parsed, and an unparseable finding is +precisely the one a human most needs to see. Surfaced, not thrown, exactly as the stale +fix-report case above is: the gate stays advisory and states the shortfall. +The !unparsed conjunct here is the SECOND of the two exits that discarded the shortfall, and +it is not redundant with the one above: that guard keys on order and stands down when a fix +report exists, so a run with a fix report and no parseable finding reaches THIS line with +rows.length 0. Both exits now decline to fire while a shortfall is outstanding, and the +result is a zero-row ledger carrying an unparsed key -- an honest record that the review +declared findings and none of them were understood, which is strictly better than the file +not existing. A genuinely clean review is untouched either way: a declared total of 0 is not +greater than order.length, so unparsed is 0 and both returns still fire. + +**Before `const escapePipes = (t) => t.replace(/\\.|\|/g, (m) => (m === '|' ? '\\|' : m));`** + +A bare | in a Source cell is escaped on render so the table stays a table. Scanned as PAIRS, +not by the preceding character: an escaped pair (backslash + anything) is kept verbatim and only +a pipe outside one is escaped. The previous form, /(^|[^\\])\|/g, CONSUMED the character before +the pipe, so adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third run +to converge) and an escaped backslash before a pipe (A\\|B) hid the pipe behind the wrong +parity and left it bare. Found by the round-3 adversarial pass, not by the property -- whose +generator then emitted at most one bare pipe, the one case the old form got right; it now +reaches adjacent pipes and both backslash parities, against an independent parity oracle. + +**Before `fs.writeFileSync(process.env.DISPOSITION_FILE, render(new Date().toISOString()));`** + +READ-MODIFY-WRITE, NO LOCK. The ledger is rendered whole from a read taken above, and nothing +serializes two writers: this step has two dispatchers (execute-phase's gate and +code-review-fix's record_disposition) plus a human the legend invites to hand-edit, so a +lost update is a real window, not a theoretical one. Same shape as #3780 (WINDOWS.md +append under parallel executors), which #4681 closed with a cross-process lock in +src/broken-windows.cts. NOT taken here: this is a shell-embedded script with no build +dependency on the compiled tree, and adopting the lock module is its own change. Residual, +stated in docs/features/code-review-pipeline.md; not reproduced as a lost update. + +```bash +# Each fenced block runs in a FRESH shell, so block 1's PADDED/REVIEW_FILE/DISPOSITION_FILE are NOT +# live here — re-derive them from the two inputs this step consumes (`PHASE_DIR`, `PHASE_NUMBER`). +# Inheriting them is not merely stale, it is EMPTY, and the failure is silent rather than loud: +# the embedded script throws on reading the empty review path, the trailing `|| echo` swallows it +# as a non-blocking skip, and no ledger is written at all. The shim preamble below is re-emitted +# for the same reason, and these three belong beside it. +# PADDED must survive a DOTTED phase number, of ANY segment count. This step is dispatched from +# exactly TWO places: `execute-phase.md` (`code_review_gate`) and `code-review-fix.md` +# (`record_disposition`). Only the second validates anything -- `code-review-fix.md`'s PADDED_PHASE validator anchors +# `^[0-9]+[A-Z]?(\.[0-9]+)*$`, an unbounded `*` widened by #4568 and a letter axis widened by #4744, so it accepts `03.1`, `23.1.2` AND `12A`. +# `execute-phase.md` applies NO shape gate at all, so this fence is not mirroring an upstream +# guarantee; it IS the guarantee. (`code-review.md`'s own PADDED_PHASE validator is identical but never +# dispatches this step. It was cited here as a caller for several rounds and is not one.) And +# `printf "%02d"` cannot format one: bash prints `invalid number` and exits 1, which under +# `set -euo pipefail` aborts this step on its FIRST line -- the loudest possible failure from +# the gate that promises never to block, and it takes the whole phase's review reporting with +# it. Pad the integer part only and carry the sub-number verbatim, so 3.1 -> 03.1, 23.1.2 -> +# 23.1.2 and 3 -> 03. The segment count is deliberately NOT bounded here: the canonical +# grammar in src/phase-id.cts (`PHASE_NUMBER_TOKEN_SOURCE`, #2128) is unbounded in segments, +# and #4568 widened the one dispatcher that validates to match it on that axis, so a guard +# narrower than that dispatcher means no ledger for a phase id the dispatcher already accepted. +# On failure NO path is built and the fence refuses by name: advisory means advisory, and it +# also means never probing a path assembled out of a value we just rejected. +# VALIDATE, THEN FORMAT -- never format and fall back on failure. `printf "%02d" abc` writes +# `00` to stdout BEFORE it fails, so a `$(printf ... || printf %s ...)` fallback CONCATENATES +# the two and yields `00abc`; `08` fails the same way as invalid octal, giving `0008.1` for a +# legitimate `08.1`. Both were driven. `${PHASE_NUMBER:-}` because an UNSET input must not trip +# `set -u` in a step that promises not to abort. Those printf failures are why this block VALIDATES +# instead of formatting; the pad itself performs NO arithmetic since round 14 (see the +# `case "${#_dig}"` line below), so it has no octal hazard to guard and needs no `10#`. `10#` +# survives in this step only where it still belongs -- on the severity COUNTS, which really are +# numbers being added. +# VALIDATE THE WHOLE VALUE, then format -- and on failure build NO path at all. +# Carrying an unusable value verbatim was the first draft and it was worse than the bug it +# replaced: PHASE_NUMBER is interpolated into a file path, so `../../etc/passwd` produced +# `${PHASE_DIR}/../../etc/passwd-REVIEW.md`, where the old `printf "%02d"` had at least +# mangled it to `00`. `code-review-fix.md`'s PADDED_PHASE validator already checks `^[0-9]+[A-Z]?(\.[0-9]+)*$` against its +# own PADDED_PHASE -- the padded form, not the raw PHASE_NUMBER this step is handed -- while +# `execute-phase.md` validates nothing at all; this step has two call sites and validates for +# itself rather than trusting either. Anything else yields an EMPTY PADDED and the blocks +# below refuse to build a path from it. +# PHASE_DIR is checked for NON-EMPTINESS ONLY. Both inputs come from the caller's init query, so +# neither is raw user input; only PHASE_NUMBER has a SHAPE (`^[0-9]+[A-Z]?(\.[0-9]+)*$`) to check +# against. A filesystem path admits `..` and symlinked parents alike, so a shape +# check here rejects working setups and proves nothing. Residual: PHASE_DIR may itself be a symlink +# and the ledger is written through it -- left alone, and not a security boundary. +_pd="${PHASE_DIR:-}" +_pn="${PHASE_NUMBER:-}" +_ok=1 +[ -n "$_pd" ] || _ok=0 +case "$_pn" in + ''|*[!0-9.A-Z]*) _ok=0 ;; # empty, or any character outside [0-9.A-Z] -- this is the traversal fence + .*|*.) _ok=0 ;; # leading or trailing dot + *..*) _ok=0 ;; # EMPTY SEGMENT. The three arms plus the letter-axis block below + # accept exactly digits[LETTER](.digits)* -- byte-congruent with the + # callers' ^[0-9]+[A-Z]?(\.[0-9]+)*$ -- rather than merely wider than + # the retired `*.*.*` arity bound, which masked `1..2` by accident. +esac +# THE LETTER AXIS. The canonical grammar (src/phase-id.cts) is digits, an OPTIONAL single uppercase +# letter, then dotted digit segments -- `12A`, `3A`, `23A.1.2`. #4744 (#4660) widened the six +# shell/markdown mirrors to it after this branch was cut, and its lint ratchet then flagged this +# step as the one letterless mirror left. The character class above admits the letter; these +# arms pin WHERE it may sit -- only as the last character of the integer part, at most once -- +# so `23a`, `A23`, `2A3`, `23AB` and `23.1A` are all refused. +if [ "$_ok" = "1" ]; then + _int="${_pn%%.*}" + case "$_pn" in *.*) _sub=".${_pn#*.}" ;; *) _sub="" ;; esac + _let="${_int##*[0-9]}" # what trails the last digit: '' or the letter + _dig="${_int%"$_let"}" + case "$_dig" in ''|*[!0-9]*) _ok=0 ;; esac # the integer part must be digits first + case "$_let" in ''|[A-Z]) ;; *) _ok=0 ;; esac # at most ONE letter, uppercase + case "$_sub" in *[!0-9.]*) _ok=0 ;; esac # no letter in any later segment +fi +# LENGTH-BOUND EACH COMPONENT SEPARATELY -- and the REASON changed at round 14, so read this rather +# than inherit it. It used to be integer overflow: the pad ran `$((10#$_int))`, bash integers wrap at +# 2^64, and a 54-digit value yielded -7908320945662590977 SILENTLY as the padded phase. The pad is a +# string pad now and converts nothing, so that overflow is unreachable and its rationale is dead. +# WHAT THE BOUND STILL DOES, stated narrowly because the obvious wider claim is FALSE: it bounds each +# SEGMENT, and NOTHING here bounds the COMPOSITE. Every segment is joined into ONE filename +# component and depth is unbounded, so a per-segment bound does not enforce a filename limit -- +# driven at round 14: thirty 8-digit segments yield a 278-character PADDED and a 294-character name +# against a NAME_MAX of 255. That is a real residual of this validator, it predates the pad change, +# and it is named here rather than papered over with a filesystem rationale the bound does not +# deliver. The bound belongs on the INTEGER PART: applied to the whole value it rejected +# `12345678.1`, whose integer part is a legal 8 digits, while accepting `1.123456` -- an accidental +# bound on the composite that was both too strict and too loose. Every later segment is bounded too, +# on the same narrow reading. +# THE LOOP IS THE POINT, and it is what makes the heading above TRUE. The earlier form bounded +# `${_pn#*.}` -- the WHOLE tail after the first dot -- which is one component only while the id +# has at most two. Once N-segment ids are accepted (see the shape arms), that form rejects +# `1.1234567.1`, whose every component is a legal 7-or-fewer digits, purely because the tail +# measures 9 characters. That is the composite bound this comment already called "too strict", +# surviving one level up. Walk the segments instead, so the rule is per-component in fact and +# not only in the heading. The loop terminates on any string the shape arms admit: each pass +# strips a leading `.`, and the no-dot pass clears $_rest. +if [ "$_ok" = "1" ]; then + _rest="$_pn" + while [ -n "$_rest" ]; do + case "$_rest" in + *.*) _seg="${_rest%%.*}"; _rest="${_rest#*.}" ;; + *) _seg="$_rest"; _rest="" ;; + esac + # The bound is on the DIGITS: a letter suffix is one character the digit bound has no + # stake in, so `12345678A` is within it exactly as `12345678` is. + case "${_seg%[A-Z]}" in ?????????*) _ok=0 ;; esac + done +fi +if [ "$_ok" = "1" ]; then + # padStart(2,'0'), EXACTLY, and as a STRING -- the canonical normalizer left-pads the digit run to a + # MINIMUM of two and otherwise preserves it, so arithmetic is the wrong tool. `printf "%02d" + # "$((10#$_dig))"` agreed on `8`/`08`/`09` and silently DISAGREED on every longer leading-zero run: + # `008` -> `08`, `0008A` -> `08A`, resolving a REVIEW.md path init never writes. Driven at round 14 + # by the property that asserts agreement with `normalizePhaseName` over generated ids; the 13-shape + # matrix that preceded it sampled no run longer than two and could not see it. Dropping the + # arithmetic also RETIRES the octal hazard `10#` existed to work around, rather than guarding it. + # The letter and any dot segments ride along verbatim, as the canonical padder does: `3A` -> `03A`. + case "${#_dig}" in 1) PADDED="0${_dig}${_let}${_sub}" ;; *) PADDED="${_dig}${_let}${_sub}" ;; esac +else + PADDED="" +fi +# REFUSE BEFORE BUILDING ANY PATH. An unusable input yields an empty PADDED above, and the +# earlier placement -- after the assignments -- meant a rejected value still had +# `${_pd}/-REVIEW.md` assembled and stat'ed before the refusal fired. Nothing is constructed +# from a value we have already rejected. +if [ -z "$PADDED" ]; then + echo "Code review disposition skipped (unusable phase number or directory: '${PHASE_NUMBER:-}')" + return 0 2>/dev/null || exit 0 +fi +REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md" +DISPOSITION_FILE="${_pd}/${PADDED}-REVIEW-DISPOSITION.md" +# The condition stated above this block is re-derived HERE rather than left to the reader. Block 1 +# computes REVIEW_STATUS and emits nothing, and its shell is gone, so nothing downstream can act on +# it: a prose-only gate on a value no later block can see is not a gate. Without this, a clean +# re-review rewrites an existing ledger it was never meant to touch. +REVIEW_STATUS="" +REVIEW_TOTAL="" +REVIEW_READ=0 # block 1's distinction, re-derived here: read-but-unparseable is not absent +if [ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]; then + REVIEW_READ=1 + _FM=$(tr -d '\r' < "$REVIEW_FILE" 2>/dev/null | LC_ALL=C awk 'NR==1{if($0!="---") exit; next} /^---$/{closed=1; exit} {buf = buf $0 "\n"} END{if (closed) printf "%s", buf}' || true) + REVIEW_STATUS=$(echo "$_FM" | LC_ALL=C grep -m1 "^status:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) + # The frontmatter total is carried into the script so the two parsers in this step can be + # RECONCILED. The counts come from the frontmatter; the rows come from `### :` heading + # matches against a closed CR|BL|WR|IN alternation. They are two independent numbers produced + # one paragraph apart, and nothing compared them: a finding the heading parser cannot match + # contributed no row, no note and no diagnostic, and the ledger then asserted `open: 3 of 3` + # over a set strictly smaller than the console line had just reported. Anchored inside the + # `findings:` mapping — see the anchoring note in block 1 — and digit-only, because a + # non-numeric total is not a number to reconcile against. + _FINDINGS_FM=$(echo "$_FM" | LC_ALL=C awk '/^findings:[[:space:]]*$/{f=1; next} f&&/^[^[:space:]]/{exit} f' || true) + # ONE PARSER FOR EVERY COUNT THIS BLOCK READS, and both halves of it are load-bearing. + # `-f2-` keeps everything AFTER the first colon: `-f2` alone takes only the SECOND FIELD, so the + # malformed `critical: 1: junk` arrives as the perfectly numeric `1`. And the ends are trimmed + # rather than `tr -d ' '`-ed, which would delete INTERNAL spaces and turn `1 0` into `10`. + # Both quirks are long-standing in the sibling reads and both were INERT here until this block + # began reconciling; each one repairs a malformed scalar into a number that then decides whether a + # shortfall is reported. An adversarial pass drove both: `critical: 1: junk` wrongly suppressed a + # real `unparsed: 2`, and a LENIENT total beside a STRICT severity was worse still -- `critical: 5 0` + # with `total: 1 0` repaired only the total, rejected the severity, skipped the contradiction check + # and INVENTED `unparsed: 7`. A field is either trustworthy or it is not; parsing one leniently and + # its sibling strictly is the shape that fabricates. + REVIEW_TOTAL=$(echo "$_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*total:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) + case "$REVIEW_TOTAL" in ''|*[!0-9]*) REVIEW_TOTAL="" ;; ?????????*) REVIEW_TOTAL="" ;; esac + # ONE FIELD, ONE TRUST MODEL, ACROSS BOTH FENCES. Block 1 withholds the whole breakdown unless the + # four counts are numeric AND `critical + warning + info == total`; this fence used to bound `total` + # for digits and length only and then hand it to the `unparsed:` reconciliation, so a REVIEW.md whose + # `findings:` block is internally inconsistent (`total: 10` beside `critical: 1, warning: 1, info: 1`) + # made block 1 print the countless form -- breakdown suppressed as untrustworthy -- while this block + # still computed an `unparsed:` shortfall from that same untrusted number. It fails in the SAFE + # direction (over-reports a possible gap rather than hiding one), which is why it is not a blocker; + # it is still two trust models for one field, one fence apart, and the weaker one is downstream. + # `blocker:` is the documented tier-equivalent of `critical:` (gsd-code-reviewer.md 'Label + # equivalence') -- the same alternation block 1 reads, because a mirror that drops it would diverge + # on exactly the reviews that use it. + # TRIM THE ENDS, NEVER `tr -d ' '`, AND THE DIFFERENCE DECIDES A SUPPRESSION. `tr -d` deletes + # INTERNAL spaces too, so a malformed `critical: 1 0` would arrive as the perfectly numeric `10`. + # Every read in this step now takes the end-trim instead -- the sibling reads were moved off `tr -d` + # in the same round, so this is no longer a divergence between blocks -- and the reason it matters + # HERE is that a value which LOOKS numeric can satisfy the sum test and SUPPRESS a real + # `unparsed:` shortfall. Suppression is the new + # behaviour, so the admission test for it is strict -- an internal space survives the trim, fails the + # digit `case` below, and the shortfall is reported. Fail-safe in the only direction that matters: + # when the frontmatter is malformed we decline to suppress, rather than trusting a repaired number. + _c_crit=$(echo "$_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*(critical|blocker):" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) + _c_warn=$(echo "$_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*warning:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) + _c_info=$(echo "$_FINDINGS_FM" | LC_ALL=C grep -E -m1 "^[[:space:]]*info:" | cut -d: -f2- | LC_ALL=C sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//' || true) + # THE CHECK IS NARROWER THAN BLOCK 1'S, DELIBERATELY, AND THE DIFFERENCE IS NOT AN OVERSIGHT. + # Block 1 withholds on `REVIEW_COUNTS_OK`, which demands ALL FOUR counts be numeric -- because it + # DISPLAYS all four, and `6 findings -- critical` is the half-filled line that rule exists to + # prevent. This fence displays none of them: it uses `total` alone, to reconcile against the number + # of headings the row parser matched. So the all-four rule does not port. Applied verbatim here it + # would blank `total` on a REVIEW.md carrying `total: 5` and no severity keys -- a review whose total + # is perfectly usable -- and SILENTLY DROP an `unparsed:` shortfall this step reports correctly today. + # That trades a safe-direction over-report for a silent under-report, which is the wrong way round + # and is the exact failure class the `unparsed:` key was added to close. + # What DOES port is the CONTRADICTION: when the three severities are all present and numeric and do + # not sum to `total`, the frontmatter disagrees with itself and `total` is not a number to reconcile + # against. Block 1 already suppresses its breakdown on that input; this fence now declines to compute + # a shortfall from it. Absent counts are not a contradiction -- there is nothing to disagree. + # AN ABSENT SEVERITY STILL BOUNDS THE SUM FROM BELOW, and that is enough to prove a contradiction + # in one direction. Counts are non-negative, so a missing one can only ADD: if the severities that + # ARE present and numeric already sum to MORE than `total`, the block disagrees with itself whatever + # the missing value is. Requiring all three before comparing missed that -- driven by an adversarial + # pass: `critical: 4`, `warning: 4`, no `info:`, `total: 5` reconciled against a total the present + # counts had already refuted. So the comparison is two-armed: EQUALITY when all three are known, + # and a LOWER BOUND when they are not. `_p_sum` accumulates only the present-and-numeric ones. + _sum_ok=1; _p_sum=0 + for _c in "$_c_crit" "$_c_warn" "$_c_info"; do + # Present AND numeric AND within the same length bound the total carries -- `10#` below needs + # digits, and bash integers wrap at 2^64. + case "$_c" in + ''|*[!0-9]*) _sum_ok=0 ;; + ?????????*) _sum_ok=0 ;; + *) _p_sum=$(( _p_sum + 10#$_c )) ;; + esac + done + # `10#` on every operand, for block 1's reason: bash infers the base from a leading zero, so + # `critical: 08` makes $(( )) fail with "value too great for base" and, under `set -e`, takes the + # whole advisory step down -- strictly worse than the stale count this check exists to prevent. + if [ -n "$REVIEW_TOTAL" ]; then + _t=$(( 10#$REVIEW_TOTAL )) + if [ "$_sum_ok" = "1" ]; then + # All three known: the sum must match exactly. + if [ "$_p_sum" -ne "$_t" ]; then REVIEW_TOTAL=""; fi + elif [ "$_p_sum" -gt "$_t" ]; then + # Not all known: only an OVERSHOOT is provable. An undershoot is the absent count's job. + REVIEW_TOTAL="" + fi + fi +fi +# Skip a clean/skipped/absent review ONLY when there is nothing to reconcile AT ALL. An EXISTING +# ledger is still brought up to date -- freezing it would leave findings showing open that the +# review no longer reports, and an unconditional skip would make the reconciliation path +# unreachable on exactly the run that needs it. +# A FIX REPORT IS THE SECOND REASON TO PROCEED: a direct `/gsd:code-review N --auto` writes no gate +# ledger and a converged loop leaves `status: clean`, so a fully fixed phase recorded nothing. +_fix_any=0 +[ -f "${_pd}/${PADDED}-REVIEW-FIX.md" ] && _fix_any=1 +# Backups count too -- a converged loop's earlier iterations live only there. An unmatched glob +# expands to the literal pattern, which `-f` rejects. +# $(printf '%s' "$PADDED") per lint-workflow-shellcheck's #4109 remedy: a bare $VAR in a `for x in` +# splits differently under bash and zsh. +for _f in "${_pd}/$(printf '%s' "$PADDED")-REVIEW-FIX.iter"*.md; do [ -f "$_f" ] && _fix_any=1; done +# The word for an empty status names WHICH empty it is, for the same reason block 1 does: a +# review that was read and could not be parsed is 'unparsed', an absent one is 'none'. +_st="${REVIEW_STATUS:-none}"; [ -z "$REVIEW_STATUS" ] && [ "$REVIEW_READ" = "1" ] && _st="unparsed" +case "$REVIEW_STATUS" in + ''|clean|skipped) + if [ ! -f "$DISPOSITION_FILE" ] && [ "$_fix_any" = "0" ]; then + echo "Code review disposition skipped (status: ${_st})" + return 0 2>/dev/null || exit 0 + fi + echo "Code review status ${_st}; reconciling the fix report and any existing disposition ledger." + ;; +esac +_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; _gsd_id_ok() { case "$("$1" runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') return 0;; *) return 1;; esac; }; _gsd_homes() { _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif _gsd_homes; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; [ -n "$_G" ] && _gsd_id_ok "$_G"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and no identity-proving gsd_run is on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; _gsd_id_ok gsd_run && GSD_IDENTITY_STATUS=ok; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi +# Built before the command for READABILITY, not as a fix. ShellCheck's SC2097/SC2098 here is a FALSE +# POSITIVE: prefix assignments take effect left to right (driven, bash and dash). +FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md" +REVIEW_FILE="${REVIEW_FILE}" DISPOSITION_FILE="${DISPOSITION_FILE}" PADDED="${PADDED}" \ +REVIEW_TOTAL="${REVIEW_TOTAL}" \ +FIX_REPORT_FILE="${FIX_REPORT_FILE}" node -e " + (function main() { + const fs = require('fs'), path = require('path'); + const norm = (s) => s.replace(/\r\n/g, '\n'); + if (fs.existsSync(process.env.DISPOSITION_FILE) && !fs.lstatSync(process.env.DISPOSITION_FILE).isFile()) { + console.log('Code review disposition skipped: ' + process.env.DISPOSITION_FILE + ' exists and is not a regular file; refusing to read or write through it.'); + return; + } + // Captures the id AND the title: the title is what tells a stale fix report apart from a + // current one, because finding ids are reused across re-reviews. + const ID_RE = /^###\s+((?:CR|BL|WR|IN)-\d+)\s*:\s*(.*)\$/; + // BL- is Critical-tier-equivalent to CR- (gsd-code-reviewer.md 'Label equivalence'). + const sectionSev = new Map(); + // PREFIX severity -- the LAST resort, an inference from the id alone. Used only when neither + // the current review's section nor a severity this ledger already RECORDED is available; the + // precedence and the reason for it are stated at sev(), defined below once identity is known. + const prefixSev = (id) => ({ CR: 'critical', BL: 'critical', WR: 'warning' }[id.split('-')[0]] || 'info'); + const headings = (text) => { + // Fenced blocks are skipped: review and fix bodies quote example findings, and a heading + // inside a fence is an illustration, not a finding. + const out = []; + let fence = null; + for (const l of norm(text).split('\n')) { + const f = l.match(/^ {0,3}(\`{3,}|~{3,})/); // >3 spaces is an indented block, not a fence + if (f) { + const ch = f[1][0], len = f[1].length; + if (!fence) { fence = { ch: ch, len: len }; out.push({ fence: true }); continue; } + // A CLOSER carries nothing but whitespace after the marker; an info string makes it + // an opener's shape, never a close. + if (ch === fence.ch && len >= fence.len && /^\s*\$/.test(l.slice(l.indexOf(f[1]) + f[1].length))) { + fence = null; out.push({ fence: true }); continue; + } + out.push({ skip: true, line: l }); continue; // a foreign marker inside a fence is content + } + if (fence) { out.push({ skip: true, line: l }); continue; } + const m = l.match(ID_RE); + out.push(m ? { id: m[1], title: m[2].trim(), line: l } : { line: l }); + } + return out; + }; + const order = [], title = new Map(); + // An ABSENT review still has a ledger to reconcile: the step's own guard proceeds when + // one exists, and throwing here would send that run to the trailing non-blocking fallback + // with the + // ledger untouched -- the freeze the reconciliation path exists to prevent. + const reviewText = fs.existsSync(process.env.REVIEW_FILE) ? fs.readFileSync(process.env.REVIEW_FILE, 'utf-8') : ''; + const SECTION_SEV = [[/^##\s+Critical Issues\s*\$/, 'critical'], [/^##\s+Warnings\s*\$/, 'warning'], [/^##\s+Info\s*\$/, 'info']]; + let curSection = null; + for (const h of headings(reviewText)) { + if (h.fence || h.skip) continue; + if (h.line !== undefined && /^##\s+/.test(h.line)) { + const hit = SECTION_SEV.find(([re]) => re.test(h.line)); + curSection = hit ? hit[1] : null; // an unrecognized ## section falls back to the prefix + } + if (h.id && order.indexOf(h.id) === -1) { + order.push(h.id); title.set(h.id, h.title); + if (curSection) sectionSev.set(h.id, curSection); + } + } + // THE FIX REPORTS THIS RUN MAY RECONCILE AGAINST. --auto overwrites REVIEW-FIX.md each iteration + // and the re-review drops what was fixed, so an iteration-1 fix is in NEITHER final artifact. + // Read the backups too, newest first. + const FIX_FINAL = process.env.FIX_REPORT_FILE; + const fixStem = path.basename(FIX_FINAL).slice(0, -3); // '-REVIEW-FIX' + const iterMarker = fixStem + '.iter'; + // String ops, not a built RegExp: every backslash is one more thing bash rewrites first. + // ONE expression, no 'return': ShellCheck lints this fence as shell and would mark the rest + // unreachable (SC2317) against a ratchet baseline. + const iterDigits = (n) => (n.indexOf(iterMarker) === 0 && n.slice(-3) === '.md') ? n.slice(iterMarker.length, -3) : ''; + const iterOf = (n) => /^[0-9]+\$/.test(iterDigits(n)) ? Number(iterDigits(n)) : null; + const fixReports = []; + if (fs.existsSync(FIX_FINAL)) fixReports.push(FIX_FINAL); + let iterFiles = []; + // Guarded: the phase directory is not guaranteed readable, and this step never aborts. + try { iterFiles = fs.readdirSync(path.dirname(FIX_FINAL)).map((n) => [iterOf(n), n]).filter((e) => e[0] !== null); } catch (e) { iterFiles = []; } + iterFiles.sort((a, b) => b[0] - a[0]); + for (const e of iterFiles) fixReports.push(path.join(path.dirname(FIX_FINAL), e[1])); + const declaredTotal = /^[0-9]+\$/.test(process.env.REVIEW_TOTAL || '') ? Number(process.env.REVIEW_TOTAL) : null; + // Against order.length -- the CURRENT review's findings -- never rows.length, which also counts + // rows carried from earlier reviews and would understate the shortfall or invent one. + const unparsed = declaredTotal !== null && declaredTotal > order.length ? declaredTotal - order.length : 0; + const unparsedNote = unparsed ? ' (' + unparsed + ' finding(s) recorded NOWHERE: the review reports ' + declaredTotal + ', but only ' + order.length + ' matched the expected heading shape \`### -NN: \`)' : ''; + if (order.length === 0 && !unparsed && !fs.existsSync(process.env.DISPOSITION_FILE) && fixReports.length === 0) return; + const prior = new Map(); + // TITLES, IN THE FRONTMATTER. Ids are reused across re-reviews (--auto renumbers), so an id alone + // does not identify a finding: driven, a prior 'CR-01 fixed' rendered a brand-new CR-01 'fixed'. + // Not a fifth table column -- the Source cell is already the hand-edited, pipe-escaping one. + const priorTitle = new Map(); + const SEV_VOCAB = ['critical', 'warning', 'info']; + const priorSev = new Map(); + // Ids whose decision could not be carried: the id now names a DIFFERENT finding. REPORTED, not + // re-homed -- rows key on the id, and two under one id is an ambiguity. The note does NOT claim the + // old row is in git: committing is gated on commit_docs. See docs/features/code-review-pipeline.md. + const reused = []; + var _fmId = null, _fmSec = null; + // Set when the ledger declares JSON scalars; without it they are bare. Load-bearing: a legacy title + // that merely LOOKED like JSON was parsed, lost its quotes, and flipped to open. + var _fmJson = false; + if (fs.existsSync(process.env.DISPOSITION_FILE)) { + for (const l of norm(fs.readFileSync(process.env.DISPOSITION_FILE, 'utf-8')).split('\n')) { + const m = l.match(/^\|\s*((?:CR|BL|WR|IN)-\d+)\s*\|\s*([^|]*?)\s*\|\s*(open|fixed|skipped|deferred)\s*\|\s*(.*?)\s*\|?\s*\$/); + if (m) { + prior.set(m[1], { d: m[3], src: m[4].replace(/\s*\(not in the current review\)\s*\$/, '') }); + // The table wins over the frontmatter (set unconditionally here, only-if-absent below), + // whichever order the two appear in the file. + if (SEV_VOCAB.indexOf(m[2]) !== -1) priorSev.set(m[1], m[2]); + } + // Frontmatter is walked in the same pass, as a SECTIONED list rather than by one line shape. + if (/^titles: json\s*\$/.test(l)) { _fmJson = true; continue; } + var msec = l.match(/^(findings):\s*\$/); + if (msec) { _fmSec = msec[1]; _fmId = null; continue; } + var mi = l.match(/^ - id: ((?:CR|BL|WR|IN)-\d+)\s*\$/); + if (mi && _fmSec) { _fmId = mi[1]; continue; } + // The frontmatter's own copy of the severity -- the fallback when the table cell is unusable. + var msv = l.match(/^ severity: (critical|warning|info)\s*\$/); + if (msv && _fmId && _fmSec === 'findings') { if (!priorSev.has(_fmId)) priorSev.set(_fmId, msv[1]); continue; } + var mkv = l.match(/^ title: (.*)\$/); + if (mkv && _fmId && _fmSec === 'findings') { + var _v = mkv[1]; + if (_fmJson) { try { _v = JSON.parse(_v); } catch (e) { /* keep the raw scalar */ } } + priorTitle.set(_fmId, _v); + continue; + } + } + } + const sameTitle = (a, b) => String(a === undefined ? '' : a).replace(/\s+/g, ' ').trim() + === String(b === undefined ? '' : b).replace(/\s+/g, ' ').trim(); + // Section headings are matched WHOLE: a prefix match would let '## Fixed Issues Verification' + // classify every finding under it as fixed. + const applied = new Map(), staleFix = []; + for (const fixPath of fixReports) { + let sect = null; + for (const h of headings(fs.readFileSync(fixPath, 'utf-8'))) { + if (h.fence || h.skip) continue; + if (/^##\s+Fixed Issues\s*\$/.test(h.line)) { sect = 'fixed'; continue; } + if (/^##\s+Skipped Issues\s*\$/.test(h.line)) { sect = 'skipped'; continue; } + if (/^##\s+/.test(h.line)) { sect = null; continue; } + if (h.id && sect && !applied.has(h.id)) { + // THREE ARMS. An id the review does not report has no title to disagree with -- not the + // stale-report case, but what a finding looks like once acted on; the old form dropped it + // silently. Reuse stays closed below. The record carries the originating report and title. + var _acted = { d: sect, src: path.basename(fixPath), t: h.title }; + if (!title.has(h.id)) applied.set(h.id, _acted); + else if (sameTitle(title.get(h.id), h.title)) applied.set(h.id, _acted); + else if (staleFix.indexOf(h.id) === -1) staleFix.push(h.id); + } + } + } + // Precedence: an applied outcome is evidence of an action on code and wins; a recorded + // non-'open' decision wins over the default. 'open' never overwrites a decision. + // Inherited only while the id names the SAME finding. An ABSENT prior title inherits: a + // pre-titles ledger has none, and refusing would reset every decision in it. + const sameFinding = (id) => !priorTitle.has(id) || !title.has(id) || sameTitle(priorTitle.get(id), title.get(id)); + const sev = (id) => sectionSev.get(id) || (priorSev.has(id) && sameFinding(id) ? priorSev.get(id) : prefixSev(id)); + const row = (id) => { + if (applied.has(id)) { const a = applied.get(id); return { id, sev: sev(id), d: a.d, src: a.src, t: title.has(id) ? title.get(id) : a.t }; } + const was = prior.get(id); + if (was && was.d !== 'open' && sameFinding(id)) return { id, sev: sev(id), d: was.d, src: was.src || 'recorded', t: title.get(id) }; + // Reused id: the NEW finding is untriaged and renders 'open'; the prior decision loses its row, + // and that is REPORTED. + if (was && was.d !== 'open' && reused.indexOf(id + '=' + was.d) === -1) reused.push(id + '=' + was.d); + return { id, sev: sev(id), d: 'open', src: '-', t: title.get(id) }; + }; + const rows = order.map(row); + const carriedIds = []; + for (const id of prior.keys()) if (order.indexOf(id) === -1 && carriedIds.indexOf(id) === -1) carriedIds.push(id); + for (const id of applied.keys()) if (order.indexOf(id) === -1 && carriedIds.indexOf(id) === -1) carriedIds.push(id); + for (const id of carriedIds) { + const act = applied.get(id), was = prior.get(id); + const d = act ? act.d : (was ? was.d : 'open'); + const src = act ? act.src : (was && was.src) || (d === 'open' ? '-' : 'recorded'); + // Title precedence: the report that DECIDED it, then the prior ledger. A carried row is absent + // from the review, so one of those two is the only record of it. + // typeof, not ||: an empty title is FALSY, and the truthy fallback discarded it -- reading back + // as a pre-format ledger and reopening the leak. + const kt = act && typeof act.t === 'string' ? act.t : priorTitle.get(id); + rows.push({ id, sev: sev(id), d: d, src: src, t: kt, carried: true }); + } + const open = rows.filter((r) => r.d === 'open').length; + const reusedNote = reused.length ? ' (' + reused.length + ' recorded decision(s) DROPPED -- the id now names a different finding, so the decision no longer has a row: ' + reused.join(', ') + ')' : ''; + const staleNote = staleFix.length ? ' (' + staleFix.length + ' fix-report entr' + (staleFix.length === 1 ? 'y titles its' : 'ies title their') + ' finding differently from the review, so ' + (staleFix.length === 1 ? 'it was' : 'they were') + ' not reconciled -- a stale report, or a re-titled one: ' + staleFix.join(', ') + ')' : ''; + if (rows.length === 0 && !unparsed && !fs.existsSync(process.env.DISPOSITION_FILE)) return; + const escapePipes = (t) => t.replace(/\\\\.|\|/g, (m) => (m === '|' ? '\\\\|' : m)); + const body = ['# Phase ' + process.env.PADDED + ': Code Review Disposition', '', '| Finding | Severity | Disposition | Source |', '|---------|----------|-------------|--------|'] + .concat(rows.map((r) => { const src = escapePipes(r.src || '-'); const mark = r.carried && !/\(not in the current review\)\s*\$/.test(src) ? ' (not in the current review)' : ''; return '| ' + r.id + ' | ' + r.sev + ' | ' + r.d + ' | ' + src + mark + ' |'; })) + .concat(['', 'Dispositions: \`open\` (recorded, not yet triaged), \`fixed\`, \`skipped\`, \`deferred\`.', 'Set \`deferred\` by hand and put the reason in the Source cell; both are preserved. A \`|\` in the reason is kept as prose and escaped on the next run.', 'Re-running the gate keeps every row it can. A row the current review no longer reports is kept and its Source cell flagged, so a finding does not leave this record silently. ONE exception: when a finding id is REUSED by a different finding, the earlier decision cannot keep a row — the id is taken — and it is dropped. A RECORDED decision (anything but \`open\`) is named on the console when that happens; a row still at \`open\` is replaced silently, because \`open\` records no decision to lose.', '']).join('\n'); + // One line: the value feeds a line-oriented record a regex re-reads. + const oneLine = (t) => String(t === undefined || t === null ? '' : t).replace(/[\r\n]+/g, ' ').trim(); + // JSON.stringify: YAML 1.2 is a JSON superset, so a colon, quote or leading '#' survives. The + // bare form emitted 'title: Parser: loses data', which a real YAML reader rejects (driven). + const yv = (t) => JSON.stringify(oneLine(t)); + const head = ['---', 'phase: ' + process.env.PADDED, 'review: ' + path.basename(process.env.REVIEW_FILE), 'titles: json', 'findings:'] + .concat(rows.map((r) => ' - id: ' + r.id + '\n severity: ' + r.sev + '\n disposition: ' + r.d + // Emitted whenever KNOWN, empty included ('### CR-01:'). Known-empty vs NOT + // KNOWN is the distinction; conflating them was a leak. Unknown stays absent. + + (typeof r.t === 'string' ? '\n title: ' + yv(r.t) : ''))) + .concat(['open: ' + open, 'total: ' + rows.length]) + // Emitted only when there IS a shortfall, so an ordinary ledger gains no noise key and the + // unchanged-run check below is unaffected on every review that parses cleanly. + .concat(unparsed ? ['unparsed: ' + unparsed] : []).join('\n'); + // Rewrite only on a real change. The timestamp is the one field that always differs, so + // stamping unconditionally would dirty the tree and produce a docs commit on every phase + // re-run with nothing to report. + const render = (stamp) => head + '\nrecorded: ' + stamp + '\n---\n\n' + body; + const stripTs = (t) => t.replace(/^recorded:.*\$/m, 'recorded:'); + const prev = fs.existsSync(process.env.DISPOSITION_FILE) ? norm(fs.readFileSync(process.env.DISPOSITION_FILE, 'utf-8')) : ''; + if (prev && stripTs(prev) === stripTs(render(''))) { + console.log('Code review disposition unchanged: ' + open + ' of ' + rows.length + ' finding(s) open' + staleNote + unparsedNote + reusedNote); + return; + } + fs.writeFileSync(process.env.DISPOSITION_FILE, render(new Date().toISOString())); + console.log('Code review disposition recorded: ' + open + ' of ' + rows.length + ' finding(s) open' + staleNote + unparsedNote + reusedNote + ' — ' + process.env.DISPOSITION_FILE); + })(); +" || echo "Code review disposition record skipped (non-blocking)." + +COMMIT_DOCS=$(gsd_run query config-get commit_docs --raw 2>/dev/null || echo "true") +# `-f` FOLLOWS a symlink, so this could hand the commit helper a link the script above just +# refused to write through -- the guard and its consumer disagreeing about the same path. +if [ "$COMMIT_DOCS" = "true" ] && [ -f "${DISPOSITION_FILE}" ] && [ ! -L "${DISPOSITION_FILE}" ]; then + gsd_run query commit "docs(${PADDED}): record code review disposition" --files "${DISPOSITION_FILE}" || true +fi +``` diff --git a/scripts/lib/platform-conformance-tier.generated.cjs b/scripts/lib/platform-conformance-tier.generated.cjs index b63674dd4..150eaba8e 100644 --- a/scripts/lib/platform-conformance-tier.generated.cjs +++ b/scripts/lib/platform-conformance-tier.generated.cjs @@ -42,6 +42,7 @@ module.exports = { "tests/claude-md.test.cjs", "tests/cline-install.test.cjs", "tests/close-phase-todos-padded-resolves.test.cjs", + "tests/code-review-fix-pipeline-regression.test.cjs", "tests/code-review-pipeline-regression.test.cjs", "tests/code-review-tier3-files-override-scoping.test.cjs", "tests/code-review.test.cjs", diff --git a/scripts/lint-workflow-shellcheck-baseline.json b/scripts/lint-workflow-shellcheck-baseline.json index 515173f22..4aa238be1 100644 --- a/scripts/lint-workflow-shellcheck-baseline.json +++ b/scripts/lint-workflow-shellcheck-baseline.json @@ -1108,5 +1108,20 @@ "file": "gsd-core/workflows/verify-work/steps/mvp-uat-framing.md", "code": "2086", "message": "Double quote to prevent globbing and word splitting." + }, + { + "file": "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", + "code": "2317", + "message": "Command appears to be unreachable. Check usage (or ignore if invoked indirectly)." + }, + { + "file": "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", + "code": "2317", + "message": "Command appears to be unreachable. Check usage (or ignore if invoked indirectly)." + }, + { + "file": "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", + "code": "2317", + "message": "Command appears to be unreachable. Check usage (or ignore if invoked indirectly)." } ] diff --git a/tests/code-review-disposition.property.test.cjs b/tests/code-review-disposition.property.test.cjs new file mode 100644 index 000000000..31f156316 --- /dev/null +++ b/tests/code-review-disposition.property.test.cjs @@ -0,0 +1,427 @@ +'use strict'; + +/** + * code-review-disposition.property.test.cjs + * + * RULESET.TESTS.property-based-testing — the disposition ledger is a + * parse/transformation contract with a render/re-parse fixed point, which is + * the textbook case for a property rather than a fixture. + * + * The step file states the contract in prose: "Re-running the gate preserves + * every disposition except `open`", and it rewrites nothing when nothing + * changed. Both are invariants over an input space that fixtures sample at a + * handful of points — id ordering, severity mix, hand-edited source cells with + * escaped AND bare pipes, carried rows from a review that no longer reports them. + * + * Four properties, all over the SHIPPED script (extracted from the step file + * and executed), never over a model of it. The header said "two" while three + * were running — the vocabulary property arrived without it, so the count is + * now stated per property rather than in a lump: + * + * idempotency — running the gate twice leaves the ledger byte-identical and + * the second run reports `unchanged`. + * round-trip — every decided disposition and its source cell survives that + * second run, i.e. render → re-parse → render is the identity + * on the decision. + * vocabulary — a disposition outside the documented enum is coerced to + * `open` rather than treated as a decision, and the `open:` + * headline agrees with the rows it renders. + * titles — a finding's TITLE survives the same cycle, through + * JSON.stringify into the frontmatter and JSON.parse back out. + * This is the contract that decides whether a reused id names + * the same finding, so a lossy round-trip does not corrupt a + * title — it DESTROYS a human's recorded triage, silently. The + * property asserts the consequence, not the JSON. + * + * Every arbitrary here is built at module scope, not in a `describe` body — + * the fast-check v4 trap that cancels a whole block. + * + * numRuns is lowered from the shared 200 because each case spawns the shipped + * script twice through the process seam; the seed stays pinned, so failures + * still reproduce. Deviating silently would be the worse trade. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const os = require('node:os'); +const path = require('node:path'); +const fc = require('./helpers/fast-check-setup.cjs'); +const { runNode, OUTCOME } = require('./helpers/process-seam.cjs'); +const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); +const { cleanup } = require('./helpers.cjs'); + +const ROOT = path.join(__dirname, '..'); +const DISPOSITION_STEP_PATH = path.join( + ROOT, 'gsd-core', 'workflows', 'execute-phase', 'steps', 'code-review-disposition.md' +); + +const RUNS = { numRuns: 40 }; + +// The SHIPPED node script, undoing exactly the two escapes its surrounding +// double-quoted shell string requires. Deliberately not a model of it. +function shippedDispositionScript() { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8').replace(/\r\n/g, '\n'); + const open = src.indexOf('node -e "'); + assert.ok(open !== -1, 'the disposition step must still embed a node -e script'); + const body = src.slice(open + 'node -e "'.length); + const end = body.indexOf('\n" || echo '); + assert.ok(end !== -1, 'the node -e script must still be closed by its || echo fallback'); + return body.slice(0, end).replace(/\\([\\$`"])/g, '$1'); +} + +const SCRIPT = shippedDispositionScript(); + +function runOnce(dir, padded) { + const res = runNode(['-e', SCRIPT], { + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + REVIEW_FILE: path.join(dir, padded + '-REVIEW.md'), + DISPOSITION_FILE: path.join(dir, padded + '-REVIEW-DISPOSITION.md'), + FIX_REPORT_FILE: path.join(dir, padded + '-REVIEW-FIX.md'), + PADDED: padded, + }, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the shipped script must run to completion'); + assert.strictEqual(res.exitCode, 0, 'the shipped script must exit 0: ' + res.stderr); + const p = path.join(dir, padded + '-REVIEW-DISPOSITION.md'); + return { + ledger: fs.existsSync(p) ? fs.readFileSync(p, 'utf8') : null, + unchanged: /disposition unchanged/.test(res.stdout), + // The reuse report is a CONSOLE note, not a ledger key — the round-trip property below + // asserts on it, and asserting `^reused:` against the ledger would have been vacuously + // true forever, which is the shape of a test that cannot fail. + stdout: res.stdout, + }; +} + +// The id prefixes gsd-code-reviewer.md's body template and its Label-equivalence +// paragraph can emit — the domain the step's own enumeration is a subset of. +const PREFIX = fc.constantFrom('CR', 'BL', 'WR', 'IN'); +const DECIDED = fc.constantFrom('fixed', 'skipped', 'deferred'); + +// The vocabulary the ledger documents, and a generator for its COMPLEMENT. +// DECIDED is drawn from the vocabulary, so no property built on it can ever present the parser +// with an out-of-vocabulary token — the round-trip property below is structurally unable to fail +// on one, which is precisely how a bare ([a-z]+) capture shipped past it. JUNK is the arbitrary +// that reaches the class DECIDED cannot: lowercase, so it stays inside the old capture's own +// character set, because a token OUTSIDE it ('Deferred') already failed safe. The unsafe half is +// the one that looks like a decision and is not. +const VOCABULARY = ['open', 'fixed', 'skipped', 'deferred']; +const JUNK = fc.stringMatching(/^[a-z]{2,10}$/).filter((s) => !VOCABULARY.includes(s)); + +// A hand-written Source cell: the one place a human writes prose into the +// ledger, so it carries the escaped pipe the rendered instruction asks for — AND the bare +// pipe a human actually types. Round 3 of #3861 found that this generator only ever emitted +// the escaped form, so the property built to stress this cell was structurally unable to +// reach the one input that broke it: a bare | failed the whole-line prior-row match, the +// finding reset to open and the reason was destroyed. A generator that reaches only the +// inputs the parser was written for is a fixture with extra steps. +// The reserved suffix is generated DELIBERATELY. The gate strips a carried marker before storing +// it, so a hand-written reason that merely ENDS in that phrase is the input most likely to be +// eaten — and a generator drawn only from innocuous characters can never produce it. Found by +// adversarial review of this file's first cut, which is the argument for putting it in. +const SOURCE_CELL = fc.stringMatching(/^[A-Za-z0-9 .,()-]{0,24}$/) + .map((s) => s.trim() || 'recorded') + .chain((s) => fc.boolean().map((withPipe) => (withPipe ? s + ' \\| see ADR-9' : s))) + .chain((s) => fc.boolean().map((barePipe) => (barePipe ? s + ' | team B to align' : s))) + // Adjacent pipes and a backslash of either parity before a pipe: the first render escape got both + // wrong while passing every input above (round 3, adversarial pass over the fix). + .chain((s) => fc.constantFrom('', ' A||B', ' C\\\\|D', ' E\\\\\\|F').map((t) => s + t)) + .chain((s) => fc.boolean().map((reserved) => (reserved ? s + ' (not in the current review)' : s))); + +const IDS = fc.uniqueArray( + fc.tuple(PREFIX, fc.integer({ min: 1, max: 99 })).map(([p, n]) => p + '-' + String(n).padStart(2, '0')), + { minLength: 1, maxLength: 6 } +); + +// A finding TITLE, which is a parsed and re-serialized value and therefore an input space, +// not decoration. Titles entered this file's contract when the ledger began carrying them to +// tell a reused id from the same finding; the generator did not follow, and every heading was +// built as the fixed string 'finding number <i>' — no colon, quote, backslash, or empty string. +// That is the same shape as this PR's round-3 blocker, where SOURCE_CELL emitted only a +// PRE-ESCAPED pipe and so could never reach the bare one that broke the parser. A generator +// drawn only from safe characters is a fixture with extra steps, and it cannot fail on the one +// input the code under test was written for. +// +// The class is chosen from what the render's own comments say the escaping is FOR: +// `:` the reason yv() exists at all — a bare `title: Parser: loses data` is not YAML. +// `"` and `\` what JSON.stringify must escape and JSON.parse must give back unchanged. +// `` the empty string — 'known-empty vs NOT KNOWN' is a distinction the render draws +// explicitly (`typeof r.t === 'string'`), and conflating them was a stated leak. +// `#` a YAML comment leader in scalar position. +// plus scalars that MIMIC the ledger's own frontmatter grammar (`findings:`, `titles: json`, +// a nested ` title: ` line), because the re-parser walks that grammar by line shape and a +// title is the one field a human-visible artifact copies verbatim out of a review. +// +// BOUND, stated rather than silently omitted: no CR or LF. A `###` heading is one line by +// definition, so a newline is not an input the heading parser can be handed at all — oneLine() +// guards the value's other producers, not this one. Widening here would generate a review this +// repo's own reviewer agent cannot emit. +const TITLE_CHARS = fc.constantFrom( + 'a', 'Z', '9', ' ', '\t', ':', '"', '\\', '#', '|', "'", '{', '}', '[', ']', ',', '-', '_', '.', + 'é', '✓', '—' +); +const TITLE = fc.oneof( + fc.array(TITLE_CHARS, { maxLength: 24 }).map((a) => a.join('')), + fc.constantFrom( + '', + 'Parser: loses data', + 'a "quoted" title', + 'a backslash \\ and a quote "', + 'title: not a key', + 'findings:', + 'titles: json', + ' title: "nested"', + ' - id: CR-99', + 'null', + '"null"', + '#3829 leading hash', + ' surrounded by space ', + 'unicode — é ✓ 中文', + 'ends with a backslash \\', + 'a title long enough to outrun a scanner that assumes short scalars: ' + 'x'.repeat(200) + ) +); + +// One TITLE per id, so every property below runs the render/re-parse cycle over the title +// contract rather than over a constant. +const FINDINGS = IDS.chain((ids) => + fc.array(TITLE, { minLength: ids.length, maxLength: ids.length }) + .map((titles) => ids.map((id, i) => ({ id, title: titles[i] }))) +); + +function reviewFor(findings) { + return ['---', 'phase: 01', 'status: issues_found', '---', ''] + .concat(findings.map((f) => '### ' + f.id + ': ' + f.title + '\n')).join('\n'); +} + +// A SECTION severity per finding, drawn independently of its id prefix. Severity entered the +// ledger's round-trip contract the moment the section began outranking the prefix (round 2, M3): +// from then on the recorded value carried information the id alone could not reproduce, and a +// carried row that re-derived it from the prefix was lossy on exactly the findings the section +// rule exists for. Every fixture used ids whose prefix happened to match their section, so the +// lossy path returned the right answer by coincidence -- the same shape as this file's own +// origin story. Drawn from the three headings gsd-code-reviewer.md emits. +const SECTION = fc.constantFrom('critical', 'warning', 'info'); +const SECTIONED = FINDINGS.chain((findings) => + fc.array(SECTION, { minLength: findings.length, maxLength: findings.length }) + .map((secs) => findings.map((f, i) => ({ id: f.id, title: f.title, section: secs[i] }))) +); +const SECTION_HEADING = { critical: '## Critical Issues', warning: '## Warnings', info: '## Info' }; +function sectionedReviewFor(findings) { + const lines = ['---', 'phase: 01', 'status: issues_found', '---', '']; + for (const sev of ['critical', 'warning', 'info']) { + lines.push(SECTION_HEADING[sev], ''); + for (const f of findings) if (f.section === sev) lines.push('### ' + f.id + ': ' + f.title, ''); + } + return lines.join('\n'); +} + +// What the ledger MUST hold for a title, derived from the heading grammar rather than copied +// from the render. ID_RE's `:\s*` eats the leading whitespace and its `.trim()` the trailing, so +// the stored value is the trimmed title; oneLine() is then the identity on it, because this +// generator emits no CR or LF (the bound above). +// +// TRIM, NOT A `\s+` COLLAPSE — the distinction is load-bearing and easy to get backwards. +// Internal runs of whitespace are collapsed by sameTitle(), which is the COMPARISON rule, and are +// preserved by oneLine(), which is the STORAGE rule. This asserts storage. Writing the collapse +// here would fail on an internal tab against entirely correct code, which is the shape of a test +// that gets weakened rather than believed the first time it goes red. +const expectedTitle = (t) => String(t).trim(); + +describe('#3829 — the disposition ledger is a render/re-parse fixed point', () => { + test('re-running the gate rewrites nothing and changes nothing', () => { + fc.assert(fc.property(FINDINGS, (findings) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-prop-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewFor(findings)); + const first = runOnce(dir, '01'); + assert.ok(first.ledger !== null, 'a review with findings must produce a ledger'); + const second = runOnce(dir, '01'); + // Idempotency: the second run is a no-op, and says so. The timestamp is + // the one field that always differs, so a gate that stamped it + // unconditionally would dirty the tree on every phase re-run. + assert.ok(second.unchanged, 'the second run must report the ledger unchanged'); + assert.strictEqual(second.ledger, first.ledger, 'the second run must not rewrite the ledger'); + } finally { + cleanup(dir); + } + }), RUNS); + }); + + test('a recorded decision and its reason survive re-rendering', () => { + fc.assert(fc.property(FINDINGS, DECIDED, SOURCE_CELL, (findings, decision, source) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-prop-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewFor(findings)); + // A human decides the first finding by hand, reason included — the case + // the Source cell exists for. + const target = findings[0].id; + fs.writeFileSync( + path.join(dir, '01-REVIEW-DISPOSITION.md'), + '| ' + target + ' | critical | ' + decision + ' | ' + source + ' |\n' + ); + const out = runOnce(dir, '01'); + const row = out.ledger.split('\n').find((l) => l.startsWith('| ' + target + ' ')); + assert.ok(row, 'the decided finding must still have a row'); + const cells = row.split(/\s\|\s/).map((c) => c.replace(/^\|\s*|\s*\|$/g, '').trim()); + // Round-trip: render -> re-parse -> render is the identity on the + // decision AND on the reason. `open` never overwrites either. + assert.strictEqual(cells[2], decision, 'the recorded disposition must survive'); + // The reason survives verbatim EXCEPT for one trailing carried marker, which the parse + // removes because it is indistinguishable from the one the render appends. That is the + // contract, not a weakened assertion: leaving a stored marker in place makes the ledger + // claim a finding is 'not in the current review' on the very run that reports it, and + // stripping unboundedly ate the cell. One occurrence, one direction, stated here so the + // trade is visible rather than discovered. + const MARK = /\s*\(not in the current review\)\s*$/; + // A bare | in the reason is kept as prose and ESCAPED on re-render, so the rendered + // cell is the escaped form of what the human wrote — same text under any markdown + // renderer, and the file converges on the second run. Before this the row simply + // failed to parse and the decision was lost, which is the defect this arbitrary now reaches. + // An INDEPENDENT oracle, deliberately not the render's own scan: a pipe is escaped iff an + // EVEN number of backslashes (including zero) immediately precedes it, counted by walking + // the string. A copy of the production regex here would agree with it when both are wrong + // (the round-3 adversarial pass refused exactly that shape), and a single-character + // look-behind agreed with the first, wrong render -- the negative control caught the mirror. + const escapePipes = (t) => { + let out = '', run = 0; + for (const ch of t) { + if (ch === '\\') { run += 1; out += ch; continue; } + if (ch === '|' && run % 2 === 0) out += '\\|'; else out += ch; + run = 0; + } + return out; + }; + assert.strictEqual( + cells[3], escapePipes(source.replace(MARK, '')), + 'the reason survives, bare pipes escaped, less at most one trailing carried marker' + ); + assert.doesNotMatch(cells[3], MARK, 'and a current finding is never marked as carried'); + } finally { + cleanup(dir); + } + }), RUNS); + }); + + test('an out-of-vocabulary disposition is coerced to open, never treated as a decision', () => { + fc.assert(fc.property(FINDINGS, JUNK, SOURCE_CELL, (findings, junk, source) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-prop-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewFor(findings)); + const target = findings[0].id; + // One transposed character is the whole input. Under ([a-z]+) this was stored as a + // decision: not the literal 'open', so it beat the default, was excluded from the open: + // headline count, and was carried forward forever — a ledger reporting a phase fully + // triaged off a typo. ADR-227's rule is that a value failing the enum check is coerced + // to the contract's safe default, and 'open' is that default. + fs.writeFileSync( + path.join(dir, '01-REVIEW-DISPOSITION.md'), + '| ' + target + ' | critical | ' + junk + ' | ' + source + ' |\n' + ); + const out = runOnce(dir, '01'); + const row = out.ledger.split('\n').find((l) => l.startsWith('| ' + target + ' ')); + assert.ok(row, 'the finding must still have a row'); + const cells = row.split(/\s\|\s/).map((c) => c.replace(/^\|\s*|\s*\|$/g, '').trim()); + assert.strictEqual(cells[2], 'open', 'a junk token is not a decision'); + // And the headline count must AGREE with the row it renders. This is the half the + // original defect actually reported wrongly: the row said one thing, `open: N` another. + assert.match(out.ledger, /^open: (\d+)$/m); + const open = Number(/^open: (\d+)$/m.exec(out.ledger)[1]); + assert.strictEqual(open, findings.length, 'every finding is open, and the count says so'); + } finally { + cleanup(dir); + } + }), RUNS); + }); + + test('a title survives the ledger round-trip, so a decision is not lost to its own title', () => { + fc.assert(fc.property(FINDINGS, DECIDED, (findings, decision) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-prop-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewFor(findings)); + const first = runOnce(dir, '01'); + assert.ok(first.ledger !== null, 'a review with findings must produce a ledger'); + const target = findings[0].id; + + // (1) THE STORED FORM. `title: <scalar>` must JSON.parse back to the trimmed heading + // title. A bare scalar would have shipped `title: Parser: loses data`, which a real YAML + // reader rejects — that is why yv() is JSON.stringify and not the value itself. + const fmLine = first.ledger.split('\n') + .find((l, i, all) => l.startsWith(' title: ') + && all.slice(0, i).reverse().find((p) => p.startsWith(' - id: ')) === ' - id: ' + target); + assert.ok(fmLine, 'the ledger must carry a title for ' + target); + assert.strictEqual( + JSON.parse(fmLine.slice(' title: '.length)), expectedTitle(findings[0].title), + 'the stored title must JSON.parse back to the title the heading carried' + ); + + // (2) THE CONSEQUENCE, which is what makes this a property and not an assertion about + // JSON. Decide the finding by EDITING THE RENDERED LEDGER IN PLACE — never by writing a + // bare row. A bare row carries no frontmatter, so priorTitle is empty, sameFinding() + // returns true through its `!priorTitle.has(id)` back-compat arm, and the title contract + // is never consulted at all: the test would pass over a completely broken round-trip. + // That collapse is the reason this property is written this way. + const decided = first.ledger.replace( + new RegExp('^\\| ' + target + ' \\| ([a-z]+) \\| open \\|', 'm'), + '| ' + target + ' | $1 | ' + decision + ' |' + ); + assert.notStrictEqual(decided, first.ledger, 'the hand-edit must actually change the row'); + fs.writeFileSync(path.join(dir, '01-REVIEW-DISPOSITION.md'), decided); + + // The review is UNCHANGED, so the id still names the same finding and the decision must be + // inherited. A lossy round-trip is observable exactly here: priorTitle would disagree with + // the heading, sameFinding() would go false, and the row would reset to `open` — a human's + // recorded triage destroyed by nothing but the shape of its own title. + const second = runOnce(dir, '01'); + const row = second.ledger.split('\n').find((l) => l.startsWith('| ' + target + ' ')); + assert.ok(row, 'the decided finding must still have a row'); + const cells = row.split(/\s\|\s/).map((c) => c.replace(/^\|\s*|\s*\|$/g, '').trim()); + assert.strictEqual(cells[2], decision, 'the decision survives its own title'); + assert.doesNotMatch( + second.stdout, /DROPPED/, + 'and no decision is reported dropped — the id still names the same finding' + ); + + // (3) FIXED POINT. Third run rewrites nothing — the re-parsed title re-renders to itself. + const third = runOnce(dir, '01'); + assert.ok(third.unchanged, 'the third run must report the ledger unchanged'); + assert.strictEqual(third.ledger, second.ledger, 'render -> re-parse -> render is a fixed point'); + } finally { + cleanup(dir); + } + }), RUNS); + }); + + test('a carried row keeps the severity its section gave it, whatever its id prefix says', () => { + fc.assert(fc.property(SECTIONED, (findings) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-prop-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), sectionedReviewFor(findings)); + const first = runOnce(dir, '01'); + assert.ok(first.ledger !== null, 'a review with findings must produce a ledger'); + // Run 2: the review reports NOTHING, so every row is carried. This is the path that + // rebuilt severity from the prefix -- sectionSev holds only the current review's + // findings, which is now none of them. + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewFor([])); + const second = runOnce(dir, '01'); + for (const f of findings) { + const row = second.ledger.split('\n').find((l) => l.startsWith('| ' + f.id + ' ')); + assert.ok(row, f.id + ' must be carried, not dropped'); + const cells = row.split(/\s\|\s/).map((c) => c.replace(/^\|\s*|\s*\|$/g, '').trim()); + assert.strictEqual(cells[1], f.section, + f.id + ': the carried row must keep the section severity, not re-infer it from the prefix'); + assert.match(row, /\(not in the current review\)\s*\|$/, 'and be marked carried'); + } + // And a THIRD run changes nothing: reading the severity back and writing it again is a + // fixed point, so the unchanged-run check is not defeated by the new field. + const third = runOnce(dir, '01'); + assert.ok(third.unchanged, 'the third run must report unchanged'); + } finally { + cleanup(dir); + } + }), RUNS); + }); +}); diff --git a/tests/code-review-fix-pipeline-regression.test.cjs b/tests/code-review-fix-pipeline-regression.test.cjs index da17e8563..e0ffda4f3 100644 --- a/tests/code-review-fix-pipeline-regression.test.cjs +++ b/tests/code-review-fix-pipeline-regression.test.cjs @@ -16,7 +16,28 @@ const { describe, test } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('node:fs'); const path = require('node:path'); -const { runNode } = require('./helpers/process-seam.cjs'); +const { runNode, runHook } = require('./helpers/process-seam.cjs'); +const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); +// Same gate the shipped fences carry everywhere in this suite: bash is not assumed on win32. +const HAS_BASH = process.platform !== 'win32'; + +// Extract the bash fence out of a named <step>, so a test can RUN it rather than grep it. +// Line-scanned rather than matched with a multiline regex: an ad-hoc fence regex trips +// local/no-adhoc-markdown-parsing, and a bare \n against readFileSync content trips +// local/no-crlf-fragile-split (Windows autocrlf yields \r\n). Same shape as bashFences() in +// tests/code-review-pipeline-regression.test.cjs, which solved this first. +function stepFence(src, name) { + const open = src.indexOf(`<step name="${name}">`); + assert.notStrictEqual(open, -1, `step "${name}" not found`); + const close = src.indexOf('</step>', open); + const lines = src.slice(open, close).replace(/\r\n/g, '\n').split('\n'); + let start = -1; + for (let i = 0; i < lines.length; i++) { + if (start === -1 && /^```bash\s*$/.test(lines[i])) { start = i + 1; continue; } + if (start !== -1 && /^```\s*$/.test(lines[i])) return lines.slice(start, i).join('\n'); + } + assert.fail(`step "${name}" carries no bash fence`); +} const { createTempDir, cleanup } = require('./helpers.cjs'); const ROOT = path.resolve(__dirname, '..'); @@ -165,21 +186,85 @@ describe('#3190 — code-review-fix --auto REVIEW.md commit + FIX_REPORT_PATH en // backups are removed so the phase directory is clean. Backup CREATION is // unchanged (out of scope); they are retained when the loop degrades. // ------------------------------------------------------------------------- - test('T6 — auto_iteration_loop removes spent .iterN.md backups on convergence', () => { + test('T6 — spent .iterN.md backups are removed on convergence, AFTER the ledger reads them', () => { + // #3190's SEMANTICS are unchanged and still asserted here: removed on convergence, retained on + // degradation, creation intact. What moved is the PLACEMENT, and it had to. This workflow keeps + // one final version of REVIEW.md and REVIEW-FIX.md rather than per-iteration copies, and the + // re-review drops a finding once it is fixed — so the backups are the only surviving record of + // what an early --auto iteration closed. Removing them at the end of the loop erased that record + // before `record_disposition` could read it, and the ledger then reported those findings as + // `open (not in the current review)`: indistinguishable from never triaged (#3861 round 5). const src = fs.readFileSync(WORKFLOW_PATH, 'utf8'); - const region = stepRegion(src, 'auto_iteration_loop'); - assert.ok( - region.includes('CONVERGED'), - 'auto_iteration_loop must track a CONVERGED flag distinguishing convergence from degradation', - ); - assert.ok( - /rm -f[\s\S]*?\.iter[\s\S]*?\*[\s\S]*?\.md/.test(region), - 'on convergence the loop must remove spent .iterN.md backups (rm -f … .iter*.md)', - ); + const loop = stepRegion(src, 'auto_iteration_loop'); + const cleanup = stepRegion(src, 'cleanup_iteration_backups'); + // Backups are still CREATED before each overwrite (the unchanged mechanism). assert.ok( - region.includes('.iter${ITERATION}.md'), + loop.includes('.iter${ITERATION}.md'), 'backup creation (cp … .iter${ITERATION}.md) must remain intact', ); + // And no longer removed inside the loop — that is the regression this test now guards. + assert.ok( + !/rm -f[\s\S]*?\.iter[\s\S]*?\*[\s\S]*?\.md/.test(loop), + 'the loop must NOT remove the backups it wrote; the ledger has not read them yet', + ); + // Removal lives in its own step, still gated on convergence. + assert.ok( + /rm -f[\s\S]*?\.iter[\s\S]*?\*[\s\S]*?\.md/.test(cleanup), + 'cleanup_iteration_backups must remove spent .iterN.md backups (rm -f … .iter*.md)', + ); + // The token greps that stood here PASSED with the branch INVERTED — the round's own review drove + // it by flipping `= "clean"` to `!=` and watching the test stay green. Checking that FINAL_STATUS, + // rm and "retained" occur SOMEWHERE says nothing about which branch each sits in. Execute it. + assert.match(cleanup, /FINAL_STATUS/, 'convergence must still be re-derived from the review status'); + // Ordering is the whole point of the move. + const recordAt = src.indexOf('<step name="record_disposition">'); + const cleanupAt = src.indexOf('<step name="cleanup_iteration_backups">'); + assert.ok(recordAt > -1 && cleanupAt > -1, 'both steps must exist'); + assert.ok(recordAt < cleanupAt, 'the ledger reads the backups BEFORE they are removed'); + }); + + test('T6b — the cleanup fence, EXECUTED: removes on clean, retains on anything else', { skip: !HAS_BASH }, () => { + // The behavioural half of T6, and the one that survives a branch inversion. Drives the shipped + // fence in both directions against real files, so an inverted condition fails here rather than + // passing a token grep — which is exactly what the round's review demonstrated of the first cut. + const src = fs.readFileSync(WORKFLOW_PATH, 'utf8'); + const fence = stepFence(src, 'cleanup_iteration_backups'); + + const drive = (status) => { + const dir = createTempDir('gsd-3861-cleanup-'); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), ['---', 'status: ' + status, '---'].join('\n')); + fs.writeFileSync(path.join(dir, '01-REVIEW.iter2.md'), 'scratch'); + fs.writeFileSync(path.join(dir, '01-REVIEW-FIX.iter2.md'), 'scratch'); + const script = [ + 'set -euo pipefail', + 'AUTO_MODE=true', + 'REVIEW_PATH="' + path.join(dir, '01-REVIEW.md') + '"', + 'FIX_REPORT_PATH="' + path.join(dir, '01-REVIEW-FIX.md') + '"', + fence, + ].join('\n'); + const res = runHook('-c', [script], { interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS }); + assert.strictEqual(res.exitCode, 0, 'the cleanup step must never abort: ' + res.stderr); + return { + review: fs.existsSync(path.join(dir, '01-REVIEW.iter2.md')), + fix: fs.existsSync(path.join(dir, '01-REVIEW-FIX.iter2.md')), + stdout: res.stdout, + }; + } finally { cleanup(dir); } + }; + + const converged = drive('clean'); + assert.strictEqual(converged.review, false, 'a converged run removes the REVIEW backup'); + assert.strictEqual(converged.fix, false, 'and the FIX-REPORT backup'); + + const degraded = drive('issues_found'); + assert.strictEqual(degraded.review, true, 'a degraded run RETAINS the REVIEW backup'); + assert.strictEqual(degraded.fix, true, 'and the FIX-REPORT backup'); + assert.match(degraded.stdout, /retained/i, 'and says so, so the trail is discoverable'); + + // An unreadable or unparseable review is not evidence of convergence — retain. + const unknown = drive(''); + assert.strictEqual(unknown.review, true, 'an unparseable status must retain, never remove'); }); }); diff --git a/tests/code-review-pipeline-regression.test.cjs b/tests/code-review-pipeline-regression.test.cjs index f5be4d2b6..ecf14b546 100644 --- a/tests/code-review-pipeline-regression.test.cjs +++ b/tests/code-review-pipeline-regression.test.cjs @@ -27,7 +27,7 @@ const assert = require('node:assert/strict'); const fs = require('node:fs'); const path = require('node:path'); const fc = require('fast-check'); -const { runHook } = require('./helpers/process-seam.cjs'); +const { runHook, runNode, OUTCOME } = require('./helpers/process-seam.cjs'); const { toLegacyResult, gitOrThrow } = require('./helpers/git-fixture.cjs'); const { PROBE_TIMEOUT_MS, GIT_TIMEOUT_MS, HOOK_FANOUT_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); const { createTempDir, createTempGitProject, cleanup, readFileNormalized } = require('./helpers.cjs'); @@ -35,6 +35,7 @@ const { foldShellContinuations, findShellFencedMatches, } = require('./helpers/shell-doc-scan.cjs'); +const os = require('node:os'); /** * A single invocation of the external `fallow` binary's `audit` @@ -44,10 +45,50 @@ const { const FALLOW_AUDIT_TIMEOUT_MS = 120000; const ROOT = path.resolve(__dirname, '..'); +// HOISTED. `const` is in the temporal dead zone until its declaration executes, and a +// `{ skip: !HAS_BASH }` option object is evaluated EAGERLY when its describe body runs -- +// so a bash-gated test added ABOVE the old mid-file declaration threw a ReferenceError +// that aborted the whole describe and CANCELLED its siblings. It caught three separate +// additions in this round alone, and the cancellation reads as a passing run in the +// summary line. Declared with the other file-level constants so placement stops mattering. +// AND NOT A PROBE — but the honest reason is narrower than the first draft of this comment claimed, +// and the correction is worth keeping. Round 5 flagged that every test exercising the shipped bash +// fences is `{ skip: !HAS_BASH }`, so block 1's severity-reporting path has no Windows-lane coverage. +// The gap is real and the count is 37 — 22 was the number of `{ skip: !HAS_BASH }` CALL SITES, and a +// skip on a `describe` cancels its subtests. +// +// This comment first justified the skip by citing `local/no-unguarded-nonportable-exec` as REQUIRING +// this exact guard. That was checked and is wrong on both halves: the rule only fires on a file that +// also chmods an exec bit with an octal literal (this file has none, so it never fires here), and +// `eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus `os.platform()`, not one. A +// constraint that exists is not a constraint that applies. +// +// What holds — and this is now MEASURED, not assumed. The measurement the previous version of this +// comment deferred has been made, on native Windows (not WSL) with Git Bash 5.2.37 / MINGW64 first on +// PATH, node v25.2.1 — the same shell family the repo's `windows-latest` lane runs: +// +// HAS_BASH left alone: 179 tests, 127 pass, 0 fail, 52 skipped +// HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped +// +// So 37 of the skips are this guard's (52 - 15; the other 15 skip for unrelated reasons), and +// unskipping them does NOT reveal a clean win: 13 pass and 24 fail. The failures cluster on exactly +// the divergence the eslint rule's subject line names — `bash -c` quoting (one surfaces as +// `unexpected EOF while looking for matching '"'`), empty captured output, and two outright +// `spawn_failed`. Flipping this constant to a runtime probe today would red the Windows lane with 24 +// failures, so the guard STAYS; what changes is that it now documents a measured gap instead of an +// assumed one. Closing it means porting the fences themselves, which is a change of its own, not a +// line in a review round. +// +// NOTE for anyone re-running this: on a WSL host `bash` on the Windows PATH resolves to +// C:\Windows\system32\bash.exe, the WSL bridge — measuring through that runs real Linux bash and +// reports a false clean. Put `C:\Program Files\Git\bin` first. +const HAS_BASH = process.platform !== 'win32'; const WORKFLOW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'code-review.md'); const PRE_PASS_STEP_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'code-review', 'steps', 'structural-pre-pass.md'); const FIXER_PATH = path.join(ROOT, 'agents', 'gsd-code-fixer.md'); const REVIEWER_PATH = path.join(ROOT, 'agents', 'gsd-code-reviewer.md'); +const EXECUTE_PHASE_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md'); +const DISPOSITION_STEP_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase', 'steps', 'code-review-disposition.md'); // --------------------------------------------------------------------------- // #4259: the T6 docs-parity site scan, hoisted out of the assertion so it can @@ -1578,6 +1619,1209 @@ describe('CONS-01..03 — external reviewer evidence consolidation (#4209)', () }); }); +// #3829 — execute-phase `code_review_gate`: surface the severity counts it +// already parses, and record a per-finding disposition. +// +// Before this change the gate extracted `status:` from REVIEW.md's frontmatter, +// discarded the `critical`/`warning`/`info`/`total` values sitting in the same +// `sed` range, and printed a message that was byte-identical for a review with +// one `info` finding and a review with a Critical. Nothing recorded what +// happened to any finding, so a phase reached `phase.complete` with Criticals +// standing and no trace they had been seen. +// +// Tested the way this file tests every other embedded parser: a pure-JS mirror +// of the shipped `node -e` script (validated against the shipped block during +// implementation), plus docs-parity assertions on the workflow .md text, which +// is itself the deployed contract. +// --------------------------------------------------------------------------- + +// Mirror of the gate's frontmatter scalar reads. The shipped block strips CR, then takes ONLY +// the first `---` ... `---` block — it does NOT use a sed range, because a sed range re-opens on +// a body `---` and runs to EOF, which leaks body lines into an OPTIONAL key's read. This mirror +// must model that extraction, not a whole-document scan: a mirror that scans the document passes +// on a fixture the shipped code fails, which is the drift this comment exists to prevent. +function parseGateCounts(reviewText) { + const lines = String(reviewText).replace(/\r/g, '').split('\n'); + // Only a CLOSED frontmatter block counts: an unterminated one must yield nothing rather than + // hand the whole review body to the reads below. + let fm = []; + if (lines[0] === '---') { + const buf = []; + let closed = false; + for (let i = 1; i < lines.length; i++) { + if (lines[i] === '---') { closed = true; break; } + buf.push(lines[i]); + } + if (closed) fm = buf; + } + // The shipped reads are `grep -m1 <key> | cut -d: -f2- | sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//'`, + // and BOTH stages matter. This mirror previously modelled `cut -d: -f2 | tr -d ' '`, which REPAIRED a + // malformed scalar into a number twice over -- `tr -d` deleting INTERNAL spaces (`1 0` -> `10`) and + // `-f2` keeping only the second colon-field (`1: junk` -> `1`). Both were retired from the shipped + // reads when an adversarial pass showed a repaired number deciding whether a shortfall was reported. + // `-f2-` keeps the whole scalar and only the ends are trimmed. The class below is POSIX + // [[:space:]] AS THE C LOCALE DEFINES IT -- space, tab, newline, VT, FF, CR -- spelled out + // literally rather than as JS `\s`, which also matches the unicode spaces C does not. That is an + // equivalence, not an approximation, ONLY because the shipped reads pin LC_ALL=C: glibc's C.UTF-8 + // classifies U+2003 as [[:space:]] and as [[:blank:]], so an unpinned shipped read trimmed a + // character this mirror keeps, and the divergence was invisible here because both sides still + // landed on the same arm for every fixture that existed. The anchors below are spelled with the + // same literal class, for the same reason -- `\s` there would match a U+2003 indent the pinned + // shipped grep does not. + const TRIM = /^[ \t\n\v\f\r]+|[ \t\n\v\f\r]+$/g; + const first = (re) => { + for (const line of fm) { + const m = line.match(re); + if (m) return line.slice(line.indexOf(':') + 1).replace(TRIM, ''); + } + return ''; + }; + // The four counts belong to the `findings:` MAPPING, so the mirror scopes to it exactly as the + // shipped awk does: select the findings: block, stop at the next column-0 key. Without this a + // top-level key later named `total:` / `info:` / `critical:` is picked up ahead of the nested + // one. `status:` stays anchored at column 0 because it IS top-level. + const findingsBlock = []; + { + let inBlock = false; + for (const line of fm) { + // The BLOCK BOUNDARY takes the literal class too, and it is not decoration: the shipped + // selector is an `awk` whose `/^findings:[[:space:]]*$/` and `/^[^[:space:]]/` resolve + // through the ambient locale exactly as grep's and sed's did. Pass 6 found the awk had been + // left unpinned while grep and sed were fixed -- `findings:<U+2003>` opened the block under + // C.UTF-8 and did not under C, so the SAME review reported a breakdown on one machine and + // the countless message on another. `\s`/`\S` here would model neither pinned side. + if (/^findings:[ \t\n\v\f\r]*$/.test(line)) { inBlock = true; continue; } + if (inBlock && /^[^ \t\n\v\f\r]/.test(line)) break; + if (inBlock) findingsBlock.push(line); + } + } + // SAME pipeline as `first` above, and the duplication is the point: BOTH helpers model the shipped + // reads, and the counts go through THIS one. A previous edit updated `first` and the comment above + // it while leaving this body on the retired `split(':')[1].replace(/ /g,'')` — the shipped block had + // moved to `-f2-` + end-trim and the mirror had not, which the parity fixtures could not see because + // both parsers landed on the countless arm for every fixture that then existed. The + // `a self-consistent repaired breakdown` fixture below is what separates them. + const firstIn = (lines, re) => { + for (const line of lines) { + const m = line.match(re); + if (m) return line.slice(line.indexOf(':') + 1).replace(TRIM, ''); + } + return ''; + }; + return { + status: first(/^status:(.*)$/), + critical: firstIn(findingsBlock, /^[ \t\n\v\f\r]*(?:critical|blocker):(.*)$/), + warning: firstIn(findingsBlock, /^[ \t\n\v\f\r]*warning:(.*)$/), + info: firstIn(findingsBlock, /^[ \t\n\v\f\r]*info:(.*)$/), + total: firstIn(findingsBlock, /^[ \t\n\v\f\r]*total:(.*)$/), + }; +} + +// ── The SHIPPED disposition builder, executed — not modelled ────────────────── +// +// Rounds 1 and 2 of the pre-filing review both refuted "the mirrors are faithful": a +// hand-written model of a shell-embedded script drifts, and when it drifts the behavioural +// tests pass while the shipped block is broken. Mutation-testing confirmed it — deleting the +// carried-row logic from the shipped file turned nothing red. +// +// So the disposition tests below run the ACTUAL script. It is pure Node (every input arrives +// through process.env), so extracting it and executing it needs no shell and works on Windows. +// The only transformation is undoing the two shell escapes the surrounding double-quoted +// `node -e "…"` string requires: \` -> ` and \$ -> $. +function shippedDispositionScript() { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8').replace(/\r\n/g, '\n'); + const open = src.indexOf('node -e "'); + assert.ok(open !== -1, 'the disposition step must still embed a node -e script'); + const body = src.slice(open + 'node -e "'.length); + const end = body.indexOf('\n" || echo '); + assert.ok(end !== -1, 'the node -e script must still be closed by its || echo fallback'); + // Undo exactly what the surrounding double-quoted shell string does, in ONE left-to-right + // pass: inside "..." a backslash is special only before $ ` " or \\, and everything else is + // literal. Doing these as separate passes (or missing \\\\ -> \\) silently hands the test a + // DIFFERENT regex from the one that ships — which is how an escaped-pipe case passed here + // while failing in the shell. + return body.slice(0, end).replace(/\\([\\$`"])/g, '$1'); +} + +// Run the shipped script against a temp phase dir and return the ledger it wrote (or null). +function runShippedDisposition({ reviewText, priorText, fixText, iterFixText, padded = '01', reviewTotal }) { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3829-')); + try { + const reviewPath = path.join(dir, padded + '-REVIEW.md'); + const dispPath = path.join(dir, padded + '-REVIEW-DISPOSITION.md'); + const fixPath = path.join(dir, padded + '-REVIEW-FIX.md'); + fs.writeFileSync(reviewPath, reviewText); + if (priorText !== undefined) fs.writeFileSync(dispPath, priorText); + if (fixText !== undefined) fs.writeFileSync(fixPath, fixText); + // The --auto loop's per-iteration backups (<NN>-REVIEW-FIX.iterN.md). Keyed by iteration + // number so a test can drive the ordering the script relies on (newest wins). + for (const [n, text] of Object.entries(iterFixText || {})) { + fs.writeFileSync(path.join(dir, padded + '-REVIEW-FIX.iter' + n + '.md'), text); + } + // Bounded by construction via the process seam — an unbounded spawn is an indefinite hang. + const res = runNode(['-e', shippedDispositionScript()], { + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + REVIEW_FILE: reviewPath, + DISPOSITION_FILE: dispPath, + FIX_REPORT_FILE: fixPath, + PADDED: padded, + // The frontmatter total block 2 derives and hands to the script, so the two parsers can + // be reconciled. Passed through here so a test can drive the shortfall path. + REVIEW_TOTAL: reviewTotal === undefined ? '' : String(reviewTotal), + }, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the shipped script must run to completion'); + assert.strictEqual(res.exitCode, 0, 'the shipped script must exit 0: ' + res.stderr); + const ledger = fs.existsSync(dispPath) ? fs.readFileSync(dispPath, 'utf8') : null; + return { ledger, stdout: res.stdout, wroteNothing: /disposition unchanged/.test(res.stdout) }; + } finally { + cleanup(dir); + } +} + +// Parse the rows out of a rendered ledger, so assertions read against real output. +function ledgerRows(ledger) { + if (ledger === null) return null; + const rows = []; + for (const line of ledger.split('\n')) { + const m = line.match(/^\|\s*((?:CR|BL|WR|IN)-\d+)\s*\|\s*([a-z]+)\s*\|\s*([a-z]+)\s*\|\s*(.*?)\s*\|$/); + if (m) rows.push({ id: m[1], severity: m[2], disposition: m[3], source: m[4] }); + } + return rows; +} + +// buildDisposition is NOT a mirror any more — it drives the SHIPPED script. +// +// Three pre-filing review rounds each refuted "the mirror is faithful", and mutation testing +// confirmed the cost: a hand-written model of a shell-embedded script drifts, and when it drifts +// the tests pass while the shipped block is broken. The model is gone; this adapter runs the real +// thing and returns the same shape the assertions below already expect. +function buildDisposition({ reviewText, priorText, fixText, padded = '01', reviewTotal }) { + const rows = ledgerRows(runShippedDisposition({ reviewText, priorText, fixText, padded, reviewTotal }).ledger); + if (rows === null) return null; + return { + rows: rows.map((r) => ({ + ...r, + carried: / \(not in the current review\)$/.test(r.source) || undefined, + source: r.source.replace(/ \(not in the current review\)$/, ''), + })), + open: rows.filter((r) => r.disposition === 'open').length, + total: rows.length, + }; +} + +const REVIEW_WITH_FINDINGS = [ + '---', + 'phase: 01', + 'findings:', + ' critical: 1', + ' warning: 2', + ' info: 1', + ' total: 4', + 'status: issues_found', + '---', + '', + '## Critical Issues', + '', + '### CR-01: SQL injection in auth', + '', + '## Warnings', + '', + '### WR-01: missing null check', + '', + '### WR-02: unused import', + '', + '## Info', + '', + '### IN-01: stale TODO', +].join('\n'); + +describe('#3829 — code_review_gate severity surfacing', () => { + test('the gate reads the counts that sit beside the status it already parsed', () => { + const parsed = parseGateCounts(REVIEW_WITH_FINDINGS); + assert.deepStrictEqual(parsed, { + status: 'issues_found', critical: '1', warning: '2', info: '1', total: '4', + }); + }); + + test('blocker: is accepted as the Critical tier-equivalent of critical:', () => { + const parsed = parseGateCounts(REVIEW_WITH_FINDINGS.replace(' critical: 1', ' blocker: 1')); + assert.strictEqual(parsed.critical, '1'); + }); + + test('a status:/total: line in the review BODY never displaces the frontmatter value', () => { + // The `sed -n '/^---$/,/^---$/p'` range re-opens on a body `---` and runs to + // EOF, so first-match (`grep -m1`) is what makes this correct — not the range. + const poisoned = REVIEW_WITH_FINDINGS + '\n\n---\n\nstatus: clean\ntotal: 999\n'; + const parsed = parseGateCounts(poisoned); + assert.strictEqual(parsed.status, 'issues_found'); + assert.strictEqual(parsed.total, '4'); + }); + + test('a REVIEW.md with no findings: block yields no counts, so the gate can fall back', () => { + const legacy = ['---', 'phase: 02', 'status: issues_found', '---', '', '# Phase 02'].join('\n'); + const parsed = parseGateCounts(legacy); + assert.strictEqual(parsed.status, 'issues_found'); + assert.strictEqual(parsed.total, ''); + assert.strictEqual(parsed.critical, ''); + }); + + + test('docs-parity: the gate states the counts rather than the countless message alone', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes('Code review: ${REVIEW_TOTAL} findings — ${REVIEW_CRITICAL} critical, ${REVIEW_WARNING} warning, ${REVIEW_INFO} info.'), + 'code_review_gate must display the per-severity breakdown it parsed' + ); + }); + + test('docs-parity: every frontmatter scalar read carries a first-match guard', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes('grep -m1 "^status:"'), + 'the status: read must keep its single-match guard (DEFECT.FRONTMATTER-SCALAR-BROAD-GREP)' + ); + }); +}); + +describe('#3829 — code_review_gate per-finding disposition record', () => { + test('every finding is recorded, defaulting to open', () => { + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, padded: '01' }); + assert.strictEqual(d.total, 4); + assert.strictEqual(d.open, 4); + assert.deepStrictEqual(d.rows.map((r) => r.id), ['CR-01', 'WR-01', 'WR-02', 'IN-01']); + assert.deepStrictEqual(d.rows.map((r) => r.severity), ['critical', 'warning', 'warning', 'info']); + }); + + test('BL- findings are recorded at the Critical tier alongside CR-', () => { + const d = buildDisposition({ + reviewText: REVIEW_WITH_FINDINGS.replace('### CR-01:', '### BL-01:'), + padded: '01', + }); + assert.strictEqual(d.rows[0].id, 'BL-01'); + assert.strictEqual(d.rows[0].severity, 'critical'); + }); + + test('--fix outcomes are reconciled from REVIEW-FIX.md with provenance', () => { + const fixText = [ + '# Phase 01: Code Review Fix Report', + '', + '## Fixed Issues', + '', + '### WR-01: missing null check', + '', + '## Skipped Issues', + '', + '### WR-02: unused import', + ].join('\n'); + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, fixText, padded: '01' }); + const byId = Object.fromEntries(d.rows.map((r) => [r.id, r])); + assert.strictEqual(byId['WR-01'].disposition, 'fixed'); + assert.strictEqual(byId['WR-01'].source, '01-REVIEW-FIX.md'); + assert.strictEqual(byId['WR-02'].disposition, 'skipped'); + assert.strictEqual(byId['CR-01'].disposition, 'open'); + assert.strictEqual(d.open, 2); + }); + + test('a finding heading outside the Fixed/Skipped sections is not a disposition', () => { + const fixText = [ + '## Fixed Issues', + '', + '### WR-01: missing null check', + '', + '## Verification', + '', + '### IN-01: mentioned while describing how the fix was verified', + ].join('\n'); + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, fixText, padded: '01' }); + const byId = Object.fromEntries(d.rows.map((r) => [r.id, r])); + assert.strictEqual(byId['WR-01'].disposition, 'fixed'); + assert.strictEqual(byId['IN-01'].disposition, 'open'); + }); + + test('a recorded decision survives a re-run — open never overwrites deferred', () => { + const priorText = [ + '| Finding | Severity | Disposition | Source |', + '|---------|----------|-------------|--------|', + '| CR-01 | critical | deferred | ships next phase |', + '| WR-01 | warning | open | - |', + ].join('\n'); + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, priorText, padded: '01' }); + const byId = Object.fromEntries(d.rows.map((r) => [r.id, r])); + assert.strictEqual(byId['CR-01'].disposition, 'deferred'); + // The source cell carries the human's stated reason and is preserved, not replaced. + assert.strictEqual(byId['CR-01'].source, 'ships next phase'); + assert.strictEqual(d.open, 3); + }); + + test('an applied --fix outcome wins over an earlier deferral', () => { + const priorText = '| CR-01 | critical | deferred | later |'; + const fixText = ['## Fixed Issues', '', '### CR-01: SQL injection in auth'].join('\n'); + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, priorText, fixText, padded: '01' }); + assert.strictEqual(d.rows[0].disposition, 'fixed'); + }); + + test('a review with no finding headings produces no record at all', () => { + const legacy = ['---', 'phase: 02', 'status: issues_found', '---', '', 'prose only'].join('\n'); + assert.strictEqual(buildDisposition({ reviewText: legacy, padded: '02' }), null); + }); + + test('docs-parity: the gate writes a REVIEW-DISPOSITION sibling, not into REVIEW.md', () => { + // Eighth src.includes() converted. It pinned the literal `${PHASE_DIR}` interpolation, so it + // went red when path construction moved to a VALIDATED local -- while the property it names, + // "a REVIEW-DISPOSITION sibling", was untouched. Assert the property: the ledger lands beside + // the review, under the derived name, and REVIEW.md itself is not written. + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-sib-')); + try { + const reviewPath = path.join(dir, '02-REVIEW.md'); + const reviewText = ['---', 'status: issues_found', '---', '', '### CR-01: a finding'].join('\n'); + fs.writeFileSync(reviewPath, reviewText); + const res = runNode(['-e', shippedDispositionScript()], { + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + REVIEW_FILE: reviewPath, + DISPOSITION_FILE: path.join(dir, '02-REVIEW-DISPOSITION.md'), + FIX_REPORT_FILE: path.join(dir, '02-REVIEW-FIX.md'), + PADDED: '02', + }, + }); + assert.strictEqual(res.exitCode, 0, res.stderr); + assert.ok(fs.existsSync(path.join(dir, '02-REVIEW-DISPOSITION.md')), + 'code_review_gate must write the disposition record to a REVIEW-DISPOSITION sibling'); + assert.strictEqual(fs.readFileSync(reviewPath, 'utf8'), reviewText, + 'and must not write into REVIEW.md — gsd-code-reviewer is its single writer'); + } finally { + cleanup(dir); + } + // Sixth src.includes() converted. It pinned the exact SOURCE LINE of the enumeration loop, so + // it went red when M3 reformatted that loop to track the enclosing section -- while the + // property it names, "enumerate every finding ID", was strictly widened rather than broken. + // A pin on a code shape reports every refactor as a regression and no regression as one. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: a', '### WR-01: b', '### IN-01: c', + '### CR-01: a duplicate id, which must not produce a second row'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review }).ledger); + assert.deepStrictEqual( + rows.map((r) => r.id), ['CR-01', 'WR-01', 'IN-01'], + 'code_review_gate must enumerate every finding ID from REVIEW.md, in order, once each' + ); + }); + + test('docs-parity: code_review_gate actually reaches the extracted step', () => { + // The step is only reachable because the parent says to read and execute it. Without this, + // every other docs-parity assertion here could pass against a file nothing loads. + const parent = fs.readFileSync(EXECUTE_PHASE_PATH, 'utf8'); + assert.ok( + /Read and execute\s+`gsd-core\/workflows\/execute-phase\/steps\/code-review-disposition\.md`/.test(parent), + 'code_review_gate must read and execute the disposition step' + ); + assert.ok( + parent.indexOf('code-review-disposition.md') < parent.indexOf('**TDD review escalation'), + 'and must do so before the TDD escalation can stop the phase' + ); + }); + + test('docs-parity: the disposition write is non-blocking', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes('Code review disposition record skipped (non-blocking).'), + 'the disposition write must never block execution flow' + ); + }); +}); + +// --------------------------------------------------------------------------- +// #3829 review round — the seven defects a cross-AI adversarial pass on the fix +// diff refuted before filing. Each is pinned here so the round's work survives. +// --------------------------------------------------------------------------- +describe('#3829 review round — frontmatter scoping, section anchoring, ledger durability', () => { + test('an OPTIONAL count absent from frontmatter is not supplied by a leaked body value', () => { + // A sed range re-opens on a body `---` and runs to EOF. First-match protects a key the + // frontmatter DOES carry; it cannot protect one it does not. Extraction must stop at the + // closing delimiter, or a body `total:` becomes the reported finding count. + const leaky = [ + '---', 'phase: 01', 'status: issues_found', '---', + '', '# Report', '', '---', '', 'total: 7', 'critical: 9', + ].join('\n'); + const parsed = parseGateCounts(leaky); + assert.strictEqual(parsed.status, 'issues_found'); + assert.strictEqual(parsed.total, '', 'a body total: must not become the finding count'); + assert.strictEqual(parsed.critical, '', 'a body critical: must not become the critical count'); + }); + + test('a CRLF-authored review parses to bare values, with no carriage return riding along', () => { + const crlf = [ + '---', 'phase: 01', 'findings:', ' critical: 1', ' warning: 0', ' info: 2', + ' total: 3', 'status: issues_found', '---', '', '### CR-01: a', + ].join('\r\n'); + const parsed = parseGateCounts(crlf); + assert.deepStrictEqual(parsed, { + status: 'issues_found', critical: '1', warning: '0', info: '2', total: '3', + }); + for (const v of Object.values(parsed)) assert.ok(!/\r/.test(v), 'no value may carry a CR'); + }); + + test('docs-parity: the gate stops at the closing delimiter and strips CR before parsing', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes(`awk 'NR==1{if($0!="---") exit; next} /^---$/{closed=1; exit}`), + 'the gate must extract only the FIRST frontmatter block, not a re-opening sed range' + ); + assert.ok( + src.includes(`tr -d '\\r' < "$REVIEW_FILE"`), + 'the gate must strip CR so a CRLF review cannot inject one into the message' + ); + }); + + test('a section heading that merely STARTS with "Fixed Issues" does not classify findings', () => { + const fixText = [ + '# Fix report', '', '## Fixed Issues Verification', '', '### IN-01: named while verifying', + ].join('\n'); + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, fixText, padded: '01' }); + const byId = Object.fromEntries(d.rows.map((r) => [r.id, r])); + assert.strictEqual(byId['IN-01'].disposition, 'open'); + }); + + test('a deferral reason written into the Source cell survives a re-run verbatim', () => { + const priorText = '| CR-01 | critical | deferred | ships next phase, see ADR-99 |'; + const d = buildDisposition({ reviewText: REVIEW_WITH_FINDINGS, priorText, padded: '01' }); + const cr = d.rows.find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred'); + assert.strictEqual(cr.source, 'ships next phase, see ADR-99'); + }); + + test('a decided finding the current review no longer reports is carried, not dropped', () => { + // --auto re-reviews and rewrites REVIEW.md; a fixed or deferred finding can vanish from it. + // Dropping the row would erase the record that it was seen — the exact failure #3829 is about. + const priorText = [ + '| CR-01 | critical | deferred | ships next phase |', + '| IN-01 | info | open | - |', + ].join('\n'); + const shrunk = ['---', 'status: issues_found', '---', '', '### WR-01: still here'].join('\n'); + const d = buildDisposition({ reviewText: shrunk, priorText, padded: '01' }); + const byId = Object.fromEntries(d.rows.map((r) => [r.id, r])); + assert.ok(byId['CR-01'], 'a deferred finding absent from the review must still be recorded'); + assert.strictEqual(byId['CR-01'].disposition, 'deferred'); + assert.strictEqual(byId['CR-01'].carried, true); + assert.strictEqual(byId['CR-01'].source, 'ships next phase'); + // M1: the untriaged row is carried TOO, and marked. Dropping it erased the record that the + // finding was ever seen -- which is #3829's complaint verbatim, reproduced by the artifact + // built to prevent it. 'Nothing was decided about it' is precisely the state that must leave + // a trace. The marker keeps it honest: the row does not claim the finding is live. + assert.ok(byId['IN-01'], 'an untriaged row for a vanished finding is carried, not deleted'); + assert.strictEqual(byId['IN-01'].disposition, 'open'); + assert.strictEqual(byId['IN-01'].carried, true); + }); + + test('the carried marker is rendered, never stored — it cannot accumulate across runs', () => { + // The marker lives in the Source cell, which is re-parsed on the next run. Storing it would + // re-append it every time: the cell grows without bound AND the file changes on every run, + // which silently defeats the unchanged-run check and restores the empty-docs-commit churn. + const shrunk = ['---', 'status: issues_found', '---', '', '### WR-01: b'].join('\n'); + let priorText = '| CR-01 | critical | deferred | ships next phase |'; + let cr; + for (let i = 0; i < 3; i++) { + const d = buildDisposition({ reviewText: shrunk, priorText, padded: '01' }); + cr = d.rows.find((r) => r.id === 'CR-01'); + // Re-render the row the way the shipped builder does, and feed it back in. + priorText = '| ' + cr.id + ' | ' + cr.severity + ' | ' + cr.disposition + ' | ' + + cr.source + (cr.carried ? ' (not in the current review)' : '') + ' |'; + } + assert.strictEqual(cr.source, 'ships next phase', 'the stored source must stay canonical'); + assert.strictEqual( + (priorText.match(/\(not in the current review\)/g) || []).length, 1, + 'the marker must appear exactly once no matter how many times the gate re-runs' + ); + }); + + test('docs-parity: the ledger is rewritten only on a real change, timestamp excluded', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes("const stripTs = (t) => t.replace(/^recorded:.*\\$/m, 'recorded:');"), + 'the gate must compare ignoring the timestamp so an unchanged run writes nothing' + ); + assert.ok( + src.includes('Code review disposition unchanged: '), + 'an unchanged run must say so rather than producing an empty docs commit' + ); + }); + + test('docs-parity: section headings are anchored whole and the prior source cell is captured', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes('/^##\\s+Fixed Issues\\s*\\$/') && src.includes('/^##\\s+Skipped Issues\\s*\\$/'), + 'fix-report section headings must be matched whole, never by prefix' + ); + // Eighth src.includes() converted -- it pinned the exact source-cell CAPTURE, and that + // capture was the round-3 defect (a bare | failed the whole match). The property it named + // is asserted behaviourally below: 'an escaped pipe in a deferral reason survives whole' + // and 'a bare pipe in a deferral reason is kept as prose'. + // Seventh src.includes() converted -- it pinned the exact comparison EXPRESSION, so it went + // red when the comparison gained whitespace normalization while the property it names was + // unchanged. Assert the property instead: a fix report naming a different finding under a + // reused id does not decide the row. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a new finding'].join('\n'); + const stale = ['## Fixed Issues', '', '### CR-01: what this id used to mean'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review, fixText: stale }).ledger); + assert.strictEqual(rows[0].disposition, 'open', + 'a fix report must name the SAME finding before its outcome is applied'); + }); +}); + +describe('#3861 round 2 — the disposition ledger is reachable in a shipped path (B1c/B1d)', () => { + const FIX_WORKFLOW = path.join(ROOT, 'gsd-core', 'workflows', 'code-review-fix.md'); + + test('code-review-fix.md reads and executes the disposition step after committing the fix report', () => { + // Without this the reconciliation logic -- the bulk of the step and nearly all of its test + // surface -- is reachable only on a RE-EXECUTION of the phase, and REQ-REVIEW-10 is unmet in + // every shipped path. execute-phase.md's gate invokes review with neither --fix nor --auto, + // so REVIEW-FIX.md cannot exist there and every row it writes is `open` by construction. + const src = fs.readFileSync(FIX_WORKFLOW, 'utf8'); + assert.match(src, /<step name="record_disposition">/, + 'the fix workflow must carry a step that records the disposition'); + assert.match(src, /gsd-core\/workflows\/execute-phase\/steps\/code-review-disposition\.md/, + 'and it must invoke the SAME step, not a second copy of the logic'); + // Ordering is load-bearing: the fix report must be on disk before the ledger claims anything + // about it. Assert the step positions rather than merely their presence. + const commitAt = src.indexOf('<step name="commit_fix_report">'); + const recordAt = src.indexOf('<step name="record_disposition">'); + const presentAt = src.indexOf('<step name="present_results">'); + assert.ok(commitAt > -1 && recordAt > -1 && presentAt > -1, 'all three steps must exist'); + assert.ok(commitAt < recordAt, 'the ledger is reconciled AFTER the fix report is written'); + assert.ok(recordAt < presentAt, 'and before results are presented'); + }); + + test('the iteration backups are removed AFTER the ledger has read them, not inside the loop', () => { + // #3861 round 5. The .iterN.md backups are the only surviving record of what an earlier --auto + // iteration fixed -- this workflow keeps ONE final version of each artifact, and the re-review + // drops a finding once it is fixed, so neither final artifact carries it. Deleting them at the + // end of the loop meant record_disposition reached a CONVERGED run with every early fix already + // erased and recorded those findings as open. + const src = fs.readFileSync(FIX_WORKFLOW, 'utf8'); + const recordAt = src.indexOf('<step name="record_disposition">'); + const cleanupAt = src.indexOf('<step name="cleanup_iteration_backups">'); + assert.ok(cleanupAt > -1, 'the backups must be removed by a named step, not inline in the loop'); + assert.ok(recordAt < cleanupAt, 'the ledger reads the backups BEFORE they are removed'); + // And the removal must not have been left behind in the loop as well. + const loopAt = src.indexOf('<step name="auto_iteration_loop">'); + const loopBody = src.slice(loopAt, src.indexOf('<step name="commit_fix_report">')); + assert.ok(!/rm -f "\$\{REVIEW_PATH%\.md\}\.iter"/.test(loopBody), + 'the loop must no longer delete the backups it just wrote'); + }); + + test('the reconciliation is reachable: gate writes all-open, the fix path resolves it', () => { + // The two call sites driven in sequence, which is the shipped order. This is the review's own + // input -> wrong output case, inverted: a phase whose findings are all fixed by a subsequent + // --fix run used to end at `open: N / total: N`, asserting that every triaged finding was + // forgotten -- worse than recording nothing, because it looks authoritative and is inverted. + const review = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 1', ' info: 0', ' total: 2', '---', '', + '## Critical Issues', '', '### CR-01: the critical one', '', + '## Warnings', '', '### WR-01: the warning one'].join('\n'); + const fixReport = ['---', 'status: complete', '---', '', + '## Fixed Issues', '', '### CR-01: the critical one', '', + '## Skipped Issues', '', '### WR-01: the warning one'].join('\n'); + + // Call site 1 -- execute-phase.md's code_review_gate. No fix report exists yet. + const gate = runShippedDisposition({ reviewText: review, reviewTotal: 2 }); + const gateRows = ledgerRows(gate.ledger); + assert.deepStrictEqual(gateRows.map((r) => r.disposition), ['open', 'open'], + 'at the gate there is no fix report, so every row is open -- correctly'); + + // Call site 2 -- code-review-fix.md's record_disposition, with the report on disk and the + // gate's ledger carried in as prior state. + const after = runShippedDisposition({ + reviewText: review, priorText: gate.ledger, fixText: fixReport, reviewTotal: 2, + }); + const rows = ledgerRows(after.ledger); + assert.deepStrictEqual( + rows.map((r) => [r.id, r.disposition]), + [['CR-01', 'fixed'], ['WR-01', 'skipped']], + 'the fix outcomes reach the ledger' + ); + assert.match(after.ledger, /^open: 0$/m, 'and the headline agrees'); + assert.match(after.ledger, /^total: 2$/m); + }); +}); + +describe('#3861 round 2 — severity comes from the section, not just the id prefix (M3)', () => { + const sevOf = (reviewText, id) => { + const rows = ledgerRows(runShippedDisposition({ reviewText }).ledger); + const row = rows.find((r) => r.id === id); + assert.ok(row, id + ' must have a row'); + return row.severity; + }; + + test("a Critical mis-numbered as WR- is recorded critical when it sits under '## Critical Issues'", () => { + // The section heading is the reviewer's OWN statement of severity, and the walker already + // visits it. Deriving from the prefix alone puts the ledger's Severity column -- the whole + // basis for triaging it -- in disagreement with the review it summarizes and with the + // frontmatter count line block 1 prints. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Critical Issues', '', '### CR-01: properly numbered', '', + '### WR-04: a critical the reviewer mis-numbered', '', + '## Info', '', '### IN-01: an info item'].join('\n'); + assert.strictEqual(sevOf(review, 'WR-04'), 'critical', 'the section outranks the prefix'); + assert.strictEqual(sevOf(review, 'CR-01'), 'critical'); + assert.strictEqual(sevOf(review, 'IN-01'), 'info'); + }); + + test('an IN- finding under ## Warnings is recorded warning', () => { + // The other direction, so the rule is not one-way: the section governs whichever way the + // prefix disagrees with it. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Warnings', '', '### IN-02: mis-numbered the other way'].join('\n'); + assert.strictEqual(sevOf(review, 'IN-02'), 'warning'); + }); + + test('the id prefix still governs when no recognized section encloses the finding', () => { + // Fallback control. A review that does not use the documented headings -- and every carried + // row from an earlier review -- must keep the prefix mapping, BL- included. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: a', '### BL-02: b', '### WR-03: c', '### IN-04: d'].join('\n'); + assert.strictEqual(sevOf(review, 'CR-01'), 'critical'); + assert.strictEqual(sevOf(review, 'BL-02'), 'critical', 'BL- stays Critical-tier-equivalent'); + assert.strictEqual(sevOf(review, 'WR-03'), 'warning'); + assert.strictEqual(sevOf(review, 'IN-04'), 'info'); + }); + + test('a lookalike section heading does not re-tier the findings under it', () => { + // Matched WHOLE, exactly as the fix-report sections are. A prefix match would let + // '## Critical Issues Verification' promote everything beneath it. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Critical Issues Verification', '', '### IN-05: not actually critical'].join('\n'); + assert.strictEqual(sevOf(review, 'IN-05'), 'info', 'an unrecognized section falls back to the prefix'); + }); + + test('a section heading inside a fenced example does not govern', () => { + // The fence walker already skips fenced content; this pins that the new section tracking + // honours it rather than reading an illustration as document structure. + const review = ['---', 'phase: 01', 'status: issues_found', '---', '', + '```markdown', '## Critical Issues', '```', '', '### IN-06: outside the fence'].join('\n'); + assert.strictEqual(sevOf(review, 'IN-06'), 'info'); + }); +}); + +describe('#3861 round 11 — a carried row keeps the severity the ledger recorded', () => { + // The ledger always WROTE a severity (table cell + frontmatter key) and nothing read it back: + // the row regex skipped the cell as [^|]*, the frontmatter walk collected only titles, and a + // carried row was rebuilt from the id prefix because sectionSev holds only the CURRENT review's + // findings. So the one case round 2's M3 fix exists for -- a Critical the reviewer mis-numbered + // WR-04 -- was recorded critical on run 1 and silently re-recorded warning on run 2, once the + // review stopped reporting it. Every carried-row test used CR-01/IN-01, whose prefix already + // matched, so the fallback returned the right answer by coincidence and no mutation could tell + // the two paths apart. Reproduced by the round-11 review by executing the shipped script twice. + const REVIEW_1 = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Critical Issues', '', '### CR-01: properly numbered', '', + '### WR-04: a critical the reviewer mis-numbered', ''].join('\n'); + // Run 2: the review no longer reports WR-04 at all. + const REVIEW_2_DROPPED = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Critical Issues', '', '### CR-01: properly numbered', ''].join('\n'); + const rowOf = (ledger, id) => { + const row = ledgerRows(ledger).find((r) => r.id === id); + assert.ok(row, id + ' must have a row'); + return row; + }; + // Run 1's ledger, with WR-04 deferred BY HAND, the way the ledger's own legend instructs. Edited + // in place so the frontmatter (and its title) is present -- a bare row would carry no title and + // sameFinding() would pass through its back-compat arm, never consulting identity at all. + const deferredLedger = () => { + const first = runShippedDisposition({ reviewText: REVIEW_1 }).ledger; + assert.strictEqual(rowOf(first, 'WR-04').severity, 'critical', 'run 1 records the section severity'); + const edited = first.replace(/^\| WR-04 \| critical \| open \| - \|$/m, '| WR-04 | critical | deferred | waiting on team A |'); + assert.notStrictEqual(edited, first, 'the hand edit must have landed'); + return edited; + }; + + test("a deferred WR-* filed under '## Critical Issues' stays critical once the review stops reporting it", () => { + const after = runShippedDisposition({ reviewText: REVIEW_2_DROPPED, priorText: deferredLedger() }); + const row = rowOf(after.ledger, 'WR-04'); + assert.strictEqual(row.severity, 'critical', 'the recorded severity survives the carry -- the prefix would say warning'); + assert.strictEqual(row.disposition, 'deferred', 'and so does the decision'); + assert.match(row.source, /^waiting on team A/, 'and the reason'); + assert.match(row.source, /\(not in the current review\)$/, 'and the row is marked carried'); + assert.match(after.ledger, /^ {4}severity: critical$/m, 'the frontmatter agrees with the table'); + }); + + test('the recorded severity also outranks the prefix while the review still reports the finding under no recognized section', () => { + // A re-review that dropped the documented headings. The current review gives no section, the + // ledger remembers what the last one said, and that is a better source than the id alone. + const plain = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: properly numbered', '### WR-04: a critical the reviewer mis-numbered'].join('\n'); + const after = runShippedDisposition({ reviewText: plain, priorText: deferredLedger() }); + assert.strictEqual(rowOf(after.ledger, 'WR-04').severity, 'critical'); + }); + + test("the current review's section still wins over the recorded value", () => { + // Precedence control: a reviewer who re-tiers a finding on re-review is the newer statement. + const retiered = ['---', 'phase: 01', 'status: issues_found', '---', '', + '## Critical Issues', '', '### CR-01: properly numbered', '', + '## Warnings', '', '### WR-04: a critical the reviewer mis-numbered'].join('\n'); + const after = runShippedDisposition({ reviewText: retiered, priorText: deferredLedger() }); + assert.strictEqual(rowOf(after.ledger, 'WR-04').severity, 'warning', 'the section is the current statement'); + }); + + test('a REUSED id does not inherit the old finding\'s severity', () => { + // Identity control, the same rule the disposition takes: ids are reused across re-reviews, so + // a brand-new WR-04 under no section starts from its own prefix, not from the finding it + // replaced. Inheriting here would put a Critical tier on an ordinary warning. + const reused = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: properly numbered', '### WR-04: an entirely different finding'].join('\n'); + const after = runShippedDisposition({ reviewText: reused, priorText: deferredLedger() }); + const row = rowOf(after.ledger, 'WR-04'); + assert.strictEqual(row.severity, 'warning', 'a different finding under a reused id is inferred from its prefix'); + assert.strictEqual(row.disposition, 'open', 'and, as before, does not inherit the decision either'); + }); + + test('a hand-mangled Severity cell falls back to the frontmatter copy, and a mangled pair to the prefix', () => { + // Both sources are enum-validated (ADR-227): a value outside critical|warning|info is not a + // severity. The table is read first because it is the surface a human edits; the frontmatter + // is the copy a human is not invited to touch. + const base = deferredLedger(); + const cellMangled = base.replace('| WR-04 | critical | deferred |', '| WR-04 | critcal | deferred |'); + assert.notStrictEqual(cellMangled, base); + let after = runShippedDisposition({ reviewText: REVIEW_2_DROPPED, priorText: cellMangled }); + assert.strictEqual(rowOf(after.ledger, 'WR-04').severity, 'critical', 'the frontmatter still knows'); + assert.strictEqual(rowOf(after.ledger, 'WR-04').disposition, 'deferred', 'a bad severity cell does not cost the decision'); + const bothMangled = cellMangled.replace(/^( {2}- id: WR-04\n {4}severity: )critical$/m, '$1critcal'); + assert.notStrictEqual(bothMangled, cellMangled, 'the frontmatter mangling must have landed'); + after = runShippedDisposition({ reviewText: REVIEW_2_DROPPED, priorText: bothMangled }); + assert.strictEqual(rowOf(after.ledger, 'WR-04').severity, 'warning', 'nothing recorded is usable, so the prefix is all that is left'); + }); + + test('a pre-severity ledger (bare rows, no frontmatter) still infers from the prefix', () => { + // Back-compat control. A ledger written by hand with no severity cell to speak of is not + // rejected; it just has nothing to carry, so the prefix governs as it always did. + const bare = '| WR-04 | | deferred | by hand |\n'; + const after = runShippedDisposition({ reviewText: REVIEW_2_DROPPED, priorText: bare }); + const row = rowOf(after.ledger, 'WR-04'); + assert.strictEqual(row.severity, 'warning'); + assert.strictEqual(row.disposition, 'deferred', 'an empty severity cell does not cost the decision'); + }); +}); + +describe('#3861 round 2 — a finding the heading parser cannot match is SURFACED, not dropped (B4)', () => { + // Two independent parsers produce two numbers one paragraph apart: the counts come from + // REVIEW.md's frontmatter, the rows from `### <ID>:` heading matches against a CLOSED + // CR|BL|WR|IN alternation. Nothing reconciled them, so a finding the alternation cannot reach + // contributed no row, no note and no diagnostic -- and the ledger asserted `open: N of N` over + // a set strictly smaller than the console line had just reported. + const REVIEW_5 = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 2', ' info: 2', ' total: 5', '---', '', + '### CR-01: a conforming finding', + '### WR-01: another conforming one', + '### WR-02: a third', + '### SEC-01: a prefix the alternation does not carry', + '#### IN-09: a heading one level too deep'].join('\n'); + + test('the shortfall is stated in the ledger frontmatter and on the console', () => { + const out = runShippedDisposition({ reviewText: REVIEW_5, reviewTotal: 5 }); + assert.match(out.ledger, /^unparsed: 2$/m, + 'the ledger must record that two findings reached no row'); + assert.match(out.ledger, /^total: 3$/m, 'and must still report the rows it does have'); + assert.match(out.stdout, /2 finding\(s\) recorded NOWHERE/, + 'the console line must say so too -- the ledger is not the only surface a human reads'); + assert.match(out.stdout, /the review reports 5, but only 3 matched/, + 'and must name both numbers, so the shortfall is checkable rather than asserted'); + }); + + test('a review whose findings all parse gains no unparsed key at all', () => { + // Negative control for the key itself. An ordinary ledger must not grow a noise key, or the + // unchanged-run check starts rewriting the file on every phase. + const clean = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---', '', + '### CR-01: the only finding'].join('\n'); + const out = runShippedDisposition({ reviewText: clean, reviewTotal: 1 }); + assert.doesNotMatch(out.ledger, /^unparsed:/m, 'nothing was dropped, so nothing is reported'); + assert.doesNotMatch(out.stdout, /recorded NOWHERE/); + }); + + test('an absent or non-numeric total reconciles nothing rather than inventing a shortfall', () => { + // The reconciliation needs a number on BOTH sides. A legacy review with no findings: block + // has no total to compare against, and reporting `unparsed: 3` there would be a fabrication. + const legacy = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: a', '### WR-01: b'].join('\n'); + const out = runShippedDisposition({ reviewText: legacy, reviewTotal: '' }); + assert.doesNotMatch(out.ledger, /^unparsed:/m); + assert.match(out.ledger, /^total: 2$/m); + }); + + test('a total SMALLER than the rows is not reported as a negative shortfall', () => { + // Boundary in the other direction: the subtraction is clamped, so a review under-reporting + // its own total cannot produce `unparsed: -1`. + const out = runShippedDisposition({ + reviewText: ['---', 'phase: 01', 'status: issues_found', 'findings:', ' total: 1', '---', '', + '### CR-01: a', '### WR-01: b'].join('\n'), + reviewTotal: 1, + }); + assert.doesNotMatch(out.ledger, /^unparsed:/m, 'no shortfall when more parsed than declared'); + assert.doesNotMatch(out.ledger, /unparsed: -/); + }); + + test('a review NONE of whose findings parse still records the shortfall on a first run', () => { + // The corner every case above misses, and the one carrying the LEAST evidence anywhere else: + // every finding in a heading shape the alternation cannot reach, on a phase with no prior + // ledger and no fix report. `order` is empty, so the early exit that stands down for "nothing + // to record" fired BEFORE the shortfall was computed -- no ledger, no console line, no + // diagnostic, for a review that declared two Criticals. A PARTIAL shortfall always reported, + // which is exactly why the total one read as covered. The shortfall is now derived above that + // exit and both exits decline to fire while one is outstanding. + const allUnmatched = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 2', ' warning: 0', ' info: 0', ' total: 2', '---', '', + '## Critical Issues', '', + '### SEC-01: a prefix the alternation does not carry', + '### SEC-02: and a second one'].join('\n'); + const out = runShippedDisposition({ reviewText: allUnmatched, reviewTotal: 2 }); + assert.notStrictEqual(out.ledger, null, + 'a review whose findings NONE parsed must still leave a record — silence here is the exact silent drop this reconciliation exists to close'); + assert.match(out.ledger, /^unparsed: 2$/m, 'and must state how many findings reached no row'); + assert.match(out.ledger, /^total: 0$/m, 'while reporting honestly that it carries no rows'); + assert.match(out.stdout, /2 finding\(s\) recorded NOWHERE/, + 'the console line must say so too — the ledger is not the only surface a human reads'); + }); + + test('a fix report present, nothing parsed, no prior ledger — the SECOND exit records it too', () => { + // The two exits that discarded the shortfall are reached by DIFFERENT inputs, so one test + // cannot pin both. The earlier one stands down as soon as a fix report exists; this input + // therefore sails past it and lands on the later `rows.length === 0` return, which had the + // identical hole. Without this case, deleting the later guard's conjunct leaves the pair + // above green and the drop returns by the other road — a surviving mutant, not a covered one. + const allUnmatched = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 2', ' warning: 0', ' info: 0', ' total: 2', '---', '', + '### SEC-01: a prefix the alternation does not carry', + '### SEC-02: and a second one'].join('\n'); + // The fix report must contribute NO row, or it never reaches the later exit: an id the + // alternation CAN match becomes a carried row, rows.length is 1, and the guard under test is + // not the one that decides. Mutation-controlled — with a CR-NN id here the later guard's + // conjunct could be deleted and this test stayed green. + const fixReport = ['## Fixed Issues', '', '### SEC-03: an unmatched id in the fix report too'].join('\n'); + const out = runShippedDisposition({ reviewText: allUnmatched, fixText: fixReport, reviewTotal: 2 }); + assert.notStrictEqual(out.ledger, null, + 'a fix report does not excuse the drop — the shortfall is still two findings recorded nowhere'); + assert.match(out.ledger, /^unparsed: 2$/m, 'and the count must survive this path too'); + }); + + test('a genuinely clean review still writes nothing — the relaxed exit is scoped to a shortfall', () => { + // Negative control for the fix itself, and the reason it is scoped rather than removed: + // relaxing that exit unconditionally would grow a zero-row ledger on every clean phase. With + // `total: 0` there is no shortfall to outstand, so both exits still fire exactly as before. + const clean = ['---', 'phase: 01', 'status: clean', 'findings:', + ' critical: 0', ' warning: 0', ' info: 0', ' total: 0', '---', '', 'No issues.'].join('\n'); + const out = runShippedDisposition({ reviewText: clean, reviewTotal: 0 }); + assert.strictEqual(out.ledger, null, 'no findings declared and none parsed — nothing to record'); + assert.strictEqual(out.stdout.trim(), '', 'and nothing to say'); + }); + + test('a clean review with an unparseable fix report still writes nothing — the later exit must keep firing', () => { + // The fourth cell of the matrix, and the only one the three tests above leave open: they pin + // the LATER exit's under-fire (it must not swallow a shortfall) and the EARLIER exit's + // over-fire (a clean review must stay unrecorded), but nothing pins the later exit's OWN + // over-fire. It is reachable because a fix report — even one naming no matchable id — makes + // the earlier exit stand down, so control arrives at the later exit with a genuinely clean + // review and no shortfall. Neutering only that exit then writes a zero-row ledger reading + // "0 of 0 finding(s) open" for a phase that had nothing to report, and all three tests above + // stay green through it. + const clean = ['---', 'phase: 01', 'status: clean', 'findings:', + ' critical: 0', ' warning: 0', ' info: 0', ' total: 0', '---', '', 'No issues.'].join('\n'); + const unparseableFix = ['## Fixed Issues', '', '### SEC-03: an unmatched id in the fix report'].join('\n'); + const out = runShippedDisposition({ reviewText: clean, fixText: unparseableFix, reviewTotal: 0 }); + assert.strictEqual(out.ledger, null, + 'a fix report that decides nothing is not a reason to record a ledger for a clean review'); + assert.strictEqual(out.stdout.trim(), '', 'and nothing to say about it'); + }); +}); + +describe('#3829 review round 2 — a review that reports nothing still reconciles the ledger', () => { + const EMPTY_REVIEW = ['---', 'phase: 01', 'status: issues_found', '---', '', 'no findings'].join('\n'); + + test('a decided row is carried when the review reports no findings at all', () => { + // Exiting early on an empty review would freeze a stale ledger showing findings as open that + // the review no longer reports — the opposite of what this record exists to do. + const priorText = [ + '| CR-01 | critical | deferred | ships next phase |', + '| WR-01 | warning | open | - |', + ].join('\n'); + const d = buildDisposition({ reviewText: EMPTY_REVIEW, priorText, padded: '01' }); + // M1: BOTH rows survive. The decided one because it was triaged, the untriaged one because + // it was not -- and an untriaged finding vanishing without trace is the defect #3829 names. + assert.strictEqual(d.total, 2, 'both rows survive; neither is silently deleted'); + assert.deepStrictEqual(d.rows.map((r) => r.id), ['CR-01', 'WR-01']); + assert.strictEqual(d.rows[0].disposition, 'deferred'); + assert.strictEqual(d.rows[0].carried, true); + assert.strictEqual(d.rows[1].disposition, 'open'); + assert.strictEqual(d.rows[1].carried, true); + assert.strictEqual(d.open, 1, 'the untriaged carried row is still open, and the count says so'); + }); + + test('a review with no findings and no prior ledger produces nothing at all', () => { + assert.strictEqual(buildDisposition({ reviewText: EMPTY_REVIEW, padded: '01' }), null); + }); + + // Ninth src.includes() retired (round 3): it pinned the guard's exact line, including the + // process.exit(0) the script no longer calls, so it went red on a change that left the property + // untouched. The property -- an empty review still reconciles an EXISTING ledger -- is asserted + // behaviourally by the round-2 describe '#3829 review round 2 — a review that reports nothing + // still reconciles the ledger' below, which drives the shipped script against a prior ledger. + +}); + +// --------------------------------------------------------------------------- +// #3829 — the SHIPPED disposition script, driven. These tests execute the real +// embedded script (see shippedDispositionScript), so a regression in the +// workflow file turns them red. The mirror-based tests above are kept for the +// pure parsing shapes; these are the ones that hold the contract. +// --------------------------------------------------------------------------- +describe('#3829 — shipped disposition script (executed, not mirrored)', () => { + const REVIEW = ['---', 'phase: 01', 'status: issues_found', '---', '', + '### CR-01: a', '', '### WR-01: b', '', '### IN-01: c'].join('\n'); + + test('every finding is recorded, defaulting to open, at the right severity', () => { + const rows = ledgerRows(runShippedDisposition({ reviewText: REVIEW }).ledger); + assert.deepStrictEqual(rows.map((r) => [r.id, r.severity, r.disposition]), [ + ['CR-01', 'critical', 'open'], ['WR-01', 'warning', 'open'], ['IN-01', 'info', 'open'], + ]); + }); + + test('--fix outcomes are reconciled, and a lookalike section heading is not one', () => { + const fixText = ['## Fixed Issues', '', '### WR-01: b', '', + '## Fixed Issues Verification', '', '### IN-01: c'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: REVIEW, fixText }).ledger); + const by = Object.fromEntries(rows.map((r) => [r.id, r])); + assert.strictEqual(by['WR-01'].disposition, 'fixed'); + assert.strictEqual(by['WR-01'].source, '01-REVIEW-FIX.md'); + assert.strictEqual(by['IN-01'].disposition, 'open', 'a lookalike heading must classify nothing'); + }); + + test("a human's deferral reason survives, and the carried marker never accumulates", () => { + // Render -> re-parse -> render, four times, exactly as consecutive phase runs would. + const shrunk = ['---', 'status: issues_found', '---', '', '### WR-01: b'].join('\n'); + let prior = ['| Finding | Severity | Disposition | Source |', + '|---------|----------|-------------|--------|', + '| CR-01 | critical | deferred | ships next phase, see ADR-99 |'].join('\n'); + let ledger; + for (let i = 0; i < 4; i++) { + ledger = runShippedDisposition({ reviewText: shrunk, priorText: prior }).ledger; + prior = ledger; + } + const cr = ledgerRows(ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred'); + assert.strictEqual(cr.source, 'ships next phase, see ADR-99 (not in the current review)'); + assert.strictEqual( + (ledger.match(/\(not in the current review\)/g) || []).length, 1, + 'the carried marker must appear exactly once however many times the gate runs' + ); + }); + + test('the ledger is a fixed point: a run that changes nothing rewrites nothing', () => { + const first = runShippedDisposition({ reviewText: REVIEW }).ledger; + const again = runShippedDisposition({ reviewText: REVIEW, priorText: first }); + assert.strictEqual(again.wroteNothing, true, 'an unchanged run must report unchanged'); + assert.strictEqual(again.ledger, first, 'and must leave the bytes alone'); + }); + + test('a finding the review no longer reports is carried whether or not it was triaged', () => { + const prior = ['| CR-01 | critical | deferred | later |', '| IN-01 | info | open | - |'].join('\n'); + const shrunk = ['---', 'status: issues_found', '---', '', '### WR-01: b'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: shrunk, priorText: prior }).ledger); + const ids = rows.map((r) => r.id); + assert.ok(ids.includes('CR-01'), 'a deferred finding must survive leaving the review'); + assert.ok(ids.includes('IN-01'), 'and so must an untriaged one -- that is the whole record'); + }); + + test('a renumbered finding leaves its old row behind, marked, rather than vanishing', () => { + // The concrete case M1 names: run 1 records CR-01 open, the re-review renumbers it to CR-02, + // and run 2's ledger used to contain neither. The old row is now carried and marked, so the + // double-count is legible rather than a silent delete -- the stated cost of the fix. + const prior = '| CR-01 | critical | open | - |'; + const renumbered = ['---', 'status: issues_found', '---', '', '### CR-02: the same finding, renumbered'].join('\n'); + const ledger = runShippedDisposition({ reviewText: renumbered, priorText: prior }).ledger; + const rows = ledgerRows(ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-02', 'CR-01']); + assert.match(ledger, /\| CR-01 \|.*\(not in the current review\) \|/, + 'the carried row is MARKED, so it does not claim the finding is live'); + }); + + test('a review reporting nothing still reconciles an existing ledger', () => { + const empty = ['---', 'status: issues_found', '---', '', 'no findings'].join('\n'); + const prior = ['| CR-01 | critical | deferred | later |', '| WR-01 | warning | open | - |'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: empty, priorText: prior }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01', 'WR-01'], 'reconciled, not truncated'); + }); + + test('a review reporting nothing with no prior ledger writes no ledger at all', () => { + const empty = ['---', 'status: issues_found', '---', '', 'no findings'].join('\n'); + assert.strictEqual(runShippedDisposition({ reviewText: empty }).ledger, null); + }); + + test('an applied outcome outranks an earlier deferral', () => { + const prior = '| CR-01 | critical | deferred | later |'; + const fixText = ['## Fixed Issues', '', '### CR-01: a'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: REVIEW, priorText: prior, fixText }).ledger); + assert.strictEqual(rows.find((r) => r.id === 'CR-01').disposition, 'fixed'); + }); + + test('a deferral on a finding STILL in the review is not reset to open', () => { + // Distinct from the carried case: this finding is present in the current review, so it takes + // the ordinary path. 'open' must never overwrite a recorded decision on that path either. + const prior = '| CR-01 | critical | deferred | ships next phase |'; + const rows = ledgerRows(runShippedDisposition({ reviewText: REVIEW, priorText: prior }).ledger); + const cr = rows.find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred'); + assert.strictEqual(cr.source, 'ships next phase'); + assert.ok(!/not in the current review/.test(cr.source), 'it is present, so it is not carried'); + }); + + test('BL- is recorded at the Critical tier, as the documented CR- equivalent', () => { + const withBlocker = ['---', 'status: issues_found', '---', '', + '### BL-01: a', '', '### WR-01: b'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: withBlocker }).ledger); + assert.deepStrictEqual(rows.map((r) => [r.id, r.severity]), [ + ['BL-01', 'critical'], ['WR-01', 'warning'], + ]); + }); + + test('a finding id whose prefix is a JS object property name produces no row', () => { + const hostile = ['---', 'status: issues_found', '---', '', + '### constructor-01: x', '', '### __proto__-02: y', '', '### CR-01: real'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: hostile }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01']); + }); +}); + +describe('#3829 review round 3 — hostile frontmatter and hand-edited ledgers', () => { + test('an UNTERMINATED frontmatter block yields no values, not the whole review body', () => { + // Stopping at "the next ---" is not enough: with no closing delimiter the scan would run to + // EOF and hand body text to every read, undoing the scoping fix entirely. + const unterminated = [ + '---', 'phase: 01', 'status: issues_found', '', + '# body', '', 'total: 777', 'critical: 66', + ].join('\n'); + const parsed = parseGateCounts(unterminated); + assert.deepStrictEqual(parsed, { status: '', critical: '', warning: '', info: '', total: '' }); + }); + + test('docs-parity: the frontmatter scan emits only when the closing delimiter was seen', () => { + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.ok( + src.includes('END{if (closed) printf "%s", buf}'), + 'an unterminated frontmatter block must yield nothing' + ); + }); + + + test('a hand-edited row missing its trailing pipe still preserves the decision', () => { + // A mangled table is already broken; silently dropping the row would lose a deferral, which + // is the exact class of loss this record exists to prevent. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a'].join('\n'); + const prior = '| CR-01 | critical | deferred | see issue 42'; + const rows = ledgerRows(runShippedDisposition({ reviewText: review, priorText: prior }).ledger); + const cr = rows.find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred'); + assert.strictEqual(cr.source, 'see issue 42'); + }); +}); + +describe('#3829 review round 3 — stale fix reports, fenced examples, hostile REVIEW.md', () => { + test('a STALE fix report does not mark a new finding of the same id as fixed', () => { + // Finding ids are reused across re-reviews. Matching on the id alone would let a fix report + // from an earlier review declare a brand-new CR-01 already fixed — the worst possible lie + // for a record whose whole job is saying what happened to a finding. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: NEW authentication bypass'].join('\n'); + const fixText = ['## Fixed Issues', '', '### CR-01: OLD null dereference'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review, fixText }).ledger); + assert.strictEqual(rows[0].disposition, 'open'); + }); + + test('a fix report naming the same finding still applies', () => { + const review = ['---', 'status: issues_found', '---', '', '### CR-01: same title'].join('\n'); + const fixText = ['## Fixed Issues', '', '### CR-01: same title'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review, fixText }).ledger); + assert.strictEqual(rows[0].disposition, 'fixed'); + }); + + test('a finding heading inside a fenced block is an example, not a finding', () => { + const review = ['---', 'status: issues_found', '---', '', '### CR-01: real', + '', '```', '### CR-77: an illustration', '```'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01']); + }); + + test('an id listed under BOTH Fixed and Skipped is not decided by row order', () => { + const review = ['---', 'status: issues_found', '---', '', '### WR-01: dup'].join('\n'); + const fixText = ['## Fixed Issues', '', '### WR-01: dup', '', + '## Skipped Issues', '', '### WR-01: dup'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review, fixText }).ledger); + assert.strictEqual(rows[0].disposition, 'fixed', 'first occurrence wins, deterministically'); + }); + + test('an escaped pipe in a deferral reason survives whole', () => { + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a'].join('\n'); + const prior = '| CR-01 | critical | deferred | wait \\| see ADR-9 |'; + const rows = ledgerRows(runShippedDisposition({ reviewText: review, priorText: prior }).ledger); + assert.strictEqual(rows[0].source, 'wait \\| see ADR-9'); + }); + + // #3861 round 3. The Source cell is the one field this ledger asks a human to hand-edit, and + // "waiting on team A | team B to align" is an ordinary thing to type there. Under the previous + // capture a bare | failed the WHOLE prior-row match: prior.get() was undefined, the row fell + // through to open with an empty Source, and the console line read "1 of 1 finding(s) open" -- + // a Critical a human explicitly deferred, with a documented reason, rendered indistinguishable + // from one never triaged, and the reason gone. Exactly the "was this ever seen" ambiguity + // #3829 exists to remove, reachable by one missing backslash. The cell is the LAST column, so + // it is now captured to the end of the line and a bare pipe is prose; the render escapes it + // so the table stays a table, and the second run converges. + test('a bare pipe in a deferral reason is kept as prose: the decision and the reason survive', () => { + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a'].join('\n'); + const prior = '| CR-01 | critical | deferred | waiting on team A | team B to align |'; + const first = runShippedDisposition({ reviewText: review, priorText: prior }); + const rows = ledgerRows(first.ledger); + assert.strictEqual(rows[0].disposition, 'deferred', 'a bare pipe must not revert the decision'); + assert.strictEqual(rows[0].source, 'waiting on team A \\| team B to align', 'the reason survives, escaped'); + assert.match(first.ledger, /^open: 0$/m, 'and the headline count agrees with the row'); + // Fixed point: the escaped form re-parses to itself, so the second run rewrites nothing. + const second = runShippedDisposition({ reviewText: review, priorText: first.ledger }); + assert.strictEqual(ledgerRows(second.ledger)[0].source, 'waiting on team A \\| team B to align'); + assert.match(second.stdout, /disposition unchanged/, 'the second run must converge'); + }); + + // Round-3 adversarial pass on the fix above. The first escape used /(^|[^\\])\|/g, which CONSUMES + // the character before the pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> + // A\|\|B, a third run to converge, breaking the advertised second-run fixed point), and an escaped + // backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare in the + // rendered table. The generator then emitted at most one bare pipe, so no property reached either; it does now. + // Round-3 adversarial pass. The script printed its verdict and then called process.exit(0) on the + // 'unchanged' branch only. Node documents process.stdout writes to pipes as ASYNCHRONOUS on + // POSIX, and an explicit exit can pre-empt a pending write, so on those lanes the caller can see + // exit 0 with no verdict line -- a hardening, not a reproduced defect: the reviewer's empty-stdout + // observation turned out to be its own sandbox (a bare console.log child printed nothing there + // either), which is stated so nobody re-reads this as evidence the drop was seen. The script now + // runs inside main() and leaves by return, so the loop drains stdout before exit. + // Shape-pinned deliberately: the property IS the absence of the call, and a behavioural test would + // have to race a pipe to fail. Comments are stripped first so a mention is not a match, and the + // match covers the dotted, bracketed and whitespace-split spellings; a call built by any other + // indirection is outside this pin and is what code review is for. + test('the shipped script never calls process.exit -- it returns, so its verdict line is never lost', () => { + const code = shippedDispositionScript().replace(/^\s*\/\/.*$/gm, ''); + assert.doesNotMatch(code, /process\s*(?:\.\s*exit\b|\[\s*['"]exit['"]\s*\])/, + 'leave main() by return; an explicit exit can drop the verdict line on a piped stdout'); + }); + + test('adjacent bare pipes and a backslash-then-pipe are escaped in ONE write, then converge', () => { + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a'].join('\n'); + const prior = '| CR-01 | critical | deferred | A||B and C\\\\|D |'; + const first = runShippedDisposition({ reviewText: review, priorText: prior }); + assert.strictEqual(ledgerRows(first.ledger)[0].source, 'A\\|\\|B and C\\\\\\|D', + 'every bare pipe is escaped on the first write, whatever precedes it'); + const second = runShippedDisposition({ reviewText: review, priorText: first.ledger }); + assert.strictEqual(second.ledger, first.ledger, 'and the escaped form is a fixed point'); + assert.match(second.stdout, /disposition unchanged/); + }); + +}); + // --------------------------------------------------------------------------- // #4665 — `--fix` must be able to act on an existing REVIEW.md even when the // incremental scope for a FRESH review is empty. check_empty_scope used to @@ -1649,3 +2893,2104 @@ describe('#4665 — check_empty_scope --fix recovery onto an existing REVIEW.md' assert.match(block, /Exit workflow/i, 'the non-recovery path must still end the workflow'); }); }); + +// --------------------------------------------------------------------------- +// #3861 round 1 — the two blockers, and the structural properties that catch them +// +// Both were invisible to every behavioural test above, for the same reason: those +// tests execute the node script through the process seam with the environment +// handed to it, so they never see the SHELL that is supposed to build that +// environment, nor the prose that tells the agent what to read. +// --------------------------------------------------------------------------- + +// Every ```bash fence in a step file, in order. +function bashFences(src) { + const out = []; + const lines = src.replace(/\r\n/g, '\n').split('\n'); + let start = -1; + for (let i = 0; i < lines.length; i++) { + if (start === -1 && /^```bash\s*$/.test(lines[i])) { start = i + 1; continue; } + if (start !== -1 && /^```\s*$/.test(lines[i])) { out.push(lines.slice(start, i).join('\n')); start = -1; } + } + return out; +} + +describe('#3861 round 1 — step-file structural contract', () => { + test('the step file never instructs the agent to read and execute ITSELF', () => { + // A step file that names its own path as something to "read and execute" is + // unbounded self-recursion at runtime, and nothing downstream bounds it. + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + const self = 'execute-phase/steps/' + path.basename(DISPOSITION_STEP_PATH); + assert.ok( + !src.includes(self), + 'code-review-disposition.md must not reference its own path — execute-phase.md is what points here' + ); + }); + + test('the disposition instruction appears exactly once', () => { + // The self-reference above arrived as a duplicated paragraph; the duplicate is + // the tell, and no linter scores markdown prose. + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + const n = src.split(/\r?\n/).filter((l) => l.startsWith('**Record a per-finding disposition.**')).length; + assert.strictEqual(n, 1, 'exactly one disposition instruction, not a pointer plus a body'); + }); + + test('every shell block derives the variables it reads — blocks do not share a shell', () => { + // Each fenced block is dispatched as its own Bash call, so a variable derived in + // block 1 is EMPTY in block 2 — and empty is silent: the write lands on a bare + // `-REVIEW-DISPOSITION.md` and the step still reports success. The step's own + // inputs (PHASE_DIR, PHASE_NUMBER) are the only values a block may inherit. + const INPUTS = new Set(['PHASE_DIR', 'PHASE_NUMBER']); + const DERIVED = ['PADDED', 'REVIEW_FILE', 'DISPOSITION_FILE']; + const fences = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8')); + assert.ok(fences.length >= 2, 'the step must still carry more than one shell block'); + for (const [i, fence] of fences.entries()) { + for (const v of DERIVED) { + if (INPUTS.has(v)) continue; + // The derivation may be INDENTED or sit inside a `case` arm -- PADDED is derived by a + // case split so a dotted phase number does not reach `printf %02d`. Anchoring the + // detector at column 0 made a real derivation invisible and the guard fired on it. The + // property being checked is unchanged: the block must ASSIGN what it reads. + // A SELF-REFERENTIAL assignment is a PASS-THROUGH, not a derivation. Block 2 prefixes its + // `node -e` with `REVIEW_FILE="${REVIEW_FILE}" DISPOSITION_FILE="${DISPOSITION_FILE}" ...` + // to put them in the child's environment, and the detector counted that as deriving them. + // So deleting block 2's REAL derivation left this guard GREEN -- on the exact defect it + // was written for. Found by negative-controlling the guard rather than by reading it. + // (Pre-existing: the original column-0 anchor matched that same line, which is at column 0.) + const fenceNoPassthrough = fence.replace( + new RegExp('(?:^|[\\s;])' + v + '="\\$\\{' + v + '\\}"', 'gm'), ' '); + // Non-empty RHS required. `REVIEW_FILE=` is an assignment TOKEN and not a derivation, and + // the reviewer evaded the earlier predicate with exactly that. This is a cheap fast-fail, + // NOT the guard's authority -- see the executed guard below, which is. + const assigned = new RegExp('(?:^|[\\s;])' + v + '=(?![\\s;#]|$)', 'm'); + const reads = new RegExp('\\$\\{?' + v + '\\b').test(fence.replace(new RegExp('(?:^|[\\s;])' + v + '=', 'gm'), '')); + if (!reads) continue; + assert.ok( + assigned.test(fenceNoPassthrough), + 'block ' + (i + 1) + ' reads ' + v + ' without deriving it — it is empty in a fresh shell' + ); + } + } + }); +}); + + +describe('#3861 round 16 — the embedded script fits a Windows command line', () => { + // Block 2 runs the record-builder as `node -e "<script>"`, so the whole script is ONE + // argv entry. Windows caps a command line at 32767 characters (CreateProcess), and Node + // surfaces the overflow as ENAMETOOLONG from spawn — the process never starts. Linux's + // ~2 MB ARG_MAX means the Linux lane cannot see this at all: when the script crossed the + // cap it stayed green on ubuntu and took out 83 tests on windows-latest in one push, + // every one of them reporting `spawn_failed` rather than anything about length. + // + // Measured on native Windows (node v25.2.1): the largest `-e` argument that still spawns + // is 32728 characters; 32729 fails. The budget below sits well under that so the next + // addition to the script has somewhere to go — long rationale belongs in the step file's + // prose, which costs the command line nothing. + // + // What this guard does NOT establish, stated so it is not read as more than it is. It counts + // the extracted JavaScript, not the command line Windows finally serializes: the 32767 cap + // applies to the whole line — executable path, quoting and backslash escaping included — so a + // quote-heavy payload expands on the way out, and the 32728 figure above is one host's + // threshold rather than the CI runner's. The budget is a practical margin with room for both + // effects, not a proof that everything it admits will spawn. + const WINDOWS_CMDLINE_CAP = 32767; + const CMDLINE_BUDGET = 24576; // 24 KiB — the cap less ~8 KiB of deliberate headroom + + test('the extracted node -e script stays well under the Windows command-line cap', () => { + const script = shippedDispositionScript(); + // Anti-vacuity: a length assertion alone passes when the extractor returns '' — which is + // exactly what a moved fence or a renamed delimiter would produce. Bound it from BELOW + // first, so a broken extractor fails here instead of reporting a comfortable 0 bytes. + assert.ok( + script.length > 4096, + 'the extractor returned ' + script.length + ' bytes — it is not reading the shipped script' + ); + assert.ok( + script.length <= CMDLINE_BUDGET, + 'the embedded node -e script is ' + script.length + ' characters; the budget is ' + + CMDLINE_BUDGET + ' and the hard Windows limit is ' + WINDOWS_CMDLINE_CAP + + '. Move long rationale out of the script and into the step file prose — it reads the ' + + 'same there and costs the command line nothing.' + ); + }); +}); +describe('#3861 round 2 — the shell-sharing guard, EXECUTED', () => { + // The textual guard above is a fast-fail, not the authority. An adversarial pass evaded it + // three ways -- an empty `REVIEW_FILE=`, a self-referential `REVIEW_FILE=$REVIEW_FILE`, and a + // commented assignment -- because a structural predicate recognises assignment TOKENS, never + // assignments that derive a usable value. No amount of regex fixes that class. + // + // So the authority moves to execution, which is the lesson this PR has now learned three times: + // run the real second fence in a FRESH shell with nothing but the step's two declared inputs, + // and require it to write the ledger at the correct derived path. Every derivation in that + // fence is load-bearing for that outcome, so no textual dodge survives it. + // RANDOM PHASES, and the randomness is the mechanism rather than decoration. Fixed fixtures + // cannot establish derivation: an adversarial pass defeated the one-phase version with + // `case ... in 7) PADDED=07 ;; *) PADDED=07 ;; esac`, and then defeated the two-phase version + // by simply adding `3.1) PADDED=03.1` to the same case. Any finite sample loses that race -- + // the enumeration just grows to cover whatever the test happens to name. + // + // A phase picked at RUN TIME raises the cost of a hardcode from two arms to the whole drawn + // domain, so an accidental loss of derivation fails on some run rather than never. + // + // STATED HONESTLY, because the first version of this comment overclaimed and was refuted: the + // domain is FINITE -- 88 integer values and 792 dotted ones -- so a mutation enumerating all + // 880 passes forever, and one covering 90 or 12345678.1 or 1.10 would still break real phases + // the draw cannot reach. This raises the bar; it does not prove derivation, and nothing short + // of reading the fence can. `Math.random()` is also unseeded, so a failure is reproducible only + // in the sense that the drawn values are printed in every assertion message below -- re-running + // draws different ones. That is the honest description of what this buys. + // + // One integer and one dotted phase per run: the dotted one additionally pins the integer-part + // split a naive `%02d` cannot express. + const rnd = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1)); + const intPhase = String(rnd(2, 89)); + const dotPhase = rnd(2, 89) + '.' + rnd(1, 9); + const pad = (v) => { + const [i, sub] = String(v).split('.'); + return String(Number(i)).padStart(2, '0') + (sub === undefined ? '' : '.' + sub); + }; + for (const [phaseNumber, padded] of [[intPhase, pad(intPhase)], [dotPhase, pad(dotPhase)]]) { + test('block 2 derives its own paths and writes the ledger for phase ' + phaseNumber, + { skip: !HAS_BASH }, () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-blk2-')); + const drawn = ' (drawn phase ' + phaseNumber + ' -> ' + padded + ')'; + try { + fs.writeFileSync(path.join(dir, padded + '-REVIEW.md'), + ['---', 'phase: ' + padded, 'status: issues_found', '---', '', + '### CR-01: a finding'].join('\n')); + const fence = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'))[1]; + assert.ok(fence && fence.includes('DISPOSITION_FILE'), 'the disposition block must be fence 2'); + // Scrub the DERIVED names from the inherited environment, so the claim "given only the + // declared inputs" is true rather than merely intended. `...process.env` is still needed + // for PATH/HOME, and an adversarial pass was right to call out the earlier wording. + const env = { ...process.env, PHASE_DIR: dir, PHASE_NUMBER: phaseNumber, RUNTIME_DIR: ROOT }; + for (const k of ['PADDED', 'REVIEW_FILE', 'DISPOSITION_FILE', 'FIX_REPORT_FILE']) delete env[k]; + const res = runHook('-c', ['set -euo pipefail\n' + fence + '\n'], { + interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS, env, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the block must run to completion'); + assert.strictEqual(res.exitCode, 0, 'advisory: it must not abort' + drawn + ': ' + res.stderr); + assert.ok(fs.existsSync(path.join(dir, padded + '-REVIEW-DISPOSITION.md')), + 'the ledger must land at the DERIVED path — a lost derivation writes elsewhere or nowhere' + drawn); + const rows = ledgerRows(fs.readFileSync(path.join(dir, padded + '-REVIEW-DISPOSITION.md'), 'utf8')); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01'], 'and it must have read the review' + drawn); + assert.ok(!fs.existsSync(path.join(dir, '-REVIEW-DISPOSITION.md')), + 'an empty PADDED must never produce a bare-named ledger' + drawn); + } finally { + cleanup(dir); + } + }); + } +}); + +// --------------------------------------------------------------------------- +// #3861 round 1, Major 4 + Minor 9 — the frontmatter reads, EXECUTED +// +// The disposition builder stopped being a mirror three rounds ago, for a reason +// this file already states: a hand model of a shell-embedded script drifts, and +// when it drifts the tests pass while the shipped block is broken. The counts +// parser kept its mirror anyway, and the argument against mirrors does not stop +// applying at the boundary between the two blocks. +// +// So the mirror loses its authority: it is now asserted AGAINST the shipped awk +// and greps, run under `set -euo pipefail` in a real shell, over every fixture +// the mirror is tested on. Divergence in either direction fails. +// +// Running the block also makes its advisory guards behavioural rather than +// textual, which retires the two `src.includes()` docs-parity assertions that +// stood in for them — Minor 9's anti-pattern, reduced by executing the thing +// the assertions were describing. +// --------------------------------------------------------------------------- + + +// Run the step's FIRST shell block — the frontmatter reads — and report WHAT IT PRINTS. +// +// This harness used to append its own `printf` of the six internal variables to the fence before +// running it, and every assertion below then read those six lines. That is a test manufacturing the +// observable it asserts on: the shipped fence emitted nothing, the tested fence emitted six lines +// because the test added them, and the whole group was green against a script that did not exist +// outside this process. It is why the "the gate reports the counts" defect shipped past a suite that +// looks like it covers exactly that surface — the green was structurally incapable of turning red +// for it. The emitter now lives in the fence (see the step file), so the harness reads the fence's +// own stdout and nothing is synthesized here. +// `plantDir` stands a DIRECTORY at the review path. It is the root-immune half of the +// unreadable-review contract: the fence's guard is `[ -f "$REVIEW_FILE" ] && [ -r ... ]`, +// and only the `-r` leg is defeated by running as root. A directory fails the `-f` leg on +// every lane and euid, so it reaches the same non-reporting arm without depending on +// permission bits at all. See the two tests at the end of this describe. +function runShippedGateCounts({ reviewText, padded = '01', writeReview = true, mode, phaseNumber, plantDir = false, extraEnv = {} }) { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-')); + try { + const reviewPath = path.join(dir, padded + '-REVIEW.md'); + if (plantDir) fs.mkdirSync(reviewPath); + else if (writeReview) fs.writeFileSync(reviewPath, reviewText); + if (mode !== undefined) fs.chmodSync(reviewPath, mode); + const fence = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'))[0]; + assert.ok(fence && fence.includes('REVIEW_COUNTS_OK'), 'the counts block must still be the first fence'); + // `set -euo pipefail` is the point, not decoration: the step is advisory, so a + // non-matching grep or an unreadable review must not take the block down. + const script = 'set -euo pipefail\n' + fence + '\n'; + const res = runHook('-c', [script], { + interpreter: 'bash', + timeoutMs: PROBE_TIMEOUT_MS, + // `Number(padded)` is deliberate for the integer case (strips the leading zero the way a + // caller's parsed phase_number does) but must not mangle a DOTTED phase: Number('03.1') is + // 3.1, which is what we want, while a non-numeric padded value would become NaN. + // `phaseNumber` overrides the derivation so a test can drive a value `padded` cannot + // express -- an unusable one. Otherwise PHASE_NUMBER is derived from `padded` as a caller's + // parsed phase_number would be. + env: { + ...process.env, + PHASE_DIR: dir, + PHASE_NUMBER: phaseNumber === undefined ? String(Number(padded)) : phaseNumber, + // `extraEnv` exists for ONE property: the reads pin LC_ALL=C, and a test that cannot set + // the ambient locale cannot prove the pin is doing anything. + ...extraEnv, + }, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the counts block must run to completion'); + return { exitCode: res.exitCode, stdout: res.stdout, stderr: res.stderr }; + } finally { + cleanup(dir); + } +} + +// Read the shipped message back into the facts it asserts. This parses the OBSERVABLE the operator +// sees — it does not reach into the fence — so a value the gate declines to report is `reported: +// false` here rather than a number this helper invented. +function readGateMessage(stdout) { + const full = /^Code review: (\d+) findings — (\d+) critical, (\d+) warning, (\d+) info\.$/m.exec(stdout); + if (full) { + return { reported: true, countsOk: '1', total: full[1], critical: full[2], warning: full[3], info: full[4] }; + } + if (/^Code review found issues\.$/m.test(stdout)) return { reported: true, countsOk: '0' }; + // The third arm (round 11): the file was READ and yielded no status. Reported -- the operator is + // told something -- but no counts and no verdict, which is what `unparsed` records. + if (/^Code review status unparsed: /m.test(stdout)) return { reported: true, countsOk: '0', unparsed: true }; + return { reported: false }; +} + +// The mirror's other half: render the two arms exactly as the shipped fence does, from the mirror's +// own parsed counts. Parity is asserted over this WHOLE STRING rather than over five intermediate +// values, which is what makes the assertion bind to something a user can see. The countsOk gate is +// modelled here because the shipped fence gates on it — digit-only, at most 8 digits, and the three +// severities must sum to the total. +function renderGateMessage(counts, phaseNumber) { + const numeric = (v) => v !== '' && /^[0-9]+$/.test(v) && v.length <= 8; + const all = [counts.total, counts.critical, counts.warning, counts.info]; + let ok = all.every(numeric); + if (ok) { + const n = (v) => parseInt(v, 10); + if (n(counts.critical) + n(counts.warning) + n(counts.info) !== n(counts.total)) ok = false; + } + // A mirror is always handed a TEXT, so the file was read by construction: an empty status here + // is the read-but-unparseable arm, never the absent one (round 11 Minor -- a malformed report + // must not read as clean). The absent arm is driven directly, not through the mirror. + if (counts.status === '') { + return 'Code review status unparsed: REVIEW.md is present but its frontmatter has no parseable status; severity counts unavailable.\n'; + } + if (counts.status === 'clean' || counts.status === 'skipped') return ''; + const head = ok + ? `Code review: ${counts.total} findings — ${counts.critical} critical, ${counts.warning} warning, ${counts.info} info.` + : 'Code review found issues.'; + return head + '\n' + `Consider running: /gsd:code-review ${phaseNumber} --fix` + '\n'; +} + +describe('#3861 round 2 — the ledger write refuses a non-regular file', () => { + // The write-safety behaviour shipped with NO regression control at all -- I hand-drove it and + // did not pin it, which the round review caught by grepping for the words. Three shapes, each + // a distinct failure mode, all driven before this test existed: + // symlink -> writeFileSync FOLLOWS it and replaced the target's contents, outside the + // phase directory, leaving the link intact so nothing looked wrong; + // FIFO -> readFileSync BLOCKED FOREVER, in a gate documented as never blocking; + // directory -> the write throws. + const mkPhase = () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-nrf-')); + fs.writeFileSync(path.join(dir, '03-REVIEW.md'), + ['---', 'status: issues_found', '---', '', '### CR-01: a finding'].join('\n')); + return dir; + }; + const runAt = (dir) => runNode(['-e', shippedDispositionScript()], { + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + REVIEW_FILE: path.join(dir, '03-REVIEW.md'), + DISPOSITION_FILE: path.join(dir, '03-REVIEW-DISPOSITION.md'), + FIX_REPORT_FILE: path.join(dir, '03-REVIEW-FIX.md'), + PADDED: '03', + }, + }); + + // Skipped on win32 per the repo's existing convention for symlink-planting tests: + // tests/settings-jsonc.test.cjs:389 skips the same class, and + // tests/unreachable-guard-drift.test.cjs:726 records why -- symlink creation requires elevated + // privileges on Windows CI. It happened to be available on the lane this round; the convention + // exists because it is not guaranteed. + test('a symlink at the ledger path is refused, and its target is untouched', + { skip: process.platform === 'win32' }, () => { + const dir = mkPhase(); + try { + const outside = path.join(dir, 'outside.txt'); + fs.writeFileSync(outside, 'ORIGINAL'); + fs.symlinkSync(outside, path.join(dir, '03-REVIEW-DISPOSITION.md')); + const res = runAt(dir); + assert.strictEqual(res.exitCode, 0, 'advisory: it refuses, it does not fail'); + assert.match(res.stdout, /not a regular file/, 'and says why'); + assert.strictEqual(fs.readFileSync(outside, 'utf8'), 'ORIGINAL', + 'the symlink target must not be overwritten'); + assert.ok(fs.lstatSync(path.join(dir, '03-REVIEW-DISPOSITION.md')).isSymbolicLink(), + 'and the link itself is left alone'); + } finally { cleanup(dir); } + }); + + test('a symlink whose target ALREADY matches is refused too — the fast path does not bypass it', + { skip: process.platform === 'win32' }, () => { + // The unchanged-run fast path read the ledger before the check, so a link whose target + // happened to match slipped through reporting `unchanged`. The check is first now. + const dir = mkPhase(); + try { + const outside = path.join(dir, 'outside.txt'); + fs.writeFileSync(outside, 'ORIGINAL'); + fs.symlinkSync(outside, path.join(dir, '03-REVIEW-DISPOSITION.md')); + runAt(dir); + const res = runAt(dir); + assert.match(res.stdout, /not a regular file/); + assert.doesNotMatch(res.stdout, /unchanged/, 'the fast path must not run ahead of the check'); + assert.strictEqual(fs.readFileSync(outside, 'utf8'), 'ORIGINAL'); + } finally { cleanup(dir); } + }); + + test('a FIFO at the ledger path is refused rather than blocking the phase forever', () => { + const dir = mkPhase(); + try { + // Through the process seam, like every other spawn in this file. + const fifoPath = path.join(dir, '03-REVIEW-DISPOSITION.md'); + runHook('-c', ['mkfifo "$1"', '_', fifoPath], { interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS }); + // GATE ON WHAT WAS CREATED, never on mkfifo's exit code. On the Windows lane mkfifo EXISTS + // and exits 0 while producing something that is not a FIFO, so an exit-code guard let this + // test run against an ordinary path: the ledger wrote normally and the assertion below + // failed for a reason that had nothing to do with the behaviour under test. Caught by CI, + // not by the local suite or five review passes -- every one of which ran on Linux. + let isFifo = false; + try { isFifo = fs.lstatSync(fifoPath).isFIFO(); } catch { isFifo = false; } + if (!isFifo) return; // no real FIFO on this platform; nothing to assert + const res = runAt(dir); + assert.strictEqual(res.outcome, OUTCOME.EXITED, + 'a FIFO must not hang the step — readFileSync blocks on one forever'); + assert.strictEqual(res.exitCode, 0); + assert.match(res.stdout, /not a regular file/); + } finally { cleanup(dir); } + }); + + test('a directory at the ledger path is refused', () => { + const dir = mkPhase(); + try { + fs.mkdirSync(path.join(dir, '03-REVIEW-DISPOSITION.md')); + const res = runAt(dir); + assert.strictEqual(res.exitCode, 0); + assert.match(res.stdout, /not a regular file/); + } finally { cleanup(dir); } + }); + + test('an ordinary ledger is still written — the refusal is not a blanket refusal', () => { + const dir = mkPhase(); + try { + const res = runAt(dir); + assert.strictEqual(res.exitCode, 0); + assert.doesNotMatch(res.stdout, /not a regular file/); + assert.ok(fs.existsSync(path.join(dir, '03-REVIEW-DISPOSITION.md'))); + } finally { cleanup(dir); } + }); +}); + +describe('#3861 round 2 — a DOTTED phase number does not break the step', () => { + // Found by the round's own adversarial review, in its MISSED section -- no finding asked about + // it. Both of this step's call sites accept `03.1` -- `code-review-fix.md`'s PADDED_PHASE validator accepts + // ^[0-9]+(\.[0-9]+)*$ (widened from `?` to `*` by #4568; this comment named the pre-#4568 form + // until round 11) and `execute-phase.md` applies no shape gate at all. (`code-review.md`'s + // own PADDED_PHASE validator is identical but never dispatches this step; it was cited here as a caller + // for several rounds and is not one.) The step reconstructed the path with `printf "%02d"`, which cannot + // format one: bash prints `invalid number` and exits 1. Under `set -euo pipefail` that aborts + // the step on its FIRST line -- the loudest possible failure from a gate that promises never to + // block, and it takes the phase's whole review report with it. + // #3861 round 5, minor 1. The PADDED derivation -- the traversal fence between an + // attacker-influenceable phase number and a file path, plus the per-component length bound -- is + // duplicated verbatim across both fences, because each fenced block runs in a fresh shell and must + // derive what it reads. Each copy is independently tested, but nothing asserted they stay in step, + // and a future edit to one could silently desync the other with the suite still green. That is the + // shared-parallel-surface shape CLAUDE.md requires a parity test for, and it is security-relevant + // validation logic rather than incidental repetition. + // + // Compared LINE BY LINE rather than through a normalizing rewrite: a normalizer would have to be + // told what may differ, and anything it was told to tolerate would stop being asserted. Exactly one + // line may differ, and the test names both of its forms. + test('the two fences derive PADDED identically, and only the refusal message may differ', () => { + const fences = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8')); + assert.strictEqual(fences.length, 2, 'the step must still carry exactly two bash fences'); + const derivationOf = (fence, which) => { + const start = fence.indexOf('_pd="${PHASE_DIR:-}"'); + const end = fence.indexOf('DISPOSITION_FILE="${_pd}/${PADDED}-REVIEW-DISPOSITION.md"'); + assert.ok(start > -1, 'block ' + which + ' must still open the derivation with _pd'); + assert.ok(end > start, 'block ' + which + ' must still close it by building the ledger path'); + return fence.slice(start, end).split('\n'); + }; + const a = derivationOf(fences[0], 1); + const b = derivationOf(fences[1], 2); + // A parity test over an empty or trivial slice passes vacuously and pins nothing. + assert.ok(a.length > 20, 'the derivation must still be the substantial block this pins'); + assert.strictEqual(a.length, b.length, 'the two derivations must have the same shape'); + const differing = a.map((line, i) => [i, line, b[i]]).filter((e) => e[1] !== e[2]); + assert.strictEqual(differing.length, 1, + 'exactly one line may differ between the two derivations; got ' + differing.length + ': ' + + JSON.stringify(differing.map((e) => [e[1], e[2]]))); + assert.match(differing[0][1], /Code review reporting skipped/, 'block 1 refuses by its own name'); + assert.match(differing[0][2], /Code review disposition skipped/, 'block 2 refuses by its own name'); + }); + + test('block 1 reports a dotted phase instead of aborting', { skip: !HAS_BASH }, () => { + const review = ['---', 'phase: 03.1', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---', '', + '### CR-01: a real finding'].join('\n'); + const out = runShippedGateCounts({ reviewText: review, padded: '03.1' }); + assert.strictEqual(out.exitCode, 0, 'advisory: a dotted phase must not abort the step'); + assert.doesNotMatch(out.stderr, /invalid number/, + 'the phase number must never reach printf %02d unsplit'); + assert.match(out.stdout, /^Code review: 1 findings — 1 critical, 0 warning, 0 info\.$/m, + 'and the review is actually found and reported'); + }); + + test('a phase number outside the documented shape builds NO path — traversal fence', { skip: !HAS_BASH }, () => { + // PHASE_NUMBER is interpolated into a file path. The first draft of the dotted-phase fix + // carried an unusable value VERBATIM, which made `${PHASE_DIR}/../../etc/passwd-REVIEW.md` + // reachable where the old `printf "%02d"` had at least mangled it to `00` -- a regression + // introduced by the fix, found by adversarially reviewing it. Both call sites accept + // ^[0-9]+(\.[0-9]+)*$ -- an UNBOUNDED segment count since #4568, anchored at + // `code-review-fix.md`'s PADDED_PHASE validator and ungated at `execute-phase.md` -- and this step has two + // call sites and validates for itself. + // `1.2.3` LEFT THIS LIST in round 11. It is a legal N-segment id at the current base, and + // asserting its refusal here is precisely what held the step narrower than both of them; + // the positive case is its own test below. What remains here is SHAPE, not arity, so the + // two malformed-dot cases that the arity guard used to mask are added explicitly. + // The LETTER AXIS joined in round 12 (#4744 / #4660): a single uppercase letter is legal only as + // the last character of the integer part, so every other placement is shape, not arity. + for (const bad of ['../../etc/passwd', 'abc', '', '-1', '3.', '.1', '+1', '3 1', '1..2', '1.2.', + '23a', 'A23', '2A3', '23AB', '23.1A']) { + const out = runShippedGateCounts({ reviewText: '', writeReview: false, phaseNumber: bad }); + assert.strictEqual(out.exitCode, 0, 'advisory: `' + bad + '` must not abort the step'); + assert.match(out.stdout, /skipped \(unusable phase number/, + '`' + bad + '` must be refused by name, not silently coerced'); + assert.doesNotMatch(out.stdout, /Code review: /, + '`' + bad + '` must not report counts read from a path built out of it'); + } + }); + + test('an N-SEGMENT phase number reports counts, exactly as its callers accept it', { skip: !HAS_BASH }, () => { + // #3861 round 11. Found by this round's own adversarial review, not by the maintainer's. + // The base range widened `code-review-fix.md`'s PADDED_PHASE validator to `^[0-9]+(\\.[0-9]+)*$` (#4568), matching the + // segment-count freedom the canonical grammar in src/phase-id.cts has carried since + // #2128. This step still carried + // `*.*.*) _ok=0` -- "more than one dot: not the documented shape" -- so `23.1.2` took the + // refusal arm, printed `skipped (unusable phase number ...)` and wrote NO ledger, for a + // phase id its own dispatcher had just produced. + // + // It degraded LOUDLY, not silently, which is exactly why nothing caught it: the + // traversal-fence test above asserted that refusal as CORRECT. An arity bound and a shape + // bound had been folded into one arm, so the test that should have failed was the test + // that encoded the bug. + // + // THREE segments and FOUR, deliberately: the retired guard was arity-shaped, so a fix that + // merely moved the bound from two dots to three would pass a three-segment-only test. + for (const phase of ['23.1.2', '1.2.3.4']) { + const review = ['---', 'phase: ' + phase, 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 2', ' info: 0', ' total: 3', '---', '', + '### CR-01: a finding'].join('\n'); + const padded = phase.replace(/^[0-9]+/, (m) => m.padStart(2, '0')); + const out = runShippedGateCounts({ reviewText: review, padded, phaseNumber: phase }); + assert.strictEqual(out.exitCode, 0, phase + ': advisory -- must not abort the step'); + assert.doesNotMatch(out.stdout, /skipped \(unusable phase number/, + phase + ': must NOT be refused -- both of this step\'s callers accept it'); + assert.match(out.stdout, /^Code review: 3 findings — 1 critical, 2 warning, 0 info\.$/m, + phase + ': the review must be found at the N-segment path and reported'); + assert.doesNotMatch(out.stderr, /invalid number/, + phase + ': the phase number must never reach printf %02d unsplit'); + } + }); + + test('a LETTER-VARIANT phase number reports counts, padded as the canonical grammar pads it', { skip: !HAS_BASH }, () => { + // #3861 round 12. #4744 (#4660) widened the six shell/markdown phase mirrors to the canonical + // grammar's letter axis after this branch was cut, and its `lint-phase-id-drift` ratchet then + // flagged this step as the one digit-only mirror left -- found by running the base range's + // modified gates against the rebased tree, not by the review. `12A` and `23A.1.2` were refused + // by name; `3A` must pad to `03A`, the letter carried verbatim after the padded digits exactly as + // src/phase-id.cts pads it. + for (const [phase, padded] of [['12A', '12A'], ['3A', '03A'], ['23A.1.2', '23A.1.2'], ['12345678A', '12345678A']]) { + const review = ['---', 'phase: ' + phase, 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---', '', + '### CR-01: a finding'].join('\n'); + const out = runShippedGateCounts({ reviewText: review, padded, phaseNumber: phase }); + assert.strictEqual(out.exitCode, 0, phase + ': advisory -- must not abort the step'); + assert.doesNotMatch(out.stdout, /skipped \(unusable phase number/, + phase + ': must NOT be refused -- its dispatcher accepts it since #4744'); + assert.match(out.stdout, /^Code review: 1 findings — 1 critical, 0 warning, 0 info\.$/m, + phase + ': the review must be found at the ' + padded + ' path and reported'); + } + }); + + test('the fence agrees with its callers across a probed set spanning both boundaries', { skip: !HAS_BASH }, () => { + // #3861 round 11. The first cut of this test was named "congruence, not merely wider" and + // probed 14 ids, none of them near the length bound. It passed, and the property it named + // was false: the shipped fence is deliberately NARROWER than the callers' regex, because + // the two `?????????*` checks bound the integer part and the suffix to 8 characters each. + // That overclaim was caught by this round's second adversarial review, which drove + // `123456789` and `1.1234567.1` — both caller-valid, both fence-refused. + // + // EXAMPLE-BASED, and the name says so rather than promising a language-level invariant. + // A finite probe set cannot prove congruence over an infinite language: the third review + // pass demonstrated this by injecting a `2) _ok=0` arm into the fence, which this test + // still passed because `2` is not in the list below. Read it as a regression pin over the + // values that actually broke, not as an exhaustive equivalence proof. + // + // It asserts two things, in two parts: + // (1) WITHIN the length bound, the fence and the callers agree exactly -- that is what + // round 11's shape fix bought, and the regression worth pinning. + // (2) BEYOND it, the fence refuses ids the callers accept. That divergence is + // PRE-EXISTING and untouched by this round (the bound predates the N-segment work + // and guards `$((10#...))` against bash's 2^64 wrap); it is pinned here so it stays + // a KNOWN narrowing rather than drifting back into an accidental one. + // The callers' regex since #4744: the letter axis is part of the agreement now. + const CALLER_RE = /^[0-9]+[A-Z]?(\.[0-9]+)*$/; + const refused = (v) => { + const out = runShippedGateCounts({ reviewText: '', writeReview: false, phaseNumber: v }); + assert.strictEqual(out.exitCode, 0, v + ': advisory -- must not abort'); + return /skipped \(unusable phase number/.test(out.stdout); + }; + + // (1) agreement, for every id whose components are each within the bound + for (const v of ['1', '03', '03.1', '23.1.2', '1.2.3.4', '12345678', '1.12345678', + '1.1234567.1', '12345678.12345678', '1.1.1.1.1.1.1.1.1.1', + '12A', '3A', '23A.1.2', '12345678A', '23a', 'A23', '2A3', '23AB', '23.1A', + '1..2', '1.2.', '.1', '3.', 'abc', '-1', '+1', '3 1', '../../etc/passwd']) { + assert.strictEqual(refused(v), !CALLER_RE.test(v), + v + ': with every component within the bound, step and callers must agree'); + } + + // (2) the remaining deliberate narrowing: a SINGLE component over 8 characters. + // This is the `$((10#...))` overflow guard and it is NOT a congruence defect -- bash + // integers wrap at 2^64, so an unbounded integer segment silently becomes a negative + // padded phase. Pinned so the narrowing stays known rather than drifting back. + for (const v of ['123456789', '1.999999999', '123456789A']) { + assert.ok(CALLER_RE.test(v), v + ': precondition -- the callers do accept this'); + assert.ok(refused(v), + v + ': the per-component 8-char bound must keep refusing this; if this flips, the ' + + 'bound changed and `$((10#...))` overflow protection needs re-deriving'); + } + }); + + test('the length bound is PER COMPONENT, not over the whole tail after the first dot', { skip: !HAS_BASH }, () => { + // #3861 round 11, C1. The bound used to read `${_pn#*.}` -- the entire suffix -- which is + // one component only while an id has at most two. The moment N-segment ids were accepted, + // that form rejected `1.1234567.1`: every component is a legal 7 digits, but the tail + // measures 9 characters. The comment above the check had promised per-component bounding + // since before this PR; the code only became untrue of it when the arity arm came out. + // + // Drives the boundary from both sides on a LATER segment, which is the part the old form + // got wrong -- an 8-char middle segment must pass and a 9-char one must fail, with the + // total length in both cases well past what the old whole-tail bound allowed. + const cases = [ + ['1.12345678.1', false, 'an 8-char middle segment is within the per-component bound'], + ['1.123456789.1', true, 'a 9-char middle segment exceeds it'], + ['1.1234567.1', false, 'the id the whole-tail bound rejected for its total length'], + ['12345678.12345678', false, 'two 8-char components, 17 characters total'], + ]; + for (const [phase, mustRefuse, why] of cases) { + const out = runShippedGateCounts({ reviewText: '', writeReview: false, phaseNumber: phase }); + assert.strictEqual(out.exitCode, 0, phase + ': advisory -- must not abort'); + assert.strictEqual(/skipped \(unusable phase number/.test(out.stdout), mustRefuse, + phase + ': ' + why); + } + }); + + test('an integer phase is still zero-padded exactly as before', { skip: !HAS_BASH }, () => { + // Negative control for the split: the ordinary path must be untouched. + const review = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---', '', + '### CR-01: a finding'].join('\n'); + const out = runShippedGateCounts({ reviewText: review, padded: '01' }); + assert.strictEqual(out.exitCode, 0); + assert.match(out.stdout, /^Code review: 1 findings/m); + }); +}); + +describe('#3861 round 1 — the counts mirror is asserted against the shipped shell', () => { + // Every fixture the mirror is exercised on above, plus the count edges. + const FIXTURES = { + 'the documented review': REVIEW_WITH_FINDINGS, + 'blocker: as the critical tier-equivalent': REVIEW_WITH_FINDINGS.replace(' critical: 1', ' blocker: 1'), + 'a body ---, status: and total: after the frontmatter': + REVIEW_WITH_FINDINGS + '\n\n---\n\nstatus: clean\ntotal: 999\n', + 'a legacy review with no findings: block': + ['---', 'phase: 02', 'status: issues_found', '---', '', '# Phase 02'].join('\n'), + 'CRLF line endings': REVIEW_WITH_FINDINGS.replace(/\n/g, '\r\n'), + 'unterminated frontmatter': ['---', 'status: issues_found', 'total: 4', '', '## Body'].join('\n'), + 'no frontmatter at all': '# Phase 01\n\nnothing here\n', + 'a zero-finding review': ['---', 'phase: 01', 'findings:', ' critical: 0', ' warning: 0', + ' info: 0', ' total: 0', 'status: issues_found', '---'].join('\n'), + // Both from the round-1 adversarial pass: the shipped pipeline collapses `1 0` to `10` where + // a trim keeps `1 0`, and truncates at a second colon where a tail capture keeps it. Neither + // is reachable from the well-formed fixtures above, which is exactly why they are here. + 'a count with an internal space': REVIEW_WITH_FINDINGS.replace(' critical: 1', ' critical: 1 0'), + // The fixture that can actually SEE a mirror/shipped parser divergence. The two above cannot: + // on them both parsers reach the countless arm, so the parity assertion holds either way. Here + // the repaired numbers are SELF-CONSISTENT (10 == 10 + 0 + 0), so the retired parser renders a + // full breakdown from a `findings:` block containing no such numbers while the shipped one + // withholds it. + 'a self-consistent repaired breakdown': + ['---', 'phase: 01', 'status: issues_found', 'findings:', ' critical: 1 0', + ' warning: 0', ' info: 0', ' total: 1 0', '---', '', '### CR-01: a'].join('\n'), + // Tab-separated scalars are valid YAML. The retired `tr -d ' '` left the tab in place and made + // every count non-numeric; the shipped trim reads them, so this fixture also pins that change. + 'tab-separated counts': + ['---', 'phase: 01', 'status: issues_found', 'findings:', ' critical:\t1', + ' warning:\t1', ' info:\t1', ' total:\t3', '---', '', '### CR-01: a'].join('\n'), + // The STATUS axis of the same class, and the fixture that pins the `first` helper. Every + // pre-existing status fixture left both parsers on the SAME arm -- `issues:found` truncates to + // `issues`, which is no more `clean` than `issues:found` is -- so the mirror's status read could + // drift from the shipped one unseen, exactly as `firstIn` did. Here truncation FLIPS the arm: + // `clean:junk` cut at the second colon is the silent `clean`, whole it reports. + 'a status whose truncation would flip the arm': + REVIEW_WITH_FINDINGS.replace('status: issues_found', 'status: clean:junk'), + // Unicode whitespace is NOT trimmed, because the reads pin LC_ALL=C. Unpinned under glibc's + // C.UTF-8 the shipped sed trimmed U+2003 and this took the silent clean arm on some machines + // and not others; the mirror never trims it. Same class as the fixture above, locale axis. + 'a status with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace('status: issues_found', 'status: clean\u2003'), + 'a count with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace(' critical: 1', ' critical: 1\u2003'), + 'a findings: opener with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace('findings:', 'findings:\u2003'), + 'a value containing a second colon': REVIEW_WITH_FINDINGS.replace('status: issues_found', 'status: issues:found'), + // POSIX [[:space:]] covers form feed and vertical tab; a [ \t] mirror does not, so the + // shipped grep matches a line the mirror rejects outright. Third counterexample, same class. + 'a key indented with a form feed': REVIEW_WITH_FINDINGS.replace(' critical: 1', '\fcritical: 1'), + // Minor 1: a TOP-LEVEL key sharing a name with a nested count. The reads were scoped to the + // frontmatter but not to the `findings:` mapping the values belong to, so `^[[:space:]]*total:` + // matched this one first and the gate reported a number from outside the breakdown. + 'a top-level total: ahead of the nested one': + ['---', 'phase: 02', 'total: 999', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---'].join('\n'), + 'a top-level info: and critical: ahead of the nested ones': + ['---', 'critical: 42', 'info: 7', 'status: issues_found', 'findings:', + ' critical: 1', ' warning: 0', ' info: 0', ' total: 1', '---'].join('\n'), + }; + + for (const [name, reviewText] of Object.entries(FIXTURES)) { + test('shipped shell and mirror agree on ' + name, { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ reviewText }); + // Parity over the WHOLE emitted message, not over five intermediate variables the test + // used to print for itself. A drift in any parsed value changes this string or the arm + // it selects, so the assertion binds to what an operator actually sees. + assert.strictEqual( + shipped.stdout, + renderGateMessage(parseGateCounts(reviewText), 1), + 'the mirror has drifted from the shipped awk/grep block' + ); + }); + } + + // The parity fixtures above run under whatever locale the suite inherits, so they can only catch a + // locale bug on a machine that happens to have it. This drives the SAME input under both locales and + // asserts the shipped fence does not care -- which is the actual property LC_ALL=C buys. Found by the + // round's fifth adversarial pass: glibc's C.UTF-8 classifies U+2003 as [[:space:]] AND [[:blank:]] + // where C and en_US.UTF-8 classify it as neither, so before the pin `status: clean<U+2003>` trimmed + // to the silent `clean` on some machines and reported on others. + for (const [name, reviewText] of Object.entries({ + 'a status with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace('status: issues_found', 'status: clean\u2003'), + 'a count with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace(' critical: 1', ' critical: 1\u2003'), + 'a unicode-space indented count key': + REVIEW_WITH_FINDINGS.replace(' critical: 1', '\u2003critical: 1'), + // Pass 6's finding, and the one input that proves the awk was missed: the BLOCK OPENER. Under + // C.UTF-8 `findings:<U+2003>` matched `/^findings:[[:space:]]*$/` and opened the mapping, under + // C it did not -- so the same review rendered a full breakdown on one machine and the countless + // message on another, with every grep and sed already pinned. + 'a findings: opener with a trailing unicode space': + REVIEW_WITH_FINDINGS.replace('findings:', 'findings:\u2003'), + 'the documented review': REVIEW_WITH_FINDINGS, + })) { + test('the shipped reads are locale-invariant on ' + name, { skip: !HAS_BASH }, () => { + const c = runShippedGateCounts({ reviewText, extraEnv: { LC_ALL: 'C', LANG: 'C' } }); + const utf8 = runShippedGateCounts({ reviewText, extraEnv: { LC_ALL: 'C.UTF-8', LANG: 'C.UTF-8' } }); + assert.strictEqual(c.exitCode, 0, 'advisory: must not abort under any locale'); + assert.strictEqual(utf8.exitCode, 0, 'advisory: must not abort under any locale'); + assert.strictEqual(utf8.stdout, c.stdout, + 'the shipped reads must not depend on the ambient locale; drop LC_ALL=C and this reds'); + // And the mirror predicts that one locale-independent answer. + assert.strictEqual(c.stdout, renderGateMessage(parseGateCounts(reviewText), 1), + 'the mirror must model the locale-pinned shipped read'); + }); + } + + test('every locale-sensitive tool in the step is pinned to LC_ALL=C', () => { + // THE CHECK THAT WOULD HAVE CAUGHT THE LAST MISS. `grep`, `sed` and `awk` all resolve + // [[:space:]] / [[:blank:]] through the ambient locale, and glibc's C.UTF-8 classifies U+2003 + // as both where C and en_US.UTF-8 classify it as neither. Pass 5 pinned the grep and sed reads + // and the round then CLAIMED the parser was locale-independent; pass 6 found the two `awk` + // mapping selectors still unpinned, because that census searched for the tools it expected + // rather than the tools that were there. Asserting the invariant over the file is that census + // in a form a future edit is far less likely to slip past. + // + // WHAT THIS IS, AND WHAT IT IS NOT. It is a REGRESSION GUARD against the accident that has now + // happened twice — a read added or edited without its pin, in a file where every existing read + // has one. It is NOT a proof, and it is deliberately not written as one. It scans TEXT, so it + // cannot see shell or JS command structure: a name that is not written literally (`$AWK "$f"`, + // a command composed as `a''wk`, a command name computed inside the embedded `node -e` block) + // is invisible to it, and so is an executable command substitution on a physical line that + // begins with `#` inside a multiline quoted argument, which the comment exemption below skips. + // + // Earlier versions of this comment tried to ENUMERATE those residuals. Three adversarial passes + // in a row then found one more each time, which is the actual lesson: the list cannot be closed, + // so a comment promising a closed list is false the moment someone is cleverer than it. The + // examples above are illustrations, not an inventory. Its two directions are NOT symmetric, and + // that asymmetry is the whole operating instruction: a report is ADJUDICABLE — read the reported + // line together with what precedes it, since the same text can be a command or an argument to one + // (` grep` is a call after `:` and a string after `printf '%s\n' \`), so the line alone does not + // always settle it — whereas SILENCE proves nothing at all, + // because the evasions above are silent and so is any evasion no one has thought of yet. So: + // investigate every report, and never read silence as proof that a new read is pinned. + // + // Driven, it does catch: unpinned, `env`-prefixed, wrongly-pinned (`LC_ALL=C.UTF-8`), + // path-qualified, line-initial, and literal-in-Node calls. Its false answers run loud rather + // than quiet — a trailing comment naming a tool, a tool name inside an awk program, or a path + // whose component starts with one (`bin/grep-wrapper`) would all trip it. That direction is the + // right one for a guard, and none of those shapes exists in the step today. (The illustration + // is deliberately not a `docs/`-prefixed path: lint-docs-guard-registration reads one of those + // as a real docs reference from this file and demands a baseline entry for it.) + // `splitLines`, not `split('\n')`: the repo's own lint bans the latter on readFileSync content + // (DEFECT.WINDOWS-CRLF-TEST-PORTABILITY), and a CRLF checkout would otherwise leave a stray + // `\r` on every line here. + const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs'); + const step = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + const offenders = []; + splitLines(step).forEach((line, i) => { + if (/^\s*#/.test(line)) return; // prose may name an unpinned form while explaining it + // Blank out the PINNED calls first, then anything left naming one of these tools is unpinned. + // Deliberately NOT keyed on a leading `|`: every call is piped today, but a guard that only + // sees pipes would wave through `awk '...' < "$f"` or `$(grep ...)`, and "correct for the + // shapes that happen to exist right now" is the exact property that let the awk selectors sit + // unpinned through a whole commit that claimed otherwise. + const rest = line.replace(/LC_ALL=C\s+(?:grep|sed|awk)\b/g, ''); + // The preceding-char class deliberately does NOT shield `/` or `.`: an earlier version did, + // and `/usr/bin/awk '...' < "$f"` sailed through it. A path-qualified call is still a call. + // `parsed`, `passed` and `awkward` stay unmatched, because their tool substring is preceded + // or followed by a word character. + const leftover = rest.match(/(?:^|[^A-Za-z0-9_-])(?:grep|sed|awk)\b/g) || []; + for (const c of leftover) offenders.push(`${i + 1}: ${c.trim()} (unpinned)`); + }); + assert.deepStrictEqual(offenders, [], + 'every grep/sed/awk in the step must be LC_ALL=C-pinned; an unpinned one makes the parse ' + + 'depend on the machine, which is what pass 5 and pass 6 each found'); + // `cut -d: -f2-` and `tr -d '\r'` are deliberately NOT pinned, and the exemption is principled + // rather than an oversight: neither resolves a character class or a collation. One splits on a + // single ASCII byte, the other deletes one literal byte. + assert.ok(/\|\s*cut -d: -f2-/.test(step), 'the cut reads are still the class-free shape'); + }); + + test('a zero-finding review reports a real breakdown, not the countless fallback', { skip: !HAS_BASH }, () => { + // Minor 6. `0` is a number, so the gate must state `0 findings — 0 critical, …` + // rather than fall back. The `case` guard rejects the empty string and non-digits; + // a guard written against truthiness would reject this and say nothing at all. + const shipped = runShippedGateCounts({ reviewText: FIXTURES['a zero-finding review'] }); + assert.strictEqual(shipped.stdout, + 'Code review: 0 findings \u2014 0 critical, 0 warning, 0 info.\n' + + 'Consider running: /gsd:code-review 1 --fix\n', + 'all four counts are numeric, so the breakdown is reported rather than withheld'); + }); + + test('a partial findings: block makes the whole breakdown unavailable', { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ + reviewText: ['---', 'findings:', ' total: 4', 'status: issues_found', '---'].join('\n'), + }); + assert.strictEqual(readGateMessage(shipped.stdout).countsOk, '0', + 'a total without the three severities is not a breakdown'); + assert.match(shipped.stdout, /^Code review found issues\.$/m, + 'the countless form is what reaches the operator'); + }); + + test('a missing REVIEW.md leaves the counts empty and does not abort', { skip: !HAS_BASH }, () => { + // Behavioural replacement for the `src.includes('if [ -f "$REVIEW_FILE" ] …')` + // assertion: under `set -e` an aborting block is what actually breaks the phase. + const shipped = runShippedGateCounts({ reviewText: '', writeReview: false }); + assert.strictEqual(shipped.exitCode, 0, 'advisory: a missing review must not abort the step'); + // An absent review yields an empty status, which is a NON-REPORTING arm: the gate says + // nothing at all rather than claiming a countless review. Asserting on the observable is + // what makes that distinction visible; the old six-line probe could not express it. + assert.strictEqual(shipped.stdout, '', 'no review, no message'); + assert.strictEqual(readGateMessage(shipped.stdout).reported, false); + }); + + test('the countless fallback requires all four counts, not just the total', { skip: !HAS_BASH }, () => { + // Fifth `src.includes()` assertion converted to a behavioural one (round 1 retired four). + // It pinned the PROSE that stated the condition, so it went red the moment the emitter moved + // into the fence and the prose was rewritten — while the behaviour it named was untouched. + // That is the pin arguing for its own conversion: the arm is now executed and observable, so + // assert the arm. Contrast is the point — a bare total takes the countless arm, the full set + // takes the breakdown arm — which a one-sided assertion could not express. + const fm = (rows) => ['---', 'findings:', ...rows, 'status: issues_found', '---'].join('\n'); + const totalOnly = runShippedGateCounts({ reviewText: fm([' total: 4']) }); + assert.match(totalOnly.stdout, /^Code review found issues\.$/m, + 'a numeric total with missing severities must not emit a half-filled breakdown'); + assert.doesNotMatch(totalOnly.stdout, /findings —/, + 'and must not emit the breakdown form at all'); + const allFour = runShippedGateCounts({ + reviewText: fm([' critical: 1', ' warning: 2', ' info: 1', ' total: 4']), + }); + assert.match(allFour.stdout, /^Code review: 4 findings — 1 critical, 2 warning, 1 info\.$/m, + 'all four present and consistent is what the breakdown arm requires'); + }); + + // The guard this pair covers is `[ -f "$REVIEW_FILE" ] && [ -r "$REVIEW_FILE" ]`, and the two + // legs need different fixtures because only one of them survives root. + // + // `chmod 0o000` does NOT make a file unreadable to root: root bypasses POSIX read permission + // bits, so `[ -r ]` stays true, the fence reads the fixture, and the assertion below sees the + // real breakdown instead of silence. That is not hypothetical here — the bench runs this suite + // as root under Docker, where this test failed with + // actual: 'Code review: 4 findings — 1 critical, 2 warning, 1 info.\n...' + // against an expected ''. CLAUDE.md names the mode-bit trick as the wrong tool for injecting an + // IO failure, and the usual remedy — monkeypatch the read to throw EACCES — does not reach this + // site: the read is performed by a spawned `bash`, not by this process, so stubbing node's `fs` + // is not on the code path at all. + // + // So the `-r` leg keeps the mode-bit fixture and declares the lanes it cannot bind on, exactly + // as tests/plan-review-convergence.test.cjs does for its own shell-side `-r` arm, and the `-f` + // leg below carries the contract on every lane INCLUDING root. Skipping the first without + // adding the second would have traded a false failure for lost coverage. + const SKIP_MODE_BITS = !HAS_BASH + ? 'POSIX-only bash fragment' + : typeof process.getuid === 'function' && process.getuid() === 0 + ? 'root bypasses the read permission bit' + : false; + + test('an unreadable REVIEW.md leaves the counts empty and does not abort', { skip: SKIP_MODE_BITS }, () => { + const shipped = runShippedGateCounts({ reviewText: REVIEW_WITH_FINDINGS, mode: 0o000 }); + assert.strictEqual(shipped.exitCode, 0, 'advisory: an unreadable review must not abort the step'); + assert.strictEqual(shipped.stdout, '', 'an unreadable review reports nothing, and does not guess'); + }); + + // Root-immune companion: a directory fails `-f` for every euid, so this binds on the bench lane + // where the test above is skipped. Same arm, same observable — silence and a zero exit. + test('a directory standing in for REVIEW.md leaves the counts empty and does not abort', { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ reviewText: REVIEW_WITH_FINDINGS, plantDir: true }); + assert.strictEqual(shipped.exitCode, 0, 'advisory: a non-regular review path must not abort the step'); + assert.strictEqual(shipped.stdout, '', 'a review path that is not a regular file reports nothing, and does not guess'); + assert.strictEqual(readGateMessage(shipped.stdout).reported, false); + }); + + // An EMPTY REGULAR FILE is a third arm, and it is NOT the missing case: `-f` and `-r` both pass, + // so the fence opens and reads the file where the missing case never gets past the guard. The + // missing-file test above cannot reach it — it passes `writeReview: false`, so no file exists at + // all — and until this test every `reviewText: ''` call in this file did the same. The PR body + // has claimed this case since round 1; it was documented as covered and was not covered. + // + // WHY the scan yields nothing, stated precisely because the obvious reading is wrong: it is NOT + // the `NR==1{if($0!="---") exit}` guard. A zero-byte file supplies awk no record at all, so that + // action never executes (NR stays 0). The output is empty because `closed` is never set and the + // END block therefore prints nothing. + // + // Since round 11 the observable is NOT the missing-file case's any more, and that is the point: + // a zero-byte REVIEW.md was read and has no parseable status, so the fence says so rather than + // staying silent -- silence is what a clean review looks like. The counts are still empty and + // nothing is guessed; what changed is that the operator is told the report could not be read. + test('an EMPTY REVIEW.md leaves the counts empty, reports the unparseable status, and does not abort', { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ reviewText: '' }); // writeReview defaults true: a real, empty file + assert.strictEqual(shipped.exitCode, 0, 'advisory: an empty review must not abort the step'); + assert.match(shipped.stdout, /^Code review status unparsed: /m, 'an empty file was read, and says so'); + assert.doesNotMatch(shipped.stdout, /^Code review: \d+ findings/m, 'and no breakdown is invented'); + assert.strictEqual(readGateMessage(shipped.stdout).unparsed, true); + }); +}); + +describe('#3861 round 11 — a malformed report does not read as clean in the counting arm', () => { + // A REVIEW.md with three criticals and an UNTERMINATED frontmatter yielded REVIEW_STATUS='', + // and the counting arm then printed nothing -- byte-identical to a clean review. Block 2 still + // said `status: none` rather than `clean`, so a careful reader could separate them downstream, + // which is the only reason this was Minor. The counting arm now distinguishes 'no findings' + // from 'could not parse', and block 2 names the same distinction ('unparsed' vs 'none'). + const UNTERMINATED = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 3', ' warning: 0', ' info: 0', ' total: 3', '', + '## Critical Issues', '', '### CR-01: a', '### CR-02: b', '### CR-03: c'].join('\n'); + + test('an unterminated frontmatter is reported as unparsed, not passed over in silence', { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ reviewText: UNTERMINATED }); + assert.strictEqual(shipped.exitCode, 0, 'advisory: a malformed review must not abort the step'); + assert.match(shipped.stdout, /^Code review status unparsed: REVIEW\.md is present but its frontmatter has no parseable status; severity counts unavailable\.$/m); + assert.doesNotMatch(shipped.stdout, /^Code review: \d+ findings/m, 'a breakdown read from an unterminated block would be the body leak this scan prevents'); + assert.doesNotMatch(shipped.stdout, /Consider running/, 'nothing here proves there are findings to fix'); + assert.strictEqual(readGateMessage(shipped.stdout).unparsed, true); + }); + + test('a closed frontmatter with no status: key is the same arm', { skip: !HAS_BASH }, () => { + const noStatus = ['---', 'phase: 01', 'findings:', ' critical: 3', ' warning: 0', ' info: 0', + ' total: 3', '---', '', '### CR-01: a'].join('\n'); + const shipped = runShippedGateCounts({ reviewText: noStatus }); + assert.match(shipped.stdout, /^Code review status unparsed: /m); + }); + + test('a review with no frontmatter at all is the same arm', { skip: !HAS_BASH }, () => { + const shipped = runShippedGateCounts({ reviewText: '# Phase 01\n\n### CR-01: a finding\n' }); + assert.match(shipped.stdout, /^Code review status unparsed: /m); + }); + + test('the absent, directory and unreadable cases stay silent — nothing was read, so nothing is described', { skip: !HAS_BASH }, () => { + // The distinction is READ-ness, not emptiness: the three guard-refused shapes have no file + // content to describe, and describing one would be the guess the guard exists to prevent. + // (The mode-bit case is pinned by its own root-aware test above.) + assert.strictEqual(runShippedGateCounts({ reviewText: '', writeReview: false }).stdout, '', 'absent'); + assert.strictEqual(runShippedGateCounts({ reviewText: UNTERMINATED, plantDir: true }).stdout, '', 'directory'); + }); + + test('block 2 names the same distinction: unparsed for a read file, none for an absent one', { skip: !HAS_BASH }, () => { + const script = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'))[1]; + const runBlock2 = (dir, padded) => runHook('-c', ['set -euo pipefail\n' + script + '\n'], { + interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS, + env: { ...process.env, PHASE_DIR: dir, PHASE_NUMBER: String(Number(padded)) }, + }); + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-unparsed-')); + try { + // No REVIEW.md, no ledger, no fix report -> skipped (status: none) + let res = runBlock2(dir, '01'); + assert.strictEqual(res.exitCode, 0); + assert.match(res.stdout, /^Code review disposition skipped \(status: none\)$/m, 'absent reads as none'); + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), UNTERMINATED); + res = runBlock2(dir, '01'); + assert.strictEqual(res.exitCode, 0); + assert.match(res.stdout, /^Code review disposition skipped \(status: unparsed\)$/m, 'a read-but-unparseable review reads as unparsed'); + assert.doesNotMatch(res.stdout, /status: none/); + } finally { + cleanup(dir); + } + }); +}); + +describe('#3861 round 12 — block 2 does not compute a shortfall from a self-contradicting findings block', () => { + // Round 12, Minor. Block 1 withholds the severity breakdown unless the four counts are numeric + // AND `critical + warning + info == total`; block 2 bounded `total` for digits and length only + // and then handed it to the `unparsed:` reconciliation. So a REVIEW.md whose `findings:` block + // disagrees with itself made block 1 print the countless form -- breakdown suppressed as + // untrustworthy -- while block 2 still computed a shortfall from that same untrusted number. + // Two trust models for one field, one fence apart, with the weaker one downstream. + // + // The fix is NARROWER than "re-apply block 1's check", deliberately: block 1 demands all four + // counts because it DISPLAYS all four. Block 2 uses `total` alone. Applying the all-four rule + // here would blank a perfectly usable `total: 5` on a review carrying no severity keys and + // SILENTLY DROP a shortfall the step reports correctly today -- trading a safe-direction + // over-report for a silent under-report. Only the CONTRADICTION ports. + const headings = ['', '### CR-01: a conforming finding', '### WR-01: another', '### WR-02: a third']; + const review = (fm) => ['---', 'phase: 01', 'status: issues_found', ...fm, '---', ...headings].join('\n'); + + const runBlock2 = (dir) => { + const script = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'))[1]; + return runHook('-c', ['set -euo pipefail\n' + script + '\n'], { + interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS, + env: { ...process.env, PHASE_DIR: dir, PHASE_NUMBER: '1' }, + }); + }; + const drive = (fm, status) => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-r12-')); + try { + const body = status === undefined ? review(fm) + : review(fm).replace('status: issues_found', 'status: ' + status); + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), body); + const res = runBlock2(dir); + assert.strictEqual(res.exitCode, 0, 'the fence is advisory and must never abort: ' + res.stderr); + const p = path.join(dir, '01-REVIEW-DISPOSITION.md'); + return { stdout: res.stdout, ledger: fs.existsSync(p) ? fs.readFileSync(p, 'utf8') : null }; + } finally { + cleanup(dir); + } + }; + + test('a contradicting findings block yields no shortfall (fails before the fix)', { skip: !HAS_BASH }, () => { + // total: 10 against 1 + 1 + 1. Three headings parse. Before the fix this rendered + // `unparsed: 7` from a number block 1 had already judged untrustworthy. + const out = drive(['findings:', ' critical: 1', ' warning: 1', ' info: 1', ' total: 10']); + assert.doesNotMatch(out.ledger, /^unparsed:/m, 'a self-contradicting total is not a number to reconcile against'); + assert.doesNotMatch(out.stdout, /recorded NOWHERE/); + assert.match(out.ledger, /^total: 3$/m, 'and the rows it does have are still reported'); + }); + + test('a CONSISTENT total still reconciles — the check narrows nothing it should not', { skip: !HAS_BASH }, () => { + // 1 + 2 + 2 == 5, three headings parse, so two are recorded nowhere and must be said. + const out = drive(['findings:', ' critical: 1', ' warning: 2', ' info: 2', ' total: 5']); + assert.match(out.ledger, /^unparsed: 2$/m); + assert.match(out.stdout, /2 finding\(s\) recorded NOWHERE/); + }); + + test('a total with NO severity keys still reconciles — absent is not a contradiction', { skip: !HAS_BASH }, () => { + // The over-reach control. There is nothing for `total: 5` to disagree WITH here, so the + // shortfall this step reports correctly today must survive the new check. + const out = drive(['findings:', ' total: 5']); + assert.match(out.ledger, /^unparsed: 2$/m, 'absent counts must not suppress a real shortfall'); + }); + + test('a partial or non-numeric breakdown is likewise not a contradiction', { skip: !HAS_BASH }, () => { + for (const fm of [ + ['findings:', ' critical: 1', ' total: 5'], + ['findings:', ' critical: x', ' warning: 1', ' info: 1', ' total: 5'], + ]) { + assert.match(drive(fm).ledger, /^unparsed: 2$/m, JSON.stringify(fm)); + } + }); + + test('`blocker:` is read as the critical tier, exactly as block 1 reads it', { skip: !HAS_BASH }, () => { + // A mirror that dropped the documented alternation would diverge from block 1 on precisely + // the reviews that use it: 1 + 1 + 1 == 3 here, so the sum agrees and nothing is suppressed. + const out = drive(['findings:', ' blocker: 1', ' warning: 1', ' info: 1', ' total: 3']); + assert.doesNotMatch(out.ledger, /^unparsed:/m, 'blocker counted as critical => the sum agrees'); + // And the contradicting twin, to prove the alternation is load-bearing rather than inert. + const bad = drive(['findings:', ' blocker: 1', ' warning: 1', ' info: 1', ' total: 9']); + assert.doesNotMatch(bad.ledger, /^unparsed:/m); + }); + + test('a severity with an INTERNAL space is malformed, and does not suppress the shortfall', { skip: !HAS_BASH }, () => { + // Found by the round's own adversarial pass, against the first version of this fix. The sibling + // reads use `cut -d: -f2 | tr -d ' '`, which deletes INTERNAL spaces too, so `critical: 1 0` + // arrives as the perfectly numeric `10`. That is long-standing in those reads and was INERT here + // until this block started reading the severities -- at which point a repaired number could + // satisfy the sum test and SUPPRESS a real `unparsed:` shortfall. The severity reads now trim the + // ends only, so the space survives, the digit check rejects it, and nothing is suppressed. + const out = drive(['findings:', ' critical: 1 0', ' warning: 0', ' info: 0', ' total: 5']); + assert.match(out.ledger, /^unparsed: 2$/m, 'a malformed severity must not license suppression'); + }); + + test('a scalar with a SECOND COLON is malformed, and the well-formed total still reconciles', { skip: !HAS_BASH }, () => { + // Second adversarial pass. `cut -d: -f2` takes only the SECOND FIELD, so `critical: 1: junk` + // arrived as the perfectly numeric `1` -- 1+0+0 != 5 read as a contradiction and SUPPRESSED a + // shortfall that was genuinely owed. `-f2-` keeps everything after the first colon, so the + // malformed scalar stays malformed, no contradiction is claimed, and `total: 5` still reconciles. + const out = drive(['findings:', ' critical: 1: junk', ' warning: 0', ' info: 0', ' total: 5']); + assert.match(out.ledger, /^unparsed: 2$/m, 'a malformed severity must not license suppression'); + }); + + test('a malformed TOTAL is not repaired into a fabricated shortfall', { skip: !HAS_BASH }, () => { + // Second adversarial pass, and the sharper of the two. The first version of this fix parsed the + // severities strictly and left `total` lenient -- so `critical: 5 0` with `total: 1 0` repaired + // ONLY the total to `10`, rejected the severity, skipped the contradiction check, and INVENTED + // `unparsed: 7` against three parsed headings. A field is either trustworthy or it is not: + // parsing one leniently and its sibling strictly is the shape that fabricates. Every count this + // block reads now goes through one parser. + const out = drive(['findings:', ' critical: 5 0', ' warning: 0', ' info: 0', ' total: 1 0']); + assert.doesNotMatch(out.ledger, /^unparsed:/m, 'a malformed total is not a number to reconcile against'); + assert.doesNotMatch(out.stdout, /recorded NOWHERE/); + }); + + test('BLOCK 1 withholds a breakdown built from REPAIRED counts', { skip: !HAS_BASH }, () => { + // The one input that distinguishes block 1's old parser from its new one, and the reason this + // test exists: the round's first negative control for the block-1 change was VACUOUS. Reverting + // block 1 to `cut -d: -f2 | tr -d ' '` left the whole suite green, because on every fixture that + // existed both parsers landed on the SAME arm -- the countless form -- so a parity assertion + // could not see the difference. + // + // A SELF-CONSISTENT repaired breakdown separates them. `critical: 1 0` / `total: 1 0` repairs to + // 10 and 10, which SUM, so the old parser reported `10 findings -- 10 critical, 0 warning, + // 0 info.` from a `findings:` block that contains no such numbers. The new parser rejects the + // scalars and takes the countless arm, which is also what block 2 does with the same bytes -- + // a console line and a ledger can no longer contradict each other on this input. + const repaired = ['---', 'phase: 01', 'status: issues_found', 'findings:', + ' critical: 1 0', ' warning: 0', ' info: 0', ' total: 1 0', '---', '', + '### CR-01: a', '### WR-01: b', '### WR-02: c'].join('\n'); + const shipped = runShippedGateCounts({ reviewText: repaired }); + assert.strictEqual(shipped.exitCode, 0, 'advisory: never abort'); + assert.doesNotMatch(shipped.stdout, /10 findings/, 'a repaired number must not be reported as a count'); + assert.doesNotMatch(shipped.stdout, /10 critical/); + assert.match(shipped.stdout, /^Code review found issues\./m, 'the countless arm is the honest one here'); + }); + + test('a status scalar carrying a second colon does not silently take the clean arm', { skip: !HAS_BASH }, () => { + // Fourth adversarial pass, MISSED. `status:` kept the retired `cut -d: -f2` after the counts moved + // off it, and it is the read where truncation costs most: the valid YAML scalar `status: clean:junk` + // arrived as the bare `clean`, so an unusable status took the CLEAN arm and suppressed both the + // report and the ledger. The whole scalar matches no arm, so the step reports instead. + const out = drive(['findings:', ' critical: 1', ' warning: 0', ' info: 0', ' total: 1'], 'clean:junk'); + assert.ok(out.ledger !== null, 'an unusable status must not suppress the ledger'); + }); + + test('a status with trailing unicode whitespace does not silently take the clean arm', { skip: !HAS_BASH }, () => { + // Fifth adversarial pass. The ledger half of the locale finding: unpinned, glibc's C.UTF-8 trimmed + // U+2003 and `status: clean<U+2003>` suppressed the ledger on exactly the machines whose locale + // said so. Pinned to C the scalar stays unusable, matches no arm, and the ledger is written. + const out = drive(['findings:', ' critical: 1', ' warning: 0', ' info: 0', ' total: 1'], 'clean\u2003'); + assert.ok(out.ledger !== null, 'an unusable status must not suppress the ledger under any locale'); + }); + + test('the parser is symmetric across all four count fields', { skip: !HAS_BASH }, () => { + // Third adversarial pass, MISSED: the fixtures exercised malformed `critical` and `total` only, + // so they did not actually pin the four-field symmetry the fix claims. Each field in turn gets + // each malformed shape; none may license suppression, because a malformed severity is not a + // disagreement and `total: 5` stays usable. + for (const key of ['critical', 'blocker', 'warning', 'info']) { + for (const bad of ['1 0', '1: junk']) { + const fm = ['findings:', ' total: 5']; + for (const k of ['critical', 'warning', 'info']) { + if (k === key || (key === 'blocker' && k === 'critical')) continue; + fm.push(` ${k}: 0`); + } + fm.push(` ${key}: ${bad}`); + assert.match(drive(fm).ledger, /^unparsed: 2$/m, `${key}: '${bad}' must not license suppression`); + } + } + }); + + test('an absent severity still proves a contradiction when the known ones OVERSHOOT', { skip: !HAS_BASH }, () => { + // Third adversarial pass. Counts are non-negative, so a missing one can only ADD: when the + // present severities already sum to MORE than `total`, the block disagrees with itself whatever + // the absent value is. Requiring all three before comparing reconciled against a total the + // present counts had already refuted. + const over = drive(['findings:', ' critical: 4', ' warning: 4', ' total: 5']); + assert.doesNotMatch(over.ledger, /^unparsed:/m, '4 + 4 > 5 is a contradiction with or without info:'); + // The other direction stays reconcilable: an UNDERshoot is exactly what the absent count explains. + const under = drive(['findings:', ' critical: 1', ' warning: 1', ' total: 5']); + assert.match(under.ledger, /^unparsed: 2$/m, 'an undershoot is the absent count\'s job, not a contradiction'); + }); + + test('a zero-padded breakdown does not take the advisory step down', { skip: !HAS_BASH }, () => { + // `10#` on every operand: bash reads a leading zero as octal, so `critical: 08` would make + // $(( )) fail with "value too great for base" and, under `set -e`, abort a step that + // promises never to block. 8 + 1 + 1 == 10, so the sum agrees and the shortfall is reported. + const out = drive(['findings:', ' critical: 08', ' warning: 01', ' info: 01', ' total: 10']); + assert.match(out.ledger, /^unparsed: 7$/m, 'octal-looking counts are read as decimal, not as an abort'); + }); +}); + +// --------------------------------------------------------------------------- +// #3861 round 1 — Minor 5, and the finding-id census the review did not ask for +// --------------------------------------------------------------------------- + +const REVIEWER_AGENT_PATH = path.join(ROOT, 'agents', 'gsd-code-reviewer.md'); + +// The alternations the shipped script uses to recognise a finding id. There are three — +// the heading matcher, the ledger row re-parser, and the frontmatter `- id:` matcher — +// and adding a prefix to only some of them is silent. The severity map below is a fourth +// copy of the same set; it is not an alternation, so it is extracted separately. +// This scan is by PATTERN, never a fixed list of sites, which is why round 5's new +// matcher was absorbed with no edit here. Do not convert it to an enumeration. +function idAlternations() { + const script = shippedDispositionScript(); + return [...script.matchAll(/\(\?:((?:[A-Z]{2}\|)+[A-Z]{2})\)-/g)].map((m) => m[1].split('|').sort().join('|')); +} + +// The FOURTH copy: the severity map's keys. It is not an alternation, so the extractor above cannot +// see it — and a set that agrees in the three regexes while mis-tiering in the map is the drift the +// guard would otherwise miss entirely. +// (Said THIRD until round 6. The extractor finds three alternations — round 5's frontmatter `- id:` +// matcher is the third — so the map has been the fourth copy since then. The count is prose only; +// nothing below reads it.) +function severityMapKeys() { + const script = shippedDispositionScript(); + const m = script.match(/\{([^}]*?)\}\[id\.split/); + assert.ok(m, 'the severity map must still be an inline object literal indexed by the id prefix'); + return [...m[1].matchAll(/([A-Z]{2}):/g)].map((x) => x[1]); +} + +describe('#3861 round 1 — stale fix reports are stated, not silently ignored', () => { + test('a fix report naming a different finding under a reused id says so', () => { + // Exact-title coupling is deliberate — ids are reused across re-reviews, so a + // stale REVIEW-FIX.md must not mark a brand-new CR-01 fixed. But failing it + // silently leaves 'open' indistinguishable from 'the report never named it', + // which is the one thing the ledger exists to tell apart. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a genuinely new finding'].join('\n'); + const fixText = ['## Fixed Issues', '', '### CR-01: the finding this id used to mean'].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText }); + const rows = ledgerRows(out.ledger); + assert.strictEqual(rows[0].disposition, 'open', 'a stale report must not decide the row'); + assert.match(out.stdout, /titles its finding differently from the review/, + 'and the mismatch must be reported, not swallowed'); + assert.match(out.stdout, /a stale report, or a re-titled one/, + 'stated as the observation it is -- the step cannot tell the two causes apart'); + assert.match(out.stdout, /CR-01/, 'naming the finding it could not reconcile'); + }); + + test('a title differing only in INTRA-LINE whitespace still reconciles (m2)', () => { + // gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that the title is + // copied byte-for-byte, so a fixer that re-spaces a title used to produce a spurious note and + // leave a genuinely-fixed row 'open'. Runs of spaces carry no information; they are collapsed. + // NAMED PRECISELY. An earlier version of this test called itself the "reflowed" case while + // substituting triple spaces, which is not a reflow -- see the bound pinned below. + const title = 'a long finding title a fixer might re-space'; + const review = ['---', 'status: issues_found', '---', '', '### CR-01: ' + title].join('\n'); + const respaced = ['## Fixed Issues', '', '### CR-01: ' + title.replace(/ /g, ' ')].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText: respaced }); + assert.strictEqual(ledgerRows(out.ledger)[0].disposition, 'fixed', + 'a re-spaced title is the same title, and the fix outcome must reach the ledger'); + assert.doesNotMatch(out.stdout, /titles its finding differently/, + 'and no spurious mismatch is reported'); + }); + + test('a title WRAPPED across lines is not reconciled — the bound, pinned deliberately', () => { + // The limit of the m2 fix, stated rather than left to be discovered. A `###` heading is ONE + // line by definition: if a fixer wraps a long title, the continuation is a separate paragraph + // and the heading parser -- correctly -- captures only the first line. Whitespace collapsing + // cannot reach across that boundary. + // + // NOT widened, and the reason is that widening is the worse defect: to reconcile a wrapped + // title the parser would have to absorb whatever follows a heading into the title, which + // silently swallows arbitrary prose and would make the stale-report check meaningless. The + // failure mode kept here is the SAFE one -- a visible mismatch note and a row left open, + // never a wrong 'fixed'. + const review = ['---', 'status: issues_found', '---', '', + '### CR-01: a long finding title that wraps'].join('\n'); + const wrapped = ['## Fixed Issues', '', '### CR-01: a long', 'finding title that wraps'].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText: wrapped }); + assert.strictEqual(ledgerRows(out.ledger)[0].disposition, 'open', + 'a wrapped heading does not reconcile -- and fails in the safe direction'); + assert.match(out.stdout, /titles its finding differently/, + 'the mismatch is reported rather than swallowed'); + }); + + test('a RE-CASED or truncated title still reports a mismatch — the strict half is kept', () => { + // The deliberate residual. Case changes and truncation are the shapes a genuinely different + // finding takes, so widening to them would trade a visible false positive for a silent false + // negative -- a stale report marking a brand-new CR-01 fixed, which is the worse direction. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: The Finding'].join('\n'); + const recased = ['## Fixed Issues', '', '### CR-01: the finding'].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText: recased }); + assert.strictEqual(ledgerRows(out.ledger)[0].disposition, 'open'); + assert.match(out.stdout, /titles its finding differently/); + }); + + test('a matching fix report reports no mismatch', () => { + // Negative control for the note itself: it must not fire on the ordinary path. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: same title'].join('\n'); + const fixText = ['## Fixed Issues', '', '### CR-01: same title'].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText }); + assert.strictEqual(ledgerRows(out.ledger)[0].disposition, 'fixed'); + assert.doesNotMatch(out.stdout, /reused id/); + }); +}); + +describe('#3861 round 1 — finding-id prefix census', () => { + test('every copy of the prefix set agrees with every other', () => { + // The set is written out FOUR times in one script — the heading matcher, the ledger + // row re-parser, the frontmatter `- id:` matcher the title tracking added, and (by its + // keys) the severity map. Adding a prefix to some of them does not error; it drops + // carried rows on the next run. + // The count is stated for the reader; nothing below depends on it. idAlternations() + // scans the script by PATTERN rather than walking a fixed site list, which is why the + // fourth site was absorbed without a change here — this comment is the only thing that + // fell behind, and a guard whose population is hand-listed is the defect it would have + // been. Do not convert this to an enumeration. + const alts = idAlternations(); + assert.ok(alts.length >= 2, 'the script must still enumerate finding-id prefixes'); + assert.strictEqual(new Set(alts).size, 1, 'the prefix enumerations have drifted apart: ' + alts.join(' vs ')); + // And the fourth copy, which is not an alternation: every prefix the regexes admit must either + // carry an explicit tier in the severity map or fall to `info` by the documented default. + // Without this, the three regexes can gain a prefix while the map silently mis-tiers it. + const mapped = new Set(severityMapKeys()); + const admitted = alts[0].split('|'); + const unmapped = admitted.filter((p) => !mapped.has(p)); + assert.deepStrictEqual( + unmapped, ['IN'], + 'only IN may rely on the info default; every other admitted prefix needs an explicit tier' + ); + // And the other direction, which a one-way check leaves open: a tier for a prefix the + // regexes never admit is dead code that reads as coverage. + const unadmitted = [...mapped].filter((p) => admitted.indexOf(p) === -1); + assert.deepStrictEqual( + unadmitted, [], + 'the severity map tiers prefixes the id regexes do not admit: ' + unadmitted.join(',') + ); + }); + + test('the prefix set covers every id shape the reviewer agent emits', () => { + // The DOMAIN is owned elsewhere — gsd-code-reviewer.md's body template and its + // Label-equivalence paragraph — so it can acquire a member without this script + // changing. An unlisted prefix is not mis-tiered, it is INVISIBLE: the finding + // never enters the order list and gets no row at all. + const agent = fs.readFileSync(REVIEWER_AGENT_PATH, 'utf8'); + // BOTH surfaces. The body template writes `### CR-01:` headings; the Label-equivalence + // paragraph defines BL in PROSE and appears in no heading at all (`### BL-` occurs zero + // times). A heading-only scan therefore passes today purely because BL happens to be + // hard-coded, and would miss the next prose-defined prefix exactly as it would miss BL. + const emitted = new Set([ + ...[...agent.matchAll(/^###\s+([A-Z]+)-\d+:/gm)].map((m) => m[1]), + ...[...agent.matchAll(/\b([A-Z]+)-\s*(?:IDs?|prefix)/g)].map((m) => m[1]), + ...[...agent.matchAll(/IDs? beginning with\s+`?([A-Z]+)-/g)].map((m) => m[1]), + ]); + assert.ok(emitted.size > 0, 'the reviewer agent must still declare its finding-id shapes'); + const known = new Set(idAlternations()[0].split('|')); + for (const prefix of emitted) { + assert.ok(known.has(prefix), prefix + '- findings would get no disposition row at all'); + } + }); +}); + +// --------------------------------------------------------------------------- +// #3861 round 1, second pass — defects found by adversarially reviewing the +// round's OWN fixes before pushing them. Every one of these was invisible to +// the maintainer's review and to the first pass above. +// --------------------------------------------------------------------------- + +describe('#3861 round 1 — fence tracking, status gating, count consistency', () => { + test('a foreign fence marker inside a fenced example does not swap example for finding', () => { + // The worst shape this file has carried: a bare fenced/not-fenced toggle treats ``` and ~~~ + // as interchangeable, so a ~~~ line inside a ``` example CLOSES the fence and the example's + // real close REOPENS one. The ledger then records the ILLUSTRATION and drops the finding — + // a confidently-written artifact that is wrong in both directions at once. + const review = ['---', 'status: issues_found', '---', '', '```', '~~~', + '### CR-77: an example inside a fence', '```', '', '### CR-01: a real finding'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01'], 'the real finding, and only it'); + }); + + test('a longer close does not require an exact-length match, per CommonMark', () => { + const review = ['---', 'status: issues_found', '---', '', '```', + '### CR-77: fenced', '````', '', '### CR-01: real'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01']); + }); + + test('the disposition block gates on review status in shell, not in prose', { skip: !HAS_BASH }, () => { + // Block 1 computes REVIEW_STATUS and emits nothing, and its shell is discarded — so a + // condition stated only in the prose between the blocks is not available to anything. A + // clean re-review would otherwise rewrite a ledger it was never meant to touch. + const fences = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8')); + assert.ok(fences.length >= 2, 'the step must still carry more than one shell block'); + assert.match(fences[1], /REVIEW_STATUS/, 'block 2 must re-derive the status it is gated on'); + assert.match(fences[1], /clean\|skipped/, 'and gate on the documented clean/skipped/empty set'); + }); + + test('an internally inconsistent breakdown is withheld, not half-rendered', { skip: !HAS_BASH }, () => { + // `total: 0` beside `critical: 1` is four valid numbers producing a self-contradicting + // line. Numeric is necessary, not sufficient. + const shipped = runShippedGateCounts({ + reviewText: ['---', 'findings:', ' critical: 1', ' warning: 0', ' info: 0', + ' total: 0', 'status: issues_found', '---'].join('\n'), + }); + assert.strictEqual(readGateMessage(shipped.stdout).countsOk, '0', + 'the counts do not sum to the total, so no breakdown'); + }); + + test('a consistent breakdown is still reported', { skip: !HAS_BASH }, () => { + // Negative control for the sum check — it must not withhold a correct breakdown. + const shipped = runShippedGateCounts({ reviewText: REVIEW_WITH_FINDINGS }); + assert.strictEqual(readGateMessage(shipped.stdout).countsOk, '1'); + }); + + test('a carried finding that REAPPEARS loses the carried marker', () => { + // The defect that storing the cell verbatim introduced, and the reason the strip is back. + // Run 1 carries CR-01 and marks it; run 2 reports CR-01 again. If the marker were permanent + // the ledger would state 'not in the current review' about a finding plainly in it — + // an artifact confidently wrong about its own contents. + const absent = ['---', 'status: issues_found', '---', '', '### WR-01: other'].join('\n'); + const back = ['---', 'status: issues_found', '---', '', '### CR-01: it came back'].join('\n'); + const prior = '| CR-01 | critical | deferred | waiting on ADR-9 |'; + const run1 = runShippedDisposition({ reviewText: absent, priorText: prior }); + assert.strictEqual( + ledgerRows(run1.ledger).find((r) => r.id === 'CR-01').source, + 'waiting on ADR-9 (not in the current review)', 'carried, and marked as such' + ); + const run2 = runShippedDisposition({ reviewText: back, priorText: run1.ledger }); + const row = ledgerRows(run2.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(row.source, 'waiting on ADR-9', 'the marker goes when the finding returns'); + assert.strictEqual(row.disposition, 'deferred', 'and the decision itself is still preserved'); + }); + + test('a carried row still does not grow its marker across runs', () => { + // The property the strip exists for, re-pinned now that it is bounded. + const review = ['---', 'status: issues_found', '---', '', '### WR-01: unrelated'].join('\n'); + const prior = '| CR-01 | critical | deferred | waiting on ADR-9 (not in the current review) |'; + const first = runShippedDisposition({ reviewText: review, priorText: prior }); + const carried = ledgerRows(first.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(carried.source, 'waiting on ADR-9 (not in the current review)'); + const second = runShippedDisposition({ reviewText: review, priorText: first.ledger }); + assert.strictEqual( + ledgerRows(second.ledger).find((r) => r.id === 'CR-01').source, + 'waiting on ADR-9 (not in the current review)', + 'exactly one marker, however many times the gate runs' + ); + }); +}); + +// --------------------------------------------------------------------------- +// #3861 round 1, third pass — the fidelity gap that let the second pass ship a +// regression the whole suite was green over. +// +// Every test above extracts the embedded script as TEXT and runs it. Bash does +// not: it expands the double-quoted `node -e "..."` argument first, so an +// unescaped backtick is COMMAND SUBSTITUTION and the script Node receives is +// not the script the tests read. That is not a hypothetical — the previous +// commit shipped exactly that, in a code comment, and 122 green tests said +// nothing because none of them ever asked bash what it would actually pass. +// --------------------------------------------------------------------------- + +describe('#3861 round 1 — the tests must run what BASH would run', () => { + test('bash expansion of the node -e argument matches what the tests extract', { skip: !HAS_BASH }, () => { + // Ask bash for the literal argument it would hand node, and compare. This is the general + // guard: it catches an unescaped backtick, an unescaped $, and any other expansion the + // extractor's two-escape undo cannot model — none of which the behavioural tests can see. + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8').replace(/\r\n/g, '\n'); + const open = src.indexOf('node -e "'); + const body = src.slice(open + 'node -e "'.length); + const end = body.indexOf('\n" || echo '); + assert.ok(end !== -1, 'the node -e script must still be closed by its || echo fallback'); + const quoted = body.slice(0, end); + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-fid-')); + try { + const out = path.join(dir, 'arg.txt'); + // `printf %s` with the SAME double-quoted string the step uses: whatever bash does to it + // on the way to node, it does here too. + const probe = 'printf %s "' + quoted + '" > ' + JSON.stringify(out) + '\n'; + const res = runHook('-c', [probe], { interpreter: 'bash', timeoutMs: PROBE_TIMEOUT_MS }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the probe must run to completion'); + assert.strictEqual( + res.stderr.trim(), '', + 'bash emitted diagnostics expanding the node -e argument — an unescaped backtick or $: ' + res.stderr + ); + assert.strictEqual( + fs.readFileSync(out, 'utf8'), shippedDispositionScript(), + 'bash hands node a DIFFERENT script than the tests exercise' + ); + } finally { + cleanup(dir); + } + }); + + // Run ONLY the status guard at the head of the disposition block — everything up to the shim + // preamble. runShippedDisposition drives the node script directly and never sees this shell at + // all, so a test written against it says nothing about the guard: it passed unchanged with the + // guard made unconditional, which is exactly the vacuity this helper exists to remove. + function runDispositionGuard({ reviewText, withLedger, withFix, withIterFix }) { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-grd-')); + try { + fs.writeFileSync(path.join(dir, '01-REVIEW.md'), reviewText); + if (withLedger) fs.writeFileSync(path.join(dir, '01-REVIEW-DISPOSITION.md'), '| CR-01 | critical | deferred | x |\n'); + if (withFix) fs.writeFileSync(path.join(dir, '01-REVIEW-FIX.md'), '## Fixed Issues\n\n### CR-01: a thing\n'); + if (withIterFix) fs.writeFileSync(path.join(dir, '01-REVIEW-FIX.iter2.md'), '## Fixed Issues\n\n### CR-01: a thing\n'); + const fence = bashFences(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'))[1]; + const cut = fence.indexOf('_GSD_SHIM_NAME='); + assert.ok(cut > 0, 'the disposition block must still open with its guard, then the shim'); + const script = 'set -euo pipefail\n' + fence.slice(0, cut) + '\nprintf "PROCEEDED\\n"\n'; + const res = runHook('-c', [script], { + interpreter: 'bash', + timeoutMs: PROBE_TIMEOUT_MS, + env: { ...process.env, PHASE_DIR: dir, PHASE_NUMBER: '1' }, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED, 'the guard must run to completion'); + return { proceeded: /PROCEEDED/.test(res.stdout), stdout: res.stdout, exitCode: res.exitCode }; + } finally { + cleanup(dir); + } + } + + test('a clean review with no ledger is skipped', { skip: !HAS_BASH }, () => { + const out = runDispositionGuard({ reviewText: ['---', 'status: clean', '---'].join('\n') }); + assert.strictEqual(out.proceeded, false, 'nothing to record and nothing to reconcile'); + assert.match(out.stdout, /skipped \(status: clean\)/); + }); + + test('a clean review with an EXISTING ledger still proceeds, to reconcile it', { skip: !HAS_BASH }, () => { + // Freezing the ledger here would leave findings showing as open that the review no longer + // reports — the case the embedded script's own reconciliation path is written for. A guard + // that skipped unconditionally would make that path unreachable on exactly the run needing it. + const out = runDispositionGuard({ reviewText: ['---', 'status: clean', '---'].join('\n'), withLedger: true }); + assert.strictEqual(out.proceeded, true, 'an existing ledger must still be reconciled'); + assert.match(out.stdout, /reconciling the fix report and any existing disposition ledger/); + }); + + test('a review reporting issues always proceeds', { skip: !HAS_BASH }, () => { + const out = runDispositionGuard({ reviewText: ['---', 'status: issues_found', '---'].join('\n') }); + assert.strictEqual(out.proceeded, true); + }); + + // #3861 round 5 — a converged `--auto` run reaches the guard with a CLEAN review and, on a direct + // /gsd-code-review invocation, no gate-written ledger. Keying the skip on the ledger alone meant a + // fully successful multi-iteration run — every finding fixed and committed — recorded nothing at all. + test('a clean review with NO ledger but a fix report still proceeds', { skip: !HAS_BASH }, () => { + const out = runDispositionGuard({ reviewText: ['---', 'status: clean', '---'].join('\n'), withFix: true }); + assert.strictEqual(out.proceeded, true, 'a fix report is a decision to record, ledger or not'); + assert.match(out.stdout, /reconciling the fix report/); + }); + + test('a clean review with NO ledger but only an ITERATION fix report still proceeds', { skip: !HAS_BASH }, () => { + // The converged loop's earlier iterations survive only as <NN>-REVIEW-FIX.iterN.md, so the + // backups have to count toward the guard exactly as the final report does. + const out = runDispositionGuard({ reviewText: ['---', 'status: clean', '---'].join('\n'), withIterFix: true }); + assert.strictEqual(out.proceeded, true, 'an iteration backup is a decision to record too'); + }); + + test('a clean review with neither a ledger nor any fix report is still skipped', { skip: !HAS_BASH }, () => { + // The widening must not become "always proceed" — the original skip is still correct when + // there is genuinely nothing to record. + const out = runDispositionGuard({ reviewText: ['---', 'status: clean', '---'].join('\n') }); + assert.strictEqual(out.proceeded, false); + assert.match(out.stdout, /skipped \(status: clean\)/); + }); + + // ── #3861 round 5 — the --auto multi-iteration reconciliation gap ────────────────────────── + // + // code-review-fix.md overwrites REVIEW-FIX.md on every iteration and DELETES the .iterN.md + // backups on convergence, so a finding fixed in iteration 1 was absent from the final fix report + // AND from the final review (it was fixed, so the re-review stopped reporting it). The step then + // fell back to the gate's `open` row and rendered `open ... (not in the current review)` — the + // same bytes a finding that vanished for an unrelated reason produces. + + test('a fix report naming a finding the review no longer reports records it FIXED, not open', () => { + // The precise site: sameTitle(undefined, h.title) is false and title.has(id) is false too, so + // the entry entered NEITHER applied NOR staleFix and was dropped in silence. + const review = ['---', 'status: issues_found', '---', '', '### WR-09: something else'].join('\n'); + const fix = ['## Fixed Issues', '', '### CR-01: the one fixed earlier'].join('\n'); + const out = runShippedDisposition({ reviewText: review, fixText: fix }); + const rows = ledgerRows(out.ledger); + const cr = rows.find((r) => r.id === 'CR-01'); + assert.ok(cr, 'the fixed finding must have a row at all'); + assert.strictEqual(cr.disposition, 'fixed', 'a committed fix must not render as open'); + assert.match(cr.source, /not in the current review/, 'and it must be marked as no longer reported'); + }); + + test('an ITERATION fix report is reconciled even when the final report has moved on', () => { + // The reviewer's scenario end to end: iteration 1 fixed CR-01, iteration 3 fixed WR-09, and the + // final REVIEW-FIX.md carries only the last iteration's scope. + const review = ['---', 'status: issues_found', '---', '', '### IN-07: still open'].join('\n'); + const fix = ['## Fixed Issues', '', '### WR-09: fixed last'].join('\n'); + const iter = { 2: ['## Fixed Issues', '', '### CR-01: fixed in iteration one'].join('\n') }; + const out = runShippedDisposition({ reviewText: review, fixText: fix, iterFixText: iter }); + const rows = ledgerRows(out.ledger); + assert.strictEqual(rows.find((r) => r.id === 'CR-01').disposition, 'fixed'); + assert.strictEqual(rows.find((r) => r.id === 'WR-09').disposition, 'fixed'); + assert.strictEqual(rows.find((r) => r.id === 'IN-07').disposition, 'open'); + }); + + test('the NEWEST fix report wins when two iterations decide the same id differently', () => { + // Reports are read newest-first, so first-occurrence-wins gives the most recent statement — + // the same precedence a duplicate id already gets WITHIN one report. + const review = ['---', 'status: issues_found', '---', '', '### IN-07: unrelated'].join('\n'); + const fix = ['## Skipped Issues', '', '### CR-01: contested'].join('\n'); + const iter = { 2: ['## Fixed Issues', '', '### CR-01: contested'].join('\n') }; + const out = runShippedDisposition({ reviewText: review, fixText: fix, iterFixText: iter }); + assert.strictEqual(ledgerRows(out.ledger).find((r) => r.id === 'CR-01').disposition, 'skipped'); + }); + + test('a REUSED id whose title differs is NOT inherited as fixed — the stale-report arm still rules', () => { + // The negative control for the arm added above. Re-review renumbers, so an earlier iteration's + // CR-01 and the current review's CR-01 can be different findings; carrying the decision across + // that boundary would render a false `fixed`, which is worse than the `open` it replaced. + const review = ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'); + const iter = { 2: ['## Fixed Issues', '', '### CR-01: an older, different finding'].join('\n') }; + const out = runShippedDisposition({ reviewText: review, iterFixText: iter }); + const cr = ledgerRows(out.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'open', 'a different finding under a reused id must stay open'); + assert.match(out.stdout, /title their finding differently|titles its finding differently/, + 'and the mismatch must be stated, not swallowed'); + }); + + // ── #3861 round 5 rework — a REUSED finding id must not inherit the old finding's decision ── + // + // Found by the round's own adversarial review, which drove it: ids are reused across re-reviews + // (the --auto loop renumbers), and row() inherited a prior decision on an id match alone. A prior + // 'CR-01 fixed' against a review reporting a brand-new CR-01 rendered the NEW finding `fixed` — a + // false decision in the artifact whose entire purpose is telling triaged from forgotten. + + test('the ledger does not promise a preservation it no longer makes', () => { + // The rendered text said "preserves every row and every disposition" while the step had gained + // an intentional drop for a reused id — shipped, user-facing text asserting something false. + // And the console must not point at git: committing is gated on commit_docs and a failed commit + // is swallowed, so under commit_docs=false the overwritten decision may exist nowhere. + // Round 5 found this test VACUOUS: it ran with no prior ledger, so no reuse ever occurred and + // the `is in git` assertion could not have failed however the console was worded. Driven through + // a real drop now, so the negative assertion is made against a console line that actually exists. + const seed = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: the original finding'].join('\n'), + priorText: '| CR-01 | critical | deferred | waiting on the vendor |\n', + }); + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: seed.ledger, + }); + assert.match(out.stdout, /decision\(s\) DROPPED/, 'precondition: this run must actually drop a decision'); + assert.doesNotMatch(out.ledger, /preserves every row and every disposition/, + 'the unqualified preservation promise must not return'); + assert.match(out.ledger, /id is REUSED/, 'and the one exception must be stated where a reader meets it'); + assert.doesNotMatch(out.stdout, /is in git/, 'the console must not assert a recovery path that may not exist'); + }); + + test('an `open` prior is replaced SILENTLY, and the shipped text says so', () => { + // Round 5: the legend and both feature docs claimed the drop is named on the console + // unconditionally. It is not — `row()` reports only a RECORDED decision (`was.d !== 'open'`). + // The behaviour is deliberate (an `open` row records no decision to lose); the text was wrong. + const seed = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: the original finding'].join('\n'), + }); + assert.strictEqual(ledgerRows(seed.ledger).find((r) => r.id === 'CR-01').disposition, 'open'); + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: seed.ledger, + }); + assert.doesNotMatch(out.stdout, /DROPPED/, 'an untriaged prior row is replaced without a report'); + assert.match(out.ledger, /a row still at `open` is replaced silently/, + 'and the legend must state that exception rather than promising an unconditional report'); + }); + + test('a dropped decision is REPORTED on the console, not lost quietly', () => { + // The `superseded:` block that preserved these was tried and withdrawn — it produced a fresh + // defect on each of three review passes. What survives is the guard (no false `fixed`) plus an + // explicit report; the prior ledger row remains in git. + const seed = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: the original finding'].join('\n'), + priorText: '| CR-01 | critical | deferred | waiting on the vendor |\n', + }); + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: seed.ledger, + }); + assert.strictEqual(ledgerRows(out.ledger).find((r) => r.id === 'CR-01').disposition, 'open'); + assert.match(out.stdout, /decision\(s\) DROPPED/, 'the drop must be stated'); + assert.match(out.stdout, /CR-01=deferred/, 'naming the id and what was decided'); + assert.doesNotMatch(out.ledger, /^superseded:/m, 'and no preservation block is written'); + }); + + test('a prior decision is NOT inherited when the id now names a different finding', () => { + const first = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: the original finding'].join('\n'), + fixText: ['## Fixed Issues', '', '### CR-01: the original finding'].join('\n'), + }); + assert.strictEqual(ledgerRows(first.ledger).find((r) => r.id === 'CR-01').disposition, 'fixed'); + assert.match(first.ledger, /title: "the original finding"/, 'the title must be recorded to make reuse detectable'); + + const second = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: first.ledger, + }); + const cr = ledgerRows(second.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'open', 'a different finding under a reused id must be untriaged'); + }); + + + + test('an EMPTY title is recorded, so it cannot read back as a pre-format ledger', () => { + // The leak that came back three passes running: while an empty title emitted no `title:` key, + // an empty-titled finding read back as legacy and inherited a decision across a reused id. + const first = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### IN-07: unrelated'].join('\n'), + iterFixText: { 2: ['## Fixed Issues', '', '### CR-01:'].join('\n') }, + }); + assert.match(first.ledger, /title: ""/, 'an empty title is still recorded, explicitly'); + const second = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: first.ledger, + }); + assert.strictEqual(ledgerRows(second.ledger).find((r) => r.id === 'CR-01').disposition, 'open', + 'an empty recorded title is a title, not an absent one'); + }); + + test('a pre-format ledger whose title merely LOOKS like JSON keeps its quotes', () => { + // Without the `titles: json` marker, JSON.parse ran on every value — so a legacy bare title + // written as "quoted whole title" lost its quotes, stopped matching, and flipped to open. + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: "quoted whole title"'].join('\n'), + priorText: [ + '---', 'phase: 01', 'review: 01-REVIEW.md', 'findings:', + ' - id: CR-01', ' severity: critical', ' disposition: deferred', + ' title: "quoted whole title"', 'open: 0', 'total: 1', 'recorded: x', '---', '', + '| CR-01 | critical | deferred | a reason |', + ].join('\n'), + }); + const cr = ledgerRows(out.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred', 'the legacy title must still identify the finding'); + assert.doesNotMatch(out.stdout, /DROPPED/, 'and no drop may be reported'); + }); + + test('a prior ledger with NO recorded title still inherits its decision (back-compat)', () => { + // A ledger written before titles were recorded carries none. Refusing to inherit there would + // reset every decision in it — the loss this guard exists to prevent, caused by the guard. + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: whatever it is called now'].join('\n'), + priorText: '| CR-01 | critical | deferred | a reason from an older ledger |\n', + }); + const cr = ledgerRows(out.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'deferred', 'an absent prior title must inherit, not reset'); + assert.strictEqual(cr.source, 'a reason from an older ledger'); + }); + + test('an iteration-derived decision cites the report it actually came from', () => { + // The Source cell hard-coded the unsuffixed <NN>-REVIEW-FIX.md, so a decision read out of an + // iteration backup cited a file that may not exist. A citation the reader cannot follow is + // worse than none. + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### IN-07: unrelated'].join('\n'), + iterFixText: { 2: ['## Fixed Issues', '', '### CR-01: fixed in iteration one'].join('\n') }, + }); + const cr = ledgerRows(out.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(cr.disposition, 'fixed'); + assert.match(cr.source, /01-REVIEW-FIX\.iter2\.md/, 'the cited report must be the one that decided it'); + }); + + // ── #3861 round 5, second rework pass — four defects the review drove out of the FIRST fix ── + + test('an iteration-only decision records the title it was decided under', () => { + // Without this the row was written with NO title -- the current review does not report the + // finding, so nothing else knows one -- and the next review reusing that id then hit the + // title-ABSENT back-compat exception and inherited the old `fixed`. The very defect the title + // machinery exists to close, surviving through the hole opened for legacy ledgers. + const first = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### IN-07: unrelated'].join('\n'), + iterFixText: { 2: ['## Fixed Issues', '', '### CR-01: the original finding'].join('\n') }, + }); + assert.match(first.ledger, /title: "the original finding"/, 'the deciding title must be recorded'); + + const second = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: a brand new finding'].join('\n'), + priorText: first.ledger, + }); + assert.strictEqual(ledgerRows(second.ledger).find((r) => r.id === 'CR-01').disposition, 'open', + 'a reused id must not inherit a decision recorded for a different finding'); + }); + + + + test('the ledger frontmatter is valid YAML even when a title contains a colon', () => { + // `title: Parser: loses data` is not YAML — a real reader returns 'bad indentation of a mapping + // entry'. The values are emitted as JSON scalars, which YAML 1.2 reads as double-quoted strings. + const yaml = require('js-yaml'); + const out = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: Parser: loses data'].join('\n'), + }); + const fm = yaml.load(out.ledger.split('---')[1]); + assert.strictEqual(fm.findings[0].title, 'Parser: loses data', 'the colon must survive the round trip'); + // And the title still round-trips through the parser as an identity, so a re-run is unchanged. + const again = runShippedDisposition({ + reviewText: ['---', 'status: issues_found', '---', '', '### CR-01: Parser: loses data'].join('\n'), + priorText: out.ledger, + }); + assert.strictEqual(again.wroteNothing, true, 'a second run must still report unchanged'); + }); + + test('a carried human reason ending in the marker is preserved, and not doubled', () => { + // Neither stripped (the previous fix's residue) nor appended twice. + const review = ['---', 'status: issues_found', '---', '', '### WR-01: unrelated'].join('\n'); + const prior = '| CR-01 | critical | deferred | defer because (not in the current review) |'; + const first = runShippedDisposition({ reviewText: review, priorText: prior }); + const carried = ledgerRows(first.ledger).find((r) => r.id === 'CR-01'); + assert.strictEqual(carried.source, 'defer because (not in the current review)'); + const second = runShippedDisposition({ reviewText: review, priorText: first.ledger }); + assert.strictEqual( + ledgerRows(second.ledger).find((r) => r.id === 'CR-01').source, + 'defer because (not in the current review)', + 'stable across runs — the render appends nothing it can already see' + ); + }); + + test('a leading-zero count is evaluated, not silently skipped', { skip: !HAS_BASH }, () => { + // Bash infers the base from a leading zero, so `critical: 08` made $(( )) fail. It does NOT + // abort — the expansion sits in an `if` condition, where set -e does not fire — which is + // worse: the sum check silently does not run, an inconsistent breakdown passes, and the + // only trace is a stray diagnostic on stderr. Asserting exit 0 alone is vacuous here; it + // was true before the fix too. Assert the check's OUTCOME and the absent diagnostic. + const padded = (total) => ['---', 'findings:', ' critical: 08', ' warning: 0', + ' info: 0', ' total: ' + total, 'status: issues_found', '---'].join('\n'); + const inconsistent = runShippedGateCounts({ reviewText: padded('9') }); + assert.strictEqual(readGateMessage(inconsistent.stdout).countsOk, '0', + '08 + 0 + 0 is 8, not 9 \u2014 the check must FIRE'); + assert.doesNotMatch(inconsistent.stderr, /value too great for base/, + 'the padded count must be read as decimal, not left to base inference'); + const consistent = runShippedGateCounts({ reviewText: padded('8') }); + assert.strictEqual(consistent.exitCode, 0, 'an advisory gate must not abort on a padded count'); + assert.strictEqual(readGateMessage(consistent.stdout).countsOk, '1', + '08 + 0 + 0 is 8, so the breakdown is consistent'); + }); + + test('a fence indented past three spaces is not a fence', { skip: !HAS_BASH }, () => { + // CommonMark: at most three leading spaces open a fence; four is an indented code block. + const src = fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8'); + assert.match(src, /\^ \{0,3\}\(/, 'the fence matcher must bound its leading whitespace'); + }); +}); + +// Placement here is now incidental. This block used to be pinned to the end of the file because +// HAS_BASH was declared mid-file and a `skip` option referencing it from an earlier block hit the +// temporal dead zone -- cancelling its neighbours rather than failing visibly. HAS_BASH is +// declared with the file's other top-level constants now, so that constraint is gone and this +// block may be moved beside its siblings whenever someone is tidying. +describe('#3861 round 1 — count validation, executed', () => { + test('all four counts are validated before the breakdown is shown', { skip: !HAS_BASH }, () => { + // Was a `src.includes()` assertion on the exact `case` line, which broke the moment that + // line grew a length bound. The behaviour is what matters and is now executed directly. + const counts = (c, w, i, t) => ['---', 'findings:', ' critical: ' + c, ' warning: ' + w, + ' info: ' + i, ' total: ' + t, 'status: issues_found', '---'].join('\n'); + const ok = (t) => readGateMessage(runShippedGateCounts({ reviewText: t }).stdout).countsOk; + assert.strictEqual(ok(counts('x', '0', '0', '0')), '0', + 'a non-numeric count withholds the whole breakdown'); + assert.strictEqual(ok(counts('1', '2', '1', '4')), '1'); + // Bash integers wrap at 2^64, so a 20-digit count reaches the sum as 0 and an inconsistent + // breakdown passes. Length-bounded, because no review reports nine digits of findings. + assert.strictEqual(ok(counts('18446744073709551616', '0', '0', '0')), '0', + 'a count long enough to wrap the sum is not a count'); + }); + + test('the count-length threshold is covered at limit-1, limit and limit+1', { skip: !HAS_BASH }, () => { + // M2. The guard is `?????????*` -- nine or more characters -- so the limit is 8 digits + // ACCEPTED, 9 REJECTED. The only cases here were 'x', single digits and a 20-digit value, + // none of which pins the boundary: dropping one `?` moves the limit to 7 digits and no test + // would have noticed. All three points are asserted, and the sum is kept consistent at each + // so the length rule is what decides the verdict rather than the sum check. + const counts = (c, w, i, t) => ['---', 'findings:', ' critical: ' + c, ' warning: ' + w, + ' info: ' + i, ' total: ' + t, 'status: issues_found', '---'].join('\n'); + const ok = (t) => readGateMessage(runShippedGateCounts({ reviewText: t }).stdout).countsOk; + const d = (n) => '1'.padEnd(n, '0'); // n digits, leading 1 so the value is exact + assert.strictEqual(d(7).length, 7); + assert.strictEqual(d(8).length, 8); + assert.strictEqual(d(9).length, 9); + assert.strictEqual(ok(counts(d(7), '0', '0', d(7))), '1', 'limit-1: 7 digits is accepted'); + assert.strictEqual(ok(counts(d(8), '0', '0', d(8))), '1', 'limit: 8 digits is accepted'); + assert.strictEqual(ok(counts(d(9), '0', '0', d(9))), '0', 'limit+1: 9 digits is rejected'); + }); +}); + +describe('#3861 round 1 — an absent review still reconciles', () => { + test('a missing REVIEW.md does not abandon an existing ledger', () => { + // The status guard proceeds when a ledger exists, so the script must tolerate the review + // being gone: reading it unconditionally threw, the trailing fallback swallowed it, and the + // ledger was left frozen — the exact freeze the reconciliation path exists to prevent. + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-3861-abs-')); + try { + const dispPath = path.join(dir, '01-REVIEW-DISPOSITION.md'); + fs.writeFileSync(dispPath, ['| CR-01 | critical | deferred | waiting on ADR-9 |', + '| WR-09 | warning | open | - |'].join('\n')); + const res = runNode(['-e', shippedDispositionScript()], { + timeoutMs: PROBE_TIMEOUT_MS, + env: { + ...process.env, + REVIEW_FILE: path.join(dir, '01-REVIEW.md'), // deliberately absent + DISPOSITION_FILE: dispPath, + FIX_REPORT_FILE: path.join(dir, '01-REVIEW-FIX.md'), + PADDED: '01', + }, + }); + assert.strictEqual(res.outcome, OUTCOME.EXITED); + assert.strictEqual(res.exitCode, 0, 'an absent review is not an error: ' + res.stderr); + const rows = ledgerRows(fs.readFileSync(dispPath, 'utf8')); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01', 'WR-09'], + 'an absent review reconciles the ledger; it does not delete rows from it'); + assert.match(rows[0].source, /not in the current review/); + } finally { + cleanup(dir); + } + }); +}); + +describe('#3861 round 1 — fence closers and one-letter prefixes', () => { + test('a line with an info string is an opener shape, never a closer', () => { + // CommonMark: a closing fence carries only whitespace after its marker. Treating an + // info-string line as a close ends the fence early and admits the example headings under it. + const review = ['---', 'status: issues_found', '---', '', '```', + '```js', '### CR-77: still inside the fence', '```', '', '### CR-01: real'].join('\n'); + const rows = ledgerRows(runShippedDisposition({ reviewText: review }).ledger); + assert.deepStrictEqual(rows.map((r) => r.id), ['CR-01']); + }); + + test('a one-letter finding prefix in the agent template is not invisible', () => { + // The domain scan required [A-Z]{2,}, so a template heading like `### C-01:` — explicit and + // parseable, not prose — contributed nothing and the guard passed over a finding shape the + // ledger would drop entirely. + const agent = fs.readFileSync(REVIEWER_AGENT_PATH, 'utf8'); + const emitted = [...agent.matchAll(/^###\s+([A-Z]+)-\d+:/gm)].map((m) => m[1]); + assert.ok(emitted.length > 0, 'the reviewer agent must still declare its finding-id shapes'); + const oneLetter = [...'### C-01: x'.matchAll(/^###\s+([A-Z]+)-\d+:/gm)].map((m) => m[1]); + assert.deepStrictEqual(oneLetter, ['C'], 'the scan must admit a single-letter prefix'); + }); +}); + +// --------------------------------------------------------------------------- +// The REVIEW.md lookup's phase-id handling, tested at the step file that owns it. +// +// These four tests used to live in `tests/nsegment-phase-grammar.test.cjs`, inside the +// `#4748` describe block, because that is where the lookup's gate was when #3829 moved the +// lookup out of `execute-phase.md`. #4781 (#4628) then removed #4748's letter-axis work from +// the grammar file, so the block that hosted them no longer exists. They are re-homed here +// unchanged in substance: they assert properties of THIS PR's step file, not of #4748's sites, +// and keeping them beside the step they guard is what stops an unrelated upstream revert from +// silently deleting this PR's own coverage. +// +// One test did NOT come along: the assertion that `execute-phase.md`'s init parse list names +// `padded_phase`. That is a property of #4748's site, not of this step — #4781 removed the +// field from that list, and carrying the assertion here would only pin someone else's revert. +// +// The MECHANISM differs from the pre-move one and the PROPERTY does not. `execute-phase.md` +// bound init's `{padded_phase}`; the step validates PHASE_NUMBER for shape and traversal and +// then pads the DIGIT RUN as a STRING, carrying an optional letter and any dot segments verbatim. +// Both refuse exactly the two shapes #4748 named: `printf "%02d"` cannot pad `03A` (prints `03`, +// exits 1) and reads an already-padded `08` as octal (prints `00`) -- and the string pad refuses +// them by not doing arithmetic at all, which also fixed a THIRD shape the arithmetic form got +// wrong (`008` -> `08`). The last two tests drive that equivalence against the canonical +// normalizer rather than asserting it. +// --------------------------------------------------------------------------- +describe('#3829 — the step\'s REVIEW.md lookup resolves a letter-suffixed phase without a shell re-pad', () => { + const { execFileSync } = require('node:child_process'); + const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs'); + const { normalizePhaseName } = require('../gsd-core/bin/lib/phase-id.cjs'); + + /** + * Pure: the indexes of every line containing `anchor`. Asserts the count so a site that is + * added, removed or renamed breaks this test loudly instead of silently narrowing what it + * covers (the step carries each anchor TWICE — one per markdown fence, each a fresh shell). + */ + function findAnchoredLineIndexes(lines, anchor, expectedCount) { + const idx = []; + lines.forEach((l, i) => { if (l.includes(anchor)) idx.push(i); }); + assert.equal( + idx.length, + expectedCount, + `expected ${expectedCount} line(s) containing ${JSON.stringify(anchor)}, found ${idx.length}`, + ); + return idx; + } + + /** Run `script` in bash with `env` merged in; never throws — returns { status, stdout, stderr }. */ + function runBash(script, env) { + try { + const stdout = execFileSync('bash', [], { + input: script, + encoding: 'utf8', + timeout: PROBE_TIMEOUT_MS, + env: { ...process.env, ...env }, + stdio: ['pipe', 'pipe', 'pipe'], + }); + return { status: 0, stdout: stdout.trim(), stderr: '' }; + } catch (e) { + return { status: e.status, stdout: String(e.stdout || '').trim(), stderr: String(e.stderr || '').trim() }; + } + } + + const stepLines = splitLines(fs.readFileSync(DISPOSITION_STEP_PATH, 'utf8')); + // TWO fences — each markdown fence is a fresh shell, so each derives and looks up for itself. + const lookups = findAnchoredLineIndexes(stepLines, 'REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"', 2); + // The derivation slice each fence runs before its lookup, taken by CONTENT rather than by a + // line number that every edit to the file would drift. + const derivStarts = findAnchoredLineIndexes(stepLines, '_pd="${PHASE_DIR:-}"', 2); + const derivations = derivStarts.map((d, n) => stepLines.slice(d, lookups[n] + 1).join('\n')); + + test('no fence re-pads the phase number with printf (fails before the fix)', () => { + // The defect shape, not the remedy: `printf "%0Nd"` applied to PHASE_NUMBER itself. + const offenders = stepLines + .map((l, n) => [n + 1, l]) + .filter(([, l]) => !/^\s*#/.test(l) && /printf\s+"%0\d*d"\s+"?\$\{?PHASE_NUMBER/.test(l)); + assert.deepEqual(offenders, [], `the step must not re-pad PHASE_NUMBER in shell: ${JSON.stringify(offenders)}`); + // And the pad it does perform is a STRING pad, not arithmetic: the canonical normalizer + // left-pads to a MINIMUM of two and otherwise preserves the run, so `$(( ))` is wrong by + // construction -- it collapsed `008` to `08` until round 14. + // Asserted as the PROPERTY, not as one spelling of it: an equivalent multi-line string pad must + // pass. What must not pass is arithmetic, or a pad that drops either carried part. + for (const d of derivations) { + const bound = d.split('\n').find((l) => /PADDED=/.test(l) && !/PADDED=""/.test(l)); + assert.ok(bound, 'the derivation must bind PADDED'); + assert.doesNotMatch(bound, /printf|\$\(\(/, `the pad must not be arithmetic: ${bound.trim()}`); + // What the pad must PRODUCE is asserted by the matrix and the property below, which execute + // it. Requiring a particular variable to appear HERE false-positives on an equivalent pad + // that routes the digit run through an intermediate -- driven, and it is why this stops at + // the defect shape. + } + }); + + test('regression control: both lookup lines are unchanged', () => { + for (const i of lookups) { + assert.equal(stepLines[i].trim(), 'REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"'); + } + }); + + test('nothing REBINDS REVIEW_FILE after the lookup', () => { + // The slices above stop AT the first anchored assignment, so on their own they cannot see a + // later line overwriting the path the fence actually consumes. Driven by an adversarial pass + // on this very test: inserting the expected lookup and then overriding it with + // REVIEW_FILE="${_pd}/WRONG-REVIEW.md" left every other assertion here green. Pin the whole + // file rather than the slice -- the only REVIEW_FILE= bindings permitted are the canonical + // lookup (once per fence) and the identity pass-through that hands it to the embedded node + // script as an env prefix. + // The predicate is deliberately wider than `^REVIEW_FILE=`: a second adversarial pass drove an + // INDENTED assignment and an `export REVIEW_FILE=...` straight through that anchor, both of + // which execute exactly like a bare one. Leading whitespace and an optional `export` are + // absorbed here so the accept-list below is what decides, not the spelling of the line. + // `+=` too: `REVIEW_FILE+=-wrong` APPENDS and executes (driven: `REVIEW_FILE=good; + // REVIEW_FILE+=-wrong` prints `good-wrong`), so an assignment-operator match that only sees + // `=` lets a real rebinding through. Third spelling found by a third adversarial pass. + const BIND_RE = /^\s*(?:export\s+)?REVIEW_FILE\+?=/; + const binds = stepLines.filter((l) => BIND_RE.test(l)).map((l) => l.replace(/^\s*(?:export\s+)?/, '')); + assert.equal(binds.filter((l) => l.startsWith('REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"')).length, 2, + 'each fence must bind the canonical lookup exactly once'); + for (const b of binds) { + const ok = b.startsWith('REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"') + || b.startsWith('REVIEW_FILE="${REVIEW_FILE}"'); + assert.ok(ok, `REVIEW_FILE is rebound to something other than the canonical lookup: ${b}`); + } + }); + + test('composition: the live derivation and lookup resolve the letter phase\'s own REVIEW.md', (t) => { + // The executable half: run the SHIPPED lines against a fixture so the validation, the padding, + // the letter carry and the path construction are exercised together. + const dir = createTempDir(); + t.after(() => cleanup(dir)); + for (const [id, status] of [['3A', 'clean'], ['8', 'issues'], ['9', 'skipped']]) { + const emitted = normalizePhaseName(id); + fs.writeFileSync(path.join(dir, `${emitted}-REVIEW.md`), `---\nstatus: ${status}\n---\n# review\n`); + for (const deriv of derivations) { + const script = [ + 'set -e', + deriv, + 'test -f "$REVIEW_FILE" || { echo "MISSING $REVIEW_FILE"; exit 3; }', + 'printf \'%s %s\' "$PADDED" "$(grep -m1 "^status:" "$REVIEW_FILE" | cut -d: -f2 | tr -d " ")"', + ].join('\n'); + const r = runBash(script, { PHASE_DIR: dir, PHASE_NUMBER: id }); + assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stdout} ${r.stderr}`); + assert.equal(r.stdout, `${emitted} ${status}`); + } + } + }); + + test('the step\'s padding agrees with the canonical normalizer across the letter matrix', (t) => { + // Driven, not argued: every shape #4748's own tests named, plus its stated regression controls. + const dir = createTempDir(); + t.after(() => cleanup(dir)); + for (const id of ['3A', '8', '9', '08', '09', '12A', '4B', '23.1.2', '03A.1.2', '1', '06', '36.14', '08.5']) { + for (const deriv of derivations) { + const r = runBash(`set -e\n${deriv}\nprintf '%s' "$PADDED"`, { PHASE_DIR: dir, PHASE_NUMBER: id }); + assert.equal(r.status, 0, `bash exited ${r.status} on ${id}: ${r.stderr}`); + assert.equal(r.stdout, normalizePhaseName(id), `padding drifted from the canonical normalizer on ${id}`); + } + } + }); + + test('property: the step\'s padding agrees with the canonical normalizer on any id it accepts', (t) => { + // The matrix above samples 13 shapes. #4748's guarantee at its ORIGINAL site was a DATAFLOW + // pin — `execute-phase.md` bound init's own `{padded_phase}`, so the lookup could not disagree + // with the canonical normalizer because it never computed anything. This step reconstructs the + // value in shell instead, so that pin is not available here and AGREEMENT is what replaces it. + // A 13-point sample cannot see a future canonical-grammar change land outside those 13 points; + // a generator can, and this is the exact blind spot rounds 3 and 5 of this PR were both about. + // Scoped to the ids the step's own guard ACCEPTS: a digit run within its 8-digit bound, an + // optional single A-Z, and optional dot segments. Milestone `N-N` forms are outside the step's + // domain, and are asserted nowhere here rather than silently passed. + const dir = createTempDir(); + t.after(() => cleanup(dir)); + // The digit run is generated as a STRING of digits, never as an integer: `String(fc.integer())` + // can never produce a LEADING ZERO, so an integer-sourced generator silently loses the `08`/`09` + // cases the finite matrix above already covered, and could never have reached `008`. That was + // this property's own first cut, and the review that caught it is why the shape is spelled out. + // Segment depth goes to four because the repo itself exercises `1.2.3.4`. The step's grammar is + // unbounded in depth; four is a bound, stated rather than implied, and it is the residual here. + const digitRun = fc.array(fc.integer({ min: 0, max: 9 }), { minLength: 1, maxLength: 8 }) + .map((ds) => ds.join('')); + const PHASE_ID = fc.tuple( + digitRun, + fc.option(fc.constantFrom(...'ABCDEFGHIJKLMNOPQRSTUVWXYZ'), { nil: '' }), + fc.array(digitRun, { maxLength: 4 }), + ).map(([run, letter, segs]) => run + letter + segs.map((s) => '.' + s).join('')); + fc.assert(fc.property(PHASE_ID, (id) => { + for (const deriv of derivations) { + const r = runBash(`set -e\n${deriv}\nprintf '%s' "$PADDED"`, { PHASE_DIR: dir, PHASE_NUMBER: id }); + assert.equal(r.status, 0, `bash exited ${r.status} on ${id}: ${r.stderr}`); + assert.equal(r.stdout, normalizePhaseName(id), `padding drifted from the canonical normalizer on ${id}`); + } + }), { numRuns: 25 }); + }); +}); diff --git a/tests/features-index-gate.test.cjs b/tests/features-index-gate.test.cjs index 461513f12..491bb0b21 100644 --- a/tests/features-index-gate.test.cjs +++ b/tests/features-index-gate.test.cjs @@ -362,6 +362,36 @@ describe('the committed docs/features/ corpus', () => { assert.equal(new Set(anchors).size, anchors.length); }); + test('no fragment declares the same REQ id twice', () => { + // WITHIN a fragment, never across the corpus. Two different features legitimately + // both carry REQ-REVIEW-01..07 — the cross-AI review feature and the code-review + // pipeline — so a corpus-wide uniqueness check would be wrong on the committed tree + // and would have to be weakened the day it first ran. A requirement list belongs to + // its feature; that is the scope of the identifier. + // + // WHY THIS EXISTS (#3829). The failure is a MERGE, not an edit. Two PRs open at once + // each append "the next" REQ number to the same list; whichever lands second is + // rebased onto a list that already used it. git merges them as different lines of one + // file and reports nothing, and neither PR's diff shows a collision — each is correct + // against the tree it was written on. Found exactly that way: #3661 took + // REQ-REVIEW-08 while #3861 also claimed it, and reverting that renumber CONSISTENTLY + // (fragment plus the regenerated projection) left `gen-features --check`, `lint:ci` + // and the pipeline suite all green with two REQ-REVIEW-08 entries standing. + // + // Nothing else in the repo reads REQ ids, so this is the only place the duplicate can + // be caught. No fragment carries one today. + const offenders = []; + for (const f of corpus.fragments) { + const ids = [...String(f.body).matchAll(/^-\s+(REQ-[A-Z0-9-]*\d)\s*:/gm)].map((m) => m[1]); + const seen = new Set(); + for (const id of ids) { + if (seen.has(id)) offenders.push(`${f.file} declares ${id} more than once`); + seen.add(id); + } + } + assert.deepEqual(offenders, []); + }); + test('groups are ordered by their lowest-ordered member', () => { const orders = corpus.groups.map((g) => g.order); assert.deepEqual([...orders].sort((a, b) => a - b), orders); diff --git a/tests/fixtures/compact-content-benchmark-baseline.json b/tests/fixtures/compact-content-benchmark-baseline.json index 3df4f5332..76931b26b 100644 --- a/tests/fixtures/compact-content-benchmark-baseline.json +++ b/tests/fixtures/compact-content-benchmark-baseline.json @@ -18,9 +18,9 @@ "reductionPct": 16.44 }, "execute-phase": { - "offTokens": 25495, - "onTokens": 23144, - "reductionPct": 9.22 + "offTokens": 25431, + "onTokens": 23080, + "reductionPct": 9.24 }, "new-project": { "offTokens": 14369, @@ -39,8 +39,8 @@ } }, "aggregate": { - "offTokens": 108922, - "onTokens": 92052, - "reductionPct": 15.49 + "offTokens": 108858, + "onTokens": 91988, + "reductionPct": 15.5 } } diff --git a/tests/fixtures/install-tree/antigravity.json b/tests/fixtures/install-tree/antigravity.json index 7ccd8cf56..0f4bea64b 100644 --- a/tests/fixtures/install-tree/antigravity.json +++ b/tests/fixtures/install-tree/antigravity.json @@ -439,6 +439,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/augment.json b/tests/fixtures/install-tree/augment.json index 8b9d86388..28827b821 100644 --- a/tests/fixtures/install-tree/augment.json +++ b/tests/fixtures/install-tree/augment.json @@ -511,6 +511,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/claude-local.json b/tests/fixtures/install-tree/claude-local.json index 34471e910..4851286c1 100644 --- a/tests/fixtures/install-tree/claude-local.json +++ b/tests/fixtures/install-tree/claude-local.json @@ -346,6 +346,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/claude.json b/tests/fixtures/install-tree/claude.json index 39c6d8781..e5baa4405 100644 --- a/tests/fixtures/install-tree/claude.json +++ b/tests/fixtures/install-tree/claude.json @@ -410,6 +410,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/cline.json b/tests/fixtures/install-tree/cline.json index 1099af2d9..36ca6956d 100644 --- a/tests/fixtures/install-tree/cline.json +++ b/tests/fixtures/install-tree/cline.json @@ -441,6 +441,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/codebuddy.json b/tests/fixtures/install-tree/codebuddy.json index baf3d5eed..c98389983 100644 --- a/tests/fixtures/install-tree/codebuddy.json +++ b/tests/fixtures/install-tree/codebuddy.json @@ -511,6 +511,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/codex.json b/tests/fixtures/install-tree/codex.json index 05e5b5dbd..5a2a598ad 100644 --- a/tests/fixtures/install-tree/codex.json +++ b/tests/fixtures/install-tree/codex.json @@ -475,6 +475,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/copilot.json b/tests/fixtures/install-tree/copilot.json index 5e040b96c..0b57dfafc 100644 --- a/tests/fixtures/install-tree/copilot.json +++ b/tests/fixtures/install-tree/copilot.json @@ -440,6 +440,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/cursor.json b/tests/fixtures/install-tree/cursor.json index a5726be4d..0b6c23b9f 100644 --- a/tests/fixtures/install-tree/cursor.json +++ b/tests/fixtures/install-tree/cursor.json @@ -439,6 +439,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/hermes.json b/tests/fixtures/install-tree/hermes.json index 32bec5ffd..d8eab7a29 100644 --- a/tests/fixtures/install-tree/hermes.json +++ b/tests/fixtures/install-tree/hermes.json @@ -439,6 +439,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/kilo.json b/tests/fixtures/install-tree/kilo.json index f490e3ed6..cc6a88a69 100644 --- a/tests/fixtures/install-tree/kilo.json +++ b/tests/fixtures/install-tree/kilo.json @@ -511,6 +511,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/kimi-code.json b/tests/fixtures/install-tree/kimi-code.json index eaa409aa1..3c3662351 100644 --- a/tests/fixtures/install-tree/kimi-code.json +++ b/tests/fixtures/install-tree/kimi-code.json @@ -440,6 +440,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/kimi.json b/tests/fixtures/install-tree/kimi.json index 86c24e84c..f4b7be131 100644 --- a/tests/fixtures/install-tree/kimi.json +++ b/tests/fixtures/install-tree/kimi.json @@ -447,6 +447,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/opencode.json b/tests/fixtures/install-tree/opencode.json index fda72eb19..6cfe42bd1 100644 --- a/tests/fixtures/install-tree/opencode.json +++ b/tests/fixtures/install-tree/opencode.json @@ -511,6 +511,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/pi.json b/tests/fixtures/install-tree/pi.json index e94a789e5..e52bc55d3 100644 --- a/tests/fixtures/install-tree/pi.json +++ b/tests/fixtures/install-tree/pi.json @@ -241,6 +241,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/qwen.json b/tests/fixtures/install-tree/qwen.json index 120f45621..5c6975ee6 100644 --- a/tests/fixtures/install-tree/qwen.json +++ b/tests/fixtures/install-tree/qwen.json @@ -439,6 +439,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/trae.json b/tests/fixtures/install-tree/trae.json index 5e6ffc526..632d28b32 100644 --- a/tests/fixtures/install-tree/trae.json +++ b/tests/fixtures/install-tree/trae.json @@ -439,6 +439,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/windsurf.json b/tests/fixtures/install-tree/windsurf.json index 09d5101ad..3027de5d4 100644 --- a/tests/fixtures/install-tree/windsurf.json +++ b/tests/fixtures/install-tree/windsurf.json @@ -367,6 +367,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/fixtures/install-tree/zcode.json b/tests/fixtures/install-tree/zcode.json index f327b6d7e..35a953506 100644 --- a/tests/fixtures/install-tree/zcode.json +++ b/tests/fixtures/install-tree/zcode.json @@ -511,6 +511,7 @@ "gsd-core/workflows/eval-review.md", "gsd-core/workflows/execute-phase.md", "gsd-core/workflows/execute-phase/detail/elaboration.md", + "gsd-core/workflows/execute-phase/steps/code-review-disposition.md", "gsd-core/workflows/execute-phase/steps/codebase-drift-gate.md", "gsd-core/workflows/execute-phase/steps/completion-reconciliation.md", "gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md", diff --git a/tests/nsegment-phase-grammar.test.cjs b/tests/nsegment-phase-grammar.test.cjs index bf4365700..2afe7e80f 100644 --- a/tests/nsegment-phase-grammar.test.cjs +++ b/tests/nsegment-phase-grammar.test.cjs @@ -336,6 +336,9 @@ const EXECUTE_PHASE = path.join(__dirname, '..', 'gsd-core', 'workflows', 'execu const COMPLETION_RECONCILIATION = path.join( __dirname, '..', 'gsd-core', 'workflows', 'execute-phase', 'steps', 'completion-reconciliation.md', ); +const CODE_REVIEW_DISPOSITION = path.join( + __dirname, '..', 'gsd-core', 'workflows', 'execute-phase', 'steps', 'code-review-disposition.md', +); const TDD_REF = path.join(__dirname, '..', 'gsd-core', 'references', 'tdd.md'); const AUTONOMOUS = path.join(__dirname, '..', 'gsd-core', 'workflows', 'autonomous.md'); const PLAN_REVIEW_CONVERGENCE = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-review-convergence.md'); @@ -456,55 +459,88 @@ describe('#4748 — the $((10#$PHASE_INT)) split sites carry a letter suffix int } }); -describe('#4748 — execute-phase.md resolves the REVIEW.md path from init\'s padded_phase, not a shell re-pad', () => { - const lines = splitLines(fs.readFileSync(EXECUTE_PHASE, 'utf8')); - const [i] = findAnchoredLineIndexes(lines, 'REVIEW_FILE="${PHASE_DIR}/${PADDED}-REVIEW.md"', 1); - const paddedLine = lines[i - 1].trim(); +describe('#4748 — the code-review gate resolves the REVIEW.md path from a letter-safe padded phase, not a shell re-pad', () => { + // #3829 moved this lookup out of `execute-phase.md` and into the lazily-read step file below. + // The inline block did not fit under ADR-857's frozen pre-phase-6 ceiling (93600), which the + // parent now clears by 139 bytes, so it cannot be restored in place. #4748's property is + // unchanged and is asserted here against the site that now performs the lookup. + // + // ONE of this block's original four assertions was a property of the INLINE site rather than of + // the lookup, and does not survive the move: the `PADDED="{padded_phase}"` literal binding. The + // step derives PADDED itself — validating PHASE_NUMBER for shape and traversal, then padding the + // digit run as a STRING and carrying the letter and dot segments verbatim — so there is no + // literal binding left to pin, and agreement with the canonical normalizer is what replaces it. + // + // The composition run DID come back, below, and an earlier cut of this block was wrong to drop it + // on the grounds that mirroring would duplicate the PR's own coverage. A STATIC assertion cannot + // hold a BEHAVIOURAL property; at the original site it could, because the property was a literal + // binding. So this block keeps deterministic ownership of #4748 by EXECUTING the step's own + // derivation over a fixed id list. The PR's fast-check property in + // `tests/code-review-pipeline-regression.test.cjs` is a different instrument over the same + // contract — generated ids rather than a fixed list — and it reaches divergences this one does + // not: a letter outside the fixed list leaves this block green. + const stepLines = splitLines(fs.readFileSync(CODE_REVIEW_DISPOSITION, 'utf8')); + const lookupIdx = findAnchoredLineIndexes(stepLines, 'REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"', 2); - test('the PADDED binding directly above the lookup reads {padded_phase} (fails before the fix)', () => { - // `printf "%02d"` cannot pad `03A` (prints `03`, exits 1) — and cannot - // even re-pad an already-padded `08` (bash reads it as octal, prints - // `00`). `init execute-phase` now emits `padded_phase` through the - // canonical normalizer, so the workflow binds it instead of re-deriving. - assert.ok(paddedLine.startsWith('PADDED='), `line above the lookup must bind PADDED: ${paddedLine}`); - assert.equal(paddedLine, 'PADDED="{padded_phase}"'); + test('no fence pads the raw phase number with printf (fails before the fix)', () => { + // `printf "%02d"` cannot pad `03A` (prints `03`, exits 1) — and cannot even re-pad an + // already-padded `08` (bash reads it as octal, prints `00`). Every binding must pad the + // DIGIT RUN, never PHASE_NUMBER itself. + const offenders = stepLines + .map((l, n) => [n + 1, l]) + .filter(([, l]) => !/^\s*#/.test(l) && /printf\s+"%0\d*d"\s+"?\$\{?PHASE_NUMBER/.test(l)); + assert.deepEqual(offenders, [], `no fence may printf-pad PHASE_NUMBER: ${JSON.stringify(offenders)}`); }); - test('regression control: the lookup line itself is unchanged', () => { - assert.equal(lines[i].trim(), 'REVIEW_FILE="${PHASE_DIR}/${PADDED}-REVIEW.md"'); - }); - - test('composition: the value init emits, substituted into the live lookup lines, resolves the letter phase\'s own REVIEW.md', (t) => { - // The model substitutes `{padded_phase}` from the init JSON, which is - // `normalizePhaseName(phase_number)` (src/init.cts). Do that substitution - // here and run the three live lines against a fixture, so the emitted - // value, the binding, the path construction and the status extraction are - // exercised together — the executable half of a `{template}` site. - const { normalizePhaseName } = require('../gsd-core/bin/lib/phase-id.cjs'); - const { createTempDir, cleanup } = require('./helpers.cjs'); - const dir = createTempDir(); - t.after(() => cleanup(dir)); - for (const [id, status] of [['3A', 'clean'], ['8', 'issues'], ['9', 'skipped']]) { - const emitted = normalizePhaseName(id); - fs.writeFileSync(path.join(dir, `${emitted}-REVIEW.md`), `---\nstatus: ${status}\n---\n# review\n`); - const script = [ - 'set -e', - paddedLine.replace('{padded_phase}', emitted), - lines[i].trim(), - lines[i + 1].trim(), - 'printf \'%s %s\' "$PADDED" "$REVIEW_STATUS"', - ].join('\n'); - assert.ok(lines[i + 1].includes('REVIEW_STATUS='), `line after the lookup must extract REVIEW_STATUS: ${lines[i + 1]}`); - const r = runBash(script, { PHASE_DIR: dir }); - assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`); - assert.equal(r.stdout, `${emitted} ${status}`); + test('no lookup pads the phase through arithmetic (fails before the fix)', () => { + // THE DEFECT SHAPE, which is what a static gate can actually hold. The canonical normalizer + // left-pads the digit run to a MINIMUM of two and otherwise PRESERVES it (`008` -> `008`), so an + // arithmetic pad is wrong by construction -- `$((10#$_dig))` collapses every longer leading-zero + // run. Deliberately NOT a pin on one spelling of the remedy: an equivalent multi-line string pad + // must pass here, and correctness is asserted by execution below rather than by shape. + for (const i of lookupIdx) { + const bound = stepLines.slice(0, i).reverse() + .find((l) => /PADDED=/.test(l) && !/PADDED=""/.test(l)); + assert.ok(bound, `the lookup at line ${i + 1} has no PADDED binding above it`); + assert.doesNotMatch(bound, /printf|\$\(\(/, + `the pad must not be arithmetic -- arithmetic collapses 008 to 08: ${bound.trim()}`); } }); - test('the workflow\'s init parse list names padded_phase, so the binding is not a literal (fails before the fix)', () => { - // A `{field}` token is substituted from the init JSON only for fields the - // workflow tells the model to parse; `phase_number` is on that list and - // `padded_phase` was not (adversarial review, claim 2). + test('composition: each fence\'s live derivation resolves the id init would emit', () => { + // #4748's property, asserted the way it has to be at THIS site. At the original site the gate + // could be static because the property was a literal binding of init's own `{padded_phase}`; + // here the step derives the value, so the property is behavioural and only execution can hold + // it. Runs the SHIPPED derivation slice of BOTH fences against the canonical normalizer. + const { normalizePhaseName } = require('../gsd-core/bin/lib/phase-id.cjs'); + const starts = []; + stepLines.forEach((l, n) => { if (l.includes('_pd="${PHASE_DIR:-}"')) starts.push(n); }); + assert.equal(starts.length, lookupIdx.length, 'each lookup must have its own derivation slice'); + const derivations = starts.map((d, n) => stepLines.slice(d, lookupIdx[n] + 1).join('\n')); + // `008` is the case the arithmetic pad got wrong and no prior fixture covered. + for (const id of ['3A', '8', '9', '08', '008', '0008A', '23A.1.2']) { + for (const deriv of derivations) { + const out = execFileSync('bash', [], { + input: `set -e\n${deriv}\nprintf '%s' "$PADDED"`, + encoding: 'utf8', + timeout: TIMEOUT, + env: { ...process.env, PHASE_DIR: '/tmp', PHASE_NUMBER: id }, + }); + assert.equal(out, normalizePhaseName(id), `the step disagreed with the normalizer on ${id}`); + } + } + }); + + test('regression control: the lookup lines themselves are unchanged', () => { + for (const i of lookupIdx) { + assert.equal(stepLines[i].trim(), 'REVIEW_FILE="${_pd}/${PADDED}-REVIEW.md"'); + } + }); + + test('the workflow\'s init parse list still names padded_phase (fails before the fix)', () => { + // A `{field}` token is substituted from the init JSON only for fields the workflow tells the + // model to parse. This one is a property of `execute-phase.md` and the move does not touch it. + const lines = splitLines(fs.readFileSync(EXECUTE_PHASE, 'utf8')); const [p] = findAnchoredLineIndexes(lines, 'Parse JSON for: `executor_model`', 1); assert.match(lines[p], /`phase_number`, `padded_phase`,/); }); diff --git a/tests/prompt-injection-scan.security.test.cjs b/tests/prompt-injection-scan.security.test.cjs index 87f70245f..75620876b 100644 --- a/tests/prompt-injection-scan.security.test.cjs +++ b/tests/prompt-injection-scan.security.test.cjs @@ -101,6 +101,24 @@ const SIZE_ONLY_WORKFLOWS = new Set([ // injection scanned. Splitting it per the progressive-disclosure pattern is the real // fix and is worth its own change. 'gsd-core/workflows/quick.md', + // ~52K after #3829's disposition ledger. Same shape as the two above: this file sat at 44,466 + // chars on next — 89% of the 50,000-char prompt-stuffing threshold — so the feature approved on + // the issue could not land in it without tripping this. The round's committed peak was 59,246 + // chars; it was cut back to ~52K before this entry was added. + // + // The overshoot is ~2.3K, and this comment is deliberately precise about that because two earlier + // versions of it were not. The first claimed the added CODE alone exceeded the line (it does not). + // The second claimed fitting under the threshold would need "essentially all remaining + // explanation" removed — also false: the gap is ~2.3K against several times that in + // round-attributed in-fence commentary, so it is reachable by cutting a fraction of it. The + // honest trade is therefore a JUDGEMENT, not an impossibility: the commentary documents logic + // that five review passes found defects in, and this file's house style is heavy in-fence + // documentation, so the maintainers may reasonably prefer the cut to the exemption. Measure with + // the scanner's own normalization (CRLF→LF, `src/security.cts`) — byte counts read ~57 high here. + // Size-only: still fully injection scanned, exactly like the two above. + // Splitting it per the progressive-disclosure pattern is the real fix and is worth its own change + // — it is already a step file extracted from execute-phase.md for that same reason. + 'gsd-core/workflows/execute-phase/steps/code-review-disposition.md', ]); // ─── Scanner ────────────────────────────────────────────────────────────────