d7b5b2c2b67b257c65a2a7f24dd0aa4105b0f5ff
3287 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d7b5b2c2b6 |
fix(#4906): migrate the ROADMAP.md Plans: line onto the PlanningDoc seam — Phase 2 (#4933)
* feat(#4906): migrate the ROADMAP.md **Plans:** line onto the PlanningDoc seam Phase 2 of epic #4906. Migrates the two writers of ROADMAP.md's Plans field onto the seam ADR-4910 locks, and deletes both bespoke regexes per Decision 2 — a correct copy of a rule the seam now owns is the same divergence risk as an incorrect one. src/phase.cts's mutateMilestonePhase carried planCountBodyPattern, one capture group, replace-to-end-of-line: this is #4852, still live before this change. src/roadmap.cts's cmdRoadmapUpdatePlanProgress carried the correct three-arm sibling (the #2853/#3584 correction) that phase.cts never adopted. Both now call findField/setFieldValue/serialize against a PlanningDoc parsed from the same milestone/phase-confined substring their existing withPhaseSection / replaceInCurrentMilestone wrappers already compute — those confinement wrappers are unchanged, only the field-write mechanism inside them moved. Verified end-to-end through the real commands against real fixtures, not against an isolated reimplementation of the classification logic: cmdRoadmapUpdatePlanProgress and cmdPhaseComplete both preserve a trailing human annotation across a real count rewrite, and both leave a bracketed human annotation (the #3584 Finding A discriminator) untouched. Found and fixed inline, in the already-merged src/planning-document.cts, rather than deferred: BOLD_FIELD_RE recognized only **Label:** (colon inside the closing bold). gsd-core/templates/roadmap.md ships every field, Plans included, as **Label**: (colon outside) -- migrating roadmap.cts onto the seam as it stood would have silently regressed real generated ROADMAP.md files back to the bug this migration exists to remove. Widened to recognize both spellings; deliberately NOT widened to a bare unbolded Label: form, which would register ordinary prose as a spurious field. roadmap.cts's writer also recognizes a bare singular/plural count (1 plan / 3 plans, no fraction) as an existing count token to overwrite, not template- placeholder or freeform prose -- the template's own single-plan-phase shape and #3584 Finding B's fix. Preserved exactly; this shape is easy to drop by only porting the more common fraction form. Two of Phase 2's three originally-cited defects turned out to be already fixed on next, independent of this epic, and are struck via a dated ADR amendment rather than silently narrowed: #4862 (stateReplaceField's own anchoring hardening already preserves sibling fields) and #4499 (spliceFrontmatter's per-key preservation already keeps block sequences byte-identical). Both reproduced against the built module before being struck, not assumed. STATE.md's field-write engine is re-scoped out of this phase entirely -- not because of its get_impact rating alone (measured the same way, the two sites THIS phase keeps are also CRITICAL, and an earlier draft of the amendment claimed otherwise without checking; corrected) but because updateCore is a multi-field transaction with frontmatter sync and preservation reconciliation that does not map onto PlanningDoc's node model, where migrating it would mean designing that model, not calling an existing seam function. Refs #4852 Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4906): preserve a trailing annotation with no em-dash separator Test authoring surfaced a real regression against an existing #3584 fixture: `**Plans**: 0/1 plans executed (11-16 are gap closure from VERIFICATION)` -- a parenthetical annotation glued on with a bare space, no em-dash -- was left completely untouched by the migrated code instead of being rewritten with the count updated and the parenthetical preserved. Root cause: planning-document.cts's parseBoldFieldLine splits a field's value from its trailing annotation only on the literal " -- " separator. An em-dash-separated annotation already lives outside `value` in `trailingSpan`, untouched by setFieldValue regardless -- that path was never broken. A parenthetical with no em-dash has nowhere to go but inside `value`, and the migrated classification required the WHOLE value to match a count-token shape exactly, so this case fell into "leave untouched." Fixed in the migrated call sites, not in the seam: prefix-match the count token against the field's current value, then re-glue whatever textual suffix follows WITHIN that value onto the new count text before writing. Correct for both shapes with no special-casing -- the em-dash case's suffix-within-value is empty by construction (the annotation already lives outside value), the parenthetical case's suffix is exactly the glued content, preserved verbatim. Deliberately not fixed by widening planning-document.cts's separator grammar to also recognize a bare-space-then-parenthesis: that seam is already-merged and already-tested, and guessing at an open-ended set of annotation shapes at the seam level is exactly what isTemplatePlaceholder already avoids by staying caller-side. "What counts as a Plans-field count token" is domain knowledge about this one field. phase.cts's writePlansField had no arm-2/arm-3 classification before this migration -- its original regex unconditionally overwrote whatever value was present. That unconditional-overwrite behavior is preserved exactly for values with no recognizable count-token prefix; only the recognized-count case gained suffix preservation, matching what phase.cts actually did before. Verified end-to-end via the real cmdRoadmapUpdatePlanProgress and cmdPhaseComplete commands against real fixtures for both the parenthetical and em-dash shapes at both sites. Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4906): cover the migrated Plans-line writers at both sites 15 cases across tests/phase.test.cjs and tests/roadmap.test.cjs, one per row of the phase test matrix, extending the existing describe blocks and helper functions those files already use for these two commands rather than building parallel fixtures. Covers: trailing-prose preservation on a real count rewrite at both sites (the #4852 regression, and the #2853/#3584 non-regression); zero-trailing- content boundary; the bracketed-template-placeholder vs bracketed-human- annotation discriminator (#3584 Finding A); a missing Plans field not crashing either command; an unrelated unreadable sibling node in the same confined section surfacing rather than corrupting the field; confinement holding across sibling phases and milestones; a round trip through the command's own read path; CRLF safety; and a parity assertion that both sites now produce identical Plans-line text for identical inputs, proving one shared mechanism rather than source-grepping for the deleted regex literals. Authoring caught a real regression before it could land silently: an existing #3584 fixture using a parenthetical annotation with no em-dash separator failed against the first version of the migration. Reported rather than edited to match the broken behavior -- see the paired fix commit. That existing test needed no changes once the fix landed; its assertion was verified independently against the real CLI before this commit. Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4906): give phase.cts's Plans writer the same arm-2/arm-3 classification as roadmap.cts Isolated adversarial review (mandatory orthogonal review pass) executed writePlansField against `[Deferred pending re-scope]` and the fresh-template placeholder wording and found the first version of this migration kept phase.cts's OLD unconditional-overwrite behavior for the no-count-prefix case instead of adopting the same isTemplatePlaceholder / arm-3-untouched classification roadmap.cts's sibling site already uses. A bracketed human annotation was being silently rewritten to a new count -- a real content-destroying regression against this phase's own design-doc behavior table, not an accepted trade-off, and exactly the kind of divergence between the two sites Decision 2's parity requirement exists to eliminate. writePlansField now runs the same template-placeholder check and "no count prefix and not a placeholder => leave untouched" branch before ever calling setFieldValue. Added two site-1 tests (rows 4 and 6 of the phase test matrix) mirroring the existing site-2 coverage for this exact discriminator. Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4906): restore plain (non-bold) Plans: line support and fix a test helper mismatch gsd-test (mandatory verification, run before every push) surfaced two real defects the migration's manual CLI checks had not caught: 1. #1163 regression: a hand-edited/pre-template ROADMAP.md can carry a PLAIN (non-bold) `Plans:` line rather than the canonical `**Plans**:`/ `**Plans:**` bold field. The old deleted regexes tolerated this shape; the seam's BOLD_FIELD_RE is deliberately bold-only (widening it would register ordinary prose like "Note: see below" as a spurious field seam-wide), so the migrated writers silently no-op'd on it instead of updating the count -- a real, previously-tested behavior lost. Fixed with a caller-side fallback in both src/roadmap.cts (where the failing #1163 test lives) and src/phase.cts (added for parity, per Decision 2 -- the two sites should not diverge on which legacy shapes they tolerate): when findField finds no boldField Plans node, look for a plain `Plans:` line directly and apply the same arm-1/2/3 classification against it. This is domain knowledge about one field's legacy tolerated shape, the same class of thing isTemplatePlaceholder already keeps caller-side rather than seam grammar. 2. tests/roadmap.test.cjs's new rows 4/5 (site 2) seeded the colon-outside spelling (`**Plans**: ...`) but their plansLineIn() helper only matched colon-inside (`**Plans:**`), so both assertions compared against `undefined` regardless of whether the write logic was correct -- a test-authoring bug, not a source defect. Fixed the helper to recognize both BOLD_FIELD_RE spellings, matching what the production code actually supports. Verified via a real gsd-test run before this fix (outcome: failed, 5 failures, all in tests/roadmap.test.cjs) and will be re-verified via a real gsd-test run on this commit before push, per this repo's non-rationalization rule: a red gate is fixed, never explained away. Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4906): backfill the Phase 2 changeset fragment's PR number pr:0 -> pr:4933, now that gh api POST /pulls has returned the real number. Refs #4906 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1e3e1f7cd8 |
enhance(#4570): allow disabling planner stall detection (#4585)
* enhance(#4570): allow disabling planner stall detection * docs(#4570): add changelog fragment Emitted-Drift-Ack-Growth: plan-phase.md — the explicit opt-out gate covers all five planner and checker spawn classes Emitted-Drift-Ack-Growth: settings-advanced.md — the toggle prompt and bounded-recovery warning expose the new setting * chore(#4570): refresh compact-content baseline * fix(#4570): preserve default-on watchdog fallback * docs(#4570): qualify the chunked-mode orchestrator rules with the toggle The two chunked-planning-mode stall-watch imperatives read as unconditional, with the opt-out stated only in a following bullet. State the PLANNER_STALL_DETECTION_ENABLED condition inline, matching the three sites already qualified in plan-phase.md. * chore(#4570): refresh compact-content baselines against rebased next * fix(#4570): sync planner stall launcher Keep the stall-detection helper aligned with the canonical runtime launcher. * chore(#4570): refresh compact baseline after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
6486647626 |
fix(#4949): cap unmeasured files per Windows conformance chunk (#4950)
* fix(#4949): cap unmeasured files per Windows conformance chunk The windows conformance CI lane keeps going red every few updates: scripts/run-tests.cjs kills a chunk at the 600s per-chunk backstop with zero failing tests, pure slowness. Both recent incidents ( |
||
|
|
af822a8024 |
fix(#4776): resolve the artifact-exists prompt under --auto (#4832)
* fix(#4776): resolve the artifact-exists prompt under --auto /gsd-ui-phase <phase> --auto stopped at 'UI-SPEC.md already exists for Phase {N}. What would you like to do?' whenever the file was on disk — which is most often after an earlier run wrote the contract as a draft and ended before its checker ran, exactly the state a re-run exists to verify. Step 4 had no --auto arm; step 9.5 in the same file has had one since it was added, which is how the drift went unnoticed. Step 4 now auto-selects Skip: the existing UI-SPEC is left untouched and the run proceeds to the checker. Skip is the only non-destructive choice — Update re-runs the researcher, which rewrites the whole contract and drops answers a person already recorded in it, and View exits without verifying anything. spec-phase.md's artifact-exists arm auto-selected 'Update it', which is the same defect with the opposite sign: an unattended run regenerating a spec nobody is watching. Per the decision recorded on #4776 — an unattended run reuses an existing artifact rather than regenerating it — it now auto-selects Skip and leaves the spec unchanged. The max-revision-iterations escalation (Force approve / Edit manually / Abandon) is deliberately untouched and pinned by a test: accepting blocking findings is a decision a person makes. Closes #4776 * chore(#4776): add changeset fragment Emitted-Drift-Ack-Growth: ui-phase.md — --auto arm added to the existing-UI-SPEC branch (#4776) Emitted-Drift-Ack-Growth: spec-phase.md — --auto arm reworded to reuse the existing SPEC (#4776) * fix(#4776): extend the reuse-as-is --auto fix to the 3 sibling files The PR's original scope claim -- that ai-integration-phase.md, eval-review.md and ui-review.md were "scoped by triage to their own issues" -- was false; no such issues existed, and it contradicted #4776's own most recent (2026-09-16) triage comment, which explicitly widened the fix to require all 5 files under one recommended fix. Applies the same reuse-as-is pattern: ai-integration-phase.md mirrors ui-phase.md's 3-way Update/View/Skip shape (auto-selects Skip); eval-review.md and ui-review.md have only Re-audit/View (auto-selects View, the only non-regenerating choice). None of the three had any prior --auto handling at all -- each has exactly one AskUserQuestion call site total, and it is the one this fix resolves, so an --auto run through any of them no longer stalls anywhere. Emitted-Drift-Ack-Growth: ai-integration-phase.md — the --auto arm reusing an existing AI-SPEC is the deliverable (#4776) Emitted-Drift-Ack-Growth: eval-review.md — the --auto arm reusing an existing EVAL-REVIEW is the deliverable (#4776) Emitted-Drift-Ack-Growth: ui-review.md — the --auto arm reusing an existing UI-REVIEW is the deliverable (#4776) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b90eef28e8 |
enhance(#3829): report code review severity counts and record a per-finding disposition (#3861)
* enhance(#3829): report code review severity counts and record a per-finding disposition `code_review_gate` extracted `status:` from REVIEW.md's frontmatter and discarded the `critical`/`warning`/`info`/`total` values sitting in the same range, so its output was byte-identical for a review with one `info` finding and a review with a Critical. Nothing anywhere recorded what happened to a finding: no file under `gsd-core/workflows/` branches on `issues_found`, and `gsd-verifier.md` has zero references to REVIEW.md. A phase therefore reached `phase.complete` with Criticals standing and no trace they had been seen. Both halves were approved on the issue; the gate stays advisory. A — severity surfacing, in `execute-phase.md`. The gate states the breakdown it already parsed, accepting `blocker:` as the documented tier-equivalent of `critical:`. The breakdown is shown only when all four counts are numeric (`REVIEW_COUNTS_OK`); otherwise the countless message stands, because gating on the total alone still emits `6 findings — critical` for a review carrying a total and nothing else. Frontmatter is extracted by an `awk` that emits only when it saw the CLOSING delimiter, after stripping CR. A `sed` range re-opens on a body `---` and runs to EOF: first-match protects a key the frontmatter always carries, but not an optional one, so a review with no `findings:` block and a body `total:` line would have reported the body's number. An unterminated block would leak the whole body the same way. Every read is guarded and `|| true`-terminated. This step is advisory, and under `set -e`/`pipefail` a non-matching `grep` exits 1 — an assignment whose command substitution fails would take the step down with it. A REVIEW.md that is missing, a directory, or unreadable now leaves the counts empty and execution continues. B — per-finding disposition, in a new lazily-read step file, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`, referenced from the gate in the established plain read-and-execute form. One row per finding ID, defaulting to `open`, and: - `fixed`/`skipped` are reconciled from REVIEW-FIX.md, whose section headings are matched WHOLE — a prefix match let `## Fixed Issues Verification` classify every finding beneath it as fixed — and only when the fix report names the SAME finding. Finding ids are reused across re-reviews, so matching on the id alone let a stale fix report declare a brand-new CR-01 already fixed. - headings inside fenced blocks are ignored; a quoted example is not a finding. - an id listed under both sections resolves by first occurrence, not row order. - a recorded disposition is preserved together with the reason in its Source cell, escaped pipes included, and a hand-mangled row missing its trailing pipe still keeps its decision. - a decided finding the current review no longer reports is CARRIED and marked; `--auto` rewrites REVIEW.md each iteration, so this is routine, and dropping the row would erase the record that it was seen. An untriaged `open` row for a vanished finding is not carried. A review reporting nothing still reconciles an existing ledger rather than freezing it. - a run that changes no disposition rewrites nothing, so a re-executed phase does not produce a docs commit whose only delta is a timestamp. The record is a sibling artifact, not a section inside REVIEW.md: `--auto`'s re-review loop rewrites REVIEW.md every iteration, so a ledger kept inside it would not survive the next pass, and REVIEW.md has a single writer that this step is not. B lives in an extracted step file because `execute-phase.md` was 91,493 bytes against a 98,304 hard cap the size-budget test calls a red line, and because `scanWiredKinds` caps a call site's dispatch-coverage region at 6000 characters — an inline version pushed the `kind == "gate"` paragraph out of that window, which silently drops `gate` from the covered set and fails `gen-capability-registry --check` while pointing at the capability rather than at the prose that displaced it. Extraction is what that size test's own message prescribes, and it leaves the file at 93,854 bytes. The tests execute the shipped script rather than modelling it. Three adversarial review rounds each refuted "the mirror is faithful", and mutation testing agreed: with a hand-written model, deleting the carried-row logic from the shipped file turned nothing red. The suite now extracts the embedded script — undoing exactly the four shell double-quote escapes — and runs it, so all ten mutations of its behaviour are caught. * chore(#3829): set changeset fragment pr to 3861 * fix(#3829): keep execute-phase.md under both size ceilings and propagate the launcher probe The first push failed `full test (macos-latest, 24, shard 3/3)`. Two things it caught that the CI-selected scope for this diff does not run, and that I therefore did not run either: 1. `execute-phase.md` is governed by TWO ceilings, not one. The XL hard cap in `tests/workflow-size-budget.test.cjs` (98304) was satisfied at 95179, but the frozen ADR-857 pre-phase-6 ceiling in `tests/claude-orchestration.test.cjs` (93600) was not. The whole budget from base is 2107 bytes, which the inline reporting half alone did not fit. That half now lives in the extracted step file alongside the disposition half, and the parent carries only the paragraph that reads and executes it — 91529 bytes, 36 over base. 2. The step file calls `gsd_run`, so it owes the hermes runtime-home probe that `tests/runtime-launcher-parity.test.cjs` (E) requires of every workflow file that does. Propagated with `node scripts/sync-runtime-launcher.cjs`, the remedy that test names. Verified with the FULL unit suite this time rather than the scoped selection — 14 shards, 0 failures — plus `npm run lint:ci`, and a re-run of the ten mutations of the shipped disposition script, all still caught. * fix(#3829): stop the disposition step instructing the agent to execute itself Blocker 1 and Minor 7 of the round-1 review are one defect. The step file carried a copy of execute-phase.md's pointer paragraph, so it named its own path as something to "read and execute" — unbounded self-recursion at runtime — and that copy is also the duplicated paragraph, sitting immediately above the full instruction it duplicates. Removing the copy resolves both. execute-phase.md remains the only surface that points here, which is what it always intended. Two structural tests guard it. Both are red against the pre-fix file: no behavioural test could see either defect, because they execute the node script through the process seam and so never read the prose that tells the agent what to load. * fix(#3829): re-derive the ledger paths in the block that uses them Blocker 2. The disposition block reads REVIEW_FILE, DISPOSITION_FILE and PADDED, all derived in the step's FIRST shell block. Each fenced block is dispatched as its own shell, so all three are empty by the time the second block runs: the ledger write lands on a bare `-REVIEW-DISPOSITION.md` path and the review read finds nothing. The step then reports success having produced no artifact — the feature's central acceptance criterion, silently unmet, with no error to notice. The tell was already in the file: the gsd_run shim preamble is re-emitted in the second block for exactly this reason. These three paths belong beside it, and now are. The guard test asserts the general property rather than the instance — every block derives what it reads, inheriting only the step's declared inputs (PHASE_DIR, PHASE_NUMBER) — so a third block added later cannot reintroduce it. Red against the pre-fix file. * test(#3829): assert the counts mirror against the shipped shell, and execute its guards Major 4, with Minor 6 and part of Minor 9. The disposition builder stopped being a mirror three rounds ago, and the reason given then was that a hand model of a shell-embedded script drifts while the tests stay green. parseGateCounts kept its mirror anyway. That argument does not stop applying at the boundary between the step's two shell blocks, so the mirror now loses its authority: it is asserted against the shipped awk and greps, run under `set -euo pipefail` in a real shell, across every fixture it is exercised on. Negative-controlled in both directions. Dropping `blocker:` from the mirror alone fails the parity test; replacing the shipped awk with the leaky `sed -n '/^---$/,/^---$/p'` range fails it on the unterminated-frontmatter fixture. Divergence in either half is now red, which is what the finding asks for. Skipped on win32, where there is no bash to compare against. Minor 6: the zero-count edge is covered — `0` is numeric, so a zero-finding review reports `0 findings — 0 critical, …` rather than falling back to the countless form. A guard written against truthiness would have failed here silently, and now cannot. Minor 9, partially: running the block makes its advisory guards behavioural, so the four `src.includes()` assertions that stood in for them are retired — a missing and an unreadable REVIEW.md are now proven not to abort under `set -e`, rather than asserted to contain a string. The remaining docs-parity assertions are kept deliberately; see the PR discussion. * test(#3829): add the render/re-parse fixed-point property for the ledger Major 3. RULESET.TESTS.property-based-testing asks for at least one fc property on a parsing/transformation contract, and the ledger is one with a fixed point stated in its own prose: re-running the gate preserves every disposition except `open`, and rewrites nothing when nothing changed. Two properties, both driving the SHIPPED script rather than a model of it: idempotency — a second run reports `unchanged` and leaves the file byte-identical. Without it, the timestamp alone dirties the tree on every phase re-run. round-trip — a hand-recorded decision AND the reason beside it survive render -> re-parse -> render, escaped pipes included. The Source cell is where a human writes why something was deferred, so losing it loses the only thing that instruction asks for. Negative-controlled per property: disabling the unchanged-check fails the first and only the first; discarding the carried source cell fails the second and only the second. numRuns is 40 rather than the shared 200 because each case spawns the shipped script twice through the process seam. The seed stays pinned, so a failure still reproduces; the deviation is stated in the file header rather than made silently. * fix(#3829): state a stale fix-report match instead of dropping it silently Minor 5, plus the finding-id census this round owes. Exact-title coupling stays — ids are reused across re-reviews, so a stale REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed. What changes is the silence. A row that stays `open` because the report named a different finding under the same id is indistinguishable, to any reader, from a row that stays open because no report mentioned it. The gate now names the ids it could not reconcile, on both report paths, and stays advisory throughout. The census (RV4, self-found — the review did not ask for this). The script enumerates finding-id prefixes in three places: the heading matcher, the ledger re-parser, and the severity map's keys. The DOMAIN those enumerate is owned elsewhere — gsd-code-reviewer.md's body template and its Label-equivalence paragraph — so it can acquire a member without this script changing. Reached: CR, BL, WR, IN — 4 of 4, all present. Not reached: none today. What follows if that changes is the payload: an unlisted prefix is not mis-tiered, it is INVISIBLE — the finding never enters the order list and gets no row at all, so the artifact silently under-reports the review it is meant to record. Adding a prefix to two of the three copies fails the same way, and additionally drops carried rows on the next run. Two guards rather than a rewrite: hoisting the alternation into one constant means rebuilding three regexes inside a double-quoted shell string, which is the exact class of edit that produced both of this round's blockers. The guards make the drift loud instead, and are negative-controlled against each of the two ways it can happen. * docs(#3829): keep the feature reference descriptive, not instructional Minor 8 — a Diataxis mode mix. "Set `deferred` by hand and put the reason in the Source cell" is a how-to instruction sitting in a reference doc. The information belongs there (a reader needs to know the field exists and what preserves it); the imperative does not. Rewritten to describe the field instead: `deferred` is the one disposition the gate never writes, and the reason recorded beside it survives re-runs. The same pass records Minor 5's new behaviour, since the reference described the title coupling but not what happens when it misses. The imperative form is kept where it belongs — inside the ledger the gate renders, which is where a reader meets the field and the only place an instruction has an audience. docs/FEATURES.md regenerated from it; `gen-features.cjs --check` is green. * fix(#3829): close six defects found by reviewing this round's own fixes None of these came from the maintainer's review. They came from adversarially reviewing the five commits above before pushing them, and two are worse than anything the round was opened to fix. 1. A foreign fence marker swapped an example for a finding. The heading scanner toggled fenced/not-fenced on ANY fence marker, so a ~~~ line inside a ``` example closed the fence and the example's real close reopened one. Driven: a review quoting ~~~ inside a fenced example produced a ledger recording CR-77, the illustration, and omitting CR-01, the actual finding. A confidently-written artifact wrong in both directions at once. The open marker's character and length are now remembered, and a fence closes only on the same character at least as long, per CommonMark. 2. The disposition block had no status gate at all. The prose above it says it runs only when the review reports issues — but block 1 computes REVIEW_STATUS, emits nothing, and its shell is discarded, so no later block could act on that condition even in principle. A prose gate on a value nothing downstream can see is not a gate, and a clean re-review would rewrite a ledger it was never meant to touch. Re-derived in block 2's own shell. 3. A numeric breakdown could still be internally false. `total: 0` beside `critical: 1` is four valid numbers rendering `0 findings — 1 critical, …`. Numeric was necessary and not sufficient; an inconsistent breakdown is now withheld for the same reason a partial one is. 4. The carried-marker strip ate hand-written prose. It removed a trailing `(not in the current review)` unboundedly and unconditionally, so a deferral reason that merely ENDED in that phrase lost it — the one field a human writes into this artifact. Now bounded to one occurrence, and only on rows the marker can legitimately be on. The no-growth property it exists for is re-pinned. 5. parseGateCounts diverged from the shipped pipeline in two ways no fixture reached. The shipped reads are `cut -d: -f2 | tr -d ' '`: `tr` removes INTERNAL spaces (`1 0` -> `10`) where `.trim()` keeps them, and `cut` takes only the second colon-field where a tail capture keeps the rest. The mirror models the pipeline now, and both counterexamples are fixtures — a parity assertion that agrees only on well-formed input asserts very little. 6. The prefix census guards were both partly vacuous. The drift guard read the two regex alternations and not the severity map, so a set could agree in both regexes while mis-tiering in the map. The domain guard scanned only `### XX-01:` headings — and BL appears in no heading at all, only in the Label-equivalence prose, so the guard passed purely because BL happened to be hard-coded and would have missed the next prose-defined prefix exactly as it missed BL. Both widened; the domain the guard now sees is BL, CR, IN, WR. Each fix fails a named test on reversion and none fires on the ordinary path. The property generator now deliberately produces the reserved suffix from (4), which a generator drawn only from innocuous characters could never reach. Also corrected: the previous commit's account of the empty-path failure. The script did not write a bare `-REVIEW-DISPOSITION.md`; it threw on reading the empty review path and the trailing `|| echo` swallowed it as a non-blocking skip. Same silent outcome, different mechanism, and the comment said the wrong one. * fix(#3829): the tests now run what bash runs — and six fixes to the fixes A second adversarial pass over the previous commit. It found a regression that commit introduced, and the reason it slipped through is the finding worth keeping. THE FIDELITY GAP. Every test here extracts the embedded script as TEXT and runs it. Bash does not: it expands the double-quoted `node -e "..."` argument first, so a backtick inside it is COMMAND SUBSTITUTION. The previous commit put one in a code comment. Bash duly ran it, failed with `+: command not found`, and handed Node a script two bytes shorter than the one 122 green tests were exercising. No behavioural test could see this, because none of them ever asked bash what it would actually pass. One now does, and it is the general guard: it catches an unescaped backtick, an unescaped $, and any other expansion the extractor cannot model. Then, in the shipped step: - A padded count silently disabled the sum check. `$((08 + …))` fails on base inference; it does not abort — the expansion sits in an `if` condition, where set -e does not fire — so the check simply never ran and an inconsistent breakdown passed with a stray diagnostic as its only trace. `10#` on every operand. - The status guard made the script's own reconciliation unreachable. A clean review with an EXISTING ledger must still be reconciled — decided rows carried, stale `open` rows dropped — or the ledger freezes showing findings as open that the review no longer reports. The guard now skips only when there is nothing to reconcile. - The carried marker is no longer stripped at parse time at all. Bounding the strip still ate a carried row's human-written reason. No-growth is a property of the RENDER, so it is enforced there: a marker already present is not appended again. Nothing is stripped, nothing doubles. - Fence openers are bounded to three leading spaces, per CommonMark. - parseGateCounts matched `[ \t]` where the shipped grep uses `[[:space:]]`, which covers form feed and vertical tab. Third counterexample of the same class, and a fixture. - The census drift guard checked only one direction, so a tier for a prefix the regexes never admit stayed green as dead code that reads as coverage. TWO OF MY OWN TESTS WERE VACUOUS, and the controls are what said so. The leading-zero test asserted exit 0 and a consistent verdict — both true before the fix. The clean-review test drove the node script directly, which never executes the shell guard at all: it passed unchanged with the guard made unconditional. Both are rewritten to test the layer the defect lives on, and both now fail when their fix is reverted. Every fix in this commit fails a named test on reversion, each mutation verified to have applied before its verdict was read. * fix(#3829): the carried marker can no longer outlive the carry A third adversarial pass. Its most important finding is a defect the SECOND pass talked me into, which is worth recording as plainly as the fix. THE MARKER BECAME A LIE. Pass 2 objected that bounding the carried-marker strip still altered a human-written reason, and proposed storing the cell verbatim instead. That objection was a preference, not a defect — its own driven output showed exactly one marker, which is correct — and adopting it created a real one: once the generated marker is stored it can never leave, so a carried finding that REAPPEARS in a later review still renders "not in the current review". The ledger then contradicts its own contents. Driven both runs. The strip is back, bounded to one occurrence and unconditional. The residual ambiguity is irreducible — a reason ending in exactly that phrase is indistinguishable from the marker — and it costs nothing real: on a carried row the render puts the phrase straight back, and on a current row the phrase was self-contradictory to begin with. The unbounded quantifier is what had to go, not the strip. The property now states that contract rather than asserting a verbatim survival the code deliberately does not provide. Also: - An ABSENT REVIEW.md abandoned the ledger it was meant to reconcile. The guard proceeds when a ledger exists, then the script read the review unconditionally, threw, and the trailing fallback swallowed it — the freeze the reconciliation path exists to prevent, reached through the door the guard opened. - Counts are length-bounded as well as digit-only. Bash integers wrap at 2^64, so a 20-digit count arrived at the sum as 0 and an inconsistent breakdown passed. - A closing fence must carry only whitespace after its marker; a line with an info string is an opener's shape and ended the fence early. - parseGateCounts matched [ \t\n\v\f\r] where the shipped grep uses [[:space:]], which under this UTF-8 locale matches EM SPACE. `\s` is the faithful model. Fourth counterexample of that class, and a fixture. - The agent-domain scan required [A-Z]{2,}, so a one-letter prefix like `C-01` — explicit and parseable, not prose — was invisible to it. AND THE FIDELITY GUARD PAID FOR ITSELF INSIDE ONE SESSION: writing this round's first draft I put backticks around a token in a code comment again, in the very commit whose subject is that mistake. The probe failed, named it, and no test of behaviour could have. Two of my own tests also had to be rewritten: one asserted things true before its fix, and one drove the node script directly where the defect lived in the shell. 383 pass across the touched files and the two size ceilings; ten lint gates green; every fix fails a named test on reversion, each mutation verified to have applied before its verdict was read. * docs(#3829): the Source reason is preserved, but not verbatim — say so Found by claim-auditing the response comment before posting it, which is the one place this would have been caught: the doc and the code were written in different commits and only a reader holding both notices they disagree. The feature reference said the hand-written reason is "preserved verbatim across re-runs". It is not, and deliberately so — a reason ending in the literal phrase "(not in the current review)" loses that trailing phrase, because it is indistinguishable from the carried marker the gate appends. The exception is stated rather than dropped, with the reason it is the better trade: storing the marker instead means it never leaves, and a carried finding that later reappears goes on claiming it is absent from the very review that reports it. A ledger wrong about its own contents beats losing a duplicated phrase, but only if the doc admits which one it chose. FEATURES.md regenerated; gen-features --check and lint:docs green. * fix(#3829): the gate now emits the counts it computes (B1a/B1b) Block 1 computed REVIEW_STATUS and the four counts and printed none of them, then the prose below asked the agent to display four of them. The shell exits at the closing fence and the agent sees only stdout, so those values were unobtainable: REQ-REVIEW-08 was unreachable in every shipped path and the fence was decorative. The rule was already stated one block down -- "a prose-only gate on a value no later block can see is not a gate" -- and applied only to block 2. It now governs the block that is this step's primary deliverable. Both arms emit, and the status gate is mechanical rather than prose: a clean/skipped/absent review prints nothing, an inconsistent or partial breakdown prints the countless form, and the full breakdown prints otherwise. Driven against the review's own case (critical: 1, warning: 9, info: 8, total: 18) with no appended emitter: Code review: 18 findings - 1 critical, 9 warning, 8 info. Consider running: /gsd:code-review 1 --fix * test(#3829): the counts harness stops manufacturing the output it asserts on (B2) runShippedGateCounts extracted the shipped fence and then APPENDED its own printf of the six internal variables before running it. Every counts assertion was green against a script that existed only inside the test process: the shipped fence emitted nothing, the tested fence emitted six lines because the test added them. That is why B1a shipped past a suite that looks like it covers exactly that surface -- the green was structurally incapable of turning red for it. The emitter now lives in the fence, so the harness reads the fence's own stdout and synthesizes nothing. Parity with the mirror moved up a level with it: renderGateMessage() renders both arms from the mirror's parsed counts and the assertion compares the WHOLE emitted message, so a drift in any parsed value changes the string or the arm it selects. Asserting on the observable is strictly stronger than asserting on five intermediates, and it can express what the old probe could not -- an absent review now reports NOTHING, which is a different fact from reporting a countless review. A fifth src.includes() assertion converted with it (round 1 retired four). It pinned the PROSE stating the countless condition, so it went red when the emitter moved into the fence while the behaviour it named was untouched -- the pin arguing for its own conversion. Negative control: reverting the shipped echo now turns 16 tests red. Before this commit the same reversion turned zero red, which is the finding. * fix(#3829): the disposition column is an enum, not any lowercase token (B3) ADR-227 requires a trust boundary to validate semantic SHAPE and to coerce a failure to the contract's safe default. The ledger is a trust boundary by construction -- the rendered instruction tells a human to hand-edit it -- and the prior-row parser captured column 3 as ([a-z]+), checked against nothing. One transposed character was enough. `| CR-01 | critical | opne | - |` is not the literal 'open', so it beat the default, was excluded from the `open:` headline count, and was carried forward forever. The ledger then reported the phase fully triaged off a typo. The asymmetry is what made this a correctness bug rather than a style point: a typo OUTSIDE [a-z] ('Deferred') already failed to match, lost the decision and reset the row to open -- safe. A typo INSIDE [a-z] was unsafe. The parser failed open in the one direction that matters. A row that fails the enum now yields no prior entry and the row falls back to 'open', by the same path the capital-D case already took. The property test could not have caught this: DECIDED is drawn from the vocabulary, so no property built on it can present an out-of-vocabulary token. Added JUNK, the arbitrary for the complement, deliberately lowercase so it stays inside the old capture's own character set -- the unsafe half is the token that LOOKS like a decision and is not. The new property also asserts the headline count agrees with the row it renders, which is the half the defect actually reported wrongly. Negative control: the new property fails against the ([a-z]+) capture and passes against the enum. * fix(#3829): a finding the heading parser cannot match is surfaced, not dropped (B4) Two independent parsers produce two numbers one paragraph apart -- the counts from REVIEW.md's frontmatter, the rows from `### <ID>:` heading matches against a closed CR|BL|WR|IN alternation -- and nothing reconciled them. A finding the alternation could not reach contributed no row, no note and no diagnostic, and the ledger then declared `open: 3 of 3` over a set strictly smaller than the console line had reported one paragraph earlier. Two findings recorded nowhere, and neither artifact said so. The PR's own argument for the closed alternation -- that an unlisted prefix produces no row rather than a MIS-CLASSIFIED one -- is the wrong trade under this repo's fail-safe rule. A dropped finding is demoted below every finding that parsed, and an unparseable finding is precisely the one a human most needs to see. Block 2 now derives the frontmatter total (anchored inside the findings: mapping, digit-and-length-bounded like block 1's) and hands it to the script, which reconciles it against the CURRENT review's matched findings -- order.length, never rows.length, which also counts carried rows and would either understate the shortfall or invent one. Surfaced exactly as the stale fix-report case already is: a non-blocking `unparsed: N` key plus the console line, both naming the two numbers so the claim is checkable. Code review disposition recorded: 3 of 3 finding(s) open (2 finding(s) recorded NOWHERE: the review reports 5, but only 3 matched the expected heading shape `### <CR|BL|WR|IN>-NN: <title>`) The key is emitted only when there IS a shortfall, so an ordinary ledger gains no noise key and the unchanged-run check is unaffected. Four tests, including three negative controls the round owed itself: a clean review gains no key, an absent/non-numeric total reconciles nothing rather than fabricating a shortfall, and a total SMALLER than the row count cannot render `unparsed: -1`. Reversion control: dropping the key turns the first red. * fix(#3829): pass --raw to the commit_docs config-get (#3763) Not from the review -- from a gate the base range added after it. #3763 lands `tests/config-get-raw-guard.test.cjs`, and this branch was its sole offender: a config-get command substitution without --raw feeds JSON.stringify output into a bash string comparison, where it silently never matches for string values. The consumer here is exactly that: if [ "$COMMIT_DOCS" = "true" ] Every other shipped call site in the tree already passes --raw (spike.md, fast.md, new-milestone.md, sketch-wrap-up.md, ...), so this is sibling convention, not a new posture. Worth recording because the two readings are both correct and they disagree: round 2's review cleared this exact line under ADR-3409 as "the safe member of that family", since `query config-get <key>` with no --pick exits 1 on absence and the fallback arm is reachable. That is still true -- --raw does not change it. The base then moved and added a gate that reads the same line for a different property. * fix(#3829): scope the count reads to the findings: mapping, not just the frontmatter (m1) `^[[:space:]]*total:` matches any indented key anywhere in the block, so a top-level key later named `total:`, `info:` or `critical:` was picked up ahead of the nested one. The block's own extensive comment is about scoping the FRONTMATTER, and the scoping stopped one level short of the mapping the values actually belong to. `status:` was never exposed -- it is anchored to column 0 because it IS top-level. The reads now run over the `findings:` block alone, selected by awk and cut at the next column-0 key. Block 2's REVIEW_TOTAL derivation (added with B4) already used that filter; this brings block 1 to it, so the two agree by construction rather than by coincidence. The mirror models the same scoping, and two fixtures drive it: a top-level `total: 999` ahead of a nested `total: 1`, and top-level `critical:`/`info:` ahead of theirs. Reversion control: unanchoring the shipped reads turns them red. * fix(#3829): severity comes from the section heading, not just the id prefix (M3) gsd-code-reviewer.md emits findings under '## Critical Issues' / '## Warnings' / '## Info', and that heading is the reviewer's own statement of a finding's severity. The walker already visits every line -- the fix-report path tracks '## ' sections -- so the signal was in hand and discarded in favour of the id prefix alone. A reviewer who mis-numbers a Critical as WR-04 while filing it under '## Critical Issues' produced a row reading 'warning'. The ledger's Severity column is the whole basis for triaging it, and it then disagreed both with the review it summarizes and with the frontmatter count line block 1 prints from findings.critical. Section first, prefix as fallback: a finding under no recognized section -- a review that does not use the documented headings, and every row carried from an earlier review -- keeps the prefix mapping, BL- included. Sections are matched WHOLE, exactly as the fix-report sections are, so '## Critical Issues Verification' does not re-tier what sits under it, and a heading inside a fenced example does not govern. Five tests: both mis-numbering directions, the prefix fallback across all four prefixes, the lookalike heading, and the fenced-example case. Reversion control: prefix-only turns the first two red. Sixth src.includes() assertion converted with it -- it pinned the exact source LINE of the enumeration loop, so it went red when that loop was reformatted while the property it names was strictly widened. It now asserts the property: every finding id, in order, once each. * fix(#3829): an untriaged row is carried too, not silently deleted (M1) The carry-forward kept a prior row only when its disposition was not 'open', so an untriaged row for a finding the current review no longer reports was dropped entirely. Combined with the reconciliation gap that left EVERY row open, a re-review deleted the whole ledger. The re-review loop rewrites REVIEW.md on every iteration, so REVIEW.md does not retain it either: run 1 records CR-01 open, the re-review renumbers it to CR-02, run 2's ledger contains neither. That is #3829's complaint verbatim -- "no trace of what happened to them" -- reproduced by the artifact built to prevent it. The old justification, "nothing was decided about it", is exactly the state #3829 says must leave a trace. Every prior row is now carried, and the carried marker is what keeps it honest: the row does not claim the finding is live, it records that it was seen and never triaged. Two costs, stated rather than discovered: a renumbered finding shows twice until the old row is triaged, and a carried untriaged row persists until decided. Both are bounded by the phase's own findings, both are legible from the marker, and both beat a silent delete. Five tests updated -- they encoded the dropped-untriaged behaviour as the contract -- plus one new test for the renumbering case M1 names. Reversion control: restoring the guard turns six red. Two self-inflicted defects caught while writing this, both by probes round 1 built: - Four unescaped backticks in a comment inside the double-quoted node -e argument, which bash ran as command substitution. The extractor-parity probe fired ("--auto: command not found"). Third time that trap has been sprung in this PR, third time the probe caught it. - The reworded ledger footer contained the literal carried-marker phrase, and the marker-accumulation assertion counts it across the whole file, so a doc line read as a second marker. The assertion was right. * test(#3829): cover the count-length threshold at limit-1, limit and limit+1 (M2) The guard is `?????????*` -- nine or more characters -- so the limit is 8 digits accepted, 9 rejected. The only cases were 'x', single digits and a 20-digit value, none of which pins the boundary. RULESET.TESTS boundary-coverage is a hard rule here and it was unmet. All three points asserted, with the sum kept consistent at each so the LENGTH rule is what decides the verdict rather than the sum check incidentally agreeing. Reversion control is the off-by-one M2 names: dropping one `?` moves the limit to 7 digits, which no test could previously notice, and now turns this one red. * feat(#3829): wire the disposition ledger into the fix path (B1c/B1d) REQ-REVIEW-09 was unreachable in every shipped path. execute-phase.md's code_review_gate invokes review with neither --fix nor --auto, so <NN>-REVIEW-FIX.md cannot exist when the gate runs and every row it writes is `open` by construction. The operator then runs /gsd:code-review N --fix by hand -- the very suggestion the step prints -- which writes REVIEW-FIX.md and never touched the ledger. A phase with 23 findings, all fixed, ended at `open: 23 / total: 23`: the artifact that exists to distinguish a triaged finding from a forgotten one asserted that 23 triaged findings were forgotten. Worse than recording nothing, because it looks authoritative and is inverted. Taking remedy (i), not (ii). Narrowing the docs to say the ledger reflects the previous phase execution is a legitimate choice, but it ships a feature whose central artifact is inert and then documents the inertness. ONE ADAPTATION, because the prescribed site does not exist. The review says to wire code-review.md's --fix/--auto path. code-review.md is not the writer (gsd-code-fixer writes the report, code-review-fix.md commits it), and more decisively it has no point that is AFTER the report exists: it delegates through code-review/steps/dispatch-fix.md, which calls Workflow(code-review-fix.md) and then exits the workflow. There is nothing downstream of that call to wire to. The site is code-review-fix.md, immediately after commit_fix_report. That is where the report is on disk and committed, it is the canonical implementation for all fix logic by dispatch-fix.md's own statement, and it additionally covers a direct invocation of that workflow -- which a wiring in code-review.md would have missed. The same step, not a second copy: it consumes PHASE_DIR and PHASE_NUMBER, both already parsed from the init JSON, and it is idempotent, so a phase that reaches the gate and then a fix run ends with one ledger reflecting both rather than two competing ones. Driven end to end: the gate writes `open: 2 of 2`, the fix path reconciles to fixed/skipped and `open: 0`. Two tests -- one pins the wiring and its ordering relative to commit_fix_report and present_results, one drives the two call sites in sequence. Reversion control: removing the step turns the first red; the second covers the reconciliation the wiring makes reachable rather than the wiring itself. * fix(#3829): a reflowed fix-report title is the same title (m2) The stale-fix-report guard compared titles with trim() equality. The strict instinct is right -- ids are reused across re-reviews, so a stale REVIEW-FIX.md must not mark a brand-new CR-01 as already fixed -- but gsd-code-fixer.md writes '### {finding_id}: {title}' under no contract that the title is copied byte-for-byte from REVIEW.md. A fixer that reflows a long title produced a spurious mismatch note, left a genuinely-fixed row 'open', and told the reader the report named a different finding. That false-positive mode was acknowledged nowhere. Whitespace is normalized, and only whitespace: a wrapped title is the same title, and it is the one divergence that carries no information. Case changes and truncation stay strict on purpose -- they are the shapes a genuinely DIFFERENT finding takes, and widening to them would trade a visible false positive for the silent false negative the strict match exists to prevent. The residual is now stated in the step rather than left to be rediscovered. The note's wording changed with it. It asserted the report "names a different finding"; both causes reach that branch and the step cannot tell them apart, so it now reports the observation -- "titles its finding differently from the review ... a stale report, or a re-titled one" -- rather than a conclusion it has not earned. Three tests: the reflow case reconciles cleanly, the re-cased case still reports, and the stale case still reports with the new wording. Reversion control: restoring the strict comparison turns the reflow test red. Seventh src.includes() converted -- it pinned the comparison EXPRESSION, so it went red when the comparison gained normalization while the property it names was unchanged. * docs(#3829): describe the flow that ships, not the one implied (m3) Both reference pages said "/gsd-code-review <N> --fix records fixed and skipped, which the gate reconciles from REVIEW-FIX.md" -- true in the abstract, materially misleading in practice, because no shipped path performed that reconciliation. With B1c/B1d wired it is now real, and the pages say WHERE it happens rather than leaving a reader to assume the in-phase gate does it: the gate runs before any fix report exists and writes all-open, and --fix is what records what happened. The round's other behaviour changes land here too, since a reference page that lags the artifact is worse than none: - the disposition column is a closed vocabulary, and a value outside it falls back to open rather than being treated as a decision - severity comes from the section heading when the review uses one, and from the ID prefix otherwise - an unparsed shortfall is stated rather than dropped - titles are compared ignoring whitespace, so a reflowed title still reconciles, and a mismatch is reported as an observation rather than as a claim that the report is stale - EVERY row is carried now, triaged or not, with the cost of the renumbered-finding double-entry stated rather than left to be found docs/FEATURES.md regenerated from the fragment; lint:generated-sync and lint:docs both exit 0. * chore(#3829): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 landed on next in #3954: the acknowledgment is a git commit trailer now, and tests/emitted-drift-acks/ no longer exists. Worth noting for anyone reading the rebase: this did NOT surface as the modify/delete conflict the migration guidance predicts. This branch ADDED its fragment rather than modifying an existing one, and the base deleted only the files that were already there, so the replay was clean and the fragment survived silently into a directory that no longer exists. Quieter than a conflict, and worse -- the gate is what catches it, not git. Two Growth keys rather than the fragment's one: round 2 wired the ledger into code-review-fix.md, so that file grew too. Both key on the bare filename, per the Growth namespace. Emitted-Drift-Ack-Growth: code-review-fix.md — #3829 review round 2, blocker 1c/1d: REQ-REVIEW-09 was unreachable in every shipped path because the in-phase gate runs before any REVIEW-FIX.md exists, so every ledger row it wrote was open and nothing ever reconciled them. This file gains one step, record_disposition, that reads and executes the same lazily-read step after commit_fix_report. It is the only point in the fix flow that is after the report is on disk: code-review.md delegates here through steps/dispatch-fix.md and exits, so it has no such point at all. Growth is one step of prose, no logic is duplicated, and the step is idempotent so the two call sites converge on one ledger. * chore(#3829): the changeset describes the round's behaviour, not round 1's It renders into CHANGELOG, so it carries the same misleading implication minor 3 was about: "the gate ... reconciling fixed/skipped from REVIEW-FIX.md" reads as though the in-phase gate does it, when the gate runs before any fix report exists. Says where it happens, and picks up the round's other user-visible changes -- carried untriaged rows, section-based severity, the disposition vocabulary, and the unparsed shortfall. * fix(#3829): a dotted phase number no longer aborts the step Found by this round's own adversarial review, in its MISSED section: no finding asked about it, and it is the most serious thing in the round after the two blockers. Both callers explicitly accept a dotted phase -- code-review.md:60 and code-review-fix.md:36 both validate ^[0-9]+(\.[0-9]+)?$ and name "03.1" in their own error text -- and both fences reconstructed the path with `printf "%02d" "${PHASE_NUMBER}"`, which cannot format one. Driven with PHASE_NUMBER=3.1: bash prints `invalid number` and exits 1, and under `set -euo pipefail` that aborts the step on its FIRST line. An advisory gate that promises never to block took the phase's entire review report down with it, and the newly wired fix-path call site inherited the same defect. Pad the integer part and carry the sub-number verbatim, so 3.1 -> 03.1 and 3 -> 03, with both arms falling back to the raw value rather than aborting. Driven: 3.1 now reads 03.1-REVIEW.md and writes 03.1-REVIEW-DISPOSITION.md; the integer path is unchanged. Two other findings from the same review, both about claims rather than code: MINOR 2's TEST WAS MIS-NAMED, and the reviewer was right to refute the claim. It called itself the "reflowed" case while substituting triple spaces, which is not a reflow. Driven: a genuinely WRAPPED heading is still not reconciled, because a `###` heading is one line by definition and the continuation is a separate paragraph. Not widened -- absorbing whatever follows a heading into the title would swallow arbitrary prose and make the stale-report check meaningless, and the kept failure mode is the safe one (a visible mismatch note, never a wrong "fixed"). The test is renamed to what it covers and the bound is now pinned by its own test. THE SHELL-SHARING GUARD DID NOT GUARD. Negative-controlling it -- rather than reading it -- showed that deleting block 2's real REVIEW_FILE derivation left it GREEN, on the exact defect it was written for. Block 2 prefixes its `node -e` with `REVIEW_FILE="${REVIEW_FILE}" ...` to put the values in the child's environment, and the detector counted that self-referential pass-through as a derivation. Pass-throughs are now excluded, and the control fires. Pre-existing, not introduced here: the original column-0 anchor matched that same line. Also worth recording: my first attempt at that control silently patched nothing and reported clean. Same lesson this PR already learned once. * fix(#3829): validate the phase number before formatting it, and make the shell guard executable Three findings from the round review's continuation pass, all confirmed by driving them. 1. MY OWN DOTTED-PHASE FIX WAS WRONG on the fallback path. `printf "%02d" abc` writes `00` to stdout BEFORE it fails, so `$(printf ... || printf %s ...)` CONCATENATES the two: `abc` became `00abc`, empty became `00`, and a legitimate `08.1` became `0008.1` because bash reads the leading zero as octal. An unset PHASE_NUMBER also aborted under `set -u` -- in the step that promises never to abort. Validate, then format: never format and fall back on failure. Driven across every edge the review named -- 3.1 -> 03.1, 3 -> 03, 08.1 -> 08.1, 09 -> 09, 1.2.3 -> 01.2.3, and abc / empty / -1 / unset carried verbatim with exit 0. 2. THE SHELL-SHARING GUARD STILL DID NOT GUARD. Excluding pass-throughs was not enough: a structural predicate recognises assignment TOKENS, never assignments that derive a usable value, so `REVIEW_FILE=`, `REVIEW_FILE=$REVIEW_FILE` and a commented-out assignment all evaded it. No regex closes that class. The authority moves to execution -- the third time this PR has learned that lesson. The real second fence now runs in a fresh shell with nothing but the step's two declared inputs and must write the ledger at the correct derived path. All four mutations are caught: empty assignment, self-reference, commented-out, and deletion. The textual check stays as a cheap fast-fail and is labelled as one. 3. THE TITLE-BOUND CORRECTION HAD NOT REACHED THE DOCS. The step comment and both docs pages still said a reflowed title reconciles, contradicting the bound pinned one commit earlier. Superseded prose left standing reads as current to anyone arriving cold, so all three surfaces are rewritten rather than annotated, and FEATURES.md regenerated. Also hoisted `HAS_BASH` to the file's other top-level constants. `const` is in the temporal dead zone until its declaration runs, and a `{ skip: !HAS_BASH }` option object is evaluated eagerly, so a bash-gated test added above the old mid-file declaration threw a ReferenceError that aborted its whole describe and CANCELLED its siblings -- while the summary line still read `fail 0`. It caught three separate additions in this round before I stopped moving tests and moved the constant. * fix(#3829): refuse an out-of-shape phase number instead of carrying it into a path Self-found while writing the prompt for the next review pass, which is the honest provenance: I asked the reviewer whether a path traversal was reachable through PHASE_NUMBER, then checked before dispatching. It was, and I had introduced it. The previous commit's fallback carried an unusable phase number VERBATIM, and PHASE_NUMBER is interpolated into a file path: PHASE_NUMBER='../../etc/passwd' -> REVIEW_FILE=/tmp/phase/../../etc/passwd-REVIEW.md The `printf "%02d"` it replaced had at least mangled that to `00`. A fix that makes a path more reachable than the bug it replaced is a regression, whatever it does for the case it was written for. Both callers already validate ^[0-9]+(\.[0-9]+)?$ (code-review.md:60, code-review-fix.md:36), so this is defense in depth rather than a live exploit -- but the step has two call sites now and should not take either caller's word for its own inputs. It validates the WHOLE value and, on failure, builds no path at all: PADDED is empty and each fence refuses by name rather than coercing. Block 1 declines to report counts read from a path made out of the bad value; block 2 declines to write, which also keeps it clear of the bare-name ledger defect round 1 closed. Driven across the shape boundary: 3.1 / 3 / 08.1 / 09 accepted; abc, empty, unset, 1.2.3, -1, 3., .1, +1, "3 1" and ../../etc/passwd all refused with exit 0 and a named diagnostic. Reversion control: restoring carry-verbatim turns the traversal test red. * fix(#3829): bound the phase number's length, and make the shell guard prove derivation Third adversarial pass. Two of its three refutations were already closed by the previous commit (the ../escape and 1/../../escape traversals, and the unset-input abort); these two were not. 1. A 54-DIGIT PHASE NUMBER WRAPPED SILENTLY. The validator accepted any all-digit value, so `$((10#$_int))` overflowed 64 bits and PADDED became `-7908320945662590977`. Length-bounded now at 8 digits, exactly as the counts already are and for the identical reason -- and the counts guard sitting twenty lines away is why this one is embarrassing rather than subtle. Driven at the boundary: 8 digits accepted, 9 rejected. The bare `${PHASE_NUMBER}` in the suggestion line is hardened to `${PHASE_NUMBER:-}` while here. The empty-PADDED guard makes it unreachable today, but it is one refactor away from an unbound-variable abort under `set -u`, in the step that promises not to abort. 2. THE EXECUTED SHELL GUARD PROVED THE FENCE WORKS, NOT THAT IT DERIVES. A single-phase probe is satisfied by a hardcode, and the review demonstrated exactly that: replacing the derivation with `case ... in 1) PADDED=01 ;; 7) PADDED=07 ;; *) PADDED=07 ;; esac` breaks every real phase and passed the entire suite. It now runs two distinct phases, 7 and 3.1 -- a hardcode cannot satisfy both, and the dotted one additionally pins the integer-part split. The claim "given only the declared inputs" was also overstated: the test spreads `...process.env` (it needs PATH and HOME). The DERIVED names are now explicitly deleted from that environment, so the claim is true rather than merely intended. Also rewrote a comment that had become false: it pinned a describe to the end of the file because of the HAS_BASH temporal-dead-zone constraint, which the hoist removed. Superseded prose left standing reads as current to anyone arriving cold. The changeset's "the gate stays advisory and never blocks" is now verified rather than asserted: both fences exit 0 under an unset PHASE_NUMBER and a traversal-shaped one. * fix(#3829): validate both inputs, refuse before building a path, and never write through a symlink Fourth adversarial pass. Four findings, all confirmed by driving them. 1. PHASE_DIR WAS NOT VALIDATED AT ALL. Unset, both fences died with `PHASE_DIR: unbound variable` under `set -u` -- the same class as PHASE_NUMBER, which I had just spent two commits fixing while its sibling input sat one line away. The step declares two inputs; it now validates two. 2. THE LENGTH BOUND WAS ON THE WRONG THING. The nine-character glob applied to the WHOLE value rather than the integer part, so it falsely rejected `12345678.1` (a legal 8-digit phase) while accepting `1.123456`. Each component is bounded on its own now; the sub-number is bounded too, since it is likewise interpolated into a filename. 3. REJECTED VALUES STILL HAD PATHS BUILT FROM THEM. The refusal guard sat AFTER the assignments, so an unusable input still assembled `${PHASE_DIR}/-REVIEW.md` and stat'ed it before refusing. The guard is now the first thing after validation, and both fences construct paths from validated locals rather than from the raw environment. 4. THE LEDGER WRITE FOLLOWED SYMLINKS. From the review's MISSED section, and the sharpest thing in it: `fs.writeFileSync` follows a symlink, so a pre-existing symlink at the ledger path replaced the contents of whatever it pointed at -- outside the phase directory, with the link left intact so nothing looked wrong. Driven, and the target's contents were gone. This PR introduces the artifact, so it owns the check: an existing ledger that is not a regular file is not a ledger, and the advisory gate says so and steps over. The executed shell guard now draws its phases AT RUN TIME. Fixed fixtures cannot establish derivation -- the review defeated the one-phase version with a hardcode, then defeated the two-phase version by adding one more arm to the same case. Any finite sample loses that race. A phase picked per run cannot be enumerated in advance; the drawn values print in every assertion message so a failure stays reproducible. Control: the review's three-value hardcode now fails on three consecutive runs. Eighth src.includes() converted -- it pinned the literal `${PHASE_DIR}` interpolation and went red when construction moved to a validated local, while "writes a REVIEW-DISPOSITION sibling" was untouched. It now asserts that property, and that REVIEW.md is not written. The changeset's "stays advisory and never blocks" is verified rather than asserted: 8 of 8 hostile-input cases across both fences exit 0 -- both inputs unset, PHASE_DIR unset, a traversal-shaped phase, and a missing phase directory. * fix(#3829): check the ledger path before reading it, and pin the write-safety behaviour Fifth adversarial pass, and the last one this round. Three fixes, three disclosed residuals. FIXED 1. A FIFO AT THE LEDGER PATH BLOCKED FOREVER. readFileSync on a FIFO never returns, so the step documented as "advisory, never blocks" blocked indefinitely -- the literal counterexample to its own headline claim. The non-regular-file check ran after that read. 2. THE UNCHANGED-RUN FAST PATH BYPASSED THE CHECK. A symlink whose target already matched the rendered ledger read through the link, reported `unchanged`, and never reached the refusal. Both fixed by the same move: the check is now the FIRST thing the script does, before any read or write of that path. Ordering was the defect, not the predicate. 3. THE COMMIT TEST FOLLOWED THE LINK the script had just refused. `[ -f ]` resolves symlinks, so the guard and its consumer disagreed about the same path and the helper could still be handed one. `[ ! -L ]` added. And the behaviour shipped with NO regression control -- I hand-drove it last commit and did not pin it, which the review caught by grepping for the words. Five tests now: symlink, symlink-with-matching-target, FIFO, directory, and an ordinary ledger as the negative control so the refusal is not a blanket one. mkfifo goes through the process seam like every other spawn here. DISCLOSED, NOT FIXED -- these are stated in the step rather than carried silently: - TOCTOU between the lstat and the write. Node exposes no portable O_NOFOLLOW write, and an attacker who can write into the phase directory mid-run already has what the check would protect. It narrows a real accident; it is not a security boundary and the docs claim none. - A hard link passes isFile() by construction. - The REVIEW.md and REVIEW-FIX.md reads still resolve symlinks. They are reads of files the operator owns, in their own phase directory. Also narrowed a comment that overclaimed. The randomized guard's domain is FINITE -- 88 integer and 792 dotted values -- so a mutation enumerating all 880 passes forever, and Math.random() is unseeded, so "reproducible" means only that the drawn values are printed on failure. Raising the bar is what it buys; proving derivation is not, and nothing short of reading the fence is. The previous comment claimed otherwise and was refuted. * test(#3829): make the write-safety controls portable to the Windows lane CI caught what neither the local suite nor five adversarial review passes could: every one of those ran on Linux. The FIFO test gated on `mkfifo`'s exit code. On the Windows lane mkfifo EXISTS and exits 0 while producing something that is not a FIFO, so the guard passed, the test ran against an ordinary path, the ledger wrote normally, and the assertion failed for a reason unrelated to the behaviour under test. It now gates on `lstatSync().isFIFO()` -- what was actually created, not what the command claimed. Control: with the shipped guard disabled the test still goes red on Linux, where the FIFO is real. The two symlink tests are skipped on win32, following this repo's existing convention for symlink-planting tests (tests/settings-jsonc.test.cjs:389 skips the same class; tests/unreachable-guard-drift.test.cjs:726 records the reason -- symlink creation requires elevated privileges on Windows CI). The privilege happened to be available on the lane this round, which is exactly why the convention is not "try it and see". * fix(#3829): a bare `|` in a deferral reason is prose, not a parse failure Review round 3, the one blocker. The Source cell is the one field this ledger asks a human to hand-edit, and "waiting on team A | team B to align" is an ordinary thing to type there. The prior-row capture admitted a pipe only when escaped, so a bare one failed the WHOLE line: prior.get() was undefined, the row fell through to `open` with an empty Source, and the console line read "1 of 1 finding(s) open" — a Critical a human explicitly deferred, with a documented reason, rendered indistinguishable from one never triaged, and the reason gone. The exact ambiguity #3829 exists to remove, reachable by one missing backslash. The Source cell is the LAST column, so it is now captured through to the end of the line, less an optional trailing pipe; a bare `|` inside it is prose. The render escapes a bare pipe on the next write so the table stays a table, and the escaped form re-parses to itself, so the second run reports `unchanged` — the fixed point holds. The ledger's own instruction line says so instead of asking the human to escape. Why the property never caught it: SOURCE_CELL only ever appended a PRE-ESCAPED pipe, so the arbitrary built to stress this cell could not reach the one input that broke it. It now also emits a bare pipe, and the round-trip expectation is the escaped form of what the human wrote. A fixed regression case drives the reviewer's exact input through two runs and asserts the decision, the reason, the headline count and convergence. Negative-controlled: both new tests fail against the previous capture. The src.includes() pin on the old capture text is retired for the behavioural case — it was pinning the defect. * fix(#3829): escape every bare pipe in one write, whatever precedes it Round 3, found by the adversarial pass over the round's own fix rather than by the review. The first escape used /(^|[^\\])\|/g, which CONSUMES the character before the pipe: adjacent bare pipes were escaped one per run (A||B -> A\||B -> A\|\|B, a third run to converge, breaking the advertised second-run fixed point), and an escaped backslash before a pipe (A\\|B) hid the pipe behind the wrong parity and left it bare in the rendered table. The property generator emits at most one bare pipe, which is the one case the old form got right, so no property reached either. Scan as pairs instead: an escaped pair (backslash + anything) is kept verbatim and only a pipe outside one is escaped. One write, then a fixed point. Regression case drives `A||B and C\\|D` through two runs; it fails against the previous escape. * fix(#3829): the script leaves by return, so an explicit exit cannot drop its verdict line Round 3, from the adversarial pass over the round's own fix. The embedded node script printed its verdict and then called process.exit(0) -- on the 'unchanged' branch only; the 'recorded' branch fell off the end. Node's "A note on process I/O" documents process.stdout writes to pipes and sockets as asynchronous on POSIX, and process.exit() as forcing exit before pending asynchronous stdout writes complete -- so on a POSIX lane the caller can see exit 0 with no verdict line. This is a hardening against that documented hazard, not a reproduced defect: the reviewer's empty-second-run stdout, which first pointed here, turned out to be its own sandbox -- a bare console.log child printed nothing there either -- and that attribution is withdrawn. The script now runs inside main() and leaves by return on all four early-exit paths, so the event loop drains stdout before the process ends. Same exit status either way, and the || echo fallback is unaffected. A structural test pins the absence of the call (comment-stripped; dotted, bracketed and whitespace-split spellings). The empty-review docs-parity pin that asserted the literal process.exit(0) line is retired -- the round-2 describe drives that property behaviourally. Also widens the property generator: SOURCE_CELL now reaches adjacent pipes and a backslash of either parity before a pipe, the two shapes the first render escape got wrong while passing every input the generator could then produce -- checked against an independent parity-walk oracle rather than a copy of the render's own scan. * fix(#3829): record what an --auto iteration fixed, instead of reporting it open Round 5's major. `record_disposition` runs once, after the whole capped-at-3 `--auto` loop converges — but this workflow keeps ONE final version of REVIEW.md and REVIEW-FIX.md rather than per-iteration copies, and deletes the .iterN.md backups on convergence. A finding fixed in iteration 1 was therefore absent from the final review (it was fixed, so the re-review stopped reporting it) AND from the final fix report (overwritten by the last iteration), so the row fell back to the gate's `open` and rendered `open ... (not in the current review)` — the same bytes a finding that vanished for an unrelated reason produces. That is the one distinction #3829 exists to make, undone by the artifact built to make it. The precise site was the two-arm `applied` construction: for an id the current review does not report, `sameTitle(undefined, h.title)` is false and `title.has(id)` is false too, so the entry entered NEITHER `applied` NOR `staleFix`. It was dropped in silence. Four changes, one defect: - A third arm. When the review does not report an id at all there is no title to disagree with, so this is not the stale-report case — it is what a finding looks like once it has been acted on. Record it. The id-reuse hazard stays closed by the arm below it: when the review DOES report the id, a title mismatch still goes to `staleFix` and is never applied, so a renumbered finding cannot inherit an earlier iteration's `fixed`. - Rows for decided ids the review no longer reports, carried and marked. A decision the ledger cannot render is a decision lost — the same silent drop the carry-forward loop already refuses for prior rows, one source over. - The .iterN.md fix-report backups are read alongside the final report, newest first, so the most recent statement about an id wins — the precedence a duplicate id already gets within one report. - The shell guard proceeds on a fix report, not only on an existing ledger. A direct `/gsd-code-review N --auto` writes no gate ledger, and a converged loop leaves `status: clean`, so a fully successful multi-iteration run recorded nothing at all. And the backups now go in `cleanup_iteration_backups`, after the ledger has read them. #3190's rule is untouched — spent scratch on convergence, retained on degradation — only the timing moved; deleting them inside the loop erased every early fix before anything read it. `CONVERGED` does not survive the loop's shell and is re-derived from the final review's status, which is exactly how the loop sets it; anything but a proven-clean review retains. Seven new regression tests plus an ordering test, all eight reversion-controlled against pre-fix code — every one fires. One is the negative control that matters: a reused id whose title differs must stay `open`, never inherit `fixed`. Residual, stated: an id appearing only in an iteration fix report takes its severity from the id prefix rather than a section heading, because `sectionSev` is built from the current review. That is the documented fallback for carried rows, not a new gap. * test(#3829): pin the two PADDED derivations against a silent desync Round 5's minor 1. Each fenced block runs in a fresh shell and must derive what it reads, so the PADDED derivation — the traversal fence between an attacker-influenceable phase number and a file path, plus the per-component length bound — is duplicated verbatim. Both copies were independently tested and nothing asserted they stay in step, which is the shared-parallel-surface shape CLAUDE.md requires a parity test for, on security-relevant validation logic rather than incidental repetition. Compared line by line rather than through a normalizing rewrite: a normalizer has to be told what may differ, and whatever it is told to tolerate stops being asserted. Exactly one line may differ — each block refuses by its own name — and the test names both forms. It also asserts the slice is substantial, since a parity test over an empty slice passes vacuously. Control: dropping one `?` from block 2's length bound, which moves that copy's limit to 7 digits while block 1 keeps 8, turns it red. That is the exact silent divergence the finding describes. One correction to the finding's own statement, since it is worth recording: the cited lines are :324 and ~:480, which are node-script lines; the derivations are at :48-83 and :211-246. And they are 35-of-36 identical rather than byte-identical — the refusal message differs, deliberately. * docs(#3829): state the PHASE_DIR trust boundary instead of carrying it Round 5's minor 2 asked that the assumption behind PHASE_DIR's validation be confirmed rather than silently carried forward at the two new call sites. It is confirmed, and the comment that stood here was wrong about it: "PHASE_DIR is the step's other declared input and gets the same treatment" describes something the code does not do. Both inputs have the SAME provenance — each caller binds them from `gsd_run query init.phase-op` (code-review-fix.md:7,17; execute-phase.md the same) — so neither is raw user input and neither is more trusted. The asymmetry is not about trust. It is that only one of them has a shape: PHASE_NUMBER carries a documented contract, `^[0-9]+(\.[0-9]+)?$`, asserted by both callers, so a value outside it is provably wrong and is refused. PHASE_DIR's contract is "a filesystem path", which admits `..`, absolute and relative forms and symlinked parents alike; no predicate separates a legitimate planning directory from an illegitimate one, so a shape check would reject working setups while proving nothing. So the emptiness check is adopted as what it actually is — the guard against `PHASE_DIR: unbound variable` aborting a step that promises never to block — and the shape check is declined, with the reason written where the next reader meets it rather than left to be re-derived. The residual is restated in place rather than left in a PR comment: PHASE_DIR may itself be a symlink and the ledger is then written through it, outside the phase directory, deterministically. Left alone deliberately — the write goes where the caller pointed. Not a security boundary, and nothing here claims one. * docs(#3829): record why HAS_BASH is a platform assumption, not a probe Round 5's minor 3 is DECLINED, and the reason is the repo's own contract rather than a judgement call — written at the constant so the next reader does not "fix" it and re-enable what the rule exists to prevent. The gap is real and confirmed: 22 tests carry `{ skip: !HAS_BASH }`, so block 1's bash severity-reporting path has no Windows-lane coverage. But `local/no-unguarded-nonportable-exec` (eslint-rules/no-unguarded-nonportable-exec.cjs, DEFECT.WINDOWS-TEST-PORTABILITY) REQUIRES this guard around `sh -c` / `bash -c` in tests, and its own remedy text names `if (process.platform !== 'win32')` as the sanctioned form, because these constructs fail under Windows Git Bash. So the constant is the repo's answer to this question, not an oversight in this PR. Swapping it for a runtime `bash` probe would light 22 tests up on a lane the rule has already determined they cannot pass — trading a legible, rule-encoded skip for a red matrix. Reversing that is the rule's decision; a change here belongs with a change there. * docs(#3829): describe how --auto's iterations reach the disposition ledger The reconciliation section described the `--fix` path accurately and said nothing about `--auto`, which is where round 5's major lived. It now states that the loop overwrites its fix report each pass, that the re-review drops a finding once it is fixed, that the gate therefore reads the per-iteration backups newest-first, and that the backups are removed after the ledger has read them rather than before. It also states the converged-with-no-ledger case: a fix report on disk is reason enough to record. FEATURES.md regenerated (176 features / 21 groups). Changeset extended to name the shipped behaviour rather than only the `--fix` half. * fix(#3829): clear lint-workflow-shellcheck, a gate the base range added Not from the review. The rebase onto `next` brought in `lint-workflow-shellcheck` (#4109), whose baseline was generated before this PR's new step file existed — so that file's findings are new by construction and `lint:ci` exited 1 on the rebased head before this round touched anything. The last green CI run predates the gate. Caught locally rather than by a red push. Three fixes and one baseline entry, split by whether the finding is real: - STRUCTURAL (not ShellCheck, not baselineable): the guard's `for _f in "…${PADDED}-REVIEW-FIX.iter"*.md` is the bare `for x in $VAR` shape that word-splits differently under bash and zsh. Wrapped in `$(printf '%s' "$PADDED")`, the linter's own prescribed remedy. - SC2097/SC2098, and this one was a genuine latent bug rather than a lint nit: `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"` sat in the same env-prefix list that sets `PADDED`, so its `${PADDED}` expanded the OUTER variable, not the one two entries earlier. Both happen to hold the same value here, which is exactly why it would have kept being wrong quietly. Built before the command now. - SC2317 ×3 is baselined, not fixed. It fires on `return 0 2>/dev/null || exit 0` — the deliberate idiom that lets a fence refuse whether it is sourced or executed — and the verdict is a false positive: the `exit 0` is reached precisely in the executed case. Rewriting a dual-mode refusal to satisfy a wrong unreachability claim trades a real behaviour for a clean report. Baseline 207 -> 210. `lint:ci` exits 0. 173 tests pass across the two touched files. * fix(#3829): a reused finding id no longer inherits the old finding's decision Found by this round's own adversarial review, which drove it rather than reasoned about it — and it refuted the arm I had named as my strongest suspicion, so it is recorded as a correction, not a discovery. Finding ids are reused across re-reviews: the --auto loop renumbers. `row()` inherited a prior decision on an id MATCH ALONE, with nothing checking it was the same finding. Driven: a prior `CR-01 fixed` row against a review reporting a brand-new CR-01 rendered the NEW finding `fixed`. A false decision in the artifact whose entire purpose is telling triaged from forgotten — the same failure mode round 4's blocker was, reached by the other door. I had argued this was closed by the stale-report arm. It is not: that arm guards the FIX-REPORT path only. The PRIOR-LEDGER path had no title check at all. - The ledger now records each finding's title, in the FRONTMATTER rather than a fifth table column: the Source cell is the field a human hand-edits and the one that must escape pipes, and a second free-text column doubles that surface for no reader benefit. - A prior decision is inherited only when the recorded title still matches. An ABSENT prior title inherits, deliberately — a ledger written before titles were recorded carries none, and refusing there would reset every decision in it, which is the loss this guard exists to prevent, caused by the guard. - A decision whose id has been reused is PRESERVED under a `superseded:` key rather than dropped. The review's driven refutation was precisely that the mismatch was surfaced while the decision was lost. It cannot keep a row — the id is taken, and two rows under one id is an ambiguity, not a record — so it is carried in the frontmatter, re-emitted every run, deduped by id+title, and named on the console. - And an iteration-derived decision now cites the report it actually came from. The Source cell hard-coded the unsuffixed `<NN>-REVIEW-FIX.md`, so a decision read out of an iteration backup cited a file that may not exist. A citation the reader cannot follow is worse than none. Also the review's finding. Five new tests. Four fail against the pre-fix step; the fifth — that a ledger with no recorded title still inherits — is a BACK-COMPAT guard and passes both ways by construction. It is not a reversion control and is not counted as one. * fix(#3829): follow the cleanup move through, and stop miscalling a converged run Three loose ends the earlier cleanup relocation left, two of them found by the round's own review and one by the suite. **#3190's own test still pinned the old placement.** T6 asserted the `.iterN.md` removal lives inside `auto_iteration_loop` — exactly what moving it broke. Its SEMANTICS are unchanged and still asserted: removed on convergence, retained on degradation, creation intact. What it now pins additionally is the ordering that forced the move — the ledger reads the backups BEFORE they are removed — and that the loop no longer removes what it just wrote. Rewritten rather than deleted: the assertion was superseded, the guarantee was not. **`CONVERGED` had become a decoy.** With the removal gone from the loop, the flag was set in two places and read in none. Deleted, and the prose that still said "the loop sets it" rewritten to what is true: the loop breaks on exactly one condition, a clean re-review, which leaves REVIEW.md at `status: clean` — and that is what `cleanup_iteration_backups` re-derives from. **A converged final iteration reported the opposite of what happened.** The post-loop message keyed on the iteration COUNTER alone, so a run that converged ON iteration 3 exited with `ITERATION == MAX_ITERATIONS` and printed "Reached maximum iterations. Remaining issues documented in REVIEW-FIX.md" over a run in which every finding was fixed. Convergence is re-derived from the review the loop left behind — the same signal the cleanup step reads, so the two cannot disagree. * docs(#3829): retract two claims this round made and could not support Both were caught by the round's own adversarial review, both were driven, and both would have reached the maintainer. Recording the retraction where the claim was made, rather than only in a PR comment. **The env-prefix "latent bug" does not exist.** An earlier commit in this round claimed that `FIX_REPORT_FILE="${_pd}/${PADDED}-REVIEW-FIX.md"`, sitting in the same `node -e` env-prefix list that sets `PADDED`, expanded the OUTER variable rather than the one two entries earlier — reading ShellCheck's SC2097/SC2098 as a defect report. Driven in bash and in dash: assignments in one prefix list take effect left to right, and the later entry DOES see the earlier one. The warning is a false positive here. The split is kept, but for readability only; the comment no longer describes it as a fix. **The HAS_BASH decline rested on a rule that does not govern these call sites.** It cited `local/no-unguarded-nonportable-exec` as REQUIRING the `process.platform !== 'win32'` guard. Checked, and wrong on both halves: the rule fires only on a file that also chmods an exec bit with an octal literal, and this file has none — so it never runs here — while `eslint-rules/lib/platform-guard.cjs` accepts four guard shapes plus `os.platform()`, not one. A constraint that exists is not a constraint that applies, and I did not check which. The decline stands on narrower and honest grounds: whether these fences PASS on the Windows lane is UNVERIFIED. What evidence there is points at divergence rather than absence — the rule's subject line is that `bash -c` constructs "fail on Windows Git Bash", and this PR already measured `mkfifo` existing on that runner, exiting 0, and creating no FIFO. So a probe would not be a clean win; it would light 22 tests on a lane whose shell semantics are known to differ and unknown in detail. That is a measurement to make deliberately, not a change to make in passing. The gap is real and is now stated as a gap. * fix(#3829): close four defects the review drove out of the first title fix The round's own adversarial review re-ran against the reworked tree and refuted two more claims. Every item below is its finding, verified before acting. **An iteration-only decision recorded no title, so the reuse guard leaked.** `applied` stored `{d, src}` and the row took its title from the current review — which does not report the finding at all. The row shipped with no title, and the next review reusing that id hit the title-ABSENT back-compat exception and inherited the old `fixed`. The exact defect the title machinery exists to close, surviving through the hole opened for legacy ledgers. `applied` now carries the title it was decided under. **A changed decision was dropped in favour of the obsolete one.** The dedupe was a has()-guard, so re-superseding a finding whose decision had since changed left the older record standing. It now replaces. **Re-spaced titles double-recorded.** The dedupe keyed on the raw title while `sameTitle()` collapses whitespace; the key now agrees with the comparison. **And the frontmatter was not valid YAML.** `title: Parser: loses data` is rejected outright by a real reader, and the `superseded:` line format was not YAML at all. Values are emitted as JSON scalars — YAML 1.2 is a JSON superset — and superseded records are properly nested. Round-tripped through js-yaml in the tests. One more, self-inflicted while fixing the above: the parse registered each carried superseded record TWICE, once at `- id:` under an empty-title key and again at `title:`. Records doubled on every run. They are collected during the walk and registered once, complete. **T6 was vacuous.** The review flipped `= "clean"` to `!=` in the cleanup and the rewritten T6 still passed — it greps for `FINAL_STATUS`, `rm` and "retained" occurring somewhere, never wiring them to a branch. T6b now EXECUTES the fence in both directions against real files. It fails on that exact mutation. **And a converged final iteration printed two success messages** — the loop's break already reported it. This branch now stays silent and exists only to withhold the degradation warning. Three CI gates the base range brought in, all tripped by this round's own text: - `/gsd-code-review` in a comment — runtime workflow artifacts take the colon form. Now `/gsd:code-review`. - The preamble-ordering parity test: my PHASE_DIR comment wrote the literal `gsd_run` before the shim preamble. Reworded. - Prompt-stuffing: the file passed 50K. I trimmed 5.8K of my own commentary first; even removing every added comment leaves the added CODE over the line, and the file entered this round at 44,523 — 89% of the budget. Added to SIZE_ONLY_WORKFLOWS with the same reasoning the two existing entries carry, and the same acknowledgement: splitting is the real fix. * test(#3829): extract the cleanup fence without an ad-hoc markdown regex T6b's helper used `/```bash\n([\s\S]*?)\n```/`, which trips two of the repo's own rules: `local/no-adhoc-markdown-parsing` (use the sectionizer, not a hand-rolled fence regex) and `local/no-crlf-fragile-split` (a bare `\n` against readFileSync content is wrong under Windows autocrlf). Line-scanned now, CRLF-normalized first — the same shape `bashFences()` in tests/code-review-pipeline-regression.test.cjs already uses, which solved this first. `npm run lint` is clean and T6b still fails on the inverted-branch mutation it exists to catch. * fix(#3829): withdraw the superseded-decision store; keep the identity guard Three adversarial passes over this round each found real defects, and passes 2 and 3 were entirely inside the `superseded:` block added in pass 1 — a second identity scheme, keyed on (id, title), living beside the row store keyed on id. Pass 3 refuted it on three separate counts: a legacy title that merely looked like JSON lost its quotes and fabricated a record; a finding that was deferred, superseded, then returned and fixed left an active row and an obsolete superseded record standing together, reporting `unchanged` forever; and my own test for the replacement path never passed the earlier ledger in, so it guarded nothing. The construct had no terminal state. It is withdrawn. **What survives is the safety property.** The ledger records each finding's title, and a recorded decision is carried forward only while the id still names the same finding. That is what stops a renumbered `CR-01` inheriting an earlier `CR-01`'s `fixed` — a false decision in the artifact whose purpose is telling triaged from forgotten, and the same class as round 4's blocker. **What is given up, and it is disclosed rather than hidden.** On a detected reuse the earlier decision loses its row. The drop is reported on the console naming the id and what had been decided, the previous ledger is committed so the row remains in git, and docs/features/code-review-pipeline.md states the limitation. Two defects from pass 3 are fixed rather than deleted, because they are in the guard and not the store: - **Known-empty and NOT-KNOWN were conflated.** `### CR-01:` yields an empty title; that is a title. While it emitted no `title:` key it read back as a pre-format ledger and inherited across a reused id — the same leak, three passes running. Emitted whenever the title is known, empty included; a carried row no source knows stays absent, which is the legacy-compatible read. Underneath it was a falsy fallback: `(act && act.t) || priorTitle.get(id)` discards `''`. Now a typeof check. - **JSON.parse ran on legacy values.** A pre-format ledger whose bare title was written `"quoted"` was parsed and lost its quotes, so the decision stopped matching. The frontmatter now declares `titles: json` and the parse is gated on it; a ledger without the marker keeps its scalars. One defect from pass 3 is NOT mine and is not fixed here: a converged run prints a success message from the loop break AND another from `present_results`. Both predate this round. My earlier claim that "the duplicate is gone" was true only of the pair I introduced; the pre-existing pair stands, and widening this round into `present_results` is not warranted. 188 tests pass. The three new tests fire against the pre-simplification step. `lint:ci` exits 0. The step file is 55,590 chars, down from a 62,220 peak. * docs(#3829): stop the ledger promising a preservation it no longer makes Fourth review pass. No machinery defects this time — both findings are claims in text this step SHIPS, which is the class this whole stack exists to prevent. **The rendered ledger still said "Re-running the gate preserves every row and every disposition."** That was true until the same round gave the step an intentional drop for a reused finding id, and then it was false in the artifact's own user-facing footer. It now states what the step does, including the one exception, where a reader actually meets it. **And the console asserted "the previous ledger is in git."** Committing the ledger is gated on `commit_docs`, and a failed commit is swallowed — so under `commit_docs=false` the overwritten decision may exist nowhere. The note reports the drop and stops there; asserting a recovery path that may not be there is the same overclaim in a smaller font. Two residuals from the same pass are DECLINED and documented rather than fixed, because both would need the second identity scheme just withdrawn: - A pre-titles ledger carries no titles, so its decisions inherit on the id alone. Refusing there resets every decision in every existing ledger, which is the loss the guard exists to prevent. - Two genuinely distinct findings sharing both an id and a title are indistinguishable to an (id, title) key. The pass also refuted the `titles: json` marker on a ledger written by `b86ea6065^`, which emitted JSON titles before the marker existed. Declined: that revision is an intermediate commit on this unpushed branch and has never been released. The PR's published head writes no titles at all, so a real ledger is either pre-titles (unmarked, bare — handled) or written by the shipped version (marked). The unmarked-JSON state cannot reach a user. Test pinned, and it fails against the pre-correction step. * docs(#3829): fix four wrong citations and one false size justification All four came out of a claim-audit of this round's own response comment — an audit of the text, not the code, which is where the remaining errors were. - **The caller citation was wrong.** The in-code note said both inputs bind from `gsd_run query init.phase-op`. `execute-phase.md:85` uses `init.execute-phase`; only `code-review-fix.md:22` uses `init.phase-op`. The substantive point is unchanged — both are orchestrator-derived, neither is raw user input — but the citation was not checked. - **A leftover "the prior row is in git."** Removed from the console note last commit, left standing in the comment two lines above it. - **The docs still carried the promise the ledger had just dropped.** The rendered footer was corrected; the same sentence in `docs/features/code-review-pipeline.md` was not. - **The SIZE_ONLY_WORKFLOWS justification was false.** It claimed the added CODE alone exceeded the threshold. Removing every round-added comment leaves 47,148 chars against a 50,000 limit, so the file CAN fit — the claim was wrong, and an exemption defended on a wrong premise is worse than no exemption. So the entry is re-justified on what is actually true, and earned first: another **10,188 chars** of this round's own commentary are cut (62,220 → 52,032, from a 44,523 baseline that was already 89% of the budget). Fitting under is possible only by stripping essentially all remaining explanation from logic three review passes found defects in. That is the wrong trade in a file whose house style is heavy in-fence documentation, and the entry says so rather than implying the file had no choice. One measurement corrected while checking: the Windows-lane skip count is **37**, not the 22 the review cited nor the 26 I first counted. Twenty-two and 26 count `{ skip: !HAS_BASH }` CALL SITES; a skip on a `describe` cancels its subtests. Forced the constant false and counted what actually skips. 236 tests pass. `lint:ci` exits 0. * fix(#3829): the drop report is conditional, and two published claims were not A fifth adversarial pass, run against the two commits that went out AFTER the fourth pass and were never reviewed, refuted three claims this round published. 1. The drop is NOT reported unconditionally. `row()` reports only a RECORDED decision (`was.d !== 'open'`); a prior row still at `open` is replaced in silence. The behaviour is right — `open` records no decision to lose — but the shipped ledger legend and BOTH feature docs asserted the report happens every time. Text corrected in all three places, which is the same defect class this round already corrected once for the preservation promise. 2. The test guarding that console wording was VACUOUS: it ran with no prior ledger, so no reuse occurred and its `is in git` assertion could not have failed however the console was worded. Driven through a real drop now, with the drop asserted as a precondition. A new test covers the `open` arm and fails on the pre-fix legend. 3. The HAS_BASH gap is now MEASURED rather than assumed, on native Windows with Git Bash 5.2.37 / MINGW64 first on PATH, node v25.2.1: HAS_BASH left alone: 179 tests, 127 pass, 0 fail, 52 skipped HAS_BASH forced true: 179 tests, 140 pass, 24 fail, 15 skipped So 37 of the skips are this guard's, confirming the count the round published — and unskipping is NOT a clean win: 24 fail, clustered on `bash -c` quoting and spawn failures, exactly the divergence the eslint rule's subject line names. The guard stays; it now documents a measured gap. The stale "the count is 22" comment is gone. 4. The size-exemption justification was wrong a second time. The overshoot is ~2.3K normalized chars, not "essentially all remaining explanation": the round's committed peak was 59,246 chars (not 62,220, which was never committed), and it entered at 44,466 chars, not 44,523 — both earlier figures mixed bytes into a character measurement. Rewritten to the numbers the scanner actually produces. Also: the shipped comment said both callers validate the phase shape without naming that they validate PADDED_PHASE, not the raw PHASE_NUMBER this step is handed. * fix(#3829): renumber this PR's two REQs, which #3661 took while the branch sat The rebase onto current `next` surfaced a REQ-number collision, not a text conflict. #3661 landed `REQ-REVIEW-08` (`workflow.code_review_point`) on `docs/features/code-review-pipeline.md` while this branch also claimed 08 and 09 for severity surfacing and the per-finding disposition. Two different requirements under one identifier is the kind of thing that reads as correct in both diffs and is wrong in the merged tree. Base numbering wins, because it shipped: `REQ-REVIEW-08` stays #3661's. This PR's two become **REQ-REVIEW-09** (severity surfacing) and **REQ-REVIEW-10** (per-finding disposition). Swept the whole tree rather than the conflict hunk — two references sat in files git merged cleanly and never flagged: - `gsd-core/workflows/code-review-fix.md:450`, the prose stating why `record_disposition` is the step's only reachable call site. - `tests/code-review-pipeline-regression.test.cjs:1782`, the comment on the test that pins that call site. `docs/FEATURES.md` is regenerated from the fragment rather than hand-edited; `node scripts/gen-features.cjs --check` is green (178 features, 21 groups) and `lint:generated-sync` exits 0. Two things stated rather than quietly carried. The `Emitted-Drift-Ack-Growth` trailer on the round-2 commit still reads `REQ-REVIEW-09` for what is now REQ-REVIEW-10 — it is a historical acknowledgment of that commit's growth, and its purpose is unaffected, so it is left rather than rewritten across 52 replayed commits. And `docs/INVENTORY-MANIFEST.json` appeared stale immediately after the replay, reporting two missing `cli_modules/` entries; that was the lane's pre-rebase build output, not manifest drift. Rebuilding in the replayed lane and re-checking shows it in sync and unmodified. Regenerating before the build would have committed the deletion of two base-added entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): reach the title round-trip with a generator that can break it Round 6's only finding. The round-5 title tracking introduced a fresh parser (the `titles: json` / ` - id:` / ` title:` frontmatter walk) and a fresh bijective contract (`JSON.stringify(oneLine(t))` out, `/^ title: (.*)$/` plus `JSON.parse` back in), and `tests/code-review-disposition.property.test.cjs` was untouched since round 4 with no reference to `title` at all. Every heading the generator built was `'### <id>: finding number <i>'` — never a colon, a quote, a backslash, or the empty string. You were right that this is the round-3 shape again, and I would rather demonstrate that than assert it. Two mutations to the shipped step, each a plausible edit rather than a contrived one: A. render `titles: raw` instead of `titles: json`, so the re-parser never JSON.parses and stores the quoted scalar as the title; B. `yv = (t) => oneLine(t)` — the bare scalar, no JSON at all. mutation A — new generator: FAIL old generator: pass (3/3) mutation B — new generator: FAIL old generator: pass (3/3) Both ship past the pre-round suite. The gap was reachable, not theoretical. What changed: - `TITLE`, a new arbitrary drawn from the class the render's own comments say the escaping is for — `:` (why `yv()` exists), `"` and `\` (what stringify/parse must round-trip), the empty string (the known-empty vs not-known distinction the render draws explicitly) — plus scalars that MIMIC the ledger's own frontmatter grammar (`findings:`, `titles: json`, a nested ` title: ` line, ` - id: CR-99`), unicode, surrounding whitespace, and one title long enough to outrun a scanner assuming short scalars. - `FINDINGS` now carries a title per id, so all four properties run the cycle over the title contract instead of over a constant. `IDS` keeps the old id-only shape it is built from. - A fourth property asserting the round trip in the two places it is observable: the stored scalar must `JSON.parse` back to the trimmed heading title, and a hand-recorded decision must survive the next run. The second half is the one that matters, and its construction is the point. The decision is made by EDITING THE RENDERED LEDGER IN PLACE, never by writing a bare row the way the existing properties do. A bare row carries no frontmatter, so `priorTitle` is empty, `sameFinding()` returns true through its `!priorTitle.has(id)` back-compat arm, and the title contract is never consulted — the property would pass over a completely broken round-trip. Both mutations above go green against the bare-row form. That collapse is why the property is written this way, and the comment says so in place. So the assertion is the consequence, not the JSON: a lossy round-trip does not corrupt a title, it makes `sameFinding()` false and resets a human's `deferred` to `open` with the reason gone — this PR's own founding failure mode, reached through the field the round-5 work added. BOUND, stated rather than quietly omitted: the generator emits no CR or LF. A `###` heading is one line by definition, so a newline is not an input the heading parser can be handed; `oneLine()` guards the value's other producers, not this one. Two things found while writing it, both corrected here rather than left: - `runOnce` now returns stdout. The reuse report is a CONSOLE note, not a ledger key, so my first draft's `assert.doesNotMatch(ledger, /^reused:/m)` was vacuously true forever — a test that cannot fail. - `expectedTitle` is a TRIM, not a `\s+` collapse. Collapsing is `sameTitle`'s COMPARISON rule; `oneLine()` is the STORAGE rule and preserves internal whitespace. The collapse form fails on an internal tab against entirely correct code, which is how a test gets weakened instead of believed the first time it goes red. The file header claimed "two properties" while three were running; it now states four, one line each. 239 tests pass across the four pipeline files, 0 skipped. `lint:ci` exits 0 (`lint-workflow-shellcheck`: 203 baseline findings, 0 new). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): the prefix census guard now says four sites, because round 5 added one Self-found, from re-deriving the round-1 finding-id census this round rather than carrying the round-1 verdict forward. The census guard's comment says the prefix set is "written out three times — the heading matcher, the ledger re-parser, and (by its keys) the severity map". That was true when it was written. Round 5's title tracking added a fourth copy: the frontmatter `- id: ((?:CR|BL|WR|IN)-\d+)` matcher that rebuilds `priorTitle`. The guard itself did not fall behind, and the reason is worth keeping visible: `idAlternations()` scans the extracted script by PATTERN rather than walking a fixed list of sites, so the new alternation was absorbed with no edit. Verified by running the extractor at this head — three alternations found, one distinct set, severity map keys `CR,BL,WR` with `IN` on the documented `info` default, 0 domain members not reached. Only the prose fell behind. Corrected, with the pattern-scan rationale stated in place so the next reader does not helpfully convert it into the hand-listed enumeration it deliberately is not — which would be exactly the defect this guard exists to catch, in the guard. Census discharge for this round: re-derived at the rebased head over the extracted shipped script, 3 enumeration sites reached, 0 not reached; the domain (the prefixes `gsd-code-reviewer.md` can emit, walked across both its heading template and its prose Label-equivalence paragraph) is unchanged since round 1 at 4 of 4. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): catch a duplicate REQ id in a fragment, since nothing did Not from your review — this is the test the round owed itself, and I would rather say why than let it look like scope creep. The renumber commit earlier in this round has no reversion control without it. I reverted that fix to check, and the first attempt LOOKED controlled: reverting only the fragment turned `gen-features --check` red. That is the generated-sync gate noticing the projection went stale, not anything noticing the collision. Reverting CONSISTENTLY — fragment plus a regenerated `docs/FEATURES.md` — is silent: gen-features --check rc=0 lint:ci rc=0 pipeline suite rc=0 with two `REQ-REVIEW-08` entries standing in one requirement list. Nothing in the repo reads REQ ids at all, so there was no second place for it to be caught. The failure this guards is a MERGE, not an edit, which is why review does not see it: two PRs open at once each append "the next" REQ number to the same list, and whichever lands second is rebased onto a list that already used it. git merges them as different lines of one file and reports nothing. Neither PR's diff shows a collision — each is correct against the tree it was written on. That is exactly how #3661 and this PR both ended up claiming REQ-REVIEW-08. Scope, stated because it is the part that could be wrong: the check is WITHIN a fragment, never across the corpus. Two different features legitimately both carry `REQ-REVIEW-01..07` — the cross-AI review feature and the code-review pipeline — so corpus-wide uniqueness would be false on the committed tree and would have to be weakened the day it first ran. A requirement list belongs to its feature; that is the scope of the identifier. It lives in `describe('the committed docs/features/ corpus')` because it is an invariant over the committed corpus, which is that block's stated job, and it pins no count — the file's own header rules out counts as shared mutable cells that every feature PR would have to edit. Control: green on the committed tree (no fragment carries a duplicate today); red on the restored collision, naming the file and the id. 85 tests pass in this file. Happy to drop this if you would rather the round stayed inside the review's four corners — but then the renumber ships uncontrolled, and I would rather put that choice in front of you than make it quietly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): finish the census comment correction, which stopped one line short Found by this round's own pre-push adversarial review, which refuted the claim the previous commit made about itself. `7f019d985` said the census comment correction was complete. It corrected one site and left two, both in the helper block twelve lines above the test it belongs to: - `severityMapKeys`' header still read "The THIRD copy: the severity map's keys". With three alternations the map is the FOURTH copy, and has been since round 5. - `idAlternations`' header said "adding a prefix to only two of them is silent", written when there were two alternations and never updated to three. This is the defect the original correction was ABOUT, committed inside the correction: a fragment of prose carries no supersession marker, so a reader landing on line 2810 gets the dead count stated as current fact, and the fixed comment eighty lines down does not reach them. Fixing one surface and leaving its neighbour is not a partial fix, it is the same fix not done. The region is now consistent end to end, and both headers say the thing that actually matters — the scan is by PATTERN, not a fixed list of sites, which is why round 5's new matcher needed no edit here and why converting it to an enumeration would reintroduce exactly the drift it guards. WHILE HERE, a disclosure that was narrower than the truth. `7a6680e8f` said the `Emitted-Drift-Ack-Growth` trailer still names REQ-REVIEW-09 for what is now REQ-REVIEW-10, and left it deliberately rather than rewrite 52 replayed commits. That is right, but it is not the whole set: the message BODIES of `c94106568` ("wire the disposition ledger into the fix path") and `06282f668` ("migrate the emitted-drift ack") both state "REQ-REVIEW-09 was unreachable in every shipped path", meaning the disposition requirement, which is now REQ-REVIEW-10. Same decision, stated at its real size: three historical references, not one. They are commit history rather than living documentation — git is the record of what was believed when — and rewriting the branch to correct a number in a message would cost every review round its correspondence to the commits it reviewed. The TREE carries no stale reference; `docs/`, the workflows and the tests all read REQ-REVIEW-09 for severity surfacing and REQ-REVIEW-10 for the disposition. Regression file: 181 tests pass, 0 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * test(#3829): the third stale count, and a disclosure that over-counted itself Both found by re-running this round's pre-push review after the last fix. It refuted the commit that claimed the region was consistent — for the second time in a row — and it was right again. **The third site.** `:2917` said "And the third copy, which is not an alternation" and `:2919` said "Without this, both regexes can gain a prefix". Written when there were two alternations; there are three, so the map is the fourth copy and it is three regexes that can drift. Worth saying how it survived two passes, because the mechanism is the point and it is the same one this PR keeps re-learning. Both earlier passes VERIFIED with a grep built from the strings I had just fixed — `THIRD copy`, case-sensitive, plus a handful of phrasings I expected. `the third copy` in lowercase matched none of them, and `both regexes` was not a phrasing I thought to look for. A grep returns what you already thought of; that is not a verification of prose, it is a re-statement of your own assumption. The region is now checked by reading it end to end, and all four count statements agree: three alternations (heading matcher, ledger row re-parser, frontmatter `- id:` matcher), with the severity map as the fourth copy. **And the disclosure over-counted.** The previous commit widened the historical REQ-REVIEW-09 references from one to three. Three is wrong. There are TWO underlying statements: - `c94106568`'s message body, and - the `Emitted-Drift-Ack-Growth` trailer on `06282f668`. I counted `06282f668` twice — once as "the trailer" and once as "a body" — when its only mention IS that trailer (`git show -s --format=%B 06282f668 | grep -c REQ-REVIEW-09` outside the trailer line: 0). Over-counting is the safe direction and it is still a wrong number in a message, which is the thing this round has been correcting all along. The decision is unchanged: both are commit history rather than living documentation, and rewriting the branch to fix a number in a message would cost every review round its correspondence to the commits it reviewed. The TREE carries no stale reference — 08 is #3661's `workflow.code_review_point`, 09 is severity surfacing, 10 is the per-finding disposition. Comment-only in one test file; no assertion, regex or extracted-script expectation moved. Regression file: 181 tests pass, 0 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mgqVNP3rqTVMgBNnBMnGH * fix(#3829): join the disposition-step dispatch so REQ-LANG-04 inheritance is provable `lint-response-language-coverage` (#2529, which landed on `next` after this PR was approved) reported `execute-phase/steps/code-review-disposition.md` as having no response-language coverage. The step does inherit it: `execute-phase.md` imports `references/execute-phase-response-language.md` and dispatches the step with `Read and execute`. The dispatch stub wrapped, leaving the verb at the end of one line and the path at the start of the next, and `namesFragmentAsEntryPoint` matches within a single line — so a genuine inheritance was unprovable to the linter. Rejoining the verb and the path restores it: `namesFragmentAsEntryPoint` goes false -> true and the lint reports `OK (165 workflows covered)`. Only line breaks move — the word stream is identical to the previous revision, and the file is unchanged at 93,390 bytes, so no growth acknowledgment is owed. This takes the third coverage form the lint documents — inheritance — rather than the inline directive the CI message names first. Where inheritance is provable the lint's own comments say a second copy "buys no coverage and adds a sentence that can drift", and the step file already sits over the prompt-stuffing threshold. Swept all 76 fragments in the catalog: this is the only one whose parent's previous line ends with a dispatch verb. The 17 others that are mentioned without a provable entry point are table-routed or bare prose references carrying no dispatch verb at all, and correctly hold the pinned inline directive instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JgX6QQmygeZnQqbc3o8RNC * chore(#3829): regenerate derived artifacts after rebase onto next The rebase onto current `next` conflicted on the 19 install-tree goldens and `docs/FEATURES.md`. Those are generated, so the conflicts were resolved arbitrarily and the generators re-run (`npm run regen:derived`) rather than hand-merged — a clean textual merge of a generated file attests the merge, never the content. Reconciled per artifact against the base's own committed copy rather than against the pre-regen tree, because the pre-regen tree is the arbitrary resolution: - all 19 `tests/fixtures/install-tree/*.json` now differ from `upstream/next` by exactly one key, `gsd-core/workflows/execute-phase/steps/code-review-disposition.md`; - `docs/FEATURES.md` differs by exactly REQ-REVIEW-09/10 and this PR's own reference section; - `docs/INVENTORY-MANIFEST.json` differs by exactly the same one step file, and needed no regeneration to get there. Nothing the base added was dropped by the arbitrary resolution: the restored entries (the `gsd-core/agents/` and `gsd-core/commands/gsd/` families, the compact templates, the `detail/elaboration.md` files, `gsd-secret-read-guard.js`) are all base-owned and came back through the generator, which is what the resolve-arbitrarily-then-regenerate discipline is for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EBjyHFtRTtHDD2tM6V6aUV * fix(#3829): repair the rebase's conflict resolution in the regression suite The rebase onto current `next` hit one add/add conflict in this file: #4209's external-reviewer-evidence describe and this PR's #3829 block were added at the same insertion point. Resolving it by keeping both sides was correct in substance and wrong in mechanics — the conflict boundary cuts through two open blocks that the SHARED trailing ` });\n});` closes, so each side carries +2 unbalanced braces on its own and concatenating them left the file with 683 `{` against 680 `}`. `node --check` fails outright, so the whole file deregistered rather than failing a test — 188 tests silently stopped existing. Rebuilt the region as a real three-way merge (ancestor |
||
|
|
76ef60ba25 |
enhance(#4836): prefer the graphify CLI for planner and researcher graph queries (#4874)
* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries The planner gets one knowledge-graph query per phase and the researcher two or three, and that single shot decides which modules the plan treats as related — and therefore how tasks are ordered into waves. It was spent on the built-in reader, which seeds by case-insensitive substring match over a node's label and description and then expands a hardcoded two hops. The phase "User Authentication" seeds on `author`, `authoring` and `unauthorized` with the same weight as `authenticate`, and when the inflated payload exceeds `--budget` the trimmer drops edges by confidence tier — so the highest-confidence tier can be discarded to fit a payload that bad seeding inflated in the first place. The graphify CLI is already a hard dependency of /gsd-graphify build, and it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context filters before traversal. Both prompts now prefer it and fall back to the built-in reader, branching on `command -v graphify` — the same degradation shape the repo already uses for Context7 to ctx7. Binary presence is a self-satisfying gate: a graph can only exist if the binary built it, so the fallback covers edge cases (a CI checkout with a committed graph, a binary since removed), not the common path. No new config key and no new tool grant — both agents already have Bash. The planner additionally runs `graphify affected`. The reference states its own goal as "which subsystems may be affected by changes in this phase", which is literally reverse traversal by relation; the built-in reader only approximates it with undirected two-hop expansion and has no equivalent verb, so `affected` is skipped on the fallback path. `graphify status` now reports `graph_path`, the resolved absolute graph location, on both the present and the missing branch. The CLI takes the graph location as `--graph`, and the prompts must not re-derive `.planning/graphs/graph.json` for it: that would point the CLI at a non-existent local mirror in exactly the umbrella multi-repo setup `graphify.graph_path` (#1825) exists to serve. For the same reason the presence gate in both prompts is now the `status` call itself rather than a bare `ls` of the default location, which was already blind to the override. Known limit, stated in both prompts rather than implied: the two paths return different shapes. `graphify query` emits prose and has no `--json` flag; the built-in emits JSON with per-edge confidence tiers and budget_met/budget_estimate. `--budget` also counts rendered output on one and estimated payload bytes on the other (#2738) — same flag name, different unit. Both are read by a model and nothing machine-parses the injected block. With graphify absent from PATH the injected context is byte-identical to before. Closes #4836 Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point. Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes. * docs(#4836): record the CLI-first graph query in the planner and researcher entries * chore(#4836): add changeset fragment * enhance(#4836): name the full domain word in the planner's query-term examples The reference's own example — phase "User Authentication" → term "auth" — is the exact collision the CLI-first path exists to avoid, and it stays a collision whenever the fallback path runs, since that path matches the term as a substring of label and description. * fix(#4836): surface graph_path on the unparseable-graph status branch graphifyStatus() returned graph_path on the exists:true and exists:false outcomes but not on the third, error, outcome (graph.json present but unparseable). The planner/researcher prompts gate CLI-first dispatch on exists, not on this outcome, so a corrupt graph file made them fall through to the CLI-first branch with the literal <graph> placeholder and no real path to substitute. * docs(#4836): note graph_path's trust boundary at the --graph interpolation graph_path is reflected verbatim into a double-quoted --graph argument the agent executes via Bash. It comes from graphify.graph_path, a config surface already trusted elsewhere, so this isn't a new trust boundary -- but it is a new injection site (no --graph flag existed on this call before). One-line caution for anyone hardening this later. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
585cab41cc |
fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices (#4920)
* fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices * docs(#4770): backfill the changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
6a4984cf69 |
fix(#4763): surface the displaced session record and pass --phase from the executor decision loop (#4919)
* test(#4763): failing-first — replaced-record payload and executor --phase pins * fix(#4763): surface the displaced session record and pass --phase from the executor decision loop state record-session keeps its last-writer-wins write (the recorded single-slot handoff design) but no longer displaces silently: when a non-empty Stopped At or authored Resume File record is replaced, the payload carries the full prior text under replacedRecord. Same-value rewrites, the insert path, and the #944 template-default DWIM are not displacements and report nothing. The executor decision loop now passes --phase "${PHASE}" to state.add-decision, matching execute-plan.md, so decisions stop inheriting whichever phase the global pointer names (#4763 case 2). advance-plan is unchanged (#3311 by-design). Emitted-Drift-Ack-Growth: gsd-executor.md — the decision loop gained its --phase guard and a comment naming why (#4763) * docs(#4763): add the changeset fragment * docs(#4763): backfill the changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
55fba5f7ce |
feat(#4917): add the PlanningDoc parse → mutate → serialize seam — Phase 1 of #4906 (#4918)
* feat(#4917): add the PlanningDoc parse -> mutate -> serialize seam Phase 1 of epic #4906, implementing ADR-4910 and its 2026-09-21 amendment. Net-new leaf module; NO call site is migrated, so nothing in the twelve absorbed issues changes behavior yet. src/planning-document.cts composes the seams that already exist rather than reimplementing them: markdown-sectionizer for structure (fences and code spans come from stripFencedCode / scanInlineCodeSpans, never a second scanner), markdown-table for tables, frontmatter for frontmatter, write-set for Result<T>. What is structural rather than conventional: - A field node carries labelSpan, valueSpan and trailingSpan separately, and the only write entry point takes a node id and writes into valueSpan. trailingSpan is readable and has no exported writer, so the #2853/#3584/#4852 rule ("the verb owns the count token ONLY") stops depending on an author remembering a third capture group. - Mutation is node-addressed. A handle is minted by the parser, so a caller cannot name a node the parser did not find. No path strings — a path is a grammar, and a grammar needs a parser. - serialize splices staged spans into the ORIGINAL buffer. A no-edit serialize is byte-identical, which is what eliminates the #4499 defect without a targeted fix, and which also makes an already-escaped table cell impossible to double-escape on a round trip. - A node that fails to parse carries its own error and span; siblings stay readable. - Per the ADR amendment, serialize REFUSES whenever any node carries a parse error, even with zero staged edits, naming the offender. One defect was found while building and fixed in place rather than deferred: PLANNING_ARTIFACTS first derived from isCanonicalPlanningFile unfiltered, so config.json, state.json, milestone.lock and skill-manifest.json were accepted and returned {ok:true, nodes:[]} — "this document records nothing" when the truth was "I have no grammar for this file". That is the exact empty-vs-error confusion this epic exists to remove (#4899, #4900), reproduced inside the seam built to prevent it. The registry now filters to .md while still DERIVING from isCanonicalPlanningFile, because hand-writing a second list is the divergence this epic is about, and parsePlanningDoc refuses a non-markdown kind at the document level per ADR-4910 section 5. New bin/lib module bookkeeping: .gitignore, eslint.config.mjs ignore (ADR-457 — lint the .cts, not the emitted .cjs), docs/INVENTORY.md row, regenerated docs/INVENTORY-MANIFEST.json, and the CONTEXT.md glossary entry. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4917): cover the PlanningDoc seam across 28 input classes 30 cases in 14 describe blocks, one per row of the phase test matrix. The two load-bearing tests are the fast-check properties (seed 20260921, numRuns 200): - a single node mutation leaves every byte outside that node's valueSpan identical to the source - serialize with zero staged edits is the identity function Both are DOCUMENT-SHAPED per CONTRIBUTING.md fixture provenance (#2371): the generator assembles arbitrary frontmatter, heading, label and value text with join(), and never calls serialize or any other function from the module under test to build a fixture. Seeding the generator from the module's own writer would make the document shape a constant, and the property could then never explore a document the writer would not itself emit. Verified the byte-range assertion is not vacuous with a control run: it passes against the real writer and FAILS against a simulated #4852 writer (one capture group, replace-to-end-of-line), which visibly drops the trailing annotation. Boundary coverage is zero / one / two staged edits. Negative space carries its own rows — bold emphasis in prose, a field-shaped line inside a fenced block, the same inside an inline code span, and a horizontal rule mid-body all correctly mint no node. Row 24 is the Generative-Fix-Divergence parity assertion: PLANNING_ARTIFACTS must not diverge from isCanonicalPlanningFile. Row 28 covers the non-markdown canonical file found during the build. Assertions are structural throughout — a typed Result / NodeRead / SerializeOutcome shape, or a byte range computed from the node's own Span. Full-string equality appears only where the contract IS byte equality. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4917): apply review findings — refuse unrepresentable values, compose the layers the seam claimed Three review passes ran against this branch: /security-review, an isolated adversarial pass, and a standards+spec pass. Five findings, all fixed here. Every one is the epic's own failure class reproduced inside the seam built to end it, which is the thing worth noticing. 1. setFieldValue accepted a value containing a newline. It survived serialization and reparsed as a REAL sibling field — one write to Plans forged a second Owner into a document that already had one. Content became structure, which defeats ADR-4910 Decision 2's "cannot reach past its own token by construction": the token boundary is a LINE boundary. 2. setFieldValue accepted a value containing the trailing separator and SILENTLY TRUNCATED it. Staged "sneaky - annotation", read back "sneaky", with the remainder reclassified as trailing prose. No error, nothing unreadable, both resulting nodes parsing perfectly. Worse than (1) because it loses the caller's own value rather than adding something visible. Both are fixed by ONE general check, deliberately not a blacklist: setFieldValue rebuilds the candidate line, re-parses it through the same field grammar, and refuses unless the value reads back identical. Blacklisting the separator would close this instance and leave the class open for whatever separator the grammar grows next. The round-trip check is ADR-4910 Decision 4 stated executably. 3. The module reimplemented two layers it claims to compose. Checklist detection hand-rolled a checkbox regex that markdown-sectionizer's iterateBullets already owns. Frontmatter span detection re-derived fence handling because frontmatter.cts's frontmatterRegion was module-private — so ADR-4910 section 1's stated layering was UNREACHABLE as written, and the first implementation routed around it silently instead of surfacing the gap. frontmatterRegion is now exported (additive only; ADR-2143 section 2's extend-never-mutate lock is inherited) and both layers are consumed. 4. The CONTEXT.md glossary entry asserted "the Frontmatter Module supplies frontmatter" while zero frontmatter.cts code was invoked. That was a false claim in the repo's vocabulary of record, written by me, and it is now true rather than edited away. 5. Adopting iterateBullets narrowed GFM coverage: it classifies only dash-prefixed task items as checkboxes, so "* [ ] x" and "+ [x] y" stopped becoming checklist nodes. Widening the sectionizer is forbidden by the inherited lock, so the task-list MARKER is interpreted in this module while bullet STRUCTURE still comes from the sectionizer. The sharpest finding was not a defect. The fast-check generator constrained values to [A-Za-z0-9 .,!?], so it could not emit an em-dash, newline, backtick, pipe or asterisk — precisely where (1) and (2) lived. The property was real, seeded and non-vacuous, and structurally blind to the module's actual bug class. The generator now spans the grammar's own metacharacters, and a new property asserts that every value setFieldValue ACCEPTS round-trips identically. Proven able to fail: against a scratch copy with the guard stripped it fails after 7 cases on newValue "\n". Recorded as a measured boundary, not fixed: a bare CR inside a field line leaves that field unrecognised. Measured — sibling fields still parse, no error node, and serialize stays byte-identical, so the worst case is an unreadable field and never a damaged document. Flagged for Phase 4's empty-vs-error census. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4917): backfill changeset pr number to 4918 Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b956bb7c67 |
fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs (#4902)
* test(#4834): failing-first launcher hijack regressions * fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs A gsd_run on PATH that cannot prove it is @opengsd/gsd-core (a foreign package, or a release older than the runtime-identity verb) is no longer accepted by the launcher snippet's PATH arm; resolution falls through to the hard error when no path-based candidate matches. The runtime-config-home arm now precedes the PATH arm, restoring the documented prefer-local-over-PATH order, so an installer-managed install wins even against a genuine global. The 16-home probe list is factored into _gsd_homes() and the identity gate into _gsd_id_ok(), keeping the per-copy delta at +141 bytes. The files whose frozen ceilings had no headroom (gsd-executor, gsd-plan-checker, gsd-verifier, gsd-planner, execute-phase, execute-plan) now load the resolver by @-include from gsd-core/references/gsd-run-resolver.md (the onboard.md pattern) instead of carrying an inline copy. Propagated to all other inlined workflow/agent copies via scripts/sync-runtime-launcher.cjs; the resolver reference re-copied byte-equal (parity B2); the hard-error text, docs/how-to/diagnose-a-foreign-gsd-tools.md, and the CONTEXT.md launcher predicate updated to match (#4834); the quick-batch row-48 guard gains the canonical-preamble sweep carve-out (#4834, per its own #3730/#2529 precedents); the compact-content benchmark baseline regenerated. Emitted-Drift-Ack-Growth: add-backlog.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-tests.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-todo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ai-integration-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-uat.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: autonomous.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: check-todos.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: cleanup.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: code-review-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: code-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: complete-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: debug.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: diagnose-issues.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: discuss-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: do.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: docs-update.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: edit-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: eval-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: explore.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: extract-learnings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: fast.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: forensics.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: graduation.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-code-fixer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debugger.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-eval-auditor.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-eval-auditor.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-intel-updater.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-intel-updater.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-project-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-project-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-research-synthesizer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-research-synthesizer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-ui-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: health.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: import.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: inbox.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ingest-docs.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: insert-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: list-seeds.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: list-workspaces.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: map-codebase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: milestone-summary.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: mvp-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: next.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: pause-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plan-review-convergence.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plant-seed.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: pr-branch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: profile-user.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: progress.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: quick-batch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: quick.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: remove-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: remove-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: resume-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: scan.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: secure-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings-advanced.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings-integrations.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ship.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sketch-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sketch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: smart-entry.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spec-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spike-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spike.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: stats.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sync-skills.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: thread.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: transition.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ui-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ui-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ultraplan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: undo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: validate-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: verify-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) * docs(#4834): backfill the changeset PR number * test(#4834): regenerate the compact-content benchmark baseline after the rebase --------- Co-authored-by: sim <sim@local> |
||
|
|
eea9247c93 |
enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select Auto-mode used to auto-select a checkpoint:decision's first <option> unconditionally, making a decision checkpoint's safety depend on option presentation order rather than an authored choice. Add an optional auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode now escalates to a human exactly like gate="blocking-human" does; present, it names the option auto-mode selects; naming an id with no matching <option id> is a hard structural-validation error at plan-parse time rather than a silent fallback to the first option. gate="blocking-human" continues to win over everything, unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys An isolated adversarial review of the auto_select work found that both new attribute regexes used \b as their left anchor, which is a word boundary, not a "start of attribute name" boundary. A decoy attribute ending in the same word (e.g. data-id="...") sitting before the real id="..." on the same <option> tag matched first, silently corrupting the extracted option id. Anchor on (?:^|\s) instead so only the real attribute name can match. Adds a regression test reproducing the exact decoy-attribute shape, plus a Unicode option-id test from the same review pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane lint-docs-guard-registration failed: the new test reads docs/reference/ plan-md.md but was not registered, so a future edit to that doc could silently desync from the test without the guard catching it on the PR that changed the doc. Registered alongside its direct precedents (precondition-element.test.cjs, reversibility-tagging.test.cjs), which read the same file for the same reason. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling The remote gsd-test run caught what local checks missed: execute-phase.md carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes of headroom before this change, and the original checkpoint:decision wording pushed it to 93772 (over the ceiling). Cascaded into failures in phase6-capstone-conformance, execute-phase-completion-reconciliation, claude-orchestration, and the compact-content drift-report test. Also caught: tests/package-legitimacy-gate.test.cjs anchors a "decision is conditional, not unconditional" safety check on the literal phrase "first option" in the decision bullet — which #4095 deliberately removes, since there is no more unconditional first-option pick. The test was asserting an assumption this change intentionally makes obsolete; re-anchored on tokens that still identify the bullet ('decision', 'auto-spawn') without weakening what the test actually verifies (the bullet must still carry a blocking-human carve-out). Also fixed a word-order mismatch between my own new test's regex and the actual doc text it was asserting against (tests/auto-select-attribute.test.cjs). Regenerated the compact-content benchmark baseline (tests/fixtures/compact-content-benchmark-baseline.json) to match the new byte counts. Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap. Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4095): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
029acd9158 |
fix(#4434): win32 unmeasured test files weigh the documented ~2.2x Windows-cost floor, not the Linux-measured mean (#4903)
* fix(#4434): win32 unmeasured test files weigh the documented ~2.2x Windows-cost floor, not the Linux-measured mean test(#4434): failing-first — win32 unmeasured-file weight must use the documented Windows-cost multiplier, not the plain table mean fix(#4434): makeFileWeigher takes a platform and applies WINDOWS_UNMEASURED_COST_MULTIPLIER (2.2) to the unmeasured-file fallback on win32 only; measured files and other platforms are unaffected chore(#4434): changeset fragment Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4434): add fast-check property coverage for the win32 unmeasured-file weight multiplier CLAUDE.md requires a fast-check property test for budget-limit logic; the prior example-based #4434 tests didn't satisfy that. Adds two seeded, bounded property tests: the unmeasured-file fallback matches the platform rule for any measured table, and a measured file's weight is platform-invariant for any measured table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4434): backfill changeset PR number (4903) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4434): the Windows unmeasured-file multiplier must not apply when there is no timings table at all Real Windows CI on this PR caught a regression my own linux-only gsd-test verification couldn't see: makeFileWeigher applied WINDOWS_UNMEASURED_COST_MULTIPLIER even when `timings` is null (missing/ corrupt/empty table), breaking the pre-#2456 "no table degrades to uniform weight 1" invariant several existing tests depend on. The multiplier now only applies to a file absent from an otherwise-loaded table — the actual #4434 mechanism (a real, loaded, Linux-measured table with unmeasured entries) — never to the no-table-at-all path. Six pre-existing tests that called makeFileWeigher/pack helpers with no explicit platform (silently inheriting whatever OS runs them) now pin an explicit 'linux' platform, since they test the platform-agnostic mean-vs- median and Object.prototype-safety invariants, not #4434's Windows behavior. One subprocess-level test is isolated from the real committed tests/test-timings.json via RUN_TESTS_TIMINGS_FILE, matching this file's existing isolation convention for cost-profile-sensitive assertions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ccfed63355 |
fix(#4823): the Current Plan reset is scoped to the Current Position section (#4898)
* test(#4823): failing-first — the Current Plan reset must not rewrite prose outside Current Position * fix(#4823): the Current Plan reset is scoped to the Current Position section — the whole-body 'Plan' fallback matched hard-wrapped prose lines starting with plan: * chore(#4823): changeset fragment * chore(#4823): backfill changeset PR number (4898) --------- Co-authored-by: sim <sim@local> |
||
|
|
bd79a97df0 |
fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error, never status missing (#4896)
* test(#4806): failing-first — unparseable VERIFICATION.md frontmatter is a parse error, not status missing / Field not found * fix(#4806): unparseable VERIFICATION.md frontmatter reports a distinct parse error — never 'missing' or 'Field not found' * test(#4806): census pin 66→67 — cmdFrontmatterGet's unparseable-frontmatter error is a new output({error}) call site * chore(#4806): backfill changeset PR number (4896) --------- Co-authored-by: sim <sim@local> |
||
|
|
e2d879f681 |
fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse (#4895)
* test(#4802): failing-first — acknowledge must refuse an unparseable-frontmatter target instead of splicing over it * fix(#4802): audit acknowledge refuses targets whose frontmatter fails to parse — never splices the marker-marked object over the file * chore(#4802): backfill changeset PR number (4895) --------- Co-authored-by: sim <sim@local> |
||
|
|
a8394713d1 |
fix(#4801): init.manager resolves archived phase directories through findPhaseInternal (#4893)
* test(#4801): failing-first — an archived phase directory must resolve and report complete in init.manager * fix(#4801): init.manager resolves archived phase directories through findPhaseInternal The private current-milestone-only scan (matchPhaseDirs over listMilestonePhaseDirs) counted archived phase directories as missing, so an archived phase with a passed verification reported no_directory/phase_complete:false. findPhaseInternal — already imported, already the shared primitive for five other init commands — searches the live directory first and falls back through the workstream-scoped archive (#2855); the private copy and its single-consumer entries list are retired. * chore(#4801): backfill changeset PR number (4893) --------- Co-authored-by: sim <sim@local> |
||
|
|
88b5775dc8 |
enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI gsd-ui-auditor is chartered to audit interaction and handed a capture driver with no interaction verb: `npx playwright screenshot` cannot click, fill, hover, press or snapshot, so a hover state, an open menu, a focus ring or a form's validation state never appears in its evidence and every Experience Design finding degrades to code reading. Implements the shape approved at triage, not a new capability: - capabilities/ui/capability.json declares `workflow.ui_interaction_capture` (boolean, default false) on the capability that already owns the auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated. - gsd-core/workflows/ui-review.md reads the key through gsd_run and hands it to the auditor as `interaction_capture:` in the spawn <config> block — the auditor carries no gsd_run resolver, so the key travels by value. - agents/gsd-ui-auditor.md gains an anchored interaction-capture section AFTER the static block. With the key on and a Chrome binary resolved it starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on an --isolated profile, opens the dev URL the static block reached, takes the a11y snapshot for element uids, captures the baseline and a Tab focus-ring state, drives the UI-SPEC's interactive components, saves console output, and stops the daemon unconditionally. Key off, no dev server, or no Chrome: one status line, and the Playwright-only static path runs exactly as before — the static fence is untouched. Needs only Bash: no MCP server, no tools: change. Chromium-only by nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so readiness is polled through evaluate_script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): bind the interaction-capture shape and containment - manifest, generated registry, config schema and config-set/loadConfig all know workflow.ui_interaction_capture as a default-off boolean, and hand-written non-booleans fall to the slice default - the orchestrator reads the key and hands it down; the auditor never grows a gsd_run dependency - the static fence stays Playwright-only and the interaction fence chrome-devtools-only, so key-off is today's path - the interaction fence runs under bash with a stub driver on PATH: key off / absent / no dev server / no Chrome invoke nothing; the happy path starts first and stops last on the [selected] pageId with the documented flags; a failed capture is removed and not counted; new_page and start failures still honour the stop-only-if-started rule; CHROME_BIN and CHROME_DEVTOOLS_MCP_VERSION overrides flow through - docs/CONFIGURATION.md row shape; registered in the docs-guard lane Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * docs(#4223): document workflow.ui_interaction_capture and its how-to - docs/CONFIGURATION.md: one row in the workflow.* table, default-off - docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the interaction-capture section adds, skips and never claims - docs/how-to/enable-ui-interaction-capture.md: turn it on, read the `**Interaction captures:**` outcomes, what it does not do, turn it off - docs/README.md: index the how-to beside live-DOM verification Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): add changeset Added-type fragment; pr: carries the issue number until the PR exists. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace; the hyphen form is retired there and the slash-command-namespace guard rejects it. docs/ keep the hyphen form by convention. Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered. Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): per-run daemon session, bounded navigation, and step failures that count Three findings from the pre-file adversarial review of the interaction fence, folded in: - `--sessionId <epoch>-<pid>` on every driver call. `start` restarts whatever daemon shares its session and --isolated isolates only the browser profile, so two concurrent audits — or an audit beside the operator's own CLI daemon — would otherwise stop each other. The CLI accepts hex and dashes only; the id is validated by the test stub. - `new_page --timeout 30000`: the one verb that takes a bound, placed before every verb that does not, so a hung page is caught first. - a failed take_snapshot or press_key now increments the failure count and is named on stdout; two clean screenshots can no longer read as `0 failed` after the step that gives the interactions their uids failed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal Second review round, both reviewers: - session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a subshell, so two audits forked from one parent in the same second shared an id and could stop each other's daemon (driven by the reviewer) - `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver under Git Bash still matches the `$` anchor, and `|| true` on the assignment so a failed new_page cannot abort the block under `set -e -o pipefail` before the unconditional stop - a failed take_snapshot removes any snapshot.txt it left or inherited from a reused directory, so stale uids never drive the interactions - `<config>` placeholder is `{interaction_capture}`, lowercase like its `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt template, not a bash heredoc - how-to: the `not captured` row no longer claims the daemon started Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges Third review round: - a new_page that prints a page line and then exits non-zero is a failed navigation, not a page id: the exit status is checked in an `if` before the output is parsed (driven by the reviewer against the previous `|| true`, which masked exactly that) - regression cases for what the last two rounds added: CRLF driver output, a stale snapshot removed on failure, partial-output new_page failure, and the whole fence under `set -e -o pipefail` (both the failed-navigation path and the happy path) - the harness whitelist gains `date`; the session-id assertion now requires all three parts, so a silently empty epoch cannot hide again Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder - the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes, over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces; the interaction section's comments are tightened to the same content in fewer bytes (23559 now). No bash changed — the fence's own tests and the real-browser run are unchanged. - .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which the post-create backfill rewrites to the PR's own number. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): set changeset fragment pr to 4477 * test(#4223): compare the fence's status path with the separator the fence uses On the windows-latest lane the happy-path case failed on `\interaction` vs `/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a literal slash, and the assertion built its expectation with path.join. Every other case in the file passed on that lane, including the CRLF and errexit/pipefail ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): drop the inert file-header allow-test-rule marker Review round 1 on #4477: the `source-text-is-the-product` marker sat at line 2, outside no-source-grep's 8-line lookahead of every readFileSync site (the first is ~60 lines down), so it suppressed nothing. It was also unnecessary: every read in this file is a .md/.json path, which the rule does not trigger on. Deleted rather than relocated — there is no site to relocate it to. Negative control: `eslint` on the file is clean without it. * chore(#4223): regenerate the platform-conformance-tier lists for the new test Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier gate after this branch opened; its two committed lists must name every file under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never been in them. Once the branch was updated against next the lists were stale and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier --check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's fragment-single-edit-propagation, which sees the same staleness as regen:derived touching files beyond the fragment edit under test. Regenerated with the repo's own generators, no hand-editing. The general tier goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry; both --check arms are clean. Verified the red is this PR's own file and not base drift: at upstream/next both generators report "list matches" (546 / 196), and our committed copies were byte-identical to next's before this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ * fix(#4223): bound, confine and trap the chrome-devtools driver fence Round 4 — three findings in one fence, interleaved on the same lines, so one commit: - Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a background job in its own process group (`set -m`) under a watchdog that kills the whole group at the ceiling — TERM, then KILL two seconds later. One pid is not enough: npm forwards SIGTERM only to its direct child, so killing `npx` alone leaves the client holding the fence's stdout and a `$(cdt … new_page)` capture blocked past the ceiling (driven against a real npx tree by the round's adversarial review; the pid-only first cut of this commit had exactly that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )` subshell: a subshell inherits bash's saved copies of the caller's stdio (the fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a watchdog outlived its kill — measured as the intermittent 30 s test run the review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 -- -pgid`, every 0.1 s) and stands down by itself once the group is empty; nothing ever signals it. The group, not the leader pid: a child can outlive the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll stood down at once and left the substitution open-ended (driven by the round's adversarial review at 6× the ceiling; a pgid cannot be reused while any member lives, which a bare pid can). The daemon `start` launches is spawned detached (its own session) and never in that group. Two platforms forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM; sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git Bash a signal to a watchdog still starting up hung the fence's `wait` for it: 18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs, while a fence slowed by xtrace, or three tests run alone, never hit it (a startup race; the mechanism is not pinned further). Polling: 0/300 slow calls and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3 full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU, msys and busybox sleep; BSD sleep documents it). A clock that cannot launch (`sleep … || exit 0`) stands the watchdog down rather than firing at once and killing a healthy call — by design that leaves a hung call unbounded, the pre-round-4 behaviour, instead of failing a healthy one. A hung call returns once its group is gone: at the ceiling, plus up to the 2 s TERM-to-KILL grace. The KILL after the grace is sent only to a group that is still alive: a pgid freed during the grace can be reused, and an unconditional KILL could hit an unrelated group (the round's review). `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog rather than either. - --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may write under the run's interaction/ directory and nowhere else. Relative, like every --filePath (unchanged from rounds 1-3): the daemon resolves both against one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and path.resolve()s both), and a relative path needs no dialect translation — an absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native daemon cannot resolve (CI's windows conformance shard caught the first cut). --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified), so the documented floor moves from ^1.8.0 to ^1.9.0, where --allowUnrestrictedPaths is deprecated. - `stop` is owed by an EXIT trap after a successful `start`, not by position (it replaces any earlier EXIT trap — none exists in this file); the explicit call keeps it in order, a flag makes the trap a no-op afterwards, and only the shell that installed the trap may act: a subshell copy of the fence state carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced locally at 3/40 under load: the second `stop` came from a subshell pid, never main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist and CI's macos conformance job showed the guard comparing empty to empty. The fence was driven under bash 3.2.57 for the injected-subshell, errexit failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy paths. A failed resize_page is a counted failed step now, not the one bare command an errexit runner could abort on. Prose in the section is tightened to pay for the mechanism: 23559 -> 24517 bytes against the 24576 DEFAULT-tier cap. Tests: the stub driver hangs as a real child tree (sh waiting on a child that holds stdout — never an exec), so a pid-only kill fails the new aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative- controlled: it blocks for the harness's whole cap on the old wrapper). A hung start and a hung capture are cut off within ceiling + grace + slack and still reach stop; an injected bare failure under errexit reaches stop through the trap, exactly once; an injected subshell call of cdt_stop issues nothing; the happy path issues exactly one stop; every driver call site names a ceiling and the only bare $CDT is the wrapper's own spawn; the start line carries --workspace with the capture directory, every --filePath lies under it, and no code line carries --allowUnrestrictedPaths. A driver whose leader exits at once while a child keeps holding the capture pipe is still cut off at the ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone (negative-controlled: the trap form kills `start` in under 20 ms). The harness EXPORTS its stub-only PATH — unexported, the exec'd watchdog fell through to bash's compiled-in default PATH and never saw the stub dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash, pid-preserving). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the accessibility tree, with entered form values) and console.txt (which can carry tokens) were committable by `git add .`. The gate now ignores `interaction/` as a directory — the next artifact type is covered by construction — and it appends whatever an existing .gitignore lacks instead of writing once. The write-once form was the same defect one step later: every project that had already run an audit would never have received the new pattern at all. Tests run the gate fence under bash: a fresh file carries every pattern; an image-only file from an earlier audit gains interaction/ and keeps its own header without duplicating present lines; a second run appends nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * test(#4223): declare the interaction-capture anchor as a comment marker The #4324 colon-token gate (slash-command-namespace) landed on next after this branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an unconvertible /gsd: command token. It is a section anchor of the same family as gsd:live-dom-families and gsd:write-continue, so it is declared in COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added gates against the merged tree; CI at ca8d2508 predates the gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
0977d0a475 |
fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it (#4892)
* test(#4797): failing-first on Windows — a .cmd/.bat shim under a spaced path must run * fix(#4797): run-with-timeout's .cmd mediation folds onto projectSpawnInvocation — spaces in the shim path no longer break it The private /d /s /c argv-array copy quoted the shim path (any space in the path) and cmd.exe /s stripped the first-and-last quote of the whole /c string, so the pre-space fragment became the program name: exit 1, empty stdout, on every repo whose path contains a space. The declared seam wraps the whole command line in one extra quote pair with windowsVerbatimArguments — the reporter verified the shape against the repro on Windows. * chore(#4797): backfill changeset PR number (4892) --------- Co-authored-by: sim <sim@local> |
||
|
|
2e14b4df17 |
fix(#4415): treat an absent worktree as removed, not as a branch mismatch (#4612)
* fix(#4415): treat an absent worktree as removed, not as a branch mismatch Claude Code removes a subagent's worktree the moment the subagent finishes with a clean tree. A gsd-executor that committed everything — SUMMARY.md included, under `commit_docs: true` — is exactly that case, so by the time the orchestrator reaches wave cleanup the directory is routinely gone while the branch it left behind is intact and mergeable. `git -C <gone> rev-parse --abbrev-ref HEAD` fails, and nothing distinguished that filesystem failure from a real branch disagreement: both reached the same `if`, so the entry blocked `branch_mismatch`, NOTHING merged, and the branch was left dangling. When the directory instead vanished after the merge landed, `git worktree remove` failed "is not a working tree" and the entry blocked `worktree_remove_failed`, leaving the branch undeleted and the operator to run `git worktree prune` + `git branch -D` + `rm -rf` by hand every wave. Disambiguated at the point of failure rather than ahead of it. A SUCCESSFUL in-worktree read still decides identity exactly as before — a present worktree on the wrong branch blocks, unchanged — and only a FAILED read consults the filesystem. Two reads can fail, and they are not the same path: * The branch read fails with the directory absent. There is no checkout for identity to come from, so it falls back to `refs/heads/<branch>` read from repoRoot; a missing ref still blocks, so an absent worktree never becomes a silent pass. The SUMMARY rescue and the dirty check are then skipped. * The branch read succeeded and the later `status` read fails with the directory now absent — the harness removed it while the repoRoot-side base, deletion and scope checks ran. Identity was already established from the checkout and the rescue has already run; only the dirty decision is skipped. Without this, a mid-entry removal still blocked `worktree_dirty` with nothing merged: the same bug, one window later. Skipping those reads is not a claim that the worktree was clean. This code cannot tell who removed the directory, and a forced or manual `rm -rf` of a DIRTY worktree would already have destroyed an uncommitted SUMMARY before cleanup ran. The narrow thing that is true either way is that a missing source cannot be read. The two reads also fail differently: the default SUMMARY finder catches the unreadable directory and returns no files, while `git -C <gone> status` errors — and that error is what surfaced as `worktree_dirty`. A rescue that genuinely FAILS still blocks, since a copy that errored part-way can mean an uncommitted SUMMARY was really lost. Teardown prunes the stale .git/worktrees admin entry rather than removing a path that is not there, re-reading presence instead of reusing the branch-step answer since the harness can act in between. For an entry accepted as ABSENT it prunes ONLY and never issues `worktree remove --force`: that entry was merged without the rescue and dirty checks, so force-removing a checkout recreated at that path would delete contents that never passed either one — strictly worse than the bug being fixed. A genuine prune failure still reports `worktree_remove_failed`, and a blocked teardown still withholds the branch delete. `git worktree prune` is repository-wide maintenance, not an entry-scoped operation. The presence probe resolves `worktree_path` against repoRoot, the way git does. `normalizeCleanupManifestEntry` takes the path from the manifest verbatim, so it can be relative, and every git call passes it as `-C <path>` with `cwd: plan.repoRoot`; a bare `fs.existsSync` would have resolved it against the PROCESS working directory instead. Those differ whenever cleanup runs from elsewhere, reachable today through gsd-tools' `--cwd` override, and the mismatch reads both ways: a present checkout reported absent — skipping the dirty check that would have blocked it — or an absent one reported present. An earlier cut resolved presence UP FRONT, before the branch read. That broke 52 existing tests: every cleanup-wave test uses a fake path that does not exist on disk and injects no `existsSync`, so all of them re-routed down the absent branch. Disambiguating at the point of failure leaves those tests reading as they did. Three rows still needed their premise stated — each stubs a git failure against a worktree that is genuinely present — and now inject `existsSync: () => true`. No assertion in any of the three changed. Fourteen rows added. Every early row held presence CONSTANT and so could not reach the windows that matter, since the bug is caused by a directory that changes state WHILE cleanup runs: removal after the branch read, a present worktree whose status fails (which must still block), removal between the clean status read and teardown, a reappeared checkout at teardown, #2852 isolation of a blocked absent entry from the entries after it, and relative-path resolution. Verified: ran the issue's own reproduction verbatim against a build of this branch — `merged_removed`, merge commit present, branch deleted, no prunable entry in `git worktree list`. The same reproduction against a build at the merge-base returns blocked/branch_mismatch, no merge, branch present, `wt1 ... prunable`. Five of the first eight rows go red against the true merge-base file; the three that stay green are the safety-preservation rows. The rows added after each review round go red against the commit that round reviewed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * chore(#4415): add changeset Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * fix(#4415): build the probe-path expectation with path.resolve, not path.join The row asserting that the presence probe resolves a relative `worktree_path` against repoRoot failed on windows-latest while the code under test was correct. On win32 `path.resolve` prepends the current drive to a drive-less absolute path (`/repo/main` -> `D:\repo\main`) and `path.join` does not, so a join-built expectation disagrees with correct behavior: expected: '\repo\main\.claude\worktrees\agent-a1' actual: 'D:\repo\main\.claude\worktrees\agent-a1' `path.resolve` is what the fix must use — it is how git resolves `-C <path>` against `cwd: plan.repoRoot` — so the expectation moves to resolve as well. Two `notEqual` rows keep that from being circular: the probe must receive neither the raw relative path nor a process-cwd resolution. Verified by mutation — dropping the repoRoot anchoring in `worktreeExists` turns the row red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRaNKCDUEacHudVDwvat8X * fix(#4415): confirm absence before skipping the rescue and dirty checks `fs.existsSync` answers false for a genuinely missing path AND for one it merely cannot traverse — EACCES on a parent directory, an unreachable mount. Verified: with a parent at mode 000, `existsSync` returns false while `statSync` throws EACCES. That distinction carries weight here, because "absent" is what lets an entry skip the SUMMARY rescue and the dirty check. An unreadable-but-present worktree read as absent, so cleanup merged over uncommitted work that the dirty check exists to refuse — and it contradicted this code's own comment that a present checkout whose git read fails stays blocked. Before this PR a failed git read blocked unconditionally, so treating unreadable as present is not a new safety rule; it is the one that was already there. The default probe becomes `statSync`, which reports WHY it failed. Only ENOENT is absence; anything else reads as present and blocks. An injected probe stays authoritative, so tests state presence directly with no hidden dependency on the real filesystem, and may throw to state that a path is unreadable. Two rows added: an unreadable worktree still blocks as branch_mismatch with no merge and no teardown, and a confirmed-ENOENT probe still takes the absent path. Verified by mutation — reverting the discrimination to the permissive `return false` turns the unreadable row RED while the ENOENT row stays green, which is what distinguishes discrimination from over-blocking. The mutation was confirmed to reach the compiled artifact the test loads. Also from this round: the row named for a checkout that "reappeared" never modeled reappearance (production probes presence once, at identification), so it is renamed to the unconditional contract it does prove; the comment crediting the notEqual rows with removing circularity is narrowed to what they actually establish; and the changeset now says only a confirmed absence takes the new path. Found by Codex full-PR review (round 3) before pushing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): source identity from git's registration, removal from the errno Maintainer review rejected the premise this fix rested on. It held that once the worktree directory is gone there is no checkout to read, so identity must fall back to `refs/heads/<branch>`. Git does not lose the binding — measured, after `rm -rf`: worktree /path/to/wt branch refs/heads/feat-x prunable gitdir file points to non-existent location The ref fallback weakened identity from "the checkout registered at this path is on this branch" to "a branch by this name exists", which let a foreign sibling branch merge. Identity now comes from `git worktree list --porcelain`, so the #3677 swap control keeps its teeth on the absent path; the new swap row is what would have caught this, and dropping the branch conjunct turns only that row red. Two defects in the first cut of the porcelain rework, both measured rather than reasoned about: `prunable` is not a removal test. With a parent directory at mode 000, git prints `prunable gitdir file points to non-existent location` for a checkout that is STILL THERE — it cannot traverse the parent, so it reports the gitdir file as missing. Treating prunable as "removed" would skip the rescue and dirty checks and merge over uncommitted work in an unreadable worktree, reintroducing the review's Major finding by another route. Each source now answers only what it can prove: porcelain for identity, `statSync`'s errno for removal. Only ENOENT is removal; EACCES/EIO blocks, as it did before this PR. `git worktree prune` is repository-wide. Measured: two removed worktrees plus ONE prune leaves neither registration behind. Reading the list per entry therefore let the first absent entry's teardown erase the identity evidence of every entry after it, merging one worktree per wave and blocking the rest as branch_mismatch — worse than the bug being fixed, since a wave of parallel executors is the normal case. The identity read is now a snapshot, captured lazily on the first entry that needs it and reused for the wave, which is both pre-prune and off the happy path. The `existsSync` probe and its dep locals are deleted; the filesystem is consulted only for the errno. The comment calling repository-wide prune "Harmless" was wrong under the new identity rule and says so now. Tests: identity and removal are stated on their own axes rather than through one present/absent boolean. Added the absent-path #3677 swap row, the two-absent-entry prune row, a bare `prunable` marker row, and a fail-safe row for an unreadable worktree list. Three mutations each kill exactly the intended rows, verified against the compiled artifact the tests load. One fixture that still stated presence through the removed `existsSync` seam was passing for the wrong reason and now states both axes. Verified: lint:ci exit 0; full suite 24/24 chunks, 37,164 tests, 0 failures; tests/worktree-safety.test.cjs 422/422. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): re-confirm absence before teardown, and prove the porcelain claim against real git Maintainer review, Major. Presence was classified once, at identification, and everything between that point and teardown — the base, deletion and scope gates, and the merge itself — is a window in which a worktree can reappear. The defence was "prune only, and a live checkout would make `branch -D` fail visibly", which holds only while prune's own staleness check is not fooled by the same filesystem-visibility gap that produced the false absence one call earlier. If it is, prune clears the admin entry, `branch -D` then SUCCEEDS, and a live, unreviewed, un-rescued worktree loses its branch. That asymmetry is the argument for the fix: the bug this PR set out to repair only ever BLOCKED, while this path could DESTROY state. Absence is now re-confirmed with `confirmedGone()` immediately before teardown — no new subprocess, just the statSync already in hand — and a reappeared directory blocks as `worktree_remove_failed` instead of reaching prune or the branch delete. The review was also right that the gap was known and unverified: the existing row said so in its own comment ("it does NOT model the reappearance transition itself"). It is modelled now, by a stat that answers "gone" at identification and "present" at teardown. Mutation-verified: removing the re-confirmation turns ONLY the new row red while the old "prune, never force-remove" row stays green, which is exactly why that row could not have caught this. Minor, same review: the #4415 block was entirely mock-based, so the factual claim the identity mechanism rests on was asserted in comments and measured out of band but never proved executably. Two real-git rows now prove it — that git keeps the path -> branch binding after the checkout is deleted and marks the entry prunable, and that it ALSO reports prunable for an unreadable worktree that is still there, which is why removal is confirmed by errno rather than by prunable. The second row skips as root, where mode 000 does not deny traversal. Minor 2 (rescueSummaryArtifacts resolving worktree_path against process.cwd() while the new code resolves against plan.repoRoot) is pre-existing and not reachable through the CLI's same-cwd invocation; left for a follow-up issue rather than widened into this PR. Verified: lint:ci exit 0; full suite 27/27 chunks, 37,739 tests, 0 failures, against the true merge-base; tests/worktree-safety.test.cjs 425/425. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * test(#4415): make the real-git rows platform-correct The Windows conformance shard caught both rows on their first push, and both failures were mine, not the code's. Path separators: git reports porcelain paths with FORWARD slashes on every platform, while `path.join` yields backslashes on win32, so `includes()` compared separator styles rather than paths and the registration assertions failed. Both sides are normalised before comparison now. Premise setup: the unreadable-worktree row establishes "git cannot traverse the parent" with mode 000, which win32 does not honour for directory traversal at all — the row would have asserted `prunable` against a perfectly readable worktree and failed for a reason unrelated to the behaviour under test. It now skips on win32 for the same reason it already skipped as root, with both reasons stated together. Verified: lint:ci exit 0; tests/worktree-safety.test.cjs 425/425 locally. The Windows shard is the real check for the separator fix, since macOS cannot reproduce it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL * fix(#4415): warn when an entry is accepted as absent, giving prunable its consumer Maintainer review round 3, both Medium findings — they close together, as the review noted. The absent path reported `merged_removed`/`ok` indistinguishably from an ordinary merge. This code cannot tell "the harness cleanly removed a finished executor" from "an operator or an external process removed this path": git keeps the path -> branch registration and `statSync` reports ENOENT in both cases. Before this path existed every anomalous absence blocked loudly, so accepting the routine case silently took the operator's only signal away from the case that is not routine. The module already carries an advisory channel for a materially less risky condition — scope conformance, a few lines below — so withholding one here was inconsistent with its own pattern. `WAVE_CLEANUP_WARNING.ACCEPTED_ABSENT_WORKTREE` is now emitted at both acceptance sites, carrying git's own `prunable` reason. Advisory, never a gate: the entry still merges. That also gives `WorktreeEntry.prunable` a consumer. It was parsed, documented as "worth surfacing to an operator", and then never read — the errno rework made it unused for the predicate and the parsing stayed behind. Quoting git's reason here is what it was for. The bare-marker test was vacuous, as the review said: it asserted `merged_removed`, which is driven by `confirmedGone` and the branch match, not by the bare-marker parsing it claimed to cover, so a regression in that parsing would not have reddened it. It now asserts the parsed value reaches the warning. A bare `prunable` line normalises to the literal 'prunable' — a truthiness signal, not a reason — so the warning reports null there rather than quoting a marker back at an operator as though git had said something. `WAVE_CLEANUP_WARNING`'s locked code set is updated deliberately, with the reason recorded in the test: the lock exists so a new advisory code is a decision rather than something that appears because a branch needed one. Verified: mutation — suppressing the warning at both sites turns both new rows red; lint:ci exit 0; full suite 27/27 chunks, 38,245 tests, 0 failures; tests/worktree-safety.test.cjs 426/426. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S4mJZpSNwoyVtRVUQoijfL --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
822934c901 |
fix(#4794): decision-coverage answers an unmeasured shape on could-not-parse; a non-file context path fails closed (#4889)
* test(#4794): failing-first — could-not-parse must answer an unmeasured shape; a directory context path fails closed * fix(#4794): could-not-parse answers an unmeasured shape (null counts, unreadable ids, no uncovered); a non-file context path fails closed * chore(#4794): backfill changeset PR number (4889) * test(#4794): skip the directory-identity probe when the platform cannot discriminate (windows runner volume collapse, measured) Two consecutive windows conformance runs failed the probe with measured identical (dev, ino) for two distinct mkdtemp directories (dev=3606225537, ino=9007199255243448 for both) — a runner-volume property, not a regression in the guard. On such a platform the guard's identity containment degrades to refuse-everything (fail-closed, documented); the probe asserts capability, so the honest response is an explicit t.skip carrying the measurement (ADR-2719 §6), not a red lane for every PR. --------- Co-authored-by: sim <sim@local> |
||
|
|
3d2cb1fb01 |
test(#4850): read freshness under the fixture git timeout in derivationIsNotMemoizedAcrossRenders (#4870)
The row asserts that deriveStateFreshness is not memoized across renders, and it did so through the hook's real git spawn bounded by the 1500 ms production timeout. On a loaded Windows runner the second spawn can exceed that bound, and readStateHeadCommits returns its designed null, which the exact-count assertion reads as a failure (null !== 10). Wrap the two freshness reads in the block's existing withSpawnSpy, forwarding to the real execFileSync with the fixture-scoped GIT_FIXTURE_TIMEOUT_MS in place of the production bound. The spawn stays real, so a genuine memoization regression (second read returning 5) still fails the row. No production file changes; STATE_FRESHNESS_GIT_TIMEOUT_MS stays 1500. Closes #4850 Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8a5166598c |
fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next (#4873)
* fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next Commit |
||
|
|
09e1110e68 |
fix(#4788): code spans in the decision bold lead-in are opaque to the separator grammar (#4883)
* test(#4788): failing-first — code spans in the decision bold lead-in are opaque to the separator grammar * fix(#4788): the decision lead-in runs are code-span-aware — a backticked span is data, never grammar * chore(#4788): backfill changeset PR number (4881) * chore(#4788): correct changeset PR number (4883, was a guessed 4881) --------- Co-authored-by: sim <sim@local> |
||
|
|
5906a24ede |
fix(#4786): plan-row detection accepts the bare planId stem — suffix-less hand-written lists tick in place (#4880)
* test(#4786): failing-first — a suffix-less hand-written plan list is ticked in place, never duplicated * fix(#4786): plan-row detection accepts the bare planId stem — a suffix-less hand-written list is ticked in place, never duplicated * chore(#4786): backfill changeset PR number (4880) --------- Co-authored-by: sim <sim@local> |
||
|
|
6dcc0428dd |
fix(#4784): Xcode gates pass -project and derive the destination from available simulators (#4879)
* test(#4784): failing-first — Xcode gates must pass -project and derive the destination from available simulators * fix(#4784): Xcode gates pass -project '$XCODEPROJ' and resolve the destination from available simulators, skipping loudly when none exist Also surfaces the -collect-test-diagnostics never / workflow.test_gate_timeout guidance on a 124 timeout (a real iOS suite measured ~600s of sysdiagnose collection after the tests passed). * test(#4784): reachability-honest test shapes — apostrophe-confused prose (the path the false positive actually reaches past the splitter) and the file's single-quote convention * fix(#4784): review fold-ins — unanchor the simulator extraction (real simctl lines end in a state suffix), escape apostrophes for the bash -c layer, assign the command variables on the skip path, behavioral extraction test The isolated reviewer's Finding 1 was critical: the first cut's $-anchored sed matched ZERO real simctl lines (every device line ends in (Shutdown)/(Booted)), so the gates would have skipped on every machine. The 46,454-test green could not see it (content assertions do not execute the pipeline); the new behavioral test runs the gate's own extraction against a real-format line. * chore(#4784): backfill changeset PR number (4879) --------- Co-authored-by: sim <sim@local> |
||
|
|
d435723c95 |
fix(#4782): claude's agents kind skips compact variants — consumed only by the non-claude persona gate (#4878)
* test(#4782): failing-first — claude install must not stage compact agent variants (and must prune stale ones) * fix(#4782): claude's agents kind skips compact variants — they are consumed only by the non-claude persona-fallback gate Emitted-Drift-Ack-Hash: agents/gsd-advisor-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ai-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-assumptions-analyzer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-code-fixer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-code-reviewer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-codebase-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-debug-session-manager.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-classifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-doc-writer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-dom-verifier.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-domain-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-eval-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-eval-planner.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-framework-selector.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-integration-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-intel-updater.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-mempalace-curator.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-nyquist-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-pattern-mapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-project-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-research-synthesizer.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-roadmapper.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-security-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-auditor.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-checker.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-ui-researcher.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies Emitted-Drift-Ack-Hash: agents/gsd-user-profiler.compact.md — #4782: claude never selects compact (the init agent-skills persona gate is non-claude-only), so the claude agents kind no longer stages name-shadowing copies * chore(#4782): changeset fragment * chore(#4782): backfill changeset PR number (4878) --------- Co-authored-by: sim <sim@local> |
||
|
|
4fe2837ab5 |
fix(#4774): plan-criteria R4 requires the pipe not to be doubled — a logical-OR fallback is handled, not swallowed (#4877)
* test(#4774): failing-first — R4 must not read a logical-OR fallback as a pipeline stage * fix(#4774): R4 requires the pipe not to be doubled — a logical-OR fallback is a handled failure, not a swallowed one Also corrects the rows-3/4 test's makeCriteriaPlan usage (second arg is the <verify> block, not a second criteria line). * chore(#4774): backfill changeset PR number (4877) --------- Co-authored-by: sim <sim@local> |
||
|
|
a87b83d485 |
fix(#4764): dep_phases extracts only Phase-prefixed references from Depends-on prose (#4876)
* test(#4764): failing-first — dep_phases must extract only Phase-prefixed references, never dates/shas/ledger ids/self * fix(#4764): dep_phases anchors phase references to their 'Phase' prose context and never emits the row's own number * fix(#4764): review fold-ins — hoist the anchored dep-reference grammar to phase-id, cover Oxford lists and hyphen ranges, repair the property test Adversarial review found: Oxford-comma lists under-extracted ('Phases 1, 2, and 3' dropped the tail member — a silent real-blocker clear, the dangerous direction); hyphen ranges ('Phases 1-3') kept only the first endpoint; the property test called fc.hexaString (absent in fast-check 4.8, threw every run) and passed the junk arbitrary unspread (vacuous guard) with no completeness assertion; planning-inspect's extractDependencyTokens carried the same whole-field scrape (generative-fix divergence). The anchored grammar now lives beside PHASE_NUMBER_TOKEN_SOURCE in phase-id.cts and both readers interpolate it. * chore(#4764): backfill changeset PR number (4876) --------- Co-authored-by: sim <sim@local> |
||
|
|
969456c46d |
fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path (#4875)
* test(#4759): failing-first — preserved-foreign hooks warning must not claim may-not-load for a commonjs package.json * fix(#4759): hooks marker warning states the will-not-load claim conditionally, like the plugin path * test(#4759): review fold-ins — assert installer output on stdout+stderr (warn is stderr), register cleanup before the spawn, add the type-less foreign case * chore(#4759): changeset fragment * chore(#4759): backfill changeset PR number (4875) --------- Co-authored-by: sim <sim@local> |
||
|
|
d36514b816 |
fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() (#4872)
* test(#4758): failing-first — rescue must resolve a relative worktree_path against repoRoot * fix(#4758): rescue resolves a relative worktree_path against repoRoot, not process.cwd() * test(#4758): review fold-ins — t.after cleanup pattern, post-resolution reader contract comment * chore(#4758): changeset fragment * chore(#4758): backfill changeset PR number (4872) * test(#4758): windows lanes key rescue fakes on resolved path identity, not verbatim strings win32 path.resolve rewrites driveless-absolute POSIX-style fixture values to the current drive, so the rescue's (correct) resolved-path handoff stopped matching verbatim string keys: #3804/#245/#2556/B7/#2852 fakes silently skipped the rescue and my seam test compared against a POSIX literal. Fakes now key on path.resolve(repoRoot, …) identity — the same semantics the code and git -C use — so every rescue test exercises the rescue on every platform. --------- Co-authored-by: sim <sim@local> |
||
|
|
63edc777e6 |
fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading (#4868)
* test(#4588): observed fork-from-HEAD must suppress the stale-origin degrade (failing first) * fix(#4588): observe fork-from-HEAD from prior harness worktrees before degrading * fix(#4588): a throwing probe git call is an inconclusive observation, not a crash * test(#4588): name the observed reason in the inconclusive-row assertions * test(#4588): hermetic state I/O in the inconclusive-row fixtures * chore(#4588): changeset fragment * test(#4588): contained cache I/O, emit-payload assertions (review fold-ins) * fix(#4588): review fold-ins — invalidation fixture, hermetic confirm row, gate comment, cache+scope rows * chore(#4588): backfill changeset PR number (4868) --------- Co-authored-by: sim <sim@local> |
||
|
|
9a41a95212 |
fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams * fix(#4717): consult the per-install runtime marker at both identity seams resolveReportedRuntime (agent_runtime) and loadConfigResolved (config.runtime) both ignored the per-install .gsd-runtime marker that resolveRuntime and the model-resolver gate already read. On a multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing misreported every Claude Code session as codex, and a shared defaults.json stamped by the first non-Claude install leaked its runtime to every other one. Seam 1: the reported-runtime ladder becomes explicit > install marker > host detection > claude. Seam 2: loadConfigResolved fills an empty config.runtime from GSD_RUNTIME then the marker, copy-on-write (the builtin-defaults branch returns a shared object). Explicit runtimes and marker-less trees are unchanged. * fix(#4717): a marker-detected runtime opts into its tier map (decision a) * fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins * chore(#4717): backfill changeset PR number (4861) --------- Co-authored-by: sim <sim@local> |
||
|
|
58c7bbb16a |
fix(#4667): rewrite codex @ includes to the codex install root (#4858)
* test(#4667): add failing-first coverage for the codex @-include rewrite Behavioral end-to-end: a real in-process install(true,'codex') into a temp CODEX_HOME must leave zero @~/.claude includes in GSD-owned .md artifacts, rewrite the issue's own example include to @~/.codex/, keep the deliberate _GSD_RUNTIME_ROOT .claude fallback chains byte-identical, and stay idempotent across a reinstall (no doubled prefix). All four are RED until the installer grows the manifest-scoped rewrite pass. * fix(#4667): rewrite codex @ includes to the codex install root Codex-installed agents and commands kept @~/.claude/gsd-core/... (and @/Users/trekkie/.claude/gsd-core/...) include references pointing into the Claude install: silent wrong-copy reads on dual-runtime machines at divergent versions, missing files on codex-only ones. Several emitters bypass the per-runtime converters, so the per-emitter fixes since #570 rotted. Adds a manifest-scoped rewrite pass in install() beside the leak scanner: for codex, every manifest-tracked .md/.toml artifact has the @-include forms rewritten to the codex root. The pass matches the exact include literal only — the _GSD_RUNTIME_ROOT/$PREFERRED_CONFIG_DIR fallback chains, prose .claude mentions, and CHANGELOG.md are untouched — and runs before the scanner, which remains the verification backstop for anything a future emitter introduces. * test(#4667): cover the HOME-anchored include form and sync the pass comment Adds behavioral coverage for the second rewrite literal (@$HOME/.claude/gsd-core/ -> @$HOME/.codex/gsd-core/) via the plan-review-convergence command, and corrects the pass's comment: the agent .tomls are generated after it and prefix themselves, so the .toml branch of the rewrite is inert by design. * test(#4667): baseline the offline upgrade against the deployed tree (sanctioned) * chore(#4667): backfill changeset PR number (4858) * test(#4667): sandbox the install home in the include-rewrite tests (#3712 guard) * test(#4667): install via subprocess with an isolated env in the include-rewrite tests --------- Co-authored-by: sim <sim@local> |
||
|
|
e1f72cd324 |
fix(#4741): the plan checkbox tick respects the superseded exclusion (#4851)
* test(#4741): a superseded plan must not be ticked from its summary (failing first) * fix(#4741): the plan checkbox tick respects the superseded-plan exclusion * fix(#4741): review fold-ins — changeset typo, dedupe planId derivation * test(#4741): exercise a halted summary on the active plan in the #2830 pin * chore(#4741): backfill changeset PR number (4851) --------- Co-authored-by: sim <sim@local> |
||
|
|
11b3091df0 |
fix(#4738): record opencode's staged skills in the install manifest (#4847)
* test(#4738): opencode manifest must record its staged skills (failing first) * fix(#4738): record opencode's staged skills in the install manifest * test(#4738): use the centralized temp-dir helper in the manifest tests * fix(#4738): retire the dead hostBehaviors vocabulary entry, tighten detector asserts, temper changeset * chore(#4738): backfill changeset PR number (4847) --------- Co-authored-by: sim <sim@local> |
||
|
|
c5629bbe74 |
fix(#4734): degrade worktree isolation when the root has no git repository (#4843)
* test(#4734): non-git root must degrade worktree isolation (failing first) * fix(#4734): degrade worktree isolation when the root has no git repository * fix(#4734): review fold-ins — 3972 ladder fixture, parity fixture, docs row, message wording * chore(#4734): backfill changeset PR number (4843) --------- Co-authored-by: sim <sim@local> |
||
|
|
8d0b6868ae |
fix(#4725): write normalization preserves tight paragraph-list shape (#4842)
* test(#4725): write normalization must not reflow untouched prose (failing first) * fix(#4725): stop write normalization injecting a blank before a list after prose * test(#4725): repair ordered-list fixture and list-spacing snapshot * test(#4725): assert whole-file prose stability, fix heading-list comment * chore(#4725): backfill changeset PR number (4842) --------- Co-authored-by: sim <sim@local> |
||
|
|
c9a5cc3e12 |
fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head. |
||
|
|
bff99a8bb5 |
fix(#4731): read hard-wrapped Goal/Requirements fields past the line break (#4826)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; isolated adversarial review round completed (MEDIUM table-bleed finding fixed with RED/GREEN evidence) and sha-pinned bench 46331/0 on the merged head. |
||
|
|
fb3e228a0d |
fix(#4724): classify Surefire/Failsafe XML as RED evidence (#4825)
* test(#4724): add failing-first coverage for Surefire XML RED evidence * fix(#4724): classify Surefire/Failsafe XML as RED evidence check tdd-red-evidence parsed only node:test TAP, so a JVM project's genuine Maven red scored INVALID_RED while hand-written synthetic TAP scored RED_EVIDENCE_OK — the gate was passable only by fabricating its input (issue #4724's measured repro). classifyRedEvidence detects Surefire/Failsafe XML (a <testsuite> element) and parses it by TAG-BOUNDARY scanning: each <testcase> owns its own tag (self-closing) or the segment up to its </testcase> closer, so the issue's warned-about spanning trap (a lazy lazy match from a green self-closing case to the next closing tag) cannot misreport names. A <failure> or <error> child marks the case failing; the target matches at class granularity (exact classname, dotted-suffix, or method name). Any parse anomaly degrades to not-failing — the module stays fail-closed and PURE (no fs/clock; report freshness remains the workflow's run-start check per the issue's implementation notes). TAP classification is byte-identical: all existing fixtures stay green. * test(#4724): pin the scanner hardening — truncation, TAP-message flip, CDATA phantom * docs(#4724): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
7d0c6339d0 |
fix(#4705): emit Antigravity-native tool names as a YAML sequence (#4822)
* test(#4705): add failing-first coverage for native Antigravity tool sequences * fix(#4705): emit Antigravity-native tool names as a YAML sequence convertClaudeAgentToAntigravityAgent and the installer's twin emitted Gemini CLI tool names as a comma-separated scalar. Antigravity's documented subagent contract (antigravity.google/docs/subagents) wants a YAML sequence of native names — view_file, grep_search, run_command, replace_file_content are the documented examples, and wrong or malformed grants can hang the subagent per Antigravity's own warning. Map values move to the native vocabulary where documented (Read -> view_file, Edit -> replace_file_content, Bash -> run_command, Grep -> grep_search); undocumented entries keep their best-known grant rather than being dropped (dropping would silently remove a restriction). The emitter writes one '- name' item per line; an agent whose every tool was filtered emits an explicit tools: [] instead of an empty scalar. Pre-existing pins updated to the native vocabulary. * test(#4705): update the #4727 map-value pin to the Antigravity-native vocabulary The #4727-era pin held the map VALUES at the Gemini CLI dialect on the belief that Antigravity speaks it; the confirmed bug #4705 (with Antigravity's own documented subagent contract) supersedes that for the four documented names. Key/shape pinning is preserved; only the values move. * docs(#4705): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
be1b76dddd |
fix(#4700): queue the headless mempalace mine on the palace lock (#4821)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase * fix(#4699): skip already-complete phases in the next_phase cascade Both next-phase scans selected the numerically lowest phase above N without consulting completion state, so completing a reopened phase persisted an already-[x] phase as STATE.md current_phase while roadmap.analyze correctly named the outstanding one (issue repro: completing 2 with phases 1 and 3 already [x] returned next_phase 03). The cascade collects the complete phase numbers from the roadmap checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in both the disk scan and the roadmap scan; a [x] checkbox row and its heading sibling both name a phase that is never next. Heading-only and checkbox-less roadmaps behave exactly as before. * test(#4699): align the negative-control expectation with the disk spelling * test(#4699): pin the STATE.md persistence and the all-later-complete tail corner Review findings: the regression never asserted STATE.md current_phase (the issue's actual harm), and the all-later-phases-[x] corner (is_last_phase true, next_phase null) was unpinned. A changeset fragment is included. * docs(#4699): backfill changeset PR number * test(#4700): add failing-first coverage for the queued headless mine * fix(#4700): queue the headless mempalace mine and surface skipped captures The capture's mine ran in the foreground with no lock handling: MemPalace wraps every mine in a per-palace lock, so any concurrent writer (two phases finishing a stage at once, a git-hook refresh mining the same palace) made it exit 1 (MineAlreadyRunning) and the onError: skip step silently dropped the capture — unlost for CONTEXT/PLAN/SUMMARY files that can be re-filed, unrecoverable for execute:wave:post problem-fix pairs. The mine now queues via --daemon --background (MemPalace #2029: the daemon holds a job refused the lock and runs it when the holder exits), and the report step gains the queued and skipped outcomes per #4700's requirement that a skipped capture never stay silent. Option 2 (write_routing.cli) is unreleased at MemPalace 3.9.0; option 3 (retry) re-enters the same lock race — both declined in the PR body. * fix(#4700): queue the wave:post problems fragment's headless mine too The issue names the execute:wave:post problem-fix pair as the unrecoverable loss (no source file to re-file later); the capture-problems fragment's headless mine ran foreground like the capture capability's did. Same fix: --daemon --background, with the lock-deferral rationale inline. * docs(#4700): backfill changeset PR number * fix(#4682): register the stale-reverification part in the capability registry The new steps/ part is a shipped workflow file; gen-capability-registry --check requires it in the committed registry. --------- Co-authored-by: sim <sim@local> |
||
|
|
d707318e0c |
fix(#4699): skip already-complete phases in the next_phase cascade (#4820)
* test(#4699): add failing-first coverage for skipping complete phases in next_phase * fix(#4699): skip already-complete phases in the next_phase cascade Both next-phase scans selected the numerically lowest phase above N without consulting completion state, so completing a reopened phase persisted an already-[x] phase as STATE.md current_phase while roadmap.analyze correctly named the outstanding one (issue repro: completing 2 with phases 1 and 3 already [x] returned next_phase 03). The cascade collects the complete phase numbers from the roadmap checkboxes (milestone-scoped, comparePhaseNum-deduped) and skips them in both the disk scan and the roadmap scan; a [x] checkbox row and its heading sibling both name a phase that is never next. Heading-only and checkbox-less roadmaps behave exactly as before. * test(#4699): align the negative-control expectation with the disk spelling * test(#4699): pin the STATE.md persistence and the all-later-complete tail corner Review findings: the regression never asserted STATE.md current_phase (the issue's actual harm), and the all-later-phases-[x] corner (is_last_phase true, next_phase null) was unpinned. A changeset fragment is included. * docs(#4699): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
2bfff17ff8 |
fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing * fix(#4682): route stale verification to the verifier regeneration path The stale routing entry sent users to /gsd-verify-work — but verify-work never rewrites VERIFICATION.md (its only write is the human_needed canonicalization), so following the advice re-ran UAT, reached the same stale check, and looped. init's projector and execute-phase's generic next_command presentation both mirror this entry, so the dead end appeared on three surfaces. The stale entry now routes to execute-phase, and execute-phase's all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping the spine under its frozen ADR-857 ceiling) mirroring the missing route: skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier — regenerating VERIFICATION.md and its digest, marked phase or not. The non-stale fall-through. Staleness detection, the digest format (#4623), every other routing entry, and the #3684 resume arms are untouched. Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682) Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682) * test(#4682): register the stale-reverification part and align projected commands The new steps/ part must be registered in the inventory manifest and the per-runtime golden install trees (regen:derived); the projected stale next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the runtime surface), the human_needed bare-report probe keeps routing to verify-work (unchanged semantics), and init-manager's recommended action follows the new command. * test(#4682): prefix the remaining stale routing assertions with the runtime surface Nine stale next_command assertions and the human_needed bare-report probe still carried the unprefixed or flipped forms from the earlier line-number edit; all now assert the shipped /gsd-execute-phase <phase> projection, with the human_needed probe reverted to its unchanged verify-work routing. * test(#4682): align the last stale projection assertions with the execute-phase route * docs(#4682): backfill changeset PR number * test(#4682): refresh the compact-content baseline after the rebase The rebase onto the #4670 squash brought verify-work.md's bounded reconciliation text into this branch; the committed compact-content baseline now reflects the post-rebase split sizes. Local --check is clean; the previous bench drift (+243) was the baseline, not the diff. * fix(#4682): carry the response_language directive in the stale-reverification part The new steps/ part is its own coverage unit for lint-response-language-coverage; it takes the shared canonical directive line like its sibling execute-phase parts. --------- Co-authored-by: sim <sim@local> |
||
|
|
651511d1e3 |
fix(#4670): bound the commit-claim window to the plan's own history (#4813)
* test(#4670): add failing-first coverage for the bounded commit-claim window * fix(#4670): bound the commit-claim window to the plan's own history The reconciliation measured plan_head_before..HEAD — a window that grows with every later plan's task and SUMMARY commits plus execute-phase's own phase-completion commit — so an honest plan flagged commit_claim_mismatch as soon as anything landed after it (real project: claims 3/5/2 measured 20/10/5). The executor now also records plan_head_after (HEAD at its measurement moment, after the last task commit, before the SUMMARY commit), and verify-work reconciles exactly against plan_head_before..plan_head_after with a merge-base ancestry check; SUMMARYs without the anchor fall back to the legacy warning path instead of an unsound BLOCKER. Both #3968 failure modes (claimed commits never made; task commits lost) still block, driven by the issue's own fixture scenarios. Emitted-Drift-Ack-Growth: verify-work.md — reconciliation gains the bounded window and legacy fallback (#4670) Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670) * fix(#4670): name the history-rewrite case and keep the executor under its cap Review findings: the BLOCKER enumeration named only the two #3968 causes, so an honest plan whose recorded window was rewritten afterwards (rebase, amend, cherry-pick) got a mislabeled diagnosis — the clause now names that case with the manual-recount remedy. The executor's growth crossed the LARGE-tier hard cap (49152), so the plan_head_after documentation is compressed to the minimal capture + frontmatter write (verify-work.md carries the semantics), the #2751 PROSE_ALLOWLIST entry is re-pointed at the shifted line (#4670 moved it from 823 to 825), the compact-content baseline is regenerated, and the changeset records the two un-established edges (shared-base waves, subrepo ledgers). Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670) * docs(#4670): backfill changeset PR number * docs(#4670): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
652796e903 |
fix(#4665): route --fix past the empty-scope exit (#4810)
* test(#4665): add failing-first contract coverage for the --fix empty-scope recovery * fix(#4665): route --fix past the empty-scope exit check_empty_scope exited the entire workflow whenever REVIEW_FILES was empty — before dispatch-fix — so with #3661's incremental scoping, a phase whose only post-review changes were planning artifacts could never run --fix against its standing REVIEW.md findings, and the skip output did not even mention the flag. The skip is now a self-contained guarded fence (explicit REVIEW_FILES emptiness check): it fires only when --fix is absent OR the phase's REVIEW.md does not exist. Otherwise the workflow proceeds directly to dispatch-fix, which delegates to code-review-fix.md — the canonical fix implementation that already documents handling an existing REVIEW.md — while the fresh-review steps (structural pre-pass, reviewer lanes, spawn_reviewer, commit_review) are skipped: nothing new to review, nothing to commit. dispatch-fix.md's route docstring is synced. Emitted-Drift-Ack-Growth: code-review.md — check_empty_scope gains the --fix recovery branch (#4665) * docs(#4665): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
85545a77a5 |
fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate verify-work.md's complete_session step flips VERIFICATION.md to passed on 'zero issues' alone, so a session whose every UAT row is blocked (a session that observed nothing) canonicalizes the report. Pins the deployed contract the fix must satisfy: the flip runs the unflagged phase uat-passed predicate inside the human_needed branch, frontmatter.set sits inside a passed==true guard, a refusal message carries the blocker count and keeps human_needed, and an indeterminate pre-check fails closed. All four new assertions are RED until the workflow grows the guard. * fix(#4663): gate the canonicalization on the uat-passed predicate complete_session flipped VERIFICATION.md to passed whenever the session recorded zero issues and the status was human_needed — but blocked rows are not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked session (one that observed nothing) rewrote the canonical report to passed. Every later reader (transition.md's preliminary check, resume paths, validate-phase, verification.status) then inherited the unearned pass while the phase-close predicate correctly refused it. The flip now runs the phase-close predicate in a new --uat-only form before canonicalizing: UAT rows evaluated (at least one pass, no pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers skipped — they must be, because the report still reads human_needed at pre-check time and that status is itself a blocking verification entry, so the full predicate could never pass there and the flip would deadlock (found by isolated review, probed). The --require-verification call stays the transition gate; the refusal branch reports the blocker count and keeps human_needed; an indeterminate pre-check fails closed. Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * test(#4663): align the canonicalize pre-check needles with the shipped line The workflow line carries a 2>/dev/null redirect the needles did not include, so both pre-check assertions fail against the committed fix (fixed-string grep verified). Reviewer-found; needle and message aligned. * fix(#4663): reword the canonicalize prose and refresh its size baseline The rationale paragraph mentioned the flagged transition-gate call by its flag, putting a --require-verification literal before the first phase uat-passed occurrence and breaking the existing ordering pin; the prose now describes it without the literal. verify-work.md's growth also drifted the committed compact-content baseline; regenerated via benchmark-compact-content.cjs --write (derived artifact, report-not-gate contract). Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * docs(#4663): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
c72fb34e9a |
fix(#4658): give the ui plan gate's evidence check a native branch (#4807)
* test(#4658): add failing-first coverage for native frontend evidence hasStaticFrontendEvidence recognised only JS-ecosystem evidence, so computeUiPlanGate could never block for a SwiftUI/Compose/Flutter/XAML project. Adds evidence-level fixtures for the four suggested markers (import-matched for .swift/.kt/.dart, extension-alone for .xaml), the reporter's non-UI Swift control case, marker-exactness and SKIP_DIRS and I/O-degrade negatives, gate-level block assertions through makeProject's new native frontendEvidence modes, and a pinned-seed fast-check property. All new assertions are RED until src/ui-frontend-evidence.cts grows the native branch. * fix(#4658): give the ui plan gate's evidence check a native branch hasStaticFrontendEvidence recognised only JS-ecosystem evidence (a root package.json UI-framework dep, or a .tsx/.jsx/.vue/.svelte file), so computeUiPlanGate could never block for a SwiftUI, Jetpack Compose, Flutter, or .NET MAUI project — the #3312 gate was structurally unreachable for them. Adds a native BFS over the same bounds and skip rules: .xaml is evidence by extension alone (the .tsx analogue), while .swift/.kt/.dart count only when their content carries the ecosystem's UI import marker (import SwiftUI / import UIKit, androidx.compose, package:flutter) — matched on the import, not the extension, so a non-UI Swift package stays silent exactly as the issue's 37-file control case requires. Marker reads are bounded to a 64 KiB prefix; any I/O failure degrades to false per the module contract. The #3718 vocabulary filter, the JS evidence rules, and the weaker-extension exclusion are untouched. * chore(#4658): regenerate the macos conformance tier list The native-evidence additions to tests/check-ui-plan-gate.test.cjs move the file into the macOS conformance tier per the classifier; the committed generated list is a derived artifact and must match the live tests/ tree (the same sync the fragment-single-edit-propagation install test enforces). * fix(#4658): accept both Dart quote styles and extract the shared bounded walk The isolated reviews' remaining findings: the Dart marker carried only the single-quote anchor, missing legal double-quoted imports (a spec-narrowing deviation); the two evidence walks duplicated the subtle MAX_WALK_ENTRIES cap semantics verbatim, so they are extracted into one walkProjectFiles BFS with a visit callback; tests now use the createTempDir helper, shared fixture literals that cannot drift from NATIVE_UI_CONTENT_MARKERS, a double-quoted Flutter import case, and drop a vacuous assertion and a mid-body re-require alias. * docs(#4658): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
5e729445d3 |
fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first proposeConsiderations regression (Danish text + English text_en must classify as its English equivalent, not land in the #1110 unclassified sentinel), proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en validation cases (empty/whitespace/non-string, unconditional under an elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a reference-doc Inputs parity test, and a fast-check property proving any cue-matching prose classifies identically under a cue-free Danish rendering plus text_en. All new assertions are RED until src/ui-consideration-probe.cts and the workflow/reference docs are updated. * fix(#4657): give the ui consideration probe a text_en language channel Element gains an optional text_en; classifyElement's own signature stays untouched (a locked, directly-tested export) and the text_en ?? text selection is pushed to the two classification call sites (proposeConsiderations, proposeElements) instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero kinds. Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the Non-English projects section (mirroring spec-phase Step 5.5) and the ELEMENTS_JSON shape comment documents the field with both zero-applicable guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT predicate (with both derived indexes regenerated), and the nav-override test expectation stay in sync. The ui-phase contract test carries the site-scoped allow-test-rule marker and its cluster is registered in the test-file-count allowlist ratchet. Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657) * docs(#4657): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
3014775a3f |
fix(#4656): expose coverage.unclassified and widen the zero-applicable guards (#4800)
* fix(#4656): expose coverage.unclassified and widen the zero-applicable guards * fix(#4656): regenerate golden coverage fixtures and update the rollup pin Emitted-Drift-Ack-Growth: spec-phase.md — #4656: guard widened to the all-unclassified case, doc claim corrected Emitted-Drift-Ack-Growth: ui-phase.md — #4656: guard widened identically * fix(#4656): sync edge-probe doc blocks and coverage pins with the new field * fix(#4656): key the mandatory confirmation on the widened guard * docs(#4656): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |