5214ad5802db4e8ae50c98a7c771da9f2fb878bc
1074 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d29b50d696 |
fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md * fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding * fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch. * chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names * fix(#4051): review fixes — em-dash description style, split audit-fix route * chore(#4051): sync skill mirrors of execute-phase/phase descriptions * chore(#4051): add changeset (pr backfill to follow) * chore(#4051): backfill PR 4289 in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
b1b7cabfb5 |
docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry Adds one `type: "eos"` entry for a Qoder host integration and regenerates docs/registries/eos-registry.md. Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop). The integration depends on @opengsd/gsd-core, negotiates the ADR-1239 host-integration handshake, and projects GSD's agents, skills, and hook scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the China edition), merging GSD's lifecycle hooks into settings.json. Every axis is sourced from Qoder's own docs per the never-infer rule. `dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a frontmatter-declared, per-agent-definition property, and GSD's two isolation negotiation models both assume a per-dispatch injection point Qoder does not expose. Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in favor of the EoS path. Closes #4123 * docs(#4123): backfill changeset pr field |
||
|
|
75ee7b0214 |
enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate One query verb computes TDD-applicability for a plan (CLI flag, plan type: tdd frontmatter, a task's tdd="true" attribute, or the workflow.tdd_mode config default), mirroring phase.mvp-mode's precedence-cascade shape. Foundation for epic #4272 Phase 2, which wires both dispatch backends to consume it instead of restating the predicate independently. Also fixes workflow.tdd_mode, workflow.research, and workflow.nyquist_validation, which never reached cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone because loadConfig() never populates config.workflow — a dead accessor found while wiring this verb's own config read, fixed inline per the no-defer rule rather than left alongside it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4273): document phase.tdd-applicable's FEATURES.md entry Add a docs/features/ fragment for the new phase.tdd-applicable query verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left untouched: it documents /gsd-* slash commands only, and the sibling verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI reference entry anywhere in docs/ either -- only inline prose mentions in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md precedent to extend. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): stop whitelisting capability-owned config keys centrally workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are each already owned by their own first-party capability's federated config schema (the tdd/research/nyquist capabilities declare them under their own capability.json `config`), resolved via isCapabilityConfigKey. Adding them to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as the prior commit in this branch did (mirroring workflow.mvp_mode, which genuinely is central-only), declares the same key in two places at once. That collision breaks capability-loader.cts's loadRegistry composition: gsd-test caught this as 84-85 unrelated failures across capability-cli/capability-command-dispatch/capability-lifecycle test files, every one showing "unknown capability: <id>" for a freshly-installed third-party capability that should have resolved fine. Verified directly (not asserted): reverting only this file, keeping the config-loader.cts tdd_mode/research/nyquist_validation flattening and the init.cts call-site fixes from the prior commit, and re-running the exact capability install + capability set repro from tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the failure with the whitelist entries present and clears it without them. loadConfig() still surfaces all three flattened values correctly with no central whitelist entry (confirmed directly against the compiled module) — the whitelist additions were never required for the #4273 fix to work; they were an incorrect over-application of the mvp_mode precedent to keys that aren't central. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4273): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f4bf449296 |
fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat cmdAuditUat admits `human_needed` OR `gaps_found`, but parseVerificationItems had a body only for the first and returned an empty array for the second — standing on a comment deferring to `plan-phase --gaps`, a different command audit-uat never reaches. Since cmdAuditUat pushes a file into `results` only when `items.length > 0`, a `gaps_found` report did not under-report: it vanished, taking its phase's `by_phase` row with it, so a clean-looking total gave the reader no cue anything was skipped. Eligibility now has one owner (the caller) and parseVerificationItems reports what the file says. The closed-entry filter could not be built on extractFrontmatter: its array-item parser keeps only each `- ` entry's FIRST line and has no notion of nested key/value objects, so an entry's `status:`/ `resolution:` siblings never reach its output and a closed entry is indistinguishable from an open one downstream. Rather than grow a competing object-list parser — or change extractFrontmatter, whose blast radius is every frontmatter consumer in the repo — this reads the raw segment BEFORE the flattening, via the existing anchored sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps` machinery that already parses exactly this `- `-opened, indentation- continued shape. The human_needed path is byte-for-byte unchanged: same reader, same display names, same numbering, no resolved-entry filtering — pinned by a test and verified by identical CLI output on base and head. parseGapsItems keeps its narrower `status: resolved` rule so no *-UAT.md behaviour moves. Closes #3850 * chore(#3850): backfill changeset pr number for #3879 * fix(#3850): one parse per entry, one fence parser, one resolved-entry rule Adversarial review on #3879: B1, B2, M3, m5, m8 and n9. B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell 5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found` report vanished from the audit exactly as it did before this fix — this issue's own symptom, on a platform the repo already has a named defect class for. `extractFrontmatter`'s BOM+fence logic is now factored out as `frontmatterRegion` and shared. One fence parser, not two. B2 — the resolved-entry skip paired two DIFFERENT parsers by array index: `parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A block sequence written at its key's indent — ordinary, legal YAML — makes them disagree about entry count, and from the first disagreement every index names a different entry, so an OPEN entry inherits a CLOSED one's resolution and is silently dropped. That is the defect this PR exists to fix, reintroduced inside the fix. Display name and sibling fields now come from ONE parse of the raw slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as `parseYamlRegion` does, so the string is byte-identical to what `extractFrontmatter` produced. The flattened array remains the #2286 GATE, but is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment. M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both readers use, rather than two copies differing only in `result`. m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an acceptance criterion #3850 does not contain: the issue has no AC section, and its suggested fix (2) states the skip unconditionally, naming a file with 14 of 16 entries resolved. That file is `human_needed`, so the asymmetry left the reporter's own scenario over-reporting by 14. m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and says the column-0 boundary rule is now a cross-module contract. n9 — the vestigial bare block is gone and its body de-indented. Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF fixture (M4 — it survived by accident, now pinned) and the unified skip rule. Fail-first verified by running the new tests against the pre-fix build: the BOM, nested-sequence and unified-skip cases are red there. * fix(#3850): read the entries as objects, not as re-parsed display text Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1 (#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml: `parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry now flattens to `test: A, resolution: R` rather than to its first line. The original mechanism existed ONLY to work around that lossy first-line flattening — it sliced the raw frontmatter segment and re-parsed each entry by hand so a `resolution:` sibling was visible at all. With a real parser upstream that workaround is obsolete, so it is deleted rather than repaired: `sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the `splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are all gone. `frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` — the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence, same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping one step before the display flattening. `flattenObjectListItem` is exposed alongside it so a caller deriving a display name produces the byte-identical string `extractFrontmatter` would have. That collapses the review's blockers into properties of the parse rather than things this fix has to get right: - B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI. - B2 (index pairing) — there is no second reader. Display name and sibling fields come from one object. - M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers. - M4 (CRLF) — js-yaml's, not ours; verified through the CLI. Also confirmed on the rebased base, per review: #3850 still reproduces on `next` after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found` fixture), so this PR is still doing work #3707 did not do. Nothing was dropped as redundant. One behaviour note: `entryField` returns a present value verbatim and treats only whitespace-only as absent. Trimming would rewrite an author's `truth:` on its way to becoming the display name. * fix(#3850): keep every frontmatter list entry at its own row Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result to objects, and filtering COMPACTS: `parseHumanVerificationItems` then numbered the survivors by their position in the compacted array. On a list mixing object and non-object entries the non-object rows disappeared outright and the rest were renumbered — #3850's own vanishing-row defect, reached through entry SHAPE instead of file STATUS. Base never had it: it walked the display array, so every row surfaced at its own position. Renamed to `frontmatterListEntries` and it no longer filters (the name now matches what it returns). Deciding what a non-object entry MEANS is a caller's judgement; dropping it is nobody's. Both readers now walk the DISPLAY array — one element per row, the array #2286 already gates on — and consult the parsed array only for "does this entry carry a closure field?". `parsedEntriesFor` owns that pairing and checks the two lengths agree before trusting an index; all-null is the correct degradation, since over-reporting a closed row is recoverable and closing the wrong one is not. Names stay byte-identical to base for every entry shape, including a nested sequence (`[nested]`, not `["nested"]`). Same class closed in the gaps reader: a non-object `gaps:` entry surfaced nothing at all and now surfaces as `unknown`, which is this module's documented fail-safe direction (`parseGapsItems`) on a false-negative bug. Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule inlined twice; `extractFrontmatter` now routes through it, so "one fence parser" is enforced rather than asserted in a comment. Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and its "shared by both readers" comment corrected — it has one call site, and the two readers differ deliberately, each mirroring its own established sibling (`parseGapsItems` vs #2286). Documented at the divergence. Tests: `B2` asserted a name substring, so it passed while the row was mis-numbered and would have passed through outright loss; it now asserts positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim, B2c the survivors' file positions across skipped rows, B2d the gaps reader. All four fail-first against the reviewed head; 332/332 green with the fix. * fix(#3850): make status authoritative, and let the two gaps readers agree Round 4 review, all five findings. Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as closure regardless of `status:`, so `status: failed` + `resolution: "attempted retry, still failing"` vanished from the report — the silently-vanishing-item defect #3850 exists to close, reached by field combination instead of file status. Closure is now per key, because the two keys have different conventions and one rule cannot serve both: `gaps:` `status: resolved` only, byte-identical to the rule `parseGapsItems` applies to a `## Gaps` markdown section, so one authored entry cannot read closed in one reader and open in the other. `human_verification:` a bare `resolution:` still closes, since that is how verifier-written entries record it — but a readable `status:` that contradicts it wins. A single unified rule was the first draft and is wrong: it closes a frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which `parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring claims it mirrors that reader's fail-safe status handling. The contradiction guard is not a judgment call about YAML. It is the rule this codebase already applies to the same field pair: `validateResolution` (probe-core.cts) rejects a populated `resolution:` on a non-resolved status outright — "a populated payload is an authoring mistake ... Reject it so the mistake surfaces." A reporter cannot throw, so it surfaces the item. Minor 1. Direct unit tests for `frontmatterListEntries` and `flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that historically co-changes with `frontmatter.cts`. They were reachable only through `uat.cts`' readers before. Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted directly. Verified unreachable through content rather than assumed: both readers enter through `frontmatterRegion`, `extractFrontmatter`'s only extra argument gates a warning, and `normalizeParsedValue`'s `value.map` is 1:1. It is a drift alarm for a future edit to either parser, so the helper is exported for tests rather than left as the one unpinned branch. Minor 3. The vestigial `const skipResolved = true` and its dead conditional are gone. Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:` entry has no `test:` in its vocabulary — the template's entries carry truth/status/reason/artifacts/missing — so it was speculative support for a field the shape does not have, and it collided with the 1..N row numbers `parseHumanVerificationItems` assigns by array position. Not reading it makes the collision impossible; an offset would have rewritten an authored value, against `entryField`'s verbatim contract. Docs, changeset and the dispatcher docstring all stated the unconditional rule and are corrected — three prior rounds here were comment/code drift. Fail-first proven: restoring the universal rule reddens all three new unit tests and both rewritten properties. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
585a8b7f1b |
fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity * fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path * fix(#3747): scope doc evidence claim to skills discovery per adversarial review * chore(#3747): add changeset * chore(#3747): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
1fe85cd43e |
chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk Repo-wide sweep (ahead of adding lint rules for these exact bug classes) found both incident patterns still live and unfixed on `next`: - scripts/run-tests.cjs's sweepProtectSet walk stopped on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. win32 dirname('D:\') is a fixed point (length 3, never satisfies `> 1`... wait, it does satisfy length>1), so a selected file living outside runTempRoot (the common case) spins the walk forever on Windows. Extracted a pure, exported computeSweepProtectSet helper that terminates on dirname(cur) === cur instead, with in-process RuleTester-style coverage for both win32 and posix paths. - tests/run-tests-temp-root.test.cjs's own #4020 regression test set only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never reads TMPDIR on Windows (only TEMP, then TMP), so the redirect silently no-oped there — masked because Windows CI died in the dirname-walk hang above before ever reaching this test. - tests/config-schema.property.test.cjs's fallow config-set test had the same TMPDIR-only pattern, direct process.env assignment this time, restored in its own finally block. Origin: #4220 and its shared root cause #4020. * feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for either shape. - local/require-full-tmpdir-triad: flags a TMPDIR environment override (direct process.env.TMPDIR assignment, or a TMPDIR property in a spawn-like call's env: object literal) not accompanied by TEMP and TMP in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows. Registered on tests/**/*.cjs, matching the require-userprofile-with-home precedent. - local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning from dirname() with no fixed-point termination guard (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a no-op at the platform root, but the value differs by platform (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped length/equality bound never fires on Windows. Registered on BOTH tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in scripts/run-tests.cjs, not tests/. Both rules join the zero-escape-hatch discipline already established for this catalog (no bespoke comment marker; PROTECTED_RULES in tests/portability-rule-disable-ban.test.cjs independently bans eslint-disable of either). ADR-1703 and its two companion contributing docs get an amendment documenting the mechanism, code examples, and the repo-wide sweep (three live instances found and fixed in the prior commit; no others found). CI test-scope selection updated so an edit to either rule or to scripts/run-tests.cjs re-runs the right suites. * fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too checkWhile bailed out early unless node.test was a LogicalExpression, so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } -- was silently skipped and never reported. That is the EXACT minimal shape of the original #4020/#4220 bug, and it is literally the shape used by this rule's own shipped RuleTester fixtures (the "equality-only bound" invalid cases), which were failing (0 errors reported, 1 expected) until this fix -- confirmed by running RuleTester directly against both fixtures, not just via a passing test-runner exit code. The conjunct-collection helper already handled a non-LogicalExpression test correctly (it pushes a single node as the sole conjunct); only the early-return gate needed to stop requiring a compound && / || test. Verified: RuleTester run directly against both previously-broken fixtures plus two new sanity cases (a guarded single-condition loop stays valid; an unrelated single-condition loop stays silent), and a fresh `npx eslint .` across the whole repo remains clean (no other single-condition dirname-walk shape exists in the tree). * fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call isSpawnLikeCallee only recognized a MemberExpression callee (child_process.spawnSync(...)) or a bare identifier in ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare -- const { spawnSync } = require('child_process'); spawnSync(...) -- has an Identifier callee named "spawnSync", which matched neither branch, so the whole env-literal check was skipped. gsd-test caught this: both "invalid: child_process.spawnSync with TMPDIR-only env" cases in tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors reported, 1 expected). Widened the bare-identifier branch to also match any of the known ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same lightweight convention this repo's other eslint-rules/*.cjs use (e.g. no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow tracing. Verified: RuleTester run directly against all 11 cases in tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that were failing), all pass; a fresh npx eslint . and npm run lint:ci across the whole repo remain clean. * fix(#4244): correct a stale escape-hatch reference in a test comment The comment on the "length comparison against another expression's length" case referenced a "// allow-dirname-walk marker" that doesn't exist -- the rule has zero comment-based escape hatches by design (ADR-1703), and an earlier draft's marker mechanism was removed before this branch's first commit. Spec-axis review caught the stale reference. No behavior change; comment-only. * chore(#4244): backfill changeset PR number (pr:0 -> pr:4246) --------- Co-authored-by: sim <sim@local> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
91ed46882a |
feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives Adds the full behavioral (tests/quick-batch.test.cjs) and property-based (tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core primitives per the phase's 35-row test matrix — task-list parsing (inline + --file, with path-confinement/symlink-escape/non-regular-file rejection), collision-safe quick-id preallocation under withPlanningLock, BATCH.json schema/validation/resume, dependency-DAG + partitionByFileOverlap wave construction, and exactly-once STATE.md completion (including the STATE-row-written-but-manifest-not-yet-updated crash window). The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a not-yet-written src/quick-batch.cts) does not exist yet — every test in both files fails at the top-level require() before any assertion runs. Five fast-check properties cover collision-freedom under lock contention, resume idempotency, exactly-once STATE completion, wave totality, and DAG-respecting wave order, per the design doc's property-based-coverage requirement. * feat(#3675): implement quick-batch core primitives Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core primitives per the phase design lock — pure/state primitives and CLI-testable core operations only, no agent dispatch, no worktree creation, no user-facing command (Phase 4/#3676's job): - parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list parsing (>=2 items required) and a --file variant strictly confined to the planning workspace root via requireSafePath, rejecting non-regular-file targets. - allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id preallocation under withPlanningLock, checked against both on-disk .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json manifests (never on-disk-only, which would miss another in-flight batch that hasn't dispatched any real quick directory yet) — replicates cmdInitQuick's own grammar rather than delegating to it (that function's 2-second granularity is not batch-safe). - computeWaves: deterministic wave construction combining dependency-DAG layering with partitionByFileOverlap (#3674), called per DAG layer over path-separator-normalized planned_files — normalization happens at this module's boundary, never inside the Phase 2 helper. - loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated JSON, wrong types, missing fields, out-of-batch dependency references, dependency cycles, a worktree path absent from disk). - resumeBatch: skips complete items, never auto-retries failed items, propagates/reverses blocked status along the DAG to a fixed point, and detects a STATE.md row that already exists for a non-complete item (the "STATE written, BATCH.json not yet updated" crash window) — completing it without re-appending. Idempotent across repeated calls. - completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion — appendQuickTaskRow (unmodified) is called at most once per quick id, gated by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks Completed" table, since appendQuickTaskRow itself carries no idempotency. BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling of .planning/quick/ — never inside it, so scanQuickTasks never misreads a batch manifest as a broken quick task. * docs(#3675): register the new quick-batch module New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row plus the regenerated docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary entry matching the convention set by the sibling File Overlap Partitioner Module (#3674) entry it sits beside. NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically sorted single-entry insertion into families.cli_modules, matching the existing file's structure) rather than via `node scripts/gen-inventory-manifest.cjs --write` — this session's MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script path from Bash, with no available Memtrace tool to route through instead. The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs --check` to confirm this hand-edit is byte-identical to the generator's own output before merging. * fix(#3675): resolve lint findings in quick-batch primitives and tests Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color array, two unnecessary `as string[]` casts TS 5.5's inferred type predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch` import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a CONTEXT.md glossary illustration that looked like a real file reference. * feat(#3675): close acceptance-criteria gaps found in review Standards- and spec-axis review (plus a self-caught race) surfaced real gaps against issue #3675's own acceptance criteria and this repo's test conventions: - BATCH.json was missing options, base_revision, per-item wave, and per-item commit — the issue's AC explicitly lists all four as things the manifest must track. Added them: createBatch persists caller-supplied batchOptions/baseRevision verbatim and assigns each item its computed wave index; completeQuickItem now persists the commit onto the item, not just the STATE.md row. All four are backward-tolerant on load (an older/hand-built manifest without them still validates). - resumeBatch had no "incompatible base divergence" check at all, despite the AC and the ADR's own "Base divergence" section requiring one. Added an opt-in currentBaseRevision comparison that fails closed with a recoverable diagnostic on mismatch, and touches nothing on refusal. - resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the only durable write path in this module that wasn't lock-protected, a real lost-update race against a concurrent completeQuickItem or another resume. Now runs inside the same lock createBatch/ completeQuickItem use. - loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no size cap (security review, Low/informational); switched to the existing safeJsonParse (1MB cap) for defense-in-depth. - Parser (parseTaskList) had only example-based tests; CLAUDE.md requires a fast-check property test for parsers. Added one plus a companion reject-property for <2 items. - The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for direct limit-1/limit/limit+1 testing without needing 46k fixture dirs. - Issue AC explicitly asks for prompt-injection-payload test coverage, distinct from the existing shell-metacharacter test; added one. - Test row 9 (FIFO skip) silently returned instead of calling t.skip(), so an unsupported platform would report a pass rather than a documented skip; fixed to bind the test-context param and skip properly. - Extracted toWaveInput to remove a 2-site production duplication of the QuickBatchItem -> computeWaves reshape (Standards-axis smell). - Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/ is user-facing even though the compiled .cjs is gitignored). * fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason gsd-test caught this: switching loadBatch to safeJsonParse changed the parse- failure message shape ("... parse error — ...") without preserving the "not valid JSON" substring row 27's own test asserts on. Re-wrap safeJsonParse's error into the original diagnostic phrasing regardless of which of its three failure modes fired. * docs(#3675): backfill changeset pr number to 4190 * fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one CI caught this on windows-latest: row 9's platform-skip only caught mkfifo throwing (command not found). On this runner mkfifo resolves to something that exits 0 without creating a file (NTFS has no FIFO concept), so execution fell through to parseTaskListFromFile against a path that doesn't exist, producing an ENOENT stat error instead of the expected "not a regular file" rejection. Check the artifact actually exists before trusting a zero exit code, and skip with a documented reason either way. --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
77dcdda534 |
enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4) * fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4) * test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift src/core-utils.cts's new #4014 import block shifted every subsequent line by 6, moving generateSlugInternal's real closing brace from line 193 to 199. tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that line number to plant a synthetic violation immediately after the function's real body; the guard script itself locates the boundary dynamically via brace-matching and needed no change. * docs(#4014): document the unreadable-directory scope signal and add changeset * docs(#4014): backfill changeset PR number to #4163 * test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff --------- Co-authored-by: sim <sim@local> |
||
|
|
b0572c0108 |
feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage golden script) as a regression safety net ahead of extracting partitionStages into a standalone module. Also adds the new module's unit and property tests (test matrix rows 1-11) against its expected public API, which does not exist yet and is added in the next commit. * feat(#3674): extract file-overlap partitioner into a shared, generic module Moves partitionStages' greedy first-fit file-overlap algorithm into a new, dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap), generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own Plan[] shape onto the generic input and back — behavior-preserving, no dependency ordering, no path normalization, no filesystem access moved or added. Enables a future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive without pulling in orchestration internals. * docs(#3674): register the file-overlap-partitioner module bookkeeping New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row (regenerated via gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry matching the convention set by similarly-scoped leaf modules (text-lines.cts, plan-dependency-graph.cts, spec-section.cts). * fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct - docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write, produced no diff (manifest is keyed by content, not row order) - tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is: tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier. * fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction The `no two plans in the same stage share a modified file` property reconstructs which physical item produced each output id via `remaining.findIndex(r => r.id === id)`. Under duplicate ids (an explicitly-supported input shape for `partitionByFileOverlap`) that reconstruction can pick the wrong physical occurrence, producing a false-positive overlap failure (observed counterexample: p0(f1), p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but misread by id-order as [[p0,p208#1],...], which do overlap). Properties (a) determinism and (b) totality already exercise duplicate ids correctly and are left unchanged; only this property's generated items are now constrained to unique ids via `fc.uniqueArray`, where the reconstruction is unambiguous. --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
7960374d15 |
fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git Behavioral coverage: extract the trailers:key token from ship.md and prove a real git commit carrying that trailer reads back via %(trailers:key=token,valueonly). gate_status contains an underscore, which git's trailer machinery cannot tokenize, so the audit read was structurally empty. * fix(#3962): rename the TDD-Audit trailer token to gate-status Underscore is not a valid git trailer token character, so %(trailers:key=gate_status,...) could never match. Renames the token at the read, the documented aggregate write, and the section's prose/table header. Self-suppression semantics (#2431) unchanged. * test(#3962): compare outcome against the seam's literal, fix doc token Review follow-ups: the round-trip test compared outcome to 'EXITED' but the process seam emits 'exited'; and the trailer rename is propagated to docs/ship-pr-body-sections.md's read/write examples. * chore(#3962): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
91d5fdff6f |
chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse advisory output alongside the existing additionalContext prose: - gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName - gsd-context-monitor.js: severity ('warning'|'critical') - gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523 pattern) - gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its findings array already existed from #3523) - gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg, distinct from the existing force-add block leg's code additionalContext stays byte-identical in every hook (verified per-hook against the pristine HEAD version across a spread of payload shapes). Migrates all 20 assertion sites named in the issue off additionalContext.includes(...)/assert.match(...) substring-matching onto the new typed fields, per CONTRIBUTING.md's prohibition on raw text matching on test outputs. Closes #3546 * test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's own doc comment says is sized for plumbing reads (rev-parse/branch/log), not write-heavy add/commit spawns. That file already documents the exact same defect class from a prior incident (PR #3323) and exports GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites - commitN just wasn't using it. Observed failure: `git commit -m filler 9` timed out under normal bench load, unrelated to any of this PR's own diff (hooks/*.js + 5 other test files). Not a flake: root-caused to the timeout bound being sized for the wrong call class, per this repo's no-flakes rule. * chore(#3546): backfill changeset PR number (#4167) --------- Co-authored-by: sim <sim@local> |
||
|
|
ff0361071d |
feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane cursor-agent exposes --model (204 selectable models) but the cursor lane declared modelArg: null / modelConfigKey: null, so review.models.cursor was rejected as an unknown config key and the #1517 reviewer-instances escape hatch silently discarded a configured model at modelExpansion. Wire the lane the same way codex already is: inject {{model}} into args right after -p, set modelArg to --model, and declare modelConfigKey as review.models.cursor plus its config schema entry. An unconfigured lane still invokes byte-identically to today. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): add changeset for review.models.cursor Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3653): update co-change surfaces that assumed cursor has no model key gsd-test surfaced three surfaces still hardcoding "cursor declares no modelConfigKey", broken by wiring review.models.cursor: - tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797 invariant test listed cursor among lanes that must own no model key. - tests/settings-integrations.test.cjs: the #3651 keyless-lane test listed cursor as keyless, including a live config-set assertion that now correctly succeeds instead of failing (swapped to qwen). - gsd-core/workflows/settings-integrations.md: the settable-keys enumeration and two prose call-outs still named cursor as keyless. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): acknowledge deliberate growth of settings-integrations.md settings-integrations.md grew 4 bytes because it now enumerates review.models.cursor as a settable key alongside the other reviewer lanes, matching the modelConfigKey wired for cursor in this PR. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer The previous commit's trailer was separated from Co-Authored-By by a blank line, splitting it into an earlier, non-trailer paragraph — git's trailer parser only recognizes the last contiguous block. Restating it here immediately adjacent to Co-Authored-By so both parse as trailers. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): backfill changeset pr number pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
4dfc46bbe7 |
enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate * feat(#3348): add context-drift pre-check gate for plan-phase Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md effective last-changed time (git commit time, falling back to mtime for uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently reuses an upstream artifact that predates a decision added to CONTEXT.md after that artifact was derived from it. Deterministic, no model call. New `gsd_run verify context-drift <phase>` command, sibling to the existing verify.codebase-drift/verify.schema-drift gates in the drift capability. Warn-only by default (workflow.context_drift_precheck), with an opt-in workflow.context_drift_action: block escape hatch. Wired at plan:pre in plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions. * fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution * fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe The #3859 empty-diff guard decides whether a submodule bump would land using `--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules` config. The real `git commit -- <paths>` that follows was never given the same override, so under a bare `diff.ignoreSubmodules=all` repo config the two calculations disagree: driven on git 2.39.5 (Debian bookworm, the linux-node24 test-matrix image), the guard correctly stands aside but the scoped commit itself then silently fails (exit 1, no error text) for a gitlink bump it had just confirmed would be recorded, surfacing as commit_failed instead of committed:true. Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the probe and the commit it protects can never diverge. Harmless when no submodule path is involved (driven: identical outcome on an ordinary scoped file, with and without the flag). * fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal * fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a current-milestone enumeration). This PR's refactor pass lifted that block into a shared helper, resolvePhaseDirByToken, also used by the new cmdVerifyContextDrift — the guard tracks exemptions by enclosing function name, so the readdirSync now lives in an unexempted function and started firing. Move the exemption to resolvePhaseDirByToken (same written reason, now covering both callers) instead of migrating to listAllPhaseDirs, which would introduce two real behavior deltas here: it catches readdirSync failures internally (old code let them throw) and sorts results by phase number before matchPhaseDirs picks matches[0] (old code used raw, OS-dependent readdirSync order). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): satisfy lint:ci — slash form, capability registry regen - docs/features/context-drift-gate.md used the deprecated /gsd: colon form; docs are never passed through the install-time slash-form converters, so lint-docs-command-form requires the hyphen form. Regenerated docs/FEATURES.md from the corrected fragment. - Regenerated gsd-core/bin/lib/capability-registry.cjs after editing capabilities/drift/capability.json (lint:generated-sync). * fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit` subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for every scoped commit call, breaking 17 position-based assertions in the commit-files pathspec regression suite that read `a[0] === 'commit'` to find the commit invocation among recorded git calls. `execGit` already accepts an `env` option merged onto `process.env` before spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0` as a per-invocation config override functionally identical to `-c key=val`, expressed via env instead of argv. Passing that env alongside the existing commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged) fixes the real commit's effective diff.ignoreSubmodules to match the empty-diff guard's probe without moving anything in argv position 0. Scoped to exactly the canScope branch, matching the probe's own preconditions and leaving no behavior change for commits the probe never evaluated. No test file changes needed — the 17 previously-failing assertions test argv[0] against the array passed into execGit, which never changes. * fix(#3348): register verify-context-drift in the check subcommand router The drift capability's new plan:pre gate declares check.query "verify.context-drift", which normalizes to `check verify-context-drift`, but no such subcommand was routed — phase6-capstone-conformance's uniform-block-field test failed with "Unknown check subcommand" for every declared gate query. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot of the drift capability's config keys. #3348 legitimately adds two new keys (workflow.context_drift_precheck, workflow.context_drift_action) for its own plan:pre context-drift gate — extend the expected set (and clarify the assertion message) without weakening the test's exactness. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction #3348 (an earlier commit on this branch, e4b80ad81) extracted cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block into the shared resolvePhaseDirByToken helper (also used by the new cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's function-scoped exemption from cmdVerifySchemaDrift to resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer contains a line the guard's detectors match, so it needs no exemption. tests/phase-locator.test.cjs's E2 test still pinned the exemption to the old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to match the same "migrated call site's exemption must move, not duplicate" pattern the test already applies to cmdRoadmapAnalyze and cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the still-exempt list and add symmetric assertions that it no longer carries the exemption while resolvePhaseDirByToken now does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs 'always exits 0 (query command contract)' included the no-phase-arg case, which contradicts the file's own earlier 'errors with usage message on missing phase arg' test (that case legitimately exits 1 via the Usage error) — drop it from the always-exits-0 cases. 'degrades to mtime comparison outside a git repo' and '...in a repo with no commits' relied on real wall-clock ordering between two back-to-back writeFileSync calls to prove CONTEXT.md is newer than RESEARCH.md; on a fast filesystem both can land in the same mtime tick, producing a tie that computeContextDrift's strict `<` correctly treats as not-stale, so stale_artifacts comes back empty. Make both tests deterministic via explicit fs.utimesSync instead of relying on timing (CONTRIBUTING.md: never assert elapsed wall-clock time). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture The "all plan:pre when-keys false" fixture explicitly disables every known workflow.* plan:pre toggle, but didn't yet know about the new workflow.context_drift_precheck key (defaults to true), so the new drift context-drift gate stayed active and broke the empty-activeHooks assertion. Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3348): backfill changeset PR number (pr:0 -> 4147) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c0fd2e3f4c |
feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route Extends tests/host-integration.test.cjs with negotiateHostCapabilities maxConcurrency negotiation coverage (test matrix rows 1-15, including a fast-check property test) and a new #3673 dispatch-capacity CLI route describe block spawning the real gsd-tools.cjs (rows 16-25). Extends tests/host-integration-validator-parity.test.cjs with an all-19-descriptor maxConcurrency presence/validity sweep (row 26) and adds a hostile-input validator test to host-integration.test.cjs (row 27). The dispatch.maxConcurrency field does not exist yet, so these tests fail. * feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration Interface (ADR-1239 Phase 1), following the existing dispatch.isolation sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES floors, and a negotiateHostCapabilities branch that passes through a positive safe integer and fails closed to 1 otherwise (no engine-side reduction, per the design doc's explicit rejection of a min(host,engine) rule). capability-validator.cjs gains parity validation for the new field (optional, positive safe integer or the "undocumented" sentinel — mirroring isolation's "added after existing descriptors" treatment). gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling of `dispatch-isolation` with no side effects: live env (GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence. All 19 capabilities/*/capability.json descriptors now declare dispatch.maxConcurrency: claude carries the one cited value (20, per code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented" (not yet researched for this axis). * docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix Updates docs/reference/host-integration-interface.md's dispatch struct entry (also backfilling the previously-undocumented isolation/backgroundDispatch sub-fields found stale in the same table) and adds fail-closed/live-transport precedence prose for the new maxConcurrency field. Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in docs/reference/host-integration-capability-matrix.md — required for tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check, which asserts every declared dispatch sub-axis is documented there. Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist and example descriptor block to mention maxConcurrency (and, likewise backfilling a stale gap, isolation/backgroundDispatch). * fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check Exports isPositiveSafeInteger from src/host-integration.cts as the single source of truth for the dispatch.maxConcurrency positive-safe-integer contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity now both call it instead of independently reimplementing the same predicate. Also removes the __proto__/constructor/prototype reserved-name branch from capability-validator.cjs's maxConcurrency check — copy-pasted from the string-enum fields above it, but unreachable for a numeric field (the generic positive-safe-integer branch already rejects any string) and absent from maxDepth, the field the code's own comment claims to mirror. --------- Co-authored-by: sim <sim@local> |
||
|
|
b848b23861 |
feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md - Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's existing render + generic contribution dispatch pattern - Inject planner-targeted contribution fragments into the planner prompt, after AGENT_SKILLS_PLANNER, matching D-08 ordering - Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions - Record quick.md's byte-growth rationale in this commit trailer for every gsd-core-verbatim runtime Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does. * test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call < agent-skills placeholder < injection block < spawn ordering) assertions, all extracted via a brace-bounded slice anchored on the literal injection instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER} and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain brace pairs). Add a capability-registry.test.cjs describe block proving the registry-wide D-07 exclusion is meaningful: at least one plan:pre contribution exists, every plan:pre contribution has a non-empty into/fragment.inline, and the registry as a whole carries at least one non-planner-into contribution. Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's real content, and buildContract() still yields exactly 5 entries. Verified red-without-Task-1 by temporarily reverting quick.md to its pre-f30de9cc content and re-running these three suites (D-08 failed as expected), then restored via git checkout and re-confirmed green. * docs(#3778): note quick planning also renders plan:pre in the tutorial The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the plan:pre hook set. Since quick.md now dispatches the same hook set (f30de9cc), the sentence understated the capability's real reach. * feat(#3778): add changeset fragment * chore(#3778): reference the upstream issue in the changeset fragment The fragment was the only one of 81 in .changeset/ without a trailing (#NNNN) reference or a bold lead-in. serializeChangelog auto-appends only the pr: field, so the rendered CHANGELOG entry carried no link back to issue #3778. * test(#3778): scope the D-07 registry assertion to what it actually proves The registry-wide non-planner check was named "D-07 exclusion is meaningful", which overclaims: it proves only that `into` takes non-planner values somewhere in the registry, not that anything is excluded at plan:pre. Every plan:pre contribution is currently into: "planner", so the filter is a forward-looking safeguard there. Narrowing the assertion to plan:pre (as review suggested) would fail today. Asserting plan:pre is all-planner would be brittle — it would break the day a legitimate non-planner plan:pre contribution lands, which is exactly when the safeguard starts doing work. So the assertion is unchanged and only the name and comment are corrected. * chore(#3778): point the changeset fragment at the upstream PR The fragment carried pr: 3, the fork staging PR. changeset lint derives the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that would read as pr-field drift. Point it at open-gsd/gsd-core#3934. * test(#3778): require contributions in Quick revision prompts * test(loop-host): require Quick auxiliary registration * fix(#3778): preserve contributions in Quick plan revisions * fix(#3778): validate Quick as a planner contribution host * fix(#3778): tighten Quick contribution contract * test(#3778): drop unnecessary source-contract exemption * fix(#3778): require Quick planner target coverage * docs(#3778): describe targeted auxiliary coverage --------- Co-authored-by: davdittrich <davdittrich@gmail.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7ca5479f97 |
docs(#3672): amend ADR-1239 for quick-batch concurrency and single-writer contract (#4146)
Locks the architecture gate the maintainer required before any /gsd:quick-batch implementation PR (epic #3344, condition #2): canonical terminology, the foreground-coordinator single-writer invariant over BATCH.json/STATE.md/roadmap metadata, the dispatch.maxConcurrency sub-field (fail-closed to 1), the GSD_DISPATCH_MAX_CONCURRENCY live-capacity transport and its precedence over the descriptor value, the gsd-tools query dispatch-capacity contract, the dispatch.isolation interaction, and backpressure/recovery/deterministic-merge rules. Docs-only; closes its own Phase 0 sub-issue, not the epic. Co-authored-by: sim <sim@local> |
||
|
|
bdfc62889b |
fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape advancePlanCore derived the value FORMAT from the field NAME, so it handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the compound `Plan: N of M`, but not the hybrid of the two: the legacy field name carrying a compound value with no Total Plans sibling. `legacyTotal` is null so the legacy branch fell through, and the compound branch reads the `Plan` field through a `^Plan:`-anchored pattern that never matches `Current Plan:`. Both produced NaN against a file whose plan numbers are plainly readable. The shape is not exotic. An agent wrote it unprompted into a project's STATE.md, believing it was the parseable form, and every subsequent run in that project inherited the failure and worked around it by hand. Track the field name and the value shape separately (`planSourceField`, `planRawValue`) so write-back targets whichever field the value came from. The legacy pair still takes precedence when both fields exist, so a stray "of N" inside Current Plan cannot override an explicit Total Plans — covered by a new test. Also replace the caller's catch-all error. It reported "Cannot parse Current Plan or Total Plans" for ANY transition failure, and named no accepted shape, so a reader learned neither what failed nor what to write. It now distinguishes "no result" from "unreadable plan position" and lists all three shapes. The existing test asserted the literal "cannot parse"; it now asserts the message names the shapes, which is the property that makes it actionable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): keep zero-padding when advancing a compound plan value The compound write-back rewrote only the leading half of "N of M", so a padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic on its own, but a plan line that looks wrong is one the next writer tidies by hand, and hand-tidying this particular line is what produced the hybrid shape the previous commit had to teach the parser to read. Pad the incremented number to the width it was written with. padStart never truncates, so a value that outgrows its padding widens correctly: 09 of 12 advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances to 3 of 6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): pass a literal field name to the compound write-back The previous commit passed `planSourceField` — a variable — as the field-name argument to `stateReplaceField`, which trips the state-write-path drift guard's `unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care: a Title-Case literal cannot collide with a lowercase or snake_case frontmatter key, so it is safe whatever the content argument is, while a variable could hold anything and therefore requires its content to be demonstrably frontmatter-stripped first. The content argument here IS stripped — `body` is `stripFrontmatter(content)` — but the guard does a narrow backward scan rather than dataflow tracking, by design, and the nearest preceding assignment to `body` is another `stateReplaceField` result. Rather than baseline a bypass or ask a future reader to re-derive that the invariant holds, dispatch on the discriminator and pass the literal. Guard goes from 1 finding to 0; its own 32 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(3784): add changeset fragment for #3785 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring Triage published six acceptance criteria; two were only half-covered. AC1 asks that the hybrid write back to the SAME field with padding preserved. The existing hybrid test used an unpadded value and asserted only `result.data`, so it proved the parse but never the write. Now asserts the written content is `05 of 06` on the original field, and that no separate `Plan:` field appears as a side effect. AC6 asks that the shared field reader not be loosened. Reading the hybrid is the transition's job; `stateExtractField('Plan')` is line-anchored and has 13+ callers, so teaching it to match a name merely ENDING in "Plan" would be the wrong fix and would silently change what those callers read. This holds by construction here — the reader is untouched — but nothing locked it in. The new test fails if anyone later reaches for that shortcut. Also drops the changeset fragment written against the auto-closed PR number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(#3784): add changeset fragment for #3791 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(#3784): write the advanced plan back to the field it was read from Review findings 2-6 on #3791 were one defect seen from several angles: the read path learned the hybrid `Current Plan: N of M` shape, the write path did not follow it. - `bumpLeadingNumber` now owns the increment for all three parse branches. Only the leading digits belong to this transition; the padding width and everything after it (` of M`, and the `\r` of a CRLF file) are the author's text and are preserved. The legacy branch wrote `String(newPlan)`, which turned `2 of 99` into `3` and `04` into `5`. - `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm only ever looked for `Plan:`, so on a hybrid file the `## Current Position` section was never reached; combined with the body-level write being single-shot and bold-preferring, a file carrying the field at both sites advanced the header and left the section a plan behind. The parameter defaults to `Plan`, so the two callers that pass no plan are unchanged. - Tests: both-sites-advance (fails without the section arm), legacy write-back content assertions (the previous test read only `data` and so could not see the lossy write), hybrid boundary at limit-1 and limit+1, a CRLF fixture, and an fc property pinning the padding-width contract. Two characterization tests pinned `**Current Plan:** 02` advancing to `3`. That dropped padding is the defect #3784 reports, so the expectation is corrected to `03` rather than the fix being narrowed around it. * fix(#3784): drop the unreachable advance-plan error branch, sync the doc Findings 1 and 8 on #3791. The `!resultData` arm could not fire: the transform callback assigns `resultData` unconditionally, only runs once STATE.md is known to exist (the missing-file case returns "STATE.md not found" upstream), and every `advancePlanCore` return path sets `data`. It was a speculative second failure mode with a message no caller could receive, and the comment beside it claimed to distinguish two things that were never two. `!resultData` stays in the condition as a type guard, which is all it ever was. `docs/json-errors.md:142` quoted the old error literal verbatim and was the sole occurrence in the tree; it now quotes the emitted one. * chore(#3784): describe the write-back fix in the changeset * fix(#3784): anchor the plan grammar and widen the schema row to match Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit. B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to `['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs --check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not surface this row, so there is nothing stale to regenerate. B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as `4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE `Status: Phase complete — ready for verification` into the user's file. Both shapes are now anchored at the start and every number comes from a capture group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER` rather than letting `data` and the persisted string disagree. Nothing on this path calls `parseInt` on a raw field value any more. The grammar keeps a trailing remainder after the total, because `Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from requiring `of <total>` to follow the leading number immediately, not from forbidding a suffix. M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there were no leading digits, so `+2` reported `advanced: true` while writing the file untouched. M2 — both section arms use replacer functions. File-derived text was being spliced into a `String.replace` replacement string, where `$&` / `` $` `` / `$'` expand: a value of `04 of 06 $&` spliced part of the document into itself. `stateReplaceField` already used a function; these now agree with it. M3 — the section arm targets the name the SECTION carries, and the body write now writes both spellings, each with its own rendering. Keying off the header's name left the other name stale in both directions: a legacy header beside a `Current Plan:` section line, and a `**Plan:**` header beside one. * fix(#3784): derive the shape error from the schema, widen the test coverage Review round 3 on #3791: B3, m1, m2, m5 and the two test nits. B3 — the accepted-shape set had two owners: the parser branches and an English list hand-written beside them in `state.cts`. Nothing coupled them, so adding a branch left the message stale and removing one left it advertising a shape that errors, with no test able to see either. The message is now built from `STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the schema instead of restating the list. `Plan: N of M` is still spelled out explicitly because no schema row owns the body-only `Plan` field — `buildStateFrontmatter` never reads it into frontmatter, so it has no key to hang a row on. m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the branch under review. It now drives both compound spellings and ranges past 99 so the width transition is covered by the property rather than one example. A second property covers the legacy pair's own preservation contract. Both were mutation-checked: dropping the padStart turns 9 tests red. m2 — degenerate boundary fixtures around the threshold (`0 of 0` is phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored grammar must refuse, including Arabic-Indic digits. m5 — `docs/json-errors.md` described rather than quoted the message, since it is now schema-derived and a verbatim quote would be a third owner. Nits — the CRLF assertion could not see a `\n` at index 0; the `!/^Plan:/m` presence proxy is now an identity assertion on the whole `## Current Position` body. * fix(#3784): give the section plan write its own flag, and stop narrowing what parses Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2. B1 — the section fallback was guarded by `!mutated`, and `mutated` is FUNCTION-wide, already set by the phase/status/lastActivity arms that `advancePlanCore` always populates. A section spelling the field bold or as a pipe-table row therefore skipped its fallback because an UNRELATED field had been refreshed, and stayed a plan behind the header — the split-brain document this arm exists to prevent. The arm now tracks its own `planWritten`. Worth recording: the reviewer's fixture does not reproduce. The body-level status write lands on the section's own `Status:` when the document has no header `Status:`, so `mutated` is still false by the time the plan arm runs and the fallback fires. The discriminating shape needs a header `Status:` to absorb that write AND a bold section plan line. The mechanism was right; the example was not, and the regression test uses the shape that actually fails. B2 — `fallbackName` chose one name by ternary. In the legacy shape both values are populated, so it always chose `Current Plan` and a `**Plan:**` section line — which base did write — got nothing. Each name is now attempted independently with its own fallback. B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`, `Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond this transition's read/write out of scope, so that narrowing was not licensed. Both grammars now carry the same trailing tolerance. The prose defect stays closed by the START anchor, not by forbidding suffixes. Major 1 — the whole-body `Plan` write is scoped to documents that declare a `Plan` field, instead of firing unconditionally where `stateReplaceField`'s first match could be prose outside `## Current Position`. Major 2 — the error message names both `Plan` spellings the parser accepts; it previously omitted the sibling-paired form, which is the same message-disagrees-with-parser drift the derivation exists to close. B4 — the changeset claimed a guarantee B1 broke; it now describes what ships. Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the CR; the trailing group is belt-and-braces, not the primary defence). All three blocker regression tests verified red against the pre-fix source. * test(#3784): pin the hybrid shape against #3807's ambiguity refusal #4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this branch's last run, on the same function. The guard sits above the parse, so a refused document is never parsed and the shape #3784 adds cannot reach the mutation — but that is a property of source ordering, so assert it as behaviour instead. Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled, the ambiguous hybrid document advances its FIRST entry's `Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect, reached through #3784's shape. Both tests go red; both go green with the guard restored. The control pins the other direction: an unambiguous hybrid section still advances, and its zero-padding still survives. * fix(#3784): advance every spelling from its own text, refuse when they disagree Round 6 review. B1 and M1 are one defect, so they are one fix. `advancePlanCore` picked one field to parse from, computed `newPlan`, then wrote BOTH spellings from that field's numbers. Two symptoms: B1 With `Plan` as the parse source, `Current Plan` was re-stamped with the number just derived from `Plan`. `Current Plan: 7` beside `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing derived for that field, no error, no diagnostic. M1 With the legacy pair winning, the `Plan:` line was re-rendered from a bare `${newPlan} of ${totalPlans}` built out of the sibling field. `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`. The changeset's claim that padding and everything after it survive was true only for whichever field happened to be the parse source. Now: every spelling is advanced from its own raw text via `bumpLeadingNumber`, so each keeps its own padding, its own total and its own trailing annotation. Differing TOTALS are preserved, not reconciled — `Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`. Differing CURRENT numbers are refused, with `reason: "ambiguous_plan_position"` and both candidates named. Same posture as #3807's multi-`Phase:` guard one field over: name the conflict, let the caller resolve it, never pick. The guard sits immediately after the parse, BEFORE the phase-complete branch — guarding only the normal advance would let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase complete" into a document whose two spellings never agreed. A field present but unreadable (`Plan: TBD`) is left exactly as authored. Refusing the whole document because an unrelated line cannot be read would be a narrowing #3784 does not license; writing a derived number over it is the fabrication B1 was filed for. The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone. It existed only so the write path could ask which field the value came from, and the write path no longer asks. M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A revision of this PR added it; base refused it. It cannot be given the schema-row + forcing-test coupling the other shapes have, because `Plan` is body-only and `buildStateFrontmatter` never reads it into frontmatter, so there is no `current_*` key to hang a row on. Parser, the spelling in `advancePlanShapeError`, and the lockstep test move together — the invariant is the lockstep, not the length of the list. N1. The whitespace narrowing (`5phases` no longer parses where `parseInt` read 5) is documented in the changeset beside the other deliberate narrowings, rather than loosened. Loosening restores the half-parse this change exists to remove. Tests: eight new cases plus a property that crosses the two spellings with agreeing and disagreeing numbers — the review noted the existing properties never did. Fail-first proven: restoring the old write path reddens seven of the eight, both new property arms, and two pre-existing padding tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * fix(#3784): report Current Plan as updated only when it was written The write became conditional in the previous commit — a `Current Plan:` that is present but unreadable is left as authored — but the `updated` push stayed unconditional, so `transitionCore` reported a field it had not touched. `reconcileReportedFields` would have caught it against the persisted bytes at the `state.cts` caller, but `transitionCore`'s own `updated` is consumed directly (milestone-lock, the transition tests) and has to be true on its own. Covers the mirror of the unreadable-spelling case: `Current Plan: TBD` beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the legacy field is the one that cannot advance. Fail-first proven. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * test(#3784): account for the new refusal in the output({error}) census `tests/io.test.cjs`' A3 census asserts the exact population of `output({error})` call sites in `src/`, per module. The `ambiguous_plan_position` refusal added a 27th to `state.cts`, so the census went red at 26/65. Updated the way #3807 updated it when it added the ambiguous-POSITION error one line above: bump the count and name the addition inline, so the next person reads why the number is what it is. The alarm did its job — it is the only gate that noticed a new user-visible error path had been introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7c116b1c17 |
fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744)
* fix(#3697): warn when the Requirements line under-selects REQ-IDs
`cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting
on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is
correct for the canonical comma list the template ships, and silently wrong
for every other form:
`RANGE-01 … RANGE-05` -> the two ENDPOINTS only; the interior IDs are
never considered, yet `requirements_updated`
reports true with zero warnings
`RANGE-01…05` -> ZERO IDs; the whole line is inert
The silence is structural: the only cross-check, `ghostReqIds`, is itself
`citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it
by construction — and to `traceabilityWriteMisses` and `requirements_updated`
with it.
Warn on both paths. This does not add range support: the selected set is
unchanged, so no existing ledger write changes. The trigger is ID-SHAPED
EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a
range operator joining two IDs — with parenthetical citations and HTML
comments stripped before the scan, so the #2334/#2339 over-warning on
`None`, on the shipped `<!-- brackets optional -->` template comment, and on
annotated lines cannot return.
Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs
(10 cases: 4 defect, 2 canonical controls, 4 negative-space controls).
Fixes #3697
* fix(#3697): rework under-selection detection onto tokens, not a free-text scan
Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb.
That review refuted 5 of 9 claims; three were false-positive classes in exactly
the category #2334/#2339 had to REMOVE:
`RANGE-01, RANGE-02 - 3 points` the bare-hyphen alternative read
`RANGE-02 - 3` as a range
`REQ-01, REQ-02 — locked per ADR-7.` the trailing period kept `ADR-7.` out
of the anchored filter, so the
unanchored substring scan reported it
as unparsed
`REQ-01, REQ-02 (see (ADR-7), then ADR-8)`
nested parens left `ADR-8)` behind
Replaces the free-text substring scan + loose range regex with three narrow,
token-based rules (R1 range-shaped token, R2 pure range operator flanked by two
selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's
CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks
nothing — it now says "selected".
Side effect: the two false NEGATIVES the same review found are now covered —
`RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`.
NOT YET DONE (see the handoff prompt): regression tests for the four false
positives, the two new true positives, and the #3697-4 tightening the review's
CLAIM 8 asked for (it currently filters on the warning's phrasing rather than
asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests
green, and a 20-case standalone harness covering every case above.
* test(#3697): pin the v2 token-detector boundary end-to-end
Six new cases + two hardenings for the review findings against v1:
- #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) —
the operator set's `to|thru|through` arm was previously untested.
- #3697-5 (new): a tight range hidden inside balanced parentheses
(`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token,
and ticks exactly RANGE-01 — the paren shave must not hide it.
- #3697-4 gains the four false-positive classes a free-text detector
produced: numeric estimate (`- 3 points`), date annotation, em-dash
citation with trailing period (`— locked per ADR-7.`), and nested
parenthetical citations.
- #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty,
not that one phrase is absent — a re-worded over-warning cannot pass.
Negative control: against the merge-base with its lib rebuilt, all 6
defect tests fail and all 10 controls pass.
* fix(#3697): close round-2 review findings — annotation false positives
Round 2 of the adversarial review (against 822a72a04) refuted five
claims; this closes the false-positive class and the cheap misses:
- R1's bare-hyphen arm now demands a full ID on BOTH sides
(`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation
(`FY-2026-08`) and a sub-numbered ID, and warning on those is the
expensive class. Tight hyphen shorthand with a live selection is the
disclosed false negative; at zero selection R3 still catches it.
- R2 requires the endpoint pair to imply an INTERIOR (same prefix,
gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop
nothing, so an annotation hyphen between adjacent IDs stays silent.
- R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty
line citing its rationale, not unparsed residue.
- Token shave: quotes/backticks now shaved from alphanumeric tokens
(`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a
bracket-only shave so `(..)` surfaces its operator.
- 256-char token cap bounds the quadratic unanchored substring test.
- Warning text mentions range expansion only when a range rule fired.
Tests: 6 new cases (22 total). Negative control against the merge-base:
8 defect tests fail, 14 controls pass.
* fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers
Round 3 of the adversarial review (against 2eb92dd0e) refuted four
claims; this closes them:
- Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the
tokenizer before R1's `\s*` can see them and under-selected silently.
A glued-fragment rule warns when an operator is glued to a full ID
with an ID-shaped neighbour on the open side and the endpoint pair
implies an interior.
- Cross-prefix pairs around a separator no longer read as ranges:
`REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and
real ranges are same-prefix by nature. `impliesInterior` now returns
false on prefix mismatch and computes the gap with BigInt (parseInt
lost precision past 2^53).
- The token shave now removes markdown emphasis markers and curly
quotes, so `**None** (per ADR-7)` reaches the placeholder gate and
`**RANGE-02..RANGE-05**` reaches R1.
- The unanchored-substring cap rises to 2048 (a markdown-link range
with a long URL cleared 256); the anchored range regexes scan
linearly and drop their cap.
Tests: 6 new cases (28 total; 377/377 file-wide). Negative control
against the merge-base: 11 defect tests fail, 17 controls pass.
* fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule
`TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment
rule read it as `to` + `REQ-05`, warning on the canonical two-ID list
`REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only
(`..`+, ellipsis, dashes): a word operator glued to an ID is an ID,
not a range spelling.
Tests: word-operator-prefixed ID control (misparse channel silent; the
fixture's ghost-ID warning legitimately fires, so the whole-channel
assertion stays with the registered controls) and an underscore-wrapped
tight-range defect case. 30 targeted cases; 379/379 file-wide; negative
control: 12 defect tests fail on the merge-base, 18 controls pass.
* fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording
- The glued-fragment TRAILING arm takes the word operators back: an ID
must end in digits, so `REQ-01through` can never be an ID — the
round-4 TOREQ collision was leading-arm-only, and symbol-only on both
arms lost the `REQ-01through REQ-05` typo class.
- A trailing run of 2+ dots survives the punctuation shave: `REQ-01..`
is a glued range operator, not sentence punctuation, and the shave
was silently eating the `REQ-01.. REQ-05` form.
- The warning now says the line "could not be parsed as" a
comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list
— the selector just cannot parse decorated tokens — and a warning
that misstates the input teaches readers to distrust it.
Tests: two new trailing-glue defect cases (32 targeted; 381/381
file-wide). Negative control: 14 defect tests fail on the merge-base,
18 controls pass.
* test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset
CONTRIBUTING bans try/finally inside test bodies (it masks failures);
the approved shape is `t.after(() => cleanup(tmpDir))`. All seven
converted tests are this PR's own additions; the file's pre-existing
instances are untouched.
* chore(#3697): add changeset fragment for the Requirements-line under-selection warning
changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a
user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed
fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per
the house format, with the (#3697) backlink.
* refactor(#3697): extract the Requirements-line detector to a testable surface
Round-3 review Blocker 1 requires a fast-check property test over this
detector (`RULESET.TESTS.property-based-testing`: modules implementing
parsing contracts must include at least one), and Blocker 2 requires
limit-1/limit/limit+1 fixtures on its 2048-char token cap
(`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while
the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test
reaches it by spawning the CLI, and a property test cannot pay a subprocess
per generated case.
So the selector and the three detection rules move to module scope as
`analyzeRequirementsLine` (pure, exported) plus
`formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This
commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and
the pre-round suite passes against it unmodified (403/403).
Two things the move makes explicit rather than incidental. The selector and
the detector tokenize the SAME line DIFFERENTLY — the selector strips only
`[` and `]`, the detector also shaves quotes, emphasis and trailing sentence
punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected
while `ADR-7` is still nameable in a warning. They now sit adjacent with the
reason written down, so they cannot drift apart silently.
And the stale citations in the moved comment are corrected. It pointed at
src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had
drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is
`gsd-core/templates/roadmap.md:32`. Both are now anchored by content.
* fix(#3697): stop the warning claiming a misparse that did not happen
Round-3 review Major 3 and Minor 4. Both are the same defect: the warning
asserted more than the evidence supported.
MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05
deferred` warned "could not be parsed ... Range forms are not expanded;
rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05,
i.e. every ID written on the line. Nothing was dropped, and the pinned
control only stayed silent because its pair was ADJACENT (gap == 1), so the
control was passing by accident of the fixture rather than by the rule.
The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a
range and as an annotation separator are textually identical, and no
token-level rule separates them; staying quiet re-opens the exact silent
under-selection #3697 is about. Deciding the ambiguity by assertion in
either direction is wrong. So it is DISCLOSED: the warning now has two
channels, chosen by whether any ID-shaped token was actually left unselected
(`droppedIdShaped`).
* something was dropped (tight range, glued fragment, inert residue)
-> "could not be parsed as a comma-separated REQ-ID list", as before.
* nothing was dropped (only the spaced-operator rule fired)
-> "contains what reads as a range between two cited REQ-IDs", stating
both readings and saying explicitly that an annotation separator means
the line is already correct.
This retires the "could not be parsed" wording for the four #3697-1 spaced
cases too, and that is a deliberate expectation change rather than a fix
counted twice: those lines never failed to parse either. They still warn,
still name the selected IDs, and still assert the endpoint-only marking is
unchanged; #3697-1 now also asserts the misparse channel stays SILENT.
MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a
citation as requirement content it had failed to read. The trigger is
correct and stays: #3697's acceptance criterion asks for a warning "when it
selects zero IDs from a line that is non-empty and is not the `TBD`
placeholder", and inferring placeholder-ness from arbitrary prose is the
free-text heuristic this detector exists to avoid. What was wrong is the
wording, so the non-range arm now says "ID-shaped text that was not
selected" and names the escape the author actually has (`TBD` / `None`).
Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse,
must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A
(tracked in ADR-12)` — must warn, must not say "Unparsed text", must not
diagnose a range, must name the placeholder escape).
Reversion control: reverting the ambiguous channel fails #3697-9 (3 named
tests); reverting the R3 wording fails #3697-10 (2 named tests).
* fix(#3697): cap every token predicate, complete the dash set, cover the boundary
Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three
are about the detector's own predicates, so they land together.
BLOCKER 2 — the 2048-char budget had no boundary coverage.
`RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit /
limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 /
2048 / 2049 against BOTH predicate families the cap guards, and each asserts
its fixture's exact length before asserting behaviour, so a mis-built
fixture fails loudly rather than passing at the wrong size. Clause (d) of
that rule — an input pushed within reserve-distance of the limit — has no
referent here: this is a hard cap with no reserve constant beside it, and
the test comment says so rather than leaving the omission to be re-derived.
NIT 6 — the cap guarded only the unanchored ID-substring regex. The
anchored range regexes were left uncapped, justified by a comment asserting
they scan linearly. The finding is right that this is informational (they
are anchored; the input is a local ROADMAP.md), but an asserted property is
cheaper to enforce than to defend, so all three predicates now share one
`short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must
now decline to classify.
SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this
code fixes at author time over a domain that grows without it, so the round
owes a census of what the enumeration reaches.
reached: `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through
NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure
dash, U+2015 horizontal bar, U+2212 minus sign
consequence: a range spelled with any of those is SILENTLY under-selected
— #3697's own defect, in the code that exists to fix it
Those five close. They are the same operator at a different codepoint and
carry none of the ASCII hyphen's collision risk, because they are not the
REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a
U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm
beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is
held to is untouched, and #3697-12 pins that.
Still NOT reached, declined with reason rather than left unstated: `→`, `~`,
`..=`, `..<`, `until`, and `up to` (two tokens, so never one operator
token). Each is a symbol or word with an independent non-range use between
two REQ-IDs — the over-warning class #2334 cost three rounds.
Reversion control: reverting the uniform cap fails #3697-B2 (limit+1);
reverting the dash set fails #3697-11 (5 named tests).
* test(#3697): add the fast-check property coverage the parser rule requires
Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md)
requires modules implementing parsing contracts to carry at least one
fast-check property test asserting a domain invariant, and the round-2 diff
had zero occurrences of `fc.` across its +311 test lines. Five properties,
1,900 generated cases:
P1 soundness of silence (boundary containment) — for ANY canonical comma
list of well-formed REQ-IDs, the selected set EQUALS the written set
and nothing warns. This is the #2334 over-warning invariant and the
#3697 under-warning invariant asserted as one statement, over
generated IDs rather than hand-picked ones. It generalises #3697-4b:
a prefix beginning with a word operator (`TORANGE-05`) is an ID, and
P1 covers that class rather than the single example.
P2 completeness — a same-prefix pair with an interior between them,
separated by any of the nine spaced operators, ALWAYS warns.
P3 the #2334 invariant — an ADJACENT pair around a separator can drop
nothing, so it stays silent however it is annotated.
P4 totality + idempotency — total over arbitrary strings, deterministic,
and the formatter agrees with the analysis on whether there is
anything to say (a warn with no text, or text with no warn, is a
channel that can go silent or noisy on its own).
P5 containment — every selected ID is ID-shaped and appears verbatim in
the input.
Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and
P5 hold against the round-2 code as well as this one — they are regression
guards, not bug-finders, and their value is that the invariants are now
stated and generatively checked rather than implied by examples. P2 is the
one that would have failed before the dash enumeration was completed.
fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from
`fc.array(...).map(join)` with the alphabet pinned to the selector's own
`[A-Z0-9]` class.
These live in tests/phase.test.cjs rather than a new
`phase.property.test.cjs`: `lint-test-file-count` caps a production module
at 2 test files and phase.cts is already at its allowlisted entry, so a new
file would trade one gate for another.
* docs(#3697): document the ROADMAP Requirements-line grammar
Round-3 review Minor 5 — the change adds net-new user-visible warning output
for a grammar constraint documented nowhere under docs/. `type: Fixed` is
docs-exempt so this does not block, but a warning about a rule the reader
cannot look up is not actionable, and that is worth fixing whether or not a
gate demands it.
Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the
existing SUMMARY artifact-check advisory it is a sibling of: the supported
comma-list form, why ranges are deliberately not expanded, that `TBD` and
`None` are the entire placeholder vocabulary, and what each of the two
warning voices means — including that the range/annotation one may be
reporting a line that is already correct.
Existing file rather than a new one, deliberately: docs/ carries generated
indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity
or index gate this change has no reason to touch.
* fix(#3697): rule-scope the warning-channel discriminator
Self-found at the round's pre-push review, against the Major 3 fix two
commits back. That fix chose the channel from a LINE-GLOBAL question — "was
any ID-shaped token left unselected?" — while the rules that produce the
warning are not line-global. The two disagree as soon as the line carries an
ID-shaped token no rule fired on:
`RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)`
`(ADR-7)` survives the selector's bracket strip, so the global test called it
a drop and sent the line to the assertive channel — putting the false "could
not be parsed ... rewrite the line" claim back on a correct line. That is
review finding Major 3 returning through a side door, and it directly
contradicts #3697-4, which pins a parenthetical citation as NOT unparsed
residue.
The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the
ambiguous channel requires that R2 fired, that no other rule did, and that
every endpoint R2 fired on was actually selected. The last conjunct is not
redundant — the detector shaves brackets and the selector does not, so R2 can
fire on a `(RANGE-02)` that was never selected, and that IS a drop:
`RANGE-01 (RANGE-02) — RANGE-05` -> assertive, correctly
`droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the
channel can ask about the endpoints the rule fired on) and the
`rangeReadingOnly` verdict.
Reversion control: against the line-global rule, #3697-9b fails. #3697-9c
passes under both rules — there the dropped token IS the R2 endpoint, so the
two agree; it is a regression guard, not a bug-finder, and is recorded as
such rather than counted as a second control.
* fix(#3697): hold every dash to the strict range shape, not just ASCII
Self-found at the round's pre-push review, and it CORRECTS a claim made two
commits back. That commit widened the range-operator set by five Unicode
dashes and asserted they "carry none of the ASCII hyphen's collision risk,
because they are not the REQ-ID separator". That reasoning was wrong. The
collision is a property of the SHAPE — `PREFIX-\d+ <dash> \d+` is also a date
(`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not
care which dash sits in the operator slot, because the ID's own separator is
still ASCII either side of it. Measured:
RANGE-01 (target FY-2026-08) silent <- pinned by #3697-4
RANGE-01 (target FY-2026‐08) WARNED <- same line, U+2010
So the widening reintroduced the #2334 over-warning class on a date
annotation. It also exposed that the inconsistency PREDATES this PR: U+2013
and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash
forms of that same date annotation warned before round 3 ever ran.
One rule for every dash: a tight range spelled with any of the eight must
carry a FULL ID on both sides, exactly as the bare hyphen already had to.
`..`, `…` and the word operators stay loose — no date or sub-number reading
exists between two numbers, so the strict shape would cost them coverage for
nothing.
The cost is a false negative, and it is one the design already accepts:
`RANGE-01, RANGE-02-05` is silent today, deliberately, and now
`RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than
opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so
R3 catches it.
Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight
dashes), #3697-13b (full-ID tight range still warns for all eight),
#3697-13c (loose operators keep their numeric endpoint), #3697-13d (the
accepted false negative is symmetric, and the bare zero-selection line still
warns).
* fix(#3697): close the round's own pre-push review findings
An adversarial cross-AI review of this round refuted 4 of its 10 claims. All
four were real. Every fix below is to code THIS round introduced.
1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1).
`REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told
the author "the line is already correct and nothing needs to change" — while
`(REQ-02)` had been dropped by the selector, which does not strip
parentheses.
The channel choice is still right, and deliberately so: `(ADR-7)` and
`(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts
the false "could not be parsed" claim back on a line carrying a citation —
the misroute fixed two commits ago. No rule can adjudicate this; the author
can. So the voice stops asserting the line is correct (it now speaks about
the SEPARATOR, which is all it has evidence about), and BOTH voices gained a
factual clause naming ID-shaped text the selector skipped, with the reason
(brackets are not stripped) and no verdict attached.
2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2).
A 2049-char range token warned before this round and went silent after it:
the "uniform cap" commit bounded the predicate and, with it, the warning.
That is #3697's own defect, introduced by the fix for a nit.
The cap bounds the WORK, not the warning. An over-cap token carrying `-` is
now recorded as unclassified (a linear `includes`, never the unanchored
regex the cap exists to keep off it) and gets its own voice: "could not be
checked ... the REQ-ID selection on this line is unverified". Unclassified
is reported, never treated as clean.
3. THE CAP WAS NOT UNIFORM (review MISSED finding).
R2 capped the operator token but not its neighbours, so
`<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over
both endpoints unbounded. The glued rule had the same hole. Every
participant is capped now.
4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5).
It drew from a bare `fc.string()`, which over 500 samples produced max
length 10 and ZERO inputs containing a REQ-ID — the loop body never executed
an assertion. A containment property that never contains anything is a green
test measuring nothing. The generator now interleaves real IDs with noise
and the property ASSERTS it saw them (>50/500), so it can never silently go
vacuous again. The free-form coverage it was actually providing survives,
honestly labelled, as #3697-P6.
The same finding refuted this round's claim that P2 distinguishes pre-round
behaviour: every operator P2 uses was already in the pre-round operator set.
P2 is a regression guard, and its comment now says so.
Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review
MISSED finding) and is corrected with the code.
Tests: #3697-9d (soft voice names the skipped ID, never claims the line is
correct), #3697-9e (over-cap token reported as unclassified, still warns),
#3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the
boundary on the PREDICATE's verdict with `warn` asserted true at every length —
asserting `warn === false` at limit+1 was itself finding 2.
* docs(#3697): describe the third voice and the dash rule
Follow-on to the review-findings commit: that commit corrected the docs' claim
about the soft voice but left two things the code now does undescribed.
- There are THREE voices, not two. The over-cap voice ("could not be checked
... unverified") arrived with the fix for the review's CLAIM 2 and had no
entry.
- Dash spellings require a full ID on both sides, and `..` / `…` / the word
operators do not. That asymmetry is deliberate and load-bearing —
`PREFIX-<digits><dash><digits>` is date- and sub-number-shaped — so a reader
hitting `REQ-01-05` and getting silence has no way to find out why. The
accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still
reported) is stated rather than left to be discovered.
Documentation only; no behaviour change.
* fix(#3697): close the continuation review's findings
A continuation of the same adversarial reviewer, run against the reworked
round, refuted 6 of 7 claims. Four were real defects in this round's own work
and are fixed here; the other two are answered rather than changed, below.
1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A).
The previous commit's message said both voices gained it. Only the soft
return appended it. The assertive voice now carries it too — and, because
that voice already names range tokens and inert residue under its own
clauses, the note is filtered to what those did not already name. A warning
that says the same token twice is one readers learn to skim.
2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A).
It read "brackets and parentheses are not stripped". Square brackets ARE
stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so
only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md,
which had inherited the same error.
3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B).
`oversizedTokens` filtered on `includes('-')`, which misses an over-cap
OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined
to classify it once capped, and nothing reported it. That is the exact
regression the field was added to close, one input over. Any token past the
cap now counts — what it contains is irrelevant when we could not read it.
4. AND THEN OVER-REPORTED ONE (review MISSED finding).
With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the
uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line
is unverified" — a contradiction inside one warning. A token the selector
took was examined end to end, so it is excluded.
Two findings are answered, not changed:
CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and
deliberate: this round does not touch what gets MARKED, and the pattern is
anchored at both ends with no nested quantifier, so it is linear. The claim
that "all predicate paths are capped" was too broad; the DETECTOR's are.
CLAIM E — `REQ-01, REQ-02<dash>05` is silent for every dash. That is the
documented, deliberate cost of holding dashes to the strict shape, and it is
symmetric with ASCII, which behaved that way before this PR. The reviewer is
right that "without losing a range spelling that should be detected" was too
strong; a bare `REQ-02<dash>05` still warns.
Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim
true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a
selected over-cap ID is never called unverified). #3697-9f is rebuilt — its
first version used the SAME id twice, so R2 could not have fired even uncapped
and it proved nothing; it now uses endpoints with a gap and fails when the
neighbour cap is removed.
Reversion control: all four fixes fail a named test when reverted in isolation
(#3697-9h, #3697-9i, #3697-9g, #3697-9f).
* fix(#3697): scope the over-cap exemption to what could actually pair
A second continuation of the same reviewer, against the reworked round,
confirmed the two claims that matter most and refuted three. This closes the
one real defect; the other two are answered below.
CLAIM J / CLAIM K (one defect, found from both directions). The previous
commit exempted EVERY selector-accepted token from `oversizedTokens`, on the
reasoning that the selector is uncapped and anchored so it examined the whole
token. True of that token's SELECTION — and not the same as "no rule was
suppressed by it". Two over-cap valid IDs either side of `..` are both
selected, so both were exempted, and R2 is capped: a line that warned before
this round went silent.
That is the third appearance of one class in this round — the cap suppresses a
check, and the suppression is not reported. Each fix for it over-corrected in
the opposite direction, which is why the rule is now stated in terms of what
was actually suppressed rather than in terms of the token: an over-cap token is
exempt only when it was selected AND nothing beside it could have paired with
it into a range (no range operator, no glued fragment, no second over-cap
token). Everything else is unexaminable and says so.
Two findings are answered, not changed:
CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule
that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01`
silent: recognise `PREFIX-a<dash>b` only when another SELECTED id on the line
shares that prefix. That refutes this round's claim that the strict-dash
trade was FORCED, and the claim is withdrawn — it is a design choice. The
choice stands for this PR: the conservative rule is what ASCII already did
before #3697, adopting a new contextual heuristic unreviewed at the end of a
round is how the last three defects in this round were made, and #3697 asks
for a warning rather than better range inference. Named here so the
alternative is on the record rather than lost.
Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not
classified at all" and warns. Selection is not bounded; only range detection
is. Corrected.
Confirmed by the same pass, and worth recording because they are the PR's
load-bearing promises: a 20,000-input comparison of the pre-extraction selector
against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has
changed), and the uncapped selector regex was measured linear from 100k to 800k
characters.
Tests: #3697-9f now asserts the range case is reported rather than silent, and
#3697-9j pins the exemption's scope in both directions. Reversion control:
restoring the blanket exemption fails both.
* fix(#3697): warn on zero selection, as the acceptance criterion asks
`Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed
SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`,
the `placeholderLed` census comment, and the advice string the command emits
to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned,
because the citation supplied the ID-shaped residue R3 required, while bare
`Deferred` did not. The claim was written into three places and never
executed once.
This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length ===
0` while the raw capture is non-empty and not `TBD`".
R3b keys on the SELECTION being empty, never on what the prose means, so it
adds no free-text heuristic. It is deliberately not gated on ID-shaped
residue the way R3 is, and the negative space is what settles that: all
fifteen #2334/#2339 fixtures are held silent by non-zero selection or by
`placeholderLed`, and not one of them by the ID-shape gate — measured, not
argued. The gate was buying no negative space while costing the acceptance
criterion.
`tokens.length > 0` keeps an empty line and a comment-only line silent: the
tokenizer strips `<!-- ... -->` before splitting, so the shipped template's
own comment cannot reach the rule.
Selection behavior is unchanged. This warns; it never invents an ID.
Also extracts `warn` to a named const (round 3 review Minor 3) — this commit
adds a disjunct to exactly that predicate, and in the return literal a later
reordering would be a TDZ ReferenceError rather than a reader-visible error.
Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the
warning names the TBD/None escape), #3697-14b (five placeholder spellings
stay whole-channel silent), #3697-14c (comment-only line stays silent).
Fail-first controls: all six #3697-14 cases fail against the pre-fix tree;
-14b and -14c pass at both ends, which is correct — they pin silence the
widening must preserve.
* fix(#3697): name the REQ-ID a glued delimiter dropped
`RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with
`requirements_updated: true` — #3697's own half-success failure mode, reached
by one wrong delimiter, and silent before this rule. It is the issue's AC-1a
("a warning whenever the line contains ID-shaped content that the tokenizer
did NOT select") at the shape most likely to be typed by accident.
Round 4 review rated this Major rather than Blocker on the ground that the
case is indistinguishable from a parenthesised citation, since `(ADR-7)` also
shaves down to a bare ID. At the RAW token level it is distinguishable, and
that is what makes the rule shippable: `REQ-01;` is shaved of a trailing
DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and
requires the token to sit outside any parenthetical.
Measured before implementing: 0 false positives and 0 false negatives across
21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first
cut without the parenthetical test scored 3 false positives — every one of
them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that
test is the rule's boundary rather than an optimisation.
Adds the delimiter census the module did not have. The range-operator domain
was already censused; the comma-substitute domain was not. Swept 26
spellings: exactly two produce a silent under-selection, `; ` and `: `. Every
other spelling either selects both IDs or selects none and already warns. The
review hand-listed the semicolon; the colon is the sibling that sweep found,
and it fails identically.
`rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing
was dropped, and must not speak for a line where something demonstrably was.
Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick
exactly the unchanged selection), #3697-15b (three citation forms stay
whole-channel silent). Fail-first control: all four #3697-15 cases fail
against the previous commit's tree; -15b passes at both ends, pinning the
boundary the widening must not cross.
* fix(#3697): give the Requirements-line warning a stable machine kind
The warning's kind existed only in the prose of its message, so every consumer
and every test had to regex an English sentence — and rewording a message
silently un-asserted the tests that pinned it. Round 4 review Major 3.
The repo already had the settled seam for exactly these semantics.
`CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a
truncated scan, which is precisely this module's third voice; and
`WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same
reason. ADR-3473 Decision 3 ("failure is a value") points the same way.
`formatRequirementsLineWarning` now returns `{ code, message }` instead of a
bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a
legitimate value, but the success arm is no longer a naked string one field
away from the shape the ADR standardises on.
The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is
a documented `string[]` in `phase complete`'s JSON output, rendered by
execute-phase.md's "If has_warnings is true" step, so re-typing its elements
would be a breaking output-contract change for a shipped command. The code is
emitted as its own additive `requirements_line_warning` field, absent
entirely when the line is clean.
Vocabulary, exported so tests key on it rather than on string literals:
`req-line-misparse`, `req-line-range-reading`, `req-line-unverified`.
Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code;
message-content assertions stay where the user-visible wording is itself under
test. #3697-16 pins the code end-to-end through the CLI's JSON for four line
shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean
line emits no kind at all, because a field present on every run carries no
information. #3697-P4 holds kind-and-message-appear-together and
kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added
later cannot ship without one.
* test(#3697): pin the divergence against the second parser of the same line
CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when
sharing constants/arrays/parsers between parallel surfaces, add a parity
assertion test that fails if they diverge." Round 4 review Major 2.
`normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP
`**Requirements:**` value — its own docblock says callers "may pass the
roadmap value through verbatim" — and diverges on four axes. Measured, not
inferred:
line phase complete gap-checker
RANGE-01..RANGE-05 [] 5 IDs
None (per ADR-7) [] ["ADR-7"]
(REQ-02) [] ["REQ-02"]
REQ-01a [] ["REQ-01a"]
REQ-01, REQ-02 both both
This pins the divergence rather than removing it, which is the review's
second option and the correct one here: unifying the two would change what
`phase complete` MARKS, and "the ledger-writing set is byte-identical to base"
is the one invariant this PR holds fixed. Every axis is now asserted in BOTH
directions, so drift on either side fails here instead of widening silently.
The range axis is a DELIBERATE disagreement and is labelled as such — #3697
declines range expansion in terms ("I am not asking for range syntax to be
supported") while gap analysis adopted it under #1269.
The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a
declared-empty line to `phase complete`, which reads the lead token, and a
one-requirement line to gap-checker, which strips parentheses first so the
citation survives its ID-shape filter. That is a citation being reported as a
requirement.
#3697-17b states the cost concretely: one line, five requirements in scope to
gap analysis and zero to phase complete. This PR is what makes that
contradiction visible, by finally giving the silent side a voice.
* fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap
Two round 4 review minors, both about a warning saying something it cannot
support.
MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so
`FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`.
`REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape
silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no
warning at all — but whenever some OTHER rule fired on a line that also
carried a date annotation, the rider told the author to "check whether any of
it is a requirement" about a date. Not a false warning, since the line was
warning anyway; false CONTENT, in the #2334 voice, through the side door.
Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped`
stays a faithful record of what the selector skipped — it is documented as a
fact that never routes — while the user-facing clause declines to assert
requirement-ness about a shape the design already ruled unadjudicable.
#3697-18b is the other half, so the filter cannot become a silencer: a
genuinely dropped REQ-ID is still named.
MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent,
so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel
would have passed. Whole-channel silence is not available on that fixture (the
pre-existing ghost-ID warning legitimately fires on the unregistered
`TORANGE-05`), so the precise assertion is that no Requirements-line warning
of ANY kind was emitted. The machine code added earlier in this round is what
makes that statable; before it, "both channels" could only have meant a second
prose regex.
* docs(#3697): record the Requirements-line seam in CONTEXT.md
CLAUDE.md names the CONTEXT.md glossary as a PR gate, and
`get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25
co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a
named seam, three warning kinds, a bound, a rule taxonomy and a deliberate
cross-parser divergence, and recorded none of it. Round 4 review Major 4.
The precedent is explicit rather than inferred: the directly analogous seam
is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound
and its boundary obligation — and that entry is the one this module's third
voice was modelled on.
Eight predicates, in the machine-oriented section beside it:
.module the two exported functions and the code vocabulary
.selector-identity citedReqIds is byte-identical to base and is the
only thing reaching the ledger — a change to what
phase.complete MARKS is outside this contract
.rules R1 / R2 / R2' / R3 / R3b / R4 / over-cap
.kinds the three codes, and why they ride beside
warnings[] rather than inside it
.cap 2048, neighbours included, and the boundary rule
.placeholder the gate that actually holds the negative space
.census-domains both open domains with their NOT-reached members
.gap-checker-divergence the four axes, pinned not unified
The changeset type is `Fixed`, which exempts this PR from the docs/
co-change requirement — but the glossary gate is separate from that
exemption, and the 2048 cap in particular is a machine-canon-shaped fact
that until now existed only inside a source comment.
`docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync
confirms all six targets in sync.
* docs(#3697): document what the command now does, in one changeset sentence
DOCS. The grammar section predated this round's two new rules, so it
under-described the behaviour it exists to make lookup-able:
- The placeholder paragraph enumerated three words; the rule is a DEFAULT.
Any wording that selects no REQ-IDs warns, and the placeholders are matched
as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty
too. The comment-only line is called out, because "any other wording" would
otherwise read as covering the shipped template's own `<!-- ... -->`.
- The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is
the quietest way to lose a requirement on this line — `requirements_updated`
reads `true` either way — so it gets its own paragraph, with the
parenthetical exemption stated beside it.
- The machine kind is documented where a consumer would look for it, with the
instruction to key on the kind rather than the wording.
- The skipped-text note no longer implies it reports date shapes; it
deliberately does not, and silently omitting that left the doc promising the
behaviour this round removed.
CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is
`**<Bold user-visible change>** — <symptom-led explanation>.` and both
canonical examples are one sentence; this fragment ran three. Now one, and
covering what the round actually delivers rather than only the range shape it
started from.
* test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint
A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The
docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/`
references in exempt test files and fails when a new one appears, so that
comment turned four green gates red — `ci-docs-guard-registry` and the
registration lint — for a file that reads no documentation at all.
Caught by diffing the full suite's failing-name set against the same suite run
at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base
(install / config-home / shadowing tests under the sandbox HOME), and exactly
these 4 did not.
Rephrased rather than baselined. Adding the path to
DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely
starts READING a new docs path — the violation text asks the author to
re-confirm the exemption still holds. Nothing here reads documentation; the
guard matched prose. Baselining would have recorded a coupling that does not
exist and made the next reader wonder what phase.test.cjs does with
CLI-TOOLS.md. The comment still names where the contract is written, just
without planting a path string.
* fix(#3697): close four defects this round's own pre-push review drove
An adversarial cross-AI review of this round, run before the push, returned 6
CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree,
and every one was a shape the author had not probed — the rules were correct
across the probe set and wrong just outside it.
(1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through
the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired:
`ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does
not reach a BARE citation. The FP probe that scored this rule 0/0 only ever
tested the parenthesised form.
Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the
module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already
demands an agreeing prefix for the same reason. Cost, stated in the census: a
dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays
silent. Same trade the strict-dash rule takes — under-report a rare shape
rather than over-report a common one. Pinned as a declared blind spot by
(2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped
REQ-01 silently: the selector strips square brackets and R4's raw scanner did
not. The bracket spelling the shipped template recommends was the one shape the
rule could not see.
(3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal
requirement id — gap-checker's `parseRequirements` accepts it from
REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement
behind a rule meant only to hide dates. Narrowed to a four-digit year segment.
The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED,
and the removal is recorded in place: its premise was refuted, it was not
inconvenient.
(4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while
reading as empty to the author, so R3b fired with nothing on screen to explain
it. Zero-width and format characters are now stripped — stripped rather than
treated as delimiters, because splitting on one would fabricate two fragments
out of one ID.
Also corrects the documentation the same review found overstated: the line is
split on commas AND whitespace, and the ID shape is matched case-insensitively,
so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was
pre-existing selector behaviour; this round is the one that asserted the docs
were true of it.
Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is
stripped, not split on), -19c (three citation forms), -19d (both bracket
spellings), -19e (both halves of the rider boundary), -19f (the declared blind
spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0
failures; lint:ci clean.
* fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop
The pre-push review's continuation refuted six of seven follow-up claims. The
first one is the one that mattered: the invisible-character fix committed in
c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from
the detector wholesale made `REQ-01<ZWSP>, REQ-02` go SILENT — the selector
really does drop REQ-01, and the strip removed the only evidence of it. The
test written alongside asserted the tokens and the empty R4 result and never
asserted `warn`, so it DOCUMENTED the bug rather than catching it; that
omission was the reviewer's own MISSED finding.
An invisible is two different questions about one character, and the fix is to
stop conflating them: absence-of-content for the empty test, DECORATION on a
token for the drop rule. Neither is a reason to delete it from the line.
R4 is generalised accordingly, because the continuation drove four more shapes
a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`,
`REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus
`**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one
class: decoration on a token the selector then cannot take. One rule, not four
patches; patching them individually is how a list stays short and wrong.
PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving
them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was
never there and broke #3697-9d's channel routing with it. A parenthesis is this
rule's citation marker.
The rider stops adjudicating an undecidable shape. `API-2-01` is a legal
requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15`
are dates — no regex separates them, and both filters this round tried scored a
miss in each direction. It now NAMES the token and states the ambiguity, which
is the same thing the two warning voices already do about a range separator.
Filtering hides a real dropped requirement; reporting it bare asks the author
whether a date is a requirement; saying "this may equally be a date" does
neither.
The census and the docs are corrected to what the code does, including the part
that is NOT complete: the prefix gate does not stop a citation that SHARES a
selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level
separates that from a real drop. A prose heuristic on "see" is the free-text
detector this module exists to avoid, so the honest move is to say so.
CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a
20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is
identical to upstream/next's for every input, and marking is untouched.
513 tests in phase.test.cjs, 0 failures; lint:ci clean.
* fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line
Third pass of the round's own pre-push review, scoped to regression-hunting
rather than further polish. Three findings, all driven, all mine.
R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed
a dropped requirement: the previous cut treated any shaved decoration as
evidence, and emphasis is not evidence. Nothing separates that line from
`**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive
signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the
token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to
the skipped-text rider, which names the id without asserting a drop, exactly as
`(REQ-02)` is handled. That is the #2334 class caught one cut before shipping.
R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper
sitting behind a delimiter. Shaves to a stable point now.
The range OPERATOR lost its invisibles handling. `REQ-01 <ZWSP>..<ZWSP> REQ-05`
went silent, because the previous commit removed the invisible strip from BOTH
the tokenizer and R4 when only R4's was wrong. An invisible is two questions
about one character: for the classification rules it is noise and is stripped
from the token; for the drop rule it is the evidence and must survive on the
raw line. Stripping in both places hid a dropped id; stripping in neither hid a
range. The reviewer's MISSED finding named the missing control — regression
tests covered invisibles inside ids and not beside operators — and #3697-19i is
that control.
UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03`
reported nothing: a running-depth counter left the unclosed `(` open through
end-of-line, so every genuine drop after it inherited citation immunity. A
parenthesis confers that immunity only as part of a MATCHED span now — an
unmatched one is a typo, not a citation.
CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode:
`citedReqIds` identical to upstream/next, marking untouched, warnings appended.
522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero
head-only failures against a probe worktree at upstream/next.
* fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright
Fourth and final pass of the round's own pre-push review. Three findings, and
they shared one root cause, so this is a narrower rule rather than a longer list
of shapes.
TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single
whitespace token, so a matched parenthetical inside it conferred immunity on the
`REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans
are now deleted from the line outright — which states what is actually meant,
that for this rule a citation is not on the line — and an UNMATCHED paren is a
typo that confers nothing. That also retires the running-depth counter whose
previous bug was the mirror image: an unclosed `(` swallowing the rest of the
line.
DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was
reported as a dropped requirement. It is a citation with sentence punctuation.
The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be
touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify;
`**REQ-01**;` does not, because outside the styling that character is
punctuation. An invisible needs no adjacency test — nobody types one on
purpose, so anywhere in the token it is corruption rather than intent.
`**REQ-01**; REQ-02` therefore goes silent, and the test row asserting
otherwise is inverted rather than deleted quietly: it was added one commit ago
on the reasoning this pass refuted, and nothing distinguishes it from
`see **REQ-7**; next topic`.
Worth recording plainly: three successive cuts of this rule fired on a
citation, and each fix was a narrower definition of EVIDENCE, never a longer
list of shapes. The list-lengthening instinct is what produced the bug each
time.
The review's last MISSED finding named the missing control — the paren tests
all surrounded matched spans with whitespace, so none covered a span sharing a
token with an id outside it. #3697-19j carries both directions now.
526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero
head-only failures against a probe at upstream/next; the invariant that
`citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary
Unicode inputs.
* fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop
Two defects, both found by this round's own pre-publication body claim-audit.
1. A DEMONSTRATED drop was discarded by the unverified voice. On
`REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in
delimiterDroppedIds and the formatter then reported `req-line-unverified`,
whose message never mentions it — the one actionable finding masked by the
token beside it. The over-cap channel now excludes a line carrying an R4
hit, exactly as rangeReadingOnly already did and for the same reason: that
voice's whole claim is that nothing could be checked, and R4 has already
checked something. The assertive channel still carries the over-cap rider,
so nothing about the cap is traded away. Pinned by #3697-19l, which fails
against the pre-fix build and nothing else does.
2. Three shipped artifacts asserted behaviour the code does not have — the
same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules
predicate, the CLI tools reference, and the warning's own advice string all
listed markdown emphasis as an R4 trigger. It is not: styling is shaved
BEFORE the test and tolerated around an id, never a trigger on its own, so
`**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger
is exactly a glued `;`/`:` or an embedded invisible.
The census predicate was wrong in a second way. Its 26-spelling separator
sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and
therefore biased: one-sided attachment drops silently for every punctuation
outside the set — `/ | & + . > \` and the full-width and non-ASCII forms
`; , ؛` all measured silent. The domain is wide open and R4 covers two
characters of it. Said plainly in all three places rather than widened
here: every previous widening of this rule first fired on a citation, so it
is not done blind at the end of a round.
Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n
one-sided separators) so the documents and the code cannot drift apart again —
which is what the round-4 blocker asked for.
tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* fix(#3697): re-sweep the separator census properly, and say what it really found
The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `"
from a 26-spelling sweep. That conclusion was forced by how the sweep was
built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the
semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for
every other separator. Different members of the domain were tested in
different shapes, so no other answer was reachable. Caught by this round's
pre-publication claim-audit of the response comment, reading the census
comment against its own swept list.
Re-swept fully crossed and driven through the built artifact: 21 separators x
{bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select
both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All
34 are one shape — a separator glued to exactly one of the two ids, e.g.
`REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and
the `;`/`:` that R4 covers.
So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census
comment, the CONTEXT.md census-domains predicate and the CLI tools reference
all say. The set is deliberately not widened here: three successive cuts of
this rule fired on a citation, and a fourth at the end of a round with no
adversarial pass is how each of those got in.
Second false passage in the same block: styling-only decoration was described
as "left to the skipped-text rider, which names the id". A rider only exists
inside a message, and a message only exists once some rule sets `warn` — so on
a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent.
Describing it as handled reads as coverage. #3697-19m already pins the silence.
so the test matches the documented claim.
tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next
Rebased onto next @
|
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
1a358ce0fd |
feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
62b0d939b6 |
feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey Add an optional `timeoutConfigKey` field to the reviewer lane descriptor, resolved in `resolveLanePlan` at invocation time and falling back to the frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes declare `review.timeouts.<slug>` on both surfaces (the descriptor and their capability.json manifest), validated by capability-validator.cjs. For the antigravity lane, the native `agy --print-timeout` flag — previously a second hardcoded literal (`540s`) independent of the outer cap — is now derived from the same resolved outer timeout in `antigravityArgv`, preserving the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is declared data, the inner one is handler-owned). The antigravity default timeoutFloorMs stays at 600s per the maintainer's disposition; users raise it through the new config key instead. * docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper Address code-review findings on the timeoutConfigKey change: extract the inline timeout-resolution logic into a named, exported, directly-tested resolveTimeoutMs helper (matching the file's existing configString/ normalizeHost convention); document the new review.timeouts.* federated config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md, and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment. * fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner gsd-test caught two design mistakes in the prior commits: 1. SpawnPlan.argv is documented and tested as fully resolved by resolveLanePlan (model/effort/output/prompt already folded in) — leaving the antigravity '{{nativeTimeout}}' marker unresolved until the runner's antigravityArgv violated that contract and broke tests that read plan.argv directly (tests/antigravity-reviewer.test.cjs, tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}' is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself via the new nativeTimeoutToken() helper, exactly like the other four. antigravityArgv reverts to its pre-#3274 four-argument form. Also missed updating capabilities/antigravity/capability.json's invoke.args to match the descriptor, which broke the manifest/descriptor parity test. 2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three lanes with neither a model flag nor a host — may own no config key beyond their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes violated it. Fix: those three keep timeoutConfigKey: null and own no review.timeouts.* key, matching their existing modelConfigKey: null. The other 9 lanes are unaffected. * chore(#3274): backfill changeset PR number (pr:0 -> 4083) --------- Co-authored-by: sim <sim@local> |
||
|
|
6eea00b707 |
enhance(#3301): tell reviewers the plan ids and total count, grade coverage (#4084)
* test(#3301): add failing-first plan coverage manifest tests Failing-first regression tests for the plan-id manifest, the updated Review Instructions, and the mechanical per-reviewer coverage check, ahead of the review.md implementation. RED baseline before the fix lands. * test(#3301): raise allow-test-rule-refs unverified ceiling for new marker Adding tests/review-plan-coverage-manifest.test.cjs's source-text-is-the-product marker grows the unverified-exemption pool by one (282 -> 283), the same documented growth path scripts/lint-allow-test-rule-refs.cjs's own failure output names. Confirmed clean via 'npm run lint:allow-test-rule-refs' locally. * feat(#3301): tell reviewers the plan ids and total count, grade coverage build_prompt now derives a plan-id manifest from each *-PLAN.md filename (stripping the -PLAN.md suffix) and appends it, with the total plan count, to both gsd-review-instructions.md and gsd-review-prompt.md. The Review Instructions prose requires one heading-verbatim section per id before any cross-plan or overall-risk content. write_reviews grades each dispatched lane's real (non-stub, non-empty) review against that same manifest and records an optional plan_coverage: frontmatter block, present only when a lane is incomplete. The match escapes regex metacharacters in the id and excludes a preceding/trailing hyphen or word character as a boundary, closing the two traps named in the issue (a decimal phase like 12.6 satisfied by 12X6-01; a threat id like T-04-07 registering as coverage of plan 04-07). CodeRabbit is exempt, since it never receives the source-grounding prompt carrying the manifest. This closes the gap where a review that silently covers only some plans in a multi-plan phase is indistinguishable from one that covers all of them. * docs(#3301): add changeset fragment * test(#3301): use t.after() instead of try/finally for cleanup CONTRIBUTING.md bans try/finally inside test bodies. Code review caught this in the new coverage-manifest test file; switch every fixture-cleanup site to the approved t.after() pattern. * test: use t.after() instead of try/finally in #3300's build_prompt tests Pre-existing try/finally-for-cleanup pattern in this file (landed for #3300) violates CONTRIBUTING.md's explicit ban on try/finally inside test bodies. Surfaced incidentally while reviewing #3301's diff, which cites this file as its extraction-pattern precedent; fixed inline per the no-defer rule rather than deferred to a separate PR. * test(#3301): anchor coverage-check extraction on the fence line, not prose `.plans-manifest.md` also appears in write_reviews' own prose ahead of the ```bash fence, so indexOf found that occurrence first and the backward-walk-to-fence-open landed on the earlier, unrelated gate-check block instead. gsd-test caught this: coverage-check tests expecting a real verdict got null, because the wrong block ran and never writes .plan-coverage-<slug>.json. Anchor on the fence-only bash assignment line instead. Emitted-Drift-Ack-Growth: review.md — #3301 adds the plan-coverage manifest and mechanical coverage check to build_prompt/write_reviews. * fix(#3301): route id escaping through the canonical pattern seam ADR-3212 (epic #3212) consolidated ~44 hand-rolled regex-escape copies into one owner, src/pattern.cts's escapeRegex, specifically to stop this exact class of duplication. My coverage-check node -e script hand-rolled the identical metachar-escape regex — invisible to eslint-rules/no-adhoc-regex-escape.cjs only because it lives inside a workflow markdown file, not a .cts/.cjs source file the shape-matching guard scans. Require the compiled seam (gsd-core/bin/lib/pattern.cjs) instead, matching the established node -e-requires-a-compiled-lib idiom already used elsewhere in this workflow (code-review.md's code-review-flags.cjs/code-review-depth.cjs calls). Verified both named traps from the issue still resolve correctly under escapeRegex's RegExp.escape-backed implementation, which differs in escaped-text shape (hex-escapes hyphens/leading chars) but not match result. * test(#3301): run coverage-check block with cwd at the repo root The block's node -e now requires ./gsd-core/bin/lib/pattern.cjs, a path relative to the repo root (correct for production, which always runs from there). The test harness ran it with cwd at the fixture's own temp dir instead, so the require failed. Add an optional cwd param to runScript (default: root, unchanged for the plan-copy-block tests) and pass the real repo root for every coverage-check call site. Manually verified end-to-end before spending another remote run: the extracted block now produces the expected {complete:true} verdict. * docs(#3301): backfill changeset pr number (pr:0 -> pr:4084) --------- Co-authored-by: sim <sim@local> |
||
|
|
4d70b4dc43 |
fix(#4068): add --merge-async to c8 coverage-merge invocations (#4069)
* test(#4068): failing-first regression guard for c8 --merge-async flag Adds a config-invariant test asserting test:coverage:unit and test:coverage:report pass --merge-async to c8, guarding against the coverage-merge OOM (release.yml finalize dry-run, exit 134/SIGABRT) regressing a third time (prior stopgap: #199). Also commits the sourced research memo backing the diagnosis. * test(#4068): rename regression test to avoid lint-test-file-count collision coverage-merge-async-flag.test.cjs's effective prefix (coverage-merge- async-flag) matched the unrelated gsd-core/bin/lib/coverage.cjs module's 2-file cap under lint-test-file-count.cjs's startsWith bucketing, failing FAIL_EXCEEDS_LIMIT (confirmed via the RED gsd-test run on 7b6e6ca9e). Renamed to c8-merge-async-flag.test.cjs -- no colliding prefix. * fix(#4068): add --merge-async to c8 coverage-merge invocations The `test:coverage:unit` script OOM-crashed (exit 134, SIGABRT) in the release.yml finalize dry-run of 1.12.0: all 1785 unit tests pass, then c8's report/merge phase crashes ~76s later against the 6144 MB heap ceiling. Root cause (verified against this repo's pinned c8@11.0.0 source, node_modules/c8/lib/report.js): the default sync merge path, Report._getMergedProcessCov(), loads every raw V8 coverage file for the whole run into memory as one array before merging. This is a recurrence of #199 (same OOM class at ~466 tests, "fixed" by raising the heap ceiling) -- the suite outgrew 6144 MB as it grew to 1785 tests, and test.yml's own coverage-gate job already independently hit and stopgap-fixed the identical class once (test.yml:614-620, 4096->8192 MB). c8 ships the upstream fix for exactly this: --merge-async (c8 v7.14.0, already inside the pinned c8@^11.0.0 range -- no dependency bump), which switches to Report._getMergedProcessCovAsync(), reading and merging one raw file at a time instead of loading them all at once. The merge arithmetic (mergeProcessCovs()) and --all's zero-coverage-file inclusion are unchanged between the two paths, so the coverage gate's accuracy is unaffected. Verified empirically (not just by source inspection): against a 223 MB synthetic raw-coverage corpus derived from this repo's actual 235 gsd-core/bin/lib/**/*.cjs files (same order of magnitude as the 358 MB figure documented in test.yml), c8's real Report class crashes with the same "Ineffective mark-compacts near heap limit" signature via the sync path at a 300 MB heap ceiling, while the async path completes at a flat ~40-50 MB peak down to a 100 MB ceiling. Full repro steps and numbers: .gsd/bug/fix-4068-coverage-merge-oom/10-diagnosis.md. Only `test:coverage:unit` and `test:coverage:report` are changed. `test:coverage` (line 150) is a non-CI dev-convenience script, never invoked as a command by any workflow. `test:coverage:unit:raw` (line 153) already runs --reporter none, which skips the merge/report phase this flag affects, entirely -- adding it there would be a no-op. Fixes #4068 * fix(#4068): add changeset fragment for the coverage-merge OOM fix The changeset fragment was generated locally (npm run changeset) but never committed -- isolated code review caught the gap (no .changeset/*.md on the branch, confirmed via git diff --name-status). * chore(#4068): backfill changeset PR number pr:0 -> pr:4069 --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
472f585f7c |
fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates `milestone complete <version>` is a one-way door — ROADMAP.md and REQUIREMENTS.md archived, every phase directory in the milestone MOVED, STATE.md rewritten — and ran unconditionally on first invocation through every invocation path, including `query milestone.complete <version>`, whose `query` meta-prefix reads as a read-only namespace but performs no filtering (#167's invocation-compatibility shim + #3243's dotted-form normalization). The gate lives on the destructive command itself, not on the `query` prefix (the prefix is an intentional invocation mechanism, not a permission boundary — restricting it would break dozens of shipped workflow callers). Without --confirm and without --dry-run the command now refuses via error() before reading anything beyond its arg checks, so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run still previews with no confirmation needed and is now documented in the usage block (it was only documented for the sibling archive-quick). --force keeps its narrow meaning — bypassing the TRUNCATED-scope and unstarted-phase guards — and does not double as the mutation opt-in. --confirm follows the existing `phases clear --confirm` idiom in the same module. complete-milestone.md's two invocations pass --confirm (the workflow has gathered explicit user intent by that step). Existing tests get --confirm appended — pre-change behavior is exactly confirmed behavior — and a #3726 regression block covers: refusal + full-tree byte-identity on both invocation forms, --force not satisfying the gate, --dry-run still passing without confirmation, and --confirm proceeding. The refusal tests fail against pre-fix code (negative control run). Fixes #3726 * docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS Cross-AI review of the fix diff (codex, pre-create) caught three shipped doc sites still instructing the now-refused bare invocation: the CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's two guard-override instructions (`--force` alone now refuses without --confirm). Localized CLI-TOOLS copies already lag the English synopsis (no --force/--dry-run either) and follow the translation pipeline, not this fix. * chore(#3726): set changeset fragment pr to 3774 * test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack Two CI reds from the --confirm gate, both this branch's own misses: - tests/qa/scenarios/milestone-rollover.json invoked `milestone complete 1.0 --force` as a JSON arg-array fixture — a caller shape the test sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never enumerated. Adds --confirm; the scenario's boundary-crossing contract is otherwise untouched. - complete-milestone.md's +420-byte --confirm note trips the emitted-attribution growth ratchet. Acknowledged as a #3726 append to the existing complete-milestone.md entry in 3409-unreachable-guard-arms.json (two ack sources may never name the same path, per that fragment's own precedent). Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed. * docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm Review Major 1: the truncated-window and unstarted-phase guard paragraphs still told the reader to "Pass `--force` to override", which now refuses (--force alone does not satisfy the confirmation gate), while the flag table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair so the file no longer contradicts itself. * docs(#3726): synopsis renders --confirm and --dry-run as alternatives Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as "a dry run still needs --confirm", the opposite of AC 3. Render the pair as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage docblock, and let the flag rows carry the rule. * test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack Rebase onto next (26 commits) surfaced three tests the gate now refuses: the #3685 write-flag contract pair in tests/milestone.test.cjs and the `milestone complete` boundary fixture in tests/state-contract.test.cjs all invoke the command bare. Each now passes --confirm (a mutating run is exactly what they assert on). The +420 byte complete-milestone.md growth ack rode on 3409-unreachable-guard-arms.json, which #3078 swept from next as fully spent — hence the modify/delete conflict. Re-filed under a fresh fragment named for this issue, never resurrecting the swept one. * test(#3726): pin the present-but-falsy arm of the confirmation gate Review Minor 1: the boundary triple covered absent and present but not present-but-falsy. The gate is an exact-token match, so --confirm=false and --confirm=0 refuse today — pinned (canonical + query forms, whole .planning/ tree byte-identical) so a future `=`-aware or prefix-matching parser cannot silently turn --confirm=false into a confirmed run of an irreversible command. * test(#3726): drop --confirm from dry-run-only invocations Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run invocations that never needed it, so each stopped standing as incidental proof that a preview needs no confirmation. Reverted to the pre-PR form; the dedicated AC-3 test carries the explicit assertion. * docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate REQ-I18N-02 (docs/features/internationalized-documentation.md) requires translations to stay synchronized with the English source. The four localized CLI-TOOLS.md guides still advertised a bare `milestone complete <version>`, which now exits 1. Render the English synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and `[--archive-quick]` flags the translations had also fallen behind on. * test(#3726): drop --confirm from the remaining preview-only invocations Round 2 reverted the --confirm appends on --dry-run-only invocations in tests/milestone.test.cjs, but four more sat in two files the sweep missed: tests/milestone-archive.test.cjs (three) and tests/milestone-window-single-owner.test.cjs (one). Each is a preview run whose whole purpose is to document that a preview mutates nothing, so `--dry-run ... --confirm` contradicted the semantics the test exists to pin. Dropping the token restores each as incidental proof that a preview needs no confirmation; the dedicated AC-3 test keeps the explicit assertion. No assertion added, relaxed, or removed — the change is four tokens. * chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer #3954 (ADR-3942) moved emitted-drift acknowledgments out of tests/emitted-drift-acks/ and into git commit trailers, and the fragment directory no longer exists on next. The reason this PR's fragment carried moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit; the fragment file is removed rather than resurrected. Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift. * fix(#3726): name --confirm in the version-required refusal The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke the command without args and the error lists what is required) stopped at `version required for milestone complete (e.g., v1.0)` — one required argument short. Discovering --confirm took a second round trip through the gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test that also asserts the version-less invocation leaves .planning/ untouched. * test(#3726): pin the milestone complete docs against a silent regression The changeset is `type: Fixed`, which the docs-required lint exempts, so nothing in CI would notice a later edit that reinstated the bare-`--force` override prose or dropped `--confirm` from the synopsis. Four tests in tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md and its four localized mirrors; the `--confirm` flag row; both guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md, by guard name (a substring match on each instruction's `--force --confirm` text); and — as an identity ratchet over the milestone-complete sections — every `--force` sentence or clause that lacks `--confirm`, so a new bare instruction in its own sentence or clause fails whatever its wording. Named residual: a bare instruction spliced into the same clause as a compliant one coalesces with it and passes the ratchet; the by-name pins are what keep the four known instructions from losing the pairing that way. The file is registered in scripts/docs-guard-registry.cjs so the pin runs on the PR that changes those docs, not only after merge. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
370cfc6680 |
enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local> |
||
|
|
519ac23ebb |
fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first) * docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts; its exit-2 block IS the contract — a post-tool hook cannot prevent a commit) and gsd-session-state.sh is registered SessionStart (session orientation, not post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the three INVENTORY locales; the issue asked for a neighbouring-row scan, which is how the session-state row was found. All other rows in the four tables verify against the surface. * fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins Adversarial review found the same two wrong rows shipped in five more files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now covers all ten shipped tables. The parity parser unioned only the Kimi mirror list, silently exempting agent-isolation-guard (registered via the dynamic preToolEvent push): probes are now parsed too, with bare hook names resolved against hooks/ ground truth and dynamic event variables resolved to their canonical (non-Gemini) events; an exact-set pin replaces the loose size guard. allow-test-rule marker carries the issue ref; unverified-ceiling 280→281 (audited: the new marker is legitimate — the suite reads product docs whose text is the contract). * fix(#3839): register the hook-table parity suite in the docs-guard lane The new suite reads ten docs/ paths, so lint-docs-guard-registration requires it in the docs-guard registry — the first GREEN bench run caught the omission (the RED run's docs-guard failures were the same signal, previously misread as marker fallout). * chore(#3839): changeset fragment (pr number backfilled after PR creation) * chore(#3839): backfill changeset PR number (4041) --------- Co-authored-by: sim <sim@local> |
||
|
|
80de48c319 |
enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs
Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.
local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.
Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.
The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.
Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.
Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.
Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.
Verification runs on the remote runner.
Refs #3914
* docs(#3914): the epic added three guards, it did not remove one
The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.
An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.
Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.
Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.
Refs #3914
* fix(#3914): make the successor actually supersede before retiring the predecessor
An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:
const exit = 'exit';
process[exit](1);
n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.
This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.
The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.
The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.
The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the two exit rules are complementary — keep both
Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.
I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:
function f(exit) { process[exit](1); } 0 vs 1
let exit='exit'; exit='exit'; process[exit]() 0 vs 1
const { exit } = ...; process[exit](1) 0 vs 1
plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.
The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.
So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.
What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.
The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.
Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the exemption term was a no-op — the net is +4
Fourth correction to this ledger, and a fourth error of the same kind.
Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in
|
||
|
|
ac7587287b |
fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true #3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue was reopened rather than noted. The section now states three things: every field is single-valued, the section is overwritten rather than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line appended in good faith is silently ignored rather than winning. Progress history belongs in ## Performance Metrics, two headings down, and the text now points there. The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before being written rather than inferred from the issue title: stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase") -> "1 of 5 (First)" The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags im with NO g, so String.match returns the first hit. Writing "first wins" without running it would have repeated the exact error I had to retract twice in this epic already. A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines and asserts it resolves independently - negative space, because a reader returning the first line of the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a last-match variant returns "9 of 9 (Appended later)" and T1 reds. No assertion checks that the document contains a sentence. That is what local/no-source-grep exists to stop, and it would pin wording that is allowed to improve. The point of the test is that if that regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly becoming a lie with nothing to notice. Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four locales. The constraint is structural, not stylistic - confirmed by running that comparison after the edit. The four locale copies are translated rather than left stale. They are not gate-enforced for prose, so "nothing fails" was available and is not the same as correct: leaving four documents asserting something the English one now contradicts is a correctness problem. Code spans and the anchor link stay untranslated - they name real tokens. The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been false the fix would have belonged in the generator, and a hand edit would have been silently reverted. One real gate failure fixed inline rather than reported: the new test's comments referenced docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and lint-docs-guard-registration failed lint:ci correctly. Registered. Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope and stays unclaimed. Closes #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): the rule I documented was false — replace it with the measured one An isolated review returned two blockers. Both mine, and the first is the worse kind: I wrote a falsifiable rule into a reference page and got it wrong. 1. "A duplicate resolves to the FIRST occurrence" is FALSE. stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence beats document order. Measured against the built reader, all intra-section: Phase: A (plain) / **Phase:** B (bold, later) -> B LATER WINS Phase: A (indented) / Phase: B (plain, later) -> B LATER WINS Phase: A (plain) / | Phase | T (table) | -> A first wins My original verification tested plain-versus-plain, saw first-wins, and generalized to all forms. Measuring one case and claiming the general rule is the same error I have had to retract twice already in this epic. It is also worse than silence. The sentence told authors an appended line is safely ignored; a bold line appended "for emphasis" silently overrides the original. Someone trusting the doc would have corrupted their own state file. And #3812 never asked for a resolution rule - it asked for single-valued, overwrite-not-append, and where history goes. The rule was my unrequested addition. Replaced with the measured truth: resolution is by FORM (bold anywhere, then plain at line-start, then table row), and only WITHIN the winning form does the first occurrence win. Both consequences stated plainly - a higher-ranked form wins regardless of position, and an indented `Phase:` is invisible to the plain form. All five claims in the new paragraph verified by execution before being written, including the two I had wrong. 2. The tests tested the wrong case and passed for the wrong reason. T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section case, which #2956 already fixed by scoping. #3812 says verbatim that #2956 "fixed the inter-section case and never addressed intra-section duplication", so the case the new prose describes was untested, and the fixtures passed because of section scoping rather than field resolution. They also called bare stateExtractField rather than the production chain, T2 could not discriminate first from last at all, and no fixture mixed forms - which is precisely why the false claim survived to review. Rewritten as four rows, all intra-section, all through the real stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first wins within a form), plain-then-bold (the bold LATER value wins - the row whose absence let the false claim ship), indented-then-plain (indented invisible), and sibling-field independence. Each proven to fail against a reader that disagrees. 3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the anchor their own heading generates. ja-JP/ko-KR kept the English heading, so theirs already resolved. 4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section designed to grow" with a bare demonstrative whose nearest referent read as Current Position - saying the opposite of the point. Rewritten so the clause attaches unambiguously to `## Performance Metrics`. 5. Cross-locale drift, flagged by the implementing agent rather than by me: after fixing EN, the four locales still stated the OLD false rule. Four documents asserting something measured to be wrong is worse than four saying nothing. All four now carry a faithful translation of the corrected paragraph, with code spans, each file's own anchor, and the ja/ko referent fix preserved. Verified: all five claims executed against the built reader; every rewritten test row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to date, so the edits stay outside the generated marker regions; locale heading parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): second false rule on the same page — scope the ranking to the section A second isolated review found a second false falsifiable claim, and the failure mode is the same one twice in a row: attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT Both times the claim covered a wider surface than what was actually executed. The fix each time was not a better sentence, it was executing the surface the sentence describes. BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false. ## Current Position / Phase: 1 of 5 + ## Archive / **Phase:** 88 bare stateExtractField(whole doc) -> "88 (other section)" PRODUCTION (slice then extract) -> "1 of 5 (in section)" #2956's section slice means production never hands another section to the matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not seen. The ranking is real but scoped: it applies WITHIN `## Current Position`. I verified against the bare function and wrote a claim about the system. Every existing test placed its bold line inside the section, which is exactly why nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive` and asserts production returns the in-section plain value, with the unscoped reader asserted to DISAGREE so the row proves the scoping rather than assuming it. BLOCKER — the changeset still shipped the ORIGINAL retracted claim. I corrected the page and left the release note saying "resolves to the first occurrence ... a second entry added in good faith is silently ignored". The note contradicted the page it announces, and the release note is what most people actually read. Rewritten to the corrected rule. MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the file", which is the vacuous direction; the surprising case, and the one the very next clause illustrates with an APPENDED bold line, is "even if it comes LAST". All four locales reproduced the inversion faithfully, so it was an EN-source defect rather than translation drift. Two sharp edges now named, both measured: a bold `**Phase:**` followed only by trailing spaces resolves to an EMPTY STRING and does not fall through to a valid plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form ranking has to say where the ranking bites. Also fixed: all five files labelled the link `## Performance Metrics` while the heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only the label's level was wrong. Every clause in the final paragraph re-verified through the PRODUCTION chain (stateCurrentPositionSlice -> stateExtractField), clause by clause, before being written: bold in another section does not win; bold in frontmatter does not win; bold appended last does win; first wins within one form; trailing-space bold yields empty. All four locales carry the same corrected rule. gen-state-md-docs --check reports all 6 targets up to date; heading counts unchanged at 20/20 across all five files, so locale heading-parity is untouched; build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3812): backfill changeset pr number Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ab69b9ce56 |
enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from |
||
|
|
dd4f179672 |
feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
12f9d1d9a0 |
enhance(#3913): docs, and the guards come down (#3994)
ADR-3889 terminal phase. Generated docs/reference/exit-codes.md from the exit-code declaration with a --check drift arm; deleted the inert soft-error-exit-zero oracle; promoted untyped-success from SMELL to VIOLATION so it can fail a build; pruned all 5 smell-baseline entries. Fixed inline: two mis-scoped oracles (routing-validity, value-hygiene), a second source behind the band table, unescaped declaration strings reaching Markdown, and a pre-existing Windows 8.3 short-name path-comparison defect. Guard ledger corrected from a claimed net -4 to a measured net -1. Closes #3913 |
||
|
|
b351c83e03 |
docs(adr): ADR-3646 — per-task external-tracker content-resolution seam (#3991)
Resolves the four conditions on the approved-feature verdict for #3646: new execute:task granularity tier below wave, hard-halt enforced via a code-side resolver seam (Lens B) rather than prose dispatch (since #3647's dispatch-reliability defect is still open), registration/validation requirements for loop-hook-dispatch.md + capability-validator.cjs, and autonomous-mode behavior via a new non-gate kind. Closes #3969 Co-authored-by: sim <sim@local> |
||
|
|
3a6c0412a9 |
enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) Extends ADR-1703's portability rule catalog with a second production-runtime rule: it flags an exact-case read of a Windows case-varying environment variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any receiver that is not process.env itself, matched via an env-shaped-receiver check to avoid colliding with ordinary `.path`-named properties elsewhere in the tree. Exports the seam's private `_envGet` as `envGet` so the rule's remediation message names a real helper, and fixes the one pre-existing violation the tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read). Closes #3624 * fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case Review findings from the code-review + isolated-adversarial passes: - extractStaticName only matched non-computed Identifier keys, so a destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8 acceptance case) silently evaded the rule. Widened to accept a Literal key regardless of computed, which is safe for MemberExpression too (its non-computed property is always an Identifier by grammar). - Added the missing RuleTester valid case for "a case-insensitive accessor call" (envGet(env, 'PATH')) from the issue's Done-when checklist. * docs: backfill changeset PR number for #3624 (PR #3976) --------- Co-authored-by: sim <sim@local> |
||
|
|
bb4f3073c0 |
docs(#3984): give every §8 rule its delivered status and its measured executor (#3986)
Surfaced by an /adr-phase-coverage audit of epic #3473 after its last sub-issue merged. §8 is the epic's declared source of truth - "where this section and the code disagree, the code is the defect" - and line 136 defines a per-rule status vocabulary. Not one of the nine rules was ever advanced past its pre-implementation status. The word Enforced appeared exactly ONCE in the document: in the sentence that defines it. Three rules read "Required - phase unassigned" for work that Phases 5, 7 and 8 demonstrably delivered, so a reader auditing coverage today would conclude three of the nine contract rules had no owner. That is the orphan shape this epic exists to eliminate, sitting in the epic's own contract, and it is the same defect class as B6's guard ledger: a status column is a claim, and it was false. Flipping all nine to Enforced would have repeated the mistake inverted. Decision 2 says a declared policy with no executor is a loud failure, so Enforced asserts an executor EXISTS - and measurement found the nine are not equally enforced. The vocabulary is now three values: Enforced a standing guard wired into lint:ci or CI fails the build; a new wrong call site cannot land Enforced (structural) the wrong call site is unrepresentable Shipped - test-covered delivered and regression-tested, but NO standing guard; a regression in covered code is caught, a new wrong call site elsewhere is not Measured per rule, with the executor cited: 8.1 Enforced lint-vendored-deps + no-external-require-in-bin 8.2 Enforced (partial) phase-enumeration-drift covers the sentinel arm; the #3357 resolver arm is test-covered only 8.3 Shipped NO guard for slug or marker re-derivation 8.4 Shipped structural for argv; no guard for the general rule 8.5 Shipped the delivering change added zero guard scripts 8.6 Enforced (structural) the transaction type; state-write-path-drift 8.7 Shipped structural via reconcileReportedFields; no guard 8.8 Enforced gen-state-md-docs --check in lint:generated-sync 8.9 Enforced (prospective) lint-fix-has-regression-tests fires on NEW fix commits; it does not assert the 19 children are covered. 17 of 19 have a test citing their number; #3364 and #3812 have none - a text match, not proof the behavior is uncovered, but not evidence it is covered either The third status value is not a euphemism for done. It names precisely where this epic's own thesis - make the wrong call site unrepresentable - is NOT yet achieved. §8.3 and §8.5 are the thinnest: nothing today stops a twelfth inline slug copy or a second silent swallow. Also corrects a claim this ADR made three days ago and #3977 has since overtaken. The ledger amendment concluded #3426/#3239 "need new detectors". The measurement behind that was right - the glob widening structurally could not reach them - but the prescription was not the best available: #3977 rerouted the test's parsers onto the ADR-2143 seam, so there is no ad-hoc parse left for a detector to find. Removing the violation beat teaching the rule to see it. Both issues are CLOSED; the roster row and the amendment now say so, and say why the earlier conclusion should not be built on. Closes #3984 Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d24e22b156 |
enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1 ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what remained was the declaration — and the pin that makes it invisible today. The census corrected two documented figures before any code changed. ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier note of mine claiming 23 was wrong and is corrected). And output({error}) is **64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2. That matters because this phase's criterion demands the pin be asserted over the enumerated population rather than sampled; asserting over a stale 60 would leave four sites unpinned while claiming full coverage, which is the shape of failure this epic exists to remove. The issue does not state the fact that shapes the design: output() never touches the exit code. Confirmed by reading it — it writes fd 1 and returns. So a declared outcome for those 64 sites had nowhere to be READ. The mapping was never the work; wiring somewhere for the declaration to land was. The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol cells, each because the module is emitted to three locations and a module-level `let` would let instances disagree, and runMain already maps a code returned by main(). A third cell inherits that solution. output() records DEGRADED for any {error} payload — key-order agnostic, which is exactly why the "42 sites" figure undercounts — and runMain projects the cell only when main() returns nothing, so an explicit return still wins. error() maps its reason through a table over the closed 25-member enum, leaving all 278 call sites untouched; 226 of them pass no reason at all. The version gate lives in error(), NOT in projectOutcome: registered names are version-invariant there, so mapping a reason straight through would make USAGE project to 64 under v1 and break the pin on its first line. projectOutcome is left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included. Proven rather than asserted. v1 is byte-identical across three real CLI paths — config-get plain, config-get --json-errors, and an output({error}) path — matching exit code and exact bytes against the pre-change build. Under GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND -> NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason, because without it a mapping where everything projects to 1 under both versions would satisfy every other assertion and the declaration would be theatre. A1 iterates all 25 enum members and A3 asserts over the measured 64-site population, so a 26th reason or a 65th site fails until it is given a mapping — the drift guard this phase needs, given ADR-2980's own count had drifted +4 unnoticed. Verification runs on the remote runner. Refs #3912 * fix(#3912): the outcome cell must never lower an exit code The remote run caught a fail-open that this phase introduced, in the phase whose entire purpose is removing fail-opens. `state validate --strict` on a missing STATE.md exited **0** where it must exit 1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set non-zero by the command was clobbered down to success. Confirmed live against a fixture, before and after. This refutes a review conclusion recorded earlier in this phase, that the cell was "fail-closed and can never mask a failure as success". It could, and did. Recording that plainly so the assumption is not repeated: the cell's danger was never only that it might add a failure — it was that projecting it unconditionally overwrites whatever decision came before. Projection is now guarded: it may set a code only when none is set, and an already-non-zero exit code always wins. The full precedence — explicit `main()` return, then an existing non-zero exitCode, then the declared outcome — is written at the projection site. A regression test drives a void return with a pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix build. The second failure was my test encoding the wrong contract, not a code defect. It asserted `output({found:false, error: undefined})` records DEGRADED because the KEY is present. `JSON.stringify` drops undefined, so the payload the user receives is `{"found":false}` — carrying no error at all, and calling that degraded would hand back exit 80 under v2 for output that reads as clean. The discriminator is a serializable error VALUE, not key presence. The test now pins `{error: undefined}` as explicitly NOT degraded, and the design doc's wording is tightened to match. Verification runs on the remote runner. Refs #3912 * docs(#3912): the versioned exit contract, and a flag defect the docs found Diataxis pass for Phase 8, plus a real fix that only surfaced because writing the how-to meant running its own examples. The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned projection this phase provides, so it gets an amendment naming #3912 / ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects DEGRADED to 80. The amendment also records the count drift rather than restating a stale figure — the ADR ratified 60 output({error}) sites in 9 modules; the AST-measured population is 64 across the same 9 (frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the enumerated 64. json-errors.md gains the outcome-declaration reference, including the precedence order a review pass got wrong and the suite refuted: an explicit main() return, then an already-set non-zero process.exitCode, then the declared outcome. Projection may only ever set a code, never lower one. A how-to is owed here and is written, not skipped. Under v1 nothing changes, so the audience is an operator opting into v2 and needing to know what the codes mean for a CI gate — a migration, which is how-to shaped. It covers turning v2 on, the code table, why 80 is "ran and reported a condition" rather than a crash, and how to split a gate that treats any non-zero as fatal. No tutorial: there is no new entry point to learn, and under the default contract a reader would be walked through observing nothing. The defect. Running the how-to's own Step 1 example returned $ gsd-tools --exit-contract=v2 state validate --strict Error: Unknown command: --exit-contract=v2 (exit 64) while the same flag trailing the subcommand worked and exited 80. The flag half-worked, by argv position. resolveContractVersion scans argv non-destructively, so the token survived into the dispatcher, which treats argv[2] as the command name. --json-errors had already solved precisely this at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv splice must happen here too, otherwise the dispatcher below sees --json-errors as an unknown command." The later flag never got the same treatment. Fixed rather than documented around: the version is resolved first — which memoizes the cell and makes an invalid value throw early — and then every --exit-contract= token is spliced out of the dispatcher's argv copy. --exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The regression test pins leading position, trailing position, agreement between the two, and a loud failure on v3 rather than a silent fall back to v1. Neither review engine would have caught this: the defect is invisible in the diff, because the diff does not touch argv handling. It surfaced only from running the documentation's own example. Writing a how-to is an execution pass. Verification runs on the remote runner. Refs #3912 * fix(#3912): the flag splice has to run before the run-with-timeout return An isolated review of the previous commit found that the fix did not deliver what it claimed, and that two of its own tests were weak. All three findings reproduced by execution before any change was made. The fix was placed below a return. main() intercepts `run-with-timeout` at gsd-tools.cjs:4436 and returns from there — above both the --json-errors block and the --exit-contract splice added in the previous commit. So the flag still died in leading position for that one command: $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..." Error: Unknown command: run-with-timeout (exit 64, child never ran) The previous commit message and the test's describe-block both claimed position-independence unconditionally. That was an overclaim, not a gap left open, and it is the part worth naming: the fix was verified by hand on the commands I happened to think of, and `run-with-timeout` returns before the code I was verifying. Both global-flag blocks now run above the interception, with a comment naming it so a later edit cannot slide them back down. Moving --json-errors up fixes the identical pre-existing bug for that flag, verified failing beforehand (exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect, same block, and a known-broken twin next to a fixed one is not a resting state. Two tests were not pulling their weight. The invalid-value test was vacuous — it passed against the pre-fix build, because `--exit-contract=v3` already exited 1 there and already printed the resolve error lazily through error() -> getContractVersion. Both its assertions held before the fix, so it pinned nothing. The real discriminator is that the pre-fix build emits BOTH "Unknown command: --exit-contract=v3" and the resolve error, while the fixed build emits only the latter; the test now asserts that absence. The leading-position and leading==trailing tests asserted proxies — "not 64", "no Unknown command", "the two agree" — none of which pin a value, and all of which would survive both positions being identically broken. With a .planning directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0 under v1 in both positions. Those numbers are pinned now. The multi-token case the descending splice loop exists for is covered too, and run-with-timeout has regression tests for both flags. The lesson is narrower than "test more". Hand-verifying the production behavior does not verify that the test would have caught its absence. The pre-fix binary has to be run against the test's own assertions. Investigated and deliberately not changed: splicing before --cwd parsing degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd: <path>", but that is pre-existing — verified on the pre-fix build via --json-errors, which already did it. This change joins the pattern rather than creating it, and both forms exit 64 on malformed input either way. Verification runs on the remote runner. Refs #3912 * chore(#3912): backfill changeset pr numbers to 3983 * test(#3912): pin the reason-table invariant as set equality, not a count A graph-backed review flagged the unchecked lookup in expectedErrorCode3912. Investigated by execution: the drift guard DOES hold — for an unmapped reason under v2 the production error() yields 1 while the table yields undefined, so the assertion fails. Not a correctness defect, and deliberately NOT made tolerant, since a tolerant lookup would destroy the guard. Two real problems remained. The guard asserted the wrong invariant: it counted the TABLE's keys at 25 rather than checking they match the ENUM's values, so a renamed member keeps the count at 25 and slips past, and a 26th member leaves the table at 25 and slips past too. Both were then caught only indirectly, by an undefined mismatch producing 'must exit undefined'. It is now a sorted set equality, so the failure names the specific missing or extra reason. And the comment above it described a '?? FAIL' fallback that does not exist anywhere in the function. It now states what the code actually does, verified by running it rather than by reading it. Refs #3912 --------- Co-authored-by: sim <sim@local> |
||
|
|
355c943b08 |
enhance(#3626): make CONTEXT.md seam claims checkable via a SEAM.*.enforced-by gate (#3975)
* feat(#3626): make CONTEXT.md seam claims checkable via SEAM.*.enforced-by gate Adds SEAM.<id>.owns / SEAM.<id>.enforced-by=lint-rule:<name>|test:<path> predicates to CONTEXT.md, generalizing the existing WORKTREE.SEAM.* shape, plus scripts/lint-seam-enforcement.cjs (wired into lint:ci) which fails when a declared single-owner seam names no existing, registered enforcement mechanism. Backs all six current module-level single-seam/ single-canonical-owner claims found in CONTEXT.md, including the Shell Command Projection Module's Windows-binary-resolution claim via #3619's local/no-private-binary-resolution rule. Scope is resolves-only per maintainer decision: the gate proves an enforcement pointer exists and is registered, not that its surface covers every file the seam claims. See docs/adr/3626-context-md-seam-claim-gate.md. Closes #3626 * fix(#3626): back the Package Identity Module's seam claim too Isolated adversarial review caught a miss in the "no grandfather list" sweep: the Package Identity Module also declares itself "Single seam owning GSD's published-package coordinates" and already names its real enforcement (scripts/lint-package-identity-drift.cjs). Backs it with SEAM.package-identity.owns/enforced-by=test:tests/package-identity.test.cjs, bringing the total to 7 backed seams. Also makes explicit, in the design doc and ADR, that function-level "single owner" sentences inside already-covered modules (STATE.md Document Module, etc.) are deliberately out of scope — a seam claim is about a module's boundary, not every function inside it. * docs(#3626): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
d98b55562c |
enhance(#3910): the raw terminator is banned by construction (#3980)
* enhance(#3910): move the last src/ terminators onto the seam Phase 6 bans the raw terminator by construction, which it cannot do while violations stand. A census found 12 sites the rule would flag; nine of the ten unsanctioned ones were owned by no phase of the epic at all — a coverage hole in the decomposition, since P0-P2 are infra, P3 the gate modules, P4 the scanners, P5 the fragments, P7 the hooks, P8 io.cts, and P6 itself only adds the rule. `src/**/*.cts` now holds exactly 2 raw exits, both inside `terminateNow`, the single sanctioned site. `io.cts`'s `error()` is the interesting one. It was first called substantive on "dozens of callers, contract risk" — asserted, not measured, and the measurement refuted it: 289 call sites, zero inside a try whose catch would swallow a throw. The real obstacle was structural instead: `terminateNow` cannot emit exit 1, because ADR-3889 §1 makes 0 and 1 unallocatable and `nameForExitCode(1)` throws. So the only route is `ExitError` under `runMain`, which sets exitCode and writes stderr only when the error carries a user message — keeping the existing stderr write and throwing a message-less ExitError is observably identical. That census was still too narrow, and running the CLI proved it. It asked whether the CALL sits in a try/catch; the two regressions that surfaced were interceptors elsewhere on the stack: - `command-routing-hub.cts`'s `dispatch()` swallowed the ExitError into a HandlerFailure, so the caller emitted a duplicated, wrong stderr line on every Hub-routed path. It now rethrows ExitError explicitly — the same shape `gsd-tools.cjs` already used at two dispatch sites, so this follows an established idiom rather than inventing one. - the profile-pipeline router's deliberately un-awaited `.catch(e => error(...))` turned an ExitError rejection into an uncaught exception; it now mirrors runMain's handling. `edge-probe` and `ui-consideration-probe` gained `runMain` wrappers because probe-core's new throwing default would otherwise have escaped them. A follow-up sweep of every dispatcher — 19 command routers, the Hub, the gsd-tools dispatch seams — found no further swallowing catch. The admitted bound: ~1260 non-rethrowing catches repo-wide were scanned structurally but not individually classified. Both real regressions were found by execution, not by reading, so the suite is the detector that matters here. `gsd-tools.cjs:253` stays a raw exit deliberately: it is the ensureRuntimeBuild bootstrap, which runs before cli-exit is required, so the seam does not yet exist. It needs a second allowlist entry, which means #3910's "single allowlist entry" criterion is unachievable as written. Verification runs on the remote runner. Refs #3910 * enhance(#3910): ban the raw terminator by construction Adds local/require-registered-exit and registers it on all four globs: src/**/*.cts, scripts/**/*.cjs, hooks/**/*.js, gsd-core/bin/**/*.cjs. Registering on the .cts glob is load-bearing, not redundant — the emitted .cjs mirrors are globally eslint-ignored, so a rule registered only on the emitted globs is blind to the sources. That is the #3496 lesson, and it is how the previous guard became invisible: n/no-process-exit was 'error' in one block yet fired zero times on all three surfaces that mattered. The dead n/no-process-exit: 'off' block for hooks is deleted in the same PR. Phase 7 migrated every hook, so the exemption now protects nothing. Two allowlist entries, not the one #3910 anticipated. terminateNow's body is detected STRUCTURALLY — a process.exit lexically inside a function of that name — rather than by a path and line number that rots. The second is gsd-tools.cjs's ensureRuntimeBuild bootstrap, an inline disable with its reason at the call site: it runs before ./lib/cli-exit.cjs is required, so the seam does not exist yet and no migration is possible. #3910's 'single allowlist entry' criterion is therefore unachievable as written, and is amended with the measurement rather than quietly missed. The rule is proven able to FAIL, per glob: four positive controls, one for each registered glob. A guard that cannot be shown to fire is not a guard. Four matching negative controls pin process.exitCode as never-flagged — conflating it with process.exit is what inflated this epic's original census 2x. An allowlist case and a near-miss (same shape, different function name) fix the structural detection in place. Verification runs on the remote runner. Refs #3910 * fix(#3910): stop the detached catch from throwing, and scope the allowlist Review findings, one of them a regression the previous fix introduced. _handlePipelineRejection called error() from inside a DETACHED .catch(). error() now throws, so that throw became an unhandled promise rejection — and on Node >=15 with --unhandled-rejections=throw, Node dumps a raw stack trace with absolute paths on top of the clean Error: line. That was impossible before this branch, because process.exit(1) terminated synchronously before any rejection machinery could observe it. The handler now writes byte-identical stderr itself, in both plain and --json-errors form, and sets exitCode in place. This was the THIRD interceptor found, and like the first two it surfaced by running the CLI rather than by reading code. The rule's terminateNow allowlist had no path constraint, so any function anywhere named terminateNow across all four globs inherited it. It now requires the structural nesting check AND a cli-exit.cts basename — still no line numbers to rot. The four per-glob positive controls only varied a filename inside RuleTester, which never resolves eslint.config.mjs. Since the rule is filename-agnostic, all four exercised identical logic and none proved the rule was WIRED — this epic's own failure mode. A registration test now asserts the rule resolves for a real path in each glob, and it is proven able to fail: removing one glob's registration flips the resolved value from [2] to undefined. Three evasions the rule cannot catch (computed member, aliasing, .call/.apply) are documented in its header and pinned by tests, labelled as known limits rather than endorsed, so a future change that starts catching them is a deliberate diff. Refs #3910 * docs(#3910): document the raw-terminator ban Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it, not hand-edited). How-To: docs/how-to/resolve-a-raw-terminator-finding.md, indexed from docs/README.md — a contributor whose code trips the rule picks among three replacements by surface (runMain/ExitError for a CLI path, terminateNow for a hook, process.exitCode where the process should drain), and needs to know why process.exitCode is correct and never flagged, since conflating the two is what inflated this epic's original census 2x. The page also names the three patterns the rule cannot catch and says plainly that using one to dodge it is a review finding, not a fix — documenting them without that sentence would read as a sanctioned workaround. docs/INVENTORY.md deliberately untouched: eslint-rules/ is not a tracked family in the manifest (verified — a regen produced a zero diff), so a hand-written row would desync the table from the family it claims to belong to. Refs #3910 * fix(#3910): a catch that sniffs the message swallows an ExitError The remote run returned 41 failures, and one of them was a live production regression rather than a test artifact. `cmdMilestoneComplete`'s unstarted-phase guard re-threw only when `e.message.startsWith('Cannot mark milestone complete:')`. `error()` used to `process.exit(1)`, uncatchable, so the guard always fired. It now throws an ExitError carrying no message, the string test fails, and the ExitError was silently swallowed — the guard stopped blocking milestone completion entirely. Proven against the real CLI: pre-fix, a milestone with an unstarted phase archived at exit 0; post-fix it is blocked at exit 1 with the intended message. That is a guard that silently stopped guarding, which is this epic's thesis appearing inside the phase meant to enforce it. Worth stating plainly: an earlier census DID examine this site, saw a `throw e`, and classified it as rethrowing. It was wrong — the rethrow is conditional, and a conditional rethrow on an inspected message is indistinguishable from an unconditional one unless you read the predicate. So the class was swept rather than patched where it was tripped over. An AST census of every CatchClause across src/, gsd-core/bin/ and scripts/ found 38 conditional rethrows. Two more had the same defect and are fixed the same way: `config.cts`'s `'No config.json'` sniff and `gsd-tools.cjs`'s `e.name === 'WindowsError'`. The remaining 25 are provably unreachable — every one wraps a bare fs, YAML, manifest-require or git-exec primitive that cannot throw ExitError — and two were scanner false positives, both explained. Each fix is an unconditional `instanceof ExitError` rethrow placed BEFORE any inspection, matching the idiom command-routing-hub and gsd-tools already used. Residual bound, stated rather than implied: zero known-reachable unfixed sites, contingent only on error() never later being called inside one of those 25 primitive try blocks. The remaining failures were harness artifacts, and the harnesses were corrected to the new contract rather than the assertions weakened. Tests that mocked `process.exit` to observe termination now catch ExitError and assert its code; tests parsing stderr as a single JSON object still assert exactly that, with their ad-hoc `node -e` scripts wrapped in runMain so it is true. milestone and phase-resolution-parity needed no test change — they were correctly written against the real bug and are what caught it. Verification runs on the remote runner. Refs #3910 * chore(#3910): backfill the changeset PR number Also reframes the fragment to lead with the user-visible change — the milestone guard blocking again — rather than the narrowest of the three fixes. Refs #3910 --------- Co-authored-by: sim <sim@local> |
||
|
|
15af0f5536 |
enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ea5efc151 |
enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local> |
||
|
|
fa41bfec5c |
enhance(#3942): the emitted-drift ack is PR-lifetime data — move it to a commit trailer (#3954)
* test(#3942): failing-first suite for the emitted-drift ack commit trailer Binds 37 input classes from the phase test matrix to the behavior ADR-3942 specifies, before any of it exists. Stubs return benign empty values rather than throwing, deliberately: several rows assert that something DOES throw (cap overflow, uncomputable commit range), and a throwing stub would turn those green for the wrong reason and destroy the red. The two rows that carry the design's load: - merge-base semantics. The range is $(git merge-base base HEAD)..HEAD, not base..HEAD, because changedPaths comes from `git diff base...HEAD` (three dot). Two-dot would let the ack set and the change set disagree about which commits are this PR's. The fixture forks a topic branch, puts a trailer on each side, and asserts only the topic-side trailer is in range. - fail-closed on an uncomputable range. With fragments a depth-1 checkout passes VACUOUSLY, every fragment reading as brand-new. With trailers the range cannot be computed at all, and returning an empty set would silently disarm the gate, so it must throw. The fixture builds a genuine shallow clone rather than simulating one. Also covers the self-inflicted case: this change's own documentation quotes the trailer syntax, so an example landing at the end of a commit message would arm a live acknowledgment keyed on the literal placeholder text. Keys carrying angle brackets or whitespace are rejected. Authored per the phase artifacts 40-design.md and 50-test-matrix.md. Not yet run on the remote runner — this commit exists to be tested. Refs #3942 * chore(#3942): move the emitted-drift ack to a commit trailer Implements ADR-3942, superseding ADR-2719 section 3 and its #2789 amendment. Sections 1, 2 and 4-7 are retained: the conservation law is unchanged, only the storage of its escape hatch moved off the working tree. An acknowledgment explains one PR's ripple, and the moment that PR merges the ripple is in the base, so it can never clear anything again. It was stored in permanent shared state anyway, and every consequence of that mismatch had to be built and then maintained. The chain is #2789 -> #2914 -> #3078 -> #3842 -> #3823 -> #3875, each fix generating the next defect, ending in a scheduled sweeper whose own first PR could not merge itself. Added parseAckTrailers + renderAckTrailer (pure) and readAckTrailers (IO shell), reading Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: trailers over the merge-base range. tests/emitted-ack-trailer.test.cjs, 37 cases, written failing-first and confirmed red before any of this existed. Changed diffEmitted takes two structurally distinct key-space maps instead of one shared paths map. That closes a latent defect: the spaces were separated by convention only, so a growth key satisfied a hash lookup by naming coincidence. staleAcks now reports which space a key was declared in. REMEDIATION teaches the trailer, per space, with its example rendered through renderAckTrailer so the taught grammar cannot drift from what the parser accepts. Removed the sweep workflow, the guard-no-ack-on-next job, the standalone linter and its lint:ci entry, the fragment directory and its three spent fragments, the legacy single-file union, and the baseAck/spentAcks mechanism -- spentness is now structural, not computed. Two range properties carry the design and are pinned by tests rather than asserted: the range is merge-base scoped, matching git diff base...HEAD, so an already-merged trailer is out of range by construction; and an uncomputable range throws instead of reading as zero acknowledgments, which is the inverse of the fragment guard's vacuous pass. Three deliberate observable changes, each disclosed in the changeset: the unread runtime field is gone, the legacy file is no longer read, and cross-space excusal no longer works. Ten open PRs carry fragments and will meet a modify/delete conflict. Measured before landing and accepted deliberately; the one-line migration is in the PR body. Verified: lint:ci exit 0. Remote runner to follow on this exact sha. Refs #3942 * fix(#3942): silent trailer collapse, lost coverage, and an unbounded cap Six findings from the orthogonal review round, all fixed in place. BLOCKER -- two trailers of the same name on one commit collapsed silently. readAckTrailers built `separator=1d` where git needs `separator=%x1d`: the `separator=` value inside a %(trailers:...) placeholder is itself a pretty-format string, so the bare hex was emitted as two literal characters and the split on \x1d never matched. Two same-name trailers therefore joined into one value with errors empty -- the first reason absorbing the second entry's key. Silent truncation, the exact class MAX_ACK_TRAILERS throws to prevent. Confirmed with od -c against real git output before and after. The failing-first matrix did not catch it because its "both spaces coexist" row uses Hash plus Growth -- different trailer NAMES -- so the value separator was never exercised. Two regression tests now cover same-name trailers directly. Coverage recovered: normalizeAckReason and INVISIBLE stayed on the live path via parseAckTrailers but lost every test when the old suite was pruned. Back under test against the current surface -- all six invisible codepoints individually, whitespace collapse, trim, CRLF, and two seeded fast-check properties. Dropping any single codepoint now fails. MAX_ACK_TRAILERS counted raw trailers before de-duplication, so one trailer carried forward across rebased commits counted once per commit and could throw on a legitimate branch. Now counts distinct entries; 100 identical repeats dedupe to one. diffEmitted validated baseline, current and changedPaths but not the new ackHash/ackGrowth, so a bad shape raised an unhandled TypeError instead of an error verdict -- the same defect shape this file documents for #2778. Docs: CONTRIBUTING and TESTING-SUITES were rewritten only in their first sections; the later passages still taught fragments, git rm and the deleted guard, contradicting the new text directly above them. Finished. Also extends lint-removed-but-needed to exempt docs/adr and docs/research. That gate fails on any docs mention of a file deleted in the same diff, which makes it impossible to document a deletion in the PR performing it -- an ADR's whole job is naming what it retired. Exemption is narrow and comes with a test proving the gate still fires for a live consumer elsewhere under docs/. A guard that cannot fail is worse than no guard. Maintainer-approved. CONTEXT.md names the retired machinery by role rather than by filename: its generated projection lands in docs/, which that gate does scan. Adds docs/how-to/acknowledge-emitted-drift.md. The required docs set is Reference and Explanation, so the task quadrant can be empty with every gate green -- and this change has a real multi-step journey, including the fragment migration ten open PRs now need. lint:ci exit 0. Refs #3942 * docs(#3942): correct the duplicate-trailer rule in CONTRIBUTING Both axes of the code review independently flagged the same passage, without seeing each other's output. It claimed two declarations of the same key are always "a hard, loudly-reported error, not a silent last-wins". That is only half true, and the missing half is the one contributors hit: identical declarations -- same key, same reason -- dedupe silently, because a trailer legitimately survives a rebase and reappears on every rebased commit. Failing there would red a branch for doing nothing wrong, which is exactly why the dedup exists. Only a same-key/different-reason pair errors, and that one is a genuine ambiguity about which explanation holds. As written, the paragraph told a contributor that a rebase-carried trailer breaks the gate -- the opposite of the behavior. CONTEXT.md's parallel entry already stated it correctly; this brings CONTRIBUTING into line. Doc-only, root-level markdown. Refs #3942 * chore(#3942): backfill changeset PR number to 3954 --------- Co-authored-by: sim <sim@local> |
||
|
|
03b7125293 |
enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks Binds the four fabrication sites found by executing the surfaces (ADR-3889 failure class (c)), each with a positive control so an over-firing fix goes red: - the blocking api-coverage.verify-pre gate certifying "no external-API integration" from a zero-byte phase scope - the assumption-delta query route scanning an unresolvable phase section as the empty string and reporting it as an examined negative - both capability fragments' probe fallbacks, which append a fabricated verdict rather than replacing, and fire on the legitimate exit-1 negative Verification runs on the remote runner. Refs #3909 * enhance(#3909): a probe that could not run no longer asserts a verdict ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a confident negative; each now reports what it could not establish. - check api-coverage.verify-pre: a phase with no plan body and no roadmap section ran detection over zero bytes and PASSED the blocking seal gate, certifying "no external-API integration" from input it never read. It now holds with scope_unavailable. The discriminator is bytes examined, never signals found, so a phase whose plans are real and simply carry no API vocabulary passes exactly as before. - query assumption-delta scan: an unresolvable phase section was scanned as the empty string and reported as an examined negative. It now returns {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded result in the payload, leaving the gsd-tools exit projection to P8. - both capability fragments: `|| echo '{"detected":false}'` appended rather than replaced, and fired on the legitimate exit-1 negative, so a correct answer and an honest skip both arrived as two concatenated objects. They now keep the probe's own payload and manufacture only an explicit probe_unavailable skip when the probe produced nothing at all. Every registered outcome is more restrictive on a blocking gate, so this can turn a false green red and never a red green. Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md seal-time outcome table, and a new how-to for the reason-code vocabulary. Verification runs on the remote runner. Closes #3909 * test(#3909): correct the stale unknown-phase assertion `unknown phase → detected:false, no throw (graceful)` scanned phase 999 against a two-phase roadmap and asserted `detected === false`. That pinned the fabrication as intended behavior: the phase does not exist, so the detector was handed the empty string and its "no core assumption changed" answer described nothing that was ever read. It now asserts the skipped-with-reason shape. The graceful-degradation contract the test was actually protecting — the query succeeds and does not throw on an unknown phase — is unchanged. Found by code review, not by the author. Refs #3909 * docs(#3909): author the FEATURES entry in its generator source `docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the per-feature fragments in `docs/features/`. The API-coverage entry was edited in the generated file, so the next regeneration silently dropped it. The text now lives in `docs/features/api-coverage-gate.md` and `docs/FEATURES.md` is regenerated from it, leaving the shipped file byte-identical and its content actually derivable. Caught by `lint:generated-sync`. Refs #3909 * test(#3909): bind the skip to "not found", and pin the discriminator The first verification run went red on one case, and the case was wrong rather than the code. `getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a missing ROADMAP.md, but for a section whose body is whitespace-only it returns the heading line alone — which is not empty. So a body-less section WAS found, and reporting `detected:false` over its heading is a real negative, not a fabrication. The test had assumed the resolver yielded `''` there. Correcting the test rather than the resolver keeps `skipped` bound to the distinction the issue asks for — found versus not found — and avoids diverging `assumption-delta scan` from `roadmap.get-phase`, which the fragment documents as sharing one resolver. Also adds the seeded property the test matrix had promised: for any plan body, the scope read back is whitespace-only exactly when the body was. That pins the gate's discriminator to bytes examined, so it cannot quietly become "no signals found", across unicode whitespace and CRLF. `docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table — surfaced by the co-change gate, not by a lint failure. Refs #3909 * chore(#3909): backfill the changeset PR number Refs #3909 --------- Co-authored-by: sim <sim@local> |