f0bb0787c9f550c19f7ccfbaf2cfb4db98218db3
429 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0bb7525a62 |
fix(#2943): rename get-library-docs -> query-docs; correct the ctx7 fallback rationale (#2963)
* test(#2943): parity guard against the nonexistent get-library-docs tool Second context7 naming drift after #2017 (which guarded the plugin-marketplace PREFIX). #2017's guard only checks tools: frontmatter lines, not prose bodies — which is where the broken tool NAME (get-library-docs) lived. The context7 MCP server registers only resolve-library-id and query-docs; get-library-docs is a stale copy from upstream's own README. Scans the shipped prose surface (agents/, gsd-core/references|workflows/, commands/gsd/, skills/) and fails if any artifact instructs an agent to call mcp__context7__get-library-docs. Excludes tests/ (a fixture may use the name as a negative input) and CHANGELOG/RELEASE-NOTES-LEGACY (history). Fails-first: 4 offenders today (gsd-executor.md:29, research-documentation-lookup.md:5, discovery-phase.md:68 & :104). * fix(#2943): rename get-library-docs to query-docs and correct the ctx7 fallback rationale The context7 MCP server registers only resolve-library-id and query-docs (verified against upstream packages/mcp/src/index.ts); get-library-docs is a stale name copied from upstream's own README. Four shipped prose sites instructed agents to call a tool the server does not register, so every research path that loaded the canonical reference either errored, fell through to the ctx7 CLI branch, or fabricated a result. - research-documentation-lookup.md, gsd-executor.md, discovery-phase.md (x2): get-library-docs -> query-docs, params context7CompatibleLibraryId/topic -> libraryId/query (the registered contract). - Same files' ctx7 CLI fallback rationale: the cited cause (anthropics/claude-code#13898 'strips MCP tools from agents with a tools: frontmatter restriction') was wrong on two counts — #13898 is closed and was never about tools: frontmatter. Rewritten to describe the real mechanism (custom subagents cannot see project-scoped .mcp.json; they only inherit user-scoped ~/.claude/mcp.json). The fallback itself is kept. - discovery-phase.md 'mode: code/info' dropped — query-docs takes libraryId + query only; the code-vs-concepts intent is now expressed via the query text. resolve-library-id is unchanged (still registered upstream). CHANGELOG and RELEASE-NOTES-LEGACY citations are historical record, left as-is. * chore(#2943): add changeset fragment (pr:0 placeholder) * test(#2943): widen parity-guard scan surface to docs/ (isolated-review finding) The isolated adversarial review flagged that SCAN_DIRS omitted docs/, which ships docs/AGENTS.md — agent-consumed prose carrying 8 mcp__context7__* refs. No false negative today (it uses only the wildcard), but a future banned-name addition there would slip through, recreating the exact drift this guard exists to prevent. Add docs/ to the scan surface, with an EXCLUDED_FILES set for historical record (docs/RELEASE-NOTES-LEGACY.md, CHANGELOG.md) that must not be rewritten to satisfy the guard. * fix(#2943): update shifted PROSE_ALLOWLIST line + acknowledge gsd-executor.md growth The gsd-test gate caught two real consequences of the rationale rewrite in agents/gsd-executor.md (the +2-line corrected mechanism description shifted line numbers below it): 1. tests/no-bare-gsd-tools-command-position.test.cjs: the legitimate 'gsd-tools query commit' descriptive mention moved from line 791 -> 793. Update the PROSE_ALLOWLIST entry to the new line (the mention is unchanged, just relocated by my edit above it). Without this the gate reports both a stale allowlist entry (791) and a new offender (793) for the same mention. 2. tests/emitted-drift-acks/2943-context7-tool-name.json: gsd-executor.md grew 95 bytes (the accurate mechanism rationale is longer than the wrong one-line #13898 attribution it replaces). Acknowledge the growth with the reason. Both are mandated by the gate, not optional. The rename itself (get-library-docs -> query-docs) is byte-neutral-ish; only the rationale rewrite grew the file. * chore(#2943): backfill changeset PR number 2963 --------- Co-authored-by: sim <sim@local> |
||
|
|
f092c6da85 |
fix(#2649): diagnose-issues + execute-plan run worktree.base-check before worktree dispatch (#2955)
* test(#2649): failing-first — diagnose-issues + execute-plan must run base-check before worktree dispatch * fix(#2649): diagnose-issues + execute-plan run worktree.base-check before dispatch diagnose-issues.md spawn_agents and execute-plan.md Pattern A spawned worktree-isolated subagents (gsd-debugger / gsd-executor) without the pre-dispatch worktree.base-check gate that execute-phase (#683/#1369) and quick (#1941) already run. Claude Code's isolation="worktree" forks from origin/HEAD, not live local HEAD; without the gate, the documented GSD steady state (commit every step locally, push only on request) hits the verify-only worktree_branch_check guard's exit-42 halt mid-investigation with no auto-degrade. Mirror the quick.md #1941 pattern: before dispatch, run `gsd_run query worktree.base-check --pick shouldDegrade`; if true, print its message + a #2649 warning to stderr and set USE_WORKTREES=false (sequential main-tree dispatch). The verify-only guard stays as a backstop in both cases. Per the triage and #2649 acceptance criterion 5, execute-plan.md's Pattern A (identified as a second site with the identical gap) is fixed in the SAME change — same bug class, same one-line gate, two workflow files — rather than filed as a separate follow-up. * fix(#2649): ack the diagnose-issues + execute-plan growth (per-PR fragment) The two workflow files grew vs next (diagnose-issues.md +1381, execute-plan.md +905) adding the #2649 base-check gate. emitted-attribution requires an ack; this is a per-PR fragment under tests/emitted-drift-acks/ (#2914 mechanism, replacing the legacy shared emitted-drift-ack.json). * test(#2649): tighten base-check ordering assertion + guard backstop survival Address code-review minors: - the ordering assertion was a loose disjunction that passed even if the base-check moved AFTER the dispatch; tighten to assert base-check < Agent() (the real invariant). - add a test that the verify-only <worktree_branch_check> backstop remains embedded in the Agent() prompt (acceptance criterion 4 — the base-check is a pre-dispatch degrade, the guard is a post-fork fail-closed backstop; both layers must survive). * changeset(#2649): diagnose-issues + execute-plan auto-degrade on stale worktree base * changeset(#2649): backfill PR number 2955 --------- Co-authored-by: sim <sim@users.noreply.github.com> |
||
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
76b7d73039 |
fix(#2733): route gate-passed spec-phase paths into the probe steps (#2779)
* fix(#2733): route gate-passed spec-phase paths into the probe steps All four gate-passed transitions in spec-phase.md said "Jump to Step 6", textually bypassing the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes. Steps 5.5/5.6 were spliced between Step 5 and Step 6 by two later feature commits and the pre-existing jumps were never re-pointed, so no jump instruction in the file reached Step 5.5 at all and both probes were unreachable dead prose. Re-point the four gate-passed jumps (lines 129, 162, 168, 170) to Step 5.5. Control then flows 5.5 -> 5.6 -> 6 as the probes' own preconditions prescribe. The max-rounds "write anyway" bypasses and the probes' own "proceed to Step 6" exits are deliberately unchanged. Add tests/spec-phase-probe-reachability.test.cjs, which derives the mandatory probe steps from the file's own headings rather than hardcoding 5.5/5.6, so a future spliced-in probe step is covered without editing the test. It also locks the two coupled constraints: the max-rounds bypass must not be redirected into a probe, and each probe must keep its own onward exit. The existing probe contract tests are untouched and still pass; both scope from the "## Step 5.5"/"## Step 5.6" heading onward and were structurally incapable of observing the upstream jump text. * chore(changeset): Fixed fragment for #2779 (spec-phase probe reachability) * fix(#2733): route Step 5.5's own soft gate into Step 5.6 Round-1 review blocker. The four upstream gate-passed jumps were re-pointed to Step 5.5, but Step 5.5's own terminal soft gate at :305 still read "proceed to Step 6" - so the COMMON path (all applicable edges resolved) skipped the prohibition-completeness probe outright. Same defect class as the four this PR already fixed, on the success path of the very step being fixed: the SPEC shipped with an empty Prohibitions section instead of an empty Edge Coverage one. Its sibling at :393 is byte-identical yet correct, because Step 6 genuinely follows Step 5.6. Position, not phrasing, is the discriminator. The guard could not see it: the transition matcher keyed only on the literal "Jump to Step", and :305 says "proceed to Step". Widened it to a verb alternation (jump/proceed/continue/go/return/skip + "to Step N", case-insensitive) and renamed it TRANSITION_RE to match what it now models. This makes the file's own docstring promise - that a future spliced-in probe is covered without editing the test - true for a step whose exit is worded differently. Verified no false positives: the two pre-existing "continue to Step 3/4" transitions are upstream of both probes but target pre-probe steps, and the max-rounds bypass block contains no step transitions at all. Fail-first verified before fixing :305 - with the widened matcher against the unfixed workflow the guard fails naming exactly "spec-phase.md:305 jumps to Step 6, skipping mandatory Step 5.6", 4 pass / 1 fail; after the fix, 5/5. The two sibling probe contract tests stay 16/16. Also from review: - STEP_HEADING_RE gains an explicit \r? before $. Without it, on a CRLF checkout `.` stops before the \r and the unanchored $ fails to match, yielding ZERO steps and vacuously passing every assertion in the file. Not live today (.gitattributes forces eol=lf) but this repo has a recurring CRLF-regex bug class, so the guard no longer leans on it. - allow-test-rule category corrected to source-text-is-the-product; the previous runtime-contract-is-the-product is not one of the six recognized categories (CONTRIBUTING.md:609-619). - changeset body given the documented bold-lead-in form. - emitted-drift ack reason updated: +8 -> +10 bytes across five transitions (31987 -> 31997), DEFAULT tier, cap 40960. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
854c93533c | chore: sync next package version to 1.9.1 | ||
|
|
42f4f184c0 |
fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context verify-summary's Pattern 1 matched any backticked path-like token with no context check, so a prose mention of a future deliverable (`shared/types.ts` in a 'next phase will add…' sentence) was checked for existence and its absence failed the verdict on a healthy phase. #2685 added shape filtering but no context check. - src/verify.cts: both extraction patterns now require a claim label on the line (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no longer matches; genuine labeled claims still do. - gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION so relative claim paths resolve against the project root, not the raw cwd (subdirectory invocation no longer manufactures missing files). Regression tests: prose mention not treated as a claim; prose-only SUMMARY passes; absent claimed file still fails. * chore(#2844): backfill changeset PR 2910 --------- Co-authored-by: Test <test@example.com> |
||
|
|
79ed181ec0 |
fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows, tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow structural pre-pass then no-op'd silently — a hard execution failure read the same as 'optional dependency absent'. (A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is preserved; cmdArgs stays an array. POSIX untouched. (B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an absent binary. Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX negative-space test guards the unchanged bash -c callers. * chore(#2667): changeset fragment * chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true) * fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth CI caught two failures on the first push: 1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout cap's process-group reap (exit 124 never fired; hit the 30s harness backstop) AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY excluded now: real PE executables spawn fine directly; only .cmd/.bat are the CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test (node.exe spawned directly, exit 0). 2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the #2667 fallow pre-pass failure-KIND case statement; acknowledge it. --------- Co-authored-by: Test <test@example.com> |
||
|
|
4232a79396 | chore: sync next package version to 1.9.0 | ||
|
|
6e0bc50142 |
fix(#2666): code-review scopes root-level + extensionless build files, cross-checks against git diff (#2895)
* test(#2666): add regression + docs-parity guards for code-review file scoper The Tier-2 SUMMARY.md extractor dropped every repository-root file (no `/`) and every extensionless build file (Dockerfile/Makefile/etc.) via an AND-joined predicate. Adds behavioral tests against the pure-function mirror plus docs-parity structural guards that bind the shipped workflow .md to the fix. RED: the docs-parity guards fail against the pre-fix shipped predicate. * fix(#2666): accept root-level + extensionless build files in code-review scope; intersect-and-warn Two coordinated edits to gsd-core/workflows/code-review.md compute_file_scope: (A) Tier-2 SUMMARY extractor: replace the AND-joined predicate `/\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw)` (which required BOTH a directory separator AND a trailing extension, silently dropping every root-level file and every extensionless build file) with a relaxed predicate that accepts any path with a trailing extension OR a known extensionless build basename (Dockerfile/Containerfile/Makefile/Justfile/Procfile). (B) Tier-3: convert the eq-zero git-diff gate into an intersect-and-warn — whenever a reliable diff base is available, cross-check the SUMMARY scope against `git diff --name-only` and warn about (then add) any changed files the SUMMARY extractor did not surface. Portable (bash 3.2, no associative arrays) so a partial SUMMARY result can no longer silently ship an incomplete review scope. * chore(#2666): changeset fragment * fix(#2666): use exact whole-line matching (grep -Fxq) in Tier-3 cross-check Adversarial review found the unanchored `case "$IN_SCOPE" in *"$file"$\\n*` substring membership test would false-match: a root-level `Dockerfile` in the diff substring-matches an already-scoped `docker/Dockerfile`, silently skipping it — reintroducing the exact class of silent-scope-loss bug this PR fixes. Switch to `grep -Fxq` (exact whole-line match). Add docs-parity guard for exact matching + the basename-collision regression. * fix(#2666): resolve gsd-test failures — paraphrase predicate in comment, ack code-review.md growth gsd-test caught 3 issues on 7469c3f22: 1. The docs-parity guard fired on the .md COMMENT which restated the buggy predicate verbatim — paraphrase the comment so it no longer contains the exact string the guard detects. 2. Cascade subtest failure from #1. 3. emitted-attribution: code-review.md grew 2814 bytes — acknowledge the deliberate #2666 growth in tests/emitted-drift-ack.json. * chore(#2666): backfill changeset PR number 2895 --------- Co-authored-by: Test <test@example.com> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
b7b5c3712c |
fix(#2762): chunked --reviews replans instead of no-op + outline resume marker written to file (#2887)
* test(#2762): chunked --reviews must replan, not no-op (outline marker + per-plan --reviews exception) * fix(#2762): chunked --reviews replans plans instead of skipping 100% + outline resume marker written to file Defect A: §8.5.1 outline resume-check greped for a marker the agent only RETURNED (never wrote to the file) → outline always re-ran (broke crash-resume). Fix: the outline agent writes ## OUTLINE COMPLETE into the file. Defect B: §8.5.2 per-plan resume-check skipped any plan with frontmatter, no --reviews exception → --reviews skipped 100% of plans (contradicted §6 'go straight to replanning'). Fix: gate the skip on --reviews being ABSENT. Crash-resume (non-reviews) still skips. Condensed adjacent §8.5 prose to keep plan-phase.md under the 94519B cap (net -33B). * chore(#2762): changeset fragment * chore(#2762): backfill changeset PR number (2887) --------- Co-authored-by: Test <test@example.com> |
||
|
|
3af1941948 |
fix(#2772): resolve four discuss-phase text inconsistencies (dead MAX_PASSES read, gate-prompts drift, circular auto_advance, answer_validation drift) (#2886)
* test(#2772): structural guards for the four discuss-phase text inconsistencies * fix(#2772): resolve four discuss-phase text inconsistencies 1. auto.md: remove the dead MAX_PASSES/max_discuss_passes config read (contradicted the mandated single-pass rule + wasted a shim invocation per auto run). 2. gate-prompts.md: context-handling options now match the actual check_existing flow (Update it | View it | Skip, not Overwrite|Append|Cancel); gray-area-option no longer mandates 'Let Claude decide' (contradicts discuss-phase.md's no-cop-out rule). 3. discuss-phase.md: auto_advance fallback ends the workflow instead of routing back to the already-run confirm_creation step (circular). 4. discuss-phase-assumptions.md: re-sync answer_validation to the parent canonical block (had drifted — lost the 'Other' empty-text branch). * chore(#2772): changeset fragment * fix+test(#2772): also fix the assumptions auto_advance circularity (review minor 1) + add positive test anchors (review minor 2) The sibling discuss-phase-assumptions.md had the identical auto_advance→confirm_creation circularity; fix it the same way (end the workflow). Add positive anchors to both auto_advance tests so a re-phrased regression can't slip past. File #2885 for the dead max_discuss_passes config still advertised in settings/registry/docs (review minor 3). * fix(#2772): keep discuss-phase.md under the 32000B #717 cap + ack assumptions growth The auto_advance fixes + the assumptions answer_validation re-sync grew both files past the emitted-attribution gate (and discuss-phase.md past the #717 32000B cap). Condense the auto_advance prose in both files (discuss-phase.md now net -11, under cap; auto.md already net -4650 from the MAX_PASSES shim removal). Add discuss-phase-assumptions.md to tests/emitted-drift-ack.json for its residual +220 (answer_validation re-sync + auto_advance fix). * chore(#2772): backfill changeset PR number (2886) --------- Co-authored-by: Test <test@example.com> |
||
|
|
dbb0a653be |
fix(#2771): advisor mode spawns registered gsd-advisor-researcher subagent instead of general-purpose (#2884)
* test(#2771): advisor mode must spawn gsd-advisor-researcher, not general-purpose * fix(#2771): spawn registered gsd-advisor-researcher subagent instead of general-purpose in advisor mode universal-anti-patterns rule 10 (injected into discuss-phase via <required_reading>) says NEVER use non-GSD agent types. The advisor mode spawned general-purpose and manually told the agent to read the def — but gsd-advisor-researcher IS registered, so spawning by type auto-loads it. Drop the manual-read prompt line (re-specifying the def is a drift risk) and use the registered type. * chore(#2771): changeset fragment (mentions follow-up #2883) * test(#2771): widen manual-read-line regex to deny phrasing variants (review minor) /read\s+@.*gsd-advisor-researcher\.md/i (case-insensitive, any 'read @' lead-in) so a drift variant like 'Read @' or 'Load @' can't sneak the manual-def-read back in. * chore(#2771): backfill changeset PR number (2884) --------- Co-authored-by: Test <test@example.com> |
||
|
|
185da024cb |
fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate The handler conflated empty-arg (caller error) with file-missing (legitimate skip), returning passed:true/skipped on an empty argument. Add: empty arg → passed:false; real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed. * fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block Handler (check-command-router.cts): split the guard — empty/missing contextPath argument is a caller error (fail closed, passed:false, mirrors #1365); a real path whose file genuinely does not exist keeps the legitimate green skip. Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block (it was set in the step-1 init block, which does not survive into the separately- spawned gate block — so the gate ran with an empty arg and silently green-skipped). * chore(#2770): changeset fragment * fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test The handler now fails closed on an empty contextPath arg, so the workflow's unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the gate with an empty arg → passed:false → exit 1, hard-halting the legitimate 'Continue without context' plan-phase path. Guard the empty-glob case: only run the gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the in-block recompute AND the empty-glob guard. * fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net +89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth. Widen the drift-guard test window (the gate invocation is now nested in the empty-glob guard, so the old 400-char window missed the glob recompute). * chore(#2770): backfill changeset PR number (2881) --------- Co-authored-by: Test <test@example.com> |
||
|
|
82ca13f5a5 |
fix(#2736): write current_phase_name from the transition intent, not the lossy prose round-trip (#2821)
* fix(#2736): intent-first current_phase_name on transitions; dash-first prose precedence Primary: completePhase (adapter) and beginPhase (via readModifyWriteStateMd options) pass the intent-held display name to syncStateFrontmatter as an authoritative override, applied after every derive/preserve/carry-forward step — so the lossy prose round-trip can never destroy a name the transition just resolved. Names containing a parenthetical (`Closer-ruling measurement (D1a)`) now land in frontmatter verbatim instead of collapsing to the parenthetical (`D1a`). Secondary (#1695 AC #3 residual): parsePhaseFromProse prefers the em-dash name when it is a genuine name (not a status keyword, not a `Milestone:` tail), else falls back to the parenthetical — satisfying both first-party writer shapes (`N — Name (aside)` and `N (Name) — EXECUTING`). Still lossy for paren-containing names, which is why the intent-first override is the primary fix. plannedPhase carries no name in its intent, so it is naturally out of scope. Fixes #2736 * docs(changeset): backfill PR number for #2736 fragment * fix(#2736): drop an unnecessary type assertion on result.data StateTransitionResult.data is already `Record<string, unknown> | undefined`, so the cast was a no-op and tripped @typescript-eslint/no-unnecessary-type- assertion (CI lint-tests red on the first push; every test lane was green). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
b36e3b7e1f |
fix(#2751): normalize bare gsd-tools command-position calls to gsd_run in shipped source (#2851)
* test(#2751): regression guard — no command-position bare gsd-tools calls Agents/workflows instructed bare `gsd-tools <verb>` invocations that fail with 'command not found' on a shim-only install (#725 fixed only the Codex conversion pipeline; the Claude-facing source shipped them verbatim). Adds a source-text guard (allow-test-rule: source-text-is-the-product) scanning agents/*.md + gsd-core/workflows/*.md for the operative shape `gsd-tools <verb> <arg>`, excluding command -v probes / resolver definitions, with a documented PROSE_ALLOWLIST for descriptive mentions that name the command without instructing literal invocation. A stale-allowlist check ensures entries stay real. RED first; source fix lands next commit. * fix(#2751): normalize bare gsd-tools calls to gsd_run in Claude-facing source The 12 command-position bare `gsd-tools <verb>` instructions across agents/ and gsd-core/workflows/ failed with 'command not found' on a shim-only install (no gsd-tools binary on PATH). #725 fixed this only for the Codex install-conversion pipeline; the Claude-facing SOURCE shipped the bare calls verbatim, and new ones kept accumulating (new-project.md:114 landed 12 days AFTER #725 closed). Rewrite each operative site to the portable `gsd_run` resolver that the same files already define (3-18x each) — a pure command-position token swap preserving all arguments, flags, --files, and surrounding prose. Every runtime now benefits from one source change instead of each needing its own converter. Touched sites (12): gsd-intel-updater (validate/snapshot/extract-exports), gsd-code-fixer (query commit), gsd-planner (learnings.query), gsd-project- researcher (websearch/research-plan/classify-confidence), gsd-phase-researcher (websearch/research-plan/classify-confidence), new-project (project-instruction- file / commit --files), new-milestone (commit --files). Preserves command -v gsd-tools probes, resolver-snippet definitions, and the 4 descriptive prose mentions that NAME the command without instructing invocation. RED @ 1bb12ba1. * chore(#2751): allow-test-rule issue ref + changeset fragment Add the (#2751) tracking ref to the source-text-is-the-product annotation per ADR-456, and the .changeset Fixed fragment (pr:0, backfilled post-PR). * fix(#2751): convert remaining command-position bare gsd-tools calls (verify-summary, windows, worktree, smart-entry, quick-tasks-append) Isolated adversarial review (Step 4) found the first pass missed genuine command-position bare calls because the regression test's hand-maintained 6-verb list silently false-passed verify-summary (the 'verify' branch matched the prefix then died on the hyphen) and omitted windows/worktree/smart-entry/ quick-tasks-append entirely. Convert these 8 additional operative sites across new-project.md, new-milestone.md, ship.md, execute-phase.md, progress.md, smart-entry.md, quick.md. * chore(#2751): backfill changeset PR number (2851) * fix(#2751): normalize allowlist paths to forward slashes — Windows path-separator false-flag The PROSE_ALLOWLIST is keyed by file:line using forward-slash paths, but path.relative() returns backslash separators on Windows, so the allowlist lookup failed and the 6 descriptive mentions were flagged as offenders on the windows-latest CI lane. Normalize rel to forward slashes before the lookup so the allowlist matches identically on every OS. --------- Co-authored-by: Test <test@example.com> |
||
|
|
0408276791 |
chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the ADR's swap amendment: a federated config slice lives inside a capabilities/<id>/capability.json, and three of the five key families had no capability directory until 5a created them. Four key families move to the lanes that use them; the central-schema removal and the federated addition land in this one commit because the exclusivity invariant fails the build on a key present in both. review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy ACROSS lanes and stay central. Two things the issue did not name, both found while building it: 1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys against manifest.validKeys only, and two of the four families (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>) were pattern-backed. That is not cosmetic: isCentralConfigKey consults those patterns and mergeFederatedConfig skips every key for which it returns true, so declaring a slice while the pattern survived would have shipped an INERT slice behind a green gate — the exact half-migrated shape the invariant exists to prevent. The gate now loads the patterns from the same manifest the runtime reads. 2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a federated key always resolves to its declared default. The three fallback guards in review.md checked only empty-or-"null", so a user who set the GLOBAL review.max_prompt_tokens would have silently lost trimming on the HTTP lanes. The guards now treat 0 as unset. D9 says review.models.<slug> is owned by "the lane whose slug it names". That is false for one lane: the shipped key is review.models.agy while the slug is antigravity. Ownership follows the lane; the key name is preserved, because renaming would break every config that sets it. Existing tests updated rather than left asserting the old world: config-get on a cleared federated key yields empty instead of not-found (what #2046 actually protects — never persisting the literal "null" — is unchanged and still asserted); the config-schema dynamic pattern representative moves to reviewer_instances; the prototype-pollution guard case moves to a surviving dynamic prefix so alert #26 keeps its coverage, with a new case asserting the old key is now rejected earlier; and Phase 2's harvest-widening inertness assertion becomes an ownership assertion, since Phase 4 is what consumes it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives A federated config key always resolves to its declared default, so an unset per-lane prompt budget needed a value the workflow could treat as 'not configured'. The first cut used 0 — which is wrong: 0 is already a LEGITIMATE per-lane budget meaning 'do not trim this lane' (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. The sentinel is now -1, which is not a valid token budget, so all three states stay distinguishable: unset falls back to global, an explicit 0 disables trimming for that lane, and an explicit N is used. Locked by three CLI round-trip tests. Surfaced by the isolated security reviewer before it crashed mid-run; verified independently against the shipped trim guard rather than taken on trust. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): update central-registration assertions and stay under the review.md cap The remote runner caught both; my local sweep missed the files. 1. tests/plan-review-convergence.test.cjs asserted the three local-server host keys are in VALID_CONFIG_KEYS. They are federated to their lane capabilities now, and the exclusivity invariant forbids a key living in both places. What #2306-local actually protects is that config-set ACCEPTS them, so that is what is asserted — via isValidConfigKey, the predicate config-set itself uses, which spans central and federated. A second assertion pins federated ownership, so a silent reversion back to the central schema fails too. 2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap is a red line, not a budget to raise. The three per-lane budget guard comments were near-identical; condensed to one terse line each. 61371 bytes, 69 to spare. Real extraction to workflows/review/modes/ is Phase 5b/6 work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs Isolated security review findings. MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse error and returned [], while its sibling loadCentralConfigKeys, reading the SAME file, writes to stderr and throws ExitError(1) on that identical failure class. Fail-open here defeats the gate this function exists to feed: with zero patterns, validateCrossCapability's pattern-collision check silently passes and an inert federated slice ships green. It was masked in the one production call site only because loadCentralConfigKeys runs first against the same path — a coincidence of ordering, not a guarantee, and this function is exported and called standalone. The two now share a contract: ENOENT is the legitimate absent case, anything else throws loudly. A single unparseable PATTERN is still skipped, which degrades to "checked less" rather than blocking every build. The branch had zero coverage; it now has two tests (malformed JSON, EISDIR). MINOR — docs/CONFIGURATION.md still listed review.models.qwen and review.models.cursor as settable, ~770 lines below this PR's own new Ownership section. Those lanes take no model flag, so they declare no model key and config-set now rejects them. Rows removed; the missing review.models.agy row added; the per-reviewer budget row corrected to name only the lanes that own a budget key, and to document that a per-lane 0 disables trimming for that lane. Also fixes a shadowed "raw" binding introduced by the fail-closed change, which made the generator unrequirable — caught immediately by its own --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2797): backfill changeset pr number to 2841 * fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next stays green, and it is currently blocking at least three unrelated PRs (#2841, #2832, #2827) with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees The existing loop comment attributes this to `git add .` rehashing O(n^2) blobs "before the object write had landed" and works around it by staging one path per iteration. That is not the cause: sequential execFileSync calls cannot race each other's object writes, and the failure persisted after that change — it simply moved to a lower commit index. The cause is that the git() helper inherited the runner's environment. A leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole operation. In every case `git commit` then cannot resolve a blob it just staged, which is precisely the error above. Verified by negative control: with GIT_DIR exported, this test fails on the unfixed helper (the git commands operate on the wrong repository entirely); with the helper stripping GIT_* it passes. The single-path staging is kept — it is genuinely less work — but it is no longer load bearing. Found while shipping #2797. Fixed in place rather than deferred: it is a defect surfaced during the work, and it is blocking other contributors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2452): build the base-advance commits empty, removing the lost-object class The base-ref guard has been failing in CI with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees It is currently red on at least three unrelated PRs (#2841, #2832, #2827) while next stays green. Two theories have now been tried and neither held. #1881 blamed `git add .` rehashing O(n^2) blobs and switched to staging one path per iteration; the failure moved from commit 32 to commit 25 and carried on. The preceding commit here made the git helper hermetic against leaked GIT_* environment — that IS a real vulnerability (with GIT_DIR exported the helper operates on the wrong repository entirely, proven by negative control) but it produces a different error than CI reports, so it is not demonstrably the cause either. Neither trigger reproduces off-CI, so this stops guessing at the trigger and removes the failure CLASS instead. The loop needs base-branch DEPTH and nothing else: no assertion reads these commits' contents, and base-side files cannot appear in `origin/base...HEAD` regardless. `--allow-empty` writes no blob and no tree, so there is no object for the index to reference and lose. It is also far less work than 60 write+hash+index cycles. The guard still proves its mechanism: the test asserts that a --depth=1 base fetch FAILS and a full fetch resolves, so a broken topology would surface immediately rather than passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Test <test@example.com> |
||
|
|
b12d4df03b |
fix(#2694): normalize CRLF before frontmatter-boundary match in code-review workflows (#2839)
* test(#2694): CRLF frontmatter boundary regression for code-review workflows The code-review / code-review-fix workflows embed inline node -e one-liners whose frontmatter boundary regex used a literal \n, silently returning null on CRLF-saved SUMMARY.md / REVIEW.md / REVIEW-FIX.md artifacts and dropping every file in that summary (acceptance: per-artifact, no warning when the phase aggregate stays non-zero). Adds: - behavioral CRLF==LF boundary extraction tests (replica of the shipped one-liner's boundary step), proving the buggy literal-\n returns null on CRLF while the fixed normalize-then-match yields a byte-identical body; - a structural-regression-guard (allow-test-rule: structural-regression-guard) that reads the two shipped workflow files and asserts every boundary site normalizes \r\n -> \n before matching, so a revert of the fix is caught. * fix(#2694): normalize CRLF before frontmatter-boundary match in code-review workflows The code-review and code-review-fix workflows embed nine inline node -e one-liners that extract YAML frontmatter via a boundary regex content.match(/^---\n([\s\S]*?)\n---/) The literal \n defeated any CRLF-saved artifact (\r between --- and the line terminator), so SUMMARY.md / REVIEW.md / REVIEW-FIX.md saved with CRLF endings silently contributed zero files (or 'unknown' status / 'invalid') with no per-artifact warning. The Tier-3 git-diff fallback only fires when the aggregate across all summaries is zero, so a single CRLF summary among LF summaries produced no signal at all. Normalize \r\n -> \n once before the existing boundary match at all nine sites (code-review.md x3, code-review-fix.md x6). Byte-identical to the LF path; mirrors the canonical src/frontmatter.cts extractFrontmatter intent (CRLF == LF at the boundary); zero risk of \r leaking into field values consumed by the inner JS or the shell grep/cut pipeline. RED @ 94d0213 (3 failures, structural guard caught the shipped-text bug, both linux-node22+24 lanes).GREEN pending. * chore(#2694): acknowledge code-review workflow growth + changeset fragment emitted-attribution (ADR-2719) reports the byte growth from the CRLF-normalize insertion in code-review.md (+69) and code-review-fix.md (+138); both are the intended #2694 fix. Adds the .changeset Fixed fragment (pr:0, backfilled post-PR). * test(#2694): mixed CRLF/LF phase yields the union of both artifacts (criterion 2) The spec-axis review flagged that acceptance criterion 2 (a phase with a mix of CRLF-affected and unaffected artifacts no longer silently drops the CRLF artifact's contribution) was only transitively satisfied. Adds an explicit mixed-phase test replicating the full shipped Tier-2 extractor (boundary + inner key_files parse) across one LF and one CRLF SUMMARY.md, asserting the union of both — plus a RED proof showing the buggy boundary drops the CRLF artifact silently (aggregate non-zero, so the Tier-3 eq-zero fallback never fired). Locks the silent-partial-masking behavior the triage named as the more serious half of the defect. * docs(changeset): backfill #2694 PR number to 2839 * fix(#2694): make the CRLF regression test itself CRLF-lint-clean CI lint-tests caught that the new test tripped local/no-crlf-fragile-split: - the frontmatter boundary regex replicas (fixed + buggy) were RegExpLiterals with a bare \n; the rule flags frontmatter-shape regexes unconditionally. Build them via new RegExp(...) (byte-identical .source to the shipped literal) so the faithful replica is not a lint violation — the buggy replica MUST keep the literal \n, that is the bug it demonstrates. - the structural guard's src.split('\n') on the readFileSync'd workflow file was genuinely CRLF-fragile; use /\r?\n/ per the rule's canonical fix. - the allow-test-rule annotation gains its (#2694) tracking ref per ADR-456. lint:ci now exit 0 (incl. lint-allow-test-rule-refs, lint-emitted-drift-ack, lint-fix-has-regression-test: PASS). |
||
|
|
6a9babda69 |
chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3. - Five reviewers GSD never installs into become lane-only role:reviewer capabilities with no runtime body, no runtimeCompat and no install surface: gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail, which is now deleted outright. - The six hosts that are ALSO reviewers gain a reviewer body alongside their runtime body. Their runtime bodies are byte-identical to next -- verified per capability against the git blob, not asserted -- so no install behaviour moves. - KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived legacy alias for one release; where a capability carries both, the body wins and the slug appears once. Alias removal is Phase 7 (#2801). THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity, claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama, opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it, and the test asserts that literal list rather than a count. kimi-code is deliberately NOT declared here. It is net-new with no invoke_reviewers leg, so declaring it now would make it selectable but not invocable -- present in --all, selected, emitting an empty section for the whole 5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a reviewer at all and gains nothing. The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field, including probe and invoke sub-fields. All eleven are byte-identical, key order included. The epic's premise is that the manifest and the core descriptor describe the same lane with NO translation layer, and Phase 2's review already caught one divergence that every other test missed. Two ADR corrections folded in, as Phases 1-3 each did: 1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798 claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable: D9 assigns review.<host>_host to lane capabilities that do not exist until THIS phase creates them, and a federated config slice must live inside capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4. 2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs bin/lib/*.cjs modules, not capability directories -- antigravity, opencode and qwen appear zero times in it -- and gen-inventory-manifest --check passes with the five new dirs and no edit. Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported LANE_SLUG_RE. The prose had not followed the code. Closes #2798 * fix(#2798): catalogue reviewer capabilities in the generated matrix The capability matrix rendered exactly two tables, feature and runtime, via renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD role, so every role:"reviewer" capability was silently dropped from the first-party catalogue. The drift guard did not catch it, and could not: --check compares generated output against the committed file, and both omitted the five lanes identically, so it reported "up to date" while five shipped capabilities were invisible in the one document that is supposed to list what ships. A guard blind to an entire role is not guarding. This phase is what exposed it -- it ships the first role:"reviewer" capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect found while working is fixed in the current change, which overrides one-concern-per-PR). Verified red-before-green: with a lane row deleted from the matrix, --check now exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so there was nothing for the guard to compare. Phase 6 (#2800) still owns enriching the matrix with lane-specific detail (slug/flag/transport columns) and the locale parity gate. This is the narrower fix: the capabilities APPEAR at all. * fix(#2798): close two hardening gaps and record three limits durably Isolated security review (5 targets, no blockers) reproduced two gaps in the new deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for reuse and carries no other validation, so it must not depend on its caller. - A whitespace-only slug passed the length>0 test verbatim and occupied a roster entry it could never match. Slugs are now trimmed before the emptiness test. A blank body correctly falls through to the legacy alias rather than DROPPING the lane, which would have been worse than the blank slug. - KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there breaks import for EVERY consumer rather than degrading selection. It is now guarded, yielding an empty roster on a malformed registry. That is a visible degradation, not a silent one: under D4 an explicitly requested reviewer that is unavailable is an ERROR, so /gsd:review --claude against an empty roster fails loudly. This also removes an asymmetry -- the sibling capability-trust module documents its collectors as TOTAL and wraps them for exactly this reason. Also records three findings that previously existed ONLY in squash-merged PR bodies, which is not a durable record: - ADR-2782 D5 gains an implementation note explaining why the resolved host is deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds the resolved host; the loader has no config resolver, so folding it in would make the loader and lifecycle compute different signatures for one manifest and re-prompt forever. The binding is split: signature covers the SHA-pinned manifest fields, the consent record stores the resolved host, and Phase 5b re-resolves at invocation -- which is where rule 4 already puts the check. A reader comparing rule 1 to the code would otherwise conclude it is unimplemented. - CONTEXT.md's capability-trust entry still described THREE executable surfaces. Phase 3 added the fourth and made that false; corrected here, since it is drift this epic introduced rather than Phase 6's new-glossary-term work. - stableJson documents the NaN/Infinity/undefined -> null signature collision and why it is unreachable (JSON grammar has no such literal, so JSON.parse throws first). Reachability rests entirely on the ingest path staying JSON.parse-only, so the note lives where someone would break it. * chore(#2798): backfill changeset pr number to 2837 |
||
|
|
46e84d5e39 |
chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat (#2823)
* chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat Phase 2 of epic #2782 under ADR-2782. Delivers D1, D2, D3, D7, D8 and the four Phase-1 vocabulary amendments (A1-A4). - VALID_ROLES gains "reviewer"; the reviewer body is admissible on role:runtime (a host that is also a reviewer keeps one manifest) and on the new role:reviewer (a lane that is not an install target). A reviewer body on role:feature is an error: declaring one is an assertion of lane-ness. - validateReviewerBody + validateLaneProbe + validateLaneInvoke: nine closed enums, a transport discriminator selecting mutually-exclusive invoke sub-shapes, bounded probes (D7), and outputArg required-iff outputChannel is file-arg and forbidden otherwise. - Absent-safe (D4.1): only `undefined` is absent. null/{}/[]/false/0 are malformed assertions and error. 39 of 39 shipped capabilities depend on this. - collectReviewerWarnings: an unknown field inside the body warns, never errors, so a forward-built manifest degrades visibly instead of failing the build. - D8 uniqueness (slug / flags / reviewsSection) lives in validateCrossCapability, so it is enforced at build time over first-party AND at load time over the merged first-party union overlay set, with first-party-wins falling out of the loader's existing ordering rather than a new provenance check. - Config harvest widened past the role==="feature" branch in both the generator and the ownership loop. The often-cited cause of the stranded reviewer config keys -- the runtime body forbidding feature-only fields -- is not the mechanism: `config` is not in FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME. The cause is two harvest sites that never read it. Verified inert: no shipped capability declares config on a non-feature role, and the generated registry is unchanged. Three ADR corrections are folded in (Phase 1 set the precedent of amending in-phase): the misattributed config-stranding cause, D3's inverted profile-membership claim, and the specified capability folder names for lm_studio / llama_cpp, which would have failed the id kebab-case invariant. Closes #2795 * chore(#2795): collapse nine enum checks into one validateEnumField helper Standards-axis review findings, both applied: - Duplicated Code: the enum-membership + enumerate-the-members error shape repeated near-verbatim at nine call sites. Routing them through one helper makes "the error names its valid members" structural rather than a convention repeated nine times, where it would drift. That property is load-bearing until Phase 6 ships the prose reference, because these errors are currently the only documentation of the vocabulary. - Speculative Generality: the isReservedName() pre-check on every enum field was inert. A VALID_* set never contains __proto__/constructor/prototype, so membership alone already rejects them, and "must be one of: ..." is more actionable than "is a reserved name". The literal guards remain where they do real work -- the key-derived write sites in the registry generator and the claim() accumulator. The reserved-name test now asserts all three reserved names are rejected via enum membership, rather than one name via a branch that no longer exists. * fix(#2795): align lane slug grammar with Phase 1 and wire the load-time diagnostic channel Spec-axis review findings, both applied. (1) The slug grammar had diverged from Phase 1's core descriptor. Phase 1 exports LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/ (leading digit permitted); the manifest validator required a leading LETTER. A slug the core descriptor accepts -- a model-named lane such as 4o-mini -- would have been rejected by the manifest validator, which is exactly the translation layer ADR-2782 exists to delete. It was inert only because all eleven shipped slugs begin with a letter, so nothing else would have caught it until a third party shipped such a lane. The grammar cannot be reduced to one definition: Phase 1's module compiles to gitignored build output, and capability-validator.cjs is a committed plain .cjs that must load on a fresh worktree before build:lib has ever run. That makes this the repo's DEFECT.GENERATIVE-FIX class, so the duplication now carries a parity assertion -- laneSlugGrammarMatchesPhase1Descriptor -- which compares both the source grammar and the accept/reject verdict for a shared input set, and fails if the two ever drift again. (2) collectReviewerWarnings had exactly one caller: the build-time generator, which only ever sees first-party in-repo manifests. The real third-party overlay loader never called it and ValidatorModule did not declare it, so ADR-2782 D4.3 -- an unknown field inside a reviewer body is ignored WITH A WARNING -- surfaced nowhere at runtime, which is precisely the case D4.3 exists for. loadRegistry now collects those diagnostics on the accept path, behind a typeof-guard (an older built validator without the function still loads) and a try/catch (ADR-1244 D2's never-crash contract outranks a diagnostic). They land in a NEW OverlayMeta.diagnostics field rather than OverlayMeta.warnings, because warnings records capabilities that were SKIPPED and a consumer treating every entry as inactive would mislabel a working lane. Covered end-to-end by overlayLaneWithUnknownFieldIsAcceptedAndDiagnosed, which drives a real global-scope overlay through loadRegistry and asserts the lane is accepted, produces no skip warning, and yields a diagnostic naming the field. * fix(#2795): make the reviewer validators honour their documented totality contract Isolated adversarial review finding (MAJOR), reproduced by execution. validateReviewerBody documents itself as "TOTAL: returns an array of error strings for ANY input and never throws", and the overlay loader contracts every validator to RETURN errors -- #1461 OVL-1 records a validator that THREW and would have crashed every consumer of loadRegistry. The contract was false at ten sites: JSON.stringify throws on a BigInt and on a circular structure, and every enum/scalar rejection path interpolated the rejected value into its own rejection message. Reading the value could throw too, before any message was built, via a throwing getter or a Proxy get/ownKeys trap. Not reachable through a capability.json today -- every ingestion path is a plain JSON.parse of file text, which cannot express any of those shapes. Fixed anyway: the contract is stated on an EXPORTED function, and a caller must not have to re-derive today's reachability analysis to know whether it holds. Two layers, because serialization safety alone is insufficient: - describeValue() renders any value without throwing, so messages stay useful (a BigInt now reads "got: 10n" rather than degrading to a generic fallback). - A structural try/catch around validateReviewerBody and collectReviewerWarnings makes the guarantee absolute rather than argued, covering read-time throws that fire before any message exists. The same review found the property test guarding this contract was FALSE CONFIDENCE, which is the more important half. fc.anything() at default constraints emits no BigInt, no circular reference, no getter and no Proxy -- 20,000 sampled draws produced zero of each -- so the test was named for a contract its generator could not reach. Even withBigInt is insufficient under whole-value fuzzing, because the defect needs an exotic value in a specifically NAMED field and random key names never land on one. The property is now field-targeted across all twelve reviewer fields, and a companion test enumerates the shapes fast-check cannot generate at all (BigInt, circular, throwing getter, symbol, function, null-prototype) across scalar positions, array-element positions, and read-time traps. Verified red-before-green: with the fix reverted both property tests fail; with it restored all 119 pass. * chore(#2795): backfill changeset pr number to 2823 |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1b41083220 |
enhance(#2151): probe interactive-control for loading and error states (#2575)
* feat(#2151): probe interactive-control for loading + error states The ui-consideration-probe taxonomy mapped interactive-control to only one consideration (long-text), so a control-only UI surface (e.g. a theme toggle) lifted no loading or error consideration — the verifier never asked what a control shows while its action is in flight or when it fails, and a spec omitting those states could PASS. Add 'interactive-control' to the loading and error entries' elements in UI_TAXONOMY so control-only surfaces are probed for in-flight and failure states. empty is deliberately excluded (a control is not data-bearing; empty would be Goodhart noise). No new categories, no cue-map change, no probe-core change — a widening within the closed shape-rooted 8 (ADR-550), an independently-versionable predicate- generator adapter (ADR-857). Regenerated the reference-doc coverage table and the 19 golden-install- parity fixtures (reference-doc hash). Regression test added first (RED: interactive-control yielded only long-text; GREEN: now error+loading+long-text). Closes #2151 * feat(#2151): add changeset for interactive-control loading/error coverage --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e48eb44003 |
fix(#1856): give the executor-worktree refusal a handoff instead of a dead end (#2727)
* test(#1856): failing-first contract for the orchestrator cwd-drift guard handoff The #48 guard correctly refuses to execute waves from an agent worktree, but the refusal is a dead end: that worktree can hold committed fixes AND uncommitted work, and "re-run from the orchestrator's own worktree" silently means abandoning them. The reporter was left choosing between continuing from a blocked worktree and losing the work. The guard is shell embedded in execute-phase.md, so these tests extract the block by a stable marker and EXECUTE it against real git fixtures — the shipped text is the runtime contract. Covers the stranded-commit and dirty-tree report, the integration commands, both agent- namespaces, commit-count boundaries 0/1/2, and the constraints the guard's own comment records: it must NOT fire on an ordinary branch, on 'agentic-refactor', or on a legitimate feature worktree under .claude/worktrees/, and must degrade cleanly with no resolvable base or a detached HEAD. RED expected: the marker does not exist, so extraction fails and every case errors. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#1856): give the executor-worktree refusal a handoff instead of a dead end The #48 cwd-drift guard correctly refuses to execute waves from an agent worktree — its comment records why ("this is how a wrong-base merge nearly shipped ~1000 files"), and that refusal is untouched here. The defect is that it was a dead end. At the moment it fires, the worktree can hold committed product work, uncommitted product and planning changes, and the live gap-planning context. Telling the user to "re-run from the orchestrator's own worktree" silently means abandoning all of it, because the orchestrator worktree cannot see commits that live only on the agent branch. The reporter was left choosing between continuing from a blocked worktree and losing five commits plus uncommitted work. The refusal now reports what is actually stranded — the commit count and log against the resolved base, and the uncommitted files — followed by the concrete integration sequence (commit here, switch to an orchestrator-safe checkout, merge or cherry-pick, re-run) and a verify command. Nothing is claimed that is not there: a clean worktree with no commits ahead prints the plain refusal with no empty sections. Every added command is diagnostic and `|| true`-guarded, so a failure degrades to the original refusal rather than crashing before the message prints. Verified: an unresolvable base still refuses cleanly. Deliberately NOT done: auto-merging or auto-cherry-picking the agent branch. That is precisely the operation #48 exists to stop the orchestrator performing from a drifted cwd, at the moment it has least confidence about which tree is which. Reporting beats acting here. The guard block carries a `gsd:guard=orchestrator-cwd-drift` marker so the new contract test can extract and EXECUTE the shipped shell against real git fixtures rather than asserting on its characters. Fixes #1856 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#1856): report true counts, add the changeset, and document the seam Review findings, all from the isolated adversarial pass: - The dirty-file list was capped at 20 with no indication, so a worktree with 27 uncommitted files reported 20 — under-informing the user about exactly what is stranded, which is the entire point of this change. Both lists now count BEFORE truncating and print "… and N more". Verified with 25 commits / 27 dirty files. - The has-commits condition was written out twice and could drift on a future edit. Collapsed to a single _WT_HAS_COMMITS flag. - The changeset fragment existed but was untracked, so it was in neither commit on this branch and the PR gate would have failed against real history. - CONTEXT.md:122 documents this exact seam ("the orchestrator runs a cwd-drift guard at execute_waves entry…") and was not extended. Now records the handoff report, that the refusal condition and exit code are unchanged, and that every added command is diagnostic and || true-guarded. Verified NOT a defect, correcting the review's premise: the guard block does break when its line endings are CRLF, but .gitattributes:2 is `* text=auto eol=lf`, which OVERRIDES core.autocrlf and forces LF on checkout on every platform including Windows — so the shipped file is LF there too, and the installer copies it through Node without translating endings. The reproduction (mine and the reviewer's) required injecting CRLF by hand. It is also not fixable from inside the script: a \r breaks the shell parse at the block's first line, before any #1856 code runs. Neither introduced nor amplified by this change. Also noted and left as-is by design: the review flagged that #1856's "offer an explicit recovery option" could be read as requiring an interactive/automated integration rather than printed instructions. Deliberate — see the commit that added the block: auto-merging is the exact operation #48 exists to prevent the orchestrator performing from a drifted cwd. Called out in the PR body for a maintainer decision rather than silently chosen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#1856): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
09477f925e |
fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading The Workflow backend emitted every agent() call with no model, so model_overrides / model_policy / model_profile were silently inert on that path while the inline path honored them (ADR-1411). Neither existing suite contained the string 'model' at all. The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor') rather than hardcoding either, so it asserts backend parity rather than a fixed string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids reaching the code generator, the #2285 composed seam, CLI config-defaulting, and a fast-check round-trip property. RED expected: no model key is emitted anywhere, and --executor-model does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): thread the resolved executor model into the Workflow backend The Workflow backend emitted every agent() call with no model at all, so model_overrides / model_policy / model_profile_overrides / model_profile were silently inert on that path while the inline path honored all of them. The model was not dropped at the last step — it was absent from the whole seam: agentOptions() took no model, EmitInput had no field to carry one, and ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could not forward one. The generated script asserted the parity it broke. VERIFY-FIRST, which #2686 flags as the question that decides the fix: the Workflow tool's agent() DOES accept a per-call model. Its documented signature is agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? }) so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled out. ADR-1143:24's option enumeration omitting `model` is an incomplete enumeration, not a decision to exclude it. - agentOptions(p, executorModel) emits `model` only when it is a non-empty string that is not "inherit" (#2517: an empty model 404s on runtimes without native tier aliases). A non-string is a malformed config: omit, never throw. - executorModel threaded through EmitInput and ResolveWaveDispatchInput. - The CLI resolves gsd-executor from project config by DEFAULT rather than requiring a flag, reading the same source the inline path reads. An orchestrator that never learns about a new flag would otherwise silently keep the old bug. --executor-model exists only to pin/override. - ADR-1411 provenance: the generated header now states which model was applied, or that none resolved and why. A fallback must be a visible value. Compatibility: when nothing resolves, the emitted options object is byte-identical to before, so every existing caller and assertion is unaffected. Behavior change (Hyrum's Law): opted-in users move from session inheritance to the catalog-resolved executor model. Adding a `model` key also changes agent() opts, which invalidates the cached prefix of any in-flight resumeFromRunId run — a one-time re-execution. Both disclosed in the changeset. Fixes #2686 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): reject script-breaking model ids and share the emit predicate The isolated adversarial review found a BLOCKER in my own provenance comment, proven by execution (the emitted script exited 42 from an injected statement). U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment in EVERY engine — the ES2019 change legalized them inside string LITERALS only. So quoteString (JSON.stringify) is sufficient for the `model: "..."` object literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in a model id closed the comment and made the rest of the line live top-level code. The value is reachable from `.planning/config.json` (model_overrides / model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any non-claude runtime — attacker-influenceable in a cloned repo. `emitWorkflowScript` now rejects a string executorModel carrying any character in UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to phaseDir/runId, which is proof the codebase knew this hazard. Rejection is ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an emit failure to the inline backend WITH that reason, so the degradation is visible. A non-string stays on the existing defensive path (omit, never throw) — that is malformed config, not an injection attempt. Also from the reviews: - The predicate deciding "is this model emittable" was duplicated between the emission and the comment asserting it. Extracted to emittableModel() so a generated comment can never claim something the generator did not do — the exact failure class #2686 was filed for. - That predicate now trims and lower-cases before comparing, closing a real #2517-class gap: " " and "INHERIT" were previously emitted verbatim. - The adversarial test was pass-always against this very vulnerability — it asserted only that JSON.stringify appeared. Replaced with the real contract (rejection) plus an execution-level check that no LineTerminator survives into the comment. A raw U+2028 had also been committed into that test's fixture array where a tab was intended; both are now explicit \u escapes. - optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated model contained — silently not testing what it claimed. Now brace- and string-aware. Stale-test corrections in tests/fix-2285-*: three assertions froze the exact options literal `{ agentType: "gsd-executor" }`. The object legitimately gained an optional additive `model` key, so they now assert the invariant they exist to protect (agentType present, isolation absent) rather than a frozen literal. The CLI-vs-pure equality test pins --executor-model on both sides; otherwise it compared a config-resolved CLI run against a pure call given no model. CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): fix the options extractor and model the rejection path Two defects in my own test helper, caught by the full matrix: - optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The emitted shape is agent("brief", { ... }), so that never matched and the helper returned an empty array, making every assertion over it vacuously true. It now anchors on agent( and takes the first balanced, string-aware {...} after it. - The fast-check property predated the security fix and asserted ok:true for any generated string. Strings carrying an unscriptable character are now rejected, so the property models the real three-way contract: unscriptable -> ok:false; trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the trimmed value. Verified locally against the built module: omit values clean, both plans carry the model on the parity path, property passes 500 runs at seed 42. Test file re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are explicit \u escapes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): scope no-control-regex on the mirrored unscriptable-char class The class is the point of the assertion — those bytes are exactly what must be rejected — so the rule is disabled at that line rather than the class weakened. UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence the mirror. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2686): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1c93df04db |
fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows (#2713)
* test(#2711): derive the omit-rule guarded set from the corpus instead of a hand list The GUARDED array was a Goodhart metric: it reported green across 15 non-compliant workflows for no better reason than that nobody had added them to it. The guard now derives its set — every workflow emitting a model="{…}" dispatch site must state the omit-on-inherit/empty rule — and asserts the derivation is non-empty so a broken scan fails rather than passes. Rule detection stays a PROPERTY check, not a template match: plan-phase.md and execute-phase.md state it in different words and both are correct. RED expected on 15 workflows: audit-milestone, code-review, code-review-fix, debug, discuss-phase-assumptions, docs-update, map-codebase, new-milestone, new-project, quick, secure-phase, ui-phase, ui-review, validate-phase, verify-work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows 15 of the 19 model=-dispatching workflows carried no omit-on-inherit/empty guidance — 43 unguarded dispatch sites. Each would emit model="" whenever the bound *_model resolved empty, which is the DEFAULT state on non-Claude runtimes: the installer writes resolve_model_ids:"omit" into ~/.gsd/defaults.json for every one of them (references/model-profiles.md:101), and resolveModelInternal returns "" for that case (src/model-resolver.cts:383-386) and "inherit" for opus-tier agents and the inherit profile (:395). Both 404 on runtimes without native tier aliases — the failure #2517 documented and fixed in one file. Each file now carries a `<!-- #2517 model-omit-on-inherit -->` blockquote naming its own bound placeholders and linking the canonical statement in references/model-profile-resolution.md, mirroring the `<!-- #2508 runtime-aware-dispatch -->` block already present in all 15. The rule text lives in the reference; the workflows carry a pointer plus the one-line instruction, so the next revision edits one file rather than fifteen. plan-phase.md and execute-phase.md are deliberately untouched — they already state the rule in their own wording, and the guard checks the property rather than a template string. No dispatch site is edited and no placeholder renamed: the #2684 binding guard reports the same 19 files / 60 placeholders / 0 findings before and after, which is the independence proof that this change is additive prose only. There is no Hyrum's-Law routing change to disclose. Placement is span-aware. An initial pass anchored to the #2508 marker, but in six files that marker sits INSIDE the Agent(prompt="…") string, so the new paragraph's literal model= landed in a dispatch call span and tripped the #2284 fail-closed Hermes projection guard (bin/install.js:3704), refusing the install. Blocks are now anchored before the opening Agent( of the span owning the first dispatch, and verified to fall inside no span. gen:golden exits 0 across all 19 runtimes. Fixes #2711 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): cite the issue number in the changeset body and tidy block placement Review findings from the two orthogonal passes: - The changeset body ended (#0). Repo convention across every prior fragment (e.g. #2617/#2693, #2608, #2605) is that the trailing (#NNN) is the ISSUE number, known at authoring time; only the frontmatter pr: field carries the 0 placeholder pending backfill. (#0) would have rendered a dead link in the published release notes. - new-milestone.md glued the inserted block directly under the preceding paragraph with no blank line, inconsistent with the other 14 insertions. - The derived-guard non-vacuity floor was >=17 against an actual derived count of 19, tolerating a silent two-file regression. Tightened to >=19. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): reword the omit block so it survives Hermes projection, and exempt quick.md by size The first block wording regressed two suites on the full matrix (4 failures on both linux-node22 and linux-node24). gen:golden passing was not sufficient evidence — it exercises the installer's own fail-closed guard, which is narrower than the dedicated tests. 1. tests/fix-2284-hermes-agent-delegate-task-projection.test.cjs asserts that the INSTALLED code-review-fix.md contains no `model=` anywhere outside a string literal — masked whole-file, not merely inside call spans. The block's backticked `model=` survived the mask. The assertion is right: on Hermes the projection strips the parameter because delegate_task has no per-call model at all, so instructing the orchestrator to "omit the model= parameter" is advice about a parameter that does not exist there. The block now says "the `model` parameter" and carries no bare `model=` token. 2. tests/prompt-injection-scan.security.test.cjs flagged quick.md at 50,164 normalized chars against a 50,000 prompt-stuffing threshold. quick.md sits just under the line on next, so any insertion trips it — the situation review.md is already documented for in SIZE_ONLY_WORKFLOWS ("sat at 49,971 chars — 29 below the threshold — so it was going to trip on whatever was added to it next"). quick.md joins it with the same justification. This is a size-finding exemption only: the file is still fully injection scanned, and every other security check still runs on it. Because the canonical block can no longer carry a literal `model=`, the guard's detector now accepts the `<!-- #2517 model-omit-on-inherit -->` marker as the canonical signal, falling back to the inline-prose property for the four files that predate it (plan-phase, execute-phase, scan, ship — all four match the legacy branch). That is strictly stronger than word-proximity matching, and it keeps the guard a property check rather than a template match. Verified: derived guard 19/19 with 0 missing; the #2684 binding guard unchanged at 19 files / 60 placeholders / 0 findings; no inserted block contains a bare model= token; the masked-projection assertion passes for code-review-fix.md; gen:golden exits 0 across all 19 runtimes; lint:ci exits 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2711): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6fb832a93d |
fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve (#2710)
* test(#2684): failing-first guard for unbound model= dispatch placeholders Extends the #2517 omit-on-inherit bucket with a behavioral binding guard: every model="{X}" in a workflow must name a field that workflow actually binds — an init-payload key (queried for real), a shell assignment, or a declared parse field. scan.md ({resolved_model}) and ship.md ({balanced_model}) substitute names nothing emits, so the orchestrator invents the value (ADR-1411). Also pins the shipped reference that seeded the placeholder and instructs the #2517-forbidden model="inherit". RED expected on scan.md, ship.md, and references/model-profile-resolution.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve scan.md:85 passed model="{resolved_model}" and ship.md:486 passed model="{balanced_model}" — names neither workflow's init payload emits. init.map-codebase emits mapper_model; init.phase-op emits no model field at all. With no source for the substitution the orchestrator invents a value, so model_overrides/model_policy are silently inert at both sites — the invisible partial application ADR-1411 prohibits. - scan.md: declare the init fields in a Parse JSON line and substitute {mapper_model}, matching its sibling map-codebase.md. - ship.md: ref.agent is only known at runtime, so resolve it per hook via query resolve-model and dispatch with {HOOK_AGENT_MODEL}, following the same NAME=$(...) → model="{NAME}" convention every other shell-resolved dispatch in the corpus uses (code-review.md, secure-phase.md, ui-phase.md). - Both sites now carry the #2517 rule: omit model= entirely when the resolved value is "inherit" or empty. A bare rename would have traded a dangling placeholder for model="", which 404s on non-Claude runtimes — an agent type absent from the profile table resolves to the empty string, which is exactly ship.md's case. - references/model-profile-resolution.md was the seam: it shipped the copy-pasteable {resolved_model} snippet scan.md inherited, used the stale Task( spelling, and instructed passing model="inherit" outright. Rewritten to teach the real binding convention and the omit rule. Two defects surfaced while fixing this and fixed inline rather than deferred: 1. ref.agent originates in a capability manifest, which may be third-party. Resolving it at runtime made this the first place that value reaches a shell command, so ship.md now validates its shape before interpolating and skips the hook when it fails — matching code-review.md's existing defense-in-depth pattern. Covered by a test that runs the shipped regex against real agent names and injection payloads. 2. The reference doc's omit example first placed a literal model= inside an Agent(...) comment, which tripped the #2284 fail-closed Hermes projection guard (bin/install.js:3704) and refused the install outright. Moved out of the call span; gen:golden is green across all 19 runtimes. Guard tests extend the existing #2517 bucket: every model="{X}" in a workflow must name a field that workflow binds — an init-payload key queried for real, a shell assignment, or a declared parse field. Behavior change (Hyrum's Law): the scan mapper now runs on the catalog-resolved model rather than the session model. The omit path is unchanged. Disclosed in the changeset body. Fixes #2684 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2684): validate capability-supplied ref.agent in-context, not in the shell The first cut of the ref.agent guard was placed after the injection point it was meant to close. It instructed the orchestrator to substitute the untrusted manifest value into a shell assignment and test it there: HOOK_AGENT="<the ref.agent value>" if [[ "$HOOK_AGENT" =~ ^[A-Za-z0-9][A-Za-z0-9._-]*$ ]]; then … Substitution happens before bash parses anything, so a manifest supplying `x"; touch /tmp/pwned; echo "` yields three statements and runs the middle one unconditionally — the regex fires afterwards and protects nothing. ship.md now requires the check to run in-context, the same way the workflow already reads activeHooks ("do NOT pipe it through a shell parser"), and to skip the hook outright on a mismatch. Only a value that has already matched ^[A-Za-z0-9][A-Za-z0-9._-]*$ ever reaches a command line. The guard test asserts the ordering — the in-context requirement and the absence of any raw shell assignment — not just that the pattern rejects metacharacters, since a pattern alone was exactly what gave false assurance here. Also corrects the empty-string explanation in references/model-profile- resolution.md. model_profile:"inherit" resolves to the literal "inherit", not "" (model-resolver.cts:395), and an unknown agent takes the empty-string path via resolve_model_ids:"omit" rather than by absence alone — the doc claimed all three produced "". The workflow instructions were already correct (omit on "inherit" OR empty); only the rationale in the citable reference was wrong. Both found by the isolated adversarial review pass for #2684. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2684): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c44241a0b |
fix(#2605): make dropped local-server reviewer lanes loud instead of silent (#2689)
* fix(#2605): make dropped lm_studio / llama_cpp reviewer lanes loud, not silent The `lm_studio` and `llama_cpp` reviewer legs in `gsd-core/workflows/review.md` carried the same empty-output defect as the claude/gemini legs (#2494, fixed in #2592) in a worse variant: when the local endpoint was unreachable or returned empty content, they wrote NOTHING to `{run_dir}/gsd-review-<leg>.md`. There was no `[ ! -s … ]` stub at all, so the file never existed, `write_reviews` omitted that reviewer's section, and the outcome was indistinguishable from the reviewer never having been selected. Two diagnostic holes are closed, because an OpenAI-compatible server fails in two ways that leave evidence in different places: - Transport failure (endpoint unreachable): curl writes to stderr and exits non-zero. Both legs used `curl -s`, which suppresses curl's ERROR text as well as the progress meter, and then discarded stderr to `/dev/null` — so there was nothing to capture even in principle. Now `-sS` with stderr to a `.err` sidecar, matching the claude/gemini/codex legs. - Application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in the response BODY, so stderr is empty and only the body is evidence. The stub appends the raw response. The `llama_cpp` leg additionally piped curl straight into `jq`, throwing the body away before anything could inspect it; the response is now captured to a variable first, as the `lm_studio` leg already did. The existing `>&2` warning is kept and now points at the stub file. Failing-first verified by extracting both shipped blocks and running them under a real bash against a stubbed curl: pre-fix, all three failure modes produce NO review file and no `.err` sidecar for both legs; post-fix, each produces a diagnosable stub, and a successful review still passes through untouched. Closes #2605 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2605): close the remaining silent-drop paths the review surfaced Four defects found by the orthogonal review of the first commit, all fixed here rather than deferred (they are the same defect class the issue is about, and three of them sit inside the code that commit touched). 1. CodeRabbit leg was still unguarded. `coderabbit review --prompt-only 2>/dev/null > file` with no `[ ! -s … ]` stub — the last leg still shaped like pre-#2494 code. A missing or unauthenticated binary left a zero-byte file that write_reviews rendered as "ran cleanly, nothing to report". It now captures stderr to a `.err` sidecar and emits the same diagnosable stub as every other leg. This is the leg the next issue in the #2494 -> #2592 -> #2605 series would have been about. 2. The budget-skip path dropped the lane just as silently. When `prepare_trimmed_prompt_for_reviewer` fails, `*_SKIP=1` bypasses the entire block — guard included — so no file was written and the only trace was a stderr warning nothing persists. All three local-server legs now write a "review skipped: prompt budget too small" stub on that path. 3. Whitespace-only replies evaded the guard. `[ ! -s … ]` counts BYTES, and command substitution strips trailing newlines but not spaces, so a reply of `" "` was written out and passed as a "successful" but vacuous review — the same indistinguishable-from-success outcome the guard exists to prevent. A `case` glob now normalizes whitespace-only content to empty. 4. `echo "$VAR"` swallowed option-like content. bash's builtin `echo` treats a value of exactly `-n`/`-e`/`-E` as a flag and writes 0 bytes, which would trip the empty guard and DISCARD a genuine reply. Switched to `printf '%s\n'`, the idiom the OpenCode leg in this same file already uses for this reason. Also brings the Ollama leg to parity while it is in hand: it always emitted a stub so it never silently vanished, but it was the least diagnosable of the three local-server legs — bare `-s`, stderr to /dev/null, and the response piped straight into jq so the error body was discarded unread. Verified by extracting all four shipped blocks and running them under a real bash against stubbed CLIs: 22 cases (7 per local-server leg x 3, plus CodeRabbit) all produce the contracted output, and a successful review still passes through untouched on every leg. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2605): regenerate golden install-parity fixtures for the review.md edit The remote test run on the prior commit failed with 19 "golden parity — <runtime>" mismatches. review.md is installed into every runtime's tree, so editing it changes its content hash in all 19 golden fixtures. Regenerated with `npm run gen:golden`; the diff is exactly one line per fixture — the gsd-core/workflows/review.md hash — and nothing else. This is the second ripple of a workflow edit, alongside tests/workflow-size-baseline.json which the first commit already updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2605): backfill changeset PR number (#2689) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c07734216f |
fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19 installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped with no citation and no evidence quote. Sourcing every axis independently against Kimi Code CLI's own docs (the issue's explicit requirement — `kimi` and `kimi-code` are distinct products) showed three values had been inherited from the Python `kimi` descriptor rather than sourced: - `embeddingMode` imperative -> declarative. Kimi Code plugins are a `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic API (docs/en/customization/plugins.md) — the same shape as `codex`. - `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md). The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not. - `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth bound is published, so the fail-closed sentinel applies over a guessed integer. `dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is registered with the host and `resolveDispatchType` maps every role onto coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance under Documentation gaps instead. Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/ backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499), and the install adapter is not selected by `embeddingMode` (install.js:543 always uses the imperative adapter). The one visible effect is the curated profile pin, which moves programmatic-cli -> declarative-cli. Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit member that has been in the closed vocabulary since kimi-code shipped. Same defect class and countermeasure as #2598: pin the corrected values and require the matrix to agree with the descriptor, because a descriptor/matrix disagreement is how the gap survived. Closes #2603 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed Surfaced by the orthogonal review of this change. `negotiateHostCapabilities` emits a sentinel-specific warning for every dispatch sub-axis carrying the documented `undocumented` value — namedDispatch, nested, background, subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell through to the numeric guard and reported `host dispatch.maxDepth is missing or not a number — treating as 0`. That message is indistinguishable from a genuinely malformed descriptor, so a correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this PR's kimi-code correction adds a seventh, which is why it is fixed here rather than left in place. The numeric guard keeps firing for genuinely malformed values; both paths still degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests, including the boundary case that the sentinel carve-out must not swallow a real malformed value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2603): backfill changeset PR number (#2687) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a3853472de |
fix(#2598): declare OpenCode subagent dispatch synchronous, not background (#2682)
* fix(#2598): declare OpenCode subagent dispatch synchronous, not background capabilities/opencode/capability.json advertised dispatch.background: true and dispatch.backgroundDispatch: true. negotiateHostCapabilities and every degradationFor / shouldFlattenDispatch consumer trusts these per-field values, so declaring a capability the host lacks OVERSTATES it — the opposite of the fail-closed posture the negotiation is built for. The issue's own citations needed checking before acting: the host-integration matrix (ADR-1239's designated deployment source-of-truth) documented `true` with NEWER evidence than the issue cited, and explicitly marked the issue's sst/opencode#5887 reference as a stale snapshot superseded by #2087. git log confirms #2087 deliberately flipped these from false to true, citing OpenCode v1.15.0/v1.17 as "background subagents enabled by default in all modes". Applying the issue as filed would, on that evidence, have REGRESSED a deliberate update. So the claim was verified against current upstream rather than either document. `packages/opencode/src/effect/runtime-flags.ts` on `dev` today reads: experimentalBackgroundSubagents: enabledByExperimental("OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS") `enabledByExperimental` falls back to the `experimental` flag and `bool()` defaults to false — the parameter is hidden from the model unless an operator opts in by env var. Upstream #29638 is still OPEN and confirms the session loop `tasks.pop()`s one subtask at a time. #2087's "default-on in all modes" reading does not hold against current dev. The issue's CONCLUSION is therefore right even though part of its evidence was superseded: concurrent dispatch cannot be relied on, so both fields are false. The matrix rows are corrected with the verified citation rather than reverted to the old #5887 quote, so the record shows why the value is false TODAY rather than re-asserting evidence that was legitimately superseded. Neighbouring sub-fields are untouched and pinned by test: namedDispatch, subagentToolkit, and isolation:'orchestrator-worktree' (which works via `opencode run --dir` at the OS process level and is unaffected — #2584 does not depend on this value either way). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2598): re-pin the dispatch contract tests to synchronous OpenCode dispatch gsd-test on the descriptor change came back FAILED (5 unique, both node versions). The failures were not incidental — they were deliberate contract-pin tests encoding #2087's decision, one named literally "background UPGRADE": tests/host-integration-descriptors.test.cjs - EXPECTED_FLATTEN[opencode] === false (background-eligible) - the derived background-eligible set pin tests/opencode-imperative-reference.test.cjs - "descriptor declares background dispatch true/true (v1.15/v1.17 upgrade)" - "background UPGRADE changes shouldFlattenDispatch: false now" So this is a recorded decision being reversed, not drift being corrected, and it is reversed on evidence: current upstream `dev` gates the capability behind OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS (default false) and upstream #29638 (OPEN) confirms the session loop still handles one subtask at a time. The issue is filed by the maintainer and explicitly directs "update golden-parity / validator fixtures as needed", which sanctions re-pinning. Behavioral consequence, verified: shouldFlattenDispatch(opencode) now returns TRUE, so GSD serializes opencode dispatch instead of trusting concurrency it cannot get. That is the correct fail-closed direction and is safe today — no shipped GSD flow drives OpenCode background waves (per the issue), and isolation:'orchestrator-worktree' is unaffected because it works at the OS process level via `opencode run --dir`, not via the native subagent. Each re-pinned test now asserts the retracted contract in the opposite direction — feeding the #2087 axes back in must still yield "would not flatten" — so a silent re-flip of either field is caught rather than merely un-asserted. lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2598): backfill changeset pr number (#2682) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d08c32048 |
fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable Every emitted script was rejected. Four invalid constructs, the first fatal on its own, so the Workflow backend could never dispatch a wave: 1. no `export const meta = {…}` first statement -> whole script rejected 2. resumeFromRunId("<id>") -> "resumeFromRunId is not defined". It is a Workflow TOOL INPUT parameter, not a script function. The run id still reaches the caller via summary.resumeRunId, to pass as that input. 3. budget(<n>) -> "budget is not a function". `budget` is a read-only object { total, spent(), remaining() } fed by the caller's token directive; a script cannot set it. Recorded as intent in a comment. 4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions". Now parallel([() => agent(…), …]) — passing agent() results directly also started every agent eagerly, before parallel() could bound concurrency. The single-plan stage had its own branch with the same parallel() defect; both branches are now one array-emitting path. Waves also emit phase() calls whose titles match meta.phases exactly, so progress groups correctly. Two secondary defects kept the script from ever being REACHED — which is why this shipped undetected: 5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no scriptable way" to introspect it and told callers to omit the flag, so gate 5 returned agent_sdk_version_unknown on every automated run while `capability state` still reported active:true. True for bash, false for Node: the router now reads the installed @anthropic-ai/claude-agent-sdk version, walking node_modules up the tree and reading package.json directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the SDK's exports map does not expose ./package.json. Precedence: explicit flag > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an unresolvable version still declines to inline. A too-old SDK now reports the truthful agent_sdk_version_below_floor instead of unknown. 6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any invocation without --runtime reported runtime_not_claude on an ordinary Claude project. Now delegates to runtime-slash.resolveRuntime. The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` snippet is removed rather than repaired: it was also shell-dependent — zsh does not word-split unquoted parameter expansions, so it collapsed to a single argv element, argValue() never matched, and the run failed into the same agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto- resolution removes the need for the construct entirely. Verified with the issue's own repro: no flags now reaches the version gate; an SDK above the floor yields backend:"workflow" with a script that parses as a real ES module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids Findings from the isolated review, all fixed. HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was already RED because of it. The registry embeds the fragment text INLINE, so the shipped/installed copy still taught the exact broken contract this PR fixes: the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown" guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of checking its exit code, so I recorded a red chain as green — checking $? now.) HIGH — three existing tests asserted the OLD broken shape and would have failed CI; none was touched by the first commit: tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…") tests/claude-orchestration.test.cjs — .includes('budget(') tests/claude-orchestration-command-router.test.cjs — .includes('budget(') Each now asserts the corrected contract: the id/pool reaches the caller via summary, and neither construct is ever CALLED. Two sibling assertions had also gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the new explanatory COMMENT contains that substring, not because anything is wired. Rewritten to assert the real property. MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked within a wave, but nothing checked wave ids across waves. That was harmless before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a matching meta.phases entry and the tool matches titles by exact string — two waves sharing an id would collapse into one progress group and misattribute the second wave's agents to the first. Rejected at validation, with tests either side of the boundary. MEDIUM — the fragment contradicted itself (its "Manifest construction" header still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never told the orchestrator to pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to a tool-invocation input, an implementer following only the fragment would have silently regressed phase-resume to a no-op. Both fixed. MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and docs/explanation/claude-orchestration-capability.md documented `resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output — teaching the bug as the feature. Updated to the real contract, including the required meta block and the thunk-array parallel() form. (The changeset is `Fixed`, so the docs gate exempts this; it is corrected because it is wrong, not because a gate demanded it.) LOW — the router's top-of-file comment still described the divergent `--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the fix; and inserting resolveInstalledAgentSdkVersion had orphaned resolveDetectionArgs' JSDoc above the wrong function. Both repaired. lint:ci now exits 0 (verified by exit code, not by reading output). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2590): backfill changeset pr number (#2681) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
920a5f3f06 |
fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep (#2673)
* fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep The reviewer/workflow config lookups resolved scalars and object fields with a `gsd_run query <cmd> … | jq … 2>/dev/null || <default>` shape. On any machine without jq (the default on Windows/Git-Bash) the jq stage fails with exit 127, the failure is swallowed by 2>/dev/null + the trailing || default, and the variable comes back EMPTY — the configured per-lane model/host/budget is silently dropped and the lane falls back to CLI defaults with no diagnostic. gsd-tools ships native flags that do the same job with no external dep: config-get <key> --raw (strips JSON quotes off a scalar) resolve-model <id> --pick model (descends an object) resolve-execution … --pick <f> (same) verification.status … --pick status Replaced every jq-piped config/model/verify lookup across review.md (×23), plan-phase.md, ship.md, debug.md (incl. the redundant boolean coercion — --raw returns true/false as bare tokens natively), autonomous.md (×2), ai-integration-phase.md (×4), and eval-review.md. The legitimate structured-JSON jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile) are untouched — only the jq-replaceable lookups moved to the native flags. Adds tests/fix-2589-config-get-no-jq.test.cjs: a source-invariant guard asserting no audited workflow pipes config-get/resolve-model/resolve-execution/verification.status to jq (fails-first on the pre-fix text, passes after). * test(#2589): update autonomous-converge jq assertion to --pick; regen golden fixtures Two test consequences of the workflow-doc edits in the prior commit: 1. tests/autonomous-converge.test.cjs pinned the OLD jq-dependent shape (`verification.status … | jq -r '.status//empty'`) as the canonical routing contract. The test's INTENT is correct (route human validation through canonical verification.status) but it over-specified the MECHANISM (the jq pipe). Updated the assertion to match the new native --pick status shape; the contract being guarded (canonical verification.status read before the human_needed branch) is unchanged. 2. The golden-install-parity fixtures (19 runtimes) record a content hash of every installed workflow .md; the 7 edited workflows changed those hashes. Regenerated via `npm run gen:golden` (the test's own failure message instructs this). Only the 7 edited workflow hashes changed in each fixture. * fix(#2589): declare jq a prerequisite for the lanes that still need it; repair test file Three defects in the first cut of the #2589 fix: 1. tests/autonomous-converge.test.cjs was a JavaScript syntax error. The regex literal /...2>\/dev/null .../ left the second slash unescaped, terminating the literal early and parsing `null` as regex flags: SyntaxError: Invalid regular expression flags The whole file failed to load, so every assertion in it — including the #1522 and #1526 guards — silently stopped running. Replaced with the string-compare form already used at line 202 for the sibling shell-snippet assertion. 2. lint:ci failed. tests/fix-2589-config-get-no-jq.test.cjs buckets into the capped `config` production module via its `config-get-...` effective prefix, making it a novel offender against the 2-file cap. The test is about workflow documents, not the config module, so it is renamed to fix-2589-workflow-jq-dependency.test.cjs (free prefix) rather than growing the allowlist with a module that does not actually need a 5th test file. 3. The fix deleted the repo's only jq-prerequisite declaration. review.md:244 ("install jq if missing") was the anchor plan-review-convergence.md cites by line number, and it went away with the jq pipes — while the ollama, lm_studio, llama_cpp, opencode, and agy lanes still hard-require jq to parse HTTP /v1/chat/completions responses, opencode's JSONL event stream, and agy's conversation cache. On a jq-less host those five lanes swallow exit 127 into empty output: the same silent-degradation class #2589 exists to close. detect_clis now probes jq alongside the other prerequisites and emits jq:available / jq:missing, and the five dependent lanes are treated as undetected when it is absent, with an install hint. The six lanes that do not need jq stay selectable. plan-review-convergence.md now cites the section by name instead of a line number that moves. Regression guards added to the renamed test file: review.md must keep the jq probe and must name all five dependent lanes, and no workflow may cite review.md by line number. Workflow-size baseline and the 19 install-parity goldens regenerated for the review.md / plan-review-convergence.md edits. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2589): decide the ship verification gate on a single verification.status read Isolated review finding (medium). Pre-fix, ship.md captured verification.status ONCE into $VERIFICATION and picked status / next_action / next_command off that cached JSON with three jq calls. --pick takes a single dot-path field, so the mechanical conversion issued three separate queries up front: three node spawns that each re-read the phase VERIFICATION.md and re-derive the commit-time vs mtime staleness comparison, on every ship — including the common passing path that never uses the two message fields. It also meant the gate's verdict and the message shown to the user were derived from three reads with no guarantee they observed the same state. The gate now reads `status` once and decides. The two message-only fields are read on the blocking path only, after PHASE_VERIFICATION_INCOMPLETE is already determined — so the passing path costs one query instead of three, and a concurrent write between reads can no longer make the gate and its message disagree, because the block/allow decision no longer depends on them. Adding a multi-field --pick to gsd-tools would have collapsed this to one query, but that changes the flag's output contract and belongs in its own change. Regression guard in tests/fix-2589-workflow-jq-dependency.test.cjs: ship.md must read verification.status exactly three times total, the block decision must follow the status read, and next_action / next_command must both appear after the blocking prose so they cannot drift back onto the passing path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * docs(#2589): document the jq prerequisite for the five reviewer lanes that need it /gsd-review's ollama, lm_studio, llama_cpp, opencode, and agy lanes parse JSON GSD does not produce (OpenAI-compatible /v1/chat/completions responses, OpenCode's JSONL event stream, Antigravity's conversation cache), so they require jq on PATH. Nothing in docs/ said so. Records which five lanes need it, which six do not, that reading configured models/hosts/budgets no longer requires jq at all, and what /gsd-review now does when jq is absent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2589): backfill changeset pr number (#2673) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
89b673e40d |
feat(#2631): planner emits estimate and plan-checker surfaces the over-budget flag (#2670)
* test(#2631): failing-first planner estimate emission and over-budget surfacing * feat(#2631): emit plan estimate and surface the over-budget split recommendation * fix(#2631): extract sizing prose to references to fit planner and plan-phase caps * fix(#2631): move estimate check to plan-checker; fix template regex and caps * fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble * fix(#2631): invoke estimate-check after the launcher preamble in plan-checker * fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference * chore(#2631): backfill changeset pr to 2670 * chore(#2631): backfill changeset pr to 2670 |
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
b897070de3 |
fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard (#2656)
* fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard The tracked build artifact gsd-core/bin/lib/api-coverage.cjs had drifted four days behind src/api-coverage.cts: PR 2551 landed the #2366 fix in the source without regenerating the compiled output, so the module that actually ships still carried none of it. Regenerate the artifact, and add scripts/lint-compiled-artifact-sync.cjs to lint:generated-sync so a tracked compiled artifact can never again silently diverge from its source. The check derives its file set from git ls-files rather than a hand-maintained list, and is regime-agnostic: if these artifacts are later untracked and gitignored per ADR-457, the tracked set becomes empty and the check passes trivially. Verified fail-first: the guard exits 1 against the previously-committed artifact (37731 bytes vs 38634 expected) and 0 after regeneration. Also fixes a defect this change surfaced in tests/no-phantom-issue-refs: its PHANTOM list still banned 2551 and 2361, but GitHub numbers issues and PRs from one shared counter and this repo has since reached 2654, so both now resolve to merged PRs. The guard was rejecting accurate citations of them — it failed this very commit for naming PR 2551 as the drift's provenance. Verified by replaying the guard's scan: 1 offender under the old list, 0 under the pruned one. Only 3182 is still a 404. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2653): backfill changeset pr number to 2656 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a40ee8a5e7 |
fix(#2479): drop the codex hook-trust bypass flag and its capability probe (#2536)
* fix(#2479): drop the codex hook-trust bypass flag and its capability probe The /gsd-review codex lane emitted a hook-trust bypass flag via a capability-probed variable (#1115). Host-harness safety classifiers deny commands carrying the flag (23/33 sampled invocations), one denial citing the probe itself as intent, while flagless retries succeeded 32/32 — the flag only bypasses persisted hook trust, a first-run condition with no steady-state value. Remove both the flag and the probe per maintainer direction (no config key). #1115's diagnosability half — stderr to .err, folded into the lane on empty output — is untouched; its version-gate becomes vacuous with no flag to gate. The regression test inverts: the literal flag is now banned file-wide in review.md (covers continuation lines, carrier variables, and probes), alongside bans on the carrier variable and any codex help-grep probe shape. Fixes #2479 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2479): add changeset for PR #2536 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2479): changeset body in house format (bold lead) + post-rebase fixture regen Review round 2 Minor: wrap the changeset lead clause in the required **bold** span (reviewer-supplied text, applied verbatim). Rebased onto current next; golden-install-parity (19 runtimes) + size baseline regenerated — delta confined to review.md's hash/size. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6ee4349272 |
fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) * chore(#2537): backfill changeset pr to 2642 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b97b3dad21 |
fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix) (#2634)
* test(#2517): guard plan/execute-phase omit model= when *_model inherit/empty * fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix) * chore(#2517): backfill changeset pr to 2634 |
||
|
|
c2b498e452 |
fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH (#2628)
* test(#2498): guard branch-creation uses --no-track (all workflows) * fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH * chore(#2498): backfill changeset pr to 2628 |
||
|
|
57b2bd8368 |
fix(#2491): finish todos/done -> todos/completed rename (14 stale refs) + guard (#2626)
* test(#2491): add todos/done rename under-sweep guard * fix(#2491): finish todos/done -> todos/completed rename (14 stale refs) * chore(#2491): backfill changeset pr to 2626 * test(#2491): fix lint-legacy-dir-name + allow-test-rule-refs (split legacy token, add issue ref) |
||
|
|
4a66d62d10 |
feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver
Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.
worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.
resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.
Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: rebuild tracked state-transition.cjs to match #2400 source
The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit
|
||
|
|
ec7978c0b4 | feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) | ||
|
|
bd18a593a9 |
fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs (#2592)
* fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs The gemini and claude blocks in gsd-core/workflows/review.md were the only two of the ten prompt-fed reviewer legs with both `2>/dev/null` and no empty-output guard. Any failure that wrote no stdout — CLI missing, unauthenticated, rate-limited, crashed — left a zero-byte review file with the only diagnostic evidence already discarded. write_reviews substitutes each file's raw content into a `## <Reviewer> Review` section with no marker separating "empty/failed" from "ran cleanly, nothing to report", so a failed lane silently degraded the advertised N-reviewer consensus to N-1 while present_results reported success. Both /gsd:review and /gsd:plan-review-convergence share the invoke_reviewers step, so both were affected. Both legs now redirect stderr to a `.err` sidecar and write a diagnostic stub with the captured stderr appended when the review file comes back empty — the same shape the codex and cursor legs already use. Scope limit, noted in the block comment: the guard only runs if the block itself completes. A host Bash-tool timeout that kills the whole block skips it, the same hard bound the OpenCode block already documents. The existing timeout guidance (#2194) covers that case and is unchanged. Regression test extracts the two dispatch blocks verbatim from the workflow and runs them under bash against a failing CLI stub, asserting the review file is non-empty and carries a diagnosable message plus the captured stderr. It fails against pre-fix review.md (4 of 5 cases). Golden install-parity fixtures and the workflow size baseline are regenerated: one review.md hash per fixture, one size entry. Fixes #2494 * fix(#2494): stamp the changeset fragment with the filed PR number The fragment was authored with the documented `pr: 0` placeholder because the PR number does not exist until `gh pr create` returns. Now that this PR is #2592, stamp it — scripts/changeset/parse.cjs requires pr > 0, so the placeholder would fail the Changeset Required check. |
||
|
|
3b15a1e3cc |
fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos (#2597)
* test(#2576): add failing regression for padded resolves_phase compare * fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos * chore(#2576): backfill changeset pr to 2597 * test(#2576): drop fast-check property test (cross-platform-fragile on Windows CI) |
||
|
|
7d298d6d4d |
fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too (#2561)
* test(#2474): update dispatch gate test for dual-gate behavior The #2772 test asserted the gate reads USE_WORKTREES_FOR_PLAN only. Update to accept the dual-gate (USE_WORKTREES + USE_WORKTREES_FOR_PLAN). * fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too The per-plan dispatch condition checked only USE_WORKTREES_FOR_PLAN (submodule-derived), ignoring the project-level USE_WORKTREES flag. Add USE_WORKTREES to the gate. Net-negative edit: compress two nearby prose lines to offset the added shell condition (93353 bytes, down from 93368). Closes #2474 * docs(#2474): backfill changeset PR number (2561) * fix: merge coverage gate into single-process check (#2474) The test:coverage:unit script chained two c8 invocations with &&: the first ran tests and wrote coverage data to .nyc_output/, the second read that data for per-file branch checks. On fast CI runners (ubuntu/24), the second process started before the filesystem flushed the first process's writes — a classic TOCTOU race that caused intermittent coverage gate failures. Replace the two-process chain with a single c8 invocation that generates both text and json-summary reports, followed by a Node script (scripts/check-coverage-gate.cjs) that reads the JSON summary once and checks both overall and per-file thresholds. No filesystem race is possible because the JSON report is fully written before the check script reads it. |
||
|
|
0ad3c5dd14 |
fix(#2279): refresh date stamps on map-codebase Update runs (#2550)
* fix(#2279): reword date stamping to overwrite existing dates on Update runs The map-codebase agent and workflow instructions only said to replace [YYYY-MM-DD] placeholders, but Update-path files already contain concrete dates from the prior run. Reword to SET the date stamps unconditionally, overwriting whatever date is already there. Closes #2279 * docs(#2279): backfill changeset PR number (2550) |