* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
1339 lines
62 KiB
JavaScript
1339 lines
62 KiB
JavaScript
// docs-guard-exempt: 'docs/DEVELOPMENT.md' is a synthetic files-list fixture entry, not read as content.
|
|
// allow-test-rule: source-text-is-the-product
|
|
// The workflow and agent .md files ARE the product: their text is loaded and
|
|
// executed/interpreted at runtime by the agent host. Testing that specific
|
|
// strings exist within these files tests the deployed contract, not an
|
|
// implementation detail. No runtime API exists to enumerate the label accept-
|
|
// list or filter-set definitions — the text IS the specification.
|
|
//
|
|
// Bug 1 (compute_file_scope) — The inline Node.js script embedded in the
|
|
// workflow .md is the parser. The test implements the identical parse logic as
|
|
// a pure JS function (mirroring lines 172-184 of code-review.md exactly) and
|
|
// asserts on its structured output. A separate docs-parity assertion checks
|
|
// that the workflow .md contains the hyphen-aware boundary regex and the
|
|
// em-dash/parenthetical stripping — both of which are the deployed contract.
|
|
//
|
|
// Bug 2 (present_results) — Tested both behaviourally (pure JS helper that
|
|
// mimics the grep|cut pipeline) and via docs-parity on the workflow .md text.
|
|
//
|
|
// Bugs 3 and reviewer contract — docs-parity only on agents/*.md: the filter-
|
|
// set definition and label-equivalence contract exist only as text in those
|
|
// files; there is no runtime enumeration API.
|
|
|
|
'use strict';
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const { runHook } = require('./helpers/process-seam.cjs');
|
|
const { toLegacyResult, gitOrThrow } = require('./helpers/git-fixture.cjs');
|
|
const { PROBE_TIMEOUT_MS, GIT_TIMEOUT_MS, HOOK_FANOUT_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
|
const { createTempDir, createTempGitProject, cleanup, readFileNormalized } = require('./helpers.cjs');
|
|
|
|
const ROOT = path.resolve(__dirname, '..');
|
|
const WORKFLOW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'code-review.md');
|
|
const PRE_PASS_STEP_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'code-review', 'steps', 'structural-pre-pass.md');
|
|
const FIXER_PATH = path.join(ROOT, 'agents', 'gsd-code-fixer.md');
|
|
const REVIEWER_PATH = path.join(ROOT, 'agents', 'gsd-code-reviewer.md');
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Pure-function implementation of the compute_file_scope Node script body.
|
|
// This mirrors the logic in code-review.md lines 172-184 exactly.
|
|
// If those lines change, this function must be updated in tandem (and the
|
|
// docs-parity assertions below will catch a mismatch at the regex level).
|
|
//
|
|
// #2666: the acceptance predicate accepts root-level paths (no `/`) and known
|
|
// extensionless build files (Dockerfile/Makefile/etc.), not only nested paths
|
|
// with a trailing extension. Prose bullets are rejected by the known-filename /
|
|
// has-extension distinction (plus the post-processing existence check backstop
|
|
// in the shipped workflow).
|
|
const KNOWN_EXTENSIONLESS_BUILD_FILES = new Set([
|
|
'dockerfile', 'containerfile', 'makefile', 'justfile', 'procfile',
|
|
]);
|
|
function isAcceptablePath(raw) {
|
|
// A trailing `.`+alphanumerics extension qualifies (root-level OR nested):
|
|
// package.json, renovate.json, .gitlab-ci.yml, AGENTS.md, app/foo.tsx
|
|
if (/\.[A-Za-z0-9]+$/.test(raw)) return true;
|
|
// Known extensionless build filename (basename, case-insensitive): Dockerfile, Makefile, …
|
|
const base = raw.split('/').pop();
|
|
if (KNOWN_EXTENSIONLESS_BUILD_FILES.has(base.toLowerCase())) return true;
|
|
return false;
|
|
}
|
|
function parseKeyFiles(yaml) {
|
|
const files = [];
|
|
let inSection = null;
|
|
for (const line of yaml.split('\n')) {
|
|
if (/^\s+created:/.test(line)) { inSection = 'created'; continue; }
|
|
if (/^\s+modified:/.test(line)) { inSection = 'modified'; continue; }
|
|
// Hyphen-aware boundary: reset inSection for ANY key: line (including key-decisions:, etc.)
|
|
if (/^\s*[\w-]+:/.test(line) && !/^\s*-/.test(line)) { inSection = null; continue; }
|
|
if (inSection && /^\s+-\s+(.+)/.test(line)) {
|
|
let raw = line.match(/^\s+-\s+(.+)/)[1].trim();
|
|
raw = raw.replace(/^['"]|['"]$/g, '');
|
|
// Order matters: parens BEFORE em-dash because em-dashes can appear inside parens
|
|
raw = raw.replace(/\s+\([^)]*\)\s*$/, '');
|
|
raw = raw.split(/\s+—\s/)[0].trim();
|
|
if (isAcceptablePath(raw)) {
|
|
files.push(raw);
|
|
}
|
|
}
|
|
}
|
|
return files;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Pure-function implementation of the present_results severity-label parser.
|
|
// Mirrors the grep -E "^\s*(critical|blocker):" | head -1 | cut -d: -f2 | xargs
|
|
// pipeline from code-review.md.
|
|
// ---------------------------------------------------------------------------
|
|
function parseFrontmatterCritical(frontmatter) {
|
|
const lines = frontmatter.split('\n');
|
|
const match = lines.find((l) => /^\s*(critical|blocker):/.test(l));
|
|
if (!match) return { critical: 0 };
|
|
const value = match.split(':').slice(1).join(':').trim();
|
|
return { critical: parseInt(value, 10) || 0 };
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// BUG 1 — SUMMARY parser: compute_file_scope must not bleed prose from
|
|
// hyphenated sections (key-decisions:, patterns-established:, etc.) into the
|
|
// file list, and must strip em-dash descriptions and parentheticals.
|
|
// ---------------------------------------------------------------------------
|
|
describe('Bug 1 — compute_file_scope SUMMARY parser', () => {
|
|
test('extracts only key-files.created and key-files.modified entries', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' created:',
|
|
' - app/foo.tsx',
|
|
' modified:',
|
|
' - lib/bar.ts',
|
|
'key-decisions:',
|
|
' - We chose RSC for performance reasons',
|
|
'patterns-established:',
|
|
' - Always validate at the boundary',
|
|
'requirements-completed:',
|
|
' - REQ-01 done',
|
|
].join('\n');
|
|
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files.sort(), ['app/foo.tsx', 'lib/bar.ts'].sort());
|
|
});
|
|
|
|
test('strips em-dash narrative from bullet: "app/foo.tsx — RSC catalogue with filters"', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' created:',
|
|
' - app/foo.tsx — RSC catalogue with topic/mode/date filters',
|
|
].join('\n');
|
|
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files, ['app/foo.tsx']);
|
|
});
|
|
|
|
test('strips parenthetical from bullet: "tests/bar.test.ts (122 lines — 17 assertions)"', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' created:',
|
|
' - tests/bar.test.ts (122 lines — 17 assertions)',
|
|
].join('\n');
|
|
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files, ['tests/bar.test.ts']);
|
|
});
|
|
|
|
test('hyphenated sections in any order produce identical results', () => {
|
|
const yamlA = [
|
|
'key-decisions:',
|
|
' - Some decision',
|
|
'key-files:',
|
|
' created:',
|
|
' - src/index.ts',
|
|
'patterns-established:',
|
|
' - Some pattern',
|
|
].join('\n');
|
|
|
|
const yamlB = [
|
|
'patterns-established:',
|
|
' - Some pattern',
|
|
'key-files:',
|
|
' created:',
|
|
' - src/index.ts',
|
|
'key-decisions:',
|
|
' - Some decision',
|
|
].join('\n');
|
|
|
|
assert.deepStrictEqual(parseKeyFiles(yamlA), parseKeyFiles(yamlB));
|
|
assert.deepStrictEqual(parseKeyFiles(yamlA), ['src/index.ts']);
|
|
});
|
|
|
|
test('prose-only bullets from key-decisions are never included in file list', () => {
|
|
const yaml = [
|
|
'key-decisions:',
|
|
' - We chose RSC for performance reasons',
|
|
' - Deferred auth to Phase 3',
|
|
'key-files:',
|
|
' created:',
|
|
' - app/page.tsx',
|
|
].join('\n');
|
|
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files, ['app/page.tsx']);
|
|
});
|
|
|
|
// #2666 — the Tier-2 extractor must NOT drop repository-root files (no `/`)
|
|
// or known extensionless build files. Pre-fix the buggy predicate
|
|
// `/\//.test(raw) && /\.[A-Za-z0-9]+$/.test(raw)` dropped every root-level
|
|
// path and every extensionless build file anywhere in the tree.
|
|
test('#2666 RED: root-level files with extensions are accepted (package.json, renovate.json, .gitlab-ci.yml, AGENTS.md)', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' modified:',
|
|
' - package.json',
|
|
' - renovate.json',
|
|
' - .gitlab-ci.yml',
|
|
' - AGENTS.md',
|
|
' - CLAUDE.md',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(
|
|
files.sort(),
|
|
['.gitlab-ci.yml', 'AGENTS.md', 'CLAUDE.md', 'package.json', 'renovate.json'],
|
|
'root-level files with extensions must not be dropped for lacking a directory separator',
|
|
);
|
|
});
|
|
|
|
test('#2666: nested extensionless build files are accepted (docker/Dockerfile, web/Makefile)', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' modified:',
|
|
' - docker/Dockerfile',
|
|
' - web/Makefile',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files.sort(), ['docker/Dockerfile', 'web/Makefile']);
|
|
});
|
|
|
|
test('#2666: root-level extensionless build files are accepted (Dockerfile, Makefile, Justfile, Containerfile, Procfile)', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' created:',
|
|
' - Dockerfile',
|
|
' - Makefile',
|
|
' - Justfile',
|
|
' - Containerfile',
|
|
' - Procfile',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(
|
|
files.sort(),
|
|
['Containerfile', 'Dockerfile', 'Justfile', 'Makefile', 'Procfile'],
|
|
);
|
|
});
|
|
|
|
test('#2666 acceptance #1: the reporter 10-file Docker+CI phase yields all 10 paths', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' created:',
|
|
' - Dockerfile',
|
|
' - .gitlab-ci.yml',
|
|
' - renovate.json',
|
|
' - AGENTS.md',
|
|
' - CLAUDE.md',
|
|
' - docs/DEVELOPMENT.md',
|
|
' - scripts/version-consistency-gate.mjs',
|
|
' - web/package.json',
|
|
' - web/version_management.md',
|
|
' - web/update-version.cjs',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(
|
|
files.sort(),
|
|
[
|
|
'.gitlab-ci.yml', 'AGENTS.md', 'CLAUDE.md', 'Dockerfile',
|
|
'docs/DEVELOPMENT.md', 'renovate.json', 'scripts/version-consistency-gate.mjs',
|
|
'web/package.json', 'web/update-version.cjs', 'web/version_management.md',
|
|
],
|
|
'the full reporter phase must scope all 10 files, including Dockerfile + root files',
|
|
);
|
|
});
|
|
|
|
test('#2666 negative-space: a path-like prose bullet with no extension and unknown basename is rejected', () => {
|
|
// `topic/mode/date filters` has a `/` but no extension and an unknown basename —
|
|
// the pre-fix predicate dropped it (good), the relaxed predicate must STILL drop it.
|
|
const yaml = [
|
|
'key-decisions:',
|
|
' - topic/mode/date filters',
|
|
'key-files:',
|
|
' created:',
|
|
' - app/page.tsx',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files, ['app/page.tsx']);
|
|
});
|
|
|
|
test('#2666 negative-space: em-dash/parenthetical stripping still works on an accepted root file', () => {
|
|
const yaml = [
|
|
'key-files:',
|
|
' modified:',
|
|
' - Dockerfile — multi-stage build',
|
|
].join('\n');
|
|
const files = parseKeyFiles(yaml);
|
|
assert.deepStrictEqual(files, ['Dockerfile']);
|
|
});
|
|
|
|
// Docs-parity: the workflow .md must contain the hyphen-aware boundary regex
|
|
// so what we tested above is actually what is deployed.
|
|
test('code-review.md contains hyphen-aware boundary regex [\\w-]+', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
// Locate the Node script block in the compute_file_scope step
|
|
const scriptStart = src.indexOf('const files = [];');
|
|
assert.ok(scriptStart !== -1, 'compute_file_scope script must contain "const files = [];"');
|
|
const scriptEnd = src.indexOf('if (files.length)', scriptStart);
|
|
const scriptSection = src.slice(scriptStart, scriptEnd);
|
|
// Must use [\\w-]+ (hyphen-aware) not \\w+ only
|
|
const hasHyphenAwareRegex = scriptSection.includes('[\\\\w-]') || scriptSection.includes('[\\w-]');
|
|
assert.ok(
|
|
hasHyphenAwareRegex,
|
|
'compute_file_scope boundary regex must be hyphen-aware ([\\w-]+), found section:\n' + scriptSection
|
|
);
|
|
});
|
|
|
|
// Docs-parity: the workflow .md must contain the em-dash and parenthetical stripping.
|
|
test('code-review.md contains em-dash split and parenthetical strip in script body', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
const scriptStart = src.indexOf('const files = [];');
|
|
const scriptEnd = src.indexOf('if (files.length)', scriptStart);
|
|
const scriptSection = src.slice(scriptStart, scriptEnd);
|
|
assert.ok(
|
|
scriptSection.includes('replace(/\\s+\\([^)]*\\)\\s*$/, \'\')'),
|
|
'Script must strip parentheticals with replace(/\\s+\\([^)]*\\)\\s*$/, \'\')'
|
|
);
|
|
assert.ok(
|
|
scriptSection.includes('split(/\\s+—\\s'),
|
|
'Script must split on em-dash to strip narrative'
|
|
);
|
|
});
|
|
|
|
// #2666 docs-parity: the shipped workflow must NOT still carry the buggy
|
|
// AND-joined predicate that required BOTH a `/` and a trailing extension —
|
|
// that predicate dropped every root-level file and every extensionless build
|
|
// file. Catches a revert of the #2666 fix.
|
|
test('#2666 docs-parity: compute_file_scope does not contain the buggy slash-and-extension predicate', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
const scriptStart = src.indexOf('const files = [];');
|
|
const scriptEnd = src.indexOf('if (files.length)', scriptStart);
|
|
const scriptSection = src.slice(scriptStart, scriptEnd);
|
|
assert.ok(
|
|
!scriptSection.includes('/\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw)'),
|
|
'compute_file_scope must not use the buggy AND-joined /\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw) ' +
|
|
'predicate (#2666) — it drops every root-level and extensionless build file. Found section:\n' +
|
|
scriptSection
|
|
);
|
|
});
|
|
|
|
// #2666 docs-parity: the shipped workflow must reference the known
|
|
// extensionless build filenames so Dockerfile/Makefile/etc. are accepted.
|
|
test('#2666 docs-parity: compute_file_scope accepts known extensionless build files (Dockerfile)', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
const scriptStart = src.indexOf('const files = [];');
|
|
const scriptEnd = src.indexOf('if (files.length)', scriptStart);
|
|
const scriptSection = src.slice(scriptStart, scriptEnd);
|
|
assert.ok(
|
|
/dockerfile/i.test(scriptSection),
|
|
'compute_file_scope must reference known extensionless build filenames (e.g. Dockerfile) ' +
|
|
'so they are not dropped (#2666). Found section:\n' + scriptSection
|
|
);
|
|
});
|
|
|
|
// #2666 docs-parity: the Tier-3 git-diff fallback must intersect with the
|
|
// SUMMARY scope and warn on dropped files (not only fire on zero Tier-2 hits).
|
|
test('#2666 docs-parity: Tier-3 intersects/warns against git diff --name-only', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
// The shipped workflow must compute git diff --name-only AND emit a warning
|
|
// when the diff contains files the SUMMARY extractor did not surface.
|
|
assert.ok(
|
|
src.includes('git diff --name-only'),
|
|
'code-review.md must run `git diff --name-only` to cross-check the SUMMARY scope (#2666)'
|
|
);
|
|
assert.ok(
|
|
/warn|missing|not surfaced|did not|not in/i.test(src),
|
|
'code-review.md must warn when git diff contains files the SUMMARY extractor dropped (#2666)'
|
|
);
|
|
});
|
|
|
|
// #2666 docs-parity: the membership test must be EXACT whole-line matching
|
|
// (grep -Fxq), not an unanchored `case` substring match — otherwise a short
|
|
// basename in the diff (root `Dockerfile`) substring-matches a longer scoped
|
|
// path (`docker/Dockerfile`) and is silently skipped, reintroducing the bug.
|
|
test('#2666 docs-parity: Tier-3 cross-check uses exact whole-line matching (grep -Fxq), not substring case', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
assert.ok(
|
|
src.includes('grep -Fxq'),
|
|
'code-review.md Tier-3 cross-check must use grep -Fxq (exact whole-line match) for membership ' +
|
|
'testing, not an unanchored `case` substring match that would skip a root `Dockerfile` ' +
|
|
'whose name appears as a suffix of an already-scoped `docker/Dockerfile` (#2666)'
|
|
);
|
|
// The unanchored substring `case "$IN_SCOPE" in` membership test must NOT be
|
|
// present — it would false-match a basename suffix. Plain substring check (no
|
|
// regex, so no CRLF-fragility): the grep -Fxq positive guard above proves the
|
|
// correct mechanism; this negative guard catches a revert to the `case` form.
|
|
assert.ok(
|
|
!src.includes('case "$IN_SCOPE"'),
|
|
'code-review.md Tier-3 must not use the unanchored `case "$IN_SCOPE"` substring membership ' +
|
|
'test (#2666) — use grep -Fxq for exact whole-line matching'
|
|
);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// BUG 2 — severity-label parser: present_results must accept both `critical:`
|
|
// and `blocker:` as Critical-tier frontmatter keys.
|
|
// ---------------------------------------------------------------------------
|
|
describe('Bug 2 — present_results severity-label parser', () => {
|
|
test('frontmatter with blocker: 8 is parsed as critical: 8', () => {
|
|
const frontmatter = [
|
|
'phase: 03-courses',
|
|
'reviewed: 2025-01-01T00:00:00Z',
|
|
'findings:',
|
|
' blocker: 8',
|
|
' warning: 2',
|
|
' info: 0',
|
|
' total: 10',
|
|
'status: issues_found',
|
|
].join('\n');
|
|
|
|
const result = parseFrontmatterCritical(frontmatter);
|
|
assert.strictEqual(result.critical, 8);
|
|
});
|
|
|
|
test('frontmatter with critical: 5 is parsed as critical: 5', () => {
|
|
const frontmatter = [
|
|
'phase: 03-courses',
|
|
'reviewed: 2025-01-01T00:00:00Z',
|
|
'findings:',
|
|
' critical: 5',
|
|
' warning: 1',
|
|
' info: 0',
|
|
' total: 6',
|
|
'status: issues_found',
|
|
].join('\n');
|
|
|
|
const result = parseFrontmatterCritical(frontmatter);
|
|
assert.strictEqual(result.critical, 5);
|
|
});
|
|
|
|
test('frontmatter with neither critical nor blocker returns 0', () => {
|
|
const frontmatter = [
|
|
'phase: 03-courses',
|
|
'findings:',
|
|
' warning: 3',
|
|
' info: 1',
|
|
' total: 4',
|
|
'status: issues_found',
|
|
].join('\n');
|
|
|
|
const result = parseFrontmatterCritical(frontmatter);
|
|
assert.strictEqual(result.critical, 0);
|
|
});
|
|
|
|
// Docs-parity: the workflow .md must contain the updated grep pattern.
|
|
test('code-review.md present_results grep accepts both critical and blocker labels', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
assert.ok(
|
|
src.includes('grep -E "^[[:space:]]*(critical|blocker):"'),
|
|
'code-review.md present_results must grep for both critical: and blocker: labels'
|
|
);
|
|
});
|
|
|
|
// Docs-parity: the workflow .md must contain the updated grep for BL- headings.
|
|
test('code-review.md present_results grep includes BL- headings alongside CR- and WR-', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
assert.ok(
|
|
src.includes('### BL-') && src.includes('### CR-') && src.includes('### WR-'),
|
|
'code-review.md present_results must grep for BL- alongside CR- and WR- headings'
|
|
);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// BUG 3 — fixer agent ID alphabet and filter sets must include BL-* alongside CR-*.
|
|
// ---------------------------------------------------------------------------
|
|
describe('Bug 3 — gsd-code-fixer BL-* inclusion in filter sets', () => {
|
|
test('finding_parser documents BL-\\d+ as Critical-tier-equivalent', () => {
|
|
const src = fs.readFileSync(FIXER_PATH, 'utf8');
|
|
const parserStart = src.indexOf('<finding_parser>');
|
|
const parserEnd = src.indexOf('</finding_parser>');
|
|
assert.ok(parserStart !== -1, 'gsd-code-fixer.md must have a <finding_parser> block');
|
|
const parserSection = src.slice(parserStart, parserEnd);
|
|
assert.ok(
|
|
parserSection.includes('BL-'),
|
|
'finding_parser block must document BL-* as a Critical-tier-equivalent ID prefix'
|
|
);
|
|
});
|
|
|
|
test('parse_findings step documents severity as "Critical (CR-* or BL-*)"', () => {
|
|
const src = fs.readFileSync(FIXER_PATH, 'utf8');
|
|
const stepStart = src.indexOf('<step name="parse_findings">');
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
assert.ok(stepStart !== -1, 'gsd-code-fixer.md must have a parse_findings step');
|
|
const stepSection = src.slice(stepStart, stepEnd);
|
|
assert.ok(
|
|
stepSection.includes('CR-* or BL-*') || stepSection.includes('CR-* and BL-*'),
|
|
'parse_findings step must describe Critical severity as "CR-* or BL-*"'
|
|
);
|
|
});
|
|
|
|
test('critical_warning filter set includes BL-* alongside CR-* and WR-*', () => {
|
|
const src = fs.readFileSync(FIXER_PATH, 'utf8');
|
|
const stepStart = src.indexOf('<step name="parse_findings">');
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
const stepSection = src.slice(stepStart, stepEnd);
|
|
|
|
const critWarningIdx = stepSection.indexOf('critical_warning');
|
|
assert.ok(critWarningIdx !== -1, 'parse_findings must define critical_warning filter');
|
|
const lineStart = stepSection.lastIndexOf('\n', critWarningIdx);
|
|
const lineEnd = stepSection.indexOf('\n', critWarningIdx);
|
|
const filterLine = stepSection.slice(lineStart, lineEnd);
|
|
assert.ok(
|
|
filterLine.includes('BL-'),
|
|
'critical_warning filter line must include BL-*: ' + filterLine.trim()
|
|
);
|
|
});
|
|
|
|
test('sort order description mentions both CR-* and BL-* for Critical tier', () => {
|
|
const src = fs.readFileSync(FIXER_PATH, 'utf8');
|
|
const stepStart = src.indexOf('<step name="parse_findings">');
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
const stepSection = src.slice(stepStart, stepEnd);
|
|
assert.ok(
|
|
stepSection.includes('BL-'),
|
|
'parse_findings sort-order description must mention BL-* as Critical-tier alongside CR-*'
|
|
);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// REVIEWER CONTRACT — gsd-code-reviewer.md must acknowledge BL-/blocker: as
|
|
// an accepted alternative to CR-/critical: (tier-equivalent).
|
|
// ---------------------------------------------------------------------------
|
|
describe('Reviewer contract — gsd-code-reviewer.md label-equivalence', () => {
|
|
test('write_review step documents blocker: as accepted alternative to critical:', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepStart = src.indexOf('<step name="write_review">');
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
assert.ok(stepStart !== -1, 'gsd-code-reviewer.md must have a write_review step');
|
|
const stepSection = src.slice(stepStart, stepEnd);
|
|
assert.ok(
|
|
stepSection.includes('blocker'),
|
|
'write_review step must acknowledge blocker: as a tier-equivalent alternative to critical:'
|
|
);
|
|
});
|
|
|
|
test('write_review step acknowledges BL- finding ID prefix as Critical-tier-equivalent', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepStart = src.indexOf('<step name="write_review">');
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
const stepSection = src.slice(stepStart, stepEnd);
|
|
assert.ok(
|
|
stepSection.includes('BL-'),
|
|
'write_review step must acknowledge BL- as a Critical-tier-equivalent finding ID prefix'
|
|
);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// BUG 4 (#2352) — compute_file_scope must tilde-expand `~/...`-prefixed
|
|
// SUMMARY.md key-files entries BEFORE the "Filter deleted files" existence
|
|
// check. Bash only tilde-expands a literal `~` written in source text, never
|
|
// one arriving as the value of an already-expanded variable — so a real file
|
|
// recorded as `~/.claude/gsd-core/workflows/verify-phase.md` was silently
|
|
// misclassified as deleted and dropped from REVIEW_FILES, and a phase whose
|
|
// every recorded file used a `~/...` path hit the empty-scope skip
|
|
// ("No source files changed ... Skipping review.") as a false negative.
|
|
//
|
|
// Tested both ways: a docs-parity assertion (cross-platform, pure fs read)
|
|
// that the normalization block exists in the deployed workflow text, and a
|
|
// behavioral test that extracts the actual "Expand tilde paths" +
|
|
// "Filter deleted files" bash blocks from code-review.md and executes them
|
|
// via a real bash subprocess against planted files under a fresh HOME.
|
|
// ---------------------------------------------------------------------------
|
|
describe('Bug 4 (#2352) — compute_file_scope tilde-path expansion', () => {
|
|
// Docs-parity: the workflow .md must contain the tilde-normalization block
|
|
// as step 1 of "Post-processing (all tiers)", ahead of the deleted-file
|
|
// filter, so what we behaviorally test below is what is actually deployed.
|
|
test('code-review.md contains a tilde-expansion block ahead of the deleted-file filter', () => {
|
|
const src = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
const postProcessingIdx = src.indexOf('**Post-processing (all tiers):**');
|
|
assert.ok(postProcessingIdx !== -1, 'code-review.md must have a "Post-processing (all tiers)" section');
|
|
|
|
const expandIdx = src.indexOf('EXPANDED_FILES=()', postProcessingIdx);
|
|
assert.ok(expandIdx !== -1, 'Post-processing must contain an EXPANDED_FILES=() tilde-expansion loop');
|
|
|
|
const caseIdx = src.indexOf('case "$file" in', postProcessingIdx);
|
|
assert.ok(caseIdx !== -1 && caseIdx < expandIdx + 400, 'tilde-expansion loop must use a case "$file" in match');
|
|
assert.ok(
|
|
src.slice(caseIdx, caseIdx + 200).includes('"~/"*)') &&
|
|
src.slice(caseIdx, caseIdx + 200).includes('${HOME}${file#\\~}'),
|
|
'tilde-expansion loop must rewrite a leading ~/ to ${HOME}/... via ${file#\\~}'
|
|
);
|
|
|
|
const deletedFilterIdx = src.indexOf('DELETED_COUNT=0', postProcessingIdx);
|
|
assert.ok(deletedFilterIdx !== -1, 'Post-processing must still contain the deleted-file filter');
|
|
assert.ok(
|
|
expandIdx < deletedFilterIdx,
|
|
'tilde-expansion loop must run BEFORE the deleted-file filter, not after'
|
|
);
|
|
});
|
|
|
|
// Extract the tilde-expansion fence and the (non-adjacent — the exclusions
|
|
// filter sits between them) deleted-file-filter fence from the
|
|
// "Post-processing (all tiers)" section of code-review.md — the exact
|
|
// snippets the runtime executes, located by content anchor rather than
|
|
// position so an intervening step doesn't silently swap in the wrong
|
|
// block — and glue them behind a synthetic REVIEW_FILES=("$@") seed for
|
|
// direct execution. The exclusions filter itself is intentionally skipped
|
|
// here: it only matches relative planning-artifact paths and is orthogonal
|
|
// to tilde expansion (see code-review.md step 2, "Apply exclusions").
|
|
function extractPostProcessingScript() {
|
|
// readFileNormalized() strips \r\n -> \n before either fence below is
|
|
// sliced out and later spawned via spawnSync('bash', ...) in
|
|
// runPostProcessing() — an un-normalized read on a Windows checkout would
|
|
// break bash mid-script (DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE, #2650).
|
|
const src = readFileNormalized(WORKFLOW_PATH);
|
|
const postProcessingIdx = src.indexOf('**Post-processing (all tiers):**');
|
|
assert.ok(postProcessingIdx !== -1, 'code-review.md must have a "Post-processing (all tiers)" section');
|
|
|
|
function fenceContaining(marker) {
|
|
const markerIdx = src.indexOf(marker, postProcessingIdx);
|
|
assert.ok(markerIdx !== -1, `expected to find "${marker}" in the Post-processing section`);
|
|
const fenceStart = src.lastIndexOf('```bash', markerIdx);
|
|
assert.ok(fenceStart !== -1 && fenceStart > postProcessingIdx, `no \`\`\`bash fence before "${marker}"`);
|
|
const bodyStart = src.indexOf('\n', fenceStart) + 1;
|
|
const fenceEnd = src.indexOf('\n```', bodyStart);
|
|
assert.ok(fenceEnd !== -1, `unterminated \`\`\`bash fence containing "${marker}"`);
|
|
return src.slice(bodyStart, fenceEnd);
|
|
}
|
|
|
|
const tildeBlock = fenceContaining('EXPANDED_FILES=()');
|
|
const deletedBlock = fenceContaining('DELETED_COUNT=0');
|
|
|
|
return [
|
|
'REVIEW_FILES=("$@")',
|
|
tildeBlock,
|
|
deletedBlock,
|
|
'printf "%s\\n" "${REVIEW_FILES[@]}"',
|
|
'echo "REVIEW_FILES_COUNT=${#REVIEW_FILES[@]}"',
|
|
'echo "DELETED_COUNT=$DELETED_COUNT"',
|
|
].join('\n');
|
|
}
|
|
|
|
function runPostProcessing(homeDir, files) {
|
|
const script = extractPostProcessingScript();
|
|
// "bash" as $0 so the real REVIEW_FILES entries land in "$@" from $1.
|
|
return toLegacyResult(
|
|
runHook('-c', [script, 'bash', ...files], {
|
|
interpreter: 'bash',
|
|
env: { ...process.env, HOME: homeDir },
|
|
timeoutMs: PROBE_TIMEOUT_MS,
|
|
})
|
|
);
|
|
}
|
|
|
|
let tmpHome;
|
|
|
|
test('setup: plant a fresh HOME with a real file', { skip: process.platform === 'win32' }, () => {
|
|
tmpHome = createTempDir('gsd-2352-home-');
|
|
fs.mkdirSync(path.join(tmpHome, '.claude', 'gsd-core', 'workflows'), { recursive: true });
|
|
fs.writeFileSync(
|
|
path.join(tmpHome, '.claude', 'gsd-core', 'workflows', 'verify-phase.md'),
|
|
'# real file\n',
|
|
'utf8'
|
|
);
|
|
});
|
|
|
|
test(
|
|
'AC1: a ~/-prefixed path to a real file survives and is not counted deleted',
|
|
{ skip: process.platform === 'win32' },
|
|
() => {
|
|
const result = runPostProcessing(tmpHome, ['~/.claude/gsd-core/workflows/verify-phase.md']);
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
assert.match(
|
|
result.stdout,
|
|
new RegExp(path.join(tmpHome, '.claude', 'gsd-core', 'workflows', 'verify-phase.md').replace(/[/\\.]/g, '\\$&')),
|
|
`expected expanded absolute path in surviving REVIEW_FILES; got: ${JSON.stringify(result.stdout)}`
|
|
);
|
|
assert.match(result.stdout, /DELETED_COUNT=0/, `expected DELETED_COUNT=0; got: ${JSON.stringify(result.stdout)}`);
|
|
assert.match(
|
|
result.stdout,
|
|
/REVIEW_FILES_COUNT=1/,
|
|
`expected the tilde path to survive into REVIEW_FILES; got: ${JSON.stringify(result.stdout)}`
|
|
);
|
|
}
|
|
);
|
|
|
|
test(
|
|
'AC2: a ~/-prefixed path to a non-existent file is still correctly excluded as deleted',
|
|
{ skip: process.platform === 'win32' },
|
|
() => {
|
|
const result = runPostProcessing(tmpHome, ['~/.claude/gsd-core/workflows/does-not-exist.md']);
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
assert.match(result.stdout, /DELETED_COUNT=1/, `expected DELETED_COUNT=1; got: ${JSON.stringify(result.stdout)}`);
|
|
assert.match(
|
|
result.stdout,
|
|
/REVIEW_FILES_COUNT=0/,
|
|
`expected the missing tilde path to be dropped; got: ${JSON.stringify(result.stdout)}`
|
|
);
|
|
}
|
|
);
|
|
|
|
test(
|
|
'AC3: a phase where every recorded file is a real ~/-prefixed path does not empty the scope',
|
|
{ skip: process.platform === 'win32' },
|
|
() => {
|
|
const result = runPostProcessing(tmpHome, ['~/.claude/gsd-core/workflows/verify-phase.md']);
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
const countMatch = result.stdout.match(/REVIEW_FILES_COUNT=(\d+)/);
|
|
assert.ok(countMatch, `expected a REVIEW_FILES_COUNT line; got: ${JSON.stringify(result.stdout)}`);
|
|
assert.ok(
|
|
Number(countMatch[1]) > 0,
|
|
'an all-tilde real-file scope must not reduce to zero (would trigger the empty-scope skip)'
|
|
);
|
|
}
|
|
);
|
|
|
|
test(
|
|
'AC4: mixed tilde + missing ordinary relative path resolve independently',
|
|
{ skip: process.platform === 'win32' },
|
|
() => {
|
|
const result = runPostProcessing(tmpHome, [
|
|
'~/.claude/gsd-core/workflows/verify-phase.md',
|
|
'this/relative/path/does-not-exist.md',
|
|
]);
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
assert.match(result.stdout, /DELETED_COUNT=1/, `expected exactly 1 deleted; got: ${JSON.stringify(result.stdout)}`);
|
|
assert.match(
|
|
result.stdout,
|
|
/REVIEW_FILES_COUNT=1/,
|
|
`expected only the tilde path to survive; got: ${JSON.stringify(result.stdout)}`
|
|
);
|
|
assert.doesNotMatch(
|
|
result.stdout,
|
|
/this\/relative\/path\/does-not-exist\.md/,
|
|
'the missing ordinary relative path must not survive into REVIEW_FILES'
|
|
);
|
|
}
|
|
);
|
|
|
|
test('teardown: remove the temp HOME', { skip: process.platform === 'win32' }, () => {
|
|
cleanup(tmpHome);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Shared diff-base extraction/execution helpers (Bug 5 #3191, Bug 6 #3503).
|
|
//
|
|
// The workflow computes "the phase's base commit" in three independent bash
|
|
// invocations (each <step> is its own shell): the Tier-3 file-scope fallback
|
|
// (compute_file_scope), the agent-context DIFF_BASE (spawn_reviewer), and the
|
|
// fallow pre-pass's --changed-since base (structural-pre-pass.md).
|
|
//
|
|
// Behavioral style follows Bug 4: extract the SHIPPED bash from the workflow
|
|
// .md files by content anchor and execute it via a real bash subprocess
|
|
// against a git fixture — so the assertion binds the deployed text, not a
|
|
// JS reimplementation. Running the real `git log` (not a regex shim) is what
|
|
// makes platform-level regex holes (the #3191 macOS `\b` no-op) visible.
|
|
// ---------------------------------------------------------------------------
|
|
|
|
// The ```bash fence containing `marker`, located after `fromIdx`.
|
|
function fenceContaining(src, marker, fromIdx = 0) {
|
|
const markerIdx = src.indexOf(marker, fromIdx);
|
|
assert.ok(markerIdx !== -1, `expected to find "${marker}" in workflow source`);
|
|
const fenceStart = src.lastIndexOf('```bash', markerIdx);
|
|
assert.ok(fenceStart !== -1, `no \`\`\`bash fence before "${marker}"`);
|
|
const bodyStart = src.indexOf('\n', fenceStart) + 1;
|
|
const fenceEnd = src.indexOf('\n```', bodyStart);
|
|
assert.ok(fenceEnd !== -1, `unterminated \`\`\`bash fence containing "${marker}"`);
|
|
return src.slice(bodyStart, fenceEnd);
|
|
}
|
|
|
|
// The Tier-3 derivation prefix: fence start up to the REVIEW_FILES branch.
|
|
function extractTier3Derivation() {
|
|
const src = readFileNormalized(WORKFLOW_PATH);
|
|
const fence = fenceContaining(src, '# Compute diff base from phase commits');
|
|
const cut = fence.indexOf('if [ ${#REVIEW_FILES[@]} -eq 0 ]');
|
|
assert.ok(cut !== -1, 'Tier-3 fence must contain the REVIEW_FILES empty-scope branch');
|
|
return fence.slice(0, cut);
|
|
}
|
|
|
|
// spawn_reviewer no longer derives its own DIFF_BASE (#4209 B3 fix: a second,
|
|
// divergent recomputation there made the external reviewer lane and the
|
|
// internal reviewer diff against different base SHAs on any re-review). It
|
|
// now reuses the value compute_file_scope's Tier-3 derivation already
|
|
// computed, so this is the SAME snippet as extractTier3Derivation() — kept
|
|
// as a distinct name so T2/T5 below still read as testing spawn_reviewer's
|
|
// contract, not just Tier 3's.
|
|
function extractSpawnReviewerDerivation() {
|
|
return extractTier3Derivation();
|
|
}
|
|
|
|
// The fallow phase-scope derivation, from the step fragment. The fragment
|
|
// carries markdown-escaped quotes (\") in this fence — an authoring
|
|
// artifact that survived #2994 fragmentization verbatim; the runtime agent
|
|
// normalizes them when transcribing, so the test does the same before
|
|
// executing. Sliced from FALLOW_SCOPE_ARGS=() (skipping the gsd-tools
|
|
// runtime resolver line above it, which exits 1 on machines without an
|
|
// installed gsd-tools and is orthogonal to the base-derivation under test)
|
|
// to just before the gsd_run invocation (which needs the real binary).
|
|
function extractFallowDerivation() {
|
|
const src = readFileNormalized(PRE_PASS_STEP_PATH);
|
|
const fence = fenceContaining(src, 'FALLOW_PHASE_START=$(git log');
|
|
const scopeStart = fence.indexOf('FALLOW_SCOPE_ARGS=()');
|
|
assert.ok(scopeStart !== -1, 'fallow fence must define FALLOW_SCOPE_ARGS=()');
|
|
const cut = fence.indexOf('gsd_run run-with-timeout');
|
|
assert.ok(cut !== -1, 'fallow fence must contain the gsd_run run-with-timeout call');
|
|
assert.ok(scopeStart < cut, 'FALLOW_SCOPE_ARGS must precede the gsd_run invocation');
|
|
return fence.slice(scopeStart, cut).replace(/\\"/g, '"');
|
|
}
|
|
|
|
// Execute a derivation snippet with PADDED_PHASE (and the fallow scope gate)
|
|
// set, echoing the values it computes between sentinels so multi-line
|
|
// PHASE_COMMITS parse cleanly.
|
|
function runDerivation(repo, snippet, phase) {
|
|
const script = [
|
|
`PADDED_PHASE=${phase}`,
|
|
// #3995: the derivations anchor on the phase's own directory, not a
|
|
// commit-subject grep — the fixture commits each phase's directory at
|
|
// its first scope commit.
|
|
`PHASE_DIR=${repo}/.planning/phases/${phase}-ctx`,
|
|
'FALLOW_SCOPE=phase',
|
|
snippet,
|
|
'echo "===PHASE_START==="',
|
|
'printf \'%s\\n\' "$PHASE_START"',
|
|
'echo "===DIFF_BASE==="',
|
|
'printf \'%s\\n\' "$DIFF_BASE"',
|
|
'echo "===FALLOW_BASE==="',
|
|
'printf \'%s\\n\' "$FALLOW_BASE"',
|
|
'echo "===END==="',
|
|
].join('\n');
|
|
// Bash FAN-OUT: the extracted snippet runs `git log` plus an `echo | tail`
|
|
// pipe — the wrong class for `PROBE_TIMEOUT_MS` (a single short CLI
|
|
// probe). Same class as the observed CI failures in
|
|
// tests/quick-branching.test.cjs (PR #3787 run 32668773524) and
|
|
// tests/worktree-safety.test.cjs (`next` run 32608945654). See
|
|
// HOOK_FANOUT_TIMEOUT_MS in ./helpers/timeouts.cjs for the class
|
|
// rationale.
|
|
return toLegacyResult(
|
|
runHook('-c', [script, 'bash'], {
|
|
interpreter: 'bash',
|
|
cwd: repo,
|
|
timeoutMs: HOOK_FANOUT_TIMEOUT_MS,
|
|
})
|
|
);
|
|
}
|
|
|
|
function parseSentinel(stdout, name) {
|
|
const m = stdout.match(new RegExp(`===${name}===\\n([\\s\\S]*?)\\n===`));
|
|
if (!m) return null;
|
|
return m[1].split('\n').map((l) => l.trim()).filter((l) => l.length > 0);
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Bug 5 (#3191) — EVERY diff-base derivation must use the same anchored,
|
|
// portable derivation.
|
|
//
|
|
// The workflow computes "the phase's base commit" in three independent bash
|
|
// invocations (each <step> is its own shell): the Tier-3 file-scope fallback
|
|
// (compute_file_scope), the agent-context DIFF_BASE (spawn_reviewer), and the
|
|
// fallow pre-pass's --changed-since base (structural-pre-pass.md). #2989
|
|
// anchored only the Tier-3 copy — and did so with `\b`, which is not a POSIX
|
|
// ERE token, so on macOS (regex(3)) that grep matches NOTHING and Tier 3
|
|
// always fails closed. The other two sites kept the original unanchored
|
|
// `--grep="${PADDED_PHASE}"`, whose oldest substring match is routinely a
|
|
// version-string/date commit from months before the phase existed.
|
|
// (#3503 later replaced the anchor itself — a subject-line conventional-
|
|
// commit scope match instead of the "[Pp]hase N" prose phrase, which GSD's
|
|
// own commits never contain; see Bug 6. The lockstep + portability +
|
|
// fail-closed contract THIS block verifies is unchanged.)
|
|
//
|
|
// Behavioral style follows Bug 4: extract the SHIPPED bash from the workflow
|
|
// .md files by content anchor and execute it via a real bash subprocess
|
|
// against a git fixture — so the assertion binds the deployed text, not a
|
|
// JS reimplementation. Running the real `git log` (not a regex shim) is what
|
|
// keeps platform-level regex holes (the #3191 macOS `\b` no-op) visible.
|
|
// ---------------------------------------------------------------------------
|
|
// Shared: the fixture phase directory every derivation anchors on (#3995).
|
|
const PHASE06_PLAN_REL = path.join('.planning', 'phases', '06-ctx', '06-PLAN.md');
|
|
|
|
// Shared history builder (was local to the #3503 describe; the #3995 rows
|
|
// reuse it). Each entry is [relPath, subject, body?]; parent dirs are created.
|
|
function buildHistory(prefix, commits) {
|
|
const repo = createTempGitProject(prefix);
|
|
const hashes = {};
|
|
for (const [file, message, body] of commits) {
|
|
fs.mkdirSync(path.dirname(path.join(repo, file)), { recursive: true });
|
|
fs.writeFileSync(path.join(repo, file), `${message}\n`);
|
|
gitOrThrow(['add', file], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['commit', '-m', message, ...(body ? ['-m', body] : [])], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
hashes[file] = gitOrThrow(['rev-parse', 'HEAD'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS }).trim();
|
|
}
|
|
return { repo, hashes };
|
|
}
|
|
|
|
describe('Bug 5 (#3191) — same anchored, portable phase-scope grep at all three diff-base sites', () => {
|
|
const SKIP_WIN32 = { skip: process.platform === 'win32' };
|
|
|
|
// Fixture: five commits whose messages exercise every false-match class
|
|
// from the issue — version string + date, another phase's plan whose scope
|
|
// number is a digit-superset, a prose "Phase N" mention in another phase's
|
|
// subject — plus the phase's real first scope commit and an unrelated HEAD.
|
|
function buildFixture(prefix, phaseCommitMessage, opts = {}) {
|
|
// opts.skipPhaseDir: the fail-closed row (T5) commits NO phase directory,
|
|
// so the directory anchor must resolve nothing.
|
|
const commitPhaseDir = opts.skipPhaseDir !== true;
|
|
const repo = createTempGitProject(prefix);
|
|
const phaseDir = path.join(repo, '.planning', 'phases', '06-ctx');
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
const commits = [
|
|
['c1.txt', 'chore: bump to v2.06.0 on 2026-01-05'],
|
|
['c2.txt', 'docs(60-01): unrelated phase-plan work'],
|
|
[commitPhaseDir ? PHASE06_PLAN_REL : 'c3.txt', phaseCommitMessage],
|
|
['c4.txt', 'chore: Phase 60 cleanup'],
|
|
['c5.txt', 'docs: touch README'],
|
|
];
|
|
const hashes = {};
|
|
for (const [file, message] of commits) {
|
|
fs.mkdirSync(path.dirname(path.join(repo, file)), { recursive: true });
|
|
fs.writeFileSync(path.join(repo, file), `${message}\n`);
|
|
gitOrThrow(['add', file], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['commit', '-m', message], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
hashes[file] = gitOrThrow(['rev-parse', 'HEAD'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS }).trim();
|
|
}
|
|
return { repo, hashes };
|
|
}
|
|
|
|
test(
|
|
'T1 + T4: Tier-3 derivation matches ONLY the phase\'s real scope commit — never a digit-substring or superset hit',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo, hashes } = buildFixture('gsd-3191-tier3-', 'docs(06): capture phase context');
|
|
try {
|
|
const result = runDerivation(repo, extractTier3Derivation(), '06');
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
const phaseStart = parseSentinel(result.stdout, 'PHASE_START');
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
// AC: the phase's real commits are a small minority of digit-containing
|
|
// commits; the derivation must resolve to an ancestor near the phase's
|
|
// actual first commit (c3^) — never the older v2.06.0/docs(06-01) hits.
|
|
assert.deepStrictEqual(
|
|
phaseStart,
|
|
[hashes[PHASE06_PLAN_REL]],
|
|
`Tier-3 anchor must resolve to the phase dir's first commit; got: ${JSON.stringify(phaseStart)}`
|
|
);
|
|
assert.deepStrictEqual(
|
|
diffBase,
|
|
[`${hashes[PHASE06_PLAN_REL]}^`],
|
|
'Tier-3 DIFF_BASE must be the phase first-commit parent'
|
|
);
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
test(
|
|
'T2: spawn_reviewer DIFF_BASE derivation uses the same anchored grep (not the bare digit)',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo, hashes } = buildFixture('gsd-3191-spawn-', 'docs(06): capture phase context');
|
|
try {
|
|
const result = runDerivation(repo, extractSpawnReviewerDerivation(), '06');
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
const phaseStart = parseSentinel(result.stdout, 'PHASE_START');
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
// Pre-fix this matches c1 and c2 as well and tail -1 picks c1 — the
|
|
// oldest unrelated match — feeding a bogus diff_base to the reviewer
|
|
// agent exactly when files: is empty (the fail-closed scenario).
|
|
assert.deepStrictEqual(
|
|
phaseStart,
|
|
[hashes[PHASE06_PLAN_REL]],
|
|
`spawn_reviewer anchor must resolve to the phase dir's first commit; got: ${JSON.stringify(phaseStart)}`
|
|
);
|
|
assert.deepStrictEqual(
|
|
diffBase,
|
|
[`${hashes[PHASE06_PLAN_REL]}^`],
|
|
'spawn_reviewer DIFF_BASE must be the phase first-commit parent'
|
|
);
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
test(
|
|
'T3: fallow phase scope derives --changed-since from the anchored grep, never an old substring match',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo, hashes } = buildFixture('gsd-3191-fallow-', 'docs(06): capture phase context');
|
|
try {
|
|
const result = runDerivation(repo, extractFallowDerivation(), '06');
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
// Pre-fix the unanchored grep's oldest match is the v2.06.0 commit, so
|
|
// FALLOW_SCOPE_ARGS resolves to --changed-since <old-unrelated-commit>
|
|
// and widens the structural pre-pass far beyond the phase.
|
|
assert.deepStrictEqual(
|
|
fallowBase,
|
|
[`${hashes[PHASE06_PLAN_REL]}^`],
|
|
`FALLOW_BASE must be the phase first-commit parent, got: ${JSON.stringify(fallowBase)}`
|
|
);
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
test(
|
|
'root commit: fallow phase scope uses a resolvable root SHA',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const repo = createTempDir('gsd-4183-fallow-root-');
|
|
try {
|
|
gitOrThrow(['init', '-b', 'main'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['config', 'user.email', 'test@test.com'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['config', 'user.name', 'Test'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['config', 'commit.gpgsign', 'false'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
|
|
const phaseFile = path.join(repo, '.planning', 'phases', '06-ctx', 'PLAN.md');
|
|
fs.mkdirSync(path.dirname(phaseFile), { recursive: true });
|
|
fs.writeFileSync(phaseFile, '# phase context\n');
|
|
gitOrThrow(['add', '.planning'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['commit', '-m', 'docs(06): initial phase context'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
const rootSha = gitOrThrow(['rev-parse', 'HEAD'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS }).trim();
|
|
|
|
fs.writeFileSync(path.join(repo, 'index.js'), 'module.exports = 1;\n');
|
|
gitOrThrow(['add', 'index.js'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
gitOrThrow(['commit', '-m', 'feat: add source'], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS });
|
|
|
|
const result = runDerivation(repo, extractFallowDerivation(), '06');
|
|
assert.equal(result.status, 0, `snippet exited ${result.status}; stderr=${result.stderr}`);
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
assert.deepStrictEqual(
|
|
fallowBase,
|
|
[rootSha],
|
|
`root-parent FALLOW_BASE regression: expected ${rootSha}, got ${JSON.stringify(fallowBase)}`,
|
|
);
|
|
assert.equal(
|
|
gitOrThrow(['rev-parse', '--verify', `${fallowBase[0]}^{commit}`], { cwd: repo, timeoutMs: GIT_TIMEOUT_MS }).trim(),
|
|
rootSha,
|
|
'FALLOW_BASE must resolve to the root commit',
|
|
);
|
|
|
|
if (process.env.CI) {
|
|
const { requireFallowBinary } = require('../gsd-core/bin/lib/fallow-runner.cjs');
|
|
const { execTool } = require('../gsd-core/bin/lib/shell-command-projection.cjs');
|
|
const audit = execTool(
|
|
requireFallowBinary({ cwd: ROOT, envPath: '' }),
|
|
['audit', '--changed-since', fallowBase[0], '--format', 'json'],
|
|
{ cwd: repo, timeout: 120000 },
|
|
);
|
|
assert.ok([0, 1].includes(audit.exitCode), `fallow root audit exit=${audit.exitCode}; stderr=${audit.stderr}`);
|
|
console.log(`fallow-root-audit normal-exit=${audit.exitCode}`);
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
},
|
|
);
|
|
|
|
test(
|
|
'T5: with no genuine phase scope commit, every derivation yields NO base (fail-closed preserved)',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo } = buildFixture('gsd-3191-closed-', 'feat: scanner core', { skipPhaseDir: true }); // no committed phase dir anywhere
|
|
try {
|
|
for (const [label, snippet] of [
|
|
['tier3', extractTier3Derivation()],
|
|
['spawn_reviewer', extractSpawnReviewerDerivation()],
|
|
['fallow', extractFallowDerivation()],
|
|
]) {
|
|
const result = runDerivation(repo, snippet, '06');
|
|
assert.equal(result.status, 0, `${label} exited ${result.status}; stderr=${result.stderr}`);
|
|
const phaseStart = parseSentinel(result.stdout, 'PHASE_START');
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
assert.deepStrictEqual(phaseStart, [], `${label}: no phase dir committed — anchor must stay empty`);
|
|
assert.deepStrictEqual(diffBase, [], `${label}: DIFF_BASE must stay empty (no bogus base)`);
|
|
assert.deepStrictEqual(fallowBase, [], `${label}: FALLOW_BASE must stay unset`);
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
// T6 docs-parity anti-revert (#3191/#3995): every diff-base derivation in
|
|
// both files must use the SAME phase-directory anchor — and no message-grep
|
|
// derivation may return (a subject carries no milestone bound; that class
|
|
// failed five times: #2989/#3191/#3503/#3995).
|
|
test('T6 docs-parity: all diff-base derivations use the identical phase-directory anchor; no --grep site remains', () => {
|
|
const sources = [
|
|
readFileNormalized(WORKFLOW_PATH),
|
|
readFileNormalized(PRE_PASS_STEP_PATH).replace(/\\"/g, '"'),
|
|
];
|
|
for (const src of sources) {
|
|
assert.ok(
|
|
src.includes('PHASE_START=$(git log --format="%H" --diff-filter=A -- "${PHASE_DIR}"'),
|
|
'each file must derive the base from the phase directory\'s first commit (#3995)'
|
|
);
|
|
}
|
|
const grepSites = [];
|
|
for (const src of sources) {
|
|
for (const m of src.matchAll(/^\s*[A-Z_]+=\$\(git log[^\n]*--grep=[^\n]*$/gm)) {
|
|
grepSites.push(m[0]);
|
|
}
|
|
}
|
|
assert.deepStrictEqual(
|
|
grepSites.filter((l) => l.includes('PHASE_SCOPE_NUM') || /phase-\)?\(/.test(l)),
|
|
[],
|
|
'no phase-scope message-grep derivation may remain — subjects carry no milestone bound (#3995)'
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('Bug 6 (#3503/#3995) — diff base keys on the phase directory, not commit subjects', () => {
|
|
const SKIP_WIN32 = { skip: process.platform === 'win32' };
|
|
|
|
const REPRO_HISTORY = [
|
|
['c1.txt', 'chore: bump to v2.06.0 on 2026-01-05'],
|
|
['c2.txt', 'feat(60-01): probe wiring', 'The EF path still uses it, fenced to Phase 06 per D-09.'],
|
|
['c3.txt', 'docs: commit message format', 'Phase headers use the form:\n\n### Phase 06 (Cluster B): Title\n\nin ROADMAP detail sections.'],
|
|
[PHASE06_PLAN_REL, 'docs(06): capture phase context'],
|
|
['c5.txt', 'feat(06-01): implement scanner core'],
|
|
['c6.txt', 'docs(phase-6): update tracking after wave 1'],
|
|
['c7.txt', 'docs: touch README'],
|
|
];
|
|
|
|
test(
|
|
'T1: prose forward-references and doc-format examples never capture the base — it resolves to the phase dir first commit at all three sites',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo, hashes } = buildHistory('gsd-3503-scope-', REPRO_HISTORY);
|
|
try {
|
|
const sites = [
|
|
['tier3', extractTier3Derivation()],
|
|
['spawn_reviewer', extractSpawnReviewerDerivation()],
|
|
['fallow', extractFallowDerivation()],
|
|
];
|
|
for (const [label, snippet] of sites) {
|
|
const result = runDerivation(repo, snippet, '06');
|
|
assert.equal(result.status, 0, `${label} exited ${result.status}; stderr=${result.stderr}`);
|
|
const phaseStart = parseSentinel(result.stdout, 'PHASE_START');
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
const dirFirst = hashes[PHASE06_PLAN_REL];
|
|
if (label !== 'fallow') {
|
|
assert.deepStrictEqual(
|
|
phaseStart,
|
|
[dirFirst],
|
|
`${label}: anchor must resolve to the phase dir's first commit; got: ${JSON.stringify(phaseStart)}`
|
|
);
|
|
}
|
|
const expected = [`${dirFirst}^`];
|
|
if (label === 'fallow') {
|
|
assert.deepStrictEqual(fallowBase, expected, `${label}: base must be the phase dir first commit's parent`);
|
|
} else {
|
|
assert.deepStrictEqual(diffBase, expected, `${label}: base must be the phase dir first commit's parent`);
|
|
}
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
test(
|
|
'T2: subject spellings are irrelevant to the directory anchor — unpadded and padded histories resolve identically',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo, hashes } = buildHistory('gsd-3503-unpadded-', [
|
|
['c1.txt', 'feat(60-01): probe wiring', 'Deferred to Phase 06 per D-09.'],
|
|
[PHASE06_PLAN_REL, 'docs(phase-6): capture phase context'],
|
|
['c3.txt', 'feat(6-01): implement scanner core'],
|
|
['c4.txt', 'test(6): persist human verification items as UAT'],
|
|
['c5.txt', 'docs: touch README'],
|
|
]);
|
|
try {
|
|
for (const [label, snippet] of [
|
|
['tier3', extractTier3Derivation()],
|
|
['spawn_reviewer', extractSpawnReviewerDerivation()],
|
|
['fallow', extractFallowDerivation()],
|
|
]) {
|
|
const result = runDerivation(repo, snippet, '06');
|
|
assert.equal(result.status, 0, `${label} exited ${result.status}; stderr=${result.stderr}`);
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
const expected = [`${hashes[PHASE06_PLAN_REL]}^`];
|
|
if (label === 'fallow') {
|
|
assert.deepStrictEqual(fallowBase, expected, `${label}: base must be the phase dir first commit's parent`);
|
|
} else {
|
|
assert.deepStrictEqual(diffBase, expected, `${label}: base must be the phase dir first commit's parent`);
|
|
}
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
test(
|
|
'T3: no committed phase dir fails closed (no silent arbitrary base)',
|
|
SKIP_WIN32,
|
|
() => {
|
|
const { repo } = buildHistory('gsd-3503-closed-', [
|
|
['c1.txt', 'chore: bump to v2.06.0'],
|
|
['c2.txt', 'feat(60-01): probe wiring', 'Deferred to Phase 06 per D-09.'],
|
|
['c3.txt', 'docs(06): capture phase context'],
|
|
['c4.txt', 'docs: touch README'],
|
|
]);
|
|
try {
|
|
for (const [label, snippet] of [
|
|
['tier3', extractTier3Derivation()],
|
|
['spawn_reviewer', extractSpawnReviewerDerivation()],
|
|
['fallow', extractFallowDerivation()],
|
|
]) {
|
|
const result = runDerivation(repo, snippet, '06');
|
|
assert.equal(result.status, 0, `${label} exited ${result.status}; stderr=${result.stderr}`);
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
assert.deepStrictEqual(diffBase, [], `${label}: DIFF_BASE must stay empty without a committed phase dir`);
|
|
assert.deepStrictEqual(fallowBase, [], `${label}: FALLOW_BASE must stay unset`);
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
|
|
// #3995: the milestone-blind repro. A PREVIOUS milestone's phase-02 commit
|
|
// exists in history with a perfectly anchored subject; the current
|
|
// milestone's phase 02 has its own directory. The old derivation's
|
|
// unbounded grep + tail -1 selected the archived milestone's commit and
|
|
// took a 7-file phase to a 3388-file scope; the directory anchor cannot.
|
|
test(
|
|
"T4 (#3995): a previous milestone's same-numbered phase commit never captures the base",
|
|
SKIP_WIN32,
|
|
() => {
|
|
const oldMilestonePhase = path.join('.planning', 'milestones', 'v1.1-phases', '02-old', '02-PLAN.md');
|
|
const currentPhase = path.join('.planning', 'phases', '02-ctx', '02-PLAN.md');
|
|
const { repo, hashes } = buildHistory('gsd-3995-milestone-', [
|
|
[oldMilestonePhase, 'feat(02-01): research-project command, workflow, and template'],
|
|
['mid.txt', 'chore: close milestone v1.1'],
|
|
[currentPhase, 'feat(02-01): current milestone phase 02 plan 01'],
|
|
['c4.txt', 'docs: touch README'],
|
|
]);
|
|
try {
|
|
for (const [label, snippet] of [
|
|
['tier3', extractTier3Derivation()],
|
|
['spawn_reviewer', extractSpawnReviewerDerivation()],
|
|
['fallow', extractFallowDerivation()],
|
|
]) {
|
|
const result = runDerivation(repo, snippet, '02');
|
|
assert.equal(result.status, 0, `${label} exited ${result.status}; stderr=${result.stderr}`);
|
|
const diffBase = parseSentinel(result.stdout, 'DIFF_BASE');
|
|
const fallowBase = parseSentinel(result.stdout, 'FALLOW_BASE');
|
|
const expected = [`${hashes[currentPhase]}^`];
|
|
if (label === 'fallow') {
|
|
assert.deepStrictEqual(fallowBase, expected,
|
|
`${label}: base must be the CURRENT phase dir's first commit, never the archived milestone's (#3995)`);
|
|
} else {
|
|
assert.deepStrictEqual(diffBase, expected,
|
|
`${label}: base must be the CURRENT phase dir's first commit, never the archived milestone's (#3995)`);
|
|
}
|
|
}
|
|
} finally {
|
|
cleanup(repo);
|
|
}
|
|
}
|
|
);
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// #4209 Phase 1 Plan 3 (Task 2) — external reviewer evidence consolidation.
|
|
// gsd-code-reviewer.md must treat <external_reviewer_evidence> as untrusted
|
|
// input: independently re-verify every claim against the actual current
|
|
// source before it can appear in REVIEW.md, fold a verified claim into the
|
|
// SAME Narrative Findings section (no second schema), and never let text
|
|
// embedded inside an evidence file act as an instruction. code-review.md's
|
|
// EXTERNAL_EVIDENCE_BLOCK must keep restating the four fixed prohibitions.
|
|
// ---------------------------------------------------------------------------
|
|
describe('CONS-01..03 — external reviewer evidence consolidation (#4209)', () => {
|
|
function loadStep(src, stepName) {
|
|
const stepStart = src.indexOf(`<step name="${stepName}">`);
|
|
assert.ok(stepStart !== -1, `agent must have a ${stepName} step`);
|
|
const stepEnd = src.indexOf('</step>', stepStart);
|
|
return src.slice(stepStart, stepEnd);
|
|
}
|
|
|
|
test('load_context parses <external_reviewer_evidence> and marks it untrusted', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepSection = loadStep(src, 'load_context');
|
|
assert.ok(stepSection.includes('external_reviewer_evidence'),
|
|
'load_context must parse the external_reviewer_evidence block');
|
|
assert.ok(/untrusted/i.test(stepSection),
|
|
'load_context must explicitly mark external reviewer evidence as untrusted data');
|
|
});
|
|
|
|
test('load_context requires independent re-verification against actual source before accepting a claim', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepSection = loadStep(src, 'load_context');
|
|
assert.ok(/re-open|reopen/i.test(stepSection) && /re-read/i.test(stepSection),
|
|
'load_context must require re-opening and re-reading the actual cited source before accepting an external claim');
|
|
assert.ok(/REJECTED|reject/i.test(stepSection),
|
|
'load_context must state that an unverifiable external claim is rejected, not included');
|
|
});
|
|
|
|
test('load_context resists prompt injection embedded inside evidence text', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepSection = loadStep(src, 'load_context');
|
|
assert.ok(/prompt-injection|prompt injection/i.test(stepSection),
|
|
'load_context must name prompt injection as a threat from evidence content');
|
|
assert.ok(/never a command|not a command/i.test(stepSection),
|
|
'load_context must state evidence text is data, never a command');
|
|
});
|
|
|
|
test('a verified external claim folds into Narrative Findings with no separate schema (CONS-03)', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const stepSection = loadStep(src, 'load_context');
|
|
assert.ok(/Narrative Findings/.test(stepSection),
|
|
'load_context must route a verified external claim into the existing Narrative Findings section');
|
|
const writeReviewSection = loadStep(src, 'write_review');
|
|
assert.ok(/external:/.test(writeReviewSection),
|
|
'write_review must document the (external: {slug}) provenance tag for a verified external finding');
|
|
assert.ok(!/## External/i.test(writeReviewSection),
|
|
'write_review must not introduce a separate External Findings section — one REVIEW.md schema only');
|
|
});
|
|
|
|
test('critical_rules restates the untrusted-evidence contract', () => {
|
|
const src = fs.readFileSync(REVIEWER_PATH, 'utf8');
|
|
const rulesStart = src.indexOf('<critical_rules>');
|
|
const rulesEnd = src.indexOf('</critical_rules>');
|
|
assert.ok(rulesStart !== -1 && rulesEnd !== -1, 'gsd-code-reviewer.md must have a critical_rules section');
|
|
const rulesSection = src.slice(rulesStart, rulesEnd);
|
|
assert.ok(/external_reviewer_evidence|external reviewer/i.test(rulesSection),
|
|
'critical_rules must restate the external-evidence-is-untrusted contract');
|
|
});
|
|
|
|
test('code-review.md restates the four fixed source-review prohibitions when handing off evidence', () => {
|
|
const workflowSrc = fs.readFileSync(WORKFLOW_PATH, 'utf8');
|
|
const blockStart = workflowSrc.indexOf('EXTERNAL_EVIDENCE_BLOCK=$(printf');
|
|
assert.ok(blockStart !== -1, 'code-review.md must build an EXTERNAL_EVIDENCE_BLOCK');
|
|
const blockEnd = workflowSrc.indexOf('\n', workflowSrc.indexOf(')', blockStart));
|
|
const blockText = workflowSrc.slice(blockStart, blockEnd);
|
|
for (const prohibition of ['no source mutation', 'no test execution', 'no background processes', 'no active polling']) {
|
|
assert.ok(blockText.includes(prohibition),
|
|
`EXTERNAL_EVIDENCE_BLOCK must restate "${prohibition}" (SAFE-03..06)`);
|
|
}
|
|
});
|
|
});
|