Files
msd-core/tests/code-review.test.cjs
Dennis Alexis Valin Dittrich 18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00

1387 lines
70 KiB
JavaScript

// allow-test-rule: source-text-is-the-product
// Reads .md/.json/.yml product files whose deployed text IS what the
// runtime loads — testing text content tests the deployed contract.
/**
* GSD Code Review Tests
*
* Validates all code review artifacts from Phases 1-4:
* - Agent frontmatter (gsd-code-reviewer, gsd-code-fixer)
* - Command structure (code-review.md, code-review-fix.md)
* - Workflow structure (code-review.md, code-review-fix.md)
* - Config key registration (workflow.code_review, workflow.code_review_depth)
* - Workflow integration points (execute-phase, quick, autonomous)
*
* Test structure:
* - CR-AGENT: Hermetic agent tests (repo files only)
* - CR-CMD: Hermetic command tests (repo files only)
* - CR-WORKFLOW: Hermetic workflow tests (repo files only)
* - CR-CONFIG: Hermetic config tests (repo files only)
* - CR-INTEGRATION: Conditional integration tests (skip if plugin dir absent)
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { scanFencedBlocks } = require('../gsd-core/bin/lib/markdown-sectionizer.cjs');
const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs');
/** Return the raw text of every ```bash fenced block in `content`. */
function extractBashBlocks(content) {
const lines = content.split(/\r?\n/);
const blocks = [];
for (const block of scanFencedBlocks(lines)) {
if (block.closeLineIdx === -1) continue;
if ((block.infoString || '').trim().toLowerCase() !== 'bash') continue;
blocks.push(lines.slice(block.openLineIdx, block.closeLineIdx + 1).join('\n'));
}
return blocks;
}
const os = require('os');
const { runGsdTools, createTempProject, createTempGitProject, cleanup } = require('./helpers.cjs');
const { runNode } = require('./helpers/process-seam.cjs');
const { escapeRegex } = require('../gsd-core/bin/lib/pattern.cjs');
const REPO_ROOT = path.join(__dirname, '..');
const GSD_TOOLS_BIN = path.join(REPO_ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
// --- Test Environment Setup ---
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
const COMMANDS_DIR = path.join(__dirname, '..', 'commands', 'gsd');
const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows');
/**
* Parse top-level (non-nested, non-escaped) Skill() invocations from a workflow .md file.
*
* Returns an array of structured objects: [{ skill, args }]
* - `skill` is the value of the `skill="..."` keyword argument
* - `args` is the value of the `args="..."` keyword argument (or null if absent)
*
* Skips occurrences inside escaped string contexts like
* prompt="... Skill(skill=\"x\", args=\"y\") ..."
* by walking the file character-by-character and tracking whether we are inside
* a double-quoted string. Escaped quotes (\") are treated as literal content.
*
* This avoids regex/.includes() text-matching: callers receive a structured list
* and assert against fields and tokenized args.
*/
function parseWorkflowSkillInvocations(content) {
const invocations = [];
let i = 0;
let inString = false;
while (i < content.length) {
const ch = content[i];
if (inString) {
if (ch === '\\' && i + 1 < content.length) {
// Skip escape sequence (e.g. \" or \\)
i += 2;
continue;
}
if (ch === '"') {
inString = false;
}
i += 1;
continue;
}
if (ch === '"') {
inString = true;
i += 1;
continue;
}
// Look for top-level "Skill(" at this position
if (content.startsWith('Skill(', i)) {
const callStart = i + 'Skill('.length;
// Find the matching close paren, respecting strings/escapes inside the call
let j = callStart;
let depth = 1;
let innerInString = false;
while (j < content.length && depth > 0) {
const c = content[j];
if (innerInString) {
if (c === '\\' && j + 1 < content.length) {
j += 2;
continue;
}
if (c === '"') innerInString = false;
j += 1;
continue;
}
if (c === '"') {
innerInString = true;
} else if (c === '(') {
depth += 1;
} else if (c === ')') {
depth -= 1;
if (depth === 0) break;
}
j += 1;
}
const callBody = content.slice(callStart, j);
const parsed = parseSkillCallBody(callBody);
if (parsed) invocations.push(parsed);
i = j + 1;
continue;
}
i += 1;
}
return invocations;
}
/**
* Parse the body of a Skill(...) call into { skill, args }.
* Body looks like: skill="name", args="value" (args optional).
* Returns null if no skill keyword is found.
*/
function parseSkillCallBody(body) {
const kwargs = {};
const isIdentChar = (c) => /[A-Za-z0-9_]/.test(c);
const isWs = (c) => /\s/.test(c);
let i = 0;
while (i < body.length) {
// Skip whitespace and commas
while (i < body.length && (isWs(body[i]) || body[i] === ',')) i += 1;
if (i >= body.length) break;
// Read identifier key
const keyStart = i;
while (i < body.length && isIdentChar(body[i])) i += 1;
const key = body.slice(keyStart, i);
if (!key) break;
// Expect '='
while (i < body.length && isWs(body[i])) i += 1;
if (body[i] !== '=') break;
i += 1;
while (i < body.length && isWs(body[i])) i += 1;
// Expect quoted value
if (body[i] !== '"') break;
i += 1;
let value = '';
while (i < body.length) {
const c = body[i];
if (c === '\\' && i + 1 < body.length) {
value += body[i + 1];
i += 2;
continue;
}
if (c === '"') {
i += 1;
break;
}
value += c;
i += 1;
}
kwargs[key] = value;
}
if (!('skill' in kwargs)) return null;
return { skill: kwargs.skill, args: 'args' in kwargs ? kwargs.args : null };
}
// Plugin directory resolution (cross-platform safe)
const PLUGIN_WORKFLOWS_DIR = process.env.GSD_PLUGIN_ROOT || path.join(os.homedir(), '.claude', 'gsd-core', 'workflows');
const PLUGIN_AVAILABLE = fs.existsSync(PLUGIN_WORKFLOWS_DIR);
// --- CR-AGENT: code review agent frontmatter ---
describe('CR-AGENT: code review agent frontmatter', () => {
test('gsd-code-reviewer.md has required frontmatter fields', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-reviewer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('name:'), 'gsd-code-reviewer missing name:');
assert.ok(frontmatter.includes('description:'), 'gsd-code-reviewer missing description:');
assert.ok(frontmatter.includes('tools:'), 'gsd-code-reviewer missing tools:');
assert.ok(frontmatter.includes('color:'), 'gsd-code-reviewer missing color:');
});
test('gsd-code-fixer.md has required frontmatter fields', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('name:'), 'gsd-code-fixer missing name:');
assert.ok(frontmatter.includes('description:'), 'gsd-code-fixer missing description:');
assert.ok(frontmatter.includes('tools:'), 'gsd-code-fixer missing tools:');
assert.ok(frontmatter.includes('color:'), 'gsd-code-fixer missing color:');
});
test('gsd-code-reviewer.md has Read, Bash, Glob, Grep, Write tools', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-reviewer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('Read'), 'gsd-code-reviewer missing Read tool');
assert.ok(frontmatter.includes('Bash'), 'gsd-code-reviewer missing Bash tool');
assert.ok(frontmatter.includes('Glob'), 'gsd-code-reviewer missing Glob tool');
assert.ok(frontmatter.includes('Grep'), 'gsd-code-reviewer missing Grep tool');
assert.ok(frontmatter.includes('Write'), 'gsd-code-reviewer missing Write tool');
});
test('gsd-code-fixer.md has Read, Edit, Write, Bash, Grep, Glob tools', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('Read'), 'gsd-code-fixer missing Read tool');
assert.ok(frontmatter.includes('Edit'), 'gsd-code-fixer missing Edit tool');
assert.ok(frontmatter.includes('Write'), 'gsd-code-fixer missing Write tool');
assert.ok(frontmatter.includes('Bash'), 'gsd-code-fixer missing Bash tool');
});
test('gsd-code-reviewer.md does not have skills: in frontmatter', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-reviewer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(!frontmatter.includes('skills:'),
'gsd-code-reviewer has skills: in frontmatter — breaks Gemini CLI');
});
test('gsd-code-fixer.md does not have skills: in frontmatter', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(!frontmatter.includes('skills:'),
'gsd-code-fixer has skills: in frontmatter — breaks Gemini CLI');
});
test('gsd-code-fixer.md rollback uses git checkout (not Write tool)', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
assert.ok(content.includes('git checkout --'),
'gsd-code-fixer rollback should use git checkout -- {file} for atomic rollback');
assert.ok(!content.includes('PRE_FIX_CONTENT'),
'gsd-code-fixer should not use PRE_FIX_CONTENT in-memory capture (use git checkout instead)');
});
test('gsd-code-fixer.md success_criteria consistent with rollback strategy (git checkout)', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own agent .md content, fixed-size author-controlled content
const successCriteria = content.match(/<success_criteria>([\s\S]*?)<\/success_criteria>/)?.[1] || '';
assert.ok(successCriteria.includes('git checkout'),
'gsd-code-fixer success_criteria must reference git checkout rollback');
assert.ok(!successCriteria.includes('Write tool with captured'),
'gsd-code-fixer success_criteria must not say Write tool for rollback');
});
test('gsd-code-fixer.md flags logic-bug fixes for human review', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
assert.ok(content.includes('requires human verification'),
'gsd-code-fixer should flag logic-bug fixes as requiring human verification');
});
test('gsd-code-reviewer.md REVIEW.md spec includes files_reviewed_list field', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-reviewer.md'), 'utf-8');
assert.ok(content.includes('files_reviewed_list'),
'gsd-code-reviewer REVIEW.md frontmatter spec must include files_reviewed_list for --auto scope persistence');
});
// #2825: gsd-code-fixer is the only writer that hand-rolls a git worktree; it
// must honor workflow.use_worktrees (the documented opt-out) like its four
// sibling writer workflows, and never rm -rf a possible Windows reparse point.
test('#2825 gsd-code-fixer.md reads workflow.use_worktrees and gates git worktree add on it', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
assert.ok(
content.includes('workflow.use_worktrees'),
'gsd-code-fixer setup_worktree must read the workflow.use_worktrees config flag (#2825)',
);
// The git worktree add must be CONDITIONAL on the flag, not unconditional.
// Locate the worktree-add line and confirm a USE_WORKTREES gate precedes it.
assert.ok(
/USE_WORKTREES=.false./.test(content) || content.includes('if [ "$USE_WORKTREES" = "false" ]'),
'gsd-code-fixer must gate worktree creation on USE_WORKTREES=false (skip when opted out) (#2825)',
);
});
test('#2825 gsd-code-fixer.md forbids rm -rf on a possible reparse point (Windows junction safety)', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
assert.ok(
/rm -rf.*reparse point|reparse point.*rm -rf|NEVER .rm -rf.|never use .rm -rf/i.test(content),
'gsd-code-fixer must forbid rm -rf on a possible reparse point/junction (#2825) — on Windows that is the delete-the-target path',
);
});
test('#2825 gsd-code-fixer.md records where verification ran (main checkout vs worktree)', () => {
const content = fs.readFileSync(path.join(AGENTS_DIR, 'gsd-code-fixer.md'), 'utf-8');
assert.ok(
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own agent .md content, fixed-size author-controlled content
/verification[\s\S]*(main checkout|worktree)|(main checkout|worktree)[\s\S]*verification/i.test(content),
'gsd-code-fixer REVIEW-FIX.md must record where verification ran (main checkout vs worktree) so a reader knows if the numbers are reproducible (#2825)',
);
});
});
// --- CR-CMD: code review command structure ---
describe('CR-CMD: code review command structure', () => {
test('code-review.md has correct frontmatter name: gsd:code-review', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('name: gsd:code-review'),
'code-review.md missing correct name in frontmatter');
});
// #2790: code-review-fix.md was consolidated into code-review.md as the --fix flag.
test('code-review.md has --fix flag absorbing code-review-fix (#2790)', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('--fix'),
'code-review.md must document --fix flag (absorbed code-review-fix)');
});
test('code-review.md references workflow: code-review.md', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('code-review.md'),
'code-review.md does not reference its workflow');
});
test('code-review.md references code-review-fix workflow via --fix (#2790)', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('code-review-fix') || content.includes('--fix'),
'code-review.md must reference code-review-fix workflow or --fix flag');
});
test('code-review.md has argument-hint in frontmatter', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('argument-hint:'),
'code-review.md missing argument-hint');
});
test('code-review.md argument-hint includes --fix flag (#2790: absorbed code-review-fix)', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('argument-hint:') && content.includes('--fix'),
'code-review.md must have argument-hint with --fix');
});
test('code-review.md has allowed-tools in frontmatter', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('allowed-tools:'),
'code-review.md missing allowed-tools');
});
test('code-review.md has allowed-tools in frontmatter (covers fix too, #2790)', () => {
const content = fs.readFileSync(path.join(COMMANDS_DIR, 'code-review.md'), 'utf-8');
const frontmatter = content.split('---')[1] || '';
assert.ok(frontmatter.includes('allowed-tools:'),
'code-review.md missing allowed-tools');
});
});
// --- CR-WORKFLOW: code review workflow structure ---
describe('CR-WORKFLOW: code review workflow structure', () => {
test('code-review.md workflow has <step name="initialize">', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('<step name="initialize">'),
'code-review.md workflow missing initialize step');
});
test('code-review.md workflow has <step name="check_config_gate">', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('<step name="check_config_gate">'),
'code-review.md workflow missing check_config_gate step');
});
test('code-review.md workflow references gsd-code-reviewer agent', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
assert.ok(content.includes('gsd-code-reviewer'),
'code-review.md workflow does not reference gsd-code-reviewer agent');
});
test('code-review-fix.md workflow has <step name="initialize">', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review-fix.md'), 'utf-8');
assert.ok(content.includes('<step name="initialize">'),
'code-review-fix.md workflow missing initialize step');
});
test('code-review-fix.md workflow references gsd-code-fixer agent', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review-fix.md'), 'utf-8');
assert.ok(content.includes('gsd-code-fixer'),
'code-review-fix.md workflow does not reference gsd-code-fixer agent');
});
test('code-review-fix.md workflow has iteration cap', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review-fix.md'), 'utf-8');
// Check for iteration logic with cap
assert.ok(content.includes('MAX_ITERATIONS') || (content.includes('3') && content.includes('iteration')),
'code-review-fix.md workflow missing iteration cap logic');
});
test('code-review.md --files path traversal guard rejects paths outside repo', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
// Guard must resolve and compare against REPO_ROOT
assert.ok(content.includes('REPO_ROOT') && content.includes('realpath'),
'code-review.md missing path traversal guard (realpath + REPO_ROOT check)');
assert.ok(content.includes('File path outside repository'),
'code-review.md missing rejection message for paths outside repo');
});
test('code-review.md uses portable while-read loop for array dedup (not mapfile)', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
// mapfile is bash 4+ only; macOS ships bash 3.2. Dedup must use portable while-read.
// Note: 'mapfile' may appear in platform_notes documentation — check bash code blocks only
const codeBlocks = extractBashBlocks(content);
const hasMapfileInCode = codeBlocks.some(block => block.includes('mapfile -t'));
assert.ok(!hasMapfileInCode,
'code-review.md bash code blocks use mapfile which is bash 4+ only — breaks macOS default bash 3.2');
assert.ok(content.includes('while IFS= read -r'),
'code-review.md should use portable while-read loop instead of mapfile');
});
test('code-review-fix.md uses portable while-read loop for array construction (not mapfile)', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review-fix.md'), 'utf-8');
const codeBlocks = extractBashBlocks(content);
const hasMapfileInCode = codeBlocks.some(block => block.includes('mapfile -t'));
assert.ok(!hasMapfileInCode,
'code-review-fix.md bash code blocks use mapfile which is bash 4+ only — breaks macOS default bash 3.2');
assert.ok(content.includes('while IFS= read -r'),
'code-review-fix.md should use portable while-read loop instead of mapfile');
});
// #3661: configurable code-review hook point (see .gsd/phase/feat-3661-code-review-hook-point/
// 40-design.md and 50-test-matrix.md Section G) — manual invocation reads
// workflow.code_review directly (independent of the automatic loop-point
// selector workflow.code_review_point) instead of gating on registry
// presence at the hardcoded execute:post point; and Tier 2/3 file scoping
// narrows to what changed since the phase's last review.
describe('#3661: configurable code-review hook point (test matrix Section G)', () => {
function extractStepBody(content, stepName) {
const re = new RegExp(`<step name="${stepName}">([\\s\\S]*?)<\\/step>`);
const m = content.match(re);
return m ? m[1] : null;
}
test('G1: checkConfigGateReadsCodeReviewConfigDirectly', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const stepContent = extractStepBody(content, 'check_config_gate');
assert.ok(stepContent, 'code-review.md missing check_config_gate step');
assert.ok(/gsd_run query config-get workflow\.code_review\b/.test(stepContent),
'check_config_gate must read workflow.code_review via gsd_run query config-get');
});
test('G2: checkConfigGateNoLongerProbesExecutePostHooks', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const stepContent = extractStepBody(content, 'check_config_gate');
assert.ok(stepContent, 'code-review.md missing check_config_gate step');
assert.ok(!stepContent.includes('render-hooks execute:post'),
'check_config_gate must no longer gate on render-hooks execute:post — manual invocation must work regardless of workflow.code_review_point');
});
test('G3: computeFileScopeDerivesLastReviewCommit', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const stepContent = extractStepBody(content, 'compute_file_scope');
assert.ok(stepContent, 'code-review.md missing compute_file_scope step');
assert.ok(
/LAST_REVIEW_COMMIT=\$\(git log --format=%H -1 -- "\$\{PHASE_DIR\}\/\$\{PADDED_PHASE\}-REVIEW\.md"/.test(stepContent),
'compute_file_scope must derive LAST_REVIEW_COMMIT from the phase REVIEW.md git history',
);
});
test('G4: tier2SkipsUnchangedSummariesSinceLastReview', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const codeBlocks = extractBashBlocks(content);
const tier2Block = codeBlocks.find(block =>
block.includes('for summary in $(printf') && block.includes('LAST_REVIEW_COMMIT'));
assert.ok(tier2Block, 'code-review.md Tier 2 SUMMARY loop must reference LAST_REVIEW_COMMIT');
assert.ok(
/git diff --quiet "\$\{LAST_REVIEW_COMMIT\}" HEAD -- "\$summary"/.test(tier2Block),
'Tier 2 must contain a `git diff --quiet "${LAST_REVIEW_COMMIT}" HEAD -- "$summary"` skip-conditional',
);
assert.ok(/\bcontinue\b/.test(tier2Block),
'Tier 2 unchanged-since-last-review guard must `continue` (skip) the summary, not just log');
});
test('G5: tier3PrefersLastReviewCommitOverPhaseStartDerivation', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const codeBlocks = extractBashBlocks(content);
// #3995 replaced the old commit-message-grep PHASE_COMMITS derivation with
// PHASE_START (git log --diff-filter=A -- "${PHASE_DIR}") — a phase number is
// only unique within a milestone, not the whole repo, so #3661's fallback
// chain rides on whichever derivation is current rather than pinning the old name.
const tier3Block = codeBlocks.find(block =>
block.includes('DIFF_BASE=""') && block.includes('PHASE_START'));
assert.ok(tier3Block, 'code-review.md Tier 3 DIFF_BASE derivation block not found');
const lastReviewIdx = tier3Block.indexOf('if [ -n "$LAST_REVIEW_COMMIT" ]; then');
const phaseStartElifIdx = tier3Block.indexOf('elif [ -n "$PHASE_START" ]; then');
assert.ok(lastReviewIdx !== -1, 'Tier 3 DIFF_BASE must check LAST_REVIEW_COMMIT');
assert.ok(phaseStartElifIdx !== -1, 'Tier 3 DIFF_BASE must fall back to PHASE_START via elif (#3995 derivation unchanged)');
assert.ok(lastReviewIdx < phaseStartElifIdx,
'LAST_REVIEW_COMMIT must be checked BEFORE the PHASE_START fallback, so a prior review narrows the diff base');
assert.ok(/DIFF_BASE="\$LAST_REVIEW_COMMIT"/.test(tier3Block),
'Tier 3 must set DIFF_BASE directly from LAST_REVIEW_COMMIT when present (no ^ parent offset)');
});
test('G6: tier2GuardIsNoOpWhenLastReviewCommitEmpty', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
const codeBlocks = extractBashBlocks(content);
const tier2Block = codeBlocks.find(block =>
block.includes('for summary in $(printf') && block.includes('LAST_REVIEW_COMMIT'));
assert.ok(tier2Block, 'code-review.md Tier 2 SUMMARY loop must reference LAST_REVIEW_COMMIT');
// The skip-conditional's guard must require LAST_REVIEW_COMMIT to be
// non-empty (`[ -n "$LAST_REVIEW_COMMIT" ]`) as the FIRST operand of an
// `&&` chain, so on a phase's first review (LAST_REVIEW_COMMIT="") the
// conditional is structurally a no-op — bash short-circuits `&&` before
// ever reaching `git diff --quiet`, so nothing can be skipped.
assert.ok(
/if \[ -n "\$LAST_REVIEW_COMMIT" \] && git diff --quiet/.test(tier2Block),
'Tier 2 skip-conditional must guard on `[ -n "$LAST_REVIEW_COMMIT" ]` as the first `&&` operand, so it is a structural no-op when LAST_REVIEW_COMMIT is empty (first review)',
);
});
});
});
// --- CR-CONFIG: config key registration ---
describe('CR-CONFIG: config key registration', () => {
test('config-set accepts workflow.code_review', () => {
const tmpDir = createTempProject();
try {
const result = runGsdTools('config-set workflow.code_review true', tmpDir);
assert.ok(result.success, `config-set should accept workflow.code_review: ${result.error}`);
} finally {
cleanup(tmpDir);
}
});
test('config-set accepts workflow.code_review_depth', () => {
const tmpDir = createTempProject();
try {
const result = runGsdTools('config-set workflow.code_review_depth standard', tmpDir);
assert.ok(result.success, `config-set should accept workflow.code_review_depth: ${result.error}`);
} finally {
cleanup(tmpDir);
}
});
test('config-get workflow.code_review returns value set via config-set', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const setResult = runGsdTools(['config-set', 'workflow.code_review', 'true'], tmpDir);
assert.ok(setResult.success, `config-set workflow.code_review failed: ${setResult.error}`);
const getResult = runGsdTools(['config-get', 'workflow.code_review'], tmpDir);
assert.ok(getResult.success, `config-get workflow.code_review failed: ${getResult.error}`);
assert.strictEqual(getResult.output, 'true',
`workflow.code_review should return "true", got ${getResult.output}`);
});
test('config-get workflow.code_review_depth returns value set via config-set', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const setResult = runGsdTools(['config-set', 'workflow.code_review_depth', 'standard'], tmpDir);
assert.ok(setResult.success, `config-set workflow.code_review_depth failed: ${setResult.error}`);
const getResult = runGsdTools(['config-get', 'workflow.code_review_depth'], tmpDir);
assert.ok(getResult.success, `config-get workflow.code_review_depth failed: ${getResult.error}`);
assert.strictEqual(getResult.output, '"standard"',
`workflow.code_review_depth should return '"standard"', got ${getResult.output}`);
});
// ── #3661: workflow.code_review_point — CLI-behavioral (50-test-matrix.md
// Section F, rows F4-F7). Uses createTempGitProject + runGsdTools + the real
// gsd-tools subprocess (via runNode), not source-grep, per the matrix's
// coverage-strategy note.
function renderHooksEnvelope(tmpDir, point) {
const result = runNode(
[GSD_TOOLS_BIN, 'loop', 'render-hooks', point, '--cwd', tmpDir],
{ cwd: REPO_ROOT, timeoutMs: 15000 },
);
assert.strictEqual(result.exitCode, 0, `Expected exit 0 for render-hooks ${point}. stderr: ` + (result.stderr || ''));
return JSON.parse(result.stdout.trim());
}
test('F4: renderHooksExecutePostActiveByDefault', (t) => {
const tmpDir = createTempGitProject();
t.after(() => cleanup(tmpDir));
const envelope = renderHooksEnvelope(tmpDir, 'execute:post');
const step = envelope.activeHooks.find((h) => h.capId === 'code-review' && h.kind === 'step');
assert.ok(step, 'Expected an active code-review step at execute:post by default. Got: ' + JSON.stringify(envelope.activeHooks));
});
test('F5: renderHooksExecuteWavePostInactiveByDefault', (t) => {
const tmpDir = createTempGitProject();
t.after(() => cleanup(tmpDir));
const envelope = renderHooksEnvelope(tmpDir, 'execute:wave:post');
const step = envelope.activeHooks.find((h) => h.capId === 'code-review' && h.kind === 'step');
assert.strictEqual(step, undefined, 'code-review step must be inactive at execute:wave:post by default (point not selected). Got: ' + JSON.stringify(envelope.activeHooks));
});
test('F6: renderHooksFlipsWithConfigSetCodeReviewPoint', (t) => {
const tmpDir = createTempGitProject();
t.after(() => cleanup(tmpDir));
const setResult = runGsdTools(['config-set', 'workflow.code_review_point', 'execute:wave:post'], tmpDir);
assert.ok(setResult.success, `config-set workflow.code_review_point failed: ${setResult.error}`);
const postEnvelope = renderHooksEnvelope(tmpDir, 'execute:post');
const wavePostEnvelope = renderHooksEnvelope(tmpDir, 'execute:wave:post');
const postStep = postEnvelope.activeHooks.find((h) => h.capId === 'code-review' && h.kind === 'step');
const wavePostStep = wavePostEnvelope.activeHooks.find((h) => h.capId === 'code-review' && h.kind === 'step');
assert.strictEqual(postStep, undefined, 'execute:post code-review step must be inactive once flipped to execute:wave:post');
assert.ok(wavePostStep, 'execute:wave:post code-review step must become active once the point is flipped');
});
test('F7: configSetRejectsOutOfEnumCodeReviewPoint', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const result = runGsdTools(['config-set', 'workflow.code_review_point', 'bogus'], tmpDir);
assert.strictEqual(result.success, false, 'config-set must reject an out-of-enum workflow.code_review_point value');
assert.match(
result.error || '',
/Invalid workflow\.code_review_point/,
`Expected an enum-rejection error, got: ${result.error}`,
);
// The rejected value must not have been persisted.
const getResult = runGsdTools(['config-get', 'workflow.code_review_point'], tmpDir);
assert.ok(getResult.success, `config-get workflow.code_review_point failed: ${getResult.error}`);
assert.notStrictEqual(getResult.output, '"bogus"', 'out-of-enum value must not be silently accepted/persisted');
});
});
// --- CR-REVIEWER-LANES: optional external source-reviewer dispatch (#4209) ---
describe('CR-REVIEWER-LANES: optional external source-reviewer dispatch (#4209)', () => {
const workflowContent = fs.readFileSync(path.join(WORKFLOWS_DIR, 'code-review.md'), 'utf-8');
test('code-review.md workflow has <step name="dispatch_reviewer_lanes">', () => {
assert.ok(workflowContent.includes('<step name="dispatch_reviewer_lanes">'),
'code-review.md workflow missing dispatch_reviewer_lanes step');
});
// #4209 (round 5 review): every prior "one fence, not two" fix in this step was verified by
// manually extracting a SUB-SLICE of the step and pre-seeding the variables that slice reads
// (e.g. CODE_REVIEW_POINT set by the test driver, not by the fence itself) — which is exactly
// why a SIBLING cross-fence bug on CODE_REVIEW_POINT itself went undetected for a full review
// round even after the EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS instance was fixed. This test
// extracts and executes the step's ENTIRE bash content as the workflow author's own execution
// model actually runs it (one process, nothing pre-seeded except genuinely external inputs),
// and additionally asserts there is exactly one fence — a structural invariant that makes any
// future accidental re-split fail loudly here instead of silently at runtime.
function extractDispatchReviewerLanesFences() {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepLines = splitLines(stepMatch[1]);
return scanFencedBlocks(stepLines)
.filter((b) => b.infoString.trim().toLowerCase() === 'bash' && b.closeLineIdx !== -1)
.map((b) => stepLines.slice(b.openLineIdx + 1, b.closeLineIdx).join('\n'));
}
test('dispatch_reviewer_lanes is exactly one bash fence (no cross-fence variable read can reappear)', () => {
const fences = extractDispatchReviewerLanesFences();
assert.equal(fences.length, 1,
`expected dispatch_reviewer_lanes to be exactly one continuous bash fence, found ${fences.length} — a split fence means any variable set in one and read in another is silently empty (this file's own documented execution model: fenced blocks do not share shell state)`);
});
test('dispatch_reviewer_lanes computes CODE_REVIEW_POINT and dispatches in the SAME process, end to end (#4209 round 5)', () => {
const [fence] = extractDispatchReviewerLanesFences();
const tmpDir = createTempGitProject();
const dispatchArgsPath = path.join(tmpDir, 'dispatch-args.txt');
try {
// `review-lane dispatch-step --explicit codex` would spawn the REAL `codex` CLI (present on
// this machine) via runner.runLane, which then blocks on interactive auth with no stdin —
// a genuine hang, not a test artifact (#4209 round 5, BL-01). This test's subject is the
// FENCE's own bash control flow (CODE_REVIEW_POINT/EXPLICIT_JOINED computed correctly and
// threaded into dispatch-step's argv) — not the external CLI dispatch-step goes on to spawn.
// `gsd_run` stays real for `config-get`/`review-lane explicit-from-argv` (what this test
// verifies) and is short-circuited ONLY for `review-lane dispatch-step`, whose argv is
// captured to a file for assertion instead of executed for real.
const driver = [
`GSD_TOOLS=${JSON.stringify(GSD_TOOLS_BIN)}`,
`DISPATCH_ARGS_PATH=${JSON.stringify(dispatchArgsPath)}`,
'gsd_run() {',
' if [ "$1" = "review-lane" ] && [ "$2" = "dispatch-step" ]; then',
' printf \'%s\\n\' "$@" > "$DISPATCH_ARGS_PATH"',
' cat >/dev/null',
' echo \'{"ok":true,"dispatched":false,"selection":{},"results":[]}\'',
' return 0',
' fi',
' node "$GSD_TOOLS" "$@"',
'}',
`REPO_ROOT=${JSON.stringify(tmpDir)}`,
'REVIEW_DEPTH=standard',
'DIFF_BASE=deadbeef',
'REVIEW_FILES=(src/foo.ts)',
'set -- --codex',
fence,
'echo "===RESULT==="',
'echo "CODE_REVIEW_POINT=[$CODE_REVIEW_POINT]"',
'echo "EXPLICIT_JOINED=[$EXPLICIT_JOINED]"',
].join('\n');
const result = require('node:child_process').spawnSync('bash', ['-c', driver], { cwd: tmpDir, encoding: 'utf8', timeout: 15000 });
assert.equal(result.status, 0, `driver failed: stdout=${result.stdout} stderr=${result.stderr}`);
assert.match(result.stdout, /CODE_REVIEW_POINT=\[execute:post\]/,
`CODE_REVIEW_POINT must be computed and survive within the SAME fence that later uses it for --point, got: ${result.stdout}`);
assert.match(result.stdout, /EXPLICIT_JOINED=\[codex\]/,
`EXPLICIT_JOINED must resolve --codex to its canonical slug within the same fence, got: ${result.stdout}`);
const dispatchArgs = splitLines(fs.readFileSync(dispatchArgsPath, 'utf-8')).filter(Boolean);
assert.ok(dispatchArgs.includes('--point'), `dispatch-step argv missing --point: ${dispatchArgs.join(' ')}`);
assert.equal(dispatchArgs[dispatchArgs.indexOf('--point') + 1], 'execute:post',
`dispatch-step must receive the SAME CODE_REVIEW_POINT the fence computed, got argv: ${dispatchArgs.join(' ')}`);
assert.ok(dispatchArgs.includes('--explicit'), `dispatch-step argv missing --explicit: ${dispatchArgs.join(' ')}`);
assert.equal(dispatchArgs[dispatchArgs.indexOf('--explicit') + 1], 'codex',
`dispatch-step must receive the SAME EXPLICIT_JOINED the fence computed, got argv: ${dispatchArgs.join(' ')}`);
} finally {
cleanup(tmpDir);
}
});
test('dispatch_reviewer_lanes passes --cap-id/--point to dispatch-step instead of resolving the trait itself', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
assert.match(stepContent, /--cap-id code-review --point "\$CODE_REVIEW_POINT"/,
'must delegate trait resolution to dispatch-step via --cap-id/--point, not scrape loop render-hooks itself (#4209 maintainer redirect: no per-workflow hand-wiring of the gate)');
assert.ok(!/loop render-hooks/.test(stepContent),
'the workflow must not call loop render-hooks itself — that belongs to dispatch-step, the reusable seam');
});
test('dispatch_reviewer_lanes delegates roster-flag matching to review-lane explicit-from-argv, not an inline node -e (#4209 RQ-02)', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
assert.match(stepContent, /review-lane explicit-from-argv -- "\$@"/,
'must delegate roster/flag matching to the shared explicit-from-argv subcommand');
assert.ok(!/mergeReviewerLanes/.test(stepContent),
'the workflow must not re-implement the roster merge inline — that duplicate is exactly what RQ-02 removed');
});
// #4209 (maintainer redirect): the trait must be enforced by the shared `dispatch-step` CLI
// itself, not trusted from a caller-passed boolean — otherwise a second capability reusing this
// seam gets zero enforcement from declaring the trait alone. These run the REAL command against
// the REAL first-party capability registry (capabilities/code-review/capability.json), not a
// stubbed value.
test('review-lane dispatch-step: --cap-id code-review --point execute:post resolves the real trait as true', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir, '--explicit', 'not-a-real-reviewer-xyz',
'--cap-id', 'code-review', '--point', 'execute:post', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
const parsed = JSON.parse(result.stdout.trim());
// An unresolvable slug still proves the trait gate was passed: TRAIT_NOT_ENABLED short-
// circuits before selection is ever attempted (dispatched:false, ok:true), whereas a real
// selection failure on a resolved (trait-enabled) dispatch is ok:false with selection.errors.
assert.notStrictEqual(parsed.reason, 'trait_not_enabled',
`expected the real code-review capability step's trait to be enabled, got: ${JSON.stringify(parsed)}`);
assert.strictEqual(parsed.ok, false, 'an unresolvable explicit lane past a passed trait gate must still be a reported failure');
} finally {
cleanup(tmpDir);
}
});
test('review-lane dispatch-step: an unknown --cap-id resolves the trait as false (fails closed, not open)', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir, '--explicit', 'codex',
'--cap-id', 'no-such-capability-xyz', '--point', 'execute:post', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
const parsed = JSON.parse(result.stdout.trim());
assert.strictEqual(parsed.dispatched, false, 'a capId whose trait is not enabled must dispatch nothing');
assert.strictEqual(parsed.reason, 'trait_not_enabled');
assert.deepStrictEqual(parsed.results, [], 'no lane may run when the trait is not enabled');
} finally {
cleanup(tmpDir);
}
});
test('review-lane dispatch-step: omitting --cap-id/--point resolves the trait as false (no context means no opt-in)', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir, '--explicit', 'codex', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
const parsed = JSON.parse(result.stdout.trim());
assert.strictEqual(parsed.reason, 'trait_not_enabled',
`a caller with no --cap-id/--point context must not be silently opted in, got: ${JSON.stringify(parsed)}`);
} finally {
cleanup(tmpDir);
}
});
test('review-lane dispatch-step: --cap-id without --point warns (misconfigured, not opted out) (#4209 RQ-03)', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir, '--explicit', 'codex',
'--cap-id', 'code-review', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
assert.match(result.stderr, /--cap-id and --point must both be given/,
`a --cap-id with no --point must warn distinctly from a correct no-context opt-out, got stderr: ${result.stderr}`);
const parsed = JSON.parse(result.stdout.trim());
assert.strictEqual(parsed.reason, 'trait_not_enabled');
} finally {
cleanup(tmpDir);
}
});
test('review-lane explicit-from-argv matches CLI flags against the merged roster (#4209 RQ-02)', () => {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'explicit-from-argv', '--', '--codex', '--agy'],
{ cwd: REPO_ROOT, timeoutMs: 15000 },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
assert.strictEqual(result.stdout.trim(), 'antigravity,codex');
});
test('review-lane explicit-from-argv resolves to empty when no known flag is present', () => {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'explicit-from-argv', '--'],
{ cwd: REPO_ROOT, timeoutMs: 15000 },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
assert.strictEqual(result.stdout.trim(), '');
});
test('dispatch_reviewer_lanes step derives explicit flags from the roster, not a hand-maintained list', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own workflow markdown, bounded author-controlled prose
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
assert.ok(stepMatch, 'dispatch_reviewer_lanes step not found');
const stepContent = stepMatch[1];
assert.ok(stepContent.includes('review-lane-descriptor.cjs'),
'dispatch_reviewer_lanes must derive flags from the canonical review-lane-descriptor roster');
assert.ok(!/\[\s*['"]--(codex|agy|gemini|claude)['"]/.test(stepContent),
'dispatch_reviewer_lanes must not hand-maintain a static reviewer-flag array (DOCS-03 / anti-pattern)');
});
test('dispatch_reviewer_lanes calls review-lane dispatch-step exactly once', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
const calls = stepContent.match(/gsd_run review-lane dispatch-step/g) || [];
assert.strictEqual(calls.length, 1,
`dispatch_reviewer_lanes must call review-lane dispatch-step exactly once, found ${calls.length}`);
});
test('dispatch_reviewer_lanes passes already-resolved repo root, depth, and base SHA (SAFE-01)', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
assert.ok(stepContent.includes('--repo-root "$REPO_ROOT"'), 'must pass already-resolved REPO_ROOT');
assert.ok(stepContent.includes('--depth "$REVIEW_DEPTH"'), 'must pass already-resolved REVIEW_DEPTH');
assert.ok(stepContent.includes('--base-sha "$DIFF_BASE"'), 'must pass already-resolved DIFF_BASE');
});
test('dispatch_reviewer_lanes explains and skips rather than silently failing when explicit lanes have no resolvable DIFF_BASE', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
assert.ok(/\[\s+\${#EXPLICIT_REVIEWER_SLUGS\[@\]}\s+-gt\s+0\s+\]\s+&&\s+\[\s+-z\s+"\$DIFF_BASE"\s+\]/.test(stepContent),
'must guard against dispatching with an empty DIFF_BASE when lanes were explicitly requested');
assert.match(stepContent, /no diff base could be resolved/,
'must explain why explicitly requested lanes did not run, rather than leaving the generic missing_provenance rejection unexplained');
});
test('dispatch_reviewer_lanes is a no-op when no explicit reviewer-lane flag is present (COMP-01)', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
assert.ok(/if\s*\[\s*\$\{#EXPLICIT_REVIEWER_SLUGS\[@\]\}\s*-gt\s*0\s*\]/.test(stepContent),
'dispatch_reviewer_lanes must gate the dispatch-step call behind a non-empty explicit selection');
});
test('spawn_reviewer prompt interpolates ${EXTERNAL_EVIDENCE_BLOCK}', () => {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="spawn_reviewer">([\s\S]*?)<\/step>/);
assert.ok(stepMatch, 'spawn_reviewer step not found');
assert.ok(stepMatch[1].includes('${EXTERNAL_EVIDENCE_BLOCK}'),
'spawn_reviewer must interpolate EXTERNAL_EVIDENCE_BLOCK into the agent prompt');
});
test('external evidence block marks findings as unverified and requires re-opening source (CONS-02)', () => {
// Scoped to the EXTERNAL_EVIDENCE_BLOCK assignment line itself (#4209 review): a whole-file
// match on `workflowContent` would still pass if UNVERIFIED and re-open/reopen appeared in
// two unrelated parts of this 1000+-line workflow, which proves nothing about the actual
// evidence block's contract. Line-filtered via splitLines (not a bare-`\n` regex spanning
// readFileSync content) so this stays CRLF-portable.
const blockLine = splitLines(workflowContent).find((l) => l.includes('EXTERNAL_EVIDENCE_BLOCK=$(printf'));
assert.ok(blockLine, 'EXTERNAL_EVIDENCE_BLOCK assignment not found');
assert.ok(/UNVERIFIED/.test(blockLine) && /re-open|reopen/i.test(blockLine),
'external evidence block must mark external findings as unverified and require the internal reviewer to re-open cited source');
});
// --- Real subprocess behavior: `review-lane dispatch-step` (fail-closed, no raw fallback) ---
test('review-lane dispatch-step is a no-op with no --explicit selection', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir,
'--cap-id', 'code-review', '--point', 'execute:post', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
const parsed = JSON.parse(result.stdout.trim());
assert.strictEqual(parsed.dispatched, false, 'no explicit selection must dispatch zero lanes');
assert.strictEqual(parsed.reason, 'no_lanes_selected');
} finally {
cleanup(tmpDir);
}
});
test('review-lane dispatch-step fails closed on an explicitly requested unknown lane (SAFE-07)', () => {
const tmpDir = createTempGitProject();
try {
const result = runNode(
[GSD_TOOLS_BIN, 'review-lane', 'dispatch-step',
'--repo-root', tmpDir, '--depth', 'standard', '--base-sha', 'deadbeef',
'--run-dir', tmpDir, '--cwd', tmpDir, '--explicit', 'not-a-real-reviewer-xyz',
'--cap-id', 'code-review', '--point', 'execute:post', '--raw'],
{ cwd: REPO_ROOT, timeoutMs: 15000, input: 'src/foo.ts\n' },
);
assert.strictEqual(result.exitCode, 0, `expected exit 0, stderr: ${result.stderr || ''}`);
const parsed = JSON.parse(result.stdout.trim());
assert.strictEqual(parsed.dispatched, false, 'an unresolvable explicit lane must plan/invoke nothing');
assert.strictEqual(parsed.ok, false, 'an explicitly requested unavailable lane must be a failure, not a silent success');
assert.deepStrictEqual(parsed.results, [], 'no lane fallback result may appear');
assert.ok(
(parsed.selection.errors || []).some((e) => e.includes('not-a-real-reviewer-xyz')),
'the unresolved slug must be named in the selection errors',
);
} finally {
cleanup(tmpDir);
}
});
// --- #4209 R4: execute the actual EVIDENCE_LIST reducer extracted from the workflow, not a
// reimplementation of it, so a regression in the real markdown fails this test (see B1: the
// reducer must warn on parsed.ok===false, not just per-lane results[]/selection.errors). ---
function extractEvidenceReducer() {
// eslint-disable-next-line local/no-unbounded-quantifier -- bounded author-controlled workflow markdown
const stepMatch = workflowContent.match(/<step name="dispatch_reviewer_lanes">([\s\S]*?)<\/step>/);
const stepContent = stepMatch[1];
const startMarker = 'if [[ "$DISPATCH_JSON" == @file:* ]]; then';
const start = stepContent.indexOf(startMarker);
assert.ok(start !== -1, 'expected the EVIDENCE_LIST reducer in dispatch_reviewer_lanes');
const end = stepContent.indexOf('\n ")', start);
assert.ok(end !== -1, 'unterminated EVIDENCE_LIST reducer fence');
return stepContent.slice(start, end + '\n ")'.length);
}
function runReducer(dispatchJson) {
const script = `DISPATCH_JSON=${JSON.stringify(dispatchJson)}\n${extractEvidenceReducer()}\necho "$EVIDENCE_LIST"`;
const result = require('node:child_process').spawnSync('bash', ['-c', script], { encoding: 'utf8', timeout: 15000 });
return { stdout: result.stdout, stderr: result.stderr, status: result.status };
}
test('EVIDENCE_LIST reducer warns on a whole-dispatch rejection (B1), not just per-lane failures', () => {
const { stdout, stderr } = runReducer(JSON.stringify({
dispatched: false, ok: false, reason: 'missing_provenance', results: [],
}));
assert.match(stderr, /external reviewer dispatch rejected \(missing_provenance\)/,
`expected a whole-dispatch rejection warning, got stderr: ${stderr}`);
assert.strictEqual(stdout.trim(), '', 'a rejected dispatch must produce no evidence lines');
});
test('EVIDENCE_LIST reducer still warns per-lane and still emits evidence for lanes that succeeded', () => {
const { stdout, stderr } = runReducer(JSON.stringify({
dispatched: true,
ok: false,
results: [
{ slug: 'codex', ok: true, reviewPath: '/tmp/gsd-review-codex.md' },
{ slug: 'agy', ok: false, reason: 'invoke_failed', detail: 'binary not found' },
],
}));
assert.match(stderr, /external reviewer lane 'agy' failed \(invoke_failed: binary not found\)/,
`expected a per-lane failure warning, got stderr: ${stderr}`);
assert.strictEqual(stdout.trim(), '- codex: /tmp/gsd-review-codex.md',
'the lane that succeeded must still produce an evidence line');
});
test('EVIDENCE_LIST reducer unwraps the @file: overflow protocol (R1)', () => {
const tmpDir = createTempGitProject();
try {
const payloadPath = path.join(tmpDir, 'dispatch-result.json');
fs.writeFileSync(payloadPath, JSON.stringify({
dispatched: true, ok: true,
results: [{ slug: 'codex', ok: true, reviewPath: '/tmp/gsd-review-codex.md' }],
}));
const { stdout, stderr } = runReducer(`@file:${payloadPath}`);
assert.strictEqual(stdout.trim(), '- codex: /tmp/gsd-review-codex.md',
`expected the @file:-wrapped payload to be unwrapped and parsed, got stdout: ${stdout} stderr: ${stderr}`);
} finally {
cleanup(tmpDir);
}
});
});
// --- CR-INTEGRATION: workflow integration points ---
describe('CR-INTEGRATION: workflow integration points', () => {
test('execute-phase.md contains code_review_gate step', { skip: !PLUGIN_AVAILABLE ? 'Plugin dir not installed' : false }, () => {
const content = fs.readFileSync(path.join(PLUGIN_WORKFLOWS_DIR, 'execute-phase.md'), 'utf-8');
assert.ok(content.includes('code_review_gate'),
'execute-phase.md missing code_review_gate step name');
});
test('execute-phase.md resolves code-review capability hook', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'execute-phase.md'), 'utf-8');
// eslint-disable-next-line local/no-unbounded-quantifier -- parses maintainer-authored workflow markdown, bounded prose, not adversarial input
const gateMatch = content.match(/<step name="code_review_gate"[^>]*>([\s\S]*?)<\/step>/);
assert.ok(gateMatch, 'execute-phase.md missing code_review_gate step');
const gateContent = gateMatch[1];
assert.ok(gateContent.includes('loop render-hooks execute:post'),
'execute-phase.md code_review_gate must resolve execute:post capability hooks');
assert.ok(gateContent.includes('ref.skill == "code-review"'),
'execute-phase.md code_review_gate must identify the code-review capability hook');
assert.ok(!gateContent.match(/config-get\s+workflow\.code_review/),
'execute-phase.md code_review_gate must not read workflow.code_review directly');
});
// #3661: the generic execute:wave:post step-dispatch contract
// (loop-hook-dispatch.md § step) invokes `Skill(skill="gsd-<ref.skill>")` with no
// phase argument, but code-review.md's `initialize` step requires one positionally
// (`PHASE_ARG="${1}"`) — without it the review reports "Phase not found" and exits.
// Caught by the orthogonal spec review; fixed with a precedented carve-out in step
// 5.75 mirroring code_review_gate's own `args="${PHASE_NUMBER}"` invocation.
test('execute-phase.md wave-post dispatch passes PHASE_NUMBER to the code-review skill', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'execute-phase.md'), 'utf-8');
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own workflow markdown, bounded author-controlled prose
const stepMatch = content.match(/5\.75\.[\s\S]*?(?=\r?\n5\.8\.)/);
assert.ok(stepMatch, 'execute-phase.md missing step 5.75 (execute:wave:post capability dispatch)');
const stepContent = stepMatch[0];
assert.ok(stepContent.includes('ref.skill == "code-review"'),
'step 5.75 must carve out ref.skill == "code-review" from the generic step-dispatch contract');
assert.ok(/Skill\(skill="gsd-code-review",\s*args="\$\{PHASE_NUMBER\}"\)/.test(stepContent),
'step 5.75 must dispatch code-review with an explicit args="${PHASE_NUMBER}", matching code_review_gate\'s invocation — the bare generic Skill(skill="gsd-<ref.skill>") form has no phase argument and code-review.md requires one');
});
test('execute-phase.md does NOT contain ls.*REVIEW.md.*head pattern', { skip: !PLUGIN_AVAILABLE ? 'Plugin dir not installed' : false }, () => {
const content = fs.readFileSync(path.join(PLUGIN_WORKFLOWS_DIR, 'execute-phase.md'), 'utf-8');
// Extract code_review_gate section to check
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own workflow .md content, fixed-size author-controlled content
const gateMatch = content.match(/<step name="code_review_gate">([\s\S]*?)<\/step>/);
if (gateMatch) {
const gateContent = gateMatch[1];
assert.ok(!gateContent.match(/ls.*REVIEW\.md.*head/),
'execute-phase.md code_review_gate uses non-deterministic glob pattern (ls | head)');
}
});
test('quick.md contains code-review invocation', { skip: !PLUGIN_AVAILABLE ? 'Plugin dir not installed' : false }, () => {
const content = fs.readFileSync(path.join(PLUGIN_WORKFLOWS_DIR, 'quick.md'), 'utf-8');
assert.ok(content.includes('code-review') || content.includes('code_review'),
'quick.md missing code-review invocation');
});
test('quick.md resolves code-review capability hook', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'quick.md'), 'utf-8');
const start = content.indexOf('**Step 6.25: Code review (auto)**');
// #2994 (pre-existing since #2994's earlier quick-verification.md extraction,
// 18ff35d20): Step 6.5's content moved into
// gsd-core/workflows/quick/steps/quick-verification.md behind a
// `<!-- gsd:section id="quick-verification" -->` marker, so the literal
// "**Step 6.5: Verification" heading text no longer follows Step 6.25 in
// this file — the marker is the correct end-of-step delimiter now (mirrors
// phase6-review-capabilities.test.cjs's identical retarget for the same move).
const end = content.indexOf('<!-- gsd:section id="quick-verification"', start);
assert.ok(start !== -1 && end !== -1, 'quick.md missing Step 6.25 code review section');
const reviewContent = content.slice(start, end);
assert.ok(reviewContent.includes('loop render-hooks execute:post'),
'quick.md code review step must resolve execute:post capability hooks');
assert.ok(reviewContent.includes('ref.skill == "code-review"'),
'quick.md code review step must identify the code-review capability hook');
assert.ok(!reviewContent.match(/config-get\s+workflow\.code_review/),
'quick.md code review step must not read workflow.code_review directly');
});
// autonomous.md tests read from the repo's canonical workflow source (WORKFLOWS_DIR),
// not the user-installed plugin dir. The plugin dir can lag behind the repo until the
// user re-installs, so asserting against it produces false negatives. The repo file
// is the source of truth and is always present in CI checkouts.
test('autonomous.md contains gsd-code-review skill invocation', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'autonomous.md'), 'utf-8');
// Parse Skill(...) invocations into structured objects and assert canonical
// hyphen form is referenced. Canonical command form is hyphen
// (gsd-code-review); colon form (gsd:code-review) is the legacy
// frontmatter-name form removed in PR #2819.
const invocations = parseWorkflowSkillInvocations(content);
const skillNames = invocations.map(inv => inv.skill);
assert.ok(skillNames.includes('gsd-code-review'),
`autonomous.md must invoke Skill(skill="gsd-code-review", ...); found skills: ${JSON.stringify(skillNames)}`);
assert.ok(!skillNames.includes('gsd:code-review'),
'autonomous.md must not use legacy colon form gsd:code-review (canonical is hyphen form)');
});
test('autonomous.md auto-fix uses consolidated gsd-code-review --fix invocation (#2790)', () => {
// After #2790, gsd-code-review-fix was absorbed into gsd-code-review as
// the --fix flag. The autonomous workflow must invoke the consolidated
// form, not the deleted gsd-code-review-fix skill.
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'autonomous.md'), 'utf-8');
const invocations = parseWorkflowSkillInvocations(content);
const skillNames = invocations.map(inv => inv.skill);
assert.ok(!skillNames.includes('gsd-code-review-fix'),
`autonomous.md must not invoke deleted gsd-code-review-fix skill (consolidated into --fix); found: ${JSON.stringify(skillNames)}`);
assert.ok(!skillNames.includes('gsd:code-review-fix'),
'autonomous.md must not use legacy colon form gsd:code-review-fix');
// Find a gsd-code-review invocation that carries the --fix flag (the
// consolidated auto-fix entry point).
const fixInvocation = invocations.find(inv => {
if (inv.skill !== 'gsd-code-review') return false;
const tokens = new Set((inv.args ?? '').split(/\s+/).filter(Boolean));
return tokens.has('--fix');
});
assert.ok(fixInvocation,
`autonomous.md must invoke Skill(skill="gsd-code-review", args="... --fix ...") for auto-fix; found: ${JSON.stringify(invocations)}`);
});
test('autonomous.md contains --auto flag on consolidated --fix invocation (#2790)', () => {
const content = fs.readFileSync(path.join(WORKFLOWS_DIR, 'autonomous.md'), 'utf-8');
// Find the gsd-code-review invocation that carries --fix (the consolidated
// auto-fix entry point), then assert --auto is one of its arg tokens.
// Tokenize via whitespace-split to avoid substring matches that could
// conflate --auto with --auto-foo.
const invocations = parseWorkflowSkillInvocations(content);
const fixInvocation = invocations.find(inv => {
if (inv.skill !== 'gsd-code-review') return false;
const tokens = new Set((inv.args ?? '').split(/\s+/).filter(Boolean));
return tokens.has('--fix');
});
assert.ok(fixInvocation, 'autonomous.md missing Skill(skill="gsd-code-review", args="... --fix ...") invocation');
const argTokens = new Set((fixInvocation.args ?? '').split(/\s+/).filter(Boolean));
assert.ok(argTokens.has('--auto'),
`autonomous.md gsd-code-review-fix args missing --auto flag; got args="${fixInvocation.args}"`);
});
});
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-2839-review-fix-transactional-cleanup.test.cjs — consolidation epic #1969 (B8 #1977)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-2839-review-fix-transactional-cleanup (consolidation epic #1969 B8 #1977)", () => {
/**
* Regression test for bug #2839
*
* /gsd-code-review-fix cleanup tail is non-transactional. If the agent is
* interrupted (system restart, OOM kill) AFTER the last fix commit but
* BEFORE `git worktree remove`, the worktree is orphaned in
* `git worktree list`, the agent's branch is left with unmerged commits,
* and STATE.md is never advanced. To anyone reading main only, the phase
* looks "ready to plan" while critical fixes sit on a dangling branch.
*
* Fix: introduce a recovery sentinel JSON at
* ${PHASE_DIR}/.review-fix-recovery-pending.json
* The sentinel is written AFTER `git worktree add` succeeds and
* REMOVED only after `git worktree remove` completes, so the cleanup
* tail is transactional from the orchestrator's perspective. If the
* process dies in between, the sentinel is left behind pointing at the
* orphan worktree and branch — a future run, /gsd-resume-work, or
* /gsd-progress can detect and complete the recovery.
*/
'use strict';
// allow-test-rule: source-text-is-the-product (see #2839)
// The gsd-code-fixer agent's working instructions ARE the product — Claude
// follows them at runtime. Structural assertions over the markdown source
// test the deployed contract. See bug-2686 for the same pattern.
const { describe, test, before } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { parseFrontmatter } = require('./helpers.cjs');
const SENTINEL_NAME = '.review-fix-recovery-pending.json';
function extractStep(content, stepName) {
const re = new RegExp(`<step\\s+name="${stepName}">([\\s\\S]*?)</step>`);
const m = content.match(re);
return m ? m[1] : null;
}
describe('bug-2839: /gsd-code-review-fix cleanup is transactional', () => {
let agentPath;
let agentContent;
let frontmatter;
before(() => {
agentPath = path.join(__dirname, '..', 'agents', 'gsd-code-fixer.md');
assert.ok(fs.existsSync(agentPath), 'agents/gsd-code-fixer.md must exist');
agentContent = fs.readFileSync(agentPath, 'utf-8');
frontmatter = parseFrontmatter(agentContent);
assert.ok(frontmatter, 'agent must have YAML frontmatter');
});
test('agent declares a recovery sentinel filename', () => {
assert.ok(
agentContent.includes(SENTINEL_NAME),
`gsd-code-fixer.md must reference the recovery sentinel ${SENTINEL_NAME} so an interrupted cleanup tail is discoverable (#2839)`
);
});
test('sentinel is written inside setup_worktree, after git worktree add', () => {
const setupStep = extractStep(agentContent, 'setup_worktree');
assert.ok(setupStep, 'setup_worktree step must exist');
assert.ok(
setupStep.includes(SENTINEL_NAME),
`setup_worktree must reference ${SENTINEL_NAME} so the sentinel is created at the start of the run (#2839)`
);
const addPos = setupStep.indexOf('git worktree add');
assert.ok(addPos !== -1, 'setup_worktree must contain `git worktree add`');
// The sentinel WRITE (not just a reference) must come after `git worktree add`.
// Earlier references are allowed (e.g. recovery check for a stale sentinel
// from a prior interrupted run). Look for an explicit write — either a
// shell `>`/`>>` redirection, a `node -e` invocation that uses
// `fs.writeFileSync(...sentinel...)`, or a `Write` tool reference.
const writeIdx = (() => {
const candidates = [
/fs\.writeFileSync\([^)]*sentinel/,
/>\s*"?\$sentinel/,
/>\s*"?\$\{sentinel\}/,
/Write the recovery sentinel/i,
];
let earliest = -1;
for (const re of candidates) {
const m = re.exec(setupStep);
if (m && (earliest === -1 || m.index < earliest)) earliest = m.index;
}
return earliest;
})();
assert.ok(
writeIdx !== -1,
'setup_worktree must explicitly describe writing the sentinel (#2839)'
);
assert.ok(
addPos < writeIdx,
'sentinel must be written AFTER `git worktree add` succeeds (#2839)'
);
});
test('sentinel records worktree path, branch, and padded_phase as JSON fields', () => {
for (const key of ['worktree_path', 'branch', 'padded_phase']) {
assert.ok(
agentContent.includes(key),
`recovery sentinel must record \`${key}\` so a future /gsd-resume-work or /gsd-progress can locate the orphan state (#2839)`
);
}
});
test('sentinel removal happens only AFTER git worktree remove succeeds', () => {
const setupStep = extractStep(agentContent, 'setup_worktree');
assert.ok(setupStep, 'setup_worktree step must exist');
const cleanupAnchor = setupStep.lastIndexOf('Cleanup tail (transactional');
assert.ok(cleanupAnchor !== -1, 'setup_worktree must document cleanup-tail section');
const cleanupSection = setupStep.slice(cleanupAnchor);
const removeIdx = cleanupSection.indexOf('git worktree remove "$wt" --force');
assert.ok(removeIdx !== -1, 'cleanup-tail must remove worktree');
// Within the cleanup-tail section, accept either a literal-filename form
// (`rm -f .../.review-fix-recovery-pending.json`) or a shell-variable form
// referring to the previously-declared `sentinel` variable
// (`rm -f "$sentinel"` / `rm -f "${sentinel}"`).
const escapedName = escapeRegex(SENTINEL_NAME);
const sentinelRemovalRe = new RegExp(
`(rm\\s+(?:-f\\s+)?[^\\n]*(?:${escapedName}|\\$\\{?sentinel\\}?)|unlink[^\\n]*(?:${escapedName}|\\$\\{?sentinel\\}?))`
);
const sentinelRemovalMatch = sentinelRemovalRe.exec(cleanupSection);
assert.ok(
sentinelRemovalMatch,
`agent must remove the sentinel file (rm or unlink ${SENTINEL_NAME}) as part of the cleanup tail (#2839)`
);
const sentinelRemovalIdx = sentinelRemovalMatch.index;
assert.ok(
removeIdx < sentinelRemovalIdx,
'cleanup ordering must be: `git worktree remove` BEFORE sentinel removal (#2839)'
);
});
test('agent documents detection of pre-existing sentinel from a prior interrupted run', () => {
const lower = agentContent.toLowerCase();
const mentionsRecovery =
lower.includes('stale sentinel') ||
lower.includes('existing sentinel') ||
lower.includes('previous sentinel') ||
lower.includes('prior run') ||
lower.includes('pre-existing sentinel') ||
lower.includes('recovery');
assert.ok(
mentionsRecovery,
'agent must describe how it handles a pre-existing sentinel from a previous interrupted run (#2839)'
);
});
test('cleanup-tail obligation is documented as transactional / atomic', () => {
const lower = agentContent.toLowerCase();
const mentionsTransactional =
lower.includes('transactional') ||
lower.includes('atomic cleanup') ||
lower.includes('cleanup tail');
assert.ok(
mentionsTransactional,
'agent must document the cleanup tail as transactional/atomic (#2839)'
);
});
});
});
}