Files
msd-core/tests/review-parallel-lanes.test.cjs
Tom Boucher fb9823e1e1 fix(#3689): refuse a ledger write when the rendered table disagrees with its JSON (#3828)
* test(#3689): failing-first coverage for the ledger table/JSON agreement guard

`.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is
its source of truth, but nothing checks the two still agree before a write
overwrites the table. `windows append` / `waive` / `fixed` therefore discard a
drifted cell silently, and erase a table-only row entirely, both at exit 0.

Adds to tests/broken-windows.test.cjs:
- five refusal cases that fail today, covering all three write commands, a
  drifted cell, a table-only row, and drift on a non-first row; each asserts the
  typed reason via GSD_JSON_ERRORS and that the file is byte-identical after
  the refusal, so a guard that refuses only after writing cannot pass
- six anti-tightening pins that must stay green: an agreeing ledger, the
  first-write ENOENT path, #2893 trailing-prose preservation, #3657 3-backtick
  fence tolerance, escaped pipes and backslashes in a description, and the
  zero-entry placeholder table
- a fast-check property pinning the round trip the guard depends on —
  extractTableRegion(renderLedger(l)) === renderTable(l.entries) — because a
  false refusal on a clean ledger would be worse than the bug

Fixtures are built by running the real CLI and then perturbing only the table,
so frontmatter and JSON stay consistent and the pre-existing counts cross-check
still passes; a hand-written ledger would pass these for the wrong reason.

Refs #3689

* fix(#3689): refuse a ledger write when the rendered table disagrees with its JSON

`.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is
its source of truth, and `writeLedgerAtomic` regenerated that table on every
`windows append` / `waive` / `fixed` without ever checking the two still agreed.
A hand-edited cell was silently reverted; a row that existed only in the table
vanished entirely. Both at exit 0, with nothing on stdout to say so.

The write seam now compares the on-disk table against
`renderTable(<entries parsed from the on-disk JSON>)` before regenerating
anything, and refuses with a typed `windows_ledger_table_drift` error naming the
drifted row ids and the remedy. Because the check sits at the single write seam,
all three commands inherit it, and the file is left byte-identical on refusal.

Deliberately not enforced in `parseLedger`: hardening the read would break
`windows status` and the ship gate on exactly the ledgers an operator needs to
inspect to diagnose the drift.

Two hazards handled explicitly, both discovered in review of the first draft:

- The pre-image read now distinguishes ENOENT from every other errno, per the
  #1950-H2 fail-closed-on-unreadable invariant `readLedgerOrNull` already
  honors. A bare catch would have let an unreadable pre-image skip the guard
  and write anyway.
- Both the entries baseline and the table extraction pass the pre-image's own
  frontmatter `total_count` to `locateJsonBlock`. Without that hint the
  no-expectation fallback binds to the LATEST fenced JSON array in the file,
  which is the operator's prose block whenever that prose contains one — the
  exact case #2893 exists for — refusing every write on a ledger that never
  drifted. A regression test covers it.

Also extends the CONTEXT.md Broken Windows Ledger glossary entry: the table is
a third projection of the same source, cross-checked at the write seam, and the
frozen REASON enum gains WINDOWS_LEDGER_TABLE_DRIFT.

Fixes #3689

* fix(#3689): bind prose preservation to the pre-image's own ledger block

Found while reviewing the table drift guard: the #2893 trailing-prose
preservation in `writeLedgerAtomic` passed `ledger.total_count` — the
POST-mutation count — as the disambiguation hint for a lookup over the
PRE-image. On an append the pre-image holds N entries while the hint says N+1,
so the hint can never match and `locateJsonBlock` falls through to its
last-array-shaped-span fallback.

When the operator's trailing prose itself contains a fenced JSON array — the
ordinary case #2893 was written to protect — that prose block wins the
fallback. The preserved region is then computed from the prose fence rather
than the ledger fence, and everything between them, including the operator's
own text above the array, is silently dropped on the next write.

Reproduced against the real CLI: a prose block reading "Operator notes above
the array, IMPORTANT DO NOT LOSE THIS TEXT." plus a fenced 3-element array came
back empty after one `windows append`.

Both the prose lookup and the drift guard now share one pre-image-derived
`preImageExpectedTotal`, taken from the pre-image's own frontmatter, so they
bind to the same and correct block. The existing trailing-prose regression test
is strengthened to assert the prose survives byte-for-byte rather than merely
that the command exited 0 — asserting only the exit code is why this was
invisible.

Refs #3689

* fix(#3689): anchor table extraction on the header row, not a line-prefix scan

Independent review found the drift guard could brick a ledger nobody had
hand-edited. `validateDescription` accepts a description containing a raw
newline, and `renderTable`'s cell escaping covers backslash and pipe but not
newlines — so such a description renders a row that physically spans two file
lines, the second of which does not begin with `|`.

`extractTableRegion` bounded the table by walking backward over the contiguous
run of `|`-prefixed lines, so it stopped at that split. In the common case
where the row's tail is the last line before the fence it returned null, and
every subsequent append/waive/fixed was refused with "table region could not be
located" — permanently, with no CLI recovery path, on a ledger that never
drifted. A false refusal is worse than the bug this guard exists to fix.

The region is now anchored on the header row `renderTable` always emits,
running from its last line-start occurrence to the end of the pre-fence text.
The boundary is the fence rather than a line prefix, so a multi-line row is
captured whole, re-renders byte-identically, and compares equal. The header
literal is hoisted to one constant both `renderTable` branches and the
extractor share, so the two surfaces cannot drift apart.

Deliberately unchanged: `cell()` and `validateDescription`. The cosmetic
corruption a newline causes in the rendered table is pre-existing, and either
escaping it or rejecting the input would change what existing ledgers render to
or what input is accepted.

Also closes a coverage gap the standards review raised: the non-ENOENT
pre-image read branch — the one that stops an unreadable file from bypassing
the guard — now has a behavioral test that injects EACCES by monkeypatching
`fs.readFileSync` for that one path and restoring it in a `finally`, never by
`chmod 0o000` (root ignores mode bits, so that would pass with zero coverage).
The #3689 property generator no longer strips newlines out of descriptions,
which is why this was invisible to it.

Refs #3689

* chore(changeset): backfill PR number for #3689 fragment

* chore(changeset): backfill PR number for #3689 fragment

* fix(#3689): terminate the header scan when the match sits at index 0

`extractTableRegion`'s backward search for the table header could loop
forever. On a rejected match at index 0 it set `searchFrom = idx - 1`, i.e.
`-1`; `String.prototype.lastIndexOf` clamps its position argument into
`[0, length]`, so the next iteration searched from 0, found the same match,
rejected it identically, and set `-1` again. The loop made no progress.

Reachable only through the exported `extractTableRegion` — `writeLedgerAtomic`
reaches it after `parseFrontmatterStrict` has already succeeded, so the
candidate region begins with the `---` frontmatter fence and a match at index 0
is impossible. Latent rather than live, but an exported `for(;;)` that can fail
to advance is not something to ship.

Confirmed by running the pre-fix compiled function on
`TABLE_HEADER_LINE + 'X\n' + <a valid json fence>` as a backgrounded child: it
was still alive after five seconds having printed nothing, and had to be killed.
Post-fix the same input returns `null` promptly — correct, since the sole
header occurrence fails the end-of-line test and no valid header exists.

A regression here would stall the suite rather than fail it, so the new test
also asserts the returned value rather than relying on termination alone. No
wall-clock assertion is involved.

Refs #3689

* test(#3034): publish the lane trace before the done-file that releases dependents

`preservesSelectionOrderParallelDespiteCompletionOrder` forces a reverse
completion order with a dependency chain rather than sleeps: each stub lane
waits on `done-<dep>` before finishing. It then ended with

    touch "$RUN_DIR/done-$slug"
    echo "end:$slug" >> "$TRACE"

Those are two unsynchronized operations in separate shell processes. A
dependent's `wait_for_file` unblocks the instant the upstream's `touch` lands,
but the upstream's own `echo` has not necessarily run — so if the upstream is
descheduled between the two, the dependent can run its whole body and append
its `end:` line first. The done-file was published before the state it signals.

Observed on the remote runner as `[end:claude, end:codex, end:gemini]` where
selection order demands `[end:claude, end:gemini, end:codex]`. The failure was
in the fixture's own self-check, before it reached the assertion #3034 exists to
make.

Not a flake and not a wall-clock margin: this branch passed the full suite twice
at 14f494644 and 90c5d7a03, and the only delta in the failing run was one added
test in tests/broken-windows.test.cjs — an unrelated module. Adding load
elsewhere in the suite was enough to invert it, which is what a real race does.

Swapping the pair establishes a genuine happens-before: anything a dependent can
observe is written before the file that releases it. A comment records why, so
the order is not tidied back.

The production path is unaffected and was independently confirmed correct —
`invoke_reviewers` joins every lane with `wait`, then aggregates by iterating
DISPATCH_SLUGS in selection order, reading per-slug result files. It consumes no
completion-order signal at all.

Refs #3034

---------

Co-authored-by: sim <sim@local>
2026-08-24 17:43:25 -04:00

577 lines
23 KiB
JavaScript

'use strict';
/**
* Failing-first tests for #3034 (opt-in parallel reviewer lanes).
*
* Design: .gsd/phase/feat-3034-parallel-reviewer-lanes/40-design.md
* Test matrix: .gsd/phase/feat-3034-parallel-reviewer-lanes/50-test-matrix.md
*
* The unit under test is the real, shipped `<step name="invoke_reviewers">`
* fenced bash block in gsd-core/workflows/review.md — extracted and EXECUTED
* (never re-typed), with the single I/O seam `gsd_run` replaced by a shell
* function stub. The feature this file exercises (an opt-in
* `review.parallel_lanes` config key that backgrounds lane dispatch and
* joins before aggregation) does not exist yet, so several tests below are
* expected to be RED against the current shipped block.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const {
createTempDir,
createTempProject,
cleanup,
readFileNormalized,
runGsdTools,
} = require('./helpers.cjs');
const { runHook } = require('./helpers/process-seam.cjs');
const { HOOK_FANOUT_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
const REPO_ROOT = path.join(__dirname, '..');
const REVIEW_MD_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'review.md');
const SELECTED_3 = 'codex,gemini,claude';
// ─── extraction (source-text-is-the-product) ──────────────────────────────
/**
* Reads review.md and extracts the fenced bash block inside
* `<step name="invoke_reviewers">`. Mirrors extractCwdGuardBash() in
* tests/worktree-cleanup.test.cjs: readFileNormalized() strips CRLF at the
* read boundary (what actually makes the captured body safe to hand to
* bash); the `\r?\n` in the fence regex below is redundant on that
* already-normalized input but kept anyway because a bare `\n` in a
* markdown-fence-shaped regex trips the local/no-crlf-fragile-split rule.
*/
function extractInvokeReviewersBash() {
const content = readFileNormalized(REVIEW_MD_PATH);
const stepMarker = '<step name="invoke_reviewers">';
const stepIdx = content.indexOf(stepMarker);
if (stepIdx === -1) {
throw new Error(`extractInvokeReviewersBash: could not find "${stepMarker}" in ${REVIEW_MD_PATH}`);
}
const afterStep = content.slice(stepIdx + stepMarker.length);
const endMarker = '</step>';
const endIdx = afterStep.indexOf(endMarker);
if (endIdx === -1) {
throw new Error(`extractInvokeReviewersBash: could not find closing "${endMarker}" after invoke_reviewers in ${REVIEW_MD_PATH}`);
}
const stepBody = afterStep.slice(0, endIdx);
const fenceRe = /```(?:bash|sh)\r?\n([\s\S]*?)```/;
const fenceMatch = fenceRe.exec(stepBody);
if (!fenceMatch) {
throw new Error(`extractInvokeReviewersBash: no \`\`\`bash fence found inside invoke_reviewers step in ${REVIEW_MD_PATH}`);
}
const block = fenceMatch[1];
if (!block.trim()) {
throw new Error('extractInvokeReviewersBash: extracted bash block is empty');
}
if (!block.includes('review-lane invoke')) {
throw new Error('extractInvokeReviewersBash: extracted block does not contain "review-lane invoke" — anchor may have drifted');
}
if (!block.includes('SELECTED_REVIEWERS')) {
throw new Error('extractInvokeReviewersBash: extracted block does not contain "SELECTED_REVIEWERS" — anchor may have drifted');
}
return block;
}
// ─── stub preamble ─────────────────────────────────────────────────────────
//
// The ONLY I/O seam the extracted block calls is `gsd_run`. Everything else
// executed by runDispatch() is the real shipped shell text. `barrier` is the
// positive, deterministic concurrency proof this suite uses in place of any
// elapsed-time assertion (CLAUDE.md bans wall-clock assertions).
const STUB_PREAMBLE = [
'arg_after() {',
' local flag="$1"; shift',
' while [ $# -gt 0 ]; do',
' if [ "$1" = "$flag" ]; then',
' printf %s "$2"',
' return 0',
' fi',
' shift',
' done',
'}',
'',
'in_list() {',
' local needle="$1" list="$2" item old_ifs="$IFS"',
' IFS=","',
' for item in $list; do',
' if [ "$item" = "$needle" ]; then',
' IFS="$old_ifs"',
' return 0',
' fi',
' done',
' IFS="$old_ifs"',
' return 1',
'}',
'',
'dep_of() {',
' local slug="$1" pair old_ifs="$IFS"',
' IFS=","',
' for pair in $STUB_DEPS; do',
' case "$pair" in',
' "$slug="*)',
' printf %s "${pair#*=}"',
' IFS="$old_ifs"',
' return 0',
' ;;',
' esac',
' done',
' IFS="$old_ifs"',
'}',
'',
'wait_for_file() {',
' local target="$1" i=0',
' while [ ! -f "$target" ] && [ "$i" -lt 200 ]; do',
' sleep 0.05',
' i=$((i + 1))',
' done',
'}',
'',
'barrier() {',
' local slug="$1" i=0 count',
' touch "$BARRIER_DIR/$slug"',
' count=$(ls -1 "$BARRIER_DIR" | wc -l)',
' while [ "$count" -lt "$LANE_COUNT" ] && [ "$i" -lt 200 ]; do',
' sleep 0.05',
' i=$((i + 1))',
' count=$(ls -1 "$BARRIER_DIR" | wc -l)',
' done',
' if [ "$count" -lt "$LANE_COUNT" ]; then',
' echo "barrier-timeout:$slug" >> "$TRACE"',
' return 1',
' fi',
' return 0',
'}',
'',
'gsd_run() {',
' if [ "$1" = "query" ] && [ "$2" = "config-get" ] && [ "$3" = "review.parallel_lanes" ]; then',
' if [ "$STUB_CONFIG_GET_FAILS" = "1" ]; then',
' return 1',
' fi',
' printf %s "$STUB_PARALLEL"',
' return 0',
' fi',
'',
' if [ "$1" = "query" ] && [ "$2" = "review-lane" ] && [ "$3" = "plan" ]; then',
' shift 3',
' local sel',
' sel="$(arg_after --selected "$@")"',
' if in_list "$sel" "$STUB_BUDGET_FAIL"; then',
' printf %s \'{"promptBudget": 10}\'',
' else',
' printf %s \'{"promptBudget": -1}\'',
' fi',
' return 0',
' fi',
'',
' if [ "$1" = "query" ] && [ "$2" = "prompt-budget" ]; then',
' shift 2',
' local out base slug',
' out="$(arg_after --output-prompt "$@")"',
' base="$(basename "$out")"',
' slug="${base#gsd-review-prompt-}"',
' slug="${slug%.md}"',
' if in_list "$slug" "$STUB_BUDGET_FAIL"; then',
' return 2',
' fi',
' : > "$out"',
' return 0',
' fi',
'',
' if [ "$1" = "query" ] && [ "$2" = "review-lane" ] && [ "$3" = "invoke" ]; then',
' shift 3',
' local slug dep pad',
' slug="$(arg_after --slug "$@")"',
' echo "start:$slug" >> "$TRACE"',
' if [ "$STUB_BARRIER" = "1" ]; then',
' barrier "$slug" || true',
' fi',
' dep="$(dep_of "$slug")"',
' if [ -n "$dep" ]; then',
' wait_for_file "$RUN_DIR/done-$dep"',
' fi',
' echo "stub review body for $slug" > "$RUN_DIR/gsd-review-$slug.md"',
' pad=""',
' if [ "$STUB_PAD_BYTES" -gt 0 ] 2>/dev/null; then',
' pad="$(head -c "$STUB_PAD_BYTES" /dev/zero | tr "\\0" "x")"',
' fi',
' if ! in_list "$slug" "$STUB_SILENT"; then',
' printf \'{"slug":"%s","pad":"%s"}\\n\' "$slug" "$pad"',
' fi',
// #3689: the done-file is a cross-process happens-before edge — a
// dependent lane unblocks the instant this file appears (wait_for_file
// above just polls for its existence), so everything a dependent may
// observe (the "end:$slug" trace line) must be written BEFORE the file
// that releases it. touch-then-echo let a descheduled upstream lose the
// race to its own dependent, inverting the #3034 completion-order trace.
' echo "end:$slug" >> "$TRACE"',
' touch "$RUN_DIR/done-$slug"',
' if in_list "$slug" "$STUB_FAIL"; then',
' return 1',
' fi',
' return 0',
' fi',
'',
' return 0',
'}',
].join('\n');
// ─── dispatch runner ───────────────────────────────────────────────────────
/**
* Build the stub env for a runDispatch() call from `opts`. Every key is
* always present (never omitted) so the generated `set -u` script never
* dereferences an unset variable — `opts.parallel === null` deliberately
* maps to the empty string, which is exactly how an unset config key reads
* back through `config-get --raw`.
*/
function buildEnv(opts, runDir, tracePath, barrierDir) {
const selected = opts.selected;
const laneCount = selected.split(',').filter((s) => s.length > 0).length;
const deps = opts.deps || {};
return {
...process.env,
SELECTED_REVIEWERS: selected,
EXPLICIT_FLAG: '',
STUB_PARALLEL: opts.parallel === null || opts.parallel === undefined ? '' : opts.parallel,
STUB_CONFIG_GET_FAILS: opts.configGetFails ? '1' : '0',
STUB_FAIL: (opts.failSlugs || []).join(','),
STUB_BUDGET_FAIL: (opts.budgetFailSlugs || []).join(','),
STUB_SILENT: (opts.silentSlugs || []).join(','),
STUB_BARRIER: opts.barrier ? '1' : '0',
STUB_DEPS: Object.entries(deps).map(([k, v]) => `${k}=${v}`).join(','),
STUB_PAD_BYTES: String(opts.padBytes || 0),
LANE_COUNT: String(laneCount),
TRACE: tracePath,
BARRIER_DIR: barrierDir,
};
}
/**
* Runs the real extracted invoke_reviewers bash block with the gsd_run stub
* spliced in front of it. `{run_dir}` is replaced globally with a real temp
* directory. Returns the dispatch outcome plus the artifacts it produced.
*/
function runDispatch(t, opts) {
const scriptDir = createTempDir('gsd-3034-script-');
const runDir = createTempDir('gsd-3034-rundir-');
const barrierDir = createTempDir('gsd-3034-barrier-');
t.after(() => {
cleanup(scriptDir);
cleanup(runDir);
cleanup(barrierDir);
});
const block = extractInvokeReviewersBash();
const tracePath = path.join(scriptDir, 'trace.log');
const env = buildEnv(opts, runDir, tracePath, barrierDir);
const scriptPath = path.join(scriptDir, 'dispatch.sh');
const script = [
'#!/usr/bin/env bash',
'set -u',
STUB_PREAMBLE,
block.split('{run_dir}').join(runDir),
].join('\n');
fs.writeFileSync(scriptPath, script, { mode: 0o755 });
const result = runHook(scriptPath, [], {
interpreter: 'bash',
cwd: runDir,
env,
timeoutMs: HOOK_FANOUT_TIMEOUT_MS,
});
const jsonlPath = path.join(runDir, 'gsd-review-lane-results.jsonl');
const jsonl = fs.existsSync(jsonlPath) ? fs.readFileSync(jsonlPath, 'utf-8') : '';
const lines = jsonl.split('\n').filter((l) => l.trim() !== '');
const trace = fs.existsSync(tracePath)
? readFileNormalized(tracePath).split('\n').filter((l) => l.trim() !== '')
: [];
return {
outcome: result.outcome,
exitCode: result.exitCode,
stderr: result.stderr,
jsonl,
lines,
trace,
runDir,
};
}
/** Parsed slug order from JSONL lines — never assert on raw JSONL text. */
function slugOrder(lines) {
return lines.map((l) => JSON.parse(l).slug);
}
function serialTrace(slugs) {
return slugs.flatMap((s) => [`start:${s}`, `end:${s}`]);
}
// ─── #1/#2 — default and explicit-disabled serial dispatch ───────────────
describe('#3034 default and explicit-disabled dispatch stays serial', () => {
test('defaultsToSerialDispatchWhenKeyUnset', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: null });
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
});
test('staysSerialWhenExplicitlyDisabled', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'false' });
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
});
});
// ─── #3/#4 — opt-in concurrency and join-before-aggregate ─────────────────
describe('#3034 opt-in concurrency', () => {
test('dispatchesLanesConcurrentlyWhenEnabled', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', barrier: true });
const timeouts = result.trace.filter((l) => l.startsWith('barrier-timeout:'));
assert.deepEqual(timeouts, [], 'no lane should hit the barrier timeout when lanes run concurrently');
});
test('joinsAllLanesBeforeAggregation', (t) => {
// barrier:true is load-bearing, not decoration. Without it the stub lanes
// finish instantly and a missing `wait` could still race to 3 lines,
// making this pass intermittently. Held at the barrier, a missing join
// deterministically aggregates ZERO lines.
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', barrier: true });
assert.equal(result.lines.length, 3, 'dispatch must return only after every lane wrote its result');
});
});
// ─── #5/#6 — selection-order preservation ─────────────────────────────────
describe('#3034 JSONL preserves selection order, not completion order', () => {
test('preservesSelectionOrderSerial', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: null });
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
});
test('preservesSelectionOrderParallelDespiteCompletionOrder', (t) => {
// codex waits on gemini; gemini waits on claude -> forces reverse
// completion order (claude, gemini, codex) while selection order stays
// codex, gemini, claude.
const result = runDispatch(t, {
selected: SELECTED_3,
parallel: 'true',
deps: { codex: 'gemini', gemini: 'claude' },
});
const endMarkers = result.trace.filter((l) => l.startsWith('end:'));
// Pin the fixture actually forced reverse completion first — otherwise
// the selection-order assertion below would pass vacuously.
assert.deepEqual(endMarkers, ['end:claude', 'end:gemini', 'end:codex']);
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
});
});
// ─── #7/#8 — PIPE_BUF boundary triple (plus one clearly-oversized case) ───
describe('#3034 oversized lane results stay intact under concurrency', () => {
test('keepsOversizedLaneResultsIntactUnderConcurrency', async (t) => {
for (const padBytes of [4095, 4096, 4097, 8192]) {
await t.test(`padBytes=${padBytes}`, (t2) => {
const result = runDispatch(t2, { selected: SELECTED_3, parallel: 'true', padBytes });
assert.equal(result.lines.length, 3);
for (const line of result.lines) {
assert.doesNotThrow(() => JSON.parse(line), `line failed to parse at padBytes=${padBytes}: ${line.slice(0, 80)}...`);
}
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
});
}
});
});
// ─── #9 — single-lane boundary ─────────────────────────────────────────────
describe('#3034 single-lane boundary', () => {
test('singleLaneParallelMatchesSerial', (t) => {
const serial = runDispatch(t, { selected: 'codex', parallel: null });
const parallel = runDispatch(t, { selected: 'codex', parallel: 'true' });
assert.equal(serial.lines.length, 1);
assert.equal(parallel.lines.length, 1);
assert.deepEqual(slugOrder(serial.lines), slugOrder(parallel.lines));
const serialMd = fs.readFileSync(path.join(serial.runDir, 'gsd-review-codex.md'), 'utf-8');
const parallelMd = fs.readFileSync(path.join(parallel.runDir, 'gsd-review-codex.md'), 'utf-8');
assert.equal(serialMd, parallelMd);
});
});
// ─── #10 — empty selection is a no-op ──────────────────────────────────────
describe('#3034 empty selection', () => {
test('emptySelectionIsANoOp', (t) => {
const result = runDispatch(t, { selected: '', parallel: 'true' });
assert.equal(result.outcome, 'exited');
assert.equal(result.exitCode, 0);
assert.deepEqual(result.lines, []);
assert.deepEqual(result.trace, []);
});
});
describe('#3034 duplicate slug in the selection', () => {
test('duplicateSlugDispatchesOnceAndWritesOneLine', (t) => {
// A slug repeated in SELECTED_REVIEWERS would otherwise put two
// concurrent background jobs on the same `>`-truncated per-slug result
// file, corrupting whichever one finishes last. DISPATCH_SLUGS
// de-duplicates before dispatch, so codex must run exactly once.
const result = runDispatch(t, { selected: 'codex,gemini,codex', parallel: 'true' });
assert.equal(result.outcome, 'exited');
assert.equal(result.trace.filter((l) => l === 'start:codex').length, 1);
assert.equal(result.trace.filter((l) => l === 'end:codex').length, 1);
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini']);
});
});
// ─── #11/#12 — lane failure does not abort siblings ────────────────────────
describe('#3034 lane failure does not abort sibling lanes', () => {
test('laneFailureDoesNotAbortSiblingLanes', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', failSlugs: ['gemini'] });
assert.equal(result.lines.length, 3, 'a failing lane still contributes its stub result line');
assert.ok(
fs.existsSync(path.join(result.runDir, 'gsd-review-gemini.md')),
'the failing lane\'s diagnostic stub .md must be preserved',
);
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
});
test('laneFailureSerialUnchanged', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: null, failSlugs: ['gemini'] });
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
assert.equal(result.lines.length, 3);
});
});
// ─── #13 — budget-too-small skip emits no result line ─────────────────────
describe('#3034 budget-too-small skip', () => {
test('budgetSkipEmitsNoResultLine', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', budgetFailSlugs: ['gemini'] });
assert.deepEqual(slugOrder(result.lines), ['codex', 'claude'], 'the budget-skipped lane contributes no result line');
const stub = fs.readFileSync(path.join(result.runDir, 'gsd-review-gemini.md'), 'utf-8');
assert.match(stub, /skipped/);
assert.match(stub, /prompt budget/);
});
});
// ─── #14/#15 — silent / absent lane contributes no line ───────────────────
describe('#3034 silent lane contributes nothing, not a blank line', () => {
test('silentLaneContributesNoLine', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', silentSlugs: ['gemini'] });
assert.deepEqual(slugOrder(result.lines), ['codex', 'claude']);
// Folded row #15 (absent result file): budgetSkipEmitsNoResultLine above
// already exercises the "invoke never started, no file at all" shape via
// its `continue`. This assertion pins the sibling shape: a lane that DID
// run and wrote nothing must not leave a stray blank JSONL line.
assert.ok(!result.jsonl.includes('\n\n'), 'an empty lane result must not leave a blank JSONL line');
});
});
// ─── #16/#17 — non-canonical truthy values stay serial ────────────────────
describe('#3034 non-canonical truthy config values stay serial', () => {
test('nonCanonicalTruthyValuesStaySerial', async (t) => {
const nearMisses = ['TRUE', 'True', '1', 'yes', 'on', ' true', 'true '];
for (const value of nearMisses) {
await t.test(`parallel_lanes="${value}"`, (t2) => {
const result = runDispatch(t2, { selected: SELECTED_3, parallel: value });
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']), `value "${value}" must not opt into parallel dispatch`);
});
}
});
});
// ─── #18 — config-get failure fails safe to serial ─────────────────────────
describe('#3034 broken config tooling fails safe to serial', () => {
test('configGetFailureFallsBackToSerial', (t) => {
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', configGetFails: true });
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
});
});
// ─── #19/#20/#21 — config-set registers review.parallel_lanes ─────────────
describe('#3034 review.parallel_lanes config key', () => {
test('configSetAcceptsAndPersistsParallelLanes', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const setResult = runGsdTools('config-set review.parallel_lanes true', tmpDir);
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
const configPath = path.join(tmpDir, '.planning', 'config.json');
const config = JSON.parse(fs.readFileSync(configPath, 'utf-8'));
assert.equal(config.review?.parallel_lanes, true);
assert.equal(typeof config.review?.parallel_lanes, 'boolean');
const getResult = runGsdTools('config-get review.parallel_lanes --raw', tmpDir);
assert.ok(getResult.success, `config-get failed: ${getResult.error}`);
assert.equal((getResult.output || '').trim(), 'true');
});
test('configSetPersistsBooleanFalse', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const setResult = runGsdTools('config-set review.parallel_lanes false', tmpDir);
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
const configPath = path.join(tmpDir, '.planning', 'config.json');
const config = JSON.parse(fs.readFileSync(configPath, 'utf-8'));
assert.equal(config.review?.parallel_lanes, false);
assert.equal(typeof config.review?.parallel_lanes, 'boolean');
});
test('rejectsUnregisteredNeighbouringKey', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
// Missing trailing "s" — proves the whitelist is load-bearing and the
// two tests above are not vacuous (they'd pass even for an unregistered
// key if config-set accepted anything).
const result = runGsdTools('config-set review.parallel_lane true', tmpDir);
assert.equal(result.success, false, 'an unregistered near-miss key must be rejected');
});
});
// ─── #22 — serial/parallel artifact equivalence ────────────────────────────
describe('#3034 serial and parallel dispatch produce equivalent artifacts', () => {
test('serialAndParallelProduceEquivalentArtifacts', (t) => {
const serial = runDispatch(t, { selected: SELECTED_3, parallel: null });
const parallel = runDispatch(t, { selected: SELECTED_3, parallel: 'true' });
assert.deepEqual(slugOrder(serial.lines), slugOrder(parallel.lines));
assert.deepEqual(
serial.lines.map((l) => JSON.parse(l)),
parallel.lines.map((l) => JSON.parse(l)),
);
for (const slug of ['codex', 'gemini', 'claude']) {
const serialMd = fs.readFileSync(path.join(serial.runDir, `gsd-review-${slug}.md`), 'utf-8');
const parallelMd = fs.readFileSync(path.join(parallel.runDir, `gsd-review-${slug}.md`), 'utf-8');
assert.equal(serialMd, parallelMd, `gsd-review-${slug}.md must be byte-identical between the two paths`);
}
});
});