* test(#3689): failing-first coverage for the ledger table/JSON agreement guard `.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is its source of truth, but nothing checks the two still agree before a write overwrites the table. `windows append` / `waive` / `fixed` therefore discard a drifted cell silently, and erase a table-only row entirely, both at exit 0. Adds to tests/broken-windows.test.cjs: - five refusal cases that fail today, covering all three write commands, a drifted cell, a table-only row, and drift on a non-first row; each asserts the typed reason via GSD_JSON_ERRORS and that the file is byte-identical after the refusal, so a guard that refuses only after writing cannot pass - six anti-tightening pins that must stay green: an agreeing ledger, the first-write ENOENT path, #2893 trailing-prose preservation, #3657 3-backtick fence tolerance, escaped pipes and backslashes in a description, and the zero-entry placeholder table - a fast-check property pinning the round trip the guard depends on — extractTableRegion(renderLedger(l)) === renderTable(l.entries) — because a false refusal on a clean ledger would be worse than the bug Fixtures are built by running the real CLI and then perturbing only the table, so frontmatter and JSON stay consistent and the pre-existing counts cross-check still passes; a hand-written ledger would pass these for the wrong reason. Refs #3689 * fix(#3689): refuse a ledger write when the rendered table disagrees with its JSON `.planning/WINDOWS.md` renders its markdown table from the fenced JSON that is its source of truth, and `writeLedgerAtomic` regenerated that table on every `windows append` / `waive` / `fixed` without ever checking the two still agreed. A hand-edited cell was silently reverted; a row that existed only in the table vanished entirely. Both at exit 0, with nothing on stdout to say so. The write seam now compares the on-disk table against `renderTable(<entries parsed from the on-disk JSON>)` before regenerating anything, and refuses with a typed `windows_ledger_table_drift` error naming the drifted row ids and the remedy. Because the check sits at the single write seam, all three commands inherit it, and the file is left byte-identical on refusal. Deliberately not enforced in `parseLedger`: hardening the read would break `windows status` and the ship gate on exactly the ledgers an operator needs to inspect to diagnose the drift. Two hazards handled explicitly, both discovered in review of the first draft: - The pre-image read now distinguishes ENOENT from every other errno, per the #1950-H2 fail-closed-on-unreadable invariant `readLedgerOrNull` already honors. A bare catch would have let an unreadable pre-image skip the guard and write anyway. - Both the entries baseline and the table extraction pass the pre-image's own frontmatter `total_count` to `locateJsonBlock`. Without that hint the no-expectation fallback binds to the LATEST fenced JSON array in the file, which is the operator's prose block whenever that prose contains one — the exact case #2893 exists for — refusing every write on a ledger that never drifted. A regression test covers it. Also extends the CONTEXT.md Broken Windows Ledger glossary entry: the table is a third projection of the same source, cross-checked at the write seam, and the frozen REASON enum gains WINDOWS_LEDGER_TABLE_DRIFT. Fixes #3689 * fix(#3689): bind prose preservation to the pre-image's own ledger block Found while reviewing the table drift guard: the #2893 trailing-prose preservation in `writeLedgerAtomic` passed `ledger.total_count` — the POST-mutation count — as the disambiguation hint for a lookup over the PRE-image. On an append the pre-image holds N entries while the hint says N+1, so the hint can never match and `locateJsonBlock` falls through to its last-array-shaped-span fallback. When the operator's trailing prose itself contains a fenced JSON array — the ordinary case #2893 was written to protect — that prose block wins the fallback. The preserved region is then computed from the prose fence rather than the ledger fence, and everything between them, including the operator's own text above the array, is silently dropped on the next write. Reproduced against the real CLI: a prose block reading "Operator notes above the array, IMPORTANT DO NOT LOSE THIS TEXT." plus a fenced 3-element array came back empty after one `windows append`. Both the prose lookup and the drift guard now share one pre-image-derived `preImageExpectedTotal`, taken from the pre-image's own frontmatter, so they bind to the same and correct block. The existing trailing-prose regression test is strengthened to assert the prose survives byte-for-byte rather than merely that the command exited 0 — asserting only the exit code is why this was invisible. Refs #3689 * fix(#3689): anchor table extraction on the header row, not a line-prefix scan Independent review found the drift guard could brick a ledger nobody had hand-edited. `validateDescription` accepts a description containing a raw newline, and `renderTable`'s cell escaping covers backslash and pipe but not newlines — so such a description renders a row that physically spans two file lines, the second of which does not begin with `|`. `extractTableRegion` bounded the table by walking backward over the contiguous run of `|`-prefixed lines, so it stopped at that split. In the common case where the row's tail is the last line before the fence it returned null, and every subsequent append/waive/fixed was refused with "table region could not be located" — permanently, with no CLI recovery path, on a ledger that never drifted. A false refusal is worse than the bug this guard exists to fix. The region is now anchored on the header row `renderTable` always emits, running from its last line-start occurrence to the end of the pre-fence text. The boundary is the fence rather than a line prefix, so a multi-line row is captured whole, re-renders byte-identically, and compares equal. The header literal is hoisted to one constant both `renderTable` branches and the extractor share, so the two surfaces cannot drift apart. Deliberately unchanged: `cell()` and `validateDescription`. The cosmetic corruption a newline causes in the rendered table is pre-existing, and either escaping it or rejecting the input would change what existing ledgers render to or what input is accepted. Also closes a coverage gap the standards review raised: the non-ENOENT pre-image read branch — the one that stops an unreadable file from bypassing the guard — now has a behavioral test that injects EACCES by monkeypatching `fs.readFileSync` for that one path and restoring it in a `finally`, never by `chmod 0o000` (root ignores mode bits, so that would pass with zero coverage). The #3689 property generator no longer strips newlines out of descriptions, which is why this was invisible to it. Refs #3689 * chore(changeset): backfill PR number for #3689 fragment * chore(changeset): backfill PR number for #3689 fragment * fix(#3689): terminate the header scan when the match sits at index 0 `extractTableRegion`'s backward search for the table header could loop forever. On a rejected match at index 0 it set `searchFrom = idx - 1`, i.e. `-1`; `String.prototype.lastIndexOf` clamps its position argument into `[0, length]`, so the next iteration searched from 0, found the same match, rejected it identically, and set `-1` again. The loop made no progress. Reachable only through the exported `extractTableRegion` — `writeLedgerAtomic` reaches it after `parseFrontmatterStrict` has already succeeded, so the candidate region begins with the `---` frontmatter fence and a match at index 0 is impossible. Latent rather than live, but an exported `for(;;)` that can fail to advance is not something to ship. Confirmed by running the pre-fix compiled function on `TABLE_HEADER_LINE + 'X\n' + <a valid json fence>` as a backgrounded child: it was still alive after five seconds having printed nothing, and had to be killed. Post-fix the same input returns `null` promptly — correct, since the sole header occurrence fails the end-of-line test and no valid header exists. A regression here would stall the suite rather than fail it, so the new test also asserts the returned value rather than relying on termination alone. No wall-clock assertion is involved. Refs #3689 * test(#3034): publish the lane trace before the done-file that releases dependents `preservesSelectionOrderParallelDespiteCompletionOrder` forces a reverse completion order with a dependency chain rather than sleeps: each stub lane waits on `done-<dep>` before finishing. It then ended with touch "$RUN_DIR/done-$slug" echo "end:$slug" >> "$TRACE" Those are two unsynchronized operations in separate shell processes. A dependent's `wait_for_file` unblocks the instant the upstream's `touch` lands, but the upstream's own `echo` has not necessarily run — so if the upstream is descheduled between the two, the dependent can run its whole body and append its `end:` line first. The done-file was published before the state it signals. Observed on the remote runner as `[end:claude, end:codex, end:gemini]` where selection order demands `[end:claude, end:gemini, end:codex]`. The failure was in the fixture's own self-check, before it reached the assertion #3034 exists to make. Not a flake and not a wall-clock margin: this branch passed the full suite twice at 14f494644 and 90c5d7a03, and the only delta in the failing run was one added test in tests/broken-windows.test.cjs — an unrelated module. Adding load elsewhere in the suite was enough to invert it, which is what a real race does. Swapping the pair establishes a genuine happens-before: anything a dependent can observe is written before the file that releases it. A comment records why, so the order is not tidied back. The production path is unaffected and was independently confirmed correct — `invoke_reviewers` joins every lane with `wait`, then aggregates by iterating DISPATCH_SLUGS in selection order, reading per-slug result files. It consumes no completion-order signal at all. Refs #3034 --------- Co-authored-by: sim <sim@local>
577 lines
23 KiB
JavaScript
577 lines
23 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Failing-first tests for #3034 (opt-in parallel reviewer lanes).
|
|
*
|
|
* Design: .gsd/phase/feat-3034-parallel-reviewer-lanes/40-design.md
|
|
* Test matrix: .gsd/phase/feat-3034-parallel-reviewer-lanes/50-test-matrix.md
|
|
*
|
|
* The unit under test is the real, shipped `<step name="invoke_reviewers">`
|
|
* fenced bash block in gsd-core/workflows/review.md — extracted and EXECUTED
|
|
* (never re-typed), with the single I/O seam `gsd_run` replaced by a shell
|
|
* function stub. The feature this file exercises (an opt-in
|
|
* `review.parallel_lanes` config key that backgrounds lane dispatch and
|
|
* joins before aggregation) does not exist yet, so several tests below are
|
|
* expected to be RED against the current shipped block.
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const {
|
|
createTempDir,
|
|
createTempProject,
|
|
cleanup,
|
|
readFileNormalized,
|
|
runGsdTools,
|
|
} = require('./helpers.cjs');
|
|
const { runHook } = require('./helpers/process-seam.cjs');
|
|
const { HOOK_FANOUT_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
|
|
|
|
const REPO_ROOT = path.join(__dirname, '..');
|
|
const REVIEW_MD_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'review.md');
|
|
|
|
const SELECTED_3 = 'codex,gemini,claude';
|
|
|
|
// ─── extraction (source-text-is-the-product) ──────────────────────────────
|
|
|
|
/**
|
|
* Reads review.md and extracts the fenced bash block inside
|
|
* `<step name="invoke_reviewers">`. Mirrors extractCwdGuardBash() in
|
|
* tests/worktree-cleanup.test.cjs: readFileNormalized() strips CRLF at the
|
|
* read boundary (what actually makes the captured body safe to hand to
|
|
* bash); the `\r?\n` in the fence regex below is redundant on that
|
|
* already-normalized input but kept anyway because a bare `\n` in a
|
|
* markdown-fence-shaped regex trips the local/no-crlf-fragile-split rule.
|
|
*/
|
|
function extractInvokeReviewersBash() {
|
|
const content = readFileNormalized(REVIEW_MD_PATH);
|
|
|
|
const stepMarker = '<step name="invoke_reviewers">';
|
|
const stepIdx = content.indexOf(stepMarker);
|
|
if (stepIdx === -1) {
|
|
throw new Error(`extractInvokeReviewersBash: could not find "${stepMarker}" in ${REVIEW_MD_PATH}`);
|
|
}
|
|
const afterStep = content.slice(stepIdx + stepMarker.length);
|
|
|
|
const endMarker = '</step>';
|
|
const endIdx = afterStep.indexOf(endMarker);
|
|
if (endIdx === -1) {
|
|
throw new Error(`extractInvokeReviewersBash: could not find closing "${endMarker}" after invoke_reviewers in ${REVIEW_MD_PATH}`);
|
|
}
|
|
const stepBody = afterStep.slice(0, endIdx);
|
|
|
|
const fenceRe = /```(?:bash|sh)\r?\n([\s\S]*?)```/;
|
|
const fenceMatch = fenceRe.exec(stepBody);
|
|
if (!fenceMatch) {
|
|
throw new Error(`extractInvokeReviewersBash: no \`\`\`bash fence found inside invoke_reviewers step in ${REVIEW_MD_PATH}`);
|
|
}
|
|
const block = fenceMatch[1];
|
|
|
|
if (!block.trim()) {
|
|
throw new Error('extractInvokeReviewersBash: extracted bash block is empty');
|
|
}
|
|
if (!block.includes('review-lane invoke')) {
|
|
throw new Error('extractInvokeReviewersBash: extracted block does not contain "review-lane invoke" — anchor may have drifted');
|
|
}
|
|
if (!block.includes('SELECTED_REVIEWERS')) {
|
|
throw new Error('extractInvokeReviewersBash: extracted block does not contain "SELECTED_REVIEWERS" — anchor may have drifted');
|
|
}
|
|
|
|
return block;
|
|
}
|
|
|
|
// ─── stub preamble ─────────────────────────────────────────────────────────
|
|
//
|
|
// The ONLY I/O seam the extracted block calls is `gsd_run`. Everything else
|
|
// executed by runDispatch() is the real shipped shell text. `barrier` is the
|
|
// positive, deterministic concurrency proof this suite uses in place of any
|
|
// elapsed-time assertion (CLAUDE.md bans wall-clock assertions).
|
|
|
|
const STUB_PREAMBLE = [
|
|
'arg_after() {',
|
|
' local flag="$1"; shift',
|
|
' while [ $# -gt 0 ]; do',
|
|
' if [ "$1" = "$flag" ]; then',
|
|
' printf %s "$2"',
|
|
' return 0',
|
|
' fi',
|
|
' shift',
|
|
' done',
|
|
'}',
|
|
'',
|
|
'in_list() {',
|
|
' local needle="$1" list="$2" item old_ifs="$IFS"',
|
|
' IFS=","',
|
|
' for item in $list; do',
|
|
' if [ "$item" = "$needle" ]; then',
|
|
' IFS="$old_ifs"',
|
|
' return 0',
|
|
' fi',
|
|
' done',
|
|
' IFS="$old_ifs"',
|
|
' return 1',
|
|
'}',
|
|
'',
|
|
'dep_of() {',
|
|
' local slug="$1" pair old_ifs="$IFS"',
|
|
' IFS=","',
|
|
' for pair in $STUB_DEPS; do',
|
|
' case "$pair" in',
|
|
' "$slug="*)',
|
|
' printf %s "${pair#*=}"',
|
|
' IFS="$old_ifs"',
|
|
' return 0',
|
|
' ;;',
|
|
' esac',
|
|
' done',
|
|
' IFS="$old_ifs"',
|
|
'}',
|
|
'',
|
|
'wait_for_file() {',
|
|
' local target="$1" i=0',
|
|
' while [ ! -f "$target" ] && [ "$i" -lt 200 ]; do',
|
|
' sleep 0.05',
|
|
' i=$((i + 1))',
|
|
' done',
|
|
'}',
|
|
'',
|
|
'barrier() {',
|
|
' local slug="$1" i=0 count',
|
|
' touch "$BARRIER_DIR/$slug"',
|
|
' count=$(ls -1 "$BARRIER_DIR" | wc -l)',
|
|
' while [ "$count" -lt "$LANE_COUNT" ] && [ "$i" -lt 200 ]; do',
|
|
' sleep 0.05',
|
|
' i=$((i + 1))',
|
|
' count=$(ls -1 "$BARRIER_DIR" | wc -l)',
|
|
' done',
|
|
' if [ "$count" -lt "$LANE_COUNT" ]; then',
|
|
' echo "barrier-timeout:$slug" >> "$TRACE"',
|
|
' return 1',
|
|
' fi',
|
|
' return 0',
|
|
'}',
|
|
'',
|
|
'gsd_run() {',
|
|
' if [ "$1" = "query" ] && [ "$2" = "config-get" ] && [ "$3" = "review.parallel_lanes" ]; then',
|
|
' if [ "$STUB_CONFIG_GET_FAILS" = "1" ]; then',
|
|
' return 1',
|
|
' fi',
|
|
' printf %s "$STUB_PARALLEL"',
|
|
' return 0',
|
|
' fi',
|
|
'',
|
|
' if [ "$1" = "query" ] && [ "$2" = "review-lane" ] && [ "$3" = "plan" ]; then',
|
|
' shift 3',
|
|
' local sel',
|
|
' sel="$(arg_after --selected "$@")"',
|
|
' if in_list "$sel" "$STUB_BUDGET_FAIL"; then',
|
|
' printf %s \'{"promptBudget": 10}\'',
|
|
' else',
|
|
' printf %s \'{"promptBudget": -1}\'',
|
|
' fi',
|
|
' return 0',
|
|
' fi',
|
|
'',
|
|
' if [ "$1" = "query" ] && [ "$2" = "prompt-budget" ]; then',
|
|
' shift 2',
|
|
' local out base slug',
|
|
' out="$(arg_after --output-prompt "$@")"',
|
|
' base="$(basename "$out")"',
|
|
' slug="${base#gsd-review-prompt-}"',
|
|
' slug="${slug%.md}"',
|
|
' if in_list "$slug" "$STUB_BUDGET_FAIL"; then',
|
|
' return 2',
|
|
' fi',
|
|
' : > "$out"',
|
|
' return 0',
|
|
' fi',
|
|
'',
|
|
' if [ "$1" = "query" ] && [ "$2" = "review-lane" ] && [ "$3" = "invoke" ]; then',
|
|
' shift 3',
|
|
' local slug dep pad',
|
|
' slug="$(arg_after --slug "$@")"',
|
|
' echo "start:$slug" >> "$TRACE"',
|
|
' if [ "$STUB_BARRIER" = "1" ]; then',
|
|
' barrier "$slug" || true',
|
|
' fi',
|
|
' dep="$(dep_of "$slug")"',
|
|
' if [ -n "$dep" ]; then',
|
|
' wait_for_file "$RUN_DIR/done-$dep"',
|
|
' fi',
|
|
' echo "stub review body for $slug" > "$RUN_DIR/gsd-review-$slug.md"',
|
|
' pad=""',
|
|
' if [ "$STUB_PAD_BYTES" -gt 0 ] 2>/dev/null; then',
|
|
' pad="$(head -c "$STUB_PAD_BYTES" /dev/zero | tr "\\0" "x")"',
|
|
' fi',
|
|
' if ! in_list "$slug" "$STUB_SILENT"; then',
|
|
' printf \'{"slug":"%s","pad":"%s"}\\n\' "$slug" "$pad"',
|
|
' fi',
|
|
// #3689: the done-file is a cross-process happens-before edge — a
|
|
// dependent lane unblocks the instant this file appears (wait_for_file
|
|
// above just polls for its existence), so everything a dependent may
|
|
// observe (the "end:$slug" trace line) must be written BEFORE the file
|
|
// that releases it. touch-then-echo let a descheduled upstream lose the
|
|
// race to its own dependent, inverting the #3034 completion-order trace.
|
|
' echo "end:$slug" >> "$TRACE"',
|
|
' touch "$RUN_DIR/done-$slug"',
|
|
' if in_list "$slug" "$STUB_FAIL"; then',
|
|
' return 1',
|
|
' fi',
|
|
' return 0',
|
|
' fi',
|
|
'',
|
|
' return 0',
|
|
'}',
|
|
].join('\n');
|
|
|
|
// ─── dispatch runner ───────────────────────────────────────────────────────
|
|
|
|
/**
|
|
* Build the stub env for a runDispatch() call from `opts`. Every key is
|
|
* always present (never omitted) so the generated `set -u` script never
|
|
* dereferences an unset variable — `opts.parallel === null` deliberately
|
|
* maps to the empty string, which is exactly how an unset config key reads
|
|
* back through `config-get --raw`.
|
|
*/
|
|
function buildEnv(opts, runDir, tracePath, barrierDir) {
|
|
const selected = opts.selected;
|
|
const laneCount = selected.split(',').filter((s) => s.length > 0).length;
|
|
const deps = opts.deps || {};
|
|
return {
|
|
...process.env,
|
|
SELECTED_REVIEWERS: selected,
|
|
EXPLICIT_FLAG: '',
|
|
STUB_PARALLEL: opts.parallel === null || opts.parallel === undefined ? '' : opts.parallel,
|
|
STUB_CONFIG_GET_FAILS: opts.configGetFails ? '1' : '0',
|
|
STUB_FAIL: (opts.failSlugs || []).join(','),
|
|
STUB_BUDGET_FAIL: (opts.budgetFailSlugs || []).join(','),
|
|
STUB_SILENT: (opts.silentSlugs || []).join(','),
|
|
STUB_BARRIER: opts.barrier ? '1' : '0',
|
|
STUB_DEPS: Object.entries(deps).map(([k, v]) => `${k}=${v}`).join(','),
|
|
STUB_PAD_BYTES: String(opts.padBytes || 0),
|
|
LANE_COUNT: String(laneCount),
|
|
TRACE: tracePath,
|
|
BARRIER_DIR: barrierDir,
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Runs the real extracted invoke_reviewers bash block with the gsd_run stub
|
|
* spliced in front of it. `{run_dir}` is replaced globally with a real temp
|
|
* directory. Returns the dispatch outcome plus the artifacts it produced.
|
|
*/
|
|
function runDispatch(t, opts) {
|
|
const scriptDir = createTempDir('gsd-3034-script-');
|
|
const runDir = createTempDir('gsd-3034-rundir-');
|
|
const barrierDir = createTempDir('gsd-3034-barrier-');
|
|
t.after(() => {
|
|
cleanup(scriptDir);
|
|
cleanup(runDir);
|
|
cleanup(barrierDir);
|
|
});
|
|
|
|
const block = extractInvokeReviewersBash();
|
|
const tracePath = path.join(scriptDir, 'trace.log');
|
|
const env = buildEnv(opts, runDir, tracePath, barrierDir);
|
|
|
|
const scriptPath = path.join(scriptDir, 'dispatch.sh');
|
|
const script = [
|
|
'#!/usr/bin/env bash',
|
|
'set -u',
|
|
STUB_PREAMBLE,
|
|
block.split('{run_dir}').join(runDir),
|
|
].join('\n');
|
|
fs.writeFileSync(scriptPath, script, { mode: 0o755 });
|
|
|
|
const result = runHook(scriptPath, [], {
|
|
interpreter: 'bash',
|
|
cwd: runDir,
|
|
env,
|
|
timeoutMs: HOOK_FANOUT_TIMEOUT_MS,
|
|
});
|
|
|
|
const jsonlPath = path.join(runDir, 'gsd-review-lane-results.jsonl');
|
|
const jsonl = fs.existsSync(jsonlPath) ? fs.readFileSync(jsonlPath, 'utf-8') : '';
|
|
const lines = jsonl.split('\n').filter((l) => l.trim() !== '');
|
|
const trace = fs.existsSync(tracePath)
|
|
? readFileNormalized(tracePath).split('\n').filter((l) => l.trim() !== '')
|
|
: [];
|
|
|
|
return {
|
|
outcome: result.outcome,
|
|
exitCode: result.exitCode,
|
|
stderr: result.stderr,
|
|
jsonl,
|
|
lines,
|
|
trace,
|
|
runDir,
|
|
};
|
|
}
|
|
|
|
/** Parsed slug order from JSONL lines — never assert on raw JSONL text. */
|
|
function slugOrder(lines) {
|
|
return lines.map((l) => JSON.parse(l).slug);
|
|
}
|
|
|
|
function serialTrace(slugs) {
|
|
return slugs.flatMap((s) => [`start:${s}`, `end:${s}`]);
|
|
}
|
|
|
|
// ─── #1/#2 — default and explicit-disabled serial dispatch ───────────────
|
|
|
|
describe('#3034 default and explicit-disabled dispatch stays serial', () => {
|
|
test('defaultsToSerialDispatchWhenKeyUnset', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: null });
|
|
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
|
|
});
|
|
|
|
test('staysSerialWhenExplicitlyDisabled', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'false' });
|
|
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
|
|
});
|
|
});
|
|
|
|
// ─── #3/#4 — opt-in concurrency and join-before-aggregate ─────────────────
|
|
|
|
describe('#3034 opt-in concurrency', () => {
|
|
test('dispatchesLanesConcurrentlyWhenEnabled', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', barrier: true });
|
|
const timeouts = result.trace.filter((l) => l.startsWith('barrier-timeout:'));
|
|
assert.deepEqual(timeouts, [], 'no lane should hit the barrier timeout when lanes run concurrently');
|
|
});
|
|
|
|
test('joinsAllLanesBeforeAggregation', (t) => {
|
|
// barrier:true is load-bearing, not decoration. Without it the stub lanes
|
|
// finish instantly and a missing `wait` could still race to 3 lines,
|
|
// making this pass intermittently. Held at the barrier, a missing join
|
|
// deterministically aggregates ZERO lines.
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', barrier: true });
|
|
assert.equal(result.lines.length, 3, 'dispatch must return only after every lane wrote its result');
|
|
});
|
|
});
|
|
|
|
// ─── #5/#6 — selection-order preservation ─────────────────────────────────
|
|
|
|
describe('#3034 JSONL preserves selection order, not completion order', () => {
|
|
test('preservesSelectionOrderSerial', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: null });
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
|
|
});
|
|
|
|
test('preservesSelectionOrderParallelDespiteCompletionOrder', (t) => {
|
|
// codex waits on gemini; gemini waits on claude -> forces reverse
|
|
// completion order (claude, gemini, codex) while selection order stays
|
|
// codex, gemini, claude.
|
|
const result = runDispatch(t, {
|
|
selected: SELECTED_3,
|
|
parallel: 'true',
|
|
deps: { codex: 'gemini', gemini: 'claude' },
|
|
});
|
|
|
|
const endMarkers = result.trace.filter((l) => l.startsWith('end:'));
|
|
// Pin the fixture actually forced reverse completion first — otherwise
|
|
// the selection-order assertion below would pass vacuously.
|
|
assert.deepEqual(endMarkers, ['end:claude', 'end:gemini', 'end:codex']);
|
|
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
|
|
});
|
|
});
|
|
|
|
// ─── #7/#8 — PIPE_BUF boundary triple (plus one clearly-oversized case) ───
|
|
|
|
describe('#3034 oversized lane results stay intact under concurrency', () => {
|
|
test('keepsOversizedLaneResultsIntactUnderConcurrency', async (t) => {
|
|
for (const padBytes of [4095, 4096, 4097, 8192]) {
|
|
await t.test(`padBytes=${padBytes}`, (t2) => {
|
|
const result = runDispatch(t2, { selected: SELECTED_3, parallel: 'true', padBytes });
|
|
assert.equal(result.lines.length, 3);
|
|
for (const line of result.lines) {
|
|
assert.doesNotThrow(() => JSON.parse(line), `line failed to parse at padBytes=${padBytes}: ${line.slice(0, 80)}...`);
|
|
}
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
|
|
});
|
|
}
|
|
});
|
|
});
|
|
|
|
// ─── #9 — single-lane boundary ─────────────────────────────────────────────
|
|
|
|
describe('#3034 single-lane boundary', () => {
|
|
test('singleLaneParallelMatchesSerial', (t) => {
|
|
const serial = runDispatch(t, { selected: 'codex', parallel: null });
|
|
const parallel = runDispatch(t, { selected: 'codex', parallel: 'true' });
|
|
|
|
assert.equal(serial.lines.length, 1);
|
|
assert.equal(parallel.lines.length, 1);
|
|
assert.deepEqual(slugOrder(serial.lines), slugOrder(parallel.lines));
|
|
|
|
const serialMd = fs.readFileSync(path.join(serial.runDir, 'gsd-review-codex.md'), 'utf-8');
|
|
const parallelMd = fs.readFileSync(path.join(parallel.runDir, 'gsd-review-codex.md'), 'utf-8');
|
|
assert.equal(serialMd, parallelMd);
|
|
});
|
|
});
|
|
|
|
// ─── #10 — empty selection is a no-op ──────────────────────────────────────
|
|
|
|
describe('#3034 empty selection', () => {
|
|
test('emptySelectionIsANoOp', (t) => {
|
|
const result = runDispatch(t, { selected: '', parallel: 'true' });
|
|
assert.equal(result.outcome, 'exited');
|
|
assert.equal(result.exitCode, 0);
|
|
assert.deepEqual(result.lines, []);
|
|
assert.deepEqual(result.trace, []);
|
|
});
|
|
});
|
|
|
|
describe('#3034 duplicate slug in the selection', () => {
|
|
test('duplicateSlugDispatchesOnceAndWritesOneLine', (t) => {
|
|
// A slug repeated in SELECTED_REVIEWERS would otherwise put two
|
|
// concurrent background jobs on the same `>`-truncated per-slug result
|
|
// file, corrupting whichever one finishes last. DISPATCH_SLUGS
|
|
// de-duplicates before dispatch, so codex must run exactly once.
|
|
const result = runDispatch(t, { selected: 'codex,gemini,codex', parallel: 'true' });
|
|
assert.equal(result.outcome, 'exited');
|
|
assert.equal(result.trace.filter((l) => l === 'start:codex').length, 1);
|
|
assert.equal(result.trace.filter((l) => l === 'end:codex').length, 1);
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini']);
|
|
});
|
|
});
|
|
|
|
// ─── #11/#12 — lane failure does not abort siblings ────────────────────────
|
|
|
|
describe('#3034 lane failure does not abort sibling lanes', () => {
|
|
test('laneFailureDoesNotAbortSiblingLanes', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', failSlugs: ['gemini'] });
|
|
assert.equal(result.lines.length, 3, 'a failing lane still contributes its stub result line');
|
|
assert.ok(
|
|
fs.existsSync(path.join(result.runDir, 'gsd-review-gemini.md')),
|
|
'the failing lane\'s diagnostic stub .md must be preserved',
|
|
);
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'gemini', 'claude']);
|
|
});
|
|
|
|
test('laneFailureSerialUnchanged', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: null, failSlugs: ['gemini'] });
|
|
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
|
|
assert.equal(result.lines.length, 3);
|
|
});
|
|
});
|
|
|
|
// ─── #13 — budget-too-small skip emits no result line ─────────────────────
|
|
|
|
describe('#3034 budget-too-small skip', () => {
|
|
test('budgetSkipEmitsNoResultLine', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', budgetFailSlugs: ['gemini'] });
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'claude'], 'the budget-skipped lane contributes no result line');
|
|
const stub = fs.readFileSync(path.join(result.runDir, 'gsd-review-gemini.md'), 'utf-8');
|
|
assert.match(stub, /skipped/);
|
|
assert.match(stub, /prompt budget/);
|
|
});
|
|
});
|
|
|
|
// ─── #14/#15 — silent / absent lane contributes no line ───────────────────
|
|
|
|
describe('#3034 silent lane contributes nothing, not a blank line', () => {
|
|
test('silentLaneContributesNoLine', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', silentSlugs: ['gemini'] });
|
|
assert.deepEqual(slugOrder(result.lines), ['codex', 'claude']);
|
|
// Folded row #15 (absent result file): budgetSkipEmitsNoResultLine above
|
|
// already exercises the "invoke never started, no file at all" shape via
|
|
// its `continue`. This assertion pins the sibling shape: a lane that DID
|
|
// run and wrote nothing must not leave a stray blank JSONL line.
|
|
assert.ok(!result.jsonl.includes('\n\n'), 'an empty lane result must not leave a blank JSONL line');
|
|
});
|
|
});
|
|
|
|
// ─── #16/#17 — non-canonical truthy values stay serial ────────────────────
|
|
|
|
describe('#3034 non-canonical truthy config values stay serial', () => {
|
|
test('nonCanonicalTruthyValuesStaySerial', async (t) => {
|
|
const nearMisses = ['TRUE', 'True', '1', 'yes', 'on', ' true', 'true '];
|
|
for (const value of nearMisses) {
|
|
await t.test(`parallel_lanes="${value}"`, (t2) => {
|
|
const result = runDispatch(t2, { selected: SELECTED_3, parallel: value });
|
|
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']), `value "${value}" must not opt into parallel dispatch`);
|
|
});
|
|
}
|
|
});
|
|
});
|
|
|
|
// ─── #18 — config-get failure fails safe to serial ─────────────────────────
|
|
|
|
describe('#3034 broken config tooling fails safe to serial', () => {
|
|
test('configGetFailureFallsBackToSerial', (t) => {
|
|
const result = runDispatch(t, { selected: SELECTED_3, parallel: 'true', configGetFails: true });
|
|
assert.deepEqual(result.trace, serialTrace(['codex', 'gemini', 'claude']));
|
|
});
|
|
});
|
|
|
|
// ─── #19/#20/#21 — config-set registers review.parallel_lanes ─────────────
|
|
|
|
describe('#3034 review.parallel_lanes config key', () => {
|
|
test('configSetAcceptsAndPersistsParallelLanes', (t) => {
|
|
const tmpDir = createTempProject();
|
|
t.after(() => cleanup(tmpDir));
|
|
|
|
const setResult = runGsdTools('config-set review.parallel_lanes true', tmpDir);
|
|
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
|
|
|
|
const configPath = path.join(tmpDir, '.planning', 'config.json');
|
|
const config = JSON.parse(fs.readFileSync(configPath, 'utf-8'));
|
|
assert.equal(config.review?.parallel_lanes, true);
|
|
assert.equal(typeof config.review?.parallel_lanes, 'boolean');
|
|
|
|
const getResult = runGsdTools('config-get review.parallel_lanes --raw', tmpDir);
|
|
assert.ok(getResult.success, `config-get failed: ${getResult.error}`);
|
|
assert.equal((getResult.output || '').trim(), 'true');
|
|
});
|
|
|
|
test('configSetPersistsBooleanFalse', (t) => {
|
|
const tmpDir = createTempProject();
|
|
t.after(() => cleanup(tmpDir));
|
|
|
|
const setResult = runGsdTools('config-set review.parallel_lanes false', tmpDir);
|
|
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
|
|
|
|
const configPath = path.join(tmpDir, '.planning', 'config.json');
|
|
const config = JSON.parse(fs.readFileSync(configPath, 'utf-8'));
|
|
assert.equal(config.review?.parallel_lanes, false);
|
|
assert.equal(typeof config.review?.parallel_lanes, 'boolean');
|
|
});
|
|
|
|
test('rejectsUnregisteredNeighbouringKey', (t) => {
|
|
const tmpDir = createTempProject();
|
|
t.after(() => cleanup(tmpDir));
|
|
|
|
// Missing trailing "s" — proves the whitelist is load-bearing and the
|
|
// two tests above are not vacuous (they'd pass even for an unregistered
|
|
// key if config-set accepted anything).
|
|
const result = runGsdTools('config-set review.parallel_lane true', tmpDir);
|
|
assert.equal(result.success, false, 'an unregistered near-miss key must be rejected');
|
|
});
|
|
});
|
|
|
|
// ─── #22 — serial/parallel artifact equivalence ────────────────────────────
|
|
|
|
describe('#3034 serial and parallel dispatch produce equivalent artifacts', () => {
|
|
test('serialAndParallelProduceEquivalentArtifacts', (t) => {
|
|
const serial = runDispatch(t, { selected: SELECTED_3, parallel: null });
|
|
const parallel = runDispatch(t, { selected: SELECTED_3, parallel: 'true' });
|
|
|
|
assert.deepEqual(slugOrder(serial.lines), slugOrder(parallel.lines));
|
|
assert.deepEqual(
|
|
serial.lines.map((l) => JSON.parse(l)),
|
|
parallel.lines.map((l) => JSON.parse(l)),
|
|
);
|
|
|
|
for (const slug of ['codex', 'gemini', 'claude']) {
|
|
const serialMd = fs.readFileSync(path.join(serial.runDir, `gsd-review-${slug}.md`), 'utf-8');
|
|
const parallelMd = fs.readFileSync(path.join(parallel.runDir, `gsd-review-${slug}.md`), 'utf-8');
|
|
assert.equal(serialMd, parallelMd, `gsd-review-${slug}.md must be byte-identical between the two paths`);
|
|
}
|
|
});
|
|
});
|