* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking answer. Three families do exactly that today; this commit pins each one RED. Measured on this tree, 2026-08-27: intel query, .planning/intel/file-roles.json nested 12000 deep -> exit 1, "Error: Maximum call stack size exceeded" searchJsonEntries / matchesInValue carry no depth parameter at all. The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it. same fixture nested 48 and 49 deep -> both return total=1 at exit 0, truncated=undefined Nothing distinguishes "searched to the bottom" from "stopped looking". query phase-plan-index, a plan whose depends_on names an unresolvable token -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] The token is never mentioned. computeDependencyLevels drops the edge with `if (!resolvedDep) continue;`, every plan becomes a root, and the tool then reports the author's correct wave: as the thing that is wrong. countPhasePlansAndSummaries with fs.readdirSync throwing EACCES -> hasContext:false, indistinguishable from a phase that simply has no CONTEXT.md. context_read_error is undefined. The shapes these tests assert against, chosen here so the implementation has a target rather than inventing one later: `truncated: boolean` on the intel query result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and `context_read_error: string | null` per analyzed phase. Deliberately green, and they must stay that way — each stops the fix from over-firing: depth 48 is found and NOT flagged truncated (the ceiling is inclusive) a shallow miss reports no truncation (noise control, N1) 10,000 siblings at depth 2 are unaffected (the bound is DEPTH, N2) a genuine wave: mismatch on a fully-resolved DAG still warns (N3) a genuinely missing directory is absent, not an error the emitted depends_on display mapping still passes an unresolved token through verbatim — already pinned by the existing #3785 test, so no duplicate was added T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the real CLI and reads the emitted JSON, because a unit assertion on computeDependencyLevels would have passed throughout #3427's life. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3885): no silent swallow, and no verdict manufactured from dropped data Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed into an output that reads as authoritative. The recursion bound, restored but NOT verbatim (src/intel.cts) MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and matchesInValue, which carried no depth parameter at all. The bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it — §8.3's "a consolidation may not delete an invariant along with the surface that held it", demonstrated. Measured before: a .planning/intel file nested 12000 deep exits 1 with "Error: Maximum call stack size exceeded". Reachable from a project document. The original returned a bare `false` at the ceiling. Restoring that verbatim would trade a crash for a silent "no match" when the truth is "I stopped looking" — the same class this epic exists to close, and ADR-3473 Decision 4 forbids it. So the bound carries a truncation signal: nesting 47 -> found, truncated false nesting 48 -> found, truncated false (the ceiling is inclusive) nesting 49 -> not found, truncated TRUE nesting 12000 -> exit 0, truncated TRUE, no RangeError A shallow document that simply has no match reports truncated FALSE — the flag means "I stopped early", never "I found nothing", or it would be noise. The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected. The dropped edge is named, and stops being blamed on the author (src/phase.cts) computeDependencyLevels dropped every unresolvable depends_on token with a bare `continue`. Each drop makes a plan a root, so the whole phase collapses to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as the thing that was wrong. Before: warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] After: warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not resolve to any plan in this phase — edge dropped, wave placement for this plan may be unreliable"] The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is deliberately not built here. The emitted depends_on display mapping still passes an unresolved token through verbatim (#3785). No artifact from failed inputs (gsd-core/workflows/review.md, #3352) A failed lane leaves no result file, so "every lane failed" is exactly "the aggregate JSONL has zero lines" — the gate condition already existed as a byproduct. REVIEWS.md is no longer written in that case, and the commit step is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT counted as a failure. Per-lane output and non-empty .err are preserved to .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that the lanes failed at all; the commit step names one file, never a glob, so the diagnostics are not swept in. Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2) Four callers collapsed an EACCES on a phase directory into [] and reported hasContext:false — byte-identical to a phase that simply has no CONTEXT.md. Each now names the directory it could not read. A genuinely missing directory stays absent rather than becoming an error, which is what keeps the fix from over-firing. Fatal errno folded into a retry set: audited, no defect found Reported as a verified negative rather than padded with a change. withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776; atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they return or rethrow the final error rather than swallowing it. estimate-cli's sole caller surfaces that rethrow as write_error in its JSON output. Manufacturing a diff to make the checkbox look worked-on is the Goodhart outcome Decision 6 exists to prevent. Disclosed: R46 (the commit step names one file, never a glob) is a real regression guard but is NOT independently failing-first — the commit fence is byte-identical pre- and post-fix, so it only fails pre-fix through its shared extraction dependency. Recorded rather than claimed as fail-first. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence Two review findings, both real, both in my own change. An isolated adversarial review found the evidence-preservation block never checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in a SEPARATE fenced block. A disk-full or unwritable phase directory therefore still destroyed the only copy of the failed lanes' output — reintroducing the exact #3352 data loss this item exists to stop, inside the fix for it. Preservation and cleanup are now one block, because each fenced block is a separate execution and a shell variable cannot carry between them. mkdir -p and each cp are exit-checked; cleanup runs only when preservation succeeded, and a failure warns naming the intact run directory. "Nothing to preserve" is not a failure and still cleans up. Driven three ways: success removes run_dir, failure leaves it intact with the warning, nothing-to-preserve removes it. The failure is induced by a file-vs-directory conflict rather than chmod 0o000, which root bypasses. The new unresolved-depends_on warning embedded a user-authored token verbatim: warnings: ["Plan 03-02: depends_on token \"evil Plan 03-01: FORGED WARNING\" does not resolve ..."] The JSON wire form is safe, and the security reviewer judged it non-exploitable for that reason. It is escaped anyway through formatDiagnosticToken — the helper #3884 added one phase earlier for exactly this class. warnings[] is an array a consumer naturally prints line by line, and not reusing the sibling fix is the generative-fix-divergence shape this epic exists to close. The same treatment is applied to context_read_error / phase_dir_read_error, which embed a phase directory path a repository can choose, and to the fs error message, which echoes the raw path itself. Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per negative space N2, disclosed rather than left to be discovered. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot" Blocker from the round-2 isolated review, and it is my own inconsistency: this phase applied "unreadable is not absent" to phase directories and left it broken in the file it was already editing. chmod 000 .planning/intel/file-roles.json gsd-tools intel query <term> -> {"matches":[],"total":0,"truncated":false} exit 0 safeReadJson swallowed every read failure and returned null, so an EACCES was byte-indistinguishable from an absent file AND from a genuine no-match. Now it separates three states: ENOENT stays silently absent, because not every project has every intel file and intelQuery loops over all of them expecting misses; EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel file previously read as "no matches" too — same defect, same fix. Threading that outcome through the other three callers found something worse than the reported case. intelDiff returned no_baseline:true for a corrupt or unreadable snapshot — not a silent failure but an actively FALSE verdict, telling the caller they never took a snapshot when they did. That is §8.5's headline case, so it is fixed and tested rather than noted. intelStatus and intelApiSurface collapsed the same way; intelApiSurface additionally printed a "not yet populated" banner that was simply untrue. Every row is failing-first, including the absent-file ones — the field is new, so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin that the fix does not OVER-fire on the ordinary absent case, which is what would turn this into noise on every project lacking an intel file. IO failure is injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as a test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): build the pathological intel fixture as text, not by stringifying a nested object The remote runner came back red on Linux with two failures, both T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on macOS. The product was never at fault. writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned. Linux's container stack is smaller than macOS's, which is the whole of the platform difference. Measured, with the same document built as JSON TEXT so nothing in the building process recurses: depth=100 rc=0 truncated=true depth=5000 rc=0 truncated=true depth=12000 rc=0 truncated=true depth=60000 rc=0 truncated=true V8 parses this shape iteratively; only stringify recurses. The bound works at every depth tried. The fixture is now built by string concatenation. That is also the more faithful input — a real deeply nested JSON document on disk is exactly what the bound guards, where a stringified object was only ever a way to produce one. The depth stays 12000. Lowering it would have made the test pass by weakening it to accommodate a fixture bug, and 12000 is a legitimate pathological input the product handles. T4 remains a genuine fail-first: rebuilt against the parent of the commit that added the bound, the string-built depth-12000 fixture still drives the CLI to rc=1 with "Error: Maximum call stack size exceeded". A comment records why the fixture is text, so it is not "simplified" back into a macOS-green / Linux-red test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3885): backfill the changeset PR number Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): normalize path separators before splicing into the workflow's bash CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the remote runner were all green. AssertionError: commit must name the single REVIEWS.md file; got: --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The harness spliced an OS-native temp path into the extracted bash, and bash consumes \U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory survived — which is the other two assertions. This is a fixture defect, not a product one, and that was checked rather than assumed. In production the phase directory is toPosixPath-normalized at every call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and the run directory is created by "mktemp -d" running inside the bash block itself (gsd-core/workflows/review.md:163), which emits POSIX-style output even under Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it. The file's pre-existing #3034 harness splices raw native paths too, but only ever inside double-quoted assignments, so it never tripped this — my new harness followed that convention faithfully into the one place where it does not hold. Both now splice through toPosixPath from shell-command-projection, the established seam, which is a no-op on POSIX and mirrors what production does. No assertion was weakened. "commit must name the single REVIEWS.md file" and "the run dir must still be destroyed" still assert exactly that; only how the fixture supplies its path changed. Nothing is skipped on Windows — a t.skip() here would have hidden the question of whether the exposure was real, which is the question that mattered. Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input produces a byte-identical shape, proving the normalization is idempotent. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): stop the harness making the deleted run dir its own cwd Windows shard 3/3 stayed red after the separator fix, on two assertions the separator fix never touched: AssertionError: the run dir must still be destroyed AssertionError: nothing to preserve is not a failure — run dir must still be removed The separators were a real bug and fixing them fixed the --files assertion. They were not this bug, and two CI cycles went into the wrong axis before I stopped converting path forms and looked at what the harness actually does. runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's working directory WAS the directory the block under test then removes with rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live process's working directory cannot be removed. So on Windows the directory survived and both assertions failed, on macOS and Linux it vanished and they passed. Nothing to do with slashes. Harness-only. Production never cd's into the run directory; every reference is by absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself (gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged. Fix: the child now runs with its cwd in an unrelated temp directory that the block under test never deletes. Neither assertion was weakened, and nothing is skipped on Windows — the tests in this file carry no platform guard and run there unconditionally, which is how this surfaced at all. Honest limit: the Windows failure mode cannot be reproduced on macOS, because POSIX permits the very thing Windows refuses. The diagnosis is grounded in that documented divergence and in the fact that only the Windows lane failed, but the green outcome on windows-latest is unverified until CI runs it. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
140 lines
6.5 KiB
Markdown
140 lines
6.5 KiB
Markdown
# How to enable parallel reviewer lanes
|
|
|
|
Cut the wall-clock cost of a multi-reviewer `/gsd-review` pass from the sum of its lanes toward
|
|
its slowest lane — without losing the rate-limit protection the sequential default exists to
|
|
provide.
|
|
|
|
> **Default-off, and deliberately so.** Reviewer lanes are dispatched one at a time because
|
|
> concurrent invocation can trip provider rate limits, and a lane lost to a rate limit is a
|
|
> cross-AI review that quietly went blind in one eye. Turning this on is you asserting that your
|
|
> providers can take the concurrency. Nothing detects that for you.
|
|
|
|
**What you need:**
|
|
|
|
- Two or more reviewer lanes that actually run on this host. With one lane there is nothing to
|
|
overlap and the setting changes nothing.
|
|
- Provider capacity for concurrent requests — separate accounts, generous quota, or local model
|
|
servers (`ollama`, `lm-studio`, `llama-cpp`) that have no external limit at all.
|
|
|
|
---
|
|
|
|
## Step 1 — Check what your review pass actually runs
|
|
|
|
Parallelism only helps if several lanes are selected. Confirm the set first:
|
|
|
|
```bash
|
|
gsd config-get review.default_reviewers
|
|
```
|
|
|
|
If that returns `Key not found`, no-flag runs use every detected reviewer and `--all` is
|
|
redundant. If it names a single reviewer, stop here — enabling the key would change nothing.
|
|
|
|
---
|
|
|
|
## Step 2 — Turn the key on
|
|
|
|
```bash
|
|
gsd config-set review.parallel_lanes true
|
|
```
|
|
|
|
Verify it took:
|
|
|
|
```bash
|
|
gsd config-get review.parallel_lanes --raw
|
|
# → true
|
|
```
|
|
|
|
**The guard is strict equality.** Only the exact value `true` opts in. `"1"`, `"yes"`, `"on"` and
|
|
`"TRUE"` are all read as *not enabled* and leave dispatch sequential. This is intentional: a
|
|
mistyped config gets the conservative behavior rather than firing concurrent requests at a
|
|
rate-limited provider. If `config-get` shows anything other than `true`, the setting is off.
|
|
|
|
---
|
|
|
|
## Step 3 — Run a review and read the result
|
|
|
|
```bash
|
|
/gsd-review --phase 3 --all
|
|
```
|
|
|
|
or, for the convergence loop:
|
|
|
|
```bash
|
|
/gsd-plan-review-convergence 3 --all
|
|
```
|
|
|
|
Open `{phase_dir}/{padded_phase}-REVIEWS.md` and check the `reviewers:` frontmatter list. Every
|
|
lane you selected must appear there. That list is the contract: lanes are joined before the file
|
|
is rendered, so a missing reviewer means that lane did not produce a review — never that
|
|
aggregation ran early.
|
|
|
|
**If every selected lane failed, `REVIEWS.md` is not written at all** (ADR-3473 §8.5, #3885) — the
|
|
run reports the failure instead of synthesizing a review artifact from zero lane results. This is
|
|
not specific to parallel dispatch (a sequential run where every lane fails behaves the same way),
|
|
but concurrency gives you more ways to lose every lane in one pass. Each lane's raw output and any
|
|
non-empty `.err` file are preserved beside the phase's artifacts before the run's own cleanup runs,
|
|
so a missing `REVIEWS.md` is diagnosable, not silent — see the table below for what a *partial*
|
|
failure (some, not all, lanes down) looks like instead.
|
|
|
|
Section order in `REVIEWS.md`, and line order in the run's `gsd-review-lane-results.jsonl`, are
|
|
unchanged from sequential dispatch. They follow reviewer-selection order, not completion order,
|
|
so a diff of two runs shows no reordering churn.
|
|
|
|
---
|
|
|
|
## Telling the outcomes apart
|
|
|
|
The single most useful habit: **an empty or stub review is a dropped lane, not a clean review.**
|
|
That is true sequentially too, but concurrency gives you more ways to drop one at once.
|
|
|
|
| What you see | What it means | What to do |
|
|
|---|---|---|
|
|
| Every selected lane in `reviewers:`, all sections populated | Working as intended | Nothing |
|
|
| A lane's section carries the "failed or returned empty output" header | The lane ran and produced nothing usable. Read the captured stderr in the stub | If it names a rate limit or quota, your provider cannot take this concurrency — see below |
|
|
| A lane reports `probe_timeout` or `host_unreachable` | The lane could not be reached at all — a local server that is down, not a concurrency effect | Start the server; unrelated to this setting |
|
|
| A lane reports `budget_too_small` | Its prompt budget cannot fit the minimum review set | Raise `review.max_prompt_tokens_per_reviewer.<slug>`; unrelated to this setting |
|
|
| A lane reports `egress_host_changed` | The lane was consented to one destination and the config now names another. It is blocked, not redirected | Re-consent deliberately; unrelated to this setting |
|
|
| A lane is missing from `reviewers:` entirely | It was never selected | Check your flags and `review.default_reviewers` |
|
|
| Several lanes stub out at once, with provider errors | The concurrency is more than your account can take | Turn the key back off, or narrow `review.default_reviewers` |
|
|
|
|
**Rate-limited lanes fail loudly.** A lane that gets throttled goes down the same path as any
|
|
other failing lane — a diagnostic stub carrying its stderr, kept distinguishable from a real
|
|
review by its header. It is not silently dropped and it does not abort its sibling lanes.
|
|
|
|
---
|
|
|
|
## Turning it back off
|
|
|
|
```bash
|
|
gsd config-set review.parallel_lanes false
|
|
```
|
|
|
|
The next pass dispatches sequentially again. Nothing else changes: no artifact written under the
|
|
parallel setting needs migrating, because the output layout is identical in both modes.
|
|
|
|
---
|
|
|
|
## What this does not speed up
|
|
|
|
**Convergence cycles stay sequential.** `/gsd-plan-review-convergence` runs
|
|
`plan-phase → review → replan → re-review`, and each cycle genuinely depends on the previous
|
|
one's output. Enabling this key makes each *review pass* inside a cycle faster; it does not
|
|
reduce the number of cycles, and it does not overlap planning with reviewing. A three-cycle
|
|
convergence run is still three sequential rounds.
|
|
|
|
**There is no concurrency bound.** Every selected lane dispatches at once. Eleven selected lanes
|
|
means eleven concurrent requests. If that is more than you want, the control is the size of your
|
|
selected set — `review.default_reviewers` or explicit flags — not a throttle on this setting.
|
|
|
|
**Reviewer instances sharing an adapter are not grouped.** Two
|
|
[`review.reviewer_instances`](../CONFIGURATION.md#reviewer-instances-for-gsd-review-1517) entries
|
|
backed by the same CLI dispatch concurrently against that one provider. If you run several
|
|
same-provider instances, you are the most likely configuration to hit a limit, and this setting
|
|
gives you no way to serialize just those.
|
|
|
|
---
|
|
|
|
**See also:** [Configuration reference](../CONFIGURATION.md#parallel-reviewer-lanes-for-gsd-review-3034)
|
|
· [Reviewer instances](../CONFIGURATION.md#reviewer-instances-for-gsd-review-1517)
|
|
· [`/gsd-review` command reference](../COMMANDS.md)
|